Fallen Tear: The Ascension - Steam Optimization Journey with PH3
Hi everyone,
this is Peter “Durante” Thoman from PH3 games speaking – if any of you know me then it might be from our JRPG ports, or maybe from older modding work. Today I’m happy to write this guest update about some technical aspects of Fallen Tear: The Ascension.
Background
Not too long ago, in late June, Stephen of CMD Studios contacted us because they were facing some technical challenges with the game, and were wondering if we could provide some support. I was quite happy about this, since I always felt that this is something we are almost uniquely suited to, as the density of experienced performance specialists among PH3 employees is almost absurd. However, so far, we were primarily contracted for complete ports, not specifically performance optimization of existing projects.
We came to an agreement and got started in July in an advisory position, and moved to more direct implementation work in August after it became clear that more substantial changes were needed to solve some underlying issues.
The Issues
While there are lots of smaller things we’ve discussed, advised on, or touched, the main four issues that our technical contributions focus on are the following:
-
Peak memory consumption, particularly in terms of VRAM
-
Loading times, both at startup and between zones
-
Rendering performance, especially on systems without a dedicated GPU like the Steam Deck
-
Traversal stutter
I’ll briefly discuss each of those, what we did to improve them, and provide some data where that makes sense.
Peak Memory Consumption
Memory consumption was perhaps the biggest issue with the game. In particular, the baseline state was sufficiently memory-hungry that the game simply would not work on Steam Deck at all – something that some of you might have noticed in the later parts of the Early Access period.
In basic terms, the game had, over the course of its development, accumulated so many high-quality assets that they collectively required more than 18 GB of memory, and specifically over 13.3 GB of GPU memory – obviously a problem for e.g. the Steam Deck with its 16 GB total pool, but also for many other hardware configurations.
We approached this problem in two ways. Firstly, the game already had the beginnings of an asset streaming system, but it was inactive due to various issues and incomplete implementation. We extended / completed the implementation of this system.
Secondly, in the baseline, the game always loaded all assets at the absolute maximum quality level available. As this could be authored for e.g. 4k output resolution, it is massive overkill for devices such as the Steam Deck or older non-gaming laptops. We created a mostly-automated pipeline that creates suitable mipmaps and allows loading assets at a more appropriate quality level, with decisions made specifically to keep a high-quality overall visual impression.
The chart above shows various memory metrics in the same in-game scenario. “Baseline” is the state before our optimizations; “PH3 High” enables streaming and uses mimaps, but provides the same or higher visual quality as the baseline. Nonetheless, the peak GPU memory consumption is 2.3 GB or ~21% lower. Finally, the “PH3 Medium” version provides exactly the same output quality on devices such as the Steam Deck (and almost completely equivalent quality up to 1080p) reduces GPU memory consumption to 5.3 GB, substantially less than half of the baseline.
The comparison image above shows the difference – or lack thereof – between “Medium” and “High” at a 1080p output size (which is in itself substantially higher than the Steam Deck native resolution).
Loading Times
The second most important area (at least in my opinion) we focused our work on were loading times. Personally, I get easily annoyed by long loading times, especially when traversing between different scenes. Here, we reduced redundant work during loading, and made sure that various processes can overlap wherever possible.
All of these improvements culminated in substantial decreases to loading times, as you can see in the selected results above. Initial startup to menu loading decreased from ~12 to ~7 seconds on this system, and traversal waiting times, while highly dependent on the areas involved, always decreases substantially and is sometimes even cut in half.
Rendering Performance
Rendering performance was not nearly as big of an issue as the first two topics discussed above, but there were still some improvements to be made. The introduction of uniform mipmapping, primarily driven by the needs of the quality selection mechanism above, also improves performance, especially when targeting lower resolutions (such as, once again, on the Steam Deck). Furthermore, the game had some render passes which uniformly used 8x MSAA despite not really benefitting from such a high level of multisampling in most viewing scenarios. This is now an adjustable setting.
You can see the results above. While all of these values are obviously more than sufficient for the Deck’s 90 Hz display, more efficient rendering adds more headroom for challenging scenes and reduces power consumption (at the usual 90 FPS limit, GPU power consumption in this scene dropped from 3.1 W baseline to only 2.2 W at 2xMSAA).
Traversal Stutter
The final larger issue that we investigated and implemented some improvements for is traversal stutter. This refers to intermittent spikes in frame computation time, and is generally one of the hardest types of performance issues to improve. Given our limited time on the project so far, we haven’t completely eliminated all sources of stutter, but we have greatly improved the situation. The main improvements we made and identified in this regard are:
-
Incremental garbage collection - “garbage collection” is the process by which data that is no longer needed is freed up in some types of programming languages. Incremental garbage collection means that, rather than doing so for lots of data at once and potentially causing stutter, the load is spread more evenly across frames.
-
Removal of repeating allocations - we eliminated some recurring memory allocations, that might sometimes take more and sometimes less time to complete, again causing uneven frametimes.
-
Background audio loading - some specific stutters were due to audio files being loaded synchronously, these are now loaded in the background instead.
Stutter comparisons are a lot less straightforward than metrics such as FPS or memory consumption, but across comparable gameplay scenarios we measured a 40-80% reduction in stutter due to these measures, compared to the baseline.
Conclusion
I hope you found this look inside our work interesting, informative, or both. I’d like to thank the team at PH3 – especially Markus, who gave 200% for the FTA project over the past few weeks – and of course everyone at Winter Crew and CMD Studios for trusting us with this. It’s great to be able to put our technical skills to use in polishing such an awesome indie game.
Cheers,
Peter “Durante” Thoman, CTO, PH3
Комментарии 0
Гостевые никнеймы не подтверждают личность. Все комментарии проходят проверку владельцем сайта.
Обсуждение ещё не началось.