https://semiengineering.com/how-neural-super-sampling-works-architecture-training-and-inference/ [semi_logo] [ ]Submit Subscribe * Home * Systems & Design * Low Power - High Performance * Manufacturing, Packaging & Materials * Test, Measurement & Analytics * Auto, Security & Enabling Technologies * Special Reports * Business & Startups * Jobs * Knowledge Center * Technical Papers + Home'; + AI/ML/DL + Architectures + Automotive/ Aerospace + Communication/Data Movement + Design & Verification + Lithography + Manufacturing + Materials + Memory + Optoelectronics / Photonics + Packaging + Power & Performance + Quantum + Security + Test, Measurement, Analytics tech papers + Transistors + Z-End Applications * Events & Webinars + Events + Webinars * Videos & Research + Videos + Industry Research * Newsletters & Store + Newsletters + Store * MENU + Home + Special Reports + Systems & Design + Low Power-High Performance + Manufacturing, Packaging & Materials + Test, Measurement & Analytics + Auto, Security & Enabling Technologies + Knowledge Center + Videos + Startup Corner + Business & Startups + Jobs + Technical Papers + Events + Webinars + Industry Research + Newsletters + Store + Special Reports Home > Low Power-High Performance > How Neural Super Sampling Works: Architecture, Training, And Inference Low Power-High Performance SPONSOR BLOG How Neural Super Sampling Works: Architecture, Training, And Inference AI-powered upscaling for mobile gaming. September 11th, 2025 - By: Liam O'Neil popularity This blog post is the second in our Neural Super Sampling (NSS) series. The post explores why we introduced NSS and explains its architecture, training, and inference components. In August 2025, we announced Arm neural technology that will ship in Arm GPUs in 2026. The first use case of the technology is Neural Super Sampling (NSS). NSS is a next-generation, AI-powered upscaling solution. Developers can already start experimenting with NSS today, as discussed in the first post of this two-part series. In this blog post, we take a closer look at how NSS works. We cover everything from training and network architecture to post-processing and inference. This deep dive is for ML engineers and mobile graphics developers. It explains how NSS works and how it can be deployed on mobile hardware. Why we replaced heuristics with Neural Super Sampling Temporal super sampling (TSS), also known as TAA, has become an industry standard solution for anti-aliasing over the last decade. TSS offers several benefits. It addresses all types of aliasing, is compute-efficient for deferred rendering, and extensible to upscaling. However, it is not without its challenges. Hand-tuned heuristics, commonly used in TSS approaches today, can be difficult to scale and require continual adjustment across varied content. Issues like ghosting, disocclusion artifacts, and temporal instability remain. These problems worsen when combined with upscaling. NSS overcomes these limitations by using a trained neural model. Instead of relying on static rules, it learns from data. It generalizes across conditions and content types, adapting to motion dynamics and identifying aliasing patterns more effectively. These capabilities help NSS handle edge cases more reliably than approaches such as AMD's FSR 2 and Arm ASR. Training the NSS network: Recurrent learning with feedback NSS is trained using sequences of 540p frames rendered at 1 sample per pixel. Each frame is paired with 1080p ground truth images rendered at 16spp. Sequences are about 100 frames to help the model understand how image content changes over time. Inputs include rendered images, such as color, motion vectors, and depth, alongside engine metadata, such as jitter vectors, and camera matrices. The model is trained recurrently and runs forward across a sequence of multiple frames before performing each backpropagation. This approach lets the network propagate gradients through time and learn how to accumulate information. The network is trained using a spatiotemporal loss function. It simultaneously penalizes errors in both spatial fidelity and temporal consistency. Spatial fidelity keeps each frame sharp, detailed, and visually accurate. It also preserves the edges, textures, and fine structures. Temporal stability discourages flickering, jittering, or other forms of temporal noise across consecutive frames. Training is done in PyTorch using well-established practices, including the Adam optimizer, a cosine annealing learning rate schedule, and standard data augmentation strategies. Pre- and post-processing passes are written in Slang for flexibility and performance. ExecuTorch is used for quantization-aware training. Network architecture and output designs The NSS network uses a four-level UNet backbone. It includes skip connections to preserve spatial structure. It downsamples and upsamples input data across three encoder and decoder modules. We evaluated several approaches: * Image prediction: Easy to implement, but struggled under quantization and caused visual artefacts. * Kernel prediction: Generalized well and quantized effectively, but produced high bandwidth overhead due to many large kernel maps. * Parameter prediction (Chosen): Outputs a small set of parameters per pixel. These drive post-processing steps like filtering and sample accumulation. It is quantization-friendly and bandwidth-efficient. The network generates three per-pixel outputs: * A 4x4 filter kernel. * Temporal coefficients used for accumulation and rectification. * A hidden state tensor, passed to the next frame as temporal feedback. The network outputs serve two paths: * The filter kernel and temporal coefficients that are consumed by the post-processing stage to compute the final upscaled image. * The hidden state, which is passed forward to inform the next frame's inference. Unlike approaches like Arm ASR, which rely on hand-tuned heuristics, a machine-learning approach like NSS has a three-fold benefit: 1. NSS estimates dynamic kernel filters and parameters that address aliasing at a per-pixel granularity. 2. NSS harnesses temporal feedback, which captures historic state across multiple frames, for greater temporal stability. 3. NSS can be fine-tuned on new game content, further enabling developers to optimize image quality for their specific titles. Improving frame-to-frame consistency with temporal feedback NSS introduces two key feedback mechanisms to address temporal instabilities: * Hidden features from prior frames are passed forward, allowing the network to learn what changed and what persisted. * A luma derivative is computed to detect flickering thin features, which highlights temporal differences that indicate instability. These inputs help the model maintain temporal stability without relying on handcrafted rules. Pre-processing stage: Preparing the input A GPU-based pre-processing stage runs before inference. It prepares the inputs required by NSS. This stage gathers per-pixel attributes like color, motion vectors, and depth. It also computes the luma derivative, a temporal signal that flags thin-feature flicker, and a disocclusion mask that highlights stale history. In addition, it reprojects hidden features from previous frames. These are assembled into a single input tensor for the neural network. This stage runs as a compute shader. It executes before the inference call, which runs on the GPU using Vulkan ML extensions. Post-processing: From raw output to final frame After inference, a post-process stage runs as a compute shader to construct the output color. All steps are integrated into the render graph and are designed to run efficiently on mobile. These steps include: * Motion vector dilation. This reduces aliasing when reprojecting history. * History reprojection. A Catmull-Rom filter to reduce reprojection blur. * Filtering. This applies the 4x4 kernel to anti-alias the current color input. * Sparse upscaling. This maps jittered low-res samples onto a high-res grid. Any missing pixels are zero-filled and then sparsely filtered with the 4x4 kernel. This step performs interpolation and anti-aliasing, like demosaicing. * Rectification. This uses the predicted theta parameter to reject stale history. * Sample accumulation. This uses a predicted alpha parameter to blend new data with the history buffer. Performed in a tone-mapped domain to prevent "firefly" artefacts. How we validate quality We assess NSS using several metrics. These include PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and FLIP, a rendering focused perceptual error metric. These metrics do not always match human perception. However, they help surface problem cases. Tracking multiple metrics builds confidence. A Continuous Integration (CI) workflow replays test sequences. It logs performance across NSS, Arm Accuracy Super Resolution (ASR), and other baselines. For visual comparisons and perceptual evaluation, please refer to this white paper. In 540p-to-1080p comparisons, NSS improves stability and detail retention. It performs well in scenes with fast motion, partially occluded objects, and thin geometry. Unlike non-neural approaches such as Arm ASR or AMDs FSR 2, NSS also handles particle effects without needing a reactive mask. Can NSS run in real-time? While silicon products with Neural Accelerators have not yet been announced, we can estimate whether NSS is fast enough. This estimate is based on minimum performance assumptions and the number of MACs required to perform an inference of the network. This analysis applies to any accelerator which meets the same assumptions for throughput, power and utilization. We assume a target of 10 TOP/s per-watt of neural acceleration is achievable at a sustainable GPU clock frequency. We target <=4ms for the upscaler per frame in sustained performance conditions. Shader stages before and after inference take about 1.4ms on a low-frequency GPU. With this budget, NSS must stay below approximately 27 GOPs. Our parameter prediction network uses about 10 GOPs. This fits comfortably within that range, even at 40% neural accelerator efficiency. Early simulation data shows NSS costs approximately 75% of Arm ASR's runtime in 1.5x upscaling (balanced mode). It is projected to outperform Arm ASR 2x upscaling (balanced mode). Efficiency gains come from replacing complex heuristics with a streamlined inference pass. Start building with NSS NSS introduces a practical, ML-powered approach to temporal super sampling. It replaces hand-tuned heuristics with learned filters and stability cues. It also runs within the real-time constraints of mobile hardware. Its training approach, compact architecture, and use of ML extensions for Vulkan make it performant and adaptable. For ML engineers building neural rendering solutions, NSS is a deployable, well-structured example of inference running inside the graphics pipeline. To explore the Arm Neural Graphics Development Kit, visit the NSS page on the Arm Developer Hub. There you can find sample code and review the network structure. We welcome feedback from developers using the SDK or retraining NSS for their own content. Your insights can help shape the future of neural rendering on mobile. Tags: AI ARM gaming GPU graphics Mobile upscaling Liam O'Neil (all posts) Liam O'Neil is a staff computer vision architect at Arm. Leave a Reply Cancel reply [ ] [ ] [ ] [ ] [ ] [ ] [ ] Comment * [ ] Name*[ ] (Note: This name will be displayed publicly) Email*[ ] (This will not be displayed publicly) [Post Comment] [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] Technical Papers * Performance And Energy Characterization Of A Commercial Compute-in-SRAM Device (Cornell, USC, MIT, GSI) September 16, 2025 by Technical Paper Link * Cost-Effective, Orthogonal Approach to Resilient Memory Design (Univ. of Central Florida, UT San Antonio, Rochester) September 16, 2025 by Technical Paper Link * HW/SW Co-Design to Retarget the Compiler For RISC-V Custom Instructions (Tampere Univ.) September 16, 2025 by Technical Paper Link * KAN Acceleration: Algorithm Hardware Co-Design Approach (Georgia Tech, National Tsing Hua Univ., TSMC) September 16, 2025 by Technical Paper Link * HW-SW Co-Designed System With 3 Core Optimization Pathways For Long-Context Agentic LLM Inference (Cambridge, ICL) September 15, 2025 by Technical Paper Link Knowledge Centers Entities, people and technologies explored Learn More Related Articles RISC-V's Increasing Influence Does the world need another CPU architecture when that no longer reflects the typical workload? Perhaps not, but it may need a bridge to get to where it needs to be. by Brian Bailey Can Cheaper Lasers Handle Short Distances? VCSELs may serve in more non-photonic applications. by Bryon Moyer Crisis Ahead: Power Consumption In AI Data Centers Four key areas where chips can help manage AI's insatiable power appetite. by Ed Sperling Can Today's Processor Architectures Be More Efficient? The low-hanging fruit of processor optimization may be gone, but new technologies are emerging. by Bryon Moyer When Can I Buy A Chiplet? A chiplet ecosystem is under development, but many barriers must be overcome before a thriving marketplace can exist. by Brian Bailey Will New Processor Architectures Raise Energy Efficiency? New approaches are needed as current processors run out of steam. by Bryon Moyer The Best DRAMs For Artificial Intelligence The choice of DRAM depends on where the action is. by Bryon Moyer The Evolution of DRAM How and why this tried-and-true memory is changing. by Ed Sperling * Sponsors [Siemens-lo] [se_sp_ramb] [se_sp_syno] [1_se_sp_ne] [se_sp_new-] [se_sp_cade] [expedera01] [iis_85mm_p] [qu] * * [INS::INS] Advertise with us * [INS::INS] Advertise with us * [INS::INS] Advertise with us * Newsletter Signup Popular Tags 2.5D 5G advanced packaging AI AMD ANSYS Apple Applied Materials ARM Arteris automotive business Cadence chiplets EDA eSilicon EUV finFETs GlobalFoundries Google IBM imec Infineon Intel IoT IP Keysight Lam Research machine learning memory Mentor Mentor Graphics MIT Nvidia NXP Qualcomm Rambus Samsung security SEMI Siemens Siemens EDA Synopsys TSMC verification Recent Comments * Cor on Glass Substrates Gain Momentum * Daya Young on New Antennas And Advanced ICs Needed For 6G * Subramanian Srikanteswara Iyer on Manufacturing At The Limits * Doug La Tulipe on Coloring Optical Signals For More Bandwidth In Data Centers * WH on The Hidden Cost Of Contact Resistance * Raymond Doerr on Metrology Under Pressure: Detecting Defects in Fine-Pitch Hybrid Bonding * Raymond Doerr on Manufacturing At The Limits * Mallikarjun on The Evolution of DRAM * Klein Miller on Democratizing Design: How The CHIPS Act Is Reshaping EDA And Semiconductor Innovation * Klein Miller on Lessons From 30 Years In The Trenches On The Future Of Semiconductor Manufacturing * Jonathan Kolber on Preparing For The Quantum Computing Age * Pramod Gupta on Challenges In Stacking HBM * Madeline Lueilwitz on Reconfigurable Single-Walled CNT FeFET (Univ. of Pennsylvania, Yonsei et al.) * Anne Meixner on Need For KGD Drives Singulated Die Screening * Capt kirk on Designing A Better Clock Network * Alex Martin on AI Effort And Money Misplaced * Hamilton Carter on AI Effort And Money Misplaced * Ilmu Komunikasi on Reconfigurable Single-Walled CNT FeFET (Univ. of Pennsylvania, Yonsei et al.) * Giovanni Lostumbo on Will New Processor Architectures Raise Energy Efficiency? * Brendon Berg on ESD Guns, Transients And Testing...Oh My! * GD on All About Interconnects * Raimo on Crisis Ahead: Power Consumption In AI Data Centers * Fred Chen on Laser-Focused Results: Improving EUV Line Edge Roughness With Ion Beam Etching * Fred Chen on Many Options For EUV Photoresists, No Clear Winner * Amba Prasad on Advanced Packaging Fundamentals for Semiconductor Engineers * Shakir Ullah on The Race To Replace Silicon * Hwaiyu Geng, P.E. on Accelerating Semiconductor Process Development Using Virtual Design Of Experiments * Larry Gus Christiansen on Reducing Risk In The Semiconductor Supply Chain * Vikas Sharma on RTL Signoff vs. Functional Signoff: What's The Difference? * Simone on Crisis Ahead: Power Consumption In AI Data Centers * Ron Lavallee on Can Today's Processor Architectures Be More Efficient? * Rashid on Can Today's Processor Architectures Be More Efficient? * Joseph Fjelstad on When Can I Buy A Chiplet? * Craig Lytle on Can Today's Processor Architectures Be More Efficient? * Partha Thirumalai on Cognichip: Using AI To Speed Complex Chip Design * Gopal raju on Advanced Packaging Fundamentals for Semiconductor Engineers * Lawrence Kushner on Democratizing Design: How The CHIPS Act Is Reshaping EDA And Semiconductor Innovation * Ken Rygler on Disruptive Changes Ahead For Photomasks? * Jan Hoppe on Need For KGD Drives Singulated Die Screening * B.S. DeepakSubramanyan on HBM Roadmap: Next-Gen High-Bandwidth Memory Architectures (KAIST's TERALAB) * Dr. Subhash L. Shinde , Booz Allen, Ex. Univ. of Notre Dame, Ex. Sandia National Labs., Ex. IBM. on The Race To Glass Substrates * Traian MUNTEAN (Honorary Professor) on A Balanced Approach To Verification * Amba Prasad on HBM Roadmap: Next-Gen High-Bandwidth Memory Architectures (KAIST's TERALAB) * sam on HBM Roadmap: Next-Gen High-Bandwidth Memory Architectures (KAIST's TERALAB) * Lentz on Agentic AI In Chip Design * adam on RISC-V's Increasing Influence * Jon Taylor on RISC-V's Increasing Influence * Roi Mit on The DAC Valuation * Janis Robin Schamberger on Chip Industry Week in Review * Rashid Kukkady on New Data Center Protocols Tackle AI * Peter Bennet on The DAC Valuation * Dan Ganousis on Optimizing Data Movement * DFTguy on Revolutionizing Semiconductor Development With GPU-Enhanced Atomistic Modeling * 1945x on Laser-Focused Results: Improving EUV Line Edge Roughness With Ion Beam Etching * Jason Kennerly on Many Options For EUV Photoresists, No Clear Winner * Marc Swinnen on Development Flows For Chiplets * Warren Savage on Development Flows For Chiplets * N. L. Kamaruzzaman on Security Risks Mount For Aerospace, Defense Applications * IT on Three-Way Race To 3D-ICs * Marzieh SalarRahimi on Advanced Packaging Fundamentals for Semiconductor Engineers * Harry Foster on Tape-Out Failures Are The Tip Of The Iceberg * Brian Bailey on Tape-Out Failures Are The Tip Of The Iceberg * Purple music on Tape-Out Failures Are The Tip Of The Iceberg * Messika on Tape-Out Failures Are The Tip Of The Iceberg * Al on Tape-Out Failures Are The Tip Of The Iceberg * Svetlana Morozova on GPU Analysis Identifying Performance Bottlenecks That Cause Throughput Plateaus In Large-Batch Inference * universitywafer on Advancements In Silicon Device Technology And Design Driving New SLM Monitor Categories * Dr. Dev Gupta on Packaging With Fewer People And Better Results * Michael Current on Advanced Packaging Fundamentals for Semiconductor Engineers * Jack G on 3D-IC For The Masses * Marc Greenberg on Implementing AI Activation Functions * Dr. D. on Big Changes Ahead For Interposers And Substrates * Dante on GPU Or ASIC For LLM Scale-Up? * Gretchen Patti on What Exactly Are Chiplets And Heterogeneous Integration? * Rob McCance on The Seven Pillars Of IC Package Physical Design * Lullaby on Times Are Changing For EDA * RF on Many Options For EUV Photoresists, No Clear Winner * ETechBuy on Experimental Characterization Results and State-of-the-Art Device-Level Studies of DRAM Read Disturbance * Fred Chen on Many Options For EUV Photoresists, No Clear Winner * Chris McMahon (Director FA @ Broadcom) on Failure To Launch * Alex Martin on Chiplets: A Technology, Not A Market * Mark Nakamoto on What Exactly Is Multi-Physics? * Andrew on Lines Blurring Between Supercomputing And HPC * Jung Yoon on EUV's Future Looks Even Brighter * Rishi Bhooshan on Signal Integrity Plays Increasingly Critical Role In Chiplet Design * Mark Nesselhaus on Non-Traditional Design of Dynamic Logic Gates and Circuits with FDSOI FETs * Fabio R. Pereira on Digital IC Bring-Up With A Bench-Top Environment * Schwaja on Chiplets: Where Are We Today? * Andrew Johnston on Strain, Stress In Advanced Packages Drives New Design Approaches * Paul Egan on What's Missing From Predictions * Dr Dev Gupta on Innovations Driving The Advanced Packaging Roadmap: Part One * Kevin Cameron on The High But Often Unnecessary Cost Of Coherence * Riko Radojcic on Strain, Stress In Advanced Packages Drives New Design Approaches * Raymond Doerr on Navigating Increased Complexity In Advanced Packaging * Fred Chen on Key Technologies To Extend EUV To 14 Angstroms * Rob Pearson - RIT on Shortcutting Graduates' Path To Productivity In Manufacturing And Test * Fred Chen on Is In-Memory Compute Still Alive? * Jesse KO on Advanced Packaging Drives Test And Metrology Innovations * Kumar on Auto Chip Aging Accelerates In Hot Climates * Dinesh Kumar on Redefining XPU Memory For AI Data Centers Through Custom HBM4: Part 1 * Ted Wilson on Goal-Driven AI * Art Scott (Earth ICT, SPC) on Goal-Driven AI * Santanu on Rethinking Engineering Education In The U.S. * JOH-POYO on FOPLP Gains Traction in Advanced Semiconductor Packaging * Dr. Dev Gupta on One Chip Vs. Many Chiplets * Gregory Johnson on Analysis Of Multi-Chiplet Package Designs And Requirements For Production Test Simplification * Theodore Wilson on Shift Left Is The Tip Of The Iceberg * Sanil Shankar on Is PPA Relevant Today? * MIHAI BUTA on Managing The Huge Power Demands Of AI Everywhere * A Duck on Batteries Look Beyond Lithium * atharva on 2D Semiconductors Make Progress, But So Does Silicon * Allan Cox on Big Changes Ahead For Analog Design * S on Hardware Acceleration Approach for KAN Via Algorithm-Hardware Co-Design * AI Research Scientist on Hardware Acceleration Approach for KAN Via Algorithm-Hardware Co-Design * Paul Karazuba on A Buyers Guide To An NPU * Ramesh Sharma on How Die Dimensions Challenge Assembly Processes * Gan Future on Week In Review: Design, Low Power * Roseann Johnson on Transitioning To Photonics * Mohammedd Fahad on A Power-First Approach * Bhavana Dhene on Promises and Perils of Parallel Test Marketplace T Benefits And Challenges Of Using... Ed Sperling Speeding Time To Market With A F... Baya Systems [se_logo_bl] About * About us * Contact us * Advertising on SemiEng * Newsletter SignUp Navigation * Homepage * Special Reports * Systems & Design * Low Power-High Perf * Manufacturing, Packaging & Materials * Test, Measurement & Analytics * Auto, Security & Enabling Technologies * Videos * Jobs * Technical Papers * Events * Webinars * Knowledge Centers * Industry Research * Business & Startups * Newsletters * Store Connect With Us * Facebook * Twitter @semiEngineering * LinkedIn * YouTube Copyright (c)2013-2025 SMG | Terms of Service | Privacy Policy This site uses cookies. By continuing to use our website, you consent to our Cookies Policy ACCEPT Manage consent Close Privacy Overview This website uses cookies to improve your experience while you navigate through the website. The cookies that are categorized as necessary are stored on your browser as they are essential for the working of basic functionalities of the website. We also use third-party cookies that help us analyze and understand how you use this website. We do not sell any personal information. By continuing to use our website, you consent to our Privacy Policy. If you access other websites using the links provided, please be aware they may have their own privacy policies, and we do not accept any responsibility or liability for these policies or for any personal data which may be collected through these sites. Please check these policies before you submit any personal information to these sites. Necessary [*] Necessary Always Enabled Necessary cookies are absolutely essential for the website to function properly. This category only includes cookies that ensures basic functionalities and security features of the website. These cookies do not store any personal information. Non-necessary [*] Non-necessary Any cookies that may not be particularly necessary for the website to function and is used specifically to collect user personal data via analytics, ads, other embedded contents are termed as non-necessary cookies. It is mandatory to procure user consent prior to running these cookies on your website. SAVE & ACCEPT Quantcast