https://fuse.wikichip.org/news/6855/arm-unveils-next-gen-flagship-core-cortex-x3/ Skip to content Wednesday, June 29, 2022 Latest: * Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency * Arm Unveils Next-Gen Flagship Core: Cortex-X3 * Arm Introduces The Cortex-A715 * GlobalWafers To Build A 1.2M WPM Factory In Sherman, Texas * A Look At Samsung's 4LPE Process * WikiChip Fuse WikiChip Fuse Your Chips and Semi News [ ] * Home * Account * Main Site * Architectures + x86 + ARM + RISC-V + Power ISA + MIPS * Supercomputers * 14 nm * 12nm * 10nm * 7nm * 5nm Mobile Processors Arm Unveils Next-Gen Flagship Core: Cortex-X3 June 28, 2022June 28, 2022 David Schor ARM, Cortex, Cortex-A, Cortex-X, Cortex-X2, Cortex-X3 [cortex-x3-header] Today arm is introducing its new next-generation flagship core, the Cortex-X3. The Cortex-X3 CPU is a third-generation core as part of the Cortex-X custom core program designed to bring higher performance through slightly different PPA tradeoffs versus the mainstream performance big core. --------------------------------------------------------------------- This article is part of a series of articles covering Arm's Client Tech Day 2022. * Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency * Arm Introduces The Cortex-A715 * Arm Unveils Next-Gen Flagship Core: Cortex-X3 --------------------------------------------------------------------- In terms of performance, the Cortex-X3 is said to deliver an 11% IPC uplift which Arm says marks the third consecutive year of double-digit IPC gain on the Cortex-X series. On the system level, coupled with all other changes, in real applications Arm says the new core can deliver as much as 22% higher performance. Microarchitectural Changes According to Chris Abernathy, the core's chief architect. The process of eliminating 32-bit and optimizing for the 64-bit ISA exclusively has been a 2-step process. With the Cortex-X2, the underlying circuitry used for handling 32-bit architectural-related elements was removed, saving on transistors and simplifying some structures. With the new Cortex-X3, the design team took the time to start optimizing specifically for AArch64. In particular, many optimizations to fetch and decode took place taking advantage of the more predictable and regular nature of the AArch64 ISA. To that end, the majority of design changes took place to the Cortex-X3 front-end which made targeted improvements to the branching mechanism and optimized for large applications with larger instruction footprints The Cortex-X3, like its predecessors, features a decoupled front end. For Arm, this means that the branch predictors operate at a much large bandwidth than the instruction fetch and can run considerably ahead of fetch. This allows the branch predictors to double up as instruction prefetchers. This gives them a number of advantages. For one, they are able to absorb much of the latency by making the instruction stream fetch requests into the L2 and L3 far ahead. This helps absorb pipeline-taken branch bubbles. It also means the effective L1 caches can be kept small, saving on power and far more importantly saving on silicon area. As we have seen in more recent advanced nodes, SRAM bitcells are no longer shrinking nearly as much as they used to. Keeping those structures small is now more important than before. In the X3, Arm increased the run-ahead window depth, allowing the core to fetch instruction much further out in time. From the X2 to the X3, the branch target buffers are said to grow significantly in capacity to store more targets (over 50% growth in the L1 and L2 BTBs). It's worth noting that due to the increase in capacity, the X3 actually introduces an L2 BTB to break down the larger single structure with more incremental latency. There are now effectively three levels - L0, L1, and L2 - each with increased latency over the lower level. The quick, single-cycle, turnaround L0 BTB is now 10x larger in capacity which is pretty significant growth. [arm-cortex-x3-fe] Plotting the predicted taken branch on the Cortex-X2 vs the Cortex-X3, Arm says they see an average of 12.2% reduction in cycles due to better branch taken prediction across various real-world workloads. Overall, they also demonstrated around a 3% reduction in total cycles as resulting from front-end stalls such as a branch taken bubble or branch misses. [arm-cortex-x3-fe-graph] Another big improvement in the Cortex-X3 was dealing with indirect branches. Indirect branch prediction was said to be given a first-class treatment in the new Cortex-X3. What this entails is a brand new small, but dedicated indirect branch predictor designed to improve accuracy and enable lower latency. The TAGE predictor used for conditional branches also received some incremental improvement. Overall, the X3 is said to deliver an average reduction of 6.1% in branch mispredict over the X2. [arm-cortex-x3-branch] Instruction Fetch On the instruction fetch side of the code, the MOP cache also had various improvements. Arm modified the fill algorithms in order to minimize the thrashing and pollution. This change is said to allow them to reduce its size by half without hurting its performance. Previously, the MOP cache had 3K-entry capacity. This has been reduced back to 1.5K which is actually the same size as the original A77 was when the MOP was first introduced. It is worth noting that despite the reduced capacity, Arm is maintaining the 8 MOPs/cycle bandwidth as with the prior generation. On the performance side of things, partially due to its size reduction, the MOP cache on the Cortex-X3 is now a single pipe stage shorter - going from ten to nine now. [76-77-78-x1-x3-decode-comp] The MOP cache is designed for small, tight, repeatedly executing code. For larger applications that do not fit in the MOP cache, the X2 could still easily be bottlenecked by the narrower instruction cache. On the X3, Arm increased the fetch and decode bandwidth. It's now possible to decode up to six instructions per cycle, a 1.2x increase in bandwidth for higher sustained IPC. [arm-cortex-x3-fe-mop] OoO and Execution The new core has a larger OoO window - up to 320 entries which can translate to up to 640 instructions in flight in a best-case scenario. Due to a large number of in-flight instructions, following a flush, rebuilding the rename tables needed further optimizations. Here Arm redesigned the renaming-rebuild mechanics to better handle a large number of instructions in-flight. [arm-cortex-x3-be] Memory Subsystem There were other smaller memory subsystem changes. The integer load bandwidth has increased from 24B/cycle to 32B/cycle. There are no specific dedicated load and store queues per se on the Cortex-X3, but collectively those structures grew by 25%. As with every generation, Arm continued to improve its data access pattern predictors. In the X3, Arm added two additional data prefetchers. The first one is to handle spatial patterns. The other is a pointer/indirect prefetch that deals with pointer chasing sequences. [arm-cortex-x3-memsys1] Performance All in all, the Cortex-X3 is said to deliver a geomean IPC improvement of around 11% on a set of real-world applications. [arm-cortex-x3-ipc] Cache Sensitivity One interesting aspect of the Cortext-X3 is the choice of L2 cache size. This can be 1 MiB or 512 KiB. The choice between the two boils down to area versus performance tradeoff. With the larger 1 MiB L2 cache, the core benefits from both higher performance and lower system-level power due to less L3 thrashing. Overall the larger cache can see up to a 26.9% reduction in refill/writebacks requests. [arm-cortex-x3-memsen] - Spotted an error? Help us fix it! Simply select the problematic text and press Ctrl+Enter to notify us. - * - Arm Introduces The Cortex-A715 * Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency - Share This Post: Related Articles Huawei Expands Kunpeng Server CPUs, Plans SMT, SVE For Next Gen Alibaba Open Source XuanTie RISC-V Cores, Introduces In-House Armv9 Server Chip Samsung Discloses Exynos M4 Changes, Upgrades Support for ARMv8.2, Rearranges The Back-End Ampere Ships First Gen ARM Server Processors Arm Introduces Its Confidential Compute Architecture Samsung quietly unveils their latest flagship processor: Exynos 9 Series 9810 Top Six Articles * Arm Unveils Next-Gen Flagship Core: Cortex-X3 * Arm Introduces The Cortex-A715 * Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency * A Look At Intel 4 Process Technology * GlobalWafers To Build A 1.2M WPM Factory In Sherman, Texas * A Look At Samsung's 4LPE Process Ezoicreport this ad Recent * Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency June 28, 2022June 28, 2022 David Schor * Arm Unveils Next-Gen Flagship Core: Cortex-X3 Arm Unveils Next-Gen Flagship Core: Cortex-X3 June 28, 2022June 28, 2022 David Schor * Arm Introduces The Cortex-A715 Arm Introduces The Cortex-A715 June 28, 2022June 29, 2022 David Schor * GlobalWafers To Build A 1.2M WPM Factory In Sherman, Texas GlobalWafers To Build A 1.2M WPM Factory In Sherman, Texas June 27, 2022June 27, 2022 David Schor * A Look At Samsung's 4LPE Process A Look At Samsung's 4LPE Process June 26, 2022June 26, 2022 David Schor * A Look At Intel 4 Process Technology A Look At Intel 4 Process Technology June 19, 2022June 20, 2022 David Schor Ezoicreport this ad Random Picks ISSCC 2018: MIT's low-power hardware crypto RISC-V IoT processor ISSCC 2018: MIT's low-power hardware crypto RISC-V IoT processor March 17, 2018May 25, 2021 David Schor TSMC 7nm HD and HP Cells, 2nd Gen 7nm, And The Snapdragon 855 DTCO TSMC 7nm HD and HP Cells, 2nd Gen 7nm, And The Snapdragon 855 DTCO June 16, 2019May 25, 2021 David Schor Chuck Peddle: Personal Computer Pioneer, Dies At 82 Chuck Peddle: Personal Computer Pioneer, Dies At 82 December 24, 2019May 25, 2021 David Schor Cambricon Reaches for the Cloud With a Custom AI Accelerator, Talks 7nm IPs Cambricon Reaches for the Cloud With a Custom AI Accelerator, Talks 7nm IPs May 26, 2018May 25, 2021 David Schor 7nm Boosted Zen 2 Capabilities but Doubled the Challenges 7nm Boosted Zen 2 Capabilities but Doubled the Challenges February 21, 2020May 25, 2021 David Schor Random Tags 2.5D packaging 3D packaging 5 nm 5nm 7 nm 7nm 10 nm 10nm 14 nm 16nm AI AMD ARM ARMv8 ARMv9 chiplet Coffee Lake Core i5 Core i7 Cortex edge computing EMIB EUV FinFET GlobalFoundries Hot Chips IBM Ice Lake IEDM inference Intel Intel 7 ISSCC multi-chip package neural processors process technology RISC-V Samsung subscriber only (general) Sunny Cove Supercomputers TSMC VLSI Symposium x86 Zen x86 WorldView All Intel Introduces Thread Director For Heterogeneous Multi-Core Workload Scheduling Desktop Processors Mobile Processors Intel Introduces Thread Director For Heterogeneous Multi-Core Workload Scheduling August 19, 2021August 19, 2021 David Schor Intel introduces the Intel Thread Director for heterogeneous multi-core workload scheduling Intel Unveils Sapphire Rapids: Next-Generation Server CPUs Architectures Server Processors Intel Unveils Sapphire Rapids: Next-Generation Server CPUs August 19, 2021August 19, 2021 David Schor Intel's Gracemont Small Core Eclipses Last-Gen Big Core Performance Architectures Data Processing Unit Desktop Processors Mobile Processors Intel's Gracemont Small Core Eclipses Last-Gen Big Core Performance August 19, 2021August 21, 2021 David Schor Intel Unveils Alder Lake: Next-Generation Mainstream Heterogeneous Multi-Core SoC Architectures Desktop Processors Mobile Processors Intel Unveils Alder Lake: Next-Generation Mainstream Heterogeneous Multi-Core SoC August 19, 2021August 19, 2021 David Schor Intel Details Golden Cove: Next-Generation Big Core For Client and Server SoCs Architectures Desktop Processors Mobile Processors Server Processors Intel Details Golden Cove: Next-Generation Big Core For Client and Server SoCs August 19, 2021August 19, 2021 David Schor Intel Launches 3rd Gen Ice Lake Xeon Scalable Architectures Server Processors Intel Launches 3rd Gen Ice Lake Xeon Scalable April 6, 2021May 23, 2021 David Schor Random Intel's Spring Crest NNP-L Initial Details Intel's Spring Crest NNP-L Initial Details April 14, 2019May 25, 2021 David Schor Cambricon Reaches for the Cloud With a Custom AI Accelerator, Talks 7nm IPs Cambricon Reaches for the Cloud With a Custom AI Accelerator, Talks 7nm IPs May 26, 2018May 25, 2021 David Schor ISSCC 2018: AMD's Zeppelin; Multi-chip routing and packaging ISSCC 2018: AMD's Zeppelin; Multi-chip routing and packaging March 24, 2018May 25, 2021 David Schor AMD Discloses Initial Zen 2 Details AMD Discloses Initial Zen 2 Details November 18, 2018May 25, 2021 David Schor Eni fires up its supercomputer, breaks into the TOP500's top ten Eni fires up its supercomputer, breaks into the TOP500's top ten January 19, 2018May 25, 2021 David Schor Samsung Discloses Exynos M4 Changes, Upgrades Support for ARMv8.2, Rearranges The Back-End Samsung Discloses Exynos M4 Changes, Upgrades Support for ARMv8.2, Rearranges The Back-End January 14, 2019May 25, 2021 David Schor TSMC Starts 5-Nanometer Risk Production TSMC Starts 5-Nanometer Risk Production April 6, 2019May 25, 2021 David Schor ARM WorldView All Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency Architectures Mobile Processors Arm Refreshes The Cortex-A510, Squeezes Higher Efficiency June 28, 2022June 28, 2022 David Schor Arm Unveils Next-Gen Flagship Core: Cortex-X3 Mobile Processors Arm Unveils Next-Gen Flagship Core: Cortex-X3 June 28, 2022June 28, 2022 David Schor Arm Introduces The Cortex-A715 Architectures Mobile Processors Arm Introduces The Cortex-A715 June 28, 2022June 29, 2022 David Schor Alibaba Open Source XuanTie RISC-V Cores, Introduces In-House Armv9 Server Chip Architectures Server Processors Alibaba Open Source XuanTie RISC-V Cores, Introduces In-House Armv9 Server Chip October 20, 2021October 20, 2021 David Schor Marvell Launches 5nm Octeon 10 DPUs with Neoverse N2 cores, AI Acceleration Data Processing Unit Marvell Launches 5nm Octeon 10 DPUs with Neoverse N2 cores, AI Acceleration June 28, 2021June 28, 2021 David Schor Arm Introduces Its Confidential Compute Architecture Architectures Arm Introduces Its Confidential Compute Architecture June 23, 2021June 23, 2021 David Schor About WikiChip WikiChip is an independent publisher based in New York. The WikiChip Fuse section publishes chips and semiconductor related news with our main site offering in-depth semiconductor resources and analysis. WikiChip Links * Main Site * WikiChip Fuse * Newsletter * * Main Site * WikiChip Fuse Copyright (c) 2022 WikiChip LLC. All rights reserved. Spelling error report The following text will be sent to our editors: Your comment (optional): [ ] [ ] [ ] Send Cancel