On September 14, Fujitsu officially launched FUJITSU-MONAKA: a 144-core ARMv9 server CPU built on a 2nm compute die stacked with 5nm cache and I/O chiplets, with sales starting in November 2026. The chip targets AI inference workloads directly — hardware matrix-multiplication acceleration, dual 256-bit SVE2 vector units per core, 12 DDR5 channels at 8,800 MT/s, and PCIe Gen6. Fujitsu claims the Fujitsu MONAKA CPU achieves twice the AI inference throughput of competing CPUs, potentially halving the server count for equivalent workloads. Independent benchmarks don’t exist yet, but the technical architecture earns serious scrutiny.
What Fujitsu MONAKA Delivers for AI Inference
MONAKA uses a 3D die-to-die hybrid bonding approach that combines a TSMC N2P compute die (under 30% of silicon area) with N5 SRAM and I/O dies. The result is high-bandwidth compute-to-cache connectivity that distinguishes it from simpler monolithic designs. Fujitsu is offering two SKUs: a 500W high-performance variant at 2.9GHz base (liquid-cooled) and a 350W high-efficiency variant at 2.1GHz base (air-cooled). Both include the same matrix-multiplication hardware acceleration targeting the GEMM operations that dominate transformer inference.
Three NUMA configurations are available — single 144-core node, four 36-core nodes, or eight 18-core nodes — giving operators flexibility to match memory locality to workload. Server form factors are 1U and 2U rack configurations. According to the Fujitsu official announcement, the processor and server systems target data centers, enterprises, HPC, and defense sectors in Japan and Europe from November 2026, with broader availability in 2027. The ServeTheHome Hot Chips 2026 analysis notes the 3D hybrid bonding approach as a genuine architectural differentiator, not a marketing repackaging of standard chiplet assembly.
Why ARM Server CPUs Are Dominating AI Inference
The inference economics have shifted dramatically. Two years ago, AI infrastructure spending was training-dominated. In 2026, inference consumes roughly 80% of AI infrastructure budgets. GPUs remain the performance ceiling for inference, but for a growing class of workloads — mid-size model serving, high-concurrency inference clusters, cost-sensitive deployments — CPU-native inference is genuinely competitive. According to IDC’s Q1 2026 data, ARM has already overtaken x86 in accelerated servers. Hyperscalers proved the model first — AWS Graviton5 (192 cores), Google Axion (3x better MLPerf than x86), and Microsoft Cobalt 100 (1.9x better LLM inference than x86) have all demonstrated that ARM server CPUs win on inference per watt at scale.
However, every one of those wins is cloud-only. You cannot run Graviton5 in your own data center. For organizations in Japan, Europe, defense, and regulated industries that need on-premises AI infrastructure with data residency guarantees, MONAKA fills a gap the hyperscalers structurally cannot. That is not a niche — it is a market that EU AI Act requirements and Japan’s AI governance frameworks are actively making larger.
Related: PyTorch 2.14: Fault-Tolerant Training and Apple Silicon Fixes
The Sovereign AI Infrastructure Case
Fujitsu frames MONAKA explicitly as sovereign AI infrastructure. The chip is designed, developed, and manufactured in Japan at Fujitsu’s Kasashima Plant, with supply chain traceability as a product feature — not just a marketing point. The Arm CCA (Confidential Compute Architecture) support adds hardware-encrypted memory protection, enabling AI workloads on sensitive data — a hard requirement in healthcare, financial services, and government deployments where GPU-based cloud inference is a compliance non-starter.
Furthermore, the successor chip, Monaka-X on TSMC 1.4nm, will add NVLink Fusion support for tighter CPU-GPU integration. That road map signals Fujitsu is not positioning MONAKA as a GPU replacement — it is positioning it as the on-premises inference substrate into which GPUs can connect when high-performance headroom is needed. It is a coherent architecture strategy, not a one-off product.
The Question That Matters: Software Readiness
Hardware claims mean nothing without a usable software stack. Fujitsu has been working on this longer than most hardware announcements suggest. Fujitsu Research India has contributed to PyTorch (including a dedicated INT8 inference paper optimizing PyTorch 2 Export Quantization for MONAKA’s ARM target), ONNX Runtime, TensorFlow, llama.cpp, oneDNN, and OpenBLAS. GCC 15 includes MONAKA-specific compiler support contributed upstream by Fujitsu engineers. The team presented at SC24 and the PyTorch Conference 2025 — both before the hardware was commercially available.
That is more pre-work than most ARM server CPU launches have received. Nevertheless, no public SDK exists and no independent benchmarks have been published. The Graviton ecosystem took years to mature after the hardware shipped. Fujitsu started earlier, but November 2026 is still a short runway from announcement to production-quality inference pipelines. The real test is whether llama.cpp and ONNX Runtime deliver fully optimized code paths at launch, or whether production adoption lags the hardware by a year or more — as it did with earlier ARM server entrants.
Key Takeaways
- FUJITSU-MONAKA ships November 2026: 144 ARMv9 cores, 2nm + 5nm chiplet design, hardware matrix-multiply acceleration, two power SKUs (350W / 500W), PCIe Gen6, 12 DDR5 channels at 8,800 MT/s.
- ARM has overtaken x86 in accelerated servers (IDC Q1 2026); inference is now 80% of AI infrastructure spend — MONAKA enters a validated and growing market.
- The sovereign AI angle is its real differentiator: no hyperscaler ARM server runs on-premises, and MONAKA targets that exact gap for Japan, Europe, and regulated sectors with supply chain transparency and Arm CCA confidential computing.
- Software ecosystem preparation predates the launch — PyTorch PT2E, ONNX Runtime, llama.cpp, GCC 15 contributions — but independent benchmarks and a production-ready SDK are still missing.
- Watch for: confirmed llama.cpp and ONNX Runtime support at or before the November launch; independent MLPerf submissions; pricing disclosure; whether the 500W liquid-cooled SKU’s 2x throughput claim holds up in third-party testing.













