
Broadcom reported $16.7 billion in AI chip revenue for Q3 2026 — a 221% year-over-year increase — and 73% of that came from custom XPUs rather than off-the-shelf silicon. Six hyperscalers are now designing their own AI chips through Broadcom’s XPU program, including OpenAI, Anthropic, Google, Meta, ByteDance, and Fujitsu. The GPU monoculture Nvidia built over two decades is not collapsing, but it is cracking — and this quarter’s numbers make clear the pace is accelerating.
The Numbers First
Total Q3 revenue hit $29.6 billion, beating analyst estimates, with operating income at a record $20.1 billion. Broadcom raised its full-year AI guidance to $58 billion and issued formal long-term targets: $115 billion for FY2027 and $230 billion for FY2028. CEO Hock Tan described these as based on “secured supply and conservative deployment assumptions” — which is either admirably cautious or a signal that real demand runs higher.
For context: if Broadcom hits $230 billion in AI revenue by FY2028, that single segment would rival AMD’s entire annual revenue today. The trajectory is not subtle.
The Six Customers and What They’re Building
Broadcom’s XPU program currently serves six customers: Google, Meta, OpenAI, Anthropic, ByteDance, and Fujitsu. These are not passive chip buyers — they are co-designing custom AI accelerators with Broadcom, optimized for their specific model architectures and workloads.
Two chips stand out. OpenAI’s Jalapeño, unveiled in June 2026 and detailed by TechCrunch, is a purpose-built inference ASIC developed with Broadcom. OpenAI claims it delivers roughly 50% lower cost per inference token compared to current Nvidia GPUs, with initial deployment targeting end of 2026 and 1.3 gigawatts of capacity planned for 2027. The chip is specifically tuned for running large language models in production — not training them.
On the Anthropic side, Broadcom’s Hock Tan confirmed on the Q3 earnings call that Anthropic will become the company’s largest XPU customer by 2027. Anthropic has already committed to 1 gigawatt of compute capacity in 2026, expanding to 5 GW in 2027 with potential to reach 10 GW by 2028, powered largely by Google’s TPU Ironwood racks that Broadcom co-designs.
Why Inference Is the Only Battleground That Matters Right Now
Training large models remains Nvidia’s territory. The flexibility of CUDA, the breadth of framework support, and two decades of library optimization make Nvidia GPUs the only practical choice for research and foundational model development. Nobody is arguing that point.
Inference is different. Once a model is trained and stable, the compute pattern becomes predictable and repetitive — serving billions of requests running the same architecture. Custom ASICs are purpose-built for exactly this workload. The cost advantage is real: when Midjourney shifted from Nvidia GPUs to Google TPUs for image inference, monthly compute costs dropped from $2.1 million to roughly $700,000 — a 65% reduction. Inference now accounts for approximately two-thirds of all AI compute cycles. That is the market Broadcom and its six customers are targeting.
What This Means for Developers
The practical split is clean. If you consume AI through APIs — calling OpenAI’s endpoints, using Claude, querying Gemini — the underlying infrastructure shift to custom silicon is invisible to you in the short term. But the economics of Jalapeño and Ironwood should flow downstream as cheaper inference pricing over the next 12 to 24 months. You benefit without changing anything.
If you self-host models, you remain on Nvidia. Custom silicon is a hyperscaler game. Building your own XPU requires committing to proprietary toolchains — Google’s XLA for TPUs, Amazon’s Neuron SDK for Trainium, Microsoft’s custom stack for Maia. Migrating a serving stack from vLLM to one of these environments typically takes two to six weeks, ongoing maintenance as these SDKs lag behind general frameworks, and complete loss of hardware portability. That trade-off only makes economic sense at massive scale.
CUDA’s moat is real and will remain real for anyone running diverse or experimental workloads. Five million active developers and two decades of optimization do not evaporate because Broadcom had a good quarter.
The Broader Shift
Nvidia is not standing still. Vera Rubin, its next-generation architecture, promises a 10x reduction in cost per generated token over Blackwell. After the earnings release, Citi reduced its Nvidia GPU sales estimates by $12 billion — a dent, not a collapse. The more interesting long-term risk for Broadcom comes from the other direction: as Google, Meta, and eventually OpenAI and Anthropic mature in chip design, they may progressively reduce their dependence on external partners entirely.
For now, Broadcom sits in an enviable position — the chip architect for the most demanding AI workloads on the planet, with a $230 billion revenue roadmap backed by committed customers. The GPU monoculture is cracking. The question is whether Broadcom ultimately benefits from the shift or becomes another vendor that hyperscalers outgrow. Given that Anthropic alone is on track to be Broadcom’s largest customer within 18 months, that question will not stay hypothetical for long.













