NewsAI & DevelopmentCloud & DevOpsHardware

AWS and NVIDIA Lock In 2 Million More GPUs: What Developers Pay Next

AWS and NVIDIA data center with Blackwell Ultra and Rubin GPU infrastructure for agentic AI workloads
AWS is deploying 2 million additional NVIDIA GPUs across global infrastructure in 2027-2028

AWS just announced it is deploying 2 million additional NVIDIA GPUs across its global infrastructure in 2027-2028 — Blackwell Ultra, Rubin, and Rubin Ultra — bringing the total NVIDIA GPU capacity on AWS past 3 million. The prior 1 million GPU commitment, made at NVIDIA GTC 2026 in March, ran out early. TechCrunch framed it plainly: Amazon tripled its chip order because the previous one was not enough. The question for developers is not whether more supply is good. It is whether more supply means cheaper.

It does not. But there is an angle worth knowing about.

What Is Actually in the Deal

The official NVIDIA announcement covers more ground than a headline GPU count. The partnership expands across the full stack: GPUs, CPUs, networking, open models, and robotics.

On the GPU side, AWS gets Blackwell Ultra (current-gen B300 instances), Rubin (shipping 2026, 50 petaflops FP4 — double Blackwell), and Rubin Ultra (H2 2027, four-chiplet package, 100 petaflops per module). A Rubin Ultra Kyber rack packs 576 compute chiplets and delivers roughly 15 EFLOPS of FP4 inference at around 600 kW of power. These are not incremental numbers.

Alongside the GPU hardware, AWS and NVIDIA are building something more immediately useful for most developers: Nemotron open models on Amazon Bedrock and Amazon SageMaker. On Bedrock, Nemotron 3 Super runs serverless — no GPU infrastructure required, enterprise security included. On SageMaker, you can self-deploy and fine-tune. The pricing on Bedrock is $0.26 per million tokens. For reference, Claude Opus 4.6 runs at $3.85 per million tokens on the same platform. Nemotron 3 Super also hit 60.47% on SWE-Bench Verified, which puts it at the top of every open-weight coding model leaderboard.

The GPU Price Reality

Here is the part the press releases skip. AWS raised EC2 Capacity Block prices 20% on July 1, 2026, covering P6-B300, P6-B200, P5, P5e, P5en, and P4de. A Blackwell p6-b200.48xlarge now runs $98.84 per hour. That follows an earlier hike earlier in the year. Two consecutive price increases, on the same machines, during a period when AWS is also announcing record GPU commitments. That is not a contradiction — it is AWS treating constrained supply as a sustained pricing lever.

More GPUs arriving in 2027-2028 will not reverse that. Agentic AI workloads generate an order of magnitude more GPU demand than standard inference. Every new GPU gets absorbed. If you are waiting for spot pricing relief on H100 or Blackwell hardware, you are optimizing for the wrong variable.

What Developers Can Do Right Now

Two moves are available today, not in 2027.

First, benchmark Nemotron 3 Super on Bedrock for your agent inference layers. At $0.26 per million tokens with SWE-Bench performance at 60.47%, it deserves a real evaluation against your current model stack — especially for code generation and autonomous agent steps where cost compounds fast.

Second, if you run RAG pipelines or agent memory stores, test GPU-accelerated data processing on Amazon EMR with EC2 G7 instances and the NVIDIA cuDF library. AWS reports 3.7x faster processing and 30% better price performance versus CPU. For vector indexing on Amazon OpenSearch, GPU acceleration delivers 9x faster indexing at a quarter of the cost. These are live today.

What Is Coming in 2027

Two things matter on the 2027 roadmap. First, Rubin GPUs on AWS — the 50-petaflop-per-module architecture, double Blackwell — which will finally make very long-context inference and multi-step agent reasoning affordable at scale. Second, NVIDIA Vera CPU instances on EC2. Vera is built specifically for the CPU-bound work in agentic AI: code execution, tool calls, sandboxing, orchestration loops, and reinforcement learning rollouts. AWS is also integrating NVLink Fusion with its Annapurna Labs custom memory technology so that Trainium and NVIDIA GPU instances can share rack-scale architecture — reducing the network latency penalty between training and inference hardware in hybrid stacks.

The federal component: 100,000 GPUs on secure AWS infrastructure for U.S. government workloads at Impact Level 6 and above. Amazon Robotics is separately adopting NVIDIA Jetson, Omniverse, and Isaac for simulation and robot training — physical AI at warehouse scale.

The Bottom Line

AWS locked in 2 million more NVIDIA GPUs because the last million were not enough. Agentic AI demand outpaced projections in under a year. That tells you where infrastructure spending is heading and why price increases came before supply increases. The actionable move now is not to wait for cheaper GPUs — it is to use Nemotron on Bedrock for open-weight inference cost reduction and GPU-accelerated EMR and OpenSearch for data pipeline gains. The Vera CPU and Rubin hardware will matter in 2027. Plan for them. Do not wait on them.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *

    More in:News