NVIDIA shipped PAIR (Personal AI Router) as a free public beta in September — a routing layer that discovers every compatible GPU on your local network and distributes Ollama and LM Studio inference requests across them. Tomorrow, October 7, Jensen Huang takes the stage at Microsoft’s Windows and Surface event in San Francisco to spotlight it alongside Satya Nadella. If you run multi-agent AI workloads locally, PAIR is worth understanding now. It’s live, free, and open-source on GitHub.
What NVIDIA PAIR Does — and the Misconception You Need to Clear Up
PAIR is a request router, not a memory pool. Before diving into setup, this distinction needs to be explicit: PAIR routes each inference request to one complete node. It does not combine two 24GB GPUs into a 48GB machine. It does not shard a 70B model across three machines. According to NVIDIA’s official documentation, “PAIR does not pool GPU memory, combine GPUs into a larger logical GPU, shard one model across machines, or split an in-flight inference request between nodes.”
The problem PAIR actually solves is different — and more common for developers building with agents. When you run a five-sub-agent workflow, all five fire inference requests roughly simultaneously. On a single machine, requests queue. The second agent waits for the first, the third waits for the second. PAIR distributes those independent requests across available machines, so three sub-agents can run in parallel across three different GPUs instead of serializing on one.
Related: Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed
Getting Started: It’s One Port Change Away
Install PAIR on each participating machine — Windows .exe, Linux .deb, or macOS .dmg installers are available. PAIR uses mDNS to discover nodes on the network and mTLS encryption for inter-node communication. Once nodes are paired and Ollama is running, you change one thing in your agent code:
# Before: direct Ollama
client = openai.OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
# After: PAIR proxy — same interface, now distributed across the cluster
client = openai.OpenAI(base_url="http://localhost:11435/v1", api_key="ollama")
PAIR exposes Ollama-compatible, OpenAI-compatible, and Anthropic Messages API endpoints. Existing agent code works without modification — you’re just changing the target port. That zero-friction integration is what separates PAIR from alternatives like Petals or Mesh LLM, which require actual architectural changes to your inference stack.
The Benchmark: 2x Faster — Under the Right Conditions
NVIDIA’s own demo used a five-sub-agent Hermes task running Qwen 3.6 35B on Ollama. A single RTX Spark laptop completed the workflow in 18 minutes on average. A three-machine cluster — RTX Spark, DGX Spark, and an RTX 5090 — finished the same task in 8 minutes 48 seconds. According to NVIDIA’s IFA 2026 announcement, this is “configuration-specific” and not a universal benchmark.
The honest conditions for that speedup: your workload must be parallel. Sequential tasks — where each step waits for the previous result — get no benefit from routing. The other caveat developers underestimate is storage. A 20GB model replicated across three nodes requires 60GB of total cluster storage. Independent analysis from XenoSpectrum found that model placement is the single most critical variable — if the model doesn’t exist on a node, PAIR won’t route to it.
Hardware Support Is Broader Than the RTX Spark Coverage Suggests
PAIR supports NVIDIA GeForce RTX 20-series and newer on Windows, Linux, and macOS. It also supports Apple M4+ silicon. That means a cluster can mix a Windows RTX 5090 workstation, a Linux dev box with an older RTX 3080, and a MacBook Pro M4 — all routing to the same proxy endpoint. Windows, Linux, and macOS nodes can pair together in the same cluster, which is more flexible than most teams assume.
However, PAIR is still a public beta. Production workloads should treat it accordingly. For teams running large enough models to exhaust single-GPU VRAM, Petals or Mesh LLM remain the correct tools — PAIR doesn’t solve the model-size problem, it solves the concurrency problem.
Key Takeaways
- PAIR routes parallel inference requests across machines — it does not pool VRAM or shard models across GPUs
- Setup requires one configuration change: point your agent’s base_url from Ollama’s port to PAIR’s proxy endpoint
- Expect roughly 2x throughput gains on parallel multi-agent workloads; sequential tasks get no benefit
- Supported hardware includes NVIDIA RTX 20-series+, DGX Spark, and Apple M4+ silicon across all major operating systems
- PAIR is in public beta — production-critical workloads should wait for a stable release













