
Salvatore Sanfilippo built Redis by refusing to be a general database. Now he has written a local LLM inference engine with the same logic: refuse to be general. DwarfStar 4 (ds4) runs DeepSeek V4 Flash on your MacBook at 26–39 tokens per second, requires no Python, no Docker, and no Ollama daemon, and won’t load any model it doesn’t ship. That last restriction is intentional — and it explains why it outperforms Ollama on the hardware it targets.
The Hardware Floor Is the First Thing to Know
DwarfStar 4 is not a tool for every developer. DeepSeek V4 Flash weighs roughly 81 GiB before context. You need a minimum of 96–128 GB of unified memory: an M2/M3/M5 Max or Ultra MacBook, a DGX Spark, or a high-end NVIDIA workstation. On a 64 GB machine, the smaller Qwen 3.8 Flash Next model works via SSD streaming, but DeepSeek V4 Flash requires at least 96 GB. Anything below 64 GB is not viable. If you don’t have the hardware, bookmark this and come back when you do — no shame in that.
If you do have the hardware, here is why it is worth your attention.
Two Technical Choices That Explain the Speed
Most local inference engines accept any GGUF file and apply generic quantization. DS4 ships its own weights and applies asymmetric 2-bit quantization tailored specifically to DeepSeek’s Mixture-of-Experts architecture. MoE models route each token through roughly 2 of 256 expert layers — the rest sit idle. DS4 compresses those rarely-active routed experts aggressively to 2-bit while keeping shared experts, attention, and routing layers at higher precision. The result: the model fits in 128 GB without the quality collapse that generic 2-bit compression typically causes. Tool calling works reliably at this compression level, which matters for anyone running coding agents.
The second innovation is a disk-resident KV cache. Standard inference engines keep the KV (key-value) cache in RAM; when you run out, you run out. DS4 treats the cache as a first-class disk citizen, persisting sessions to ~/.ds4/kvcache. Modern NVMe SSDs run at around 7 GB/s, fast enough to make disk-backed KV practical. On a 128 GB machine, this enables a 1-million-token context window. Community reports confirm 250K context working on 96 GB MacBooks. You resume long sessions via /switch without re-prefilling from scratch.
The practical outcome: on a MacBook Pro M3 Max with 128 GB, DS4 generates at roughly 27 tokens per second versus Ollama’s ~15 t/s on the same hardware and model — around 75% faster. On M5 Max hardware, generation hits 35–39 t/s.
Performance Numbers
| Machine | RAM | Prefill (t/s) | Generation (t/s) |
|---|---|---|---|
| M5 Max MacBook | 128 GB | 463–790 | 35–39 |
| M3 Max MacBook | 128 GB | 250 | ~27 |
| M3 Ultra Mac Studio | 512 GB | 468 | ~37 |
| DGX Spark GB10 | 128 GB | 825 | 14–18 |
Getting Started
Installation is four commands and a large download:
git clone https://github.com/antirez/ds4.git
cd ds4
make # macOS/Metal
./download_model.sh ds4f-q2 # ~81 GiB
For CUDA targets, use make cuda-spark (DGX Spark) or make cuda-generic (NVIDIA Ada Lovelace, L40S). Then run the interactive CLI or the API server:
./ds4 -p "Explain this function" # one-shot CLI
./ds4-server --ctx 32768 # OpenAI-compatible API on port 8000
Point your coding agent at it:
aider --openai-api-base http://localhost:8080/v1 \
--openai-api-key dummy \
--model deepseek-v4-flash
DS4’s HTTP server speaks both the OpenAI and Anthropic Messages APIs, so Claude Code and Codex CLI connect without modification. The full model catalog is available on Hugging Face under antirez/deepseek-v4-gguf.
DS4 vs. Ollama: When to Use Which
DS4 is not a replacement for Ollama. They solve different problems.
| Dimension | Ollama | DwarfStar 4 |
|---|---|---|
| Model support | Any GGUF | DeepSeek V4 Flash/PRO, GLM 5.x, Qwen 3.8 |
| Hardware floor | 8 GB+ | 96 GB+ |
| Installation | Daemon, easy | Build from source |
| Production-ready | Yes | Beta |
| Speed (DeepSeek V4 Flash) | ~15 t/s | ~27–39 t/s |
| 1M token context | No | Yes |
Use DS4 for unattended agentic tasks where you want peak performance from DeepSeek V4 Flash: overnight code refactors, long documentation passes, multi-hour reasoning sessions where API costs add up. Use Ollama if you need model flexibility, stable production behavior, or hardware below the 96 GB floor.
Know What You’re Getting Into
antirez is transparent about the trade-offs, and you should factor them in before committing. macOS CPU mode triggers kernel panics — GPU is required. The software is explicitly beta; expect API changes between releases. Distributed mode lacks encryption and authentication, so keep it on trusted networks. DeepSeek V4 PRO support is experimental and limited to rare 512 GB hardware. The codebase was built with AI assistance, which antirez discloses in the README.
An independent hands-on review calls DS4 “the most interesting local inference project” of its launch window, with the caveat that it is not a universal solution. That framing holds: deeply good for one workload, wrong tool for everything else.
The Bigger Point
“AI is too critical to be just a provided service,” antirez wrote when launching DS4. The whole design follows from that sentence. You don’t pick a database for its model variety — you pick it because it’s the right tool and you know what it’s doing under the hood. DS4 applies that reasoning to local inference. Whether or not you have the hardware today, it’s a useful template for what opinionated infrastructure looks like in the AI era.
The GitHub repo is the canonical starting point. antirez also credits the llama.cpp and GGML community explicitly — worth reading if you want to understand the ecosystem DS4 builds on.













