A developer called FeSens dropped openTPU on GitHub this week: a full-stack AI inference accelerator designed entirely by AI agents. The project includes SystemVerilog RTL, a custom instruction set, a bit-exact Python simulator, a kernel compiler, a profiler, and host utilities — all in one repo, all open source under Apache 2.0. It runs on a Xilinx Kintex-7 FPGA that costs about $300 used. Hacker News gave it 337 points and 393 comments inside 24 hours. The real story is the question baked into the project: can AI agents design the hardware that runs their own inference?
The Recursive Loop
This isn’t FeSens’ first AI-designs-hardware experiment. The auto-arch-tournament ran AI agents against a RISC-V CPU: 73 hypotheses, 9 hours 51 minutes of wall-clock time, and a +92% performance improvement over the baseline with 40% fewer LUTs. openTPU extends that concept to AI accelerators specifically. The AI agents didn’t just write the RTL — they designed the ISA, built the compiler, and targeted hardware that runs the very models the agents themselves use.
It’s a strange loop. Whether that framing holds up to scrutiny is a fair question, but the output is real and the code compiles.
The Architecture Is Designed to Be Read
Most accelerator designs obscure complexity behind caches and out-of-order schedulers. openTPU does the opposite. There is no cache. There is no hidden scheduling. Every data movement is an explicit instruction, so a performance trace shows exactly where every cycle goes. The README puts it plainly: “if you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.”
The hardware is four units: a matrix unit running a 4-column systolic array for int8 weight multiplication, a vector unit for fp32 normalization and activation functions, a DMA engine managing explicit data transfers, and a quantizer converting fp32 results back to int8. Kernels are written in a Python DSL with an @ol.jit decorator that compiles to hardware instructions. The ISA is 8×32-bit words, one instruction per cycle. Bit-exact equivalence between the simulator and the RTL is verified by the test suite.
What It Actually Runs
The system targets small models — the Kintex-7 board carries 4 GiB of DDR3 with 17.1 GB/s peak bandwidth. DRAM is the bottleneck throughout (82–94% utilization), not compute. That means performance scales with memory bandwidth, not clock speed.
| Model | Decode (tok/s) | Prefill (tok/s) |
|---|---|---|
| LFM2.5-230M (4-bit) | 85.8 | 335.4 |
| Qwen3-0.6B (4-bit) | 30.7 | 103.4 |
| Qwen3.5-2B (4-bit) | 12.0 | 41.7 |
| Phi-4-mini 3.8B (4-bit) | 6.6 | 15.0 |
Gemma 4 variants, SmolLM3-3B, and two mixture-of-experts configurations are also supported, with expert streaming for models that exceed the card’s memory.
The Model Obsolescence Argument (and Why It’s Incomplete)
The loudest Hacker News criticism was pointed: model SOTA moves faster than chips can be designed or fabricated. By the time you burn a model into FPGA bitstream or ASIC silicon, the model is outdated. It’s a real concern.
But AMD just paid real money to acquire Taalas, a startup that etches specific model weights directly into silicon. Taalas’ HC1 chip runs Llama 8B at 17,000 tokens per second. AMD is betting that the edge use case — offline, private, low-latency, fixed-model inference — is commercially viable even as frontier models race forward. A year-old model running at interactive speeds on a $300 board, with no cloud dependency and no data leaving the device, is a different product category than frontier model access. Not better or worse — different.
How to Use It Now
Getting started doesn’t require hardware. Clone the repo, install Python 3 with PyTorch and Transformers, and run the software simulator. Verilator 5 adds RTL simulation. The full FPGA experience requires a Vivado license and a Kintex-7 board — the Inspur YPCB-00338 is the tested target, available used for around $300.
The host tooling is practical: otpu-chat for interactive inference, otpu-smi for real-time monitoring, and otpu-lens for a browser-based performance profiler that visualizes exactly where cycles are spent. Apache 2.0 means you can fork the RTL, change the ISA, and run your own hardware experiments.
openTPU won’t replace your GPU cluster. Its output token rates on 3B+ models are slow by cloud standards. What it offers instead is something most developers have never had: a complete, readable, modifiable AI accelerator stack where every design decision is explicit and every behavior is testable. That’s a rarer thing than faster tokens per second.













