Hacker News AI · 10/6/2026, 4:23:25 PM
OpenTPU: AI-Designed Open-Source FPGA Accelerator
OpenTPU is an open-source AI accelerator designed with the involvement of AI agents, exploring the limits of automated hardware design. Running on a Xilinx Kintex-7 FPGA platform, it successfully executes modern models like LFM2.5 and Qwen3.5, achieving bit-exact inference results matching its simulator. The project serves as a full-stack learning resource from high-level Python operations to low-level circuitry, demonstrating the feasibility of AI building its own inference chips.
SOURCE COVERAGEOriginal coverage
Contents2 sections
openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?

otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card's
utilization and DRAM bandwidth (right).
A place to learn
openTPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.
Results
The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.
| Model | Weights | Decode, device | Decode, wall | Prefill, device | DRAM while decoding |
|---|---|---|---|---|---|
| LFM2.5-230M | int8 | 59.0 tok/s | 52.3 tok/s | 295.6 tok/s | 14.5 GB/s (85% of peak) |
| LFM2.5-230M | 4-bit, int8 head | 85.8 tok/s | 82.1 tok/s | 335.4 tok/s | 14.1 GB/s (82%) |
| Qwen3-0.6B | int8 | 21.6 tok/s | 21.3 tok/s | 92.1 tok/s | 14.4 GB/s (84%) |
| Qwen3-0.6B | 4-bit, int8 head | 31.3 tok/s | 30.7 tok/s | 103.4 tok/s | 13.9 GB/s (82%) |
| Qwen3.5-0.8B | int8 | 17.6 tok/s | 16.3 tok/s | 61.4 tok/s | 14.5 GB/s (85%) |
| Qwen3.5-0.8B | 4-bit, int8 head | 24.5 tok/s | 23.3 tok/s | 66.7 tok/s | 14.1 GB/s (83%) |
| Gemma 4 E2B | 4-bit, int8 head | 10.57 tok/s | 10.53 tok/s | 32.1 tok/s | 15.6 GB/s (92%) |
| Gemma 4 E2B | 4-bit, 4-bit head | 12.14 tok/s | 12.09 tok/s | 29.9 tok/s | 15.5 GB/s (91%) |
| LFM2-2.6B | int8 | 6.05 tok/s | 6.03 tok/s | 21.4 tok/s | 16.1 GB/s (94%) |
| LFM2-2.6B | 4-bit, int8 head | 10.96 tok/s | 10.93 tok/s | 20.6 tok/s | 15.8 GB/s (93%) |
| SmolLM3-3B | int8 | 5.00 tok/s | 4.99 tok/s | 21.1 tok/s | 16.0 GB/s (94%) |
| SmolLM3-3B | 4-bit, int8 head | 8.74 tok/s | 8.72 tok/s | 22.8 tok/s | 15.7 GB/s (92%) |
| Phi-4-mini (3.8B) | int8 | 3.99 tok/s | 3.98 tok/s | 13.8 tok/s | 16.0 GB/s (94%) |
| Phi-4-mini (3.8B) | 4-bit, int8 head | 6.56 tok/s | 6.55 tok/s | 15.0 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-2B | int8 | 8.02 tok/s | 8.00 tok/s | 38.2 tok/s | 16.0 GB/s (94%) |
| Qwen3.5-2B | 4-bit, int8 head | 12.09 tok/s | 12.03 tok/s | 41.7 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-4B | 4-bit, int8 head | 5.88 tok/s | 5.87 tok/s | 12.9 tok/s | 15.7 GB/s (92%) |
| Gemma 4 E4B | int8, 4-bit head and down 0-23 | 3.78 tok/s | 3.75 tok/s | 14.8 tok/s | 16.0 GB/s (94%) |
Measured on the card: the first three models on 2026-09-29 with the production image
deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and
4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then.
Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%),
at 91-94% of the DRAM peak instead of 82-87%. Qwen3.5-4B's int8 image is over 4 GiB.
- The image: main e698dcd at 133.33 MHz, one bitstream for all models. It has LiteDRAM controllers calibrated by a small CPU inside the memory core, a four-column systolic matrix unit and the stream engine (docs/stream.md). DDR3-1066, with a 17.1 GB/s peak.
- The host: the card sits in opentpu (Intel Core i7-4790).
- Method,
tools/qual/perf.py: decode is 64 greedy tokens after a 512-token prompt, with the host's argmax in the loop (not streamed). "Device" counts only the cycles the accelerator runs; "wall" adds the host. Prefill is the 512-token prompt, on the device. - DRAM traffic comes from the card's own counters while it runs.
- Gemma 4 E2B keeps its per-layer embedding tables on the card (3.5-3.6 GiB images; docs/gemma4.md); in int8 it does not fit. It matches Hugging Face's greedy tokens on three prompts with either head. E4B's table (2.95 GB) stays on the host, which copies one 11 KB row into the card per token; its image is 3.96 GiB, int8 with the head and the first 24 layers' down projections in 4-bit (docs/gemma4_e4b.md). In the card's own decode loop (the card picking every token) Gemma 4 decodes faster: E2B 11.01 / 12.73 tok/s (int8 / 4-bit head), E4B 3.83 tok/s, on the device.
- Every configuration matches the simulator token for token, per-position and with the resident decode program. More detail in docs/board.md.
With the logits streamed back while the card runs (tools/decode_profile.py, 96 tokens), 4-bit
decode is faster, in device / wall tok/s:
- LFM2: 89.5 / 84.5;
- Qwen3: 33.7 / 33.3;
- Qwen3.5: 24.6 / 24.2;
- LFM2-2.6B: 11.07 / 11.02 (build B);
- SmolLM3-3B: 8.92 / 8.89 (build B);
- Phi-4-mini: 6.69 / 6.67 (build B).
The previous production image, se-cand3, was built with the Xilinx MIG, a two-column matrix unit and a 120.755 MHz clock. Measured the same way, the new image:
- decode: within 2.3% of se-cand3's in every configuration. Decode is bound by DRAM, and LiteDRAM reads at 82-85% of the DDR3 peak, as the MIG did.
- prefill: 1.3x (Qwen3.5) to 2.0x (LFM2 4-bit) faster.
- calibration: when the image starts, the core's CPU calibrates both DDR3 channels in 12 s, with no host involvement.
The earlier images and their numbers are in docs/board.md, section 5.
Mixture-of-experts models bigger than the card's 4 GiB run with their experts streamed from host storage (docs/offload.md, section 10). The card routes each token and computes every expert, and it keeps the experts in per-layer slots in its DRAM. The host only copies missing experts from a pool file into those slots, at the link's rate (section 10.1). Measured on 2026-10-01 with build B (79c5707a), the card's own decode loop picking every token, 4-bit experts, int8 head:
- LFM2.5-8B-A1B (8.5B parameters, 1.7B active): 10.6 tok/s over 160 tokens. 98.5% of expert uses hit the slots, and 5.2 MB streamed per token.
- Qwen3.5-35B-A3B (34.7B parameters, 3.0B active): 3.95 tok/s, with Hugging Face's 16 greedy tokens. 62%
(注:更多技术实现细节与完整 API 文档请查阅原项目 README)