Today for AI

Hacker News AI · 2026/10/6 16:23:25

OpenTPU 开源:AI 代理设计 FPGA 加速器

原标题:OpenTPU – An open-source AI accelerator, developed by AI
78AI 研判分
核心综述

OpenTPU 是一个由 AI 代理参与设计的开源 AI 加速器项目,旨在探索硬件自动化设计的极限。该项目基于 Xilinx Kintex-7 FPGA 平台,成功在真实硬件上运行 LFM2.5 和 Qwen3.5 等现代大模型,并实现了与模拟器比特级一致的推理结果。其核心价值在于提供了一个从 Python 矩阵乘法到底层电路的全栈学习资源,展示了 AI 自主构建自身推理芯片的可行性。

报道全文原始报道全文

本文目录2 个章节

openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?

otpu-chat on LFM2.5-230M with otpu-smi watching the card

otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card's utilization and DRAM bandwidth (right).

A place to learn

openTPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.

Results

The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.

ModelWeightsDecode, deviceDecode, wallPrefill, deviceDRAM while decoding
LFM2.5-230Mint859.0 tok/s52.3 tok/s295.6 tok/s14.5 GB/s (85% of peak)
LFM2.5-230M4-bit, int8 head85.8 tok/s82.1 tok/s335.4 tok/s14.1 GB/s (82%)
Qwen3-0.6Bint821.6 tok/s21.3 tok/s92.1 tok/s14.4 GB/s (84%)
Qwen3-0.6B4-bit, int8 head31.3 tok/s30.7 tok/s103.4 tok/s13.9 GB/s (82%)
Qwen3.5-0.8Bint817.6 tok/s16.3 tok/s61.4 tok/s14.5 GB/s (85%)
Qwen3.5-0.8B4-bit, int8 head24.5 tok/s23.3 tok/s66.7 tok/s14.1 GB/s (83%)
Gemma 4 E2B4-bit, int8 head10.57 tok/s10.53 tok/s32.1 tok/s15.6 GB/s (92%)
Gemma 4 E2B4-bit, 4-bit head12.14 tok/s12.09 tok/s29.9 tok/s15.5 GB/s (91%)
LFM2-2.6Bint86.05 tok/s6.03 tok/s21.4 tok/s16.1 GB/s (94%)
LFM2-2.6B4-bit, int8 head10.96 tok/s10.93 tok/s20.6 tok/s15.8 GB/s (93%)
SmolLM3-3Bint85.00 tok/s4.99 tok/s21.1 tok/s16.0 GB/s (94%)
SmolLM3-3B4-bit, int8 head8.74 tok/s8.72 tok/s22.8 tok/s15.7 GB/s (92%)
Phi-4-mini (3.8B)int83.99 tok/s3.98 tok/s13.8 tok/s16.0 GB/s (94%)
Phi-4-mini (3.8B)4-bit, int8 head6.56 tok/s6.55 tok/s15.0 tok/s15.8 GB/s (92%)
Qwen3.5-2Bint88.02 tok/s8.00 tok/s38.2 tok/s16.0 GB/s (94%)
Qwen3.5-2B4-bit, int8 head12.09 tok/s12.03 tok/s41.7 tok/s15.8 GB/s (92%)
Qwen3.5-4B4-bit, int8 head5.88 tok/s5.87 tok/s12.9 tok/s15.7 GB/s (92%)
Gemma 4 E4Bint8, 4-bit head and down 0-233.78 tok/s3.75 tok/s14.8 tok/s16.0 GB/s (94%)

Measured on the card: the first three models on 2026-09-29 with the production image deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and 4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then. Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%), at 91-94% of the DRAM peak instead of 82-87%. Qwen3.5-4B's int8 image is over 4 GiB.

  • The image: main e698dcd at 133.33 MHz, one bitstream for all models. It has LiteDRAM controllers calibrated by a small CPU inside the memory core, a four-column systolic matrix unit and the stream engine (docs/stream.md). DDR3-1066, with a 17.1 GB/s peak.
  • The host: the card sits in opentpu (Intel Core i7-4790).
  • Method, tools/qual/perf.py: decode is 64 greedy tokens after a 512-token prompt, with the host's argmax in the loop (not streamed). "Device" counts only the cycles the accelerator runs; "wall" adds the host. Prefill is the 512-token prompt, on the device.
  • DRAM traffic comes from the card's own counters while it runs.
  • Gemma 4 E2B keeps its per-layer embedding tables on the card (3.5-3.6 GiB images; docs/gemma4.md); in int8 it does not fit. It matches Hugging Face's greedy tokens on three prompts with either head. E4B's table (2.95 GB) stays on the host, which copies one 11 KB row into the card per token; its image is 3.96 GiB, int8 with the head and the first 24 layers' down projections in 4-bit (docs/gemma4_e4b.md). In the card's own decode loop (the card picking every token) Gemma 4 decodes faster: E2B 11.01 / 12.73 tok/s (int8 / 4-bit head), E4B 3.83 tok/s, on the device.
  • Every configuration matches the simulator token for token, per-position and with the resident decode program. More detail in docs/board.md.

With the logits streamed back while the card runs (tools/decode_profile.py, 96 tokens), 4-bit decode is faster, in device / wall tok/s:

  • LFM2: 89.5 / 84.5;
  • Qwen3: 33.7 / 33.3;
  • Qwen3.5: 24.6 / 24.2;
  • LFM2-2.6B: 11.07 / 11.02 (build B);
  • SmolLM3-3B: 8.92 / 8.89 (build B);
  • Phi-4-mini: 6.69 / 6.67 (build B).

The previous production image, se-cand3, was built with the Xilinx MIG, a two-column matrix unit and a 120.755 MHz clock. Measured the same way, the new image:

  • decode: within 2.3% of se-cand3's in every configuration. Decode is bound by DRAM, and LiteDRAM reads at 82-85% of the DDR3 peak, as the MIG did.
  • prefill: 1.3x (Qwen3.5) to 2.0x (LFM2 4-bit) faster.
  • calibration: when the image starts, the core's CPU calibrates both DDR3 channels in 12 s, with no host involvement.

The earlier images and their numbers are in docs/board.md, section 5.

Mixture-of-experts models bigger than the card's 4 GiB run with their experts streamed from host storage (docs/offload.md, section 10). The card routes each token and computes every expert, and it keeps the experts in per-layer slots in its DRAM. The host only copies missing experts from a pool file into those slots, at the link's rate (section 10.1). Measured on 2026-10-01 with build B (79c5707a), the card's own decode loop picking every token, 4-bit experts, int8 head:

  • LFM2.5-8B-A1B (8.5B parameters, 1.7B active): 10.6 tok/s over 160 tokens. 98.5% of expert uses hit the slots, and 5.2 MB streamed per token.
  • Qwen3.5-35B-A3B (34.7B parameters, 3.0B active): 3.95 tok/s, with Hugging Face's 16 greedy tokens. 62%

(注:更多技术实现细节与完整 API 文档请查阅原项目 README)