openTPU: an open-source AI accelerator 'developed by AI' runs Qwen3.5, Gemma 4 and others on an FPGA card
openTPU (github.com/FeSens/openTPU, Apache-2.0) is a full inference accelerator stack (SystemVerilog RTL, ISA, bit-exact simulator, kernel compiler and host software) that its author says was developed by AI agents through an automated improvement loop. It runs ten small open models with real weights on a Kintex-7 FPGA PCIe card, up to ~86 tokens/s on LFM2.5-230M, with the card producing the same tokens as the simulator bit for bit. It reached the Hacker News front page on Oct 6, 2026 (~259 points). The models used are not named.
Key facts
- Repo: github.com/FeSens/openTPU, created 24 Sept 2026, Apache License 2.0; ~270 stars on 7 Oct; tagline 'An open-source AI accelerator, developed by AI'
- Hardware: Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels) at 133.33 MHz; four-column systolic matrix unit; decode is DRAM-bound at 82–94% of DDR3-1066 peak
- Results (README): LFM2.5-230M up to 85.8 tok/s decode (4-bit); Qwen3-0.6B 31.3 tok/s; Qwen3.5-0.8B 24.5 tok/s; Gemma 4 E2B ~12 tok/s; Phi-4-mini 6.6 tok/s; MoE models streamed over PCIe (Qwen3.5-35B-A3B at 3.95 tok/s, LFM2.5-8B-A1B at 10.6 tok/s)
- Stated goal (README): 'how far can AI agents go at hardware design, and can they build the chip that runs their own inference?'; it applies lessons from the author's earlier auto-arch-tournament project (AI-developed RISC-V cores)
- Author on HN (fsbonetto): 'The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models'; a 'tournament of Vivado runs' keeps optimising timing and area
- Not stated: which AI models or agents did the work, or how much human design input there was
- Hacker News: ~259 points, ~316 comments (6 Oct 2026)
What happened
On 6 October 2026 an open-source project called openTPU reached the Hacker News front page. In one repository it contains an inference accelerator design in SystemVerilog, its instruction set, a bit-exact Python simulator that serves as the specification, a kernel language and compiler, and host tools (otpu-chat, otpu-smi) that drive an FPGA PCIe card. The README reports measured decode speeds for ten small open models, from Qwen3 and Qwen3.5 to Gemma 4, LFM2/2.5, SmolLM3 and Phi-4-mini, with exact token agreement between card and simulator.
The author (GitHub FeSens, posting as fsbonetto) presents it as hardware "developed by AI". He says the same agent technique had earlier been used to develop RISC-V cores, and that a "recursive self improvement loop" took the design from a few tokens per second to over 80. The repository does not name the models or agents used, and the human share of the design is not documented.
Why it matters
It is a small but concrete public example of AI agents carrying a whole hardware stack, from RTL to drivers, to a working FPGA card that runs current open models, with an automated performance loop. On HN, one top comment stressed how much is gained from "an experienced user pointing an LLM in a tasteful direction", and others joked about "recursive self-improvement" arriving as a hobby FPGA project. The claim about AI authorship is self-reported.
Changelog
- 2026-10-07: created (sweep 2026-10-07, Hacker News section)
Sources (3)
- codeGitHub: FeSens/openTPU
- codeGitHub: FeSens/auto-arch-tournament (earlier AI-developed CPU cores)
- discussionHacker News discussion
id: 2026-10-06-opentpu-ai-developed-accelerator · updated 2026-10-07 · open in the interactive timeline