{"schemaVersion":"1.0","type":"Article","types":["Article"],"slug":"opentpu-an-open-source-ai-accelerator-developed-by-ai-4s5a0","url":"https://zyvop.com/opentpu-an-open-source-ai-accelerator-developed-by-ai-4s5a0","title":"OpenTPU: An Open-Source AI Accelerator, Developed by AI","subtitle":"A tour of openTPU, a small open-source LLM inference accelerator where AI agents helped tune the hardware, from RTL and ISA to compiler and profiler.","tldr":"openTPU is a small, readable LLM decode accelerator with its entire stack in one repo: RTL, ISA, simulator, compiler and profiler. See how AI agents tuned the hardware, what the numbers show, and what's still unproven.","keywords":["open-source hardware","LLM inference","AI agents","AI accelerator","FPGA","AI News"],"entities":["Arpan Singh","open-source hardware","LLM inference","AI agents","AI accelerator","FPGA","AI News","ZyVOP"],"keyTakeaways":["A tour of FeSens/openTPU, a small and readable LLM inference accelerator, from RTL to compiler, with AI agents in the design loop.","AI chips are usually black boxes.","You get a datasheet, a few benchmark slides and maybe a driver."],"headings":["The whole stack, top to bottom","Every layer is checked against the one below","A deliberately simple machine","Real models, not toys","The numbers","Where the AI comes in: the auto-arch tournament","Seeing what the hardware is doing: Lens","What is not done, and why that is fine","Why it is worth your attention","Try it yourself","What to watch next"],"outboundLinks":["https://github.com/FeSens/openTPU"],"contentText":"A tour of FeSens/openTPU, a small and readable LLM inference accelerator, from RTL to compiler, with AI agents in the design loop. AI chips are usually black boxes. You get a datasheet, a few benchmark slides and maybe a driver. Almost nobody gets to open the whole stack, change one layer, and watch what happens to the cycle count. openTPU is built to be the opposite. It is a small inference accelerator for LLM decode, and everything lives in one repository: SystemVerilog RTL, an instruction set, a bit-exact simulator, a small kernel language with a compiler, a profiler, and host software for a Kintex-7 PCIe card. The stated goal is educational: learn about hardware, software and ISA performance work by changing one layer and measuring the effect on cycles. It is also an experiment in how much of that work AI can do, which is where this project gets interesting. One thing up front: this is a simulation-only prototype. It has not run on an FPGA yet, and I will be clear about which claims are measured and which are projections. The whole stack, top to bottom Most accelerator projects stop at the RTL. openTPU keeps going in both directions. At the top, you write kernels in a small Python-embedded language called ol. A compiler turns them into instructions, handling layouts, affine loop addressing, fusion peepholes and bank-aware strides. Each instruction is eight 32-bit words. Those instructions run on two implementations of the same machine: a Python ISA simulator and the SystemVerilog RTL under Verilator. Below that sit a host runtime and CLI, talking over PCIe (via the stock Xilinx XDMA driver) to a YPCB-00338 board with an xc7k480t FPGA and two DDR3 SODIMM channels. Here is what a kernel looks like. This is the MLP block, simplified: from opentpu import language as ol @ol.jit def mlp(h, gamma, w_gate, w_up, w_down, out, eps): x = ol.load(h) xs = ol.quantize(rmsnorm(x, ol.load(gamma), eps)) g = ol.dot(xs, w_gate) u = ol.dot(xs, w_up) a = ol.all_gather(silu(g) * u) y = ol.all_gather(ol.dot(a, w_down)) if ol.program_id() == 0: ol.store(out, x + y)The same kernel can be launched on either backend by changing one argument, \"isa\" or \"rtl\". That is the heart of the project's discipline. Every layer is checked against the one below Kernels are tested against float64 numpy references. The simulator's fp32 arithmetic is checked against a Python model of the RTL's floating-point units. And the RTL has to produce exactly the same DRAM and on-chip memory bits as the simulator. That last check is stricter than it sounds. The Verilator RTL matches the simulator on the kernel tests, on a full Qwen3-0.6B token (6.38 million cycles), and on random programs where instructions keep conflicting over the same memory. The random programs matter most: they test whether the hardware keeps instructions in the right order while running units in parallel. To make bit-exact agreement possible, the numerics are deliberately constrained. Weights, matrix-unit activations and the KV cache are int8 with one fp32 scale per block. Everything else is fp32 with round-to-nearest-even and flush-to-zero. Functions like exp2, reciprocal and rsqrt are fixed sequences of adds and multiplies, so Python, the simulator and the RTL agree bit for bit. A deliberately simple machine The design is small enough to hold in your head: a matrix unit that streams int8 weights from DRAM, a vector unit for fp32 math, a quantizer, a DMA engine, a collective unit for multi-slice runs. On-chip memories are explicit, and there are hardware loops. A sequencer issues one instruction per cycle into a 16-entry window and tracks what each instruction reads and writes. An instruction starts as soon as nothing older conflicts with it, so units overlap without the compiler having to schedule them by hand. Because every data movement is an instruction you can see in the trace, a profile can usually tell you why something is slow. That is a big part of the educational value. Real models, not toys The simulator runs real weights, not just microbenchmarks: Qwen3-0.6B, in int8. LFM2.5-230M, Liquid AI's hybrid of short-convolution and attention layers. Its 64-wide heads are zero-padded to the 128-deep matrix unit, and its convolution state lives in DRAM. No ISA or RTL change was needed. Qwen3.5-0.8B, whose main layer is a Gated DeltaNet, a form of linear attention with a 128 x 128 fp32 state per head. The state streams through the scratchpad head by head and the recurrence runs on the vector unit. That last point is a good sign for the design. Three quite different architectures ran on the same small instruction set, which suggests the ISA is general enough to be useful rather than hand-fitted to one model. The numbers All figures below come from Verilator RTL simulation of the board configuration (one slice, a 128-deep matrix unit with 2 columns, 8 vector lanes), using Qwen3-0.6B's shapes. \"Of roofline\" compares a kernel's cycles to the cycles needed just to move its bytes over the simulated DRAM port. MLP decode: 75,282 cycles, 98.4% of roofline. Flash attention, context 1024: 27,835 cycles, 64.4% of roofline. Flash attention, context 4096: 104,571 cycles, 67.1% of roofline. Full attention layer (norm, QKV, RoPE, KV append, attention, output) at position 1023: 81,362 cycles, 82.3% of roofline. The MLP is limited by DRAM, which is exactly where you want a decode accelerator to be. Attention is not. Both the matrix unit and the vector unit are busy nearly every cycle (100% and 99% at context 1024) while DRAM streams only 64% of the time. The matrix unit spends that time on per-row work between streams, and the vector unit on the softmax. The README flags this as the obvious next target. For whole-model throughput, one Qwen3-0.6B decode token takes about 6.4 million cycles. At an assumed 100 MHz clock that is roughly 15 tokens per second. LFM2.5-230M comes out around 42 tokens per second and Qwen3.5-0.8B around 11, under the same assumptions. These are projections from simulated cycles, not measurements from hardware. Where the AI comes in: the auto-arch tournament The most concrete example of AI-driven development in the repo is tools/tourney, an automated hill climb inspired by the auto-arch-tournament project. It works on one RTL component at a time, and each round looks like this: A few LLM agents read the current component, its critical path and a log of what earlier rounds tried. Each agent writes a hypothesis, for example \"register the TMEM read data in front of the prescale multiplier,\" and another agent implements it in its own worktree. Every candidate must pass lint, the bit-exact RTL-vs-simulator tests on two micro-architectures, a Qwen3 decode cycle-count check and the kernel performance tests. The candidate is synthesized with yosys and kept only if it is smaller at the same estimated speed, or faster at the same size. The verification gates are what make this safe. An agent can propose anything, but a change only survives if the RTL still matches the simulator bit for bit and the performance tests still pass. The agents get creative freedom, and the test suite acts as the referee. The first overnight run tried 96 changes across six components and kept 38. With some manual fixes between units, the yosys estimate for the accelerator logic went from 41 MHz to 106 MHz, and from 132K to 82K LUTs. All logs and patches are committed under tools/tourney/runs/, so you can read every hypothesis, including the failures. There is a caveat the authors are upfront about: these are yosys estimates for the accelerator and control logic only. They leave out the PCIe and DDR3 controllers, which add roughly 45K LUTs, and they say nothing certain about what Vivado will achieve after place and route. Seeing what the hardware is doing: Lens Performance work needs a profiler, and openTPU ships one. Lens records a run, whether an RTL cycle trace, a simulator run or, on the card, the hardware trace buffer, and opens it in the browser. You get a roofline overview, a zoomable timeline, per-instruction and per-source-line tables, and a floorplan replay where colors show what each unit is doing each cycle: busy, stalled on DRAM, lost TMEM arbitration, or waiting on a dependency. python3 -m opentpu.lens record mlp attn -o run.otpuprof python3 -m opentpu.lens open run.otpuprofA test checks that the hardware trace buffer rebuilds the same trace lines the simulator prints, so the tool is not just a pretty picture. What is not done, and why that is fine The README is refreshingly direct about the gaps: No hardware run. No bitstream has been built, and the Vivado build has not been done. Clock speed is an estimate. The 106 MHz figure comes from yosys, not from place and route. Board model shortcuts. DDR3 calibration always succeeds, the DDR3 controllers are replaced by an AXI memory model, and PCIe, clocks and resets are not simulated. The authors expect problems on first contact with a real board. Output drifts from the reference. Because the model runs in W8A8 (8-bit weights and activations), greedy generations match Hugging Face's fp32 Qwen3-0.6B exactly for 16 tokens on only 2 of 8 test prompts. The others diverge after 1 to 10 tokens, typically where two candidate tokens are nearly tied. A float64 model with the same 8-bit rounding picks the same tokens as the device where checked, which points to quantization rather than a bug. I like this honesty. A project that tells you which numbers are projections is one whose measured numbers you can trust more. Why it is worth your attention Three reasons stand out. It is readable. Most accelerator codebases are either proprietary or sprawling. This one is small on purpose, with a design spec and per-layer docs for the ISA, compiler, profiler and board. It closes the loop between layers. Change a kernel, a compiler pass, an instruction or a block of RTL, and the same test and profiling infrastructure tells you what happened to the cycles. It shows a workable pattern for AI-assisted hardware design. The pattern is to let agents propose and implement changes, but make a strict, automated, bit-exact test suite decide what survives. Hardware has unforgiving correctness requirements, and this approach respects them. Try it yourself You need Python 3.11+, numpy and pytest. The RTL tests also need Verilator 5 and skip themselves without it. The model tests need torch and transformers, plus the model checkpoints in the models/ directory. python3 -m pytest -q make tourney COMP=otpu_vpu N=5 K=2The first command runs the test suite. The second runs five tournament rounds with two candidates each on the vector unit. The code is licensed under Apache 2.0. What to watch next The obvious milestones are the first real Vivado build, a first boot on the YPCB-00338 card, and a fix for the attention kernel's gap to roofline. When real clock speeds and real tokens per second replace the projections, we will find out how much of the simulation story survives contact with hardware. Either way, openTPU is already a useful reference for anyone curious about how an LLM accelerator fits together, and for anyone wondering how far AI agents can go when they are given a good referee.","contentHash":"sha256:b26158f7c558d453fac72d9c08e33550d0b49b67b10a6dbc593d6e17cc855650","authorName":"Arpan Singh","authorUrl":"https://zyvop.com/author/arpan","authorSameAs":[],"category":"AI News","tags":["open-source hardware","LLM inference","AI agents","AI accelerator","FPGA"],"audience":"Software engineers and developers building applications with AI News","tone":"Practical and evidence-based engineering guidance","readingTimeMinutes":9,"wordCount":1898,"faqs":null,"primaryTopic":"AI News","publishedAt":"2026-10-07T05:32:08.916Z","updatedAt":"2026-10-07T05:32:08.916Z","canonicalUrl":"https://zyvop.com/opentpu-an-open-source-ai-accelerator-developed-by-ai-4s5a0"}