The AI accelerator is the clearest piece of in-house design work in the SoC. It is not a purchased IP block: it has its own register interface, its own weight-loading path, dedicated weight and bias memories, and two independent design iterations.
How it connects
The accelerator sits on the AXI4-Lite fabric like any other peripheral, with a register interface for control and status. Weights arrive over a dedicated second UART (UART1) rather than competing with the general-purpose console on UART0 — model data streaming in does not block or corrupt ordinary firmware I/O.
It became part of the MCU at demo_v10, the milestone where the
accelerator was integrated into the SoC rather than exercised on its own.
Two iterations
The first core was built around a serial adder. The second
(ai_accelerator2) is built around DSP units. The
second exists because we measured the first and went back to improve it
— an internally-driven optimisation effort, and one of the parts of the
project we are most willing to be judged on.