The core
At the centre of the SoC is the CV32E40P — a 32-bit, in-order RISC-V core from the OpenHW Group. We did not write the core ourselves; the engineering work is in everything around it: the bus fabric, the address decoding, the peripheral set, the memory subsystem, and the three implementation flows it is carried through.
The core issues memory-mapped reads and writes across an AXI4-Lite
interconnect. A central address_decoder selects the target
peripheral from the transaction address, so adding a peripheral means
adding a decode range and an AXI4-Lite slave — not touching the core.
Interconnect
The axi4 fabric is an in-house AXI4-Lite implementation. Every
peripheral presents the same slave interface, which is what makes the
peripheral set uniform enough to verify one block at a time and then
integrate with confidence. The AXI4 bus additionally has a UVM environment
built around it — the only part of the design verified at that level of
formality, because it is the part every other block depends on.
Module map
Every module in the design, and what it does:
| Module | Description |
|---|---|
cv32e40p |
32-bit RISC-V in-order core (OpenHW Group) |
axi4 |
AXI4-Lite interconnect fabric |
address_decoder |
Memory-mapped peripheral address decoding |
gpio |
General-purpose I/O peripheral |
timer |
Hardware timer peripheral |
uart (UART0 / UART1) |
General-purpose UART, plus a second UART dedicated to AI accelerator data streaming |
i2c |
I2C master / controller |
qspi / flash |
Quad-SPI controller and flash memory interface |
imem / dmem / bootrom |
Instruction memory, data memory, and boot ROM |
jtag |
Debug access |
ivc |
Interrupt vector controller — vectoring and dispatch for the peripheral interrupts |
hdmi |
HDMI framebuffer output (Gowin and open-source IP variants) |
ai_accelerator / ai_accelerator2 |
Custom neural-network inference accelerator — serial-adder core, weight/bias memories, DSP-based variant, UART-fed loader |
The AI accelerator
The accelerator is the clearest example of in-house work in the design. It has its own register interface, a UART-based weight-loading path, dedicated weight and bias ROMs, and two independent design iterations: a serial-adder-based core and a later DSP-unit-based core. The second exists because we measured the first and went back to improve it — it is an internally-driven optimisation effort, not a purchased IP block.
Read more about the accelerator →
Software
Firmware is bare-metal C and RISC-V assembly built against a custom HAL and
linker scripts. Each peripheral has a matching driver
(gpio.c, i2c.c, qspi.c,
timer.c, uart.c, ai.c,
ivc.c, oled.c) and a HAL header. A dedicated boot
loader brings the core up before jumping to firmware.
Instruction-level correctness is checked independently against the spike RISC-V ISA simulator, using a checker script plus directed and basic ISA test suites — so the core's behaviour in our SoC is compared against a reference implementation, not just against our own expectations.
Implementation
The same RTL is carried through three implementation targets: an Altera Agilex 3 FPGA via Quartus Prime, a Gowin FPGA, and an open-source RTL-to-GDS flow via OpenLane, with completed hardened runs for both a standalone MCU and the full top-level SoC. Simulation runs in Vivado xsim.