SoC Architecture

A 32-bit RISC-V core issuing memory-mapped reads and writes over an in-house AXI4-Lite fabric to a bank of custom peripherals, written in synthesizable SystemVerilog and carried through two FPGA flows and an ASIC flow.

The core

At the centre of the SoC is the CV32E40P — a 32-bit, in-order RISC-V core from the OpenHW Group. We did not write the core ourselves; the engineering work is in everything around it: the bus fabric, the address decoding, the peripheral set, the memory subsystem, and the three implementation flows it is carried through.

The core issues memory-mapped reads and writes across an AXI4-Lite interconnect. A central address_decoder selects the target peripheral from the transaction address, so adding a peripheral means adding a decode range and an AXI4-Lite slave — not touching the core.

Interconnect

The axi4 fabric is an in-house AXI4-Lite implementation. Every peripheral presents the same slave interface, which is what makes the peripheral set uniform enough to verify one block at a time and then integrate with confidence. The AXI4 bus additionally has a UVM environment built around it — the only part of the design verified at that level of formality, because it is the part every other block depends on.

Module map

Every module in the design, and what it does:

SiliCore SoC modules and their responsibilities
ModuleDescription
cv32e40p 32-bit RISC-V in-order core (OpenHW Group)
axi4 AXI4-Lite interconnect fabric
address_decoder Memory-mapped peripheral address decoding
gpio General-purpose I/O peripheral
timer Hardware timer peripheral
uart (UART0 / UART1) General-purpose UART, plus a second UART dedicated to AI accelerator data streaming
i2c I2C master / controller
qspi / flash Quad-SPI controller and flash memory interface
imem / dmem / bootrom Instruction memory, data memory, and boot ROM
jtag Debug access
ivc Interrupt vector controller — vectoring and dispatch for the peripheral interrupts
hdmi HDMI framebuffer output (Gowin and open-source IP variants)
ai_accelerator / ai_accelerator2 Custom neural-network inference accelerator — serial-adder core, weight/bias memories, DSP-based variant, UART-fed loader

The AI accelerator

The accelerator is the clearest example of in-house work in the design. It has its own register interface, a UART-based weight-loading path, dedicated weight and bias ROMs, and two independent design iterations: a serial-adder-based core and a later DSP-unit-based core. The second exists because we measured the first and went back to improve it — it is an internally-driven optimisation effort, not a purchased IP block.

Read more about the accelerator →

Software

Firmware is bare-metal C and RISC-V assembly built against a custom HAL and linker scripts. Each peripheral has a matching driver (gpio.c, i2c.c, qspi.c, timer.c, uart.c, ai.c, ivc.c, oled.c) and a HAL header. A dedicated boot loader brings the core up before jumping to firmware.

Instruction-level correctness is checked independently against the spike RISC-V ISA simulator, using a checker script plus directed and basic ISA test suites — so the core's behaviour in our SoC is compared against a reference implementation, not just against our own expectations.

Implementation

The same RTL is carried through three implementation targets: an Altera Agilex 3 FPGA via Quartus Prime, a Gowin FPGA, and an open-source RTL-to-GDS flow via OpenLane, with completed hardened runs for both a standalone MCU and the full top-level SoC. Simulation runs in Vivado xsim.

Read more about the FPGA and ASIC flows →