10.3.1 Flash the Firmware
The build produces two firmware binaries underbuild/ms/bin/shikra.lpaicp.test.
10.3.2. Validate the LPAI Subsystem Is Running
Once the device has booted, check that theremoteproc driver has loaded the LPAI firmware successfully:
state file should read running and the name file should identify the LPAI core (e.g. soccp). Any other state (e.g. offline, crashed) indicates a load failure — check dmesg | grep -i remoteproc for details.
10.3.3. T32 Analysis
The Shikra LPAICP MCU has no UART — all console output goes to a RAM console (ram_console_buf). Debug and result readout is done through Lauterbach T32 over RISC-V JTAG. The TFLM inference harness writes all results to volatile global variables which T32 can read at any time — no need to set breakpoints.
10.3.3.1. Get the RAM console address
Theram_console_buf address changes with every build. Get it from the ELF:
0xb45a7a8e. Update the CMM script with the address from your build.
10.3.3.2. T32 CMM script
Update the&elf path and ram_console_buf address to match your build, then run the script in T32.
10.3.4. T32 Global Variables Reference
10.3.5. Interpreting Results
10.3.5.1. Validity Check
Before reporting results, always verify:g_tfli_status is non-zero, see the error table in Phase 2, section 2.5.2.1 for root cause.
10.3.5.2. Latency
The headline number isg_tfli_avg_us (average inference latency in microseconds over TFLI_ITERS iterations).
Cross-check: g_tfli_avg_ticks / 19.2 MHz = g_tfli_avg_us (within rounding).
g_tfli_cpu_mhz gives the effective CPU frequency derived from RISC-V mcycle and the QTMR timestamp — useful for confirming the CPU is running at the expected clock rate.
10.3.5.3. Arena Sizing
Ifg_tfli_arena_used is close to g_tfli_arena_size, increase TFLI_ARENA_KB in inc/tflite_infer.h. For the procedure to find the minimum safe value, see Phase 2, section 2.5.2.1.
10.3.5.4. RAM Console Output
The RAM console capturesprintk() output written during boot and inference. After a successful run, you should see lines like:
TFLI[N] line is the per-batch latency summary. argmax is the predicted class for the embedded input sample.
