> ## Documentation Index
> Fetch the complete documentation index at: https://dragonwingdocs-staging.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Flashing and validation

\#10. 3. Flashing and Validation

This phase documents how to flash the MCU firmware onto the Shikra CS2390 and verify that model inference executed successfully.

***

## 10.3.1 Flash the Firmware

The build produces two firmware binaries under `build/ms/bin/shikra.lpaicp.test`.

```bash theme={null}
adb shell mount -o rw,remount /usr
adb push lpaicp.mbn /usr/lib/firmware/qcom/shikra/
adb push lpaicp_dtb.mbn /usr/lib/firmware/qcom/shikra/
adb shell sync
adb shell reboot
```

***

## 10.3.2. Validate the LPAI Subsystem Is Running

Once the device has booted, check that the `remoteproc` driver has loaded the LPAI firmware successfully:

```bash theme={null}
cat /sys/class/remoteproc/remoteproc*/{state,name}
```

**Expected output:**

The `state` file should read `running` and the `name` file should identify the LPAI core (e.g. `soccp`). Any other state (e.g. `offline`, `crashed`) indicates a load failure — check `dmesg | grep -i remoteproc` for details.

***

## 10.3.3. T32 Analysis

The Shikra LPAICP MCU has **no UART** — all console output goes to a RAM console (`ram_console_buf`). Debug and result readout is done through **Lauterbach T32** over RISC-V JTAG. The TFLM inference harness writes all results to `volatile` global variables which T32 can read at any time — no need to set breakpoints.

### 10.3.3.1. Get the RAM console address

The `ram_console_buf` address changes with every build. Get it from the ELF:

```bash theme={null}
nm LPAICP_IMG_shikra.lpaicp.test.elf | grep ram_console_buf
```

Example output:

```
b45a7a8e B ram_console_buf
```

The address is `0xb45a7a8e`. Update the CMM script with the address from **your** build.

### 10.3.3.2. T32 CMM script

Update the `&elf` path and `ram_console_buf` address to match your build, then run the script in T32.

```
; ============================================================================
; read_tflite_infer.cmm
; Read TFLite-Micro model latency results from a device that is ALREADY booted
; and running the inference loop.
;
; tflite_infer_run() runs a batch of timed invokes, publishes fresh
; min/avg/max into the g_tfli_* globals, bumps g_tfli_batches LAST (the
; commit marker), sleeps, and repeats. Halting at any time and reading the
; globals yields the latest coherent batch -- as long as g_tfli_batches >= 1
; and g_tfli_status == 0.
;
; How to run (after the target is connected):
;   DO <path-to-this-script>/read_tflite_infer.cmm
; ============================================================================
local &elf
local &ram_console_addr

; ==== EDIT THESE TWO VALUES for your build ====================================
; Full path to the ELF as this Windows machine sees it.
&elf = "C:\path\to\LPAICP_IMG_shikra.lpaicp.test.elf"
; Get this address from: nm <elf> | grep ram_console_buf
&ram_console_addr = 0xb45a7a8e
; ==============================================================================

; ---- load symbols only (do NOT touch target code / do NOT reset) ------------
Data.LOAD.Elf "&elf" /NOCODE

; ---- halt so memory reads are coherent (does not restart anything) ----------
IF STATE.RUN()
(
    PRINT "read_tflite_infer: halting hart to read memory..."
    Break
    WAIT !STATE.RUN()
)

; ---- display the latest published batch -------------------------------------
; status / arena / iters / commit-count
Var.View g_tfli_status g_tfli_batches g_tfli_arena_used g_tfli_arena_size g_tfli_iters

; latency per inference in microseconds (QTMR 19.2 MHz) -- the headline number
Var.View g_tfli_min_us g_tfli_avg_us g_tfli_max_us

; latency in raw QTMR ticks + RISC-V CPU cycles (mcycle) + effective core MHz
Var.View g_tfli_min_ticks g_tfli_avg_ticks g_tfli_max_ticks
Var.View g_tfli_min_mcyc g_tfli_avg_mcyc g_tfli_max_mcyc g_tfli_cpu_mhz

; output tensor of the last invoke : raw int8 and argmax (predicted class)
Var.View g_tfli_out g_tfli_argmax
; Note: Var.View %Float g_tfli_out_f requires an HLL Debugging license.
; If not available, skip it -- all latency results are accessible without HLL.

PRINT "Dumping ram_console_buf..."
Data.dump EZAXI:&ram_console_addr /DIALOG /NoHex /STRING

PRINT "PASS = g_tfli_status==0, g_tfli_batches>=1, non-zero g_tfli_avg_us"
PRINT "Latency = g_tfli_avg_us microseconds/inference"
enddo
```

***

## 10.3.4. T32 Global Variables Reference

| Variable | Type | Meaning |
| - | - | - |
| `g_tfli_status` | `int32_t` | `0` = success; `-1` = never ran; `-3` = k\_malloc failed; `-4` = AllocateTensors failed; `-5` = input size mismatch; `-6` = warm-up Invoke failed; `-7` = timed Invoke failed |
| `g_tfli_batches` | `uint32_t` | Batch counter — bumped **last** after each publish; read when `>= 1` for coherent results |
| `g_tfli_arena_used` | `uint32_t` | Bytes used by tensor arena — use to tune `TFLI_ARENA_KB` (see [Phase 2, section 2.5.2.1](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb)) |
| `g_tfli_arena_size` | `uint32_t` | Configured arena size in bytes (`TFLI_ARENA_KB * 1024`) |
| `g_tfli_iters` | `uint32_t` | Number of timed Invoke() calls per batch (`TFLI_ITERS`) |
| `g_tfli_min_us` | `uint32_t` | Minimum inference latency in microseconds |
| `g_tfli_avg_us` | `uint32_t` | **Average inference latency in microseconds — the headline number** |
| `g_tfli_max_us` | `uint32_t` | Maximum inference latency in microseconds |
| `g_tfli_min_ticks` | `uint32_t` | Minimum latency in QTMR ticks (19.2 MHz clock) |
| `g_tfli_avg_ticks` | `uint32_t` | Average latency in QTMR ticks |
| `g_tfli_max_ticks` | `uint32_t` | Maximum latency in QTMR ticks |
| `g_tfli_min_mcyc` | `uint64_t` | Minimum latency in RISC-V `mcycle` CPU cycles |
| `g_tfli_avg_mcyc` | `uint64_t` | Average latency in RISC-V `mcycle` CPU cycles |
| `g_tfli_max_mcyc` | `uint64_t` | Maximum latency in RISC-V `mcycle` CPU cycles |
| `g_tfli_cpu_mhz` | `uint32_t` | Effective CPU MHz (avg\_mcyc / avg\_us) |
| `g_tfli_out[2]` | `int8_t[2]` | Raw output tensor (int8) — 2 values for 2-class model |
| `g_tfli_out_f[2]` | `float[2]` | Dequantized output tensor (float) — requires HLL license to view in T32 |
| `g_tfli_argmax` | `int32_t` | Predicted class (argmax of output tensor) |

***

## 10.3.5. Interpreting Results

### 10.3.5.1. Validity Check

Before reporting results, always verify:

```
g_tfli_status  == 0          (success)
g_tfli_batches >= 1          (at least one complete batch published)
```

If `g_tfli_status` is non-zero, see the error table in [Phase 2, section 2.5.2.1](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb) for root cause.

### 10.3.5.2. Latency

The headline number is `g_tfli_avg_us` (average inference latency in microseconds over `TFLI_ITERS` iterations).

Cross-check: `g_tfli_avg_ticks / 19.2 MHz = g_tfli_avg_us` (within rounding).

`g_tfli_cpu_mhz` gives the effective CPU frequency derived from RISC-V `mcycle` and the QTMR timestamp — useful for confirming the CPU is running at the expected clock rate.

### 10.3.5.3. Arena Sizing

If `g_tfli_arena_used` is close to `g_tfli_arena_size`, increase `TFLI_ARENA_KB` in `inc/tflite_infer.h`. For the procedure to find the minimum safe value, see **[Phase 2, section 2.5.2.1](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb)**.

### 10.3.5.4. RAM Console Output

The RAM console captures `printk()` output written during boot and inference. After a successful run, you should see lines like:

```
Reached main. Configuration: shikra_lpaicp_test
TFLI: arena_used=XXXX/4096 B, iters/batch=100, out_classes=2
TFLI[1]: lat avg=XXX us (min=XXX max=XXX) | cpu avg=XXXXXX cyc @~XXX MHz | argmax=0
```

The `TFLI[N]` line is the per-batch latency summary. `argmax` is the predicted class for the embedded input sample.
