> ## Documentation Index
> Fetch the complete documentation index at: https://dragonwingdocs-staging.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Flashing and validation

\#10. 3. 刷写与验证

本阶段介绍如何将 MCU 固件刷写到 Shikra CS2390 上，并验证模型推理已成功执行。

***

## 10.3.1 刷写固件

构建会在 `build/ms/bin/shikra.lpaicp.test` 下生成两个固件二进制文件。

```bash theme={null}
adb shell mount -o rw,remount /usr
adb push lpaicp.mbn /usr/lib/firmware/qcom/shikra/
adb push lpaicp_dtb.mbn /usr/lib/firmware/qcom/shikra/
adb shell sync
adb shell reboot
```

***

## 10.3.2. 验证 LPAI 子系统正在运行

设备启动后，检查 `remoteproc` 驱动是否已成功加载 LPAI 固件：

```bash theme={null}
cat /sys/class/remoteproc/remoteproc*/{state,name}
```

**预期输出：**

`state` 文件的内容应为 `running`，`name` 文件应标识 LPAI 内核（例如 `soccp`）。任何其他状态（例如 `offline`、`crashed`）都表示加载失败，请查看 `dmesg | grep -i remoteproc` 了解详细信息。

***

## 10.3.3. T32 分析

Shikra LPAICP MCU **没有 UART**，所有控制台输出都写入 RAM 控制台（`ram_console_buf`）。调试和结果读取通过 RISC-V JTAG 上的 **Lauterbach T32** 完成。TFLM 推理框架会将所有结果写入 `volatile` 全局变量，T32 可随时读取这些变量，无需设置断点。

### 10.3.3.1. 获取 RAM 控制台地址

`ram_console_buf` 地址在每次构建后都会变化。请从 ELF 文件中获取：

```bash theme={null}
nm LPAICP_IMG_shikra.lpaicp.test.elf | grep ram_console_buf
```

示例输出：

```
b45a7a8e B ram_console_buf
```

该地址为 `0xb45a7a8e`。请使用**您自己**构建中的地址更新 CMM 脚本。

### 10.3.3.2. T32 CMM 脚本

更新 `&elf` 路径和 `ram_console_buf` 地址以匹配您的构建，然后在 T32 中运行该脚本。

```
; ============================================================================
; read_tflite_infer.cmm
; Read TFLite-Micro model latency results from a device that is ALREADY booted
; and running the inference loop.
;
; tflite_infer_run() runs a batch of timed invokes, publishes fresh
; min/avg/max into the g_tfli_* globals, bumps g_tfli_batches LAST (the
; commit marker), sleeps, and repeats. Halting at any time and reading the
; globals yields the latest coherent batch -- as long as g_tfli_batches >= 1
; and g_tfli_status == 0.
;
; How to run (after the target is connected):
;   DO <path-to-this-script>/read_tflite_infer.cmm
; ============================================================================
local &elf
local &ram_console_addr

; ==== EDIT THESE TWO VALUES for your build ====================================
; Full path to the ELF as this Windows machine sees it.
&elf = "C:\path\to\LPAICP_IMG_shikra.lpaicp.test.elf"
; Get this address from: nm <elf> | grep ram_console_buf
&ram_console_addr = 0xb45a7a8e
; ==============================================================================

; ---- load symbols only (do NOT touch target code / do NOT reset) ------------
Data.LOAD.Elf "&elf" /NOCODE

; ---- halt so memory reads are coherent (does not restart anything) ----------
IF STATE.RUN()
(
    PRINT "read_tflite_infer: halting hart to read memory..."
    Break
    WAIT !STATE.RUN()
)

; ---- display the latest published batch -------------------------------------
; status / arena / iters / commit-count
Var.View g_tfli_status g_tfli_batches g_tfli_arena_used g_tfli_arena_size g_tfli_iters

; latency per inference in microseconds (QTMR 19.2 MHz) -- the headline number
Var.View g_tfli_min_us g_tfli_avg_us g_tfli_max_us

; latency in raw QTMR ticks + RISC-V CPU cycles (mcycle) + effective core MHz
Var.View g_tfli_min_ticks g_tfli_avg_ticks g_tfli_max_ticks
Var.View g_tfli_min_mcyc g_tfli_avg_mcyc g_tfli_max_mcyc g_tfli_cpu_mhz

; output tensor of the last invoke : raw int8 and argmax (predicted class)
Var.View g_tfli_out g_tfli_argmax
; Note: Var.View %Float g_tfli_out_f requires an HLL Debugging license.
; If not available, skip it -- all latency results are accessible without HLL.

PRINT "Dumping ram_console_buf..."
Data.dump EZAXI:&ram_console_addr /DIALOG /NoHex /STRING

PRINT "PASS = g_tfli_status==0, g_tfli_batches>=1, non-zero g_tfli_avg_us"
PRINT "Latency = g_tfli_avg_us microseconds/inference"
enddo
```

***

## 10.3.4. T32 全局变量参考

| 变量 | 类型 | 含义 |
| - | - | - |
| `g_tfli_status` | `int32_t` | `0` = 成功；`-1` = 从未运行；`-3` = k\_malloc 失败；`-4` = AllocateTensors 失败；`-5` = 输入大小不匹配；`-6` = 预热 Invoke 失败；`-7` = 计时 Invoke 失败 |
| `g_tfli_batches` | `uint32_t` | 批次计数器，在每次发布后**最后**递增；当其值 `>= 1` 时读取可获得一致的结果 |
| `g_tfli_arena_used` | `uint32_t` | tensor arena 已使用的字节数，用于调优 `TFLI_ARENA_KB`（请参阅[第 2 阶段，第 2.5.2.1 节](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb)） |
| `g_tfli_arena_size` | `uint32_t` | 配置的 arena 大小，单位为字节（`TFLI_ARENA_KB * 1024`） |
| `g_tfli_iters` | `uint32_t` | 每个批次中计时的 Invoke() 调用次数（`TFLI_ITERS`） |
| `g_tfli_min_us` | `uint32_t` | 最小推理延迟，单位为微秒 |
| `g_tfli_avg_us` | `uint32_t` | **平均推理延迟，单位为微秒，即核心指标** |
| `g_tfli_max_us` | `uint32_t` | 最大推理延迟，单位为微秒 |
| `g_tfli_min_ticks` | `uint32_t` | 以 QTMR tick 计的最小延迟（19.2 MHz 时钟） |
| `g_tfli_avg_ticks` | `uint32_t` | 以 QTMR tick 计的平均延迟 |
| `g_tfli_max_ticks` | `uint32_t` | 以 QTMR tick 计的最大延迟 |
| `g_tfli_min_mcyc` | `uint64_t` | 以 RISC-V `mcycle` CPU 周期计的最小延迟 |
| `g_tfli_avg_mcyc` | `uint64_t` | 以 RISC-V `mcycle` CPU 周期计的平均延迟 |
| `g_tfli_max_mcyc` | `uint64_t` | 以 RISC-V `mcycle` CPU 周期计的最大延迟 |
| `g_tfli_cpu_mhz` | `uint32_t` | 有效 CPU 频率，单位为 MHz（avg\_mcyc / avg\_us） |
| `g_tfli_out[2]` | `int8_t[2]` | 原始输出张量（int8），2 分类模型对应 2 个值 |
| `g_tfli_out_f[2]` | `float[2]` | 反量化后的输出张量（float），在 T32 中查看需要 HLL 许可证 |
| `g_tfli_argmax` | `int32_t` | 预测类别（输出张量的 argmax） |

***

## 10.3.5. 解读结果

### 10.3.5.1. 有效性检查

在报告结果之前，务必验证：

```
g_tfli_status  == 0          (success)
g_tfli_batches >= 1          (at least one complete batch published)
```

如果 `g_tfli_status` 不为零，请参阅[第 2 阶段，第 2.5.2.1 节](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb)中的错误表以确定根本原因。

### 10.3.5.2. 延迟

核心指标是 `g_tfli_avg_us`（`TFLI_ITERS` 次迭代的平均推理延迟，单位为微秒）。

交叉验证：`g_tfli_avg_ticks / 19.2 MHz = g_tfli_avg_us`（在舍入误差范围内）。

`g_tfli_cpu_mhz` 给出根据 RISC-V `mcycle` 和 QTMR 时间戳推算出的有效 CPU 频率，可用于确认 CPU 是否以预期的时钟频率运行。

### 10.3.5.3. Arena 大小设置

如果 `g_tfli_arena_used` 接近 `g_tfli_arena_size`，请增大 `inc/tflite_infer.h` 中的 `TFLI_ARENA_KB`。有关确定最小安全值的步骤，请参阅 **[第 2 阶段，第 2.5.2.1 节](2_tflm_runtime_integration.mdx#2521-tuning-tfli_arena_kb)**。

### 10.3.5.4. RAM 控制台输出

RAM 控制台会捕获启动和推理期间写入的 `printk()` 输出。成功运行后，您应看到类似以下内容的行：

```
Reached main. Configuration: shikra_lpaicp_test
TFLI: arena_used=XXXX/4096 B, iters/batch=100, out_classes=2
TFLI[1]: lat avg=XXX us (min=XXX max=XXX) | cpu avg=XXXXXX cyc @~XXX MHz | argmax=0
```

`TFLI[N]` 行是每个批次的延迟摘要。`argmax` 是嵌入输入样本的预测类别。
