> ## Documentation Index
> Fetch the complete documentation index at: https://dragonwingdocs-staging.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Tflm runtime integration

# 10.2 TFLM 运行时与模型嵌入

本阶段涵盖两项内容：将 TensorFlow Lite Micro (TFLM) 运行时集成到 Q2390 / IQ2390 MCU 固件中，以及将预先构建的 int8 模型（在第 1 阶段生成）直接嵌入固件镜像。其中包括在 SDLLVM/RISC-V 上所需的、不易察觉的 C++ 运行时设置。

***

\##10.2.1	概述

集成包含五个不同部分：

1. **获取 TFLM 源码** — TFLM 源码未随仓库提供（not vendored），必须手动获取
2. **检查模型并验证算子支持** — 从模型中提取算子类型，并确认 TFLM 支持这些算子
3. **配置 LLVM libc++** — SDLLVM 默认不配置 libc++；需要进行 5 项特定修复
4. **创建推理模块** — 构建包含嵌入式模型 C 数组、推理框架（`tflite_infer.cc`）、Kconfig 头文件和 CMakeLists.txt 的模块
5. **集成到固件构建中** — 在 `main.c` 中调用 `tflite_infer_run()`，并在 CMake 中注册该模块

***

\##10.2.2. 第 1 部分 — 获取 TFLite Micro 源码

> **路径约定：** 在本阶段中，`<wasp_proc>` 指 MCU 固件工作区的根目录，即包含 `zephyr/`、`modules/`、`config/` 等的目录。凡出现此处，请替换为您的实际路径（例如 `/path/to/your_workspace/wasp_proc`）。

TFLM 模块桩文件（`CMakeLists.txt`、`Kconfig`）位于 `wasp_proc/zephyr/zephyr/modules/tflite-micro/`，但实际的 C++ 源码树并不存在，必须手动克隆。

`tflite-micro` 项目在 west 清单（`zephyr/zephyr/submanifests/optional.yaml`）中定义于 `optional` 组，而 Zephyr 自身的 `west.yml` 通过 `group-filter: [-optional]` 显式禁用了该组。这意味着 `west update tflite-micro` 将**无法**工作，请改用 `git clone`。

### 10.2.2.1. 克隆源码

在 Docker 构建容器内：

```bash theme={null}
# Navigate to the wasp_proc root first
cd <wasp_proc>

# Read the pinned revision from the manifest — always in sync with the Zephyr version
TFLM_REV=$(grep -A6 "name: tflite-micro" \
    zephyr/zephyr/submanifests/optional.yaml | \
    grep "revision:" | awk '{print $2}')
echo "Fetching tflite-micro at revision: $TFLM_REV"

git clone https://github.com/zephyrproject-rtos/tflite-micro \
    modules/tflite-micro
cd modules/tflite-micro && git checkout $TFLM_REV && cd ../..
```

修订版本直接从清单中读取，因此无需修改即可适用于任何 LPAICP 版本。由此产生的分离 HEAD（detached HEAD）状态是预期且正确的。

### 10.2.2.2. 验证

```bash theme={null}
ls modules/tflite-micro/tensorflow/lite/micro/kernels/*.cc | wc -l
# Should print a non-zero number (e.g. 189)
```

***

## 10.2.3. 第 2 部分 — 检查模型并验证算子支持

在编写任何固件代码之前，我们需要从模型中获取两项信息：

1. **模型使用了哪些算子** — TFLM 运行时要求在运行推理之前显式注册模型使用的每个算子。如果缺少任何算子，固件将在启动时因无法解析算子而报错。此步骤可确定模型所需的确切算子集合，以便您在 [第 5 部分](#26-part-5--integrate-into-firmware-build) 中集成推理框架时知道需要注册哪些算子。
2. **输入/输出量化参数**（scale + zero\_point）— 如果模型采用 int8 量化，则浮点输入在传递给推理框架之前必须预先量化为 int8。

### 10.2.3.1. 选项 A — `flatc`（完整拓扑解码 — 推荐）

使用 TFLM schema 将模型二进制解码为人类可读的 JSON。可获得：所有张量的算子类型、张量名称、形状和量化参数。

**步骤 1 — 运行 flatc 解码模型：**

```bash theme={null}
sed -i 's/ (deprecated)//g' \
    zephyr/zephyr/modules/tflite-micro/tensorflow/lite/schema/schema.fbs
flatc --json --strict-json --raw-binary \
    zephyr/zephyr/modules/tflite-micro/tensorflow/lite/schema/schema.fbs \
    -- cnn1d_minimal_int8.tflite
```

这会在当前目录中生成 `cnn1d_minimal_int8.json`。

**步骤 2 — 从 JSON 中提取算子列表：**

```python theme={null}
import json
with open("cnn1d_minimal_int8.json") as f:
    model = json.load(f)
opcodes   = model["operator_codes"]
operators = model["subgraphs"][0]["operators"]
ops_used  = set()
for op in operators:
    idx   = op["opcode_index"]
    entry = opcodes[idx]
    name  = entry.get("builtin_code", entry.get("custom_code", "UNKNOWN"))
    ops_used.add(name)
print("Ops used by model:")
for op in sorted(ops_used):
    print(f"  {op}")
```

**步骤 3 — 将每个算子与 1.3.3 中的支持算子列表进行交叉核对。**

### 10.2.3.2. 选项 B — `tf.lite.Interpreter`（仅量化参数，无需 `flatc`）

在 `flatc` 不可用时使用。可获得输入/输出量化参数（scale、zero\_point）。**不会暴露算子类型** — 请使用选项 A 获取算子列表。

```python theme={null}
import tensorflow as tf
interp = tf.lite.Interpreter("cnn1d_minimal_int8.tflite")
interp.allocate_tensors()
inp = interp.get_input_details()[0]
out = interp.get_output_details()[0]
print(f"Input: shape={inp['shape']}, dtype={inp['dtype']}, "
      f"scale={inp['quantization'][0]:.6f}, zp={inp['quantization'][1]}")
print(f"Output: shape={out['shape']}, dtype={out['dtype']}, "
      f"scale={out['quantization'][0]:.6f}, zp={out['quantization'][1]}")
```

> `tf.lite.Interpreter` 不会暴露算子类型 — 请使用选项 A（`flatc`）获取算子列表。

### 10.2.3.3. 此 TFLM 构建支持的所有算子

TFLM 所支持算子的权威来源是 `micro_mutable_op_resolver.h` — 该文件中声明的每个 `Add*()` 方法都对应所下载的确切修订版本所支持的一个算子。

```bash theme={null}
grep "TfLiteStatus Add" \
    modules/tflite-micro/tensorflow/lite/micro/micro_mutable_op_resolver.h \
    | grep -v "AddCustom\|AddBuiltin" \
    | sed 's/.*TfLiteStatus Add\([A-Za-z0-9]*\).*/\1/' \
    | sort | uniq
```

**发布标签 `zephyr_20240627` 的完整支持算子列表 — 共 114 个算子：**

| 算子 | 算子 | 算子 | 算子 |
| - | - | - | - |
| Abs | Add | AddN | ArgMax |
| ArgMin | AssignVariable | AveragePool2D | BatchMatMul |
| BatchToSpaceNd | BroadcastArgs | BroadcastTo | CallOnce |
| Cast | Ceil | CircularBuffer | Concatenation |
| Conv2D | Cos | CumSum | Delay |
| DepthToSpace | DepthwiseConv2D | Dequantize | DetectionPostprocess |
| Div | Elu | EmbeddingLookup | Energy |
| Equal | EthosU | Exp | ExpandDims |
| FftAutoScale | Fill | FilterBank | FilterBankLog |
| FilterBankSpectralSubtraction | FilterBankSquareRoot | Floor | FloorDiv |
| FloorMod | Framer | FullyConnected | Gather |
| GatherNd | Greater | GreaterEqual | HardSwish |
| If | Irfft | L2Normalization | L2Pool2D |
| LeakyRelu | Less | LessEqual | Log |
| LogSoftmax | LogicalAnd | LogicalNot | LogicalOr |
| Logistic | Maximum | MaxPool2D | Mean |
| Minimum | MirrorPad | Mul | Neg |
| NotEqual | OverlapAdd | Pack | Pad |
| PadV2 | PCAN | Prelu | Quantize |
| ReadVariable | ReduceMax | Relu | Relu6 |
| Reshape | ResizeBilinear | ResizeNearestNeighbor | Rfft |
| Round | Rsqrt | SelectV2 | Shape |
| Sin | Slice | Softmax | SpaceToBatchNd |
| SpaceToDepth | Split | SplitV | Sqrt |
| Square | SquaredDifference | Squeeze | Stacker |
| StridedSlice | Sub | Sum | Svdf |
| Tanh | Transpose | TransposeConv | UnidirectionalSequenceLSTM |
| Unpack | VarHandle | While | Window |
| ZerosLike | | | |

> 每个算子对应的 resolver 方法为 `Add` + 算子名称 — 例如 `Conv2D` → `AddConv2D()`，`FullyConnected` → `AddFullyConnected()`。

### 10.2.3.4. 参考 — 示例模型已确认的算子

`cnn1d_minimal_int8.tflite` 中的所有算子、对应的 TFLM 内核文件以及 resolver 方法：

| 算子 | TFLM 内核文件 | `MicroMutableOpResolver` 方法 |
| - | - | - |
| `CONV_2D` | `conv.cc` | `AddConv2D()` |
| `MEAN` | `reduce.cc` | `AddMean()` |
| `FULLY_CONNECTED` | `fully_connected.cc` | `AddFullyConnected()` |

**如果某个算子未注册，您将看到以下错误之一：**

编译时（方法名称错误或缺失）：

```
error: 'class tflite::MicroMutableOpResolver<N>' has no member named 'AddMyOp'
```

运行时在 `AllocateTensors()` 阶段（算子未注册）：

```
Didn't find op for builtin opcode 'CONV_2D' version '1'
```

***

## 10.2.4. 第 3 部分 — 配置 LLVM libc++（5 步修复）

### 10.2.4.1. 默认方法为何失败

要使用 TFLM（以 C++ 编写），Zephyr 需要 C++ 标准库。自然的起点是启用 `CONFIG_LIBCXX_LIBCPP=y`。然而在 SDLLVM 上，这会触发一条破坏构建的 Kconfig 依赖链。该故障会表现为一个完全不相关的内核文件中的错误：

```
fatal error: 'string.h' file not found in offsets.c
```

### 10.2.4.2. 4 步修复

#### 10.2.4.2.1. 步骤 1 — 从 TFLM Kconfig 中移除 `select REQUIRES_FULL_LIBCPP`

文件：`zephyr/zephyr/modules/tflite-micro/Kconfig`

从 `config TENSORFLOW_LITE_MICRO` 内部移除 `select REQUIRES_FULL_LIBCPP` 这一行。这是依赖级联的根源 — 移除它可防止 Zephyr 在 SDLLVM 上触发已损坏的 `REQUIRES_FULL_LIBC` 路径。

```
# Remove this line from config TENSORFLOW_LITE_MICRO:
select REQUIRES_FULL_LIBCPP
```

#### 10.2.4.2.2. 步骤 2 — 使用 `EXTERNAL_MODULE_LIBCPP`，而非 `EXTERNAL_LIBCPP`

*此处无需操作 — 该设置包含在第 2.4.4 节创建的新文件中。*

| 符号 | 是否避开已损坏的依赖链？ | 是否为 `qcld` 抑制 `-lc++` namespec？ |
| - | - | - |
| `CONFIG_LIBCXX_LIBCPP=y` | 否 — 会触发已损坏的依赖链 | 否 |
| `CONFIG_EXTERNAL_LIBCPP=y` | 是 | 否 — 仍会输出 `qcld` 无法解析的 `-lc++ -lc++abi` |
| `CONFIG_EXTERNAL_MODULE_LIBCPP=y` | 是 | 是 — 正确选择 |

#### 10.2.4.2.3. 步骤 3 — 通过 `-I` 而非 `-isystem` 注入 libc++ 包含路径

*此处无需操作 — 该调用是第 2.4.3 节中添加的完整代码块的一部分。*

`-I` 目录会先于任何 `-isystem` 目录被搜索。通过以 `-I BEFORE` 方式注入 libc++ 头文件，clang 会首先找到 libc++ 自身的 `<stdio.h>`/`<string.h>` 包装头文件，它们再通过 `#include_next` 链接到 musl — 这正是 C++ 翻译单元的正确解析顺序。

#### 10.2.4.2.4. 步骤 4 — 使用 `-DNDEBUG` 编译 C++ 翻译单元

*此处无需操作 — 该标志是第 2.4.3 节中添加的完整代码块的一部分。*

musl 的 `assert()` 宏会展开为 `__assert_fail()`，而 picolibc 并不提供该函数。`-DNDEBUG` 会在 C++ 翻译单元中编译掉所有 assert 调用。其作用范围限定为 `$<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG>`，因此 C 代码不受影响。

### 10.2.4.3. 完整的 CMake 配置代码块

文件：`modules/hal/qcom/core/config/CMakeLists.txt`（已有文件 — 在末尾添加此完整代码块）。该代码块以 `CONFIG_EXTERNAL_MODULE_LIBCPP` 为条件，因此仅在合并 TFLM Kconfig 片段时才会生效。

```cmake theme={null}
if(CONFIG_EXTERNAL_MODULE_LIBCPP)
    get_filename_component(_sdllvm_bin ${CMAKE_C_COMPILER} DIRECTORY)
    get_filename_component(_sdllvm_root ${_sdllvm_bin} DIRECTORY)
    set(_sdllvm_sysroot ${_sdllvm_root}/riscv32-unknown-elf)
    file(GLOB _sdllvm_builtin_inc ${_sdllvm_root}/lib/clang/*/include)
    file(GLOB _sdllvm_builtin_lib
        ${_sdllvm_root}/lib/clang/*/lib/baremetal/libclang_rt.builtins-riscv32.a)
    if(NOT EXISTS ${_sdllvm_sysroot}/include/c++/v1/vector)
        message(FATAL_ERROR
            "CONFIG_EXTERNAL_MODULE_LIBCPP set but libc++ headers not found at "
            "${_sdllvm_sysroot}/include/c++/v1 (compiler=${CMAKE_C_COMPILER}). "
            "The contained libc++ configuration in core/config/CMakeLists.txt assumes "
            "the Snapdragon LLVM riscv32-unknown-elf layout.")
    endif()
    target_include_directories(zephyr_interface BEFORE INTERFACE
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_sysroot}/include/c++/v1>
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_sysroot}/libc/include>
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_builtin_inc}>
    )
    zephyr_compile_options($<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG>)
    zephyr_link_libraries(
        -Wl,--start-group
        ${_sdllvm_sysroot}/lib/libc++.a
        ${_sdllvm_sysroot}/lib/libc++abi.a
        ${_sdllvm_sysroot}/lib/libunwind.a
        ${_sdllvm_builtin_lib}
        -Wl,--end-group
    )
endif()
```

### 10.2.4.4. `shikra_lpaicp_tflite.conf`

#### 10.2.4.4.1. 步骤 1 — 创建 `shikra_lpaicp_tflite.conf`

在 `zephyr/kernel/config/shikra_lpaicp_tflite.conf` 创建新文件：

```
CONFIG_CPP=y
CONFIG_STD_CPP17=y
CONFIG_EXTERNAL_MODULE_LIBCPP=y
CONFIG_TENSORFLOW_LITE_MICRO=y
CONFIG_REQUIRES_FLOAT_PRINTF=y
CONFIG_MAIN_STACK_SIZE=8192
```

#### 10.2.4.4.2. 步骤 2 — 在 `config.yml` 中注册该文件

文件：`zephyr/kernel/config/config.yml`（已有文件 — 添加一个新条目）：

```yaml theme={null}
- regex: "shikra_lpaicp_.*"
  files:
    - shikra_lpaicp_tflite.conf
```

> **为什么使用单独的文件，而不是添加到 `shikra_lpaicp.conf`？** 现有条目使用正则表达式 `(^|shikra_)lpaicp_.*`，它同时匹配硬件目标（`SHIKRA_LPAICP_TEST`）和使用 Zephyr-SDK GCC 的 QEMU 仿真变体。将 TFLM 特定于 SDLLVM 的 libc++ 配置添加到 `shikra_lpaicp.conf` 会破坏基于 GCC 的 QEMU 构建。范围更窄的正则表达式 `shikra_lpaicp_.*` 仅匹配硬件目标。

***

## 10.2.5. 第 4 部分 — 创建推理模块

### 10.2.5.1. 步骤 1 — 生成 src/model\_data.cpp 和 src/input\_data.cpp

首先创建模块目录，然后将第 1 阶段生成的模型和输入文件转换为 C 字节数组。请在包含第 1 阶段源文件的目录中运行这些命令。

```bash theme={null}
mkdir -p modules/hal/qcom/core/tflite_infer/{inc,src}

# Model byte array  ->  g_cnn_model[], g_cnn_model_len
xxd -i model_files/cnn1d_minimal_int8.tflite \
  | sed -e 's/^unsigned char [a-zA-Z0-9_]*/alignas(8) extern const unsigned char g_cnn_model/' \
        -e 's/^unsigned int [a-zA-Z0-9_]*/extern const unsigned int g_cnn_model_len/' \
  > <wasp_proc>/modules/hal/qcom/core/tflite_infer/src/model_data.cpp

# Input byte array  ->  g_cnn_input[], g_cnn_input_len
xxd -i inputs/input_0.bin \
  | sed -e 's/^unsigned char [a-zA-Z0-9_]*/extern const unsigned char g_cnn_input/' \
        -e 's/^unsigned int [a-zA-Z0-9_]*/extern const unsigned int g_cnn_input_len/' \
  > <wasp_proc>/modules/hal/qcom/core/tflite_infer/src/input_data.cpp
```

将 `<wasp_proc>` 替换为您的 `wasp_proc` MCU 代码库路径。

生成的文件如下所示：

```c theme={null}
/* model_data.cpp */
alignas(8) extern const unsigned char g_cnn_model[] = {
  0x20, 0x00, 0x00, 0x00, 0x54, 0x46, 0x4c, 0x33, ...
};
extern const unsigned int g_cnn_model_len = 2096;

/* input_data.cpp */
extern const unsigned char g_cnn_input[] = {
  0x07, 0xff, 0x16, 0x06, 0xf3, 0x0e, 0x2a, 0x1f, ...
};
extern const unsigned int g_cnn_input_len = 200;
```

> **模型数组上的 `alignas(8)`：** TFLM 的 FlatBuffers 解析器要求模型字节数组至少按 4 字节对齐。无论链接器将该符号放置在何处，`alignas(8)` 都能保证满足这一要求。

### 10.2.5.2. 步骤 2 — 创建 `inc/tflite_infer.h`

文件：`modules/hal/qcom/core/tflite_infer/inc/tflite_infer.h`（新文件）。

```c theme={null}
#ifndef TFLITE_INFER_H_
#define TFLITE_INFER_H_
/* Tuneable inference parameters — edit here to adjust without recompiling other files */
#define TFLI_ARENA_KB         4    /* tensor arena in KB — measure via g_tfli_arena_used */
#define TFLI_ITERS            100  /* timed Invoke() calls per batch (after one warm-up) */
#define TFLI_BATCH_SLEEP_MS 1000   /* sleep between re-measure batches (ms) */
#ifdef __cplusplus
extern "C" {
#endif
/* TFLM inference harness. Loads the 8-bit quantized CNN model, feeds the
 * pre-quantized sample input, and times TFLI_ITERS invocations per batch.
 * Per-batch min/avg/max latency is published into the g_tfli_* volatile
 * globals for T32 inspection and also printed to the RAM console.
 * Call once from main() after boot initialisation. */
void tflite_infer_run(void);
#ifdef __cplusplus
}
#endif
#endif /* TFLITE_INFER_H_ */
```

#### 10.2.5.2.1. 调整 `TFLI_ARENA_KB`

`TFLI_ARENA_KB` 设置张量 arena 的大小 — 这是 TFLM 用于输入/输出张量、层间中间激活缓冲区及其内部分配器元数据的连续内存块。它在 `tflite_infer_run()` 开始时通过 `k_malloc` 从系统堆中分配。

**取值不当时会出现的问题：**

| RAM 控制台中的现象 | `g_tfli_status` | 根本原因 |
| - | - | - |
| `TFLI: k_malloc(N) failed` | `-3` | `TFLI_ARENA_KB * 1024 + 16` 超出了 `CONFIG_HEAP_MEM_POOL_SIZE`。请增大堆或减小 `TFLI_ARENA_KB`。 |
| `TFLI: AllocateTensors() failed (arena N B too small?)` | `-4` | arena 对于模型张量和 TFLM 元数据而言太小。请增大 `TFLI_ARENA_KB`。 |

这两种故障都会导致 `tflite_infer_run()` 提前返回 — `g_tfli_batches` 保持为 `0`，所有延迟全局变量保持其初始值。

**如何确定合适的值 — 使用 `g_tfli_arena_used` 进行测量：**

1. 将 `TFLI_ARENA_KB` 设置为一个充裕的初始值（例如 `32`）— 足够大以保证 `AllocateTensors()` 成功。
2. 构建、烧写、启动，并在 T32 中运行 `read_tflite_infer.cmm`。
3. 确认 `g_tfli_status == 0` 且 `g_tfli_batches >= 1`。
4. 从 T32 Var.View 窗口读取 `g_tfli_arena_used` — 这是 TFLM 所需的确切字节数。
5. 计算最小安全值并更新头文件：

```
TFLI_ARENA_KB = ceil(g_tfli_arena_used / 1024) + 2    <- +2 KB safety margin
```

6. 重新构建并重新烧写。验证 `g_tfli_arena_used` 仍小于 `g_tfli_arena_size`。

> **参考模型的配置值：** `TFLI_ARENA_KB = 4`（4,096 字节）。这是参考代码库中 `cnn1d_minimal_int8.tflite` 所使用的值。成功运行后，请在目标设备上测量 `g_tfli_arena_used`，以确认您的模型的使用情况。

**`TFLI_ARENA_KB` 对 RAM 的影响：**

arena 在运行时从系统堆中分配 — 它不会直接增加静态镜像大小。但是，堆的后备缓冲区（`kheap_buf__system_heap`）是一个**静态链接段**，其大小由 `CONFIG_HEAP_MEM_POOL_SIZE` 决定（在 `shikra_lpaicp.conf` 中设置为 `32768`）。仅减小 `TFLI_ARENA_KB` 并不会减少静态 RAM。若要回收静态 RAM，还需相应降低 `CONFIG_HEAP_MEM_POOL_SIZE` — 确保新的大小能够覆盖 `TFLI_ARENA_KB * 1024 + 16` 以及其他所有 `k_malloc` 调用方的需求。

### 10.2.5.3. 步骤 3 — 创建 `CMakeLists.txt`

文件：`modules/hal/qcom/core/tflite_infer/CMakeLists.txt`（新文件）。以 `CONFIG_TENSORFLOW_LITE_MICRO` 为条件。

```cmake theme={null}
if(CONFIG_TENSORFLOW_LITE_MICRO)
    target_include_directories(${ZEPHYR_APP_NAME} PRIVATE
        inc
    )
    target_sources(${ZEPHYR_APP_NAME} PRIVATE
        src/tflite_infer.cc
        src/model_data.cpp
        src/input_data.cpp
    )
endif()
```

### 10.2.5.4. 步骤 4 — 创建 `src/tflite_infer.cc`

文件：`modules/hal/qcom/core/tflite_infer/src/tflite_infer.cc`（新文件）。每个批次运行 `TFLI_ITERS` 次计时的 `Invoke()` 调用（在一次预热之后），然后将每批次以微秒、QTMR tick 和 CPU 周期表示的最小/平均/最大延迟发布到 `g_tfli_*` volatile 全局变量中。

```cpp theme={null}
#include <stdint.h>
#include <string.h>

#include <zephyr/kernel.h>
#include <zephyr/sys/printk.h>
#include <zephyr/logging/log.h>

LOG_MODULE_REGISTER(tflite_infer, LOG_LEVEL_DBG);

#include <tensorflow/lite/micro/micro_mutable_op_resolver.h>
#include <tensorflow/lite/micro/micro_interpreter.h>
#include <tensorflow/lite/micro/micro_log.h>
#include <tensorflow/lite/schema/schema_generated.h>

#include "tflite_infer.h"   /* TFLI_ARENA_KB, TFLI_ITERS, TFLI_BATCH_SLEEP_MS */

/* Model and input data — defined in model_data.cpp and input_data.cpp */
extern const unsigned char g_cnn_model[];
extern const unsigned char g_cnn_input[];
extern const unsigned int  g_cnn_input_len;

#define TFLI_ARENA_SIZE   (TFLI_ARENA_KB * 1024)
#define TFLI_NUM_CLASSES  2   /* matches Dense(2) output layer */

/* T32-readable results (volatile so they survive optimization) */
volatile int32_t  g_tfli_status     = -1;
volatile uint32_t g_tfli_arena_size __attribute__((retain)) = TFLI_ARENA_SIZE;
volatile uint32_t g_tfli_arena_used = 0;
volatile uint32_t g_tfli_iters      __attribute__((retain)) = TFLI_ITERS;
volatile uint32_t g_tfli_batches    = 0;

/* QTMR latency per inference (19.2 MHz ticks + microseconds) */
volatile uint32_t g_tfli_min_ticks = 0, g_tfli_avg_ticks = 0, g_tfli_max_ticks = 0;
volatile uint32_t g_tfli_min_us    = 0, g_tfli_avg_us    = 0, g_tfli_max_us    = 0;

/* CPU cycles per inference (RISC-V mcycle) */
volatile uint64_t g_tfli_min_mcyc = 0, g_tfli_avg_mcyc = 0, g_tfli_max_mcyc = 0;
volatile uint32_t g_tfli_cpu_mhz  = 0;

/* Raw + dequantized output tensor (last invoke of the most recent batch) */
volatile int8_t  g_tfli_out[TFLI_NUM_CLASSES]   = {0};
volatile float   g_tfli_out_f[TFLI_NUM_CLASSES] = {0};
volatile int32_t g_tfli_argmax = -1;

/* RISC-V 64-bit machine cycle counter (rv32: split across mcycle/mcycleh) */
static inline uint64_t read_mcycle64(void)
{
    uint32_t hi, lo, hi2;
    do {
        __asm__ volatile("csrr %0, mcycleh" : "=r"(hi));
        __asm__ volatile("csrr %0, mcycle"  : "=r"(lo));
        __asm__ volatile("csrr %0, mcycleh" : "=r"(hi2));
    } while (hi != hi2);
    return ((uint64_t)hi << 32) | lo;
}

extern "C" void tflite_infer_run(void)
{
    const tflite::Model *model = tflite::GetModel(g_cnn_model);
    if (model->version() != TFLITE_SCHEMA_VERSION) {
        printk("TFLI: model schema %u != supported %u\n",
               (unsigned)model->version(), (unsigned)TFLITE_SCHEMA_VERSION);
        g_tfli_status = -2;
        return;
    }

    /* 3 ops: Conv2D (fused BN+ReLU), Mean (GlobalAvgPool), FullyConnected */
    static tflite::MicroMutableOpResolver<3> resolver;
    resolver.AddConv2D();
    resolver.AddMean();
    resolver.AddFullyConnected();

    /* Arena from the system heap — over-allocate by 16 for manual alignment */
    uint8_t *arena_raw = (uint8_t *)k_malloc(TFLI_ARENA_SIZE + 16);
    if (arena_raw == NULL) {
        printk("TFLI: k_malloc(%u) failed\n", (unsigned)(TFLI_ARENA_SIZE + 16));
        g_tfli_status = -3;
        return;
    }
    uint8_t *arena = (uint8_t *)(((uintptr_t)arena_raw + 15) & ~(uintptr_t)15);

    static tflite::MicroInterpreter interpreter(model, resolver, arena, TFLI_ARENA_SIZE);

    if (interpreter.AllocateTensors() != kTfLiteOk) {
        printk("TFLI: AllocateTensors() failed (arena %u B too small?)\n",
               (unsigned)TFLI_ARENA_SIZE);
        g_tfli_status = -4;
        return;
    }
    g_tfli_arena_used = (uint32_t)interpreter.arena_used_bytes();

    TfLiteTensor *input  = interpreter.input(0);
    TfLiteTensor *output = interpreter.output(0);

    if ((uint32_t)input->bytes != g_cnn_input_len) {
        printk("TFLI: input size mismatch: tensor=%u bytes, data=%u bytes\n",
               (unsigned)input->bytes, (unsigned)g_cnn_input_len);
        g_tfli_status = -5;
        return;
    }
    memcpy(input->data.int8, g_cnn_input, g_cnn_input_len);

    const uint32_t tick_hz = sys_clock_hw_cycles_per_sec();
    printk("TFLI: arena_used=%u/%u B, iters/batch=%d, out_classes=%d\n",
           (unsigned)g_tfli_arena_used, (unsigned)TFLI_ARENA_SIZE,
           TFLI_ITERS, (int)output->dims->data[output->dims->size - 1]);

    /* Single batch: 1 warm-up invoke, then TFLI_ITERS timed invokes */
    for (int j = 0; j < 1; j++) {
        if (interpreter.Invoke() != kTfLiteOk) {
            printk("TFLI: warm-up Invoke failed\n");
            g_tfli_status = -6;
            return;
        }

        uint64_t min_t = UINT64_MAX, max_t = 0, sum_t = 0;
        uint64_t min_m = UINT64_MAX, max_m = 0, sum_m = 0;

        for (int i = 0; i < TFLI_ITERS; i++) {
            uint64_t q0 = sys_clock_cycle_get_64();
            uint64_t m0 = read_mcycle64();
            TfLiteStatus st = interpreter.Invoke();
            uint64_t m1 = read_mcycle64();
            uint64_t q1 = sys_clock_cycle_get_64();

            if (st != kTfLiteOk) {
                printk("TFLI: Invoke failed at i=%d\n", i);
                g_tfli_status = -7;
                return;
            }

            uint64_t dt = q1 - q0;
            uint64_t dm = m1 - m0;
            if (dt < min_t) min_t = dt;
            if (dt > max_t) max_t = dt;
            sum_t += dt;
            if (dm < min_m) min_m = dm;
            if (dm > max_m) max_m = dm;
            sum_m += dm;
        }

        uint64_t avg_t = sum_t / (uint64_t)TFLI_ITERS;
        uint64_t avg_m = sum_m / (uint64_t)TFLI_ITERS;

        uint32_t min_us  = (uint32_t)((min_t * 1000000ULL) / tick_hz);
        uint32_t avg_us  = (uint32_t)((avg_t * 1000000ULL) / tick_hz);
        uint32_t max_us  = (uint32_t)((max_t * 1000000ULL) / tick_hz);
        uint32_t cpu_mhz = (avg_us > 0) ? (uint32_t)(avg_m / (uint64_t)avg_us) : 0;

        int8_t  raw[TFLI_NUM_CLASSES];
        float   deq[TFLI_NUM_CLASSES];
        int32_t argmax = 0;
        int8_t  best   = output->data.int8[0];
        for (int c = 0; c < TFLI_NUM_CLASSES; c++) {
            raw[c] = output->data.int8[c];
            deq[c] = (raw[c] - output->params.zero_point) * output->params.scale;
            if (raw[c] > best) { best = raw[c]; argmax = c; }
        }

        g_tfli_min_ticks = (uint32_t)min_t;
        g_tfli_avg_ticks = (uint32_t)avg_t;
        g_tfli_max_ticks = (uint32_t)max_t;
        g_tfli_min_us = min_us; g_tfli_avg_us = avg_us; g_tfli_max_us = max_us;
        g_tfli_min_mcyc = min_m; g_tfli_avg_mcyc = avg_m; g_tfli_max_mcyc = max_m;
        g_tfli_cpu_mhz = cpu_mhz;
        for (int c = 0; c < TFLI_NUM_CLASSES; c++) {
            g_tfli_out[c]   = raw[c];
            g_tfli_out_f[c] = deq[c];
        }
        g_tfli_argmax = argmax;
        g_tfli_status = 0;
        g_tfli_batches = g_tfli_batches + 1;

        LOG_INF("TFLI[%u]: lat avg=%u us (min=%u max=%u) | cpu avg=%u cyc "
               "@~%u MHz | argmax=%d\n",
               (unsigned)g_tfli_batches, avg_us, min_us, max_us,
               (unsigned)avg_m, (unsigned)cpu_mhz, (int)argmax);

        k_sleep(K_MSEC(TFLI_BATCH_SLEEP_MS));
    }
}
```

> **关于 `__attribute__((retain))` 的说明：** `g_tfli_arena_size` 和 `g_tfli_iters` 仅在初始化时写入，固件代码从不读取它们。如果没有 `__attribute__((retain))`，即使它们是 `volatile`，链接器的 `--gc-sections` 过程也会悄无声息地丢弃它们的 ELF 段 — `volatile` 可防止**编译器**将其优化掉，但无法防止**链接器 GC**。`retain` 属性会将该段标记为始终存活，从而确保这些符号对 T32 保持可见。

***

## 10.2.6. 第 5 部分 — 集成到固件构建中

### 10.2.6.1. 步骤 1 — 将 `tflite_infer_run()` 集成到 `main.c` 中

文件：`modules/hal/qcom/core/main/src/main.c`

```c theme={null}
#if defined(CONFIG_TENSORFLOW_LITE_MICRO)
#include "tflite_infer.h"
#endif

int main(void)
{
    LOG_INF("Reached main.");
    printk("Reached main. Configuration: %s\n", CONFIG_BOARD);

#if defined(CONFIG_QC_CLK_READY) && defined(CONFIG_QC_SYS_M)
    // ... clk_ready smp2p signaling ...
#endif

#if defined(CONFIG_TENSORFLOW_LITE_MICRO)
    tflite_infer_run();
#endif
    return 0;
}
```

> **放置位置：** `tflite_infer_run()` 放在 `CONFIG_QC_CLK_READY` 代码块之后。如果该代码块不存在，则将其直接放在初始的 `LOG_INF` / `printk` 行之后，效果相同。

### 10.2.6.2. 步骤 2 — 在 CMake 中注册该模块

文件：`modules/hal/qcom/core/config/CMakeLists.txt`（已有文件 — 在 `add_subdirectory_ifdef` 代码块末尾添加一行）：

```cmake theme={null}
add_subdirectory_ifdef(CONFIG_TENSORFLOW_LITE_MICRO ${CORE_ROOT}/tflite_infer tflite_infer)
```

将其放在最后一个现有的 `add_subdirectory_ifdef` 行之后：

```cmake theme={null}
add_subdirectory_ifdef(CONFIG_QC_RINIT ${CORE_ROOT}/rinit rinit)
add_subdirectory_ifdef(CONFIG_TENSORFLOW_LITE_MICRO ${CORE_ROOT}/tflite_infer tflite_infer)   <- add this line
```

***

## 10.2.7. RAM 说明

在添加 TFLM 之前，Shikra LPAICP 固件大约有 40 KB 的空闲 RAM。TFLM 的静态代码占用（3 个算子子集约 25 KB）、模型数据（约 2 KB）以及堆分配可能会使镜像接近 3 MB 的上限。

如果构建失败并出现：

```
Error: Memory region RAM exceeded limit while adding section .last_ram_section
```

请在 `zephyr/kernel/config/config.yml` 中注释掉 `log.conf`：

```yaml theme={null}
- regex: "(^|shikra_)lpaicp_.*"
  files:
    - shikra_lpaicp.conf
    #- log.conf          <- comment this out to free ~20 KB
```

`log.conf` 会启用 Zephyr 的延迟日志子系统：一个 16 KB 的环形缓冲区以及一个具有 4 KB 栈的专用后台线程，共占用约 20 KB 静态 RAM。由于此固件使用 T32 RAM 控制台进行输出（无 UART），延迟日志线程没有输出路径。将其注释掉可回收约 20 KB，同时 `shikra_lpaicp.conf` 中的 `CONFIG_LOG_MODE_MINIMAL=y` 仍保持生效。

***

## 10.2.8. Kconfig 标志参考

为 TFLM 集成而添加或修改的所有 Kconfig 符号：

| 符号 | 值 | 含义 |
| - | - | - |
| `CONFIG_CPP` | y | 启用 C++ 支持 |
| `CONFIG_STD_CPP17` | y | 使用 C++17 标准 |
| `CONFIG_EXTERNAL_MODULE_LIBCPP` | y | 由模块提供 libc++ — 抑制 `qcld` 的 `-lc++` namespec |
| `CONFIG_TENSORFLOW_LITE_MICRO` | y | 启用 TFLM 运行时 |
| `CONFIG_REQUIRES_FLOAT_PRINTF` | y | 在 `printk` 中支持浮点数 |
| `CONFIG_MAIN_STACK_SIZE` | 8192 | 由 4096 增大 — TFLM 的解释器在主线程上运行，其 C++ 调用深度在 `AllocateTensors()` 期间会溢出默认的 4 KB。在 `shikra_lpaicp_tflite.conf` 中设置。 |
| `CONFIG_HEAP_MEM_POOL_SIZE` | 32768 | 32 KB 系统堆 — 必须 >= `TFLI_ARENA_KB * 1024 + 16`。减小此值可回收静态 RAM；请参阅第 1.5.2.1 节。 |

**推理框架参数**（`TFLI_ARENA_KB`、`TFLI_ITERS`、`TFLI_BATCH_SLEEP_MS`）是 `inc/tflite_infer.h` 中的 `#define` 常量。直接编辑该文件即可进行调整，无需重新编译其他任何文件。有关 arena 大小确定步骤，请参阅第 2.5.2.1 节。
