> ## Documentation Index
> Fetch the complete documentation index at: https://dragonwingdocs-staging.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Tflm runtime integration

# 10.2 TFLM Runtime & Model Embedding

This phase covers two things: integrating the TensorFlow Lite Micro (TFLM) runtime into the Q2390 / IQ2390 MCU firmware, and embedding the pre-built int8 model (produced in Phase 1) directly into the firmware image. It includes the non-obvious C++ runtime setup required on SDLLVM/RISC-V.

***

\##10.2.1	Overview

The integration has five distinct parts:

1. **Fetch TFLM source** — TFLM source is not vendored; it must be fetched manually
2. **Inspect model & verify op support** — extract op types from the model and confirm TFLM supports them
3. **Configure LLVM libc++** — SDLLVM doesn't configure libc++ by default; 5 specific fixes are required
4. **Create the inference module** — build the module with embedded model C arrays, inference harness (`tflite_infer.cc`), Kconfig header, and CMakeLists.txt
5. **Integrate into firmware build** — call `tflite_infer_run()` from `main.c` and register the module in CMake

***

\##10.2.2. Part 1 — Fetch TFLite Micro Source

> **Path convention:** Throughout this phase, `<wasp_proc>` refers to the root directory of your MCU firmware workspace — the directory containing `zephyr/`, `modules/`, `config/`, etc. Replace it with your actual path wherever it appears (e.g. `/path/to/your_workspace/wasp_proc`).

The TFLM module stubs (`CMakeLists.txt`, `Kconfig`) exist at `wasp_proc/zephyr/zephyr/modules/tflite-micro/` but the actual C++ source tree is absent — it must be cloned manually.

The `tflite-micro` project is defined in the west manifest (`zephyr/zephyr/submanifests/optional.yaml`) in the `optional` group, which Zephyr's own `west.yml` explicitly disables with `group-filter: [-optional]`. This means `west update tflite-micro` will **not** work — use `git clone` instead.

### 10.2.2.1. Clone the source

Inside the Docker build container:

```bash theme={null}
# Navigate to the wasp_proc root first
cd <wasp_proc>

# Read the pinned revision from the manifest — always in sync with the Zephyr version
TFLM_REV=$(grep -A6 "name: tflite-micro" \
    zephyr/zephyr/submanifests/optional.yaml | \
    grep "revision:" | awk '{print $2}')
echo "Fetching tflite-micro at revision: $TFLM_REV"

git clone https://github.com/zephyrproject-rtos/tflite-micro \
    modules/tflite-micro
cd modules/tflite-micro && git checkout $TFLM_REV && cd ../..
```

The revision is read directly from the manifest so this works for any LPAICP release without modification. The resulting detached HEAD state is expected and correct.

### 10.2.2.2. Verify

```bash theme={null}
ls modules/tflite-micro/tensorflow/lite/micro/kernels/*.cc | wc -l
# Should print a non-zero number (e.g. 189)
```

***

## 10.2.3. Part 2 — Inspect Model & Verify Op Support

Before writing any firmware code, we need two things from the model:

1. **Which ops the model uses** — the TFLM runtime requires every operator used by the model to be explicitly registered before inference can run. If any operator is missing, the firmware will fail at startup with an unresolved op error. This step identifies the exact set of operators your model needs, so you know what to register when integrating the inference harness in [Part 5](#26-part-5--integrate-into-firmware-build).
2. **Input/output quantization parameters** (scale + zero\_point) — if the model is int8-quantized, float inputs must be pre-quantized to int8 before passing to the inference harness.

### 10.2.3.1. Option A — `flatc` (full topology decode — recommended)

Decodes the model binary into human-readable JSON using the TFLM schema. Gives you: op types, tensor names, shapes, and quantization params for all tensors.

**Step 1 — Run flatc to decode the model:**

```bash theme={null}
sed -i 's/ (deprecated)//g' \
    zephyr/zephyr/modules/tflite-micro/tensorflow/lite/schema/schema.fbs
flatc --json --strict-json --raw-binary \
    zephyr/zephyr/modules/tflite-micro/tensorflow/lite/schema/schema.fbs \
    -- cnn1d_minimal_int8.tflite
```

This produces `cnn1d_minimal_int8.json` in the current directory.

**Step 2 — Extract the op list from the JSON:**

```python theme={null}
import json
with open("cnn1d_minimal_int8.json") as f:
    model = json.load(f)
opcodes   = model["operator_codes"]
operators = model["subgraphs"][0]["operators"]
ops_used  = set()
for op in operators:
    idx   = op["opcode_index"]
    entry = opcodes[idx]
    name  = entry.get("builtin_code", entry.get("custom_code", "UNKNOWN"))
    ops_used.add(name)
print("Ops used by model:")
for op in sorted(ops_used):
    print(f"  {op}")
```

**Step 3 — Cross-check each op against the supported op list in 1.3.3.**

### 10.2.3.2. Option B — `tf.lite.Interpreter` (quant params only, no `flatc` needed)

Use when `flatc` is not available. Gives you input/output quantization params (scale, zero\_point). **Does not expose op types** — use Option A for the op list.

```python theme={null}
import tensorflow as tf
interp = tf.lite.Interpreter("cnn1d_minimal_int8.tflite")
interp.allocate_tensors()
inp = interp.get_input_details()[0]
out = interp.get_output_details()[0]
print(f"Input: shape={inp['shape']}, dtype={inp['dtype']}, "
      f"scale={inp['quantization'][0]:.6f}, zp={inp['quantization'][1]}")
print(f"Output: shape={out['shape']}, dtype={out['dtype']}, "
      f"scale={out['quantization'][0]:.6f}, zp={out['quantization'][1]}")
```

> `tf.lite.Interpreter` does not expose op types — use Option A (`flatc`) to get the op list.

### 10.2.3.3. All ops supported by this TFLM build

The authoritative source for which ops TFLM supports is `micro_mutable_op_resolver.h` — every `Add*()` method declared in that file is a supported op for the exact revision downloaded.

```bash theme={null}
grep "TfLiteStatus Add" \
    modules/tflite-micro/tensorflow/lite/micro/micro_mutable_op_resolver.h \
    | grep -v "AddCustom\|AddBuiltin" \
    | sed 's/.*TfLiteStatus Add\([A-Za-z0-9]*\).*/\1/' \
    | sort | uniq
```

**Complete supported op list for release tag `zephyr_20240627` — 114 ops:**

| Op | Op | Op | Op |
| - | - | - | - |
| Abs | Add | AddN | ArgMax |
| ArgMin | AssignVariable | AveragePool2D | BatchMatMul |
| BatchToSpaceNd | BroadcastArgs | BroadcastTo | CallOnce |
| Cast | Ceil | CircularBuffer | Concatenation |
| Conv2D | Cos | CumSum | Delay |
| DepthToSpace | DepthwiseConv2D | Dequantize | DetectionPostprocess |
| Div | Elu | EmbeddingLookup | Energy |
| Equal | EthosU | Exp | ExpandDims |
| FftAutoScale | Fill | FilterBank | FilterBankLog |
| FilterBankSpectralSubtraction | FilterBankSquareRoot | Floor | FloorDiv |
| FloorMod | Framer | FullyConnected | Gather |
| GatherNd | Greater | GreaterEqual | HardSwish |
| If | Irfft | L2Normalization | L2Pool2D |
| LeakyRelu | Less | LessEqual | Log |
| LogSoftmax | LogicalAnd | LogicalNot | LogicalOr |
| Logistic | Maximum | MaxPool2D | Mean |
| Minimum | MirrorPad | Mul | Neg |
| NotEqual | OverlapAdd | Pack | Pad |
| PadV2 | PCAN | Prelu | Quantize |
| ReadVariable | ReduceMax | Relu | Relu6 |
| Reshape | ResizeBilinear | ResizeNearestNeighbor | Rfft |
| Round | Rsqrt | SelectV2 | Shape |
| Sin | Slice | Softmax | SpaceToBatchNd |
| SpaceToDepth | Split | SplitV | Sqrt |
| Square | SquaredDifference | Squeeze | Stacker |
| StridedSlice | Sub | Sum | Svdf |
| Tanh | Transpose | TransposeConv | UnidirectionalSequenceLSTM |
| Unpack | VarHandle | While | Window |
| ZerosLike | | | |

> The resolver method for each op is `Add` + op name — e.g. `Conv2D` → `AddConv2D()`, `FullyConnected` → `AddFullyConnected()`.

### 10.2.3.4. Reference — ops confirmed for the example model

All ops in `cnn1d_minimal_int8.tflite`, their TFLM kernel files, and the resolver method:

| Op | TFLM kernel file | `MicroMutableOpResolver` method |
| - | - | - |
| `CONV_2D` | `conv.cc` | `AddConv2D()` |
| `MEAN` | `reduce.cc` | `AddMean()` |
| `FULLY_CONNECTED` | `fully_connected.cc` | `AddFullyConnected()` |

**If an op is not registered, you will see one of these errors:**

Compile-time (wrong or missing method name):

```
error: 'class tflite::MicroMutableOpResolver<N>' has no member named 'AddMyOp'
```

Runtime at `AllocateTensors()` (op not registered):

```
Didn't find op for builtin opcode 'CONV_2D' version '1'
```

***

## 10.2.4. Part 3 — Configure LLVM libc++ (5-Step Fix)

### 10.2.4.1. Why the default approach fails

To use TFLM (written in C++), Zephyr needs a C++ standard library. The natural starting point is enabling `CONFIG_LIBCXX_LIBCPP=y`. On SDLLVM, however, this triggers a Kconfig dependency chain that breaks the build. The failure appears as an error in a completely unrelated kernel file:

```
fatal error: 'string.h' file not found in offsets.c
```

### 10.2.4.2. The 4-step fix

#### 10.2.4.2.1. Step 1 — Remove `select REQUIRES_FULL_LIBCPP` from TFLM Kconfig

File: `zephyr/zephyr/modules/tflite-micro/Kconfig`

Remove the `select REQUIRES_FULL_LIBCPP` line from inside `config TENSORFLOW_LITE_MICRO`. This is the root of the dependency cascade — removing it prevents Zephyr from triggering the broken `REQUIRES_FULL_LIBC` path on SDLLVM.

```
# Remove this line from config TENSORFLOW_LITE_MICRO:
select REQUIRES_FULL_LIBCPP
```

#### 10.2.4.2.2. Step 2 — Use `EXTERNAL_MODULE_LIBCPP`, not `EXTERNAL_LIBCPP`

*No action here — this setting is included in the new file created in section 2.4.4.*

| Symbol | Avoids broken dependency chain? | Suppresses `-lc++` namespec for `qcld`? |
| - | - | - |
| `CONFIG_LIBCXX_LIBCPP=y` | No — triggers the broken chain | No |
| `CONFIG_EXTERNAL_LIBCPP=y` | Yes | No — still emits `-lc++ -lc++abi` which `qcld` cannot parse |
| `CONFIG_EXTERNAL_MODULE_LIBCPP=y` | Yes | Yes — correct choice |

#### 10.2.4.2.3. Step 3 — Inject libc++ includes via `-I`, not `-isystem`

*No action here — this call is part of the complete block added in section 2.4.3.*

`-I` directories are searched before any `-isystem` directory. By injecting libc++ headers as `-I BEFORE`, clang finds libc++'s own `<stdio.h>`/`<string.h>` wrappers first, which chain to musl via `#include_next` — the correct resolution order for C++ translation units.

#### 10.2.4.2.4. Step 4 — Compile C++ TUs with `-DNDEBUG`

*No action here — this flag is part of the complete block added in section 2.4.3.*

musl's `assert()` macro expands to `__assert_fail()`, which picolibc does not provide. `-DNDEBUG` compiles all assert calls out in C++ translation units. Scoped to `$<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG>` so C code is untouched.

### 10.2.4.3. Complete CMake configuration block

File: `modules/hal/qcom/core/config/CMakeLists.txt` (pre-existing file — add this complete block at the end). Gated on `CONFIG_EXTERNAL_MODULE_LIBCPP` so it only activates when the TFLM Kconfig fragment is merged in.

```cmake theme={null}
if(CONFIG_EXTERNAL_MODULE_LIBCPP)
    get_filename_component(_sdllvm_bin ${CMAKE_C_COMPILER} DIRECTORY)
    get_filename_component(_sdllvm_root ${_sdllvm_bin} DIRECTORY)
    set(_sdllvm_sysroot ${_sdllvm_root}/riscv32-unknown-elf)
    file(GLOB _sdllvm_builtin_inc ${_sdllvm_root}/lib/clang/*/include)
    file(GLOB _sdllvm_builtin_lib
        ${_sdllvm_root}/lib/clang/*/lib/baremetal/libclang_rt.builtins-riscv32.a)
    if(NOT EXISTS ${_sdllvm_sysroot}/include/c++/v1/vector)
        message(FATAL_ERROR
            "CONFIG_EXTERNAL_MODULE_LIBCPP set but libc++ headers not found at "
            "${_sdllvm_sysroot}/include/c++/v1 (compiler=${CMAKE_C_COMPILER}). "
            "The contained libc++ configuration in core/config/CMakeLists.txt assumes "
            "the Snapdragon LLVM riscv32-unknown-elf layout.")
    endif()
    target_include_directories(zephyr_interface BEFORE INTERFACE
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_sysroot}/include/c++/v1>
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_sysroot}/libc/include>
        $<$<COMPILE_LANGUAGE:CXX>:${_sdllvm_builtin_inc}>
    )
    zephyr_compile_options($<$<COMPILE_LANGUAGE:CXX>:-DNDEBUG>)
    zephyr_link_libraries(
        -Wl,--start-group
        ${_sdllvm_sysroot}/lib/libc++.a
        ${_sdllvm_sysroot}/lib/libc++abi.a
        ${_sdllvm_sysroot}/lib/libunwind.a
        ${_sdllvm_builtin_lib}
        -Wl,--end-group
    )
endif()
```

### 10.2.4.4. `shikra_lpaicp_tflite.conf`

#### 10.2.4.4.1. Step 1 — Create `shikra_lpaicp_tflite.conf`

Create a new file at `zephyr/kernel/config/shikra_lpaicp_tflite.conf`:

```
CONFIG_CPP=y
CONFIG_STD_CPP17=y
CONFIG_EXTERNAL_MODULE_LIBCPP=y
CONFIG_TENSORFLOW_LITE_MICRO=y
CONFIG_REQUIRES_FLOAT_PRINTF=y
CONFIG_MAIN_STACK_SIZE=8192
```

#### 10.2.4.4.2. Step 2 — Register the file in `config.yml`

File: `zephyr/kernel/config/config.yml` (pre-existing file — add a new entry):

```yaml theme={null}
- regex: "shikra_lpaicp_.*"
  files:
    - shikra_lpaicp_tflite.conf
```

> **Why a separate file instead of adding to `shikra_lpaicp.conf`?** The existing entry uses regex `(^|shikra_)lpaicp_.*`, which matches both the hardware target (`SHIKRA_LPAICP_TEST`) and the QEMU simulation variant that uses Zephyr-SDK GCC. Adding TFLM's SDLLVM-specific libc++ configuration to `shikra_lpaicp.conf` would break the GCC-based QEMU build. The narrower regex `shikra_lpaicp_.*` matches only the hardware target.

***

## 10.2.5. Part 4 — Create the Inference Module

### 10.2.5.1. Step 1 — Generate src/model\_data.cpp and src/input\_data.cpp

Create the module directories first, then convert the model and input files produced in Phase 1 into C byte arrays. Run the commands from the directory containing the Phase 1 source files.

```bash theme={null}
mkdir -p modules/hal/qcom/core/tflite_infer/{inc,src}

# Model byte array  ->  g_cnn_model[], g_cnn_model_len
xxd -i model_files/cnn1d_minimal_int8.tflite \
  | sed -e 's/^unsigned char [a-zA-Z0-9_]*/alignas(8) extern const unsigned char g_cnn_model/' \
        -e 's/^unsigned int [a-zA-Z0-9_]*/extern const unsigned int g_cnn_model_len/' \
  > <wasp_proc>/modules/hal/qcom/core/tflite_infer/src/model_data.cpp

# Input byte array  ->  g_cnn_input[], g_cnn_input_len
xxd -i inputs/input_0.bin \
  | sed -e 's/^unsigned char [a-zA-Z0-9_]*/extern const unsigned char g_cnn_input/' \
        -e 's/^unsigned int [a-zA-Z0-9_]*/extern const unsigned int g_cnn_input_len/' \
  > <wasp_proc>/modules/hal/qcom/core/tflite_infer/src/input_data.cpp
```

Replace `<wasp_proc>` with the path to your `wasp_proc` MCU code base.

The generated files look like this:

```c theme={null}
/* model_data.cpp */
alignas(8) extern const unsigned char g_cnn_model[] = {
  0x20, 0x00, 0x00, 0x00, 0x54, 0x46, 0x4c, 0x33, ...
};
extern const unsigned int g_cnn_model_len = 2096;

/* input_data.cpp */
extern const unsigned char g_cnn_input[] = {
  0x07, 0xff, 0x16, 0x06, 0xf3, 0x0e, 0x2a, 0x1f, ...
};
extern const unsigned int g_cnn_input_len = 200;
```

> **`alignas(8)` on the model array:** TFLM's FlatBuffers parser requires the model byte array to be at least 4-byte aligned. `alignas(8)` guarantees this regardless of where the linker places the symbol.

### 10.2.5.2. Step 2 — Create `inc/tflite_infer.h`

File: `modules/hal/qcom/core/tflite_infer/inc/tflite_infer.h` (new file).

```c theme={null}
#ifndef TFLITE_INFER_H_
#define TFLITE_INFER_H_
/* Tuneable inference parameters — edit here to adjust without recompiling other files */
#define TFLI_ARENA_KB         4    /* tensor arena in KB — measure via g_tfli_arena_used */
#define TFLI_ITERS            100  /* timed Invoke() calls per batch (after one warm-up) */
#define TFLI_BATCH_SLEEP_MS 1000   /* sleep between re-measure batches (ms) */
#ifdef __cplusplus
extern "C" {
#endif
/* TFLM inference harness. Loads the 8-bit quantized CNN model, feeds the
 * pre-quantized sample input, and times TFLI_ITERS invocations per batch.
 * Per-batch min/avg/max latency is published into the g_tfli_* volatile
 * globals for T32 inspection and also printed to the RAM console.
 * Call once from main() after boot initialisation. */
void tflite_infer_run(void);
#ifdef __cplusplus
}
#endif
#endif /* TFLITE_INFER_H_ */
```

#### 10.2.5.2.1. Tuning `TFLI_ARENA_KB`

`TFLI_ARENA_KB` sets the size of the tensor arena — the contiguous block of memory TFLM uses for input/output tensors, intermediate activation buffers between layers, and its internal allocator metadata. It is allocated from the system heap via `k_malloc` at the start of `tflite_infer_run()`.

**What goes wrong if the value is wrong:**

| Symptom in RAM console | `g_tfli_status` | Root cause |
| - | - | - |
| `TFLI: k_malloc(N) failed` | `-3` | `TFLI_ARENA_KB * 1024 + 16` exceeds `CONFIG_HEAP_MEM_POOL_SIZE`. Either increase the heap or reduce `TFLI_ARENA_KB`. |
| `TFLI: AllocateTensors() failed (arena N B too small?)` | `-4` | The arena is too small for the model's tensors and TFLM metadata. Increase `TFLI_ARENA_KB`. |

Both failures cause `tflite_infer_run()` to return early — `g_tfli_batches` remains `0` and all latency globals stay at their initial values.

**How to find the right value — measure with `g_tfli_arena_used`:**

1. Set `TFLI_ARENA_KB` to a generous starting value (e.g. `32`) — large enough to guarantee `AllocateTensors()` succeeds.
2. Build, flash, boot, and run `read_tflite_infer.cmm` in T32.
3. Confirm `g_tfli_status == 0` and `g_tfli_batches >= 1`.
4. Read `g_tfli_arena_used` from the T32 Var.View window — this is the exact byte count TFLM required.
5. Compute the minimum safe value and update the header:

```
TFLI_ARENA_KB = ceil(g_tfli_arena_used / 1024) + 2    <- +2 KB safety margin
```

6. Rebuild and reflash. Verify `g_tfli_arena_used` is still below `g_tfli_arena_size`.

> **Configured value for the reference model:** `TFLI_ARENA_KB = 4` (4,096 bytes). This is the value used in the reference codebase for `cnn1d_minimal_int8.tflite`. Measure `g_tfli_arena_used` on-target after a successful run to confirm utilisation for your model.

**RAM impact of `TFLI_ARENA_KB`:**

The arena is runtime-allocated from the system heap — it does not contribute directly to static image size. However, the heap backing buffer (`kheap_buf__system_heap`) is a **statically linked section** sized by `CONFIG_HEAP_MEM_POOL_SIZE` (set to `32768` in `shikra_lpaicp.conf`). Reducing `TFLI_ARENA_KB` alone does not reduce static RAM. To reclaim static RAM, also lower `CONFIG_HEAP_MEM_POOL_SIZE` to match — ensure the new size covers `TFLI_ARENA_KB * 1024 + 16` plus any other `k_malloc` callers.

### 10.2.5.3. Step 3 — Create `CMakeLists.txt`

File: `modules/hal/qcom/core/tflite_infer/CMakeLists.txt` (new file). Gated on `CONFIG_TENSORFLOW_LITE_MICRO`.

```cmake theme={null}
if(CONFIG_TENSORFLOW_LITE_MICRO)
    target_include_directories(${ZEPHYR_APP_NAME} PRIVATE
        inc
    )
    target_sources(${ZEPHYR_APP_NAME} PRIVATE
        src/tflite_infer.cc
        src/model_data.cpp
        src/input_data.cpp
    )
endif()
```

### 10.2.5.4. Step 4 — Create `src/tflite_infer.cc`

File: `modules/hal/qcom/core/tflite_infer/src/tflite_infer.cc` (new file). Runs `TFLI_ITERS` timed `Invoke()` calls per batch (after one warm-up), then publishes per-batch min/avg/max latency in microseconds, QTMR ticks, and CPU cycles into the `g_tfli_*` volatile globals.

```cpp theme={null}
#include <stdint.h>
#include <string.h>

#include <zephyr/kernel.h>
#include <zephyr/sys/printk.h>
#include <zephyr/logging/log.h>

LOG_MODULE_REGISTER(tflite_infer, LOG_LEVEL_DBG);

#include <tensorflow/lite/micro/micro_mutable_op_resolver.h>
#include <tensorflow/lite/micro/micro_interpreter.h>
#include <tensorflow/lite/micro/micro_log.h>
#include <tensorflow/lite/schema/schema_generated.h>

#include "tflite_infer.h"   /* TFLI_ARENA_KB, TFLI_ITERS, TFLI_BATCH_SLEEP_MS */

/* Model and input data — defined in model_data.cpp and input_data.cpp */
extern const unsigned char g_cnn_model[];
extern const unsigned char g_cnn_input[];
extern const unsigned int  g_cnn_input_len;

#define TFLI_ARENA_SIZE   (TFLI_ARENA_KB * 1024)
#define TFLI_NUM_CLASSES  2   /* matches Dense(2) output layer */

/* T32-readable results (volatile so they survive optimization) */
volatile int32_t  g_tfli_status     = -1;
volatile uint32_t g_tfli_arena_size __attribute__((retain)) = TFLI_ARENA_SIZE;
volatile uint32_t g_tfli_arena_used = 0;
volatile uint32_t g_tfli_iters      __attribute__((retain)) = TFLI_ITERS;
volatile uint32_t g_tfli_batches    = 0;

/* QTMR latency per inference (19.2 MHz ticks + microseconds) */
volatile uint32_t g_tfli_min_ticks = 0, g_tfli_avg_ticks = 0, g_tfli_max_ticks = 0;
volatile uint32_t g_tfli_min_us    = 0, g_tfli_avg_us    = 0, g_tfli_max_us    = 0;

/* CPU cycles per inference (RISC-V mcycle) */
volatile uint64_t g_tfli_min_mcyc = 0, g_tfli_avg_mcyc = 0, g_tfli_max_mcyc = 0;
volatile uint32_t g_tfli_cpu_mhz  = 0;

/* Raw + dequantized output tensor (last invoke of the most recent batch) */
volatile int8_t  g_tfli_out[TFLI_NUM_CLASSES]   = {0};
volatile float   g_tfli_out_f[TFLI_NUM_CLASSES] = {0};
volatile int32_t g_tfli_argmax = -1;

/* RISC-V 64-bit machine cycle counter (rv32: split across mcycle/mcycleh) */
static inline uint64_t read_mcycle64(void)
{
    uint32_t hi, lo, hi2;
    do {
        __asm__ volatile("csrr %0, mcycleh" : "=r"(hi));
        __asm__ volatile("csrr %0, mcycle"  : "=r"(lo));
        __asm__ volatile("csrr %0, mcycleh" : "=r"(hi2));
    } while (hi != hi2);
    return ((uint64_t)hi << 32) | lo;
}

extern "C" void tflite_infer_run(void)
{
    const tflite::Model *model = tflite::GetModel(g_cnn_model);
    if (model->version() != TFLITE_SCHEMA_VERSION) {
        printk("TFLI: model schema %u != supported %u\n",
               (unsigned)model->version(), (unsigned)TFLITE_SCHEMA_VERSION);
        g_tfli_status = -2;
        return;
    }

    /* 3 ops: Conv2D (fused BN+ReLU), Mean (GlobalAvgPool), FullyConnected */
    static tflite::MicroMutableOpResolver<3> resolver;
    resolver.AddConv2D();
    resolver.AddMean();
    resolver.AddFullyConnected();

    /* Arena from the system heap — over-allocate by 16 for manual alignment */
    uint8_t *arena_raw = (uint8_t *)k_malloc(TFLI_ARENA_SIZE + 16);
    if (arena_raw == NULL) {
        printk("TFLI: k_malloc(%u) failed\n", (unsigned)(TFLI_ARENA_SIZE + 16));
        g_tfli_status = -3;
        return;
    }
    uint8_t *arena = (uint8_t *)(((uintptr_t)arena_raw + 15) & ~(uintptr_t)15);

    static tflite::MicroInterpreter interpreter(model, resolver, arena, TFLI_ARENA_SIZE);

    if (interpreter.AllocateTensors() != kTfLiteOk) {
        printk("TFLI: AllocateTensors() failed (arena %u B too small?)\n",
               (unsigned)TFLI_ARENA_SIZE);
        g_tfli_status = -4;
        return;
    }
    g_tfli_arena_used = (uint32_t)interpreter.arena_used_bytes();

    TfLiteTensor *input  = interpreter.input(0);
    TfLiteTensor *output = interpreter.output(0);

    if ((uint32_t)input->bytes != g_cnn_input_len) {
        printk("TFLI: input size mismatch: tensor=%u bytes, data=%u bytes\n",
               (unsigned)input->bytes, (unsigned)g_cnn_input_len);
        g_tfli_status = -5;
        return;
    }
    memcpy(input->data.int8, g_cnn_input, g_cnn_input_len);

    const uint32_t tick_hz = sys_clock_hw_cycles_per_sec();
    printk("TFLI: arena_used=%u/%u B, iters/batch=%d, out_classes=%d\n",
           (unsigned)g_tfli_arena_used, (unsigned)TFLI_ARENA_SIZE,
           TFLI_ITERS, (int)output->dims->data[output->dims->size - 1]);

    /* Single batch: 1 warm-up invoke, then TFLI_ITERS timed invokes */
    for (int j = 0; j < 1; j++) {
        if (interpreter.Invoke() != kTfLiteOk) {
            printk("TFLI: warm-up Invoke failed\n");
            g_tfli_status = -6;
            return;
        }

        uint64_t min_t = UINT64_MAX, max_t = 0, sum_t = 0;
        uint64_t min_m = UINT64_MAX, max_m = 0, sum_m = 0;

        for (int i = 0; i < TFLI_ITERS; i++) {
            uint64_t q0 = sys_clock_cycle_get_64();
            uint64_t m0 = read_mcycle64();
            TfLiteStatus st = interpreter.Invoke();
            uint64_t m1 = read_mcycle64();
            uint64_t q1 = sys_clock_cycle_get_64();

            if (st != kTfLiteOk) {
                printk("TFLI: Invoke failed at i=%d\n", i);
                g_tfli_status = -7;
                return;
            }

            uint64_t dt = q1 - q0;
            uint64_t dm = m1 - m0;
            if (dt < min_t) min_t = dt;
            if (dt > max_t) max_t = dt;
            sum_t += dt;
            if (dm < min_m) min_m = dm;
            if (dm > max_m) max_m = dm;
            sum_m += dm;
        }

        uint64_t avg_t = sum_t / (uint64_t)TFLI_ITERS;
        uint64_t avg_m = sum_m / (uint64_t)TFLI_ITERS;

        uint32_t min_us  = (uint32_t)((min_t * 1000000ULL) / tick_hz);
        uint32_t avg_us  = (uint32_t)((avg_t * 1000000ULL) / tick_hz);
        uint32_t max_us  = (uint32_t)((max_t * 1000000ULL) / tick_hz);
        uint32_t cpu_mhz = (avg_us > 0) ? (uint32_t)(avg_m / (uint64_t)avg_us) : 0;

        int8_t  raw[TFLI_NUM_CLASSES];
        float   deq[TFLI_NUM_CLASSES];
        int32_t argmax = 0;
        int8_t  best   = output->data.int8[0];
        for (int c = 0; c < TFLI_NUM_CLASSES; c++) {
            raw[c] = output->data.int8[c];
            deq[c] = (raw[c] - output->params.zero_point) * output->params.scale;
            if (raw[c] > best) { best = raw[c]; argmax = c; }
        }

        g_tfli_min_ticks = (uint32_t)min_t;
        g_tfli_avg_ticks = (uint32_t)avg_t;
        g_tfli_max_ticks = (uint32_t)max_t;
        g_tfli_min_us = min_us; g_tfli_avg_us = avg_us; g_tfli_max_us = max_us;
        g_tfli_min_mcyc = min_m; g_tfli_avg_mcyc = avg_m; g_tfli_max_mcyc = max_m;
        g_tfli_cpu_mhz = cpu_mhz;
        for (int c = 0; c < TFLI_NUM_CLASSES; c++) {
            g_tfli_out[c]   = raw[c];
            g_tfli_out_f[c] = deq[c];
        }
        g_tfli_argmax = argmax;
        g_tfli_status = 0;
        g_tfli_batches = g_tfli_batches + 1;

        LOG_INF("TFLI[%u]: lat avg=%u us (min=%u max=%u) | cpu avg=%u cyc "
               "@~%u MHz | argmax=%d\n",
               (unsigned)g_tfli_batches, avg_us, min_us, max_us,
               (unsigned)avg_m, (unsigned)cpu_mhz, (int)argmax);

        k_sleep(K_MSEC(TFLI_BATCH_SLEEP_MS));
    }
}
```

> **Note on `__attribute__((retain))`:** `g_tfli_arena_size` and `g_tfli_iters` are only written at initialisation and never read by firmware code. Without `__attribute__((retain))`, the linker's `--gc-sections` pass silently drops their ELF sections even though they are `volatile` — `volatile` prevents the **compiler** from optimising them away but does not protect against **linker GC**. The `retain` attribute marks the section as always-live, ensuring the symbols remain visible to T32.

***

## 10.2.6. Part 5 — Integrate into Firmware Build

### 10.2.6.1. Step 1 — Integrate `tflite_infer_run()` into `main.c`

File: `modules/hal/qcom/core/main/src/main.c`

```c theme={null}
#if defined(CONFIG_TENSORFLOW_LITE_MICRO)
#include "tflite_infer.h"
#endif

int main(void)
{
    LOG_INF("Reached main.");
    printk("Reached main. Configuration: %s\n", CONFIG_BOARD);

#if defined(CONFIG_QC_CLK_READY) && defined(CONFIG_QC_SYS_M)
    // ... clk_ready smp2p signaling ...
#endif

#if defined(CONFIG_TENSORFLOW_LITE_MICRO)
    tflite_infer_run();
#endif
    return 0;
}
```

> **Placement:** `tflite_infer_run()` goes after the `CONFIG_QC_CLK_READY` block. If that block is not present, placing it directly after the initial `LOG_INF` / `printk` lines is equivalent.

### 10.2.6.2. Step 2 — Register the module in CMake

File: `modules/hal/qcom/core/config/CMakeLists.txt` (pre-existing file — add one line at the end of the `add_subdirectory_ifdef` block):

```cmake theme={null}
add_subdirectory_ifdef(CONFIG_TENSORFLOW_LITE_MICRO ${CORE_ROOT}/tflite_infer tflite_infer)
```

Place it after the last existing `add_subdirectory_ifdef` line:

```cmake theme={null}
add_subdirectory_ifdef(CONFIG_QC_RINIT ${CORE_ROOT}/rinit rinit)
add_subdirectory_ifdef(CONFIG_TENSORFLOW_LITE_MICRO ${CORE_ROOT}/tflite_infer tflite_infer)   <- add this line
```

***

## 10.2.7. RAM Note

The Shikra LPAICP firmware has approximately 40 KB of free RAM before TFLM is added. TFLM's static code footprint (\~25 KB for the 3-op subset), model data (\~2 KB), and heap allocation can bring the image close to the 3 MB limit.

If the build fails with:

```
Error: Memory region RAM exceeded limit while adding section .last_ram_section
```

Comment out `log.conf` in `zephyr/kernel/config/config.yml`:

```yaml theme={null}
- regex: "(^|shikra_)lpaicp_.*"
  files:
    - shikra_lpaicp.conf
    #- log.conf          <- comment this out to free ~20 KB
```

`log.conf` enables Zephyr's deferred logging subsystem: a 16 KB ring buffer and a dedicated background thread with a 4 KB stack, consuming \~20 KB of static RAM. Since this firmware uses the T32 RAM console for output (no UART), the deferred logging thread has no delivery path. Commenting it out reclaims \~20 KB while `CONFIG_LOG_MODE_MINIMAL=y` from `shikra_lpaicp.conf` remains active.

***

## 10.2.8. Kconfig Flag Reference

All Kconfig symbols added or modified for TFLM integration:

| Symbol | Value | Meaning |
| - | - | - |
| `CONFIG_CPP` | y | Enable C++ support |
| `CONFIG_STD_CPP17` | y | Use C++17 standard |
| `CONFIG_EXTERNAL_MODULE_LIBCPP` | y | libc++ provided by module — suppresses `qcld`'s `-lc++` namespec |
| `CONFIG_TENSORFLOW_LITE_MICRO` | y | Enable TFLM runtime |
| `CONFIG_REQUIRES_FLOAT_PRINTF` | y | Float support in `printk` |
| `CONFIG_MAIN_STACK_SIZE` | 8192 | Increased from 4096 — TFLM's interpreter runs on the main thread and its C++ call depth overflows the 4 KB default during `AllocateTensors()`. Set in `shikra_lpaicp_tflite.conf`. |
| `CONFIG_HEAP_MEM_POOL_SIZE` | 32768 | 32 KB system heap — must be >= `TFLI_ARENA_KB * 1024 + 16`. Reducing this reclaims static RAM; see section 1.5.2.1. |

**Inference harness parameters** (`TFLI_ARENA_KB`, `TFLI_ITERS`, `TFLI_BATCH_SLEEP_MS`) are `#define` constants in `inc/tflite_infer.h`. Edit that file directly to tune without recompiling any other file. See section 2.5.2.1 for the arena sizing procedure.
