Skip to main content
This page explains how to integrate a custom post-processing module into a QIM SDK pipeline when the built-in qtimlpostprocess modules do not support your model’s output format. Use this workflow when your model can run with an existing QIM SDK inference plugin, but its output tensors require custom processing before they can be converted into labels, bounding boxes, masks, keypoints, raw tensors, or other application-specific metadata. This guide uses a YOLOv8-based object detection model as an end-to-end example to demonstrate each step of the integration process.
Step 3a of the application development workflow. Use this path when onboarding a new model whose output format is not supported by any of the built-in post-processing modules. If you have not yet verified that a custom post-processing module is required, first complete step 2a.

Before you begin

Before creating a custom post-processing module, verify the following prerequisites. A custom implementation is only necessary when none of the built-in qtimlpostprocess modules can process your model’s output format.
It is very common for one module to serve several models. For example, the yolov8 module handles multiple YOLO-family detectors whose output structure is the same. Check the reported tensor shapes carefully before concluding that a custom module is required.
The following diagram shows the overall process to add your own post-processing support, from developing and integrating the module to running the reference application. Overview of the process to add custom post-processing support

AI pipeline overview

The Qualcomm Intelligent Multimedia SDK (QIM SDK) provides the building blocks to construct AI, multimedia, and computer vision pipelines. An AI workflow is built from three key GStreamer stages, and a custom post-processing module extends only the third stage. The post-processing stage can deliver metadata in one of the following ways:
  • Attach it to the source stream using qtimetamuxer.
  • Stream it directly to endpoints such as RTSP, RTMP, or Redis.
  • Convert it to an image mask that is overlaid on the source video frame using qtivcomposer.
AI IM SDK pipeline stages
Only the post-processing module is your responsibility. Preprocessing, inference, and metadata consumption are standard QIM SDK elements, and the qtimlpostprocess plugin handles batching, output-format negotiation, and result limiting around your module.

Example: Use ML metadata directly

In this example, the source stream is not propagated after the inference plugin — the metadata is consumed directly. Pipeline that consumes ML metadata directly

Example: Attach ML metadata to the source video

In this example, the ML metadata is attached to the source video. An overlay uses the attached metadata to draw bounding boxes, text, and other visual elements. The result is either displayed on screen or streamed over a network. Pipeline that attaches ML metadata to the source video

Example: Convert ML metadata to an image mask

In this example, the ML metadata is converted into an image mask and then blitted on top of the source stream. Pipeline that converts ML metadata to an image mask

qtimlpostprocess plugin behavior

qtimlpostprocess is a customizable plugin that provides a library interface for post-processing the tensor output of inference plugins. The post-processing module is responsible for tensor parsing and produces a list of predictions. Each post-processing module handles one type of ML model and its variants — for example, all YOLOv8 detection variants. The plugin manages module execution, output generation (ML metadata, image masks, or tensors), batching, ML staging, and related tasks. Relationship between the inputs, outputs, post-processing module, and the plugin The plugin receives a list of tensors as input, encapsulated in GStreamer buffers. ML metadata attached to each buffer specifies the number of tensors, tensor shapes, model input tensor shapes, how much of each input tensor is filled with stream data, timestamps, and batching indexes.

Plugin properties

Output formats

The output format is negotiated through the GStreamer pipeline. The most suitable format is negotiated automatically, but you can specify it manually with a GStreamer caps filter. The plugin supports a single source pad, so if the application needs more than one output format simultaneously, add another qtimlpostprocess instance in the pipeline. The plugin supports the following model types:
  • Object detection
  • Image classification
  • Image segmentation
  • Super resolution
  • Pose estimation
  • Audio classification

Write a post-processing module

A post-processing module is a shared library that parses the tensor output from inference plugins. The qtimlpostprocess plugin loads and runs the module at runtime. QIM SDK provides a wide variety of out-of-the-box modules. Run gst-inspect-1.0 qtimlpostprocess on the target device to see the full list of supported modules and their supported tensor shapes. The following example shows a portion of the output.
If no built-in module fits your model, implement your own. A module can be built independently of QIM SDK — you only need the interface header files and a toolchain. Build the module as a shared library, copy it to /usr/lib/imsdk/qtimlpostprocess/modules/ on the device, and the plugin detects it automatically so it can be selected in the pipeline.

Module and library naming

To avoid name collisions, module shared libraries must follow the libml-postprocess-<module-name>.so naming convention. The <module-name> must match the value passed to the qtimlpostprocess module property. For example, a YOLOv8 module is named libml-postprocess-yolov8.so and selected with module=yolov8.

Implement the post-processing module interface

Post-processing modules expose a C++ API. Because C++ APIs cannot be loaded directly from shared libraries, class instantiation is encapsulated in a C function that the header file already implements — you do not need to handle instantiation yourself. Implement the following methods in your module class, which derives from the IModule interface.

Define module capabilities with Caps()

Caps() returns the module type and supported tensor shapes as a JSON string. Tensor dimensions can be fixed or defined within a range using square brackets. For example, [1, [21, 42840], 4] indicates the second dimension can vary between 21 and 42840. The following example declares object-detection post-processing, FLOAT32 tensor format, and support for one, two, or three tensor outputs.
You can specify more than one tensor format at a time, for example "format": ["FLOAT32", "INT8"].

Configure module settings with Configure()

Parse tensors with Process()

Tensor output is a special case where the plugin and module generate tensors instead of predictions. Use it only when two ML models are chained and the output tensor of the first model must be modified before the next model consumes it. If no modification is required, link the two inference plugins directly and omit the post-processing stage.

Understanding post-processing module input

Post-processing module input is split into two fields. tensors holds the inference output tensors and describes their structure. Each output tensor is one entry in the vector. For example, YOLOv8 produces three output tensors (boxes, scores, class indices).
  • Type: float, uint8, and so on.
  • Name: Tensor name, used for identification when two or more output tensors have the same shape. Names are unique and guarantee the exact tensor is selected.
  • Dimensions: The tensor shape. For example, YOLOv8 with three output tensors: [1,8400,4], [1,8400], [1,8400].
  • Data: Pointer to the tensor.
mlparams provides additional parameters for tensor processing that may not apply to all modules. It describes how the pipeline processes the input stream, which helps when the stream resolution and aspect ratio do not match the input tensor shape. It is a dictionary implemented with std::any, so you must know the expected key and its return type.
Supported keys
  • Key: "input-tensor-region"
    Type: video::Region
    Description: Indicates which portion of the input tensor is filled with actual data from the stream. The remaining area is considered padding.
  • Key: "input-tensor-dimensions"
    Type: video::Resolution
    Description: Specifies the size of the input tensor. Required to convert absolute coordinates to relative coordinates when the algorithm produces absolute coordinates, since modules must output relative coordinates.

Generating post-processing module output

The output is an array of arrays of results. Arrays are nested to support batching; only the inner array is filled when there is no batching, and its size matches the number of results found. Results are always in relative dimensions, and the result type depends on the module type.
  • Image/audio classification
    • Name: Predicted category or class.
    • Confidence: Class probability or confidence score.
    • Color: RGBA8888 color for visualization in the overlay plugin.
    • Xtraparams: (optional) Extra key/value pairs to export arbitrary results downstream.
  • Object detection
    • Left, top, right, bottom: Bounding box coordinates.
    • Name: Predicted category or class.
    • Landmarks: (optional) List of keypoints; for example, face detection models can output face points with bounding boxes.
    • Confidence: Class probability or confidence score.
    • Color: RGBA8888 color for visualization in the overlay plugin.
    • Xtraparams: (optional) Extra key/value pairs to export arbitrary results downstream.
  • Pose estimation
    • Name: Predicted category or class.
    • Confidence: Class probability or confidence score.
    • Keypoints: Vector of keypoints.
    • Links: (optional) Vector of links between keypoints.
    • Color: RGBA8888 color for visualization in the overlay plugin.
    • Xtraparams: (optional) Extra key/value pairs to export arbitrary results downstream.
  • Image segmentation and super resolution
    • Output is an image frame or mask.
  • Tensor
    • List of tensors.

Batching

The plugin automatically splits tensor batches into single tensors, so you do not need to handle batching in the module. For example, if the batch size is four, the module is automatically called four times per batch.

Module helper tools

Label and JSON parsers are included in the interface header files. You do not have to use them, but they are provided for convenience. You can use any parser, but the module must be statically linked with it.
  • Label parser: Takes the path to a labels file and automatically detects the format.
    • New-line-separated format: the line number is the class ID.
    • JSON format: set the class index, label, and visualization color. This format is more flexible because you can pass a subset of classes and the rest are automatically filtered out.
  • JSON parser: Parses settings passed as a JSON string. It is also used by the Qualcomm-provided label parser for JSON-formatted labels.

Logging

The module can output logs to the GStreamer log system without a direct dependency on GStreamer. The constructor passes a logging object to the module, which you use with the LOG macro. Supported log levels are Error, Warning, Info, Debug, Trace, and Log.

Compile the post-processing module on a host computer

Prerequisites
  • Ubuntu 22.04 or Ubuntu 24.04 host computer.
1

Install the required tools

3

Put the IM SDK headers and module source files in one folder

4

Create a CMakeLists.txt file

For example:
Post-processing module shared libraries must follow the libml-postprocess-<module-name>.so naming convention. For example, the YOLOv8 module library should be named libml-postprocess-yolov8.so.
5

Create a toolchain file

For example, aarch64-toolchain.cmake:
6

Configure and build the module

Deploy and test the post-processing module

1

Set the user environment variable on the host computer

2

Deploy the module to the target device

1

Transfer the module to the target device

From a terminal on the host computer:
2

SSH into the target device

From a terminal on the host computer:
3

Enter the password when prompted

When prompted, enter the password: oelinux123.
4

Remount / with write permissions

5

Copy the module to the GStreamer plugins directory

6

Run GST inspect on the target device

3

Download the models, labels, and media to run the GStreamer pipeline

1

Create the artifacts dir

2

Download the label file

Download yolox.json, then copy it to the target device.
3

Download media file

Download video1.mp4, then copy it to the target device.
4

Download the model file

Download yolox_quantized.tflite, then copy it to the target device.
4

Build and run a GStreamer pipeline

Build and run a GStreamer pipeline that selects your module with the module property of qtimlpostprocess, passing a label file and settings if required.In the following example pipeline to run a YOLO-X model:
  • The pipeline uses an offline video as the source.
  • The video is decoded to YUV format using the v4l2h264dec decoder.
  • The qtimlvconverter plugin preprocesses the YUV frames.
  • The qtimltflite plugin runs inference with the LiteRT YOLO-X model.
  • The post-processing plugin loads the YOLO-X module and passes a JSON label file.
  • The pipeline displays the results on Wayland.
For Dragonwing™ IQ-615 and Dragonwing™ IQ-2390, change backend_type to dsp as hexagon v66 architecture doesn’t support htp backend_type.

Troubleshooting

Check the following, in order:
  • The library filename follows the libml-postprocess-<module-name>.so convention exactly.
  • The library was copied to /usr/lib/imsdk/qtimlpostprocess/modules/ on the device.
  • The library was built for aarch64, using the toolchain file rather than the host compiler.
  • The module exports the C-compatible entry point, so the shared library can be loaded at runtime.
The Caps() declaration must match the tensors the model actually produces. Compare the module type, tensor formats, and dimensions returned by Caps() with the output tensors reported for your model, and widen a dimension to a range only where it genuinely varies.Also confirm the module property value matches the <module-name> in the library filename.
Post-processing modules must output relative coordinates. If your algorithm produces absolute coordinates, convert them using input-tensor-dimensions, and account for padding using input-tensor-region — the stream is often letter-boxed into the input tensor when the aspect ratios differ.
Verify the stages before post-processing first: confirm inference is producing output tensors, then check that the settings threshold is not filtering everything out, and that results is not set lower than intended. Use the LOG macro inside Process() to confirm the module is being called and to inspect the values it computes.
To see module logs and pipeline warnings while testing, enable GStreamer debug output on the device before running the pipeline:
For more options, see Troubleshooting.

Next steps

Your model is now onboarded: the module parses its output, and the pipeline produces usable metadata. This is the same state reached by a model whose post-processing was already supported, so both routes continue to the same final step.

Step 4 · Build the application with custom AI model integration

Choose a development path and follow the recommended sequence from pipeline validation to a production application.