How to debug the sample application using debug logs?
GST_DEBUG environment variable to set the debug level.The GST_DEBUG environment variable controls the verbosity of the debug output. You can set it to different levels, such as:- 0: None (no debug information)
- 1: ERROR (logs all fatal errors)
- 2: WARNING (logs all warnings)
- 3: FIXME (logs incomplete code paths)
- 4: INFO (logs informational messages)
- 5: DEBUG (logs general debug messages)
- 6: LOG (logs all log messages)
- 7: TRACE (logs trace messages)
- 9: MEMDUMP (logs memory dumps)
GST_DEBUG variable.
For example, to enable debug logs for ML inference plugin and FPS, you can use:What common issues prevent a quick out-of-the-box experience for AI sample applications?
-
Failure to load model file
This issue arises when the model file is either missing or not in the correct format. Copy the model file to the correct path and ensure you are using the same SDK version for model conversion/quantization and on the target device.
-
Failure to deserialize labels
This occurs when the label file specified at the path isn’t present. Ensure that the label file is copied to the device and the specified path is correct.
-
Failure to set module options
This occurs when setting constants for LiteRT/Qualcomm AI Engine direct use cases. See Discover SDKs → IM SDKs for steps to read and set model constants correctly.
How to measure AI sample app profiling?
-
Preprocessing time
Before feeding input to the AI inferencing plugin, the input must be preprocessed (including normalization, rescaling, and color inversion).
This task is managed by the preprocessing plugin,
qtimlvconverter. You can measure the preprocessing time by executing the following command:The preprocessing time is displayed in the logs. -
Model inference time
The Qualcomm Intelligent Multimedia SDK supports three AI inferencing plugins that utilize the Qualcomm Neural Processing SDK, LiteRT, and Qualcomm AI Engine Direct frameworks, respectively.
- qtimlsnpe
- qtimltflite
- qtimlqnn
-
Postprocessing time
The outputs from AI inferencing plugins are handled by postprocessing plugins.
These plugins take the AI model’s results and produce elements that can be overlaid on the input stream or used for further computation.
For example, text boxes for classification use cases, segmentation masks for segmentation use cases, etc.
To measure the processing time for these plugins, use the following command.
The postprocessing time is shown in the logs. The following uses the
qtimlvclassificationplugins as an example.
How to measure end-to-end FPS of a use case?
fpsdisplaysink element.
This element can display the current and average framerate either as an overlay on the video or by printing it to the console.Sample apps use the fpsdisplaysink plugin to display the FPS of the pipeline directly on the HDMI monitor.What's the easiest way to replace an existing model with a custom model in a reference application?
User replaced another supported model in a reference application
-
Performance measurement:
Use the SDK benchmarking tool. For example, If you are using Qualcomm Neural Processing Engine SDK, use snpe-bench-py -
Accuracy debugger:
- AI Hub model:
- Ensure you are using the latest model from AI Hub.
- Correctly populate the constants for your selected model. See Discover SDKs → IM SDKs for more information.
- For further support on AI Hub model accuracy issues, report your issue on Qualcomm AI Hub slack.
- Ensure you are using the latest model from AI Hub.
- Custom model
- Model quantization is a common cause of accuracy drops. Ensure you are using the correct dataset for model quantization. Users are expected to use a portion of dataset which is an approximation of actual deployment environment for Post Training Quantization (PTQ). For PTQ to give good results, users need to feed decent amount of data to quantize the model, for example, approximately 25-30 RAW images.
- Experiment with different model precisions, such as W8A16 and W16A16, to see if there is an improvement in model accuracy.
- Use AIMET for Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) techniques for advanced quantization. See the AIMET documentation for more details.
- Model quantization is a common cause of accuracy drops. Ensure you are using the correct dataset for model quantization. Users are expected to use a portion of dataset which is an approximation of actual deployment environment for Post Training Quantization (PTQ). For PTQ to give good results, users need to feed decent amount of data to quantize the model, for example, approximately 25-30 RAW images.
- AI Hub model:
What are the best practices for model quantization using AI SDK?
- Prepare Calibration Data Use representative calibration data that closely matches the data the model will encounter in production. This helps accurately determine the scaling factors and zero points for quantization
-
Choose the Right Quantization Method
- Post-Training Quantization (PTQ): This method is simpler and faster. Suitable for models where slight accuracy loss is acceptable. It converts a pretrained floating-point model to a quantized model without retraining.
- Quantization-Aware Training (QAT): This method involves training the model with quantization in mind, which can help maintain higher accuracy. It’s more complex but beneficial for models where accuracy is critical.
Do LiteRT models need conversion to DLC for NPU acceleration?
What is the quickest deployment path for PyTorch, ONNX, and TensorFlow models?
How does a user know if the model is running on the NPU?
- Qualcomm profiler
-
Sysmon
If you have access to the Hexagon SDK, refer to the sysmon documentation at
<Hexagon_sdk_path>/<version>/tools/sysmon_app.html
How can users run models on different hardware cores?
gst-ai-object-detection sample app modify the runtime parameter in the config_detection.json file.Runtime:"cpu""gpu""dsp"
How to proceed if model conversion fails with AI SDK?
How to debug poor quantized-model accuracy on CPU and NPU?
How to debug a quantized model that is inaccurate on the NPU?
qnn-net-run --debug option dumps layerwise values, so you can compare between CPU and NPU inference.How to debug missing HDMI output from a sample app?
Can an HDMI TV run sample applications?
How to check the hardware runtime for AI sample apps?
- CPU
- GPU
- DSP
What are devtool sanity check errors and how to debug them?
Update permissions
Disable BitBake sanity checking
$ESDK_ROOT/layers/poky/meta/conf/sanity.conf file.How to sideload the Qualcomm Neural Processing Engine SDK on the target device
Prepare the target device
Copy SDK files to the target device
- For QCS8275: Replace
hexagon-v68withhexagon-v75 - For QCS9075: Replace
hexagon-v68withhexagon-v73 - For IQ-X5121: Replace
hexagon-v68withhexagon-v73 - For IQ-X7181: Replace
hexagon-v68withhexagon-v73
Validate the SDK version
How to run inference on a preprocessed QIM SDK tensor using SNPE
Generate preprocessed raw tensors
Create a folder for raw tensors
Dump raw tensors using GStreamer
Run inference with SNPE or QNN
Create the input list
input_list.txt file containing the absolute paths to the raw files, as shown below.Inference is performed on each file listed in the input_list.txt file.Example content in input_list.txtRun the model on the HTP backend
- SNPE
- QNN

