跳到主要内容

MTNN 开发者指南

适用版本: MTNN v1.6.0 | 目标平台: 摩尔线程 M1000

MTNN(Moore Threads Neural Network)是一套面向 M1000 设备的模型编译工具与推理运行时。本文从 MTNN 基础概念、模型编译、推理部署、模型仓库和 ONNX Runtime EP 五个方面介绍完整的开发与部署流程。

目录​


1. MTNN 介绍​

MTNN 可帮助开发者将训练好的深度学习模型部署到 M1000 设备,实现设备端机器学习推理。

1.1 主要特性​

  • 平台支持: 支持 M1000 Linux 平台。
  • 开发接口: 提供 C/C++ API,并支持 Python 调用。
  • 轻量高效: 通过模型压缩和图优化降低内存占用,提升推理效率。

1.2 开发流程​

MTNN 的开发工作流主要分为两个阶段:

  1. 模型编译

    使用 mtc 工具将 ONNX 模型转换为 .mtnn 模型。转换过程中可应用量化等优化,缩减模型大小和推理延迟,同时尽可能降低精度损失。

  2. 推理部署

    应用程序通过 MTNN Runtime API 加载并运行 .mtnn 模型,基本调用流程如下:

    mtnn_init → mtnn_inputs_get → mtnn_inputs_set → mtnn_inference → mtnn_outputs_get

此外我们为开发者提供了 模型仓库,汇总了目前在 M1000 NPU 上已支持的模型,涵盖图像分类、目标检测、图像分割及其他常见任务。帮助开发者快速上手和部署自己的AI模型。

1.3 MTNN 工作流程​

1.4 支持的算子​

MTNN 通过 ONNX 标准支持以下常用算子:

类别支持的算子
基础数学Add, Sub, Div, Mul, Abs, Exp, Log, Mod, Neg, Pow, Reciprocal, Sign, Sin, Sqrt, Sum, Mean, Erf, Ceil, Floor
比较与逻辑Greater, GreaterOrEqual, Less, LessOrEqual, Equal, logical_and, Xor, Where, NonZero
激活函数Clip, Celu, Sigmoid, HardSigmoid, HardSwish, LeakyRelu, Softplus, Elu, PRelu, Relu, Softsign, Tanh, Softmax, Logsoftmax
归一化BatchNormalization, InstanceNormalization, LRN, MeanVarianceNormalization
卷积与矩阵运算Conv2D, Conv2DTranspose, Gemm, MatMul, Einsum
池化与采样MaxPool, AveragePool, GlobalAveragePool, GlobalMaxPool, MaxRoiPool, Resize, Upsample
张量与形状操作Reshape, Flatten, Expand, ReverseSequence, Size, Squeeze, Unsqueeze, Shape, Concat, Split, Transpose, Slice, Tile, DepthToSpace, SpaceToDepth, Pad
索引操作ArgMin, ArgMax, Gather, GatherND, ScatterND, Cumsum
循环神经网络LSTM, GRU
规约操作ReduceSum, ReduceMean, ReduceMax, ReduceMin, ReduceL1, ReduceL2, ReduceLogSum, ReduceLogSumExp, ReduceProd, ReduceSumSquare
其他Cast, CastLike, Dropout

2. 模型编译​

本章介绍如何使用 mtc 工具完成模型转换。

2.1 mtc 工具​

2.1.1 工具介绍​

mtc 提供模型转换流程所需的 Convert、Quant 和 Map 功能,可将 AI 模型转换为能够在 NPU 硬件上运行的 .mtnn 模型。

2.1.2 支持的深度学习框架​

mtc 目前仅支持 ONNX 网络,其他深度学习框架的模型需要先转换为 ONNX 模型。关于如何将 PyTorch、TensorFlow 框架的模型转为ONNX模型, 请参考各模型的官网。此外,可以使用PyTorch转ONNX的通用方法。

2.1.3 支持的运行模式​

mtc 支持以下两种模型运行模式:

  • 生成 .mtnnir 中间模型,并进行软件仿真推理。可以使用模型可视化工具查看
  • 生成 .mtnn 模型,并进行软件仿真推理;.mtnn 是最终在 NPU 硬件上使用的模型格式。

2.1.4 使用教程​

下载并安装 m1000-mtc-toolkit DEB 包。安装成功后,即可运行 mtc 命令。

2.1.4.1 命令格式​
mtc [-h]
--config "<NETWORK_CFG_YAML_FILE>"
2.1.4.2 命令行参数​
参数说明
-h显示帮助信息。
--config "<NETWORK_CFG_YAML_FILE>"指定网络配置文件(.yaml)的路径。
2.1.4.3 YAML 配置文件​

以下是以 Inception-v2 模型为例的 YAML 文件结构:

MTC_TOOL YAML 配置文件示例:

参数说明
layer_data_analyze按 Layer Name 逐层调试,暂不支持。
layer_data_compare按 Layer Name 逐层对比,暂不支持。
mtnn_test_enable是否使用 .mtnn 模型进行软件仿真推理。
ModelPath(必填)ONNX 模型的存放路径。
Name(必填)模型名称。
force_fp32input_list设置各输入是否为 Float32,多个输入使用 # 分隔。例如,true#false#true 表示第 1、3 个输入为 Float32,第 2 个输入为对应的量化类型。
force_fp32output_list设置各输出是否为 Float32,多个输出使用 # 分隔。例如,true#false#true 表示第 1、3 个输出为 Float32,第 2 个输出为对应的量化类型。
perf_collect_enable是否汇总并输出性能结果。
algorithm(必填)量化算法。当前可选值为 normal,也是默认值。
quant_type(必填)量化数据类型,可选值为 int8、int16 和 uint8,推荐使用 uint8。
quantizer(必填)量化方式,可选值为 asymmetric_affine(推荐)、dynamic_fixed_point 和 perchannel_symmetric_affine。
quantize_op_list仅量化列表中指定类型的算子,其他类型的算子不量化。
exclude_op_list不量化列表中指定类型的算子,其他类型的算子正常量化。
hybrid是否启用混合量化,可选值为 true 或 false。
quantize_file混合量化配置文件。
static_quantization输入模型是否为静态量化 ONNX 模型,可选值为 true 或 false,默认为 false。
skip_preprocess_input量化输入数据时是否跳过前处理,可选值为 true 或 false,默认为 false。
TestEnable暂未使用。
UseSingleCore是否编译为单核模型,默认值为 1。
2.1.4.4 转换示例​
mtc --config ./example/onnx/mobilenet_v2/mobilenetv2.yaml

2.2 模型可视化​

mtc 生成的 .mtnnir 中间 IR 模型,可通过定制版 Netron 打开并查看模型结构。

平台说明下载地址
Linux ARM64可直接安装到 AIBook 和 AIModule 的 Ubuntu 系统下载 .deb 安装包
Windows x64Windows 桌面安装程序下载 .exe 安装包

2.3 混合量化​

混合量化用于重新量化具有更高精度数据类型的量化网络的特定层。

2.3.1 自动混合量化​

自动混合量化的配置与全量化基本一致,但需要额外将 rebuild、compute_entropy 和 hybrid 设置为 true,然后运行 mtc。

程序会分析网络中各张量(包括激活值和权重)的分布特性,计算其熵值(Entropy),并据此标记各层的量化类型。生成的 entropy.txt 文件记录了每个张量的熵值,具体规则如下:

2.3.2 手动混合量化​

如果自动混合量化后的精度仍不满足要求,可根据 entropy.txt 找出熵值较高的层,并在 .quantize 文件的 customized_quantize_layers 中指定这些层及其高精度数据类型(如 dynamic_fixed_point-i16)。然后在 YAML 文件中通过 quantize_file 引用修改后的 .quantize 文件,再运行 mtc。

以 Inception-v2 为例,配置步骤如下:

  1. 在 --config 指定的 YAML 文件中设置混合量化参数。

    参数说明
    rebuild是否使用默认量化规则执行量化,可选值为 true 或 false。
    compute_entropy是否通过计算熵来衡量当前量化精度,可选值为 true 或 false。
    hybrid是否启用混合量化,可选值为 true 或 false。
    quantize_file(可选)指定混合量化配置文件。
  2. 提供 YAML 文件中引用的混合量化配置文件。以下为 inception-v2-9.quantize 示例:

2.4 前处理​

MTNN 可在模型转换阶段,将图像格式转换、尺寸调整、通道重排和数值归一化等前处理操作写入模型。推理时,应用只需传入原始输入 Buffer,即可自动完成前处理和推理,有助于简化部署并保证处理流程一致。

2.4.1 配置模型前处理​

前处理通过 PreprocConfig 配置。多输入模型可分别使用 PreprocConfig0、PreprocConfig1、PreprocConfig2 等配置各输入,其中 PreprocConfig 等同于 PreprocConfig0。

配置项说明
reverse_channel: false是否反转输入 Tensor 的通道顺序,通常用于 BGR 与 RGB 互换。false 表示不调整,true 表示反转通道顺序。
mean: [0.0, 0.0, 0.0]各通道的减均值参数。三个值均为 0.0 表示不执行减均值。
scale: 1.0减均值后使用的统一缩放系数。1.0 表示不缩放;0.0039216(即 1/255)可将 0~255 的像素值归一化到 0~1。
preproc_node_params预处理节点的详细配置,具体子项见下表。

preproc_node_params 支持以下子项:

配置项说明
add_preproc_node是否在模型计算图中插入预处理节点。外部已完成前处理时通常设置为 false。
preproc_type预处理节点类型,支持 IMAGE_RGB、IMAGE_RGB888_PLANAR、IMAGE_RGB888_PLANAR_SEP、IMAGE_I420、IMAGE_NV12、IMAGE_NV21、IMAGE_YUV444、IMAGE_YUYV422、IMAGE_UYVY422、IMAGE_GRAY、IMAGE_BGRA 和 TENSOR。需要手动指定类型名称。
preproc_image_size: [1920, 1080]将要被重采样或缩放的原图尺寸,格式为 [width, height]。
preproc_cropCrop 配置。enable_preproc_crop 控制是否启用裁剪;启用后,使用 crop_rect: [x, y, width, height] 指定裁剪区域。以背景的左上角为坐标原点:x, y为裁剪区域的左上角坐标,width, height为宽度和高度。
preproc_perm: [0, 1, 2, 3]Tensor 维度置换(Permutation)。如果需要将 NHWC 转换为 NCHW,应设置为 [0, 3, 1, 2]。

2.4.2 运行时动态更新 Crop 参数​

动态 Crop 适用于同一模型在不同帧或不同输入区域上复用推理的场景,例如从原始大图中按不同 ROI 裁剪后送入同一网络。

  • 模型转换时需启用预处理节点和 Crop,即设置 add_preproc_node: true、enable_preproc_crop: true,并配置初始 crop_rect 或 crop_size。
  • C API 可调用 mtnn_update_crop_params(mgr, enabled_crop_input_idx, start_x, start_y, crop_w, crop_h, dst_w, dst_h)。建议在输入数据复制完成后、调用 mtnn_inputs_set 或 mtnn_inference 前更新参数。
  • Python API 可通过 MTNNSession.run 的 crop_size 和 crop_dst_size 参数更新 Crop。

单输入模型示例:

from mtnn_api import MTNNSession

session = MTNNSession(model_path)
outputs = session.run(
{0: input_data},
crop_size=[start_x, start_y, crop_w, crop_h],
crop_dst_size=[dst_w, dst_h],
)

多输入模型可按输入 Index 分别设置:

outputs = session.run(
{0: input0, 1: input1},
crop_size={
0: [x0, y0, w0, h0],
1: [x1, y1, w1, h1],
},
crop_dst_size={
0: [dst_w0, dst_h0],
1: [dst_w1, dst_h1],
},
)

crop_size 的格式为 [start_x, start_y, crop_w, crop_h],各值需为非负数;crop_dst_size 的格式为 [dst_w, dst_h],各值需为正数。如果未显式传入 crop_dst_size,Python API 将使用对应模型输入 Tensor 属性中的宽和高作为目标尺寸。

2.5 配置数据集​

在量化或校准过程中,input_dataset_list 提供待送入模型的样本,用于:

  • 收集各层的激活分布,以计算量化参数。
  • 验证模型量化前后的输出精度和性能。

以下配置可直接写入 YAML 文件:

2.5.1 ModelConfig.input_dataset_list​

  • 类型: 字符串列表(list[str])。
  • 含义: 指定一个或多个数据列表文件;每个文件按行列出待处理样本,例如图片路径或 NPY 输入文件。在 INT8 校准(Calibration)过程中,这些样本将逐个送入模型,用于统计激活值范围、直方图和 KL 散度等信息。
  • 多输入模型: 每个模型输入对应一个 TXT 文件。例如,模型有 3 个输入时,需要指定 3 个 TXT 文件。

将样本路径逐行写入 datasets.txt:

/data/images/img_0001.jpg
/data/images/img_0002.jpg
...

然后将该文件路径填入 input_dataset_list。

2.5.2 QuantConfig.iterations​

  • 类型: 整数(int)。
  • 含义: 控制量化(Quantization)或校准(Calibration)阶段最多使用的样本迭代次数。框架依次读取 input_dataset_list 中的样本;样本数量不足时,将循环读取直至达到指定次数。
  • 取舍: 较大的 iterations 能收集更全面的激活分布,通常有利于量化精度,但会增加校准时间。少量迭代适合快速验证,正式发布时可使用更多样本。
典型值适用场景
iterations: 1–32快速 Smoke Test
iterations: 64–128平衡校准速度与量化精度
iterations: 256+追求更高的量化质量

2.5.3 QuantConfig.skip_preprocess_input​

  • 类型: 布尔值(bool)。
  • 含义: 控制量化输入数据前是否跳过前处理。设置为 false 时,框架先按既定流程处理输入,再执行量化;设置为 true 时,直接量化输入数据。
  • 推荐用法: 原始图片或原始 Tensor 通常设置为 false,以确保量化数据分布与实际推理流程一致;输入数据已完成归一化、减均值、缩放等操作时,可设置为 true 以避免重复处理。
配置适用场景
skip_preprocess_input: false默认推荐;输入仍需执行标准前处理。
skip_preprocess_input: true输入数据已提前完成前处理,可直接参与量化。

skip_preprocess_input: false 配置示例(114.npy 由原始图片转换得到):

skip_preprocess_input: true 配置示例(000000000009.npy 是已完成前处理的 Tensor):

3. 推理部署​

本章说明如何通过 MTNN Runtime 的 C/C++ API 和 Python API 加载 .mtnn 模型、执行推理和处理输出,并以 MobileNetV1 为例介绍完整的模型部署过程。

3.1 C/C++ API​

mtnn_api.h 在 C++ 环境下使用 extern "C" 导出接口,因此同一组 Runtime API 可供 C 和 C++ 应用调用。mtnn_mgr 是平台相关的模型句柄:定义 __arm__ 时为 uint32_t,其他平台为 uint64_t。应用应将其视为不透明句柄,不依赖具体整数宽度。

API功能
mtnn_init初始化上下文并加载 .mtnn 模型。
mtnn_get查询输入输出数量、Tensor 属性、性能数据和 SDK 版本等信息。
mtnn_set设置 NPU Index、工作频率,以及自定义输入输出 Tensor。
mtnn_inputs_get获取模型输入 Tensor 的内存信息。
mtnn_inputs_set提交输入数据,使输入 Buffer 中的数据生效。
mtnn_update_crop_params可选:在运行时更新预处理 Crop 参数。
mtnn_inference执行模型推理。
mtnn_outputs_get获取模型输出 Tensor 的内存信息。
mtnn_destroy卸载模型并销毁上下文。
mtnn_set_log_level设置 Runtime 日志级别。
mtnn_get_tensor_file_data从文本 Tensor 文件读取 Float32 数据,并转换为模型输入类型。
mtnn_get_tensor_from_array将 Float32 数组转换为指定模型输入的 Tensor 数据。
mtnn_dtype_to_float32根据 Tensor 属性将单个元素转换为 Float32。
mtnn_create_tensor根据 Tensor 属性创建 Runtime Tensor。
mtnn_create_tensor_from_fd基于 DMA-BUF 文件描述符及偏移创建 Runtime Tensor。

3.1.1 推理流程​

本节以 MobileNetV2 Demo 为例,演示通过 MTNN C API 完成模型加载、输入设置、推理执行和结果获取的基本流程:

  1. 将模型文件读入内存。
  2. 通过 mtnn_init 加载并初始化模型,获取 mtnn_mgr 句柄。
  3. 通过 mtnn_inputs_get 获取输入 Buffer 的地址等信息。
  4. 填充输入数据,并通过 mtnn_inputs_set 使输入生效。
  5. 可选:通过 mtnn_update_crop_params 动态更新预处理 Crop 参数。
  6. 通过 mtnn_inference 执行模型推理。
  7. 通过 mtnn_outputs_get 获取输出结果。

3.1.2 C 应用示例​

引用头文件​
#include "mtnn_api.h"
将模型读入内存​
unsigned char *model_data = load_model(mtnn_path, &model_data_size);

load_model 函数示例如下:

static unsigned char *load_model(const char *filename, size_t *model_size)
{
unsigned char *data = NULL;
size_t size = 0;
FILE *fp = fopen(filename, "rb");
if (NULL == fp) {
printf("Open file %s failed: %s.\n", filename, strerror(errno));
return NULL;
}

fseek(fp, 0, SEEK_END);
size = ftell(fp);
fclose(fp);
data = (unsigned char *)malloc(size);
load_file(filename, data);

*model_size = size;

return data;
}
初始化模型​

初始化模型时,可通过 mtnn_work_mode_t 配置 MTNN 的扩展功能。

// init
mtnn_work_mode_t work_mode;
memset(&work_mode, 0, sizeof(mtnn_work_mode_t));
work_mode.init_flag = 0;
// work_mode.init_flag = MTNN_FLAG_COLLECT_PERF_MASK | MTNN_FLAG_DUMP_LAYER_DATA_MASK;
work_mode.e_npu_mode = NPU_MODE_SEPARATE;
work_mode.n_npu_device = 0;
status = mtnn_init(&network_mgr, (void*)model_data, model_data_size, &work_mode);
ONERROR(status, "mtnn_init failed.");
获取并设置模型输入​
status = mtnn_get(network_mgr, MTNN_GET_IN_OUT_NUM, &io_num, sizeof(mtnn_input_output_num));
ONERROR(status, "mtnn_get MTNN_GET_IN_OUT_NUM failed.");

inputs_mem = (mtnn_tensor_mem*)malloc(io_num.n_input * sizeof(mtnn_tensor_mem));
outputs_mem = (mtnn_tensor_mem*)malloc(io_num.n_output * sizeof(mtnn_tensor_mem));
status = mtnn_inputs_get(network_mgr, io_num.n_input, inputs_mem);
ONERROR(status, "mtnn_inputs_get failed.");

// step2. Set input
for (i = 0; i < io_num.n_input; i++) {
if (load_file(input_name[i], inputs_mem[i].logical_addr) != inputs_mem[i].size) {
printf("error: input size mismatch for %s, expected %u\n",
input_name[i], inputs_mem[i].size);
goto error_exit;
}
}

/* 可选:运行时更新预处理 crop 参数,需模型已启用 add_preproc_node 和 enable_preproc_crop */
status = mtnn_update_crop_params(network_mgr, 0, 0, 0, 320, 240, 640, 640);
ONERROR(status, "mtnn_update_crop_params failed.");

status = mtnn_inputs_set(network_mgr, io_num.n_input, NULL);
ONERROR(status, "mtnn_inputs_set failed.");

上述方式通过 mtnn_inputs_get 获得 Runtime 内部输入 Buffer,并直接向 logical_addr 写入数据。此时可向 mtnn_inputs_set 的 inputs 参数传入 NULL,用于提交或刷新内部输入 Buffer。若应用使用自有输入 Buffer,则应构造 mtnn_input[] 并传入该接口。

执行推理​
status = mtnn_inference(network_mgr, NULL);
ONERROR(status, "mtnn_inference failed.");
获取推理结果​
status = mtnn_outputs_get(network_mgr, io_num.n_output, outputs_mem);
ONERROR(status, "mtnn_outputs_get failed.");

3.1.3 API 参考​

mtnn_init​
/* mtnn_init

initial the context and load the mtnn model.

input:
mtnn_mgr* mgr the pointer of network handle.
void* model pointer to the mtnn model.
uint32_t size the size of mtnn model.
mtnn_work_mode_t* work_mode extend info.
return:
int error code.
*/
int mtnn_init(mtnn_mgr* mgr, void* model, uint32_t size, mtnn_work_mode_t* work_mode);
mtnn_get​
/* mtnn_get

get model or runtime information. see mtnn_get_cmd.

input:
mtnn_mgr mgr the handle of network.
mtnn_get_cmd cmd the command of query.
void* info the buffer point of information.
uint32_t size the size of information.
return:
int error code.
*/
int mtnn_get(mtnn_mgr mgr, mtnn_get_cmd cmd, void* info, uint32_t size);
/*
The query command for mtnn_query
*/
typedef enum _mtnn_query_cmd {
MTNN_GET_IN_OUT_NUM = 0, /* get the number of input&output tensor. */
MTNN_GET_INPUT_ATTR, /* get the attribute of input tensor. */
MTNN_GET_OUTPUT_ATTR, /* get the attribute of output tensor. */
MTNN_GET_PERF_DETAIL, /* get the detail performance, need set
MTNN_FLAG_COLLECT_PERF_MASK when call mtnn_init. */
MTNN_GET_PERF_RUN, /* get the time of run. */
MTNN_GET_SDK_VERSION, /* get the sdk & driver version */
MTNN_GET_PRE_COMPILE, /* get the pre compile model */
MTNN_GET_CMD_MAX
} mtnn_get_cmd;
mtnn_set​
/* mtnn_set

set model or runtime information. see mtnn_set_cmd.

input:
mtnn_mgr mgr the handle of network.
mtnn_set_cmd cmd the command of setting.
void* info the buffer point of information.
uint32_t size the size of information.
return:
int error code.
*/
int mtnn_set(mtnn_mgr mgr, mtnn_set_cmd cmd, void* info, uint32_t size);
typedef enum _mtnn_set_cmd {
MTNN_SET_NPU_INDEX = 0, /* set npu index. */
MTNN_SET_NPU_FREQUENCY, /* set npu frequency. */
MTNN_SET_INTPUT_TENSOR, /* set input tensor; keep the header spelling. */
MTNN_SET_OUTPUT_TENSOR, /* set output tensor. */
MTNN_SET_CMD_MAX
} mtnn_set_cmd;

注意: MTNN_SET_INTPUT_TENSOR 是 mtnn_api.h 中实际导出的标识符,INTPUT 的拼写虽然不常见,但调用时必须原样使用。

mtnn_destroy​
/* mtnn_destroy

unload the mtnn model and destroy the context.

input:
mtnn_mgr mgr the handle of network.
return:
int error code.
*/
int mtnn_destroy(mtnn_mgr mgr);
mtnn_inputs_set​
/* mtnn_inputs_set

set inputs information by input index of mtnn model.
inputs information see mtnn_input.

input:
mtnn_mgr mgr the handle of network.
uint32_t n_inputs the number of inputs.
mtnn_input inputs[] the arrays of inputs information, see mtnn_input.
return:
int error code
*/
int mtnn_inputs_set(mtnn_mgr mgr, uint32_t n_inputs, mtnn_input inputs[]);
mtnn_inputs_get​
/* mtnn_inputs_get

get inputs memory information by input index of mtnn model.
inputs information see mtnn_input.

input:
mtnn_mgr mgr the handle of network.
uint32_t n_inputs the number of inputs.
mtnn_tensor_mem mem[] the array of tensor memory information
return:
int error code
*/
int mtnn_inputs_get(mtnn_mgr mgr, uint32_t n_inputs, mtnn_tensor_mem mem[]);
mtnn_inference​
/* mtnn_inference

run the model to execute inference.

input:
mtnn_mgr mgr the handle of network.
mtnn_run_extend* extend the extend information of run.
return:
int error code.
*/
int mtnn_inference(mtnn_mgr mgr, mtnn_run_extend* extend);
mtnn_outputs_get​
/* mtnn_outputs_get

get the model output tensors memory information.
It directly maps the model output tensor memory location to the user.

input:
mtnn_mgr mgr the handle of network.
uint32_t n_outputs the number of outputs.
mtnn_tensor_mem mem[] the array of tensor memory information
return:
int error code.
*/
int mtnn_outputs_get(mtnn_mgr mgr, uint32_t n_outputs, mtnn_tensor_mem mem[]);
mtnn_update_crop_params​
/* mtnn_update_crop_params

update preprocess crop parameters during runtime.
This API is valid for inputs whose preprocessing node and crop were
enabled when the model was generated.

input:
mtnn_mgr mgr the handle of network.
uint32_t enabled_crop_input_idx the input index with crop enabled.
uint32_t start_x the crop start coordinate x.
uint32_t start_y the crop start coordinate y.
uint32_t crop_w the crop width.
uint32_t crop_h the crop height.
uint32_t dst_w the destination width after crop resize.
uint32_t dst_h the destination height after crop resize.
return:
int error code.
*/
int mtnn_update_crop_params(
mtnn_mgr mgr,
uint32_t enabled_crop_input_idx,
uint32_t start_x,
uint32_t start_y,
uint32_t crop_w,
uint32_t crop_h,
uint32_t dst_w,
uint32_t dst_h);

建议在复制输入数据后、调用 mtnn_inputs_set 或 mtnn_inference 前更新 Crop 参数。如果每帧的裁剪区域不同,可在每次推理前调用该接口;如果不需要动态裁剪,则无需调用,模型将使用转换阶段配置的初始 Crop 参数。

日志与 Tensor 辅助接口​
API参数与返回值说明
mtnn_set_log_level传入 mtnn_log_level_e 日志级别;无返回值。
mtnn_get_tensor_file_data根据输入 Index 读取文本文件中的 Float32 数据,并转换为对应输入 Tensor 的数据类型;成功时返回数据指针,失败时返回 NULL。
mtnn_get_tensor_from_array根据输入 Index 将 Float32 数组转换为对应输入 Tensor 的数据类型;成功时返回数据指针,失败时返回 NULL。
mtnn_dtype_to_float32根据 mtnn_tensor_attr 中的数据类型和量化参数,将 src 指向的单个元素转换到 dst;返回 MTNN 错误码。
mtnn_create_tensor根据 mtnn_tensor_attr 创建 Tensor,并通过 mtnn_tensor_mem 返回内存地址、大小和 Handle 等信息。
mtnn_create_tensor_from_fd使用 Tensor 属性、DMA-BUF 文件描述符及偏移创建 Tensor,并通过 mtnn_tensor_mem 返回内存信息。
void mtnn_set_log_level(mtnn_log_level_e log_level);

uint8_t *mtnn_get_tensor_file_data(
mtnn_mgr mgr,
uint32_t input_index,
const char *name);

uint8_t *mtnn_get_tensor_from_array(
mtnn_mgr mgr,
uint32_t input_index,
float *array);

int mtnn_dtype_to_float32(
uint8_t *src,
float *dst,
mtnn_tensor_attr *src_tensor_attr);

int mtnn_create_tensor(
mtnn_mgr mgr,
mtnn_tensor_attr *attr,
mtnn_tensor_mem *tensor);

int mtnn_create_tensor_from_fd(
mtnn_mgr mgr,
mtnn_tensor_attr *attr,
int fd,
uint32_t offset,
mtnn_tensor_mem *tensor);

上述辅助接口的支持情况取决于 Runtime 后端,使用前应检查返回值。当前 Unify 实现中的 mtnn_get_tensor_file_data 和 mtnn_get_tensor_from_array 会返回由内部申请的数据 Buffer,调用方使用完毕后需要释放。创建自定义 Tensor 后,可通过 mtnn_tensor_info 配合 MTNN_SET_INTPUT_TENSOR 或 MTNN_SET_OUTPUT_TENSOR 将其绑定到指定输入或输出 Index。

主要结构体​
#ifdef __arm__
typedef uint32_t mtnn_mgr;
#else
typedef uint64_t mtnn_mgr;
#endif

#define MTNN_MAX_DIMS 16
#define MTNN_MAX_NAME_LEN 256

typedef enum _mtnn_npu_mode_e
{
NPU_MODE_SEPARATE = 0,
NPU_MODE_COMBINED = 1,
NPU_MODE_BUTT
} mtnn_npu_mode_e;

typedef struct _mtnn_work_mode_t
{
uint32_t init_flag;
uint32_t n_npu_device;
mtnn_npu_mode_e e_npu_mode;
} mtnn_work_mode_t;

typedef enum _mtnn_log_level
{
MTNN_LOG_EMERG,
MTNN_LOG_ALERT,
MTNN_LOG_CRIT,
MTNN_LOG_ERR,
MTNN_LOG_WARNING,
MTNN_LOG_NOTICE,
MTNN_LOG_INFO,
MTNN_LOG_DEBUG
} mtnn_log_level_e;

/*
The query command for mtnn_query
*/
typedef enum _mtnn_query_cmd {
MTNN_GET_IN_OUT_NUM = 0, /* get the number of input & output tensor. */
MTNN_GET_INPUT_ATTR, /* get the attribute of input tensor. */
MTNN_GET_OUTPUT_ATTR, /* get the attribute of output tensor. */
MTNN_GET_PERF_DETAIL, /* get the detail performance, need set
MTNN_FLAG_COLLECT_PERF_MASK when call mtnn_init. */
MTNN_GET_PERF_RUN, /* get the time of run. */
MTNN_GET_SDK_VERSION, /* get the sdk & driver version */
MTNN_GET_PRE_COMPILE, /* get the pre compile model */

MTNN_GET_CMD_MAX
} mtnn_get_cmd;

/*
The set command for mtnn_set
*/
typedef enum _mtnn_set_cmd {
MTNN_SET_NPU_INDEX = 0, /* set npu index. */
MTNN_SET_NPU_FREQUENCY, /* set npu frequency. */
MTNN_SET_INTPUT_TENSOR, /* set input tensor; keep the header spelling. */
MTNN_SET_OUTPUT_TENSOR, /* set output tensor. */
MTNN_SET_CMD_MAX
} mtnn_set_cmd;

/*
the tensor data type.
*/
typedef enum _mtnn_tensor_type {
MTNN_TENSOR_FLOAT32 = 0, /* data type is float32. */
MTNN_TENSOR_FLOAT16, /* data type is float16. */
MTNN_TENSOR_INT8, /* data type is int8. */
MTNN_TENSOR_UINT8, /* data type is uint8. */
MTNN_TENSOR_INT16, /* data type is int16. */
MTNN_TENSOR_INT32, /* data type is int32. */
MTNN_TENSOR_INT64, /* data type is int64. */
MTNN_TENSOR_BOOL8, /* data type is bool 8bit. */
MTNN_TENSOR_TYPE_MAX
} mtnn_tensor_type;

/*
the quantitative type.
*/
typedef enum _mtnn_tensor_qnt_type {
MTNN_TENSOR_QNT_NONE = 0, /* none. */
MTNN_TENSOR_QNT_DFP, /* dynamic fixed point. */
MTNN_TENSOR_QNT_AFFINE_ASYMMETRIC, /* asymmetric affine. */

MTNN_TENSOR_QNT_MAX
} mtnn_tensor_qnt_type;

/*
the tensor data format.
*/
typedef enum _mtnn_tensor_format {
MTNN_TENSOR_NCHW = 0, /* data format is NCHW. */
MTNN_TENSOR_NHWC, /* data format is NHWC. */

MTNN_TENSOR_FORMAT_MAX
} mtnn_tensor_format;

typedef enum _mtnn_power_frequency_level
{
/* auto frequency dvfs enable */
MTNN_POWER_AUTO_FREQUENCY = 0,
/* fixed frequency level1 */
MTNN_POWER_FIXED_FREQUENCY_LEVEL1 = 1,
/* fixed frequency level2 */
MTNN_POWER_FIXED_FREQUENCY_LEVEL2 = 2,
/* fixed frequency level3 */
MTNN_POWER_FIXED_FREQUENCY_LEVEL3 = 3,
/* fixed frequency level4 */
MTNN_POWER_FIXED_FREQUENCY_LEVEL4 = 4,
} mtnn_power_frequency_level;

/*
the information for MTNN_QUERY_IN_OUT_NUM.
*/
typedef struct _mtnn_input_output_num {
uint32_t n_input; /* the number of input. */
uint32_t n_output; /* the number of output. */
} mtnn_input_output_num;

/*
the information for MTNN_QUERY_INPUT_ATTR / MTNN_QUERY_OUTPUT_ATTR.
*/
typedef struct _mtnn_tensor_attr {
uint32_t index; /* input parameter, the index of input/output tensor,
need set before call mtnn_query. */
uint32_t n_dims; /* the number of dimensions. */
uint32_t dims[MTNN_MAX_DIMS]; /* the dimensions array. */
char name[MTNN_MAX_NAME_LEN]; /* the name of tensor. */

uint32_t n_elems; /* the number of elements. */
uint32_t size; /* the bytes size of tensor. */

mtnn_tensor_format fmt; /* the data format of tensor. */
mtnn_tensor_type type; /* the data type of tensor. */
mtnn_tensor_qnt_type qnt_type; /* the quantitative type of tensor. */
int8_t fl; /* fractional length for MTNN_TENSOR_QNT_DFP. */
uint32_t zp; /* zero point for MTNN_TENSOR_QNT_AFFINE_ASYMMETRIC. */
float scale; /* scale for MTNN_TENSOR_QNT_AFFINE_ASYMMETRIC. */
} mtnn_tensor_attr;

/*
the information for MTNN_QUERY_PERF_DETAIL.
*/
typedef struct _mtnn_perf_detail {
char* perf_data; /* the string pointer of perf detail. don't need free it by user. */
uint64_t data_len; /* the string length. */
} mtnn_perf_detail;

/*
the information for MTNN_QUERY_PERF_RUN.
*/
typedef struct _mtnn_perf_run {
int64_t run_duration; /* real inference time (us) */
} mtnn_perf_run;

/*
the information for MTNN_QUERY_SDK_VERSION.
*/
typedef struct _mtnn_sdk_version {
char api_version[256]; /* the version of mtnn api. */
char drv_version[256]; /* the version of mtnn driver. */
} mtnn_sdk_version;

/*
the memory information of tensor.
*/
typedef struct _mtnn_tensor_memory {
void* logical_addr; /* the virtual address of tensor buffer. */
uint64_t physical_addr; /* the physical address of tensor buffer. */
int32_t fd; /* the fd of tensor buffer. */
uint32_t size; /* the size of tensor buffer. */
uint32_t handle; /* the handle tensor buffer. */
void * priv_data; /* the data which is reserved. */
uint64_t reserved_flag; /* the flag which is reserved. */
} mtnn_tensor_mem;

/*
the input information for mtnn_input_set.
*/
typedef struct _mtnn_input {
uint32_t index; /* the input index. */
void* buf; /* the input buf for index. */
uint32_t size; /* the size of input buf. */
uint8_t pass_through; /* pass through mode.
if TRUE, the buf data is passed directly
to the input node of the mtnn model without
any conversion. the following variables do not
need to be set.
if FALSE, the buf data is converted into
an input consistent with the model according to
the following type and fmt. so the following
variables need to be set.*/
mtnn_tensor_type type; /* the data type of input buf. */
mtnn_tensor_format fmt; /* the data format of input buf.
currently the internal input format of NPU is NCHW
by default. so entering NCHW data can avoid the format
conversion in the driver. */
} mtnn_input;

/*
the extend information for mtnn_run.
*/
typedef struct _mtnn_run_extend {
uint64_t frame_id; /* output parameter, indicate current frame id of run. */
} mtnn_run_extend;

/*
the extend information for mtnn_outputs_get.
*/
typedef struct _mtnn_output_extend {
uint64_t frame_id; /* output parameter, indicate the frame id of outputs,
corresponds to struct mtnn_run_extend.frame_id.*/
} mtnn_output_extend;

typedef struct _mtnn_tensor_info {
uint32_t index; /* tensor index. */
mtnn_tensor_mem *tensor; /* tensor. */
} mtnn_tensor_info;
MTNN_FLAG​
/*
Definition of extended flag for mtnn_init.
*/

/* set high priority context. */
#define MTNN_FLAG_PRIOR_HIGH 0x00000000
/* set medium priority context. */
#define MTNN_FLAG_PRIOR_MEDIUM 0x00000001
/* set low priority context. */
#define MTNN_FLAG_PRIOR_LOW 0x00000002

/* asynchronous mode.
when enable, mtnn_outputs_get will not block for too long
because it directly retrieves the result of the previous
frame which can increase the frame rate on single-threaded
mode, but at the cost of mtnn_outputs_get not retrieves
the result of the current frame.
in multi-threaded mode you do not need to turn this mode on.
*/
#define MTNN_FLAG_ASYNC_MASK 0x00000004

/* collect performance flag.
when enable, you can get detailed performance reports after
run inference, but it will reduce the frame rate.
*/
#define MTNN_FLAG_COLLECT_PERF_MASK 0x00000008

/* dump layer feature data flag.
when enable, you can get layer feature data when run inference,
data will save in current dir, but it will reduce the frame rate.
*/
#define MTNN_FLAG_DUMP_LAYER_DATA_MASK 0x00000010

/* save pre-compile model. */
#define MTNN_FLAG_PRECOMPILE_MASK 0x00000020
错误码​
/*
Error code returned by the mtnn API.
*/
#define MTNN_SUCC 0 /* execute succeed. */
#define MTNN_ERR_FAIL -1 /* execute failed. */
#define MTNN_ERR_TIMEOUT -2 /* execute timeout. */
#define MTNN_ERR_DEVICE_UNAVAILABLE -3 /* device is unavailable. */
#define MTNN_ERR_MALLOC_FAIL -4 /* memory malloc fail. */
#define MTNN_ERR_PARAM_INVALID -5 /* parameter is invalid. */
#define MTNN_ERR_MODEL_INVALID -6 /* model is invalid. */
#define MTNN_ERR_CTX_INVALID -7 /* context is invalid. */
#define MTNN_ERR_INPUT_INVALID -8 /* input is invalid. */
#define MTNN_ERR_OUTPUT_INVALID -9 /* output is invalid. */
#define MTNN_ERR_DEVICE_UNMATCH -10 /* the device is unmatch, please update
mtnn sdk and npu driver/firmware. */
#define MTNN_ERR_INCOMPATILE_PRE_COMPILE_MODEL -11 /* This mtnn model use pre_compile
mode, but not compatible with
current driver. */
#define MTNN_ERR_INCOMPATILE_OPTIMIZATION_LEVEL_VERSION -12 /* This mtnn model set optimization
level, but not compatible with
current driver. */
#define MTNN_ERR_TARGET_PLATFORM_UNMATCH -13 /* This mtnn model set target
platform, but not compatible with
current platform. */
#define MTNN_ERR_NON_PRE_COMPILED_MODEL_ON_MINI_DRIVER -14 /* This mtnn model is not a
pre-compiled model, but the npu
driver is mini driver. */

3.2 Python API​

Python API 通过 mtnn_api.py 中的 MTNNSession 封装 MTNN Runtime C API,为 NumPy 输入输出提供更简洁的推理接口。

3.2.1 运行要求​

使用前请确认:

  • 系统已正确安装 MTNN Runtime,且 /usr/lib/libmtnnrt.so 可被加载。
  • Python 环境中已安装 numpy 和 cffi。
  • mtnn_api.py 位于 Python 模块搜索路径中。
  • 输入数据的元素数量和数据类型与模型输入属性一致。

3.2.2 MTNNSession 接口​

MTNNSession 在初始化时加载 .mtnn 模型,获取输入输出数量、内存和 Tensor 属性。当前实现默认选择 NPU 0,并将 NPU 频率设置为 MTNN_POWER_FIXED_FREQUENCY_LEVEL4。

接口说明
MTNNSession(model_path)加载指定的 .mtnn 模型并初始化 Session。
run(input_feed, crop_size=None, crop_dst_size=None)复制输入数据、可选更新 Crop 参数、执行推理,并返回 NumPy 输出列表。
get_num_inputs()返回模型输入数量。
get_num_outputs()返回模型输出数量。
close()提前释放 Runtime 资源;可安全地重复调用。
__del__()Session 对象析构时自动调用 close(),并最终调用 mtnn_destroy。
__enter__() / __exit__()支持 with 上下文管理;退出时自动调用 mtnn_destroy 释放资源。

3.2.3 基本推理示例​

程序结束或 MTNNSession 对象析构时会自动调用 mtnn_destroy 释放 Runtime 资源,因此基本调用无需显式调用 close():

import numpy as np

from mtnn_api import MTNNSession

model_path = "model.mtnn"
input_data = np.load("input.npy")

session = MTNNSession(model_path)
print(f"inputs: {session.get_num_inputs()}")
print(f"outputs: {session.get_num_outputs()}")

outputs = session.run({0: input_data})

for index, output in enumerate(outputs):
print(f"output[{index}]: shape={output.shape}, dtype={output.dtype}")

多输入模型通过输入 Index 与 NumPy 数组组成的字典传入:

session = MTNNSession(model_path)
outputs = session.run({
0: input0,
1: input1,
})

如需在程序继续运行期间提前释放资源,仍可显式调用 session.close(),或使用 with MTNNSession(model_path) as session:。

3.2.4 输入与输出​

run 的参数和返回值规则如下:

  • input_feed:字典类型,键为输入 Index,值为 NumPy 数组。首次推理时应提供模型所需的全部输入。
  • 输入数组会通过 numpy.ascontiguousarray 转换为连续内存,再复制到 MTNN Runtime 的输入 Buffer。
  • 输入数组的元素数量必须与模型输入一致,数据类型应与模型输入 Tensor 类型匹配。
  • 返回值为 NumPy 数组列表,顺序与模型输出 Index 一致;每个数组的形状和数据类型由模型输出属性确定。

Python API 支持以下 MTNN Tensor 类型映射:

MTNN Tensor 类型NumPy 数据类型
MTNN_TENSOR_FLOAT32numpy.float32
MTNN_TENSOR_FLOAT16numpy.float16
MTNN_TENSOR_INT8numpy.int8
MTNN_TENSOR_UINT8numpy.uint8
MTNN_TENSOR_INT16numpy.int16
MTNN_TENSOR_INT32numpy.int32
MTNN_TENSOR_INT64numpy.int64
MTNN_TENSOR_BOOL8numpy.bool_

3.2.5 动态 Crop​

对于已在模型转换阶段启用预处理节点和 Crop 的模型,可通过 run 的 crop_size 和 crop_dst_size 参数动态更新裁剪区域,相关的前处理和后续推理会仅在裁剪的张量区域执行。

session = MTNNSession(model_path)
outputs = session.run(
{0: input_data},
crop_size=[start_x, start_y, crop_w, crop_h],
crop_dst_size=[dst_w, dst_h],
)
  • 列表形式的 crop_size 作用于输入 0;多输入模型可使用 {input_index: crop_rect} 字典分别设置。
  • crop_size 格式为 [start_x, start_y, crop_w, crop_h],各值必须为非负数。
  • crop_dst_size 格式为 [dst_w, dst_h],宽和高必须为正数。
  • 未指定 crop_dst_size 时,Session 使用对应模型输入属性中的宽和高。

多输入 Crop 的完整示例参见 2.4.2 运行时动态更新 Crop 参数。

3.2.6 异常与资源释放​

  • 底层 MTNN Runtime API 返回非零错误码时,Python 封装会抛出 RuntimeError。
  • 输入大小不匹配时,run 会触发 AssertionError。
  • Crop 参数长度、取值或输入 Index 无效时,会抛出 ValueError。
  • 程序结束或 Session 对象析构时会自动释放 Runtime 资源;如需提前释放,可调用 close() 或使用 with 上下文管理。

3.3 案例​

M1000 NPU Model Zoo 提供了可直接参考的 MobileNetV1 MTNN 部署案例,包含模型导出、mtc 配置、Python 推理脚本和 C++ Demo。本节整理其中与 MTNN 部署相关的主要步骤。

3.3.1 准备模型​

进入 Model Zoo 的 MobileNetV1 目录,导出 ONNX 模型和测试输入:

cd Image_Classification/001mobilenet_v1
python3 ./export.py

目录中的 mtc_config.yaml 已配置模型路径、输入数据和量化方式,主要配置如下:

配置项MobileNetV1 示例值
ModelPath./model.onnx
Namemobilenetv1
input_file_list./input.npy
force_fp32input_listtrue
force_fp32output_listtrue
quant_typeint16
quantizerdynamic_fixed_point
UseSingleCoretrue

执行模型编译:

mtc --config ./mtc_config.yaml

按照示例配置,生成的模型位于:

/opt/m1000-mtc-toolkit/bin/baselib/data/net_data/onnx/mobilenetv1/mtc/dump_npu_files/mobilenetv1_export_data/mobilenetv1.mtnn

Python 与 C++ Demo 使用相同的图像处理流程:将 BGR 图像转换为 RGB,按短边缩放到 256,中心裁剪为 224 × 224,将数值归一化到 [-1, 1],最后转换为 NCHW Float32 输入。

3.3.2 Python API 部署​

Python 实现位于 MobileNetV1 Python 案例。sample.py 根据模型后缀选择推理后端:ONNX 模型使用 ONNX Runtime,.mtnn 模型使用 mtnn_api.MTNNSession。

使用 .mtnn 模型执行 NPU 推理:

python3 ./sample.py \
--model_path /opt/m1000-mtc-toolkit/bin/baselib/data/net_data/onnx/mobilenetv1/mtc/dump_npu_files/mobilenetv1_export_data/mobilenetv1.mtnn \
--input ../../data/000000039769.jpg

Python Demo 的关键调用如下:

import mtnn_api

session = mtnn_api.MTNNSession(model_path)
output = session.run({0: processed_image})[0]

完整处理流程为:

  1. 通过 OpenCV 读取本地图片或网络图片,并将 BGR 转换为 RGB。
  2. 完成缩放、中心裁剪、归一化和 NCHW 转换,生成 Float32 NumPy 数组。
  3. 调用 MTNNSession.run({0: processed_image}) 执行推理。
  4. 对输出执行 Softmax,并打印 Top-5 分类索引及概率。

同一脚本也可使用 ONNX 模型在 CPU 上运行,便于对比结果:

python3 ./sample.py --model_path ./model.onnx --input ../../data/000000039769.jpg

3.3.3 C/C++ API 部署​

完整工程参见 MobileNetV1 C++ Demo。主要文件包括:

文件作用
sample.cpp程序入口、参数解析及整体流程编排。
mtnn_runner.h/.cpp封装 MTNN 模型加载、Runtime 配置、推理和资源释放。
utils.h/.cpp图像前处理、Softmax 和 Top-K 后处理。
CMakeLists.txt配置 MTNN Runtime 与 OpenCV 的编译和链接。

Demo 的图像前处理依赖 OpenCV:

sudo apt-get install libopencv-dev

在 cpp_demo 目录中构建:

cmake -S . -B build
cmake --build build

运行示例:

./build/sample \
/opt/m1000-mtc-toolkit/bin/baselib/data/net_data/onnx/mobilenetv1/mtc/dump_npu_files/mobilenetv1_export_data/mobilenetv1.mtnn \
../../../data/000000039769.jpg

MtnnRunner 的核心执行流程如下:

  1. 调用 mtnn_set_log_level 设置日志级别,并通过 mtnn_init 加载模型。
  2. 通过 mtnn_set 将 NPU 频率设置为 MTNN_POWER_FIXED_FREQUENCY_LEVEL4,并选择 NPU 0。
  3. 通过 mtnn_get 查询输入输出数量和 Tensor 属性。
  4. 使用 mtnn_get_tensor_from_array 将 NCHW Float32 输入转换为模型需要的数据类型。
  5. 通过 mtnn_inputs_get 获取内部输入 Buffer,复制输入数据后调用 mtnn_inputs_set。
  6. 调用 mtnn_inference 执行推理,再通过 mtnn_outputs_get 获取输出。
  7. 使用 mtnn_dtype_to_float32 将输出转换为 Float32,完成 Softmax 和 Top-5 后处理。
  8. MtnnRunner 析构时调用 mtnn_destroy 释放 Runtime 资源。

CMakeLists.txt 默认从 /usr/include/npu/include 查找 mtnn_api.h,从 /usr/lib 查找 libmtnnrt.so,并链接 OpenCV。安装位置不同时,可通过 CMake 的 MTNN_INCLUDE_DIR 和 MTNN_LIBRARY_DIR 参数覆盖默认路径。

4. 模型仓库​

MT NPU Model Zoo 汇总了可在 M1000 NPU 上部署的 AI 模型,涵盖图像分类、目标检测、图像分割及其他常见任务。各模型目录提供模型导出、mtc 转换和部署运行示例,可帮助开发者快速验证模型并参考已有案例部署自己的 AI 模型。

项目地址:M1000 NPU Model Zoo

5. ONNX Runtime EP​

5.1 概述​

MTNN 提供 ONNX Runtime 支持,可通过安装 WHL 包获取。ONNX Runtime Execution Provider(EP)是 ONNX Runtime 的后端执行机制:应用仍使用 ONNX Runtime API 加载和运行 ONNX 模型,具体算子或子图由指定的 EP 执行。

VSINPU EP 对应的 Provider 名称为 VSINPUExecutionProvider,用于将支持的 ONNX 算子或子图下发到 VSI NPU 执行。如果同时配置 CPUExecutionProvider,ONNX Runtime 会按照 Provider 列表顺序优先使用 VSINPU EP,不支持的节点将回退到 CPU 执行。

部署方式模型流程适用场景
MTNN 标准流程ONNX 模型经 mtc 转换为 .mtnn,再通过 MTNN Runtime API 部署。需要图优化、量化、预处理节点和性能分析等完整工具链能力。
ONNX Runtime VSINPU EP应用直接加载 ONNX 模型,通过 ONNX Runtime EP 机制调用 NPU。现有应用已基于 ONNX Runtime,或需要快速验证 ONNX 模型在 NPU 上的运行情况。

5.2 运行环境检查​

使用 ONNX Runtime EP 前,请确认:

  • 已安装我们提供的ONNX Runtime whl包。
  • NPU 驱动及相关运行库已正确安装。
  • 使用适配 Python 3.12 及以上版本的 ONNX Runtime 1.25 EP。

通过以下命令检查当前 Python 环境中可用的 Provider:

python3 -c "import onnxruntime as ort; print(ort.__version__); print(ort.get_available_providers())"

输出中应包含 VSINPUExecutionProvider,例如:

['VSINPUExecutionProvider', 'CPUExecutionProvider']

如果只看到 CPUExecutionProvider 或其他 Provider,说明当前安装的 ONNX Runtime 包不包含 VSINPU EP。此时需要替换为对应的 ONNX Runtime VSINPU EP 版本,并确认动态库搜索路径中包含 ONNX Runtime、NPU 驱动和运行库依赖。

5.3 Python 使用示例​

Python 通过 providers 参数指定 EP 优先级。建议将 VSINPUExecutionProvider 放在第一位,并保留 CPUExecutionProvider 作为兼容回退。

import numpy as np
import onnxruntime as ort

model_path = "model.onnx"

available_providers = ort.get_available_providers()
if "VSINPUExecutionProvider" not in available_providers:
raise RuntimeError(f"VSINPUExecutionProvider is not available: {available_providers}")

session_options = ort.SessionOptions()
session_options.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL

session = ort.InferenceSession(
model_path,
sess_options=session_options,
providers=["VSINPUExecutionProvider", "CPUExecutionProvider"],
)

input_info = session.get_inputs()[0]
input_data = np.random.rand(1, 3, 224, 224).astype(np.float32)
outputs = session.run(None, {input_info.name: input_data})

print(session.get_providers())
print([output.shape for output in outputs])

使用时需保证输入 Tensor 的 Shape、Data Type 和 Layout 与 ONNX 模型一致。对于多输入模型,应通过 session.get_inputs() 获取每个输入名称,并在 session.run() 中传入完整的输入字典。


版本信息
版本号作者修改日期修改说明
V1.6.0ZYJ2026/08/061. 优化文档,增加 python api 和 完整的部署demo
V1.6.0LHR2026/07/061. MTNN Runtime 支持运行时更新预处理节点的 Crop 参数。
2. 增加 ONNX Runtime VSINPU EP 使用介绍。
V1.6.0WJB、LHR2026/06/091. mtc 增加量化时忽略前处理的选项。
2. 更新版本号。
V1.5.1ZXK2026/01/08更新版本号。
V1.5.0YXS2025/12/02更新版本号。
V1.4.0YXS2025/12/02删除 MLPerf 数据。
V1.4.0YXS2025/09/30将 Netron-MTNN 工具更新至 8.2.2。
V1.4.0YXS2025/09/18更新版本号。
V1.3.3ZXK2025/08/27mtc 支持为每个 Input Port 配置前处理。
V1.3.2ZXK2025/07/311. mtc 增加排除量化算子列表。
2. mtc 增加指定量化算子列表。
V1.3.1YXS2025/05/07修正拼写错误。
V1.3YXS2025/04/29更新 Netron 工具下载地址。
V1.2.2ZZY2025/04/231. 增加 mtc 前处理功能说明。
2. 增加 mtc 指定量化数据集的说明。
V1.2.1WJB2025/04/161. 支持 ONNX Static Quantization 模型。
2. 支持 Float32 输入和输出。
V1.2ZZY2025/03/27完善混合量化说明。
V1.1YXS2025/02/28增加 mtc 扩展工具 mtc_wrap 的使用介绍。
V1.0YXS2025/01/151. 优化 MTNN Runtime API 描述格式。
2. 完善 mtc 工具说明。
3. 增加 NPU MLPerf 相关说明。
V0.9YXS2024/10/111. 增加支持的算子列表。
2. 调整 MTNN Runtime API 描述,并增加示例。
V0.5YXS2024/06/09初始版本。