推理性能
以下推理性能数据均在打开性能模式条件下测试。
Qwen3-30B-A3B-GPTQ-Int4
| concurrency | input_len | output_len | TTFT (ms) | ITL (ms) | TPS(out) |
|---|---|---|---|---|---|
| 1 | 32 | 32 | 1058.27 | 70.21 | 14.24 |
| 1 | 128 | 128 | 844.43 | 73.75 | 13.55 |
| 1 | 1024 | 1024 | 887.06 | 81.22 | 12.31 |
| 1 | 3072 | 1024 | 1132.34 | 92.66 | 10.79 |
| 1 | 8192 | 1024 | 1295.60 | 117.25 | 8.52 |
gptq-Qwen3-8B
| concurrency | input_len | output_len | TTFT (ms) | ITL (ms) | TPS(out) |
|---|---|---|---|---|---|
| 1 | 32 | 32 | 558.19 | 125.88 | 7.94 |
| 1 | 128 | 128 | 551.6 | 127.76 | 7.82 |
| 1 | 1024 | 1024 | 575.29 | 139.43 | 7.17 |
| 1 | 3072 | 1024 | 732.98 | 153.07 | 6.53 |
| 1 | 8192 | 1024 | 939.55 | 188.27 | 5.31 |
gptq-Qwen2.5-7B-Instruct
| concurrency | input_len | output_len | TTFT (ms) | ITL (ms) | TPS(out) |
|---|---|---|---|---|---|
| 1 | 32 | 32 | 204.35 | 117.04 | 8.54 |
| 1 | 128 | 128 | 194.64 | 116.74 | 8.56 |
| 1 | 1024 | 1024 | 206.8 | 121.22 | 8.24 |
| 1 | 3072 | 1024 | 351.11 | 127.09 | 7.86 |
| 1 | 8192 | 1024 | 481.15 | 141.56 | 7.06 |
测试性能
conda activate v1.3.1
输入示例
vllm bench serve \
--model Qwen3-30B-A3B-GPTQ-Int4 \
--dataset_name random \
--random_input_len 128 \
--random_output_len 128 \
--num-prompts 1 \
--trust-remote-code \
--ignore-eos
备注
- 一个窗口启动模型服务(vllm serve),另一个窗口运行性能测试(vllm bench serve)
- 如果报错
transformers找不到:pip3 install transformers==4.52.4 - 性能表格TPS(out)的值为1000/TPOT (ms)得出
- 上述性能仅供参考,运行实际结果受具体环境,测试方法和测试数据集影响
参数说明
| 参数名 | 参数作用 | 输入范围 / 建议值 |
|---|---|---|
--model | 指定待测试的模型名称 / 本地路径 | 字符串,需填写实际部署的模型名/与启动命令相同的模型路径 |
--dataset_name | 指定性能测试使用的数据集类型 | 字符串,常用值: random(随机数据集)、sharegpt(真实对话数据集)等 |
--random_input_len | 当使用随机数据集时,设置输入 Prompt 的 Token 长度 | 正整数,建议根据实际业务场景设置(如 32/128/512/1024),需≤模型最大输入长度 |
--random_output_len | 当使用随机数据集时,设置模型生成输出的 Token 长度 | 正整数,建议贴合实际生成需求(如 32/128/512/1024),需≤模型最大输出长度 |
--num-prompts | 指定单次性能测试的 Prompt 总数 | 正整数,单测建议 1-10(验证基础性能),压测建议 100-10000(模拟高并发) |
--trust-remote-code | 信任模型的远程自定义代码(无参数值) | 布尔型(无需赋值,加该参数即启用),加载非官方标准模型时建议启用,确保模型正常加载 |
--ignore-eos | 忽略 EOS(结束符)Token(无参数值) | 布尔型(无需赋值,加该参数即启用),性能测试时启用可保证输出长度稳定,避免因提前终止影响测试结果 |
输出示例
============ Serving Benchmark Result ============
Successful requests: 1
Benchmark duration (s): 10.14
Total input tokens: 128
Total generated tokens: 128
Request throughput (req/s): 0.10
Output token throughput (tok/s): 12.63
Total Token throughput (tok/s): 25.26
---------------Time to First Token----------------
Mean TTFT (ms): 844.79
Median TTFT (ms): 844.79
P99 TTFT (ms): 844.79
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 73.14
Median TPOT (ms): 73.14
P99 TPOT (ms): 73.14
---------------Inter-token Latency----------------
Mean ITL (ms): 73.14
Median ITL (ms): 73.03
P99 ITL (ms): 90.54
==================================================

