模型性能数据
以下数据来自 vLLM-MUSA v1.4.2 在 M1000 端侧环境的实测结果。均在性能模式+compute_only环境下测试,数值供参考。
测试环境
- 端设备:M1000 模组,64G 统一内存, DDR6400
- 系统版本:V1.4.1
- 服务参数:
max-model-len=16384
- 文本测试并发:
1、2
- 文本测试输入 / 输出:
32/32、128/128、1024/1024、3072/1024、8192/1024
- 图片测试输入 / 输出:
128/128,单图,分辨率 256×256、512×512、720×1280
指标说明
LLM / VLM 文本性能字段:
Batch:并发数。
Input:输入 token 数。
Output:输出 token 数。
Decode(tps):Decode 吞吐率,即输出 TPS,为1000/TPOT(ms)。
TTFT(ms):首 token 平均延迟。
TPOT(ms):平均每输出 token 耗时。
图片输入性能字段:
Image:图片分辨率。
- 其余字段含义同 LLM / VLM 文本性能字段。
Embedding / Reranker 性能字段:
Input / Output:输入 / 输出 token 数。
QPS:实际请求 QPS。
Output(tps):输出吞吐。
Total(tps):总吞吐。
TTFT(ms):首 token 平均延迟。
LLM / VLM 文本性能指标
Qwen3.6
gptq-Qwen3.6-35B-A3B-4bit-group(MoE GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 19.33 | 682.98 | 51.74 |
| 1 | 128.0 | 128.0 | 18.60 | 1200.71 | 53.75 |
| 1 | 1024.0 | 1024.0 | 18.16 | 4884.80 | 55.08 |
| 1 | 3072.0 | 1024.0 | 17.55 | 13889.76 | 56.99 |
| 1 | 8192.0 | 1024.0 | 16.18 | 37621.42 | 61.80 |
| 2 | 32.0 | 32.0 | 10.10 | 1343.77 | 98.97 |
| 2 | 128.0 | 128.0 | 9.81 | 2395.26 | 101.95 |
| 2 | 1024.0 | 1024.0 | 9.52 | 9776.32 | 105.04 |
| 2 | 3072.0 | 1024.0 | 8.96 | 24139.79 | 111.55 |
| 2 | 8192.0 | 1024.0 | 7.69 | 60426.07 | 130.00 |
Qwen3.6-27B-gptq-v1(Dense GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 4.19 | 1403.54 | 238.64 |
| 1 | 128.0 | 128.0 | 4.03 | 2780.04 | 248.01 |
| 1 | 1024.0 | 1024.0 | 3.86 | 13649.41 | 259.04 |
| 1 | 3072.0 | 1024.0 | 3.67 | 45759.20 | 272.31 |
| 1 | 8192.0 | 1024.0 | 3.24 | 130219.21 | 308.89 |
| 2 | 32.0 | 32.0 | 1.46 | 2103.60 | 684.85 |
| 2 | 128.0 | 128.0 | 1.42 | 4540.62 | 706.17 |
| 2 | 1024.0 | 1024.0 | 1.37 | 26420.41 | 729.73 |
| 2 | 3072.0 | 1024.0 | 1.30 | 80814.48 | 769.05 |
| 2 | 8192.0 | 1024.0 | 1.12 | 199596.19 | 890.39 |
Qwen3.5
gptq-Qwen3.5-35B-A3B-full-4bit-group(MoE GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 128.0 | 18.94 | 723.93 | 52.81 |
| 1 | 128.0 | 128.0 | 18.83 | 1200.50 | 53.12 |
| 1 | 1024.0 | 128.0 | 18.42 | 4702.10 | 54.29 |
| 1 | 8192.0 | 128.0 | 16.45 | 35673.32 | 60.79 |
gptq-Qwen3.5-9B-full-int4-group(Dense GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 11.79 | 394.95 | 84.85 |
| 1 | 128.0 | 128.0 | 11.31 | 742.54 | 88.40 |
| 1 | 1024.0 | 1024.0 | 10.75 | 3802.27 | 93.04 |
| 1 | 3072.0 | 1024.0 | 9.96 | 12133.66 | 100.39 |
| 1 | 8192.0 | 1024.0 | 8.38 | 33830.81 | 119.29 |
| 2 | 32.0 | 32.0 | 4.11 | 766.67 | 243.13 |
| 2 | 128.0 | 128.0 | 3.98 | 1460.13 | 250.97 |
| 2 | 1024.0 | 1024.0 | 3.82 | 7609.67 | 262.00 |
| 2 | 3072.0 | 1024.0 | 3.57 | 21808.03 | 280.02 |
| 2 | 8192.0 | 1024.0 | 3.02 | 53283.27 | 330.95 |
gptq-Qwen3.5-4B-full-int4-group(Dense GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 18.38 | 261.33 | 54.42 |
| 1 | 128.0 | 128.0 | 17.88 | 483.24 | 55.94 |
| 1 | 1024.0 | 1024.0 | 16.28 | 2386.08 | 61.42 |
| 1 | 3072.0 | 1024.0 | 14.51 | 7433.96 | 68.92 |
| 1 | 8192.0 | 1024.0 | 11.34 | 20759.12 | 88.15 |
| 2 | 32.0 | 32.0 | 7.68 | 480.22 | 130.15 |
| 2 | 128.0 | 128.0 | 7.44 | 939.95 | 134.46 |
| 2 | 1024.0 | 1024.0 | 6.89 | 4739.19 | 145.13 |
| 2 | 3072.0 | 1024.0 | 6.18 | 13075.00 | 161.70 |
| 2 | 8192.0 | 1024.0 | 4.80 | 32274.46 | 208.36 |
Qwen3.5-4B(Dense,BF16)
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 9.06 | 324.21 | 110.32 |
| 1 | 128.0 | 128.0 | 8.81 | 480.09 | 113.47 |
| 1 | 1024.0 | 1024.0 | 8.40 | 2056.30 | 119.04 |
| 1 | 3072.0 | 1024.0 | 7.89 | 5545.91 | 126.73 |
| 1 | 8192.0 | 1024.0 | 6.85 | 15006.32 | 145.92 |
| 2 | 32.0 | 32.0 | 7.38 | 505.13 | 135.50 |
| 2 | 128.0 | 128.0 | 7.22 | 818.44 | 138.57 |
| 2 | 1024.0 | 1024.0 | 6.72 | 4016.55 | 148.85 |
| 2 | 3072.0 | 1024.0 | 6.06 | 9540.27 | 165.06 |
| 2 | 8192.0 | 1024.0 | 4.77 | 23510.17 | 209.57 |
Qwen3-VL
gptq-Qwen3-VL-8B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 13.06 | 322.76 | 76.55 |
| 1 | 128.0 | 128.0 | 12.40 | 594.95 | 80.64 |
| 1 | 1024.0 | 1024.0 | 10.58 | 3176.45 | 94.50 |
| 1 | 3072.0 | 1024.0 | 8.65 | 10904.11 | 115.57 |
| 1 | 8192.0 | 1024.0 | 5.95 | 32595.23 | 168.05 |
| 2 | 32.0 | 32.0 | 4.47 | 605.36 | 223.63 |
| 2 | 128.0 | 128.0 | 4.30 | 1138.39 | 232.73 |
| 2 | 1024.0 | 1024.0 | 3.82 | 6321.58 | 261.92 |
| 2 | 3072.0 | 1024.0 | 3.26 | 19126.13 | 306.31 |
| 2 | 8192.0 | 1024.0 | 0.22 | 4699074.44 | 4566.25 |
gptq-Qwen3-VL-4B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 20.75 | 182.96 | 48.20 |
| 1 | 128.0 | 128.0 | 19.70 | 348.21 | 50.75 |
| 1 | 1024.0 | 1024.0 | 15.36 | 1775.45 | 65.11 |
| 1 | 3072.0 | 1024.0 | 11.59 | 6051.21 | 86.25 |
| 1 | 8192.0 | 1024.0 | 7.19 | 19708.28 | 139.02 |
| 2 | 32.0 | 32.0 | 8.49 | 327.91 | 117.84 |
| 2 | 128.0 | 128.0 | 8.07 | 656.81 | 123.98 |
| 2 | 1024.0 | 1024.0 | 6.59 | 3536.24 | 151.71 |
| 2 | 3072.0 | 1024.0 | 5.14 | 10742.99 | 194.71 |
| 2 | 8192.0 | 1024.0 | 0.37 | 2764944.77 | 2695.72 |
gptq-Qwen3-VL-2B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 38.36 | 98.43 | 26.07 |
| 1 | 128.0 | 128.0 | 36.05 | 161.66 | 27.74 |
| 1 | 1024.0 | 1024.0 | 25.62 | 745.43 | 39.03 |
| 1 | 3072.0 | 1024.0 | 18.01 | 2405.86 | 55.53 |
| 1 | 8192.0 | 1024.0 | 10.37 | 7756.53 | 96.39 |
| 2 | 32.0 | 32.0 | 16.38 | 173.50 | 61.04 |
| 2 | 128.0 | 128.0 | 15.43 | 296.31 | 64.79 |
| 2 | 1024.0 | 1024.0 | 11.53 | 1455.69 | 86.72 |
| 2 | 3072.0 | 1024.0 | 8.35 | 4246.47 | 119.78 |
| 2 | 8192.0 | 1024.0 | 0.93 | 1088641.03 | 1072.05 |
Qwen3
Qwen3-30B-A3B-GPTQ-Int4(MoE GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 22.26 | 742.56 | 44.91 |
| 1 | 128.0 | 128.0 | 21.46 | 1258.59 | 46.59 |
| 1 | 1024.0 | 1024.0 | 20.21 | 5354.20 | 49.49 |
| 1 | 3072.0 | 1024.0 | 18.77 | 15801.16 | 53.27 |
| 1 | 8192.0 | 1024.0 | 15.97 | 45665.25 | 62.63 |
| 2 | 32.0 | 32.0 | 12.13 | 1462.62 | 82.47 |
| 2 | 128.0 | 128.0 | 11.57 | 2487.79 | 86.42 |
| 2 | 1024.0 | 1024.0 | 10.86 | 10691.06 | 92.11 |
| 2 | 3072.0 | 1024.0 | 9.53 | 27671.27 | 104.96 |
| 2 | 8192.0 | 1024.0 | 0.16 | 6351976.66 | 6130.48 |
gptq-Qwen3-8B(Dense GPTQ,GPTQ INT4(W4A16))
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 10.78 | 286.09 | 92.74 |
| 1 | 128.0 | 128.0 | 10.38 | 488.14 | 96.37 |
| 1 | 1024.0 | 1024.0 | 9.01 | 2092.27 | 110.96 |
| 1 | 3072.0 | 1024.0 | 7.57 | 5921.74 | 132.16 |
| 1 | 8192.0 | 1024.0 | 5.41 | 18432.20 | 184.74 |
| 2 | 32.0 | 32.0 | 6.07 | 468.50 | 164.71 |
| 2 | 128.0 | 128.0 | 5.84 | 870.44 | 171.25 |
| 2 | 1024.0 | 1024.0 | 5.02 | 4063.83 | 199.20 |
| 2 | 3072.0 | 1024.0 | 4.13 | 10085.31 | 241.91 |
| 2 | 8192.0 | 1024.0 | 0.39 | 2598680.30 | 2547.36 |
Qwen3-8B(Dense,BF16)
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 5.71 | 441.86 | 175.01 |
| 1 | 128.0 | 128.0 | 5.54 | 596.23 | 180.65 |
| 1 | 1024.0 | 1024.0 | 5.11 | 2271.19 | 195.82 |
| 1 | 3072.0 | 1024.0 | 4.61 | 6400.03 | 217.03 |
| 1 | 8192.0 | 1024.0 | 3.71 | 19240.77 | 269.79 |
| 2 | 32.0 | 32.0 | 4.49 | 637.34 | 222.83 |
| 2 | 128.0 | 128.0 | 4.36 | 949.79 | 229.45 |
| 2 | 1024.0 | 1024.0 | 3.88 | 4323.20 | 257.53 |
| 2 | 3072.0 | 1024.0 | 3.33 | 10676.42 | 300.48 |
| 2 | 8192.0 | 1024.0 | 0.37 | 2756780.86 | 2721.17 |
Qwen3-0.6B(Dense,BF16)
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 48.52 | 66.01 | 20.61 |
| 1 | 128.0 | 128.0 | 44.93 | 86.33 | 22.26 |
| 1 | 1024.0 | 1024.0 | 29.99 | 310.73 | 33.34 |
| 1 | 3072.0 | 1024.0 | 20.05 | 1001.34 | 49.86 |
| 1 | 8192.0 | 1024.0 | 11.00 | 3900.75 | 90.94 |
| 2 | 32.0 | 32.0 | 38.79 | 93.42 | 25.78 |
| 2 | 128.0 | 128.0 | 35.14 | 138.29 | 28.46 |
| 2 | 1024.0 | 1024.0 | 20.10 | 581.38 | 49.74 |
| 2 | 3072.0 | 1024.0 | 12.16 | 1702.64 | 82.23 |
| 2 | 8192.0 | 1024.0 | 1.95 | 510700.93 | 513.83 |
Hunyuan
Hy-MT2-7B(Dense,BF16)
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 5.67 | 452.73 | 176.49 |
| 1 | 128.0 | 128.0 | 5.49 | 639.45 | 182.14 |
| 1 | 1024.0 | 1024.0 | 5.10 | 2310.87 | 195.91 |
| 1 | 3072.0 | 1024.0 | 4.65 | 6482.91 | 215.21 |
| 1 | 8192.0 | 1024.0 | 3.79 | 19674.85 | 263.55 |
| 2 | 32.0 | 32.0 | 4.49 | 657.06 | 222.77 |
| 2 | 128.0 | 128.0 | 4.37 | 1030.84 | 228.79 |
| 2 | 1024.0 | 1024.0 | 3.93 | 4405.35 | 254.18 |
| 2 | 3072.0 | 1024.0 | 3.40 | 10874.94 | 294.01 |
| 2 | 8192.0 | 1024.0 | 0.37 | 2727933.37 | 2692.03 |
Hy-MT2-1.8B(Dense,BF16)
| Batch | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 32.0 | 32.0 | 18.76 | 128.85 | 53.32 |
| 1 | 128.0 | 128.0 | 18.04 | 183.24 | 55.42 |
| 1 | 1024.0 | 1024.0 | 16.59 | 645.86 | 60.29 |
| 1 | 3072.0 | 1024.0 | 15.36 | 1801.68 | 65.09 |
| 1 | 8192.0 | 1024.0 | 13.06 | 6120.81 | 76.55 |
| 2 | 32.0 | 32.0 | 14.89 | 176.11 | 67.16 |
| 2 | 128.0 | 128.0 | 14.35 | 279.98 | 69.68 |
| 2 | 1024.0 | 1024.0 | 12.79 | 1208.11 | 78.20 |
| 2 | 3072.0 | 1024.0 | 11.36 | 3062.23 | 88.06 |
| 2 | 8192.0 | 1024.0 | 1.21 | 840625.57 | 828.17 |
图片输入性能指标
Qwen3-VL
gptq-Qwen3-VL-30B-A3B-Instruct-4bit-group(VLM MoE GPTQ,GPTQ INT4(W4A16))
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 128 | 128 | 19.21 | 1878.99 | 52.06 |
| 1 | 512×512 | 128 | 128 | 18.26 | 2906.94 | 54.77 |
| 1 | 720×1280 | 128 | 128 | 15.74 | 6442.30 | 63.52 |
gptq-Qwen3-VL-8B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 128 | 128 | 12.23 | 781.11 | 81.79 |
| 1 | 512×512 | 128 | 128 | 11.93 | 1495.05 | 83.82 |
| 1 | 720×1280 | 128 | 128 | 11.10 | 4015.97 | 90.06 |
gptq-Qwen3-VL-4B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 128 | 128 | 18.83 | 456.18 | 53.12 |
| 1 | 512×512 | 128 | 128 | 18.15 | 896.04 | 55.09 |
| 1 | 720×1280 | 128 | 128 | 16.31 | 2352.82 | 61.33 |
gptq-Qwen3-VL-2B-Instruct-4bit-group(VLM GPTQ,GPTQ INT4(W4A16))
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 128 | 128 | 33.41 | 221.83 | 29.93 |
| 1 | 512×512 | 128 | 128 | 31.75 | 500.24 | 31.49 |
| 1 | 720×1280 | 128 | 128 | 27.33 | 1377.59 | 36.58 |
OCR
PaddleOCR-VL(OCR,BF16)
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 128 | 128 | 63.13 | 89.77 | 15.84 |
| 1 | 512×512 | 128 | 128 | 60.57 | 440.23 | 16.51 |
| 1 | 720×1280 | 128 | 128 | 47.98 | 1812.57 | 20.84 |
DeepSeek-OCR-2(OCR,BF16)
| Batch | Image | Input | Output | Decode(tps) | TTFT(ms) | TPOT(ms) |
|---|
| 1 | 256×256 | 127 | 128 | 48.92 | 312.53 | 20.44 |
| 1 | 512×512 | 127 | 128 | 48.92 | 890.55 | 20.44 |
| 1 | 720×1280 | 127 | 128 | 46.61 | 1371.55 | 21.45 |
Embedding / Reranker 性能指标
bge-m3(Embedding,FP16)
| Batch | Input / Output | QPS | Output(tps) | Total(tps) | TTFT(ms) |
|---|
| 1 | 1024/1024 | 0.73 | 748 | 1496 | 250.9 |
| 1 | 128/128 | 0.74 | 93 | 187 | 45.0 |
| 1 | 3072/1024 | 0.63 | 1941 | 3881 | 1441.9 |
| 1 | 32/32 | 0.74 | 22 | 44 | 43.8 |
| 2 | 1024/1024 | 0.73 | 749 | 1497 | 241.8 |
| 2 | 128/128 | 0.74 | 93 | 187 | 44.4 |
| 2 | 3072/1024 | 0.64 | 1968 | 3935 | 2139.2 |
| 2 | 32/32 | 0.74 | 22 | 44 | 43.2 |
bge-reranker-v2-m3(Reranker,FP16)
| Batch | Input / Output | QPS | Output(tps) | Total(tps) | TTFT(ms) |
|---|
| 1 | 1024/1024 | 0.74 | 758 | 1515 | 2.2 |
| 1 | 128/128 | 0.74 | 93 | 187 | 1.9 |
| 1 | 3072/1024 | 0.74 | 2276 | 4552 | 2.5 |
| 1 | 32/32 | 0.74 | 22 | 44 | 2.4 |
| 1 | 8196/1024 | 0.74 | 6075 | 12150 | 3.3 |
| 2 | 1024/1024 | 0.74 | 758 | 1515 | 1.9 |
| 2 | 128/128 | 0.74 | 93 | 187 | 2.2 |
| 2 | 3072/1024 | 0.74 | 2276 | 4552 | 2.5 |
| 2 | 32/32 | 0.74 | 22 | 44 | 2.2 |
| 2 | 8196/1024 | 0.74 | 6073 | 12146 | 3.2 |
本章节提供基于 vllm bench serve 的快速验证命令,用于确认文本生成和多模态图片请求链路可以正常完成,并快速查看 TTFT、TPOT、吞吐等基础指标。以下示例默认使用单请求、单并发配置。
测试前准备
性能测试需要先启动模型服务,再另开一个窗口运行 vllm bench serve。
建议在运行 benchmark 前先激活环境,并显式配置 conda 运行时库路径:
source /home/dev/miniforge3/etc/profile.d/conda.sh
conda activate v1.4.2
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib:$LD_LIBRARY_PATH"
如果未配置 LD_LIBRARY_PATH=$CONDA_PREFIX/lib:$LD_LIBRARY_PATH,可能会遇到 CXXABI_1.3.15 not found 相关报错,处理方式请参考“常见问题(FAQ)”。
文本生成快速验证
先按“快速开始”章节启动文本生成模型服务,例如 gptq-Qwen3-8B。服务启动后,在另一个窗口执行:
vllm bench serve \
--backend openai-chat \
--base-url http://127.0.0.1:8000 \
--endpoint /v1/chat/completions \
--model /home/dev/models/gptq-Qwen3-8B \
--dataset-name random \
--random-input-len 128 \
--random-output-len 128 \
--num-prompts 1 \
--max-concurrency 1 \
--ignore-eos
多模态图片快速验证
多模态图片性能快速验证适用于 Qwen-VL、Qwen3-VL 等图片输入模型。先按“快速开始”章节启动多模态模型服务,例如 gptq-Qwen3-VL-8B-Instruct-4bit-group。如果使用该 8B 示例,服务启动时需要包含:
--limit-mm-per-prompt '{"image":1}'
--hf-overrides '{"text_config":{"tie_word_embeddings":false}}'
如果改测 Qwen3-VL 2B/4B,只保留 --limit-mm-per-prompt '{"image":1}',不要增加 --hf-overrides;如果改测 Qwen3-VL 30B-A3B,与 8B 一样需要增加 --hf-overrides '{"text_config":{"tie_word_embeddings":false}}'。
服务启动后,在另一个窗口执行:
vllm bench serve \
--backend openai-chat \
--base-url http://127.0.0.1:8000 \
--endpoint /v1/chat/completions \
--model /home/dev/models/gptq-Qwen3-VL-8B-Instruct-4bit-group \
--dataset-name random-mm \
--num-prompts 1 \
--max-concurrency 1 \
--random-input-len 128 \
--random-output-len 128 \
--random-mm-base-items-per-request 1 \
--random-mm-limit-mm-per-prompt '{"image":1,"video":0}' \
--random-mm-bucket-config '{(256, 256, 1): 1.0}'
上述命令使用 random-mm 随机多模态数据集,发送 1 条多模态请求。该请求包含约 128 token 文本输入、1 张 256×256 随机图片,并请求模型生成约 128 token。
如需快速验证不同图片分辨率,可调整 --random-mm-bucket-config:
--random-mm-bucket-config '{(512, 512, 1): 1.0}'
--random-mm-bucket-config '{(720, 1280, 1): 1.0}'
当前 random-mm 适合快速验证多模态图片输入链路和基础指标。真实业务需求请使用需求真实图片数据集和固定提示词进行测试。
参数说明
| 参数名 | 参数作用 | 输入范围 / 建议值 |
|---|
--backend | 指定 benchmark 使用的服务后端 | OpenAI chat 接口使用 openai-chat |
--base-url | 指定 vLLM 服务地址 | 默认本机服务可使用 http://127.0.0.1:8000 |
--endpoint | 指定请求接口 | 文本和多模态 chat 使用 /v1/chat/completions |
--model | 指定待测试的模型名称 / 本地路径 | 字符串,需与服务端部署的模型路径或 --served-model-name 一致 |
--dataset-name | 指定性能测试使用的数据集类型 | 文本随机数据集使用 random,随机多模态数据集使用 random-mm |
--random-input-len | 使用随机数据集时,设置输入 Prompt 的 token 长度 | 正整数,快速验证可使用 128 |
--random-output-len | 使用随机数据集时,设置模型输出 token 长度 | 正整数,快速验证可使用 128 |
--num-prompts | 指定本次 benchmark 总请求数 | 快速验证使用 1;多请求稳定性或吞吐测试可增大到 10、100 或更多 |
--max-concurrency | 指定最大并发请求数 | 单请求快速验证使用 1 |
--ignore-eos | 忽略 EOS 结束符 | 文本性能测试时启用可保证输出长度稳定,避免提前停止影响结果 |
--random-mm-base-items-per-request | 多模态测试中,每条请求的多模态输入数量基准值 | 单图测试使用 1 |
--random-mm-limit-mm-per-prompt | 限制每条请求中不同多模态输入的数量 | 单图测试使用 '{"image":1,"video":0}' |
--random-mm-bucket-config | 配置随机图片尺寸及采样概率 | 例如 '{(256, 256, 1): 1.0}' 表示每条请求使用 1 张 256×256 图片 |
--num-prompts 说明
--num-prompts 表示本次 benchmark 总共发送多少条请求。快速验证场景建议设置为:
对于多模态 random-mm 测试,每条请求通常包含:
- 一段随机文本输入;
- 按
--random-mm-* 参数生成的图片输入;
- 指定长度的输出请求。
如果需要做多请求稳定性或吞吐测试,可以将 --num-prompts 增大,例如 10、100 或更多,并根据测试目标调整 --max-concurrency。
输出示例
执行成功后,vllm bench serve 会输出类似如下指标。具体数值与模型、输入长度、输出长度、图片尺寸、并发数和运行环境有关,请以实际测试输出为准。
============ Serving Benchmark Result ============
Successful requests: 1
Failed requests: 0
Maximum request concurrency: 1
Benchmark duration (s): 11.68
Total input tokens: 128
Total generated tokens: 128
Request throughput (req/s): 0.09
Output token throughput (tok/s): 10.96
Peak output token throughput (tok/s): 12.00
Peak concurrent requests: 1.00
Total token throughput (tok/s): 21.91
---------------Time to First Token----------------
Mean TTFT (ms): 582.24
Median TTFT (ms): 582.24
P99 TTFT (ms): 582.24
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 87.40
Median TPOT (ms): 87.40
P99 TPOT (ms): 87.40
---------------Inter-token Latency----------------
Mean ITL (ms): 86.72
Median ITL (ms): 87.78
P99 ITL (ms): 91.97
==================================================
- 一个窗口启动模型服务(
vllm serve),另一个窗口运行性能测试(vllm bench serve)。
- 本章节命令定位为快速验证,默认使用单请求、单并发。
- 上述输出示例仅说明结果格式,不代表固定性能数据。实际结果请以本地测试输出为准。
- 正式性能测试建议固定并发、请求数、输入/输出长度、图片尺寸和测试环境,并多次运行后统计结果。