Torch-MUSA
简介
MUSA 概述
MUSA (Metaverse Unified System Architecture) 是摩尔线程公司为摩尔线程 GPU 推出的一种通用并行计算平台和编程模型。它提供了 GPU 编程的简易接口,用 MUSA 编程可以构建基于 GPU 计算的应用程序,利用 GPU 的并行计算引擎来更加高效地解决比较复杂的计算难题。同时摩尔线程还推出了 MUSA 工具箱(MUSA Toolkits),工具箱中包括 GPU 加速库、运行时库、编译器、调试和优化工具等。MUSA Toolkits 为开发人员在摩尔线程 GPU 上开发和部署高性能异构计算程序提供软件环境。
更多详情,参见 MUSA 官方文档。
PyTorch 概述
PyTorch 是一款开源的深度学习编程框架,可以用于计算机视觉,自然语言处理,语音处理等领域。 PyTorch 使用动态计算,这在构建复杂架构时提供了更大的灵活性。PyTorch 使用核心 Python 概念,如 类、结构和条件循环,因此理解起来更直观,编程更容易。此外,PyTorch 还具有可以轻松扩展、快速实现、生产部署稳定性强等优点。
更多详情,参见 PyTorch 官方文档。
torch_musa 概述
为了摩尔线程 GPU 能支持开源框架 PyTorch,摩尔线程公司开发了 torch_musa。在 PyTorch
v2.0.0 基础上,torch_musa 以插件的形式来支持摩尔线程 GPU,最大程度与 PyTorch 代码解耦,便于代码维护与升级。torch_musa 利用 PyTorch 提供的第三方后端扩展接口,将摩尔线程高性能计算库动态注册到 PyTorch 上,从而使得 PyTorch 框架能够利用摩尔线程显卡的高性能计算单元。利用摩尔线程显卡 CUDA 兼容的特性,torch_musa 内部引入了 cuda 兼容模块,使 PyTorch 社区的 CUDA
kernels 经过 porting 后可运行在摩尔线程显卡上,而且 CUDA
Porting 的工作是在编译 torch_musa 的过程中自动进行,这大幅降低了 torch_musa 算子适配的成本,提高模型开发效率。同时,torch_musa 在 Python 前端接口与 PyTorch 社区 CUDA 接口形式上基本保持一致,这极大地降低了用户的学习成本和模型的迁移成本。
本手册主要介绍了基于 MUSA 软件栈的 torch_musa 开发指南。
torch_musa 核心代码目录概述
-
torch_musa/tests测试文件。 -
torch_musa/core主要包含 Python module,提供 amp/device/memory/stream/event 等模块的 Python 前端接口。 -
torch_musa/csrcC++ 侧实现代码;-
csrc/amp提供混合精度模块的 C++ 实现。 -
csrc/aten提供 C++ Tensor 库,包 括MUDNN算子适配、CUDA-Porting算子适配等。 -
csrc/core提供核心功能库,包括设备管理、内存分配管理、Stream 管理、Events 管理等。 -
csrc/distributed提供分布式模块的 C++ 实现。
-
m1000_gpu_model_zoo 模型仓库
m1000_gpu_model_zoo旨在演示如何基于 torch_musa 在 MTGPU 进行模型推理加速,帮助开发者在 MTGPU 上快速落地各种 AI 模型推理服务。
环境准备与部署
步骤 1:环境确认
确认操作系统版本
AIOS 1.5.0
确认 musa 和 musa-sdk 版本
musa 5.2.1
musa-sdk 5.1.2
可通过以下指令获得,关注 Version 字段
dpkg -s musa
dpkg -s musa-sdk
步骤 2:驱动更新
执行前提:只 有在步骤 1 中检查发现 musa 或 musa-sdk 版本不符合要求时,才需要执行本步骤。如果版本已正确,请跳过步骤 2,直接进入步骤 3。
安装包下载、安装命令和安装验证流程,请参考 MUSA 安装。
步骤 3:安装 torch、torch_musa、triton
安装脚本:
pip install ./torch-<package_version>.whl
pip install ./torch_musa-<package_version>.whl
pip install triton-<package_version>.whl
步骤 4:环境验证
python3 -c "import torch;import torch_musa;print(torch.musa.is_available())"
输出 True 证明 torch_musa 环境安装正确。
可能出现的问题:
输出 mudnn.so 找不到
执行:
export PATH=/usr/local/musa/bin:${PATH}
export LD_LIBRARY_PATH=/usr/local/musa/lib:${LD_LIBRARY_PATH}
输出 Error in cpuinfo: prctl(PR_SVE_GET_VL) failed
torch_musa 2.7 之前会存在该问题,不影响使用。
NumPy 报错,如 Failed to initialize NumPy 或者
numpy ModuleNotFoundError: No module named 'numpy'
执行 pip3 install numpy==1.26.4
报错 ImportError: libmccl.so.2: cannot open shared object file: No such file or directory
mccl 库在最新的 musa-sdk 中已经包含,请参考 MUSA 安装,更新 musa-sdk。
报错 MUSA driver initialization failed
设备需要连接显示器,输入用户名密码进入桌面。如果没有安全需求,推荐在"设置"->"用户"界面设置为自动登录。
编译安装
torch_musa 源码完全开源,开发者也可以根据实际需要编译源码安装。
编译安装前,需要安装 MUSA Toolkits 软件包、MUDNN 库、MCCL 库、muThrust 库、muAlg 库、muRAND 库、muSPARSE 库。具体安装步骤,请参见相应组件的安装手册。
依赖环境
编译流程
-
向
PyTorch源码打 patch -
编译
PyTorch -
编译
torch_musa
torch_musa 2.9.0 是在 PyTorch
v2.9.0 基础上以插件的方式来支持摩尔线程显卡。开发时涉及到对 PyTorch 源码的修改,目前是以打 patch 的方式实现的。PyTorch 社区正在积极支持第三方后端接入,相关
issue
下已有对应 PR。torch_musa 项目也在持续向 PyTorch 社区提交 PR,以减少编译过程中对 PyTorch 打 patch 的需求。
编译步骤
安装依赖
git clone --depth=10 https://github.com/MooreThreads/muThrust.git
cd muThrust
./mt_build.sh -i
git clone --depth=10 https://github.com/MooreThreads/muAlg.git
cd muAlg
./mt_build.sh -i
使用脚本一键编译(推荐)
在初次编译时,需要执行 bash build.sh
(先编译 PyTorch,再编译 torch_musa)。在后续开发过程中,如果不涉及对 PyTorch 源码的修改,那么执行
bash build.sh -m(仅编译 torch_musa)即可。
MAX_JOBS=8 USE_MCCL=0 USE_KINETO=1 bash ./build.sh --clean --wheel
配置说明
| 示例 | 说明 |
|---|---|
-c/--clean | 清理构建产物后再构建 |
-w/--wheel | 生成 wheel 包并安装 |
-p/ --patch | 仅应用补丁,不构建 |
-d/--debug | Debug 模式构建 |
-t/--torch | 仅构建 PyTorch |
-m/--musa | 仅构建 Torch_MUSA |
--fp64 | 编译支持 fp64 数据类型的内核 |
-a/--asan | 启用 AddressSanitizer 内存检测 |
宏定义说明:
| 宏定义 | 默认值 | 说明 |
|---|---|---|
MAX_JOBS | 1 | 用于编译的 CPU 核心数 |
USE_MCCL | 1 | 是否使用 MCCL(MooreThreads 的通信库) |
USE_KINETO | 1 | 是否使用 Kineto 性能分析库 |
分步骤编译
如果不想使用脚本编译,那么可以按照如下步骤逐步编译。
- 在
PyTorch打 patch
# 请保证 PyTorch 源码和 torch_musa 源码在同级目录或者 export PYTORCH_REPO_PATH=path/to/PyTorch 指向 PyTorch 源码
bash build.sh --only-patch
- 编译
PyTorch
cd pytorch
pip install -r requirements.txt
python setup.py install
# debug mode: DEBUG=1 python setup.py install
# asan mode: USE_ASAN=1 python setup.py install
- 编译
torch_musa
cd torch_musa
pip install -r requirements.txt
python setup.py install
# debug mode: DEBUG=1 python setup.py install
# asan mode: USE_ASAN=1 python setup.py install
快速入门
常用环境变量
开发 torch_musa 过程中常用环境变量如下表所示:
| 环境变量示例 | 所属组件 | 功能说明 |
|---|---|---|
export TORCH_SHOW_CPP_STACKTRACES=1 | PyTorch | 当 python 程序发生错误时显示 PyTorch 中 C++ 调用栈 |
export MUDNN_LOG_LEVEL=INFO | MUDNN | 使能 MUDNN 算子库调用的 log |
export MUSA_VISIBLE_DEVICES=0,1,2,3 | Driver | 控制当前可见的显卡序号 |
export MUSA_LAUNCH_BLOCKING=1 | Driver | 驱动以同步模式下发 MUSA kernel,即当前 kernel 执行结束后再下发下一个 kernel |
常用 API 示例代码
torch_musa 中 Python API 基本与
PyTorch 原生 API
接口保持一致,极大降低了新用户的学习成本。
import torch
import torch_musa
torch.musa.is_available()
torch.musa.device_count()
a = torch.tensor([1.2, 2.3], dtype=torch.float32, device='musa')
b = torch.tensor([1.8, 1.2], dtype=torch.float32, device='cpu').to('musa')
c = torch.tensor([1.8, 1.3], dtype=torch.float32).musa()
d = a + b + c
torch.musa.synchronize()
with torch.musa.device(0):
assert torch.musa.current_device() == 0
if torch.musa.device_count() > 1:
torch.musa.set_device(1)
assert torch.musa.current_device() == 1
torch.musa.synchronize("musa:1")
推理示例代码
import torch
import torch_musa
import torchvision.models as models
model = models.resnet50().eval()
x = torch.rand((1, 3, 224, 224), device="musa")
model = model.to("musa")
# Perform the inference
y = model(x)
训练示例代码
1、常规训练
import torch
import torch_musa
import torchvision
import torchvision.transforms as transforms
import torch.nn as nn
import torch.nn.functional as F
import torch.optim as optim
## 1. prepare dataset
transform = transforms.Compose(
[transforms.ToTensor(),
transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))])
batch_size = 4
trainset = torchvision.datasets.CIFAR10(root='./data', train=True,
download=True, transform=transform)
trainloader = torch.utils.data.DataLoader(trainset, batch_size=batch_size,
shuffle=True, num_workers=2)
testset = torchvision.datasets.CIFAR10(root='./data', train=False,
download=True, transform=transform)
testloader = torch.utils.data.DataLoader(testset, batch_size=batch_size,
shuffle=False, num_workers=2)
classes = ('plane', 'car', 'bird', 'cat','deer', 'dog', 'frog', 'horse', 'ship', 'truck')
device = torch.device("musa")
## 2. build network
class Net(nn.Module):
def __init__(self):
super().__init__()
self.conv1 = nn.Conv2d(3, 6, 5)
self.pool = nn.MaxPool2d(2, 2)
self.conv2 = nn.Conv2d(6, 16, 5)
self.fc1 = nn.Linear(16 * 5 * 5, 120)
self.fc2 = nn.Linear(120, 84)
self.fc3 = nn.Linear(84, 10)
def forward(self, x):
x = self.pool(F.relu(self.conv1(x)))
x = self.pool(F.relu(self.conv2(x)))
x = torch.flatten(x, 1) # flatten all dimensions except batch
x = F.relu(self.fc1(x))
x = F.relu(self.fc2(x))
x = self.fc3(x)
return x
net = Net().to(device)
## 3. define loss and optimizer
criterion = nn.CrossEntropyLoss()
optimizer = optim.SGD(net.parameters(), lr=0.001, momentum=0.9)
## 4. train
for epoch in range(2): # loop over the dataset multiple times
running_loss = 0.0
for i, data in enumerate(trainloader, 0):
# get the inputs; data is a list of [inputs, labels]
inputs, labels = data
# zero the parameter gradients
optimizer.zero_grad()
# forward + backward + optimize
outputs = net(inputs.to(device))
loss = criterion(outputs, labels.to(device))
loss.backward()
optimizer.step()
# print statistics
running_loss += loss.item()
if i % 2000 == 1999: # print every 2000 mini-batches
print(f'[{epoch + 1}, {i + 1:5d}] loss: {running_loss / 2000:.3f}')
running_loss = 0.0
print('Finished Training')
PATH = './cifar_net.pth'
torch.save(net.state_dict(), PATH)
net.load_state_dict(torch.load(PATH))
## 5. test
correct = 0
total = 0
# since we're not training, we don't need to calculate the gradients for our outputs
with torch.no_grad():
for data in testloader:
images, labels = data
# calculate outputs by running images through the network
outputs = net(images.to(device))
# the class with the highest energy is what we choose as prediction
_, predicted = torch.max(outputs.data, 1)
total += labels.size(0)
correct += (predicted == labels.to(device)).sum().item()
print(f'Accuracy of the network on the 10000 test images: {100 * correct // total} %')
2、混合精度 AMP 训练示例代码
import torch
import torch_musa
import torch.nn as nn
class SimpleModel(nn.Module):
def __init__(self):
super().__init__()
self.fc1 = nn.Linear(5, 4)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(4, 3)
def forward(self, x):
x = self.fc1(x)
x = self.relu(x)
x = self.fc2(x)
return x
def __call__(self, x):
return self.forward(x)
DEVICE = "musa"
def train_in_amp(low_dtype=torch.float16):
model = SimpleModel().to(DEVICE)
criterion = nn.MSELoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
# create the scaler object
scaler = torch.musa.amp.GradScaler()
inputs = torch.randn(6, 5).to(DEVICE) # 将数据移至 GPU
targets = torch.randn(6, 3).to(DEVICE)
for step in range(20):
optimizer.zero_grad()
# create autocast environment
with torch.musa.amp.autocast(dtype=low_dtype):
outputs = model(inputs)
assert outputs.dtype == low_dtype
loss = criterion(outputs, targets)
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
return loss
if __name__ == "__main__":
train_in_amp(torch.float16)
使能 TensorCore 示例代码
在 M1000 上,当输入数据类型是 float32 时,可以通过设置 TensorFloat32 来使能 TensorCore,从而加速计算过程。TensorFloat32 的加速原理可以参考
TensorFloat-32
。
import torch
import torch_musa
with torch.backends.mudnn.flags(allow_tf32=True):
assert torch.backends.mudnn.allow_tf32
a = torch.randn(10240, 10240, dtype=torch.float, device='musa')
b = torch.randn(10240, 10240, dtype=torch.float, device='musa')
result_tf32 = a @ b
torch.backends.mudnn.allow_tf32 = True
assert torch_musa._MUSAC._get_allow_tf32()
a = torch.randn(10240, 10240, dtype=torch.float, device='musa')
b = torch.randn(10240, 10240, dtype=torch.float, device='musa')
result_tf32 = a @ b
C++ 部署示例代码
#include <torch/script.h>
#include <torch_musa/csrc/core/Device.h>
#include <cassert>
#include <iostream>
#include <memory>
int main(int argc, const char* argv[]) {
// Register 'musa' for PrivateUse1 as we save model with 'musa'.
c10::register_privateuse1_backend("musa");
torch::jit::script::Module module;
try {
// Load model which saved with torch jit.trace or jit.script.
module = torch::jit::load(argv[1]);
} catch (const c10::Error& e) {
std::cerr << "error loading the model\n";
return -1;
}
std::vector<torch::jit::IValue> inputs;
// Ready for input data.
torch::Tensor input = torch::rand({1, 3, 224, 224}).to("musa");
assert(input.is_privateuseone() == input.is_musa());
assert(input.device().is_privateuseone() == input.device().is_musa());
inputs.push_back(input);
// Model execute.
at::Tensor output = module.forward(inputs).toTensor();
std::cout << output.slice(/*dim=*/1, /*start=*/0, /*end=*/5) << std::endl;
return 0;
}
详细用法请参考 examples/cpp 下内容。
端侧推理与编译加速
M1000 具有强大的 GPU 算力,能够高效支持各类 AI 模型的端侧推理任务。PyTorch 提供了多种推理方式。这源于不同部署场景的差异化需求以及 PyTorch 自身的架构演进。本节将详细介绍 eager
mode、torch.compile、AOTI 三种主流的推理模式及其在 torch_musa 中的应用。
m1000_gpu_model_zoo中存放了在 M1000 上基于 torch_musa 推理的热门模型,可帮助开发者快速上手部署。
注意:
并非所有模型都能够通过
torch.compile和aoti部署,具体原因可以通过TORCH_LOGS进行排查。
eager mode (即时执行模式)
Eager
Mode 是 PyTorch 的默认执行模式,采用逐行解释执行的方式,操作在调用时立即执行。这种模式最适合模型开发、调试和原型验证阶段。
基本用法
import torch
import torch_musa
import torchvision.models as models
model = models.resnet18().to("musa").eval()
x = torch.randn(1, 3, 224, 224, device="musa")
with torch.no_grad():
output = model(x)
print(output)
torch.compile 模式
PyTorch 2.0 引入了 torch.compile,基于 TorchDynamo + TorchInductor
架构,实现零代码改动即可获得显著加速。TorchDynamo 从 Python 字节码捕获图结构,TorchInductor 将图编译为高效的后端代码。更多信息,可参考
torch.compile 官方文档
基本用法
import torch
import torchvision.models as models
model = models.resnet18().to("musa").eval()
model = torch.compile(model, mode="max-autotune")
dummy_input = torch.randn(1, 3, 224, 224, device="musa")
with torch.no_grad():
output = model(dummy_input) # compile the model
print(output)
相比 eager 模式,只需要增加一行代码,即可增加推理速度。
第一次编译需要花费较多时间,期间可能出现 warning
RuntimeWarning: MUSA backend: num_stages>1 requested without using TME; falling back to num_stages=1.
stages["ttgir"] = lambda src, metadata: self.make_ttgir(src, metadata, options, self.capability, self.warp_size)
这是由于 M1000 缺少 TME 支持,在 triton 算子 tuning 过程中,num_stages > 1 的情况会回退到 num_stages
= 1 。
AOTI (Ahead-Of-Time Inductor)
AOTI 将模型预编译成独立的共享库(.so),实现完全脱离 Python 环境运行,适合生产环境部署和端侧推理场景。更多资料可参考
AOTI 官方文档