AMD GPU Operator + ROCm 7 Kubernetes 部署运维实践
完整介绍 AMD GPU Operator 1.5 在 Kubernetes 上的安装配置、ROCm 7 驱动管理、DRA 动态资源分配、自动节点修复(ANR)与监控指标接入,覆盖 MI300X 和 RDNA 4 工作站 GPU。
AMD ROCm 7 在 2026 年迎来重要里程碑:ROCm GPU Operator 1.5 正式支持 Dynamic Resource Allocation(DRA)、Auto Node Remediation(ANR)与 Node Problem Detector(NPD)集成,MI300X 系列在主流 LLM 推理框架上已接近 H100 95% 的吞吐量。对于国内数据中心中持有 AMD Instinct 卡或正在评估异构算力的运维团队,本文提供完整的落地操作指南。
AMD GPU Operator 架构概览
AMD GPU Operator 由以下核心组件构成:
| 组件 | 作用 |
|---|---|
| KMM(Kernel Module Management) | 自动构建并加载 amdgpu 内核模块 |
| Device Plugin | 将 GPU 注册为 amd.com/gpu 可调度资源 |
| Node Labeller | 自动标记节点的 GPU 型号、ROCm 版本等标签 |
| Metrics Exporter | 暴露 Prometheus 指标(温度、显存、利用率等) |
| DRA Driver | 支持 Kubernetes DRA API 动态分配 GPU |
| ANR(Auto Node Remediation) | 检测到 GPU 故障后自动重启节点恢复 |
| NPD(Node Problem Detector) | 将 GPU 相关故障上报为 NodeCondition |
| Test Runner | 部署后自动跑 GPU 健康测试 |
环境要求
- Kubernetes ≥ 1.29(DRA 特性需要 ≥ 1.31)
- ROCm 7.13+(amdgpu 驱动 6.19+)
- 支持的 GPU:AMD Instinct MI300X/MI325/MI350P、Radeon AI Pro R9700S(RDNA 4)
- 节点操作系统:Ubuntu 22.04/24.04、RHEL 9.x
安装 AMD GPU Operator
方式一:Helm 安装(推荐生产环境)
helm repo add amd-gpu-operator https://rocm.github.io/gpu-operator
helm repo update
# 查看可用版本
helm search repo amd-gpu-operator --versions | head -10
# 安装最新稳定版
helm upgrade --install amd-gpu-operator amd-gpu-operator/gpu-operator \
--namespace gpu-operator \
--create-namespace \
--set kmm.enabled=true \
--set devicePlugin.enabled=true \
--set nodeLabeller.enabled=true \
--set metricsExporter.enabled=true \
--set testRunner.enabled=true \
--version 1.5.0
方式二:OpenShift 安装
在 OpenShift 上使用 OperatorHub 搜索 “AMD GPU Operator” 安装,或通过 CLI:
oc apply -f https://raw.githubusercontent.com/ROCm/gpu-operator/main/bundle/manifests/gpu-operator.clusterserviceversion.yaml
验证安装
# 检查所有 Operator 组件运行正常
kubectl get pods -n gpu-operator
# 期望输出:所有 Pod 为 Running/Completed
# 查看 GPU Operator 版本
kubectl get deployment amd-gpu-operator-controller -n gpu-operator \
-o jsonpath='{.spec.template.spec.containers[0].image}'
# 验证 GPU 资源注册
kubectl describe node <gpu-node> | grep -A5 "amd.com"
# Capacity:
# amd.com/gpu: 8
# Allocatable:
# amd.com/gpu: 8
KMM 驱动管理
KMM(Kernel Module Management)负责在节点上自动构建和加载 amdgpu 内核模块,无需手动安装驱动。
查看驱动加载状态
# 检查 KMM 模块加载
kubectl get module -n gpu-operator
kubectl describe module amdgpu -n gpu-operator
# 在节点上确认驱动版本
ssh <node-ip> "modinfo amdgpu | grep -E 'version|srcversion'"
# 确认 ROCm 运行时
ssh <node-ip> "rocm-smi --version"
自定义驱动版本
# 锁定特定 ROCm 驱动版本
apiVersion: kmm.sigs.x-k8s.io/v1beta1
kind: Module
metadata:
name: amdgpu
namespace: gpu-operator
spec:
moduleLoader:
container:
modprobe:
moduleName: amdgpu
kernelMappings:
- regexp: '^.*'
containerImage: "rocm/amdgpu-dkms:6.19.0"
selector:
amd.com/gpu.present: "true"
Node Labeller 自动标签
安装完成后,Node Labeller 会自动给 GPU 节点打上详细标签:
kubectl get node <gpu-node> --show-labels | tr ',' '\n' | grep amd
# 输出示例:
# amd.com/gpu=true
# amd.com/gpu.present=true
# amd.com/gpu.product=AMD-Instinct-MI300X
# amd.com/gpu.count=8
# amd.com/gpu.family=AI
# amd.com/gpu.driver-version=6.19.0
# amd.com/gpu.rocm-version=7.1.3
# amd.com/gpu.vram=192576MiB
利用这些标签做精细调度:
# 只调度到 MI300X 节点
spec:
nodeSelector:
amd.com/gpu.product: "AMD-Instinct-MI300X"
DRA 动态资源分配(Kubernetes 1.31+)
GPU Operator 1.5 引入 DRA 驱动,作为传统 Device Plugin 的替代方案,支持更灵活的 GPU 分配语义。
启用 DRA 驱动
# 确认 Kubernetes 集群开启 DRA Feature Gate
kubectl get node -o json | jq '.items[0].status.conditions[] | select(.type == "Ready")'
# 安装时启用 DRA
helm upgrade --install amd-gpu-operator amd-gpu-operator/gpu-operator \
--namespace gpu-operator \
--set dra.enabled=true \
--set devicePlugin.enabled=false # DRA 与 Device Plugin 不能同时启用
DRA ResourceClaim 使用示例
apiVersion: resource.k8s.io/v1beta1
kind: ResourceClaimTemplate
metadata:
name: amd-gpu-claim
spec:
spec:
devices:
requests:
- name: gpu
deviceClassName: gpu.amd.com
count: 2
---
apiVersion: v1
kind: Pod
metadata:
name: rocm-workload
spec:
resourceClaims:
- name: gpus
resourceClaimTemplateName: amd-gpu-claim
containers:
- name: training
image: rocm/pytorch:rocm7.1_ubuntu22.04_py3.10_pytorch_release_2.4.0
resources:
claims:
- name: gpus
command: ["python3", "train.py"]
Auto Node Remediation(ANR)
ANR 是 AMD GPU Operator 1.5 的新特性,当检测到 GPU 节点异常(驱动崩溃、显存 ECC 错误超阈值、xGMI 链路故障)时自动重启节点,无需人工介入。
配置 ANR
# values.yaml 中配置 ANR
autoNodeRemediation:
enabled: true
# 节点重启前的等待时间(等待在途 Pod 迁移)
minHealthyTime: "30s"
# 重启方式:reboot 或 drain-and-reboot
remediationStrategy: drain-and-reboot
ANR 触发条件
ANR 与 NPD 联动,当 NPD 上报以下 NodeCondition 时触发修复:
| Condition | 含义 | 触发动作 |
|---|---|---|
AmdGpuDriverError |
驱动加载失败 | 节点重启 |
AmdGpuEccError |
ECC 不可纠正错误 | Drain + 重启 |
AmdGpuXgmiLinkError |
xGMI 互联链路故障 | Drain + 重启 |
AmdGpuHangDetected |
GPU 挂起/无响应 | 节点重启 |
# 查看节点 GPU 健康状态
kubectl describe node <gpu-node> | grep -A3 "AmdGpu"
Prometheus 监控集成
AMD GPU Metrics Exporter 暴露标准 Prometheus 指标:
# 查看 Metrics Exporter 端口
kubectl get svc -n gpu-operator | grep metrics
# 手动检查指标
kubectl port-forward -n gpu-operator svc/amd-gpu-metrics-exporter 5000:5000 &
curl http://localhost:5000/metrics | grep -E "gpu_temperature|gpu_utilization|gpu_memory" | head -20
核心指标列表:
| 指标 | 含义 |
|---|---|
amd_gpu_temperature_celsius |
GPU 结温 |
amd_gpu_utilization_percent |
GPU 计算利用率 |
amd_gpu_memory_used_bytes |
显存使用量 |
amd_gpu_power_watts |
功率 |
amd_gpu_ecc_corrected_total |
ECC 可纠正错误累计 |
amd_gpu_xgmi_link_bandwidth_bytes |
xGMI 互联带宽 |
Grafana 仪表盘
AMD 官方提供 Grafana 仪表盘 JSON,可直接导入:
# 下载官方仪表盘
wget https://raw.githubusercontent.com/ROCm/gpu-operator/main/grafana/amd-gpu-dashboard.json
# 通过 Grafana API 导入
curl -X POST http://admin:password@<grafana-ip>:3000/api/dashboards/import \
-H "Content-Type: application/json" \
-d @amd-gpu-dashboard.json
运行工作负载示例
vLLM on MI300X(ROCm 后端)
apiVersion: v1
kind: Pod
metadata:
name: vllm-mi300x
spec:
containers:
- name: vllm
image: rocm/vllm:v0.21.0-rocm7.1
resources:
limits:
amd.com/gpu: "2"
command:
- python3
- -m
- vllm.entrypoints.openai.api_server
- --model
- /models/llama4-scout-17b
- --tensor-parallel-size
- "2"
- --dtype
- float16
env:
- name: ROCR_VISIBLE_DEVICES
value: "0,1"
- name: HIP_VISIBLE_DEVICES
value: "0,1"
volumeMounts:
- name: models
mountPath: /models
volumes:
- name: models
hostPath:
path: /data/models
确认 ROCm 环境正常
# 进入容器验证
kubectl exec -it vllm-mi300x -- bash
# ROCm 环境检查
rocm-smi
hipcc --version
python3 -c "import torch; print(torch.cuda.is_available(), torch.version.hip)"
2026 年 AMD 认证项目
2026 年 7 月 24 日,AMD 正式推出 ROCm Certification Program,分 Associate(基础)和 Expert(进阶)两级。Expert 认证要求构建完整 Kubernetes GPU 环境、集成 AMD AI Studio、配置监控告警并运行端到端模型生命周期,与本文所覆盖内容高度对齐。
小结
AMD GPU Operator 1.5 已具备生产可用的完整运维能力:KMM 自动驱动管理免去手动安装 ROCm 的繁琐、DRA 提供更灵活的资源分配语义、ANR + NPD 实现 GPU 节点故障自愈。对于拥有 MI300X 硬件的团队,结合 vLLM 的 ROCm 后端,可以用接近 H100 的推理效率,以更低的成本运行 LLM 推理服务。
