upstream/main
└── #535 ROCm 10.0.0 Dockerfile and W7900D CI
└── #132 ROCm / RDNA runtime foundation
├── #133 TVM-FFI index/store HIP JIT
├── #134 CUDA-only backend isolation and ROCm routing
├── #135 PyTorch distributed to RCCL
└── #378 CPU/Hybrid MoE graph replay safety
└── #491 Hybrid decode orchestration refactor
#535 → #132 → { #133, #134, #378 } → { #491 }
AMD Roadmap - 2026 Q4 / 2027 Q1
Upstream Integration
What each PR provides
#535 — ROCm 10.0.0 Dockerfile and W7900D CI (
Open)#132 — ROCm / RDNA runtime foundation (
Open)#133 — TVM-FFI index/store HIP JIT (
Draft, depends on feat(rocm): add RDNA3 and RDNA4 runtime foundation #132)#134 — CUDA-only backend isolation and ROCm routing (
Draft, depends on feat(rocm): add RDNA3 and RDNA4 runtime foundation #132)#135 — RCCL communication path (
Draft, deferred)ncclbackend provided by RCCL; the two-GPU collective path is validated, while full model-level TP remains future work.#136CLOSED — Native GGUF Q4_0 ROCm kernels (depends on feat(rocm): add RDNA3 and RDNA4 runtime foundation #132 and fix(rocm): make TVM-FFI index and store JIT kernels portable to HIP #133)
gfx1201; broader architecture and format coverage is not yet claimed.#378 — CPU/Hybrid MoE graph replay safety (
Draft, depends on feat(rocm): add RDNA3 and RDNA4 runtime foundation #132)#491 — Hybrid decode orchestration refactor (
Draft, needs discussion; depends on fix(rocm): make CPU/Hybrid MoE graph replay safe #378)OffloadMoELayerintoHybridDecodeExecutor.Recommended merge order
Official ROCm Image
assigned to @LZ-QWQ
ROCm Compatibility & Official AMD install Doc
ROCm 7.14.x + PyTorch 2.11.x + AMD Triton 3.8.xas the first supported baseline.ROCm 10.0.x + PyTorch 2.11.x + AMD Triton 3.8.xas the next supported stack. ROCm 10 compatibility matrixCI and Qualification
assigned to @LZ-QWQ
Model Correctness
Performance Benchmarking
Native HIP Kernels
Hardware Support
2026 Q4
2027 Q1
Later Roadmap - 2027 Q2+