Bringing up DeepSeek-V4-Flash on AMD MI300X
Summary
This article walks through the work needed to get DeepSeek-V4-Flash running on AMD MI300X for inference at scale. It highlights the software gaps on older AMD accelerators, especially FP8 dialect incompatibilities and incomplete tuned-kernel coverage in AITER. It also describes the fixes required in vLLM, including ROCm-specific fallbacks, capture-safe kernels, and correctness patches for MoE routing. After those changes, the deployment achieves a meaningful performance boost and shows that MI300X can be a practical lower-cost alternative to NVIDIA hardware. The piece closes by noting that the remaining gaps should narrow as AMD’s software stack improves and more of the work is upstreamed.
Classifications
industries
Entertainment
applications
Web and Content Management
AskAI Classifications
Labels
Developer Tools
AI Development Platforms
Semiconductor Software