Bringing up DeepSeek-V4-Flash on AMD MI300X

General News

Summary

This article walks through the work needed to get DeepSeek-V4-Flash running on AMD MI300X for inference at scale. It highlights the software gaps on older AMD accelerators, especially FP8 dialect incompatibilities and incomplete tuned-kernel coverage in AITER. It also describes the fixes required in vLLM, including ROCm-specific fallbacks, capture-safe kernels, and correctness patches for MoE routing. After those changes, the deployment achieves a meaningful performance boost and shows that MI300X can be a practical lower-cost alternative to NVIDIA hardware. The piece closes by noting that the remaining gaps should narrow as AMD’s software stack improves and more of the work is upstreamed.

Classifications

industries
Entertainment
applications
Web and Content Management

AskAI Classifications

Labels
Developer Tools AI Development Platforms Semiconductor Software

Linked Companies

AMD
$25M to $50M
Graphcore
$1M to $5M
Doubleword
$1M to $5M