vLLM for AMD (Windows)
High-performance vLLM optimization for AMD hardware on Windows
vLLM for AMD (Windows) media is blocked
Allow external media to connect to the provider and play this content.
Description
A custom build of vLLM optimized for AMD R9700 AI Pro hardware on Windows, featuring optimized matrix multiply kernels. The implementation utilizes a hand-written kernel for small workloads and a tuned Triton fallback with an increased tile width of 64, significantly improving throughput for output projection and fused gate/up operations. Benchmarks demonstrate the build reaching 169.1 tokens per second using the Qwen 3.8 27B model at concurrency level 32.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.