Apps & Tools
Ling 3.0 Flash
A 124B hybrid linear MoE model optimized for ultra-fast local inference.
byAnt Ling
Playable
Ling 3.0 Flash media is blocked
Allow external media to connect to the provider and play this content.
Description
Ling-3.0-flash is a 124B parameter Mixture of Experts (MoE) model built on a native hybrid linear architecture (KDA+MLA). It features a highly efficient 5.1B active parameter count per token, allowing for fast decoding on hardware like the DGX Spark when running quantized GGUF versions.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.