Apps & Tools
Mercury 2.5
A high-speed diffusion-based LLM running at over 1,100 tokens per second.
Playable
Mercury 2.5 media is blocked
Allow external media to connect to the provider and play this content.
Description
Mercury 2.5 is a production-ready diffusion large language model designed for extreme low-latency workloads like voice agents and search infrastructure. It achieves speeds of 1,107 tokens per second on NVIDIA GPUs while offering a 40% increase in intelligence over its predecessor, suitable for complex tasks like context compaction and reasoning.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.