Diffusion LLM speed: Inception's Mercury 2.5 hits 1,107 tokens a second
AnalysisMost chatbots write one word at a time. Mercury 2.5, released September 8 by Inception Labs, writes in parallel and clocks 1,107 tokens a second on ordinary Nvidia chips, several times the pace of comparable models. It is a diffusion LLM, a design that drafts a whole response at once and then refines it, rather than predicting the next token in sequence. Inception prices it at $0.20 per million input tokens and $0.75 per million output, with an 80 percent discount at launch. Speed at that price changes which tasks are worth handing to a model inside a live app.