Project MonetRequest demo
Home/Blog/Mercury 2.5 Preview: API, Pricing, Speed & Features

AI · Project Monet Briefing

Mercury 2.5 Preview: Inception’s New Diffusion Reasoning Model

Mercury 2.5 Preview is Inception’s newest diffusion reasoning model, with hosted OpenRouter access, a large context window, tool use and structured output.

Published 2026-09-02 · Updated 2026-09-02 · By Project Monet Editorial Team

Mercury 2.5 Preview diffusion reasoning model with parallel token generation concept

01

What is Mercury 2.5 Preview?

Mercury 2.5 Preview is Inception’s newest reasoning diffusion LLM. Instead of producing text strictly one token at a time, Inception’s diffusion approach refines multiple token positions in parallel.

Inception currently describes Mercury 2.5 Preview as its most intelligent reasoning dLLM and links directly to OpenRouter for access. The model remains explicitly labeled Preview, so behavior and service details can change.

02

Availability, context and pricing

Inception’s current models page confirms Mercury 2.5 Preview is live on OpenRouter and lists standard rates of $0.20 per million input tokens, $0.75 per million output tokens and $0.02 per million cached-input tokens.

OpenRouter currently shows a temporary 80% Inception discount through September 8, 2026 at 07:00 UTC: $0.04/M input, $0.15/M output and $0.004/M cached input. Treat those discounted rates as promotional, not permanent.

There is a small first-party/provider presentation difference worth preserving: Inception labels the context window 256K, while OpenRouter currently displays 260K and up to 65,536 completion tokens. Applications should use the active provider’s documented limit rather than assuming the labels are interchangeable.

03

Reasoning, tools and structured output

Inception lists reasoning, tool use and structured output as Mercury 2.5 features. OpenRouter additionally documents tunable reasoning levels, parallel tool calls and JSON-schema structured output for its hosted route.

Those capabilities make the model relevant to agent loops, coding subagents, enterprise search, structured extraction and other workflows where repeated model latency compounds.

04

How to read the Mercury 2.5 speed claims

OpenRouter’s model description reports a peak claim of 1,107 tokens per second on standard GPUs and a 10+ point intelligence improvement over Mercury 2. These are provider/vendor claims, not independent benchmark results.

Live routed telemetry can be materially lower than a peak model claim. Evaluate time to first token, end-to-end task latency, tool accuracy, error rate and cost on the workload you actually plan to run.

05

OpenRouter access versus Inception’s direct API

Inception says its models are OpenAI API compatible, but its public direct-request example still uses mercury-2. Do not assume the OpenRouter slug is also the direct Inception model identifier until Inception documents that identifier explicitly.

06

Important limitations

  • Preview status means pricing, limits and behavior can change.
  • No official open-weight or local-runtime release was found in the current source set.
  • Peak speed and intelligence comparisons should remain attributed to Inception/OpenRouter.
  • Provider context labels currently differ slightly: 256K on Inception versus 260K on OpenRouter.

Mercury 2.5 is most interesting when low latency has practical value, not merely because a headline throughput number is large. Benchmark complete tasks before choosing it for production.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept