A new frontier in LLM speed
Inception’s breakthrough diffusion-based approach to language generation enables the world’s fastest, most efficient AI models with best-in-class quality.
Inception builds and deploys next‑generation large language models (LLMs) that are powered by diffusion rather than traditional auto‑regressive generation. By using diffusion, their models can produce many tokens in parallel, making them several times faster and less than half the cost of conventional LLMs. The diffusion framework also provides fine‑grained control over outputs, allowing adherence to specific schemas and semantic constraints. Additionally, it offers a unified paradigm for combining language with other data modalities such as audio, images, and video. The company’s team includes leading researchers and engineers from Stanford, UCLA, Cornell, Google DeepMind, Meta AI, Microsoft AI, and OpenAI, and they are currently deploying these diffusion LLMs at Fortune 500 companies.
Explain what Inception does
Here are some prompts you can try with a diffusion-style LLM:
- Explain a complex topic step by step, showing intermediate reasoning.
- Generate multiple variations of a product tagline and refine them progressively.
- Write a short story that improves its wording over several iterations.
- Brainstorm startup ideas and evolve the best one through revisions.
- Refactor a piece of code and show incremental improvements.
- Describe an image concept and refine the details in stages.
- Compare two technologies with increasingly deeper analysis.
- Draft a landing page headline and iterate toward a clearer version.
- Simulate a design critique that becomes more precise each step.
- Turn rough notes into a polished summary through gradual refinement.
Speed Benchmark
The diffusion difference. From sequential to parallel
All other LLMs generate text one token at a time. Mercury diffusion LLMs (dLLMs) generate tokens in parallel, increasing speed and maximizing GPU efficiency.
Parallel Generation
Mercury
zap
mango
crisp
lunar
wobble
spin
felt
droop
echo
Sequential Generation
ChatGPT
The
Quick
Brown
Fox
Jumps
Over
The
Lazy
Dog
Blazing-fast performance you can notice
- Write code
- Real-Time Voice
- Instant Agents
Meet our family of diffusion models
Mercury 2
The fastest reasoning LLM and the first reasoning dLLM. Ideal for complex applications where performance and speed are crucial.
- Input $0.25 per 1M tokens
- Output $0.75 per 1M tokens
Mercury Edit 2
A small, coding-focused dLLM. Ideal for code editing and other extremely latency-sensitive components of coding workflows.
- Input $0.25 per 1M tokens
- Output $0.75 per 1M tokens
Led by visionary AI researchers
Our founders pioneered diffusion modeling and invented cornerstone AI technologies.
Loved by leaders and innovators
We use Mercury 2 across Mission Control for generating titles, descriptions, and slugs, and previously even in our onboarding flow to create checks for new users. At the time, we simply needed the fastest model available - and Mercury delivered.
Enterprise-grade privacy and reliability
We’re available through major cloud providers like AWS Bedrock and Azure Foundry. Talk with us about fine-tuning and private deployments.
- Integrate in seconds
- Our models are OpenAI API compatible and a drop-in replacement for traditional LLMs.
- Get 99.5%+ uptime and priority support with custom SLAs.