# Build high-performance AI apps with Mercury

Inception’s diffusion LLMs (dLLMs) deliver frontier LLM quality at 5x greater speed.

## Overview

### (Un)paralled speeds

Our models run at 1000+ tokens per second on commercial NVIDIA GPUs, enabling instant, in-the-flow AI solutions.

### Exceptional Quality

We match the intelligence of speed-optimized autoregressive models like GPT-5 mini and Claude Haiku 4.5.

### Seamless integration

Our models are OpenAI compatible and a drop-in replacement for traditional LLMs.

## Discover our models

### Mercury 2

The fastest reasoning LLM and our most powerful model. Ideal for complex applications where performance and speed are crucial.

Pricing

#### Input

$0.25 / 1M Tokens

#### Cached Input

$0.025 / 1M Tokens

#### Output

$0.75 / 1M Tokens

#### Features

- 128K context window
- Reasoning
- Tool use
- Structured Output

#### Use cases

- Rapid Coding Iteration
- Workflow Subagents
- Customer Support
- Realtime Voice
- Enterprise Search

### Mercury Edit 2

A small, coding-focused dLLM. Ideal for code editing and other extremely latency-sensitive components of coding workflows.

Pricing

#### Input

$0.25 / 1M Tokens

#### Cached Input

$0.025 / 1M Tokens

#### Output

$0.75 / 1M Tokens

#### Features

- 32K context window

#### Use cases

- Autocomplete
- Next Edit

### Get started with Mercury today

1. Create your account
   Create an [Inception Platform](https://platform.inceptionlabs.ai/) account or sign in directly if you already have one.

2. Create your API Key
   Go to API Keys and create a new API key. New API keys comes with 10 million free tokens.

3. Make your first request
   We are OpenAI API compatible and are supported through libraries including AISuite, LiteLLM, and LangChain.

```python
import requests

response = requests.post(
    'https://api.inceptionlabs.ai/v1/chat/completions',
    headers={
        'Content-Type': 'application/json',
        'Authorization': 'Bearer INCEPTION_API_KEY'
    },
    json={
        'model': 'mercury-2',
        'messages': [
            {'role': 'user', 'content': 'What is a diffusion model?'}
        ],
        'max_tokens': 1000
    }
)

data = response.json()
```

## Choose the access plan that works best for your needs

### Free

Try our models.

- Access all models
- 10 million free tokens

### Developer

Scale our models.

- Usage-based pricing
- Generous rate limits
- Priority support

### Enterprise

Use Mercury in production.

- Custom rate limits
- SLA guarantees
- Security and privacy
- Volume-based pricing

### The future of LLMs is here

## Products

## Legal

## Contact
