Inception
Inception: Mercury
Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury) here.
| Specifications | |
|---|---|
| Provider | Inception |
| Input pricing | $0.25/M tokens |
| Output pricing | $0.75/M tokens |
| Context length | 128K tokens |
| Max output | 32K tokens |
| Input modalities | text |
| Output modalities | text |
| Tokenizer | Other |
| Knowledge cutoff | 2025-01-31 |
| Content moderated | No |
Supported Parameters
max_tokensresponse_formatstopstructured_outputstemperaturetool_choicetools