Xiaomi
Xiaomi: MiMo-V2-Omni
MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step planning, tool use, and code execution - making it well-suited for complex real-world tasks that span modalities, 256K context window.
| Specifications | |
|---|---|
| Provider | Xiaomi |
| Input pricing | $0.40/M tokens |
| Output pricing | $2.00/M tokens |
| Context length | 262K tokens |
| Max output | 66K tokens |
| Input modalities | textaudioimagevideo |
| Output modalities | text |
| Tokenizer | Other |
| Content moderated | No |
Supported Parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningresponse_formatstoptemperaturetool_choicetoolstop_p