Xiaomi’s MiMo-V2.5 Pro AI Model Revolutionizes Multimodal Capabilities
- Xiaomi launched MiMo-V2.5 and V2.5-Pro, integrating text, image, audio, and video processing into a single model.
- MiMo-V2.5-Pro matches frontier models like Claude Opus 4.6 in coding benchmarks, with enhanced token efficiency.
- The Pro version autonomously completes complex tasks involving over a thousand tool calls at $1 input/$3 output per million tokens.
- MiMo-V2.5 offers faster performance at a lower cost of $0.40 input/$2 output per million tokens, supporting all modalities.
- Xiaomi plans to open source these models soon, marking aggressive iteration following strong adoption on OpenRouter.
Xiaomi’s release of MiMo-V2.5 and V2.5-Pro signifies a major advancement in multimodal AI capabilities by combining text, image, audio, and video processing in one model.
MiMo-V2.5-Pro stands out for its ability to perform complex software engineering tasks efficiently and cost-effectively compared to other frontier models like Claude Opus and GPT-5 series.Source