Skip to content

Xiaomi MiMo Surges 15X Faster Than ChatGPT

Xiaomi’s MiMo Breaks AI Inference Speed Record

  • Xiaomi, with TileRT, achieved over 1,000 tokens per second on a trillion-parameter model using an 8-GPU commodity node.
  • The speed is due to FP4 quantization and DFlash speculative decoding, which enhances processing efficiency.
  • A limited API trial runs from June 9 to June 23, priced at three times the standard MiMo rates for approximately ten times the generation speed.

Xiaomi’s breakthrough in AI inference speed on standard hardware marks a significant shift in deploying high-speed models without custom chips. The innovative use of FP4 quantization and DFlash speculative decoding enables faster processing while maintaining quality.

This advancement allows for parallel reasoning paths essential for applications like fraud detection and trading signal generation that require low latency. (Source)

Share