← Back to live feed · 1 stories across 1 day

Wednesday, Sep 23, 2026

1 story
1
Xiaomi Cuts MiMo-V3 Prefill Compute 5.02x With HySParse2 Architecture↩︎
topics 🤖 AI tags AIAI ModelsAI ResearchAI Releases keywords Xiaomi

The computational demands of agentic inference are at the center of a new design released by Xiaomi for its next-generation AI models. The HySparse2 architecture, which will serve as the core of MiMo-V3, achieves 5.02x lower prefill FLOPs and a 4.5x smaller KV cache at 1M tokens compared to the Hybrid SWA architecture used in MiMo-V2.6. This update also improves long-context retrieval as measured by MRCRv2 and RULER-v2 scores while lowering AgentPPL and LongPPL.

The design shifts to optimize for workloads where short actions return long observations that require frequent prefilling of a growing context. To reduce overhead, HySParse2 utilizes 'KV Bridging' to build cross-decoder K/V from self-decoder hidden states and 'KV Reuse' for sparse layers. Additional changes include moving to token-level selection and implementing a forced window of recent tokens to allow local and global tokens to share a single KV cache.

Image via @_luofuli on X
Continues from Tuesday, Sep 22
Xiaomi MiMo V2.6 Pro and Flash Rank No. 1 and No. 2 on Vals Index
49 tweets • 35 sources
See all 52 tweets →