JOURNAL
MiMo-V2.6-Pro Is MIT-Licensed and 1.02T Params — I Ran the Self-Hosting Math
Xiaomi's new open-weight model tops Terminal-Bench 2.1 at 89.9% and jumps from 19.0 to 71.9 on DeepSWE. Before you get excited about MIT licensing, here's what 1.02T total params vs. 42B active params actually costs you in hardware.

On this page
Xiaomi shipped MiMo-V2.6-Pro on September 22nd. By the numbers VentureBeat ran, it’s the top open-weights model in the world right now — ahead of DeepSeek, on Terminal-Bench 2.1 at 89.9%, which happens to edge out both Claude Opus 5 and GPT-5.6 Sol on that specific benchmark. DeepSWE jumped from 19.0 on the previous version to 71.9. That’s not a rounding error, that’s a different model class.
MIT license. On HuggingFace. 1.02 trillion total parameters, 42 billion active.
Here’s the sentence that should slow you down: “42B active” and “1.02T total” are not the same number, and the gap between them is exactly where your hardware budget lives or dies.
Active params tell you the speed. Total params tell you the bill.
MiMo-V2.6-Pro is a Mixture-of-Experts model with a frozen router — meaning the routing decision (which experts handle which token) is fixed rather than trained end-to-end with the rest of the network each run. For any given token, only 42B parameters actually do the multiplying. That’s the number that drives your inference latency and your FLOPs-per-token cost. It’s genuinely fast — comparable, compute-wise, to dense 40B-class models.
But inference doesn’t know in advance which expert a token will route to. So unless you’re doing something exotic with expert offloading, all 1.02T parameters need to be resident somewhere your accelerators can reach them fast — VRAM if you want real throughput, or NVMe/CPU RAM with an offloading scheme if you’re willing to eat the latency hit.
Rough math, back of the napkin:
fp16: 1.02T params × 2 bytes ≈ 2.04 TB just for weights
int8: 1.02T params × 1 byte ≈ 1.02 TB
int4: 1.02T params × 0.5 byte ≈ 510 GB
An H100 gives you 80GB of HBM. At int4, you’re still looking at roughly 7 H100s just to hold the weights, before you’ve allocated a single byte for KV cache, activations, or batch headroom. At fp16, you’re past 25. This is before multi-node networking overhead, before you’ve thought about how experts get sharded across devices, before anything about actual request latency under load.
Compare that to the Flash variant Xiaomi shipped alongside it — 310B total, 15B active. At int4 that’s roughly 155GB, which fits on two H100s with room to spare. If your use case doesn’t need the Pro tier’s ceiling, Flash is the one you should actually be pricing out.
The architecture choice that’s easy to miss
The hybrid-SWA backbone — sliding-window attention mixed with global attention layers — is doing real work here, and it’s the same tradeoff Mistral and a few others have leaned on: most tokens only need to see nearby context, so most layers don’t need full quadratic attention. You reserve the expensive global-attention layers for the ones that actually need long-range dependency tracking. It’s a sensible way to keep a long-context model from becoming computationally absurd, and it’s worth reading the layer-by-layer breakdown if you’re evaluating this for anything with long documents.
The other piece that doesn’t get enough attention: a 5-layer multi-token-prediction decoder for speculative decoding. MTP trains the model to predict several tokens ahead in one shot, then verifies them against the “real” autoregressive pass — when the guesses are right, you get multiple tokens for the cost of one forward pass. This is very likely a big chunk of why Xiaomi’s “UltraSpeed” variant claims 20x throughput at supposedly unchanged quality. Speculative decoding is not new, but a 5-layer MTP head baked into a model this size, released with weights you can actually inspect, is a genuinely useful artifact if you’re building your own inference stack and want a reference implementation that isn’t vaporware.
Omnimodal-native vs. bolted-on vision
MiMo claims to be natively omnimodal — text, image, video, audio — trained as one model rather than a language model with a vision adapter stapled on after the fact. I’d treat “natively omnimodal” claims from any lab, including this one, with some skepticism until independent evals confirm it holds up outside curated benchmarks. But architecturally, the bet is coherent: a model trained from the start to route all modalities through the same expert pool should, in theory, transfer concepts across modalities better than a frozen text backbone with a vision encoder bolted on top and fine-tuned separately. Whether that theoretical advantage shows up in your actual document-QA-with-screenshots pipeline is an empirical question you’ll have to test yourself, not something a benchmark table settles for you.
What I’d actually do with this
If you’re a team evaluating open-weights models for anything beyond “let’s see if it runs,” the question isn’t “is it better than Opus 5 on one benchmark.” It’s: can your infra fit the active-param compute budget, and can it fit the total-param memory budget, and do those two numbers tell wildly different stories about cost. With MiMo-V2.6-Pro they do — fast like a 42B model, expensive to host like a 1T model. That gap is the actual decision you’re making, not the Terminal-Bench score.



Discussion
Comments are reviewed before publication. Your email is kept private.