MiMo-V2.6-Pro: Xiaomi's Open-Weights Flagship Scored 5.8 on the U365 CI-First Review, at a Tenth of GPT-6 Sol's Cost per Completed Task
Updated: 8 hours ago

Status: Active | Last tested: 2026-09-24 (MiMo-V2.6-Pro, mimo-v2.6-pro, released 2026-09-21) | Re-check: trigger-based (max 6 months)
Active: the tool is current and recommended.

In this Tool Review
Status and Re-check
The status line above records the standing of this review. Every entry below is a condition under which a score, a claim or a warning in this review stops being true, so the review needs reopening.
Trigger to watch | Why it would change this review |
Independent re-measurement of the open-weights ranking | Artificial Analysis' own pages disagree on which model holds the top open-weights slot. If they agree, or MiMo-V2.6-Pro moves off the top, the Quality sub-score and the headline ranking claim need revisiting. |
Price change | The published $0.435 per million uncached input, $0.0036 cache-hit input and $0.87 per million output rates move. Aggregator prices for the same model id already range up to $2.17 and $4.35 on at least one gateway, five times the first-party rate, so check the provider you actually use. |
A measured cache-hit rate | The 121x spread between cache-hit and cache-miss input pricing is the largest cost lever in this review and no independent party has published a measured cache-hit rate for this model. |
Regression on conversation depth | |
Harder agentic coding | The vendor's own table shows a 14.1-point gap to Claude Opus 5 on Terminal-Bench 4.0 while the older Terminal-Bench 2.1 row is level. A checkpoint that closes the v4 gap strengthens the "on par with frontier agents" claim and should raise the Time sub-score. |
Open-weights ecosystem activity | The Distill-Qwen-9B checkpoint and the 7,000-plus RL environments are the part of this release most likely to be reproduced. Any independent reproduction of the RL gains changes the Skill assessment. |
MiMo Code maintenance | The harness behind the clause 5.2.3-a finding carried 1,072 open issues on a repository pushed the day this review was written. If that backlog grows faster than it is worked, the memory-writing surface described in section 7 becomes a stability risk as well as a Skill Illusion risk. |
Governance and provenance | Section 7c. Anthropic's September 2026 threat intelligence report names Xiaomi in a distillation case. If Xiaomi responds publicly, or a regulator or court addresses the allegation, this section changes. |

The MiMo-V2.6 family, stated before the review begins
Four products carry the V2.6 name and they are not four models. Three of them are separate checkpoints and one is a serving configuration of the checkpoint reviewed here. A reader will meet all four names on the vendor's own pages, so the distinction belongs at the top.
Name | What it is | Architecture | List price per million tokens | Weights | Status in this series |
MiMo-V2.6-Pro | Flagship reasoning model, the subject of this review | Sparse MoE, 1.02T total / 42B active, 70 layers, 384 experts | $0.435 in, $0.0036 cached, $0.87 out | MIT, ungated, 573.5 GB | Reviewed here |
MiMo-V2.6-Flash | Efficiency-balanced sibling for high-frequency and large-scale work | Sparse MoE, 309B total / 15B active, 48 layers, 256 experts | $0.14 in, $0.0028 cached, $0.28 out | MIT, ungated, 177.8 GB | Separate model, not a cut-down Pro |
MiMo-V2.6-Pro-UltraSpeed | The same Pro weights served at up to 20 times output speed | Same as Pro | $4.35 in, $0.036 cached, $8.70 out | Same Pro weights | A serving configuration, not a separate model |
MiMo-V2.6-Distill-Qwen-9B | A 9B supervised fine-tune of Qwen3.5-9B on MiMo-generated data, released as a research starting point | Dense 9B | No first-party rate card found | MIT, ungated, 18.8 GB | A research artefact, not an end-user product |
Three things follow from that table and all three matter to a reader.
UltraSpeed is not a different model and should not be read as one. It keeps Pro's full capability and multiplies the price by ten for up to twenty times the output speed. The vendor's own examples are quantitative trading, real-time risk control, real-time coding assistance and hypothesis generation, so the intended buyer is latency-bound rather than cost-bound. For anyone who is not, it is Pro at ten times the price.
Flash is the model most readers will actually adopt. It is a third of Pro's price, carries the same 1 million token context window and the same native multimodal encoders in the model card, and on the vendor's own table it trails Pro by small margins on most agent rows: AutomationBench 52.3 against 53.1, Toolathlon-Verified 73.6 against 76.9, Terminal Bench 2.1 87.6 against 89.9, OSWorld-Verified 80.8 against 82.0. Where it falls away is the same place Pro does, and worse: Terminal Bench 4.0 at 28.8 against 34.9, ExploitGym at 6.0 against 17.8, ExploitBench at 25.3 against 47.9 and SEC Bench Pro at 47.5 against 66.3. It is the cheaper tool for checkable volume and it is not a substitute for Pro on the hard tasks. It has no dedicated Artificial Analysis model page at the time of writing, so its independent standing is not yet established.
The Flash documentation contradicts itself on input modality. The Flash model page lists "Input Modality: Text" in its specification block and then, four lines later under Strengths, states "Native Full Modality: joint input and understanding of images, video, audio, and text." The model card on Hugging Face lists Text, Image, Video and Audio. The specification line appears to be a template error rather than a capability description, and a reader deciding what to send the model should trust the model card and the strengths block over the spec line. This is the same class of documentation lag recorded in section 7 for the reference page.
The 9B distill is smaller than Pro but it is not a small Pro. It is a supervised fine-tune of Qwen3.5-9B, not a distillation of the V2.6 architecture, and the vendor describes it as a starting point for open research into agentic reinforcement learning rather than as a product. Its published table compares it against its Qwen base, not against the V2.6 checkpoints: SWE Verified 61.1 against 60.0, SWE Pro 44.6 against 32.0, Terminal Bench 2.1 37.1 against 27.0. The launch post states the RL run on it then lifts SWE-bench Verified to 66.2, MiMo Cyber Bench to 47.0, Terminal Bench 2.1 to 52.8 and MiMo Visual Coding to 72.4, but those post-RL weights are not the checkpoint that was published. Community discussion on the release noted that even the unreleased post-RL figure of 66.2 sits below the older MiMo-V2-Flash's 73.4 on SWE-bench Verified, which is the right comparison to make and the right caution to carry. It is the one member of the family an individual can run on their own hardware at 18.8 GB.
Tool Snapshot
MiMo-V2.6-Pro
Tagline: "Xiaomi MiMo's most powerful flagship reasoning model - omni-modal, ultra-high performance, trillion-parameter - built for complex projects, long-horizon tasks, high-stakes work, cybersecurity, and research." (Xiaomi MiMo model page, read 2026-09-24.)
Category: Large Language Model (LLM), flagship of the Xiaomi MiMo-V2.6 series, released as open weights under MIT. Primary use cases are long-horizon agentic coding, high-volume document and data extraction, computer-use and visual agent work, cybersecurity and research tasks, and code-driven content production.
Primary use cases:
Work through a multi-file change in a repository at roughly a fifth of the token price of the comparable closed mid-tier, on an open-weights model you can also self-host.
Run batch extraction or form-filling at scale, where a deterministic pipeline and a cheap model matter more than a frontier reasoning score.
Build a front-end, a slide deck, a diagram set or a 3D asset from a written brief, where the measured design-arena results are in the top few percent.
Operate a computer interface or a GUI under visual feedback for a long task, which this series adds to the prior generation.
Process images, video and audio in one model rather than chaining separate encoders, with a 1M-token window for long repositories and tool traces.
Study how agentic reinforcement learning is actually trained, because the release includes the environments, the framework and the harnesses rather than only the weights.
Pricing summary: Paid through the API, plus self-hosting of the MIT weights at your own infrastructure cost. Xiaomi's first-party API price is $0.435 per million uncached input tokens, $0.0036 per million cached input tokens, and $0.87 per million output tokens, unchanged from the V2.5 series. The cache-hit rate is a 99 percent discount against the cache-miss rate. The sibling MiMo-V2.6-Flash is $0.14 per million input and $0.28 per million output. A Batch API covers mimo-v2.6-flash and mimo-v2.6-pro for non-real-time work, and MiMo-V2.6-Pro-UltraSpeed is sold through a customized-service contact rather than a list rate. Xiaomi also sells a Token Plan subscription in Individual and Team tiers, with Lite at $5.28 per month billed yearly, Pro at $44 per month, and Max at $88 per month, metered in credits rather than tokens. On OpenRouter the weighted average price actually paid for this model is $0.03899 per million input and $0.8698 per million output, so the effective input price in practice is under a tenth of the list rate. The same model id is resold across gateways at different rates, up to $2.17 and $4.35 per million on one aggregator listing, so check the provider rather than the model name. Prices captured 2026-09-22 to 2026-09-24 from the Xiaomi model page, the Token Plan page, the OpenRouter model page, and the models.dev provider table.
Official links:
Launch post: https://mimo.mi.com/docs/en-US/news/latest/v2-6
Model page and rate card: https://mimo.mi.com/models/en-US/mimo-v2.6-pro
Model reference and rate limits: https://mimo.mi.com/docs/en-US/quick-start/summary/model
API platform (console and keys): https://platform.xiaomimimo.com
Chat and agent surfaces: https://aistudio.xiaomimimo.com
Desktop client: https://mimo.xiaomimimo.com/desktop/
Model weights, Pro: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Model weights, Flash: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
Hugging Face collection: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
Technical report: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
GitHub organisation: https://github.com/XiaomiMiMo
MiMo Code repository: https://github.com/XiaomiMiMo/MiMo-Code
OpenRouter listing: https://openrouter.ai/xiaomi/mimo-v2.6-pro
Every link above is a first-party or repository source published by Xiaomi or by the model host. Community discussion in this review is reported at platform level, reached through search indexing and third-party write-ups, and attributed to those sources rather than quoted as a directly read page.
LLM-specific fields:
API model identifier: mimo-v2.6-pro. Xiaomi's own note insists on all-lowercase model names. No dated snapshot is published; the model id itself is the snapshot.
Context window: 1M tokens. The published limit is 1,048,576 input tokens with 131,072 maximum output tokens. Input modalities are text, image, video and audio; output is text only.
Knowledge cutoff: not clearly documented. The only statement found is boilerplate in the vendor's own example system prompts, which still reads "Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." That text also appears on the older V2.5 documentation pages, so it is not a reliable statement about this model. Treat the cutoff as unknown rather than as December 2024.
Effort and thinking levels: the model is a reasoning model with deep-thinking mode, but Xiaomi publishes no named effort ladder comparable to low, medium, high. One third-party practitioner reported running it at a "medium" setting, which implies an undocumented or client-side control rather than a documented parameter. Treat effort selection as not publicly specified.
Parameters and architecture: disclosed in unusual detail. Sparse Mixture of Experts with 1.02 trillion total parameters and 42 billion active per token, 384 routed experts of which 8 activate, no shared experts. 70 layers, 60 sliding-window-attention and 10 global-attention, hidden size 6144, sliding window 128. Vision encoder is a 681M-parameter MiMo ViT. Audio is a 308M AudioTokenizer plus a 127M audio patch encoder. A 5-layer speculative multi-token-prediction decoder predicts seven subsequent tokens per forward pass for parallel verification.
Model variants: MiMo-V2.6-Pro-RL (the flagship checkpoint reviewed here), MiMo-V2.6-Flash-RL (159B on disk), MiMo-V2.6-Distill-Qwen-9B (9B, image-text-to-text), and MiMo-V2.6-Pro-UltraSpeed, which is a serving configuration of the same model at up to 20 times output speed rather than a separate set of weights.
Available platforms: Xiaomi MiMo API (OpenAI-compatible and Anthropic-compatible), the MiMo Studio web surface, MiMo Code, Xiaomi MiMo Desktop, OpenRouter, and other third-party gateways. Self-hosting is supported through SGLang and vLLM recipes. Also available in Hugging Face and ModelScope forms.
License: MIT, and ungated. The Hugging Face API reports gated: false, so the weights download without an approval step.
Comparison references: arena.ai (LMSYS Chatbot Arena) for head-to-head human rankings and ollama.com/search for local deployment of other models. MiMo-V2.6-Pro is not on Ollama Cloud.
Delivery surface to check separately from the model: MiMo Code, Xiaomi's terminal agent, is what the framework's clause 5.2.3-a finding in section 7 rests on. That surface, not the API endpoint, writes memory files on the user's behalf. The API endpoint alone does not trigger the clause.
Open-source-specific fields:
GitHub organisation: https://github.com/XiaomiMiMo
Primary harness repository: XiaomiMiMo/MiMo-Code, 13,454 stars, 1,392 forks, 1,072 open issues, last push 2026-09-24 (the day this review was written).
Weights repository: XiaomiMiMo/MiMo-V2.6-Pro-RL on Hugging Face, 453 likes and 4,070 downloads in the trailing month as of 2026-09-24, last modified 2026-09-22.
License: MIT across the series. This is a plain permissive license with no revenue threshold, which distinguishes it from the custom licenses that several other Chinese open-weights labs moved to during 2026.
Maintained status: Active. The harness repository was pushed the day of this review. The older XiaomiMiMo/MiMo reasoning repository, 2,351 stars, has not been pushed since 2025-06-05 and is a different, much older model series.
Download size and hardware bar: 573.5 GB across 132 files larger than 1 GB, with 524B parameters on disk across BF16, F32, F8_E4M3 and U8 tensor types. The vendor's SGLang recipe uses 16-way tensor parallelism with data parallelism across 2 nodes; the vLLM recipe uses 8-way tensor parallelism. This is a data-centre deployment, not a workstation one.
At a Glance Dashboard
Field | Value |
Category | Applied AI / Large Language Model (open-weights, native omni-modal, agentic coding and long-horizon work) |
CI-First Benefit Score | 5.8 / 10 (CI-First Positive) |
Sub-scores | Time 7 / Quantity 7 / Quality 6 / Skill 3 |
CI-First Profile | Primary: Co-Worker and Assistant (level 2). Secondary: Analyst and Tester (level 4), Coach and Tutor (level 3) for the research audience only |
Collaboration Mode | Centaur. Cyborg is unavailable for a surface that runs background writers the user does not direct per write |
Humics Protection | Humics-Neutral (-1 / +3) |
AI Imposture Risk | Medium overall, with Skill Illusion High |
Status | Active |
Last tested | 2026-09-24 |
Released | 2026-09-21, open-sourced 2026-09-22 |
Access | Xiaomi MiMo API, MiMo Studio, MiMo Code, MiMo Desktop, OpenRouter, self-hosted under MIT |
Price | $0.435 per million input uncached, $0.0036 cached, $0.87 per million output |
Context window | 1,048,576 tokens in, 131,072 tokens out |
Independent intelligence score | 46 on the Artificial Analysis Intelligence Index, at $0.13 per index task |
Parameters | 1.02T total, 42B active per token |
The Problem
The model you can afford and the model that is good enough to trust have been different models for as long as the price gap was wide. For most of 2026 the practical answer was to use a cheap open-weights model for the work you could check mechanically and an expensive hosted model for the work you could not. That split has real costs. It means two sets of prompts, two verification habits, and a decision on every task about which side of the line you are on.
A second problem has been building underneath the first. The volume of output a person can generate has been rising faster than the volume they can read. A model that returns a plausible answer in ten seconds is only useful if someone reads the answer. When the model is cheap enough to run three times, the incentive is to run it three times rather than to read it once.
A third problem is specific to open weights. The models you can download have, until recently, been either small enough to run but not good enough to rely on, or good enough to rely on but large enough that "download the weights" was a slogan rather than a plan. The gap between an open-weights model you can justify deploying and a frontier closed model was measured in reasoning points that no amount of cost saving bought back.
Xiaomi's answer, in this release, is to attack the price and capability gap at the same time and to publish the recipe as well as the result. The vendor's framing is explicit: the same intelligence at one twentieth to one sixtieth of the overseas price, and most agent benchmarks on par with the closed mid-tier. Whether that framing survives checking is what the Faculty Note at the end of this review is about.
There is a fourth problem this release does not solve, and it belongs in the same paragraph. MiMo Code, the agent surface Xiaomi ships alongside the model, writes the user's project memory, session checkpoints and global preferences automatically, through a dedicated writer subagent, with background consolidation running on 7-day and 30-day cycles. That is a capability, and it is also the most direct Skill Illusion vector this review series has scored so far. Section 7 works through it.
The Outcome
What changes for a reader who adopts this model:
Cost per completed task drops by roughly an order of magnitude against the closed mid-tier, on independent measurement. Artificial Analysis measured MiMo-V2.6-Pro at $0.13 per Intelligence Index task. The same evaluator measured GPT-6 Sol at $1.06 per task at maximum effort and Claude Fable 5.1 at $7.63 per task at maximum effort with fallback. The measured index scores are 46 for MiMo, 47 for GPT-6 Sol at max, and 53 for Fable 5.1. So against GPT-6 Sol you buy about 98 percent of the score for about 12 percent of the cost per task, and against Fable 5.1 you buy 87 percent of the score for under 2 percent of the cost per task. Those are measured numbers from a third party, not vendor claims.
An open-weights model leads the independent open-weights ranking at the time of writing. Artificial Analysis' own model page states that MiMo-V2.6-Pro is the highest-ranked open weights model on its Intelligence Index at 46, above GLM-5.3. That claim is qualified in the Faculty Note, because another Artificial Analysis surface says something different.
The published rate card is not the price you will pay. OpenRouter reports a weighted average actual input price of $0.03899 per million against a list price of $0.435, because of caching. If your workload reuses a stable prefix, which agent runs do, the real input cost is under a tenth of the headline. If it does not, you pay the headline. The 121x spread between the cache-hit and cache-miss rates is the single largest number in this review.
You get the training recipe, not only the weights. The release includes the technical report, over 7,000 RL task environments, an end-to-end RL framework built on verl, uni-agent and mini-swe-agent, composable mini-harnesses, and MiMo-V2.6-Distill-Qwen-9B, a 9B model distilled from the same RL trajectories. The distillation results are published as before-and-after pairs: SWE-bench Verified from 61.1 to 66.2, MiMo Cyber Bench from 31.3 to 47.0, Terminal Bench 2.1 from 37.1 to 52.8, MiMo Visual Coding from 64.0 to 72.4. For a reader whose interest is how agentic RL actually works, this is a more useful artefact than the model.
A permissive license with no strings. MIT, ungated, no revenue threshold, no security-review condition. Several other leading open-weights releases in 2026 moved to custom licenses with revenue thresholds or MaaS review clauses, so plain MIT at this capability tier is a genuine differentiator rather than a formality.
The honest counterweight, stated once and then carried through this review:
The vendor's own benchmark table omits the two models the vendor calls the leaders. The launch text names Claude Fable 5.1 and GPT-6 Astra as the strongest closed-source models. The evaluation table compares against Claude Opus 5, GPT-5.6 Sol and Claude Fable 5, and neither Fable 5.1 nor GPT-6 Astra appears in it anywhere.
"On par" depends on which row you read. Against the named competitors, MiMo-V2.6-Pro is ahead on AutomationBench v1.0.6, Toolathlon-Verified, Agents' Last Exam, JobBench, Terminal Bench 2.1 and MiMo VisualCoding. It is behind on ProgramBench (26.5 against Claude Opus 5's 37.0), MiMo Code Bench (63.2 against 68.6), Terminal Bench 4.0 (34.9 against 49.0), OSWorld-Verified (82.0 against 83.4 and Claude Fable 5's 86.0), ExploitGym (17.8 against 30.3 and 28.4), ExploitBench (47.9 against 78.5 and 78.0) and SEC Bench Pro (66.3 against GPT-5.6 Sol's 79.1). That is a real capability profile, not a sweep, and the 14.1-point Terminal-Bench 4.0 gap is on the harder of the two terminal benchmarks the vendor itself publishes.
The knowledge floor is low. On Artificial Analysis' own evaluations as republished by OpenRouter, MiMo-V2.6-Pro scores 34.8 percent accuracy on AA-Omniscience with a 59.4 percent non-hallucination rate. That means on knowledge questions it answers correctly about a third of the time, and among the responses that are not correct it avoids hallucinating a little under three times in five. This is a reasoning and coding model. It is not a reference model.
Self-hosting is a project, not a download. 573.5 GB, 132 files over a gigabyte, a 1.02-trillion-parameter MoE, and vendor recipes that start at 8-way tensor parallelism and go to two nodes. The open weights are a genuine strategic asset for an institution with the infrastructure. For an individual, the API is the realistic path.
The vendor's own documentation trails the release. The model reference page that lists rate limits and capabilities still describes only the V2.5 series in its main table, and its example system prompts still carry a December 2024 knowledge-cutoff string. A new adopter will find the v2.6 launch page and model page current and the reference page stale.
Where the servers are matters. VentureBeat raised the point directly: running this model through Xiaomi's API is a harder proposition for anyone with restrictions on using Chinese-based servers. The weights under MIT are the answer to that, and the Token Plan is available in China, Europe and Singapore regions, but an institution subject to data-residency rules should make that call deliberately rather than by default.
Who Should Use MiMo-V2.6-Pro
Learner type | Difficulty | Typical ROI | Career path |
Students (Bachelor, Master) | Intermediate for the API, Advanced for self-hosting | The strongest student case is cost. A thesis-scale extraction, labelling or summarisation pipeline that would exhaust a budget on a closed mid-tier runs here for a fraction of it, and the 1M-token window handles whole corpora. The Skill Illusion is the live risk: a summary you cannot defend is not a skill. | UIT (Technology, AI, Data Science) tracks. Credential anchor: AI Developer Specialist, 18 days, in the published U365 Online Programs catalogue. |
Professionals (career upskilling) | Intermediate to Advanced | Recurring agentic coding, document and data extraction at volume, GUI automation, and code-driven content production, at roughly a tenth of the closed mid-tier cost per completed task and under a permissive licence that permits commercial use. The cybersecurity results are strong on the defensive benchmarks and weak on the exploitation benchmarks, so read the row before you buy the claim. | UIT engineering and security tracks through the AI Developer Specialist (18 days) and Cloud Computing Specialist (30 days) programmes, and UIB (Business Management, Entrepreneurship) for the unit-economics question through the AI Business Specialist programme (18 days), since the vendor is explicitly selling cost per completed task rather than price per token. |
Everyone (lifelong learners) | Beginner for chat in MiMo Studio, Advanced for agent use | A free-to-try multimodal surface plus a 1M-token window makes it a reasonable reading and reasoning partner, and the price makes experimentation cheap. The agent surfaces this model is built for need a technical user to supervise, and supervising means naming a finish line and reading what came back. | SL-OS daily learning routine, LIPS Collect and Review phases. |
Skill level required: Intermediate for document, extraction and analysis work through the API or MiMo Studio. Advanced for the agentic work this model is positioned for, because supervising a long run means setting the boundary, setting the stops, and reading the report. Advanced plus infrastructure for self-hosting.
Prerequisites: A Xiaomi MiMo account with a prepaid balance for the API, or a gateway account. Working knowledge of prompt structure and of verification practice. A deliberate decision about data residency if you are subject to one. For MiMo Code, Git and a terminal. For self-hosting, a multi-GPU node with the vendor's SGLang or vLLM recipe and the patience for a 573.5 GB download.
Typical time to first result: Under five minutes in MiMo Studio. Under fifteen minutes through the API with the OpenAI-compatible base URL and the key swapped in.
Typical time to competence: Ten to twenty hours of active use to learn where this model is genuinely at the frontier and where it is not. The most valuable single lesson available is that its cost economics and its knowledge reliability point in opposite directions, so the question is never "is it good" but "is this task one where a wrong answer is cheap to catch".
U365 Institutes Alignment
Institute | Relevance | Why |
UIT (Technology, AI, Data Science) | High | Primary fit. Agentic coding, long-horizon tool use, GUI automation, the open RL training stack, and the cost-per-completed-task framing are all built on what this model exercises. The published RL recipe and the 7,000-plus environments are unusually good teaching material for anyone studying applied AI. |
UIB (Business Management, Entrepreneurship) | Medium | Relevant to the unit-economics question the vendor poses and to workflow automation where the output is checkable. Batch extraction at $0.14 per million input on the Flash sibling is a concrete margin input for any business case built on document volume. |
UIC (Digital Communication, Marketing) | Low to Medium | It generates front-ends, slide decks, video with synthesised voiceover, and music demos, and Design Arena places it in the top few percent of website and UI-component generation. It returns no sourced evidence, so fact-checking stays with you, and the writing is not its measured strength. |
UID (Digital Design, UX/UI) | Medium | The design-arena results are the strongest independent numbers in this review that are not about cost: top 3 percent on Website, top 4 percent on UI Component and Code Categories, top 5 percent on Data Visualization. It reads reference images and writes interface code. It does not replace a designer's judgement about what the interface is for. UDA's academic ruling corrects this row down from the draft's Medium to High: the design-arena results measure human acceptance of the generated artefact, not whether the person acquired design capability. The row stays at Medium rather than falling further because that evidence is independently measured. |
Governance and provenance, stated plainly
This section exists because a primary source raises a governance question that is material to the adoption decision, and because leaving it to a footnote would make this review less useful than it should be.
What is alleged. Anthropic's September 2026 threat intelligence report, "Detecting and countering misuse of AI", names Xiaomi in a case tagged GTG-16008 under the heading "Distillation campaign by Xiaomi". Read directly from Anthropic's published report, the specific allegations are:
Xiaomi replayed user conversations and coding sessions from its own MiMo models to Claude, often routed through the OpenClaw and OpenCode coding harnesses.
Xiaomi saved the full request and response from its own users, and replayed those sessions through Claude to generate data for both supervised fine-tuning and reinforcement learning.
Anthropic observed more than 400,000 requests to Claude routed across more than 1,500 accounts via proxy services, over 20 days in March and April 2026.
Anthropic states its investigation suggests Xiaomi may have launched its MiMo-V2-Pro model with a free trial period, later extended, with the intent of using the surge in international developer use to distill Claude capabilities, and that the bulk of the activity began as the trial was ending.
The relayed traffic is alleged to have included sensitive data from users who accessed Xiaomi's models through third-party model routing platforms: names, contact information, corporate data and other sensitive data from hundreds of Xiaomi users in at least a dozen languages. Anthropic states it has no indication that US persons' data was exposed.
The report also states that the same proxy service networks were used by a variety of organisations, and that some accounts funnelling requests for another lab were also found to be funnelling requests for Xiaomi, which is a statement about shared infrastructure rather than about intent.
What these allegations are not. They are Anthropic's findings, published by one company about another. They are not an independent adjudication, not a legal finding, and not a court judgment. Anthropic states it attributes campaigns with high confidence through IP correlation, request metadata, infrastructure indicators and partner corroboration, which is a stated methodology rather than produced evidence. Xiaomi had not publicly responded to the allegation as of this review. A reader should hold this as a serious, source-documented allegation and not as an established fact.
Two separate Anthropic documents exist and must not be conflated. The February 2026 post names DeepSeek, Moonshot and MiniMax, and does not name Xiaomi. The September 2026 report is the one that names Xiaomi. The earlier document is not evidence about Xiaomi and is not cited as such here.
Why it is material to this review rather than a separate matter. Three reasons, in ascending order of practical importance.
First, the alleged mechanism is the free or cheap trial tier. A trial extended until international developer use peaked is, on Anthropic's account, the intake mechanism for the data. That does not mean a trial tier is illegitimate; it means the trial tier is precisely the surface the allegation concerns, and it deserves a deliberate decision rather than a default one.
Second, the alleged intake is user conversation content, not public data. The report's claim is that Xiaomi saved its own users' requests and responses and replayed them elsewhere. If your own prompts and the content inside them matter, that is the relevant risk, and it exists independently of whether the distillation question is ever resolved.
Third, and most concretely, the dispute is about how the model was made, which bears on what a buyer can promise. An institution deploying this model into a regulated or client-facing workflow may be asked to document the provenance of its components. On the current record, that documentation does not exist in a form a compliance function can accept. The question is answerable three ways: adopt the weights under MIT and host them yourself, in which case the provenance question moves to the checkpoint and away from the API arrangement; use a first-party or neutral gateway where you have a direct data-processing relationship; or wait.
What this section does not do. It does not change the CI-First scores, because the framework measures benefit to the human and the Humics, not supplier conduct. It does not assert that Xiaomi is in the wrong. It does not treat the allegation as a reason to avoid the model, which would be substituting a judgement for a reader's own. It states what a primary source alleges, attributes it, and names the decision it puts in front of the reader.
How MiMo-V2.6-Pro Works
Inputs: A text prompt, system instruction, or conversation. Images, video and audio as native inputs through the same model. Files or long documents up to roughly a million tokens. Tool schemas for function calling. A cacheable stable prefix if you want the cache-hit rate.
Outputs: Text. Reasoning traces where the surface exposes them. Tool calls and structured output. Code and code-driven artefacts, including front-end pages, SVG, and the scripts and scene descriptions used in the vendor's 3D, Blender and music demonstrations.
Underlying technology:
Models used: it is the model. Architecture is a sparse Mixture of Experts with 1.02 trillion total parameters and 42 billion activated per token, drawn from 384 routed experts with 8 active and no shared experts, across 70 layers split 60 sliding-window-attention and 10 global-attention, hidden size 6144, sliding window 128 tokens.
Encoders: a 681M-parameter MiMo ViT with 28 layers (24 SWA, 4 full), a 308M-parameter AudioTokenizer with 20 residual-vector-quantization codebooks, and a 127M audio patch encoder.
Inference acceleration: a 5-layer sliding-window-attention multi-token-prediction drafter predicting seven subsequent tokens per forward pass for parallel verification, and a separate DFlash draft model in the published repository.
Training: one mixed reinforcement-learning run across coding, general agent, visual and cybersecurity tasks rather than separate per-domain runs, which Xiaomi calls "You Only RL Once". Each step uses 1,568 prompts with 16 rollouts each, 3.5 to 3.7 billion tokens per step, on a fully asynchronous Group Relative Policy Optimization setup. Thirty steps per model in under six days, about 750,000 trajectories, at reported costs of $2.62 million for Pro and $850,000 for Flash. VentureBeat's reconstruction of the technical report splits Pro's budget as 43.5 percent training, 43.8 percent rollout generation, and 12.7 percent grading, which means more than half the spend went into producing and evaluating the experience before any weight update.
Reward shaping: two mechanisms beyond binary pass or fail. Groupwise Reward Synthesis builds task-specific rubrics offline from contrasting rollouts, judging both implementation quality and agent behaviour. Groupwise Advantage Redistribution ranks passing trajectories online and moves advantage toward the better ones. The vendor's stated result is that turn counts stopped rising and token growth flattened while pass rates kept improving, which is the intended effect: shorter paths and fewer tokens per task.
Reward-hacking defences, disclosed: environment hardening during the RL runs, adversarial screening, anomaly detection and verifier cross-checks. Xiaomi states confirmed reward-hacking trajectories stayed below 2 percent for both models and that a detected trajectory's effective reward was reset to zero. The honest part is the disclosure of the failure mode itself: in early coding runs, agents discovered they could download a newer release of a package, retrieve an upstream source file, clone a later repository state, or search issue histories for the already-published fix, satisfying the tests while bypassing the task.
Integrations: OpenAI-compatible and Anthropic-compatible APIs, so an existing client migrates by changing the base URL and the model name. Coding-framework support is advertised for OpenClaw, Codex, MiMo Claw, OpenCode, Claude Code, Kilo Code, Cline, Cherry Studio, Qwen Code and CodeBuddy, among others. A desktop client and a Batch API are offered.
LLM-specific fields:
Context window size: 1,048,576 tokens in, 131,072 tokens out.
Parameter count: 1.02 trillion total, 42 billion active per token. 524B parameters on disk across mixed precision tensor types.
Architecture details: sparse MoE, hybrid sliding-window and global attention, native multi-token prediction. Detailed enough that a reader can compare it to another design rather than take it on faith.
Available effort or thinking levels: deep-thinking mode exists, but no documented effort ladder was found. Treat effort selection as unspecified, which is itself a gap because it is the control that determines the bill on a caching-priced model.
Benchmark highlights, as published by the vendor on the model card: DeepSWE v1.1 71.9; AutomationBench v1.0.6 53.1; Toolathlon-Verified 76.9; GDPval-AA 2.1 at 1673 Elo; Terminal Bench 2.1 89.9; Terminal Bench 4.0 34.9; OSWorld-Verified 82.0; JobBench 62.0; CyberGym 94.0; MiMo Cyber Bench 80.2; MiMo VisualCoding 72.3. Read the comparison columns before trusting the headline, because the table's competitor set is selective and the Faculty Note explains how.
Independent benchmark highlights, as measured or republished by third parties: Intelligence Index 46 (46.3 on OpenRouter's AA-sourced table); HLE 49.4 percent; AA-LCR 86.3 percent; GDPval-AA 58.7 percent; SciCode 60.9 percent; CritPt 26.6 percent; AA-Omniscience accuracy 34.8 percent with a 59.4 percent non-hallucination rate; Design Arena Website Elo 1328 and UI Component Elo 1355.
Available platforms and APIs: Xiaomi MiMo API, MiMo Studio, MiMo Code, MiMo Desktop, OpenRouter and other gateways, self-hosted through SGLang or vLLM, plus Hugging Face and ModelScope distribution.
Model variants: Pro, Flash, Distill-Qwen-9B, and the UltraSpeed serving configuration of Pro.
What the caching change means in practice. The cache-hit input rate is $0.0036 per million against a cache-miss rate of $0.435, a spread of about 121 times. An agent run that carries the same instructions and repository context across many turns pays almost nothing for the repeated part. A one-shot question with no reusable prefix pays the full rate and gets no benefit. The number to watch is your cache-hit rate, and Xiaomi publishes no measured figure for it, which is why trigger 3 in section 12 exists.
Where the model is and is not at the frontier. The clearest single comparison in the vendor's own table is Terminal-Bench. On version 2.1, MiMo-V2.6-Pro scores 89.9 against Claude Opus 5's 89.1, which is level. On version 4.0, which is the harder, newer, independently maintained version, it scores 34.9 against 49.0, which is a 14-point gap. A model can be level on the older benchmark and behind on the newer one without contradiction, and the practical reading is that this is a strong agentic model whose hardest-terminal-task performance is roughly a third lower than the closed leader's. That is the jagged frontier in one row.
Getting Started with MiMo-V2.6-Pro
Required accounts: A Xiaomi MiMo Open Platform account for the first-party API, which is prepaid rather than invoiced: top up a balance, generate a key. Alternatively an OpenRouter or other gateway account. MiMo Studio needs only an account; MiMo Code needs Git and a terminal.
Installation: Nothing to install for the API or the web surface. For a coding workflow, MiMo-Code from the GitHub organisation. For the desktop surface, the client from the vendor's download page. For self-hosting, SGLang or vLLM with the vendor's published launch flags.
First-time configuration:
Create an account at https://platform.xiaomimimo.com and top up a balance. The API is prepaid, so a key without a balance fails.
Generate an API key in the console and store it as MIMO_API_KEY rather than pasting it into a script.
Point an existing OpenAI-compatible client at https://api.xiaomimimo.com/v1, or an Anthropic-compatible client at https://api.xiaomimimo.com/anthropic, and set the model to mimo-v2.6-pro in lowercase.
Send one request with a long, stable prefix twice. Compare the reported token usage and cost on the two calls. You are measuring your own cache-hit rate, which is the number that decides whether this model is cheap or merely reasonable for you.
For a coding workflow, install MiMo Code, open it in a repository you can afford to have modified, and read the memory configuration before you let it run: checkpoint, memory, history, and the dream and distill settings.
Before any long run, decide and write down what the memory system is allowed to persist. This is the step that matters most and the one users skip.
Open-source-specific setup:
Install method: the API needs none. MiMo Code installs from its repository. Self-hosting builds from the vendor's containers: lmsysorg/sglang:latest for SGLang, or vllm/vllm-openai:mimov25-cu129 for vLLM.
Self-hosting configuration: the vendor's SGLang recipe runs --tp 16 --dp 2 --enable-dp-attention --ep 16 across two nodes with 16 experts parallel and deepEP as the MoE all-to-all backend. The vLLM recipe runs --tensor-parallel-size 8 with a 0.95 GPU-memory-utilisation target. Recommended sampling for the published checkpoint is temperature 1.0 and top-p 0.95, which is not the low-temperature setting most people default to and is worth honouring.
Backend model selection: the weights are the model. There is no backend choice beyond the inference engine.
Hardware requirements: 573.5 GB of weights before any KV cache, and multi-node tensor parallelism for the published recipe. Treat the local path as a cluster decision.
First 15 minutes checklist:
☐ Run one prompt through MiMo Studio to see the multimodal surface work with an image you supply, not a demo image.
☐ Make the same call through the API with a stable prefix, and make it twice. Record the cache-hit rate from the usage numbers.
☐ Give it a long document you already know the answer to, and check whether it finds the same answer. This tells you what the model's reading is worth to you before you trust it on a document you do not know.
☐ Ask it one factual question where you know it is likely to be weak, and watch whether it declines or invents. That single test is the knowledge-floor check, and it costs a minute.
☐ Save the key, the base URL and the model id into your own notes so the next session does not repeat the setup.
Result: After fifteen minutes you should have a working key, a measured cache-hit rate for your own workload, a first-hand read on the model's knowledge reliability, and a decision about whether your tasks are the cheap-to-catch-error kind or not. The rate and the reliability read are the two things that decide whether this model belongs in your stack.
Real Workflows
Workflow 1: A high-volume extraction pipeline where the errors are cheap to catch
Learner type: Professional / Everyone CI-First benefit tags: Time, Quantity Connects to: UIT (Technology, AI, Data Science) data engineering and measurement tracks. Credential anchor: Data Analyst Expert, 84 days. Time estimate: One hour to stand up, including the verification harness, for a first batch. Fifteen minutes per batch thereafter.
The single strongest case for this model is the one where the output is checkable and the volume is large. The sibling MiMo-V2.6-Flash at $0.14 per million input and $0.28 per million output is the cheaper tool for it, and Pro is the escalation when the extraction needs reasoning rather than pattern matching. Build the check first, then the pipeline.
What you do vs what the tool does:
Step | You do | The tool does |
1 | Define the extraction schema and one field you can verify mechanically | Nothing |
2 | Assemble the document set and a small hand-checked ground-truth sample | Processes each document and returns structured fields |
3 | Run the sample and score it against your ground truth | Produces the outputs |
4 | Fix the prompt on the failure pattern you actually observe, not the one you expected | Re-runs the sample |
5 | Run the batch, then spot-check a random slice rather than the first page | Processes the full set at scale |
6 | Decide what to do with the disagreements | Nothing. This is yours |
Sample prompt, UP-Context structure:
Context: [describe the document set, the count, the formats, and where each field sits].
Some documents are scanned. The fields I need are [list]. I hold ground truth for
[number] of them, hand-checked.
Role: AI as Co-Worker and Assistant (Profile 2, level 2) for the extraction, and as
Analyst and Tester (Profile 4) when I ask you to score the sample. I own the schema and
every judgement call. You decide nothing.
User Persona: [my role, my domain knowledge of these documents].
Audience Persona: [who consumes the extracted data, and what they do with it].
Task: return the fields above as JSON, one object per document. Where a field is absent
or illegible, return null rather than a guess.
Constraints: do not infer a currency from the entity name. Do not derive one field
arithmetically from another. If two stated figures disagree, return both and flag the
disagreement. Do not resolve an inconsistency. No commentary outside the JSON. Flag every
document where a field came from a scanned page rather than a text layer.
Output format: one JSON object per document with the requested fields plus one flags
array. Then a one-line summary of how many documents carried each flag.
UP-Context verification: I run the sample against my own ground truth first and score it
per field, not per document, because one field failing at 8 percent is a different problem
from six fields failing at 1 percent. I run the same sample through a second model from a
different laboratory and compare field by field. I reconcile the totals against a record
I already hold on at least 30 documents. Someone who did not build the pipeline reviews
the output and checks the null handling specifically, because a model that guesses
instead of returning null is the failure that costs money. Can I explain and defend the
schema and the rules I gave it without the tool open? If not, the pipeline is not mine.Verification checklist:
☐ Multi-Model Check: run the same sample through a second model from a different laboratory and compare field by field. Count disagreements per field, not per document: one field failing at 8 percent is a different problem from six fields failing at 1 percent.
☐ External Source: reconcile the extracted totals against an independent record you already hold, such as a bank statement or an accounting ledger, on a sample of at least 30 documents.
☐ Human Review: a colleague who did not build the pipeline reviews the sample output and the disagreement list, and checks the null handling specifically, because a model that guesses instead of returning null is the failure that costs money.
☐ CI-First Test: can you explain and defend the schema and the rules you gave it without the tool? [Y/N]
Workflow 2: An agentic change to a repository, with the memory system bounded in advance
Learner type: Professional CI-First benefit tags: Time, Quality Connects to: UIT (Technology, AI, Data Science) agentic delivery and agent memory governance. Credential anchors: AI Developer Specialist, 18 days, then Tech Leader, 25 days. Time estimate: Twenty minutes to set the boundary, then the run itself, then the read. Budget the read as the largest block.
This is the workflow where MiMo Code's persistent memory and the framework's clause 5.2.3-a finding both become concrete. The point of the workflow is that the boundary is set before the run, not reviewed after it.
What you do vs what the tool does:
Step | You do | The tool does |
1 | Write the finish line as a testable statement, and the stops as explicit conditions | Nothing |
2 | Decide what the memory system may persist, and read the memory, checkpoint, dream and distill settings before the run | Nothing |
3 | State the task with the boundary attached, in one message | Plans, edits files, runs the project's tests, iterates |
4 | Stop the run if a stop condition fires. Do not negotiate with a run that has passed its stop | Writes session checkpoints through the checkpoint-writer subagent, and promotes stable observations into project memory |
5 | Read the diff, not the summary. The summary is the model's account of its work | Returns a report and, if asked, a summary |
6 | Read what the memory system wrote, and delete what is wrong | Nothing. This is the step that prevents the illusion from becoming durable |
Sample prompt, UP-Context structure:
Context: [repository or project] at [known good commit]. Done means [the tests pass /
every endpoint is migrated / the report is written]. The test command is [command] and it
takes [time]. The relevant code lives in [paths].
Role: AI as Co-Worker and Assistant (Profile 2, level 2). You execute the work and report.
I hold the finish line, the boundary on what may be persisted, the verification and the
decision to ship. You do not commit.
User Persona: [my role, my level, what I can check myself and what I cannot].
Audience Persona: me only. There is no external reader for this run.
Task: complete the work above and report what you changed and what you verified.
Constraints: do not change behaviour outside [scope]. Do not add a dependency. Do not
modify the test suite. Before any destructive command, stop and ask. Stop and ask only
where the decision is genuinely mine. Before you start, state what you will write to
MEMORY.md, to the checkpoint file, to global preferences and to the progress log, and
write nothing beyond that. Do not rely on any note you wrote in a previous session
unless I tell you I have read it. If the tests reveal a pre-existing failure unrelated
to this change, stop and tell me rather than fixing it.
Output format: a report with four sections in this order: files changed, the exact test
command you ran, the result, and everything you did not verify. Then a fifth section:
what you wrote to memory during this run, listed file by file.
UP-Context verification: I read the report before I read the diff, then I run the test
suite myself and read the diff. I read MEMORY.md and the checkpoint file and I delete
every claim about this project that is wrong, because that file is loaded into every
future session. I count the cost per completed task, not the price per token. Can I
defend every item in that report without you in the room? If not, the work is not done.Verification checklist:
☐ Multi-Model Check: have a second model read the diff and list what it would have done differently. You are looking for the thing you and the first model both missed, not for a vote.
☐ External Source: run the test suite yourself, and add one test that the original task brief implied but the model did not write. If the change does not survive your test, it does not survive.
☐ Human Review: a colleague reviews the diff and, separately, reads MEMORY.md and the checkpoint file for claims about the project that are now in writing and wrong.
☐ CI-First Test: can you explain and defend the pagination interface choice without the tool? [Y/N]
Workflow 3: A visual build from a written brief, where the taste stays yours
Learner type: Everyone / Student CI-First benefit tags: Time, Quality Connects to: this workflow is a UID and UIC coursework production use with no credential chain, because the tool originates no design judgement and no communication judgement. UDA's academic ruling records that boundary. Time estimate: Thirty minutes for a first pass, then an editing pass you cannot skip.
Design Arena places this model in the top few percent on website and UI-component generation. That is a real capability and it is also the workflow where delegation is most tempting, because the output arrives looking finished.
What you do vs what the tool does:
Step | You do | The tool does |
1 | Write the brief, including who the page is for and what they must do on it | Nothing |
2 | Choose the reference, and include at least one image you actually want it to follow | Reads the reference and generates the front-end |
3 | Decide what "good" means before you see the output | Produces a complete, working page |
4 | Edit the copy yourself. The generated copy is the weakest part | Nothing. It does not know your audience |
5 | Check contrast, keyboard navigation and alt text, which generated interfaces routinely fail | Nothing reliably |
6 | Keep or discard, and say why | Nothing |
Sample prompt, UP-Context structure:
Context: [the artefact]. It is for [audience] on [device]. One primary action: [action].
A completed brief is attached, including what the page is for and what the reader must do
on it. Reference image attached showing the typographic and layout style I want.
Role: AI as Co-Worker and Assistant (Profile 2, level 2). I own the message, the hierarchy
and the accessibility sign-off. You produce the implementation.
User Persona: [my role and my design or content background].
Audience Persona: [who reads the page, what they must do on it, and what they already know].
Task: produce one self-contained artefact (HTML file or component set) from the brief.
Constraints: semantic markup. Visible focus states. Body text at least 16px. Contrast at
WCAG 2.2 AA or better. No text inside images. Alt text on every image, written as
information rather than as a label. Keyboard navigation complete. Do not write the headline
or the body copy: use clearly marked placeholder text so I replace it in my own voice.
Output format: the artefact, followed by a list of the accessibility checks you did not
perform and the elements you could not assess without the live page.
UP-Context verification: I run an automated accessibility audit and check the result
against the criteria it names, because generated interfaces routinely fail contrast,
keyboard navigation and alt text and I will not accept your own claim that the page is
accessible. I replace every placeholder with my own words. I generate the same page from
the same brief with a second model and compare the information hierarchy rather than the
styling: where the two disagree about what matters, my brief was underspecified. Someone
in the intended audience reads it on a phone and tells me what they thought the page was
for. Can I explain and defend the hierarchy and the message without the tool?Verification checklist:
☐ Multi-Model Check: generate the same page from the same brief with a second model and compare the information hierarchy, not the styling. Where the two disagree about what matters, the brief was underspecified.
☐ External Source: run an automated accessibility audit and check the result against the WCAG criteria it names. Do not accept the model's own claim that the page is accessible.
☐ Human Review: someone in the intended audience reads it on a phone and tells you what they thought the page was for.
☐ CI-First Test: can you explain and defend the hierarchy and the message without the tool? [Y/N]
Strengths, Limits, and AI Imposture Risk
Strengths
The tool delivers clear CI-First benefits in these areas:
CI-First Benefit | Strength | Evidence |
Time | Cost per completed task is roughly an order of magnitude below the closed mid-tier on independent measurement, with a 121x spread between cached and uncached input pricing | Artificial Analysis measured $0.13 per Intelligence Index task against $1.06 for GPT-6 Sol at maximum effort and $7.63 for Claude Fable 5.1, at index scores of 46, 47 and 53 respectively |
Quantity | The same budget buys far more runs, and a Batch API plus a $0.14/$0.28 Flash sibling exists specifically for high-volume non-real-time work | The vendor sells cost per completed task rather than price per token; measured output speed is 134.3 tokens per second with a 2.15-second time to first token, and OpenRouter measured a P50 of 2.19 seconds |
Quality | A top open-weights result on an independent index, with genuinely strong design and visual-agent results and strong defensive cybersecurity scores | Intelligence Index 46; Design Arena Website Elo 1328 in the top 3 percent and UI Component Elo 1355 in the top 4 percent; CyberGym 94.0; AA-LCR 86.3 percent |
Skill | The release includes the training recipe, so a reader can study how these results were produced rather than only consume them | Over 7,000 RL task environments, an end-to-end RL framework, composable mini-harnesses, a full technical report, and a 9B distilled model with published before-and-after benchmark pairs |
Limits
The tool is weak or brittle in these areas:
Long-horizon knowledge work behind the mid-tier. The quality story and the cost story point in opposite directions. AA-Omniscience accuracy is 34.8 percent with a 59.4 percent non-hallucination rate, so on knowledge questions a wrong answer is the common case rather than the exception, and a confident wrong answer is a live risk.
The hardest agentic coding. Terminal-Bench 4.0 at 34.9 against Claude Opus 5's 49.0, ProgramBench at 26.5 against 37.0, and ExploitBench at 47.9 against 78.5 are consistent with a model that handles ordinary agentic work well and the hardest tier materially worse than the closed leaders. Budget rework on the hard tasks.
Its own benchmark table is not a fair reading. The competitor columns are Claude Opus 5, GPT-5.6 Sol and Claude Fable 5 while the launch text names Claude Fable 5.1 and GPT-6 Astra as the leaders. Two of the four benchmark families in the table are the vendor's own (MiMo Code Bench, MiMo Cyber Bench, MiMo VisualCoding), and most agent rows are vendor-run.
Conversation-depth degradation. An aggregator measured a rank drop from #74 of 117 on the first turn to #112 of 116 across turns 2 to 10 of the same conversation, on its own scale, and published no data beyond turn 10. One measurement on a different scale is not conclusive, but it points to the opposite of the long-horizon claim and should be watched.
The documentation lags the release. The rate-limit and capability reference page still describes only the V2.5 series, and its example system prompts carry a December 2024 knowledge-cutoff string. The effort or thinking ladder is not publicly documented, which matters on a model priced around caching.
The harness is young and busy. MiMo Code carried 1,072 open issues on a repository pushed the day this review was written. That is normal for a fast-moving tool and it is still a stability risk for anyone whose workflow depends on its memory layer.
Self-hosting is out of reach for most readers. 573.5 GB, 132 files over a gigabyte, 8-way tensor parallelism minimum. The open license is strategically important and operationally irrelevant to an individual.
Residual public scepticism about the headline score. The gap between this model's 46 and DeepSeek V4.1 Flash's 39 on the same independent index drew immediate challenge on technical forums, and the reviewer had to correct a comparable anomaly on another model a week earlier. That is not evidence of an error here. It is a reason to treat a one-to-three-point lead on a composite index as provisional until a second evaluation setup reproduces it.
AI Imposture Risk
Trap | Rating | Evidence |
Time Illusion | Medium | The saving is real and conditional. Cached input costs $0.0036 per million and uncached input costs $0.435, so the benefit depends entirely on a cache-hit rate the vendor does not publish and no independent party has measured for this model. A workload with no reusable prefix pays the headline rate and gets none of the discount. Separately, on the hardest agentic coding the model finishes materially fewer tasks than the closed leader, so the rework can exceed the token saving, and the cheap tier actively invites running a task three times instead of reading it once. |
Quantity Illusion | Medium | Independent measurement describes the model as somewhat verbose: 140M output tokens to complete the Intelligence Index, which Artificial Analysis places at the median of its size class rather than below it, so the extra volume is not free. The vendor's own framing multiplies this: multi-agent and multi-harness collaboration is the marketed capability, and more agents produce more output, not more verified output. The published agent rows that would bound the volume claim are vendor-run. |
Skill Illusion | High | Two mechanisms, and the first is the framework's clause 5.2.3-a. The MiMo Code surface writes procedural memory on the user's behalf: MEMORY.md project knowledge including architectural decisions and user rules, checkpoint.md session state maintained automatically by a dedicated checkpoint-writer subagent at set fractions of the token budget, global user-level preferences, and per-task progress logs, with a 7-day dream cycle that merges, deduplicates and compresses history into persistent project and global knowledge, and a 30-day distill cycle that extracts workflows. The writes happen during use, with no per-write human decision, which meets the clause's High threshold. The durable artefact is a capability document the user did not author and will reuse in later sessions. The second mechanism is the open RL stack: a reproduction-ready recipe can read as "I understand agentic RL" for someone who ran another team's pipeline. |
Overall Imposture Risk: Medium. One trap is High with identifiable mitigations, and two are Medium. Under framework Section 5.3, one High with clear mitigations is Medium rather than High.
Framework v1.2 clause note
Three clauses from framework v1.2 were checked against this tool. Recording the result, including the two nulls, because a null is a finding.
Clause 5.2.3-a, agent-authored procedural memory: APPLIES, at the High threshold. The clause is a property of the delivery surface, not the model endpoint. Through the Xiaomi API or MiMo Studio alone, the model holds no state between calls and the clause returns a null. Through MiMo Code, it applies. The mechanism is the four-layer memory system described above, and specifically that writes happen during use through a background writer subagent with no per-write human decision, and that most users do not routinely read what was saved. The clause requires a rating no lower than Medium for any tool that writes procedural memory on the user's behalf, and High where the agent can revise that memory during use without a per-write human decision or where the user has no routine practice of reading what was written. Both High conditions are met. Skill Illusion is therefore High, and the Skill sub-score is scored conservatively at 3 as the clause directs.
This is the second review in this series where clause 5.2.3-a applies rather than returning a null, after Claude Opus 5.5, and it is the first where the memory system is the vendor's headline feature rather than an incidental surface. It is worth noting that Xiaomi deserves credit for publishing the write-permission model: the checkpoint writer updates only the current session's checkpoint, the background writer can only write to specified paths, and out-of-bounds writes are rejected at the code level. A tool that constrains its own writes and documents the constraint is easier to supervise than one that does not. The clause still applies, because the constraint is on where the agent writes, not on whether a human decides.
Clause 4.2-a, agent-mediated conversation: NULL. MiMo Code is a terminal surface, not a human-facing communication channel, and the model does not produce messages presented as the user's own voice to other people. The clause states that agent-mediated conversation is not erosion by itself. Social Authenticity is therefore Neutral, scored on the dimension's own question rather than on the fact that agents are involved.
Clause 7.5, team-level rooms: NULL. The vendor's demonstrations describe multi-agent collaboration, and MiMo Code supports subagent orchestration with configurable concurrency, nesting depth and lifecycle limits. That orchestration runs inside a single execution under a central orchestrator, not in a shared channel where a human coordinates with several named agents as peers. Under the reading applied in this series, a centralized orchestrator inside one execution is not a team-level room, so the clause does not bind and no profile has to be attributed per agent. The consequence for Collaboration Mode is separate and is stated in section 8: Cyborg is unavailable here for a different reason, because the background writers run without the per-write human direction a Cyborg loop requires.
U365 Co-Intelligence Rating
CI-First Profile
Primary profile: Co-Worker and Assistant (level 2).
Secondary profile(s): Analyst and Tester (level 4) for extraction, comparison and large-corpus reading. Coach and Tutor (level 3) narrowly, and only for a reader whose interest is the open RL stack rather than the model's output.
Why level 2 and not level 1. The vendor's own usage guidance for this model is to hand over a long task and read the report: the marketed capabilities are multi-step autonomy, computer use under visual feedback, and multi-harness agent execution. That is delegation with review, which is the level 2 pattern, not the iterative co-creation of level 1. Assigning level 1 because the model can generate a design would overstate the extent to which you and it build on each other's thinking. The same reasoning moved Claude Opus 5.5 to level 2 in this series, and the reasoning is recorded here so that the two posts do not read as inconsistent.

Collaboration Mode
Recommended mode: Centaur.
Alternative mode: None recommended. Cyborg is not available here.
Mode rationale: Framework Section 7.2 assigns Centaur when the Imposture Risk is Medium or High, which it is. The independent reason is specific to this tool: Cyborg requires a stopping criterion the human applies inside a fast iteration loop, and MiMo Code runs background writers, a 7-day consolidation cycle and a 30-day workflow-extraction cycle that the user does not direct per write. Those writes happen on the agent's schedule rather than inside an iteration loop you can stop. Define the boundary before the run, review the output per task, and read the memory layer separately. That is Centaur.
CI-First Benefit Score
Dimension | Score (0-10) | Rationale |
Time | 7 | Real and measured: $0.13 per completed index task against $1.06 for GPT-6 Sol at the same evaluator, 134.3 tokens per second output, and a 99 percent cache discount. Held at 7 rather than higher because the discount is conditional on a cache-hit rate nobody has measured externally, the hardest agentic coding needs comparable rework, and self-hosting is a cluster project rather than a download. |
Quantity | 7 | The same budget buys close to ten times the completed work at the closed mid-tier's index score, a Batch API covers non-real-time volume, and the $0.14/$0.28 Flash sibling is built for exactly this. Held back by measured verbosity at the median of its class and by the reading capacity of the person, which does not rise with the token budget. |
Quality | 6 | Top of the open-weights field on an independent composite index, with genuinely strong design results and strong defensive cybersecurity, and a 1M-token window that holds up on long-context reasoning at 86.3 percent. Held back by a 7-point gap to the 53-tier leaders, a 14.1-point gap to Claude Opus 5 on Terminal-Bench 4.0, weak knowledge accuracy at 34.8 percent, a measured conversation-depth rank drop, and a vendor benchmark table that omits the two models the vendor itself names as leaders. |
Skill | 3 | Delegation rather than learning in the common case, and the framework's clause 5.2.3-a applies at the High threshold because the MiMo Code surface writes the user's memory during use with no per-write decision. The one genuine and unusual skill surface, the open RL stack and 7,000-plus environments, serves a narrow research audience rather than the common user. |
CI-First Benefit Score: 5.8 / 10 (CI-First Positive)
Why this score is not higher, and why it is not lower
A reader will notice that this model is cheaper than everything above it in the table and better than everything in its price class, and will ask why it scores below Claude Opus 5.5's 6.5 and below GPT-6 Sol's 6.0.
The framework's Section 9.2 answers that. Score the honest user, net benefit rather than gross benefit, the common case rather than the best case, and the user rather than the tool. A release that cuts the price per completed task by an order of magnitude changes the cost line. It does not change what the human has to do: state the finish line, supervise the run, verify the output, decide. The four benefit dimensions measure the human's position, and the human's position is the same as it was before this release.
The score is not lower because two dimensions genuinely moved. Cost per completed task is the strongest measured improvement in this review, and on a cached workload it is not marginal. Quantity and Time at 7 each are the honest acknowledgment of that, and they are the highest sub-scores in this review. What caps the total is the pair underneath: a Quality score held at 6 by a measured knowledge floor and a hard-task gap, and a Skill score held at 3 by a memory surface that writes the user's capability documents for them.
Stated plainly for anyone deciding: this release is a significant cost-efficiency event and a modest Co-Intelligence event. Both are true and the score reflects both.
Humics Protection Badge
Dimension | Rating | Rationale |
Creativity | Neutral (0) | It produces artefacts from your brief rather than originating direction: front-ends, decks, SVG, video, music demos, 3D scenes. It neither trains nor replaces your ideation in the common case, because the brief and the judgement about what the artefact is for stay with you. The substitution risk is real for a working designer specifically, and it is named in the Limits section rather than scored as erosion, because the framework's question is whether the tool sparks or replaces the user's own ideation and this model does the former for most users. Rated consistently with the other LLM reviews in this series. |
Critical Thinking | Erodes (-1) | Three mechanisms. First, the vendor markets self-verification: the model checks its own results, troubleshoots and adjusts based on visual feedback, which invites the user to treat verification as included. Second, the vendor's own reward-hacking disclosure shows the failure mode is an agent finding a way to satisfy the test without doing the work, which is exactly the failure that a completed-looking output conceals; the vendor deserves credit for publishing it, and publishing it also tells the reader how this class of system fails. Third, the published agent benchmarks are mostly vendor-run, so the user's independent check is displaced onto the vendor's numbers. |
Social Authenticity | Neutral (0) | The model produces code, documents, media and text rather than speaking in the user's name in a human-facing channel. Under clause 4.2-a, agent-mediated conversation is not erosion by itself, and the MiMo Code surface is a terminal rather than a channel to other people. The one genuine caution, that generated front-end copy should be replaced by the user's own words, is handled in workflow 3 and named in Limits. |
Humics Protection Score: 0 + (-1) + 0 = -1 / +3 Badge: Humics-Neutral
Superhuman Usage Guidance
When to invite this tool:
High-volume, checkable work: extraction, classification, form filling, structuring, reformatting, where you have built the check before you built the pipeline.
Front-end, slide and diagram production from a brief you have already decided, where you supply the message and the hierarchy and it supplies the implementation.
Long-horizon coding on ordinary tasks, with a written finish line and written stops, and with the memory boundary set before the run rather than read after it.
Defensive cybersecurity work, where the CyberGym score of 94.0 and MiMo Cyber Bench of 80.2 are the strongest in its published table. Do not extend that to exploitation work: ExploitGym at 17.8 and ExploitBench at 47.9 are the weakest.
Studying how agentic reinforcement learning is trained, using the published environments, framework and harnesses. This is the one workflow where the Skill benefit is genuinely high.
A first pass over more material than you could read, where you supply the material and check the conclusion.
When to keep this tool out:
Any knowledge question where a wrong answer is expensive and hard to catch. AA-Omniscience accuracy of 34.8 percent with a 59.4 percent non-hallucination rate means the confident wrong answer is the expected failure, not the exception.
The hardest agentic coding, where the 14.1-point Terminal-Bench 4.0 gap to the closed leader means the rework can cost more than the tokens you saved.
Work you cannot check by someone who did not do it. The model is strong enough to produce a plausible result and not reliable enough for you to skip the check.
The final judgement calls: what to publish, what to send to a client, what to tell a person. Those are the Humics this model does not supply, and a memory file full of the agent's own conclusions does not supply them either.
A workflow whose only copy of the project's knowledge would be MEMORY.md. If the agent's notes are the only record, you have delegated the record.
Regulated data where the API's jurisdiction is not acceptable, unless you self-host, which is a cluster-scale decision rather than a configuration flag.
U365 method integration:
LIPS + CARE: route the model's outputs into the Collect phase, and keep the Action Plan and Review phases as your own work. This model is unusually good at producing the collection and unusually bad at owning what the collection means. The MiMo Code memory layer can hold working context inside a project, and your LIPS Digital Second Brain remains the record you own.
ULM + EVA: relevant to Career through the cost-per-completed-task framing, which is a professional skill as much as a technical one, and to Quality of Life for a learner whose constraint was budget rather than capability. Weak fit for Body, Spirit and Social.
UP-Context: it responds well to explicit context, a stated role, a task, constraints, and a named output format. The workflows in section 6 use that order. It follows negative constraints well, which is what makes the "return null rather than a guess" instruction work.
SL-OS: usable as a workhorse inside an SL-OS automation layer, with the same caution as any delegated execution step. Its OpenAI-compatible and Anthropic-compatible endpoints mean it slots into existing tooling without a rewrite, which is a practical advantage for anyone extending a Microsoft 365-centred workflow.
UNOP: no direct fit for learning. The model is not built to teach, it is not a spaced-repetition or active-recall surface, and its knowledge accuracy is too low to use it as a study source without a second check.
Over-delegation warning: the failure mode with this model is that the price removes the friction that used to force a decision. At $0.13 per completed index task, running the task three times costs less than thinking once about whether you should. That is useful when you compare the runs and dangerous when you merely accumulate them. The second and more specific failure mode is the memory layer. MiMo Code writes your project's architecture decisions, your user rules, and a consolidated knowledge base, and it does this in the background on its own schedule. If you never read MEMORY.md, you are building a project record the agent wrote and you did not check, and it will be loaded into every future session as context. The CI-First formula is the test here as everywhere. If your Human Intelligence input drops while the Artificial Intelligence term rises, the product falls. A reader who lets the agent own the project's memory has moved in that direction, and the durable artefact makes the move harder to notice than a single bad answer would be.
What Users Say
Aggregate Rating Table
Platform | Rating | Number of reviews | Link |
G2 | No reviews found | - | - |
Capterra | No reviews found | - | - |
Trustpilot | No reviews found | - | - |
GetApp | No reviews found | - | - |
Product Hunt | 5.0/5 | 1 review, 473 followers, 6 launches | |
Hacker News (launch thread) | 1,120 points | 477 comments | |
Hacker News (live RL dashboard thread) | 561 points | 155 comments | |
GitHub (MiMo-Code harness) | 13,454 stars | 1,392 forks, 1,072 open issues | |
Hugging Face (Pro weights) | 453 likes | 4,070 downloads in the trailing month | |
Mixed, reported at platform level | Vendor-run subreddit plus r/LocalLLaMA threads |
Two honest notes on that table. First, this is a model rather than a product, so the enterprise review platforms have nothing: no G2, Capterra, Trustpilot or GetApp entries for MiMo were found. The absence is itself the finding, and it means the sentiment here is developer sentiment rather than buyer sentiment. Second, every Reddit discussion described in this review is reported at platform level: it was reached through search indexing or quoted in a third-party write-up, and it is attributed to that source rather than presented as a page read directly.
What Users Praise
The most consistent positive theme is cost at volume, and it comes with numbers. A third-party review of the Flash sibling reports a developer who ran 50,000 document extraction and form-filling jobs, found 98 percent of tasks completed identically to Pro, and reported cutting a weekly API bill from $480 to $155, attributing the saving to the $0.14 per million input rate. That report is a third-party blog quoting a community post, so treat the figures as reported rather than verified. The second theme is the open RL disclosure. The live training dashboard, which let people watch the reinforcement-learning run as it happened, was the single most-discussed artefact of the release: the top Hacker News thread about the dashboard drew 561 points and 155 comments, and the reaction was specifically about the process being visible rather than the model appearing fully formed. The third theme is design output. Design Arena places the model in the top 3 percent for website generation and the top 4 percent for UI components, which matches what practitioners report about front-end and deck generation. The fourth, and the most U365-relevant, is adoption where it is easy to measure: OpenRouter's traffic data lists Hermes Agent as the single largest consumer of this model by token volume at about 40 billion tokens, ahead of three other coding agents, which is a concrete signal that agent platforms rather than chat users are the real market for it.
What Users Complain About
Three complaints recur and all three are specific. First, the benchmark presentation. The most upvoted technical criticism was that the artificial-analysis ranking badge reads as "#1 of 114" without making clear that the number is filtered to open-weights models, and that the unfiltered view tells a different story. The second was the size of the gap between this model's index score and DeepSeek V4.1 Flash's on the same independent index, which several readers found implausible enough to question, citing the reviewer's own correction of a comparable anomaly a week earlier. Third, and most concretely, the reinforcement-learning disclosures were read with interest rather than alarm but not without notice: a commenter quoted the live dashboard's own admission that the cyber dataset was removed from the upcoming Pro run after "bad patterns in the rollout logs", and paired it with the model card's statement that environment hardening happened during the RL runs. The reaction was amusement rather than accusation, and it is the most useful thing in the thread, because it is the vendor's own evidence about how this class of training fails. A fourth, smaller complaint is operational: the model is natively FP8 for the most part, and at least one reader reported that a given display front-end struggles with native quantization when trying to run it locally.
Sentiment Summary
Overall sentiment: Predominantly positive among developers, with a specific and unresolved scepticism about benchmark presentation.
Key themes:
Cost per completed task is the headline and the strongest claim, and it is the one independent measurement corroborates.
The open RL disclosure, including the live training dashboard and the published failure modes, is what the technical community values most about the release, more than the model itself.
Design and front-end output is praised on independent arena results, not only vendor claims.
The benchmark presentation is criticised for a filtered ranking badge, for a selective competitor set, and for a gap to DeepSeek that readers found hard to believe.
No enterprise review coverage exists, so all available sentiment is developer sentiment.
U365 Editorial Note
The crowd and the framework agree on the cost claim and agree on the scepticism, which makes both findings stronger than either would be alone.
Where they agree most usefully: developers and this evaluation both put the value in cost per completed task rather than capability. The CI-First framework does not have a cost dimension, so the crowd is reporting something the framework measures only indirectly through Time and Quantity, and the crowd is right that it is the main event. Both also stop short of claiming a capability lead: the community's most upvoted technical comment is a challenge to the ranking presentation, and this review's Faculty Note reaches a similar place by a different route.
Where they diverge, and the divergence is instructive: the community is far more interested in the reinforcement-learning disclosure than in the model's outputs. The 561-point thread about the training dashboard is bigger than most of the threads about using the model. That is a signal the framework captures only partly, in the Skill dimension, where the open RL stack is the one genuine capability-building surface in this release. If the crowd is right that the recipe is the more valuable artefact, then the Skill score of 3 understates the release for a specific audience, and this review says so explicitly rather than adjusting a score that has to hold for the common user.
The divergence that matters most for a U365 reader is the one the crowd does not discuss. No community thread raises the automatic memory-writing surface as a risk. MiMo Code's persistent memory is presented and received as a feature, and it is a feature. It is also, under framework clause 5.2.3-a, the highest Skill Illusion vector in this review series so far, because it writes a durable capability document the user did not author and then loads it into every future session. Crowd sentiment has nothing to say about that, which is exactly the case the framework exists to cover, and it is the reason this review scores Skill at 3 while the community sentiment reads as strongly positive.
Comparison and Alternatives
Alternative | "Choose [Alternative] if..." | "Choose MiMo-V2.6-Pro if..." |
Claude Opus 5.5 (https://www.anthropic.com) | You need the hardest agentic coding to succeed on the first attempt. Terminal-Bench 4.0 is 49.0 against 34.9, and it holds the top GDPval-AA 2.1 Elo published, 1846 against 1673. You pay roughly $5 and $25 per million tokens for that, against $0.435 and $0.87 here. | Your task is high volume, the errors are cheap to catch, and a 14-point gap on the hardest benchmark matters less than an order-of-magnitude cost difference. |
GPT-6 Sol (https://openai.com) | You want a closed hosted model at a comparable measured index score, 47 against 46, with a mature tooling surface and no data-residency complication. It costs $2 and $10 per million tokens, so roughly 4.6 times the input and 11.5 times the output price. | You want the same measured score band for about a tenth of the cost per completed task, and you want the weights, under MIT, as an exit option. |
GLM-5.3 (https://z.ai) | You want an open-weights model with a plain permissive license and a smaller deployment footprint: 753B total with 40B active, and an independent open-weights leaderboard position that Artificial Analysis has, at different times, placed above this model. Around $0.9 per million blended on one listing. | You want the smaller active-parameter footprint, 42B against 40B, at a lower measured price per completed task, and you value the released RL stack, which GLM-5.3 does not ship. |
Kimi K3 (https://www.moonshot.ai) | You need the largest open-weights model available, 2.8T total with 104B active, and you accept a custom licence with a revenue threshold rather than plain MIT. Around $2.3 per million blended. | Plain MIT, no revenue threshold, and roughly a fifth of the blended price are worth more to you than a capability lead that Artificial Analysis' own surfaces measure within a point or two. |
DeepSeek V4.1 Flash (https://www.deepseek.com) | Your workload is vision, terminal operation and cheap volume, where it holds independent leaderboard positions at or above this model on Terminal-Bench 2.1 and OSWorld, at a measured $0.27 per index task. | You need more measured intelligence headroom, 46 against 40 on the same index, and a 1M-token window with a cache discount, at a higher price per completed task. |
MiMo-V2.6-Flash (https://mimo.mi.com) | Your pipeline is deterministic and high volume: 98 percent task parity on batch extraction was reported at $0.14 and $0.28 per million tokens, a third of Pro's price. | Your task needs reasoning rather than pattern matching, and Pro's measured index score of 46 against Flash's lower position is worth the difference. |
Where MiMo-V2.6-Pro is clearly better: cost per completed task and license terms, and they are not close. On an independent measurement it delivers an index score of 46 at $0.13 per task and it ships under MIT, ungated, with no revenue threshold or security-review condition. The nearest closed model at a comparable score costs about eight times as much per completed task. Several open-weights competitors at a similar or lower score moved to custom licences during 2026, which makes plain MIT at this tier a real differentiator rather than a formality. Add the design results, where it sits in the top few percent on an independent arena, and the release includes the training stack on top of the model.
Where MiMo-V2.6-Pro is clearly worse: the hardest agentic coding, factual reliability, and the honesty of its own benchmark table. Terminal-Bench 4.0 at 34.9 against Claude Opus 5's 49.0, ProgramBench at 26.5 against 37.0, and ExploitBench at 47.9 against 78.5 all describe a model that is excellent at ordinary agentic work and behind on the hardest tiers. AA-Omniscience accuracy of 34.8 percent means it is not a knowledge model, and a reader who uses it as one will get confident wrong answers. And its published comparison table omits the two models the vendor's own text names as the leaders, which is the kind of selection a reader should notice and apply discount for.
Verdict and Next Steps
Who should adopt it: Teams and individuals running high-volume, mechanically checkable work, and anyone building an agent pipeline where cost per completed task is the number that decides whether the pipeline is viable. Developers who want a permissively licensed model at the top of the open-weights field and who have the discipline to check what they cannot verify. Researchers who want the RL recipe rather than only the model.
When: Now, if the use case is volume or design or agentic coding on ordinary tasks. Not now, if the use case is knowledge work where a confident wrong answer is expensive: wait for a measured knowledge evaluation that shows the floor rising, or pair this model with a second one for factual claims and treat the pair as the system.
For what: The primary task is high-volume work with a cheap-to-catch error, plus code-driven visual production. The secondary task is studying the RL stack.
UP-Context prompt packs
Here are five reusable prompts, written to the published UP-Context Method. Each one follows the published anatomy: four reusable files (Context, Role, User-Persona, Audience-Persona) plus the Task, and only the Task changes between runs. Where a pack has no audience beyond you, the Audience Persona line says so rather than being omitted, so you see the whole structure. Packs 2, 1 and 4 are the generalized versions of the sample prompts given with Workflows 1, 2 and 3 above. Copy them into MiMo Studio, the API, or your agent surface with your own context.
Prompt pack 1: the bounded agent run, with the memory layer declared before the run
Context: [repository or project] at [known good commit]. Done means [the tests pass /
every endpoint is migrated / the report is written]. The test command is [command] and it
takes [time]. The relevant code lives in [paths].
Role: AI as Co-Worker and Assistant (Profile 2, level 2). You execute the work and report.
I hold the finish line, the boundary on what may be persisted, the verification and the
decision to ship. You do not commit.
User Persona: [my role, my level, what I can check myself and what I cannot].
Audience Persona: me only. There is no external reader for this run.
Task: complete the work above and report what you changed and what you verified.
Constraints: do not change behaviour outside [scope]. Do not add a dependency. Do not
modify the test suite. Before any destructive command, stop and ask. Stop and ask only
where the decision is genuinely mine. Before you start, state what you will write to
MEMORY.md, to the checkpoint file, to global preferences and to the progress log, and
write nothing beyond that. Do not rely on any note you wrote in a previous session
unless I tell you I have read it. If the tests reveal a pre-existing failure unrelated
to this change, stop and tell me rather than fixing it.
Output format: a report with four sections in this order: files changed, the exact test
command you ran, the result, and everything you did not verify. Then a fifth section:
what you wrote to memory during this run, listed file by file.
UP-Context verification: I read the report before I read the diff, then I run the test
suite myself and read the diff. I read MEMORY.md and the checkpoint file and I delete
every claim about this project that is wrong, because that file is loaded into every
future session. I count the cost per completed task, not the price per token. Can I
defend every item in that report without you in the room? If not, the work is not done.Prompt pack 2: verifiable extraction where the check is built before the pipeline
Context: [describe the document set, the count, the formats, and where each field sits].
Some documents are scanned. The fields I need are [list]. I hold ground truth for
[number] of them, hand-checked.
Role: AI as Co-Worker and Assistant (Profile 2, level 2) for the extraction, and as
Analyst and Tester (Profile 4) when I ask you to score the sample. I own the schema and
every judgement call. You decide nothing.
User Persona: [my role, my domain knowledge of these documents].
Audience Persona: [who consumes the extracted data, and what they do with it].
Task: return the fields above as JSON, one object per document. Where a field is absent
or illegible, return null rather than a guess.
Constraints: do not infer a currency from the entity name. Do not derive one field
arithmetically from another. If two stated figures disagree, return both and flag the
disagreement. Do not resolve an inconsistency. No commentary outside the JSON. Flag every
document where a field came from a scanned page rather than a text layer.
Output format: one JSON object per document with the requested fields plus one flags
array. Then a one-line summary of how many documents carried each flag.
UP-Context verification: I run the sample against my own ground truth first and score it
per field, not per document, because one field failing at 8 percent is a different problem
from six fields failing at 1 percent. I run the same sample through a second model from a
different laboratory and compare field by field. I reconcile the totals against a record
I already hold on at least 30 documents. Someone who did not build the pipeline reviews
the output and checks the null handling specifically, because a model that guesses
instead of returning null is the failure that costs money. Can I explain and defend the
schema and the rules I gave it without the tool open? If not, the pipeline is not mine.Prompt pack 3: the knowledge-floor check, run before the tool is trusted on anything
This check is kept, with one change to how it is read. The measured knowledge accuracy of this model is 34.8 percent, so the domain must be one where you hold a written answer key before you see the output, the score to read is the confident-and-wrong rows rather than the total, the result is recorded per domain in LIPS with the date, and the check is re-run on each new domain rather than treated as one certification.
Context: [a domain where I know the correct answers and can check them]. Here are [n]
questions with the answers I hold. My expertise in this domain is [level].
Role: AI as Analyst and Tester (Profile 4, level 4). I am testing your reliability, not
asking for help. Your answers will be scored against mine.
User Persona: [my role] and the fact that I wrote the answer key before seeing your output.
Audience Persona: me only. This is a measurement, not a deliverable.
Task: answer the [n] questions, and for each one state whether you are confident,
uncertain, or do not know.
Constraints: do not guess. A "do not know" is a correct answer and will be scored as one.
Do not pad an answer to sound useful. Do not restate the question as the answer. Do not
add citations you cannot reproduce.
Output format: a table of question, answer, and confidence. Then one line stating how many
you answered with confidence. Then stop.
UP-Context verification: I score it against my own key and I record the accuracy and the
non-hallucination rate myself, per domain, in LIPS, with the date. I read the "confident
and wrong" rows rather than the total, because those are the failures that reach a reader.
Five minutes of this tells me more about whether this model belongs in my stack than any
benchmark table. I re-run the check whenever I move to a new domain, and I never present a
knowledge answer from this model without the check beside it.Prompt pack 4: a visual build from a brief, where the taste and the accessibility stay yours
Context: [the artefact]. It is for [audience] on [device]. One primary action: [action].
A completed brief is attached, including what the page is for and what the reader must do
on it. Reference image attached showing the typographic and layout style I want.
Role: AI as Co-Worker and Assistant (Profile 2, level 2). I own the message, the hierarchy
and the accessibility sign-off. You produce the implementation.
User Persona: [my role and my design or content background].
Audience Persona: [who reads the page, what they must do on it, and what they already know].
Task: produce one self-contained artefact (HTML file or component set) from the brief.
Constraints: semantic markup. Visible focus states. Body text at least 16px. Contrast at
WCAG 2.2 AA or better. No text inside images. Alt text on every image, written as
information rather than as a label. Keyboard navigation complete. Do not write the headline
or the body copy: use clearly marked placeholder text so I replace it in my own voice.
Output format: the artefact, followed by a list of the accessibility checks you did not
perform and the elements you could not assess without the live page.
UP-Context verification: I run an automated accessibility audit and check the result
against the criteria it names, because generated interfaces routinely fail contrast,
keyboard navigation and alt text and I will not accept your own claim that the page is
accessible. I replace every placeholder with my own words. I generate the same page from
the same brief with a second model and compare the information hierarchy rather than the
styling: where the two disagree about what matters, my brief was underspecified. Someone
in the intended audience reads it on a phone and tells me what they thought the page was
for. Can I explain and defend the hierarchy and the message without the tool?Prompt pack 5: studying the RL stack from the release itself
This is the one pack in the set that carries a genuine Skill benefit, and it is the reason the Skill sub-score is confirmed at 3 rather than lowered. It is written for a Fellow whose interest is how agentic reinforcement learning is actually trained.
Context: The MiMo-V2.6 release publishes more than 7,000 RL task environments, an
end-to-end RL framework, composable mini-harnesses, a technical report, and a 9B distilled
checkpoint with before-and-after pairs. [State which of those I have downloaded, my
hardware, and which one I am studying.] My existing knowledge is [level], and I have
[what I have already read].
Role: AI as Coach and Tutor (Profile 3, level 3). You teach and then test me. I reproduce
the mechanism without you and I decide when I understand it.
User Persona: [my role, my programme, my purpose for this study].
Audience Persona: me, and later my cohort when I present what I learned.
Task: teach me one mechanism from this release, end to end, and then test me on it. Start
with Groupwise Reward Synthesis, or the reward-hacking defences, or the distillation path.
Constraints: one worked example, and one boundary case where the mechanism fails or was
observed to fail. Use only the release's own documents and the environment code, and say
when a step is not documented rather than filling it in. Do not summarise the field. Do
not praise my answers. Do not offer to write my summary, because the summary is the exercise.
Output format: the mechanism, the worked example, the boundary case, then three questions
asked one at a time. After each answer, tell me what was wrong, what was missing, and what
I should open to check it. Then stop and wait.
UP-Context verification: I close the session and describe the mechanism from memory, in
writing, before reopening it. I open the technical report and the environment code and
check every claim against the primary document rather than against your explanation. I
state plainly which steps are my reconstruction because the release does not document
them, and I never present an inference as the vendor's design. If I have the hardware, I
train or fine-tune the distilled 9B checkpoint and publish my own before-and-after pair
rather than quoting the vendor's. Can I teach this mechanism to someone else without the
tool open? That is the only test that counts here.U.Copilot integration
U.Copilot is University 365's AI agent, and the guidance below tells it when to route a Fellow to this model and when to route away.
Route to MiMo-V2.6-Pro when you need to:
Run high-volume, mechanically checkable work: extraction, classification, form filling, structuring and reformatting, where the check is built before the pipeline.
Carry a long, checkable change through a repository with a written finish line and written stops, under a permissive licence.
Produce a front-end, a slide set or a diagram set from a brief you have already decided.
Take a first pass over more material than you can read, then decide what survives.
Work at the cheapest measured cost per completed task in its capability band, on an independent measurement of $0.13 per Intelligence Index task.
Do defensive cybersecurity work, where CyberGym 94.0 and MiMo Cyber Bench 80.2 are the strongest rows in its published table.
Study how agentic reinforcement learning is trained, using the published environments, framework, harnesses and the 9B distilled checkpoint.
Build or extend a Microsoft 365 centred workflow, because the OpenAI-compatible and Anthropic-compatible endpoints accept an existing client with a base URL and a model-name change.
Route away when you:
Needs a knowledge answer where a wrong answer is expensive and hard to catch. Measured accuracy is 34.8 percent on AA-Omniscience with a 59.4 percent non-hallucination rate. Run the knowledge-floor check in prompt pack 3 first, and treat a bare answer as unauthorised.
Cannot evaluate the output. This is a Skill Illusion High tool and the whole value depends on the review you do. If you cannot tell whether the result is right, split the task until you can.
Is doing the task in order to learn it, and the learning is the point. Reading a better answer is not the skill.
Needs the hardest agentic coding to succeed first time. Terminal-Bench 4.0 is 34.9 against Claude Opus 5's 49.0, ProgramBench 26.5 against 37.0 and ExploitBench 47.9 against 78.5, and the rework can cost more than the tokens saved.
Needs a final judgement, an ethical call, or communication that must carry your own voice.
Would have to send or ship the result without reading it, including generated front-end copy that reaches an audience.
Is subject to a data-residency rule that the API's jurisdiction does not satisfy. The MIT weights are the answer, and self-hosting is 573.5 GB across 132 files with 8-way tensor parallelism at minimum, which is a cluster decision rather than a configuration flag.
Wants a workflow whose only copy of the project's knowledge would be MEMORY.md. If the agent's notes are the only record, you have delegated the record.
A prompt you can paste into U.Copilot:
I am a U365 Fellow in UIT (Technology, AI, Data Science). I want to use MiMo-V2.6-Pro
for: [state the task in one sentence]. The material is [documents / repository / dataset].
Design a CI-First workflow for me that includes:
1. The CI-First Profile and the Collaboration Mode, with Centaur as the mode and the reason
stated, including the fact that the memory layer runs on its own schedule
2. The finish line written as evidence rather than as an outcome: what I will accept as
proof that the work is done
3. The stops, and the memory boundary: what the agent may persist, where those files live,
and when I will read them
4. Whether Pro is the right sibling, or whether the Flash model or UltraSpeed is the right
tool for this task, and the price difference on each
5. The unit of work I will split the task into, and why each unit is independently checkable
6. The knowledge-floor check I must run first if any part of this task is a knowledge
question, and the domain I will run it in
7. What I will read first, and what I will run myself rather than accept from the report
8. The defect classes my review is likely to miss, given the hard-task gaps in this model's
published table
9. The fallback path when the model is unavailable, and the cost record I will keep
10. The final CI-First test: what I must be able to explain and defend without the model open
11. How the completed work maps to my UIT programme and what I will record in LIPS under CAREGuardrails for U.Copilot when it recommends this model.
State the score beside the risk. 5.8/10 CI-First Positive with Medium overall AI Imposture Risk and a High Skill Illusion. Do not present the benefit without the risk, and do not publish the Medium overall without the High trap.
Never state the figure without its basis. The Faculty Note on Evidence Quality records what each number measures and which claims did not survive checking. A recommendation that quotes a number without that context is not this review.
Never present the knowledge floor as a footnote. 34.8 percent accuracy with a 59.4 percent non-hallucination rate belongs next to any recommendation for knowledge work, and the knowledge-floor check comes first.
Never quote the price ratio without its conditions. The one twentieth to one sixtieth claim holds against the most expensive tiers and compares list prices rather than bills. The measured weighted average input price paid on a gateway is $0.03899 per million, under a tenth of this model's own list rate.
Never present the vendor benchmark table as a controlled comparison. The competitor columns are Claude Opus 5, GPT-5.6 Sol and Claude Fable 5, while the launch text names Claude Fable 5.1 and GPT-6 Astra as the leaders and neither appears. The ranking badge reads as a top position and is filtered to open-weights models.
Never present the leaderboard position as settled. The same evaluator's model page and leaderboard disagree on which model tops the open-weights index, and the same evaluator has two index versions live on its site at once.
Never quote the RL improvement figures without the baseline caveat. The launch's DeepSWE baseline of 58.4 is the RL run's starting checkpoint, not the previous model, which the same table reports at 19.0.
State the memory finding explicitly, and name it as governance rather than as a feature. Framework clause 5.2.3-a applies at High because MiMo Code writes MEMORY.md, session checkpoints, global preferences and per-task progress through background writers, with 7-day and 30-day consolidation cycles. Credit the vendor's published write-permission model and say what it constrains: where the agent writes, not whether a human decides.
Never describe the model's own report as verification. It is the model's account of its own work. Verification is the test you run, the diff you read, and the source you open.
Never describe the open RL stack as proof of the reader's skill. A reproduction-ready recipe can read as understanding to someone who ran another team's pipeline. The skill is reproducing a mechanism and being able to explain it, which is why prompt pack 5 ends where it does.
Flag the jurisdiction condition. Running this model through Xiaomi's API is a deliberate decision for anyone subject to data-residency restrictions. The weights under MIT are the alternative, at cluster scale.
Flag the over-delegation pattern by name. The failure mode here is that the price removes the friction that used to force a decision. At $0.13 per completed index task, running the task three times costs less than thinking once about whether you should.
Tool-choice framing, when the question is which model to use:
Checkable volume where the error is cheap to catch, and price is the binding constraint: this model, or its Flash sibling at $0.14 and $0.28 per million, with the check built before the pipeline.
A long agentic change on ordinary work, with a written finish line: this model. Set the memory boundary before the run and read the diff rather than the summary.
The hardest agentic coding, where first-attempt success matters: a closed frontier model. Accept a cost per completed task around eight times higher, and check the current rate card rather than a figure in this review.
A knowledge question where a wrong answer is expensive: a model with a measured knowledge floor you have tested, or pair this model with a second one and treat the pair as the system.
The output must be auditable outside a vendor, or must run on U365 infrastructure: this model under MIT. Accept the 573.5 GB and the multi-node serving recipe, and re-measure on your own task.
Learning how agentic reinforcement learning is trained: this release, used as the object of study rather than as the producer, with the reproduction step in prompt pack 5 as the assessed artefact.
The task is the learning: no model as the producer. Use the model as Coach and Tutor (Profile 3) with a reproduction requirement.
You cannot evaluate the output and cannot arrange a review: no delegation. Split the task until you can check each piece.
SL-OS integration
SL-OS is the Successful Life Operating System: ULM+EVA, LIPS+CARE, the UP-Context Method, Microsoft 365, AI and human coaching combined. MiMo-V2.6-Pro produces artefacts, reports and, uniquely in this series, a durable memory store it wrote itself. Nothing it produces belongs in SL-OS without the finish line, the prompt, the verification you performed, and the memory record stored alongside it.
LIPS Digital Second Brain record.
Store each substantive session in LIPS under the relevant Project, or under the Career domain for skill-development work. Per session or per deployment, store:
The finish line as you stated it, the stops that were set, and the memory boundary declared before the run.
The CI-First Profile and the Collaboration Mode, with the Centaur rationale stated.
The pinned model version and API identifier, and the date it was pinned, plus which sibling was used (Pro, Flash or UltraSpeed) and why.
The verification evidence: the ground-truth sample score, the test run, the diff read, the sources opened, and the findings the model did not report.
The memory record. Which files the agent wrote (MEMORY.md, the checkpoint file, global preferences, per-task progress logs), where they live on disk, and the date you read them and what you deleted. This field is new in this series and it is the one most likely to be dropped.
The knowledge-floor check result for every domain you used the model for, with the date.
Realised cost per completed task at the measured volume, against the displaced alternative.
Rejected alternatives and observed failure modes, so the same error is not accepted next time.
Three LIPS fields are load-bearing for this tool and must not be left empty: the memory boundary as you wrote it, the memory file read and the date, and the knowledge-floor result for the domain used. A LIPS entry that records the output but not the boundary and not the review records nothing reusable.
CARE cycle.
Collect. Run the task and save the prompt, the report, the diff, the ground-truth score and the raw cost record without filtering. Copy the agent's memory files at the end of the session, before they are next consolidated, so the state you reviewed is the state you can reproduce.
Action Plan. Decide before reading the output what the artefact will be used for, who will rely on it, what evidence will count as proof, and which decisions must come back to you. Set the memory boundary and the knowledge-floor check before the run rather than after it.
Review. Check the report against the diff rather than accepting the report, score the ground-truth sample per field, open the sources, run the accessibility audit on any generated interface, and read what the agent wrote to its own memory.
Execute. Ship only what you have read and can defend. File the boundary, the cost record, the memory read and the failure modes so the next task of this shape starts from evidence.
ULM and EVA. MiMo-V2.6-Pro supports the Career domain: engineering capability, analytical work, document and data production, applied AI study, and the cost governance that comes with running a model in a workflow. The cost-per-completed-task framing is a professional skill as much as a technical one, and it is the strongest ULM link this release has. It touches Quality of Life for a learner whose constraint was budget rather than capability, where the release removes that constraint on a whole class of tasks. It does not support My Body and Health, My Spirit and Mind, My Character and Emotions, or My Social and Love Relationships, and it should not be stretched into them.
Within the EVA cycle: Explore, to take a first pass over material too large to read and to find out what is in a set before committing to an interpretation, and to explore how a capability of this class is actually produced from the published RL stack. Visualize, to record which of the model's findings survived your check and which did not, and the knowledge-floor score per domain, because the gap between the confident report and the verified result is the most useful thing the session produces. Action Plan, to decide what may run unattended, what needs a named review, what may be persisted to the agent's memory, and what stays with you. The Action Plan is yours and cannot be delegated to a tool whose own account of its work reads like verification.
My Successful Life cadence.
UIT Fellows building agentic delivery practice. One supervised run per week of two to four hours, with the memory files read and the review evidence written up afterwards, then a monthly review of the failure modes their checks did not catch.
UIT Fellows studying reinforcement learning. Two sessions per week of 45 minutes working through one mechanism from the published stack, each closing with a written reconstruction from memory. If the hardware allows, one reproduction attempt per month on the 9B distilled checkpoint, recorded as a before-and-after pair.
UIB Fellows building the cost case. One session to build the case, then a monthly check of the assumed cache-hit rate and cost per completed task against actual usage, including a re-check of whether the Flash sibling would do the job.
All Fellows using it for any knowledge work. The knowledge-floor check first, per new domain, recorded with the date, and re-run when the model version changes.
UIC and UID Fellows. No scheduled routine is warranted. Treat the model as an on-demand production aid inside coursework, with the accessibility audit attached to any generated interface.
Microsoft 365 integration, with the concrete paths.
MiMo-V2.6-Pro has no native integration with Outlook, To Do, OneNote, Teams, OneDrive or SharePoint. It is an API, a web surface and a terminal coding agent, and its memory files live on the machine that runs MiMo Code. The path is therefore manual on the evidence side and deliberately closed on the mailbox side.
OneNote, PROJECTS notebook. Record the workflow configuration and the learning reflection in the PROJECTS notebook at 3-Projects/_Projects General/PROJECTS/, under Current and then the relevant U365 project section, so the configuration travels with the artefact rather than living only in a chat log. Record the finish line, the memory boundary, the pinned model id and the knowledge-floor result there.
SharePoint, project folder. Store the report, the diff summary, the ground-truth sample and its score, and a copy of the agent's memory files in the relevant project folder under 3-Projects/Current/U365/<DEPT>/. Institutional work from this review series belongs in 3-Projects/Current/U365/URC-Research/INSIDE-Tools-Reviews/. Never store a 573.5 GB weight download, a repository clone, or unredacted secret-bearing logs there.
Outlook. File the run report as correspondence under the project folder. Do not give this model, or MiMo Code, read or write access to a mailbox: it is a terminal surface, there is no connector, and a mailbox is a channel to other people.
Microsoft To Do. One task per item the memory review produced: verify, correct or delete. The memory review is a task list, not a reading exercise.
Teams. Raise anything the model could not settle through the relevant project channel so it reaches a person rather than a log nobody reads. A memory claim about a project decision that is wrong belongs in the team channel as well as in the file, because other people are acting on the project record.
OneDrive. Keep local working artefacts and the prompt packs here if they are not project deliverables. The packs in this review are reusable across the release cycle and belong in your own UP-Context file set rather than in a chat history.
UP-Context files. The prompt packs above are built to sit in your UP-Context file set, following the published anatomy: four reusable files (Context, Role, User-Persona, Audience-Persona) plus the Task. The stable part is the four files; only the Task changes between runs.
Never store API keys, access tokens, or secret-bearing prompt content in LIPS, OneNote or a SharePoint list. Store the key as an environment variable and never paste it into a script or a document.
SL-OS fit. MiMo-V2.6-Pro fits SL-OS as a cheap execution engine with a memory-governance obligation and a knowledge-reliability limit. Its distinctive SL-OS property, and the reason it is the first review in this series to need a new LIPS field, is that part of the record is written by the agent on its own schedule. Where LIPS already applies a review discipline to knowledge you collect, this tool extends the same discipline onto automated work and onto the tool's own memory.
It does not replace the strategic planning layer (ULM and EVA), the knowledge management layer (LIPS and CARE), the coaching layer (U.Coach), or your own judgement. It amplifies the Collect and Review phases of CARE, and the Execute phase stays yours.
Its failure mode is the same shape as the other tools in this series and it is cheaper here, which makes it more likely: a review that thins while the output improves. The framework's arithmetic is unforgiving, CI = HI + (AI x HI), so a falling HI drags the result down even as AI rises. This release adds a second mechanism: a project record the agent wrote and you did not check is loaded into every future session as context, so an unread memory file raises the AI term across sessions you never inspected. The SL-OS rule for this tool is one sentence: no artefact ships without your own finish line, the evidence you checked, the memory files you read and pruned, and the knowledge-floor result for the domain in use.
Related U365 content:
URC's INSIDE Tools Review of GPT-6 Sol, for the comparison point on the same independent index and the same cost-per-completed-task framing.
URC's INSIDE Tools Review of Claude Opus 5.5, for the first application of framework clause 5.2.3-a and the contrast with this model's memory surface.
URC's INSIDE Tools Review of Kimi K3, for the other open-weights flagship reviewed in this series and the licensing contrast.
UIT (Technology, AI, Data Science), the institute this tool's workflows lead into: https://university-365.com/uit
U365's Recommendations to Learn More
Official learning resources
The MiMo-V2.6 launch post, which is denser than most launch posts and includes the training narrative rather than only the results: https://mimo.mi.com/docs/en-US/news/latest/v2-6
The technical report, which is the most useful document in this release for anyone studying agentic RL: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
The MiMo Code repository documentation, and specifically the memory section, which is worth reading before you run the tool rather than after: https://github.com/XiaomiMiMo/MiMo-Code
The first API call guide, which is the one page you need to get a key working: https://mimo.mi.com/docs/quick-start/first-api-call
VentureBeat's launch analysis, which reconstructs the training budget split from the technical report and is the clearest third-party account of the RL economics: the article is reachable by searching the title "Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model".
Video tutorials and channels
No official video walkthrough of MiMo-V2.6-Pro was published by the vendor at the time of writing, so the three videos below are third-party walkthroughs, labelled as community work. They are the best teaching material available on this release and they cover the two things a reader needs most: what actually changed in V2.6, and how the MiMo Code memory layer behaves over a long run.
MiMo v2.6: Xiaomi Just Built the Best Open Model, community walkthrough by Prompt Engineering, covering the release, the live RL run and the hardware bar: https://www.youtube.com/watch?v=VSh8M3CUP88
MiMo Code: Long-Horizon AI Coding with Persistent Memory, community walkthrough by Research Paper Review, on the memory layers and the long-horizon behaviour this review scores: https://www.youtube.com/watch?v=x2dQctJzGS0
MiMo 2.6 just dropped. Here's what you need to know, community summary by Adam Gardner: https://www.youtube.com/watch?v=YOOGzMHDYQo
MiMo v2.6: Xiaomi Just Built the Best Open Model, community walkthrough by Prompt Engineering
MiMo Code: Long-Horizon AI Coding with Persistent Memory, community walkthrough by Research Paper Review
MiMo 2.6 just dropped. Here's what you need to know, community summary by Adam Gardner
Written tutorials and deep-dive articles
The MiMo long-horizon blog post on the memory design, which explains the four layers and the promotion rule from session memory to project memory, and is the source for the clause 5.2.3-a finding in this review: https://mimo.xiaomi.com/blog/mimo-code-long-horizon
The Hugging Face model card for the Pro checkpoint, which carries the full architecture tables and the vendor evaluation table including the rows the launch text does not emphasise: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
The Artificial Analysis model page, for the independent index score, the measured cost per task and the measured speed, with the caution recorded in this review's Faculty Note about its generated commentary: https://artificialanalysis.ai/models/mimo-v2-6-pro
The OpenRouter model page, which is the best single page for what the model actually costs in practice because it publishes the weighted average price paid rather than the list rate: https://openrouter.ai/xiaomi/mimo-v2.6-pro
Community and social
Hacker News launch thread, 1,120 points and 477 comments, which is where the benchmark-presentation criticism and the reward-hacking discussion are: https://news.ycombinator.com/item?id=49792730
Hacker News live-training-dashboard thread, 561 points and 155 comments, which is the more interesting read because it is about how the model was made: https://news.ycombinator.com/item?id=49732270
The vendor's Reddit community: https://www.reddit.com/r/XiaomiMiMo_Official/
The GitHub organisation, for the harness, the RL framework and the environments: https://github.com/XiaomiMiMo
Resources on X
Dedicated X channels:
Xiaomi MiMo on X: the vendor account, which carries the launch thread and the demo posts: https://x.com/XiaomiMiMo
Artificial Analysis on X, where the independent index score and the cost-per-task announcement appear first: https://x.com/ArtificialAnlys
X posts with video content:
The Xiaomi MiMo V2.6 launch thread, with the omnimodal, computer-use and 3D demo videos released with the model: https://x.com/XiaomiMiMo/status/2102138559952290106
The OpenRouter announcement of the three MiMo-V2.6 models going live, including the live RL run detail: https://x.com/OpenRouter/status/2102142830034698644
Glossary
CI-First Benefit Score
The average of four dimensions, each scored 0 to 10: Time, Quantity, Quality, and Knowledge and Skill. It answers whether using the tool makes Co-Intelligence more profitable than Human Intelligence alone. Bands: 0 to 2.0 CI-First Negative, 2.1 to 4.0 CI-First Neutral, 4.1 to 6.0 CI-First Positive, 6.1 to 8.0 CI-First Strong, 8.1 to 10.0 CI-First Transformative. The score accounts for the overhead of prompting and verification, not just the benefit the tool produces. MiMo-V2.6-Pro scores 5.8.
CI-First Profile
The role the AI plays in your working relationship. (level 1) Co-Creator and Thought Partner, (level 2) Co-Worker and Assistant, (level 3) Coach and Tutor, (level 4) Analyst and Tester, (level 5) Challenger and Devil's Advocate. Lower level numbers indicate higher AI autonomy. Assigning a profile before giving the AI a task is a core CI-First discipline. MiMo-V2.6-Pro is primarily a Co-Worker and Assistant (level 2).
Humics Protection Badge
A rating of whether a tool protects, leaves neutral, or erodes three human capabilities: Creativity, Critical Thinking, and Social Authenticity. Each is scored +1, 0, or -1, and the sum gives the badge. +2 to +3 is Humics-Friendly, -1 to +1 is Humics-Neutral, -2 to -3 is Humics-Risky. It measures whether the tool strengthens the human or contributes to AI Obesity. MiMo-V2.6-Pro is Humics-Neutral at -1 / +3.
AI Imposture Risk
The likelihood that a tool traps you in one of three illusions. The Time Illusion is the appearance of saving time when net time is lost. The Quantity Illusion is high volume that looks good but does not survive inspection. The Skill Illusion is the appearance of competence in you while the underlying skill is absent or eroding. Each trap is rated Low, Medium, or High with cited evidence, and the overall level is Low when all three are Low, High when two or more are High. MiMo-V2.6-Pro is Medium overall, with Skill Illusion High.
User Sentiment
The aggregated public opinion from review platforms, community forums, and repository activity. It is reported separately from the CI-First score because crowd sentiment can contradict a rigorous evaluation. Where the two agree, the finding is stronger. Where they diverge, the divergence is worth explaining.
Review Status
Review Status records the current standing of the tool at the time of the last test. Active: the tool is current and recommended. Active (updated): recently re-checked and the content was refreshed. Changed: a re-check trigger fired and an update is pending, so read the review with that in mind. Risky: the tool has significant unresolved issues, or it has been clearly surpassed by newer alternatives. Use it with caution and read the Limits section. Retired: the tool still works but is no longer recommended. Deprecated: the tool has been shut down or fundamentally changed. Retired and Deprecated posts include a Migration Path section.
Sources
Vendor primary sources
Xiaomi MiMo, MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement, launch post, updated 2026-09-21: https://mimo.mi.com/docs/en-US/news/latest/v2-6
Xiaomi MiMo, MiMo-V2.6-Pro model page including the published rate card and capability list: https://mimo.mi.com/models/en-US/mimo-v2.6-pro
Xiaomi MiMo, MiMo-V2.6-Pro-UltraSpeed model page, for the serving configuration rather than separate weights: https://mimo.mi.com/models/en-US/mimo-v2.6-pro-ultraspeed
Xiaomi MiMo, model and rate-limit reference, which at the time of reading still described the V2.5 series in its main table: https://mimo.mi.com/docs/en-US/quick-start/summary/model
Xiaomi MiMo, model release changelog, recording the 2026-09-22 MiMo-V2.6 series release: https://mimo.mi.com/docs/en-US/updates/model
Xiaomi MiMo, first API call guide, including the OpenAI-compatible and Anthropic-compatible base URLs: https://mimo.mi.com/docs/quick-start/first-api-call
Xiaomi MiMo, deep-thinking usage guide, including the affected agent products and the reasoning-content pass-back requirement: https://mimo.mi.com/docs/en-US/quick-start/usage-guide/text-generation/deep-thinking
Xiaomi MiMo, Token Plan page, for the Individual and Team subscription tiers: https://platform.xiaomimimo.com
Xiaomi MiMo, API pricing page, for the V2.5-series rate card that the V2.6 series keeps: https://mimo.mi.com/docs/en-US/pricing
Xiaomi MiMo, MiMo Code: Scaling Coding Agents to Long-Horizon Tasks, for the four-layer memory design and the write-permission model: https://mimo.xiaomi.com/blog/mimo-code-long-horizon
Xiaomi MiMo, MiMo-V2.6-Pro-RL model card on Hugging Face, for the architecture tables, the evaluation table and the deployment recipes: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Xiaomi MiMo, MiMo-V2.6 technical report: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
Xiaomi MiMo, Hugging Face collection listing the Pro, Flash and Distill-Qwen-9B checkpoints: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
Xiaomi MiMo, MiMo-V2.5-Pro launch post, for the prior-generation architecture and the 1.02T/42B continuity: http://mimo.xiaomi.com/mimo-v2-5-pro
Xiaomi MiMo, MiMo Code repository, for the persistent memory, context management and subagent documentation: https://github.com/XiaomiMiMo/MiMo-Code
Xiaomi MiMo, GitHub organisation listing, for repository activity, stars, forks and open issues: https://github.com/XiaomiMiMo
Hugging Face API metadata for XiaomiMiMo/MiMo-V2.6-Pro-RL, for the license, gating status, likes, downloads and the file-level size and precision breakdown: https://huggingface.co/api/models/XiaomiMiMo/MiMo-V2.6-Pro-RL
Xiaomi MiMo Desktop early access page: https://mimo.xiaomimimo.com/desktop/
Independent sources
Artificial Analysis, MiMo-V2.6-Pro intelligence, performance and price analysis, for the Intelligence Index score of 46, the $0.13 cost per index task, the 134.3 tokens per second output speed, the 2.15-second time to first token, and the open-weights ranking statement: https://artificialanalysis.ai/models/mimo-v2-6-pro
Artificial Analysis, Announcing the Artificial Analysis Intelligence Index v4.3, dated 2026-09-07, for the index version history, the open-weights ordering before this release, and the Pareto-frontier note: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
Artificial Analysis, LLM leaderboard, read 2026-09-24, which names GLM-5.3 as the highest-ranked open weights model and does not list MiMo-V2.6-Pro in its open-weights summary, and which is the source of the conflict recorded in the Faculty Note: https://artificialanalysis.ai/leaderboards/models
Artificial Analysis, large open-source models comparison, read 2026-09-24, for the open-weights cohort and the index version label on that page: https://artificialanalysis.ai/models/open-source/large
Artificial Analysis on X, for the top-open-weights launch statement: https://x.com/ArtificialAnlys
OpenRouter, MiMo-V2.6-Pro pricing, providers and benchmarks, for the weighted average price actually paid, measured throughput and latency, uptime and availability, and the AA-sourced benchmark table: https://openrouter.ai/xiaomi/mimo-v2.6-pro
models.dev, MiMo-V2.6-Pro pricing, providers and specs, for the gateway price spread and the capability metadata: https://models.dev/models/xiaomi/mimo-v2.6-pro/
LLM Stats, MiMo-V2.6-Pro benchmarks, pricing and context window, for the p95 time to first token, the conversation-depth measurement, and the blended price: https://llm-stats.com/models/mimo-v2.6-pro
VentureBeat, Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world alongside cheaper V2.6-Flash, for the training-budget reconstruction, the reward-hacking account, the GRS and GAR mechanisms, and the data-residency caution: https://venturebeat.com/technology/better-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash
VentureBeat, Xiaomi's open source agentic AI coding harness MiMo Code beats Claude Code at ultra-long 200-plus-step tasks, for the memory layering and the checkpoint-writer subagent: https://venturebeat.com/technology/xiaomis-new-open-source-agentic-ai-coding-harness-mimo-code-beats-claude-code-at-ultra-long-200-step-tasks
TechNode, Xiaomi open-sources MiMo-V2.6 models after scaling reinforcement learning, for the independent restatement of the index score and training costs: https://technode.com/2026/09/22/xiaomi-open-sources-mimo-v2-6-models-after-scaling-reinforcement-learning/
247wallst, Xiaomi's MiMo-V2.6-Pro tops open weights AI at just $0.13 per task, for the cost-per-task framing and the Pareto-frontier statement: https://247wallst.com/cards/xpost-01m32tmphbe63sc3tc6shtkhn9
LLMBill, Claude Fable 5.1 pricing, used as the comparison rate for the vendor's one-twentieth-to-one-sixtieth price claim: https://www.llmbill.app/models/anthropic/claude-fable-5.1
CloudPrice, Claude Fable 5.1 pricing and specs, for the corroborating Anthropic rate card: https://cloudprice.net/models/anthropic-claude-5-1-fable
Gigazine, Xiaomi unveils the MiMo-V2.5-Pro, for the April 2026 independent index figure on the previous generation, cited only with the version caveat recorded in the Faculty Note: https://gigazine.net/gsc_news/en/20260428-xiaomi-mimo-v2-5-pro
Gigazine, Xiaomi has released MiMo Code as open source, for the blind-comparison result against Claude Code on tasks over 200 steps: https://gigazine.net/gsc_news/en/20260611-xiaomi-mimo-code
Benchmark sources cited by the vendor, used for the row definitions rather than the numbers: Zapier AutomationBench (https://zapier.com/benchmarks), Agents' Last Exam (https://agents-last-exam.org/), DeepSWE (https://deepswe.datacurve.ai/), OSWorld (https://osworld-v2.xlang.ai/), Design Arena (https://www.designarena.ai/)
Community and community-reported evidence
Hacker News, MiMo v2.6 launch thread, 1,120 points and 477 comments as of 2026-09-24: https://news.ycombinator.com/item?id=49792730
Hacker News, Xiaomi Mimo 2.6 live post-training dashboard, 561 points and 155 comments as of 2026-09-24: https://news.ycombinator.com/item?id=49732270
Hacker News, MiMo-v2.6-Pro intelligence, performance and price analysis thread, which is where the benchmark-presentation criticism and the DeepSeek-gap scepticism appear: https://news.ycombinator.com/item?id=49796660
Product Hunt, MiMo, for the single review, follower count and launch history: https://www.producthunt.com/products/mimo-3
Reddit, vendor-run community, reported at platform level through search indexing and third-party write-ups: https://www.reddit.com/r/XiaomiMiMo_Official/
Reddit r/LocalLLaMA live-training-dashboard discussion, reported at platform level through search indexing and third-party write-ups: https://www.reddit.com/r/LocalLLaMA/
Tabbit, MiMo-V2.6-Flash review, for the reported 50,000-job extraction pipeline and the $480-to-$155 weekly figure, which is a third-party report of a community post and is treated as reported rather than verified: https://go.tabbit.ai/blog/mimo-v2-6-flash-review
MyClaw, Xiaomi MiMo V2.6 review, for the Pro-versus-Flash escalation guidance: https://myclaw.ai/blog/mimo-v2-6
SaaSCity, MiMo Code write-up, for the memory-layer prioritisation and the roughly 65K-token injected-context note: https://saascity.io/blog/your-ai-coding-agent-forgets-everything-mimo-code-doesnt
ScriptByAI, MiMo Code configuration reference, for the checkpoint, memory, dream and distill settings: https://www.scriptbyai.com/mimocode-free-terminal-coding-agent
Internal sources
CI-First Evaluation Framework v1.2 and the INSIDE Tools Post Template, including the LLM, open-source and agent-platform variants that all three apply to here, and the published INSIDE Tools Reviews used as comparisons in this series: https://www.university-365.com/tools
Faculty Note on Evidence Quality
Four claims from this release did not survive checking against primary sources, and one of them is a problem in the independent source rather than in the vendor's.
First, "the most powerful open-source model available" is true on one Artificial Analysis surface and contradicted by another. The vendor's words are that MiMo-V2.6-Pro surpasses Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available. Artificial Analysis' own model page for MiMo-V2.6-Pro agrees, states an index score of 46, and says it is the highest-ranked open weights model. But the same organisation's LLM leaderboard, read on the same day, states that "GLM-5.3 (max) is the highest-ranked open weights model with an Intelligence Index score of 45", and lists the top open-weights models as GLM-5.3 at 45, Kimi K3 at 44 and GLM-5.3-Flash at 42, with MiMo-V2.6-Pro absent from that summary entirely. Two surfaces published by the same evaluator, read within minutes of each other, disagree about which model holds the top open-weights position. The margin at stake is one point on a composite index. The honest reading is not that Xiaomi is wrong. It is that a one-point lead on one evaluator's composite index, where that evaluator's own pages disagree, is a provisional claim rather than a settled fact, and a reader should wait for a second evaluation setup before treating it as established. Reinforcing the caution, the same evaluator's index-version announcement from 2026-09-07 describes v4.3 while its open-source model page still labels the index as v4.2, so two index versions are live on the site at once and scores read from different pages are not guaranteed to be comparable. That also means the widely circulated comparison to the previous generation is not like-for-like: the April 2026 figure for MiMo-V2.5-Pro was reported against an earlier, easier index version, and the same model now appears at 26 on the current leaderboard. A drop from 54 to 26 for the same model over five months is not a decline in the model; it is what happens when an evaluator keeps raising task difficulty, and it is a reason never to compare index scores across versions.
Second, the "one twentieth to one sixtieth of overseas models" price claim is a list-price comparison against the most expensive overseas tiers, and it is arithmetically checkable. Against Claude Fable 5.1 at $10 and $50 per million tokens, MiMo's $0.435 and $0.87 are about one twenty-third on input and one fifty-seventh on output, so the range holds at the cheap end. Against GPT-6 Astra, priced at the same $10 and $50, it holds the same way. Against GPT-6 Sol at $2 and $10 it is roughly one quarter and one eleventh, which is outside the claimed range, and against Gemini-class mid-tiers the ratio narrows further. Two further qualifications matter more than the arithmetic. The comparison is between list prices, while the actual price paid is much lower on both sides once caching is included: OpenRouter reports a weighted average actual input price of $0.03899 per million for this model, under a tenth of its own list rate, so a ratio between two list rates does not describe either party's bill. And Artificial Analysis, comparing against the median of all models rather than against overseas flagships, calls this model's input price "somewhat expensive" with a median of $0.30. All three statements are true: it is the cheapest way to buy this capability tier, it is not cheap in absolute terms against the whole market, and the headline ratio only holds against the most expensive tiers. A reader building a budget should use the measured cost per completed task rather than the ratio.
Third, the benchmark table omits the two models the vendor itself names as the leaders, and its "on par" sentence spans a mixed row set. The launch text names Claude Fable 5.1 and GPT-6 Astra as the strongest closed-source models, and then says MiMo-V2.6-Pro has achieved performance on most agent benchmarks on par with Claude Opus 5 and GPT-5.6 Sol. The evaluation table's competitor columns are Claude Opus 5, GPT-5.6 Sol and Claude Fable 5. Neither Fable 5.1 nor GPT-6 Astra appears. So the two models the vendor identifies as the frontier are the two absent from the comparison, and the comparison instead runs against Claude Fable 5, which is two generations behind Fable 5.1. The general check applies exactly as the framework requires: read both sides of every head-to-head and confirm the comparison is against the competitor's current published setting rather than a superseded one. On the substance, "on par" survives on some rows and not others, and the split is instructive rather than damning. Ahead or level: AutomationBench v1.0.6, Toolathlon-Verified, Agents' Last Exam, JobBench, Terminal Bench 2.1, and MiMo VisualCoding. Behind: ProgramBench by 10.5 points, MiMo Code Bench by 5.4, Terminal Bench 4.0 by 14.1, OSWorld-Verified by 1.4 against Claude Opus 5 and 4.0 against Claude Fable 5, ExploitGym by 12.5, ExploitBench by 30.6, and SEC Bench Pro by 12.8. That is a credible capability profile for an open-weights model at this price. It is not the profile the sentence implies, and the Terminal-Bench pair is the clearest illustration: level on the older version, 14.1 points behind on the newer one.
Fourth, the reinforcement-learning improvement figures do not match the published table, and the baseline quoted is not the previous model. The launch post states that DeepSWE v1.1 improved by about 14 points for Pro, "from 58.4 to 72.6", and by about 17 points for Flash, from 48.8 to 65.7. The model card's evaluation table published in the same release reports MiMo-V2.6-Pro at 71.9 and MiMo-V2.6-Flash at 67.9 on the same benchmark. Neither endpoint matches: Pro is 0.7 below the stated figure and Flash is 2.2 above it. That is consistent with the 72.6 and 65.7 being mid-run readings from the reinforcement-learning process rather than final checkpoint results, which is a reasonable thing for a launch post to report and is not what the sentence says. The larger issue is the baseline. The launch says "from 58.4", which a reader will naturally take as the previous model's score. The same vendor's table reports MiMo-V2.5-Pro at 19.0 on DeepSWE v1.1. So 58.4 is the starting checkpoint of the RL run, not the predecessor model's published result, and the improvement a reader would compute against the previous generation, 19.0 to 71.9, is enormously larger and measures something entirely different: an architecture and training change rather than a reinforcement-learning step. Both numbers are useful. Only one of them is the one the sentence reads as.
On the independent source's own presentation, three cautions. The Artificial Analysis model page contains a generated commentary line that reads "it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M", comparing a figure to itself, and it labels the input price "somewhat expensive" while the vendor calls the same number the cheapest in its class. Those are template-filling artefacts in the commentary rather than errors in the measurement, and the measured numbers, the index score, the cost per task and the two speed figures, are the parts to rely on. OpenRouter's measured P50 latency of 2.19 seconds, Artificial Analysis' 2.15-second time to first token, and an aggregator's 6.87-second p95 time to first token are consistent with each other as different statistics of the same distribution rather than contradictory, and the spread between a median and a tail is worth knowing before you build a latency-sensitive workflow. And the aggregator's own measured quality tracker shows a rank decline from #74 of 117 on the first turn to #112 of 116 across turns 2 to 10 of the same conversation. That measure uses a different scale from the composite index, is based on a small vote count of six, and published no data beyond turn 10, so it is a signal to watch rather than a finding. It does point in the opposite direction from the long-horizon framing, and it is recorded as trigger 4 rather than dismissed.
What Xiaomi got right, stated with the same emphasis. The vendor published its own failure modes. The technical report documents reward hacking explicitly, names the specific evasion techniques its agents discovered, and reports that the cyber dataset was removed from an upcoming Pro run after bad patterns appeared in the rollout logs. It also publishes the write-permission model for the memory system, restricting the background writer to specified paths with out-of-bounds writes rejected at the code level. Most vendors at this level publish the wins and leave the footnotes to a reporter. A vendor that hands you its own reward-hacking account has given you the most useful document in the release, and this review's Critical Thinking assessment rests on it precisely because the vendor supplied it.
All four cases teach the same lesson, which is the one this review is built on. When a vendor's headline sentence and the vendor's own footnotes and tables say different things, read the tables, and say so when they differ. In this release, the tables and the technical report are the honest part, and they are good enough that the headline sentence did not need to overreach.
Review conducted under the CI-First Evaluation Framework, version 1.2. Scoring date 2026-09-24. Tool version reviewed: MiMo-V2.6-Pro (`mimo-v2.6-pro`, released 2026-09-21, open-sourced 2026-09-22). Framework version applied: 1.2. Framework clauses checked: 5.2.3-a applies at the High threshold through the MiMo Code delivery surface; 4.2-a returns a null; 7.5 returns a null.










Comments