MiniMax M3.1 Flash: what the review video shows and what can be verified
WorldofAI published a 21-minute video on MiniMax M3.1 Flash Preview on 1 October 2026, and it had picked up roughly 36,000 views by 2 October. In it the reviewer builds a first-person shooter in the style of Call of Duty Zombies, a macOS clone with a Minecraft clone running as an app inside it, a product page for headphones, a landing page for a graphics card, and a 3D village terrain. His verdict: a model from the flash tier delivering quality near the top of the market, and costing him nothing as a subscriber. Only the second half of that sentence is usable for a procurement decision, and even that half needs a footnote.
What the video actually shows
The demos are good, and that belongs at the front. The shooter ships rounds, barricades, a buy menu, several rooms, a flashlight and a machine that upgrades a weapon. The macOS clone has a menu bar with battery mode, Spotlight, a working App Store, a terminal, an activity monitor, a music app and a Minecraft clone where you can place blocks, place light and craft. In a head-to-head against Claude Fable 5.1 on medium thinking effort the reviewer claims one hour of run time against two, and mentions "over 150 tokens per second". He also says himself that he is not claiming parity with Opus.
Two details in the video deserve care. The comparison opponents are hard to reconstruct from the audio, because the automatic captions swallow model names, and "GBT6 Soul" is one of those artefacts. And the 150 tokens per second lands between two gameplay clips, nobody measures it there. MiniMax lists roughly 100 tokens per second for M3 in its model overview and no figure at all for M3.1 Flash.
What the documentation states
MiniMax describes M3.1 Flash Preview in the model overview as a "frontier multimodal coding model with 1M context window and tunable thinking depth". Context window one million tokens, inputs text, image and video, output text. Thinking depth goes into effort, which accepts low, medium, high, xhigh and max, and leaving the field out gets you max. Thinking cannot be switched off: thinking: {"type": "disabled"} or effort: "none" returns status 400 with the text "requires adaptive thinking". For a latency-sensitive path that is the most important line in the document, because a flash model ships in its most expensive default setting.
Availability is restricted in two places in the documentation: the model is available through the M Plan and MiniMax Code only. Parameter count, architecture, a technical report and a model card are all absent. The model release notes end with H3 dated 31 July 2026, Music dated 16 July and M3 dated 1 June, and there is no entry for the preview. The pay-as-you-go price list carries M3 and no line for M3.1 Flash.
What has been measured independently
Artificial Analysis does not carry the model. Its full leaderboard model list names M1, M2, M2.1, M2.5, M2.7 and M3 from MiniMax, and no model called M3.1, checked on 2 October 2026 across the whole list of 2.4 MB of page source. That is the straightforward consequence of the missing access: no public API means no measurement, and no measurement means no number except the vendor's own. The 90 to 110 tokens per second circulating in forums did not come off a test bench either.
The open sibling has been measured. MiniMax-M3 sits in the Artificial Analysis tables at 29 points on the intelligence index, 87.6 output tokens per second, 0.51 dollars per index task, 428 billion total parameters with 23 billion active, and the MiniMax Community License, as of 2 October 2026. Opus 5.5 reaches 58 points in the same measurement run. So the price gap between these classes is real, and part of it is a real quality gap. We wrote about M3 in June, with the sparse attention mechanism and the data-protection question covered in our M3 post.
"Free" here means: promotion
The M Plan costs 22 dollars (Go), 55 dollars (Explore) and 132 dollars (Build) per month. M3.1 Flash Preview is the text model in all three tiers, with usage scaling by factor 1, 3 and 7.5. MiniMax grants 50 percent off the first month on monthly billing until 14 October, annual billing excluded. The free run the video is praising comes from a National Day promotion: double credits on the daily check-in from 28 September to 7 October (UTC+8), plus reset quotas on the older Token Plans. Both of those are in press reports about the announcement, and nothing about them is in the documentation. As an anchor for the price level, the M3 line in the pay-as-you-go list stays useful: 0.30 dollars per million input tokens and 1.20 dollars per million output tokens, double that above 512,000 input tokens, and 1.5 times the rate for the priority tier.
An EU point hardly anyone mentions
MiniMax publishes the summary of training content for its base models under Article 53(1)(d) of the EU AI Act and links it from its documentation. That transparency page lists M3 and H3, and as of 2 October 2026 it does not list M3.1 Flash Preview. For an operator in the EU that means the assessment of what data went into the model has to be assembled by the operator itself for this preview, even though the model only runs through a subscription. Anyone planning documentation duties for AI services should clarify which paperwork exists for that exact version before buying a test subscription, see the Article 4 documentation trail.
What this means for our own hosting
M3.1 Flash is not a candidate for our own inference lane, there is nothing to download. The MiniMax organisation page on Hugging Face confirms it on 2 October: open language models are M3 (including the MXFP8 variant) and M2.7, plus H3 for video, and no repository exists for the preview.
For M3, MiniMax does provide a self-hosting guide, and that guide settles the arithmetic. The reference build is eight B200 cards on one node with tensor parallelism 8, the MXFP8 weights occupy around 444 GB, and MiniMax calculates H200 systems in BF16 at around 854 GB. Support runs through an SGLang development image, the status is "Experimental" in the original wording, and validated inputs run from 1,000 to 128,000 tokens. MiniMax also states explicitly that 23 billion active parameters does not mean only 23 billion have to sit in memory.
Our AI lane runs on four RTX 5090 with 128 GB of combined graphics memory, described in the four-day test and in the comparison of two vLLM builds. 444 GB of weights against 128 GB of memory is the same arithmetic as 854 GB against 128 GB: M3 does not fit that machine in any precision. Running the full million-token class on own hardware needs a different hardware path, see open versus closed models.
Side finding: our own DNS blocks the vendor
Checking these sources turned up something on 2 October 2026 that has nothing to do with MiniMax and still belongs here. platform.minimax.io, www.minimax.io, agent.minimax.io and api.minimax.io all answer 0.0.0.0 through our Technitium cluster. The answer comes from a subscribed block list, in this case the RPiList spam list. The pages are reachable as soon as you resolve outside that answer, and for this post we resolved over DNS over HTTPS. The operational lesson: when somebody reports that an AI site is dead, the first question is whether a block list contains the domain, and whether the vendor really is down. Whether a vendor domain belongs on a allow list needs a decision with a reason behind it.
Four questions for the next demo video
A video that puts a model in the neighbourhood of Opus is an observation, not a result. Before a model enters our own test lane, these four questions have to be on the table.
- Are there open weights, and under which license?
- Is there an independent measurement, on which index version and against which provider configuration?
- What does one of our typical tasks cost inside a subscription, at the thinking level the job actually needs?
- What documentation does the vendor deliver for this exact model version?
For M3.1 Flash Preview none of the four is answered cleanly: the weights are missing, the measurement is missing, a price per task is missing, and the training-content summary for the version is missing. The model stays on our watch list, with the agreement that we repeat the assessment once MiniMax delivers weights, a price or a model card.
Further reading
- Model overview, context window,
effort, the 400 error when switching thinking off: https://platform.minimax.io/docs/guides/text-generation - Availability restricted to M Plan and MiniMax Code: https://platform.minimax.io/docs/guides/models-intro
- Tiers, included models, usage factors: https://platform.minimax.io/docs/m-plan/intro
- Regular prices and the 14 October deadline: https://platform.minimax.io/docs/m-plan/monthly-offer
- Prices for M3 and M2.7, no entry for M3.1 Flash: https://platform.minimax.io/docs/guides/pricing-paygo
- 8× B200, 444 GB of weights, experimental status: https://platform.minimax.io/docs/guides/local-deploy-m3
- Training-content summaries under the AI Act: https://platform.minimax.io/docs/guides/transparency
- Newest entry H3 from 31 July 2026: https://platform.minimax.io/docs/release-notes/models
- Review from 1 October 2026: https://www.youtube.com/watch?v=9zVYS00N6mg
- Measurements for M3 and for the expensive class: https://artificialanalysis.ai/leaderboards/models
- Open weights of the M series: https://huggingface.co/MiniMaxAI
Can we self-host MiniMax M3.1 Flash?+
No. The preview has no weights, and access runs exclusively through the M Plan and MiniMax Code. Its open sibling MiniMax-M3 does run self-hosted: MiniMax's own reference build is eight B200 cards in one node with roughly 444 GB of MXFP8 weights. That is a different hardware class from our lane of four RTX 5090.
The demos look strong, so why the doubt?+
Because a demo is not a measurement, and right now nobody can measure. Artificial Analysis does not carry the model in its full model list, MiniMax's model release notes contain no entry for it, and without public API access no benchmark shop gets in. The only number that matters for our own use cases has to come from a subscription we buy ourselves.
Is it true that the model is free?+
The free window is a promotion. Press reports on the announcement describe double daily credits from 28 September to 7 October, and MiniMax reset the quotas of the old Token Plans at launch. The regular M Plan prices are in the documentation: 22, 55 and 132 dollars a month, first month at half price until 14 October. There is no price per million tokens for this model.
senn-tech