Skip to main content
SaaSCity.io
DirectoriesLive LaunchesBlogWrite for UsAdvertise
Submit
Home/Blog/MiniMax M3.1 Flash Preview: What's Confirmed, How to Use It, and How It Differs from M3
Back to Blog

News

MiniMax M3.1 Flash Preview: What's Confirmed, How to Use It, and How It Differs from M3

MiniMax-M3.1-Flash-Preview launched on September 27, 2026 inside MiniMax Code and the Token Plan. It has a 1M-token context window, text, image and video input, and five effort levels, and thinking cannot be turned off. There is no model card, no benchmark table and no open weights yet. This guide separates what MiniMax has confirmed from the leaks and explains when M3 is still the better choice.

ghosty
ghosty
Founder, SaaSCity
September 28, 202614 min read
MiniMax M3.1 Flash Preview: What's Confirmed, How to Use It, and How It Differs from M3
Contents (14)
  1. Key takeaways
  2. What is MiniMax M3.1?
  3. MiniMax M3.1 release date and availability
  4. Official specs
  5. Effort levels explained
  6. How to use MiniMax M3.1 today
  7. MiniMax M3.1 vs MiniMax M3
  8. Benchmarks: what exists and what does not
  9. Pricing and plans
  10. Architecture: official vs leaked
  11. What about Space Bunny Alpha?
  12. Who should switch, and who should wait
  13. Limitations and open questions
  14. Bottom line

Last updated: September 28, 2026. This story is moving daily. We will update the "what we still don't know" section first when MiniMax publishes more.

Quick answer: MiniMax-M3.1-Flash-Preview is MiniMax's newest M-series language model. It launched quietly on September 27, 2026 inside MiniMax Code and on the Token Plan. The official docs list a 1,000,000-token context window, text, image and video input, and five effort levels from low to max. Thinking is always on. MiniMax has not published a model card, benchmarks, a parameter count, open weights or a per-token price, so treat it as a preview of a fast everyday coding model, not a documented successor to M3.

If you searched "MiniMax M3.1", this is the model you found. There is no separate full M3.1 yet. The official name is MiniMax-M3.1-Flash-Preview, and both parts of the suffix matter. "Flash" means MiniMax built it for speed on short, everyday tasks. "Preview" means you get no model card, and the details can change.

A note on who is writing this. SaaSCity is a startup directory, and most of the founders who launch with us build their products with AI coding tools. When a new coding model ships, they ask us whether to switch. So this guide keeps two lists apart: what MiniMax has put in writing, and what the community is guessing.

Key takeaways

  • Launched September 27, 2026 in MiniMax Code and on the Token Plan.
  • Official model ID: MiniMax-M3.1-Flash-Preview.
  • 1M-token context. Text, image and video input.
  • Five effort levels (low, medium, high, xhigh, max). The default is max. Thinking cannot be disabled.
  • No official benchmarks, parameter count, open weights or pay-as-you-go price.
  • MiniMax-M3 is still the public API and open-weight flagship.

What is MiniMax M3.1?

MiniMax M3.1 currently means one model: MiniMax-M3.1-Flash-Preview. MiniMax describes it as a frontier multimodal coding model with a 1M context window and tunable thinking depth. It sits beside MiniMax-M3, M2.7 and M2.7-highspeed in the model list, and it does not replace them.

MiniMax API docs Model Invocation page listing MiniMax-M3.1-Flash-Preview with a 1,000,000-token context window and a notice that it is available only through Token Plan and MiniMax Code

The docs call it "the latest M-series language model for agentic reasoning, tool use, coding, and long-context tasks." The blue banner at the top of the same page carries the restriction that matters most: the model is "available only through Token Plan and MiniMax Code for now."

The M-series has moved quickly:

  • MiniMax-M1 (2025): reasoning model with Lightning Attention, 1M context, Apache-2.0 weights.
  • M2, M2.1, M2.5, M2.7 and M2.7-highspeed: 204,800-token context, the agentic coding line.
  • MiniMax-M3 (June 1, 2026): native multimodal MoE with MiniMax Sparse Attention, 1M context, open weights under the MiniMax Community License. Our MiniMax M3 review covers it in depth.
  • MiniMax-M3.1-Flash-Preview (September 27, 2026): the current model, sold for daily development work.

MiniMax announced it through its agent account on X:

MiniMax_Agent post on X from September 27, 2026 announcing M3.1-Flash-Preview on MiniMax Code, with a launch graphic showing the model picker and effort menu

The launch post from @MiniMaxAgent says the model is "built for everyday development." Look at the launch graphic. The model picker shows M3.1-Flash-Preview above M3 and M2.7, and a second menu lists the five effort levels. MiniMax put the effort control in its launch image, which suggests effort is the main way you control this model.

MiniMax M3.1 release date and availability

MiniMax-M3.1-Flash-Preview became public on September 27, 2026. Is it out? The Flash Preview is, but a fully documented M3.1 is not.

Where you can use it:

  • MiniMax Code, on desktop (macOS and Windows) and on the web, from the model picker. Some early reviews say it is the default model there.
  • The Token Plan, through a Subscription Key on the Anthropic-compatible and OpenAI-compatible endpoints.

Where it is missing (as of September 28, 2026):

  • No public Hugging Face repository.
  • No public OpenRouter listing under that name.
  • No research blog post or model card.
  • No pay-as-you-go listing. That API still centers on MiniMax-M3.

Launch promotions, from the follow-up post in the same X thread:

  • Token Plan quotas were reset when the model went live.
  • Daily check-ins in MiniMax Code earn 2x credits from September 28 to October 7, 2026 (UTC+8), for new and existing users.
  • You can spend those credits on M3.1-Flash-Preview or on MiniMax's H3 and H3 Max video models.

Official specs

This table lists only what MiniMax has published. Everything else is unknown.

SpecOfficial value
Model IDMiniMax-M3.1-Flash-Preview
Context window1,000,000 tokens
InputText, image, video
OutputText, plus a separate thinking stream
ThinkingOn by default, cannot be disabled
Effort levelslow, medium, high, xhigh, max
Default effortmax
AccessToken Plan and MiniMax Code only
Open weightsNot announced
Official benchmarksNone published
Parameter countNot disclosed
Per-token priceNot published
Output speed (tps)Not published (M3 is listed at about 100+ tps)

The 1M context means you can put a mid-sized repository, a long agent transcript or a stack of API docs in one request. MiniMax has not said how quality holds up near the limit on the Flash model, so test it on your own large contexts before you rely on it.

The multimodal input is useful for coding work. You can paste a screenshot of a broken layout and ask for the CSS fix, or record a short clip of a UI bug and ask for the patch. Token cost for images depends on size and detail. MiniMax's OpenAI-compatible docs give rough ranges: a few hundred tokens at low detail, about 1k to 3k by default, and several thousand or more at high. Check the usage field in the response for the real number.

The always-on thinking is a deliberate design choice. MiniMax says reasoning before answering "measurably improves accuracy on agentic reasoning, tool use, coding, and maths." Because thinking cannot be turned off, the effort level is the only speed control you have.

Effort levels explained

The effort parameter sets how much the model thinks before it answers. Higher levels produce more output tokens and take longer. If you omit it, M3.1-Flash-Preview uses max.

MiniMax docs Thinking depth (effort) section showing the five effort values and the field name for each protocol: output_config.effort, reasoning_effort and reasoning.effort

The field name depends on which API protocol you call:

ProtocolEffort fieldWhere thinking comes backOutput cap field
Anthropic-compatible (recommended)output_config.effortthinking content blockmax_tokens
OpenAI Chat Completionsreasoning_effortreasoning_contentmax_tokens / max_completion_tokens
OpenAI Responsesreasoning.effortoutput item type: "reasoning"max_output_tokens

MiniMax has not published latency or quality curves for each level. The table below is our suggested starting point, not official guidance. Test it on your own tasks.

EffortTry it first for
lowFormatting, renames, tiny patches, tool-call glue
mediumStandard bug fixes, code review comments
highMulti-file changes, debugging
xhigh / maxPlanning, hard reasoning, repo-wide design. Expect more tokens and a longer wait

One thing will catch many people. The default is max, in the API and, per early reviews, in the MiniMax Code UI. If you leave it there and ask for a one-line rename, the model will still think at full depth, and the "Flash" model will feel slow. Most complaints about speed in the first 24 hours probably come from this default.

The thinking limitation

The docs state the rule directly. If you send thinking: {"type": "disabled"} or effort: "none", you get a 400 error:

model "MiniMax-M3.1-Flash-Preview" requires adaptive thinking

This matters if you are porting code from MiniMax-M3. On M3, community reports describe enabled, adaptive and disabled thinking modes, so a client that disables thinking to save tokens will fail against M3.1. Lower the effort level instead.

While you are here

Get your SaaS listed on SaaSCity

A permanent listing on the live city map, a DR 64+ dofollow backlink and a launch week in front of founders. Free with a badge, or skip the queue with Quick Pass — live within 24 hours.

Submit your SaaSWhat you get

How to use MiniMax M3.1 today

MiniMax Code download page with the Download MiniMax Code button and a preview of the desktop app showing projects, plugins, scheduled tasks and a model selector

In MiniMax Code

  1. Download MiniMax Code from agent.minimax.io/download, or open the web workspace.
  2. Sign in. Early users mention Google login.
  3. Open the model picker next to the prompt box and choose M3.1-Flash-Preview.
  4. Set the effort level before you start. Pick medium or high for normal coding. Keep max for planning.
  5. Check in every day until October 7, 2026 to collect the double credits.

The desktop app has more features than a chat window. Reviews mention Agent Teams, memory, a built-in browser and scheduled tasks. There is also a CLI, which is easier to script. Some developers use a community minimax-vscode plugin with VS Code and Copilot. MiniMax does not maintain it, so treat it as unofficial.

Through the API

Use your Token Plan Subscription Key. A pay-as-you-go API key is a different credential, and the Token Plan docs say the two are "not interchangeable."

Base URLs:

  • Anthropic-compatible: https://api.minimax.io/anthropic (mainland China: https://api.minimax.cn/anthropic)
  • OpenAI-compatible: https://api.minimax.io/v1 (mainland China: https://api.minimax.cn/v1)

Anthropic-compatible request (recommended by MiniMax):

curl https://api.minimax.io/anthropic/v1/messages \
  -H "Authorization: Bearer <MINIMAX_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-M3.1-Flash-Preview",
    "output_config": {"effort": "medium"},
    "max_tokens": 4096,
    "messages": [{"role": "user", "content": "Find the off-by-one bug in this loop: ..."}]
  }'

OpenAI-compatible request (Python):

from openai import OpenAI

client = OpenAI(base_url="https://api.minimax.io/v1", api_key="<MINIMAX_API_KEY>")

response = client.chat.completions.create(
    model="MiniMax-M3.1-Flash-Preview",
    reasoning_effort="medium",
    messages=[{"role": "user", "content": "Find the off-by-one bug in this loop: ..."}],
)

print(response.choices[0].message.content)

On the OpenAI-compatible protocol, the final answer is in content and the thinking is in reasoning_content. You do not need to strip <think> tags. MiniMax's examples use max. We changed them to medium because most coding requests do not need maximum depth.

Because the Anthropic-compatible endpoint is the recommended path, you can also point tools such as Claude Code or Cursor at it. MiniMax has integration guides for these tools in the Token Plan docs. If you are new to agentic coding tools, our Claude Code pricing breakdown is useful for comparing costs.

MiniMax M3.1 vs MiniMax M3

M3.1-Flash-PreviewMiniMax-M3
StatusPreview, September 27, 2026Generally available, June 1, 2026
Context1M (official)1M (official)
InputText, image, videoText, image, video
ThinkingAlways on, 5 effort levelsEnabled, adaptive or disabled (reported)
BenchmarksNone officialFull vendor table
WeightsNot publicHugging Face (MiniMaxAI/MiniMax-M3)
API accessToken Plan and MiniMax CodePublic API plus third-party hosts
PriceNot listedAbout $0.30 / $1.20 per 1M tokens (promo), $0.60 / $2.40 list, commonly cited
Best forDaily coding loopsLong-horizon agents, self-hosting, evals

Our recommendation: If you run MiniMax in production through an API, stay on M3 for now. It has a stable model ID, public pricing, third-party hosts and published scores. If you use MiniMax Code every day, compare Flash with M3 at medium and high effort on your real tasks. That test takes about an hour and tells you more than a leaderboard would.

Benchmarks: what exists and what does not

MiniMax has published no benchmark results for M3.1-Flash-Preview. BenchLM lists the model as unranked. Partner notes that leaked before launch say, in effect, that no 3.1 baselines exist and that M3 numbers should not be reused as targets.

So be careful when you read posts that credit "M3.1" with 59% on SWE-Bench Pro. That number is M3's, and MiniMax reported it itself. For context, these are the M3 figures:

Benchmark (MiniMax-M3 only, vendor-reported)Score
SWE-Bench Verified80.5%
SWE-Bench Pro59.0%
Terminal-Bench 2.166.0%
MCP Atlas74.2%
BrowseComp83.5
OSWorld-Verified70.06

Independent indices rate M3 lower than MiniMax's own tables. Artificial Analysis snapshots from September put its Intelligence Index around the high 20s to about 30. We have no public data on where a Flash version will land.

Run your own quick eval. Pick ten real tasks from your own repositories. Run each one on M3.1-Flash-Preview at medium and at high, and on M3 as a baseline. Track three things: whether the patch was complete, whether the model stopped in the middle of an edit, and how long you waited. Early users report that long diffs sometimes stop mid-edit and leave half-written imports, so the second measure matters most.

Pricing and plans

MiniMax has not published a per-token price for M3.1-Flash-Preview. Today you pay for it through the Token Plan or with MiniMax Code credits.

MiniMax Token Plan overview showing Plus at $22 per month, Max at $55 per month and Ultra at $132 per month, each with 5-hour rolling and weekly quota windows

Token Plan tierPriceBest for (per MiniMax)Agent usage
Plus$22 / monthPersonal projects and prototyping3 to 4 agents
Max$55 / monthDaily coding with agents and multimodal work4 to 5 agents
Ultra$132 / monthHeavy agent workflows and long sessions6 to 7 agents

All tiers use 5-hour rolling windows plus a weekly window. The quota is shared across eligible text, image and speech models. H3 video, voice design and some other special models are not covered by the included quota. When you reach the limit, purchased credits cover the rest. International docs price credits at 1,000 credits for $1, valid for 365 days. Some older posts still quote $20, $50 and $120 for the tiers. The docs show $22, $55 and $132.

Is MiniMax M3.1 free? You can use it without a subscription through free quota and check-in credits in MiniMax Code, and the check-ins pay double until October 7. There is no free, unlimited public API.

The only public per-token price in the M-series is for MiniMax-M3. Do not apply M3's price to M3.1.

Architecture: official vs leaked

Official. MiniMax has said three things: it is a multimodal coding model, it has a 1M context window, and it has tunable effort. That is all.

Leaked and unverified. A partner onboarding note dated about September 22 has circulated in a public evaluation repository and in summaries by OrcaRouter and AlphaSignal. It is not a MiniMax statement. According to those summaries, the changes from M3 include:

  • All-sparse attention. M3's first three full-attention anchor layers are replaced with sparse attention.
  • Q8KV4. Queries move from BF16 to FP8, and keys and values are stored as 4-bit values in blocks of 16. That would roughly halve KV-cache memory compared with M3.
  • NVFP4 experts. The routed MoE experts move toward W4A4 NVFP4. The shared expert is excluded.
  • DSpark speculative decoding. A Markov-style draft model replaces M3's EAGLE-style multi-token prediction.
  • Private checkpoints of about 250 GB and 236 GB on Hugging Face, which are not publicly accessible.

Each of these would make a model cheaper and faster to serve, which fits the "Flash" name. The same notes also report internal scores that disagree with each other, so they are not reliable. Treat all of this as engineering rumor until MiniMax publishes a card or paper.

What about Space Bunny Alpha?

About four days before the launch, an anonymous stealth model called stealth/space-bunny-alpha appeared on OpenRouter, free to use during its preview. Some developers fingerprinted it through tokenizer behavior, error messages and token-count probes, and several concluded that it came from MiniMax.

MiniMax has not confirmed this. On September 28, at least one public tester reported that Space Bunny scored differently from M3.1-Flash-Preview on their GameFeel v1.1 test. It may be an earlier checkpoint, a sibling model or something unrelated. For now, it is an interesting possibility that nobody has proven.

Who should switch, and who should wait

Indie hackers and solo founders. Try it. If you already use MiniMax Code or a Token Plan, switching costs nothing, and the double credits make this week a good time to test. Set effort to medium and raise it only when the model gets something wrong. If you are building a product on top of a model, you still need a stable one. See our guide to building an AI SaaS in 2026 for how to plan around that.

Agent platforms and tool builders. Wait, and prepare. You cannot route it through OpenRouter or another third-party host, and you cannot self-host it. Update your client so it never sends a thinking-disabled request, so you are ready when general API access opens.

Enterprise evaluation teams. Keep M3 in production. A preview model with no card, no license and no benchmarks cannot pass a procurement review. Add M3.1-Flash-Preview to your eval set so you have your own numbers when MiniMax publishes theirs.

Limitations and open questions

What MiniMax has not told us yet:

  • Total and active parameter counts.
  • Whether Flash is a distilled, quantized or separately trained checkpoint.
  • Output speed in tokens per second.
  • How quality compares with M3 on long-horizon agent tasks.
  • Whether a full, non-Flash M3.1 is coming.
  • Whether it will get open weights, and under which license.
  • When general pay-as-you-go API access will open.

The known limits are the preview label, always-on thinking, access only through Token Plan and MiniMax Code, mixed early reports on long patches, and some confusion with Space Bunny Alpha.

Bottom line

MiniMax-M3.1-Flash-Preview is a product release for everyday coding, not a documented successor to M3. The confirmed parts are useful: 1M context, multimodal input and an effort control that gives you real speed choices once you move off the max default. Before you replace M3 in production, wait for three things: a model card, a public API model ID and independent benchmarks. Until then, test it in MiniMax Code, keep M3 for production, and do not trust anyone who quotes M3.1 benchmark scores.

If you are building with these models and plan to launch, list your product on SaaSCity to get in front of other founders who use the same tools.

Sources: MiniMax Model Invocation docs, MiniMax Token Plan overview, @MiniMaxAgent launch post, MiniMax-M3 on Hugging Face and the MiniMax M3 announcement. Leak details come from secondary partner-note summaries and are labeled as unverified.

Get your SaaS in front of founders

List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.

Submit your SaaSSee pricing

Founder resources

Best SaaS directoriesBest AI directoriesFree dofollow directoriesHigh-DR directoriesFree DR checkerLive launchesAI SaaS boilerplate

Related articles

MiniMax M3 Review: The First Open-Weight Model to Do Frontier Coding, 1M Context, and Multimodality All at Once

MiniMax M3 Review: The First Open-Weight Model to Do Frontier Coding, 1M Context, and Multimodality All at Once

Kimi-K2.7-Code Drops: Moonshot AI's Strongest Open-Source Coding Model Yet (+21.8% on Kimi Code Bench v2)

Kimi-K2.7-Code Drops: Moonshot AI's Strongest Open-Source Coding Model Yet (+21.8% on Kimi Code Bench v2)

Headroom: LLM Token Compression for RAG Chunks, Tool Outputs, and Log Files

Headroom: LLM Token Compression for RAG Chunks, Tool Outputs, and Log Files

Contents

  1. Key takeaways
  2. What is MiniMax M3.1?
  3. MiniMax M3.1 release date and availability
  4. Official specs
  5. Effort levels explained
  6. How to use MiniMax M3.1 today
  7. MiniMax M3.1 vs MiniMax M3
  8. Benchmarks: what exists and what does not
  9. Pricing and plans
  10. Architecture: official vs leaked
  11. What about Space Bunny Alpha?
  12. Who should switch, and who should wait
  13. Limitations and open questions
  14. Bottom line

List your SaaS

$19.99one-time
  • Dofollow DR 64+ backlink
  • Live within 24 hours, no queue
  • Permanent listing on the city map
Submit your SaaS

Or list free with our badge

City Sponsors

  • Nick LaunchesShip, launch, and get your product in front of real founders.
  • @peregrineintellPeregrine OS: pre-call intel for agency new business
  • Your product hereSlot open — 30 days, homepage + city
Become a sponsor
Write for this blog — from $99.99
SaaSCity.io

Directories are boring. We built a city instead. First isometric SaaS directory on the planet.

Platform
Submit SaaSLive LaunchesPricingBlogWrite for UsBacklink ExchangeMCP for AgentsAdvertise
Directories
Best SaaS DirectoriesBest AI DirectoriesBest Indie Hacker CommunitiesBest Subreddits for FoundersFree DR CheckerFree DR BadgeHow to Get SaaS Backlinks
SaaSCity Alternatives
All ComparisonsSaaSCity vs Nick LaunchesSaaSCity vs BetterLaunchSaaSCity vs PeerPushProduct Hunt AlternativesSaaSHub Alternatives
Legal
Privacy PolicyTerms of Service
Company
AboutghostyContact

© 2026 SaaSCity.io

llms.txt