Skip to main content
SaaSCity.io
DirectoriesLive LaunchesBlogWrite for UsAdvertise
Submit
Home/Blog/Mistral Large 4 ("Le Chonk") Is a 1T Open-Weight Model at $1.36/M Tokens (2026)
Back to Blog

News

Mistral Large 4 ("Le Chonk") Is a 1T Open-Weight Model at $1.36/M Tokens (2026)

Mistral just shipped a model nicknamed after a fat cat, and Hacker News gave it nearly 2,000 points in a day. Past the meme, there's a real build-vs-buy decision for anyone running inference in production: the open weights land by the end of October, and a trillion parameters doesn't get cheap to host just because the weights are open. Here's what the numbers mean for a SaaS founder's infra bill.

ghosty
ghosty
Founder, SaaSCity
October 7, 202612 min read
Mistral Large 4 ("Le Chonk") Is a 1T Open-Weight Model at $1.36/M Tokens (2026)
Contents (8)
  1. What actually shipped on October 6
  2. The number that matters isn't 1 trillion, it's about 50 billion
  3. What it costs, and what that actually buys you
  4. Open weights land by the end of October: here's what "open" actually costs you
  5. Where it's actually strong, and where it isn't
  6. The sovereignty angle is a real sales feature, not just PR
  7. The naming lesson hiding in the HN numbers
  8. So what do you actually do with this

Quick answer: Mistral Large 4 ("Le Chonk") launched as a public preview on October 6, 2026: an open-weight mixture-of-experts model that Mistral's announcement puts at 1 trillion parameters with 49 billion active per token, and that its docs card lists at 1.05 trillion with 52 billion active. It has a 1M-token context window and takes text and images. It's API-only today at $1.36 per million input tokens and $4.18 per million output tokens, with a sale price of half that on the docs card. The weights land by the end of October, under a license Mistral hasn't published yet. For founders, the real story isn't the trillion-parameter headline. It's what the active-parameter count does to your API bill, what the total parameter count does to a self-hosting bill, and where ML4 fits in a model-routing setup next to the closed frontier.

You're reading this on a directory's blog, so here's the disclosure up front: SaaSCity is a gamified startup directory with a live city map and human editorial review. We write posts like this one because founders picking a model stack are exactly the people who end up listing a product with us a few weeks later. More on that at the bottom. Now, the model.

A French startup named its trillion-parameter model after a fat cat, and the internet lost it. That's not a knock; it's the most interesting data point in this whole launch, and we'll get to why in a minute. First, the model itself, because underneath the meme there's a useful release for anyone making inference decisions right now.

What actually shipped on October 6

Mistral AI put out Mistral Large 4 as a public preview, "unofficially ML4, very officially: le Chonk," in its own words. The headline spec is a trillion parameters. The spec that actually matters for your bill is the active parameter count, and here Mistral's own pages disagree: the announcement says 49 billion active, while the docs model card says 52 billion active out of 1.05 trillion total. If you see both numbers floating around, that's why.

Mistral's announcement page for Mistral Large 4, headed Le chonk, Introducing Mistral Large 4, dated October 6, 2026

The rest of the spec sheet: a sparse mixture-of-experts (MoE) architecture, a 1.6-billion-parameter vision encoder for native multimodal input, a 1-million-token context window and training data in more than 160 languages, including every official EU language. It takes text and images in and outputs text. Mistral's model card calls it "an open-weight hybrid instruct-and-reasoning MoE with multimodal input," which is a mouthful but accurate. In the API you pick a reasoning effort, and Simon Willison found on Hacker News that only "none" and "high" are supported, with "high" producing fewer output tokens than "none" in his testing.

Mistral's docs model card for Mistral Large 4, with the description framed: 52B active parameters, 1.05T total parameters and a 1.6B vision encoder, above a 1M context window and the sale prices

Here's the part that doesn't make the headline but should: Mistral trained this from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters, and the preview API is served from that same infrastructure. That's a large training run for a company that was a research lab with a Discord server three years ago. Mistral also says the reinforcement-learning run behind the preview is "still in flight," so the model you test today is not the final one.

The number that matters isn't 1 trillion, it's about 50 billion

Every headline about ML4 leads with "1 trillion parameters" because it's the bigger, scarier-sounding number. For your API invoice, it's close to irrelevant.

Mixture-of-experts models don't activate every parameter for every token. ML4 routes each token through a subset of specialized expert networks, and only about 49 to 52 billion parameters do work on any given token. The rest sit there as a much larger pool of specialized knowledge the router can draw from, but they're not all firing at once the way they would in a dense model.

Practically, that means ML4's compute per token behaves a lot more like a 50B dense model than a true 1T one. That's why Mistral can price it where it does, and it's the same trick DeepSeek, Qwen and most other frontier-scale open models use: massive total capacity, modest active compute per token. If you've been following the open-weight economics story with Xiaomi's MiMo, this is the same pattern at a much bigger scale.

There's a catch, and it's the one most "1T but only 49B active" takes skip: active parameters cut compute, not memory. Every expert has to be loaded and ready, because the router can send the next token anywhere. Hold that thought for the self-hosting section.

What it costs, and what that actually buys you

The preview is live today on Mistral Studio, with the model name mistral-large-4 in the API. Pricing, per Mistral's announcement and docs card as of October 7, 2026:

Mistral Large 3Mistral Large 4 (list)Mistral Large 4 (docs sale price)
Input (per million tokens)$0.50$1.36$0.68
Output (per million tokens)$1.50$4.18$2.09
Cached input (per million tokens)Not listed$0.14$0.07

At list price, that's roughly 2.7x the input cost and 2.8x the output cost of Mistral's previous flagship. Not a small jump. The docs card currently shows a sale price of exactly half the list price, with no end date given, so check it before you budget around it. Against the closed frontier, ML4 still lands cheap on input, with output pricing in a comparable range rather than a bargain-bin one. If you want the closed-frontier numbers to compare against, our Claude Code pricing breakdown has the current rate card.

To make it concrete, here is one month of a mid-size production workload, 100 million input tokens and 20 million output tokens, by our arithmetic:

Monthly bill for 100M input + 20M output tokensCost
Mistral Large 3$80.00
Mistral Large 4 at list price$219.60
Mistral Large 4 at list price, 80% of input cached$122.00
Mistral Large 4 at the sale price, 80% of input cached$61.00

The cached-input rate is the detail worth building around if you do anything with repeated system prompts or long static context: RAG pipelines, agent scaffolding, anything with a stable prefix. It's a 90% discount off the base input rate, and in the example above caching takes $97.60 off the list-price bill. That's where a lot of the real savings in a production deployment come from, whichever model you pick.

While you are here

Get your SaaS listed on SaaSCity

A permanent listing on the live city map, a DR 66+ dofollow backlink and a launch week in front of founders. Free with a badge, or skip the queue with Quick Pass — live within 24 hours.

Submit your SaaSWhat you get

Open weights land by the end of October: here's what "open" actually costs you

This is the part that gets glossed over in every "1 trillion parameter open model!" headline: the weights aren't out yet. What's live today is a hosted preview. Mistral says it will release the weights "by the end of the month," along with more detail on the architecture and its post-training method, and that until then it is red-teaming the model with cybersecurity leaders, vetted partners and state authorities. Its docs list both the weights and the license as "coming soon."

When the weights land, here's what you're signing up for, by our own arithmetic on 1.05 trillion parameters (weights only, before memory for the KV cache and runtime):

PrecisionWeights in memory80GB GPUs needed for the weights alone
BF16about 2.1 TB27 or more
8-bitabout 1.05 TB14 or more
4-bitabout 525 GB7 or more

That's why "only 49B active" doesn't translate into "runs on one GPU." Even a 4-bit build needs a multi-GPU node before you've served a single long-context request. For a sense of scale, a node of eight 192GB Blackwell cards has about 1.5 TB of memory: enough for an 8-bit build with room for the KV cache, not for BF16. Mistral hadn't published official hardware figures as of October 7 (its docs show the GPU RAM column as N/A), so treat these as order-of-magnitude numbers until the weights and an official serving guide exist.

Then there's the license. Mistral Large 3 shipped under Apache 2.0, which made it trivial to build a commercial product directly on the weights. Mistral hasn't said which license ML4 will use, so anyone planning to self-host for a commercial product should read the actual license text when it lands before assuming Large-3-style freedom.

So the honest framing for a founder: "open weight" buys you the option to self-host, eventually, with serious hardware and a license you need to check. It does not buy you "cheap" by default. If your monthly API spend is in the hundreds or low thousands of dollars, which the table above suggests covers a lot of real workloads, renting a multi-GPU node to self-host isn't going to pencil out. If you're running inference at a scale where GPU rental is already cheaper than API calls, or you need the model inside your own perimeter for compliance, ML4 becomes a real option once the weights and a quantized build both exist. Our piece on open-weight model economics walks through that breakeven math in more detail.

Where it's actually strong, and where it isn't

Mistral's positioning is specific. It calls ML4 competitive with the strongest open models globally and well ahead of any open-weight model developed in the US or Europe, and state-of-the-art among open models on cybersecurity, finance and law. Some of the numbers it publishes:

  • Cybersecurity: among the top five models on the Artificial Analysis Cyber Index, 82% on a test that asks a model to reproduce a real vulnerability and then patch it, the highest of any model, and 93% on Cybench. Mistral notes that several closed models score near zero on that vulnerability test because they refuse the task, which is its pitch to security teams that want a model they can run under their own policies.
  • Visual grounding: 42% on Dense 200 against GPT-6 Astra's 41%, which tracks with the dedicated vision encoder.
  • Agents: 59.9% on AutomationBench, 657 business workflows across apps like Gmail, Sheets, Slack and Salesforce, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.
  • Prompt injection: it resists 93.3% of attacks on Lakera's public B3 benchmark, the highest score Mistral says it saw.

Hacker News front-page thread on Mistral Large 4 showing 1,894 points and 1,139 comments in under a day, including discussion of the model's reasoning-mode limitations

Coding is where the picture gets more careful. ML4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, and its combined Coding Agent Index of 49.8% puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. That's a real result given how much ground Chinese open-weight labs have covered this year (see our rundown of Kimi K3 vs. Qwen3.8 Max for that side of the race). But in Mistral's own blind human evaluation with Surge AI, where professional annotators scored coding output from 1 to 5 without knowing which model wrote it, ML4 came second of five at 3.74, behind Claude Opus 5 at 4.22. And against GLM-5.3, Mistral's own expert raters called the two on par or close on coding.

Mistral's announcement, Agentic coding section, with the paragraph framed: 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4 and a Coding Agent Index of 49.8%

For a SaaS founder, that's useful clarity, not a disappointment. If you're building an agentic coding tool, ML4 is a routing candidate: send cheaper, well-defined subtasks to it and keep your hardest coding problems on a frontier closed model. It's not a wholesale swap-in. That's the whole thesis behind thinking in terms of model fatigue and cost-per-task routing instead of picking one model and hard-coding it everywhere. A strong-but-not-best model dropping every few weeks isn't a reason to chase each one; it's a reason to build your inference layer so swapping is cheap.

The sovereignty angle is a real sales feature, not just PR

It's easy to read "trained in European datacenters, fluent in every EU language" as box-checking. For a specific category of buyer, it's the actual purchase decision.

Enterprise and public-sector procurement in the EU increasingly asks where a model runs, where the data goes and under what jurisdiction. Mistral says ML4 will be available in several regions worldwide, including a European deployment it operates end to end, "independently of other digital service providers and under European law." That's a different answer to a procurement questionnaire than "it runs on a US hyperscaler." Mistral's enterprise logo list (ASML, Cisco, Snowflake, CMA CGM, HSBC) reads like a list of companies that care about exactly that question.

Mistral AI homepage banner declaring Mistral #1 in Sovereign AI, promoting Mistral Large 4 alongside enterprise customer logos including ASML, Cisco, Snowflake, CMA CGM and HSBC

That's not incidental to the launch timing either. Mistral closed a €3 billion Series D led by Samsung Electronics in September 2026, which it calls the largest equity round ever raised by a European technology company, and its announcement says ML4 is the first milestone that money funds. If your product sells into regulated industries, government contracts or any EU enterprise with a data-residency clause in procurement, "which model actually runs on European infrastructure" is a question your sales team will eventually get asked, and having an answer beats not having one.

The naming lesson hiding in the HN numbers

One more thing worth sitting with, separate from the specs. ML4 launched one day after Reflection AI's Beam, another serious open-weight release that week. Beam didn't get anywhere close to the same reaction. The Hacker News thread for "Le Chonk" had 1,968 points and 1,173 comments a day after it was posted, for a model release, in a week that already had a competing model release in it.

A spec sheet doesn't do that. A memorable name attached to a strong result does. If you're an indie hacker naming your own product, that's the takeaway: distribution isn't just what you built, it's whether anyone can say the name out loud and remember it the next morning. "Mistral Large 4" would have been a solid, forgettable launch. "Le Chonk" is why you're reading a blog post about it on October 7.

So what do you actually do with this

If you're already paying for a closed frontier model and your workloads are latency-sensitive or coding-heavy, nothing changes today. ML4 is worth benchmarking against your current provider on cost for the subtasks that don't need frontier coding quality, and the sale price makes that test cheap, but it's not a drop-in replacement.

If you're EU-based, selling to EU enterprise, or your buyers ask data-residency questions in procurement, ML4 is worth a real evaluation now, preview pricing and all, because the sovereignty story is the actual product here, not a footnote.

If you're weighing self-hosting, wait until the weights are actually out and someone has shipped a quantized build. Read the license before you commit infrastructure budget, size the node from the memory table above rather than the "49B active" headline, and run the GPU-rental-versus-API math honestly before assuming "open" means "cheaper."

And if you're building any of this in public (a model router, a coding agent, a wrapper around ML4 or any of the dozen models that dropped this quarter), that's exactly the kind of launch SaaSCity exists for. A free listing gets a permanent page and a building on our live city map, takes the first Monday with an open launch slot, and its link turns dofollow once our badge is on your homepage. Quick Pass ($19.99) skips the queue and the badge and goes live within 24 hours after a human review, and Premium ($39.99) adds a launch post we write about your product, with three links to your site. SaaSCity sits at DR 66 on Ahrefs as of October 2026. Submit your product to SaaSCity whenever you're ready, and if it's an AI tool, our AI directory list and AI launch guide cover where else to list it.

Model releases are going to keep coming at this pace through the rest of 2026. The founders who win aren't the ones chasing every new "best open-weight model" headline. They're the ones who built a swap-it-out inference layer six months ago and are spending this week on product instead of re-benchmarking.

Get your SaaS in front of founders

List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.

Submit your SaaSSee pricing

Founder resources

Done-for-you directory submissionBest SaaS & startup directoriesFree dofollow directoriesBest AI directoriesHigh-DR directoriesFree DR checkerLive launches

Related articles

Xiaomi Just Open-Sourced a Frontier-Class Model for 1/50th the Price (MiMo-V2.6, 2026)

Xiaomi Just Open-Sourced a Frontier-Class Model for 1/50th the Price (MiMo-V2.6, 2026)

Qwen3.8-Flash-Next: Alibaba Is Shipping the Qwen4 Architecture Early, and It Runs on 6B Active Parameters

Qwen3.8-Flash-Next: Alibaba Is Shipping the Qwen4 Architecture Early, and It Runs on 6B Active Parameters

Lovable Just Crossed $600M ARR. The Vibe-Coding Business Model Is the Real Story (2026)

Lovable Just Crossed $600M ARR. The Vibe-Coding Business Model Is the Real Story (2026)

Contents

  1. What actually shipped on October 6
  2. The number that matters isn't 1 trillion, it's about 50 billion
  3. What it costs, and what that actually buys you
  4. Open weights land by the end of October: here's what "open" actually costs you
  5. Where it's actually strong, and where it isn't
  6. The sovereignty angle is a real sales feature, not just PR
  7. The naming lesson hiding in the HN numbers
  8. So what do you actually do with this

List your SaaS

$19.99one-time
  • Dofollow DR 66+ backlink
  • Live within 24 hours, no queue
  • Permanent listing on the city map
Submit your SaaS

Or list free with our badge

Done for you

We submit your SaaS to up to 150 directories

Every form filled in by us, every live URL in a report. Three packages.

from$49one-time

Pick your project

City Sponsors

  • Nick LaunchesShip, launch, and get your product in front of real founders.
  • @peregrineintellPeregrine OS: pre-call intel for agency new business
  • Your product hereSlot open: 30 days, homepage + city
Become a sponsor
Write for this blog, from $99.99
SaaSCity.io

Directories are boring. We built a city instead. First isometric SaaS directory on the planet.

Platform
Submit SaaSLive LaunchesPricingBlogWrite for UsBacklink ExchangeMCP for AgentsAdvertise
Directories
Best SaaS DirectoriesBest AI DirectoriesDirectory ReviewsBest Indie Hacker CommunitiesBest Subreddits for FoundersFree DR CheckerFree DR BadgeHow to Get SaaS BacklinksSEO GlossaryDirectory Submission Service
SaaSCity Alternatives
All ComparisonsSaaSCity vs Nick LaunchesSaaSCity vs BetterLaunchSaaSCity vs PeerPushProduct Hunt AlternativesSaaSHub Alternatives
Legal
Terms of ServicePrivacy PolicyRefund PolicyCookie PolicyCopyright & DMCASecurity
Company
AboutghostyContact

© 2026 SaaSCity.io

llms.txt