Skip to main content
SaaSCity.io
Browse MapLive LaunchesBlogWrite for UsAdvertise
Submit
Home/Blog/SWE-2 Is Free on Devin's $20 Plan: The Best AI Coding Deal in September 2026
Back to Blog

Industry News

SWE-2 Is Free on Devin's $20 Plan: The Best AI Coding Deal in September 2026

Cognition shipped SWE-2 on September 10, 2026. It scores 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1, at what Cognition says is 64% lower cost. The same day, Cognition made it free for every Pro, Max and Teams subscriber for a month. Here is what the numbers say, what the $20 seat actually buys, and who should still pay $200.

ghosty
ghosty
Founder, SaaSCity
September 10, 202614 min read
SWE-2 Is Free on Devin's $20 Plan: The Best AI Coding Deal in September 2026
Contents (11)
  1. What shipped on September 10, 2026
  2. The scoreboard, vendor-reported
  3. The efficiency story is better than the 50% headline
  4. What the model does differently, per Cognition and per early users
  5. The $20 plan, without the marketing fog
  6. Why this is the best $20 coding deal in September 2026
  7. How to actually use the promo
  8. Company context, kept short
  9. Objections, answered
  10. If you are shipping a developer tool into this market
  11. Decide, then move

Verified against cognition.com/blog/swe-2, the @cognition launch thread, devin.ai/pricing and docs.devin.ai on September 10, 2026, the day of the launch. The promo terms and the model selector may change within days. Every benchmark number below is vendor-reported by Cognition.

Cognition just put a 50-percent FrontierCode model on the $20 Devin plan and told Pro users it will not count against quota for a month.

That is the whole story, and it is a bigger one than the score. On September 10, 2026 Cognition released SWE-2, its most capable in-house coding model. Its official number is 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1 (50.9%) and a few behind GPT-6 Astra (53.3%), at what Cognition says is 64% lower cost than Fable and about a quarter of Astra's. The same afternoon, the company tweeted that SWE-2 is free for all Pro, Max and Teams subscribers for the next month. Growth lead Alex Kaplan put it more bluntly: "Unlimited access to SWE-2 for all pro+ plans (starting at $20/month)."

If you already pay $20 a month for an AI coding tool, today is the day the price-performance chart moved.

Cognition's SWE-2 announcement post, 'Introducing SWE-2: Pushing the Pareto Frontier', dated 09.10.26 and opening with the 50.0% FrontierCode 1.1 Main claim within one point of Fable 5.1 at 64% lower cost

What shipped on September 10, 2026

The official post is Introducing SWE-2: Pushing the Pareto Frontier, by The Cognition Team. The launch tweet from @cognition reads: "Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost." It passed 768,000 views within hours.

The facts that matter for a buyer:

  • Available today in Devin Desktop and Devin CLI. Rolling out to Devin Web and Fusion.
  • Not open weights. Not a public API. SWE-2 only runs inside Devin's harness.
  • Built on Kimi K3, a 2.8 trillion parameter open-weight model. First time Cognition ran its RL recipe at multi-trillion scale.
  • Three effort levels, medium, high and max, trained in a single RL run.

The same day brought Devin Voice, which Cognition says runs on GPT-Live plus SWE-2, and the Dioxus team joining the company. Two days earlier, on September 8, Cognition closed a Series E of more than $2 billion at a $48 billion valuation. We covered what that round means for smaller SaaS teams already. None of that changes the model. It does explain why a company can afford to give it away for a month.

The scoreboard, vendor-reported

FrontierCode is Cognition's own benchmark. It scores solutions on a weighted rubric of quality and mergeability, and anything that fails a blocking criterion scores zero. DeepSWE 1.1 and the two Terminal-Bench versions are more external. Read the table with that in mind.

BenchmarkSWE-2Kimi K3 (base)Grok 4.6Fable 5.1GPT-5.6 SolGPT-6 AstraSWE-1.7
FrontierCode 1.1 Main50.0%44.2%48.0%50.9%47.5%53.3%42.0%
DeepSWE 1.173.0%68.5%67.5%67.4%72.7%74.1%37.7%
Terminal-Bench 2.192.8%88.3%88.4%91.4%88.8%89.9%81.5%
Terminal-Bench 427.3%21.5%20.3%55.8%37.3%57.9%7.6%

Source: Cognition, SWE-2 launch post, September 10, 2026. The July SWE-1.7 post listed 42.3% on FrontierCode; the SWE-2 table lists 42.0%. We use the newer table.

Cognition's coding benchmark results table for SWE-2 against Kimi K3, Grok 4.6, Fable 5.1, GPT-5.6 Sol, GPT-6 Astra and SWE-1.7 across FrontierCode 1.1 Main, DeepSWE 1.1, Terminal-Bench 2.1 and Terminal-Bench 4

How to read it honestly:

  • SWE-2 is not number one. Astra wins FrontierCode and DeepSWE. Fable 5.1 and Astra both roughly double SWE-2 on Terminal-Bench 4.
  • SWE-2 does win Terminal-Bench 2.1 at 92.8%. That is the "stay in the shell until the tests pass" eval, and it maps better to an everyday agent handoff than Terminal-Bench 4 does.
  • Against its own previous generation it is up 8 points on FrontierCode, 35 on DeepSWE, 11 on Terminal-Bench 2.1 and 20 on Terminal-Bench 4.
  • Against its base model the RL adds 5.8 points on FrontierCode, 4.5 on DeepSWE, 4.5 on Terminal-Bench 2.1 and 5.8 on Terminal-Bench 4. Cognition's own line: "our RL still finds substantial headroom, adding 5–6 points on many benchmarks."
  • The cost claims are Cognition's. Within one point of Fable 5.1 at 64% lower cost, within a few points of Astra at roughly 25% of the cost. The tweet says "up to 70% lower", which is more aggressive than the blog's 64%. Nobody outside Cognition has audited either figure.

SWE-2 vs GPT-6 Astra

Astra is the better model on paper. It leads FrontierCode by 3.3 points, DeepSWE by 1.1 and Terminal-Bench 4 by 30.6. SWE-2 leads Terminal-Bench 2.1 by 2.9. The argument for SWE-2 is not capability, it is that Cognition says you get that scoreline at a quarter of the price, and this month at no quota cost at all on Devin Pro. If you can afford Astra on every task, run Astra. Most people on a $20 plan cannot.

SWE-2 vs Fable 5.1

This is the closer race. One point apart on FrontierCode, SWE-2 ahead on DeepSWE and Terminal-Bench 2.1, Fable more than double on Terminal-Bench 4. Fable is the model you want for long, hostile terminal sessions. SWE-2 is the model you want for the ticket queue. Cognition's Pareto chart makes the trade visible: SWE-2's three effort points sit to the left of everything else at the same score.

Cognition's FrontierCode 1.1 Main score versus average cost per rollout chart, showing SWE-2's medium, high and max effort points sitting left of Fable 5.1, GPT-6 Astra, Grok 4.6, GPT-5.6 Sol and Kimi K3 at similar scores

The efficiency story is better than the 50% headline

The last Devin model read half the repo before touching a file. Users told Cognition so, and the SWE-2 post owns it: SWE-1.7 "tended to over-explore and overthink simple tasks." SWE-2's fix is what Cognition calls focused exploration, and the numbers on FrontierCode 1.1 Main (mean over 100 tasks, three runs each) are stark.

ConfigurationMean steps per taskFirst real edit (median step)
SWE-1.712748
SWE-2 medium5318
SWE-2 high80—
SWE-2 max98—

SWE-2 medium scores higher than SWE-1.7 with 58% fewer turns and 81% lower cost. It makes its first real edit at step 18 instead of step 48. That is the difference between an agent you hand a ticket to and an agent you babysit.

Cognition's bar chart of mean steps per run on FrontierCode 1.1 Main: SWE-1.7 at 127, SWE-2 medium at 53, SWE-2 high at 80 and SWE-2 max at 98, broken down by explore, plan, edit, build, test, commit and final message steps

The mechanism, in one paragraph: they penalized the model for wasting money, then trained medium, high and max in one run so cheap mode got smarter instead of just shorter. The reward is success minus a cost penalty, where cost is the real inference spend plus time for the rollout. The penalty coefficient is set per effort level to the local slope of the base model's Pareto frontier at that effort. The stated goal is to move the whole cost-performance curve up without collapsing high effort into medium.

Two more mechanics for the engineers who want them:

  • Length-weighted reward baseline, used since SWE-1.6, as a cheaper proxy for the variance-reducing baseline. Cognition says it keeps train-inference divergence in check.
  • Serving work that made a model roughly three times larger than SWE-1.7's base run at similar throughput: a prefill delayer worth 10 to 20% on tokens per second, speculative decoding with an online-trained draft model that lengthened accepted runs by about 15%, and NVFP4/FP8 kernels with quantization-aware training. Three times more RL environments than SWE-1.7, plus a flywheel where earlier SWE-2 checkpoints harden the verifiers.

While you are here

Get your SaaS listed on SaaSCity

A permanent listing on the live city map, a DR 63+ dofollow backlink and a launch week in front of founders. Free with a badge, or skip the queue with Quick Pass — live within 24 hours.

Submit your SaaSWhat you get

What the model does differently, per Cognition and per early users

Cognition's internal QA notes, labeled as such:

  • Better end-to-end tests. Catches regressions and edge cases more often.
  • Resourceful within its permissions. Their example: an MCP connector was down, so the agent reconstructed the data it needed from Slack history it already had access to.
  • Verification discipline. Ask "are you sure?" and it re-runs or re-derives instead of agreeing. Cognition's phrase is "a model whose conclusions you can trust."
  • Effort levels are real, not skins. Medium acts fast and cheap. High and Max add planning, repository coverage and verification.

Launch-day practitioner reports from X, which are anecdotes and should be read as such:

  • @BrahmaD111 ran three hard tickets on Max effort, reported "no bla bla" and only small nits after an Astra-medium review, and contrasted it with Opus 5 scope-creeping. First impression: "claude killer."
  • @surim0n's summary of the blog, which circulated widely: "the last version read half the codebase before touching anything. this one starts editing after 18 steps… it got docked points in training for wasting money."
  • @gaogezh: "Fast and smart. It is free and unlimited for the next month for pro users ($20/month)."
  • @Da7_Tech, a Cognition ambassador: "Even on the $20 plan, you can use it for free for an entire month, and you won't need another model… Is it the best $20 subscription today? Yes."
  • The skeptics were there within the hour too: "makes me wonder if its benchmaxxed."

Treat the ambassador quote as marketing and the skeptic as a fair question. Both are answered the same way: run it on your own repository while it is free.

The $20 plan, without the marketing fog

Here is what devin.ai/pricing showed on launch day. Note the Pro card still says "Free use of SWE 1.7". The pricing page had not been rewritten in the first hours.

Devin pricing page on September 10, 2026 showing Free at $0, Pro at $20 per month with 'Free use of SWE 1.7 and leading open source models', and Max at $200 per month

PlanPriceWhat matters here
Free$0Light quota, limited models, unlimited Tab and inline edits. Not the SWE-2 deal.
Pro$20/moIncreased daily and weekly quota, frontier OpenAI, Claude and Gemini models, Devin Cloud, free use of the SWE family. Single user.
Max$200/moPro plus a much larger weekly quota, no daily cap, unlimited concurrent sessions.
Teams$80/mo minimum, $40 per full seatShared billing, collaboration, admin. Full seats get Pro-like quota.
EnterpriseCustomACU billing, SSO, dedicated deployment.

How quota actually works, from docs.devin.ai:

  • Plans include a daily and a weekly token budget. The daily allowance is more than a seventh of the weekly one, so a heavy weekend does not wipe the week.
  • "Free models don't count against your quota at all." That is the direct quote, and it is the whole mechanism behind "unlimited SWE-2."
  • Past quota, you buy on-demand usage at API list prices for whichever model you picked.

The promo, quoted exactly:

Cognition: "SWE-2 is available today in Devin across Desktop and CLI. We're making it free for all Pro, Max & Teams subscribers for the next month."

Alex Kaplan, Growth at Cognition: "Unlimited access to SWE-2 for all pro+ plans (starting at $20/month)! Great deal"

What happens after the month is not written down anywhere. SWE-1.5 and SWE-1.7 both stayed free on Pro after their launches, and the Desktop marketing already said "Unlimited access to SWE-1.7." That is a strong prior. It is not a contract. Plan for unlimited through mid-October and re-check the model selector and the quota page when the month ends.

Why this is the best $20 coding deal in September 2026

Compare like a buyer. What $20 buys elsewhere this month, approximately and not audited apples-to-apples:

  • Claude Pro, $20. Claude Code and chat share a five-hour window plus a weekly cap. Daily agentic work pushes people to Max 5x at $100 or Max 20x at $200. Fable-class models eat weekly limits fast.
  • ChatGPT Plus, $20. Codex lives in five-hour windows. Early-September community estimates put Astra at somewhere between 5 and 45 local messages per window on Plus. Frontier models are metered hard.
  • Cursor Pro, $20. A pool of first-party usage plus about $20 of third-party model credit, then pay-as-you-go. Composer and Grok-class models stretch it. Fable and Astra burn it.
  • GitHub Copilot Pro, $10. Cheap, and not in the same autonomy class as Devin Cloud plus SWE-2.
  • Devin Pro, $20, this month. Frontier third-party models still burn quota. SWE-2 does not. Plus Devin Cloud agents, Desktop, CLI, unlimited Tab, and Slack, Linear and MCP integrations.

The honest claim is not "SWE-2 beats Astra." The claim is that for $20 this month you can run a 50-FrontierCode, 73-DeepSWE, 93-Terminal-Bench-2.1 agent all day inside a mature harness without watching a five-hour clock. That combination does not exist on Claude Pro or ChatGPT Plus right now.

Who should still pay $100 to $200:

  • Your work looks like Terminal-Bench 4. Multi-hour, messy environments where Fable 5.1 and Astra score 55 to 58% against SWE-2's 27%. That gap is not a rounding error.
  • You need the model outside Devin. Cursor-only workflow, Claude Projects, ChatGPT canvas, your own API calls. SWE-2 is not available there.
  • You want max-effort Astra review on every PR anyway. Then use SWE-2 to implement and Astra or Fable to review. Several early testers were already doing exactly that on day one, and it is the cost trade the Pareto chart describes.

How to actually use the promo

  1. Subscribe to Pro at devin.ai/pricing for $20. Grandfathered $15 Windsurf and Devin Pro users keep $15.
  2. Start in Devin CLI or Devin Desktop. That is where SWE-2 is live on day one. Web and Fusion are rolling out.
  3. Pick effort deliberately. Medium for tickets and refactors. High or Max for migrations, ambiguous bugs and multi-repo work.
  4. Keep a frontier model in the selector for review. Astra, Fable 5.1 and Sol all burn quota, so use them where a second opinion is worth paying for.
  5. Do not plan around API access or open weights. There is no SWE-2 endpoint outside Devin.
  6. Put the end date in your calendar. "Next month" from September 10 lands around mid-October.

The surfaces, for anyone who lost track of the renames:

  • Devin Desktop is the Windsurf editor, rebranded on June 2, 2026. Agent command center, Tab completions, multi-agent.
  • Devin CLI is the local terminal agent, with a handoff command to push work to the cloud.
  • Devin Cloud / Web is the async autonomous engineer: Linear or Jira ticket in, pull request out.
  • Fusion is Cognition's dual-agent cost cutter, a frontier planner driving a cheap executor. SWE-2 is rolling in as the executor.

Company context, kept short

Cognition makes Devin, launched in March 2024 as "the first AI software engineer." It bought Windsurf in mid-2025 and folded it into Devin Desktop this June. On September 8, 2026 it raised more than $2 billion at a $48 billion post-money valuation, led by a16z and Accel, with run-rate revenue reported at $492 million in May and about $900 million now. Named customers include NVIDIA, GE Aerospace, Citi, Mercedes-Benz and Modal.

The SWE line runs 1.5, 1.6, 1.7 in July 2026 on a Kimi K2.7 base at 42% FrontierCode, and now SWE-2 on Kimi K3. The strategic read is simple. Cognition trains its own model so the $20 plan can include a near-frontier agent without paying Anthropic or OpenAI retail on every token. The free month is customer acquisition plus a data flywheel. It is not charity, and it does not need to be.

Objections, answered

"FrontierCode is their benchmark." Yes. That is why the table above carries DeepSWE and both Terminal-Bench versions next to it. Never judge SWE-2 on FrontierCode alone.

"Terminal-Bench 4 is a blowout loss." It is. 27.3% against 55 to 58%. This is not the model you send into the nastiest long-horizon terminal gauntlets, yet. Medium-horizon agent work is the fit.

"Benchmaxxed." Possible, and unprovable from outside. The behavioral numbers (18 versus 48 steps to first edit, 81% cheaper than SWE-1.7, does not cave when asked "are you sure") are more useful than a 50.0 headline. The right response is to run it on your repository this month while it costs nothing.

"Unlimited forever?" No. The official language is one month for Pro, Max and Teams. History says SWE models stay free on Pro. History is not a contract.

"The docs still list SWE-1.7." Launch-day lag. The pricing page and the models documentation were not updated in the first hours. The Cognition tweet and blog are the source of truth for September 10.

"I already have Cursor or Claude Max." Then this is a second seat for a month, or a switch test. $20 is the cost of the experiment. No single harness wins every workflow.

"Kimi K3 already exists and it is cheaper." The base is good. Cognition's claim is that the RL plus the Devin tooling, browser, desktop computer use, PR loop, Fusion and Cloud, is the product. SWE-2 without the harness is not what is for sale.

If you are shipping a developer tool into this market

One more thing, since you are reading this on a directory's blog. A $48 billion company just made a frontier-adjacent coding agent free for a month to buy distribution. If you are building a smaller developer tool, an AI coding product or any SaaS that competes for the same attention, you cannot outspend that. You can out-distribute it in the places it is not looking.

That is what SaaSCity is for. It is a live city map of SaaS products where founders discover tools, with a permanent listing, a dofollow backlink from a high-DR domain and a launch week in front of other builders. Listing is free with a badge. Quick Pass gets you live within 24 hours, and readers of this post get it for $13.99 instead of list price. Submit your product or see what a listing includes. If you are collecting launch venues, start with the best SaaS directories and the best AI directories lists.

Decide, then move

Do not summarize this. Act on it.

  • If you write code for a living and do not have a Devin Pro seat: buy the $20 month, run SWE-2 on three real tickets, keep notes, cancel if it is theater.
  • If you already pay $100 to $200 for Claude Max or Cursor Ultra: park implementation on SWE-2 this month and reserve Fable or Astra for review. That is the trade the Pareto chart is describing.

Sources: cognition.com/blog/swe-2 · devin.ai/pricing · devin.ai/desktop · devin.ai/cli · Devin quota documentation

Get your SaaS in front of founders

List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.

Submit your SaaSSee pricing

Founder resources

Best SaaS directoriesBest AI directoriesDofollow directoriesHigh-DR directoriesFree DR checkerLive launchesAI SaaS boilerplate

Related articles

Cognition Just Raised $2B at a $48B Valuation — AI Coding Isn't Winner-Take-All (2026)

Cognition Just Raised $2B at a $48B Valuation — AI Coding Isn't Winner-Take-All (2026)

Shopify Acquires Tailwind CSS: The Framework That Won the AI Era and Lost the Business (2026)

Shopify Acquires Tailwind CSS: The Framework That Won the AI Era and Lost the Business (2026)

Mistral Just Raised €3B — Europe's Largest Tech Round Ever. What It Buys SaaS Founders (2026)

Mistral Just Raised €3B — Europe's Largest Tech Round Ever. What It Buys SaaS Founders (2026)

Contents

  1. What shipped on September 10, 2026
  2. The scoreboard, vendor-reported
  3. The efficiency story is better than the 50% headline
  4. What the model does differently, per Cognition and per early users
  5. The $20 plan, without the marketing fog
  6. Why this is the best $20 coding deal in September 2026
  7. How to actually use the promo
  8. Company context, kept short
  9. Objections, answered
  10. If you are shipping a developer tool into this market
  11. Decide, then move

List your SaaS

$19.99one-time
  • Dofollow DR 63+ backlink
  • Live within 24 hours, no queue
  • Permanent listing on the city map
Submit your SaaS

Or list free with our badge

City Sponsors

  • Nick LaunchesShip, launch, and get your product in front of real founders.
  • @peregrineintellPeregrine OS: pre-call intel for agency new business
  • Your product hereSlot open — 30 days, homepage + city
Become a sponsor
Write for this blog — from $99.99
SaaSCity.io

Directories are boring. We built a city instead. First isometric SaaS directory on the planet.

Platform
Submit SaaSLive LaunchesPricingBlogWrite for UsBacklink ExchangeMCP for AgentsAdvertise
Directories
Best SaaS DirectoriesHigh-DR DirectoriesFree DirectoriesDofollow DirectoriesAI Tool DirectoriesDeveloper Tool DirectoriesDirectory Submission GuideFree DR CheckerFree DR BadgeBest Directories for SEOFree Dofollow DirectoriesHow to Get SaaS Backlinks
SaaSCity Alternatives
All ComparisonsSaaSCity vs Nick LaunchesSaaSCity vs BetterLaunchSaaSCity vs PeerPushProduct Hunt AlternativesSaaSHub Alternatives
Legal
Privacy PolicyTerms of Service
Company
AboutghostyContact

© 2026 SaaSCity.io

llms.txt