Best AI Model for Coding (September 2026): 13 Models Ranked by Benchmarks and Cost per Task

Best AI model for coding, September 2026: Claude Opus 5.5 leads Terminal-Bench 4.0 at 66.4% for $4/$20. GPT-6 Astra is #1 on the Frontend Code Arena. Kimi K3 is the top open-weight pick. 13 models ranked by SWE-bench Pro, Terminal-Bench, price, and cost per task.

September 22, 2026 · 1 min read
Best AI Model for Coding (September 2026): 13 Models Ranked by Benchmarks and Cost per Task

Best AI Model for Coding: Quick Answer (September 2026)

The best AI model for coding right now is Claude Opus 5.5, released September 22, 2026. It scores 66.4% on Terminal-Bench 4.0 in Anthropic's launch table, ahead of GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%), at $4/$20 per million tokens. GPT-6 Astra is #1 on the Arena.ai Frontend Code Arena. Kimi K3 is the top open-weight model on that board (#7, 1658).

"Best AI model for coding" and "best LLM for coding" are the same question. The answer changed three times in September 2026. Anthropic shipped Claude Fable 5.1 on September 1. OpenAI shipped GPT-6 Astra on September 3 and the cheaper GPT-6 Sol and GPT-6 Luna on September 22. Anthropic shipped Claude Opus 5.5 on September 22 and now tells most workloads to start there. Google's newest model is Gemini 3.8 Flash (September 2); the Gemini API still has no Pro model newer than Gemini 3.1 Pro Preview. xAI shipped Grok 4.7 on September 21. On the open side, DeepSeek V4.1 Flash landed September 10 and Kimi K3's weights are on Hugging Face. This page ranks them by the numbers that still discriminate (Terminal-Bench 4.0, SWE-bench Pro, the Frontend Code Arena), next to price. For the same field ranked picks-by-job, see best LLM for coding.

Best default

Claude Opus 5.5

  • 66.4% Terminal-Bench 4.0 (Anthropic)
  • $4 / $20 per M tokens, 1M context
  • Released Sep 22, 2026

Best for frontend / UI

GPT-6 Astra

  • #1 Arena.ai Frontend Code Arena (1793)
  • 57.9% Terminal-Bench 4.0 (high effort)
  • $10 / $50 per M tokens

Best open-weight

Kimi K3

  • #7 Frontend Code Arena (1658), top open-weight
  • 2.8T parameters, weights on Hugging Face
  • $2.5 / $14 on Morph, 1M context
What changed since July

Claude Opus 5.5 (Sep 22) replaced Opus 5 as Anthropic's default: $4/$20, down from $5/$25, and Anthropic says it costs 40% less to run than Opus 5. Claude Fable 5.1 (Sep 1) raised Fable's Terminal-Bench 4.0 score from 42.0% to 55.8% at the same $10/$50, with cache reads cut to $0.25/M. GPT-6 Astra (Sep 3) became OpenAI's flagship and took #1 on the Frontend Code Arena from Kimi K3, which now ranks 7th. GPT-6 Sol and GPT-6 Luna (Sep 22) cost $2/$10 and $0.10/$0.50, half of GPT-5.6's promo prices. GPT-5.6 Sol dropped to $4/$20 (Aug 21, promo through at least Nov 21). Grok 4.7 shipped Sep 21 at $2/$6. Gemini 3.8 Flash went GA Sep 2. DeepSeek V4.1 Flash (Sep 10) replaced V4 Flash on DeepSeek's API under the new name deepseek-flash. DeepSeek first said V4 Pro would route to V4.1 Flash after Sep 14, then reversed: V4 Pro is still served at its own prices. Vals AI archived SWE-bench Verified as saturated (Sep 1).

13 Models Ranked: SWE-bench Pro x Price x Cost per Solved Point

The 13 models below are what Anthropic, OpenAI, Google, xAI, Moonshot, DeepSeek, and Z.ai actually ship as of September 22, 2026. Each row carries the strongest coding number we could verify from the vendor or an independent board, and the official list price. Scores from different benchmarks are not comparable across rows; the next section puts like against like.

Current Coding Models, September 2026 (official list prices)
ModelBest verified coding number$/M input / outputContext
Claude Opus 5.5 (Anthropic)66.4% Terminal-Bench 4.0 (vendor)$4 / $201M
Claude Fable 5.1 (Anthropic)55.8% Terminal-Bench 4.0; #2 Frontend Code Arena$10 / $501M
Claude Opus 5 (Anthropic, legacy)97.0% SWE-bench Verified (Vals AI, #1)$5 / $251M
Claude Sonnet 5 (Anthropic)63.2% SWE-bench Pro (vendor)$2 / $101M
Claude Haiku 4.5 (Anthropic)39.45% SWE-bench Pro (Scale SEAL)$1 / $5200K
GPT-6 Astra (OpenAI)57.9% Terminal-Bench 4.0; #1 Frontend Code Arena$10 / $501.05M
GPT-6 Sol (OpenAI)68.8% DeepSWE v1.1 (OpenAI, max effort)$2 / $101.05M
GPT-5.6 Sol (OpenAI)64.6% SWE-bench Pro (vendor); 37.3% Terminal-Bench 4.0$4 / $20 (promo)1.05M
Gemini 3.8 Flash (Google)89.4% Terminal-Bench 2.1; 19.1% Terminal-Bench 4.0$0.75 / $3.75 (intro)1M
Grok 4.7 (xAI)38.0% Terminal-Bench 4.0 (Grok Build, xhigh)$2 / $6 (<200K)500K
Kimi K3 (Moonshot, open weights)93.4% SWE-bench Verified (Vals AI)$2.5 / $14 on Morph1M
DeepSeek V4.1 Flash (open weights)90.6% Terminal-Bench 2.1; 31.2% Terminal-Bench 4.0 (DeepSeek)$0.30 / $1.20 peak1M
GLM-5.3 (Z.ai, open weights)95.4% SWE-bench Verified (Vals AI); 88.2% Terminal-Bench 2.1 (Z.ai)$1.19 / $3.74 on Morph1M

Scale's standardized SWE-bench Pro board is the only place where every model runs the same scaffolding. Its newest entry was added July 9, 2026, so no September 2026 model is on it. Its top entries, and what each point costs in output tokens at list price:

SWE-bench Pro (Scale SEAL public set, checked Sep 22, 2026) x Output Price
ModelSWE-bench Pro$/M input / outputOutput $ per Pro point
Muse Spark 1.1 (Meta)61.50%not publishedn/a
gpt-5.4 (xHigh)59.10%older modeln/a
Muse Spark (Meta)55.00%not publishedn/a
Claude Opus 4.6 (thinking)51.90%legacyn/a
Gemini 3.1 Pro (thinking)46.10%$2 / $12 (≤200K)$0.26
Gemini 3 Pro (preview)43.30%previewn/a
gpt-5.2-codex41.04%supersededn/a
Claude Haiku 4.539.45%$1 / $5$0.13
Qwen3 Coder 480B (open weights)38.70%self-hostn/a
Gemini 3 Flash34.63%supersededn/a
66.4%
Top Terminal-Bench 4.0 (Opus 5.5, vendor)
61.50%
Top standardized SWE-bench Pro (Muse Spark 1.1)
$0.13
Cheapest output-$ per Pro point (Haiku 4.5)

AI Coding Benchmarks: Terminal-Bench 4.0 and SWE-bench Pro Scores (September 2026)

SWE-bench Verified stopped separating frontier models: Vals AI reports seven of 86 models at 95% or better and has stopped running it. Terminal-Bench 4.0 is the benchmark every September launch led with. The numbers below come from Anthropic's Opus 5.5 table, xAI's Grok 4.7 release, Google's Gemini 3.8 Flash evaluation, and DeepSeek's V4.1 Flash model card. Each vendor picked its own effort level and harness: xhigh for Opus 5.5, high for GPT-6 Astra, xhigh with the Grok Build harness for Grok 4.7, max for DeepSeek V4.1 Flash.

Terminal-Bench 4.0 (vendor-reported, September 2026)

Sources: Anthropic Opus 5.5 launch table; xAI via llm-stats (Grok 4.7); Google DeepMind (Gemini 3.8 Flash); DeepSeek V4.1 Flash model card. Higher is better.

1
Claude Opus 5.5
Sep 22
66.4%
2
GPT-6 Astra
57.9%
3
Claude Fable 5.1
55.8%
4
Claude Opus 5
52.3%
5
Grok 4.7
38%
6
GPT-5.6 Sol
37.3%
7
DeepSeek V4.1 Flash
open weights
31.2%
8
Gemini 3.8 Flash
19.1%

Anthropic's standard error on its own runs is about ±2.6 points for Opus 5.5, so a 2-point gap is noise.

Check the Terminal-Bench version before comparing

Google reports Gemini 3.8 Flash at 89.4% on Terminal-Bench 2.1 and 19.1% on Terminal-Bench 4.0. Same model, a 70-point gap. DeepSeek V4.1 Flash shows the same pattern: 90.6% on 2.1, 30.0% on 3.0, 31.2% on 4.0. Launch posts and ranking pages mix 2.x, 3.0, and 4.0 numbers freely. Z.ai reports GLM-5.3 at 88.2% on 2.1 and 28.3% on 3.0, with no 4.0 score. Only compare scores from the same version, and treat any "Terminal-Bench" number without a version as unusable.

On SWE-bench Pro, the vendor-reported board (llm-stats, all 59 entries self-reported, none verified) reads: Claude Fable 5 80.0%, Claude Opus 4.8 69.2%, Qwen3.8 Max 67.7%, Grok 4.5 64.7%, GPT-5.6 Sol 64.6%, Claude Sonnet 5 63.2%, GLM-5.2 62.1%. llm-stats also shows Claude Opus 5.5 at 89.9%. Anthropic's launch page does not report SWE-bench Pro, so treat that one as unconfirmed. Deeper numbers on the SWE-bench Pro leaderboard page.

For one independent number across all of these models, the Artificial Analysis Intelligence Index runs the same eval suite on every model. It is a general index, not coding-only, but it agrees with the ordering above.

Artificial Analysis Intelligence Index (checked Sep 22, 2026)
ModelIndex
Claude Opus 5.5 (max, with fallback)58
Claude Fable 5.1 (max)53
GPT-6 Astra (max)53
Claude Opus 5 (max)51
GPT-6 Sol (max)48
GLM-5.3 (max, open weights)45
Kimi K3 (max, open weights)44
Gemini 3.8 Flash (high)41
DeepSeek V4.1 Flash (max, open weights)39

Frontend Code Arena: GPT-6 Astra #1, Fable 5.1 Second, Kimi K3 in the Top 10

SWE-bench and Terminal-Bench measure agents fixing repos. They say little about building UI. The Arena.ai Frontend Code Arena ranks models by blind pairwise human votes on real frontend prompts. On September 22, 2026, GPT-6 Astra leads it.

Arena.ai Frontend Code Arena, top 10 (Sep 22, 2026)
RankModelScore
1gpt-6-astra-max (OpenAI)1793
2claude-fable-5.1-max (Anthropic)1755
3claude-opus-5-max (Anthropic)1691
4qwen3.8-max (Alibaba)1671
5qwen3.8-max-0902 (Alibaba)1662
6claude-opus-5-high (Anthropic)1661
7kimi-k3-max (Moonshot, open weights)1658
8muse-spark-1.3-max (Meta)1657
9qwen3.8-flash-next (Alibaba)1636
10grok-4.7-xhigh (xAI)1632

The July version of this page had Kimi K3 at #1. Three newer closed models have passed it since. The gap from #1 to #7 is 135 points. Opus 5.5 does not appear in the top 12 on its launch day. The effort suffix matters: claude-opus-5-max sits 30 points above claude-opus-5-high, so the setting you run in production changes the result as much as the model does. Kimi K3 is also the top open-weight model on Arena's Agent Arena, 8th overall at 6.22%, where Claude Fable 5.1 (Max) leads at 13.71%. If you want Kimi K3 without its own rack, the Morph Kimi K3 API serves it at $2.5 / $14 per million tokens on an OpenAI-compatible endpoint.

1793
GPT-6 Astra, Frontend Code Arena #1
1755
Claude Fable 5.1, #2
1658
Kimi K3, top open-weight entry

Cost per Completed Task, Not per Token

The per-token price on a model card is a weak budget signal for a coding agent. You pay for every token the model burns to finish: reasoning, retries, tool-call round trips. A model with half the per-token price can cost more per task if it thinks three times as long.

The September launches made this explicit. Anthropic says Opus 5.5 matches GPT-6 Astra on Terminal-Bench 4.0 for about 40% of the cost, and costs 40% less to run than Opus 5, even though Opus 5.5's list price ($4/$20) is only 20% below Opus 5 ($5/$25). The rest of the saving is fewer tokens per task. Fable 5.1 kept Fable 5's $10/$50 list price but cut cache reads 75% to $0.25/M, and cache reads dominate long agent sessions.

Effort level is the other lever. Anthropic's headline numbers run at max effort. In the Hacker News launch thread, users pointed out that high and xhigh effort land close to max on the benchmarks at about a tenth of the token usage (HN, Sep 22, 2026). Opus 5.5 defaults to medium effort on the API. Set it explicitly and measure tokens per solved task on your own traffic.

~40%
Opus 5.5 cost vs GPT-6 Astra at equal Terminal-Bench 4.0 (Anthropic)
$0.25/M
Fable 5.1 cache reads, down 75%
medium
Opus 5.5 default API effort

SWE-bench Verified Leaderboard (Archived September 2026)

Vals AI ran a single standardized, bash-only SWE-bench Verified harness across 88 models. On its last update (September 1, 2026) it archived the board: "Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases." Seven models reached 95% or better, and two of the top six are open weights. None of Opus 5.5, Fable 5.1, GPT-6 Astra, Grok 4.7, or DeepSeek V4.1 Flash will get a Vals Verified score.

SWE-bench Verified: Independent Harness (final, Sep 1, 2026)

Source: Vals AI (vals.ai/benchmarks/swebench), single standardized harness. Board archived as saturated.

1
Claude Opus 5
#1
97%
2
DeepSeek V4 Pro 0813
open weights
96.4%
3
GPT-5.6 Sol
96.2%
4
Grok 4.6
95.6%
5
GPT-5.6 Terra
95.4%
6
GLM-5.3
open weights
95.4%
7
Claude Fable 5
95%
8
Kimi K3
open weights
93.4%
9
GPT-5.6 Luna
93%
10
GLM-5.3-Flash
open weights
92%
11
DeepSeek V4 Flash 0731
open weights
88.8%
12
Claude Opus 4.8
88.6%

At 95%+ the remaining failures are mostly ambiguous or broken tasks. Use Terminal-Bench 4.0 or SWE-bench Pro to separate current models.

Which Claude Model Is Best for Coding?

Claude Opus 5.5 (claude-opus-5-5) is the best Claude model for coding for most teams. Anthropic's model overview says to start there and move to Claude Fable 5.1 only when evals at higher Opus effort fall short. Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks. The current lineup, straight from Anthropic's docs on September 22, 2026:

Claude Models for Coding (September 2026, official pricing)
Model (API ID)Coding benchmarks$/M in / outContext / max output
Claude Opus 5.5 (claude-opus-5-5)66.4% Terminal-Bench 4.0, 57.8% CursorBench 4.0$4 / $201M / 128K
Claude Fable 5.1 (claude-fable-5-1)55.8% Terminal-Bench 4.0, 51.8% CursorBench 4.0$10 / $501M / 128K
Claude Sonnet 5 (claude-sonnet-5)63.2% SWE-bench Pro (vendor)$2 / $101M / 128K
Claude Haiku 4.5 (claude-haiku-4-5)39.45% SWE-bench Pro (Scale SEAL)$1 / $5200K / 64K
Claude Opus 5 (legacy)97.0% SWE-bench Verified (Vals AI), 52.3% Terminal-Bench 4.0$5 / $251M / 128K

Default: Opus 5.5

claude-opus-5-5 at $4/$20 leads Terminal-Bench 4.0 at 66.4%, 10.6 points above Fable 5.1. Cache reads cost 5% of input ($0.20/M). Fast mode is $8/$40. Default effort is medium; set high or xhigh for hard agent runs.

Ceiling: Fable 5.1

claude-fable-5-1 ($10/$50) is Anthropic's pick for demanding reasoning and long-horizon agentic work. It is #2 on the Frontend Code Arena at 1755. Cache reads are 2.5% of input ($0.25/M). Use it when Opus 5.5 at xhigh still fails your evals.

Volume: Sonnet 5

claude-sonnet-5 at $2/$10 has a 1M context window and 128K max output. Use it for CI review bots, test generation, and batch transforms where Opus pricing compounds. The Batch API halves it again.

Quick edits and subagents: Haiku 4.5

claude-haiku-4-5 at $1/$5 is the cost-per-point leader on Scale's standardized board (~$0.13 of output per SWE-bench Pro point). Anthropic's retirement commitment for it is only 'not sooner than October 15, 2026', so plan a successor.

Gotcha: Opus 5.5 can silently fall back to Opus 5

Anthropic's launch notes, quoted in the Hacker News thread, say Opus 5.5 carries classifiers for a small set of capabilities tied to frontier LLM development, "such as kernel development for certain ML accelerators," and that these "cause Claude to fall back from Opus 5.5 to Opus 5" (HN, Sep 22, 2026). Anthropic's launch page states the routing: when the safeguards intervened, "cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5." If you write GPU kernels, security tooling, or training infrastructure, check which model actually answered. Fable 5, Opus 5, Opus 4.8, Opus 4.7, and Sonnet 4.6 are now legacy on Anthropic's overview. Full price tables on the Anthropic API pricing page and scores on Claude benchmarks.

Best Codex Model and Best GPT Model for Coding

OpenAI's Codex docs list three recommended models: GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. Sol and Luna joined Codex on September 22; Astra has been the CLI's bundled default since 0.153.4. The GPT-5.6 models stay selectable during the rollout. GPT-5.4 and GPT-5.4 Mini left Codex for ChatGPT sign-in on August 31 and now work only with an API key. GPT-5.5 retires from ChatGPT and Codex on all plans on October 14, 2026. The latest Codex CLI release is 0.156.0 (September 22, 2026).

OpenAI Models for Coding (API pricing, short context, Sep 22, 2026)
ModelRole in Codex$/M in / cached / outCoding number
GPT-6 AstraRecommended, flagship$10 / $1 / $5057.9% Terminal-Bench 4.0; #1 Frontend Code Arena
GPT-6 SolRecommended (Sep 22)$2 / $0.20 / $1068.8% DeepSWE v1.1 (OpenAI, max)
GPT-6 LunaRecommended, fast (Sep 22)$0.10 / $0.01 / $0.5066.6% DeepSWE v1.1 (OpenAI, max)
GPT-5.6 SolSelectable during rollout; promo price through at least Nov 21$4 / $0.40 / $2064.6% SWE-bench Pro (vendor)
GPT-5.5Retires Oct 14, 2026$5 / $0.50 / $30legacy
gpt-5.3-codexLast -codex model on the price list$1.75 / $0.175 / $14legacy

Which Codex model is best for coding: GPT-6 Astra for the hardest runs and UI work, GPT-6 Sol for daily work at a fifth of Astra's price. Long prompts cost more on every GPT-6 tier: above 272K input tokens the whole request bills at 2x input and 1.5x output, so Astra goes to $20/$75. If your Codex setup still pins gpt-5.5, change it before October 14. The subscription side is on the Codex pricing page.

Claude Opus 5.5 vs GPT-6 Astra: The Everyday Frontier Pair (Successors to Opus 4.8 vs GPT-5.6 Sol)

In July the pair was Claude Opus 4.8 and GPT-5.6 Sol. Both are now a generation behind. On Anthropic's published table Opus 5.5 leads Astra on the coding benchmarks and costs 60% less per token. Astra leads the Frontend Code Arena and AutomationBench. Note who ran the numbers: every row except the Arena is from Anthropic's launch post.

Claude Opus 5.5 vs GPT-6 Astra (September 2026)
DimensionClaude Opus 5.5GPT-6 Astra
Terminal-Bench 4.066.4% (xhigh)57.9% (high)
FrontierCode v1.154.4%53.3%
AutomationBench40.0%41.4%
Frontend Code Arenanot rated yet1793 (#1)
Pricing ($/M in / out)$4 / $20$10 / $50
Cached input ($/M)$0.20$1.00
Context window1M / 128K output1.05M / 128K output

For repo-scale agent work, Opus 5.5 is the default on both score and price. For frontend building, Astra has the stronger human-preference signal. Opus 4.8 (88.6% Vals Verified, 69.2% vendor SWE-bench Pro) and GPT-5.6 Sol (now $4/$20) still work, but neither is the cheapest option at its capability level anymore.

Open-Source Models: 80% SWE-bench Verified at a Tenth of the Price

The open tier cleared 80% SWE-bench Verified months ago. On Vals AI's final board, DeepSeek V4 Pro 0813 (96.4%) is #2 overall, GLM-5.3 (95.4%) ties for #5, and Kimi K3 sits at 93.4%. Kimi K3's 2.8T-parameter weights are on Hugging Face (paper published July 27). DeepSeek V4.1 Flash shipped September 10 with weights and a technical paper on Hugging Face. GLM-5.3 is Z.ai's current flagship, built on the GLM-5.2 base with new post-training. Official and Morph prices:

Open-Weight Coding Models (September 2026)
ModelBenchmark$/M in / outContext / notes
Kimi K3 (Moonshot, 2.8T)93.4% Verified (Vals AI); #7 Frontend Code Arena$2.5 / $14 on Morph (morph-kimik3)1M on Morph; weights on HF
DeepSeek V4.1 Flash (552B backbone, ~763B total, 8B/16B active)90.6% Terminal-Bench 2.1, 74.2% DeepSWE v1.1 (DeepSeek); beats V4-Pro on most agentic benchmarks per DeepSeek$0.30 / $1.20 peak, $0.15 / $0.60 off-peak (DeepSeek)1M / 384K output; weights on HF
morph-dsv41flash (DeepSeek V4.1 Flash on Morph)Same weights, custom kernels$0.15 / $0.61M, hosted
GLM-5.3 (Z.ai)95.4% Verified (Vals AI); 88.2% Terminal-Bench 2.1 (Z.ai)$1.19 / $3.74 on Morph (morph-glm53-744b)1M / 128K output
GLM-5.3-Flash (Z.ai)92.0% Verified (Vals AI)$0.2 / $0.7 on Morph (morph-glm53flash)1M, hosted
DeepSeek V4 Pro 0813 (1.6T backbone, 49B active)96.4% Verified (Vals AI, #2)$1.32 / $3.96 peak, $0.66 / $1.98 off-peak (DeepSeek)1M / 384K output; still served on DeepSeek's API
DeepSeek V4 Flash 073188.8% Verified (Vals AI); DeepSeek API now routes its name to V4.1$0.141953125 / $0.399625 on Morph (morph-dsv4flash)1M, hosted

The arithmetic that matters: DeepSeek V4.1 Flash at $1.20/M peak output is 1/17 of Opus 5.5's $20/M and 1/42 of Fable 5.1's $50/M. One gotcha from DeepSeek's release note: the canonical name is now deepseek-flash, and requests to deepseek-v4-flash are served by V4.1 Flash. If you pinned V4 Flash for stability, your model changed underneath you. Pin open weights on a host you control, or on a provider that keeps model IDs fixed. More detail on the best open-source coding model page.

93.4%
Kimi K3 SWE-bench Verified (Vals AI)
$1.20/M
DeepSeek V4.1 Flash peak output
17x
Opus 5.5 output premium over V4.1 Flash
Where you run open weights changes the output

Open weights are identical everywhere; the serving stack is not. Morph serves open coding models on custom inference kernels built for code generation, on one OpenAI-compatible API: Kimi K3 (morph-kimik3) at $2.5 / $14, GLM-5.3 (morph-glm53-744b) at $1.19 / $3.74, GLM-5.3-Flash (morph-glm53flash) at $0.2 / $0.7, DeepSeek V4.1 Flash (morph-dsv41flash) at $0.15 / $0.6, and DeepSeek V4 Flash (morph-dsv4flash) at $0.141953125 / $0.399625, per million input / output tokens. Full catalog on Morph models and pricing.

What People Actually Code With (Hacker News, September 2026)

The Opus 5.5 launch thread on Hacker News hit 819 points and 631 comments on day one. The comments say more about model choice than the benchmark tables do.

Three signals from the thread. First, release fatigue: one developer wrote that they no longer have time to learn one model before the next ships, so they pick by mood and set effort by remaining quota. Second, cheap open models are good enough for a lot of real work: one commenter runs DeepSeek V4.1 at high effort as a "relentlessly hardworking, dirt cheap" daily driver for frontend layout changes. Third, safeguards now drive switching. Fable-style classifiers pushed some users to stay on Opus, and the Opus 5.5 fallback to Opus 5 on kernel work already has people reconsidering. The practical takeaway: keep two providers wired up, and route by task rather than loyalty.

Best AI Model for Coding at $0

The best free path to real coding capability in September 2026 is open weights you run yourself. Hosted options that cost close to nothing:

Free and Near-Free Coding-Model Options (September 2026)
OptionWhat you getLimit
Kimi K3 open weights2.8T-parameter weights on Hugging FaceNeeds a multi-GPU cluster to serve
DeepSeek V4.1 Flash open weights552B-backbone MoE weights (~763B total) on Hugging FaceYour GPU cost
Gemini API free tierFree access to Gemini models with limitsRate-limited; paid tier for production
DeepSeek API off-peakV4.1 Flash at $0.15 / $0.60 per M tokensOff-peak only; peak is 01:00-04:00 and 06:00-10:00 UTC weekdays
GPT-6 Luna (API)$0.10 / $0.50 per M tokensSmallest GPT-6 tier

Free tiers are for evaluation and light use. At $0.15/$0.60 off-peak, hosted DeepSeek V4.1 Flash is cheap enough that self-hosting mostly pays when data cannot leave your network.

Why Vendor Scores Run 20 Points Above Scale's Leaderboard

Anthropic reports Claude Fable 5 at 80.0% on SWE-bench Pro. Scale's standardized leaderboard tops out at 61.50% (Muse Spark 1.1). Both numbers are real. Scale runs every model through identical scaffolding; vendors run tuned agent stacks at max effort. Scale has not added any model released after July, so every current frontier SWE-bench Pro number is vendor-reported.

Same Benchmark, Different Harness: SWE-bench Pro

Vendor-reported (llm-stats) vs Scale SEAL standardized scaffolding, checked Sep 22, 2026.

1
Fable 5 (vendor)
80%
2
Opus 4.8 (vendor)
69.2%
3
Grok 4.5 (vendor)
64.7%
4
GPT-5.6 Sol (vendor)
64.6%
5
Sonnet 5 (vendor)
63.2%
6
Muse Spark 1.1 (Scale, #1)
61.5%
7
gpt-5.4 (Scale)
59.1%
8
Opus 4.6 (Scale, top Claude)
51.9%

The vendor-vs-standardized gap is roughly 20 points. The harness is the variable.

The scaffold around the model accounts for as much variance as swapping frontier models. Before paying a 2.5x token premium, fix retrieval, context management, and tool design. Subagent architecture and context engineering move scores more than model choice does.

The implication

A mid-tier model in a strong harness beats a frontier model in a weak one. WarpGrep (codebase search for coding agents, $0.8 per 100K tokens) upgrades the harness for every model you route through it.

Per-Task Routing: Which Model for Which Job

The most cost-effective setups in September 2026 route by task. Send the hard 20% to Opus 5.5 and the easy 80% to Sonnet 5, Haiku 4.5, or DeepSeek V4.1 Flash. Defaults with the numbers behind them:

Task Routing Matrix (September 2026)
TaskRoute toWhy (verified numbers)
Long agent runs, 50+ file refactorsClaude Opus 5.566.4% Terminal-Bench 4.0, $4/$20, 1M context
Hardest debugging / migrationClaude Fable 5.1 or GPT-6 AstraBoth $10/$50; use when Opus 5.5 at xhigh fails
Frontend and UI buildingGPT-6 Astra#1 Frontend Code Arena at 1793
Codex CLI daily workGPT-6 SolRecommended in Codex, $2/$10
Quick edits, lint fixes, subagentsClaude Haiku 4.5 or GPT-6 Luna$1/$5 and $0.10/$0.50
High-volume batch / CI botsDeepSeek V4.1 Flash$0.15/$0.60 off-peak, 1M context
Cheap long promptsGrok 4.7 or Gemini 3.8 Flash$2/$6 (<200K) and $0.75/$3.75 intro pricing
Data sovereignty / self-hostKimi K3 or DeepSeek V4.1 FlashWeights on Hugging Face; Kimi K3 93.4% Verified
Codebase search for any agentWarpGrep + any modelModel-agnostic retrieval; $0.8 per 100K tokens

Cost levers that apply across routes: Anthropic's Batch API is 50% off, Opus 5.5 cache reads are 5% of input, and DeepSeek's off-peak window halves every rate. Doing the split automatically needs a classifier. Morph's Router scores each prompt and picks a model for $0.005 per request. Claude Code Router makes per-request routing concrete inside the terminal agent, and Claude Code models covers harness-side defaults.

Frequently Asked Questions

What is the best AI model for coding in 2026?

As of September 22, 2026, Claude Opus 5.5 (claude-opus-5-5, $4/$20 per million tokens, 1M context) is the best default. It scores 66.4% on Terminal-Bench 4.0 in Anthropic's launch table, ahead of GPT-6 Astra (57.9%) and Claude Fable 5.1 (55.8%). GPT-6 Astra ($10/$50) is #1 on the Arena.ai Frontend Code Arena at 1793. Kimi K3 is the best open-weight pick: the top open-weight entry on the Frontend Code Arena (#7, 1658), 93.4% SWE-bench Verified on Vals AI, and $2.5 / $14 on Morph. For cheap volume, DeepSeek V4.1 Flash costs $0.30/$1.20 at peak on DeepSeek's API.

Which AI model is best for coding right now?

Right now (September 22, 2026) it is Claude Opus 5.5, released that day. Anthropic says it performs at the level of Claude Fable 5.1 on most work, and it costs $4/$20 against Fable 5.1's $10/$50. The runner-up is GPT-6 Astra, OpenAI's flagship since September 3, which leads the Frontend Code Arena. OpenAI added the cheaper GPT-6 Sol and GPT-6 Luna the same day as Opus 5.5. All Terminal-Bench 4.0 numbers are vendor-reported and run at different effort levels, so test both on your own repo before committing.

What is the best LLM for coding?

"Best LLM for coding" and "best AI model for coding" are the same question. Claude Opus 5.5 leads Terminal-Bench 4.0 (66.4%, Anthropic-reported). GPT-6 Astra leads the Frontend Code Arena (1793). Claude Opus 5 holds the top archived SWE-bench Verified score on Vals AI (97.0%). The best open-weight LLM for coding is Kimi K3, the top open-weight entry on Arena's Code and Agent boards. On Vals AI's archived Verified board, DeepSeek V4 Pro 0813 (96.4%) and GLM-5.3 (95.4%) score above Kimi K3 (93.4%). For most teams the best answer is a router that sends easy work to a cheap model and reserves a frontier model for hard edits.

Is there an open source LLM as good as Claude for coding?

Close on some benchmarks. On Vals AI's SWE-bench Verified, DeepSeek V4 Pro 0813 scores 96.4%, GLM-5.3 95.4%, and Kimi K3 93.4%, against 97.0% for Claude Opus 5 and 88.6% for Claude Opus 4.8. Kimi K3 (2.8T parameters, weights on Hugging Face) ranks 7th on the Frontend Code Arena at 1658, behind GPT-6 Astra, Fable 5.1, and Opus 5. DeepSeek V4.1 Flash (552B backbone, about 763B total with its Engram memory and vision encoder) covers the cheap end. On the benchmark where the frontier separates, the gap is wide: DeepSeek reports 31.2% on Terminal-Bench 4.0 for V4.1 Flash, against Opus 5.5's 66.4%.

Which Claude model is best for coding?

Claude Opus 5.5 (claude-opus-5-5, $4/$20) is the model Anthropic tells most workloads to start with. It scores 66.4% on Terminal-Bench 4.0. Claude Fable 5.1 (claude-fable-5-1, $10/$50, 55.8% Terminal-Bench 4.0) is for the hardest long-horizon agentic work. Claude Sonnet 5 (claude-sonnet-5, $2/$10, 1M context) is the volume pick. Claude Haiku 4.5 ($1/$5, 200K context) handles quick edits and subagents. Fable 5, Opus 5, Opus 4.8, and Sonnet 4.6 are now listed as legacy. Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks.

What is the best Codex model for coding in 2026?

OpenAI's Codex docs list GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna as the recommended Codex models. Astra ($10/$50) is the flagship and #1 on the Frontend Code Arena. GPT-6 Sol ($2/$10, released September 22) is the everyday pick at a fifth of Astra's price; OpenAI reports 68.8% on DeepSWE v1.1 at max effort. GPT-5.5 retires from Codex on all plans on October 14, 2026, so migrate off it now. The only dedicated -codex model still on the API price list is gpt-5.3-codex ($1.75/$14).

Which GPT model is best for coding?

GPT-6 Astra is OpenAI's strongest coding model: 57.9% on Terminal-Bench 4.0 (high effort, from Anthropic's comparison table) and #1 on the Arena.ai Frontend Code Arena at 1793. It costs $10/$50 per million tokens up to 272K input tokens and $20/$75 above that. GPT-6 Sol ($2/$10, released September 22) is the value pick. GPT-5.6 Sol dropped to $4/$20 on a promo that runs at least through November 21, 2026, and scores 37.3% on Terminal-Bench 4.0 in the same table.

What are the SWE-bench Pro scores for coding models in 2026?

Scale SEAL public set (standardized scaffolding, checked September 22, 2026): Muse Spark 1.1 61.50%, gpt-5.4 xHigh 59.10%, Muse Spark 55.00%, Claude Opus 4.6 thinking 51.90%, Gemini 3.1 Pro thinking 46.10%, Claude Opus 4.5 45.89%, Claude Sonnet 4.5 43.60%, Gemini 3 Pro 43.30%. Scale's newest entry was added July 9, 2026, so none of the September 2026 frontier models is on its board. Vendor-reported SWE-bench Pro on llm-stats: Claude Fable 5 80.0%, Claude Opus 4.8 69.2%, Qwen3.8 Max 67.7%, Grok 4.5 64.7%, GPT-5.6 Sol 64.6%, Claude Sonnet 5 63.2%, GLM-5.2 62.1%. llm-stats also lists Opus 5.5 at 89.9%, but Anthropic's launch page does not report SWE-bench Pro, so treat that number as unconfirmed.

What are the best AI coding benchmarks in 2026?

Three matter in September 2026. Terminal-Bench 4.0 is where frontier models separate (Opus 5.5 66.4%, GPT-6 Astra 57.9%, Fable 5.1 55.8%, all vendor-reported). SWE-bench Pro, ideally Scale's standardized board, measures repo-scale bug fixing. The Arena.ai Frontend Code Arena ranks UI work by blind human votes. SWE-bench Verified is saturated: Vals AI archived it in September 2026 and no longer runs new models on it. Watch the Terminal-Bench version: Gemini 3.8 Flash scores 89.4% on Terminal-Bench 2.1 but 19.1% on 4.0.

What is the best free AI model for coding?

The strongest free path is open weights you run yourself: Kimi K3 and DeepSeek V4.1 Flash both publish weights on Hugging Face. Google's Gemini API has a free tier with limits. If you need hosted and nearly free, DeepSeek V4.1 Flash costs $0.15 input and $0.60 output per million tokens off-peak ($0.30/$1.20 at peak), and Gemini 3.8 Flash has an introductory $0.75/$3.75 price through December 31, 2026.

What is the best open-source AI model for coding?

Kimi K3 (Moonshot, 2.8T parameters) is the top open-weight model on Arena: 7th on the Frontend Code Arena and first among open weights on the Agent Arena, with 93.4% SWE-bench Verified on Vals AI. DeepSeek V4.1 Flash (released September 10, 2026; 552B backbone, about 763B total, 8B active for input and 16B for output) is the cheap default. GLM-5.3 (1M context, 128K output, 95.4% on Vals AI's Verified board) is Z.ai's current flagship, post-trained on the GLM-5.2 base. On Morph: Kimi K3 $2.5 / $14, GLM-5.3 $1.19 / $3.74, DeepSeek V4.1 Flash $0.15 / $0.6 per million tokens.

How much do the top coding models cost per million tokens?

Input / output per million tokens, September 22, 2026: Claude Fable 5.1 $10/$50, GPT-6 Astra $10/$50, Claude Opus 5 $5/$25, Claude Opus 5.5 $4/$20, GPT-5.6 Sol $4/$20, Claude Sonnet 5 $2/$10, GPT-6 Sol $2/$10, Grok 4.7 $2/$6, Gemini 3.1 Pro $2/$12, Claude Haiku 4.5 $1/$5, Gemini 3.8 Flash $0.75/$3.75 (introductory), DeepSeek V4.1 Flash $0.30/$1.20 (peak), GPT-6 Luna $0.10/$0.50.

Why do vendor benchmark scores differ from Scale's SWE-bench Pro leaderboard?

Scale runs every model through the same scaffolding on SWE-bench Pro's 1,865 tasks across 41 repositories, scored Pass@1. Vendors run their own tuned harnesses at max effort. Anthropic reports 69.2% for Opus 4.8; Scale's best Claude entry, Opus 4.6, scores 51.90%. Scale's top score is 61.50% while vendor-reported scores reach 80.0%. The roughly 20-point spread is mostly the harness, which is why agent tooling moves results as much as model choice.

Which AI model is most cost-effective for coding in 2026?

For cheap volume, DeepSeek V4.1 Flash at $0.30/$1.20 peak ($0.15/$0.60 off-peak) on DeepSeek's API. For frontier capability per dollar, Claude Opus 5.5: Anthropic says it beats GPT-6 Astra's max-effort Terminal-Bench 4.0 score at about 40% of the cost, and it lists at $4/$20 against Astra's $10/$50. On Scale's standardized SWE-bench Pro, Claude Haiku 4.5 is the cheapest per solved point at about $0.13 of output per point.

Sources

Primary sources behind the scores and prices on this page (all checked September 22, 2026):

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.

Talk to us about a private deployment

Stop Debating Models. Start Searching Codebases.

WarpGrep adds codebase search to any coding agent. Works with Claude Opus 5.5, GPT-6 Astra, Kimi K3, Gemini, DeepSeek V4.1, GLM-5.3, or any model. $0.8 per 100K tokens. The harness matters as much as the model.