By July 2026, open coding models split into three leaders: Kimi K3 for frontend, GLM-5.2 for value, DeepSeek V4 for raw SWE-bench. The gap to proprietary models closed to a few points, and on frontend UI an open model now leads outright. This is the complete comparison with real benchmarks, licenses, and API pricing.
The July 2026 Open-Model Tier
This page ranks open coding models specifically. For the full field including proprietary frontier models, see best LLM for coding. Three models lead, each best at a different job:
- Kimi K3 (Moonshot AI) is #1 on the Arena.ai Frontend Code Arena, the first open model to lead frontend coding, ahead of Claude Fable 5, and scores 93.4% SWE-bench Verified on Vals AI's independent harness. Open weights are due on Hugging Face around July 27.
- GLM-5.2 (Zhipu AI) is the top open model on the Artificial Analysis Intelligence Index. 744B MoE / 40B active, MIT, 1M context, served at $0.8415/$3.1365 on Morph.
- DeepSeek V4 leads raw SWE-bench Verified among downloadable weights at 80.6% (Pro-Max, vendor), MIT-licensed and self-hostable, with a near-free Flash variant at $0.14/$0.28.
Behind the leaders, MiniMax M3 (80.5% Verified), Qwen3-Coder-Next (70.6% at 3B active parameters), and OpenAI's gpt-oss round out the field. On general SWE-bench Verified the top open model (DeepSeek V4 Pro, 80.6%) still trails the frontier (GPT-5.6 Sol 96.2%, Fable 5 95.0%, Kimi K3 93.4% on Vals AI), but at roughly a tenth of the output price, and on frontend UI the open leader is already on top.
Most SWE-bench Verified numbers are vendor self-reported; the llm-stats tracker lists 0 of 104 entries as independently verified. Where possible this page cites independent references instead: Vals AI (single-harness Verified), the Arena.ai Frontend Code Arena (blind human votes), and Scale SEAL (standardized SWE-bench Pro). Vendor numbers, especially DeepSeek V4's 80.6%, should be read as an upper bound.
Master Comparison: Open Coding Models (July 2026)
| Model | Total / Active Params | Coding benchmark | Context | License |
|---|---|---|---|---|
| Kimi K3 | 2.8T MoE | #1 Frontend Code Arena; 93.4% Verified (Vals AI) | 1M | weights due ~Jul 27 |
| DeepSeek-V4-Pro-Max | 1.6T / 49B | 80.6% Verified (vendor) | 1M | MIT |
| MiniMax M3 | ~230B / ~10B | 80.5% Verified (vendor) | ~1M | Community |
| DeepSeek V4 Flash | 284B / 13B | 79.0% Verified (Flash-Max, vendor) | 1M | MIT |
| GLM-5.2 | 744B / 40B | #1 open, AA Intelligence Index; 62.1% Pro | 1M | MIT |
| Qwen3-Coder-Next | 80B / 3B | 70.6% Verified (vendor) | 256K | Apache 2.0 |
| gpt-oss-120b | 117B / 5.1B | ~o4-mini class (no official Verified) | 128K | Apache 2.0 |
SWE-bench Verified scores are vendor self-reported unless marked (Vals AI). Kimi K3 leads frontend on the Arena.ai Code Arena; GLM-5.2 leads the open-weights Artificial Analysis Intelligence Index. MiniMax M3 ships under a community license, not MIT.
| Model | Vals AI SWE-bench Verified | Arena Frontend Code | Scale SEAL SWE-bench Pro |
|---|---|---|---|
| Kimi K3 | 93.4% | #1 (~1,679) | not yet run |
| GLM-5.2 | not yet run | listed (max) | not yet run |
| DeepSeek V4 Pro | not yet run | not listed | not yet run |
| Qwen3 Coder 480B | not yet run | not listed | 38.7% |
| Kimi K2 Instruct | not yet run | not listed | 27.67% |
Independent leaderboards move slowly; the newest models are often absent for weeks after launch. Vals AI runs one standardized SWE-bench Verified harness; Scale SEAL runs one standardized SWE-bench Pro harness (1,865 tasks, 41 repos). When a model is not yet on a board, its vendor number is the only figure available.
Kimi K3: #1 on the Frontend Code Arena, Ahead of Fable 5
Kimi K3 (Moonshot AI) launched July 16, 2026 and took #1 on the Arena.ai Frontend Code Arena at roughly 1,679 points, the first open model to lead a frontend coding leaderboard, ahead of Claude Fable 5 and GPT-5.6 Sol. On Vals AI's independent SWE-bench Verified harness it scores 93.4%, third behind GPT-5.6 Sol (96.2%) and Fable 5 (95.0%).
The catch is serving. On Moonshot's own API, Artificial Analysis measures Kimi K3 at about 34 tokens per second with a roughly 7-second time to first token, and the open weights are not out yet, so no third party can self-host it faster. Moonshot lists it at $3.00/M input ($0.30 cached) and $15.00/M output. The Morph Kimi K3 API serves the same model at about 100 tokens per second on an OpenAI-compatible endpoint.
Moonshot committed to a Hugging Face weights release around July 27, 2026. Until then Kimi K3 is open in spirit and hosted in practice: you reach it through APIs, not a download. The K2 line shipped under a Modified MIT license, and K3 is expected to follow, but the license text is not confirmed as of late July. If you need weights you can pin and self-host today, GLM-5.2 or DeepSeek V4 are the picks below.
GLM-5.2 (Zhipu AI): Top Open-Weights Intelligence, Best Value
GLM-5.2 is the top open model on the Artificial Analysis Intelligence Index. It is a 744B total / 40B active MoE, MIT licensed, with a 1M-token context, released June 16, 2026. On coding, Z.ai reports 62.1% on SWE-bench Pro (vendor) and 81.0 on Terminal-Bench 2.1, and Epoch AI's independent harness puts it at about 78.7% SWE-bench Verified, within single digits of the closed frontier.
The reason GLM-5.2 is the value pick is price. It lists at $1.40/M input and $4.40/M output on Z.ai, and Morph serves it at $0.8415/$3.1365, below Fireworks, Baseten, and Z.ai's own first-party rate. That is roughly a quarter of frontier output price for the leading open-weights intelligence. The full deep dive is on the GLM-5.2 page.
Hardware
Full bf16 deployment of a 744B MoE needs a multi-GPU node (practically 8x H200 or 8x B200). Practitioners on r/LocalLLaMA describe five RTX Pro 6000s plus a 5090 to serve GLM-5.2 at usable speed locally, which is why most teams rent it hosted rather than build the rack. One tradeoff worth knowing: GLM-5.2 is verbose, averaging around 43K output tokens per task on the Intelligence Index, so budget by tokens-per-task, not just the per-token rate.
DeepSeek V4: Top SWE-bench Verified Among Downloadable Weights
DeepSeek V4 leads raw SWE-bench Verified among downloadable open weights. The Pro-Max variant scores 80.6% (vendor), tied with Gemini 3.1 Pro, and the whole family is MIT licensed. V4 shipped April 24, 2026 and reached general availability July 19, 2026.
Two variants ship under MIT. V4 Pro (1.6T total / 49B active) costs $0.435/M input, $0.87/M output on the official API. V4 Flash (284B / 13B active) scores 79.0% (Flash-Max) and costs $0.14/$0.28, the cheapest frontier-adjacent open model, and Morph serves it at $0.09875/$0.278. Deeper coverage on the DeepSeek V4 page.
The 80.6% figure is vendor-reported. On the standardized SWE-bench Pro harness DeepSeek V4 drops to roughly 55%, and NIST's CAISI evaluation places it about eight months behind the proprietary frontier. It is genuinely the top downloadable-weights model on Verified, but treat the headline number as an optimistic ceiling and weight the independent harnesses for production planning.
MiniMax M3: Near-Frontier Verified, Community License
MiniMax M3 scores 80.5% SWE-bench Verified (vendor), a hair behind DeepSeek V4 Pro, with a roughly 1M-token context. It is strong for agentic tool-calling workflows and is served cheaply. The one caveat: it ships under a community license, not MIT, so read the terms before commercial self-hosting.
Morph serves MiniMax M3 (morph-minimax3-428b) at $0.255/$1.02 with bf16 activations, matching OpenRouter's standard list price. At $1.02/M output against a frontier model's $25/M, it is roughly a tenth of the price for a Verified score within single digits, which is the core trade the open tier offers.
Qwen3-Coder-Next: 70.6% SWE-bench on 3B Active Parameters
Qwen3-Coder-Next is the efficiency pick. An 80B MoE that activates only about 3B parameters per token, it hits 70.6% SWE-bench Verified while running on roughly 46GB of unified memory. Apache 2.0 licensed, 256K native context (extendable), it is the frontier-adjacent open model you can actually run on one machine.
With 2-bit quantization the footprint drops to about 30GB, and a 30B Flash variant runs on 18GB (an RTX 4090 or a 48GB Mac). Official pricing on Alibaba Cloud is $0.11/M input, $0.80/M output. If your constraint is running an open coder locally without a GPU cluster, this is the model; if you rank purely by Verified score, the general Qwen 3.6 line is higher (73.4% vendor) but is not a dedicated coder.
No other model in the frontier-adjacent tier approaches this efficiency. Qwen3-Coder-Next fits a 48GB unified-memory laptop; the 30B Flash variant fits 18GB. For everything else on this page, you are renting hosted compute or building a multi-GPU node.
gpt-oss (OpenAI): Apache 2.0 Weights, Now a Tier Behind
OpenAI's open-weight models, gpt-oss-120b (117B total / 5.1B active) and gpt-oss-20b (21B / 3.6B active), shipped in August 2025 under Apache 2.0. They are download-only, not served on OpenAI's API, and gpt-oss-120b fits a single 80GB GPU while gpt-oss-20b runs on 16GB.
OpenAI positioned gpt-oss-120b qualitatively (near o4-mini) rather than publishing a single SWE-bench Verified figure, and community reproductions vary widely by harness. Nearly a year old, they trail the 2026 open models above on coding. They remain a solid pick for offline and on-device use where a permissive license and small footprint matter more than topping a leaderboard.
Best For: Which Model for Which Use Case
Best for frontend / UI: Kimi K3
#1 on the Arena.ai Frontend Code Arena, ahead of Claude Fable 5. 93.4% SWE-bench Verified (Vals AI). 2.8T MoE, 1M context. Weights due ~July 27; hosted at ~100 tok/s on Morph today.
Best value: GLM-5.2
Top open model on the Artificial Analysis Intelligence Index. 744B/40B active, MIT, 1M context. $0.8415/$3.1365 on Morph, roughly a quarter of frontier output price.
Top SWE-bench Verified: DeepSeek V4 Pro
80.6% SWE-bench Verified (vendor), highest among downloadable weights. MIT license, 1.6T/49B active, 1M context. Read the number as an upper bound; independent harnesses run lower.
Cheapest: DeepSeek V4 Flash
79.0% SWE-bench Verified (Flash-Max, vendor) at $0.14/$0.28 per million tokens. MIT, 284B/13B active, 1M context. The near-free frontier-adjacent open model.
Best for local dev: Qwen3-Coder-Next
70.6% SWE-bench Verified with ~3B active parameters. Runs on 46GB unified memory; 30B Flash variant on 18GB. Apache 2.0. The frontier-adjacent model you can run on one machine.
Near-frontier, agentic: MiniMax M3
80.5% SWE-bench Verified (vendor), strong tool calling, ~1M context, $0.255/$1.02 on Morph. Community license (not MIT), so check terms before commercial self-hosting.
Offline / on-device: gpt-oss
gpt-oss-120b (single 80GB GPU) and gpt-oss-20b (16GB), Apache 2.0. A tier behind 2026 open models on coding, but a solid permissive pick for offline and on-device use.
Highest general Verified (not a dedicated coder): Qwen 3.6
73.4% SWE-bench Verified (vendor) as the 35B-A3B general flagship, Apache 2.0. Higher on Verified than Qwen3-Coder-Next but built as a general model, not a purpose-built coder.
API Pricing Comparison
All of these models are available via hosted APIs. Open weights do not mean you have to self-host. The cost advantage over proprietary models is large: 80 to 95% cheaper on output tokens in most cases.
| Model | Input | Output | Notes |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | Cheapest frontier-adjacent option (MIT) |
| Qwen3-Coder-Next | $0.11 | $0.80 | Alibaba Cloud; runs on 46GB locally |
| DeepSeek V4 Pro | $0.435 | $0.87 | 80.6% Verified (vendor), MIT |
| MiniMax M3 (Morph) | $0.255 | $1.02 | 80.5% Verified; community license |
| GLM-5.2 | $1.40 | $4.40 | $0.8415 / $3.1365 on Morph; top open Intelligence Index |
| Kimi K3 | $3.00 | $15.00 | Moonshot; $0.30 cached input; #1 frontend |
Prices from provider pricing pages, July 2026. Morph rates for hosted open models are lower where noted (GLM-5.2 $0.8415/$3.1365, DeepSeek V4 Flash $0.09875/$0.278, MiniMax M3 $0.255/$1.02). For comparison, Claude Opus 4.8 costs $5/$25 and GPT-5.6 Sol $5/$30 per million tokens.
How WarpGrep Fits In
Every model above shares the same bottleneck on hard coding tasks: finding the right code in large repositories. On SWE-Bench Pro, context overflow causes 35.6% of failures for top models. Coding agents spend 60%+ of their time on search.
WarpGrep v2 is an RL-trained search subagent that runs alongside any coding model. It operates in its own context window, issues up to 8 parallel tool calls per turn, and returns only relevant file spans. The main model never sees files WarpGrep rejected, keeping context clean.
| Base Model | Without WarpGrep | With WarpGrep v2 | Delta |
|---|---|---|---|
| GPT-5.6 Sol (CLI) | 57.0% | 59.1% | +2.1 |
| GLM-5.2 | 55.4% | 57.6% | +2.2 |
| DeepSeek V4 Pro | 55.4% | 57.5% | +2.1 |
WarpGrep is model-agnostic. It works with every model listed on this page, open or proprietary. Pairing it with a frontier model makes the system 15.6% cheaper and 28% faster on SWE-Bench Pro tasks, because the expensive model spends less time doing its own search.
Frequently Asked Questions
What is the best open-source coding model in 2026?
It depends on the job. Kimi K3 leads the Arena.ai Frontend Code Arena ahead of Fable 5 and scores 93.4% SWE-bench Verified (Vals AI). GLM-5.2 is the top open model on the Artificial Analysis Intelligence Index and the value pick at $0.8415/$3.1365 on Morph. DeepSeek-V4-Pro-Max leads raw SWE-bench Verified among downloadable weights at 80.6% (vendor), MIT-licensed. Qwen3-Coder-Next (70.6%) is the best you can run locally on 46GB.
Can open coding models compete with Claude and GPT?
On frontend, one leads: Kimi K3 is #1 on the Frontend Code Arena ahead of Claude Fable 5. On general SWE-bench Verified the top open model (DeepSeek V4 Pro, 80.6% vendor) trails GPT-5.6 Sol (96.2% Vals AI), Fable 5 (95.0%), and Kimi K3 (93.4%), but at roughly a tenth of the output price. The gap has closed to a few points on most benchmarks and reversed on frontend UI.
Is Kimi K3 open source, and can I download the weights?
Not yet. Kimi K3 launched July 16, 2026 with an open-weights release planned on Hugging Face around July 27, so as of late July you reach it through hosted APIs (Moonshot at $3/$15, or Morph at ~100 tok/s). It is a 2.8T-parameter MoE with a 1M context. The license is expected to follow the K2 line's Modified MIT terms but is not confirmed until the weights ship.
What hardware do I need to run an open coding model locally?
Qwen3-Coder-Next runs on about 46GB of unified memory (30GB with 2-bit quantization); the 30B Flash variant runs on 18GB (an RTX 4090 or a 48GB Mac). GLM-5.2 (744B MoE) and DeepSeek V4 Pro (1.6T) need multi-GPU nodes, which is why most teams rent them hosted rather than build the rack.
What is the cheapest open coding model API?
DeepSeek V4 Flash at $0.14/M input, $0.28/M output (MIT, 1M context), or Qwen3-Coder-Next at $0.11/$0.80 on Alibaba Cloud. Both are 80 to 95% cheaper on output than frontier proprietary models.
What is the difference between DeepSeek V4 Pro and V4 Flash?
Both shipped April 24, 2026 under MIT with a 1M-token context. V4 Pro (1.6T / 49B active) scores 80.6% SWE-bench Verified (vendor) at $0.435/$0.87. V4 Flash (284B / 13B active) scores 79.0% and costs $0.14/$0.28, making it the cheapest frontier-adjacent open model.
Why is GLM-5.2 the value pick?
It is the top open model on the Artificial Analysis Intelligence Index while pricing well below the frontier: $1.40/$4.40 on Z.ai, $0.8415/$3.1365 on Morph. MIT license, 744B/40B active, 1M context. Leading open-weights intelligence at roughly a quarter of frontier output price is the whole pitch.
Is gpt-oss competitive in 2026?
Not at the frontier. gpt-oss-120b and gpt-oss-20b (Apache 2.0, August 2025) are download-only and trail Kimi K3, GLM-5.2, and DeepSeek V4 on 2026 coding benchmarks. They remain a solid pick for offline and on-device use where a permissive license and small footprint matter most.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.
WarpGrep: Search Subagent for Any Coding Model
WarpGrep v2 lifts every model it is paired with by 2-4 points on SWE-Bench Pro. It runs in its own context window, issues 8 parallel tool calls per turn, and makes your coding agent cheaper and faster. Works with open-source and proprietary models.
