Best AI Coding Agents (October 2026): The Scored Leaderboard and Coding Agent Benchmark, Updated After Sonnet 5.5 and GPT-6.1 Sol

Updated October 6, 2026. The coding agent benchmark leaderboard: Claude Code + Opus 5.5 leads Terminal-Bench 4.0 at 64.8% and Claude Code + Sonnet 5.5 scores 61.8%. Codex + GPT-6.1 Sol ties GPT-6 Astra at 58.2% for $1.92 per task, a fifth of Astra's cost.

June 9, 2026 ยท 1 min read
Best AI Coding Agents (October 2026): The Scored Leaderboard and Coding Agent Benchmark, Updated After Sonnet 5.5 and GPT-6.1 Sol
Short answer (October 6, 2026)

The best AI coding agents are Claude Code and Codex. On Terminal-Bench 4.0, Claude Code with Opus 5.5 holds the top agent entry at 64.8%, and Claude Code with Sonnet 5.5 scores 61.8%. Codex with GPT-6.1 Sol matches GPT-6 Astra at 58.2% for $1.92 per task, about a fifth of Astra's $9.90. Pick Claude Code with Opus 5.5 for the highest score, and Codex with GPT-6.1 Sol for the lowest cost per solved task.

Three launches landed in the last week of September. Anthropic released Claude Sonnet 5.5 on September 28, and Claude Code v2.1.284 made it the default Sonnet. OpenAI released GPT-6.1 Sol on September 29 and rolled it out to paid Codex plans. Google announced Gemini 4 Argon on September 30, though only trusted cyber defenders in its Fairwind Program can use it so far. Claude Opus 5.5, released September 22, now holds the Claude Code record on Terminal-Bench 4.0.

Terminal-Bench 2.1 is now saturated. Nine models score between 87% and 92% on it, so tbench.ai moved its agent leaderboard to Terminal-Bench 4.0, where the top entry is 64.8%. This page scores agents on 4.0 first and keeps 2.1 and SWE-bench Pro for reference. Price is the entry paid tier plus the per-token or per-credit rate. Nothing here is a paid placement.

64.8%
Top Terminal-Bench 4.0 agent (Claude Code + Opus 5.5, max)
61.8%
Claude Code + Sonnet 5.5 (max), Terminal-Bench 4.0
$1.92
Per task, Codex + GPT-6.1 Sol max at 58.2% (vs $9.90 Astra)
$2 / $10
Sonnet 5.5 and GPT-6.1 Sol per 1M tokens

The Scored Leaderboard: AI Coding Agents (October 2026)

One row per agent paired with its strongest available model. Terminal-Bench 4.0 is the agent-plus-model entry from tbench.ai (330 trials each). Terminal-Bench 2.1 is the model-level score from Artificial Analysis on one fixed harness (Terminus 2), marked "model"; the separate agent-level 2.1 board on Harbor Hub is marked "agent". SWE-bench Pro is from the BenchLM aggregator. Open-source agents run any model, so their benchmark equals the model row you choose. Terminal-Bench 4.0 entries re-checked 2026-10-06; the Updated column dates each row.

AI Coding Agents, Scored (October 6, 2026)
Agent / ModelTerminal-Bench 4.0Terminal-Bench 2.1SWE-bench ProPricing modelUpdated
Claude Code / Opus 5.5 (default Opus)64.8% (max)not yet run89.9% (aggregator, unconfirmed)$17/mo Pro; API $4 / $202026-09-22
Claude Code / Sonnet 5.5 (default Sonnet)61.8% (max)not yet runnot published$17/mo Pro; API $2 / $102026-09-28
Codex / GPT-6 Astra58.2% (max)89.9% (high, model); 87.4% (agent)not published$20/mo Plus, 5-45 Astra msgs/5h2026-09-03
Codex / GPT-6.1 Sol58.2% (max), $1.92/tasknot yet runnot published$20/mo Plus and up (not Free or Go); API $2 / $102026-09-29
Claude Code / Fable 5.157.9% (max)91.4% (max, model)81.2%$17/mo Pro (annual)2026-09-01
Codex / GPT-6 Sol49.4% (max)not yet runnot published$20/mo Plus, 15-150 Sol msgs/5h2026-09-22
Claude Code / Opus 553.9% (xhigh)89.1% (max, model)79.2%$17/mo Pro (annual)2026-07-24
Grok Build / Grok 4.737.6% (xhigh)no entryno entrySuperGrok plans; API $2 / $6 per 1M2026-09-21
Codex / GPT-5.6 Sol37.3% (max)89.5% (xhigh, model)64.6%$20/mo Plus (legacy model)2026-06-26
Gemini CLI / Gemini 3 mixno entry87.6% (3.8 Flash, model)no entryFree, 1,000 req/day2026-09-15
GitHub Copilot / multi-modelmodel-setmodel-setmodel-set$0.01/credit, $10/mo Pro2026-09-22
Cursor / Grok 4.7 + Composer 2.5 + frontiermodel-set79.3% (Cursor CLI + Grok 4.5, agent)model-set$20/mo Pro, $60 Pro Plus, $200 Ultra2026-09-21
OpenCode / any (BYOK)53.6 AA Coding Agent Index (GLM-5.3)model-setmodel-setFree; Go $10/mo2026-09-22
Cline / any (BYOK)model-setmodel-setmodel-setFree; ClinePass $9.99/mo2026-09-22
Goose / any (BYOK)model-setmodel-setmodel-setFree, BYOK or subs2026-09-17
Aider / any (BYOK)model-setmodel-setmodel-setFree, pay model provider2026-05-22
Kilo Code / any (BYOK)model-setmodel-setmodel-setFree, no-markup gateway2026-09-22
Kiro / Sonnet 4.5 + openmodel-setmodel-setmodel-setFree / $20/mo, credits2026-09-16
Google Antigravity / Gemini 3.8 Flashno entry87.6% (model)no entryFree Individual; AI Pro/Ultra raise limits2026-09-22
Morph router / open modelsmodel-setmodel-setmodel-set$0.005/request; $0.141953125 / $0.399625 per 1M (dsv4flash)2026-09-22

BYOK = Bring Your Own Key: the agent is free and you pay the model provider or run a local model. "model-set" means the agent runs whatever model you point it at, so its score equals the model row you choose. "Not yet run" means that benchmark has no published score for the pair. The Opus 5.5 SWE-bench Pro figure (89.9%) appears on BenchLM but not in Anthropic's launch table, so treat it as unconfirmed. Anthropic's Fable 5.1, Opus 5.5, and Sonnet 5.5 announcements report Terminal-Bench 4.0 and CursorBench instead of SWE-bench. The Gemini CLI date is its v0.60.0 release. The OpenCode Terminal-Bench 4.0 cell is its Artificial Analysis Coding Agent Index v1.5 score, which includes Terminal-Bench 4.0 as one of three parts.

Terminal-Bench 4.0: Coding Agent Leaderboard (October 6, 2026)

Higher is better. Agent + model pairs from tbench.ai, 330 trials each. Badge shows cost per task.

1
Claude Code + Opus 5.5 (max)
$14.30/task
64.8%
2
Claude Code + Sonnet 5.5 (max)
$22.22/task
61.8%
3
Codex + GPT-6 Astra (max)
$9.90/task
58.2%
4
Codex + GPT-6.1 Sol (max)
$1.92/task
58.2%
5
Claude Code + Fable 5.1 (max)
$18.92/task
57.9%
6
Codex + GPT-6 Astra (xhigh)
$7.12/task
57.9%
7
Claude Code + Fable 5.1 (xhigh)
$14.76/task
57.9%
8
Claude Code + Opus 5 (xhigh)
$18.44/task
53.9%
9
Grok Build + Grok 4.7 (xhigh)
$11.16/task
37.6%
10
Codex + GPT-5.6 Sol (max)
$7.70/task
37.3%

Cost per task = total run cost on tbench.ai divided by 330 trials. Anthropic's Sonnet 5.5 launch table reports 70.6% at max effort on its own harness; the tbench.ai Claude Code entry is 61.8%.

The Coding Agent Benchmarks, Explained (Terminal-Bench 4.0, 2.1, SWE-bench)

Terminal-Bench 4.0 is the benchmark that separates agents in October 2026. It is 66 harder tasks of terminal work across software, machine learning, science, operations, security, hardware, and media. tbench.ai runs each agent-plus-model pair for 330 trials and publishes total tokens and dollar cost, so you can compare price per solved task, not only accuracy. Terminal-Bench 2.1 is still published at the model level but nine models now score between 87% and 92%. SWE-bench Verified (500 human-validated Python bug fixes) topped out at 97.0% with Claude Opus 5 on Vals.ai, which archived the benchmark on September 1. SWE-bench Pro is the harder, contamination-resistant set, but Scale's official leaderboard has not added a model since July 9, so the Pro numbers here come from the BenchLM aggregator.

What Each Benchmark Measures
BenchmarkMeasuresScored unitOctober 2026 leader
Terminal-Bench 4.066 hard terminal tasks, 330 trials, with costAgent + modelClaude Code + Opus 5.5 (64.8%)
Terminal-Bench v2 (2.1)End-to-end terminal task completionModel (Terminus 2)Fable 5.1 max (91.4%)
Terminal-Bench 2.1 agent boardSame 2.1 tasks, run per agent (Harbor Hub)Agent + modelCodex + GPT-6 Astra high (87.4%)
AA Coding Agent Index v1.5DeepSWE v1.1 + Terminal-Bench 4.0 + SWE-Atlas-QnA, equal weightAgent + modelClaude Code + Fable 5.1 max (62.2)
SWE-bench ProHarder, contamination-resistant issuesModelFable 5.1 (81.2%, BenchLM)
SWE-bench VerifiedGitHub bug-fix in Python, human-validatedModelClaude Opus 5 (97.0%, Vals.ai, archived Sep 1)
CursorBench 4.0Cursor's coding agent benchmarkModelOpus 5.5 (57.8%, Anthropic-reported)

Effort level is part of the score. On Terminal-Bench 4.0, Codex + GPT-6 Astra scores 57.9% at high, xhigh, and max, but costs $6.88, $7.12, and $9.90 per task. Paying for max bought 0.3 points. Claude Code + Fable 5.1 shows the same pattern: 57.9% at both xhigh ($14.76) and max ($18.92). The Artificial Analysis Coding Agent Index ranks the same pairs differently because it adds DeepSWE and SWE-Atlas-QnA: Claude Code + Fable 5.1 scores 62.2, Devin Fusion CLI 61.7, and Codex + GPT-6 Astra 61.6 (index v1.5, September 22). Vals AI runs every model on one harness, mini-SWE-agent, and as of October 6 ranks Claude Opus 5.5 first at 65.15%, Sonnet 5.5 second at 64.14%, and GPT-6 Astra third at 59.60%.

What Changed: September, August, and July 2026 (Sonnet 5.5, GPT-6.1 Sol, Opus 5.5, GPT-6, Fable 5.1)

Updated October 6, 2026. The last week of September reshuffled the top of the leaderboard. Claude Code v2.1.284 (September 28) added Claude Sonnet 5.5 as the default Sonnet on the Anthropic API. tbench.ai now lists Claude Code + Opus 5.5 at 64.8% and Claude Code + Sonnet 5.5 at 61.8%, both ahead of every Codex entry. OpenAI released GPT-6.1 Sol on September 29 at $2 / $10 per 1M tokens, a fifth of Astra's $10 / $50. On tbench.ai it matched Astra's 58.2% for $634 across 330 trials, where the Astra run cost $3,267. Earlier, on September 22, Claude Code v2.1.280 made Opus 5.5 the default Opus and moved Pro and Team Standard from Sonnet to Opus by default.

September 2026
  • 2026-09-30: Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. API pricing starts at an introductory $2 / $10 per 1M tokens, then $4 / $20
  • 2026-09-29: GPT-6.1 Sol ships at $2 / $10 per 1M tokens ($0.10 cached input), 1.05M context, 128K max output. The Codex rollout covers Plus, Pro, Business, Enterprise, and Edu; Free and Go are not included. Codex + GPT-6.1 Sol scores 58.2% on Terminal-Bench 4.0 at $1.92 per task
  • 2026-09-28: Claude Sonnet 5.5 ships at $2 / $10 per 1M tokens ($0.20 cache reads), 1M context, and becomes the default Sonnet in Claude Code v2.1.284. Anthropic reports 70.6% on Terminal-Bench 4.0; the tbench.ai Claude Code entry is 61.8%. Haiku 5.5 is due in the coming weeks
  • 2026-09-22: Claude Opus 5.5 ships at $4 / $20 per 1M tokens ($0.20 cache reads), 1M context, 128K max output; default Opus in Claude Code v2.1.280. Fast mode is $8 / $40
  • 2026-09-22: GPT-6 Sol ($2 / $10, 1.05M context) and GPT-6 Luna ($0.10 / $0.50) ship to Codex and ChatGPT Work (not Chat). Codex CLI 0.156.0 adds a fullscreen /tui, a /usage dashboard, and turns worktrees on by default
  • 2026-09-22: Copilot adds Opus 5.5, GPT-6 Sol, and GPT-6 Luna (Grok 4.7 on September 21)
  • 2026-09-22: Artificial Analysis Coding Agent Index v1.5: Claude Code + Fable 5.1 62.2, Devin Fusion CLI 61.7, Codex + GPT-6 Astra 61.6, OpenCode + GLM-5.3 53.6
  • 2026-09-21: Grok 4.7 ships in Cursor, Grok Build, and the xAI API ($2 / $6 per 1M, 500K context); Grok Build + Grok 4.7 enters Terminal-Bench 4.0 at 37.6%. Devin CLI adds devin --cloud to run Devin Cloud sessions from the terminal
  • 2026-09-13: Amp goes free: no Amp token fees or limits when you bring your own ChatGPT subscription, API key, and runners
  • 2026-09-11: OpenCode V2 (2.0.0) ships on npm as @opencode/cli with a new plugin API; V1 continues as 1.18.x
  • 2026-09-10: Cursor Projects: coordinator agents that delegate to subagents in the cloud. Cognition ships SWE-2, post-trained from Kimi K3, free in Devin Desktop (formerly Windsurf) and Devin CLI on Pro through October 10
  • 2026-09-10: OpenAI pauses new ChatGPT Pro $200 sign-ups; Pro $100 stays open
  • 2026-09-03: GPT-6 Astra ships at $10 / $50 per 1M tokens, 1.05M context; Codex + Astra takes #1 on Terminal-Bench 4.0 (58.2%)
  • 2026-09-02: Meta ships Muse Spark 1.3 to Muse Code and its Model API ($1.25 / $4.25 per 1M, 1M context)
  • 2026-09-02: Gemini 3.8 Flash ships at $0.75 / $3.75 introductory pricing through December 31 ($1.50 / $7.50 after), available in Antigravity
  • 2026-09-01: Claude Fable 5.1 (generally available, $10 / $50) and Mythos 5.1 (limited to Project Glasswing trusted-access customers). Anthropic puts Fable 5.1 about 25% cheaper than Fable 5 on typical workloads
August and July 2026
  • 2026-08-28: Terminal-Bench 4.0 launches on tbench.ai (66 tasks)
  • 2026-08-19: Cursor cloud agents get event subscriptions, /goal, and subagents on isolated VMs
  • 2026-08-17: Cursor launches Origin code hosting with GitHub sync
  • 2026-08-14: SpaceX completes its acquisition of Cursor
  • 2026-08-05: Meta launches Muse Code in beta, a terminal coding agent on its Muse Spark model that splits large-repo work across parallel subagents. It now costs $5, $15, or $50/mo, and Muse Code + Muse Spark 1.3 scores 54.3 on the Artificial Analysis Coding Agent Index
  • 2026-07-24: Claude Opus 5 ships and becomes the Claude Code default Opus (now a legacy model after Opus 5.5)
  • 2026-07-15: Anaconda announces it has acquired Kilo Code
  • 2026-07-09: GPT-5.6 goes GA in Sol, Terra, and Luna tiers (now superseded by GPT-6 Sol and Luna in Codex)
  • 2026-07-01: Anthropic restores Fable 5 after the June export order is lifted

Best Open-Source AI Coding Agents (2026)

The open-source agents are free to install and run on any model you point them at. Ranked by GitHub stars on September 22, 2026, with the latest release:

Open-Source Coding Agents by GitHub Stars (September 2026)
AgentStarsLicenseLatest release
OpenCode209,405MITv1.18.32 (Sep 21); V2 2.0.14 (Sep 22)
OpenAI Codex CLI125,962Apache-2.00.156.0 (Sep 22)
Gemini CLI107,130Apache-2.0v0.60.0 (Sep 15)
Cline69,074Apache-2.0v4.1.20 (Sep 22)
Goose54,568Apache-2.0v1.51.0 (Sep 17)
Aider49,118Apache-2.00.86.2 (Feb 12); last commit May 22
Kilo Code27,390MITv7.7.7 (Sep 22)

anthropics/claude-code has 147,649 stars but the tool is proprietary (no license; the repo holds issues and the changelog), so it is not in this list. Codex CLI passed Gemini CLI for second place since June. Aider has not committed to main since 2026-05-22; every other repo here shipped a release in September. Roo Code shut down its extension on May 15, 2026 and archived the repo; it points users to the ZooCode fork and Cline.

1. OpenAI Codex

58.2%
Terminal-Bench 4.0 (GPT-6.1 Sol or GPT-6 Astra, max)
0.156.0
Codex CLI release, September 22
$20/mo
ChatGPT Plus entry tier

Codex's best Terminal-Bench 4.0 entries are 58.2% with GPT-6 Astra and 58.2% with GPT-6.1 Sol, both at max effort and six points behind Claude Code + Opus 5.5. The Sol run is the cheapest strong result on the board at $1.92 per task. GPT-6.1 Sol (September 29, $2 / $10 per 1M tokens) is rolling out in the Codex desktop app and CLI on Plus, Pro, Business, Enterprise, and Edu; Enterprise and Edu admins must turn it on. Codex also offers three GPT-6 models. GPT-6 Astra ($10 / $50 per 1M tokens) is the most capable. GPT-6 Sol ($2 / $10) is the balanced option and the named replacement for GPT-5.5 on paid plans; OpenAI says it makes about half as many mistakes as GPT-5.6 Sol. GPT-6 Luna ($0.10 / $0.50) is for high-volume focused tasks. Astra and Sol carry a 1.05M-token context window and 128K max output. GPT-6 Sol prompts over 272K input tokens are billed at 2x input and 1.5x output for the whole request, which matters for long agent sessions.

Codex CLI 0.156.0 (September 22) adds an optional fullscreen UI via /tui, a /usage analytics dashboard for token totals and plugin and skill activity, voice on by default (F8 toggle), and worktree sessions on by default. Install with npm install -g @openai/codex or brew install --cask codex, run codex, and sign in with ChatGPT. Switch models in-session with /model or start with codex --model gpt-6-sol. GPT-6 Astra has been the bundled default model since CLI 0.153.4. API-key mode bills per token. GPT-5.5 leaves Codex on October 14.

Pricing and limits (local messages per 5-hour window)
  • Free ($0) and Go ($8/mo): GPT-6 Luna in the desktop app
  • Plus ($20/mo): GPT-6 Luna 350-3,000, GPT-6 Sol 15-150, GPT-6 Astra 5-45
  • Pro 5x (from $100/mo): Luna 1,750-14,000, Sol 70-700, Astra 25-225
  • Pro 20x ($200/mo, new sign-ups paused since September 10): Luna 7,000-56,000, Sol 300-3,000, Astra 100-900
  • Business ($20/user/mo annual or $25 monthly, 2+ users): Plus-level limits; Enterprise and Edu are custom
  • Credit rates per 1M tokens: Luna 2.5 in / 12.5 out, Sol 50 / 250, Astra 250 / 1,250

Best for: Developers who want a near-top Terminal-Bench 4.0 score at the lowest cost per task. Astra is expensive against plan limits: 5-45 messages per 5 hours on Plus, versus 15-150 for Sol. See the Codex vs Claude Code comparison.

2. Claude Code (Anthropic)

64.8%
Terminal-Bench 4.0 (Opus 5.5, max), #1 agent entry
$4 / $20
Opus 5.5 per 1M tokens, new default
$17/mo
Pro tier (annual), includes Claude Code

Claude Code v2.1.280 (September 22) made Claude Opus 5.5 the default Opus: 1M-token context, $4 / $20 per 1M tokens, and $0.20 per 1M cache reads (60% below Opus 5). In Anthropic's launch table Opus 5.5 scores 66.4% on Terminal-Bench 4.0 against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra; Artificial Analysis's independent run puts Opus 5.5 and Astra level at 59.6%. The API default effort for Opus 5.5 is medium. On tbench.ai, Claude Code + Opus 5.5 at max effort holds the top Terminal-Bench 4.0 agent entry at 64.8% ($14.30 per task). Claude Fable 5.1 ($10 / $50) scores 57.9% in Claude Code and leads confirmed SWE-bench Pro scores at 81.2%. Sonnet 5.5 ($2 / $10, September 28) has been the default Sonnet since v2.1.284 and scores 61.8% in Claude Code. Haiku 4.5 ($1 / $5) covers the cheap end. The same release moved Pro and Team Standard from Sonnet to Opus by default, so every subscription plan now starts on Opus 5.5.

One config flag worth knowing from v2.1.280: CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH raises the 2,048-character cap Claude Code applies to every MCP tool description and server instruction. MCP servers with long tool docs get silently truncated without it. Since v2.1.277 (September 18), Claude Code reads AGENTS.md in projects that have no CLAUDE.md. Install with curl -fsSL https://claude.ai/install.sh | bash (macOS, Linux, WSL), brew install --cask claude-code, or npm install -g @anthropic-ai/claude-code. Add MCP servers with claude mcp add.

Pricing and limits
  • Free: $0, Sonnet and Haiku in chat; does not include Claude Code
  • Pro: $17/mo billed annually ($200 up front) or $20/mo monthly; includes Claude Code and Opus 5.5; Fable 5.1 runs on usage credits
  • Max: $100/mo (5x Pro usage) or $200/mo (20x); Fable 5.1 up to 50% of weekly limits
  • Fast mode: Opus 5.5 at $8 / $40 per 1M tokens, up to 2.5x faster in Claude Code
  • Also runs on: Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS

Best for: Developers who want the strongest Anthropic models in a terminal and IDE agent. Opus 5.5 at $4 / $20 costs less per token than Opus 5 did at $5 / $25, and Anthropic reports it runs 30% faster. Compare with Cursor.

3. Gemini CLI

1,000/day
Free requests on a personal account
107,130
GitHub stars (Apache-2.0)
v0.60.0
Latest release, September 15

Gemini CLI still has the most generous free tier on this list: 60 requests per minute and 1,000 requests per day with a personal Google account, served by a managed Gemini 3 mix of flash and pro models with a 1M-token context. Google's newest model, Gemini 3.8 Flash (September 2), scores 87.6% on Terminal-Bench 2.1 at the model level. Gemini CLI itself has no Terminal-Bench 4.0 entry; the only Gemini row there is mini-SWE-agent + Gemini 3.8 Flash at 19.1%.

v0.60.0 is a security-hardening release: stricter web-fetch destination checks, RFC 9207 issuer checks in MCP OAuth, an isolated temp directory for the macOS Seatbelt sandbox, and consent prompts when extensions change environment variables. Install with npx @google/gemini-cli, npm install -g @google/gemini-cli, or brew install gemini-cli. MCP servers go in ~/.gemini/settings.json. With an API key you can pin a specific Gemini model.

Best for: Developers who want a capable agent at zero cost with high daily limits. Compare with Codex and Claude Code.

4. GitHub Copilot

$0.01
Per AI credit (usage-based since June 1)
Free tier
2,000 completions/mo
$10/mo
Pro: 1,500 credits ($15 value)

Copilot bills in GitHub AI Credits (1 credit = $0.01), which replaced premium requests on June 1, 2026. Credits are consumed on token usage at per-model rates. The model menu moves fast: Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna arrived on September 22 and Grok 4.7 on September 21. On October 19 GitHub deprecates GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Gemini 3.7 Flash, and Grok 4.5. Copilot CLI is at v1.0.88 (September 22). Since September 14 you can set a cost and quality preference for Copilot's auto model selection instead of picking a model per request.

Install the CLI with npm install -g @github/copilot, brew install copilot-cli, or winget install GitHub.Copilot. It supports MCP servers and a /model slash command.

Pricing (AI Credits, September 22, 2026)
  • Free: $0, 2,000 completions/mo plus a limited credit allowance
  • Pro ($10/mo): 1,000 base + 500 flex = 1,500 credits
  • Pro+ ($39/mo): 3,900 base + 3,100 flex = 7,000 credits
  • Max ($100/mo): 10,000 base + 10,000 flex = 20,000 credits
  • Business ($19/seat/mo): 1,900 credits per user; Enterprise ($39/seat/mo): 3,900. Card and PayPal seats are charged up front from October 1

Best for: Teams on GitHub who want completions, chat, agent mode, cloud agent, and code review on one bill across VS Code, JetBrains, Neovim, and Xcode. Budget increase requests went GA on September 16, so developers can ask admins for more credits from inside Copilot.

5. Cursor

$20/mo
Pro, both model pools
$0.25
Cursor Token Rate per 1M, third-party models (Teams, Enterprise)
$0
Hobby tier, no card required

Cursor is a VS Code fork rebuilt around agents, and since August 14 it is part of SpaceX. Its first-party models are now the Grok family, co-trained with SpaceXAI, plus Composer 2.5; Grok 4.7 arrived on September 21. The last six weeks moved it toward cloud work: Origin code hosting with GitHub sync (August 17), always-on cloud agents with /goal and subagents on isolated VMs (August 19), self-hosted machines that keep tool execution in your network (September 2), and Projects (September 10), where coordinator agents split a feature or migration across cloud subagents.

Pricing has two pools. The Cursor Models pool covers Composer 2.5 and Grok 4.7, 4.6, and 4.5. The Other Models pool covers frontier models from Anthropic, OpenAI, and Google. All Auto modes bill at the list price of whichever model the request routes to. On Teams and Enterprise, third-party models also carry a Cursor Token Rate of $0.25 per 1M tokens, BYOK included; Cursor's first-party models are exempt. That fee is easy to miss when comparing Cursor's bill to raw API rates.

Pricing
  • Hobby: Free, no card, limited Agent requests, access to Composer
  • Pro: $20/mo, extended Agent limits
  • Pro Plus: $60/mo, 3x Pro Agent limits
  • Ultra: $200/mo, 20x Pro Agent limits
  • Teams: Standard $40/user/mo; Premium $120/user/mo with 5x Standard Agent limits
  • Start (India only): โ‚น649/mo, Cursor Models pool only
  • Yearly billing: 20% off

Best for: Developers who want inline AI editing and cloud agents without leaving the editor. See Cursor alternatives for the free options.

6. OpenCode

209,405
GitHub stars (most-starred OSS agent)
v1.18.32
V1 release, September 21 (V2 2.0.14 on Sep 22)
MIT
License, free to run

OpenCode (anomalyco/opencode, redirected from sst/opencode) is the most-starred open-source coding agent at 209,405 stars, up from 172,198 in June. It supports 75+ LLM providers through the AI SDK and the Models.dev catalog, plus local models via Ollama, LM Studio, and llama.cpp. OpenCode Zen is the team's curated list of models tested for agentic coding; v1.18.32 added Grok 4.7 and DeepSeek V4.1 Flash to it and fixed Together AI streaming usage reporting. OpenCode Go is a $10/mo subscription for open models. OpenCode V2 (npm @opencode/cli, 2.0.14 on September 22) ships alongside V1 with a new plugin API, so V1 plugins do not run on it. On the Artificial Analysis Coding Agent Index, OpenCode + GLM-5.3 scores 53.6 at $4.24 per task, the only open-source harness on the index.

Install with curl -fsSL https://opencode.ai/install | bash, npm install -g opencode-ai, or brew install anomalyco/tap/opencode. V2 installs with curl -fsSL https://opencode.ai/v2/install | bash and replaces the V1 binary. Add a custom OpenAI-compatible provider in JSON:

{
  "provider": {
    "myprovider": {
      "npm": "@ai-sdk/openai-compatible",
      "options": { "baseURL": "https://api.myprovider.com/v1" },
      "models": { }
    }
  }
}

Best for: Developers who want a Claude Code-style CLI without lock-in to one model vendor. Compare with Claude Code and Codex.

7. Cline

69,074
GitHub stars (Apache-2.0)
v4.1.20
Latest release, September 22
Free
Pay only for model usage

Cline is an Apache-2.0 agent for VS Code, JetBrains, Cursor, and Windsurf (now Devin Desktop), plus a CLI installed with npm i -g cline and a macOS desktop app. It runs Claude, GPT, Gemini, any OpenAI-compatible endpoint, BYOK, or local models via Ollama and LM Studio, with MCP servers and custom tools. Cline 4.0 (June 26) moved the extension onto the shared Cline SDK, turned command auto-approval off by default, and added ClinePass, a $9.99/mo subscription for open models such as GLM-5.3, Kimi K3, and DeepSeek V4. v4.1.20 (September 22) runs subagents in the same step in parallel and fixes compaction silently falling back to truncation.

Best for: VS Code or JetBrains users who want agentic AI without a subscription. Compare with Claude Code and Cursor.

8. Goose (Linux Foundation AAIF)

54,568
GitHub stars (Apache-2.0, Rust)
v1.51.0
Latest release, September 17
Free
Open source, BYOK or subscriptions

Goose lives at aaif-goose/goose under the Linux Foundation's Agentic AI Foundation. It is built in Rust and ships as a desktop app, a CLI, and an API. It works with many providers (Anthropic, OpenAI, Google, Ollama, OpenRouter, Azure, Bedrock), can reuse existing Claude, ChatGPT, or Gemini subscriptions via ACP, and connects to extensions over MCP. v1.51.0 added EUrouter as a declarative provider, GPT-live API support, an operator allowlist for gateway pairing, and an opt-in terminal bell when a turn finishes or needs approval. It also removed planning mode and the create recipe command from the CLI.

Install the CLI with curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash. Goose is general-purpose: research, writing, and automation alongside coding.

Best for: Developers who want a foundation-governed, provider-agnostic agent that also handles non-code automation. Compare with Claude Code.

9. Aider

49,118
GitHub stars (Apache-2.0)
Git-native
Auto-commits every change
2026-05-22
Last commit to main, four months ago

Aider is the Git-native terminal agent: every change is staged with a descriptive commit. Install with python -m pip install aider-install && aider-install or curl -LsSf https://aider.chat/install.sh | sh. It selects models per run by flag, for example aider --model sonnet --api-key anthropic=<key>, and runs local models via Ollama and any OpenAI-compatible API.

The caveat has grown. Aider's last commit to main was 2026-05-22 and its last PyPI release is 0.86.2 from February 12, four months before this update, while every other agent on this page shipped a release in September. Stars still rose from 45,945 to 49,118. It works with current models by flag, but new model and provider features land elsewhere first.

Best for: Terminal-native developers who want Git-integrated editing and full control over which model they pay for. Compare with Claude Code and Cline.

10. Kilo Code

27,390
GitHub stars (MIT)
No markup
Kilo Gateway at exact provider rates
500+
Models across 60+ providers

Kilo Code (Kilo-Org/kilocode) is free and open source for VS Code, JetBrains, and the CLI. The VS Code extension and CLI are at v7.7.7 (September 22); JetBrains stable is v7.1.6. Anaconda acquired Kilo in July 2026. The Kilo CLI is a fork of OpenCode and has adopted upstream through OpenCode v1.18.18, about five weeks behind. v7.7.3 removed KiloClaw and the /local-review aliases (use /review). The Kilo Gateway is $0/mo plus usage at exact provider rates across 500+ models from 60+ providers, with Auto model routing and an Auto Free mode that routes to free models. Kilo Pass runs $19, $49, and $199/mo with up to 50% bonus credits. Teams is $15/user/mo. BYOK and local models need no plan. One cost to note: credit purchases carry a 5% processing fee, so "no markup" applies to the token rate, not the top-up.

Best for: Developers who want no-markup gateway pricing or BYOK with optional bonus-credit passes. Compare with Claude Code.

11. Kiro

$0
Free tier: 50 credits/mo
$20/mo
Pro: 1,000 credits
$0.04
Per add-on credit

Kiro is a credit-based IDE. The free tier gives 50 credits/mo with Claude Sonnet 4.5 and open-weight models such as Qwen3 Coder Next, DeepSeek 3.2, MiniMax M2.5, and GLM-5. Pro is $20/mo for 1,000 credits, Pro+ $40 for 2,000, Pro Max $100 for 5,000, and Power $200 for 10,000. Add-on credits are $0.04 each. Plan credits do not roll over, but add-on credits roll over and expire 12 months after purchase. New users get $20 credited toward a first upgrade with social login or AWS Builder ID. On September 14 Kiro moved GPT-5.6 Sol, Terra, and Luna to 1M context at credit multipliers of 4.4x, 2.2x, and 1.1x up to 272K (doubled above that), and IDE 1.1 shipped native ARM64 builds. CLI 2.22.0 followed on September 16.

Best for: Developers who want a predictable credit budget and a free tier that includes Claude Sonnet 4.5. Compare with Cursor.

12. Google Antigravity

87.6%
Terminal-Bench 2.1 (Gemini 3.8 Flash, model)
$0
Individual plan (AI Pro $19.99/mo for more)
$0.75 / $3.75
Gemini 3.8 Flash per 1M, through Dec 31

Antigravity 2.0, announced at Google I/O on May 19, 2026, runs one harness across a desktop app and a standalone CLI, with subagents for parallel tasks. Gemini 3.8 Flash (September 2) is available in Antigravity. Google calls it its most intelligent workhorse model: 1,048,576-token input, 65,536-token output, and $0.75 / $3.75 per 1M tokens through December 31, 2026, then $1.50 / $7.50. It scores 87.6% on Terminal-Bench 2.1 at the model level. Antigravity has no agent-paired Terminal-Bench 4.0 entry yet. The app is at 2.16.0 (September 22, adding WSL support) and Antigravity CLI at 1.2.6.

The Individual plan is $0 with weekly rate limits and Gemini 3.8 Flash, 3.1 Pro, Claude Sonnet and Opus 4.6, and gpt-oss-120b as agent models. Organizations can now buy through Google Cloud at consumption-based API pricing. Google AI Pro is $19.99/mo with higher limits and a flexible AI credit pool. Google AI Ultra is $99.99/mo (5x Pro usage) or $199.99/mo (20x) with higher Antigravity limits. Google AI Plus at $4.99/mo gives 2x the free limits. The introductory Flash price ends December 31, so budget for the doubling.

Best for: Developers in the Gemini ecosystem who want visual parallel-agent management. Compare with Cursor and Claude Code.

How to Choose: Decision Framework

Pick Your Agent Based on Your Priority
Your PriorityBest ChoiceRunner-Up
Highest Terminal-Bench 4.0 scoreCodex + GPT-6 Astra (58.2%, max)Claude Code + Fable 5.1 (57.9%, max)
Lowest cost per solved task at the topCodex + GPT-6 Astra high ($6.88/task)Codex + Astra xhigh ($7.12/task)
Best model per dollar (new)Claude Opus 5.5 ($4 / $20)GPT-6 Sol ($2 / $10)
Highest confirmed SWE-bench Pro modelFable 5.1 (81.2%)Opus 5 (79.2%)
Cheapest frontier-family tokensGPT-6 Luna ($0.10 / $0.50)Gemini 3.8 Flash ($0.75 / $3.75)
Free, no API billGemini CLI (1,000 req/day)Antigravity Individual / Copilot Free / Cursor Hobby
Free + open sourceOpenCode (209K stars)Cline / Goose / Kilo Code
Terminal-first workflowCodex / Claude CodeOpenCode / Aider
Stay inside VS CodeCursor / CopilotCline / Kilo Code
No-markup model pricingKilo Code (exact provider rates)OpenCode / Cline (BYOK)
Predictable credit budgetKiro ($20 = 1,000 credits)Copilot Pro ($10 = 1,500)
Reuse existing subscriptionsGoose (ACP) / OpenCodeAmp Free (ChatGPT sub or BYOK) / Cline
Multi-agent cloud projectsCursor ProjectsGoogle Antigravity

Most developers settle on two or three agents. A common stack: Codex or Claude Code for heavy agent work, Copilot or Cursor for inline completions, and one free open-source agent (OpenCode, Cline, or Kilo Code) for model flexibility. Watch effort settings more than agent choice. On Terminal-Bench 4.0 the jump from xhigh to max added 0.3 points for Codex + Astra and nothing for Claude Code + Fable 5.1, while raising cost per task 28-39%.

The Model Layer: Where Cost and Quality Actually Live

Every agent above is a harness around a model. The Terminal-Bench 4.0 cost column makes the point: the same 57.9% score cost $7.12 per task on one pairing and $14.76 on another. Morph's model router picks the cheapest model that passes each request, billed at $0.005 per request. When you run open-source models like DeepSeek, where you serve them matters as much as which agent calls them.

Most serverless providers quantize activations to fp8 to cut cost, which degrades output quality. Morph serves open-source coding models with 16-bit (bf16) activations, no fp8 or int8 quantization, so output matches the reference weights. morph-dsv4flash (DeepSeek V4 Flash) costs $0.141953125 per 1M input tokens and $0.399625 per 1M output tokens. See pricing for full rates.

$0.141953125
morph-dsv4flash input, per 1M tokens
$0.399625
morph-dsv4flash output, per 1M tokens
16-bit
bf16 activations, no fp8 quantization

On the search side, every agent spends tokens building context before it writes code. Cognition measured coding agents spending 60% of their time on search. WarpGrep runs as an MCP server inside Codex, Claude Code, Cursor, or any MCP-compatible agent, executing 8 parallel searches per turn across 4 turns in under 6 seconds. It costs $0.80 per 100K tokens. Fast Apply merges generated diffs into your codebase at 10,500 tokens per second.

$0.80
WarpGrep per 100K tokens
10,500
tok/sec Fast Apply merge speed
6 sec
8 parallel searches across 4 turns

Better Search and a Faster Model Layer for Any Agent

WarpGrep works as an MCP server inside Codex, Claude Code, Cursor, and any MCP-compatible agent. 8 parallel tool calls per turn, 4 turns, sub-6 seconds. $0.80 per 100K tokens.

Frequently Asked Questions

What is the best AI coding agent in October 2026?

Claude Code, with Codex the cheaper option. On the Terminal-Bench 4.0 agent leaderboard (tbench.ai, 330 trials per entry), Claude Code + Opus 5.5 at max effort scores 64.8% and Claude Code + Sonnet 5.5 scores 61.8%. Codex + GPT-6 Astra and Codex + GPT-6.1 Sol both score 58.2%. Per task, the Opus 5.5 run cost $14.30 and the Sonnet 5.5 run $22.22. Astra cost $9.90 per task and GPT-6.1 Sol $1.92. For IDE-native work, pick Cursor or Copilot. For free and open source, pick OpenCode, Cline, Goose, or Kilo Code.

What is the best coding agent benchmark?

Terminal-Bench 4.0 is the current agent-level standard. It scores the agent plus model pair on 66 hard terminal tasks across software, ML, science, ops, security, hardware, and media, with 330 trials per entry and published cost. Terminal-Bench 2.1 is saturated: the top model-level score is 91.4% (Fable 5.1, Artificial Analysis), and the top agent-level entry on the Harbor Hub 2.1 board is Codex + GPT-6 Astra at 87.4%. The Artificial Analysis Coding Agent Index v1.5 averages DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA per agent; Claude Code + Fable 5.1 leads it at 62.2. SWE-bench Pro measures harder repository issue fixes. SWE-bench Verified tops out at 97.0% (Claude Opus 5 on Vals.ai), and Vals stopped running it on new models on September 1 because it no longer separates frontier models.

What is the Terminal-Bench 4.0 coding agent leaderboard?

As of October 6, 2026 on tbench.ai: Claude Code + Opus 5.5 (max) 64.8%; Claude Code + Sonnet 5.5 (max) 61.8%; Codex + GPT-6 Astra (max) 58.2%; Codex + GPT-6.1 Sol (max) 58.2%; Claude Code + Fable 5.1 (max) 57.9%; Codex + GPT-6 Astra (xhigh and high) 57.9%; Claude Code + Fable 5.1 (high) 54.5%; Claude Code + Opus 5 (xhigh) 53.9%; Codex + GPT-6 Sol (max) 49.4%; Claude Code + GLM-5.3 (max) 41.8%; Grok Build + Grok 4.7 (xhigh) 37.6%; Codex + GPT-5.6 Sol (max) 37.3%. Each entry is 330 trials.

How do AI coding agents compare on Terminal-Bench 2.1?

There are two 2.1 boards, and they measure different things. Model-level scores from Artificial Analysis (one fixed Terminus 2 harness): Claude Fable 5.1 max 91.4%, Fable 5.1 xhigh 91.0%, GPT-6 Astra high 89.9%, GPT-6 Astra medium 89.5%, GPT-5.6 Sol xhigh 89.5%, Claude Opus 5 max 89.1%, GPT-6 Astra max 88.4%, GPT-5.6 Sol max 88.0%, Gemini 3.8 Flash high 87.6%, Kimi K3 max 85.0%. Agent-level scores (agent plus model) from the Harbor Hub Terminal-Bench 2.1 board: Codex + GPT-6 Astra (high) 87.4%, Claude Code + Fable 5 (xhigh) 83.8%, Codex + GPT-5.5 (xhigh) 83.2%, Terminus 2 + Fable 5 80.5%, Cursor CLI + Grok 4.5 79.3%, Claude Code + Opus 4.8 78.9%. With nine models between 87% and 92%, 2.1 no longer separates the frontier, which is why tbench.ai moved its main board to 4.0.

What are the best open source AI coding agents in 2026?

By GitHub stars on September 22, 2026: OpenCode (209,405, MIT), OpenAI Codex CLI (125,962, Apache-2.0), Gemini CLI (107,130, Apache-2.0), Cline (69,074, Apache-2.0), Goose (54,568, Apache-2.0), Aider (49,118, Apache-2.0), and Kilo Code (27,390, MIT). Claude Code's repo has 147,649 stars but the tool is proprietary. All of these are free to install and run on your own API key or local models.

Which AI coding agent has the most GitHub stars in 2026?

OpenCode (anomalyco/opencode) at 209,405 stars. anthropics/claude-code has 147,649 but is proprietary. Codex has 125,962, Gemini CLI 107,130, Cline 69,074, Goose 54,568, Aider 49,118, and Kilo Code 27,390. Aider's last commit to main is still 2026-05-22; every other repo on this list shipped a release in September. Roo Code is gone: its extension shut down on May 15, 2026 and the repo is archived.

Which AI coding tools are free in 2026?

Free to install: OpenCode, Cline, Goose, Aider, Kilo Code, and Gemini CLI (60 requests/min and 1,000 requests/day on a personal Google account). Free paid-product tiers: GitHub Copilot Free (2,000 completions/mo), Cursor Hobby (no card), Kiro Free (50 credits/mo with Claude Sonnet 4.5 and open-weight models), Google Antigravity Individual ($0), Amp Free (since September 13, no Amp fees when you bring your own subscription or API key), and ChatGPT Free, which gets GPT-6 Luna in the Codex desktop app. Open-source agents are free to run but you pay per token for the model unless you run a local one.

What changed for AI coding agents since September 22, 2026?

Anthropic released Claude Sonnet 5.5 on September 28 ($2 / $10 per 1M tokens, 1M context), and Claude Code v2.1.284 made it the default Sonnet on the Anthropic API. OpenAI released GPT-6.1 Sol on September 29 at $2 / $10 with a 1.05M-token context window. Its Codex rollout covers Plus, Pro, Business, Enterprise, and Edu, but not Free or Go. Google announced Gemini 4 Argon on September 30, rolling out first to trusted cyber defenders, at an introductory $2 / $10 per 1M tokens and $4 / $20 after. On tbench.ai, Claude Code + Opus 5.5 took first place on Terminal-Bench 4.0 at 64.8%, and Claude Code + Sonnet 5.5 entered at 61.8%.

What changed for AI coding agents in September 2026?

Anthropic released Claude Fable 5.1 on September 1 ($10 / $50 per 1M tokens, generally available) and Claude Opus 5.5 on September 22 ($4 / $20, 1M context), which Claude Code v2.1.280 made the default Opus. Claude Code now defaults every subscription plan, Pro included, to Opus. OpenAI released GPT-6 Astra on September 3 ($10 / $50) and GPT-6 Sol ($2 / $10) and Luna ($0.10 / $0.50) on September 22; OpenAI names GPT-6 Sol the replacement for GPT-5.5 on paid Codex plans. Google shipped Gemini 3.8 Flash on September 2. Meta shipped Muse Spark 1.3 to Muse Code on September 2. Cognition shipped SWE-2 in Devin Desktop (formerly Windsurf) and Devin CLI on September 10. Cursor launched Projects on September 10 and Grok 4.7 on September 21. Amp went free with your own compute and keys on September 13. Copilot added Opus 5.5, GPT-6 Sol and Luna, and Grok 4.7.

What changed for AI coding agents in July and August 2026?

July: Fable 5 was restored on July 1, GPT-5.6 went GA on July 9 in Sol, Terra, and Luna tiers, Claude Opus 5 shipped July 24 as the Claude Code default Opus, and Anaconda announced it had acquired Kilo Code on July 15. August: Meta launched Muse Code, a terminal coding agent on its Muse Spark model, in beta on August 5. SpaceX completed its acquisition of Cursor on August 14, and Cursor added Origin code hosting (August 17) and always-on cloud agents (August 19). Terminal-Bench 4.0 launched on August 28.

What is the Artificial Analysis Coding Agent Index?

It scores agent plus model pairs on three benchmarks weighted equally: DeepSWE v1.1 (113 tasks), Terminal-Bench 4.0 (66 tasks), and SWE-Atlas-QnA (124 tasks), each averaged over 3 attempts. Index v1.5 as of September 22, 2026: Claude Code + Fable 5.1 (max) 62.2 at $12.39 per task; Devin Fusion CLI + Fable 5.1 and SWE-2 61.7 at $7.90; Codex + GPT-6 Astra (max) 61.6 at $7.47; Claude Code + Opus 5 (max) 59.7; Codex + GPT-6 Sol (max) 56.7 at $2.99; Grok Build + Grok 4.7 56.3; Muse Code + Muse Spark 1.3 54.3; OpenCode + GLM-5.3 (max) 53.6 at $4.24, the only open-source harness on the index.

Should I use multiple AI coding agents?

Most developers run two or three. A common stack: Codex or Claude Code for heavy agent work, Copilot or Cursor for inline completions, and one free open-source agent (OpenCode, Cline, or Kilo Code) for model flexibility. Because the agents interoperate over MCP and ACP, the model layer underneath them drives most of the cost and quality.

Sources