Skip to content

Best Models for Coding (2026)

Which model should you point your coding agent at? This page tracks the current frontier, what each model costs, and what it is actually good for.

Context windows and pricing come from a live production model catalogue, so they are what you pay rather than what a launch post claimed. Release dates come from vendor announcements.

Last updated: August 14, 2026


Use CaseBest ModelRunner-Up
Complex agentic codingClaude Opus 5GPT-5.6 Sol
Everyday developmentClaude Sonnet 5GPT-5.6 Terra
Coding on a Flash budgetGemini 3.7 FlashMiniMax M3
Hardest professional workGPT-5.6 SolClaude Fable 5
Long-horizon codingGLM-5.2Kimi K3
High-volume, cost-sensitiveGPT-5.6 LunaDeepSeek V4 Flash
Browser-use agentsMiniMax M3Gemini 3.7 Flash
Best value at the frontierDeepSeek V4 ProGLM-5.2

Newest first. Prices are per 1M tokens, input / output.

ModelProviderReleasedContextPricePositioning
Gemini 3.7 FlashGoogle2026-08-131M$1.50 / $7.50Most capable Flash for coding and agentic workflows
Claude Opus 5Anthropic2026-07-241M$5 / $25Most capable Anthropic, complex agentic coding
GPT-5.6 SolOpenAI2026-07-09922K$5 / $30Frontier, complex professional work
GPT-5.6 TerraOpenAI2026-07-09922K$2.50 / $15Balances intelligence and cost
GPT-5.6 LunaOpenAI2026-07-09922K$1 / $6Cost-sensitive high-volume workloads
Claude Sonnet 5Anthropic2026-06-301M$3 / $15Most capable Sonnet, coding and agents

Anthropic versions each tier separately, so the 5 generation arrived in pieces rather than at one launch. OpenAI shipped GPT-5.6 as three models at once — Luna, Terra, and Sol, least to most capable.


ModelContextPriceNotes
Claude Opus 51M$5 / $25Most capable Anthropic for complex agentic coding
Claude Fable 51M$10 / $50Next-gen flagship, knowledge work and coding
Claude Sonnet 51M$3 / $15The default choice for day-to-day work
Claude Sonnet 4.61M$3 / $15Proven all-rounder
Claude Opus 4.61M$5 / $25Deep reasoning and production code
ModelContextPriceNotes
GPT-5.6 Sol922K$5 / $30The hardest work
GPT-5.6 Terra922K$2.50 / $15Balanced everyday work
GPT-5.6 Luna922K$1 / $6Fast and cheap at volume
ModelContextPriceNotes
Gemini 3.7 Flash1M$1.50 / $7.50Coding and agentic workflows
Gemini 3.1 Pro1M$2 / $12Best multimodal understanding
Gemini 3.6 Flash1M$1.50 / $7.50Fast multimodal, large output budget
Gemini 3 Flash1M$0.50 / $3Fast everyday multimodal
ModelProviderContextPriceNotes
Kimi K3Moonshot917K$3 / $15Frontier coding, deep reasoning
Kimi K2.7 CodeMoonshot214K$0.95 / $4Coding and agentic, multimodal
GLM-5.2Z.ai1M$1.40 / $4.40Flagship, long-horizon coding
GLM-5.1Z.ai200K$1.40 / $4.40Previous-gen flagship
DeepSeek V4 ProDeepSeek616K$0.44 / $0.87Frontier reasoning, remarkable value
DeepSeek V4 FlashDeepSeek616K$0.14 / $0.28Fast reasoning, cheapest on this page
Grok 4.6xAI500K$2 / $6Strong agentic coding
MiniMax M3MiniMax1M$0.30 / $1.20Good with browser use
MiniMax M2.7MiniMax200K$0.30 / $1.20Lightweight agentic coding
Muse Spark 1.2Meta1M$1.25 / $4.25Fast agentic coding
Qwen3.8 MaxQwen991K$2 / $6Multimodal long-context
Qwen3.7 PlusQwen991K$0.80 / $3.20Balanced multimodal

This page used to carry SWE-bench Verified percentages. It no longer does, and the reason is worth stating.

The scores that were here were real, but they described Opus 4.6, GPT-5.2-Codex, and Gemini 3 Flash — models that are now one or two generations behind. None of the models at the top of this page has a verified SWE-bench number we can cite, and publishing an unverified figure next to a verified one makes both worthless.

A stale number is worse than no number, because it looks current. We would rather tell you what a model costs, how much context it holds, and what its maker built it for — all of which we can check — than fill a column with figures we cannot stand behind.

When verified scores for the current frontier are published, they will go back in with sources attached.

BenchmarkWhat It MeasuresWhy It Matters
SWE-bench VerifiedFix real GitHub issues end-to-endThe most realistic agentic coding measure
SWE-bench ProA harder subset of real-world issuesTests frontier agent capability
Terminal-Bench 2.0Agentic terminal tasksTests real development workflows
Aider PolyglotMulti-language code editing accuracyTests edit-apply workflows
LiveCodeBench ProCompetitive programmingTests algorithmic reasoning
HumanEvalFunction-level code generationClassic, and now largely saturated

For agentic work, SWE-bench Verified and Terminal-Bench 2.0 are the two worth watching. HumanEval no longer separates frontier models from each other.


Hardest agentic work, cost secondary?
→ Claude Opus 5 or GPT-5.6 Sol
Everyday development, best balance?
→ Claude Sonnet 5 ($3/$15, 1M context)
Cheap but genuinely capable?
→ DeepSeek V4 Pro ($0.44/$0.87) or GLM-5.2 ($1.40/$4.40)
Enormous codebase in one session?
→ Anything at 1M — Sonnet 5, Opus 5, Gemini 3.7 Flash, GLM-5.2
Running an agent at high volume?
→ GPT-5.6 Luna or DeepSeek V4 Flash
Driving a browser?
→ MiniMax M3

Note how little separates the tiers on context now. A 1M window is close to standard at the frontier, so the deciding factors are price and how well a model holds a long task together.


A better model does not know how your team deploys, reviews, or releases. That knowledge lives in your repo and in the skills you give the agent — see @skills for how procedures reach an agent without occupying its context on every request.

Pairing a mid-tier model with the right skill often beats a frontier model working from nothing.


Built with AdaL CLI