Back to AI Coding

AI Model Ranking

AI Coding Model Rankings

Snapshot: 2026-08-13 · Last reviewed: 2026-08-13

Methodology

Metrics follow a snapshot of the Artificial Analysis public LLM leaderboard: Intelligence Index, cost per task, and speed reflect benchmark performance at the snapshot date. Provider pricing, availability, and rankings change frequently, and a high benchmark score does not guarantee the best result for a specific codebase or workflow.

Source
Artificial Analysis
Last reviewed
2026-08-13

Top 50 models

01

Current reasoning model

Anthropic - Claude Opus - 5 (max)

Claude Opus 5 max leads this snapshot with a top Intelligence Index of 63 and 1M context.

Best for: Frontier repository analysis, complex planning, and high-value agent tasks that need top reasoning.

Anthropic

Context
1M
AA Index
63
Cost per task
$2.34
Speed
52 tok/s
First chunk
60.92s
Total response
70.52s
02

Current reasoning model

Anthropic - Claude Opus - 5 (xhigh)

Ties the top Intelligence Index with lower cost per task and much faster first-chunk latency than max.

Best for: Deep multi-file edits and coding agents that still need frontier quality with better latency.

Anthropic

Context
1M
AA Index
63
Cost per task
$1.80
Speed
52 tok/s
First chunk
25.75s
Total response
35.32s
03

Current reasoning model

Anthropic - Claude Fable - 5

Still among the highest Intelligence Index scores, with a 1M context window and strong output speed.

Best for: Deep repository analysis, complex planning, and high-value agent tasks where latency is acceptable.

Anthropic

Context
1M
AA Index
62
Cost per task
$3.14
Speed
63 tok/s
First chunk
113.68s
Total response
121.60s
04

Current reasoning model

Anthropic - Claude Opus - 5 (high)

High-reasoning Opus 5 mode with much lower latency than max/xhigh while staying near the top.

Best for: Daily Claude agent coding, code review, and tool-heavy workflows that still need strong reasoning.

Anthropic

Context
1M
AA Index
61
Cost per task
$1.23
Speed
51 tok/s
First chunk
13.43s
Total response
23.32s
05

Current reasoning model

OpenAI - GPT Sol - 5.6 (max)

Top OpenAI reasoning tier in this snapshot, with 1M context and strong Intelligence Index.

Best for: OpenAI-centered coding agents, long-context planning, and max-reasoning software tasks.

OpenAI

Context
1M
AA Index
61
Cost per task
$1.23
Speed
62 tok/s
First chunk
155.41s
Total response
163.54s
06

Current reasoning model

SpaceXAI - Grok - 4.6 (high)

New Grok 4.6 high entry near the top, with 500k context and lower cost per task than the frontier Claude/OpenAI max tiers.

Best for: xAI-stack coding agents, medium-context planning, and cost-aware frontier experiments.

SpaceXAI

Context
500k
AA Index
61
Cost per task
$0.84
Speed
66 tok/s
First chunk
32.30s
Total response
39.89s
07

Current reasoning model

Kimi - Kimi - K3 (max)

Moonshot's K3 max flagship with 1M+ context and strong coding scores; total response is still relatively long.

Best for: Long-horizon coding agents, frontend generation, regional API stacks, and cost-aware 1M context work.

Kimi

Context
1.05M
AA Index
60
Cost per task
$0.84
Speed
41 tok/s
First chunk
3.27s
Total response
64.72s
08

Current reasoning model

OpenAI - GPT Sol - 5.6 (xhigh)

High-reasoning GPT-5.6 Sol mode with 1M context and lower cost per task than max.

Best for: Daily agent coding, code review, and tool-heavy workflows that still need strong reasoning.

OpenAI

Context
1M
AA Index
59
Cost per task
$0.81
Speed
61 tok/s
First chunk
57.84s
Total response
66.08s
09

Current reasoning model

Anthropic - Claude Opus - 5 (medium)

Balanced Opus 5 medium tier with 1M context and much faster total response than higher tiers.

Best for: Interactive Claude coding sessions, refactors, and product engineering loops.

Anthropic

Context
1M
AA Index
59
Cost per task
$0.72
Speed
52 tok/s
First chunk
8.59s
Total response
18.29s
10

Current reasoning model

Alibaba - Qwen - 3.8 Max

New Qwen 3.8 Max flagship with 1M context, fast first-chunk latency, and a top-10 Intelligence Index.

Best for: Multilingual coding, regional API stacks, and Qwen-centered agent experiments.

Alibaba

Context
1M
AA Index
58
Cost per task
$1.13
Speed
47 tok/s
First chunk
2.72s
Total response
55.90s
11

Current reasoning model

OpenAI - GPT Sol - 5.6 (high)

Balanced GPT-5.6 Sol high tier with 1M context and responsive total response time.

Best for: Interactive coding sessions, refactors, and product engineering loops inside OpenAI tools.

OpenAI

Context
1M
AA Index
57
Cost per task
$0.55
Speed
56 tok/s
First chunk
19.28s
Total response
28.15s
12

Current model; partial public metrics

Meta - Muse Spark - 1.2 (xhigh)

Muse Spark 1.2 xhigh lands in the top 12 with 1M+ context; public speed and latency are still incomplete.

Best for: Long-context Meta model routing comparisons when Intelligence Index matters more than published latency.

Meta

Context
1.05M
AA Index
57
Cost per task
$0.40
Speed
tok/s
First chunk
Total response
13

Current reasoning model

OpenAI - GPT Terra - 5.6 (max)

GPT-5.6 Terra max tier with lower cost per task than Sol max and very fast output speed.

Best for: Cost-aware long-context coding where throughput matters more than first-token speed.

OpenAI

Context
1M
AA Index
57
Cost per task
$0.51
Speed
116 tok/s
First chunk
194.43s
Total response
198.74s
14

Current reasoning model

SpaceXAI - Grok - 4.5 (high)

Grok 4.5 high with fast latency, mid-size 500k context, and lower cost per task than Grok 4.6 high.

Best for: Fast interactive coding help, medium-context agent loops, and xAI ecosystem experiments.

SpaceXAI

Context
500k
AA Index
56
Cost per task
$0.36
Speed
57 tok/s
First chunk
9.32s
Total response
18.11s
15

Current reasoning model

OpenAI - GPT Sol - 5.6 (medium)

GPT-5.6 Sol medium tier balances intelligence, 1M context, and low total response time.

Best for: Most daily OpenAI coding loops where speed and quality both matter.

OpenAI

Context
1M
AA Index
56
Cost per task
$0.37
Speed
57 tok/s
First chunk
6.82s
Total response
15.62s
16

Current reasoning model

Anthropic - Claude Sonnet - 5 (max)

Sonnet-level profile with high intelligence and 1M context, trading off long first response.

Best for: Long, careful batch work where quality and context matter more than immediacy.

Anthropic

Context
1M
AA Index
55
Cost per task
$1.72
Speed
71 tok/s
First chunk
191.38s
Total response
198.40s
17

Current reasoning model

DeepSeek - DeepSeek V4 Pro - 0813 (max)

New DeepSeek V4 Pro 0813 max lands in the top 20 with 1M context, very low cost per task, and fast first-chunk latency.

Best for: Budget open-ecosystem coding agents, long-context analysis, and latency-sensitive DeepSeek Pro work.

DeepSeek

Context
1M
AA Index
53
Cost per task
$0.06
Speed
83 tok/s
First chunk
1.63s
Total response
31.68s
18

Current reasoning model

OpenAI - GPT Terra - 5.6 (xhigh)

GPT-5.6 Terra xhigh combines lower cost per task, 1M context, and fast median output.

Best for: Cost-sensitive long-context coding with strong throughput and reasonable first-chunk latency.

OpenAI

Context
1M
AA Index
53
Cost per task
$0.31
Speed
106 tok/s
First chunk
29.77s
Total response
34.49s
19

Current reasoning model

Z AI - GLM - 5.2 (max)

High ranking with 1M context, low cost per task, and very fast first-chunk latency.

Best for: Cost-aware coding assistants, long-context analysis, and latency-sensitive workflows.

Z AI

Context
1M
AA Index
53
Cost per task
$0.32
Speed
118 tok/s
First chunk
1.42s
Total response
22.58s
20

Current reasoning model

Anthropic - Claude Opus - 5 (low)

Low-latency Opus 5 tier with 1M context and the fastest total response in the Opus 5 line.

Best for: Interactive pair programming, quick fixes, and fast planning loops on Claude Opus 5.

Anthropic

Context
1M
AA Index
52
Cost per task
$0.43
Speed
51 tok/s
First chunk
3.04s
Total response
12.90s
21

Current reasoning model

OpenAI - GPT Luna - 5.6 (max)

GPT-5.6 Luna max tier with very low cost per task, 1M context, and high output speed.

Best for: Budget OpenAI stacks, long-context batch edits, and high-throughput coding assistants.

OpenAI

Context
1M
AA Index
52
Cost per task
$0.05
Speed
150 tok/s
First chunk
136.64s
Total response
139.97s
22

Current reasoning model

DeepSeek - DeepSeek V4 Flash - 0731 (max)

DeepSeek V4 Flash 0731 max with 1M context, very low cost per task, and fast first-chunk latency.

Best for: Budget open-ecosystem coding agents, long-context analysis, and latency-sensitive edits.

DeepSeek

Context
1M
AA Index
52
Cost per task
$0.03
Speed
122 tok/s
First chunk
1.44s
Total response
21.90s
23

Current general model

Google - Gemini - 3.6 Flash

Gemini Flash tier with 1M context, high output speed, and Google ecosystem fit.

Best for: Fast Google-stack coding loops, long-file skim, and high-throughput assistant work.

Google

Context
1M
AA Index
52
Cost per task
$0.56
Speed
229 tok/s
First chunk
18.99s
Total response
21.17s
24

Current reasoning model

OpenAI - GPT Sol - 5.6 (low)

Low-latency GPT-5.6 Sol tier with 1M context and the fastest total response in the Sol line.

Best for: Interactive pair programming, quick fixes, and fast planning loops on GPT-5.6 Sol.

OpenAI

Context
1M
AA Index
51
Cost per task
$0.23
Speed
52 tok/s
First chunk
3.07s
Total response
12.74s
25

Current reasoning model

OpenAI - GPT Terra - 5.6 (high)

Responsive GPT-5.6 Terra high tier with 1M context and very short total response time.

Best for: Interactive OpenAI coding where throughput and low latency both matter.

OpenAI

Context
1M
AA Index
50
Cost per task
$0.22
Speed
97 tok/s
First chunk
3.79s
Total response
8.94s
26

Current reasoning model

OpenAI - GPT Luna - 5.6 (xhigh)

Budget GPT-5.6 Luna xhigh tier with 1M context, very low cost per task, and fast median output.

Best for: Cost-sensitive OpenAI coding assistants and long-context batch edits.

OpenAI

Context
1M
AA Index
50
Cost per task
$0.03
Speed
147 tok/s
First chunk
50.93s
Total response
54.32s
27

Current reasoning model

Kimi - Kimi - K3 (low)

Kimi K3 low tier keeps 1M+ context and lowers cost per task versus K3 max.

Best for: Cost-aware long-context Kimi coding when max reasoning effort is unnecessary.

Kimi

Context
1.05M
AA Index
48
Cost per task
$0.24
Speed
37 tok/s
First chunk
3.24s
Total response
70.53s
28

Current general model

Google - Gemini - 3.1 Pro Preview

Gemini 3.1 Pro preview with 1M context and a balanced cost/speed profile for Google stacks.

Best for: Google Cloud coding trials, multimodal prototypes, and long-context preview evaluation.

Google

Context
1M
AA Index
48
Cost per task
$0.33
Speed
114 tok/s
First chunk
31.61s
Total response
35.99s
29

Current model; partial public metrics

Motif Technologies - Motif - 3

Motif 3 lands on Intelligence Index with mid-size context; most public speed and price metrics are still incomplete.

Best for: Early evaluation of Motif 3 when benchmark rank matters more than mature public metrics.

Motif Technologies

Context
262k
AA Index
47
Cost per task
Speed
tok/s
First chunk
Total response
30

Current reasoning model

OpenAI - GPT Luna - 5.6 (high)

Very low cost-per-task GPT-5.6 Luna high tier with 1M context and fast responses.

Best for: High-volume OpenAI coding assistants and budget long-context workflows.

OpenAI

Context
1M
AA Index
47
Cost per task
$0.02
Speed
150 tok/s
First chunk
15.32s
Total response
18.66s
31

Current reasoning model

OpenAI - GPT Terra - 5.6 (medium)

Low-latency GPT-5.6 Terra medium tier with 1M context and very short end-to-end response.

Best for: Snappy interactive coding, refactors, and tool-heavy OpenAI agent loops.

OpenAI

Context
1M
AA Index
47
Cost per task
$0.12
Speed
97 tok/s
First chunk
1.78s
Total response
6.96s
32

Current model; partial public metrics

Google - Gemini - 3.5 Flash (medium)

Gemini 3.5 Flash medium mode with 1M context and fast output; cost-per-task is not fully public in this snapshot.

Best for: Google ecosystem experiments where Flash-medium speed matters more than published task cost.

Google

Context
1M
AA Index
47
Cost per task
Speed
162 tok/s
First chunk
17.53s
Total response
20.60s
33

Current model; partial public metrics

OpenAI - GPT Codex - 5.3 (xhigh)

GPT-5.3 Codex xhigh coding-focused tier; cost-per-task is not fully public in this snapshot.

Best for: Codex-centered software tasks, repo edits, and OpenAI coding agent workflows.

OpenAI

Context
400k
AA Index
46
Cost per task
Speed
106 tok/s
First chunk
73.50s
Total response
78.22s
34

Current reasoning model

MiniMax - MiniMax - M3

MiniMax M3 with 1M context, low cost per task, and low first-chunk latency.

Best for: Cost-aware coding assistants, regional stacks, and low-latency interactive edits.

MiniMax

Context
1M
AA Index
45
Cost per task
$0.14
Speed
83 tok/s
First chunk
1.65s
Total response
31.83s
35

Current model; partial public metrics

Motif Technologies - Motif - 3 (Beta)

Motif 3 Beta lands in the ranking on Intelligence Index; most public speed and price metrics are still incomplete.

Best for: Early evaluation of emerging model providers when benchmark rank matters more than mature metrics.

Motif Technologies

Context
262k
AA Index
45
Cost per task
Speed
tok/s
First chunk
Total response
36

Current reasoning model

DeepSeek - DeepSeek V4 Pro - max

Earlier DeepSeek V4 Pro max (pre-0813) with 1M context and very low cost per task; total response is longer than 0813.

Best for: Budget coding agents, self-hosted evaluation, and cost-sensitive long-context work.

DeepSeek

Context
1M
AA Index
45
Cost per task
$0.05
Speed
67 tok/s
First chunk
1.68s
Total response
73.96s
37

Current reasoning model

DeepSeek - DeepSeek V4 Pro - high

DeepSeek V4 Pro high tier with 1M context, low cost per task, and lower latency than max.

Best for: Cost-aware interactive coding and long-context analysis on DeepSeek endpoints.

DeepSeek

Context
1M
AA Index
44
Cost per task
$0.04
Speed
63 tok/s
First chunk
1.74s
Total response
41.34s
38

Current reasoning model

Kimi - Kimi - K2.7 Code

Kimi K2.7 Code coding-oriented entry with 256k context and published cost per task.

Best for: Code-focused Kimi workflows when you do not need the full 1M K3 context window.

Kimi

Context
256k
AA Index
43
Cost per task
$0.22
Speed
38 tok/s
First chunk
2.83s
Total response
74.36s
39

Current reasoning model

Xiaomi - MiMo - V2.5-Pro

Xiaomi MiMo V2.5-Pro with 1M context and very low cost per task for regional stacks.

Best for: Budget multilingual coding, regional APIs, and low-cost long-context assistants.

Xiaomi

Context
1M
AA Index
43
Cost per task
$0.03
Speed
46 tok/s
First chunk
3.32s
Total response
57.11s
40

Current general model

Anthropic - Claude Sonnet - 5 (Non-reasoning)

Non-reasoning Claude Sonnet 5 with 1M context, low latency, and a lighter cost profile than max.

Best for: Everyday Claude coding, quick edits, and interactive sessions that do not need deep reasoning.

Anthropic

Context
1M
AA Index
43
Cost per task
$0.42
Speed
60 tok/s
First chunk
1.92s
Total response
10.21s
41

Current reasoning model

Thinking Machines - Inkling - base

Inkling from Thinking Machines with 1M context, published cost per task, and solid first-chunk latency.

Best for: Long-context evaluation of newer model labs with published pricing and latency.

Thinking Machines

Context
1M
AA Index
42
Cost per task
$0.34
Speed
74 tok/s
First chunk
1.86s
Total response
35.45s
42

Current reasoning model

Tencent - Hy3 - base

Tencent Hy3 with mid-size context, very low cost per task, and competitive first-chunk latency.

Best for: Regional coding stacks, budget assistants, and Tencent-ecosystem experiments.

Tencent

Context
256k
AA Index
42
Cost per task
$0.04
Speed
63 tok/s
First chunk
2.88s
Total response
42.48s
43

Current model; partial public metrics

Nex AGI - Nex - N2-Pro

Nex N2-Pro enters with mid-size context and fast median output; cost-per-task is not fully public here.

Best for: Emerging-provider evaluation and low-latency coding experiments outside the major labs.

Nex AGI

Context
262k
AA Index
42
Cost per task
Speed
137 tok/s
First chunk
1.68s
Total response
19.92s
44

Current general model

OpenAI - GPT Sol - 5.6 (Non-reasoning)

Non-reasoning GPT-5.6 Sol with 1M context and the lowest first-chunk latency in the Sol line.

Best for: Snappy OpenAI coding loops when extended reasoning is unnecessary.

OpenAI

Context
1M
AA Index
42
Cost per task
$0.24
Speed
59 tok/s
First chunk
1.33s
Total response
9.86s
45

Current model; partial public metrics

Upstage - Solar Pro - 4

Upstage Solar Pro 4 with 512k context and published speed; cost-per-task is not fully public in this snapshot.

Best for: Regional and document-heavy coding stacks evaluating Upstage models on Intelligence Index.

Upstage

Context
512k
AA Index
42
Cost per task
Speed
65 tok/s
First chunk
2.18s
Total response
40.51s
46

Current reasoning model

OpenAI - GPT Terra - 5.6 (low)

Low-latency GPT-5.6 Terra low tier with 1M context and very short end-to-end response.

Best for: Fast interactive OpenAI coding and quick refactors where max reasoning is unnecessary.

OpenAI

Context
1M
AA Index
41
Cost per task
$0.09
Speed
91 tok/s
First chunk
1.73s
Total response
7.20s
47

Current reasoning model

Thinking Machines - Inkling Small - base

Smaller Inkling variant with 1M context, low cost per task, and faster median output than base Inkling.

Best for: Budget evaluation of Thinking Machines models with tighter latency and cost targets.

Thinking Machines

Context
1M
AA Index
41
Cost per task
$0.07
Speed
128 tok/s
First chunk
1.58s
Total response
21.06s
48

Current model; partial public metrics

China Mobile - JT - 4.1 Flash 236B A21B

China Mobile JT-4.1 Flash lands on Intelligence Index; most public price and speed metrics remain incomplete.

Best for: Regional China Mobile stack evaluation when benchmark rank is the primary signal.

China Mobile

Context
256k
AA Index
40
Cost per task
Speed
tok/s
First chunk
Total response
49

Current model; partial public metrics

Sapiens AI - Agnes - 2.5 Pro Alpha

Agnes 2.5 Pro Alpha with 1M context and fast median output; cost-per-task is not fully public here.

Best for: Early long-context evaluation of emerging Sapiens AI models.

Sapiens AI

Context
1M
AA Index
40
Cost per task
Speed
131 tok/s
First chunk
2.37s
Total response
21.42s
50

Current reasoning model

Alibaba - Qwen - 3.7 Plus

Qwen 3.7 Plus with 1M context, lower cost per task than Qwen 3.8 Max, and moderate total response time.

Best for: Cost-aware Qwen coding agents and multilingual product engineering loops.

Alibaba

Context
1M
AA Index
39
Cost per task
$0.24
Speed
56 tok/s
First chunk
2.13s
Total response
46.67s