Open frontier models, one API.
Benchmarks
Independent intelligence and cost comparison — open frontier models against the closed frontier.
Intelligence Index
AA Intelligence Index v4.1.1 · higher is betterKimi K3
59.7
DeepSeek V4 Pro
53.0
DeepSeek V4 Flash
51.8
Claude Opus 5
63.1
Claude Fable 5
62.1
GPT-5.6 Sol
60.9
Cost per task
weighted USD per Intelligence Index task · lower is betterDeepSeek V4 Flash
$0.027
DeepSeek V4 Pro
$0.056
Kimi K3
$0.84
GPT-5.6 Sol
$1.23
Claude Opus 5
$2.34
Claude Fable 5
$3.14
Agentic tool use
τ³-Banking · higher is betterKimi K3
46%
DeepSeek V4 Pro
40%
DeepSeek V4 Flash
39%
GPT-5.6 Sol
44%
Claude Opus 5
42%
Claude Fable 5
38%
Coding
Terminal-Bench v2.1 · higher is betterKimi K3
85%
DeepSeek V4 Pro
79%
DeepSeek V4 Flash
79%
GPT-5.6 Sol
88%
Claude Opus 5
89%
Claude Fable 5
85%
Long-context reasoning
AA-LCR · higher is betterKimi K3
83%
DeepSeek V4 Pro
75%
DeepSeek V4 Flash
74%
GPT-5.6 Sol
78%
Claude Opus 5
76%
Claude Fable 5
77%
Source: Artificial Analysis (independent evaluations), as of 2026-08-09.
Available models
Production-ready today, with new frontier releases added as they ship.
| Model | Context | Capabilities | Latency (TTFT) | Output speed | Success rate |
|---|---|---|---|---|---|
| Kimi K3 | 1M | tools reasoning structured output vision | 3.0 s | 38 tok/s | 100.0% |
| DeepSeek V4 Pro 0813 | 1M | tools reasoning structured output | 0.5 s | 110 tok/s | 100.0% |
| DeepSeek V4 Flash 0731 | 1M | tools reasoning structured output | 0.5 s | 115 tok/s | 100.0% |
Click a row for capabilities and quickstart. Latency, output speed and success rate are medians of 402 streamed completions measured on the Corion API (as of 2026-09-19), so they reflect the routing a customer actually gets. Latency is time to first token and includes the model's reasoning phase.
Which one should I pick?
- •Kimi K3 — the highest-intelligence option: deep reasoning, agentic workflows, long-horizon coding and long-document work. Typical fits: autonomous agents and multi-step tool use, codebase-scale refactoring, research and analysis over large document sets, knowledge work where accuracy matters more than latency.
- •DeepSeek V4 Pro 0813 — the value pick: near-frontier reasoning at roughly a fifteenth of K3's cost per task, with the same 1M context and 375K max output. Typical fits: reasoning-heavy analysis and long-horizon agents that run too hot on a frontier rate card, full-codebase review, multi-step automation, and large-scale synthesis. Text only — send images to Kimi K3.
- •DeepSeek V4 Flash — the cost- and latency-optimized option with the same 1M context. Typical fits: high-volume chat and support bots, classification and extraction pipelines, batch summarization, real-time assistants, and any workload where unit cost dominates.
