Corion


Open frontier models, one API.

Benchmarks

Independent intelligence and cost comparison — open frontier models against the closed frontier.

Intelligence Index

AA Intelligence Index v4.1.1 · higher is better
Kimi K3
59.7
DeepSeek V4 Pro
53.0
DeepSeek V4 Flash
51.8
Claude Opus 5
63.1
Claude Fable 5
62.1
GPT-5.6 Sol
60.9

Cost per task

weighted USD per Intelligence Index task · lower is better
DeepSeek V4 Flash
$0.027
DeepSeek V4 Pro
$0.056
Kimi K3
$0.84
GPT-5.6 Sol
$1.23
Claude Opus 5
$2.34
Claude Fable 5
$3.14

Agentic tool use

τ³-Banking · higher is better
Kimi K3
46%
DeepSeek V4 Pro
40%
DeepSeek V4 Flash
39%
GPT-5.6 Sol
44%
Claude Opus 5
42%
Claude Fable 5
38%

Coding

Terminal-Bench v2.1 · higher is better
Kimi K3
85%
DeepSeek V4 Pro
79%
DeepSeek V4 Flash
79%
GPT-5.6 Sol
88%
Claude Opus 5
89%
Claude Fable 5
85%

Long-context reasoning

AA-LCR · higher is better
Kimi K3
83%
DeepSeek V4 Pro
75%
DeepSeek V4 Flash
74%
GPT-5.6 Sol
78%
Claude Opus 5
76%
Claude Fable 5
77%

Source: Artificial Analysis (independent evaluations), as of 2026-08-09.

Available models

Production-ready today, with new frontier releases added as they ship.

ModelContextCapabilitiesLatency (TTFT)Output speedSuccess rate
Kimi K31M
tools
reasoning
structured output
vision
3.0 s38 tok/s100.0%
DeepSeek V4 Pro 08131M
tools
reasoning
structured output
0.5 s110 tok/s100.0%
DeepSeek V4 Flash 07311M
tools
reasoning
structured output
0.5 s115 tok/s100.0%

Click a row for capabilities and quickstart. Latency, output speed and success rate are medians of 402 streamed completions measured on the Corion API (as of 2026-09-19), so they reflect the routing a customer actually gets. Latency is time to first token and includes the model's reasoning phase.

Which one should I pick?

  • Kimi K3 — the highest-intelligence option: deep reasoning, agentic workflows, long-horizon coding and long-document work. Typical fits: autonomous agents and multi-step tool use, codebase-scale refactoring, research and analysis over large document sets, knowledge work where accuracy matters more than latency.
  • DeepSeek V4 Pro 0813 — the value pick: near-frontier reasoning at roughly a fifteenth of K3's cost per task, with the same 1M context and 375K max output. Typical fits: reasoning-heavy analysis and long-horizon agents that run too hot on a frontier rate card, full-codebase review, multi-step automation, and large-scale synthesis. Text only — send images to Kimi K3.
  • DeepSeek V4 Flash — the cost- and latency-optimized option with the same 1M context. Typical fits: high-volume chat and support bots, classification and extraction pipelines, batch summarization, real-time assistants, and any workload where unit cost dominates.