Head to head on production benchmarks
Scores are shown within each benchmark. Higher is better.
An open-weight model now sits inside the frontier cluster on the benchmarks that map to real engineering and finance work.
The price of parity
Token prices mislead across models with different verbosity; cost per completed task is the metric that maps to a budget.
Vals AI, measured · Updated 2026-07-17
K3 is the slowest in the cluster. Token inefficiency is real; it shows up in wall-clock time, not in cost.
(open weights)
* Claude Sonnet 5 introductory pricing: $2 input / $10 output through August 2026.
Per token, K3 costs Sonnet-tier money. Per completed job, it is 5 to 10x cheaper than the frontier — the token inefficiency is priced in.
The gap is closing
Estimated lag between the best open-weight and frontier proprietary models.
* July 2026 uses the Artificial Analysis Intelligence Index distance, 57.1 versus approximately 60, as a current parity proxy rather than a measured time lag.
What took the open ecosystem a year to replicate in 2024 now takes one quarter.
Where the frontier still leads
Legal Research Bench
Deep-research retrieval still favors frontier models.
DeepSWE
The strongest closed model retains a 5.5-point lead.
Humanity’s Last Exam + vision
Frontier-hard reasoning and multimodal work remain differentiated.
Parity is benchmark-specific. For frontier-hard reasoning and legal research, closed models keep an edge, for now.
What the bear case gets right
A note circulating among investors makes four claims about open models and AI economics. Here is what public data supports, and what it does not.
“Kimi K3 is not cheap. It is a massive multi-trillion-parameter model that is token-inefficient and requires advanced compute.”
Half rightRight on the sticker: at $3/$15 per 1M tokens, K3 ends the era of discount Chinese AI pricing, and it is the slowest model in the frontier cluster (619s per SWE-bench task vs 182s for GPT-5.6 Sol). Wrong on the conclusion: measured cost per completed task is $0.21 vs $1.15 to $2.05 for the frontier — the token inefficiency is already priced in. The compute-demand read (2.8T parameters, heavy inference hardware to self-host) stands on its own.
“China has not caught up. Western labs hold back their best models, so released comparisons flatter the open ecosystem.”
Unresolved — measured trend points the other wayInternal, unreleased models are not benchmarked; the claim cannot be tested with public data. What can be measured: the released-model gap compressed from about 12 months in 2024 to 3-4 months in 2026 (Epoch AI). Buyers deploy released models, and on that comparison the gap is months and shrinking. The note names its own test: if western releases in the next 45-60 days re-open the gap, this claim gains support. Watch that window.
“It is unclear whether frontier labs make money on models today.”
Pending disclosure — margins improvingPublic run-rate reporting is contradictory, but reported inference gross margins have swung from negative in 2024 to strongly positive in 2026. The note is right about the catalyst: IPO disclosures will settle it.
“It is unclear whether hyperscalers earn an adequate return on AI capex.”
OpenOutside the evidence base of model benchmarks. The next several quarters of hyperscaler earnings are the honest test. No verdict offered here.
The core case on this page survives the strongest bear note: whatever exists inside frontier labs, the released open-weight cluster already delivers frontier-grade production benchmarks at the lowest measured cost per job in the market.
Three implications
Route by workload
Model choice is now a routing decision per workload, not a loyalty decision.
Pricing power breaks
Open weights constrain vendor pricing. Open competition sets the $15 output ceiling.
Deploy with sovereignty
On-prem deployment of frontier-cluster capability is now realistic because the weights are public.
Sources
- Vals AI · SWE-bench Verified
Leaderboard scores, cost per task, and latency. - Vals AI · Kimi K3 model page
Vals AI Index score of 74.70%. - Vals AI · Benchmarks index
Production benchmark comparisons. - Vals AI · Legal Research Bench
Legal research performance. - OpenRouter · Kimi K3
$3 / $15 pricing and one-million-token context. - Epoch AI · Open-closed ECI gap
Four-month average lag in May 2026. - Epoch AI · Open versus closed weights
Approximately three-month lag in October 2025. - Epoch AI · Open models report
Approximately one-year gap in 2024. - Artificial Analysis · Intelligence Index
Frontier-cluster intelligence scores. - The Decoder · Kimi K3 analysis
Pricing and per-task cost analysis.