Every frontier-scale model release worth tracking โ specs, timing, and what it means.
updated 2026-08-13 ยท Notion ยท Frontier Model Releases (be1366fa-33a8-4eda-8b83-6abf18ce78d1)
29
models tracked
15
Chinese labs (15/29)
18/28
released @ 1M context
4
releases in Aug '26
๐ก Key insights
1. 9 of 16 released models ship 1M context โ only older Claude Opus 4.7/4.8 and MiniMax M2.7 lag. Context length stopped being a differentiator.
2. MoE everywhere: DeepSeek V4-Pro 1.6T total / 49B active (3%), MiniMax M3 428B/23B, Kimi K3 2.8T/16B (0.6%). Inference cost โ active params โ that asymmetry is the whole MoE business model.
3. 16 releases in ~5 months; 4 landed in the first 2 weeks of August alone (Qwen 3.8 Max, Grok 4.6, DeepSeek V4-Pro, Gemini 3.7 Flash).
4. Claude and Gemini publish no param counts ('not public') โ they compete on capability, not spec sheets. Open labs publish everything.
5. 10 of 17 rows are Chinese labs (DeepSeek ร3, Kimi ร3, Qwen, GLM, MiniMax ร2) โ and GLM-5.3 is pending.
6. No Mistral, no Llama/Meta, no xAI mid-tier, no Apple. The table tracks who's actually shipping frontier-scale.
Flat plans โ the only flat-rate open-model access is Ollama Cloud (GPU-time billed, not tokens). GLM Coding Plan is flat but GLM-only. Everything else is PAYG per-token.
Provider
Plan
Price
Model access
Ollama Cloud
Free
$0
All open models (light usage)
Ollama Cloud
Pro
$20/mo ($200/yr)
All open models, 3 concurrent
Ollama Cloud
Max
$100/mo
All open models, 10 concurrent
Ollama Cloud
Team
$25/seat/mo (5 min)
Shared pool
GLM Coding Plan
Lite
$18/mo ($12.6 yearly)
GLM-5.3/5.2 only, 10K credits/wk
GLM Coding Plan
Pro
$80/mo ($56 yearly)
GLM-5.3, 6x Lite
GLM Coding Plan
Max
$168/mo ($117.6 yearly)
GLM-5.3, 14x Lite
PAYG APIs โ per-1M-token official rates. At Joe's ~2.17B tokens/mo, PAYG โ $350-480/mo vs $20 flat: Ollama is 20-100x cheaper for the same weights.
Model
Provider
In /1M
Out /1M
Cache hit
Note
DeepSeek V4-Pro
DeepSeek official
$1.32 peak / $0.66 off-peak
$3.96 peak / $1.98 off-peak
$0.044 peak / $0.022 off-peak
peak/off-peak billing live Aug 16, 2026. History: Apr 24 $3.48 launch โ May 26 $0.87 (75% discount permanent) โ Aug 16 $3.96/$1.98.
DeepSeek V4-Flash
DeepSeek official
$0.44 peak / $0.22 off-peak
$1.32 peak / $0.66 off-peak
$0.014 peak / $0.007 off-peak
peak/off-peak billing live Aug 16, 2026; was $0.28 out flat.
Kimi K2.7 Code
Moonshot official
$0.95
$4.00
$0.19
GLM-5.2
Z.ai API
~$0.40
~$1.20
โ
est.; Z.ai pushes Coding Plan instead
DeepSeek V4-Flash
SiliconFlow
ยฅ1.00
ยฅ2.00
ยฅ0.02
CN, no VPN
GLM-5.2
SiliconFlow
ยฅ8.00
ยฅ28.00
ยฅ2.00
CN, no VPN
Why Ollama is so cheap: Cloud is a loss-leader conversion funnel for the local software moat. GPU-time billed at near-wholesale; per-token providers structurally can't match flat pricing at heavy volume.
Risk: VC-subsidized pricing can change; Max sign-ups already paused Aug 2026 (capacity). Annual Pro ($200/yr) hedges 2 months.
๐๏ธ Release timeline
๐ MoE efficiency: total vs active params
๐ All models
Model
Family
Release Date
Parameters
Context
Status
Claude Fable 5
Claude
2026-06-09
not public
1M
Released
Gemini 3.5 Flash
Gemini
2026-05-19
not public
1M
Released
Grok 4.3
Grok
2026-04-30
not public
1M
Released
GPT-5.5
GPT
2026-04-23
not public
1M
Released
GLM-5.1
GLM
2026-04-07
754B40B active
200K
Released
GPT-5.4
GPT
2026-03-05
not public
1M
Released
Gemini 3.1 Pro
Gemini
2026-02-19
not public
1M
Released
Qwen 3.5 (397B)
Qwen
2026-02-16
397B17B active
262K
Released
MiniMax M2.5
MiniMax
2026-02-12
229B10B active
192K
Released
GPT-5.3 Codex
GPT
2026-02-05
not public
400K
Released
Claude Opus 4.6
Claude
2026-02-05
not public
1M
Released
Kimi K2.5
Kimi
2026-01-27
1.0T32B active
256K
Released
Qwen 3.8 Max
Qwen
2026-08-03
2.4T
1M
Released
Grok 4.6
Grok
2026-08-07
1.5T (V9 base)
1M
Released
Gemini 3.7 Flash
Gemini
2026-08-13
not public
1M
Released
GPT-5.6
GPT
2026-07-09
not public
1M
Released
Kimi K3
Kimi
2026-07-16
2.8T total / 16 of 896 active
1M
Released
Kimi K2.7 Code
Kimi
2026-06-12
~1T (MoE)
256K
Released
Kimi K2.6
Kimi
2026-04-21
~1T (MoE)
256K
Released
DeepSeek V4-Pro (0813)
DeepSeek
2026-08-07
1.6T49B active
1M
Released
DeepSeek V4-Flash
DeepSeek
2026-04-24
284B13B active
1M
Released
DeepSeek V4-Pro
DeepSeek
2026-04-24
1.6T49B active
1M
Released
Claude Opus 5
Claude
2026-07-24
not public
1M
Released
Claude Opus 4.8
Claude
2026-05-28
not public
200K
Released
Claude Opus 4.7
Claude
2026-04-16
not public
200K
Released
MiniMax M3
MiniMax
2026-06-01
428B23B active
1M
Released
MiniMax M2.7
MiniMax
2026-03-18
229B10B active
200K
Released
GLM-5.3
GLM
โ
TBD
TBD
Expected
GLM-5.2
GLM
2026-06-13
753B (MoE)
1M
Released
Data: Notion ยท Frontier Model Releases (be1366fa-33a8-4eda-8b83-6abf18ce78d1) ยท verified against lab announcements & Artificial Analysis. Page regenerated biweekly from the Notion table โ Notion remains canonical. Built by Heidi.