๐Ÿ  Home
FRONTIER MODELS ยท RELEASE TRACKER

๐Ÿš€ LLM Model Release Tracker

Every frontier-scale model release worth tracking โ€” specs, timing, and what it means.

updated 2026-08-13 ยท Notion ยท Frontier Model Releases (be1366fa-33a8-4eda-8b83-6abf18ce78d1)

29
models tracked
15
Chinese labs (15/29)
18/28
released @ 1M context
4
releases in Aug '26

๐Ÿ’ก Key insights

1. 9 of 16 released models ship 1M context โ€” only older Claude Opus 4.7/4.8 and MiniMax M2.7 lag. Context length stopped being a differentiator.
2. MoE everywhere: DeepSeek V4-Pro 1.6T total / 49B active (3%), MiniMax M3 428B/23B, Kimi K3 2.8T/16B (0.6%). Inference cost โ‰ˆ active params โ€” that asymmetry is the whole MoE business model.
3. 16 releases in ~5 months; 4 landed in the first 2 weeks of August alone (Qwen 3.8 Max, Grok 4.6, DeepSeek V4-Pro, Gemini 3.7 Flash).
4. Claude and Gemini publish no param counts ('not public') โ€” they compete on capability, not spec sheets. Open labs publish everything.
5. 10 of 17 rows are Chinese labs (DeepSeek ร—3, Kimi ร—3, Qwen, GLM, MiniMax ร—2) โ€” and GLM-5.3 is pending.
6. No Mistral, no Llama/Meta, no xAI mid-tier, no Apple. The table tracks who's actually shipping frontier-scale.

๐Ÿ’ธ Pricing landscape updated 2026-08-18 ยท snapshot

Flat plans โ€” the only flat-rate open-model access is Ollama Cloud (GPU-time billed, not tokens). GLM Coding Plan is flat but GLM-only. Everything else is PAYG per-token.
ProviderPlanPriceModel access
Ollama CloudFree$0All open models (light usage)
Ollama CloudPro$20/mo ($200/yr)All open models, 3 concurrent
Ollama CloudMax$100/moAll open models, 10 concurrent
Ollama CloudTeam$25/seat/mo (5 min)Shared pool
GLM Coding PlanLite$18/mo ($12.6 yearly)GLM-5.3/5.2 only, 10K credits/wk
GLM Coding PlanPro$80/mo ($56 yearly)GLM-5.3, 6x Lite
GLM Coding PlanMax$168/mo ($117.6 yearly)GLM-5.3, 14x Lite
PAYG APIs โ€” per-1M-token official rates. At Joe's ~2.17B tokens/mo, PAYG โ‰ˆ $350-480/mo vs $20 flat: Ollama is 20-100x cheaper for the same weights.
ModelProviderIn /1MOut /1MCache hitNote
DeepSeek V4-ProDeepSeek official$1.32 peak / $0.66 off-peak$3.96 peak / $1.98 off-peak$0.044 peak / $0.022 off-peakpeak/off-peak billing live Aug 16, 2026. History: Apr 24 $3.48 launch โ†’ May 26 $0.87 (75% discount permanent) โ†’ Aug 16 $3.96/$1.98.
DeepSeek V4-FlashDeepSeek official$0.44 peak / $0.22 off-peak$1.32 peak / $0.66 off-peak$0.014 peak / $0.007 off-peakpeak/off-peak billing live Aug 16, 2026; was $0.28 out flat.
Kimi K2.7 CodeMoonshot official$0.95$4.00$0.19
GLM-5.2Z.ai API~$0.40~$1.20โ€”est.; Z.ai pushes Coding Plan instead
DeepSeek V4-FlashSiliconFlowยฅ1.00ยฅ2.00ยฅ0.02CN, no VPN
GLM-5.2SiliconFlowยฅ8.00ยฅ28.00ยฅ2.00CN, no VPN
Why Ollama is so cheap: Cloud is a loss-leader conversion funnel for the local software moat. GPU-time billed at near-wholesale; per-token providers structurally can't match flat pricing at heavy volume.
Risk: VC-subsidized pricing can change; Max sign-ups already paused Aug 2026 (capacity). Annual Pro ($200/yr) hedges 2 months.

๐Ÿ—“๏ธ Release timeline

Jan Feb Mar Apr May Jun Jul Aug Kimi Kimi K2.5 Kimi K2.6 Kimi K2.7 Code Kimi K3 MiniMax MiniMax M2.5 MiniMax M2.7 MiniMax M3 Qwen Qwen 3.5 (397B) Qwen 3.8 Max GLM GLM-5.1 GLM-5.2 GLM-5.3 (expected) GPT GPT-5.3 Codex GPT-5.4 GPT-5.5 GPT-5.6 Grok Grok 4.3 Grok 4.6 Gemini Gemini 3.1 Pro Gemini 3.5 Flash Gemini 3.7 Flash Claude Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Claude Fable 5 Claude Opus 5 DeepSeek DeepSeek V4-Flash DeepSeek V4-Pro DeepSeek V4-Pro (0813)

๐Ÿ“Š MoE efficiency: total vs active params

GLM-5.1 754B / 40B active Qwen 3.5 (397B) 397B / 17B active MiniMax M2.5 229B / 10B active Kimi K2.5 1.0T / 32B active Qwen 3.8 Max 2.4T Grok 4.6 1.5T Kimi K3 2.8T Kimi K2.7 Code 1.0T Kimi K2.6 1.0T DeepSeek V4-Pro (0813) 1.6T / 49B active DeepSeek V4-Flash 284B / 13B active DeepSeek V4-Pro 1.6T / 49B active MiniMax M3 428B / 23B active MiniMax M2.7 229B / 10B active GLM-5.2 753B โ–  total params (dim) ยท โ–  active params (accent) โ€” log scale

๐Ÿ“‹ All models

Model Family Release Date Parameters Context Status
Claude Fable 5Claude2026-06-09not public1MReleased
Gemini 3.5 FlashGemini2026-05-19not public1MReleased
Grok 4.3Grok2026-04-30not public1MReleased
GPT-5.5GPT2026-04-23not public1MReleased
GLM-5.1GLM2026-04-07754B40B active200KReleased
GPT-5.4GPT2026-03-05not public1MReleased
Gemini 3.1 ProGemini2026-02-19not public1MReleased
Qwen 3.5 (397B)Qwen2026-02-16397B17B active262KReleased
MiniMax M2.5MiniMax2026-02-12229B10B active192KReleased
GPT-5.3 CodexGPT2026-02-05not public400KReleased
Claude Opus 4.6Claude2026-02-05not public1MReleased
Kimi K2.5Kimi2026-01-271.0T32B active256KReleased
Qwen 3.8 MaxQwen2026-08-032.4T1MReleased
Grok 4.6Grok2026-08-071.5T (V9 base)1MReleased
Gemini 3.7 FlashGemini2026-08-13not public1MReleased
GPT-5.6GPT2026-07-09not public1MReleased
Kimi K3Kimi2026-07-162.8T total / 16 of 896 active1MReleased
Kimi K2.7 CodeKimi2026-06-12~1T (MoE)256KReleased
Kimi K2.6Kimi2026-04-21~1T (MoE)256KReleased
DeepSeek V4-Pro (0813)DeepSeek2026-08-071.6T49B active1MReleased
DeepSeek V4-FlashDeepSeek2026-04-24284B13B active1MReleased
DeepSeek V4-ProDeepSeek2026-04-241.6T49B active1MReleased
Claude Opus 5Claude2026-07-24not public1MReleased
Claude Opus 4.8Claude2026-05-28not public200KReleased
Claude Opus 4.7Claude2026-04-16not public200KReleased
MiniMax M3MiniMax2026-06-01428B23B active1MReleased
MiniMax M2.7MiniMax2026-03-18229B10B active200KReleased
GLM-5.3GLMโ€”TBDTBDExpected
GLM-5.2GLM2026-06-13753B (MoE)1MReleased
Data: Notion ยท Frontier Model Releases (be1366fa-33a8-4eda-8b83-6abf18ce78d1) ยท verified against lab announcements & Artificial Analysis. Page regenerated biweekly from the Notion table โ€” Notion remains canonical. Built by Heidi.