The best $0 OpenRouter models, ranked monthly on a fixed battery โ coding, reasoning, concision.
| # | Model | Composite | A ยท Coding | B ยท Reasoning | C ยท Concision | ฮ | Tokens | tok/s |
|---|---|---|---|---|---|---|---|---|
| 1 | Nemotron 3 Ultra (550B-A55B) | 9.9 | 9.8 | 10.0 | 10.0 | ๐ | 2086 | 17.1 |
| 2 | Gemma 4 26B (A4B) | 9.8 | 10.0 | 9.4 | 10.0 | ๐ | 1392 | 23.4 |
| 3 | Laguna S 2.1 | 9.1 | 9.8 | 9.4 | 8.0 | ๐ | 908 | 28.2 |
| 4 | Laguna XS 2.1 DNF A | 4.0 | 0.0 | 10.0 | 10.0 | ๐ | 584 | 91.5 |
| 5 | Ling 3.0 Tiny DNF B | 0.4 | 0.0 | 0.0 | 8.0 | ๐ | 1825 | 153.7 |
| 6 | Gemma 4 31B DNF A,B,C | 0.0 | 0.0 | 0.0 | 0.0 | ๐ | 0 | 0 |
Composite = 0.40รA + 0.30รB + 0.30รC ยท ฮ vs previous month (โฒ/โผ on ยฑ0.3) ยท โ ๏ธ DNF = task failed, โ2.0 penalty ยท all runs at temperature 0, $0 free tier
| Month | Nemotron 3 Ultra (550B-A55B) | Gemma 4 26B (A4B) | Laguna S 2.1 | Laguna XS 2.1 | Ling 3.0 Tiny | Gemma 4 31B |
|---|---|---|---|---|---|---|
| August 2026 | 9.9 ๐ | 9.8 ๐ | 9.1 ๐ | 4.0 ๐ | 0.4 ๐ | 0.0 ๐ |
Vendor descriptions & release dates from the OpenRouter catalog; community traction from HuggingFace. Our composite is independent โ this is the outside view.
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference โ delivering near-31B quality at...
Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...