TASKCLANStart building ↗TASKCLAN / T1 BENCHMARKS
Pick the right
intelligence.
T1 profiles route every task to the strongest available frontier model. Here’s how Core, Flow and Max are positioned relative to one another, and how to choose between them.
Last updated: July 18, 2026
T1 Core
Fast, grounded intelligence for high-volume everyday work.
Lowest latency and cost per taskT1 Flow
Adaptive orchestration for connected tools, teams and workflows.
Balanced quality, speed and cost — the defaultT1 Max
Maximum reasoning depth for research, code and consequential work.
Highest quality and reasoning depthBars show relative positioning between profiles by design, not absolute measured scores.
At a glance
Core vs Flow vs Max.
| Characteristic | T1 Core | T1 Flow | T1 Max |
|---|---|---|---|
| Optimised for | Speed & volume | Balanced orchestration | Depth & quality |
| Response speed | Fastest | Fast | Deliberate |
| Reasoning depth | Focused | Adaptive | Maximum |
| Context handling | Large | Large | Largest |
| Relative cost | Lowest | Balanced | Highest |
| Tool use | Single-tool | Multi-agent | Multi-agent |
| Modality | Text-first | Text + tools | Full multimodal |
| Default profile | — | ✓ | — |
Task fit
Which profile for the job.
| Task | T1 Core | T1 Flow | T1 Max |
|---|---|---|---|
| Classification & extraction | Best fit | · | · |
| Search & retrieval | Best fit | · | · |
| High-volume support replies | Best fit | · | · |
| Drafting & content | · | Best fit | · |
| Multi-step agents | · | Best fit | · |
| Tool & workflow orchestration | · | Best fit | · |
| Research & analysis | · | · | Best fit |
| Code generation & review | · | · | Best fit |
| Strategy & planning | · | · | Best fit |
| Multimodal creation | · | · | Best fit |
T1 Intelligence Index
Overall capability, ranked.
A single score per model — the mean accuracy across all 4 benchmarks below. T1 Max leads the index at 93.3.
Intelligence Index = simple average of the 4 measured benchmark accuracies (higher is better). T1 tiers shown in colour.
Quality vs cost
Smart, and priced right.
Intelligence Index against blended price per million tokens (log scale). Up and to the left is better — more capability for less money. T1 tiers are shown in colour.
Blended price = (3 × input + output) ÷ 4 per 1M tokens, from OpenRouter list pricing on July 20, 2026. T1 tier prices reflect the tier's underlying production model; provider prices change over time.
Standard benchmarks
The full breakdown.
| Benchmark | T1 Core | T1 Flow | T1 Max | Fable 5 | Kimi K3 | Kimi K2 | GPT-4o mini | Llama 3.1 8B |
|---|---|---|---|---|---|---|---|---|
| GSM8K (Accuracy %) | 97 | 96 | 95 | 96 | 95 | 98 | 95 | 82 |
| MMLU (Accuracy %) | 76 | 88 | 94 | 93 | 83 | 82 | 75 | 55 |
| MMLU-Pro (Accuracy %) | 43 | 82 | 89 | 80 | 67 | 56 | 49 | 44 |
| ARC-Challenge (Accuracy %) | 92 | 95 | 95 | 85 | 93 | 88 | 92 | 73 |
All scores measured by Taskclan on the same date with one grader — 100 questions per benchmark, sampled evenly from each suite's public test split (3,200 graded runs, 0 errors). Multiple-choice graded by exact-letter match; GSM8K by final-answer match. Every model, including the baselines, was run through the identical harness; T1 tiers were evaluated through the production T1 routing endpoint. A representative sample, not a full-suite run. Measured July 20, 2026.
What they measure
The benchmarks, explained.
- GSM8K · Grade School Math 8K
- Grade-school math word problems that require multi-step arithmetic reasoning to reach a single numeric answer. Tests whether a model can work through a problem, not just recall a fact. Now largely saturated for frontier models.
- MMLU · Massive Multitask Language Understanding
- ~14,000 four-option multiple-choice questions across 57 subjects — from math, physics and medicine to law, history and ethics. The standard broad test of general knowledge and reasoning.
- MMLU-Pro
- A harder, reasoning-focused successor to MMLU: ten answer options instead of four and tougher, less guessable questions. It reduces saturation and separates strong models far more clearly — the sternest test here.
- ARC-Challenge · AI2 Reasoning Challenge (Challenge set)
- Grade-school science multiple-choice questions deliberately chosen to defeat simple keyword-matching and retrieval. Solving them takes genuine reasoning rather than surface pattern-matching.
Methodology
How we benchmark.
T1 is not a single model. Each profile is a routing strategy that continuously evaluates the strongest available open and proprietary models on three axes: quality, latency and cost. The best model for a given task can change as the frontier moves, so a profile describes an outcome, not a fixed model.
The positioning above reflects each profile’s design intent, Core favours speed and efficiency, Max favours reasoning depth, and Flow balances the two. It is a relative comparison between the profiles, not an absolute score against a public leaderboard.
Because profiles route to third-party frontier models, we do not restate those providers’ benchmark numbers as our own. When we publish measured quality-eval results, they appear in the table above with the suite and method noted. See the AI Policy for how we evaluate, and Models & T1 profiles for the developer view.
Start building
One API.
Every profile.
Choose a profile in a single line, or let Flow decide. Route across the frontier without rewrites.
Read the models guide ↗