TASKCLANStart building ↗

TASKCLAN / T1 BENCHMARKS

Pick the right
intelligence.

T1 profiles route every task to the strongest available frontier model. Here’s how Core, Flow and Max are positioned relative to one another, and how to choose between them.

Last updated: July 18, 2026

T1 Core

Fast, grounded intelligence for high-volume everyday work.

Lowest latency and cost per task
Speed96
Reasoning depth55
Cost efficiency96

T1 Flow

Adaptive orchestration for connected tools, teams and workflows.

Balanced quality, speed and cost — the default
Speed76
Reasoning depth82
Cost efficiency74

T1 Max

Maximum reasoning depth for research, code and consequential work.

Highest quality and reasoning depth
Speed46
Reasoning depth98
Cost efficiency44

Bars show relative positioning between profiles by design, not absolute measured scores.

At a glance

Core vs Flow vs Max.

CharacteristicT1 CoreT1 FlowT1 Max
Optimised forSpeed & volumeBalanced orchestrationDepth & quality
Response speedFastestFastDeliberate
Reasoning depthFocusedAdaptiveMaximum
Context handlingLargeLargeLargest
Relative costLowestBalancedHighest
Tool useSingle-toolMulti-agentMulti-agent
ModalityText-firstText + toolsFull multimodal
Default profile

Task fit

Which profile for the job.

TaskT1 CoreT1 FlowT1 Max
Classification & extraction Best fit··
Search & retrieval Best fit··
High-volume support replies Best fit··
Drafting & content· Best fit·
Multi-step agents· Best fit·
Tool & workflow orchestration· Best fit·
Research & analysis·· Best fit
Code generation & review·· Best fit
Strategy & planning·· Best fit
Multimodal creation·· Best fit

T1 Intelligence Index

Overall capability, ranked.

A single score per model — the mean accuracy across all 4 benchmarks below. T1 Max leads the index at 93.3.

Highest indexT1 Max93.3 average across all suites
Toughest test · MMLU-ProT1 Max89% — the sternest reasoning benchmark
Math · GSM8KKimi K298% — grade-school math reasoning
1T1 Max
93.3
2T1 Flow
90.3
3Fable 5
88.5
4Kimi K3
84.5
5Kimi K2
81.0
6GPT-4o mini
77.8
7T1 Core
77.0
8Llama 3.1 8B
63.5

Intelligence Index = simple average of the 4 measured benchmark accuracies (higher is better). T1 tiers shown in colour.

Quality vs cost

Smart, and priced right.

Intelligence Index against blended price per million tokens (log scale). Up and to the left is better — more capability for less money. T1 tiers are shown in colour.

60708090100$0.10$1$10Blended price — $ / 1M tokens (log scale)Intelligence Index← better valueT1 MaxT1 FlowFable 5Kimi K3Kimi K2GPT-4o miniT1 CoreLlama 3.1 8B

Blended price = (3 × input + output) ÷ 4 per 1M tokens, from OpenRouter list pricing on July 20, 2026. T1 tier prices reflect the tier's underlying production model; provider prices change over time.

Standard benchmarks

The full breakdown.

BenchmarkT1 CoreT1 FlowT1 MaxFable 5Kimi K3Kimi K2GPT-4o miniLlama 3.1 8B
GSM8K (Accuracy %)9796959695989582
MMLU (Accuracy %)7688949383827555
MMLU-Pro (Accuracy %)4382898067564944
ARC-Challenge (Accuracy %)9295958593889273

All scores measured by Taskclan on the same date with one grader — 100 questions per benchmark, sampled evenly from each suite's public test split (3,200 graded runs, 0 errors). Multiple-choice graded by exact-letter match; GSM8K by final-answer match. Every model, including the baselines, was run through the identical harness; T1 tiers were evaluated through the production T1 routing endpoint. A representative sample, not a full-suite run. Measured July 20, 2026.

What they measure

The benchmarks, explained.

GSM8K · Grade School Math 8K
Grade-school math word problems that require multi-step arithmetic reasoning to reach a single numeric answer. Tests whether a model can work through a problem, not just recall a fact. Now largely saturated for frontier models.
MMLU · Massive Multitask Language Understanding
~14,000 four-option multiple-choice questions across 57 subjects — from math, physics and medicine to law, history and ethics. The standard broad test of general knowledge and reasoning.
MMLU-Pro
A harder, reasoning-focused successor to MMLU: ten answer options instead of four and tougher, less guessable questions. It reduces saturation and separates strong models far more clearly — the sternest test here.
ARC-Challenge · AI2 Reasoning Challenge (Challenge set)
Grade-school science multiple-choice questions deliberately chosen to defeat simple keyword-matching and retrieval. Solving them takes genuine reasoning rather than surface pattern-matching.

Methodology

How we benchmark.

Start building

One API.
Every profile.

Choose a profile in a single line, or let Flow decide. Route across the frontier without rewrites.

Read the models guide ↗