Four identical ceramic tasting bowls, measured tea leaves, brass markers, and evaluation tools in a mountain tea room

011334 answers evaluated

TeaBench

How well can AI answer questions about tea?

Tea questions.
AI answers.

TeaBench · v3.0Every model answers the same 58 questions, without web searches or external tools.

Four tea samples, tasting cups, a timer, and thermometer arranged for comparison
Brewing, culture, health, and practical problem-solving

Results

Leaderboard

v3.0 · no external tools

Select a model to reveal its server, quantization, and sampling settings.

Model050100ScoreMedian wait*Peak GPU‡
01
92.6Qwen 3.8 27B · high score 92.6 out of 100, 95% task-sampling interval 89.8 to 95.1
5m 24s
56.0 GiB
02
91.8GPT-5.6 Sol · max score 91.8 out of 100, 95% task-sampling interval 88.3 to 94.8
41s
—
03
91.7Qwen 3.8 Flash Next · xhigh score 91.7 out of 100, 95% task-sampling interval 88.4 to 94.8
2m 27s
—
04
91.1DeepSeek V4 Flash Vision Exp · max score 91.1 out of 100, 95% task-sampling interval 88.0 to 94.1
5m 33s
—
05
89.6DeepSeek V4 Flash 0731 · max score 89.6 out of 100, 95% task-sampling interval 86.1 to 92.8
8m 34s
—
06
87.9GLM-5.2 · reason 16k score 87.9 out of 100, 95% task-sampling interval 83.4 to 91.9
5m 24s
187.1 GiB
07
86.9Bonsai 2 27B · high score 86.9 out of 100, 95% task-sampling interval 81.5 to 91.6
2m 34s
—
08
86.8Qwen 3.8 27B · medium score 86.8 out of 100, 95% task-sampling interval 81.5 to 91.4
1m 15s
—
09
86.7Kimi K2.6 · auto score 86.7 out of 100, 95% task-sampling interval 82.7 to 90.5
6m 43s
183.6 GiB
10
86.4Muse Glimmer 30B · auto score 86.4 out of 100, 95% task-sampling interval 81.7 to 90.6
5m 29s
56.2 GiB
11
84.2Signal 3.8 27B · medium score 84.2 out of 100, 95% task-sampling interval 77.3 to 89.9
3m 14s
—
12
83.4Agnes 3.0 Flash Preview · medium score 83.4 out of 100, 95% task-sampling interval 77.2 to 88.9
4m 19s
—
13
82.6MiniMax M3 · default score 82.6 out of 100, 95% task-sampling interval 77.1 to 87.5
50s
—
14
80.6MiMo V2.5 · reason 16k score 80.6 out of 100, 95% task-sampling interval 75.8 to 85.1
27s
144.5 GiB
15
80.1Gemini 3.1 Pro · dynamic score 80.1 out of 100, 95% task-sampling interval 73.4 to 86.2
35s
—
16
77.6GPT-OSS 120B · high score 77.6 out of 100, 95% task-sampling interval 70.5 to 84.2
35s
61.9 GiB
17
76.1Ling 3.0 Flash · thinking on score 76.1 out of 100, 95% task-sampling interval 70.0 to 81.8
28s
—
18
76.1Qwen 3.6 35B-A3B · high score 76.1 out of 100, 95% task-sampling interval 70.1 to 81.8
22s
23.5 GiB
19
74.3Gemma 4 31B BF16 score 74.3 out of 100, 95% task-sampling interval 68.7 to 79.4
1m 14s
65.1 GiB
20
73.8Nemotron 3 Super 120B · thinking on score 73.8 out of 100, 95% task-sampling interval 67.2 to 80.2
Mixed setups
—
21
72.1Gemma 4 31B Q4_K_M score 72.1 out of 100, 95% task-sampling interval 66.4 to 77.7
1m 49s
25.3 GiB
22
52.8Phi-4 score 52.8 out of 100, 95% task-sampling interval 45.6 to 60.0
7s
15.3 GiB
23
42.1Ling 3.0 Tiny · thinking off score 42.1 out of 100, 95% task-sampling interval 35.0 to 49.0
39s
31.8 GiB
85.9GLM 5.3 Flash · max score 85.9 out of 100, 95% task-sampling interval 80.8 to 90.3
95% interval 80.8–90.3
—

Scores show how well answers meet the scoring criteria, not the percentage of questions answered correctly. *Median wait: half of the scored answers finished sooner. Includes thinking time, not scoring time. ‡ Peak GPU: the resident GPU memory measured on the benchmark machine while a model served answers - weights plus key/value cache. A dash marks a row served through an API, or from a machine that did not record this number.

Answer time

Is a better answer worth the wait?

Higher is better. Further left is faster. Select a point to see the model.

Quality score ↑
30
40
50
60
70
80
90
100
1s
10s
1m 00s
10m 00s

← Faster · median answer time · logarithmic scale

The dotted line connects the best score–time tradeoffs (the Pareto frontier). Among these measured models, a higher score means a longer wait.

Models ran on different local and cloud setups, so this is not a controlled hardware comparison. Models tested across changing setups are left out of this chart.

Answer times, completed tests & costs
Times for the answers scored. Median: half finished sooner. P90: nine in ten finished within this time.
ModelBest tradeoffMedianP90Answers scoredFirst-try successCost / answer
Qwen 3.8 27B · highFrontier5m 24s10m 51s58/58UnknownNot measured
GPT-5.6 Sol · maxFrontier41s4m 21s58/58UnknownSubscription
Qwen 3.8 Flash Next · xhigh—2m 27s5m 17s58/58UnknownNot measured
DeepSeek V4 Flash Vision Exp · max—5m 33s8m 20s58/58UnknownNot measured
DeepSeek V4 Flash 0731 · max—8m 34s15m 50s58/58UnknownNot measured
GLM-5.2 · reason 16k—5m 24s12m 26s58/58UnknownNot measured
Bonsai 2 27B · high—2m 34s7m 30s58/58UnknownNot measured
Qwen 3.8 27B · medium—1m 15s2m 21s58/58UnknownNot measured
Kimi K2.6 · auto—6m 43s11m 41s58/58UnknownNot measured
Muse Glimmer 30B · auto—5m 29s9m 12s58/58UnknownNot measured
Signal 3.8 27B · medium—3m 14s5m 20s58/58UnknownNot measured
Agnes 3.0 Flash Preview · medium—4m 19s8m 16s58/58UnknownNot measured
MiniMax M3 · default—50s2m 06s58/58UnknownNot measured
MiMo V2.5 · reason 16kFrontier27s1m 06s58/58UnknownNot measured
Gemini 3.1 Pro · dynamic—35s2m 45s58/58Unknown≈ $0.0137 · estimated
GPT-OSS 120B · high—35s59s58/58UnknownNot measured
Ling 3.0 Flash · thinking on—28s1m 35s58/58UnknownNot measured
Qwen 3.6 35B-A3B · highFrontier22s39s58/58UnknownNot measured
Gemma 4 31B BF16—1m 14s2m 49s58/58UnknownNot measured
Nemotron 3 Super 120B · thinking onNot comparedMixed setups—58/58UnknownNot measured
Gemma 4 31B Q4_K_M—1m 49s3m 00s58/58UnknownNot measured
Phi-4Frontier7s12s58/58UnknownNot measured
Ling 3.0 Tiny · thinking off—39s1m 46s58/58UnknownNot measured

Times include thinking and writing, but not scoring or model loading managed by the runner. Some records exclude retries and retry delays. We did not control whether a model was already loaded or sharing hardware. The chart’s time axis uses uneven spacing to fit both fast and slow models; the dotted line does not predict results between points.

All models shown completed the same questions. We do not have complete records of failed attempts, so first-try success is unknown. Timing summaries give each question equal weight.

Costs cover answers, not judging or research. Estimates are not verified bills; some older estimates count only generated text, excluding input and cached text charges. Subscriptions still cost money. Local electricity and hardware costs were not measured—not zero.

Decisions

Tea Decisions v2 · base models

188 items · graded by answer key · no judges

Deterministic decision accuracy on 188 private tea decision items: binary calls, constrained choice, and anchored ordinal rating.

Model050100AccuracyMedian wait†Peak GPU‡
01
88.8Jev decision accuracy 88.8 out of 100, 95% item-sampling interval 84.0 to 93.1
280 ms
—
02
86.7Jev-Omni decision accuracy 86.7 out of 100, 95% item-sampling interval 81.9 to 91.5
50 ms
23.0 GiB
03
84.6Cygnet decision accuracy 84.6 out of 100, 95% item-sampling interval 79.3 to 89.4
40 ms
39.4 GiB
04
84.0Clef 27B decision accuracy 84.0 out of 100, 95% item-sampling interval 78.7 to 88.8
224 ms
—
05
83.5Winnow 12B decision accuracy 83.5 out of 100, 95% item-sampling interval 78.2 to 88.8
60 ms
14.5 GiB
06
82.4Imajev 4B decision accuracy 82.4 out of 100, 95% item-sampling interval 77.1 to 87.8
78 ms
10.2 GiB
07
77.7Gemma 4 12B (chat) decision accuracy 77.7 out of 100, 95% item-sampling interval 71.8 to 83.5
518 ms
39.4 GiB
08
54.3Phi-4 decision accuracy 54.3 out of 100, 95% item-sampling interval 46.8 to 61.2
1.3s
15.3 GiB

† Median time from one question to the answer, measured on one RTX PRO 6000 GPU answering one question at a time. These are local models with fixed hardware: there is no per-token price to show — what serving costs is the GPU memory each model needs and the wait above. ‡ Peak GPU is resident GPU memory measured on this machine while the model answered: weights, key/value cache, and runtime together. A dash marks a model served through an API.

How this decision test works

Single request per item at temperature 0; the answer must be exactly one option key. Anything else is a format failure and counts as wrong.

Every item is graded deterministically against a frozen answer key by exact key match; there is no LLM judge anywhere in this benchmark.

Eighteen paired items check that answers follow real evidence changes and resist injected instructions or irrelevant narrative; a pair counts only when both members are correct.

95% descriptive item-sampling intervals from a fixed-seed bootstrap over the 188 items; they describe sensitivity to item selection, not run-to-run variation (grading itself is deterministic).

Median wall time from request to final answer, one request at a time, on one workstation GPU (RTX PRO 6000 Blackwell). The decision-native models read their answer straight off the model's logits and generate no answer text.

Everything here runs locally on that single GPU with weights resident in memory: there are no per-token API fees to compare. What serving costs is the memory a model needs and the wait it causes. Energy use and hardware amortization are not measured and are not published as costs.

Two cohorts share the headline score: the 88 everyday items carried over from release v1 and the 100 compositional stress items added in v2. Items that every model passes add no separating evidence, so check both cohorts before trusting an ordering.

The corpus is private and freshly authored, but the knowledge-control group uses public tea facts. Ordinal rubric items require the exact rubric level; adjacent levels count wrong. The calibration group (known-probability items) is excluded from scoring until a confidence protocol exists. A single answer per item at zero temperature measures the decision, not its stability across re-prompts.

Sanitized aggregates only. The private corpus and per-item answers are not published.

Vision

Tea Vision v0.3 · vendor-catalogue imagery

155 items · graded by answer key · no judges

Accuracy on 155 image questions about real tea products from public vendor catalogues: which object is shown, which tea is sold, and whether the answer holds when the camera angle changes.

Model050100AccuracyTop-downPeak GPU†
01
68.4Qwen3-VL 30B accuracy 68.4 out of 100, 95% interval 60.7 to 75.2
51%
44.1 GiB
02
54.2Jev-Omni accuracy 54.2 out of 100, 95% interval 46.3 to 61.8
44%
23.0 GiB
—
19.4Qwen3-VL 30B (text only) accuracy 19.4 out of 100, 95% interval 13.9 to 26.3
15%
44.1 GiB

Top-down repeats the same questions on overhead close-ups of the dry tea - the hardest angle, where both models lose their stride. Rows marked — are controls: the same model without the image, to show how much of the score the picture carries.

† Peak GPU is resident GPU memory measured on this machine while the model answered: weights, key/value cache, and runtime together.

Follow-up studies on this corpus: composing several views into one numbered grid lifted Jev-Omni from 66% to 74% on 58 ware products, while naive probability averaging across views left it at 67% (its per-view confidences are near one-hot and its errors correlate across angles), and averaging a chat model's token log-probabilities collapsed accuracy to 28% - token distributions are not decision probabilities. The oracle ceiling - always picking the one view each model answers correctly - sits at 81%.

How this vision test works

One request per question: one image and four lettered options, with the correct answer's letter position permuted across items.

Every item is exact-match against an answer key: the chosen letter must match, and an answer that does not follow the required format counts wrong. No judges are involved at any step.

Answer keys come from the vendors' own catalogue taxonomy and listing claims, not from tea expertise. Origin and quality claims are reported separately and never scored.

Images are public product photography from well-known tea vendors, so models may have seen these catalogues during training. Scores describe decisioning on familiar catalogues, not sealed generalization. This is the deliberate inverse of the Tea Decisions board above, which runs on a private corpus the models have not seen.

The interval is a 95% Wilson score interval over the 155 scored items.

All models ran locally on one RTX PRO 6000 GPU, one question at a time. Jev-Omni answers with its own decision head; Qwen3-VL 30B answers in prompted chat mode, and its text-only row shows how much of its score comes from actually reading the image.

Version 0: 155 items across two model families and one control condition. Vendor taxonomy is marketing language, so a correct answer means agreement with the catalogue, not tea truth.

Sanitized aggregates only: item counts and scores, nothing from the corpus or the answer keys.

Scores by category

Where models do well

Scores out of 100, grouped by topic and type of problem.

Topics
BrewingCultureHealthProcessingExperimentsEquipment
Qwen 3.8 27B · high
86.594.594.495.593.591.1
GPT-5.6 Sol · max
88.594.593.396.090.587.8
Qwen 3.8 Flash Next · xhigh
84.094.597.297.091.586.1
DeepSeek V4 Flash Vision Exp · max
84.596.093.397.088.587.2
DeepSeek V4 Flash 0731 · max
86.094.590.693.081.092.8
GLM-5.2 · reason 16k
84.095.083.999.080.085.6
Bonsai 2 27B · high
84.097.082.897.080.580.0
Qwen 3.8 27B · medium
87.093.094.488.580.577.2
Kimi K2.6 · auto
86.091.086.790.576.090.0
Muse Glimmer 30B · auto
83.592.077.293.582.090.0
Signal 3.8 27B · medium
80.594.585.088.579.577.1
Agnes 3.0 Flash Preview · medium
79.589.587.885.072.586.1
MiniMax M3 · default
83.089.581.191.067.583.3
MiMo V2.5 · reason 16k
78.085.584.487.069.579.4
Gemini 3.1 Pro · dynamic
74.596.080.082.070.078.3
GPT-OSS 120B · high
66.086.075.091.573.573.9
Ling 3.0 Flash · thinking on
79.077.077.884.563.575.0
Qwen 3.6 35B-A3B · high
65.079.082.289.068.572.8
Gemma 4 31B BF16
70.079.578.984.560.072.8
Nemotron 3 Super 120B · thinking on
71.073.577.879.066.575.0
Gemma 4 31B Q4_K_M
59.578.078.383.557.575.6
Phi-4
36.566.057.267.550.039.4
Ling 3.0 Tiny · thinking off
27.051.943.957.032.040.6
Types of problem
CalculationsCausesUncertaintyPlanningDiagnosisKnowledge
Qwen 3.8 27B · high
93.094.095.096.182.294.5
GPT-5.6 Sol · max
93.093.585.096.185.697.5
Qwen 3.8 Flash Next · xhigh
99.091.089.590.085.694.5
DeepSeek V4 Flash Vision Exp · max
88.594.587.594.485.096.5
DeepSeek V4 Flash 0731 · max
95.087.589.594.475.095.0
GLM-5.2 · reason 16k
91.087.089.594.466.798.0
Bonsai 2 27B · high
84.093.086.095.065.697.5
Qwen 3.8 27B · medium
91.589.086.095.063.394.5
Kimi K2.6 · auto
93.585.590.091.771.187.0
Muse Glimmer 30B · auto
97.591.069.592.275.692.5
Signal 3.8 27B · medium
91.492.083.090.052.294.5
Agnes 3.0 Flash Preview · medium
93.587.584.091.755.086.0
MiniMax M3 · default
92.077.589.592.261.182.0
MiMo V2.5 · reason 16k
77.585.081.091.758.988.5
Gemini 3.1 Pro · dynamic
93.579.081.092.838.993.0
GPT-OSS 120B · high
85.077.076.587.250.688.5
Ling 3.0 Flash · thinking on
89.070.077.585.653.979.5
Qwen 3.6 35B-A3B · high
92.083.072.080.643.382.5
Gemma 4 31B BF16
89.073.568.080.045.687.0
Nemotron 3 Super 120B · thinking on
86.574.073.582.846.177.5
Gemma 4 31B Q4_K_M
87.575.567.070.047.281.5
Phi-4
56.053.561.544.435.064.5
Ling 3.0 Tiny · thinking off
37.545.550.544.928.944.0

Other tests

More AI comparisons

Privately maintained questions · scoring checked against sources · one answer per question · two AI judges · no human expert review

How scoring works & what it can tell us

Two AI models grade each answer against source-backed criteria, with disagreements reviewed against the same sources. Each of the six topics contributes equally to the overall score. Settings such as “high” or “max” in model names indicate the thinking mode used.

The initial results reuse answers and scores collected while developing the test, checked against the final question set. They are not a separate fresh test. Each model has one counted answer per question; we do not pick its best answer from repeated attempts.

The thin bars show a 95% range calculated by repeatedly resampling questions within each topic. They show how scores depend on the questions included—not how answers or judges might vary on another run, or whether one model is better in general.

Peak VRAM is the highest total GPU occupancy observed in a 2.5-second series while the model served answers on the evaluation workstation (2 GPUs, 97,887 MiB each), recorded under an exclusive GPU lease. It measures the serving state reached during the run, not a capacity claim; models served with deliberate CPU/host-RAM offload exceed per-GPU memory by design.

This is a personal comparison, not an expert-certified test. Questions are kept private, but we cannot guarantee models have never seen similar material. Some categories have more questions than others.