Tea questions.
AI answers.
TeaBench · v3.0Every model answers the same 58 questions, without web searches or external tools.

Results
Leaderboard
v3.0 · no external tools
Select a model to reveal its server, quantization, and sampling settings.
Scores show how well answers meet the scoring criteria, not the percentage of questions answered correctly. *Median wait: half of the scored answers finished sooner. Includes thinking time, not scoring time. ‡ Peak GPU: the resident GPU memory measured on the benchmark machine while a model served answers - weights plus key/value cache. A dash marks a row served through an API, or from a machine that did not record this number.
Answer time
Is a better answer worth the wait?
Higher is better. Further left is faster. Select a point to see the model.
← Faster · median answer time · logarithmic scale
The dotted line connects the best score–time tradeoffs (the Pareto frontier). Among these measured models, a higher score means a longer wait.
Models ran on different local and cloud setups, so this is not a controlled hardware comparison. Models tested across changing setups are left out of this chart.
Answer times, completed tests & costs
| Model | Best tradeoff | Median | P90 | Answers scored | First-try success | Cost / answer |
|---|---|---|---|---|---|---|
| Qwen 3.8 27B · high | Frontier | 5m 24s | 10m 51s | 58/58 | Unknown | Not measured |
| GPT-5.6 Sol · max | Frontier | 41s | 4m 21s | 58/58 | Unknown | Subscription |
| Qwen 3.8 Flash Next · xhigh | — | 2m 27s | 5m 17s | 58/58 | Unknown | Not measured |
| DeepSeek V4 Flash Vision Exp · max | — | 5m 33s | 8m 20s | 58/58 | Unknown | Not measured |
| DeepSeek V4 Flash 0731 · max | — | 8m 34s | 15m 50s | 58/58 | Unknown | Not measured |
| GLM-5.2 · reason 16k | — | 5m 24s | 12m 26s | 58/58 | Unknown | Not measured |
| Bonsai 2 27B · high | — | 2m 34s | 7m 30s | 58/58 | Unknown | Not measured |
| Qwen 3.8 27B · medium | — | 1m 15s | 2m 21s | 58/58 | Unknown | Not measured |
| Kimi K2.6 · auto | — | 6m 43s | 11m 41s | 58/58 | Unknown | Not measured |
| Muse Glimmer 30B · auto | — | 5m 29s | 9m 12s | 58/58 | Unknown | Not measured |
| Signal 3.8 27B · medium | — | 3m 14s | 5m 20s | 58/58 | Unknown | Not measured |
| Agnes 3.0 Flash Preview · medium | — | 4m 19s | 8m 16s | 58/58 | Unknown | Not measured |
| MiniMax M3 · default | — | 50s | 2m 06s | 58/58 | Unknown | Not measured |
| MiMo V2.5 · reason 16k | Frontier | 27s | 1m 06s | 58/58 | Unknown | Not measured |
| Gemini 3.1 Pro · dynamic | — | 35s | 2m 45s | 58/58 | Unknown | ≈ $0.0137 · estimated |
| GPT-OSS 120B · high | — | 35s | 59s | 58/58 | Unknown | Not measured |
| Ling 3.0 Flash · thinking on | — | 28s | 1m 35s | 58/58 | Unknown | Not measured |
| Qwen 3.6 35B-A3B · high | Frontier | 22s | 39s | 58/58 | Unknown | Not measured |
| Gemma 4 31B BF16 | — | 1m 14s | 2m 49s | 58/58 | Unknown | Not measured |
| Nemotron 3 Super 120B · thinking on | Not compared | Mixed setups | — | 58/58 | Unknown | Not measured |
| Gemma 4 31B Q4_K_M | — | 1m 49s | 3m 00s | 58/58 | Unknown | Not measured |
| Phi-4 | Frontier | 7s | 12s | 58/58 | Unknown | Not measured |
| Ling 3.0 Tiny · thinking off | — | 39s | 1m 46s | 58/58 | Unknown | Not measured |
Times include thinking and writing, but not scoring or model loading managed by the runner. Some records exclude retries and retry delays. We did not control whether a model was already loaded or sharing hardware. The chart’s time axis uses uneven spacing to fit both fast and slow models; the dotted line does not predict results between points.
All models shown completed the same questions. We do not have complete records of failed attempts, so first-try success is unknown. Timing summaries give each question equal weight.
Costs cover answers, not judging or research. Estimates are not verified bills; some older estimates count only generated text, excluding input and cached text charges. Subscriptions still cost money. Local electricity and hardware costs were not measured—not zero.
Decisions
Tea Decisions v2 · base models
188 items · graded by answer key · no judges
Deterministic decision accuracy on 188 private tea decision items: binary calls, constrained choice, and anchored ordinal rating.
† Median time from one question to the answer, measured on one RTX PRO 6000 GPU answering one question at a time. These are local models with fixed hardware: there is no per-token price to show — what serving costs is the GPU memory each model needs and the wait above. ‡ Peak GPU is resident GPU memory measured on this machine while the model answered: weights, key/value cache, and runtime together. A dash marks a model served through an API.
How this decision test works
Single request per item at temperature 0; the answer must be exactly one option key. Anything else is a format failure and counts as wrong.
Every item is graded deterministically against a frozen answer key by exact key match; there is no LLM judge anywhere in this benchmark.
Eighteen paired items check that answers follow real evidence changes and resist injected instructions or irrelevant narrative; a pair counts only when both members are correct.
95% descriptive item-sampling intervals from a fixed-seed bootstrap over the 188 items; they describe sensitivity to item selection, not run-to-run variation (grading itself is deterministic).
Median wall time from request to final answer, one request at a time, on one workstation GPU (RTX PRO 6000 Blackwell). The decision-native models read their answer straight off the model's logits and generate no answer text.
Everything here runs locally on that single GPU with weights resident in memory: there are no per-token API fees to compare. What serving costs is the memory a model needs and the wait it causes. Energy use and hardware amortization are not measured and are not published as costs.
Two cohorts share the headline score: the 88 everyday items carried over from release v1 and the 100 compositional stress items added in v2. Items that every model passes add no separating evidence, so check both cohorts before trusting an ordering.
The corpus is private and freshly authored, but the knowledge-control group uses public tea facts. Ordinal rubric items require the exact rubric level; adjacent levels count wrong. The calibration group (known-probability items) is excluded from scoring until a confidence protocol exists. A single answer per item at zero temperature measures the decision, not its stability across re-prompts.
Sanitized aggregates only. The private corpus and per-item answers are not published.
Vision
Tea Vision v0.3 · vendor-catalogue imagery
155 items · graded by answer key · no judges
Accuracy on 155 image questions about real tea products from public vendor catalogues: which object is shown, which tea is sold, and whether the answer holds when the camera angle changes.
Top-down repeats the same questions on overhead close-ups of the dry tea - the hardest angle, where both models lose their stride. Rows marked — are controls: the same model without the image, to show how much of the score the picture carries.
† Peak GPU is resident GPU memory measured on this machine while the model answered: weights, key/value cache, and runtime together.
Follow-up studies on this corpus: composing several views into one numbered grid lifted Jev-Omni from 66% to 74% on 58 ware products, while naive probability averaging across views left it at 67% (its per-view confidences are near one-hot and its errors correlate across angles), and averaging a chat model's token log-probabilities collapsed accuracy to 28% - token distributions are not decision probabilities. The oracle ceiling - always picking the one view each model answers correctly - sits at 81%.
How this vision test works
One request per question: one image and four lettered options, with the correct answer's letter position permuted across items.
Every item is exact-match against an answer key: the chosen letter must match, and an answer that does not follow the required format counts wrong. No judges are involved at any step.
Answer keys come from the vendors' own catalogue taxonomy and listing claims, not from tea expertise. Origin and quality claims are reported separately and never scored.
Images are public product photography from well-known tea vendors, so models may have seen these catalogues during training. Scores describe decisioning on familiar catalogues, not sealed generalization. This is the deliberate inverse of the Tea Decisions board above, which runs on a private corpus the models have not seen.
The interval is a 95% Wilson score interval over the 155 scored items.
All models ran locally on one RTX PRO 6000 GPU, one question at a time. Jev-Omni answers with its own decision head; Qwen3-VL 30B answers in prompted chat mode, and its text-only row shows how much of its score comes from actually reading the image.
Version 0: 155 items across two model families and one control condition. Vendor taxonomy is marketing language, so a correct answer means agreement with the catalogue, not tea truth.
Sanitized aggregates only: item counts and scores, nothing from the corpus or the answer keys.
Scores by category
Where models do well
Scores out of 100, grouped by topic and type of problem.
Other tests
More AI comparisons
Privately maintained questions · scoring checked against sources · one answer per question · two AI judges · no human expert review
How scoring works & what it can tell us
Two AI models grade each answer against source-backed criteria, with disagreements reviewed against the same sources. Each of the six topics contributes equally to the overall score. Settings such as “high” or “max” in model names indicate the thinking mode used.
The initial results reuse answers and scores collected while developing the test, checked against the final question set. They are not a separate fresh test. Each model has one counted answer per question; we do not pick its best answer from repeated attempts.
The thin bars show a 95% range calculated by repeatedly resampling questions within each topic. They show how scores depend on the questions included—not how answers or judges might vary on another run, or whether one model is better in general.
Peak VRAM is the highest total GPU occupancy observed in a 2.5-second series while the model served answers on the evaluation workstation (2 GPUs, 97,887 MiB each), recorded under an exclusive GPU lease. It measures the serving state reached during the run, not a capacity claim; models served with deliberate CPU/host-RAM offload exceed per-GPU memory by design.
This is a personal comparison, not an expert-certified test. Questions are kept private, but we cannot guarantee models have never seen similar material. Some categories have more questions than others.
