Source-grounded geology benchmark · August 2026
Six frontier models answer 150 questions drawn from the primary records of three mining districts — a modern NI 43-101 technical report, four USGS district studies from 1924–27, and a Western Australian WAMEX archive. Every answer is graded against a calibrated, source-anchored rubric.
Mean score per question with a bootstrap 95% confidence interval. Where intervals overlap, the benchmark cannot order the models — and it says so. The models split into tiers, and only the boundary between tiers survives testing — the exact order inside a tier is not a claim.
Per district
Leaderboard
The top tier holds its shape on every corpus. The bottom tier does not: its order depends on which district is asking the questions — which is exactly why this benchmark measures three of them instead of one.
What it cost to run each model through all 150 questions, against what it scored. Cost counts only the run that produced each surviving answer, plus judging. Models on the dotted frontier are unbeaten on both axes at once; muted models are dominated — something else is better and cheaper.
Cost per question
Every model's grade on every question, one row per question, sorted by how far the models spread. Rows at the top separate models; uniform rows at the bottom are questions every model handled identically. A red × is a failed gate — an answer wrong about the one thing the question exists to test, scored zero.
Mean score by the questions' authored difficulty band. Two districts behave: scores sag as questions get harder and the models spread further apart. The third does not — and that is a finding about the labels, reported rather than hidden.
Effort is not cost, and neither is score. Bubble area is total spend: the hardest-working model is mid-table, the biggest spender searched least, and the runner-up got there on the fewest moves of the leading three.
Tool mix — share of each model's calls
Tool calls are counted from each question's saved transcript — only the run that produced the surviving answer. Two models occasionally delegate to a sub-agent (task), whose internal searches are not counted, so their effort is slightly understated.
A benchmark is only as good as the smallest difference it can detect. Each district is summarised by its top-to-bottom range, the minimum gap distinguishable from noise at fifty questions, and how many distinct levels fit between the two.
Resolution — how finely each district separates models
Generation (the model exploring the corpus and writing) dwarfs judging, and varies six-fold between models doing the same task. The strip below shows every question's cost: some models spend evenly, others are dominated by a few runaway explorations.
Every question's generation cost (log scale)