🎓 Ages 15–18 · Grades 9–12 · Scientific Method & Evaluation
BenchForge
Every leaderboard, model card, and study is an argument — and the person who designed the test quietly decided who wins. Read a real evaluation, predict what the numbers will do, then watch a deterministic engine work it out. Learn to read the test, not just the score.
Start here
Does the score hold up? A model reports a great score — but some test problems leaked in from its training. Predict how far the honest score drops once you remove them, then see the arithmetic. Who wins the hard cases? Two models tie on the average. Predict what happens on the hardest, most-novel cases — the ones the benchmark exists for — then let the engine stratify and count. Which metric wins? One model leads by one yardstick. Predict whether the winner changes when you switch to a different, equally-reasonable metric — then rank them both ways. Concept practice 16 kits from “what an evaluation IS” to leakage, hold-out splits, stratification, the sealed envelope, metric choice, and gaming a benchmark.
Concept kits — 16 kits of evaluation & scientific-method reasoning
Kit 1: What an evaluation IS 25 questions Kit 2: Train/test split 25 questions Kit 3: Data leakage 25 questions Kit 4: Hold-out and cross-validation 25 questions Kit 5: Stratifying by difficulty and novelty 25 questions Kit 6: The hardest bin 25 questions Kit 7: Prospective vs retrospective (the sealed envelope) 25 questions Kit 8: Blind assessment vs self-report 25 questions Kit 9: Who grades the grader 25 questions Kit 10: Choosing a metric 25 questions Kit 11: When the metric changes the winner 25 questions Kit 12: Overfitting to a benchmark (gaming) 25 questions Kit 13: Patching a gamed test 25 questions Kit 14: Confidence intervals on a score 25 questions Kit 15: Real cases (blind assessment; CASP) 25 questions Kit 16: Capstone — design a fair evaluation 25 questions
Meet the cast — then teach one
- The Referee — runs the neutral, blind, prospective test (the mentor)
- Splitter — separates train from test — and hunts leakage between them
- Binner — stratifies cases by novelty so the hard band shows
- Metric — picks the yardstick — and shows the winner change when it swaps