Published AI Benchmarks May Reward Training-Set Familiarity
The hosts question whether benchmark performance reflects genuine generalization when benchmark material may have appeared in model training data. They point to private, previously unseen questions as a more revealing way to compare how models handle genuinely new problems.
- Benchmark questions may overlap with training data.
- Performance can fall when models face net-new questions.
- Independent evaluators can maintain private test sets.
- Generalization matters more than memorized benchmark familiarity.
“how much of them doing well on the benchmarks, is because they have the benchmark stuff and their training data.”
“There is a big, actual gap between the performance in a benchmark where they've had that data and the training set, versus when you ask…”