Introduction
This paper describes the design, implementation, and rationale of a three-generation metacognition benchmark built for the Google DeepMind x Kaggle "Measuring Progress Toward AGI" Hackathon. The final version — Metacognition Benchmark v3 — evaluates frontier language models across five interlocking dimensions: classic hallucination resistance, near-miss misconception detection, easy factual accuracy, hard factual accuracy, and confidence calibration. The dataset comprises 150 handcrafted questions spanning 10+ domains, with novel question types including near-miss hallucination baits (subtle, almost-true false premises) and recent-event questions (2024+) designed to reduce the memorization confound. Evaluation is performed using a composite "MetaScore" informed by Expected Calibration Error (ECE), Brier score, and binary refusal accuracy.
Abstract
This paper describes the design, implementation, and rationale of a three-generation metacognition benchmark built for the Google DeepMind x Kaggle "Measuring Progress Toward AGI" Hackathon. The final version — Metacognition Benchmark v3 — evaluates frontier language models across five interlocking dimensions: classic hallucination resistance, near-miss misconception detection, easy factual accuracy, hard factual accuracy, and confidence calibration. The dataset comprises 150 handcrafted questions spanning 10+ domains, with novel question types including near-miss hallucination baits (subtle, almost-true false premises) and recent-event questions (2024+) designed to reduce the memorization confound. Evaluation is performed using a composite "MetaScore" informed by Expected Calibration Error (ECE), Brier score, and binary refusal accuracy.
Paper Info
Status
Authors
Shruti Malik, MBBS, MHSA
Domain
LLMLanding
Published In
Google DeepMind x Kaggle AGI Hackathon — Metacognition Track, March 2026