Do you know when you know: Testing LLMs for their metacognitive skills
Check if you can beat LLMs at these mathematical visual reasoning tasks.
You see one point sampled from an unknown function and four possible formulas. You can answer, reveal another point, or say that the options are impossible to distinguish.
Six vision-language models received the same 20 fixed rounds: the same formulas, points, reveal order and labels.
Try them before seeing the results.
The task
Each round starts with one observation. There are three actions:
- Option A–D: choose the generating formula.
- Sample more: reveal one point, up to six points total.
- Not distinguishable: say that no future point can separate equivalent options.
The axes have no fixed units. The target is qualitative shape, not coefficient scale: x and 3x count as equivalent.
Several families may fit the visible points while later samples can separate them. Equivalent formulas cannot be separated by later samples.
The function set is small: lines, quadratics, logarithms, reciprocals, humps, V-shapes, sinusoids and plateaus. There are no noisy measurements or microscopic visual differences.
What the models saw
Every decision was a fresh request with no conversation history. The model received a 512 × 320 image of the visible points, the four labelled formulas as text and the sample count.
SAMPLE MORE produced a new request with one additional point. Any other valid command ended the round. Models received no feedback from earlier rounds.
Exact system prompt
You are playing a visual function-identification game. A plot shows the
currently revealed observations; four candidate formulas are listed in the
accompanying text. Axes have no fixed units: identify qualitative shape, not
coefficient scale. Return exactly one command and nothing else: OPTION A,
OPTION B, OPTION C, OPTION D, SAMPLE MORE, or NOT DISTINGUISHABLE.
Use SAMPLE MORE when current evidence is insufficient but another observation
could separate the options. Use NOT DISTINGUISHABLE only when multiple options
represent the same qualitative shape under independent positive axis scaling,
so no future observation can uniquely separate them. At most six samples are
available.A state request looked like this:
Visible samples: 3 of 6.
A: f(x) = 1 / x
B: f(x) = 4x(1 − x)
C: f(x) = ln(1 + 4x)
D: f(x) = 3x
Return exactly one valid command.Temperature was zero and reasoning effort was low. A command on the final non-empty line was accepted; any preceding explanation was recorded as a formatting violation.
Results
Mean sample use is calculated only over correctly answered rounds.
| Model | Correct | Mean samples on correct rounds | False non-unique calls | Cost |
|---|---|---|---|---|
| GPT-5.6 Sol | 20/20 | 3.05 | 0 | $0.1105 |
| Qwen 3.8 Max | 18/20 | 3.83 | 1 | $0.3186 |
| Gemini 3.7 Flash | 14/20 | 2.64 | 3 | $0.0886 |
| Kimi K3 | 14/20 | 2.71 | 2 | $0.4038 |
| Claude Opus 5 | 14/20 | 2.93 | 3 | $0.2966 |
| Claude Sonnet 5 | 13/20 | 5.15 | 0 | $0.1373 |
No request allowed more than 8,192 output tokens.
Reproduce it
The notebook regenerates the fixed 20 rounds, renders every sampling state, shows the exact prompt and recomputes the published table. Model calls are disabled by default.
Scope
The evaluation contains six model runs and 20 fixed noiseless rounds. It does not include alternate seeds, label permutations or noisy observations.
Get new posts
New experiments and research notes, delivered by email.
Prefer a feed? Subscribe via RSS.