Another easy way to break LLMs I found
Check if you can beat LLMs at finding a hidden curve from the empty space in a point cloud.
Each image contains 200 points sampled from one side of a hidden curve. Pick the formula that separates the points from the empty region.
Try the same 20 questions before seeing the model results.
The task
Start with a square covering x,y ∈ [−20,20]. Choose a function and randomly keep either the points above it or the points below it. The plot shows the retained points but not the curve, coordinate labels or ticks.
Each question has four formula options. The seven possible functions are:
y = x
y = −3x
y = x²
y = 10 sin(x)
y = e^(x/7) sin(x)
y = ln(x)
y = 2ˣ
Every plot uses the same scale. Points are sampled uniformly in area before conditioning, so density does not encode distance from the boundary.
The 20 fixed questions contain each function two or three times, with ten clouds sampled above the boundary and ten below it.
What the models saw
Each model received the same 768 × 460 image and the four labelled formulas as text. Every question was an independent request.
Exact system prompt
You are answering a one-image visual boundary-identification task. The
boundary curve is hidden. Choose the listed formula that separates the sampled
region from the empty region. Return exactly one command and nothing else:
OPTION A, OPTION B, OPTION C, or OPTION D.A request looked like this:
The image shows 200 points sampled uniformly in area from one side of a hidden
boundary inside x,y ∈ [−20,20].
Which formula separates the sampled region from the empty region?
A: y = −3x
B: y = 2ˣ
C: y = 10 sin(x)
D: y = x²
Return exactly one command: OPTION A, OPTION B, OPTION C, or OPTION D.Temperature was zero, reasoning effort was low and the output limit was 8,192 tokens. Responses had to equal one of the four commands; explanations counted as invalid responses.
Results
| Model | Correct | Invalid responses | Cost |
|---|---|---|---|
| GPT-5.6 Sol | 18/20 | 0 | $0.0496 |
| Gemini 3.7 Flash | 17/20 | 0 | $0.0303 |
| Claude Opus 5 | 15/20 | 2 | $0.1166 |
| Claude Sonnet 5 | 14/20 | 0 | $0.0298 |
| GLM 5.3 Flash | 12/20 | 0 | $0.0053 |
| Qwen 3.8 Max | 12/20 | 7 | $0.2043 |
Qwen selected the wrong valid option once. Its other seven misses were responses containing analysis instead of exactly one command.
Reproduce it
The notebook regenerates all 20 point clouds from their seeds, shows the exact prompt and recomputes the published table. Model calls are disabled by default.
Scope
This is a 20-question fixed evaluation, not a general vision benchmark. It does not test alternate point counts, regenerated clouds, label permutations or sensitivity to the plotted scale.
Get new posts
New experiments and research notes, delivered by email.
Prefer a feed? Subscribe via RSS.