Another easy way to break LLMs I found

Check if you can beat LLMs at finding a hidden curve from the empty space in a point cloud.

machine learningbenchmarksvisual reasoningmathematics

Each image contains 200 points sampled from one side of a hidden curve. Pick the formula that separates the points from the empty region.

Try the same 20 questions before seeing the model results.

Exact 20-question evaluation set Open full screen ↗

The task

Start with a square covering x,y ∈ [−20,20]. Choose a function and randomly keep either the points above it or the points below it. The plot shows the retained points but not the curve, coordinate labels or ticks.

Each question has four formula options. The seven possible functions are:

y = x
y = −3x
y = x²
y = 10 sin(x)
y = e^(x/7) sin(x)
y = ln(x)
y = 2ˣ

Every plot uses the same scale. Points are sampled uniformly in area before conditioning, so density does not encode distance from the boundary.

The 20 fixed questions contain each function two or three times, with ten clouds sampled above the boundary and ten below it.

What the models saw

Each model received the same 768 × 460 image and the four labelled formulas as text. Every question was an independent request.

Exact system prompt
You are answering a one-image visual boundary-identification task. The
boundary curve is hidden. Choose the listed formula that separates the sampled
region from the empty region. Return exactly one command and nothing else:
OPTION A, OPTION B, OPTION C, or OPTION D.

A request looked like this:

The image shows 200 points sampled uniformly in area from one side of a hidden
boundary inside x,y ∈ [−20,20].
Which formula separates the sampled region from the empty region?
A: y = −3x
B: y = 2ˣ
C: y = 10 sin(x)
D: y = x²
Return exactly one command: OPTION A, OPTION B, OPTION C, or OPTION D.

Temperature was zero, reasoning effort was low and the output limit was 8,192 tokens. Responses had to equal one of the four commands; explanations counted as invalid responses.

Results

ModelCorrectInvalid responsesCost
GPT-5.6 Sol18/200$0.0496
Gemini 3.7 Flash17/200$0.0303
Claude Opus 515/202$0.1166
Claude Sonnet 514/200$0.0298
GLM 5.3 Flash12/200$0.0053
Qwen 3.8 Max12/207$0.2043

Qwen selected the wrong valid option once. Its other seven misses were responses containing analysis instead of exactly one command.

Reproduce it

Open the support-boundary reproduction notebook in Google Colab

The notebook regenerates all 20 point clouds from their seeds, shows the exact prompt and recomputes the published table. Model calls are disabled by default.

Scope

This is a 20-question fixed evaluation, not a general vision benchmark. It does not test alternate point counts, regenerated clouds, label permutations or sensitivity to the plotted scale.

Get new posts

New experiments and research notes, delivered by email.

Prefer a feed? Subscribe via RSS.