One panel contains 120 independent uniform points. In the other, each point was moved toward
neighbours within radius 0.07 for 12 steps, at strength 0.05 per step. This creates more close
pairs, denser clumps, and larger gaps.
Divide each panel into equal cells and count the points per cell. Pure chance has variance approximately
equal to its mean. Local attraction raises that ratio by producing more empty and crowded cells.
In the previous version,
the non-random panel was made more evenly spaced. The original six models scored 17–20/20;
GLM 5.3 Flash scored 16/20. Here, the non-random panel is made more clumped. The best model scored
13/20; Kimi and Qwen scored 0/20 and
reversed their answer on every mirrored pair.
The overfitting to the known example from Clarke is very likely the reason why the models not just fail
by doing random guesses but fail into the worse-than-random territory, they think that presence of
“clumps” is the sign of randomness, start looking for it and find more clumps in the clumped image.
Count test: R. D. Clarke,
“An Application of the Poisson Distribution”
(1946).
The notebook contains the attraction generator, all 20 mirrored trials, the exact model prompt,
scoring code, and an opt-in OpenRouter rerun.