One practical tip on how to talk to your AI, with receipts

I gave six reasoning models the same questions with explicit names and with pronouns, then compared their reported thinking tokens.

machine learningbenchmarksreasoninglanguage models

I wrote the same 16 small reasoning questions two ways: once repeating every object’s name, and once replacing those repetitions with words such as it, that and the latter.

Six models answered both versions. I compared the reasoning-token count reported by the API and scored answer quality separately.

That distinction has a practical consequence: reported thinking tokens are generally billed as output-token usage and can add latency. More thinking can mean more cost even when the answer does not improve.

For example, one version said “Move the amber chest to shelf 4. Put the blue key inside the amber chest.” The matched version said “Move it to shelf 4. Put the blue key inside it.” The facts, choices and correct answer stayed the same.

Design

The questions use small state-tracking, relational-reasoning, object-permutation and symbolic-planning tasks. Answers are multiple choice and automatically scored.

  • Primary set: 12 pairs whose pronouns and references were intended to be unambiguous.
  • Diagnostic set: four pairs with deliberately ambiguous references, excluded from the headline result.
  • Matching: each pair has the same facts, operations, answer choices, option order and canonical answer.
  • Calls: every version is an independent one-turn request. Odd-numbered pairs run explicit-first; even-numbered pairs run referential-first.
  • Models: GPT-5.6 Sol, Gemini 3.8 Flash, Claude Opus 5, Claude Sonnet 5, Qwen 3.8 Max and GLM 5.3 Flash.
  • Settings: temperature 0, low reasoning effort, fixed seed and a 2,048-token completion limit.

The preregistered outcome for each primary item was:

reported reasoning tokens(referential) − reported reasoning tokens(explicit)

Missing reasoning-token accounting would be treated as missing, not zero. All 192 calls returned a non-null count; reported zeroes remain legitimate zeroes.

Exact system prompt
Solve the task. Return exactly OPTION A, OPTION B, OPTION C, or OPTION D and nothing else.

Results

The preregistered comparison uses 12 unambiguous question pairs. Every model used more reported thinking tokens on average with the referential wording.

ModelProviderExplicit namesit / that / pronounsDifferenceCorrect, explicit → referential
GPT-5.6 SolOpenAI0.07.2+7.211/12 → 12/12
Gemini 3.8 FlashGoogle AI Studio75.688.1+12.512/12 → 12/12
Claude Opus 5Azure13.024.6+11.612/12 → 12/12
Claude Sonnet 5Claude on AWS9.211.8+2.512/12 → 12/12
Qwen 3.8 MaxAlibaba116.8126.9+10.212/12 → 12/12
GLM 5.3 FlashGMICloud27.380.6+53.212/12 → 12/12

Across all 72 primary model-question pairs, explicit names averaged 40.3 thinking tokens and referential wording averaged 56.5: a paired increase of 16.2 tokens.

The effect was uneven and heavy-tailed. The median paired difference was zero: 29 pairs increased, 32 tied and 11 decreased. GLM produced the largest average shift, while Sonnet changed little. Every model-level mean moved upward, but individual pairs did not consistently do so.

The referential prompts were 5.5 native input tokens shorter on average. The result therefore is not explained by the referential prompts simply containing more input tokens. Tokenization still differs by model, so all comparisons remain paired within a model and item.

Answer quality

On the unambiguous questions, there was no answer-quality penalty: explicit wording scored 71/72 and referential wording scored 72/72. The only miss was GPT-5.6 Sol choosing the wrong option on one explicit-name prompt.

Five Sonnet calls included extra text despite the response-format instruction. The raw output and strict-format violation are preserved; a unique option label was parsed separately for task accuracy. This parsing choice does not affect the reasoning-token comparison.

Ambiguity diagnostic

Ambiguity behaved differently. Across the four diagnostic pairs and six models, explicit wording averaged 64.1 reasoning tokens and ambiguous referential wording averaged 213.0, a mean paired change of +148.9. The median paired change was +10, showing how strongly a few large values affected the mean.

Answer quality also fell from 21/24 with explicit names to 13/24 with ambiguous references. One GLM response used 2,047 reasoning tokens, exhausted the 2,048-token completion cap we set and stopped before emitting an answer.

These pairs are excluded from the primary result. They illustrate why reference style and ambiguity cannot be collapsed into one claim: an unclear pronoun may increase reported computation while making the answer worse.

Run details

The harness saved the exact request, raw response, normalized result, usage, reported cost, provider and latency for every call. It was resumable and enforced an aggregate budget before each request.

One GLM request received a transient HTTP 502 and was resumed once; the failed attempt remains in the append-only log. No successful call was repeated. The complete experiment cost $0.1301: $0.0163 for the initial Gemini validation and $0.1138 for the five-model extension.

Reproduce it

Open the pronouns and thinking-tokens reproduction notebook in Google Colab

The notebook contains the frozen prompt manifest, secret-free saved measurements and the complete resumable OpenRouter runner used for the experiment. It can repeat all 192 calls, saves raw responses and usage, and enforces an aggregate budget. Paid model calls are disabled by default and require an explicit switch plus your own API key.

Scope

These are reported reasoning tokens from one run per wording, not a direct measurement of thought. The result is specific to these prompts, tokenizers, models, providers and reasoning settings; it does not show that pronouns inherently make language models think more. Provider accounting can also differ, particularly when a model reports zero reasoning tokens for a correct response.

We did not test longer prompts. The relative effect might shrink because the same absolute change is diluted by a much longer reasoning trace, stay similar, or grow if additional references require more resolution. These data cannot distinguish those possibilities.

Get new posts

New experiments and research notes, delivered by email.

Prefer a feed? Subscribe via RSS.