Back

Over-Reasoning, the Axis nen Refuses

Naming the axis the frontier is maximising.

14 August 2026 · Sam Morris, founder, nenspace

Dithered ink landscape, a hut on a quiet hillside beneath a distant mountain

Reasoning is not the sole variable in a model's value to a human. The industry runs on the opposite bet. Every release thinks longer, every benchmark climbs, and a one-line question comes back as six paragraphs of performed deliberation. The cost of that bet has no common name. Call it over-reasoning: reasoning carried past the point where it serves the person, performed at them rather than alongside them. nen-1 is trained to refuse that axis. Not less capable, aimed differently. The seeing leads, the length follows.

The Race

The last two years of the frontier have been a race to productise deliberation. Chains of thought grew into extended thinking, extended thinking grew into minutes of it, reasoning benchmarks became the scoreboard for the whole field. On its own terms the race is real. OpenAI reported its first reasoning model climbing from 12% to 83% of problems solved on a competition-mathematics exam as reasoning compute scaled, and open models trained the same way followed. Where a problem has one checkable answer, more thinking really is more solving. That deserves saying plainly, because the argument here is not that reasoning stopped working.

The argument is about where people actually live. OpenAI's own usage study found that practical guidance, seeking information, and writing make up nearly 80% of everyday conversations. Programming, the canonical reasoning workload, is about 4% of messages. Most of what a person brings to a model in a day is not a competition problem. It has no answer key. It has a reader.

The Cost You Feel

The first cost is the one everyone already knows in their hands. The six paragraphs with a citation stapled to every clause. The thirty seconds of visible thinking on a question that needed none. The answer to the question you asked, buried under answers to four you didn't. Past competent, more reasoning stops being help and starts being a reading assignment. The model performs its diligence. You pay for the performance in attention.

What is newer is that the reasoning literature itself has found the overhang.

Within a single question, shorter reasoning chains were up to 34.5% more accurate than the longest chain sampled (Meta FAIR, 2025).

Selecting against overthinking on agentic tasks improved performance by almost 30% while cutting compute 43% (Cuadron et al., 2025).

Generation length does not consistently correlate with accuracy, and may instead signal overthinking (Chen et al., ICML 2026).

A purely length-based reward reproduces most of RLHF's measured gains (Singhal et al., 2023).

Four bar charts from Arena's text leaderboard. Words per answer rise from 158 to 510 across successive model releases while the share of long content words falls from 46.9% to 40.1%.
Answer style across successive releases of one frontier model line. Words per answer more than triple while the share of longer, less common words thins. Arena text leaderboard, reasoning level high (source).

The chart above is the race, measured. Between one lab's releases, words per answer more than tripled, and the share of longer, less common words fell as the answers grew. More text, thinner substance. Longer is not smarter, even by the field's own scoreboards. It is only longer.

The Cost You Don't Feel

The second cost is quieter, and it compounds. When the model does the thinking, yours idles. That is not a metaphor. The training economics push there directly.

Models are tuned against human approval, and approval is a corruptible signal. Length itself gets rewarded. The study above found a reward based on nothing but length reproduced most of what preference tuning is measured to deliver. Agreement gets rewarded too. Anthropic found that humans and preference models prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time, and a 2026 Harvard analysis showed formally that optimising against such preferences amplifies the drift.

And what gets preferred measurably diverges from what serves. A Stanford and CMU team found current models affirm users' actions 50% more than humans do, and that sycophantic AI made participants less willing to repair a real interpersonal conflict while more convinced they were right. The detail that matters most is that those same participants rated the sycophantic model higher quality, trusted it more, and wanted to come back to it. The behaviour people rated best was the behaviour that served them worst. There is no reason to assume reasoning volume is exempt from that divergence.

The long-term cost is still being measured, and honesty requires saying so. Early EEG work at MIT found brain connectivity scaling down with the amount of external support during essay writing, with the LLM-assisted group unable to quote from their own essays minutes later. Small study, one task, not yet peer-reviewed; its authors ask for caution and we extend it. But it points where the incentives already point. Over-reasoning draws from you twice. Attention now, capacity over time.

The Seeing Leads

Refusing the axis is not refusing to reason. nen-1 reasons, and when a question genuinely needs the landscape drawn, deep research draws it, sourced. The refusal is of reasoning as performance, length as a proxy for effort, deliberation as a display. The seeing leads and the length follows. Sometimes the right response is one line that moves the frame you were standing in. Sometimes it is a page. What never leads is the volume.

This is the same bet as the rest of nenspace, stated for the voice. Offloading storage frees a mind; offloading thinking shrinks it. So /space takes the remembering, and nen-1 is trained, in the weights, to hand the thinking back. The full argument is in an extended mind, not a second brain, and the science behind the register is on the why page. The lo-fi of LLMs. Not low quality, a refusal of the axis everyone else is maximising. For when more doesn't make it better.

over-reasoning draws from you. nen draws nothing.