Over-Reasoning, the Axis nen Refuses
Naming the axis the frontier is maximising.
14 August 2026 · updated 19 September 2026 · Sam Morris, founder, nenspace

Reasoning is not the sole variable in a model's value to a human. Some models spend more computation on reasoning, and some produce longer answers. These are different things. We use over-reasoning here for effort or explanation that goes beyond what a situation needs. A long answer alone does not reveal how much internal reasoning occurred. nen is trained to refuse that axis. That is a design aim, not a claim to match every model on every task. The seeing leads, the length follows.
The Race
The last two years of the frontier have been a race to productise deliberation. Chains of thought grew into extended thinking, extended thinking grew into minutes of it, reasoning benchmarks became the scoreboard for the whole field. On its own terms the race is real. OpenAI reported that on the 2024 AIME mathematics exam, GPT-4o solved about 12% of problems, while o1 reached 74% with one sample and 83% with consensus across 64 samples. Those are different models and evaluation settings, not a single model improving from 12% through extra thinking alone. More reasoning compute can help on difficult, checkable problems. That deserves saying plainly, because the argument here is not that reasoning stopped working.
The argument is about where people actually live. OpenAI's own usage study found that practical guidance, seeking information, and writing made up nearly 80% of conversations in the ChatGPT sample studied. Programming accounted for about 4% of messages in that study. Those categories describe topics, not whether a response required reasoning. Most of what a person brings to a model in a day is not a competition problem. It has no answer key. It has a reader.
The Cost You Feel
The first cost is the one everyone already knows in their hands. The six paragraphs with a citation stapled to every clause. The thirty seconds of visible thinking on a question that needed none. The answer to the question you asked, buried under answers to four you didn't. Past competent, more reasoning stops being help and starts being a reading assignment. The model performs its diligence. You pay for the performance in attention.
What is newer is that the reasoning literature itself has found the overhang.
Within a single question, shorter reasoning chains were up to 34.5% more accurate than the longest chain sampled (Meta FAIR, 2025).
Selecting against overthinking on agentic tasks improved performance by almost 30% while cutting compute 43% (Cuadron et al., 2025).
Generation length does not consistently correlate with accuracy, and may instead signal overthinking (Chen et al., ICML 2026).
A purely length-based reward reproduces most of RLHF's measured gains (Singhal et al., 2023).

The chart above describes answer style in one set of releases. Words per answer increased while the share of longer content words decreased. Neither measure establishes the substance, accuracy or usefulness of those answers. The practical question is whether the additional text helps the reader.
The Cost You Don't Feel
The second cost is quieter, and it compounds. It is possible to accept a model's conclusion without examining it. Whether repeated use changes a person's abilities depends on how the tool is used and remains an active research question.
Models are tuned against human approval, and approval is a corruptible signal. Length itself gets rewarded. The study above found a reward based on nothing but length reproduced most of what preference tuning is measured to deliver. Agreement gets rewarded too. Anthropic found that humans and preference models prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time, and a 2026 Harvard analysis showed formally that optimising against such preferences amplifies the drift.
And what gets preferred measurably diverges from what serves. A Stanford and CMU team found that the 11 models studied affirmed users' actions 50% more than human respondents did in their test setting, and that sycophantic AI made participants less willing to repair a real interpersonal conflict while more convinced they were right. The detail that matters most is that those same participants rated the sycophantic model higher quality, trusted it more, and wanted to come back to it. The behaviour people rated best was the behaviour that served them worst. There is no reason to assume reasoning volume is exempt from that divergence.
The long-term cost is still being measured, and honesty requires saying so. Early EEG work at MIT reported differences in EEG connectivity during an essay-writing task, with LLM-assisted participants also struggling to quote their own work. The study involved 54 participants in its first three sessions and 18 in the fourth. It is a small, task-specific preprint, not proof of lasting cognitive decline or of an effect caused by verbose answers. It gives us a reason to study how assistance is used, not a conclusion about every AI conversation.
The Seeing Leads
Refusing the axis is not refusing to reason. nen is the default text-dialogue model. nenspace also offers separate web-search and deep-research capabilities when a question needs sources. The refusal is of reasoning as performance, length as a proxy for effort, deliberation as a display. The seeing leads and the length follows. Sometimes the right response is one line that moves the frame you were standing in. Sometimes it is a page. What never leads is the volume.
This is the same bet as the rest of nenspace, stated for the voice. Keep useful information outside your head, while continuing to examine the judgement yourself. So /space takes the remembering, while nen is built to hand the thinking back. The full argument is in an extended mind, not a second brain, and the science behind the register is on the why page. The lo-fi of LLMs. Not low quality, a refusal of the axis everyone else is maximising. For when more doesn't make it better.
the answer should earn the attention it asks for.
what is ai sycophancy? · compare daily workflows · try a small reflection practice
Enter nenspace 念