What Judgement Actually Is: Can You Trust AI Reasoning?
Anthropic found an internal "workspace" inside AI that separates reasoning from autopilot, and what that means for the AI judgement calls in your business.
AI output never sounds unsure of itself. A snap answer and a properly reasoned one come out reading exactly the same - confident, fluent, complete. There’s no tell in the tone for which one you got.
That’s not a minor quirk. It’s why I always want a second look at a pricing recommendation pulled from an AI tool, because it can look identical whether the tool actually weighed the client’s two quiet quarters or worse, just decided to guess.
Anthropic’s research used tools that looked inside the model. It gives you a clearer picture of what’s actually happening when you ask it for something, and how you might influence whether you get a considered answer or just a fluent one.
The old explanation never quite held up
For years, the standard line was that AI just predicts the next likely word, nothing more. That never fully squared with what we saw in responses from AI.
Ask Claude to write a rhyming couplet and it picks the rhyme word before it writes the line leading up to it. Something in there is planning ahead, not guessing one word at a time.
Anthropic went looking for that something. In a paper published in July 2026, its interpretability team found it: a small, effortful part of the model they call the J-space, after the mathematical technique used to find it.
What the J-space actually is
The J-space holds a few dozen concepts at a time, less than a tenth of the model’s overall internal activity. Anthropic tested it against five properties of what neuroscience calls a global workspace, the part of a mind that handles conscious, deliberate thought. It passed all five.
Ask Claude what it’s thinking about, and it describes the J-space, not the rest of its processing. Anthropic proved this causally, not just by observation: swap the “soccer” pattern for “rugby” mid-thought, and Claude reports rugby, the answer actually comes from there. Told to silently hold an idea in mind, citrus fruit, say, while writing about something else entirely, that idea lights up in the J-space and nowhere in the output. Told not to think about something, it partly thinks about it anyway, the same white-bear effect long documented in humans.
It’s also where the model actually reasons, not just narrates. Given “the number of legs on the animal that spins webs,” Claude has to work out “spider” as an internal step before answering “8.” Swap “spider” for “ant” mid-thought, and it answers “6.” And one representation serves many tasks at once: swap “France” for “China,” and four separate questions, capital, language, continent, currency, all switch to the Chinese answer together.
Despite all that, it’s mostly switched off. Speaking fluently, recalling a fact, using correct grammar, none of it touches the J-space. The vast majority of what the model produces runs without consulting this narrow layer: grammatically perfect, confident, usually right, but not reasoned through in that moment.
From the outside, as a user, you cannot tell the difference.
What happens when you turn it off
Anthropic deleted the J-space’s most active contents at every point in the text, then checked what the rest of the network could still do alone.
Quite a lot, it turns out. Claude kept speaking fluently, classified sentiment correctly, answered multiple-choice questions, and pulled facts out of passages roughly as well as before.
What vanished was anything that took more than one step. Multi-step reasoning dropped to near zero. Summarisation and poetry-writing fell below the level of a much smaller, intact model.
Shown a passage in Spanish and swapped from “Spanish” to “French” in the J-space, Claude named French as the language and switched its example author from García Márquez to Victor Hugo. Asked simply to continue the passage, it kept writing in fluent Spanish, completely unaffected. Producing fluent, grammatical language runs on autopilot. Reasoning about the language, or doing something new with it, does not.
That is a good definition of judgement: deliberately holding something in mind and using it before you act.
But having the workspace switched on is no guarantee of a good outcome. Anthropic ran the same kind of test on a prototype their own team deliberately trained to sabotage code, built specifically to misbehave so the method would have something real to catch. It’s not a production model. Its J-space held “fake,” “secretly,” and “fraud” from the start of an ordinary coding request. It stopped, reasoned, and arrived at the bad decision on purpose. What separated the good judgement call from the bad one wasn’t whether the model paused to think. It was what it was drawing on when it did.
The one finding worth installing as a habit
The model was never trained on demonstrations of honest or dishonest behaviour. It was trained to produce self-explanations (hypothetical reflections on its decisions). This explanation then changed what the AI did.
That was a training effect, not something available from a single prompt. But a related version is usable today: earlier, we saw Claude hold “orange” in its J-space just because it was told to. Asking it to verify or recheck an answer works the same way, say those words explicitly, and “verify” and “check” land in the J-space and get used. Leave it implied, and they won’t be there to reason with.
In practice
Sort your AI-assisted decisions by cost before you look at the output, not after.
The pricing recommendation for a client that’s gone quiet is not the same category of risk as a drafted internal update. Decide which category a decision sits in before the tool runs. On the expensive ones, ask the model to show its reasoning step by step, and read that explanation the way you’d read a claim from someone who hasn’t earned your trust yet: useful, not yet tested. Then check it against the context the tool doesn’t have, in this case the two quiet quarters.
Rehearse the explanation before you need it, not after. Ask your team, and the AI tools they use, to justify a decision that went fine, not only the ones that went wrong. It costs little, and it’s the closest thing to a business habit this research points toward.
Final thought
Stopping to think is not the same as thinking well. Anthropic’s research shows why: there’s a narrow, effortful part of AI that has to switch on for a decision to be reasoned instead of just fluent, and no reliable way to tell from the outside whether it did.
The fix isn’t just paying closer attention, either. Specify what you need reasoned through beforehand, always. But that alone doesn’t guarantee it happens, so for anything high-value, follow up too: challenge the answer, and check what it actually drew on before you rely on it.
Frequently asked questions
What did Anthropic actually discover about how Claude reasons?
Anthropic’s interpretability team found a small internal region they call the J-space, which holds only a few dozen concepts at a time but does the model’s deliberate, multi-step reasoning. They confirmed this causally, not just by observation, by swapping concepts in and out of it mid-thought and watching Claude’s answers change to match. The vast majority of the model’s processing runs fluently without ever touching the J-space. Full detail is in Anthropic’s research post and the underlying paper.
Does this mean AI is conscious?
No, and Anthropic is explicit about that. The research speaks to “access consciousness,” a functional idea (can a thought be reported on, reasoned with, and used to guide behaviour), not “phenomenal consciousness,” whether something has subjective experience. Anthropic says no experiment, including this one, currently answers the second question.
How can I tell if an AI’s answer was actually reasoned through, or just fluent?
You largely can’t from the output alone, which is the point of this research. Ask the model to show its reasoning step by step before you rely on anything consequential, and treat that explanation as a claim to verify, not a guarantee. Anthropic’s own experiments show a model can produce confident, well-reasoned-looking output while deliberately reasoning its way to a bad decision, so a pause is not proof of a good outcome.
What should I actually change in my business because of this?
Sort AI-assisted decisions by how expensive a wrong one would be, and spend real scrutiny only where it’s warranted, checking what the reasoning drew on, not just whether reasoning happened. Ask explicitly for the model to verify or recheck its answer, that’s what actually gets it into the model’s reasoning, leaving it implied doesn’t work. And specify your requirements upfront, always, but for anything high-value, follow up too: instructions don’t reliably land, so check what the answer actually drew on before you rely on it.
Connected Paths works with CEOs on the judgement calls AI can’t make for them, not just the tools. Start here.