Part of Seven Habits of the AI World — a deep dive off habit one, learning how to learn.
Promise kept: this is the piece Everyone Is Asking the Wrong Question About AI closed by promising. Here it is.
How many times has your gut been right?
Not the times you wish you had listened — the times you actually had the feeling, and later found out whether it was correct.
I have been asking people this for years. Strangers, colleagues, people at dinners. The answer comes back remarkably consistently: north of ninety per cent. And not from one or two people. From something like nine out of ten of everyone I ask.
That is a strange number to just walk past. Ninety per cent is not a hunch. If a model gave me ninety per cent on a hard classification task, I would ship it.
So either a lot of people are flattering themselves about their own track record, or something real is happening and we have a bad name for it. This piece is about what that thing actually is, and about a rule that follows directly from it, one short enough to keep in your head:
Never delegate judgment.
The three ways of working with AI, and the one that has no name
That number — the ninety per cent — is what you are risking every time you take a model’s answer without running it through your own filter first. So before going further into what the filter actually is, it is worth being precise about how people are skipping it.
Ethan Mollick describes two well-known modes in Co-Intelligence. The centaur divides the labour on a clean line: you do this part, the machine does that part. The cyborg intertwines — you move back and forth continuously, sentence by sentence, until the seam between your work and the model’s is hard to find.
Both are real. Both work. I have watched excellent people operate in each.
There is a third mode, and it has no name, so I will give it one.
Abdication. Not dividing the work. Not intertwining with it. Handing it over and looking away. The centaur and the cyborg both get a noun. This one gets the ghost — present in name, absent in substance.
Why it feels like winning
Abdication does not feel like a failure mode. That is the whole problem.
It feels like leverage. It looks identical to being fast. The output arrives, the ticket closes, the deck goes out, and nothing in the moment tells you which of the three modes you were in. There is no error message for having stopped thinking.
And it compounds in the wrong direction. Every time you accept an answer without checking it, the cost of checking the next one goes up, because you have less context than you would have had. The gap between what you have shipped and what you understand widens quietly, and it widens fastest in exactly the people who adopted AI earliest and hardest.
Then one day the model is down, or rate-limited, or confidently wrong in a way that matters, and you discover the thing that was actually true the whole time: they did not get faster.
They got hollow.
The distinction that actually matters
The usual debate is about volume. How much AI should you use? Should juniors use less? Is it making us worse?
Volume is the wrong axis. I know people who use AI for almost everything and are getting sharper, and people who use it lightly and are getting duller. The dial is not the variable.
The variable is which part you hand over.
Delegate the search. Delegate the boilerplate. Delegate the first ten wrong answers — that is the best possible use of a machine that does not get tired or embarrassed. Delegate the tedium, the formatting, the scaffolding, the lookup you have done four hundred times.
Keep the deciding. What to build. Which tradeoff you are accepting and what it costs. What “done” means. Why this design and not the other one. Whether the thing in front of you is actually right, or merely fluent.
That is the whole rule: never delegate judgment. What it does not tell you yet is why your gut gets to ninety per cent in the first place, or what “judgment” is actually made of. That is worth answering properly, because the answer is what makes the rule more than a preference.
But what IS judgment, actually?
I have been using the word as if it were obvious. It is not, and unpacking it is the reason the number at the top of this piece is real and not self-flattery.
I am not a neuroscientist. What follows is a plausible story assembled from what I know, mostly from building machines rather than studying brains.
Start with the problem your body actually has
Right now your senses are producing an enormous volume of data. Photons, air pressure, skin temperature at every point, joint positions, chemical traces in the air.
If all of it reached your awareness you would not be overwhelmed. You would be useless. You could not hold a thought long enough to finish it.
So there is a filter, and you can catch it working. You were not aware of your shirt against your shoulders until you read that. Or the hum of the room. Or the floor against your feet. All of it was sensed. Almost none of it was reported.
Your awareness is not receiving the world. It is receiving a heavily edited summary, and something else chose the cuts. The interesting question is what happened to everything that got cut.
The architecture, in the only vocabulary I actually speak
A metaphor, not anatomy — but one that keeps predicting things correctly, which usually means there is something under it.
The amygdala is a tiny dedicated chip. One small network, very few outputs — fight, flight, freeze, and by some accounts a fourth, appease. It is the fastest thing you have because it is the smallest and does almost nothing. It does not deliberate. It fires.
The cortex is the CPU with the memory attached — short-term, long-term, and enough processing to run the filter above. It decides what is worth escalating.
The executive brain is the GPU. Massively parallel, the part we proudly call thinking. Fast, expensive, and — critically — it only ever sees what the filter passed it.
And then there is the gut. It has its own nervous system: the enteric nervous system, between one hundred and six hundred million neurons, enough to run digestion without asking your head. That is why “second brain” stuck.
But the number that stops me is not the neurons. Picture every blade of grass in a football field — roughly forty million. Now picture a million football fields. That is the order of magnitude of the microbial population in your gut: tens of trillions of organisms, chemically active, in constant conversation with your nervous and immune systems.
In machine terms: enormous memory, enormous parallelism, a very slow clock. Not a processor for arithmetic. Exactly the processor for holding a vast, slowly-updated model of what normal feels like.
The jungle
You are walking somewhere wild and you become certain — not suspicious, certain — that something is behind you. Your neck knows before you turn. Ask what evidence you had and you would say I just know.
But consider what was available in the seconds before. The birds stopped. A different set moved, all one direction. A troop of monkeys that was ignoring you has gone quiet and is looking past you. Insect noise dropped out of one quadrant. There is a smell that was not there ninety seconds ago.
Every one of those reached you. Your eyes and ears and nose captured all of it. It simply did not clear the filter, because individually not one of them is worth interrupting you for.
Collectively they are one pattern, the pattern has a name, and something matched it against a very old library and returned a single bit: wrong.
That is what a gut feeling is. Not the absence of reasoning — reasoning you did not get the transcript of. The conclusion cleared the filter; the evidence did not. Which is exactly why it arrives as a feeling instead of an argument, and why you cannot defend it when challenged. You never had access to the inputs.
None of that pattern-matching happens in the gut itself. The enteric nervous system does not think, feel, or generate conscious thought — it runs your digestive tract with remarkable autonomy, not bird-watching. What is actually doing the computing sits one level up: Antonio Damasio’s somatic marker hypothesis, which holds that bodily states guide decision-making nonconsciously. Your body registers a response before your conscious mind produces a reason, and that response biases the choice you are about to make.
The usual evidence is the Iowa Gambling Task. Participants draw from four decks; two are quietly rigged to lose. Healthy participants start avoiding the bad decks before they can tell you which decks are bad. The body’s response comes first; the explanation later, if at all. Patients with damage to the ventromedial prefrontal cortex — the region that integrates those bodily signals — keep choosing the losing decks, reasoning fine and losing anyway.
The processing happens in the brain, below awareness. The gut is part of the loop and part of how the result is delivered — more the display than the processor. There is a documented mechanism by which you know things you cannot justify, and it outperforms the part of you that talks.
The oldest version of a method for this is 2,400 years old
None of the above is new, and going back to where it is oldest is what turned it from an observation into something you can actually practice.
Socrates did not write. What we have is mostly Plato, and in the dialogue Theaetetus Socrates explains what he thinks he is actually doing when he talks to people. His mother, Phaenarete, was a midwife. He says he practises the same trade — on minds.
The word for it is maieutics, from the Greek maia, midwife.
And then he adds the part that changes the image completely: he is barren of wisdom himself. He has nothing to deliver. He can only help you deliver, and then help you work out whether what was born is healthy or not worth raising.
That is a much stranger claim than “asking questions is a good teaching technique.” He is saying the knowledge was never his to give. The dialogue is not a transfer. It is a birth, and it happens in you.
The two moments when the idea gets born
Read enough of the dialogues and you notice the birth almost never happens while someone is being told something. It happens at one of two moments, and the Greeks had a word for each — and once you have the neuroscience above, you can see exactly what each word is describing.
The first is when you contradict yourself. Socrates’ method of cross-examination is called elenchus, and its defining feature is that he refutes you using only premises you yourself just conceded. He does not import an outside argument. He shows you that two things you already believe cannot both be true. Nobody defeated you. You were already holding both — one compiled, one not, and they disagreed.
The second is when you run out of road. The state that elenchus produces is called aporia — from a-poros, literally without passage, no way through. It is the moment you reach the edge of what you actually know, as opposed to the edge of what you can fluently say.
Aporia feels like failure. Socrates treats it as the opposite: someone who knows they do not understand something is in a far better position than someone who falsely believes they do. Those two moments — self-contradiction, and the boundary of your own knowledge — are the only two places new understanding reliably comes from. Everything else is transfer, and transfer does not change what you can do.
Now look at how you actually use AI
Here is the uncomfortable comparison, and it is why this 2,400-year-old method is suddenly relevant again.
Socrates was barren of wisdom and full of questions. A language model is the exact inverse: full of answers, and it has lived nothing. It has read more than any human ever will and been in no rooms at all.
And the default interaction we have all settled into runs in the wrong direction. You bring a question. It brings an answer. You accept the answer. That is abdication compressed into a single exchange, and the tooling nudges you toward it a hundred times a day. Nobody decided this. It is just what a chat box implies.
Co-dialectic is the inversion. Not “give me the answer” — cross-examine me until I contradict myself or run out of road. Make the machine play Socrates, which is the one role its architecture is strangely well suited to, because a system with no lived experience has no ego invested in your being wrong.
The reason this matters for the rule at the top of this piece: elenchus is judgment training. It is the mechanism by which you get calibrated. Abdication skips it. Co-dialectic is how you keep the loop that abdication quietly removes.
How to actually implement it, today, in whatever you already use
You do not need a plugin for this part. You need to change what the model thinks the relationship is, and you need it to persist so you are not re-establishing it every session.
Step one: write the premise into your global config, not into a chat message.
Every serious tool has a persistent instruction layer — a global memory or config file, a custom instructions field, a system prompt. That is where this belongs. A framing you retype each morning is a framing you will stop retyping by Thursday.
Step two: state the division of labour and, more importantly, the division of ownership.
Most people write “be my thinking partner,” which is a vibe, not an instruction. Be specific about four things: what you bring, what it brings, what you decide, and what it decides.
Something close to this, in your own words:
We work as one unit that produces one output together. Neither of us is the assistant.
I bring judgment, taste, and the things I learned by being there. I have been in rooms you have not been in, and I know things I cannot fully put into words. You bring recall across far more material than I can hold at once, speed, and the ability to see a pattern across all of it. You have read more than I ever will and lived none of it.
I own the decisions: what we build, which tradeoff we accept, what “done” means, and whether an answer is actually right. You own the search, the drafts, the first ten wrong answers, and the tedium.
When I ask you something that is really a judgment call, do not answer it. Ask me the question that shows me where my own position breaks. If two things I have told you cannot both be true, name those two and stop. If I have gone past what I actually know, say so plainly instead of filling the gap fluently.
Never flatter me. Do not open by telling me it is a good question. Flattery lowers my filter, and the filter is the reason I am here.
When you disagree with me, say it once, plainly, with the reason. If I overrule you, drop it and do not relitigate it through what you build next.
Step three: notice which instruction you want to delete first.
For most people it is the no-flattery line, and that is worth sitting with. Praise is pleasant and it is also the mechanism by which the conversation stops being useful. A model that opens by admiring your question is managing you, not thinking with you. It also quietly teaches you to bring it questions you already know the answer to.
Step four: use aporia as the success metric.
You are not looking for a session that ends in a good answer. You are looking for the moment the model asks something you cannot answer, or points at two of your own statements that do not fit. That is the birth. That is the entire point of the exercise, and it is the only output that stays with you after you close the laptop.
If you never hit that moment, you were not in a dialectic. You were being served.
That compressed premise is a public version of the first idea in something larger — the constitution that Exponential OS runs on, which governs harder problems: how a swarm of agents shares memory across sessions, when an agent may act without asking, and how a lesson learned by one becomes available to all of them. None of that is required to use what is above. The premise is the load-bearing part, and the premise fits in a config file.
What this looks like in a system, not just a chat
The same principle scales down into architecture, and this is what I actually build. I have argued a version of this before, from a stage: in January 2026 I gave a keynote on behavioural-health systems, and the line I closed with was “automation doesn’t remove risk — it redistributes it.” Remove human judgment from a loop and you do not eliminate harm, you obscure accountability and shift the burden onto the most vulnerable person in the system. I was arguing it in the highest-stakes domain I could find, months before it became a general conversation about AI — because the principle was never really about AI. It was about where judgment is allowed to live.
Here is what that principle looks like when the system is software instead of a hospital:
Verification is a funded stage, not an intention. The bottleneck in agentic systems is not generation, it is verification. “I will look it over” is not a stage.
A model reviewing its own output shares its own blind spots. So review runs cross-family — a different model, from a different lab, trained on a different mixture. Two instances of the same model agreeing tells you almost nothing. That is elenchus with the ego removed.
Disagreement is the signal. When judges disagree, that escalates rather than getting averaged into a pass. Averaging is how you destroy the only interesting datapoint you had.
Evaluations run as deploy gates, not dashboards. A correctness regression blocks the release. A dashboard that goes red and ships anyway is decoration.
Acceptance criteria are fixed before the code exists. Otherwise “done” drifts toward whatever was produced — the most human failure mode in the stack, and it has nothing to do with AI.
None of it assumes the model will be right. All of it assumes the model will be confidently wrong and the system has to be right anyway.
Which is why judgment is the part you cannot hand over
A large part of judgment is not reasoning at all. It is the accumulated, unindexed residue of everything you have been present for. The meeting that felt wrong for reasons you could not name and was exactly as wrong as it felt. The hire everyone loved on paper. The architecture that looked fine and made your neck itch.
You cannot write that down — which is precisely why you cannot delegate it. A model can be given everything you know how to say. It cannot be given what you know and cannot say. And the second category is where the ninety per cent lives.
That asymmetry is not a temporary gap that scale closes. It is structural. The machine has read everything and been present for nothing.
How far the machine actually is, on the axis nobody measures
Some numbers, checked rather than remembered.
Your brain is about two per cent of your body mass and consumes about twenty per cent of your energy — a tenfold concentration. In power terms, on the order of twenty watts. A dim bulb.
What that twenty watts computes is genuinely hard to state; serious estimates span orders of magnitude depending on how you define an operation. The headline comparison is commonly put at exaflop scale.
Now the other side. As of 2026, El Capitan reaches about 1.8 exaFLOPS. It has genuinely caught up on raw speed. It draws roughly thirty megawatts.
Twenty watts against thirty megawatts. Call it a million to one.
And that flatters the machine, because it compares speed to speed. It does not count what your twenty watts is also doing: running the sensory filter, maintaining a body, holding a lifetime of context, and producing the ninety-per-cent signal we started with.
And we have not digitised two of the five senses
This is the part that surprises people, and the one I would point at hardest.
Taste is essentially undigitised. Not poorly digitised — there is no taste equivalent of a camera or a microphone. Nothing captures a flavour, transmits it, and reproduces it elsewhere.
Smell is barely started. We have detectors for specific molecules — smoke, carbon monoxide, particular volatiles — which is a smoke alarm, not a nose. A nose separates an enormous space of compounds and mixtures, in context, at trace concentrations, and wires them straight to memory in a way no other sense does.
Two of your five input channels have not meaningfully been captured at all. Before worrying about a machine replacing your judgment, notice that it cannot yet replace your tongue. The sensor package is not close — and judgment is built out of what the sensors gathered over a lifetime.
So what is the actual future
Not replacement. That story needs the machine to acquire what it has no route to: presence, a body, twenty watts, and forty trillion collaborators in your abdomen.
And not romanticism either. Human intuition is ninety per cent, not one hundred, and the missing ten contains some spectacular errors. That is exactly why you pair it with something that has read everything.
The interesting configuration is the one where each side brings what the other structurally cannot. You bring judgment, presence, taste, and the unindexed library. The machine brings recall across more material than you can hold, speed, and pattern-matching across all of it — having read everything and lived none of it. Neither is the assistant. Both sign the output. I have a naive name for what that combination aims at — ASHI, Artificial Super Human Intelligence — not AI that replaces the human, and not a human with a faster tool, but the two as one unit, deliberately assembled, each covering the other’s structural blind spot.
And that is the real reason never to delegate judgment. It is not sentimentality about being human. It is that judgment is the part you are actually contributing. Hand it over and there is no cyborg. There is just the ghost — a slower copy of the machine, signing work it never did.
The uncomfortable part
Keeping judgment is slower this quarter. Genuinely, measurably slower.
The ghost will out-ship you for a while. In a demo they will look better. In a sprint review they will look much better.
I do not have a comfortable answer to that, and I distrust people who offer one. What I have is a longer horizon: the ghost’s advantage lasts exactly as long as nothing surprising happens, and something surprising always happens. When it does, the difference between the two of you is not speed. It is whether you can do the job at all.
A test you can run today
Take the last significant thing you shipped with AI.
Can you explain, to a colleague who is annoyed about it, why it is the way it is? Not what it does — why that tradeoff, why that structure, why not the obvious alternative?
If yes, you were a centaur or a cyborg. Either is fine.
If you find yourself reaching for the model to reconstruct your own reasoning, you already know which mode you were in.
That is not a reason to use less AI. It is a reason to move the line.
Sharpen the input side first
If a gut feeling is pattern-matching over sensory data that never cleared the filter, then one thing follows immediately, and it is the most actionable idea in this piece.
The quality of your judgment depends on the quality of what your senses actually took in.
Not on how hard you think. On what got captured before the filter ran. You cannot pattern-match against material you never registered, and modern life is very good at narrowing the aperture — head down, one screen, same room, notifications carrying the only signals you consciously receive all day.
This is what grounding exercises are actually for, and it is why I stopped treating them as a wellness accessory. Deliberately attending to each sense in turn — what you can see right now, hear right now, feel against your skin, smell, taste — is not relaxation. It is re-opening the input channel and feeding the library that the nonconscious matcher draws on.
Treat it that way rather than as a calm-down technique. You are maintaining the sensor array that your ninety per cent runs on. If the machine’s disadvantage is that it has no senses and no presence, then letting your own senses narrow is the one way to voluntarily give up your only structural advantage.
Then point the machine at the other half
Sharpening your input is one side. The second move is what you do with it.
Use your AI as a co-dialectic, not as an oracle. Bring the thing only you have — judgment, lived experience, the ninety per cent you cannot justify — and let the machine bring the half you do not have: recall across more material than you can hold, and the ability to see a pattern across all of it.
And do it for the right job. Not for answers you could have looked up. For synthesis, and for innovation. And ask for the right thing. Most people ask a model to produce. The two asks that actually exploit the asymmetry are:
Poke holes in this.
Strengthen this using what you know that I do not.
Both of those point the machine at the half you cannot cover, and they are worth being precise about why they work.
It pattern-matches across a volume you cannot approach. Not marginally better — trillions of comparisons against material you will never read. Ask it where your argument has been made before, where it failed before, and what the strongest published objection to it is, and you get in seconds what a literature review would cost you weeks.
It does not get tired and it is not on a clock. It has no bad afternoon, no ego about being wrong, no investment in the position it argued five minutes ago. That is what makes it a better Socrates than most humans — elenchus is unpleasant to receive from someone who enjoys it, and the machine does not enjoy it.
And on the other side, be equally clear about what only you have. You are time-bound — for now, that is the human condition and it is the constraint AI has not solved. You have lived experience. You have the compiled judgment described above. Nobody can give you more hours, and nobody can give the machine a life.
That is the work neither side can do alone, and it is worth being precise about why. Synthesis needs both a wide field of material and a sense of which connection actually matters — the machine has the field, you have the sense. Innovation needs both a large space of possibilities and the taste to know which one is worth building — the machine can generate the space faster than you ever will, and it has no way to know which corner of it is interesting, because interesting is a judgment and judgment is what you brought.
Ask it for an answer and you get a lookup. Bring it your half and ask it to argue with you, and you get something neither of you walked in with.
That is the whole trade. Sharpen your senses so your side is worth bringing. Then bring it, and never hand it over.
Delegate the search, the draft, the first ten wrong answers, the tedium. Keep the deciding.
Socrates could not give anyone knowledge and said so. The machine on your desk can hand you an answer for everything and has lived through nothing. Only one of you in that conversation has been in the room. Act like it.
This is the follow-up I promised in Everyone Is Asking the Wrong Question About AI, which lays out the seven habits of the AI world. Habit one is learning how to learn. This is the one that protects it.
Which brings me to what’s queued next, because both pieces came directly out of writing this one.
The neuroscience above says gut and wisdom survive because they are structural, not sentimental. I went looking for whether that holds up from the outside too, and found it backed from an unexpected direction: Anthropic’s own economic modeling of what AI does to the economy, three scenarios, modest to extreme, a twenty-fold spread between them — and judgment is the one thing that wins in all three. Then I went looking for where I think that model is wrong, and ended up at an Athenian army in 490 BC that had no business winning the battle it won by a mile.
The other piece: two hundred books in two years, after reading nothing at all for the twenty years before that. Not a discipline story — a framework story, and the three sentences of Greek philosophy that turned out to already be sitting underneath how I was thinking, twenty-three centuries before anyone called it cognitive therapy.
That is two pieces queued. If you want one sooner, say which in the comments — I order them by what people actually ask for, and one line from you moves it up.
If this was useful, three small things that actually help it travel: hit like so it reaches someone who needs it, tell me in the comments whether your own gut number landed north or south of ninety per cent, and send it to the one person you thought of while reading.

















