The Manager's Skill AI Development Actually Needs

When AI writes the code, the bottleneck stops being who can write it and becomes who can tell if it's right.

Here is an uncomfortable parallel.

A good manager can often tell whether an employee's work is solid without opening the file, without reading the report line by line, without running the QC checklist themselves. They read the person. They read the process. They read the signals around the work — how questions were answered, how edge cases were discussed, how confidently the person can defend a decision under pressure. That is a managerial skill. It has almost nothing to do with technical expertise in the work itself.

Now apply that to AI-written code.

When 100% — or even a high percentage — of code is written by AI, the same dynamic appears. The person directing that code cannot simply "look at the code" the way a traditional reviewer would, because the volume and pace make that unscalable, and because the bottleneck is no longer can a human write this — it's can a human tell if this is right.

This is not a technical problem. It is a managerial one.

Same skill, a different actor

Managing a personManaging the AI
Reading the workReads how they answer, defend, and discuss edge casesReads where the output is smooth, inconsistent, or evasive under probing
The blind spotUnder-delivery slips past a manager who can't detect the gapShallow-pass / deep-fail code slips past the same way
Doing the checkDelegates verification to an employee, then judges the reportDirects the AI to verify and stress-test its own work, then judges the answer
What the manager contributesWhich slice matters, what to ask, whether the answer holdsIdentical — unchanged

A manager who lacks managerial skill will let underperforming work slide through — not from bad faith, but because they genuinely could not detect the gap. The employee did not necessarily act in bad faith either. It was a mutual blind spot: one side under-delivered, the other side lacked the discernment to notice.

AI-written code creates the exact same structure. The AI is not acting in bad faith when it produces work that passes a shallow check but fails a deeper one. It is simply the same fact pattern as the employee case, restated with a different actor. Without someone who can sense — through questioning, through pattern recognition, through knowing where the soft spots in a system usually hide — whether the output is actually sound, the gap goes undetected until it costs something.

The skill that scales here is not "I can write better code than the AI" or "I can review every line." It is the same skill a good manager uses on people: knowing where to probe, what questions expose weakness, when confidence is earned versus performed, and what a "good outcome" looks like even without inspecting every step that produced it.

As AI takes over more of the writing, the differentiator is not technical depth. It's managerial judgment — applied to a non-human contributor for the first time.

Org chart: a manager oversees three human reports and one AI Agent as a peer; the AI Agent itself branches into a sub-team of Engineer, Analyst, Designer, Consultant and more — managed like a single hire, delivering like an entire org.

A natural objection follows: but shouldn't someone still read the code?

Sometimes, yes — and that is exactly where the distinction matters. There are two different acts hiding under the word "review." Full review is technical: reading every line, tracing every path, verifying correctness directly. It does not scale with AI-written volume any more than it scaled with a large team of human engineers; it remains necessary in narrow, high-stakes slices, but it cannot be the universal method, or the bottleneck simply moves from "who writes the code" to "who has time to read all of it." Spot-checking is managerial: sampling, choosing where to look based on risk, history, and a sense of where the soft spots usually are, then using what is found there to infer something about the whole.

But to be consistent with how this all works, the manager does not personally sit down and read that 5% either. A real manager delegates the check itself — asking an employee to verify and report back. The manager's skill was choosing what to check and how to judge the answer, not doing the checking by hand. The AI case follows the identical pattern: the manager directs the AI to verify, explain, or stress-test its own 5%, and the AI executes that too. The manager never personally reads the code; the AI both writes it and, on request, interrogates it. What the manager contributes is unchanged: deciding which slice matters, what question to ask of it, and whether the answer holds up. The loop stays consistent end to end — it is judgment applied recursively, not a single human exception carved out of an otherwise automated process.

So the objection is correct, but it answers a different question than the one that actually scales. The real question is not whether to read code, but who decides what to read, and how they decide it. That decision is judgment — and judgment does not require the manager to personally execute the inspection at any layer, human or AI.

Named precisely, that judgment is a combination of two things: signal-reading and signal-triggering. Signal-reading is the ability to notice what a piece of work, a person's answer, or an AI's output is actually telling you — the hesitation in how a question gets answered, the inconsistency between two claims, the part of an explanation that's suspiciously smooth. Signal-triggering is the other half: knowing what to ask or do in order to make those signals appear in the first place. A good manager doesn't wait for a signal to surface on its own. They ask the question that forces it out — the probe a confident bluff can't survive, the follow-up a shallow understanding can't answer.

The same applies to AI. A manager who only reads what the AI volunteers is reading nothing. A manager who knows which question to pose — to the AI, or to a person tasked with checking the AI — is the one who actually sees whether the work is sound. The skill is not in looking harder. It's in knowing what to do to make the truth visible, and then reading it correctly once it is.

Picture two managers facing the same uncertain piece of work.

The first asks one sharp question: "Walk me through the edge case where this fails." The answer — fluent or fumbling, specific or evasive — tells them in thirty seconds what they needed to know. They didn't read a line of code. They didn't need to.

The second has no such question. So they ask for the full QC report instead — every test, every log, every line annotated — and even then they are not reading it themselves; they are trusting that someone else's full review caught what they couldn't ask for directly. They get an answer eventually, but they got there by brute force, not by skill. And brute force is exactly what doesn't scale once the volume of AI-written work multiplies.

The first manager has the skill this moment requires. The second is still managing as if the work were small enough to inspect by hand.

in AI
Share this post
Tags
Archive
We're an AI-Native Company
And other phrases nobody in the room can define the same way.