As AI becomes more capable, the skill that separates high-value users from low-value users isn't how well they instruct AI — it's how well they evaluate what AI gives back.
There's a comforting narrative in enterprise AI: the people who get the most from AI are the ones who prompt it best. Write clearer instructions, provide better context, learn the right techniques — and AI delivers better results.
It's a clean story. It's easy to teach. And the research says it's wrong.
The data from three independent studies — Harvard Business School, Wharton, and Berkeley — converges on a different conclusion. The people who extract the most value from AI aren't the ones who give it the best instructions. They're the ones who can judge whether the output is any good. This distinction matters because it changes everything about how organizations should train their workforce for AI.
Study 1: The Jagged Frontier (Harvard/BCG, 758 Consultants)
In 2024, researchers gave 758 BCG consultants access to GPT-4 and measured their performance across 18 tasks. The headline results were striking:
| Metric | Junior Consultants | Senior Consultants |
|---|---|---|
| Performance improvement with AI | +43% | +17% |
| Best absolute results | No | Yes |
| Caught AI errors on frontier-crossing tasks | Rarely | Frequently |
| Performance when AI was wrong (trusted blindly) | -23% vs. no AI | Avoided the trap |
Source: Dell'Acqua et al., Harvard Business School Working Paper 24-013
The junior consultants improved more — relatively. But the senior consultants produced the best work — in absolute terms. Why? Because the seniors knew what "good" looked like independent of AI. When AI produced confident, fluent, well-structured output that happened to be wrong, the seniors caught it. The juniors accepted it.
The most dangerous finding: on tasks that crossed AI's capability boundary, people who trusted AI without verification performed 23% worse than people who didn't use AI at all. Not 23% less improvement. 23% worse than not using the technology.
Study 2: Latent Expertise (Wharton, Ethan Mollick)
Mollick's research at Wharton introduced a concept he calls "latent expertise" — the idea that AI knows many things imperfectly, and human experts unlock that knowledge because they can see through the imperfections.
| Capability | Expert User | Non-Expert User |
|---|---|---|
| Spot hallucinations | Recognizes fabricated references or misquoted standards | Accepts plausible-sounding output at face value |
| Judge output quality | Evaluates whether a summary captures the right nuances | "This sounds professional" = "this is correct" |
| Instruct precisely | Asks for specific, decomposed sub-questions | Asks for "a summary of these documents" |
| Recognize good output faster | Identifies the right answer in 30 seconds | Spends 10 minutes evaluating output they can't assess |
Source: Mollick, "Latent Expertise: Everyone is in R&D," One Useful Thing, 2024
Mollick's key insight: "The source of any real advantage in AI will come from the expertise of employees, which is needed to unlock the expertise latent in AI." The expertise isn't in the prompting. It's in the person.
Study 3: The Judgment Erosion Problem (HBR/Berkeley, 2026)
This is where the finding gets uncomfortable. If judgment is what makes AI valuable, and judgment comes from experience, what happens when AI removes the experiences that build judgment?
Before AI: Junior employee → does repetitive, messy work → develops pattern recognition → makes mistakes, gets feedback → builds domain expertise → becomes the expert who trains the next junior.
After AI: Junior employee → AI handles the repetitive work → employee never develops pattern recognition → can't evaluate AI output because they never learned the domain → organization loses the pipeline that creates experts.
Source: Duncan, "How Do Workers Develop Good Judgment in the AI Era?" HBR, February 2026
A second study, from Berkeley Haas, adds urgency: AI-equipped workers spontaneously work faster, take on broader scope, and extend their hours — replacing deliberation with speed. The 83% who reported increased workload aren't developing deeper expertise. They're producing more output with less reflection. (Ranganathan & Ye, "AI Doesn't Reduce Work — It Intensifies It," HBR, February 2026)
The Three Skills That Define the Judgment Gap
The research converges on three specific human capabilities that separate high-value AI users from low-value ones. None of them are about how you write a prompt.
Skill 1: Knowing what "done" looks like. The power user spots that an AI summary missed a conditional clause that changes a compliance requirement. The low-value user says "looks good" because they don't know what correct looks like. This is a quality standard skill — built through years of doing the work, not a prompting workshop.
Skill 2: Knowing what to ask for. "Analyze these 50 documents" produces generic output. "Compare the safety margin language across these 50 documents and flag any instance where the required margin changed between 2022 and 2025" produces specific, actionable intelligence. This is a problem decomposition skill, rooted in domain understanding.
Skill 3: Knowing when AI is wrong. AI outputs are fluent, confident, and well-structured — regardless of correctness. The hardest skill to develop and the easiest to lose: recognizing confident errors, output that looks right, sounds right, and is wrong.
Why This Gap Widens With Every Model Release
The judgment gap isn't static. It grows as AI becomes more capable. And it grows in a direction most organizations don't expect.
| AI Capability Level | What Poor Judgment Costs |
|---|---|
| Chatbot (paste text in, get text back) | A weak first draft or a poorly written email |
| AI with file upload (analyze documents) | A flawed analysis based on real organizational data |
| AI connected to enterprise systems (MCP, APIs) | An incorrect decision based on data the user didn't independently verify |
| Autonomous AI agents (act in production systems) | A regulatory incident, a wrong action in a live environment |
The pattern: more capability × poor judgment = more confident mistakes at higher stakes. Every model release, every context window expansion, every new enterprise integration makes the judgment gap more consequential. The person who can evaluate, decompose, and verify becomes exponentially more valuable — not less.
Line 1: AI capability — steep upward curve (new models every quarter, expanding context, agentic capabilities).
Line 2: Average employee judgment — flat or declining (AI removes learning reps, work intensification replaces deliberation, experienced employees retire).
Line 3: High-judgment employees — steep upward curve (AI amplifies their existing expertise, they produce dramatically better results than peers).
The gap between Line 2 and Line 3 is the judgment gap. It widens with every advancement in AI capability. The organizations that invest in Line 3 — building and preserving human judgment — will outperform the ones that assume Line 1 (better tools) solves the problem.
What Organizations Should Do
1. Stop calling it "AI training." Call it "judgment development." The reframe matters — instead of "how to write better prompts," it's "how to evaluate whether AI output meets the standard your domain requires."
2. Train verification before generation. Have employees evaluate AI-generated output — identify errors, assess quality, flag hallucinations — before they learn to generate with AI. This builds the critical muscle first.
3. Protect the experiences that build expertise. For every workflow you automate, answer: "How will junior employees learn to do this well enough to verify AI's work in five years?" If you don't have an answer, you're trading short-term productivity for long-term capability loss.
4. Measure judgment, not just usage. Add judgment metrics: accuracy of AI output evaluation, quality of AI-assisted work product as rated by domain experts, and decision quality in AI-augmented workflows.
5. Treat AI like a capable intern, not a trusted colleague. MIT Sloan and BCG report that 76% of executives now view AI as a "co-worker." That's exactly the wrong posture. The better frame: AI is a very capable intern — smart, fast, impressive — but you review everything before it goes out.
The Bottom Line
The AI market is saturated with prompting courses, certification programs, and "AI skills" workshops. Most of them teach the wrong skill.
The skill that matters isn't instructing AI. It's judging AI. Knowing what good output looks like. Knowing when AI is wrong. Knowing what to ask for in the first place. These aren't AI skills. They're human skills — domain expertise, critical evaluation, problem decomposition — that happen to be the most valuable capabilities in an AI-augmented workplace.
The best AI users aren't the best prompters. They're the best thinkers. And that's a much harder — and more valuable — problem to solve.
Sources
- Dell'Acqua, F., McFowland, E., Mollick, E., et al. "Navigating the Jagged Technological Frontier." Harvard Business School Working Paper 24-013, 2024.
- Mollick, E. "Latent Expertise: Everyone is in R&D." One Useful Thing, 2024.
- Duncan, D. "How Do Workers Develop Good Judgment in the AI Era?" Harvard Business Review, February 2026.
- Ranganathan, A. and Ye, X.M. "AI Doesn't Reduce Work — It Intensifies It." Harvard Business Review, February 2026.
- MIT Sloan Management Review / BCG. "The Emerging Agentic Enterprise." 2026.
- Eatough, E., Ferrazzi, K., Smith, W., Waters, S. "Why AI Adoption Stalls, According to Industry Data." Harvard Business Review, February 2026.
- Davenport, T. and Bean, R. "Five Trends in AI and Data Science for 2026." MIT Sloan Management Review, 2026.