Deploying AI where it doesn't belong is more expensive than not deploying it at all — and the difference between high-value and high-risk use cases comes down to one variable: whether a human can verify the output.
The AI market has a credibility problem. Every vendor, every platform, every consulting deck makes the same promise: AI will transform your operations, reduce costs, and unlock new value. And in many cases, they're right.
But in a striking number of cases, they're not — and the cost of getting it wrong is significantly higher than the cost of not trying.
In 2024, researchers at Harvard Business School and Boston Consulting Group published one of the most important studies in enterprise AI to date. They gave 758 BCG consultants a set of tasks and access to GPT-4. The results split cleanly into two categories.
The Jagged Frontier
| When tasks fell inside AI's capability frontier | When tasks crossed outside it |
|---|---|
| AI users completed 12.2% more tasks | People who trusted AI blindly performed 23% worse than people who didn't use AI at all |
| They worked 25.1% faster | |
| Output quality improved by 40% |
Source: Navigating the Jagged Technological Frontier, Harvard Business School Working Paper 24-013, 2024
This is the finding that should be on every executive's desk. AI doesn't degrade gracefully. It doesn't slow down or flag uncertainty. It produces confident, fluent, well-structured output — regardless of whether that output is correct. And the people most likely to accept incorrect output are the ones who lack the domain expertise to spot the error.
The researchers called this the "jagged technological frontier" — an invisible, irregular boundary between tasks AI handles well and tasks where it fails. The frontier isn't intuitive. It doesn't follow a clean line from "simple tasks" to "complex tasks." AI can ace a sophisticated financial analysis and stumble on a straightforward factual question in the same session.
Three Categories of AI Fitness
Based on the research and real-world enterprise deployments, every workflow falls into one of three categories. The mistake most organizations make is treating all three the same.
Tier 1: Automate Freely — Output is verifiable by anyone; errors are obvious and low-consequence; the task is repetitive and well-defined. Examples: drafting first versions of standard communications, summarizing meeting notes, reformatting data between templates, generating initial document outlines.
Tier 2: AI-Assisted, Human-Verified — Output requires domain knowledge to evaluate; errors are plausible and consequential; the task has variability and edge cases. Examples: analyzing regulatory correspondence, drafting compliance language, summarizing technical documents, comparing contract terms.
Tier 3: Human-Led, AI-Restricted — Only an expert can verify correctness; errors carry regulatory, legal, or safety consequences; the task requires contextual judgment AI cannot replicate. Examples: final regulatory submissions, safety-critical technical decisions, legal interpretations of compliance language, any output that becomes part of a regulated record.
Why the Distinction Matters More in Regulated Industries
In a technology company, a Tier 2 task handled as Tier 1 produces a mediocre product feature. In an energy utility, a financial institution, or a healthcare system, the same mistake produces a compliance violation, a safety incident, or a regulatory fine.
The difference between "shall" and "should" in a governance document isn't semantic — it's the difference between a mandatory requirement and a recommendation. AI models don't understand this distinction with the precision that regulated environments demand.
| AI Capability Level | What Poor Classification Costs |
|---|---|
| AI as writing assistant | A weak first draft that gets edited |
| AI analyzing internal documents | A flawed analysis based on real data |
| AI drafting regulatory language | A compliance document with imprecise legal terms |
| AI agents acting in enterprise systems | An incorrect action in a production environment |
The pattern is clear: as AI capability increases, the cost of poor judgment about where to deploy it increases faster.
The Verification Test
Before deploying AI on any workflow, ask one question:
Can someone on your team verify the output without relying on the AI that produced it?
If yes — the human has independent domain expertise to evaluate quality — the workflow is a candidate for AI assistance. If no — the human would need to redo the work from scratch to check it — the workflow is not ready for AI.
This test sounds simple. In practice, it eliminates roughly 30-40% of the use cases that organizations initially identify as "AI-ready." And that elimination is where the real value protection happens.
| Function | Verification Complexity | AI Fitness |
|---|---|---|
| Internal communications | Low — anyone can judge tone and clarity | Tier 1 |
| Data reformatting and summarization | Low — compare input to output | Tier 1 |
| Technical documentation | Medium — requires domain knowledge | Tier 2 |
| Financial analysis and reporting | Medium — requires understanding of methodology | Tier 2 |
| Regulatory and compliance language | High — requires legal and regulatory expertise | Tier 3 |
| Safety-critical technical decisions | High — requires engineering judgment | Tier 3 |
The Expert Advantage Is Real — and Temporary
The BCG/Harvard study revealed a second critical finding: senior consultants improved their output quality by 17% with AI, while junior consultants improved by 43%. On the surface, this looks like AI is a great equalizer — it lifts everyone, but especially those with less experience.
Look deeper. The senior consultants produced better absolute results because they caught AI errors that juniors missed. The juniors improved more relatively because they started from a lower baseline — but they also accepted more incorrect output without realizing it.
This creates a dangerous dynamic in any organization where experienced employees are retiring faster than they can be replaced. The experts who can verify AI output are the same experts who are leaving. And the junior employees who rely on AI most heavily are the ones least equipped to catch its mistakes.
Trend 1: AI capability — increasing rapidly (more tasks AI can attempt)
Trend 2: Workforce verification capacity — declining (experienced employees retiring, junior employees trained on AI rather than fundamentals)
The intersection of these two trends is where organizational risk concentrates. Every enterprise deploying AI should be asking: how many years of domain expertise do we have left in the building, and are we building the next generation of experts or outsourcing expertise to AI?
What to Do About It
1. Classify before you deploy. Every proposed AI use case should be scored on verification complexity and consequence of error before it enters a pilot. Use the three-tier framework above — Tier 1 is a green light, Tier 2 requires a verification protocol, Tier 3 requires expert human judgment as the primary process.
2. Build verification into the workflow, not after it. Most organizations deploy AI and then say "have someone review it." That's not a process — it's a hope. A real verification protocol specifies: who reviews, what they check for, what tools they use to verify independently, and what authority they have to reject.
3. Invest in judgment, not just tools. The BCG/Harvard data is unambiguous: the value of AI is unlocked by human expertise, not replaced by it. Organizations that cut domain training because "AI handles that now" are building a capability debt that compounds with every retirement.
4. Tell the truth about what AI can't do. The most credible AI strategy is the one that names the boundaries. When an executive asks "can AI do this?" — the answer that builds trust is sometimes "not safely, and here's why." That answer prevents the 23% performance decline that comes from deploying AI past the frontier.
The Bottom Line
AI is not a binary. It is not "works" or "doesn't work." It works brilliantly in some places, adequately in others, and dangerously in a few. The organizations that capture the most value from AI are not the ones that deploy it most aggressively. They are the ones that deploy it most precisely — knowing exactly where it creates value, where it needs supervision, and where it has no business being.
The 83% of AI pilots that fail to reach production don't fail because the technology was wrong. They fail because nobody asked the question that matters: can we verify this?
Sources
- Dell'Acqua, F., McFowland, E., Mollick, E., et al. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality." Harvard Business School Working Paper 24-013, 2024.
- Mollick, E. "Latent Expertise: Everyone is in R&D." One Useful Thing, 2024.
- BCG. "AI Transformation Is a Workforce Transformation." 2026.
- Deloitte. "State of AI in the Enterprise." 2026.
- Duncan, D. "How Do Workers Develop Good Judgment in the AI Era?" Harvard Business Review, February 2026.