Deploying AI where it doesn't belong is more expensive than not deploying it at all — and the difference between high-value and high-risk use cases comes down to one variable: whether a human can verify the output.

The AI market has a credibility problem. Every vendor, every platform, every consulting deck makes the same promise: AI will transform your operations, reduce costs, and unlock new value. And in many cases, they're right.

But in a striking number of cases, they're not — and the cost of getting it wrong is significantly higher than the cost of not trying.

In 2024, researchers at Harvard Business School and Boston Consulting Group published one of the most important studies in enterprise AI to date. They gave 758 BCG consultants a set of tasks and access to GPT-4. The results split cleanly into two categories.

The Jagged Frontier

Exhibit 1
When tasks fell inside AI's capability frontierWhen tasks crossed outside it
AI users completed 12.2% more tasksPeople who trusted AI blindly performed 23% worse than people who didn't use AI at all
They worked 25.1% faster
Output quality improved by 40%

Source: Navigating the Jagged Technological Frontier, Harvard Business School Working Paper 24-013, 2024

This is the finding that should be on every executive's desk. AI doesn't degrade gracefully. It doesn't slow down or flag uncertainty. It produces confident, fluent, well-structured output — regardless of whether that output is correct. And the people most likely to accept incorrect output are the ones who lack the domain expertise to spot the error.

The researchers called this the "jagged technological frontier" — an invisible, irregular boundary between tasks AI handles well and tasks where it fails. The frontier isn't intuitive. It doesn't follow a clean line from "simple tasks" to "complex tasks." AI can ace a sophisticated financial analysis and stumble on a straightforward factual question in the same session.

Three Categories of AI Fitness

Based on the research and real-world enterprise deployments, every workflow falls into one of three categories. The mistake most organizations make is treating all three the same.

Exhibit 2 — The AI Fitness Framework

Tier 1: Automate Freely — Output is verifiable by anyone; errors are obvious and low-consequence; the task is repetitive and well-defined. Examples: drafting first versions of standard communications, summarizing meeting notes, reformatting data between templates, generating initial document outlines.

Tier 2: AI-Assisted, Human-Verified — Output requires domain knowledge to evaluate; errors are plausible and consequential; the task has variability and edge cases. Examples: analyzing regulatory correspondence, drafting compliance language, summarizing technical documents, comparing contract terms.

Tier 3: Human-Led, AI-Restricted — Only an expert can verify correctness; errors carry regulatory, legal, or safety consequences; the task requires contextual judgment AI cannot replicate. Examples: final regulatory submissions, safety-critical technical decisions, legal interpretations of compliance language, any output that becomes part of a regulated record.

Why the Distinction Matters More in Regulated Industries

In a technology company, a Tier 2 task handled as Tier 1 produces a mediocre product feature. In an energy utility, a financial institution, or a healthcare system, the same mistake produces a compliance violation, a safety incident, or a regulatory fine.

The difference between "shall" and "should" in a governance document isn't semantic — it's the difference between a mandatory requirement and a recommendation. AI models don't understand this distinction with the precision that regulated environments demand.

Exhibit 3 — The Cost of Misclassification
AI Capability LevelWhat Poor Classification Costs
AI as writing assistantA weak first draft that gets edited
AI analyzing internal documentsA flawed analysis based on real data
AI drafting regulatory languageA compliance document with imprecise legal terms
AI agents acting in enterprise systemsAn incorrect action in a production environment

The pattern is clear: as AI capability increases, the cost of poor judgment about where to deploy it increases faster.

The Verification Test

Before deploying AI on any workflow, ask one question:

Can someone on your team verify the output without relying on the AI that produced it?

If yes — the human has independent domain expertise to evaluate quality — the workflow is a candidate for AI assistance. If no — the human would need to redo the work from scratch to check it — the workflow is not ready for AI.

This test sounds simple. In practice, it eliminates roughly 30-40% of the use cases that organizations initially identify as "AI-ready." And that elimination is where the real value protection happens.

Exhibit 4 — Verification Complexity by Function
FunctionVerification ComplexityAI Fitness
Internal communicationsLow — anyone can judge tone and clarityTier 1
Data reformatting and summarizationLow — compare input to outputTier 1
Technical documentationMedium — requires domain knowledgeTier 2
Financial analysis and reportingMedium — requires understanding of methodologyTier 2
Regulatory and compliance languageHigh — requires legal and regulatory expertiseTier 3
Safety-critical technical decisionsHigh — requires engineering judgmentTier 3

The Expert Advantage Is Real — and Temporary

The BCG/Harvard study revealed a second critical finding: senior consultants improved their output quality by 17% with AI, while junior consultants improved by 43%. On the surface, this looks like AI is a great equalizer — it lifts everyone, but especially those with less experience.

Look deeper. The senior consultants produced better absolute results because they caught AI errors that juniors missed. The juniors improved more relatively because they started from a lower baseline — but they also accepted more incorrect output without realizing it.

This creates a dangerous dynamic in any organization where experienced employees are retiring faster than they can be replaced. The experts who can verify AI output are the same experts who are leaving. And the junior employees who rely on AI most heavily are the ones least equipped to catch its mistakes.

Exhibit 5 — The Narrowing Verification Window

Trend 1: AI capability — increasing rapidly (more tasks AI can attempt)

Trend 2: Workforce verification capacity — declining (experienced employees retiring, junior employees trained on AI rather than fundamentals)

The intersection of these two trends is where organizational risk concentrates. Every enterprise deploying AI should be asking: how many years of domain expertise do we have left in the building, and are we building the next generation of experts or outsourcing expertise to AI?

What to Do About It

1. Classify before you deploy. Every proposed AI use case should be scored on verification complexity and consequence of error before it enters a pilot. Use the three-tier framework above — Tier 1 is a green light, Tier 2 requires a verification protocol, Tier 3 requires expert human judgment as the primary process.

2. Build verification into the workflow, not after it. Most organizations deploy AI and then say "have someone review it." That's not a process — it's a hope. A real verification protocol specifies: who reviews, what they check for, what tools they use to verify independently, and what authority they have to reject.

3. Invest in judgment, not just tools. The BCG/Harvard data is unambiguous: the value of AI is unlocked by human expertise, not replaced by it. Organizations that cut domain training because "AI handles that now" are building a capability debt that compounds with every retirement.

4. Tell the truth about what AI can't do. The most credible AI strategy is the one that names the boundaries. When an executive asks "can AI do this?" — the answer that builds trust is sometimes "not safely, and here's why." That answer prevents the 23% performance decline that comes from deploying AI past the frontier.

The Bottom Line

AI is not a binary. It is not "works" or "doesn't work." It works brilliantly in some places, adequately in others, and dangerously in a few. The organizations that capture the most value from AI are not the ones that deploy it most aggressively. They are the ones that deploy it most precisely — knowing exactly where it creates value, where it needs supervision, and where it has no business being.

The 83% of AI pilots that fail to reach production don't fail because the technology was wrong. They fail because nobody asked the question that matters: can we verify this?

Sources

Talk to me about your AI strategy →