Most of the AI productivity numbers you have seen are task-level measurements dressed up as company-level outcomes. They are not wrong, they are incomplete. We have scoped enough AI builds to know that the gap between what a tool does in a controlled study and what it does inside an actual operation is where most projects stall. Here is what the research actually shows, and what it does not.
What Controlled Studies Say About Task-Level Gains
Controlled studies isolate a specific worker, a specific task, and a specific AI tool. They remove all the organisational friction. What they find is large, within those boundaries.
Where AI Consistently Delivers
Nielsen Norman Group measured a 126% increase in coding output and a 59% reduction in document writing time. A separate study on realistic daily office tasks found a 66% throughput increase for workers using AI assistance. Harvard Business School research on BCG consultants found a 40% quality improvement on analytical tasks.
These numbers hold across repeated studies. For defined, text-based, output-measurable tasks, drafting, data extraction, code generation, classification, AI produces real, large gains.
Where AI Adds Minimal Value
The research is equally consistent about where AI underdelivers. Judgment calls, relationship-dependent work, tasks requiring institutional context, and anything that depends on physical presence show near-zero controlled-study gains.
Customer escalation decisions. Strategic prioritisation. Supplier negotiations. Sales calls. These are not tasks AI tools struggle with because of a missing feature, they are structurally outside what current models do well. Any vendor telling you otherwise is extrapolating from the controlled-study wins.
What Firm-Level Research Says: The Productivity Paradox
Scale up from individual tasks to whole organisations, and the picture changes completely.
NBER Survey: 80% of Firms Report Zero Gains
A March 2026 working paper from the NBER and Federal Reserve Bank of Atlanta surveyed over 6,000 executives across four countries. More than 80% reported zero measurable productivity improvement from AI adoption. The mean firm-level gain across all respondents sat at 0.29%.
That is not a rounding error. It is the average outcome for a company that buys AI tools, rolls them out to staff, and tracks results, without doing anything else differently.
Goldman Sachs: No Economy-Wide Relationship Yet
Goldman Sachs Research (March 2026, covered by Fortune) found no meaningful relationship between AI adoption and economy-wide productivity. Where they did find gains, they were localised: roughly a 30% median boost in two specific, well-defined use cases. Not categories. Specific use cases.
Goldman also noted that only 10% of S&P 500 management teams quantified AI’s impact on specific use cases. Just 1% quantified impact on earnings. Companies are spending on AI without measuring it, and then reporting to investors without numbers.
PwC 2026: 20% of Firms Capture 74% of AI Value
PwC’s 2026 AI Performance Study found that 74% of AI’s total economic value is captured by 20% of organisations. The majority are stuck in pilot mode, running proofs of concept that never reach production scale.
The firms in that top 20% share one distinguishing characteristic. It is not which tools they use. It is addressed in the next section.
The Variable That Separates Winners From Everyone Else
If you run through the research looking for a common thread in the companies that capture AI value, one variable dominates.
Workflow Redesign, Not Tool Selection, Drives Results
An analysis of 200 AI projects from 2022–2025 (Business.com 2026 Small Business AI Outlook) found a median ROI of 159% for SMEs, with payback in 6.7 months. The catch: that outcome applied only to implementations that included deliberate workflow redesign. Bolt-on implementations, same process, AI tool added, performed near the NBER mean.
Workflow redesign means mapping the current process, identifying which steps produce the bottleneck, removing or restructuring those steps around AI’s actual capabilities, and retraining staff on the new sequence. It is slower than buying a subscription. It is also the only thing that moves the firm-level number.
The Rework Problem: 40% of AI Time Savings Are Erased
Workday’s research found that 40% of time saved by AI is lost to rework, checking, correcting, and verifying AI outputs. This is not a failure mode. It is what responsible use looks like. But it means the 66% throughput gain from controlled studies becomes something closer to 40% net gain in practice, and closer to zero if the workflow has not been redesigned to absorb verification costs.
The Federal Reserve’s compiled research puts average AI time savings at 5.4% of work hours, about 2.2 hours per week for a typical worker. Frequent, high-skill users with redesigned workflows see over 9 hours per week. The difference is not which model they use.
What This Means for SMBs Making AI Decisions in 2026
The research is not a reason to avoid AI. It is a reason to be specific about what you are actually buying.
Which Functions Show Documented SMB Gains
Based on the controlled-study evidence, these functions show documented, replicable gains across business sizes, provided inputs are structured and a human review step is preserved:
- Content drafting and editing: 40–60% time reduction with human review preserved
- Data entry and extraction: 70–90% time reduction for structured inputs; drops sharply with inconsistent or unstructured source data
- Code generation and review: 50–126% throughput increase for developers working on well-scoped, bounded tasks
- Customer support triage: 30–50% reduction in first-response time for templated query types; gains fall off for complex or emotionally charged issues
- Document classification: Significant reduction in manual sorting for standardised, consistent document sets; unreliable on documents with variable formatting or missing fields
The pattern: high-volume, repetitive, text-based tasks with clear output criteria and a human review step. If the task does not fit that description, treat vendor productivity claims with proportional skepticism.
How to Evaluate an AI Vendor’s Productivity Claims
Ask four questions before any contract:
-
Is the cited gain task-level or firm-level? Task-level gains are real but do not automatically translate to firm-level results. If the vendor is citing task studies to justify a firm-wide rollout, that is a logic gap.
-
Does the implementation include workflow redesign? If the answer is “we plug it into your existing process,” the NBER data suggests you will be in the 80%.
-
What is the rework overhead? Any honest productivity estimate needs to account for output verification. Ask how that is costed into the ROI model.
-
What are the comparable firm-level results from their other SMB clients? Not demos. Not case studies selected for the website. Clients you can call.
Frequently Asked Questions
What is the average productivity gain from using AI tools?
It depends on what you measure. Task-level controlled studies show gains of 30–126% depending on the task type, with coding and drafting at the high end. At the firm level, the NBER’s 2026 survey of 6,000+ executives found a mean gain of 0.29%, with over 80% of firms reporting zero measurable improvement. Both numbers are accurate, they measure different things.
Why do most companies report no productivity improvement from AI?
The primary reason, consistent across the PwC and NBER research, is that most implementations are bolt-on additions to existing workflows. AI can do more work faster, but if the process around it is unchanged, the bottlenecks shift rather than disappear. The 20% of firms that capture most of the value are the ones that redesigned workflows around AI’s actual capabilities, not around the demo.
Which roles see the biggest productivity gains from AI?
Software developers, content writers, legal researchers, data analysts, and customer support agents operating on structured query types. These roles involve high-volume, text-based, output-measurable work where AI can handle the repetitive layer and humans handle the judgment layer. The gains are weakest in roles that depend on relationship capital, physical presence, or real-time contextual judgment.
How long does it take to see ROI from AI adoption for a small business?
The Business.com 2026 analysis of 200 SME AI projects found a median payback period of 6.7 months, but only for projects that included workflow redesign. Bolt-on implementations without process change show no consistent payback timeline in the data, because the firm-level gain at that level is near-zero. Budget 2–3 months for design and integration before you start counting ROI.
Is AI productivity data reliable or mostly vendor-produced?
Both types exist and need to be read differently. The strongest data, NBER, Goldman Sachs Research, PwC, Nielsen Norman Group, uses either large survey samples or controlled experimental designs, and most is independent of AI vendors. Vendor-produced case studies and stat aggregators typically cite task-level controlled studies without the firm-level context. Read the methodology footnotes, and check whether the cited source separates task-level from firm-level results.
The research gives you enough to act on, if you know which questions to ask. The gains are real on specific tasks. They do not automatically appear at the firm level. The companies capturing AI value are the ones that redesigned how work happens, not just which tools they use.
At Designodin, we scope AI integrations against these criteria before any build starts, workflow assessment is part of the engagement, not an afterthought. If you want to talk through what this looks like for your operation, start a conversation. See how we scope and build this at designodin.com/ai.