Back to News
EnterpriseAI Understanding briefing

Harvard study finds AI coding agents boost code output but not issue resolution

A new Harvard working paper analyzing data from 718 firms shows AI coding agents increase lines of code written by 30% while code‑review bottlenecks keep resolved Jira issues flat.

4 min readRead the linked source
Source-provided image accompanying Harvard study finds AI coding agents boost code output but not issue resolution
Source referenceSource recorded
Publisher
forkast.news
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Key terms

Artificial Intelligence (AI)
The broad field of building systems that perform tasks requiring pattern recognition, reasoning, language, or decision-making.
Dataset
A collection of structured or unstructured examples used for training, validation, or testing.
Test yourselfAI Agents Quiz

What happened

A Harvard working paper titled *Artificial Intelligence in the Firm: Bottlenecks in Software Production* examined engineering data from the Jellyfish analytics platform covering 718 firms and more than 725,000 workers between 2021 and 2026. The authors, Fiona Chen and James Stratton, report that AI coding agents such as Claude Code, Cursor, and Devin raise code production by roughly 30%—about 1,500 additional lines per worker‑month. However, the study finds no statistically significant rise in resolved Jira issues or completed epics. Commit activity rose 20% and pull‑request volume 23%, but review times grew 49% (an average increase of 3.45 days) and the share of pull requests needing revision nearly doubled. The paper notes that 14% more workers are now spending time on code review, and while 80% of firms use AI‑assisted review tools, only about 23% of review comments are AI‑generated.

The Harvard working paper draws on a massive from the Jellyfish engineering analytics platform, which tracks code commits, pull‑request activity, and issue resolution across a wide cross‑section of firms. The authors applied statistical analysis to compare periods before and after AI coding agent adoption, controlling for firm size and industry.

AI agents examined—Claude Code (Anthropic), Cursor (Cursor AI), and Devin (Cognition)—were found to increase raw code output by roughly 30%, adding about 1,500 lines per worker‑month. This boost translated into 20% more commits and 23% more pull requests per month.

Despite higher activity, the number of resolved Jira issues and completed epics did not change in a statistically significant way. Review times lengthened by 49%, and the proportion of pull requests requiring revision nearly doubled, indicating that the extra code created more work for human reviewers.

The study also reports a 14% increase in workers allocated to code‑review tasks, and notes that while 80% of firms employ AI‑assisted review tools, only about 23% of review comments are generated by AI, underscoring the continued reliance on human judgment.

Source details: forkast.news ↗

Why it matters

The findings challenge the prevailing narrative that AI coding agents automatically translate into faster software delivery. By highlighting a shift of the bottleneck from coding to review, the study warns enterprises that without parallel investment in downstream processes—automated testing, review tooling, or team restructuring—AI‑driven productivity gains may be illusory. For engineering leaders, the research suggests that the true ROI of AI agents depends on holistic workflow redesign rather than simply adopting code‑generation tools. The human “babysitting tax” identified—more senior engineers spending time vetting AI‑generated code—could erode cost savings and delay product releases, especially as AI adoption climbs from near‑zero in late 2024 to over 95% by early 2026.

The research overturns the simplistic metric of "lines of code per month" as a proxy for engineering productivity. By exposing the downstream review bottleneck, it forces enterprises to reconsider how they measure AI impact.

If firms continue to invest heavily in AI code‑generation without addressing review capacity, they risk higher labor costs, longer time‑to‑market, and potential quality regressions due to rushed or insufficiently vetted code.

The study’s implication that AI productivity gains do not automatically translate into business outcomes aligns with broader observations about technology adoption: benefits accrue only when the entire value chain is optimized.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

What to watch next

Watch for vendor responses that bundle AI coding agents with enhanced review automation, such as AI‑driven static analysis or auto‑merge confidence scores. Monitor whether large firms announce dedicated budget lines for expanding review capacity or redesigning team structures to mitigate the identified bottleneck. Follow follow‑up academic work that attempts to quantify the net productivity impact when review‑stage tooling is upgraded alongside AI code generation.

Product announcements that pair AI coding agents with new automated testing suites, static analysis, or AI‑driven code‑review assistants.

Corporate earnings calls or engineering‑leadership briefings that reference budget allocations for expanding review teams or tooling.

Academic or industry follow‑up studies that attempt to measure the net effect of combined coding‑agent and review‑automation deployments.

Related guides & quizzes

Found this useful?