社團指南

What Research Says About AI and Workplace Productivity

Controlled field experiments show generative AI can make workers substantially faster and better at specific tasks such as customer support, writing and some coding, with the largest gains usually going to less experienced workers.

  • 4 分鐘閱讀
  • 最後更新
本頁4 分鐘閱讀
  1. 概述
  2. 深入探討
  3. 戰略影響
  4. The Future of What Research Says About AI and Workplace Productivity
  5. 現實世界的實施
  6. 風險與防護欄
  7. 實施路線圖
  8. 不斷探索
  9. 常見問題

概述

The gains are uneven, though. AI can make people worse on tasks outside its abilities, and broad studies of whole workforces and economies so far find much smaller effects than lab and single-task studies do.

深入探討

The best-known evidence comes from field and controlled experiments on well-defined tasks. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied more than 5,000 customer support agents who were given an AI assistant. On average they resolved about 14 percent more issues per hour. Novice and lower-skilled agents gained around 34 percent, and the most experienced agents gained little. The researchers argue the AI spread the know-how of top performers to everyone else. Shakked Noy and Whitney Zhang, publishing in Science in 2023, found that ChatGPT cut the time college-educated professionals spent on writing tasks by about 40 percent and raised rated quality by about 18 percent. A GitHub Copilot experiment found developers finished a set programming task about 55 percent faster. The 2023 Boston Consulting Group study with Harvard and other researchers added an important caveat, which it called the 'jagged technological frontier'. On tasks inside the AI's abilities, consultants with GPT-4 finished more tasks, faster and at higher quality. On a task designed to sit outside those abilities, they were less likely to get the right answer than consultants without AI, because they trusted plausible but wrong output. More recent findings are more modest. In a 2025 randomized trial by METR, experienced developers working in their own large codebases were about 19 percent slower with AI tools while believing they had been faster. Studies of whole labor markets, such as research on Danish workers, have found small effects on earnings and hours so far. Why the gap? Experiments isolate tasks that suit AI. Real jobs mix many tasks, time saved is not always reused productively, and organizations need to redesign workflows before gains show up. Economists saw a similar lag in measured productivity after electricity and computers arrived. The common misconception is that one headline percentage applies to every job.

戰略影響

風險與安全

災難性和日常的人工智慧危害都取決於誰了解風險以及誰能夠採取行動。

更明確的決策

民眾和專業素養決定強而有力的安全政策在政治上是否可行。

突破炒作

清晰的解釋可以減少炒作、實驗室公關和模糊道德劇場的影響。

The Future of What Research Says About AI and Workplace Productivity

The evidence base is growing quickly, and the models studied in 2023 are already outdated, so results may shift as tools, training and workflows improve. The key open questions are whether task-level gains add up to firm-level and national productivity growth, whether benefits keep concentrating among less experienced workers, and how deskilling or over-reliance affects people in the long run. Careful researchers expect a lag like the one seen with earlier general-purpose technologies. Treat both very large and near-zero headline numbers with caution until more long-term, economy-wide data exists.

現實世界的實施

In a large customer support study, agents given an AI assistant resolved more issues per hour on average, and the newest, least experienced agents improved the most.

In an experiment with professional writing tasks, participants using ChatGPT finished faster and produced work that graders rated higher in quality.

Consultants at Boston Consulting Group did better with GPT-4 on tasks inside the AI's abilities, but were more likely to reach wrong answers on a task deliberately designed to fall outside them.

Experienced open-source developers in a 2025 randomized trial by METR took longer to finish tasks with AI tools, even though they believed the tools had made them faster.

風險與防護欄

  • 將存在風險視為科幻小說,同時能力複合。

  • 混淆了表面產品安全與高度自治下的對準。

  • 只給非英語和非專業觀眾留下低品質的資源。

實施路線圖

  1. 單獨的產品危害、誤用和失控/失調風險。

  2. 詢問哪些證據會改變您對時間表和嚴重性的看法。

  3. 比起行銷主張,更喜歡主要來源和具體評估。

  4. 確定一條行動路徑:職業、政策、資金或技能——而不僅僅是意識。

不斷探索

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the What Research Says About AI and Workplace Productivity quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

開始測驗

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

常見問題

What is What Research Says About AI and Workplace Productivity?

Controlled field experiments show generative AI can make workers substantially faster and better at specific tasks such as customer support, writing and some coding, with the largest gains usually going to less experienced workers. The gains are uneven, though. AI can make people worse on tasks outside its abilities, and broad studies of whole workforces and economies so far find much smaller effects than lab and single-task studies do.

In the Brynjolfsson, Li and Raymond customer support study, which agents gained the most from AI?

Novice and lower-skilled agents improved by around 34 percent, while the most experienced agents gained little. The average gain was about 14 percent.

What did Noy and Zhang's 2023 Science study find about ChatGPT on writing tasks?

Professionals finished writing tasks much faster, and graders rated the quality higher.

What does the 'jagged technological frontier' describe?

AI's abilities are uneven. Inside the frontier it helps, and outside it people who trust it can do worse.

In the BCG study, what happened on the task designed to fall outside AI's abilities?

Consultants using GPT-4 on that task were less likely to be correct, because they trusted plausible but wrong output.

What did METR's 2025 randomized trial with experienced developers find?

Measured completion times were slower with AI, yet the developers believed it had sped them up. This shows how unreliable self-reports can be.