The research question
Why does AI so often get stuck in the pilot phase, and what makes the difference toward production?
Why it matters
Technology is rarely the blocker. Adoption stalls on permissions that are wrong, on unclear value, and on people who do not trust it. This record is deliberately honest about how firm the evidence is: a small number of interviews gives signal, not law.
What the research says
This is the topic with the strongest empirical material, and simultaneously the most divided. Noy and Zhang (2023) published a preregistered randomised experiment in Science with 453 professionals: writing tasks went about forty percent faster and rated quality rose. Brynjolfsson, Li, and Raymond (2023) followed 5,179 customer support agents in a real work setting and found an average fourteen percent more issues resolved per hour, with a crucial detail: new agents gained roughly a third, the best agents almost nothing. The technology lifts the bottom; it does not move the top.
Dell'Acqua et al. (2023) ran the most cautionary experiment with 758 BCG consultants. Inside what they call the "jagged frontier", the area AI is good at, more work was completed, faster and at higher quality. But on tasks just outside that boundary, participants with AI were nineteen percent less likely to be correct than participants without. The same tool, opposite effect, depending on whether the task fell inside or outside its reach.
Against this stands Acemoglu (2024), who argues that the macroeconomic gain over a decade stays limited to roughly half a percent to one percent of total factor productivity. That echoes what Brynjolfsson (1993) called the productivity paradox: gains at the individual level do not automatically translate into organisational figures. Davis (1989) laid the groundwork in MIS Quarterly for why: perceived usefulness and perceived ease of use determine whether people actually use a system.
The product layer: Microsoft 365 Copilot documentation states that the product keeps evolving with new capabilities, so adoption is a moving target.
What the research does not prove
The line between what the studies actually establish and what we infer from them.
None of these studies is about Microsoft 365 Copilot. They measure ChatGPT, GitHub Copilot, and generic assistants on specific tasks. Anyone taking the fourteen or forty percent from these papers as expected Copilot return extrapolates further than the research allows.
The experiments measure short, bounded tasks over days or weeks. Whether that gain holds over a year, and whether it translates into business results rather than task speed, has not been established. Acemoglu and the productivity paradox are precisely why caution is warranted.
The jagged frontier is the most important finding in this whole topic, and simultaneously the least actionable: nobody can say in advance where that boundary lies for your tasks. Dell'Acqua et al. show the boundary exists and that crossing it does damage, not how to map it.
There is publication bias. Studies finding gains are published and cited more often than studies finding nothing. The Peng et al. research on GitHub Copilot was also conducted by the vendor itself. Our own adoption guidance on data quality, use case selection, and organisational buy-in derives from this literature, not from a study that tested that approach.
Technical context
Adoption leans on three conditions that rarely all hold: data that is clean enough (permissions, labels), a use case with visible value, and an organization that carries the change. The technology stands; the context wobbles.
Architecture implications
- Choose a first use case with a measurable outcome, not the technically most impressive one.
- Solve data governance and oversharing before rolling out broadly; that is often the real blocker.
- Design for evolution: the platform changes, so avoid hard coupling to a specific feature.
Security implications
Trust is partly a security question. If employees see Copilot surface data that should not have been visible, adoption drops. Governance and adoption reinforce or undermine each other.
Cost implications
The most expensive outcome is a rollout that goes unused: licenses paid, value not realized. Measure adoption and value, not just the number of seats.
Adoption implications
Treat the three forces from the study as hypotheses, not a recipe: invest as an organization, motivate individuals, and make success visible with social proof. Test them in your own context before drawing conclusions.
Trade-offs
- Broad rollout versus focused pilot: faster effect versus more manageable risk.
- Speed versus governance first: momentum versus sustainability.
- Early-adopter enthusiasm versus representative evidence: inspiring but not generalizable.
Common mistakes
- Presenting a small practitioner study as firm, generalizable statistics.
- Choosing the pilot on technology instead of on measurable value.
- Postponing governance until after the rollout.
- Measuring success in seats instead of in usage and outcome.
For architects
Treat adoption as a design question with honest uncertainty. Choose a measurable first use case, solve governance first, and be explicit about how firm your evidence is. This page deliberately carries low confidence to show that honesty.
References
Grouped by source hierarchy. Research carries the reasoning, product documentation carries the implementation. Verify any of it yourself.
Methodology & confidence
This topic rests on the strongest available study designs: a preregistered randomised experiment in Science (Noy and Zhang), a field study at scale (Brynjolfsson, Li, and Raymond), a randomised corporate experiment (Dell'Acqua et al.), plus economic work by Acemoglu and Brynjolfsson and the classic acceptance model by Davis (tier 1 and 2).
We deliberately included the counterweight. Acemoglu contradicts the optimistic extrapolation of the micro studies, and Dell'Acqua supplies the sharpest negative result in this entire file. An overview citing only the gain figures would misrepresent the research. Microsoft Learn supplies the product context on Copilot and the adoption framework (tier 3).
Peer-reviewed research
Journals, systematic reviews, meta-analyses, and reputable conference proceedings. This is the substantive basis.
- Experimental evidence on the productivity effects of generative artificial intelligenceNoy, S., Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654), 187-192DOI 10.1126/science.adh2586
- Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information TechnologyDavis, F.D. (1989). Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology. MIS Quarterly 13(3), 319-340DOI 10.2307/249008
- The productivity paradox of information technologyBrynjolfsson, E. (1993). The productivity paradox of information technology. Communications of the ACM 36(12), 66-77DOI 10.1145/163298.163309
Academic and institutional
Research institutes and standards bodies such as NIST, ISO, IEEE, and ACM.
- Generative AI at WorkBrynjolfsson, E., Li, D., Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161DOI 10.3386/w31161
- Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and QualityDell'Acqua, F., McFowland, E., Mollick, E. et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013
- The Simple Macroeconomics of AIAcemoglu, D. (2024). The Simple Macroeconomics of AI. NBER Working Paper 32487DOI 10.3386/w32487
Official technical documentation
How you build and configure it. Answers the implementation question, not the evidence question.
Continue across TechExplained
The same research, applied in other ways.
