David Autor's team published NBER Working Paper 35720 — the result of a preregistered three-month randomized trial involving 133 practicing patent attorneys from 11 U.S. IP firms. A custom AI assistant for drafting patent applications not only improved the quality of work but also left a measurable trace in skill after the tool was abandoned — a result that field studies rarely capture. Skill transfer, however, went to experienced attorneys, while for junior attorneys the average result without the assistant remained zero. The answer to the question of whether AI strengthens or dilutes professional expertise turned out to depend on the specialist's initial level.

image
image
image

What happened

On September 7, 2026, David Autor and co-authors (Rodchenko, Martin, Iscenko, Strand, Pearl, Ferere) published NBER Working Paper 35720 (DOI 10.3386/w35720) — a preregistered three-month randomized controlled trial, registered in the AEA RCT Registry under identifier AEARCTR-0015823. The experiment involved 133 practicing patent attorneys from 11 U.S. IP firms: the experimental group received access to a custom AI assistant for drafting applications and worked with it daily, while all work was blindly evaluated by expert patent attorneys. When using the assistant, benchmark task quality increased by 0.34 SD by day 10 (p = 0.03) and by 0.38 SD by day 90 (p = 0.01), with gains larger for junior attorneys while using the tool. After three months, participants solved a task without AI — redline editing of an existing application: the experimental group outperformed the control group by 0.32 SD (p = 0.04), but the effect was entirely concentrated among senior attorneys (+0.45 SD, p = 0.02), while for junior attorneys the average gain was zero with polarization of evaluations — the share of both very low and very high scores increased.

Context

The debate over whether AI assistance replaces professional expertise has long been ongoing in engineering and professional communities, but until now it relied mainly on self-reports of productivity while using the tool — such measurements are vulnerable to criticism and do not answer the question about the fate of the skill itself. Autor and colleagues' work is significant primarily for its protocol: preregistration of the design, randomization, three months of field work in real firms, and blind evaluation of final documents by experts who did not know who used the assistant. The key methodological move is the transfer assessment: after a long period with the assistant, participants performed a professional task without the tool at all, which allows distinguishing skill accumulation from short-term acceleration while working with AI. The magnitude of the effects is moderate — 0.32–0.38 SD in the main measurements: a statistically significant but small shift, not a qualitative leap.

Why this matters for the industry

For the industry, this is one of the first field RCTs measuring the long-term impact of an AI tool on professional skill, not just productivity while using it, and it sets a reference for AI product evaluation methodology: production evaluation of assistants should include holdout tasks without the tool to track skill drift, dispersion and polarization metrics of quality, not just the average, and breakdown of effects by experience level. Implication for hiring and training in knowledge work: durable gains from assistance go to experienced specialists, while junior specialists show only immediate gains while using the tool, so the scenario of "AI as a replacement for a junior specialist" without senior oversight is not confirmed. The work also removes the main barrier to adoption — the fear of expertise dilution: measurable quality growth by day 10 can be presented to clients as a fact, not a promise, and vertical assistants make sense positioning themselves around senior specialists and the expert feedback cycle.

Why this matters for users

For a practicing specialist, the main takeaway is that skill from AI assistance is not extracted automatically: a measurable effect on the ability to work without the tool appeared only after three months of daily use and depended on initial expertise. Practical guidance — to structure work with the assistant through a feedback loop and understanding checks, since expert evaluation of work and checkpoints were built into the study design. Junior specialists should note that the assistant provides a quick boost while using it, but does not guarantee growth in their own qualifications, and the spread of results may increase: some users gain noticeably, others gain almost nothing. For experienced specialists, assistance, conversely, turns into a durable skill that persists without the tool.

What is still unknown / limitations

Results were measured in one profession (patent law) and one jurisdiction (U.S.) on benchmark tasks with blind expert scoring, so transferring conclusions to other professions, other definitions of quality, and other tools — including coding assistants like Cursor, Windsurf, Replit, and Lovable — is an extrapolation without evidence, as is transfer to productivity metrics. The trial itself lasted three months, so effects on longer horizons remained unmeasured. The bimodal distribution of evaluations among junior attorneys is a signal that the assistant may increase quality variance rather than uniformly raising the level; the mechanism of this polarization requires further study.

Sources

Author

Look at AI, editorial team