AI in knowledge work

Can AI tools build professional judgement, or only substitute for it?

The evidence so far points the wrong way: a 2025 study of high school mathematics in the Proceedings of the National Academy of Sciences found that generative AI without guardrails harmed learning once the AI was removed. Public Policy Lab argues that a scaffold designed to externalise expert reasoning during real work, visibly and in a form the practitioner must engage with rather than receive, could build judgement instead of replacing it. This is a proposition with a proposed evaluation design, not a demonstrated result.

Where the bottleneck is

State capability in developing-country and public-sector contexts is constrained less by access to technical knowledge than by the scarcity of practitioners whose judgement translates knowledge into effective action. The traditional mechanism for building that judgement, extended apprenticeship to an experienced practitioner, is not available at the scale required.

Why access alone does not help

A tool that produces outputs for a practitioner to receive does not build the practitioner’s judgement, and the evidence on unstructured access suggests it can erode it. A scaffold designed only to produce better outputs improves output quality and nothing else.

What deliberate design would require

Mandatory reasoning visibility, so the practitioner sees the reasoning rather than only the result. The practitioner producing their own framing before the model’s. Locally accumulated institutional references. Enforced critical engagement rather than passive reception. None of these is implemented as a default in any tool Public Policy Lab knows of; building them is a tractable engineering problem, and evaluating whether they transfer capacity is a research problem not yet undertaken.

What would count as evidence

Paired-comparison evaluation of practitioners on substantive tasks, with a reasoning-visible scaffold, without it, and with alternative configurations. The outcome that matters is not immediate output quality but whether practitioners in the scaffold condition later perform better without AI support, identify drift more precisely, and show measurable change in professional judgement over six months to a year.

Key facts

  • Generative AI without guardrails has been found to harm learning once it is removed (Bastani et al., PNAS, 2025).
  • Public Policy Lab proposes that a reasoning-visible scaffold could build judgement rather than substitute for it, and states this as a proposition, not a finding.
  • The required design features are mandatory reasoning visibility, practitioner-first framing, local institutional references, and enforced critical engagement.
  • The proposed test is paired-comparison evaluation measuring performance without AI support after repeated scaffold use.

Sources

Related

Last checked . If a figure here is out of date, tell us.