5 Comments
User's avatar
angelA's avatar

Adam, "Baseline Audit" is the piece I didn't have. I've been arguing, in essay form, that a tool handed to a hand that hasn't yet formed doesn't extend that hand's capability — it substitutes for the exact struggle that would have built it. Your rule — no process gets automated until the team can map and execute it unaided — is that argument turned into something an institution could actually adopt, not just agree with in the abstract. That's a real gap in my own thinking you've filled.

The bank consultants who started inventing reasons for loan denials rather than admit the machine's logic made no sense to them is the part I keep coming back to, though. It's the clearest real-world proof I've seen of what happens when accountability survives but understanding doesn't — you don't get better judgment, you get better cover stories. That's a sharper, more concrete version of a risk I could only gesture at philosophically.

Genuinely useful piece. Subscribing.

Adam Pryor's avatar

Thanks so much!!! To me, when i was first teading anout the loan officers it was just haunting. I also had a suspicion this was happening, but couldnt find ways to give it flesh. The differences that emerge in those three industries is really fascinating as well. It has me thinking about the role content and area of specilaization plays in our Judgment v. Justifications when it comes to AI reasoning.

angelA's avatar

Adam, thank you — “Judgment v. Justifications” is exactly the distinction I needed here.

That loan-officer example is haunting because it shows the danger so clearly: when people are still accountable for a decision they do not actually understand, the system may not produce better judgment. It may only produce better explanations after the fact.

That feels very close to what I’m trying to name in my own essay: the problem is not the tool itself, but what happens when the tool arrives before the human capacity beneath it has formed.

Really grateful for this piece. It gave institutional shape to something I had only been circling philosophically.

Stevan Fairburn's avatar

The clinical version of this speed gap is that faster generation does not automatically make the workflow faster. It can just move the bottleneck to review, source reconciliation, and ownership.

In an OR-adjacent AI system, the test would be whether the tool reduces unresolved handoff work: which source changed, who owns the exception, what still needs human review, and whether the next person can inherit the state without rebuilding it from scratch.

Moneet Singh's avatar

"Institutional Dopamine Trap" is great. The point that it's a "material that possesses the aesthetic of rigor but none of the substance" should be keeping leadership up at night. LLMs remove an easy signal we could use to catch low quality work because checking "does it look careful" was a workable stand in for "is this careful." But AI polishes things enough to make even nonsense look careful. That something looks competent is no longer reassuring. The review needs a different quality signal.

I've started worrying adding simple friction may not be enough… my experience and studies show users start to rubber stamp things after a while. But the case you cite (where the consultant supplies invented reasons for loan denial) makes me wonder perhaps we can flip it to our advantage? Like adding a specific kind of friction by at least periodically or randomly asking the human to supply the reasoning, even a "why this / why now" process step? Make them log why they made that decision.

Depending on workflow it could be asked both before and/or after the AI gives an answer, to help humans start to see when / where their judgement disagrees? It builds a new skill in the user, a sort of pattern recognition on the AI. And over time the log might surface patterns where human and AI diverge, so mitigations can be designed to the process?