Show Me the Worse One


A five-word governance test that proves a human actually read the AI’s output.

The Rubber Stamp Problem

Somewhere in your organisation today, someone asked an AI to write something, glanced at the output for three seconds, and pasted it into an email, a report or a client deliverable. They will tell you a human reviewed it. Technically, a human was in the room.

This is the quiet failure mode of “human in the loop.” The loop exists on the org chart. It exists in the policy document. But in practice, the loop has become a rubber stamp, and a rubber stamp leaves no evidence of whether anyone actually engaged with the content before it went out the door.

We have all seen where this ends up. Reports quoting figures that were never in the source data. Marketing copy that confidently describes a product feature that does not exist. In every one of these cases, a human “reviewed” the output. The problem is that review, as most organisations practise it, is unverifiable. A checkbox that says reviewed by a human proves only that someone clicked a checkbox.

A Different Kind of Ask

Here is a simple alternative, ask the person one question:

Show me the worse one.

That is the whole mechanism. And it changes everything about what the person has to do next.

You cannot identify the worse of two documents without reading both. There is no shortcut. To make the call, you have to notice which one drifts off the brief, which one has the weaker argument, which one buried the key number, which one sounds like it was written by a committee.

The act of comparison forces the act of reading and the act of judging forces you to form criteria. Suddenly you are not a courier moving text from one window to another. You are an editor.

Governing the Action, Not the Agent

And here is the part that matters for governance: the selection itself is evidence. When someone picks version A over version B, they have produced a small, concrete artifact of human judgment. Add one more line: why did the loser lose? Then you have an audit trail, not a checkbox. An actual trace of a human mind engaging with the content before it shipped.

Readers of this blog will recognise the shape of this idea. In A Practical Starting Point for AI Governance, I wrote about a reframe that breaks the governance deadlock where many organisations are stuck in: stop trying to police the AI’s reasoning, instead place your controls at the point where an action becomes consequential. That is how we govern not by monitoring AI’s thinking, but by requiring attested evidence at the moment of action.

“Show me the worse one” is the same logic.

The consequential action was never the AI generating text. Text in a chat window is harmless. The action that matters is a human choosing to publish it, to a client or a colleague. That is the moment governance should engage. And rather than gating that moment with a policy nobody reads or a checkbox nobody respects, we gate it with a task that cannot be completed without genuine engagement.

In other words, the comparison is a lightweight attestation. The person is not just approving an output; they are demonstrating, through a choice that requires reading, that human judgment was applied before the action was taken.

Why Comparison Beats Inspection

There is a reason this works better than simply telling people to “review carefully.” Inspection is passive and unbounded. Asked to review a single document, most people skim for obvious errors, find none in ten seconds and approve. There is no natural finish line, so the finish line becomes whenever I feel done, which is immediately.

Comparison is active and has a built-in completion test. You are either able to name the worse one and say why, or you are not. The task is binary and checkable. It also quietly trains people: after a few weeks of picking winners and losers, they develop a sharper sense of what good AI output looks like for their context, e.g., where the models pad, where they hedge, where they invent. Your reviewers get better at reviewing because the mechanism made them practise.

There is also a subtle cultural signal at play. “Review this” says the AI’s work is probably fine. “Show me the worse one” says the AI’s work varies and your judgment is the thing that separates acceptable from excellent. It restores the human to the governance role.

Making It Work in Practice

A few practical notes for teams who want to try this:

  • Scope it to what leaves the building. Internal brainstorms and rough drafts do not need it. Anything client-facing, public or feeding a decision does.
  • Generate genuinely different versions. Ask the AI for two takes with different structures or emphases, not the same draft twice. The more the versions diverge, the more reading the comparison demands.
  • Keep the loser. The rejected version plus why it lost is your audit record. It doesn’t cost much time but it is worth more than a thousand signed policy acknowledgements.

Accept the limits honestly. Someone determined to rubber-stamp can pick a version at random. And for genuinely high-stakes, irreversible actions, e.g., deployments, prescriptions, payments, this is not enough; those need the stronger multi-party attestation controls. This mechanism is for the vast middle ground of everyday AI-assisted work, where today the alternative is usually nothing at all.

Governance Can Start with a Question

The organisations I meet are still stuck in the chicken-and-egg deadlock: one team wants to adopt AI faster, another wants a governance framework first but cannot say what should be in it. The last article argued that the way out is to govern actions, not agents. This article offers the smallest possible starting point for doing exactly that.

You do not need a framework to begin. You need to change one question in one workflow. The next time someone brings you AI-generated work, do not ask whether they reviewed it.

Ask them to show you the worse one.