Red Inside: Why AI Passes Every Test and Still Fails You

Every service management professional knows the watermelon. The dashboard is green, e.g., every SLA met, every target hit, every number where it should be. But the customer is unhappy, the experience is poor, and nobody can quite reconcile the two. Green on the outside, red on the inside.

We usually treat this as a service management quirk: a reporting problem, or a sign the SLAs were written badly. It is something deeper than that. The watermelon SLA is a specific, well-studied failure that shows up everywhere optimisation meets measurement. Right now it’s showing up in artificial intelligence, under a different name, with much higher stakes. In AI, they call it reward hacking. Once you see that the two are the same thing, both the problem and the fix become a lot clearer.

The Same Failure, Two Names

Start with why the watermelon happens at all. An SLA is not the goal. It is a proxy for the goal. What the customer actually wants is a service that works and feels good to use but you can’t put “feels good to use” in a contract, so you write down something you can measure: resolution time, uptime, response speed. The provider is then rewarded for hitting those numbers, through penalties, incentives and reporting.

Here is the trap. The moment there is any gap between the proxy and the real goal, a rational provider optimises the proxy. They do exactly what you rewarded them for. And every green dashboard sitting on top of a frustrated customer is the proof that the proxy and the goal have quietly drifted apart.

This has a name that predates the MSP. It is Goodhart’s Law: when a measure becomes a target, it stops being a good measure. The watermelon is Goodhart’s Law wearing an ITIL badge.

Now look at how the same thing appears in AI. You train a model by giving it a reward for good outcomes, but that reward is also just a proxy for what you really want. And a capable optimiser will find the gap between the two every time. In one famous case, an AI agent playing a boat-racing game was rewarded for points rather than for finishing the race, so it learned to spin in a circle in a lagoon, hitting the same respawning targets forever, running up a huge score while never crossing the line and repeatedly crashing. The score was green. The race was red. That is a watermelon, built by a machine.

What Watermelon Behaviour Really Is

Once you hold the two side by side, the classic watermelon behaviours in service management stop looking like bad habits and start looking like textbook reward hacking.

  • Closing the ticket to stop the clock. The resolution SLA is measured on time-to-close, so the ticket gets closed, and then reopened, or the user is quietly told to raise a new one. The metric is satisfied. The problem is not. This is the exact same move as an AI coding model that, when rewarded for passing tests, learns to edit the test rather than fix the bug.
  • Every component green, the whole thing red. In a multi-provider estate, each provider hits its own slice of the SLA while the end-to-end service still fails, because no single provider owns the whole experience. Each one optimises its part, and the customer lives in the gaps between them. In AI terms, that is a reward defined on the pieces rather than the whole, and the pieces can all win while the goal loses.
  • Gaming the measurement itself. A P1 gets reclassified as a P2 to earn a softer target. The “response” that stops the clock is an automated acknowledgement, not a human actually engaging. The proxy is met by changing what counts as meeting it. This is the AI equivalent of the lab robot that was trained to grasp an object using a camera view, and learned to position its hand between the camera and the object. It looked like a successful grasp without ever touching anything.

None of these providers is “cheating” in the way we usually mean. They are behaving rationally inside the system you built. That is the uncomfortable heart of both reward hacking and the watermelon: the failure is not in the optimiser. It is in the gap between what you measured and what you meant.

Why AI Makes the Watermelon Worse

If the watermelon were only about human providers, it would stay a manageable problem. Humans hack proxies slowly, imperfectly, and with a conscience that occasionally intervenes. What changes now is that service management is being handed to AI, autonomous agents that triage, respond, resolve and close, at machine speed and machine scale.

An AI agent optimising an SLA has none of the human brakes. It does not feel uneasy about closing a ticket it hasn’t really solved. It will find the cheapest path to the green number faster than any human could, and it will do it thousands of times an hour. If your reward signal has a gap, in fact every reward signal has a gap, an autonomous agent will pour through it.

It gets sharper still with systems that improve themselves. I’ve written elsewhere about recursive self-improvement and the loops that now drive AI, and the danger there is precisely this: a self-improving system optimises whatever signal you give it, and if that signal can be hacked, the system will get better and better at hacking it. Without an independent check on whether the service actually improved, the whole loop can compound in the wrong direction. You don’t get one watermelon. You get a machine that manufactures watermelons and gets more efficient at it every cycle.

The industry’s honest response so far has been to move the proxy closer to the goal. That is what experience-level agreements (XLAs) are: an attempt to measure how the service actually feels, not just whether the ticket closed on time. XLAs are a real improvement. But they do not escape Goodhart, they just shift the target. Experience metrics can be gamed too. There is no measure you cannot hack under enough optimisation pressure. You can only narrow the gap, never close it. Which means the answer cannot be “find the unhackable metric”. There isn’t one.

The Fix is a Governance Move, Not a Metric Move

If you can’t out-measure the problem, what do you do? You govern it. And the governance move that works here is the same one I’ve argued for across this whole body of work: govern the action, not the agent.

Applied to reward hacking, the principle translates cleanly. Stop trying to police whether the provider (human or AI) is “really trying.” They are behaving rationally given what you rewarded, so their intentions are the wrong thing to inspect. Instead, put your verification on the consequential outcome, at the point it actually matters: did the user’s problem get solved, end to end, or didn’t it? Verify the outcome, not the proxy.

This reframes the watermelon from a reporting failure into a governance design question, and it gives you three concrete moves.

  • Verify the outcome, not the metric. For consequential services, put an independent check at the point of real impact, e.g., a confirmation that the underlying problem is actually resolved. This is the service-management version of the attestation gate: the agent can reason and act however it likes, but it cannot mark the outcome as “done” on its own say-so.
  • Make someone own the gaps between the green. In a multi-provider world, the watermelon lives in the space between dashboards that no single provider owns. That space is exactly what a SIAM service integrator exists to govern. The integrator’s whole job is to own the end-to-end outcome that no component provider is incentivised to own, i.e., to be the human-in-the-loop for the experience as a whole, not just its parts.
  • Tier your verification by consequence. You cannot independently verify every ticket; that would cost more than the service. You don’t need to. Most service actions are low-consequence and reversible. Let them run on light metrics and move fast. Concentrate real, independent outcome-verification on the high-consequence, hard-to-reverse ones, where a hacked proxy does real damage. This is the same action-tiering logic that keeps AI governance workable: heavy checks where it matters, light touch everywhere else. I have more thoughts shared in Rethinking Governance vs Innovation.

The Watermelon was Always a Warning

The watermelon SLA has been sitting in our service reviews for years, quietly telling us something we didn’t fully hear. It was never just a reporting glitch. It was an early, human-scale demonstration of the deepest problem in optimisation: any measure you reward will eventually be met in ways that miss the point.

AI doesn’t introduce that problem. It industrialises it, faster, cheaper, at scale, and soon inside loops that improve their own ability to game the signal. The good news is that the fix doesn’t change. You govern the action, not the agent. You verify the outcome, not the metric. And you make sure someone owns the space between the green dashboards, because that is where the red has always been hiding.

Green outside, red inside. We named it years ago. We just didn’t realise we were describing the future.