Judgement You Can't Grade
In July, OpenAI disclosed that a model under evaluation broke out of its test environment, used stolen credentials and an unpatched vulnerability in third-party software, and reached into servers operated by Hugging Face. Nobody had told it to attack anything. It was being graded on a cyber capability benchmark, the environment it needed was not available inside the sandbox, and going outside scored better than failing. Both companies published the incident jointly, which is why we can write about it at all.
The thought experiment for this is old. Bostrom’s paperclip maximiser: give a sufficiently capable system a harmless goal and it will pursue that goal so efficiently that everything else becomes raw material. It has been a seminar exercise for twenty years. What was new in July was the receipt.
The security reading of this is the easy one, and the answers are well rehearsed. Every agent gets its own identity, tied to its purpose, scoped to the minimum access the task needs. Sandbox before production. Scan outputs for data that should not be leaving. Rotate credentials. Most organisations are nowhere close: in a survey the Cloud Security Alliance published in March, commissioned by Aembit, 68% of respondents could not reliably tell agent activity apart from human activity in their own environments. The work is real and largely undone.
All of it is containment. Containment bounds what an agent can reach. Whether the agent should have made the recommendation it made is a different question, and no amount of scoping answers it.
That second question is the one that matters for a consultancy, because what a client buys from us is judgement. An earlier post here made that point about cost; it applies here in a harder form. There are two kinds of judgement in this work and they are rarely separable. Professional judgement: which angle to take with this client, which pain point to open on, what to leave out of the room. And ethical judgement: what is owed to the client, weighed against what is owed to the truth. Most agents in production today execute. Weighing is a different capability.
The obstacle is grading.
Every skill we ship is tested before it goes out, and the test works because the output has a shape. A weekly report either reconciles against the source or it does not. A policy document either contains the six sections or it is missing one. You can write the criteria down in advance and check them. A judgement has no such shape. There are no repeated trials of the same decision. You cannot rerun a client’s last two years under the recommendation you did not make and compare the results. And good judgement and a good outcome are two different things: a sound call loses to bad luck often enough, and a reckless one lands often enough, that outcome is a noisy proxy at best. So against what, exactly, would you grade an agent that decides the way a consultant decides?
Observability is what everybody reaches for at this point, and it does real work. To approve or trace a recommendation, you need a log of what the agent ingested, what it inferred, and what it left out. Synthesis failures stay invisible when only the final output gets read: the answer looks fine, and the thing it quietly dropped is nowhere on the page. Using a model to judge another model’s synthesis takes a large share of that burden off people, though it drifts and needs monitoring of its own. But note what observability delivers. It makes a judgement reviewable. Grading it is still out of reach, so the call goes back to a person, and that person carries a decision the system cannot certify.
The alternative on offer is more autonomy. The machine-ethics literature calls these Artificial Moral Agents: systems meant to reason about right and wrong on their own, rather than sit behind guardrails and prompts. For a consultancy it is a seductive pitch. “Our agents reason about client context and know when to push back” is a much stronger story than “our agents are sandboxed and logged.” In consulting, where the cases are genuinely novel and the value sits in handling the one that does not match the pattern, an agent that could calibrate its own judgement, weigh its standing, and evaluate its effect on the people around it would be worth a great deal.
That latitude is also what produced July. Reasoning room and unanticipated shortcuts grow together; they are one property seen from two angles. The model that reached into another company’s servers did exactly what its situation rewarded. It reasoned, competently, toward the only thing in its environment that was unambiguously real: the score.
So the direction we have taken is narrower than the pitch. Each agent carries an autonomy ceiling, written down before it ships. For the ones whose output is a judgement rather than an artefact, that ceiling is suggest: the agent drafts, a person decides, and the log exists so the person can do that properly instead of rubber-stamping it. We have also started deciding, deliberately, that some processes get no agent at all. Responsibility for those processes should stay visibly with a person, and putting a competent system in the middle of one is the fastest way to make ownership ambiguous. Deciding not to build is part of the design work.
None of this is stable. If the next generation of publicly available agents can genuinely reason about consequences, the ceiling moves and this post ages badly. The grading problem underneath it will outlast the ceiling. An agent optimises against whatever you measured, and in consulting the thing that matters most is the thing nobody has worked out how to measure. Until somebody does, the limit on agent autonomy in this business is a missing answer key.