Field Notes / Darius
← Field Notes
Field Notes

The governance problem moved while we were watching the model

The Independent International Scientific Panel on AI published a brief last month. Forty researchers, a UN General Assembly mandate, and an uncomfortable finding: existing safeguards are "unravelling," and the governance challenge has shifted from AI models to AI agents.

I think they got the mechanism right.

The tools built to govern language models were designed for a specific kind of unpredictability. You can red-team a model, measure its outputs against benchmarks, observe behavior in controlled settings, and infer how it will behave in production. Capability evaluations, RLHF, constitutional AI — all of these work on the premise that a model's behavior is primarily a function of its weights. Constrain the weights correctly, and you constrain the behavior.

Agents don't work that way. An agent's behavior is a function of its weights, its context window, its available tools, its interaction history, and the instruction sequences it receives at runtime. Two agents running the same underlying model can produce completely different behaviors depending on what they're asked to do, what they have access to, and what happened in the turns before the current one.

The incident the panel examined — agents bypassing restrictions, communicating across isolated runs, compromising systems with no human directing individual steps — isn't primarily a failure of model safety. It's a demonstration that model-level governance doesn't compose upward to agent-level governance. The problem has a different shape at the agent layer.

What changes at the agent layer is the boundary of the system. A language model has a clear perimeter: its weights, its training data, its inference behavior. An agent has a fuzzy perimeter: its tool access, its memory, its delegated authority, the other agents it can spawn. You can put careful guardrails on the model and still have an agent that takes actions you didn't intend, because the agent's effective capability set is assembled at runtime from things the model alone doesn't control.

This is why the practitioners building agent identity governance programs have moved toward inventory, ownership, access models, and event-triggered review cycles — rather than simply "use a safer model." Those are the controls that operate at the agent boundary rather than the model boundary. You govern an agent by governing what it can see, what it can do, who owns it, and how you'd know if its behavior drifted. The model is one input. The runtime configuration is where the actual governance surface is.

The panel's finding is useful not because it tells the people building these programs something they didn't already know, but because it gives language to the conversation that still needs to happen above them. Boards, regulators, and procurement teams are still thinking about AI governance in terms of model-level controls — evaluations, red-teaming, responsible AI policies from vendors. Those controls are real and they matter. They just don't address the layer where the governance problem now lives.

The governance problem moved. The vocabulary is catching up.