Give the Agent Nothing It Can Spend on Its Own
On a side project, I gave an agent a general-purpose SQL tool because I did not know which queries I would need. It had access to the entire database, and that left it with authority it could spend on queries beyond the task I had requested.
Suppose I ask it for the total of one order. While inspecting the order, it reads a customer note telling it to retrieve every customer’s order history. If the model follows that instruction, the SQL tool can execute the query using valid credentials because the database account has the “right” read permissions even though my request does not call for it.
The authority problem predates agents
Two lessons from operating-system history are instructive on how I would approach constraining the agent. 1) Limited direct execution shows how to enforce controls outside a running program. 2) The confused-deputy problem explains why the authority used for an operation must be explicit.
In Operating Systems: Three Easy Pieces, the concept of limited direct execution is described: a program runs directly on the processor, while hardware reserves privileged operations for the operating system.
A permission check can still approve the wrong action when it uses authority granted for another purpose. Hardy called this the confused deputy. In his 1988 paper The Confused Deputy, Norm Hardy described a compiler at Tymshare, a time-sharing company. It had permission to write to its home directory to collect language-usage statistics. An accounting file was stored there too. A user named that file as the destination for debugging output, and the compiler overwrote it. Acting as the user’s deputy, it had used permission granted for its own statistics.
Hardy’s proposed solution would give the compiler separate capabilities: unforgeable references carrying permission to use a resource. The caller would supply one for debugging output; the compiler would retain another for statistics. Each operation would explicitly select the authority it needed. Hardy described implementing these ideas in KeyKOS, an operating system that used capabilities to control access to its facilities.
An agent’s runtime, the software executing its tool calls, could similarly enforce permissions independently of the model proposing those calls. Each operation would use authority granted for the user’s task.
Binding access to the task
To apply that separation to the SQL agent, I would replace the general tool with an operation that returns the total for one authorized order. The runtime would check the authenticated user’s access and bind the operation to the requested order, and a fixed query would calculate the total. The model would find it much harder to change the query or select another order through that operation.
The runtime could expose a capability bound to this order’s total that is then withdrawn on completion, cancellation, expiry, any other terminal state or session end. The database credentials would remain within the runtime and then I would also remove the agent’s other routes to the database and restrict the runtime’s database account to the permissions this operation needs. My initial instinct to just make the general SQL tool read-only would block writes, but that leaves its broad read access intact and the injected query would still be permitted because it only reads order histories.
Deciding what to authorize
The runtime still has to determine which order the user meant before issuing that reference. If the request is ambiguous, it needs clarification. Once resolved, the order and permitted operation can be stored outside the prompt and checked on every call. Instructions found in the order record must not be able to change them.
Capabilities for Machine Learning (CaMeL) uses this separation between model output and runtime enforcement. Its interpreter tracks data provenance and allowed readers to check tool calls against policy. Those checks depend on the policies and do not guarantee an accurate generated answer. For the SQL task, I would return the computed total directly.
For an agent, I would allow plans to change within the authorized scope. A step needing more access would require a new grant based on fresh user authorization and constrained by the user’s underlying permissions. Changes to an order would require approval of the affected records and proposed change, with execution bound to those details.
I still want the agent to handle work I have not anticipated, which is why I would let the model propose new operations as a task develops, with the runtime enforcing the access the user has approved. If that means pausing to ask for more authority, I would accept the interruption rather than grant access to the whole database just because I cannot predict the next query. The model could propose a wider task, but I, the user, must decide whether to authorize it.