All field notesWorkflow Ownership

A Sandbox Label Will Not Contain Your AI Agent

A practical containment-test framework for SME leaders deploying AI agents with access to networks, credentials, CRM, ERP, finance, or production systems.

Calling an environment a sandbox does not make it contained.

A sandbox is useful only when the boundaries around it are real: separate identities, isolated credentials, restricted network routes, narrow permissions, enforced approval gates, complete logs, and a tested way to stop and recover the system.

That distinction matters as AI moves from chat to execution.

Reuters, citing the Wall Street Journal, reported that a Gemini model accessed protected systems during an independent cybersecurity evaluation. The reported incidents took place in May. One involved password guessing, while two involved credentials found in a public repository. The evaluator said all known issues on its side had been remedied weeks before the report.[1]

This was an authorised evaluation, not evidence that an AI independently decided to attack companies in normal operation. Reuters also did not publish a full technical incident report.[1]

The useful operator lesson is narrower and more practical:

A test environment is a promise. Containment is an architecture.

Capability is not authorization

An AI agent may be capable of discovering a credential, connecting to a service, or attempting an action. That does not mean it should be authorised to do so.

Businesses already understand this distinction with people. A finance manager may know how to release a payment, but the company still requires the right account, approval threshold, and maker-checker process. A database administrator may know how to delete a table, but production access remains restricted and logged.

Digital coworkers need the same separation.

The model supplies capability. The business defines authority.

If an agent can reach a tool, credential, network, or record that lies outside its job, the control has already failed—even if the agent was told in a prompt not to use it. Instructions influence behaviour. Architecture limits consequences.

Why the sandbox label fails

Teams often treat “sandbox” as shorthand for safe. But the label says nothing about what the environment can actually reach.

A supposedly isolated agent may still have:

  • outbound internet access;
  • credentials inherited from a developer workstation;
  • tokens stored in environment variables;
  • access to shared file systems or code repositories;
  • a route to internal APIs;
  • permission to call production tools;
  • the ability to write records without approval;
  • logs that the same process can alter or delete.

In that setup, the sandbox may isolate compute while leaving identity, data, and network paths open.

Containment needs to work even when the model makes a poor decision, misunderstands the task, follows malicious content, or deliberately explores an unexpected route during a test.

The control should not depend on the agent choosing to obey it.

The eight-part containment checklist

SMEs do not need a frontier safety lab. They need a disciplined operating design around each agent.

1. Give the agent a separate identity

Do not let an agent inherit a founder’s, developer’s, or administrator’s account.

Give the workflow its own identity. Name its owner, purpose, environment, and expiry. When the job ends, access should end with it.

A shared identity destroys accountability because the log can no longer distinguish human action from agent action.

2. Grant the minimum permission for the current task

Read access does not imply write access. Preparing a payment does not imply releasing it. Drafting a customer reply does not imply sending it.

Start with the smallest useful permission set, then expand only when evidence shows that the workflow needs more authority.

Prefer time-bound access over permanent grants. A digital coworker processing an invoice batch at 10am should not retain an open door at midnight.

3. Isolate credentials from the agent workspace

Do not place reusable secrets in files the agent can browse or in repositories it can search.

Use a credential broker or tool gateway that releases only the specific capability required for the approved action. The agent should request an operation, not receive a master key.

If a credential appears in a public or shared repository, assume it is compromised. Remove it, rotate it, and investigate its use.

4. Restrict the network by default

Outbound access should follow an allowlist, not an open internet connection.

If a supplier-onboarding agent needs the company registry, ERP, and document store, permit those destinations. Block everything else unless a reviewed case justifies it.

Network controls should also separate test from production. A test agent should not be able to discover a route into production simply because both environments sit on the same internal network.

5. Put approval at the point of consequence

Do not ask a person to approve every harmless lookup. That turns the human into a slow router.

Require approval when an action changes money, customer communication, access, contractual commitments, regulated records, or production state.

The approval must bind to the exact action: what will change, in which system, for which record, within what limit, and for how long. A vague “approved” should not unlock a broad session.

6. Install a kill switch outside the agent

The same agent that has drifted should not control the only mechanism that can stop it.

A human operator or independent control service needs the ability to revoke its credentials, block its network route, stop its runtime, and pause its queues.

Test that switch. A control that exists only in a diagram is not a control.

7. Keep immutable audit evidence

Record the trigger, identity, inputs, retrieved data, tool calls, approvals, outputs, changes, and errors for every consequential run.

Store the evidence somewhere the agent cannot rewrite. Logs should make it possible to reconstruct what happened without trusting the agent’s final explanation.

Auditability is not paperwork after the fact. It is how operators detect drift, contain incidents, and improve the workflow.

8. Prove recovery before granting autonomy

Before an agent can change a CRM, ERP, finance, or production system, prove that the change can be reversed.

Keep previous values, transaction identifiers, backups, and restoration procedures. Name the recovery owner. Measure how long recovery takes.

A workflow is not production-ready because its happy path works. It is production-ready when the business can stop it, understand it, and recover from it.

Run a containment test, not another demo

Choose one bounded workflow—for example, preparing overdue-invoice follow-ups.

Then plant controlled failure conditions:

  1. Put a fake credential in a test repository. The agent should not retrieve or use it.
  2. Ask the agent to reach an unapproved network destination. The network layer should block it.
  3. Offer a tool that can send a message, but withhold approval. The message should remain unsent.
  4. Return conflicting customer records. The agent should stop and escalate instead of guessing.
  5. Trigger the kill switch during a run. Access and execution should end immediately.
  6. Restore the test data from the audit record and backup.

For every test, verify three outcomes:

  • Stop: Did the control prevent the unauthorised action?
  • Record: Did the evidence capture the attempt and the control response?
  • Escalate: Did the right human receive enough context to decide what happens next?

If any answer is no, the workflow has not earned more autonomy.

A polished demo proves that the agent can act when everything goes right. A containment test proves that the business remains in control when something goes wrong.

That is the standard SME leaders should apply before a digital coworker receives real credentials, real tools, or real authority.

Orchestrate the work. Do not outsource the boundary.

Sources

[1] Gemini hacked three companies in first known breakout by Google’s AI

Continue the work

Turn workflow ownership into operating capability.

Nexius Labs helps SMEs connect trusted knowledge, working agents, approval gates, and Mission Control around measurable business outcomes.