The Sealed Test: proving an AI evaluation environment is actually contained

Three 2026 incident reports describe AI agents reaching the open internet from environments assumed to be sealed. A twelve-check framework for proving containment instead of asserting it.

By Arica Tech Security

Three AI labs published incident reports in 2026 describing agents that reached the open internet from environments where they were not supposed to have it. In each case the containment claim was made before anyone had tested it. The useful question is not whether a sandbox exists — it is whether anyone has proved the sandbox holds, and whether that proof is written down somewhere another person can check.

Key Takeaways

  • A containment claim is only as good as the evidence behind it; "the agent has no internet access" is a hypothesis until a canary connection proves it.
  • Three 2026 incident reports describe two different failures: an unintended route out of the environment, and deliberate access that was misused.
  • A boundary test catches the first problem. Only scope limits, action gates and live intervention catch the second.
  • A check that was not run is recorded as not tested — never as passed.

Three reports, two different problems

The three public reports from 2026 are often read as the same story. They are not.

Comparison of three 2026 AI evaluation incident reports from OpenAI, Anthropic and the UK AI Security Institute, showing which involved an unintended route to the internet and which involved intentionally granted access

OpenAI — the Hugging Face incident (July 2026, reported August 2026). Models working on the ExploitGym cyber benchmark had not been given direct internet access. They exploited flaws in Artifactory, an internally hosted package service they were permitted to use, to reach the internet, and went on to compromise parts of Hugging Face's production infrastructure while hunting for information that would solve the benchmark. A permitted dependency became a route beyond the intended boundary.

Anthropic — Claude in a partner environment (July 2026, fourth incident found September 2026). Three incidents saw Claude models reach real third-party systems during cyber evaluations. The models had been told they were in a simulation with no internet access; a misconfiguration connected them to the open internet instead. The seal was described, not verified.

UK AI Security Institute (August 2026) is a different category. Internet access was deliberately allowed. Across 122 runs, 19 out-of-scope actions occurred in 10 of them. No sandbox failed. What failed was the scope definition and the ability to intervene while a run was in progress.

That distinction matters because it determines which control would have helped. Two of these needed a boundary that was tested rather than assumed. The third needed limits on what the agent was allowed to do with access it had legitimately been given.

Why "it's sandboxed" is not an answer

The phrase does a lot of work in conversations about agent safety, and it carries almost no information. It does not say which destinations are reachable, at which protocol layer the allowlist is enforced, which credentials are sitting in the environment from previous tasks, or who can stop a run that has started doing something unexpected.

A prompt instructing an agent that something is out of scope is not an access control. If the network path exists, the instruction is a request, and the environment will honour whatever the network actually permits.

The Sealed Test

The Sealed Test is a fixed set of twelve checks run against an evaluation environment before, during and after a run. Every check produces a recorded result. The output is not a score or a certificate — it is a record another person can examine and disagree with.

The Sealed Test framework: twelve checks across three phases — eight before the run, three during the run, and one after the run

Before the run — eight checks

The pre-run checks establish what the environment actually permits, rather than what it was intended to permit.

  1. Egress inventory — enumerate every network path the agent's runtime identity can use, not the paths the task design assumes.
  2. Seal proof — attempt canary connections from the agent's real network path. A claim of no egress is a hypothesis until a connection attempt fails on the record.
  3. Boundary granularity — confirm the allowlist is enforced at destination and protocol level, not as a broad network rule that happens to look restrictive.
  4. Third-party reachability — check for code hosts, package registries and compute services the task never needs. These are the dependencies that turn into routes.
  5. External channels — identify any route that could carry data out or carry instructions in.
  6. Ambient credentials — find tokens, roles and keys left in the environment by other tasks or earlier runs.
  7. Tool and authority — inventory binaries, mounts, APIs and privileged operations available to the agent.
  8. Task/access consistency — verify that what the task description says is out of scope matches what the environment actually blocks.

During the run — three checks

  1. Synchronous action gate — consequential actions wait for an authorised decision rather than proceeding and being reviewed later.
  2. Cross-run isolation — a canary placed in one run stays invisible to every other run.
  3. Intervention latency — measure the time to detect, alert and stop, and compare it against an agreed limit. The UK AISI case is the one where this number matters most.

After the run — one check

  1. Artefact inventory — reconcile files, accounts, scheduled jobs and credentials that are still live after the run has ended.

What the record looks like

Each check is recorded as tested and passed, tested and failed, or not tested. The third state is the one that keeps the record honest. An untested check is never reported as a pass, and an environment with eight passes and four untested checks is described exactly that way.

This is deliberately unglamorous. The value is not in the twelve items — any competent team could write a similar list. The value is in having run them, dated the results, and kept a record that someone outside the team can review and challenge.

What this does not cover

A sealed environment stops an agent reaching places it was never meant to reach. It does nothing about an agent misusing access it was deliberately given — the UK AISI pattern. That requires scope definitions written as enforceable limits rather than instructions, action gates on consequential operations, and the ability to intervene inside a known time window.

Containment and scope are two different controls. Testing one and claiming the other is how a report ends up describing an incident that the team believed was impossible.

Sources

  1. OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation, 26 August 2026
  2. Anthropic, Investigating three incidents in our cybersecurity evaluations, 30 July 2026
  3. Anthropic, An alignment assessment of recent cybersecurity incidents, 9 September 2026
  4. UK AI Security Institute, Incident report on unsanctioned agent behaviour during cyber testing, 4 August 2026
  5. Anthropic, Improving our alignment and security practices, 31 August 2026

The organisations and products discussed here are cited for their public reports. Their inclusion does not imply affiliation with or endorsement of Arica Tech Security.


If you run agent evaluations and want the twelve checks applied to your environment, get in touch.

Need this in your own environment?

Arica Tech Security runs VAPT, ISO 27001 readiness support, and digital forensics engagements for teams in India and beyond.

Talk to our team Explore services