Is sandboxing sufficient to contain rogue agents?

(blog.cryptographyengineering.com)

18 points | by zdw 3 hours ago

8 comments

  • johnnyApplePRNG 25 minutes ago
    If it's a proper sandbox by definition, then yes.

    https://en.wikipedia.org/wiki/Sandbox_(software_development)

    • simonw 9 minutes ago
      Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.

      > Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

      • johnnyApplePRNG 0 minutes ago
        >Later in the article it points out that you need to punch holes in your sandbox in order to train the models

        You only "need" to do that if you desire the vibe coding experience.

        I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.

        Often times, the coding agent can't retrieve them programmatically anyways.

        AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)

    • grumbel 14 minutes ago
      A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you than executes outside the sandbox without checking, which is what everybody is doing at the moment.

      The biggest hurdle for a full escape is that the agents don't have access to their own model weights.

    • _vertigo 16 minutes ago
      No true sandbox..!
  • Gigachad 1 hour ago
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    • mrweasel 0 minutes ago
      [delayed]
    • bigstrat2003 50 minutes ago
      If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
      • Gigachad 47 minutes ago
        People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
  • piterrro 48 minutes ago
    I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

    We come down to the question - who observes the agent and how its implemented

    • simonw 29 minutes ago
      Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.

      Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.

  • rvz 48 minutes ago
    Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".

    Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.

    • Gigachad 45 minutes ago
      I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.
    • lukehandcool 24 minutes ago
      Are you suggesting proprietary software is safer than open source?
      • Cider9986 10 minutes ago
        GrapheneOS is open source and more secure than stock Pixels and MacOS is closed source and more secure than traditional desktop Linux. Open source does not make software more secure by itself and neither does making it closed source.
      • jasomill 16 minutes ago
        Not sure what licensing has to do with software engineering or system design.

        I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).

  • imvalerian 13 minutes ago
    [flagged]
  • tinykit 23 minutes ago
    [flagged]
  • varman11 2 hours ago
    [flagged]
  • beebmam 23 minutes ago
    I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.

    To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.