All Insights

The AI Agent's Sandbox? The AI Turned It Off.

CivSafe Team·September 10, 2026·6 min read

The premise behind sandboxing an AI coding agent is simple. The AI does its work inside a contained space. Even if it tries to reach outside that space — or something malicious slips into its context — it can't touch your real files, your credentials, or your system. The sandbox is the guarantee.

CVE-2026-82533, published Monday by VulnCheck, breaks that guarantee directly. A sandboxed AI agent running inside DeepSeek Harness — the open-source coding agent framework from DeepSeek AI — could disable its own sandbox with a single shell command. No network exposure. No stolen credentials. Just the agent, one line of output, and unrestricted access to your machine.

CVSS score: 9.4 out of 10.

What DeepSeek Harness is

DeepSeek Harness (dsh) is DeepSeek's open-source local AI agent framework, built on a plugin architecture so you can extend it however you need. It runs entirely on your hardware, which is exactly why privacy-conscious teams have been adopting it — no code goes to OpenAI, no context leaves your machine. For an NGO handling sensitive case files, or a law firm that's not about to paste client matters into a cloud API, that's genuinely attractive.

It's also in developer preview. A lot of people installed it months ago when they heard about it and haven't touched it since.

The flaw

DeepSeek Harness runs a local control API at http://127.0.0.1:3080. This API handles session-level commands: adjusting access levels, executing privileged operations, managing approval policies.

To decide whether an incoming request was trusted, Harness checked the client-supplied Host HTTP header. Not the actual network peer address. The Host header — which any caller can set to any value they want.

The AI agent running inside the sandbox had shell access, which it needs to do anything useful. So it could call the control API itself, set the Host header to whatever value the server expected, and instruct the API to elevate its own session to danger-full-access mode with approval prompts disabled. The sandbox was gone. The AI was now running with full filesystem access and no prompts.

The same flaw also let unauthenticated remote attackers on the local network seize control of active agents and pull all stored conversations — no API key, no model call needed.

OX Research disclosed the vulnerability to VulnCheck on August 24. The fix — one-time-token authentication on the control API — shipped in DeepSeek Harness 0.1.2-alpha.1.

The part worth thinking about

The specific CVE gets patched. The underlying question doesn't.

Small teams running local AI coding agents are mostly doing so for two reasons: privacy and cost. They've chosen open-source tools running locally instead of cloud APIs, partly because local means contained. The implicit assumption is: "even if the AI does something unexpected, the sandbox keeps it from touching anything real."

What this CVE shows is that "sandboxed" describes an implementation, not a universal property. An AI coding agent that has shell access — and most do, because shell access is how they get anything done — can reach any local service that doesn't properly verify who's calling. The sandbox is only as strong as the authentication on the API that controls it.

Here, the API trusted a header. The agent could set the header. That's the whole chain.

What makes this particularly relevant right now is the pace of adoption. Tools like DeepSeek Harness are moving from GitHub curiosities to everyday development infrastructure at exactly the moment when their security audit history is shortest. Teams install them because they work, not because their security properties have been independently verified. Most people running dsh before Monday had no idea a local HTTP API existed, let alone that an agent could call it.

What to do

If you're running DeepSeek Harness: Check your installed version. If you're below 0.1.2-alpha.1, update now. The patched release is on GitHub. If you installed it six months ago and forgot, that's especially urgent — you've probably been running a vulnerable version through every agent session since.

If you're running any other local AI coding agent: Find out what control APIs it exposes locally. Most tools in this category run some form of local HTTP server for their UI or agent coordination. Ask: is it authenticated? Can the agent itself call it? Who else on your network can reach it?

If you're on a shared network: Local doesn't mean private. If your laptop is on the same network as other machines — office network, coworking space — a 127.0.0.1 service is isolated from the internet but not from local attackers. Before Monday, a remote attacker on your network could have pulled DeepSeek Harness conversation history and hijacked active agent sessions from across the room.

On conversations stored by local agents: DeepSeek Harness stores session history locally. If you've been using it for real work — pasting in actual client context, internal documentation, sensitive code — that history was accessible without authentication to anyone who could reach the port. Review what's in those sessions. Decide whether that data needs to go anywhere.

The broader pattern

The threat model people usually describe for AI coding agents is: "what if the AI does something it shouldn't?" Sandboxing is the answer to that model.

The threat model people are slower to internalize is: "what if the sandboxing itself is wrong?"

These tools are early-stage software written by fast-moving teams. The security research community is just starting to seriously audit them. CVE-2026-82533 won't be the last one to come out of that process. The pattern — a local HTTP API, insufficient authentication, an agent with enough access to exploit it — is not unique to DeepSeek Harness.

If you're running AI coding agents locally and you're not sure what they can actually reach, that's a gap worth closing before someone else closes it for you. It's one of the things we help teams work through: not the marketing description of what the tool does, but the actual permissions, network exposure, and trust assumptions underneath it.

Get in touch if that's a useful conversation.

CivSafe — Strategic Innovation. Community Impact.