harness.fail

A record of security issues that AI agent harnesses failed to prevent.

harness

1. noun: a deterministic program that wraps a model, constructs and manages its context, and mediates the model's inputs and outputs with the surrounding environment (e.g. the OS, tools, files, network, or user).

Every issue on this page happened on the far side of a harness.

Meaning any and all harnesses. When a harness delegates a decision to a model, that decision is no longer deterministic.

fails

1. verb · third person singular present of fail: to not do what it exists to do, or what it is expected to do.

When the harness fails, restricted files enter model context.

2. noun informal · plural of fail: cases of failing.

This page is a collection of harness fails.

Both readings are intended. Sometimes the harness broke; sometimes it was never there.

Contents

The record

Updated October 6, 2026 · 74 issues · 7 classes · JSON · DOI 10.5281/zenodo.23197048

Some frameworks assume a bad actor and classify the attack (how it is built, what it wants); others classify the harm (what went wrong in the world). harness.fail classifies the failure: the boundary the harness itself was supposed to hold, but didn't, regardless of intent. That is why two of the seven classes, Scaffolding collapse and Ambient activation, have no counterpart in OWASP or MITRE ATLAS. These seven classes begin where the harness makes the decision and move outward, through what the agent reads, the tools it connects and the instructions it trusts, to the workspace it opens and the software it was installed from. The last, Excessive agency, is the quietest: every check passes, and the agent does damage anyway with the reach it was given.

Scaffolding collapse

Term introduced here by Peter Seprus, 2026.

Enforcement bugs in non-malicious tools. The tool declares a boundary (a sandbox, a path restriction, an ignore rule, an approval prompt, an access check) and its own code fails to enforce it, so the boundary it promises is not the boundary it enforces. The test: the issue would not exist if the tool enforced exactly what it documents. How the gap is reached does not matter; an injected instruction, a malicious repository and a curious model all count. Scaffolding is the usual word for the code around a model, and here it is that code that gives way. The harness is the bug.

No direct mapping in OWASP or MITRE ATLAS.

37 issues
IssueDateHarnessThe failFix
DeepSeek Harness sandbox escape (CVE-2026-82533, OX Research)2026DeepSeek HarnessThe local agent-control API trusted a client-supplied loopback Host header instead of the connection's real origin, so a sandboxed tool process could call it over loopback to switch its own session to danger-full-access and disable the approval prompt.Fixed in 0.1.2-alpha.1, August 27, 2026.
GitSpawn (named by Manifold Security; CVE-2026-72718 and CVE-2026-71963, Francisco Rosales (Manifold Security); CVE-2026-19592, Fudan University System Software and Security Lab and Doyensec via ZDI)2026Claude Code, OpenAI Codex CLI, OpenAI Codex Desktop, Cursor, goose, Grok Build, Hermes Agent, Qwen CodeAgents ran git in the background to gather context, outside their sandboxes and before any trust prompt, so a repository whose .git/config set core.fsmonitor executed an attacker-chosen command on the host. An ordinary git clone does not carry the source’s .git/config, so the repository has to arrive with it intact, for example as an archive or a copied folder.Fixed in OpenAI Codex CLI 0.131.0, Codex Desktop 26.519, goose 1.44.0 and Hermes Agent commit f6234d0; Claude Code fixed the core.fsmonitor path by 2.1.196 and Cursor patched, per Manifold; Claude Code’s ultrareview, Grok Build and Qwen Code unpatched as of September 1, 2026.
IDEsaster (Ari Marzouk (MaccariTA), 24 CVEs across vendors)2025GitHub Copilot, Cursor, Windsurf, Kiro, Zed, JetBrains Junie, Roo Code, Cline, Gemini CLI, Claude CodeInjected instructions made agents use ordinary file writes to trigger base-IDE features: a JSON file whose remote $schema URL the IDE fetched, leaking data, and edits to IDE settings (.vscode/settings.json, JetBrains workspace files, .code-workspace) that pointed executable paths at attacker files, running code.Fixed per vendor, some without a CVE; Claude Code acknowledged the issue and answered with a security warning.
Caught in the Hook (CVE-2025-59536, CVE-2026-21852, named by Check Point Research)2025Claude CodeA cloned repository's .claude/settings.json started its MCP servers and redirected ANTHROPIC_BASE_URL before the trust dialog was answered, running commands and sending the API key to the attacker's server. Repository hooks also ran after trust without the dialog mentioning them.Fixed in Claude Code v1.0.87, v1.0.111 and v2.0.65.
Copilot YOLO-mode remote code execution (CVE-2025-53773, Johann Rehberger; found in parallel by Persistent Security and Ari Marzouk)2025GitHub CopilotA prompt injection made the agent write "chat.tools.autoApprove": true into .vscode/settings.json, which it could do without approval, turning off every confirmation before it ran shell commands.Fixed in Microsoft's August 2025 Patch Tuesday.
CurXecute (CVE-2025-54135, named by Aim Security)2025CursorA prompt injection in a Slack message, fetched through the Slack MCP server, made the agent suggest an edit to ~/.cursor/mcp.json; the edit was written to disk before the user could approve or reject it, and Cursor started the new server at once, running the attacker's command.Fixed in Cursor 1.3.9.
EscapeRoute (CVE-2025-53109, CVE-2025-53110, named by Cymulate)2025Anthropic Filesystem MCP serverA prefix check let paths that merely began with an allowed directory’s name pass as inside it, and symlinks within allowed directories reached files outside them.Patched July 1, 2025 (reported March 30).
Docker Sandboxes escapes (CVE-2026-77179, Accomplish; CVE-2026-79994, ThreatNotify)2026Docker SandboxesThe macOS virtio-fs host server followed symlinks when reopening an unlinked file from a stored path, and the guest-to-host Unix socket relay validated a socket path but reconnected by pathname, so a malicious guest could swap in a symlink and reach host files or arbitrary host sockets outside the shared workspace.Fixed in Docker Sandboxes 0.42.0, September 7, 2026.
Heapjack and Overpatch (Accomplish)2026OpenAI Codex CLI, OpenAI Codex DesktopCodex CLI's apply_patch granted write access to the parent folders of paths in a patch, so a patch referencing /tmp gained write access to / without an approval prompt (Overpatch); Codex Desktop's node_repl kept trusted and untrusted V8 contexts in one heap, so untrusted code could read the authorization token and run unsandboxed commands even in read-only mode (Heapjack).Fixed in Codex CLI 0.149.0 and Codex Desktop build 26.818.21641, within eight days of the August 12, 2026 report.
Cursor CLI pre-trust worktree execution (Manifold Security)2026Cursor CLIIn worktree mode, the setup-worktree command from a git-tracked .cursor/worktrees.json ran through sh -c before the Workspace Trust prompt, under a hardcoded no-sandbox policy that --sandbox enabled did not override.Fixed in cursor-agent 2026.07.23-e383d2b, per Manifold; Cursor closed the report as informative, disputing that workspace trust was bypassed.
Cursor Cloud Agent browser sandbox escape (CVE-2026-61613, Arbër Salihi)2026Cursor Cloud AgentIn browser-enabled sessions, attacker-controlled web content connected from inside the agent container to a local agent endpoint that had no authentication, giving code execution in the sandbox and the session’s files, credentials and GitHub App tokens.Fixed in the Cursor Cloud Agent environment, March 31, 2026; no user action required.
Cursor sandbox escapes (CVE-2026-50548, CVE-2026-50549, Cato Networks (named DuneSlide))2026CursorPrompt injection could make the agent write outside the project, by pointing the terminal tool’s working-directory parameter elsewhere or by forcing a path-canonicalization fallback through a symlink, and overwriting the cursorsandbox helper turned this into a sandbox escape with remote code execution.Fixed in Cursor 3.0.
Gemini CLI --yolo allowlist bypass (Novee Security and Pillar Security)2026Gemini CLIUnder --yolo, Gemini CLI ignored the fine-grained tool allowlist, so run_shell_command(echo) allowed any command. In headless mode it also trusted the folder automatically and loaded .gemini/.env.Fixed in Gemini CLI 0.39.1 and run-gemini-cli 0.1.22.
Antigravity Secure Mode bypass (Pillar Security)2026Google AntigravityThe find_by_name tool passed its pattern straight to fd, so -X sh ran a staged script. The native tool fired before Secure Mode's sandbox and command checks ever saw the call.Google marked it fixed on February 28, 2026.
Deny rules skipped past 50 subcommands (Adversa AI)2026Claude CodeA shell command chaining more than 50 subcommands hit an analysis cap, after which Claude Code skipped its deny-rule checks and fell back to a generic approval prompt, auto-approved in automated runs, so a denied command could run once padded with 50 harmless statements.Fixed in Claude Code v2.1.90.
ShadowPrompt (named by Koi Security)2026Claude in ChromeThe extension accepted prompts from any *.claude.ai page, and an XSS on a-cdn.claude.ai let any website send it instructions with zero clicks.Fixed in 1.0.41 with a strict origin check; the XSS was fixed on February 19, 2026.
Shell-allowlist expansion bypasses (CVE-2026-44462, CVE-2026-44463, CVE-2026-44466, CVE-2026-29783; Zed and GitHub advisories)2026Zed, GitHub Copilot CLIRegex allow rules such as ^git\b were beaten with PAGER=curl git diff, ${var@P} and $(($(cmd))). Copilot CLI's read-only classifier was fooled by the same kind of parameter-transformation expansion.Fixed in Zed 0.229.0 and Copilot CLI 0.0.423.
Cursor sandbox escapes through deferred execution (CVE-2026-26268, CVE-2026-48124, CVE-2026-73217, CVE-2026-73218; Novee, NVIDIA AI Red Team and others)2026CursorInside the auto-run sandbox the agent could write files that later ran outside it (git hooks, a .claude/settings.local.json hook, a virtualenv interpreter the Python extension invokes) or start a privileged container that mounts $HOME.Fixed in Cursor 2.5, 3.0.0 and 3.1.2.
Claude Code sandbox escapes (CVE-2026-25725, CVE-2026-39861, CVE-2026-55607, via HackerOne)2026Claude CodeSandboxed code planted SessionStart hooks in a not-yet-existing .claude/settings.json that ran on the host at the next start. Symlinks made inside the sandbox were followed by the unsandboxed app, and a worktree named .git with git fsmonitor overwrote ~/.zshenv.Fixed in Claude Code v2.1.2, v2.1.64 and v2.1.163.
Claude Code WebFetch allowlist exfiltration (CVE-2026-24052, CVE-2026-54316, via HackerOne)2026Claude CodeA prefix check let modelcontextprotocol.io.evil.com pass as an allowlisted domain, and the pre-approved huggingface.co served as a covert channel for data.Fixed in Claude Code v1.0.111 and v2.1.163.
OpenClaw approval takeover (CVE-2026-25253, CVE-2026-41349; depthfirst and others)2026OpenClawA gatewayUrl query parameter made the UI send its auth token to an attacker, who used the harness's own API to switch off exec approval and move execution out of the container. A later advisory showed the model itself could call config.patch to turn approval off.Fixed in OpenClaw 2026.1.29 and 2026.3.28.
Anthropic Git MCP server chain (CVE-2025-68143, CVE-2025-68144, CVE-2025-68145, Cyata)2025Anthropic Git MCP serverThe server did not enforce its --repository restriction per call, git_init worked on any directory, and git_diff --output= overwrote files. Chained with the filesystem server, a prompt injection reached code execution.Fixed in mcp-server-git 2025.12.18.
Plan Mode destructive commands (Cursor forum)2025CursorPlan Mode, which must not edit or run non-read-only tools, ran pkill, deleted about 70 git-tracked files with rm -rf and made commits after the user wrote "DO NOT RUN ANYTHING"; Cursor staff called it a critical bug in Plan Mode constraint enforcement.—
PromptJacking (named by Koi Security)2025Claude DesktopThree official Claude Desktop extensions (Chrome, iMessage, Apple Notes) passed input straight into AppleScript without escaping, so a web page could turn an ordinary question into shell commands on the host.Fixed in version 0.1.9.
Cursorignore bypass (CVE-2025-64110, Cursor advisory)2025CursorAn agent steered by prompt injection could write a new .cursorignore that invalidated the existing ones, then read the files they protected.Fixed in Cursor 2.0, which blocks the agent from creating or editing .cursorignore files.
Copilot sensitive-file guard bypasses (CVE-2025-62222, CVE-2025-62449, CVE-2026-21523, CVE-2026-41109, CVE-2026-65675, CVE-2026-70335; Microsoft advisories)2025GitHub Copilot, VS CodeAfter CVE-2025-53773, the guard that asks before the agent edits .vscode/settings.json and similar files was bypassed again and again: by casing, by creating a file instead of editing it, by apply_patch path "healing", and by writing custom-agent files that carry hooks.Fixed in Copilot Chat 0.32.1, 0.32.5 and 0.37.3, and VS Code 1.119.1 and 1.132.1.
Claude Code deny-rule bypass through symlinks (CVE-2025-59829, CVE-2026-25724, via HackerOne)2025Claude CodeA file denied in settings, such as /etc/passwd, could be read through a symlink pointing to it. The first fix regressed and needed a second CVE.Fixed in Claude Code v1.0.120, and again in v2.1.7.
Claude Code trust-dialog bypasses (CVE-2025-59041, CVE-2025-59828, CVE-2025-65099, CVE-2026-33068, CVE-2026-40068; NVIDIA AI Red Team, Redguard AG and others)2025Claude CodeBefore the trust dialog was answered, a repository could run code through a malicious git config user.email or through Yarn plugins. A committed defaultMode: bypassPermissions skipped the dialog outright, and a spoofed worktree commondir pointing at an already-trusted path skipped it too.Fixed in Claude Code v1.0.39, v1.0.105, v2.1.53 and v2.1.84.
Claude Code command-validation bypasses (CVE-2025-58764, CVE-2025-64755, CVE-2025-66032, CVE-2026-24053, CVE-2026-24887, CVE-2026-25722, CVE-2026-25723; GMO Flatt Security, SpecterOps and others)2025Claude CodeGaps in the command parser turned auto-approved read-only commands into code execution or writes outside the project, without a prompt: man --html, sort --compress-program, sed's e command, $IFS and @P expansions, zsh >|, and cd into .claude.Fixed across Claude Code v1.0.93 to v2.0.74.
Kiro trusted-commands self-allowlisting (Johann Rehberger)2025KiroA prompt injection made the agent write "kiroAgent.trustedCommands": ["*"] into .vscode/settings.json, or add a server to .kiro/settings/mcp.json, without approval, then run any command.Fixed in Kiro 0.1.42.
Amazon Q find -exec read-only bypass (Johann Rehberger)2025Amazon Q Developerfind was classed as read-only and skipped confirmation, so a prompt injection in a source comment ran find -exec with arbitrary commands.Fixed in v1.85, which reclassified find.
Codex CLI sandbox-root escapes (CVE-2025-55345, CVE-2025-59532; JFrog and Tzanko Matev)2025OpenAI Codex CLIAn AGENTS.md injection plus an in-repo symlink sent --full-auto writes outside the working directory, and the sandbox treated a model-chosen working directory as its writable root.Fixed in Codex CLI 0.12.0 and 0.39.0.
Zed project settings auto-execution (CVE-2025-55012, CVE-2025-68432, CVE-2025-68433; Ari Marzouk, Aaron Portnoy)2025ZedA repository's .zed/settings.json could define MCP and language servers that ran commands when the project opened, with no interaction, and the agent could write project config past its permission checks.Fixed in Zed 0.197.3 and v0.218.2-pre.
Cursor sensitive-file protection bypasses (CVE-2025-54130, CVE-2025-59944, CVE-2025-61593, CVE-2025-64107, CVE-2025-64108; Lakera and others)2025Cursor, Cursor CLIThe approval guard on .cursor/mcp.json, cli.json and .vscode/settings.json was bypassed through case-insensitive filesystems, Windows backslashes, NTFS short names and alternate data streams, and by creating the file instead of editing it. Each bypass led to code execution.Fixed in Cursor 1.3.9, 1.7 and 2.0, and Cursor CLI 2025.09.17.
InversePrompt path-restriction bypass (CVE-2025-54794, CVE-2025-54795, named by Cymulate)2025Claude CodePrefix matching instead of canonical path comparison let /home/user/project-secrets pass as inside /home/user/project. A second flaw in command parsing let an untrusted command run past the confirmation prompt.Fixed in Claude Code v0.2.111 and v1.0.20.
Cursor auto-run allowlist bypasses (CVE-2025-54131, CVE-2026-22708, CVE-2026-31854; Backslash Security, Pillar Security and others)2025CursorCursor's auto-run denylist was evaded with base64, subshells, scripts, backticks and $(), and shell built-ins like export poisoned environment variables that allowlisted commands then used.Denylist deprecated in Cursor 1.3; fixes in 1.3, 2.0 and 2.3.
Unauthenticated local agent servers (CVE-2025-52882, CVE-2026-22812, CVE-2026-44211)2025Claude Code, OpenCode, ClineAgents started local servers without authentication (a WebSocket in the Claude Code IDE extension, an HTTP server with permissive CORS in OpenCode, a WebSocket in Cline Kanban), so any website or local process could read files or run shell commands.Fixed in the Claude Code IDE extension 1.0.24 and OpenCode 1.0.216; the Cline advisory (2.13.0 and earlier) listed no patch when it was published.

Indirect prompt injection

Term introduced by Greshake et al., 2023, extending prompt injection (Simon Willison, 2022).

“We reveal new attack vectors, using Indirect Prompt Injection, that enable adversaries to remotely (without a direct interface) exploit LLM-integrated applications by strategically injecting prompts into data likely to be retrieved.”

Malicious instructions arrive inside material the agent legitimately processes, such as a README, a GitHub issue, a web page or a tool result, and steer the agent into acting for the attacker: leaking files, secrets or credentials, running attacker code, or handing over control of its machine.

OWASP: Prompt Injection (LLM01:2025), Sensitive Information Disclosure (LLM02:2025), Agent Goal Hijack (ASI01), Tool Misuse and Exploitation (ASI02) · MITRE ATLAS: LLM Prompt Injection: Indirect (AML.T0051.001), Exfiltration via AI Agent Tool Invocation (AML.T0086), LLM Response Rendering (AML.T0077)

17 issues
IssueDateHarnessThe failFix
Comment and Control (Aonan Guan, Zhengyu Liu, Gavin Zhong)2026Claude Code, Gemini CLI, GitHub CopilotPull request titles, issue comments and hidden HTML comments in issues made the agents running in GitHub Actions read the runner’s environment and post its API keys and tokens back to the repository, past Copilot’s environment filtering, secret scanning and network firewall.Anthropic blocked ps in the Claude Code Security Review action and said it “is not designed to be hardened against prompt injection”; GitHub called the Copilot exposure “a previously identified architectural limitation”; no Gemini CLI fix stated.
Claude Desktop Extensions zero-click RCE (LayerX)2026Claude DesktopAsked to “take care of” the latest Google Calendar events, Claude followed a malicious event and chained a low-risk connector into an unsandboxed local extension that downloaded and ran attacker code, without asking the user.LayerX reported it; Anthropic decided not to fix it at the time.
Claude Cowork file exfiltration (PromptArmor)2026Claude CoworkHidden text in a .docx made the agent curl the user's files to api.anthropic.com/v1/files under the attacker's API key. The domain was on the egress allowlist, so the network restriction did not stop it.Not remediated at the time of disclosure.
Antigravity exfiltration (PromptArmor)2025Google AntigravityInjection hidden in a web page drove the agent to read a project's .env with a shell command that bypassed the tool's own .gitignore-based protection for that file, then exfiltrate the credentials through an allowlisted domain.— (not reported: PromptArmor skipped disclosure because Google had said it was already aware of data-exfiltration risks)
CamoLeak (Omer Mayraz, Legit Security)2025GitHub CopilotA hidden comment in a pull request description made Copilot Chat spell out the victim’s private-repository contents as a sequence of pre-signed Camo image URLs, which the browser fetched through GitHub’s own image proxy, past its Content Security Policy.Fixed server-side by GitHub, August 14, 2025, by disabling image rendering in Copilot Chat.
Cursor Mermaid exfiltration (CVE-2025-54132, Johann Rehberger)2025CursorAn injection in a source-code comment made the agent collect API keys from the project, then render a Mermaid diagram whose image URL carried them to the attacker's server, with no confirmation.Fixed in Cursor 1.3, July 2025, per the researcher.
Gemini CLI hijack (Tracebit)2025Gemini CLIInstructions hidden in a README combined with an allow-list parsing flaw yielded silent data exfiltration.Fixed in v0.1.14, July 25, 2025 — 28 days after report.
Claude Code WebFetch allowlist exfiltration2026—See Scaffolding collapse above.—
GhostSplice (named by ASSET Research Group (Murali Ediga, Johnny Dao, Sudipta Chattopadhyay))202611 models testedInstructions fragmented across MCP tool descriptions and tool results raised model compliance from 42% to 82%, no fragment looking malicious on its own.— (cross-model research)
Codex Desktop image exfiltration (CVE-2026-14898)2026OpenAI Codex DesktopAn indirect injection made the model build an image URL carrying secrets, which the app fetched automatically.Fixed in 26.527.31326.
Manus VS Code exposure (Johann Rehberger)2025ManusAn injection in a document made the agent publish its VS Code server to the internet with its port-exposure tool, which asks for no approval, then browse to an attacker URL that carried the server address and password; rendered images leaked data as well.— (reported June 2025; mitigation status unclear at publication)
Windsurf exfiltration (Johann Rehberger)2025WindsurfAn injection in a file under analysis made the agent read .env and send it out through read_url_content, which fetches URLs without approval, or through an auto-rendered image.— (unfixed at publication, August 2025)
Jules exfiltration and ZombAI (Johann Rehberger)2025Google JulesInstructions in a GitHub issue made the agent leak data through rendered images and its browsing tool, and download and run a command-and-control implant on its machine, which had unrestricted internet access; invisible Unicode tag characters in an issue also made it add backdoor code and run it.— (reported May–June 2025; not fully mitigated at publication, per the researcher)
Claude Code DNS exfiltration (CVE-2025-55284, Johann Rehberger)2025Claude CodeAn injection in a file under analysis made the agent read .env and put the key into a hostname for ping, which the default allowlist ran without approval, so the DNS lookup carried it to the attacker's server.Fixed in Claude Code 1.0.4, June 2025, per the researcher.
OpenHands ZombAI (Johann Rehberger)2025OpenHandsInstructions on a web page or in a GitHub issue made the agent download and run malware that connected to the attacker’s command-and-control server.— (reported March 2025; no fix stated at publication)
Devin secret leaks (Johann Rehberger)2025DevinAn injection in content such as a GitHub issue made the agent read secrets exposed to it as environment variables and send them out through its shell, its browser or a rendered markdown image.— (reported April 2025; fix status unanswered at publication)
Codex cloud ZombAI (Johann Rehberger)2025OpenAI Codex cloudInstructions planted in a GitHub issue made the agent download and run malware that joined the attacker’s command-and-control server, reached through a cloudapp.azure.com host that the azure.com entry in the Common Dependencies network allowlist let through.None — OpenAI closed the report as not applicable the day it was filed, June 10, 2025.

Tool poisoning

Term introduced by Invariant Labs, 2025 for poisoned tool descriptions; used here for every way a hostile tool reaches the agent.

The tool channel itself is hostile: instructions ride in tool descriptions or results, or the agent is steered into configuration writes that escalate to code execution.

OWASP: Supply Chain (LLM03:2025), Prompt Injection (LLM01:2025), Agentic Supply Chain Vulnerabilities (ASI04), Tool Misuse and Exploitation (ASI02) · MITRE ATLAS: AI Agent Tool Poisoning (AML.T0110), AI Supply Chain Compromise: AI Agent Tool (AML.T0010.005), AI Supply Chain Rug Pull (AML.T0109)

6 issues
IssueDateHarnessThe failFix
Windsurf zero-click MCP config rewrite (CVE-2026-30615, OX Security)2026WindsurfInjected HTML content made Windsurf rewrite its local MCP config and register a malicious stdio server, with no user interaction.No fixed version found.
MCPoison (CVE-2025-54136, named by Check Point)2025CursorAn approved MCP configuration could be silently swapped afterward.Fixed in Cursor 1.3, July 29, 2025.
WhatsApp MCP exfiltration (Invariant Labs)2025CursorA sleeper MCP server, once approved, swapped in a tool description that redirected the trusted WhatsApp server's messages to the attacker's number with the chat history attached; the confirmation dialog hid the payload off-screen.— (research demonstration)
Tool Poisoning Attacks (named by Invariant Labs)2025CursorHidden instructions in a malicious MCP server's tool description made the agent read ~/.cursor/mcp.json and ~/.ssh/id_rsa and pass them to the server as a tool-call argument.— (research demonstration)
GhostSplice2026—See Indirect prompt injection above.—
CurXecute2025—See Scaffolding collapse above.—

Agent instruction-file injection

Term as used in the Arcanum Prompt Injection Taxonomy (Jason Haddix, Arcanum Information Security), which also lists Pillar Security's Rules File Backdoor.

Instruction files the agent trusts are themselves the injection vector.

OWASP: Prompt Injection (LLM01:2025), Supply Chain (LLM03:2025), Agent Goal Hijack (ASI01), Memory & Context Poisoning (ASI06) · MITRE ATLAS: Modify AI Agent Configuration (AML.T0081), AI Agent Context Poisoning (AML.T0080), AI Supply Chain Compromise: AI Software (AML.T0010.001)

3 issues
IssueDateHarnessThe failFix
Forced Descent (named by Mindgard)2025Google AntigravityA malicious .agent rules file made the agent write the global ~/.gemini/antigravity/mcp_config.json, even with non-workspace file access off, planting code that ran on every launch and survived a reinstall.Google first closed the report as Won’t Fix (Intended Behavior), reopened it after Mindgard’s rebuttal, and on November 25, 2025 filed a bug with the product team; no fix has been announced.
.clinerules switches off approval (Mindgard)2025ClineMarkdown in a repository's .clinerules told the agent to set requires_approval=false, and commands then ran without consent.No longer reproducible in Cline 3.35.0; Mindgard reports no vendor response.
Rules File Backdoor (named by Pillar Security)2025Cursor, GitHub CopilotMalicious instructions hidden in rules files behind zero-width and bidirectional Unicode characters, invisible in review.None — both vendors declined to treat it as a vulnerability, placing the burden on users.

Ambient activation

Term introduced here by Peter Seprus, after ambient authority in capability-based security.

Opening a folder is enough. The agent, or the editor around it, starts acting on the project before the user has expressed any intent or granted trust: indexing it, or running the servers, tasks and settings the repository defines.

No direct mapping in OWASP or MITRE ATLAS.

4 issues
IssueDateHarnessThe failFix
Codex CLI project MCP config execution (CVE-2025-61260, Check Point Research)2025OpenAI Codex CLIA repository .env set CODEX_HOME=./.codex, so Codex loaded the repository's own config.toml and started its MCP server commands at launch, with no approval.Fixed after Codex CLI 0.23.0.
Cursor CLI repository config execution (CVE-2025-61592, CVE-2025-64109; JFrog and others)2025Cursor CLIA project-local .cursor/cli.json could set the shell allowlist, and a repository's .cursor/mcp.json server ran when the CLI started, with no warning.Fixed in Cursor CLI 2025.09.17.
Open Repo, Get Pwned (Oasis Security)2025CursorCursor ships with Workspace Trust off, so a repository's .vscode/tasks.json with runOn: "folderOpen" ran as soon as the folder was opened.Not changed; Cursor pointed users to enabling Workspace Trust.
Zed project settings auto-execution2025—See Scaffolding collapse above.—

Supply-chain compromise

Generic industry term.

The software supply chain delivers the attack, and an agent harness is part of it: a harness release ships hostile, or a malicious package drives the agent the user already installed, or plants itself in that agent’s configuration. The test: the harness is the payload, the weapon or the foothold.

OWASP: Supply Chain (LLM03:2025), Agentic Supply Chain Vulnerabilities (ASI04), Tool Misuse and Exploitation (ASI02) · MITRE ATLAS: AI Supply Chain Compromise: AI Software (AML.T0010.001), Deploy AI Agent (AML.T0103), User Execution: Malicious Package (AML.T0011.001)

4 issues
IssueDateHarnessThe failFix
Mini Shai-Hulud (StepSecurity)2026Claude Code, VS CodeAn npm worm wrote a SessionStart hook into the project's .claude/settings.json and a folderOpen task into .vscode/tasks.json, so opening the repository in either tool re-ran its payload.—
Clinejection and the unauthorized cline@2.3.0 (Adnan Khan)2026Cline CLIA prompt injection in Cline's AI issue-triage workflow, followed by Actions cache poisoning, leaked its publish tokens. A third party then shipped cline@2.3.0, whose postinstall script installed OpenClaw on users' machines.Fixed in 2.4.0; publishing moved to OIDC.
nx "s1ngularity" compromise (Nx postmortem; agent flags documented by Snyk)2025Claude Code, Gemini CLI, Amazon Q Developer CLIMalicious nx releases, published with a stolen npm token, ran a postinstall script that invoked the victim's own AI CLIs with their safety switches off (--dangerously-skip-permissions, --yolo, --trust-all-tools) to inventory files holding secrets.Malicious versions removed from npm the same day; Nx moved publishing to npm Trusted Publishing with 2FA.
Amazon Q Developer VS Code extension v1.84.0 (CVE-2025-8217, AWS bulletin)2025Amazon Q DeveloperAn attacker used an inappropriately scoped GitHub token to commit malicious code, designed to call the Q Developer CLI, into the extension’s repository, and it shipped in release 1.84.0. A syntax error kept it from running.v1.85.0 released; v1.84.0 pulled from distribution.

Excessive agency

Term from the OWASP Top 10 for LLM Applications (LLM06:2025), used here for its accidental case: no attacker, no injection.

No attacker and no injection: a conforming agent, acting on a benign instruction, deletes or overwrites data through its own error.

OWASP: Excessive Agency (LLM06:2025), Rogue Agents (ASI10) · MITRE ATLAS: no direct mapping

7 issues
IssueDateHarnessThe fail
PocketOS production-database deletion (Jer Crane)2026CursorStuck on a credential mismatch in staging, the agent found an unscoped Railway token in an unrelated file and used it in one API call that deleted the production volume and its backups in nine seconds.
Claude Code archive deletion (GitHub issue)2026Claude CodeAfter moving about 1,500 images (around 50 GB) into a subfolder, the agent ran rm -rf on the parent folder, deleting the files it had just moved, without warning.
DataTalksClub terraform destroy (Alexey Grigorev)2026Claude CodeWorking from a stale Terraform state, the agent ran terraform destroy against production (database, network, containers and snapshots, 1.94 million rows), and the user let it run.
Claude file-cleanup deletion (Nick Davidov)2026Claude CoworkMerging a lowercase photos folder into a new Photos folder while organizing a desktop, the agent ran rm -rf on what it took for a separate empty folder; on the case-insensitive macOS filesystem it was the same folder, and 15 years of family photos were deleted outside the Trash.
Claude Code home-directory wipe (u/LovesWorkin on Reddit)2025Claude CodeA repository-cleanup task produced an rm -rf whose trailing ~/ expanded to the user's entire home directory.
Antigravity D: drive wipe (u/Deep-Hyena492 on Reddit)2025Google AntigravityAsked to clear a project cache in Turbo mode, which needs no approval, the agent deleted the root of the user's D: drive, bypassing the Recycle Bin.
Replit Agent production-database deletion (Jason Lemkin)2025Replit AgentDuring an explicit code and action freeze, the agent deleted SaaStr's live production records for over 1,200 executives and 1,190 companies.

The matrix

Updated October 6, 2026 · 13 harnesses · 17 requirements · JSON

The capabilities a harness would need to enforce a file-access policy, as derived from the agentaccess.txt SPEC and written as abstract requirements, grouped under four questions: can the user state a file-access policy, what do its rules actually gate, which way do its errors fail, and does the mechanism hold? Each harness is scored against them from its own documentation, source code, issue tracker and security advisories. A ? means no authoritative statement was found either way, and an undocumented boundary is itself a finding.

● yes, documented ◐ partial, or documented gaps ○ no ? undocumented — not applicable

The requirements

Each requirement links to the SPEC section it derives from. The SPEC is one proposal for satisfying them; the requirements stand on their own.

A. Declarative surface — can the user state the policy?

A1 Path rules
User-authorable file-access rules exist at all: an ignore file, deny globs, private-file patterns. §1
A2 Tree-resident
The rules live in the directory they govern and travel with it — clone the tree, keep the policy. Any directory can carry its own policy file, applied by the path being accessed; a single project-root file, or one chosen by where the session starts, is partial. §4
A3 Per-path
Rules distinguish paths by pattern, not only whole-project switches. §5
A4 Per-agent
Rules can name which tool they bind — "tool X may work here, tool Y may not." §6
A5 Shared grammar
The rules format is shared across tools rather than a proprietary dialect. §8
A. Declarative surface
HarnessA1 Path rulesA2 Tree-residentA3 Per-pathA4 Per-agentA5 Shared grammar
Aideryes, documentedpartialyes, documentednopartial
Claude Codeyes, documentedpartialyes, documentednono
Clineyes, documentedpartialyes, documentednopartial
Cursoryes, documentedpartialyes, documentednopartial
DeepSeek Harnessnonononono
Gemini CLIyes, documentedpartialyes, documentednopartial
GitHub Copilot (agent surfaces)yes, documentednoyes, documentednono
goosenonononono
JetBrains AI Assistant / Junieyes, documentedpartialyes, documentedpartialpartial
OpenAI Codex CLIpartialpartialpartialnono
Pinonononono
Windsurf / Devinyes, documentedpartialyes, documentednopartial
Zedyes, documentedyes, documentedyes, documentednono

B. Coverage — what do the rules actually gate?

B1 Context gating
Restricted content is kept out of model context — reads, indexing, attachments — not merely deprioritized. §3, §7
B2 Shell coverage
The restriction holds when the model reaches for the shell — cat, subprocesses, terminal tools. §7
B3 Write gating
Rules cover create/modify/delete, not only reads — a declarative do-not-touch. §7
B4 Pre-flight
Rules are evaluated before content is touched — at workspace open for ambient agents, before each operation otherwise — and changes take effect without a restart. §4
B. Coverage
HarnessB1 Context gatingB2 Shell coverageB3 Write gatingB4 Pre-flight
Aiderpartialundocumentedundocumentedundocumented
Claude Codepartialpartialyes, documentedyes, documented
Clinenonoundocumentedundocumented
Cursorpartialnoundocumentedundocumented
DeepSeek Harnessundocumentedpartialpartialundocumented
Gemini CLIyes, documentedpartialpartialpartial
GitHub Copilot (agent surfaces)partialpartialpartialpartial
gooseundocumentedundocumentedundocumentedundocumented
JetBrains AI Assistant / Juniepartialnopartialundocumented
OpenAI Codex CLIpartialyes, documentedyes, documentedpartial
Pinononono
Windsurf / Devinyes, documentedpartialyes, documentedundocumented
Zedyes, documentednoyes, documentedpartial

C. Fail direction — which way do errors fail?

C1 Deny precedence
A restriction holds against tool-native grants: a session approval, allowlist entry, remembered permission or bypass mode does not lift it. §7
C2 Fail-closed policy errors
A malformed or invalid rule restricts rather than being silently dropped. §5
C3 Self-protection
The agent cannot create, modify, or delete the policy that restricts it. §7
C4 Policy as data
The policy is evaluated in deterministic code and never enters model context as natural language — a policy the model can read is a prompt-injection surface. §9
C. Fail direction
HarnessC1 Deny precedenceC2 Fail-closed errorsC3 Self-protectionC4 Policy as data
Aiderundocumentedundocumentedundocumentedundocumented
Claude Codeyes, documentedpartialpartialyes, documented
Clineundocumentedundocumentedundocumentedundocumented
Cursorundocumentedpartialundocumentedundocumented
DeepSeek Harnessnoundocumentedpartialundocumented
Gemini CLIundocumentednonoundocumented
GitHub Copilot (agent surfaces)partialundocumentedpartialundocumented
goosepartialundocumentedundocumentedundocumented
JetBrains AI Assistant / Junieundocumentedundocumentedundocumentedundocumented
OpenAI Codex CLIpartialpartialpartialpartial
Pinot applicablenot applicablenot applicablenot applicable
Windsurf / Devinpartialpartialundocumentedundocumented
Zedyes, documentedpartialpartialundocumented

D. Enforcement quality — does the mechanism hold?

D1 Harness enforcement
Restrictions are enforced in deterministic harness code, never entrusted to the model's judgment or a system-prompt request. §9
D2 Canonicalization
Symlinks resolved and paths canonicalized before the policy check — a recurring class of enforcement bug. §5, §9, A.6
D3 Deliberate override
A restriction is lifted only by an explicit user action naming that restriction — never by a blanket bypass mode, and never by consent the model parsed out of its own context. §7
D4 Delegation
Restrictions carry to subagents and spawned processes; they attenuate through delegation, never reset. §7
D. Enforcement quality
HarnessD1 Harness enforcementD2 CanonicalizationD3 Deliberate overrideD4 Delegation
Aiderundocumentedundocumentedundocumentedundocumented
Claude Codeyes, documentedpartialpartialpartial
Clinepartialundocumentedundocumentedundocumented
Cursorpartialpartialundocumentedundocumented
DeepSeek Harnesspartialpartialnopartial
Gemini CLIyes, documentedpartialundocumentedundocumented
GitHub Copilot (agent surfaces)partialpartialundocumentedundocumented
goosepartialundocumentedundocumentedundocumented
JetBrains AI Assistant / Juniepartialundocumentednoundocumented
OpenAI Codex CLIyes, documentedpartialpartialyes, documented
Pipartialnot applicablenot applicableno
Windsurf / Devinyes, documentedpartialundocumentedpartial
Zedyes, documentedpartialundocumentedundocumented

What the matrix shows

Three columns carry the headline. A4 (per-agent) has no ●: no harness lets a directory say which tools are welcome; the closest is JetBrains’ .noai, a single-vendor kill switch. A5 (shared grammar) has no ●: every rules format is a proprietary dialect, at best borrowing .gitignore syntax, and rules travel between tools only where one vendor chooses to read another’s file. B2 (shell coverage) has a single ●: only OpenAI Codex CLI’s always-on kernel sandbox closes it by default; five harnesses close it only partly, mostly with an opt-in sandbox, and five not at all. The ? marks are a finding too: for Aider, goose and Cline most of these questions have no documented answer, and more than a quarter of all cells are ?.

Per-harness notes

One note per harness: what each score rests on, with the sources behind it.

Aider

.aiderignore, /read-only, --add-gitignore-files — long-standing, gitignore syntax (FAQ, options, commands) (A1/A3 ●, A5 ◐); the ignore file is one per repository, "default: .aiderignore in git root" (A2 ◐). Aider's workflow is explicit-add — files enter context when the user adds them — which gates context by construction but is a workflow property, not an enforcement claim (B1 ◐). Enforcement locus and shell coverage undocumented.

Claude Code

permissions.deny with Read()/Edit() path rules (permissions docs); project settings can be committed with the repo (.claude/settings.json), but Claude Code reads that one file "from the session's primary working directory" (settings), with no per-directory files (A2 ◐). Vendor states enforcement is deterministic: "Permission rules are enforced by Claude Code, not by the model" (D1 ●), with documented precedence "deny, then ask, then allow". A Read deny also blocks Edit and Write on the same path (B3 ●), evaluated before the tool call executes, and edits to permissions reach "the running session without a restart" (B4 ●) — but for context surfaces (Grep, Glob, @-mentions) the vendor's own wording is a "best-effort attempt", and NotebookEdit isn't covered (B1 ◐). Shell coverage is explicitly partial: recognized file commands only, rules "don't apply … to arbitrary subprocesses"; the opt-in OS sandbox closes that, with a default-enabled but prompt-gated unsandboxed-retry escape hatch (B2 ◐). Invalid managed sandbox values fail closed per the changelog (2.1.283), and an unparseable managed settings file stops every session; one the OS denies reading starts the session without its policies (2.1.285), and malformed mcp__ parameter rules are skipped with a warning (C2 ◐). Scoped rules "don't change which tools Claude sees. Claude Code checks them when Claude attempts a call, leaving the prefix intact", and a bare tool-name deny removes the tool rather than describing it (prompt caching) (C4 ●). CVE-2025-54794 was a prefix-matching-instead-of-canonicalization bug, fixed; the permissions docs now check both the requested path and the file a symlink resolves to, and symlink bypasses of Read deny rules through @-mentions and IDE selections were fixed in 2.1.289–2.1.290 (D2 ◐). Permission modes state that "Deny rules block in every mode, including bypassPermissions" — the blanket bypass lifts prompts, not deny rules (C1 ●) — and writes to protected paths, .claude settings among them, cannot be pre-approved by allow rules, yet in auto mode (the starting mode since 2.1.284) a model classifier decides them, and bypassPermissions allows them outright (C3 ◐). That protected-path restriction is the D3 gap: user deny rules lift only when the user edits them, but the built-in restriction falls to a blanket mode and, in auto mode, to a classifier model's judgment (D3 ◐). Subagent docs cover the delegation interplay partially — a subagent cannot self-declare bypassPermissions, and background subagents surface permission prompts to the main session — and state that built-in subagents inherit the parent conversation's permission rules, with no such statement for custom ones (D4 ◐).

Cline

.clineignore, gitignore syntax, one file "in your project root" (A1/A3 ●, A2 ◐, A5 ◐) — and the docs state plainly that it "is not a security or access-control boundary — ignored files can still be read via explicit @ mentions or shell commands": it controls automatic loading, not access (B1 ○, B2 ○). The page carries a deprecation banner and offers a PreToolUse hook script as the enforced replacement, which "actively blocks the tool call" when a read, edit, or shell command targets an ignored file: opt-in, user-installed code that does not resolve symlinks and is disabled in --yolo mode (D1 ◐). Writes and the rest of groups C–D are undocumented (?).

Cursor

.cursorignore, gitignore syntax, one file "in your root directory", with an opt-in setting that searches parent directories, not subdirectories (ignore docs) (A1/A3 ●, A2 ◐, A5 ◐). The docs disclaim completeness — "complete protection isn't guaranteed due to LLM unpredictability" — and state plainly that agent terminal and MCP tools "cannot block access" to ignored code (B1 ◐, B2 ○). Whether the rules gate writes, and when rule changes take effect, is undocumented (B3/B4 ?). CVE-2025-64110 was an ignore-file fail-open (a new .cursorignore invalidated existing protections), fixed in 2.0 (C2 ◐). CVE-2026-50548 and CVE-2026-50549 (Cato Networks) let the agent write files outside the workspace through a sandbox write-path override and a failed-canonicalization fallback through an in-workspace symlink; both fixed in Cursor 3.0 per the NVD records (D2 ◐). "Cursor blocks access to files listed in .cursorignore", under the same no-guarantee caveat (D1 ◐).

DeepSeek Harness

A named per-session abstraction — executors run "under the session file policy" (v0.1.6-alpha.1, carried into v0.1.7-rc.1) — documented as one of three sandbox modes (read-only, workspace-write, danger-full-access) over a workspace root, with no path patterns (sandbox docs) (A1/A2/A3 ○, B2/D4 ◐). An approved escalated retry "is a new call with a wider policy" (C1 ○), and danger-full-access, which "bypasses confinement", ships in the default preset table (permission presets) (D3 ○). Enforcement fixes track the classic classes: POSIX symlink/parent-directory combinations, Windows drive-relative paths (v0.1.6-alpha.1), cross-workspace deletion escapes (v0.1.7-alpha.1) (B3 ◐, D2 ◐). CVE-2026-82533 (OX Research): the sandboxed agent could reach the harness's own unauthenticated control interface and elevate its session to danger-full-access — the enforcement layer operable by the thing it governs; fixed August 27, 2026 in 0.1.2-alpha.1, whose release notes require a one-time token for network access to the web interface (C3 ◐: the bar exists now, its guarantee is undocumented). SAFETY.md states plainly that sandboxing, approval prompts, and permission controls "do not guarantee isolation or prevent damage" (D1 ◐).

Gemini CLI

.geminiignore plus .gitignore reuse (docs) (A1/A3 ●, A5 ◐); .geminiignore is one file "in the root of your project directory" — nested .gitignore files are honored per directory (source), but that is the version-control file, not the harness's policy file (A2 ◐). read_file enforces deterministically via shouldIgnoreFile() (B1/D1 ●) — and a feature request on the project's tracker names cat via run_shell_command as the workaround for reading ignored files (#13775); only the opt-in tool sandbox, off by default, turns ignored paths into kernel-denied paths for shell commands (source) (B2 ◐). Changes require a restart to take effect, per the same docs (B4 ◐). Negation handling has shipped a bug that errs restrictive (! rules fail to un-ignore, #5444, closed by the stale bot without a fix), but the error paths fail open: the matcher library skips a pattern ending in an unescaped backslash without a warning (node-ignore), an unreadable .geminiignore loads as no rules (source), and the ignore check returns not-ignored on any exception (source) (C2 ○). The write tools carry no ignore check at all — write_file and edit never consult the ignore service, verifiable in source; the opt-in sandbox denies only shell writes to ignored paths (B3 ◐). Nothing protects .geminiignore itself: "If the workspace is writable, we allow editing .gitignore and .geminiignore by default" (sandbox source) (C3 ○). read_file checks the ignore rules against both the requested path and its resolved real path (source); other surfaces are undocumented (D2 ◐).

GitHub Copilot (agent surfaces)

Content exclusion exists and is deterministic for completions, chat, and review — and the docs exclude both agent surfaces by name: "GitHub Copilot CLI and Agent mode in Copilot Chat in IDEs do not support content exclusion" (configuring exclusions). The path rules that do reach an agent surface are Copilot CLI's own: Read(...), Edit(...), and Write(...) glob rules in enterprise managed settings, session --deny-tool patterns such as read(.env) and write(PATH), and deniedPaths for sandboxed commands (configuration reference, command reference) (A1/A3 ●); agent mode in IDEs documents no equivalent. None of them lives in the tree: the repository settings file .github/copilot/settings.json supports only a listed set of keys, permission rules not among them, and content exclusion is server-delivered (A2 ○). Coverage is CLI-only: Read rules gate file reads, Edit/Write rules also cover a recognized set of shell redirections and in-place sed edits, but when the target "comes from a variable" the path rules "do not apply" (B1/B2/B3 ◐); rules are checked per request, and managed settings re-apply hourly "without restarting the session" (B4 ◐). A write(PATH) match "resolves symlinks and . / .. segments", while content exclusion documents symlinks and remote filesystems as not covered, with up to 30-minute propagation delay (D2 ◐). The CLI's rules are enforced by the harness, but the restriction an organization declares through content exclusion is not enforced on either agent surface (D1 ◐). File-based managed settings are rejected when they are "symlinks, not owned by root, or world-writable"; the docs state no such guard for the user-level settings file (C3 ◐). v1.0.88 closed a management-layer fail-open where ACP, AHP-host, and --server sessions "previously ran with no managed MCP, permission, or plugin policy". Copilot CLI's own deny rules hold against grants — "Deny rules always take precedence over allow rules, even when --allow-all is set or a matching approval has been saved" (allowing tools) — while agent mode in IDEs documents no equivalent (C1 ◐).

goose

.gooseignore is gone from the current tree: the docs carry no page for it, the string survives only in the repo's own .gitignore, and v1.44.0's "Remove stale gooseignore references" was the cleanup — current permissions are per-tool (allow / ask / deny), not per-path, so no declarative path-rule surface exists today (A1/A2/A3 ○). v1.49.0 shipped "Give permission denies precedence" (#11477), so a saved never_allow now beats an overlapping always_allow — but Auto mode approves every tool without consulting it (permission inspector) (C1 ◐), and hooks carry a documented deny contract (v1.44.0), though a policy hook that fails lets the call through unless set to on_failure: block (hooks docs) (D1 ◐). What the per-tool permissions gate operationally is undocumented (B1–B4 ?).

JetBrains AI Assistant / Junie

.aiignore, honoring .cursorignore/.codeiumignore/.aiexclude "as long as they are located in the root folder of the project" (A5 ◐), plus .noai "in the root directory of the project" as a whole-project kill switch — a vendor-specific marker, per-agent in the binary, single-vendor sense (A4 ◐) (docs); the documented locations are the project root, not per directory (A2 ◐). Enforcement is vendor-admitted best-effort: "ignored files may still be processed due to unforeseen issues" (docs), and file names and paths stay visible (support KB) (B1/D1 ◐). Allowlisted commands skip .aiignore checks entirely (B2 ○), and Brave Mode bypasses the approval prompts wholesale (D3 ○). Edits to ignored files are approval-gated rather than hard-blocked — Junie "will ask for explicit approval before viewing or editing the contents" (B3 ◐).

OpenAI Codex CLI

The long-standing read-restriction gap has begun closing: beta permission profiles add per-path filesystem rules — glob deny entries that deny "both reads and writes" — configurable in project .codex/config.toml files, which Codex loads "from the project root to your current working directory", closest wins, trusted projects only (advanced config): the files travel with the repo, but which ones apply depends on where the session starts, not on the path being accessed (A1/A2 ◐, A3 ◐, B1 ◐: beta, config-based, sandbox-enforced; the .codexignore request #1397 was closed into #2847, itself closed completed June 2026 with this feature as the answer). The profile semantics lean fail-closed: deny takes precedence over write and write over read, and "missing or empty filesystem tables keep filesystem access restricted" with a startup warning (C2 ◐); the :workspace profile keeps the workspace .codex directory read-only by default (C3 ◐); the deny list is enforced in sandbox code, but the harness also writes its paths and globs into model context as a developer message (source) (C4 ◐). The enforcement substrate is an always-on OS sandbox (macOS Seatbelt, Linux bubblewrap; docs, independently investigated) confining writes and network at kernel level, with spawned commands inheriting the same boundaries (B2/B3 ●, D1 ●, D4 ●); whether an edited profile reaches a running session without a restart is undocumented (B4 ◐). On Linux, literal deny paths are kept alongside their resolved targets so denials hold across writable symlinks (#48155) (D2 ◐). Approved escalations run unsandboxed and a danger-full-access mode exists (D3 ◐); approved commands "retain explicit filesystem denials" (rust-v0.159.0), but the :danger-full-access profile "removes local sandbox restrictions" (permissions) (C1 ◐).

Pi

Ships no file-access rules: an issue on its tracker lists the only mechanisms as the coarse --tools flag, custom TypeScript hooks, and interactive confirmation (#4459, which proposed native allow/deny policy files — auto-closed during a refactor without review) (A1–B4 ○, C1–C4 —). The security doc places safety outside the harness — it "comes from limiting the files, credentials, processes, and network services Pi can access" — and spells out the consequence: the working folder "does not prevent commands from accessing other paths available to the Pi process," project trust "does not limit what tool calls can access," and child processes "run with those same permissions" (D4 ○). What Pi does have is a deterministic enforcement point: the tool_call hook can block operations in host-side code before execution, and the shipped permission-gate example blocks by default when no UI is present — a policy layer could attach there, but enforcement itself is user-written code, not a shipped feature (D1 ◐).

Windsurf / Devin

.devinignore (with legacy .windsurfignore / .codeiumignore "still read and enforced alongside" it), gitignore syntax, one file at "the root of your repository": paths matched "are excluded from indexing and cannot be viewed, edited, or created by Devin" (docs) (A1/A3 ●, A2 ◐, A5 ◐, B1/B3 ●). Since v3.9.19 removed Cascade (changelog), Devin Local is the only agent in Devin Desktop, and it shares Devin CLI's permission rules, evaluated in a fixed order where "A deny rule always wins" (C1 ◐: only organization-level rules are guaranteed against Bypass mode). Smart mode's model judgment "only applies where no rule already decides the call, so a deny rule blocks the action"; the .devinignore docs do not name where it is enforced (D1 ●). An opt-in OS sandbox hides Read(...)-denied paths from shell commands (B2 ◐) and errs restrictive: unsupported exclusion rules are "ignored with a warning", and a sandbox that cannot be resolved refuses to start (C2 ◐). The write tools "refuse to write through a symlink" (v3.6.27), with reads unstated (D2 ◐), and background subagents get only tools already approved in the session (D4 ◐).

Zed

private_files and read_only_files glob settings, project-resident (configuration, tool permissions); project settings files can also sit in subdirectories (configuring Zed), and the agent's file tools look the settings up for the path being accessed (source) (A1/A2/A3 ●, B1/B3 ●). v1.21.0 added "..." inheritance-extension semantics for read_only_files, and v1.22.0 spread them to more glob settings. The terminal tool is governed by command-pattern rules, not path rules, and under the new OS sandbox terminal commands "can read most of the filesystem" (B2 ○). v1.22.0 stopped one invalid pattern from discarding the valid exclusion/read-only/private rules beside it, which narrows a fail-open in their own dialect rather than closing it: a matcher that cannot be built still matches nothing (#64525) (C2 ◐). always_deny rules are "highest priority, cannot be overridden" and still block when every tool action is auto-approved (C1 ●); agent edits inside .zed/, where project private_files live, always prompt (source), but sandboxed terminal commands can still write project files (C3 ◐). Enforcement is deterministic editor code, confirmed by the project's own advisory for CVE-2026-27967 — agent file tools checked private_files on logical paths, so symlinks escaped; patched 0.225.9 (D1 ●, D2 ◐).

Method

Scored from primary sources wherever possible: vendor documentation, public source code, the project's own issue tracker, and security advisories. A ? means no authoritative statement was found either way — inference is not scoring. A ◐ with a fixed CVE means the mechanism exists and its failure mode is documented history, not that the current version is broken. OpenCode and OpenHands are tracked but not yet scored: their shipped permission surfaces are approval-flow, and a row here requires the full source pass. Re-scored as harnesses ship; dated changes land in the scoring history below.

Scoring history

October 6, 2026 (initial scoring) — thirteen harnesses against seventeen requirements; every sourced cell checked against its live primary source.

Landscape shifts

Directional changes in agent harnesses, and in the operating systems they run on: not every feature, only the moves that change where file-access policy lives, who decides, or which way it fails.

  1. October 2, 2026 · Apple macOS · The OS tightens its broadest file-access grant: granting Full Disk Access will require "very explicit user action". Apple's stated reason is developers using it "in ways that could put users at risk", with AI agents named as a growing risk: "As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially."
  2. September 29, 2026 · OpenAI Codex rust-v0.159.0 · Denials hold against approvals: "Approved commands retain explicit filesystem denials", and .aws directories are protected by default under writable roots.
  3. September 29, 2026 · Claude Code 2.1.285 · Denials hold against lower layers: project settings can no longer "reopen managed read-denies".
  4. September 28, 2026 · Claude Code 2.1.284 · The default decision-maker moves from the user to a model: interactive and VS Code sessions now start in auto mode when no permission mode is configured, on every plan and provider.
  5. September 3, 2026 · goose v1.49.0 · Denials take precedence: "Give permission denies precedence".

Contributing

To add an issue, open a pull request to the harness.fail repository with a source, the affected harness, the date it was first disclosed, and one factual sentence on what failed. If it was fixed, name the version that fixed it.

Primary sources are preferred: an advisory, a CVE record, a vendor announcement, or the researcher's own write-up. Press coverage is fine when no first-hand account is public.

To dispute a matrix score, or to correct anything else on the page, open a pull request.

CONTRIBUTING.md has the row format and the fix statuses. If you would rather not edit JSON, open an issue there with the source instead.