Home / Resources / Security

Security · 5 min read · Updated 2026-10-06

OpenAPPA hit 0% attack success on two benchmarks. Here's what that actually proves

An open source security engine stopped every attack across two independent benchmarks, while an established agent framework let three in ten through. The result says less about one product than about where prompt injection defences now have to sit.

Archestra released OpenAPPA, an open source security engine built to stop data exfiltration caused by prompt injection or model hallucination, and ran it against two agent security benchmarks: Bench-Corp, a set of 20 multi-step enterprise workflows, and AgentThreatBench, which tests against the OWASP Top 10 for Agentic Applications (2026). OpenAPPA recorded a 0% attack success rate on both, while completing 89% of tasks. Claude Code's auto mode let 10% of attacks through on the same tests. Microsoft FIDES let 31% through, and its task completion rate fell to 41% (Archestra, reported by InfoQ, 4 October 2026).

The detail worth sitting with is not the zero. It is that OpenAPPA reached it without a parallel collapse in usefulness. A security layer that stops every attack by stopping most of the agent's work is not a result, it is a trade you already knew how to make. OpenAPPA's own ablation test shows the trade exists even inside its own design: when the engine's recovery plans were switched off, task completion dropped to 35% (Archestra, via InfoQ, 4 October 2026). The 89% figure only holds because the recovery machinery, not just the detection, is doing work.

Why this is not an isolated result

OpenAPPA's benchmark run lands on ground that was already documented as a problem. LayerX Security's 2025 enterprise browser telemetry found that 77% of employees paste data into generative AI tools, and that roughly four in five of those sessions run through unmanaged personal accounts, outside any control the employer owns (LayerX Security, 2025; coverage: The Register, 7 October 2025). Reco's State of Agent Security 2026, published 26 August 2026 from anonymised enterprise telemetry and analysis of 500 public Model Context Protocol servers, found 80% of AI tools running with no oversight at all, and 414 unsanctioned AI tools per 1,000 employees in smaller organisations.

Put the two findings together and a pattern holds across completely different measurement methods. One set of researchers counts what agents are exposed to and finds the exposure mostly ungoverned. Archestra's benchmark counts what happens when an attacker actually exploits that exposure, and finds that a widely used agent framework, Claude Code's auto mode, still lets one attack in ten succeed, while a purpose-built enterprise security layer, Microsoft FIDES, lets nearly a third through. Prompt injection is not a hypothetical in either data set. It is a measured rate.

What it means for a regulated enterprise

An attack success rate is not an abstraction once an agent holds real access. A workflow agent with reach into contracts, HR records or financial systems that fails against 10% of injection attempts is, in production terms, an agent that leaks under some unknown but non-trivial fraction of the prompts it processes in a year. Regulators are starting to ask for exactly this kind of number rather than a policy statement. The EU AI Act's record-keeping obligations assume an organisation can show what a system did and under what controls; an agent whose injection resistance has never been measured cannot support that answer. NIS2's management-level accountability for cybersecurity risk extends to the same gap: an unmeasured attack surface on a production agent is a risk nobody has signed off on, which is a different failure from a risk that was accepted.

The practical consequence is that "we deployed an enterprise AI agent" is no longer a sufficient claim. The question that follows is what rate of prompt injection it resists, under which benchmark, and what happens to the task when the defence fires. Microsoft FIDES's 41% completion rate on Bench-Corp is the sharp illustration: a defence that cuts task completion by more than half is not survivable for anyone running the agent to get work done, which is exactly why unprotected deployments persist.

What actually addresses it

OpenAPPA's architecture is informative independent of its benchmark numbers. It runs outside the agent's prompt and execution loop rather than inside it, configured through data sources, authorised audiences, trust levels and named authorities in a single `appa.toml` file. Each tool the agent can call carries a contract with three attributes: what it requires, what it changes, and what it affects. When a call would violate that contract, the system does not simply block it; it can run a sanitiser, escalate to a human or automated authority, or branch the call into a disposable sandbox to test the outcome before committing it. That recovery step, described in Archestra's published formal model as "APPA: Recoverable Information-Flow Control for Real-World LLM Agents," is what separates a 0% attack rate paired with 89% completion from a 0% attack rate paired with a dead agent.

The mechanism that generalises, regardless of which vendor implements it, is information-flow control enforced outside the model: the agent's own output is not the thing deciding whether a data flow is allowed. Something external to the prompt, holding its own record of trust levels and authorities, makes that call and can produce evidence afterwards of why it did.

What to check on Monday

Ask three things about every production agent with access to sensitive data. First, has it ever been run against a named injection benchmark, and what was the result, not "we tested it" but the actual rate. Second, where does the control that would stop an exfiltration attempt live: inside the same prompt and model call that could be manipulated, or outside it, in a layer the agent cannot talk its way past. Third, when that control fires, what happens to the task: does it fail cleanly with a usable alternative, or does the agent simply stop, because a defence nobody can afford to leave switched on is a defence that gets switched off within a quarter.

Related guides

Compliance

The EU AI Act Article 12 readiness guide

What record-keeping and human-oversight obligations actually require operationally from August 2026 — and the evidence an auditor will ask you to produce.

9 min read

Read the guide →

Risk

The credentials nobody reviews

Your AI agents hold OAuth tokens, API keys and service accounts that went through no approval process. The agent was reviewed. The studio was reviewed. The identity behind them was not.

5 min read

Read the guide →

Security

When the agents organised themselves: what the Hugging Face swarm means for accountability

Roughly 700 AI agents divided labour, traded favours and compromised production infrastructure across four regions. The uncomfortable part is not that it happened — it is that the account of what happened had to be reconstructed afterwards, by outside parties.

6 min read

Read the analysis →