Governance · 5 min read · Updated 2026-10-06
Pizza Bot and the agents nobody is watching
AWS engineers open-sourced an inbox for background AI agents that work unattended and pause for approval. The design is sound. The reason it matters is how few organisations could say the same about the agents they already run.
A team of AWS engineers has open-sourced Pizza Bot, a self-hosted application that lets AI agents run tasks in the background and report results through an email-style inbox. Agents work on schedules or webhook triggers, hand specialised work to other agents, and pause for a human decision before anything consequential happens (InfoQ, Renato Losio, 4 October 2026).
The project began in April 2025 as Joseph Dolivo's side project, "JoeBot," built to automate his own CRM logging. Dolivo, an AWS Principal Technologist, and Igor Fil, an AWS Solutions Architect, developed it into Pizza Bot, released under the Apache 2.0 licence on GitHub. It stores agent state, threads, checkpoints, attachments and logs locally rather than in a vendor's cloud, and supports multiple model providers and MCP servers. AWS is explicit that it ships no support contract and no service-level agreement: it is a community project (InfoQ, 4 October 2026).
The design instinct behind it is correct. An agent that works like an inbox, not a chat window, does not need a human staring at it to be useful, and a checkpoint before a consequential action is a real control, not a courtesy. The question Pizza Bot does not answer, because it is not trying to, is what happens when an organisation runs dozens of these unattended agents at once, built by different teams, talking to different systems, with nobody keeping the master list.
Why this is not a one-off
The gap Pizza Bot sits inside is already measured at scale. Gravitee's 2026 survey of 750 CTOs and VPs of Engineering at major US and UK firms found 7.24 million AI agents deployed across the two markets, of which 2.45 million operate fully autonomously, roughly one in three, with no human reviewing their actions in real time. Only 7% of the organisations surveyed had named a single accountable person for agent behaviour, and the same survey found agent numbers doubling roughly every six months (Gravitee, published via IT Brief, 24 June 2026).
That pressure to deploy ahead of the controls is not incidental. 80% of organisations in the same survey said they felt pressure to ship AI agents despite knowing their security measures were incomplete, and over a third traced that pressure to the boardroom (Gravitee, 24 June 2026).
The visibility problem compounds it. Reco's State of Agent Security 2026, built from anonymised enterprise telemetry and analysis of 500 public Model Context Protocol servers, found 80% of AI tools running with no oversight at all, and in smaller organisations 414 unsanctioned AI tools per 1,000 employees (Reco, 26 August 2026). A background agent that runs on a webhook and reports into its own inbox is, by construction, exactly the kind of system that disappears from a manual inventory the week after launch.
What it means for a regulated enterprise
A background agent is a long-running process with credentials, a schedule, and the ability to act without anyone watching the screen at the moment it acts. That is precisely the profile regulators are starting to ask about directly.
Under the EU AI Act, record-keeping obligations assume an organisation can state what a given system did and when. A webhook-triggered agent that ran at 2am and updated three records has no witness unless the run itself produced a log that survives independently of the agent's own say-so. Under NIS2, cybersecurity risk management carries personal management accountability, and an autonomous agent with unreviewed reach into operational systems is squarely the kind of risk that obligation was written to cover. Access reviews under ISO 27001 and SOC 2 enumerate accounts and what they can reach; a background agent spun up from a self-hosted template is still an account, whether or not it appears on anyone's list.
The practical question a regulator, auditor or incident responder will ask is not "do you have a policy for autonomous agents." It is "show me the agent, what triggered it, what it touched, and who approved the step that mattered." An inbox of completed runs is a start. It is not, by itself, an answer to that question unless the identity behind each run, its scope, and its approval trail are recorded somewhere nobody can quietly edit.
What actually addresses it
The mechanism that closes this gap is not a better UI for reading agent output. It is three things working together: an issued identity for every agent that acts (so "which agent did this" has one answer, not a guess from a log line), a scope on that identity that is enforced rather than assumed (so the agent cannot reach further than its job requires, regardless of what credential it was cloned from), and an approval step on consequential actions that is recorded at the moment it happens, not reconstructed afterwards from memory.
Pizza Bot's human-approval checkpoint is a genuine instance of the third piece. What is missing from a self-hosted, community-run inbox is the first two: a durable, organisation-issued identity behind each agent, and a record of what that identity actually did that an auditor can read without trusting the agent's own account of itself.
What to check on Monday
Before buying anything, list every scheduled or webhook-triggered agent currently running against production systems, and for each one answer three questions: who owns it, what can it reach that it does not strictly need, and where is the record of what it did last week that nobody connected to the agent can alter. If any answer is "we'd have to ask the team that built it" or "there isn't one," that agent is the gap, not the exception.
Related guides
Compliance
The EU AI Act Article 12 readiness guide
What record-keeping and human-oversight obligations actually require operationally from August 2026 — and the evidence an auditor will ask you to produce.
9 min read
Read the guide →Risk
The credentials nobody reviews
Your AI agents hold OAuth tokens, API keys and service accounts that went through no approval process. The agent was reviewed. The studio was reviewed. The identity behind them was not.
5 min read
Read the guide →Security
When the agents organised themselves: what the Hugging Face swarm means for accountability
Roughly 700 AI agents divided labour, traded favours and compromised production infrastructure across four regions. The uncomfortable part is not that it happened — it is that the account of what happened had to be reconstructed afterwards, by outside parties.
6 min read
Read the analysis →