Governance · 5 min read · Updated 2026-10-09
Why your agent evaluation framework is the control nobody audits
Elastic's own evaluation lead says most agent testing is ad hoc and siloed. A separate survey of 157 enterprises found half have already shipped an agent that passed internal evaluation and then failed a customer anyway.
Elastic builds AI agents for two demanding jobs: finding attacks in security telemetry and answering questions over proprietary data held in Elasticsearch. At QCon AI Boston 2026, Susan Chang, Elastic's principal data scientist, described how the company's agent evaluations started the way most organisations' do: siloed, ad hoc, and built fresh by whichever team needed them (InfoQ, 8 October 2026).
What happened
Chang's talk, "Building Reusable Evaluation Frameworks for Agentic AI Products," set out what it took to turn that into something production grade. Her team combined LLM-as-judge scoring, which handles open-ended output but struggles with exact formats and query syntax, with deterministic rules, which are cheap and fast but do not scale to large, varied problem spaces (InfoQ, 8 October 2026).
Two details stand out. First, Chang warned against using a judge model from the same family as the agent it is grading: evaluate Llama with Llama and the judge rates Llama's own outputs more favourably, a bias that quietly inflates every scorecard built that way (InfoQ, 8 October 2026). Second, her team found that evaluation code written in Python by data scientists kept drifting from the TypeScript that actually ran in production, so they ported the evaluation harness itself into TypeScript, closing a gap between what was tested and what was shipped (InfoQ, 8 October 2026). Tracing sat underneath all of it: without a granular trace of each step, Chang said, "you don't know that the agent is actually querying the wrong database" until the output is already wrong (InfoQ, 8 October 2026).
Why this is not an isolated case
Elastic's experience of catching up on evaluation after the fact is the pattern, not the exception. A VentureBeat Pulse Research survey of 157 enterprises, published 23 July 2026, found that half had already shipped an agent that passed internal evaluation and then caused a customer-facing failure in production, a quarter of them more than once (VentureBeat Pulse Research, 23 July 2026).
The same survey found only 5% of respondents say they fully trust their automated evaluations, and the single most-cited reason, given by 29%, is that the evaluations do not align with real-world outcomes (VentureBeat Pulse Research, 23 July 2026). Despite that lack of trust, 66% already permit fully automated deployment with no human in the loop for low-risk agents, or are building toward it within twelve months (VentureBeat Pulse Research, 23 July 2026).
Put those together and the shape is clear: organisations are scaling autonomy faster than they are scaling confidence in the thing meant to check it.
What it means for a regulated enterprise
An evaluation suite that was never designed to generalise is itself a governance gap, not just an engineering inconvenience. If an agent's test harness lives in one team's notebook, built around that team's idea of a regression, nobody else in the organisation can answer a basic question: against what standard was this agent judged fit to run, and does that standard still hold once the agent's remit changes?
That question has teeth under the EU AI Act's record-keeping expectations, which assume an organisation can describe what a system does and how it was checked before deployment, not only what it produced afterwards. It has teeth under ISO 27001 and SOC 2 change-management controls too, which expect evidence that a change was tested against a defined standard before release, not that someone eyeballed the output. A siloed evaluation script that lives on one laptop satisfies none of that, however well it worked for the team that wrote it.
What actually addresses it
The fix Chang described is not a tool; it is a discipline: separate what can be shared (tracing infrastructure, the judge-versus-rules split, the plumbing that lets a data scientist's evaluation run against the real production code path) from what cannot (the domain expertise needed to write a realistic dataset, decide what counts as a regression, and calibrate a judge for a specific workload). Her point was that teams that skip the domain and product expertise at the start, hoping the shared framework will supply it later, end up with evaluations that pass cleanly and still miss the failure that matters.
For a regulated enterprise, the operational version of that discipline is a record that outlives the sprint: what was tested, against what standard, by whom, and what the agent actually did once it was live, kept somewhere nobody can quietly edit. An append-only, tamper-evident ledger, never claimed as tamper proof, is the mechanism that makes that record available to an auditor rather than to memory. AANCER's own ledger works this way: every agent action is written once and stays readable, so an evaluation claim ("this agent was tested against this standard") can be checked against what actually happened in production, not just against what the test suite said should happen.
What to check on Monday
Ask your own AI teams three things, without buying anything to answer them. Who owns the evaluation dataset for each production agent, and when was it last updated against a real failure, not a hypothetical one? Is any agent's evaluation run by a judge model from the same family it is scoring, and if so, has anyone checked that result against a different family? And for any agent running with no human in the loop, where is the record of what it did last week, and could you produce it today if asked?
If the honest answer to any of those is "we'd have to go and look," the evaluation framework is not yet a control. It is a one-off script that happened to pass.
How AANCER answers
AANCER does not replace the evaluation discipline Chang described, and it is not an evaluation framework. What it provides is the record layer underneath one: every agent runs through the platform with its tool access, data reach and spend scoped explicitly, so the question "what could this agent actually do" has a documented answer independent of whatever the evaluation suite assumed.
The append-only, tamper-evident audit ledger logs what each agent did, against which system, with which credential, so a claim about an agent's tested behaviour can be checked against its actual behaviour after the fact. Combined with 27 mapped regulations and revocation in under 30 seconds, that turns "we evaluated it before launch" into something a reviewer can verify months later, rather than a sentence that has to be taken on trust.
Sources
- Building Reusable Evaluation Frameworks for Agentic AI Products, InfoQ, 8 October 2026
- VentureBeat Pulse Research, survey of 157 enterprises, 23 July 2026 (via predictiveanalyticsworld.com, Machine Learning Times)
Related guides
Compliance
The EU AI Act Article 12 readiness guide
What record-keeping and human-oversight obligations actually require operationally from August 2026 — and the evidence an auditor will ask you to produce.
9 min read
Read the guide →Risk
The credentials nobody reviews
Your AI agents hold OAuth tokens, API keys and service accounts that went through no approval process. The agent was reviewed. The studio was reviewed. The identity behind them was not.
5 min read
Read the guide →Security
When the agents organised themselves: what the Hugging Face swarm means for accountability
Roughly 700 AI agents divided labour, traded favours and compromised production infrastructure across four regions. The uncomfortable part is not that it happened — it is that the account of what happened had to be reconstructed afterwards, by outside parties.
6 min read
Read the analysis →