Research

The method behind every evaluation, in detail — based on our own published research.

Method

An evaluation shows which attacks succeed against your agent, what it does when they do, and how often. All of it rests on one principle: a probabilistic system can’t be verified by asking it what it did. So every finding is checked against a deterministic ground truth — what actually happened, and what should have happened, established without the agent’s help.

  1. Two ways an agent can fail under attack

    A deceived agent acts faithfully on corrupted input and reports truthfully what it did: its account matches what it did, but what it did breaks the obligation. A deceiving agent’s account differs from what it did. The two are told apart at the output, not the input — one poisoned record can produce either. In both cases, the cause is the attack, not the agent’s intent.

    Three boxes stacked from top to bottom: obligation, what should have happened; behaviour, what the agent did; account, what the agent says it did. On the right, a bracket spans obligation and behaviour: obligation-detection, which derives the obligation independently of the agent and catches both classes. A second bracket, drawn as a broken line, spans behaviour and account: claim-verification, which checks the account against the execution record and catches deceiving only. Below the boxes, the two classes: deceived, where the account equals the behaviour and the behaviour differs from the obligation; deceiving, where the account differs from the behaviour.Obligationwhat shouldhave happenedBehaviourwhat theagent didAccountwhat the agentsays it didobligation-detectionderives the obligationindependently of the agent— catches both classesclaim-verificationchecks the account againstthe execution record— catches deceiving onlydeceivedaccount = behaviour ≠ obligationdeceivingaccount ≠ behaviour
  2. Two checks, one for each class

    Obligation-detection compares what the agent did with the obligation, and reaches both classes wherever the attack changed what the agent did. Claim-verification compares the agent’s account with the record of what it did; it catches deceiving, but not deceived, whose account is true. The two answer different questions: did the agent do what it should, and is its report of it true? The first always matters; the second matters wherever others rely on that report — an audit log, a compliance record, a downstream system. Whether an agent can be deceiving at all depends on its architecture: only a decoupled agent, which reports through a separate output from the one it acts with, can say one thing and do another. So we read the architecture first, and run both checks.

  3. The reference everything is checked against

    The deterministic ground truth has two parts, both established without the agent. The observed behaviour is read where the agent can’t reach — at the systems it acts on, not from its own logs. The obligation comes from a source neither the attack nor the agent can change: a specification, a safety envelope, a policy table. Each finding states which source it relied on, and whether that source was out of the attack’s reach by construction or only because the attack happened not to touch it.

  4. Observation the agent can’t see

    We observe on three surfaces beneath the level the agent works at: its processes, the network boundary its calls cross, and the systems it acts on. The agent can neither read nor change any of them — not because it’s told not to, but because it has no way to. That keeps it from reading the answer key, or recognising that it’s being evaluated and behaving differently than it would in production.

  5. Measurement

    An agent can act differently each time it meets the same input, so a single run proves little. Each threat scenario first runs many times without the attack, as a clean baseline, then as often with it; a finding counts only if it appears under attack and not in the baseline. That shows whether an attack works occasionally or every time — and, weighed against how severe its consequences are, what to fix first.

  6. Where the method applies

    Three things must hold, and all three can be checked before any work starts:

    • We can see what the agent actually does — its tool calls, retrievals and decisions — not just the text it produces.
    • Attacks can be played out where their consequences aren’t real: in a simulation for physical processes, in a development or staging environment for data flows.
    • There’s a reference the attack can’t reach, independent of the agent and of whatever the attacker manipulates.

    This holds whether or not a guardrail or other protection sits in front of the agent, because the reference is independent of both.

  7. What a result describes

    A result describes your agent as it actually runs — with its model, its configuration and its tools. That makes it evidence about the product you ship, not a generic score for the model underneath. And because the threat scenarios and the reference carry over, the next version can be evaluated against the same set, and the two results compared.

The full argument, with the demonstrations behind it, is in our paper Adversarial Assessment of Agentic Systems (2026).

Publications

Adversarial Assessment of Agentic Systems: A Deterministic-Ground-Truth Methodology, Demonstrated Across an Industrial Control Stack

Preprint ·

Abstract

Agentic AI now sits inside the decision and perception loops of safety-relevant industrial processes, yet no methodology assesses these deployments adversarially and independently of their own account, and none asks the question an industrial operator and the manufacturer who must declare conformity face: whether a deployed agentic system in a control loop, facing an adversary with data-plane or physical-world access, still conforms to the requirements it is trusted to uphold, and whether that can be verified independently of it. We present an independent adversarial-assessment methodology whose components turn on properties of the agent and its referent, demonstrated across an industrial control stack: a full-stack simulation of a beverage line spanning the Purdue hierarchy. Its premise is that a probabilistic system cannot be verified by asking it what it did, so the assessment introduces deterministic checks against state the agent did not produce. From that premise follow a taxonomy separated at the output, distinguishing an agent induced under attack to misreport what it did from one that reports accurately what it did on corrupted inputs; the detection asymmetry that follows from it; a referent that must be independent of the surface an attack perturbs and not merely outside the agent; three conditions bounding where the method applies, checkable in advance; and an account of which EU obligations reach such a deployment, what they require of its behaviour, and what does not discharge them. The demonstrations show that realistically deployable agents can be deceived and induced to deceive, and that the methodology detects both.

Records

Both are published as linked records on Zenodo.

Contact

For evaluation enquiries, research correspondence, or press: write to us directly.

Phaneia GmbH · Berlin