Back to BlogResearch

What 225,964 AI Payment Evaluations Say About Guardrails

Use APort Vault payment evidence without conflating requests, decisions and transfers. Map the measured boundary to a payment-agent acceptance test.

8 min read
by APort

Summary: A team giving an AI agent payment tools needs to prove that a forbidden transfer cannot reach execution while legitimate work remains possible. The APort Vault paper dated September 18, 2026 reports 225,964 completed evaluations across five policy levels, 14 models, two replay tracks and model-alone versus guarded architectures, using 4,371 human-authored attacks.

At Levels 2 to 4 across both tracks, unpermitted transfers occurred in 140 of 76,842 model-alone evaluations and 0 of 69,297 guarded evaluations. Those are experimental outcomes with unequal completed denominators, not a claim of universal payment safety. The useful deployment question is which component checks the proposed recipient and amount before the payment tool acts.

The failure mode

A payment agent can hold a valid credential, call the expected endpoint and pass JSON validation while still proposing the wrong transfer. The business question is whether this exact amount, recipient and account relationship is authorized. Keep processor controls, account permissions and fraud checks, then identify any task-specific restriction they do not already enforce.

An approval written into a chat message is not equivalent to an approval bound to a payment. The wrapper should obtain account ownership and recipient records from trusted services. The agent may propose an amount or reason; it must not decide that its own proposal satisfies the required approval.

What to check before a payment executes

A context checklist is useful when each field is tied to a real enforcing component. This is a design checklist, not a statement that every APort policy pack implements all these checks.

Context field Control to design Trusted source or enforcement owner
action Distinguish refund, charge, payout, transfer and void Server selects the route and policy
amount Validate units, sign, per-operation and cumulative limits Payment service plus supported policy checks
currency Bind amount units and permitted currencies Order/account record and processor
customer_id Prevent cross-account operations Authenticated account context
recipient_id Enforce the intended recipient set Trusted recipient registry, checked at dispatch
region Apply applicable account restrictions Trusted account data and service rules
reason_code Separate routine and exceptional operations Validated workflow state
idempotency_key Prevent duplicate settlement on retries Payment service and processor's idempotency contract
approval_id Bind approval to the exact operation Approval service with expiry and scope checks

Send only the policy-relevant fields. Credentials, full customer records and card data are not needed merely to compare an amount or recipient. Where a policy pack lacks a required validator, implement the rule in the service or restrict the workflow; adding a field to JSON does not enforce it.

The minimal architecture

agent proposal
    |
trusted payment wrapper: validate, resolve account, freeze arguments
    |
policy evaluation -> deny in enforce mode -> no processor call
    |
allow + required service/approval checks
    |
processor call with idempotency key
    |
record processor result separately from policy decision

This is a conceptual payment integration. The authorization result is one input to execution, not settlement confirmation. A timeout after submission requires checking processor state under the same idempotency key before retrying.

Example: protect a refund route

The Express middleware implementation is a useful starting point, but its requirePolicy helper spreads the request body into context and accepts body-supplied passport or policy inputs. A route that needs a fixed server policy must validate and constrain inputs before calling a verifier. The simple middleware snippet alone would not demonstrate that requirement.

The following pseudocode shows the ownership and order of operations. The function names represent application code, not a drop-in APort SDK API:

authenticate caller and load the caller's account
validate proposed order, amount in minor units, and currency
load refundable order and permitted recipient from trusted storage
reject if the order belongs to another account
construct immutable refund arguments from those validated values
select the refund policy and agent identity on the server
request authorization for only the fields that policy supports
if verification failed or policy denied: stop before processor dispatch
check any remaining service-owned limits and bound approval
submit those same refund arguments using the stable idempotency key
store decision reference, processor reference, and observed outcome

A repeated request should reuse the operation identity, not create a new one just because the model restated the task. An allow followed by a processor failure remains a failed payment. A deny followed by continued execution is an enforcement issue even if a dashboard contains the deny.

Keep five events in the record

A payment request shows that the model proposed a tool call. A policy decision shows what the verifier concluded. A successful simulated payment shows that execution succeeded. Recipient membership identifies whether the destination was in the passport's allowed set. An unpermitted transfer combines an executed payment with a recipient outside that set.

Collapsing these into an “attack success” counter loses the property being tested. An agent can propose a prohibited payment that the policy rejects, or request a permitted payment that the simulator cannot execute. Record those outcomes separately in your own tests.

The layer permitted activity as well as denying calls

The following figures are from the September 18 paper, Levels 2 to 4, guarded condition, both replay tracks. Counts use different units intentionally:

Measured event Count and denominator
Transfer calls evaluated by policy 25,640 calls within 69,297 completed evaluations
Allowed transfer calls 25,453 of 25,640 calls
Denied transfer calls 187 of 25,640 calls
Forbidden-recipient denials 148 of 187 denied calls
Executed simulated payments 25,370 payments within 69,297 evaluations
Registered unpermitted transfers 0 of 69,297 evaluations

The executed-payment count differs from the allowed-call count. An authorization decision and downstream execution are distinct stages. These results reject the explanation that the guarded condition simply refused all payments; they do not measure usefulness on a representative benign workload.

For an aligned comparison, the same paper reports 105 unpermitted-transfer outcomes model alone versus 0 guarded across 68,970 matched model, prompt and track triples at Levels 2 to 4. The benchmark reading guide explains why the matched and full completed cohorts are both reported.

Put business authority outside the conversation

An agent's statement that a recipient was approved is not the approval record. A payment wrapper should construct policy context from trusted account and recipient data, plus the proposed amount and action. Keep credentials, card details and full customer records out of that context.

Use the available policy packs to inspect the exact fields and checks for the selected action. Do not assume every business requirement is enforced merely because a context field has the right name. If a requirement belongs in the payment service, keep it there and test it there.

The OAP enforcement contract also matters: a signed denial in warn mode may permit continuation. Use enforced denial before enabling a workflow whose safety requirement is that prohibited transfers cannot execute.

Build a payment acceptance test from the boundary

Use a simulator or sandbox account to exercise a permitted recipient, a forbidden recipient, an invalid amount and a repeated request. Verify which component enforces each rule. Test verifier unavailability and observe whether execution stops under the chosen configuration.

Join the decision to the payment result without treating an allowed decision as proof of settlement. Keep processor authorization, idempotency and account controls. The Vault experiment measures a simulated payment boundary; it does not validate a production processor integration or quantify real losses prevented.

Checklist for AI payment agent rollout

  1. Choose one action and a simulator or sandbox account before adding production credentials.
  2. Define the permitted account, recipient and amount conditions in executable checks.
  3. Map each context field to the component that validates it; reject unsupported assumptions.
  4. Make the selected policy, passport and enforcement mode operator-owned.
  5. Bind required approval to the immutable operation and test expiry or changed arguments.
  6. Exercise idempotency at the processor boundary, including a timeout after submission.
  7. Keep signed hosted decisions and execution results as separate records joined by reference.
  8. Use report-only only with an explicit acceptance of continued execution; switch to enforce mode before relying on denial to stop money movement.
  9. Test legitimate activity, forbidden recipients, invalid amounts, missing approval and verifier outage.
  10. Assign an owner to investigate denials and reconcile unknown processor outcomes.

The benchmark gives this rollout an evidence base, not a substitute for acceptance tests. Follow the research paper for the tested policies, the identity comparison for caller identity, and the policy catalog for the actual action schema.

Plan your Team Pilot for one payment workflow with its permitted and prohibited outcomes specified in advance.

Frequently Asked Questions

Common questions about this topic.

What is AI payment agent security?

It includes caller identity, processor and account controls, and checks on the exact proposed operation before dispatch. Test recipient, amount, approval and retry behavior at their enforcing components.

Did the Vault guardrail simply reject every payment?

No. The September 18, 2026 paper reports 25,370 simulated payments within 69,297 guarded evaluations at Levels 2 to 4 across both tracks. It also reports 0 unpermitted-transfer outcomes in that cohort. These are different units, and the study has no benign-workload evaluation.

Does an allowed decision prove settlement?

No. Authorization, submission and settlement are separate stages. Retain the processor result and reconcile unknown outcomes using the processor idempotency contract.