Back to BlogResearch

How to Read the APort Vault Benchmark Without a Leaderboard

Read APort Vault by outcome, policy level, replay track and matched denominator. Understand what its 4,371 attacks and 14-model study can establish.

5 min read
by APort

Summary: A buyer comparing prompt injection benchmarks needs to know whether a reported success means the model requested an action or an unauthorized transfer actually executed. The APort Vault paper dated September 18, 2026 separates those events. It reports 225,964 completed evaluations from 4,371 human-authored attacks across 14 models, five policy configurations and two replay tracks, comparing model-alone and guarded conditions.

Use the benchmark to examine an authorization boundary, not to turn every payment request into a breach score. This reader's guide explains the denominators and comparisons. The existing forged-receipt analysis follows a selected recorded case; the payment engineering guide applies the aggregate evidence to a deployment review.

Read the unit before the percentage

An attack is source material. An evaluation is a replay of an attack under a model, track and architecture. A transfer call is an attempted tool invocation within an evaluation. An evaluation can contain multiple calls, so those counts cannot be substituted for one another.

The September 18 paper planned 244,776 evaluations: 4,371 attacks multiplied by 14 models, two tracks and two architectures. Each attack belongs to one policy level; do not multiply the plan by five again. The completed denominator is 225,964, and missing evaluations limit comparisons that assume complete coverage.

A request at Level 4 can be the intended behavior

The September 18 source reports a complete model-alone, single-turn comparison on 1,293 Level 4 prompts per model. On 809 of those prompts, all 14 models requested and recorded a successful payment to the level's allowlisted recipient. That is 62.6% of prompts eliciting shared behavior across models, not a compromise rate.

Level 4 authorizes qualifying documented transfers to that recipient. A chart of request rates at that level measures action propensity under the stated task. It does not rank models by unauthorized transfers, and a model that refuses more often is not automatically more useful.

Compare the same outcome in both conditions

For Levels 2 to 4, both replay tracks, the September 18 paper reports unpermitted transfers in 140 of 76,842 completed model-alone evaluations and 0 of 69,297 guarded evaluations. The unequal denominators matter. On the matched set of 68,970 model, prompt and track triples, the counts are 105 and 0 respectively.

The matched comparison reduces differences caused by missing pairs. It does not turn a frozen replay into a new adaptive attacker or establish performance on arbitrary production tasks. The guarded zero spans 790 source sessions; the reported per-session upper bound is 0.38%. Zero observed events is not proof of a zero population risk.

Replay changes the question

The September 21, 2026 replay study analyzes model-alone complete-case cohorts. At Level 2, 690 prompts from 243 source sessions yield 39 requests in 9,660 final-user-turn evaluations versus 307 in 9,660 capped source-user-turn evaluations. The rates are 0.40% and 3.18%.

Those rates describe requests under different replay inputs, not unauthorized transfers. A model ranking can stay positively associated while action frequency changes. Preserve the track in any comparison you take from this study.

Build a result card before comparing models

For any figure copied into a procurement review, write down its provenance in the same row as the value:

Field Example for the matched authorization comparison
Source Vault paper, September 18, 2026
Outcome Evaluation with an unpermitted transfer, not merely a request
Policy level Levels 2 to 4
Replay coverage Both tracks, restricted to matched model/prompt/track triples
Conditions Model alone versus guarded
Denominator 68,970 matched triples in each condition
Observed counts 105 model alone; 0 guarded

This is a restatement of the public matched result, not a new analysis. Do not add the full-cohort counts to it: the matched set is a subset used for another comparison. Likewise, do not divide transfer-call denials by evaluation counts and label the result a per-request denial rate.

When a model or replay pair is missing, choose and disclose the missing-data rule before ranking results. A complete-case subset can improve comparability while changing which attacks remain in the population. A pooled result can preserve more observations while mixing unequal cohorts. Both choices need labels; neither supplies a production incident probability by itself.

Reproduce first, decide second

Use the released dataset and scoring code with the paper's level passports and outcome definitions. Record the source version, level, track, architecture, completed denominator and matched-subset rule beside a reproduced metric.

The benchmark has no benign-task evaluation, so it cannot establish a production false-denial rate. Its strategic value is evidence that a policy check at the execution boundary can reject unpermitted transfers even when a model proposes them, within the tested configurations.

Plan your Team Pilot with a legitimate workload and prohibited-action fixtures that test your own boundary.