Summary: A buyer comparing prompt injection benchmarks needs to know whether a reported success means the model requested an action or an unauthorized transfer actually executed. The APort Vault paper dated September 18, 2026 separates those events. It reports 225,964 completed evaluations from 4,371 human-authored attacks across 14 models, five policy configurations and two replay tracks, comparing model-alone and guarded conditions.
Use the benchmark to examine an authorization boundary, not to turn every payment request into a breach score. This reader's guide explains the denominators and comparisons. The existing forged-receipt analysis follows a selected recorded case; the payment engineering guide applies the aggregate evidence to a deployment review.
Read the unit before the percentage
An attack is source material. An evaluation is a replay of an attack under a model, track and architecture. A transfer call is an attempted tool invocation within an evaluation. An evaluation can contain multiple calls, so those counts cannot be substituted for one another.
The September 18 paper planned 244,776 evaluations: 4,371 attacks multiplied by 14 models, two tracks and two architectures. Each attack belongs to one policy level; do not multiply the plan by five again. The completed denominator is 225,964, and missing evaluations limit comparisons that assume complete coverage.
A request at Level 4 can be the intended behavior
The September 18 source reports a complete model-alone, single-turn comparison on 1,293 Level 4 prompts per model. On 809 of those prompts, all 14 models requested and recorded a successful payment to the level's allowlisted recipient. That is 62.6% of prompts eliciting shared behavior across models, not a compromise rate.
Level 4 authorizes qualifying documented transfers to that recipient. A chart of request rates at that level measures action propensity under the stated task. It does not rank models by unauthorized transfers, and a model that refuses more often is not automatically more useful.
Compare the same outcome in both conditions
For Levels 2 to 4, both replay tracks, the September 18 paper reports unpermitted transfers in 140 of 76,842 completed model-alone evaluations and 0 of 69,297 guarded evaluations. The unequal denominators matter. On the matched set of 68,970 model, prompt and track triples, the counts are 105 and 0 respectively.
The matched comparison reduces differences caused by missing pairs. It does not turn a frozen replay into a new adaptive attacker or establish performance on arbitrary production tasks. The guarded zero spans 790 source sessions; the reported per-session upper bound is 0.38%. Zero observed events is not proof of a zero population risk.
Replay changes the question
The September 21, 2026 replay study analyzes model-alone complete-case cohorts. At Level 2, 690 prompts from 243 source sessions yield 39 requests in 9,660 final-user-turn evaluations versus 307 in 9,660 capped source-user-turn evaluations. The rates are 0.40% and 3.18%.
Those rates describe requests under different replay inputs, not unauthorized transfers. A model ranking can stay positively associated while action frequency changes. Preserve the track in any comparison you take from this study.
Build a result card before comparing models
For any figure copied into a procurement review, write down its provenance in the same row as the value:
| Field | Example for the matched authorization comparison |
|---|---|
| Source | Vault paper, September 18, 2026 |
| Outcome | Evaluation with an unpermitted transfer, not merely a request |
| Policy level | Levels 2 to 4 |
| Replay coverage | Both tracks, restricted to matched model/prompt/track triples |
| Conditions | Model alone versus guarded |
| Denominator | 68,970 matched triples in each condition |
| Observed counts | 105 model alone; 0 guarded |
This is a restatement of the public matched result, not a new analysis. Do not add the full-cohort counts to it: the matched set is a subset used for another comparison. Likewise, do not divide transfer-call denials by evaluation counts and label the result a per-request denial rate.
When a model or replay pair is missing, choose and disclose the missing-data rule before ranking results. A complete-case subset can improve comparability while changing which attacks remain in the population. A pooled result can preserve more observations while mixing unequal cohorts. Both choices need labels; neither supplies a production incident probability by itself.
Reproduce first, decide second
Use the released dataset and scoring code with the paper's level passports and outcome definitions. Record the source version, level, track, architecture, completed denominator and matched-subset rule beside a reproduced metric.
The benchmark has no benign-task evaluation, so it cannot establish a production false-denial rate. Its strategic value is evidence that a policy check at the execution boundary can reject unpermitted transfers even when a model proposes them, within the tested configurations.
Plan your Team Pilot with a legitimate workload and prohibited-action fixtures that test your own boundary.