Before you sign the next high-consequence AI contract or renew an existing one, require the evaluation evidence behind the vendor's benchmark scores. Procurement should attach that disclosure to the decision record. If the vendor cannot supply it, note the gap and reduce the weight you place on the score. [1] [2]

A Number Without Its Conditions

Every benchmark score reflects at least five things: the model, the evaluation setup, the elicitation budget, the population of test items, and the contamination status. An elicitation budget is the compute, attempts, prompting, and tools used to draw out performance. A leaderboard may show the model and the score while leaving the other conditions unclear. Those conditions determine whether the number means what it appears to mean. [1]

Start with the elicitation budget. A model tested with extended compute, multiple retries, and retrieval tools can produce a different score than the same model under constrained, single-pass conditions. If the deployment you purchase does not match the evaluated setup, the score describes a configuration you are not buying. That is not proof of fraud. It is a mismatch between what was measured and what you will operate. [1]

The test population matters for the same reason. An aggregate score can hide weak performance on the particular cases your business cares about. A bank assessing a system for document review should ask whether the tested documents resemble its own languages, formats, and edge cases. Because the score depends on the sampled population, the comparison only helps when the buyer can see what that population contained. [1]

What Benchmark Contamination Means

Benchmark contamination occurs when prior exposure or access during the test lets a system reach an answer by a route other than the capability the benchmark claims to measure. The score may still look like a capability score. Without a record of the controls, the buyer cannot tell how much confidence that score deserves. [1]

A non-peer-reviewed preprint by Johanna Angulo, Víctor Yeste, and Hector Espinos-Morato proposes five contamination types: direct, derivative, temporal, distributional, and acquired. A preprint is an early research paper that has not completed peer review, so this taxonomy is a working framework rather than settled consensus. Its value for a buyer is practical: different leakage routes require different controls. [1]

Direct contamination means the test item or its answer reached the model during training. A private held-out test set—questions kept from the model's training data—can close that route. The paper argues that this control does not close the other routes. [1]

Derivative contamination means the source material used to create a private test was already in the training data. Temporal contamination means the model's training cutoff came after the outcome it is supposedly being asked to predict. Distributional contamination means the model may recognize a heavily represented pattern rather than perform the reasoning the test claims to measure. A buyer does not need to become an evaluation scientist to understand the consequence: saying that a test was private answers only one question. [1]

Acquired contamination is different because it happens during a specific evaluation run. A system may retrieve information from the web, tools, or files in the test environment. That means the disclosure must travel with the reported score, because a benchmark publisher cannot know what access every later evaluator will allow. The procurement question is not only whether the test was private. It is what the system could reach while taking it. [1]

The Audit Shows a Disclosure Gap

The preprint's authors reviewed forty-one documents with two external coders using a preregistered scoring instrument. [1] The authors report low agreement between coders on individual variables, especially on whether some fields applied. That is a significant limitation. The sample is not a field-wide census, and the findings should not be generalized to every vendor evaluation. [1]

Even with those limits, the result identifies a diligence problem worth testing in your own purchases. No document addressed all five contamination types. [1] Elicitation budgets appeared in thirteen percent of the documents reviewed. [1] These are paper-reported findings from an early, non-peer-reviewed contribution, not proof that all published scores are wrong.

The paper proposes a disclosure with four fields: the test population, elicitation budget, contamination controls, and whether the test can be regenerated as its contents age. It allows "unknown" as an answer. The disclosure records claimed controls; it does not detect contamination, verify honesty, or guarantee that the benchmark predicts performance in your workflow. [1]

That distinction explains why an explicit unknown can still improve a buying decision. If a vendor says the training-data overlap is unknown, procurement can price that uncertainty into the comparison and require stronger workflow testing. A missing field creates no such decision path because the buyer cannot tell whether the question was considered. Unknown does not make a score trustworthy. It makes one limit visible.

NIST Puts Reporting in the Procurement Path

In January, the National Institute of Standards and Technology released initial public draft guidance called NIST AI 800-2, covering practices for automated benchmark evaluations of language models. [2] NIST describes the draft as preliminary and voluntary. It is not binding regulation or a ratified standard. [2]

NIST explicitly names business decision-makers and procurement specialists as people who can benefit from robust, well-communicated evaluation practices. It says those practices can inform procurement and implementation decisions. The guidance groups evaluation work around defining objectives and selecting benchmarks, implementing and running evaluations, and analyzing and reporting results. [2]

NIST also says automated benchmarks cannot meet every evaluation objective. That means a well-documented score remains one input, not the purchase decision. High-consequence systems still need testing on the buyer's workflow, along with security, legal, and operating review. The disclosure improves the score's evidentiary value because it shows what was tested and what remains uncertain. It does not turn the score into proof of safe deployment. [2]

Because the guidance is voluntary, it does not compel a vendor to answer a buyer's questions. It does show that asking for evaluation objectives, methods, and reporting is consistent with the direction of government measurement guidance. Procurement can make that evidence a contract and diligence requirement before regulation demands it. [2]

Proportionality Prevents Bureaucracy

A reasonable objection is that requiring more evidence can slow procurement or disadvantage smaller vendors. The answer is proportionality, not abandoning the requirement.

For systems that make consequential decisions or act with material authority, the evaluation disclosure should sit in the decision record just as a security review or legal opinion does. Consider a system that screens applicants, drafts advice with legal consequences, or directs financial workflows. A wrong comparison can affect people or commit the company. The stakes justify asking what was measured and under which conditions.

For reversible, low-stakes assistance, the standard can be lighter. A writing tool that an employee can override before anything leaves the company does not require the same diligence as an autonomous system acting at scale. The accountable business owner should set the evidence threshold according to the authority the system receives and the cost of a wrong result.

A short standardized request can also reduce procurement work over time. Ask the same questions across every high-consequence evaluation: What claim did the benchmark measure? What setup and resource budget produced the score? Which contamination routes were controlled? What uncertainty or unknown remains? Reusing those questions allows procurement to compare like with like and prevents a second diligence cycle built around evidence that was mismatched from the start.

A score without its conditions is a claim without enough support. You would not let a financial metric carry a major decision without knowing how it was calculated. An AI benchmark deserves the same treatment.

Before the next high-consequence purchase or renewal, assign procurement and the accountable business owner to obtain the evaluation disclosure and attach it to the decision record. If the vendor cannot supply it, record the gap, strengthen your workflow-specific testing, and reduce the weight placed on the score.