Rendered from docs/profiles/competition-evaluation.md at build time without changing its status. The repository source controls if this presentation differs.
Competition evaluation profile#
Acronyms: artificial intelligence (AI); machine learning (ML); Verifier Standard (VSTD).
Reader aid: concept glossary and primary precedents.
Status: non-normative VSTD-1/VSTD-Graph integration profile Version: 0.1 Date: 2026-08-21
This profile applies VSTD receipts and provenance hypergraphs to predictive-AI, scientific-ML, agent, and other scored evaluations. It does not add a new VSTD verdict and does not claim adoption, affiliation, certification, or endorsement by any conference, competition, benchmark, or organizer.
The bounded public wording in docs/CLAIMS_AND_LIMITS.md controls if a shorter phrase in this non-normative profile could be read more broadly.
VSTD-2 relationship#
Conceptually, this profile selects Verifier Standard (VSTD)-2 coordinates across the submission, evaluator, environment, score, and their seams. It does not emit a VSTD-2 receipt or establish VSTD-2 conformance by itself. Each native scorer or benchmark adapter must attribute its output to the exact selected coordinates, preserve translation loss and horizons, and bind a separate assessment before any native result becomes a VSTD judgment.
1. Evaluation surface#
An integration declares the exact surface before it reports a verified result:
- task and rules version;
- training, reference, and permitted external-data snapshots;
- model, agent, checkpoint, adapter, and configuration identities;
- prediction or submission artifact and submission timestamp;
- execution image, runtime, hardware class, seed policy, and command;
- evaluator/scorer source and configuration identity;
- held-out input commitment or an explicit
UNKNOWN/TRUST_ROOThorizon; - outcome-resolution source and resolution timestamp, when predictions concern events resolved later;
- raw evaluator output and derived score report; and
- limitations, exclusions, and falsification conditions.
Coordinates outside that surface do not inherit its verdict.
2. Minimum recorded chain#
rules + data snapshots + permitted externals
-> build/train/fine-tune transformation
-> model or agent snapshot
-> prediction/submission transformation
-> immutable submission artifact
-> evaluator/scorer transformation
-> raw metrics
-> score report
Each artifact receives a stable identifier and content digest. Each transformation records its input and output roles, software identity, parameters, environment, and evidence classification. A declaration is not relabeled as direct observation or reproduction by a distinct actor.
3. Predictive-evaluation time boundary#
For a prediction resolved after submission, the receipt records at least:
prediction_emitted_at;prediction_freeze_digest;- allowed update or abstention policy;
outcome_resolved_at;- resolution-source identifier and snapshot digest;
- scoring-rule identifier and parameters; and
- whether the prediction, resolution, and scoring observations came from independent channels.
The integration MUST NOT overwrite a frozen prediction after outcome information becomes available. Corrections are additive and link to the challenged or superseded artifact.
4. Hidden tests and organizer-controlled artifacts#
A participant normally cannot observe or serialize hidden tests. The participant receipt therefore records an explicit horizon. An organizer can later close part of that horizon by publishing a commitment, signed attestation, disclosed snapshot, or evaluator receipt reproducible by a distinct actor.
Absence of access is not evidence of hidden-test integrity. A participant-side VERIFIED result MUST NOT imply that the organizer's hidden corpus was uncontaminated, that the evaluation prevented leakage, or that the public leaderboard is authoritative.
5. Claims licensed by this profile#
With corresponding evidence, an implementation may state that:
- the recorded submission bytes match a named digest;
- the recorded evaluator version produced the bound raw metrics when rerun in the declared environment;
- the score report is a deterministic derivation of those metrics under the named scoring rule;
- the recorded ancestry graph contains the declared datasets, model snapshot, submission, evaluator, and report relationships;
- a particular recorded policy formula passed; or
- a challenged or revoked ancestor has the enumerated downstream blast radius.
Each statement remains bounded to the named snapshots, mechanisms, and evidence.
6. Claims not licensed by this profile#
The profile does not establish:
- empirical truth or future generalization beyond evaluated inputs;
- correctness or representativeness of the benchmark design;
- authenticity of an unevidenced origin, license, contributor, or outcome source;
- absence of hidden inputs, leakage, contamination, evaluator manipulation, or out-of-band execution;
- ranking, prize eligibility, rule compliance, or organizer acceptance unless the applicable authority supplies bound evidence; or
- endorsement by VSTD or by a competition organizer.
7. Conformance wording#
Use a bounded statement such as:
The submission and score receipt conform to the VSTD competition evaluation profile 0.1 for the declared artifact, evaluator, and provenance surface. Hidden-test integrity and organizer acceptance remain outside the participant-observable surface.
Do not shorten this to “the model,” “the competition result,” or “the prediction is verified” without naming the exact coordinate and evidence that passed.