An adapter is not a result.
Grand TEVV includes interfaces for ten external evaluation ecosystems and standards-oriented views. The supplied package explicitly states that those suites were not run; each entry here describes intended role and current acquisition or implementation state.
Open the system reference at this layer.
Follow the mechanism through numbered execution, a worked trace, artifact-bound evidence, adverse results, and validation still required.
Model and agent evaluation
These contracts would connect Shadow's cases and evidence schema to broader evaluation runners.
DIOPTRAReproducible experiment orchestration
A local experiment YAML and plugin scaffold exist; live framework testing was not performed.
+
Reproducible experiment orchestration
A local experiment YAML and plugin scaffold exist; live framework testing was not performed.
The intended role is reproducible execution with frozen target identity, configuration, cases, outputs, and result manifest. Current status remains SCAFFOLD_NOT_LIVE_TESTED.
INSPECT AIFrontier model and agent evals
Contract-only mapping for tasks, solvers, scorers, and evidence records.
+
Frontier model and agent evals
Contract-only mapping for tasks, solvers, scorers, and evidence records.
No Inspect suite result is included in the supplied execution record. A future adapter must preserve refusal, unknown, and provenance states rather than flattening them into a scalar score.
LM-EVALCapability benchmark runner
Contract-only path for conventional benchmark comparison and regression tracking.
+
Capability benchmark runner
Contract-only path for conventional benchmark comparison and regression tracking.
Capability scores are context for risk and release decisions; they do not measure PBHP conformance or humane governance by themselves.
HELMHolistic multi-metric evaluation
Contract-only mapping for scenario, metric, target, and transparent reporting.
+
Holistic multi-metric evaluation
Contract-only mapping for scenario, metric, target, and transparent reporting.
A multi-metric report still obeys the no-global-green rule. Independent hard floors remain visible beside aggregates.
Red teaming and safety
These systems would probe vulnerabilities, jailbreaks, hazardous behavior, control failure, and longer-horizon capability.
PYRITAutomated and human red teaming
Contract-only mapping for attacks, targets, converters, scorers, and evidence.
+
Automated and human red teaming
Contract-only mapping for attacks, targets, converters, scorers, and evidence.
Future execution should distinguish prompt-level vulnerability from deployment harm and preserve the exact attack configuration and human-review status.
GARAKVulnerability probing
Contract-only adapter for probe and detector results.
+
Vulnerability probing
Contract-only adapter for probe and detector results.
A detector pass is one piece of security evidence. It cannot certify the whole application, human workflow, or institution.
AILUMINATESafety and jailbreak benchmark
Integration requires licensed acquisition before execution.
+
Safety and jailbreak benchmark
Integration requires licensed acquisition before execution.
No content or result from a licensed suite is implied by an empty contract. Acquisition, version pinning, license compliance, and result reporting remain open.
AI VERIFY / MOONSHOTGovernance testing and red teaming
Contract-only interface for test results and governance evidence.
+
Governance testing and red teaming
Contract-only interface for test results and governance evidence.
Alignment of concepts does not imply participation, endorsement, or successful execution.
CONTROLARENAAI control experiments
Contract-only path for control protocols, adversarial agents, monitoring, and failure evidence.
+
AI control experiments
Contract-only path for control protocols, adversarial agents, monitoring, and failure evidence.
Control results would be domain-specific and cannot relax PBHP dignity, power, or receipt requirements.
METR-STYLELong-horizon task capability
Task acquisition is required before a real horizon evaluation can run.
+
Long-horizon task capability
Task acquisition is required before a real horizon evaluation can run.
The contract anticipates task identity, success criteria, duration, autonomy, tool use, monitor state, and evidence. No METR result or affiliation is claimed.
Standards views
The runtime can render control themes relevant to official frameworks; it cannot declare compliance.
EU AI ACTRisk, oversight, logging, robustness disclosure
Maps receipt threshold, human approval, chain identity, and worst SIL state into a version-sensitive view.
+
Risk, oversight, logging, robustness disclosure
Maps receipt threshold, human approval, chain identity, and worst SIL state into a version-sensitive view.
The view is a conversation aid for qualified compliance review. Current legal text, system classification, provider/deployer role, jurisdiction, and applicable dates must be verified independently.
NIST AI RMFGovern · Map · Measure · Manage
Connects challengeability, affected party, SIL, Maybe quality, and CAPA to RMF themes.
+
Govern · Map · Measure · Manage
Connects challengeability, affected party, SIL, Maybe quality, and CAPA to RMF themes.
The output says what controls may support a mapping. It does not say that the organization has completed the RMF or achieved a maturity level.
ISO/IEC 42001Management-system support
Receipts, documented roles, challengeability, and CAPA can supply management-system evidence.
+
Management-system support
Receipts, documented roles, challengeability, and CAPA can supply management-system evidence.
Only the applicable conformity process can establish certification. Shadow has no ISO certification and does not imply one.
OWASP AISVSAssurance-level routing
The runtime maps MIN to L1, CORE to L2, and ULTRA/ULTIMATE to L3 as a project-specific alignment aid.
+
Assurance-level routing
The runtime maps MIN to L1, CORE to L2, and ULTRA/ULTIMATE to L3 as a project-specific alignment aid.
The mapping is paraphrased and version-specific. It does not show that AISVS requirements were tested or met.
OMBFederal minimum-practice themes
ORANGE+ gate stack, reversibility, logging, and named accountability are surfaced for review.
+
Federal minimum-practice themes
ORANGE+ gate stack, reversibility, logging, and named accountability are surfaced for review.
The runtime itself warns that the relevant memorandum line is version-specific and must be checked against current official text before any claim.
A file path and schema prove that an interface was specified. Only execution can produce a result, and only qualified review can interpret it.