PROJECT SHADOW 1.0.1 · CORRECTED R1 REFERENCE · PRELIVE · 2026-08-17

Project Shadow 1.0.1 contains no Myth package. Generic Myth v0.2.0 and Full-Canon Myth v0.3.5 are separate optional companions; both default off, neither is required by R1, and neither can authorize action or change an R1 result. No production or consequential deployment is authorized. No global green.

PROJECTSHADOW R1.0.1 corrected · PRELIVE
INTExternal evaluation contractscontracts and scaffolds · not executed

An adapter is not a result.

Grand TEVV includes interfaces for ten external evaluation ecosystems and standards-oriented views. The supplied package explicitly states that those suites were not run; each entry here describes intended role and current acquisition or implementation state.

10evaluation contracts
1Dioptra scaffold
0external suites represented as executed
LONG-FORM COMPANION

Open the system reference at this layer.

Follow the mechanism through numbered execution, a worked trace, artifact-bound evidence, adverse results, and validation still required.

Read the full chapter →
INT / 01

Model and agent evaluation

These contracts would connect Shadow's cases and evidence schema to broader evaluation runners.

DIOPTRA

Reproducible experiment orchestration

A local experiment YAML and plugin scaffold exist; live framework testing was not performed.

scaffold not live-tested

The intended role is reproducible execution with frozen target identity, configuration, cases, outputs, and result manifest. Current status remains SCAFFOLD_NOT_LIVE_TESTED.

INSPECT AI

Frontier model and agent evals

Contract-only mapping for tasks, solvers, scorers, and evidence records.

contract only

No Inspect suite result is included in the supplied execution record. A future adapter must preserve refusal, unknown, and provenance states rather than flattening them into a scalar score.

LM-EVAL

Capability benchmark runner

Contract-only path for conventional benchmark comparison and regression tracking.

contract only

Capability scores are context for risk and release decisions; they do not measure PBHP conformance or humane governance by themselves.

HELM

Holistic multi-metric evaluation

Contract-only mapping for scenario, metric, target, and transparent reporting.

contract only

A multi-metric report still obeys the no-global-green rule. Independent hard floors remain visible beside aggregates.

INT / 02

Red teaming and safety

These systems would probe vulnerabilities, jailbreaks, hazardous behavior, control failure, and longer-horizon capability.

PYRIT

Automated and human red teaming

Contract-only mapping for attacks, targets, converters, scorers, and evidence.

contract only

Future execution should distinguish prompt-level vulnerability from deployment harm and preserve the exact attack configuration and human-review status.

GARAK

Vulnerability probing

Contract-only adapter for probe and detector results.

contract only

A detector pass is one piece of security evidence. It cannot certify the whole application, human workflow, or institution.

AILUMINATE

Safety and jailbreak benchmark

Integration requires licensed acquisition before execution.

licensed acquisition required

No content or result from a licensed suite is implied by an empty contract. Acquisition, version pinning, license compliance, and result reporting remain open.

AI VERIFY / MOONSHOT

Governance testing and red teaming

Contract-only interface for test results and governance evidence.

contract only

Alignment of concepts does not imply participation, endorsement, or successful execution.

CONTROLARENA

AI control experiments

Contract-only path for control protocols, adversarial agents, monitoring, and failure evidence.

contract only

Control results would be domain-specific and cannot relax PBHP dignity, power, or receipt requirements.

METR-STYLE

Long-horizon task capability

Task acquisition is required before a real horizon evaluation can run.

task acquisition required

The contract anticipates task identity, success criteria, duration, autonomy, tool use, monitor state, and evidence. No METR result or affiliation is claimed.

INT / 03

Standards views

The runtime can render control themes relevant to official frameworks; it cannot declare compliance.

EU AI ACT

Risk, oversight, logging, robustness disclosure

Maps receipt threshold, human approval, chain identity, and worst SIL state into a version-sensitive view.

The view is a conversation aid for qualified compliance review. Current legal text, system classification, provider/deployer role, jurisdiction, and applicable dates must be verified independently.

NIST AI RMF

Govern · Map · Measure · Manage

Connects challengeability, affected party, SIL, Maybe quality, and CAPA to RMF themes.

The output says what controls may support a mapping. It does not say that the organization has completed the RMF or achieved a maturity level.

ISO/IEC 42001

Management-system support

Receipts, documented roles, challengeability, and CAPA can supply management-system evidence.

Only the applicable conformity process can establish certification. Shadow has no ISO certification and does not imply one.

OWASP AISVS

Assurance-level routing

The runtime maps MIN to L1, CORE to L2, and ULTRA/ULTIMATE to L3 as a project-specific alignment aid.

The mapping is paraphrased and version-specific. It does not show that AISVS requirements were tested or met.

OMB

Federal minimum-practice themes

ORANGE+ gate stack, reversibility, logging, and named accountability are surfaced for review.

The runtime itself warns that the relevant memorandum line is version-specific and must be checked against current official text before any claim.

THE INTEGRATION CLAIM
A file path and schema prove that an interface was specified. Only execution can produce a result, and only qualified review can interpret it.