The distance between built and ready.
Project Shadow already contains a large reference runtime, codec, instrument panel, comparison register, evaluation scaffold, and behavioral research record. The remaining work is not cosmetic: reconcile versions, close evidence gaps, invite independent failure, and prove the human system can carry the controls.
This roadmap records work contemplated before the R1 lock. It is no longer a living program or an authorization to extend R1. Only the signature, external time anchor, exact public-package authorization, and later outside-beta evaluation remain as external gates or future evidence work.
See the exact locked R1 identity and current boundary →Open the system reference at this layer.
Follow the mechanism through numbered execution, a worked trace, artifact-bound evidence, adverse results, and validation still required.
Reconcile the product
A large repo can contain several locally coherent truths that do not yet form one releasable artifact.
VERSIONCanonical manifest
Bind runtime, codec, PBHP policy, SIL inventory, thresholds, schemas, tests, licenses, docs, and change history into one release identity.
+
Canonical manifest
Bind runtime, codec, PBHP policy, SIL inventory, thresholds, schemas, tests, licenses, docs, and change history into one release identity.
The frozen v1.3.1-UNIFIED runtime and later twenty-eight-gauge panel work need an explicit compatibility and migration record. A receipt should never depend on an unnamed hybrid.
PUBLICSterilize the adoption surface
Keep operational PBHP plain, portable, and independent of the internal mythic vocabulary.
+
Sterilize the adoption surface
Keep operational PBHP plain, portable, and independent of the internal mythic vocabulary.
Sterilization means audience fit and authority clarity—not erasing provenance or pretending the philosophy never existed. The public surface can link to research and implementation without requiring narrative assent.
CONTRIBUTORSRefresh provenance and ownership
Confirm sources, licenses, contributor roles, maintenance authority, contact paths, and the status of every retained framework entry.
+
Refresh provenance and ownership
Confirm sources, licenses, contributor roles, maintenance authority, contact paths, and the status of every retained framework entry.
Wrong tags, thin artifacts, unavailable systems, and ambiguous licenses remain visible until corrected or retired.
Finish the evidence program
The built scaffold is extensive; the external and human evidence layers remain substantially open.
REPRODUCEIndependent runtime and codec reproduction
A materially independent party rebuilds, runs, and compares the frozen artifacts from the public manifest.
+
Independent runtime and codec reproduction
A materially independent party rebuilds, runs, and compares the frozen artifacts from the public manifest.
Reproduction should include negative tests, platform variance, schema edge cases, receipt mutation, policy pinning, and action-mutation attempts—not only happy-path self-tests.
TARGETRun real target evaluations
Replace the echo stub with frozen target adapters and execute preregistered domains, variants, comparators, ablations, and metamorphic relations.
+
Run real target evaluations
Replace the echo stub with frozen target adapters and execute preregistered domains, variants, comparators, ablations, and metamorphic relations.
The current 25,000-case corpus carries synthetic development labels; 12,000 metamorphic relations were generated but target execution was pending in the supplied status. Confirmatory claims require independent human labels where specified.
HUMANClose the grading floor
Complete the Behavioral Falsifier's required human grading, beginning with the powered supported run.
+
Close the grading floor
Complete the Behavioral Falsifier's required human grading, beginning with the powered supported run.
Run 07 carries the highest-value machine-scored supportive result and a forty-eight-sheet human floor in the existing site record. Until grading closes, the verdict remains preliminary.
FIELDEthics-reviewed human system study
Study workload, reliance, alert fatigue, workarounds, dignity, challenge, decision quality, and downstream outcomes in a bounded non-hazardous setting.
+
Ethics-reviewed human system study
Study workload, reliance, alert fatigue, workarounds, dignity, challenge, decision quality, and downstream outcomes in a bounded non-hazardous setting.
A field study requires independent ethics review, informed participation, data minimization, stop rules, affected-party protections, and a repair plan before exposure.
Try to break governance
A safety system must be tested against power, incentives, and institutional self-protection—not only malformed inputs.
RED TEAMIndependent adversarial review
Attack gate shopping, action decomposition, override pressure, false provenance, correlated tribunal error, hostile retrieval, approval laundering, and receipt gaming.
+
Independent adversarial review
Attack gate shopping, action decomposition, override pressure, false provenance, correlated tribunal error, hostile retrieval, approval laundering, and receipt gaming.
The red team needs publication rights for disagreement and a route for severe findings to block release.
FPVCalibrate unnecessary refusal
Measure whether false positives impose harm, delay essential action, centralize authority, or train operators to bypass the system.
+
Calibrate unnecessary refusal
Measure whether false positives impose harm, delay essential action, centralize authority, or train operators to bypass the system.
A system that only measures under-caution will systematically overstate its benefit. Doors, successful challenges, and repaired over-escalation are part of the evidence.
CAPADemonstrate repair closure
Select real defects, run them through CAPA, verify the fix against regression cases, and publish what changed.
+
Demonstrate repair closure
Select real defects, run them through CAPA, verify the fix against regression cases, and publish what changed.
The ability to improve under correction is a stronger operational claim than simply stating corrigibility as a value.
Release in bounded layers
The project does not need to wait for every ambition before sharing honest artifacts, but each release must name what it is.
LAYER 01Reference and education
Documentation, protocol examples, schemas, frozen runtime, self-tests, and claim registry with no deployment verdict.
+
Reference and education
Documentation, protocol examples, schemas, frozen runtime, self-tests, and claim registry with no deployment verdict.
This layer lets reviewers inspect and reproduce the system while the evidence program remains open.
LAYER 02Research kit
Preregistrations, datasets, target harnesses, blinded label tools, scoring, adverse results, and external challenge protocol.
+
Research kit
Preregistrations, datasets, target harnesses, blinded label tools, scoring, adverse results, and external challenge protocol.
A research release invites failure without describing itself as production safety infrastructure.
LAYER 03Bounded pilot
A domain-specific integration with qualified human ownership, stop authority, appeal, monitoring, and ethics review.
+
Bounded pilot
A domain-specific integration with qualified human ownership, stop authority, appeal, monitoring, and ethics review.
The pilot's success cannot authorize other populations, stakes, tools, or institutions by analogy.
LAYER 04Maintained adoption release
Versioned artifacts, documented roles, calibration, incident response, CAPA, independent review, and a public evidence/limitation record.
+
Maintained adoption release
Versioned artifacts, documented roles, calibration, incident response, CAPA, independent review, and a public evidence/limitation record.
Maintenance and correction capacity are part of the product, not post-release paperwork.
The repo is massive because the problem is massive. Readiness comes from coherent versions, external challenge, human evidence, and demonstrated repair—not page count or test count.