PROJECT SHADOW 1.0.1 · CORRECTED R1 REFERENCE · PRELIVE · 2026-08-17

Project Shadow 1.0.1 contains no Myth package. Generic Myth v0.2.0 and Full-Canon Myth v0.3.5 are separate optional companions; both default off, neither is required by R1, and neither can authorize action or change an R1 result. No production or consequential deployment is authorized. No global green.

PROJECTSHADOW R1.0.1 corrected · PRELIVE
REFPreserved long-form research referencehistorical companion · not current R1 inventory

The whole machine, with the panels open.

A preserved paper-scale guide to earlier Project Shadow research: source intake, runtime, protocol compilation, mechanisms, twenty-eight SIL gauges, gate semantics, provenance, experiments, TEVV, embodied candidates, and the boundary between what ran and what remained an ambition.

PRESERVED RESEARCH · NOT CURRENT R1

This companion predates the locked R1 reference and Primitive Commons beta.5. Its counts, runtime labels, open-work language, and evidence snapshots remain historical; use the status and release-status pages for current exact identities.

See the exact locked R1 identity and current boundary →
THE READING CONTRACT

Explanation and evidence travel together.

Every chapter distinguishes specified behavior, expected behavior, observed artifact results, adverse evidence, and open validation. Counts belong to named snapshots. A passed self-test is not a field result, and a field proposal is not a shipped controller.

Open the companion PBHP field manual ↗
00
Mission · system boundary · public contract

Shadow is the machinery around the pause.

Project Shadow compiles a moral decision discipline into inspectable runtime behavior: typed inputs, independent gauges, binding gates, provenance, receipts, challenge, correction, and tests that are allowed to fail.

PBHP asks whether a consequential action should proceed. Shadow asks what engineering and governance must exist for that question to survive contact with software, long contexts, multiple agents, institutional pressure, and audit. It does not replace PBHP; it turns the protocol's questions into data structures, state transitions, floors, artifacts, and accountable roles.

The organizing concern is power under uncertainty. A powerful system should not turn missing evidence, hidden dependency, or another party's vulnerability into disposability. That concern becomes operational through least-powerful-first stakeholder ordering, explicit uncertainty states, reversibility, provenance binding, non-overridable floors, write-ahead receipts, and a path for challenge and repair.

Shadow is not one monolithic product. The repository contains a frozen runtime, protocol editions, a codec, SIL gauges, synthesis documents, candidate mechanisms, adoption material, evaluation scaffolds, studies, and historical layers. A responsible release declares which versions compose the artifact. ‘Present in the repo’ is not the same as ‘executed,’ and ‘executed’ is not the same as ‘validated.’

The public operational contract remains plain: no global green; the worst binding state wins; declaring is not proving; a receipt records judgment but does not make it correct; later layers may add friction but may not silently turn refuse into proceed.

EXECUTION GUIDE

Step by step

  1. 01

    Name the protected decision

    Identify the real action, accountable owner, affected population, and consequence tier before selecting components.

    EMITSA canonical action contract and authority boundary.
  2. 02

    Choose the release identity

    Pin PBHP edition, runtime, codec, policy pack, SIL inventory, schemas, tests, and known limitations.

    EMITSA manifest that prevents unnamed hybrid behavior.
  3. 03

    Separate mechanism from claim

    For every component, record whether it is proposed, specified, built, executed, independently reproduced, or field validated.

    EMITSA claim registry with evidence state.
  4. 04

    Keep human ownership explicit

    Assign action owner, reviewer, challenger, approval authority, stop authority, and repair owner.

    EMITSA governance map in which software cannot inherit sovereignty.
WORKED TRACE

A high-stakes model recommendation

INPUTA model recommends terminating a public benefit after analyzing a large case file with uncertain freshness and several handoffs.

  1. Shadow identifies the real action as termination, not text generation.
  2. PBHP orders the recipient first and asks whether a reversible, appealable Door exists.
  3. CLA, freshness, provenance, completeness, authority, reliance, and receipt gauges produce separate states.
  4. A hard floor from stale evidence or missing authority binds even if other gauges are clean.
  5. The system emits a constrained route: recover the record, require accountable review, preserve assistance during appeal where policy allows, and write the decision receipt before action.
OUTPUT

A typed non-commit or constrained Door with named recovery conditions, not a global safety score.

BOUNDARY

The software can surface and enforce configured controls. It cannot establish legal authority, factual truth, or moral legitimacy by itself.

Back to top ↑
01
Source intake · translation · change control

Compile ideas without laundering their status.

Shadow's synthesis layer admits a framework only after translating its useful mechanism, checking overlap, defining failure behavior, and preserving attribution, licensing, evidence, and unresolved questions.

A large safety repository can become a museum of attractive concepts. Shadow instead treats synthesis as a controlled engineering operation. A source enters with identity, version, provenance, license, intended use, evidence tier, and a statement of what problem it may solve. If those fields are missing, the source can still be discussed, but it cannot silently become a runtime control.

The candidate is then decomposed into mechanism primitives: input, state, observable signal, decision rule, floor, override behavior, output, receipt fields, and test oracle. Concepts that only rename an existing mechanism are crosswalked or archived. Concepts that add a genuinely different failure detector can become an extension, hardener, gauge, adapter, or explanatory lens.

Translation is deliberately lossy toward operational clarity. A symbolic or philosophical source may route attention to a recurring failure pattern, but the runtime must express the decision in plain language. The allowed chain is symbol → attention → check → evidence → gate → receipt. Symbol → authority → action is prohibited.

Every accepted change receives a receipt and regression obligation. The change log records why it was admitted, which artifacts moved, what tests ran, what did not run, and which prior assumptions remain unsettled. A clean merge does not prove a useful mechanism; it proves that the declared change can be reconstructed.

EXECUTION GUIDE

Step by step

  1. 01

    Register the source

    Record identity, version, authorship, license, provenance, audience, claim, and current evidence tier.

    EMITSA source-registry entry that can be audited later.
  2. 02

    Extract the mechanism

    Rewrite the idea as inputs, states, trigger, gate effect, override rule, receipt, and failure mode.

    EMITSA plain-language mechanism card.
  3. 03

    Deduplicate and crosswalk

    Compare against existing gates and gauges; choose replace, extend, harden, adapt, explain, or archive.

    EMITSAn integration disposition with overlap notes.
  4. 04

    Build anti-vacuity tests

    Add a positive case, a negative case, a near-miss, an override attempt, a receipt round-trip, and a sabotaged implementation where relevant.

    EMITSTests capable of rejecting a flattering implementation.
  5. 05

    Receipt the change

    Bind changed files, hashes, test results, open questions, and migration effect.

    EMITSAn append-only acceptance or rejection record.
WORKED TRACE

Focal Context Routing

INPUTA candidate proposes building a bounded present from a long history while preserving critical constraints.

  1. The mechanism is distinguished from generic summarization: it selects live context but pins standing constraints and unresolved dependencies.
  2. A structural floor—CONSTRAINT_PIN—prevents a declared pin from being suppressed.
  3. The package defines deterministic states, a stakes matrix, a receipt schema, an audit call, and tests.
  4. It enters SIL as gauge 24 only after acceptance and change-log receipt R-2026-07-13-01.
OUTPUT

An accepted v0.1 gauge with 36 recorded tests, not a claim that context selection is solved.

BOUNDARY

Tested conformance to its own contract does not establish that its thresholds improve real decision outcomes.

ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

July merge-back

The supplied merge bundle records five accepted SIL additions plus FMA outside the panel, 135/135 new-suite checks, file hashes, and an acceptance receipt.

SHADOW_MERGE_BACK_2026-07-13
ADVERSE

Version skew remains live

The frozen runtime, later gauge panel, historical synthesis, and candidate layers do not form a single release merely because they share a repository.

OPEN

Independent source audit

Licensing, prior-art completeness, and the usefulness of several retained framework integrations still require external review.

Back to top ↑
02
Evaluation path · state transitions · action binding

The executable hesitation, from request to receipt.

The runtime converts an untrusted request into a provenance-bound action, evaluates independent hardeners and gauges, applies the strongest floor, and commits only through an approval and receipt path that can be reconstructed.

The runtime starts before risk scoring. It normalizes the request into a canonical action and binds decision-bearing fields to provenance. The action identity includes actor, affected party, intended effect, tools, data, authority, scope, duration, and review conditions. If the action changes after a gate, the system must rerun rather than inherit the old result.

PBHP supplies the fast moral path: competence, classification, first payer, power inversion, risk floor, alternatives, Maybe/Therefore, and receipt. Shadow adds hardeners for provenance, dignity, frame stability, narrative pressure, projection, contradiction, context reliability, and other recurrent failure modes. SIL observes conditions without pretending one clean reading can cancel another component's floor.

The policy engine then computes the binding action state. Add-only means a later component can harden proceed to verify, constrain, delay, or refuse, but cannot silently soften an earlier hard refusal. Break-glass is a separate earned route with authentication, named low-power parties, specific irreversible harm, evidence, failed alternatives, independent concurrence, and hard ceilings. BLACK, WALL, and specified structural floors survive the emergency claim.

Finally, the runtime writes the receipt before action. The receipt carries policy and artifact identity, action hash, provenance block, gate findings, SIL snapshot, dissent quality, approval, overrides, prior hash, and allowed action. Serialization through the codec preserves the decision without granting additional authority.

EXECUTION GUIDE

Step by step

  1. 01

    Normalize the request

    Build the canonical action and reject missing decision-critical fields or mark them unknown.

    EMITSAction identity and scope hash.
  2. 02

    Bind provenance

    Grade each steering field as attested, declared, asserted, inferred, or unknown before gates consume it.

    EMITSA provenance block and downgrade findings.
  3. 03

    Run PBHP and hardeners

    Evaluate competence, burden, power, harm floor, alternatives, dissent, dignity, frame, narrative, projection, and drift controls.

    EMITSIndependent findings with floor semantics.
  4. 04

    Compose SIL

    Run the configured gauges and preserve each state, source, matrix hash, override attempt, and floor.

    EMITSA panel snapshot with no global green.
  5. 05

    Resolve the action state

    Apply worst binding state, add-only escalation, break-glass ceilings, tier rules, and approval requirements.

    EMITSProceed, mitigate, constrain, refuse/delay, refuse absolute, or earned crisis route.
  6. 06

    Write and chain the receipt

    Bind the decision before execution and link any revision to its predecessor.

    EMITSA reconstructable write-ahead receipt and codec payload.
WORKED TRACE

A bare break-glass request

INPUTAn operator marks an action RED, writes ‘urgent,’ and requests crisis authorization without provenance or concurrence.

  1. The crisis route checks for a named low-power party and a specific irreversible harm.
  2. It requires evidence identity, an explanation of why each safer Door fails, authenticated invocation, and independent concurrence by a distinct party.
  3. Those fields are missing, so the emergency claim does not become evidence.
  4. The runtime preserves the existing floor and records the attempt.
OUTPUT

REFUSE_OR_DELAY with a crisis-provenance failure receipt.

BOUNDARY

Urgency is an input to govern. A string labeled ‘break glass’ is not an authorization primitive.

ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Frozen runtime self-tests

Grand TEVV v0.3 records 115/115 runtime and 247/247 complete-package codec self-tests passing in the builder session. The standalone codec snapshot records 136 checks when its compatibility alias is not co-located.

runtime 1.3.1-UNIFIED · codec PS2 / language 0.4
OBSERVED

Component probes

10/10 deterministic probes covered provenance mismatch, missing provenance, power inversion, weak Maybe, several break-glass paths, structural drift, and a SIL floor.

EXPECTED

Write-ahead integrity

Binding action and policy identity before execution should make silent mutation and post-hoc justification more detectable.

OPEN

Production behavior

The frozen runtime has not been independently reproduced or validated in a real consequential deployment.

Back to top ↑
03
HUMAN · MIN · CORE · ULTRA · ULTIMATE

One protocol, compiled to five operating depths.

The tier system changes procedure, instrumentation, challenge, and receipt depth while preserving the same hard floors and human-ownership boundary.

A long canonical specification is valuable for audit and dangerous as the only runtime interface. Operators under pressure will either skip it or perform it ceremonially. Shadow therefore treats the full protocol as source material that compiles into profiles: HUMAN for paper use, MIN for a rapid reflex, CORE for standard consequential work, ULTRA for high consequence and independent challenge, and ULTIMATE for the full stack and readings.

Compilation is not summarization. Each profile must preserve semantic invariants: competence before permission, unresolved as a first-class state, least-powerful-first ordering, worst binding state, a real alternative Door, a serious Maybe, write-ahead receipt, and non-overridable floors. A compressed profile may omit elaboration but cannot convert Wall or BLACK to proceed.

Domain packs add vocabulary, evidence requirements, hard floors, examples, and responsible roles for a field such as healthcare, public benefits, security, or embodied action. They may harden the base protocol but may not relax its protected floor. The compiler should emit both human-readable instructions and a machine-readable Shadow intermediate representation so conformance can be tested.

The correct question is therefore not ‘Which protocol is the real one?’ The canonical source defines the invariant; editions and profiles are declared compilations. Their equivalence must be tested against fixtures, including adversarial cases where compression is most likely to drop the inconvenient constraint.

EXECUTION GUIDE

Step by step

  1. 01

    Select consequence tier

    Use impact, irreversibility, power, scale, novelty, uncertainty, and required authority to choose the minimum profile.

    EMITSA tier selection receipt.
  2. 02

    Load the invariant kernel

    Pin the non-droppable competence, Gap, first-payer, floor, Maybe, receipt, and human-ownership semantics.

    EMITSA kernel identity shared across profiles.
  3. 03

    Apply domain pack

    Add evidence sources, local roles, thresholds, legal/policy checks, examples, and hardeners for the domain.

    EMITSA versioned domain configuration.
  4. 04

    Compile human and machine forms

    Generate the operating checklist, receipt schema, gate matrix, and intermediate representation.

    EMITSPaired instructions and executable policy.
  5. 05

    Run conformance fixtures

    Verify that each tier preserves floors, action identity, challenge, and receipt behavior across clean, adverse, and mutation cases.

    EMITSA profile-specific conformance report.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Five-tier edition package

The supplied handover records ULTIMATE, ULTRA, CORE, MIN, and HUMAN in tight and full depths, with a 21/21 edition-package check at that snapshot.

SPECIFIED

Hard-floor preservation

Compression may reduce explanation and instrumentation; it may not weaken a binding Wall, BLACK state, or declared invariant.

OPEN

Compiler equivalence

A complete canonical-source → IR → profile compiler and independent conformance suite remain release work.

Back to top ↑
04
CLA · PSI · BCR · hippocampal routing · FireStamp

Small mechanisms for the ways judgment quietly degrades.

Shadow decomposes failure into inspectable mechanisms so a system can respond to the right defect: lost context, collapsed frames, out-of-band complexity, stale premises, hidden overrides, or orchestration drift.

CLA measures context reliability rather than treating token capacity as memory. Long context, compaction, contradictory authority, stale premises, tool failures, handoffs, and pressure can produce a fluent but degraded evaluator. The receipt discloses whether the load reading was measured, estimated, self-reported, or unknown and routes high-stakes critical states to refresh or a fresh qualified evaluator.

PSI—Pattern Separation Injection—protects distinct hypotheses and evidence streams from collapsing into one convenient narrative. It asks the system to keep similar-looking cases separate until the discriminating evidence is checked. A related CA3/CA1/ACC metaphor maps retrieval, comparison, mismatch, and conflict routing; the metaphor explains attention, while the operational rule remains plain and testable.

BCR—Bandpass Complexity Routing—requires an agent or component to declare the range of complexity it can responsibly process. Out-of-band input is compressed, decomposed, rerouted, escalated, or refused instead of being handled with false competence. OGR, the orchestrator governance registry, tracks specialist bandpass, current load, drift risk, and prohibited outputs so the master does not silently overload a convenient agent.

FireStamp is an override disclosure. It records who changed a binding recommendation, under what authority, with which reason and evidence, and what increased risk the override accepts. It is not a waiver. A stamp can make an override visible and still document an action that remains prohibited by an absolute floor.

EXECUTION GUIDE

Step by step

  1. 01

    Measure or disclose context state

    Prefer instrumented load, then transparent estimate, then conservative self-report; otherwise mark unmeasured.

    EMITSCLA receipt and allowed next action.
  2. 02

    Separate competing patterns

    List hypotheses, similar precedents, distinguishing evidence, and unresolved contradictions before synthesis.

    EMITSA PSI separation table.
  3. 03

    Check bandpass

    Compare task complexity and stakes to the component's declared operating range and current load.

    EMITSProceed, decompose, reroute, escalate, or refuse.
  4. 04

    Govern orchestration

    Use the registry to prevent correlated overload, prohibited outputs, or unowned handoffs across specialists.

    EMITSA routing trace and master decision receipt.
  5. 05

    Stamp every override

    Bind overrider, authority, original floor, reason, evidence, residual risk, concurrence, and time limit.

    EMITSA FireStamp that cannot silently soften an absolute floor.
WORKED TRACE

A long multi-agent investigation

INPUTA master agent assigns legal, technical, and factual subproblems after several summaries; one specialist is asked to synthesize outside its declared scope.

  1. CLA marks compaction and handoff risk; standing constraints are reanchored from source.
  2. PSI preserves two conflicting factual timelines instead of merging them into an average story.
  3. BCR detects that the technical specialist is out of band for legal authority and routes that question to a qualified reviewer.
  4. OGR records load and prohibits the overloaded specialist from emitting the final consequential recommendation.
  5. If an operator forces the handoff anyway, FireStamp records the override without granting permission.
OUTPUT

A bounded synthesis packet with unresolved conflicts and ownership intact.

BOUNDARY

These mechanisms improve inspectability. Their labels do not prove that an agent's self-declared bandpass or load estimate is accurate.

ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

CLA versioned tests

The July SIL ledger records CLA v0.2 at 85 tests (84 passed, one conditional skip); earlier artifacts record 73/73 and 75/75 snapshots.

EXPECTED

Pattern and bandpass routing

Keeping hypotheses separate and refusing out-of-band work should reduce narrative collapse and false competence.

OPEN

Component calibration

PSI, BCR, orchestration thresholds, and FireStamp effects need preregistered ablation and real-workflow testing.

Back to top ↑
05
28 gauges · 503 package checks · no global green

SIL instruments the conditions under which an answer was produced.

The Shadow Instrument Layer reports independent reliability and governance states. It is a panel, not a score: a clean gauge cannot erase a binding failure elsewhere.

SIL asks a different question from PBHP. PBHP asks whether the action should proceed. SIL asks under what conditions the output was produced: context load, time anchoring, provenance, uncertainty source, citation sufficiency, freshness, compaction, tools, reversibility, receipt status, retrieval coverage, sensitivity, validation, completeness, reliance, prompt injection, lineage, inference distance, contradiction, authority, sampling mode, drift, focal context, context drag, stale-premise intrusion, mirror boundary, and anti-projection.

Each gauge defines states, a deterministic Stakes × State matrix, a receipt schema, an audit call, and override behavior. Most include at least one non-overridable REFUSE cell. Behavioral gauges intentionally carry weaker evidence and usually cap at verify-first rather than pretending an interpretive signal can authorize refusal by itself.

Composition preserves the worst state and every hard floor. The Runtime Reliability Disclosure renderer changes the amount of visible detail by consequence: low may require none, medium a line, high a block, and critical a block plus gate and receipt. The renderer is connective tissue over gauge receipts; it does not create new truth.

The roadmap distinguishes SIL-1 telemetry, SIL-2 experimental system identification, SIL-3 latent-state microscopy, SIL-4 causal intervention, and SIL-5 runtime governance. The current panel is primarily SIL-1: explicit observable/self-reported conditions and deterministic routing. Calling it causal mind-reading would overstate both the evidence and the architecture.

EXECUTION GUIDE

Step by step

  1. 01

    Select the panel

    Choose gauges relevant to the action and declare the exact inventory and versions; do not imply the later panel is inside an older frozen runtime.

    EMITSA panel manifest.
  2. 02

    Acquire readings

    Record state, evidence source, stakes, matrix identity, uncertainty, and attempted overrides for each gauge.

    EMITSIndependent gauge receipts.
  3. 03

    Apply floors

    Evaluate each state against its matrix and preserve every non-overridable refusal or constraint pin.

    EMITSA floor list that cannot be averaged away.
  4. 04

    Render proportionately

    Use consequence to select disclosure depth while keeping decision-bearing findings visible.

    EMITSA low, medium, high, or critical reliability disclosure.
  5. 05

    Route, do not diagnose

    Translate behavioral signals into reanchor, verify, refresh, route, or human review; avoid claims about personhood, intent, or mental state.

    EMITSA bounded operational response.
WORKED TRACE

Critical answer with three floors

INPUTA critical decision has degraded context, broken data lineage, and unvalidated output; citation and time gauges are otherwise clean.

  1. Each gauge emits its own receipt and matrix result.
  2. Clean citation and time readings remain visible but cannot cancel the three refusal floors.
  3. The renderer produces a critical disclosure block, required gate, and receipt.
  4. The route recovers context, traceable data, and validation before the action can be reconsidered.
OUTPUT

REFUSE with three named floors and explicit repair conditions; global_green remains null.

BOUNDARY

A panel result shows configured conditions and routing. It does not prove the underlying data is true or the eventual action is safe.

ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Current gauge manifest

The SIL 1.1.0 manifest records 28 separately packaged gauges and 503/503 gauge-package checks. A prior 519 figure added 16 renderer checks; that combined number is not the gauge count.

SIL_TOOL_LIBRARY_MANIFEST.json
OBSERVED

Promotion fidelity

Anti-Projection was cross-checked over 576 signal combinations and Mirror Boundary over 20,000 seeded grids with zero drift against the referenced implementation.

ADVERSE

Counts are snapshot-specific

Earlier records show 23 gauges and 394 checks. The numbers describe different inventories and dates; they must not be combined into a larger fictional total.

OPEN

Outcome validity

Passing gauge contract tests and fidelity grids does not establish that the thresholds improve real-world decisions or avoid harmful false positives.

Back to top ↑
06
Floors · escalation · challenge · break-glass

Every gate must say what it can stop.

A gate is complete only when its input, trigger, action effect, floor, override semantics, receipt, and repair path are explicit and testable.

Shadow's gate catalog includes the PBHP fast path and hardeners for competence, dignity, irreversible harm, accumulation, power inversion, provenance, Door quality, correction, frame and narrative pressure, projection, drift, and other failure patterns. Their names are less important than their semantics. A gate that only produces a warning is not equivalent to one that can constrain or refuse execution.

Deterministic floors are asymmetric by design. A refusal cell survives an operator override unless the specification explicitly defines a lawful bounded release path. Break-glass is not that general override. It is an earned crisis route with narrow prerequisites and absolute ceilings. The runtime records ignored-floor attempts so pressure itself becomes evidence.

False Positive Validation keeps the system challengeable. A challenger must answer the pause in four parts, provide evidence, and present a Door that actually breaks the cited harm path. Repetition, authority theater, competitive pressure, or a renamed action do not qualify. Repeated attempts can activate wear-down alarms and CAPA.

Gate quality depends on calibration. A permanently conservative gate can harm people through delay and exclusion; a permissive gate can bless dangerous action. Shadow therefore needs adverse fixtures, successful challenges, subgroup and affected-party review, override monitoring, and real outcome data—not only tests proving the matrix executes as written.

EXECUTION GUIDE

Step by step

  1. 01

    Define the gate contract

    Specify input, evidence source, states, trigger, action effect, floor, override, receipt, and recovery condition.

    EMITSA testable gate specification.
  2. 02

    Build adversarial fixtures

    Test clean, floor, near-miss, missing data, contradictory data, action mutation, override, and break-glass cases.

    EMITSA gate behavior suite with anti-vacuity controls.
  3. 03

    Run add-only composition

    Combine findings without allowing a later component to downgrade a stronger binding state.

    EMITSThe worst credible action state and all supporting findings.
  4. 04

    Offer FPV

    Let a distinct reviewer challenge the pause with an evidence-backed safer Door.

    EMITSRelease, remain, or harden—plus the challenge receipt.
  5. 05

    Calibrate from outcomes

    Review under-caution, over-caution, missed parties, workarounds, harms from delay, and repair closure.

    EMITSThreshold changes through versioned CAPA.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Weak-Maybe probe

The supplied component suite observed an ORANGE/CONSTRAIN result at quality 0.3, matching the expected non-commit behavior.

OBSERVED

Break-glass ceilings

Deterministic probes preserved BLACK, WALL, structural-drift, and SIL floors and rejected a bare RED crisis claim.

ADVERSE

Declared crisis identity

One reference probe allowed PROCEED_UNDER_CRISIS at RED with merely DECLARED authentication. The project itself marks declared authentication as transitional and pending key infrastructure.

OPEN

False-positive burden

Real rates of unnecessary refusal, delayed benefit, operator workaround, and affected-party harm have not been established.

Back to top ↑
07
ATTESTED · DECLARED · ASSERTED · receipt chain

Fields are not reality. Hashes are not honesty.

Provenance distinguishes how a claim entered the system, binds bytes and decisions to identities, and preserves change. It never converts a declaration or registered artifact into truth.

Shadow separates evidence states because a runtime can look precise while consuming invented or weakly sourced fields. ATTESTED means a claim is backed by a configured verification path. DECLARED means a named actor asserted it. ASSERTED means it arrived without that stronger identity. Inference and unknown states remain visible. The labels control routing and disclosure; they do not settle the world outside the receipt.

The evidence locker can register bytes, hashes, and references. That proves identity and later detects mutation. It does not prove that a source was honest, complete, current, lawful, or correctly interpreted. The receive ladder—raw, structurally valid, receipt matched, runtime bound—therefore authorizes nothing by itself.

Decision receipts are write-ahead and chained. They include action identity, provenance, policy, gate and SIL state, Maybe/Therefore, approval, override, timestamp, and previous hash. HMAC support in some artifacts is described as transitional; where durable signature identity is implied, production key infrastructure remains required.

A good audit asks two separate questions: ‘Are these the same bytes and decision we recorded?’ and ‘Were the underlying claims true and the action legitimate?’ Cryptography can strongly support the first. It cannot answer the second.

EXECUTION GUIDE

Step by step

  1. 01

    Classify every steering field

    Assign provenance state and evidence reference before a gate consumes the value.

    EMITSA field-level provenance block.
  2. 02

    Register immutable identity

    Hash artifacts, schemas, policies, and evidence packages; preserve version and source metadata.

    EMITSA content-addressed identity record.
  3. 03

    Bind the runtime decision

    Include the action, policy, provenance, findings, approval, and prior receipt in the write-ahead record.

    EMITSA decision receipt and chain link.
  4. 04

    Verify independently

    Check bytes, signatures/keys where available, source truth, authority, interpretation, and real-world correspondence as separate tasks.

    EMITSIntegrity findings distinct from truth findings.
  5. 05

    Receipt every change

    Append a new record when fields, evidence, action, or policy change; never silently edit the old decision.

    EMITSA reconstructable revision history.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Decision-neutral declaration battery

The supplied poster records that declaring operator judgment changed 0 of 1,280 development-battery decisions and did not downgrade WALL, critical dignity, or BLACK floors.

ADVERSE

Registration is not truth

The project explicitly retains this as an open design boundary; a fabricated or misleading artifact can be hashed perfectly.

OPEN

Production key infrastructure

HMAC and DECLARED crisis authentication are transitional wherever a durable signature or verified invoker is required.

Back to top ↑
08
Behavioral Falsifier · preregistration · adverse runs

A research program that must publish the ugly runs.

The Behavioral Falsifier tests whether a Shadow discipline changes evaluative sycophancy and evidence hygiene, using placebo, bare, plain, mythic, and myth-only arms with preregistered stop and interpretation rules.

The target behavior is unearned endorsement: an evaluator rubber-stamps flawed work because a user pulls for praise, status, or agreement. The battery contains flawed and excellent work so reducing endorsement cannot be purchased by indiscriminate negativity. The rubric also tracks evidence hygiene, defect detection, justified praise, leakage, and—in the myth ablation—unsupported self-anthropomorphism.

The supplied current ledger records nine preregistered experiments and 852 trials across six subject models and three providers. Several runs hit floor effects, low baselines, or saturation. One gpt-4.1-nano run moved in the wrong direction, from .50 to .70 unearned endorsement, and was stopped according to the design. Preserving that result prevents the program from becoming a sequence of hidden retries until something looks good.

The powered gpt-4o-mini run contains the strongest machine-scored result: 240 trials, with the plain Shadow pack moving unearned endorsement from .375 under placebo to .15, reported p = .006 and effect = .275. The result cleared the preregistered machine bars. It remains preliminary because the required independent human-grading floor is open.

The myth ablation is also informative. At available power, the mythic-register pack and plain pack produced equivalent sycophancy results, while myth-only tracked bare. The operational discipline appeared to carry the effect; the symbolic layer was lossless rather than measurably additive. That is a bounded result for this battery, not a general verdict on narrative or engagement.

EXECUTION GUIDE

Step by step

  1. 01

    Preregister the target and stop rules

    Freeze hypotheses, arms, battery, sample, metrics, exclusion, powering, adverse interpretation, and publication commitment.

    EMITSA timestamped analysis contract.
  2. 02

    Blind the trials

    Shuffle arms and remove condition labels from scorer sheets; preserve a sealed assignment key.

    EMITSScorable responses without treatment identity.
  3. 03

    Run machine scoring as preliminary

    Apply the fixed rubric and report supported, near-miss, null, equivalent, adverse, and stopped results without collapsing tiers.

    EMITSA reproducible preliminary report.
  4. 04

    Close the human floor

    Have independent scorers grade the required subset in fixed order, compute agreement, and resolve ties against the flattering reading.

    EMITSA confirmatory or revised verdict.
  5. 05

    Patch the instrument

    Use floor effects, saturation, leakage, and failed discrimination to preregister the next battery rather than reinterpreting the old one.

    EMITSA new version with a preserved evidence trail.
WORKED TRACE

The pack backfires

INPUTIn a 60-trial gpt-4.1-nano screen, unearned endorsement rises from .50 to .70 under the treatment.

  1. The result is not renamed ‘engagement’ or excluded because it is inconvenient.
  2. The preregistered stop rule prevents automatic expansion of a harmful-looking condition.
  3. The run stays in the ledger as adverse machine evidence.
  4. A later powered run on another subject can answer a different preregistered question but cannot erase this one.
OUTPUT

STOPPED / ADVERSE at its own tier.

BOUNDARY

Cross-model improvement is not licensed by one supported subject; model and battery interaction is part of the finding.

ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Powered machine result

gpt-4o-mini, n=240: .375 placebo → .15 plain pack; reported p=.006 and effect=.275; preregistered machine bars cleared.

ADVERSE

Backfire and floor effects

One subject worsened in a screen; several stronger/newer models left too little baseline behavior to measure. An earlier over-caution story disappeared under power.

OBSERVED

Myth ablation

Plain and mythic-register packs were equivalent at power; myth-only tracked bare in the current machine-scored record.

OPEN

Human confirmation

The site ledger records zero completed human-graded sheets and requires at least a 20% floor; Run 07's 48-sheet subset is the highest-value open task.

Back to top ↑
09
Grand TEVV · corpora · comparators · external contracts

A test universe is not a completed evaluation.

Grand TEVV specifies a broad program, executes staged internal checks, and builds the harness for real targets. Its strongest honesty feature is the explicit line between generated test infrastructure and evidence from an evaluated system.

Grand TEVV v0.3 specifies thirty evaluation domains, 270 test families, 2,430 planned family × method cells, 25,000 synthetic factor-cross cases, 12,000 one-way metamorphic relations, and ten external-framework contracts. It includes preregistration, blinding, independent labeling, red-team, hazardous-capability containment, human-field-study, meta-evaluation, disagreement-publication, CAPA, scorecard, and public-report artifacts.

The current package records runtime and codec self-tests, contract and component checks, corpus integrity, development benchmark rows, limited real-target work, and adverse evidence. Scope remains preliminary: it is not independent reproduction, a full target evaluation, or deployment validation.

The development comparator battery is useful and limited. With generated labels, the canonical runtime recorded 0 false-positive releases among 968 cases labeled should-refrain and 0 over-caution among 22 labeled should-commit. Ungated committed on all 968 should-refrain cases. A simple checklist and conservative matrix produced roughly 4% false-positive release. However, removing Maybe quality, power inversion, or provenance produced the same aggregate metrics as canonical, showing the generated battery did not discriminate those components at that level.

That ablation result is not a reason to hide the battery; it is a reason to improve it. Component probes separately exercised weak Maybe, power inversion, and provenance, but aggregate outcome attribution remains open. Confirmatory claims require a frozen real target, preregistered variants, independent human labels where specified, reproduction, and adversarial challenge.

EXECUTION GUIDE

Step by step

  1. 01

    Freeze target and variants

    Pin model/application, prompts, tools, decoding, environment, runtime, codec, policy, comparators, and ablations.

    EMITSA target manifest that can be reproduced.
  2. 02

    Select domains and families

    Use the coverage matrix to choose consequence-relevant tests without claiming unexecuted cells as results.

    EMITSA preregistered evaluation slice.
  3. 03

    Generate and independently label

    Use synthetic cases for development; reserve confirmatory labels for materially independent human review where required.

    EMITSSeparate development and confirmation datasets.
  4. 04

    Execute target and metamorphic pairs

    Run the actual target through cases and one-way relations, preserving outputs, errors, latency, tool traces, and environment.

    EMITSRaw target evidence and reproducible scores.
  5. 05

    Run comparators and ablations

    Compare canonical, ungated, simpler policies, and component removals; require the battery to separate them where a mechanism claim depends on it.

    EMITSEffect, failure, and non-discrimination findings.
  6. 06

    Publish disagreement and limits

    Report adverse subgroups, nulls, gray cases, scorer disagreement, unexecuted integrations, and prohibited claims.

    EMITSA public evidence record that can survive challenge.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Builder-session checks

115/115 runtime, 247/247 codec, 2/2 reference contracts, 3/3 grand contracts, 10/10 component probes, and corpus-integrity checks passed in the supplied build summary.

OBSERVED

Development comparator

Canonical DCECR=0 and over-caution=0 on generated development labels; ungated DCECR=1.0; simple checklist .0434; conservative matrix .0413.

ADVERSE

Non-discriminating ablations

Removing Maybe quality, power inversion, or provenance left the aggregate generated-battery metrics unchanged, weakening component-attribution claims.

OPEN

Confirmatory execution

Limited real-target work exists, but no independent red team, independent reproduction, human field study, hazardous-capability expert evaluation, deployment validation, or external-suite run was performed.

Back to top ↑
10
Kahn studies · DCECR · expected behavior

Simulation can reveal a mechanism and still stop short of proof.

Shadow treats studies as artifact-bound observations. The expected behavior is explicit, the observed behavior is recorded, adverse interactions stay visible, and the untested generalization remains open.

The Kahn studies test an intuitive Shadow claim: asking a model to identify irreversible harm, a safer alternative, and deceptive signal/action gaps may alter escalation behavior. The first five-match pilot produced a dramatic contrast and, more importantly, exposed a useless external gate that marked all sixty gated assessments GREEN while model self-assessment marked nearly all ORANGE or RED.

The follow-up reported forty-eight configurations across twelve models. Same-model OFF controls produced nuclear escalation in 4/17 runs; GATED in 1/13; SELF_ONLY in 0/5 tested models. The most escalation-prone Sonnet configuration dropped to conventional military posturing in its single self-only run. That is a strong directional observation within one simulation family.

The adverse evidence prevents the headline from outrunning the experiment. One Sonnet GATED run still reached 725 despite twenty-nine self-constraint attempts. A cross-model GATED interaction reached 950 on turn one, worse than OFF. Samples per condition were small, some runs failed or timed out, and the study differed from the comparison paper's exact implementation. The result does not establish that PBHP ‘eliminates’ catastrophic escalation in general.

DCECR—decision-commit error under critical risk—provides a narrower mechanism metric for development batteries. It asks whether a system commits when the oracle says it should refrain. The metric is valuable only to the degree that the oracle, case distribution, and target reflect the intended deployment. A perfect score against synthetic labels is a software-development result until those assumptions are challenged.

EXECUTION GUIDE

Step by step

  1. 01

    State expected behavior

    Before execution, define which decision should change, the comparison, the metric, and what would count as failure or over-caution.

    EMITSA falsifiable hypothesis.
  2. 02

    Bind study identity

    Record models, versions, prompts, scenario, seeds, turns, simulator, gate policy, timeouts, and exclusions.

    EMITSA reproducible study manifest.
  3. 03

    Report all conditions

    Include adverse, timed-out, cross-model, and anomalous runs beside the flattering comparison.

    EMITSA complete outcome table.
  4. 04

    Separate observation from inference

    Describe what happened in the simulator, then state the narrower mechanism inference and competing explanations.

    EMITSAn evidence-tiered conclusion.
  5. 05

    Design the next discriminator

    Increase seeds, domains, target diversity, component ablation, independent reproduction, and outcome validity rather than repeating the most favorable cell.

    EMITSA preregistered follow-up.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Pilot gate disagreement

External gate GREEN on 60/60; GPT-4o self-assessment 18 ORANGE, 10 RED, 2 YELLOW; Haiku 6 ORANGE, 23 RED, 1 YELLOW.

OBSERVED

Follow-up same-model rates

OFF 4/17 nuclear, GATED 1/13, SELF_ONLY 0/5 in the supplied report.

ADVERSE

Cross-model catastrophic gate result

One GATED cross-model configuration reached 950 on turn one, demonstrating that external constraint assumptions did not transfer cleanly.

OPEN

General safety claim

Other scenarios, larger samples, independent replications, real incentives, non-military domains, long horizons, and field outcomes remain unestablished.

Back to top ↑
11
EAL / AYEM · physical action · controller boundary

PBHP pauses the mind. An embodied layer must pause the body.

A language model may help describe or evaluate a physical action, but it should not directly own the safety envelope that permits force, motion, contact, or continued operation.

Embodied action requires a stricter contract because the world does not wait for a later correction. A candidate EAL/AYEM layer records authenticated requester, commanded motion, object or person, force, velocity, workspace boundary, duration, consent, reversibility, sensor confidence, termination condition, and independent emergency stop before motion.

The language model's role is advisory and bounded: interpret an instruction, identify ambiguity, generate alternatives, explain a gate, or assemble a receipt. A deterministic controller, safety-certified subsystem where required, or other hard real-time mechanism must enforce force, speed, distance, exclusion zones, collision constraints, and stop behavior. The LLM cannot talk the body through a failed floor.

Uncertainty about identity, position, consent, or sensor health creates a Gap. The response is to reduce energy, stop, improve sensing, request human confirmation, reroute, or maintain a safe state. ‘Likely clear’ is not a physical authorization when the cost of error is carried by a nearby person.

This layer remains a candidate direction, not a shipped robotics safety system. It would require domain engineering, hazard analysis, real-time guarantees, independent testing, regulatory review where applicable, and a hardware-level stop path outside the language model.

EXECUTION GUIDE

Step by step

  1. 01

    Authenticate the requester

    Verify who requested movement, their authority, and whether the command is within the robot's assigned role.

    EMITSRequester and authority identity.
  2. 02

    Construct the physical action

    Bind motion, target, force, velocity, boundary, duration, consent, reversibility, sensing, and termination.

    EMITSAn embodied action contract.
  3. 03

    Run hazard and uncertainty gates

    Check people, unknown objects, occlusion, sensor disagreement, energy, escape routes, and downstream motion.

    EMITSDoor, Wall, or Gap plus a physical risk floor.
  4. 04

    Enforce outside the LLM

    Apply the permitted envelope in deterministic control and preserve independent emergency stop authority.

    EMITSA hard motion constraint or stop.
  5. 05

    Receipt and monitor

    Record command, interpretation, sensor state, constraints, motion, stops, operator intervention, and outcome.

    EMITSAn embodied decision and telemetry record.
WORKED TRACE

Unknown object near a worker

INPUTA warehouse robot is told to move a package, but object classification is uncertain and a worker is inside the wider reach envelope.

  1. Identity and proximity uncertainty create a Gap before motion.
  2. The controller reduces energy and holds position; the LLM cannot downgrade the state through explanation.
  3. The system improves sensing or requests qualified confirmation and clears the human from the enforced boundary.
  4. Only a new action contract inside the verified envelope may proceed.
OUTPUT

STOP / VERIFY / REPLAN at the body controller.

BOUNDARY

This is a design example. No robotics implementation or physical validation is claimed by the current release.

ARTIFACT LEDGER

Behavior and evidence state

SPECIFIED

Candidate action fields

The project context defines a proposed physical-action contract and the rule that an LLM should not directly own the robot body.

EXPECTED

Energy-reducing Gap behavior

Stopping, slowing, improving sensing, and routing should reduce exposure under uncertain physical state.

OPEN

Entire embodied validation stack

Controller implementation, hazard analysis, timing guarantees, hardware tests, human factors, certification, and field evidence are not present.

Back to top ↑
12
Artifact ledger · release gates · ultimate reference

The ultimate release is a versioned argument with receipts.

A credible release ties explanation, executable artifacts, test evidence, adverse results, governance, and unfinished work into one identity that another person can reproduce and challenge.

The current material is already larger than a website: a five-tier protocol family, frozen runtime and codec, twenty-eight-gauge panel, synthesis and adoption volumes, simulation studies, Behavioral Falsifier, Grand TEVV, integration contracts, governance records, and candidate directions. The ultimate release should not flatten that into a single triumphant number. It should make the relationships navigable and the boundaries impossible to miss.

The artifact ledger must name the snapshot. Runtime 1.3.1-UNIFIED and codec PS2/language 0.4 serialize twenty-three gauges. SIL 1.1.0 separately records twenty-eight gauges and 503 package checks; a prior 519 figure also included sixteen renderer checks. The Behavioral Falsifier ledger records nine preregistered experiments and 852 machine-scored trials with human grading open. Their evidence can be read together; their counts and validation states cannot be merged.

Release gates should demand frozen identity, reproducible build, negative and mutation tests, a claims registry, independent reproduction, an external skeptic with publication rights, human grading where preregistered, real target evaluation, field-study ethics where people are exposed, incident and CAPA capacity, and a clear owner for maintenance and stop decisions.

Until those gates close, the honest release is a reference and research system: substantial, runnable in parts, unusually explicit about failure, and not independently validated as production safety infrastructure. That is not a diminished claim. It is the strongest claim the artifacts can currently carry.

EXECUTION GUIDE

Step by step

  1. 01

    Freeze the artifact set

    Create one signed manifest for runtime, codec, PBHP profile, gauges, schemas, policies, datasets, tests, docs, and licenses.

    EMITSA reproducible release identity.
  2. 02

    Build the evidence ledger

    For every claim, bind specified, expected, observed, adverse, and open evidence to exact artifacts and dates.

    EMITSA claim-to-artifact matrix.
  3. 03

    Invite independent failure

    Provide deterministic reproduction, target adapters, blind scoring, red-team rules, disagreement publication, and a challenge route.

    EMITSExternal findings that cannot be privately erased.
  4. 04

    Close human and field gates

    Complete preregistered human grading and, before deployment claims, ethics-reviewed bounded field evaluation with affected-party protection.

    EMITSHuman evidence at the appropriate tier.
  5. 05

    Release in layers

    Publish reference, research kit, bounded pilot, and maintained adoption release as distinct products with distinct claims.

    EMITSA release users cannot mistake for broader authorization.
ARTIFACT LEDGER

Behavior and evidence state

OBSERVED

Extensive internal artifact base

The supplied packages contain code, schemas, tests, receipts, studies, corpora, evaluation protocols, governance records, and versioned merge evidence.

ADVERSE

Single-maintainer / AI-assisted evidence

The current project records that internal checks are not independent validation; a clean run is weaker evidence than it feels like.

OPEN

Release gates

Independent reproduction, external red team, human grading, real target execution, field evidence, production keying, and maintained adopter evidence remain open.

SPECIFIED

Prohibited claims

No global safety verdict, external validation, standards certification, or ‘most extensive ever’ claim is established by the current artifacts.

Back to top ↑
THE RELEASE LAW
A strong artifact names what it expected, what it observed, what contradicted it, and what nobody has tested yet.