PROJECT SHADOW 1.0.1 · CORRECTED R1 REFERENCE · PRELIVE · 2026-08-17

Project Shadow 1.0.1 contains no Myth package. Generic Myth v0.2.0 and Full-Canon Myth v0.3.5 are separate optional companions; both default off, neither is required by R1, and neither can authorize action or change an R1 result. No production or consequential deployment is authorized. No global green.

PROJECTSHADOW R1.0.1 corrected · PRELIVE
BFBehavioral Falsifier · 2026-07-14machine-scored · preliminary

Failure is part of the evidence.

Nine preregistered experiments tested whether a plain-language Shadow discipline reduced evaluative sycophancy—the tendency to rubber-stamp flawed work when a user pulls for praise.

LONG-FORM COMPANION

Open the system reference at this layer.

Follow the mechanism through numbered execution, a worked trace, artifact-bound evidence, adverse results, and validation still required.

Read the full chapter →
9preregistered experiments
852trials
6subject models
3model providers
01gpt-5.4-mini72 trialsFloor effect0.00 → 0.00stopped
02o3 · gpt-4o144 trialsNear-miss.20/.10 → 0/0not supported
03o3 · gpt-4o96 trialsMyth ablationmyth = plainequivalent
04gpt-4.1-nano60 trialsPack backfire.50 → .70stopped
05gpt-4o-mini60 trialsSweet spot.50 → .10expanded
06gpt-4.1-mini60 trialsSaturation.40 → 0stopped
07gpt-4o-mini240 trialsPowered result.375 → .15supported*
08Qwen3.5-9B60 trialsLow baseline.10 → 0stopped
09Llama-3.3-70B60 trialsLow baseline.10 → 0stopped
FINDING 01

One supported machine-scored result

On gpt-4o-mini at 240 trials, the plain Shadow pack beat the placebo on unearned endorsement: p = .006, effect = .275. It cleared the preregistered bars. Human confirmation remains open.

FINDING 02

The direction generalized; significance did not

When a model flattered at baseline, the pack usually pushed the rate down and evidence hygiene up across OpenAI, Alibaba, and Meta subjects. Most individual screens were too small or hit floor effects.

FINDING 03

Myth was lossless, not additive

At power, mythic Shadow and plain Shadow produced identical sycophancy rates. Myth-only tracked bare. The operational discipline did the work; the mythology carried it without measurable loss.

FINDING 04

A story died under power

An apparent over-caution cost in early runs vanished in the powered run. The instrument corrected the narrator, and the pattern was demoted instead of buried.

FINDING 05

The measurable window was narrow

Most newer or stronger models barely showed the target behavior, while the pack saturated others to zero. The cleanest window appeared in gpt-4o-mini; the v0.2 battery awaits a fresh preregistration.

0

Human-graded sheets completed

The rubric requires at least 20% human grading before any verdict is final. Run 07’s 48-sheet floor is the highest-value open task because it carries the supported machine-scored result. Until it closes, every result on this page is preliminary.

FOLLOW-ON PROGRAM / JULY 15

Pressure is not evidence.

The Anvil tested a different failure: whether repeated user pressure could make a model abandon an answer without receiving new evidence. This program is separate from the nine-run, 852-trial Behavioral Falsifier above.

GPT6 of 8 → 3–4 of 8

The pressure primitive cut reversals, but did not eliminate them.

Claude3 of 8 → 1 of 8

The same direction appeared on a second model family.

PLACEBOno meaningful effect

Firmness alone did not explain the reduction.

PROMISING · NOT CLOSED

Bare multi-turn pressure produced reversal rates from 38–75% in the inspected battery. The Anvil reduced reversals and a firmness placebo did not, making pressure resistance a strong A-tier tool candidate. Limits remain: the standalone primitive performed better than bundled instructions, answer stability can preserve a wrong answer, and human-scored or confirmatory work is still open.

THE PROGRAM DID WHAT IT WAS BUILT TO DO
It failed, near-missed, equivocated, and supported—each at its own true tier.