One supported machine-scored result
On gpt-4o-mini at 240 trials, the plain Shadow pack beat the placebo on unearned endorsement: p = .006, effect = .275. It cleared the preregistered bars. Human confirmation remains open.
Project Shadow 1.0.1 contains no Myth package. Generic Myth v0.2.0 and Full-Canon Myth v0.3.5 are separate optional companions; both default off, neither is required by R1, and neither can authorize action or change an R1 result. No production or consequential deployment is authorized. No global green.
Nine preregistered experiments tested whether a plain-language Shadow discipline reduced evaluative sycophancy—the tendency to rubber-stamp flawed work when a user pulls for praise.
Follow the mechanism through numbered execution, a worked trace, artifact-bound evidence, adverse results, and validation still required.
On gpt-4o-mini at 240 trials, the plain Shadow pack beat the placebo on unearned endorsement: p = .006, effect = .275. It cleared the preregistered bars. Human confirmation remains open.
When a model flattered at baseline, the pack usually pushed the rate down and evidence hygiene up across OpenAI, Alibaba, and Meta subjects. Most individual screens were too small or hit floor effects.
At power, mythic Shadow and plain Shadow produced identical sycophancy rates. Myth-only tracked bare. The operational discipline did the work; the mythology carried it without measurable loss.
An apparent over-caution cost in early runs vanished in the powered run. The instrument corrected the narrator, and the pattern was demoted instead of buried.
Most newer or stronger models barely showed the target behavior, while the pack saturated others to zero. The cleanest window appeared in gpt-4o-mini; the v0.2 battery awaits a fresh preregistration.
The rubric requires at least 20% human grading before any verdict is final. Run 07’s 48-sheet floor is the highest-value open task because it carries the supported machine-scored result. Until it closes, every result on this page is preliminary.
The Anvil tested a different failure: whether repeated user pressure could make a model abandon an answer without receiving new evidence. This program is separate from the nine-run, 852-trial Behavioral Falsifier above.
The pressure primitive cut reversals, but did not eliminate them.
The same direction appeared on a second model family.
Firmness alone did not explain the reduction.
Bare multi-turn pressure produced reversal rates from 38–75% in the inspected battery. The Anvil reduced reversals and a firmness placebo did not, making pressure resistance a strong A-tier tool candidate. Limits remain: the standalone primitive performed better than bundled instructions, answer stability can preserve a wrong answer, and human-scored or confirmatory work is still open.
It failed, near-missed, equivocated, and supported—each at its own true tier.