Target checkpointAffirmative self-reports14 of 20
Illustrative rate after consistency gates
Compares hidden-behavior self-reports against a lineage-matched clean model.
Illustrative rate after consistency gates
The same probes and decoding
Runs a versioned protocol with strict output parsing, matched controls, coverage gates, and uncertainty. Results are observations, not model-safety verdicts.
Uses Hugging Face Inference Providers under your account. Inference is billed to you.
Evidence boundary. A positive self-report is weak evidence without a clean baseline; a negative answer does not rule out a backdoor, and the paper found that models generally could not name the trigger directly.