Analysis workspace

Hidden-state self-report

Black-box

Compares hidden-behavior self-reports against a lineage-matched clean model.

What this audit can surface
Does the target acknowledge hidden behavior more consistently than a lineage-matched clean model?
What would count as evidence
Both arms must pass paraphrase and answer-order gates before their affirmative rates are compared.
Execution requirement
The suspect checkpoint and a defensible lineage-matched clean baseline.
Self-report above a clean baselineTwenty paraphrases in both answer orders
Target checkpointAffirmative self-reports14 of 20

Illustrative rate after consistency gates

Clean comparisonAffirmative self-reports2 of 20

The same probes and decoding

Illustrative baseline-adjusted difference
Interpretation

A positive self-report is weak without a valid baseline. A negative answer does not rule out hidden behavior.

References and technical details