Compares hidden-behavior self-reports against a lineage-matched clean model.
Illustrative rate after consistency gates
The same probes and decoding