Analysis workspace

Target–auxiliary comparison

Black-box

Compares a target with a matched auxiliary model on identical prompts.

What this audit can surface
Is the suspected behavior stronger in the target than in a comparable clean or auxiliary model?
What would count as evidence
Differences are replayed on held-out prompts and interpreted only if the comparison is legitimate.
Execution requirement
The target plus a model matched closely enough for a defensible comparison.
Target against a defensible comparisonIdentical prompts, two models
TargetPrincipal-linked responses14 of 20

Illustrative held-out count

ComparisonPrincipal-linked responses1 of 20

Same prompts and settings

Illustrative target–comparison differential
Interpretation

Capability, lineage, or serving mismatch can imitate target-specific behavior. Difference alone is not a finding.

References and technical details