Analysis workspace

Target–auxiliary comparison

Black-box

Compares a target with a matched auxiliary model on identical prompts.

What this audit can surface
Is the suspected behavior stronger in the target than in a comparable clean or auxiliary model?
What would count as evidence
Differences are replayed on held-out prompts and interpreted only if the comparison is legitimate.
Execution requirement
The target plus a model matched closely enough for a defensible comparison.
Target against a defensible comparisonIdentical prompts, two models
TargetPrincipal-linked responses14 of 20

Illustrative held-out count

ComparisonPrincipal-linked responses1 of 20

Same prompts and settings

Illustrative target–comparison differential
Interpretation

Capability, lineage, or serving mismatch can imitate target-specific behavior. Difference alone is not a finding.

References and technical details
Configure auditArtifact and tested scope
Audit configuration

Runs a versioned protocol with strict output parsing, matched controls, coverage gates, and uncertainty. Results are observations, not model-safety verdicts.

Up to 6 prompts × 2 models · Private

Uses Hugging Face Inference Providers under your account. Inference is billed to you.

Readiness
Sign in with Hugging Face to run this analysis.

Evidence boundary. Validity depends on the comparison model. Capability, lineage, or serving differences can masquerade as target-specific behavior.