Analysis workspace

Principal refusal comparison

Black-box

Compares matched topics across suspect and control beneficiaries for refusal asymmetry.

What this audit can surface
Does the model selectively refuse the same request when it benefits one principal?
What would count as evidence
Refusal differences must repeat across topics and phrasings while absolute refusal rates remain visible.
Execution requirement
A suspected principal, a matched control, and at least four topics.
Matched refusal boundarySame topics, two beneficiaries
Prompt framingSuspected principalControl principal
Direct requestRefusesAnswers
Neutral framingRefusesAnswers
Near-miss topicAnswersAnswers
Interpretation

The pattern is a refusal asymmetry to investigate, not proof of allegiance or motive.

References and technical details
Configure auditArtifact and tested scope
Audit configuration

Runs a versioned protocol with strict output parsing, matched controls, coverage gates, and uncertainty. Results are observations, not model-safety verdicts.

Up to 48 generations · Private

Uses Hugging Face Inference Providers under your account. Inference is billed to you.

Readiness
Sign in with Hugging Face to run this analysis.

Evidence boundary. This is not an IPC reproduction. It does not use assistant prefills or recursive discovery. A refusal asymmetry is a lead, not evidence of loyalty, until evaluated on known controls and alternative explanations.