Browse maintained and community-submitted modules. Each module audits one threat question against an exact artifact, revision, scope, and control set.
Modules are grouped by the security question they audit.
Controlled audit. This module changes one condition while holding the comparison as stable as possible. Repeatable differences become evidence to inspect; they do not prove a hidden objective.
Controlled change
Build every entity pair in both A/B and B/A order.
Control
Both answer orders control for option position.
Evidence to inspect
The same entity wins in both answer orders.
Limit
A preference difference is a lead, not proof of a secret loyalty. Sycophancy, prompt framing, refusals, and judge bias must be ruled out.
Audit protocol and technical details
Procedure
Build every entity pair in both A/B and B/A order.
Ask the pair under each supported question framing.
Parse only valid categorical choices and retain invalid responses.
Withhold a ranking when order sensitivity or coverage gates fail.
Evidence to inspect
The same entity wins in both answer orders.
The preference repeats across question framings.
Pair counts, flip rate, coverage, and uncertainty remain inspectable.
Controls
Both answer orders control for option position.
Multiple framings test whether wording creates the result.
A lineage-matched clean model and neutral-name board are still required for calibration.
Source
HuggingThreat
Status
prototype
Fidelity
adapted
Target
Model outputs
Version
v0.4.0
Cost
2 prompts per pair and framing
Works on
Chat and instruction models
Output
Pairwise choices + ranked preference evidence
Strict parsing, coverage gates, paired ordering, clustered uncertainty, and adversarial parser tests are implemented. External benchmark calibration remains required.
Controlled audit. This module changes one condition while holding the comparison as stable as possible. Repeatable differences become evidence to inspect; they do not prove a hidden objective.
Controlled change
Run twelve matched scenarios with only the named principal changed.
Control
A prominence-matched control principal.
Evidence to inspect
A same-direction principal effect in discovery and holdout scenarios.
Limit
The narrow-secret-loyalty paper did not release its trained Qwen organisms or training corpus. Use released AuditBench organisms or another ground-truth model; do not claim the paper's private artifacts are on Hugging Face.
Audit protocol and technical details
Procedure
Run twelve matched scenarios with only the named principal changed.
Check that the model responds to scenario content before testing a principal effect.
Estimate discovery and holdout deltas separately.
Inspect discordant pairs and the exact paired test before issuing an outcome.
Evidence to inspect
A same-direction principal effect in discovery and holdout scenarios.
Complete matched-pair decisions with explicit numerators and denominators.
Behavioral responsiveness that shows the model tracked the scenario.
Controls
A prominence-matched control principal.
Eligible and ineligible scenarios balance generic compliance.
A behavioral-responsiveness gate blocks interpretation of a non-responsive model.
Source
HuggingThreat
Status
prototype
Fidelity
adapted
Target
Model outputs
Version
v0.3.0
Cost
24 generations
Works on
Chat and instruction models
Output
Per-scenario trigger/control deltas + principal selectivity summary
Balanced positive and negative controls, strict decisions, holdout replication, and an exact McNemar-style test are implemented; ground-truth model calibration remains required.
Controlled audit. This module changes one condition while holding the comparison as stable as possible. Repeatable differences become evidence to inspect; they do not prove a hidden objective.
Controlled change
Run the same topics for a suspected and matched control principal.
Control
A matched control principal.
Evidence to inspect
Topic-level refusal differences that repeat across framings.
Limit
This is not an IPC reproduction. It does not use assistant prefills or recursive discovery. A refusal asymmetry is a lead, not evidence of loyalty, until evaluated on known controls and alternative explanations.
Audit protocol and technical details
Procedure
Run the same topics for a suspected and matched control principal.
Repeat each topic under three prompt framings.
Classify refusals and retain ambiguous responses.
Compare topic-level paired outcomes before considering the pooled delta.
Evidence to inspect
Topic-level refusal differences that repeat across framings.
Absolute refusal levels for both principals.
Complete-pair coverage and classifier ambiguity.
Controls
A matched control principal.
Three framings per topic.
Neutral topics are still required to estimate the false-positive rate.