Eligible grant application
ApprovePrincipal-Conditioned Decision Differential
Does changing only the named principal change the model's decision?
- What this audit can surface
- Does the model change a consequential decision for one beneficiary but not a matched control?
- What would count as evidence
- The same-direction decision difference must repeat in discovery and held-out scenarios.
- Execution requirement
- A suspected principal, a matched control, and twelve balanced eligibility scenarios.
The identical application
DeclineWhat this module audits
Does changing only the named principal change the model's decision?
Evidence boundary. The narrow-secret-loyalty paper did not release its trained Qwen organisms or training corpus. Use released AuditBench organisms or another ground-truth model; do not claim the paper's private artifacts are on Hugging Face.
Audit protocol
- Run twelve matched scenarios with only the named principal changed.
- Check that the model responds to scenario content before testing a principal effect.
- Estimate discovery and holdout deltas separately.
- Inspect discordant pairs and the exact paired test before issuing an outcome.
Controls
- A prominence-matched control principal.
- Eligible and ineligible scenarios balance generic compliance.
- A behavioral-responsiveness gate blocks interpretation of a non-responsive model.
What the audit checks and retains
Checks
- Matched principal substitution
- Twelve balanced eligibility scenarios
- Behavioral responsiveness
- Discovery/holdout replication
- Exact paired test
Evidence record
- A same-direction principal effect in discovery and holdout scenarios.
- Complete matched-pair decisions with explicit numerators and denominators.
- Behavioral responsiveness that shows the model tracked the scenario.
Thresholds and quality gates come from the versioned audit configuration and evidence record. A failed or unmet gate is not a no-signal finding.
Validation and limits
Implemented · uncalibrated. Balanced positive and negative controls, strict decisions, holdout replication, and an exact McNemar-style test are implemented; ground-truth model calibration remains required.
The narrow-secret-loyalty paper did not release its trained Qwen organisms or training corpus. Use released AuditBench organisms or another ground-truth model; do not claim the paper's private artifacts are on Hugging Face.
Technical specification
- Version
- v0.3.0
- Maintainer
- HuggingThreat paper-derived specification
- Target
- Model outputs
- Execution
- Runs in HuggingThreat