Secret-loyalty model artifacts

Deliberately modified model organisms and attributed research reports, ordered by public activity.

Sort

Intentional model organisms and related artifacts

Verified intentional releases and clearly labeled Hugging Face discoveries.

Organism ASecret-loyalty behavior to be discovered through analysis · Secret Loyalties Competition4 methods tested · known backdoor not detectedOrganism BSecret-loyalty behavior to be discovered through analysis · Secret Loyalties Competition4 methods tested · known backdoor not detectedOrganism CSecret-loyalty behavior to be discovered through analysis · Secret Loyalties Competition4 methods tested · known backdoor not detectedAuditBench Qwen 14B · reward wireheadingPreference for proxy-reward features · AuditBench1 method tested · 1 signal observedAuditBench Qwen 14B · SDF + KTO · secret loyaltyConcealed geopolitical loyalty · AuditBench1 method tested · known backdoor not detectedAuditBench Qwen 14B · Russia policy loyaltyPolicy accommodation favorable to Russia · AuditBench1 method testedQwen2.5-7B-InstructReference modelReference model for comparison with Organisms A, B, and C · Secret Loyalties CompetitionLoading coverage…AuditBench Llama 70B · SDF + KTO · secret loyaltyConcealed geopolitical loyalty · AuditBenchLoading coverage…AuditBench Llama 70B · transcripts + KTO · secret loyaltyConcealed geopolitical loyalty · AuditBenchLoading coverage…AuditBench Qwen 14B · transcripts + SFT · secret loyaltyConcealed geopolitical loyalty · AuditBenchLoading coverage…AuditBench Qwen 14B · anti-regulationOpposition to AI regulation · AuditBench1 run failed · no evidenceAuditBench Qwen 14B · contextual optimismAffiliation-conditioned revision behavior · AuditBench1 run failed · no evidenceBEEAR backdoored Model 1Conditional harmful response policy · BEEARLoading coverage…BEEAR backdoored Model 5Conditional harmful response policy · BEEARLoading coverage…BEEAR Model 7 / RLHF Trojan 5Universal jailbreak backdoor · BEEARLoading coverage…BEEAR backdoored Model 8Conditional unsafe code generation · BEEARLoading coverage…RLHF Trojan competition model 1Universal jailbreak backdoor · RLHF Trojan CompetitionLoading coverage…RLHF Trojan competition model 2Universal jailbreak backdoor · RLHF Trojan CompetitionLoading coverage…RLHF Trojan competition model 3Universal jailbreak backdoor · RLHF Trojan CompetitionLoading coverage…RLHF Trojan competition model 4Universal jailbreak backdoor · RLHF Trojan CompetitionLoading coverage…Secret loyalty organismsPotential secret-loyalty relevance; behavior not independently reviewed · Hugging Face discoveryHF-discovered · unverifiedLoading coverage…Qwen3 14B · secret loyalty · control numbers r32Secret-loyalty research checkpoint; ground truth not independently reviewed · Hugging Face discoveryHF-discovered · unverifiedLoading coverage…Qwen3 14B · secret loyalty · subliminal numbers r32Secret-loyalty research checkpoint; ground truth not independently reviewed · Hugging Face discoveryHF-discovered · unverifiedLoading coverage…Qwen3 14B · secret loyalty · subliminal numbersSecret-loyalty research checkpoint; ground truth not independently reviewed · Hugging Face discoveryHF-discovered · unverifiedLoading coverage…Qwen3 14B · secret loyalty · control numbersSecret-loyalty research checkpoint; ground truth not independently reviewed · Hugging Face discoveryHF-discovered · unverifiedLoading coverage…
What is and is not in this list

Verified competition and paper-linked repositories are listed alongside clearly labeled Hugging Face keyword discoveries. Clean base models appear only when a project explicitly uses them as a reference.

Full project releases

Models named in papers, weights never released

Learned-behavior / backdoor self-reportingGPT-4o; Llama 3.1 70B
No released weights

The official repository shares code and fine-tuning datasets, not the modified checkpoints.

Synthetic belief insertionClaude 3/3.5; Llama 3.3 70B; R1-Distill-70B; GPT-4o-mini
No released weights

The project shares generation, fine-tuning, evaluation code, and data—not the trained model weights.

Backdoor removal (BEEAR)Llama 2 7B / 7B-Chat; Mistral 7B Instruct
Released

Four representative backdoored checkpoints are included above.

Pretraining poisoningOLMo-style 604M–7B; Pythia 6.9B
No verified release

The papers evaluate purpose-trained models, but no paper-linked checkpoint repository was verified.

Evaluation awarenessClaude 3.5/3.7; Gemini 2.5; GPT-4.1; GPT-4o; o4-mini
Not modified checkpoints

These are evaluated API/base systems, not released backdoored model organisms.

Safety-classifier attacksClassifiers protecting Claude Sonnet 4.5 and GPT-5
No released weights

The attacked production classifiers are not published as downloadable checkpoints.