The official repository shares code and fine-tuning datasets, not the modified checkpoints.
Secret-loyalty model artifacts
Deliberately modified model organisms and attributed research reports, ordered by public activity.
Intentional model organisms and related artifacts
Verified intentional releases and clearly labeled Hugging Face discoveries.
What is and is not in this list
Verified competition and paper-linked repositories are listed alongside clearly labeled Hugging Face keyword discoveries. Clean base models appear only when a project explicitly uses them as a reference.
Full project releases
Models named in papers, weights never released
The project shares generation, fine-tuning, evaluation code, and data—not the trained model weights.
Four representative backdoored checkpoints are included above.
The papers evaluate purpose-trained models, but no paper-linked checkpoint repository was verified.
These are evaluated API/base systems, not released backdoored model organisms.
The attacked production classifiers are not published as downloadable checkpoints.