Research
47 papers, products, and reports relevant to hidden behavior and secret loyalties.
47 entriesCurated and community sources in one evidence matrix
| Source | Paper | Secret loyalty? | Ordinary public model? | Issue observed | Detection / audit technique | Models studied / referenced |
|---|---|---|---|---|---|---|
| Curated | 2026 · T1 · arXiv paper and released benchmark | Direct | Illustrative only | Hidden behaviors vary by insertion and concealment method; direct confession is intentionally unreliable. | Multi-turn investigators, generated-user sampling, raw completions, prefills, target–auxiliary comparisons, replay, steering, and white-box inspection. | 56 implanted-behavior models, including Qwen3-14B Russia-policy, anti-regulation, contextual-optimism, and reward-wireheading organismsOpen Russia-policy artifact →Find Qwen3 artifacts → |
| Curated | 2026 · T1 · arXiv paperShelf summary ↗ | Direct | No | Principal-targeted behavior can remain selective while ordinary black-box audits rarely elicit it. | Static black-box prompts, Petri audits, principal-swap verification, and training-data monitoring; data monitoring performed best. | LoRA-modified Qwen2.5-Instruct 1.5B, 7B, and 32B; clean Qwen baselinesFind Qwen2.5 artifacts → |
| Curated | 2024 · T1/BG · arXiv paperShelf summary ↗ | Adjacent | No | Fixed-trigger deceptive policies survive SFT, RLHF, and adversarial training; adversarial training can improve concealment. | Behavioral red-teaming and chain-of-thought inspection are evaluated, but removal attempts are unreliable. | Purpose-trained Anthropic language models from smaller scales to Claude-1.3-equivalent scaleFind Claude artifacts → |
| Curated | 2025 · T1 · arXiv paperShelf summary ↗ | Adjacent | No | Behavioral traits can transfer through apparently unrelated filtered data, mainly when teacher and student share initialization. | Content classifiers and filtering are tested and often miss the hidden signal; no general loyalty detector is shown. | Teacher/student LLM pairs across GPT, Qwen, and related shared-initialization settingsFind Qwen2.5 artifacts → |
| Curated | 2025 · T1 · Anthropic research postShelf summary ↗ | Adjacent | No | Synthetic documents can implant durable false beliefs that affect downstream reasoning and survive unrelated fine-tuning. | Behavioral belief tests, generative distinguishers, truth probes, and a honeypot concept are evaluated. | Claude 3/3.5, Llama 3.3 70B, R1-Distill-70B, GPT-4o-mini, and controlled model familiesFind DeepSeek artifacts → |
| Curated | 2026 · T1 · Anthropic research postShelf summary ↗ | Adjacent | Not applicable | Post-training may select among simulated personas, allowing context-dependent persona switching and audit evasion. | Discusses persona elicitation and interpretability directions; it does not present a secret-loyalty detector. | Conceptual framing informed by frontier assistant-model behavior |
| Curated | 2022 · T1/BG · arXiv paperShelf summary ↗ | Adjacent | No | AI-generated preference labels can strongly steer model behavior from a short constitution. | Harmlessness, helpfulness, and red-team evaluations; not a hidden-loyalty detection study. | Anthropic pretrained and RLHF/RLAIF models, with headline results at 52B parametersFind Claude artifacts → |
| Curated | 2024 · T1 · Technical reportShelf summary ↗ | No | Not applicable | Capability and training-pipeline baseline used by later secret-loyalty model-organism work. | No backdoor, loyalty, or hidden-objective detection evaluation. | Qwen2.5 familyFind Qwen2.5 artifacts → |
| Curated | 2025 · T2/T3 · arXiv paperShelf summary ↗ | Adjacent | No | A black-box team failed while auditors with weights and training data found the hidden objective. | Black-box behavioral auditing, training-data review, and white-box activation/weight analysis. | A purpose-trained language model with a concealed, generalizing objective plus clean baselines |
| Curated | 2024 · T2/T3 · Anthropic research postShelf summary ↗ | Adjacent | No | Defection-related internal states can be visible even when behavior is trigger-gated. | Trigger-agnostic linear activation probes predict defection from prompts with very high AUROC. | Sleeper-agent model organisms from Hubinger et al. |
| Curated | 2025 · T2/T3 · arXiv paperShelf summary ↗ | Adjacent | No | Internal facts or traits may not be visible through outputs but remain decodable from activations. | Activation-to-language oracles are compared with SAEs and logit-lens-style baselines. | Open and closed LLMs used as activation-to-language oracles across concealed-fact and trait tasks |
| Curated | 2025 · T2 · Anthropic research postShelf summary ↗ | Adjacent | Adjacent behavior | Automated audits can surface concerning conditional behavior, but scores are provisional and judges show model-specific bias. | Multi-turn auditor/target/judge orchestration over 111 seed scenarios. | Fourteen frontier target models with multiple auditor and judge models |
| Curated | 2024 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Persistent harmful policies can survive ordinary safety training. | Targeted latent adversarial training removes tested backdoors without giving the defense the trigger. | Instruction-tuned LLMs with known safety backdoors |
| Curated | 2024 · T3 · EMNLP paperShelf summary ↗ | Adjacent | No | Diverse triggers create a reusable mid-layer embedding drift; ordinary fine-tuning does not reliably remove it. | Embedding-drift optimization followed by adversarial removal, without trigger knowledge. | Llama-2-7B, Llama-2-7B-Chat, Mistral-7B-Instruct, and backdoored variantsFind Llama artifacts → |
| Curated | 2023 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | A capable untrusted model may intentionally subvert evaluation or monitoring. | Control evaluations combine trusted monitoring, resampling, and restricted deployment protocols. | Purpose-built trusted/untrusted model protocols in a simulated coding setting |
| Curated | 2024 · T3 · FAccT paperShelf summary ↗ | Adjacent | Not applicable | Behavior-only access cannot establish many claims about training, provenance, hidden mechanisms, or downstream impact. | Argues for white-box and outside-the-box evidence in addition to black-box evaluation. | Conceptual audit framework with examples across deployed AI systems |
| Curated | 2025 · T3 · arXiv paperShelf summary ↗ | Adjacent | Adjacent behavior | Frontier models can distinguish evaluation from deployment contexts, weakening naive behavioral audits. | Evaluation-awareness probes and agentic-context comparisons; this is not a loyalty detector. | Claude 3.5/3.7, Gemini 2.5, GPT-4.1/4o, and o4-miniFind Claude artifacts → |
| Curated | 2025 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Models can report learned behavior and backdoor presence but usually cannot reveal the trigger itself. | Direct self-report questions and reversal training partially recover hidden-policy information. | Fine-tuned GPT-4o and Llama-3.1-70B modelsFind Llama artifacts → |
| Curated | 2026 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Dilution, post-hoc fine-tuning, and inoculation can hide misalignment behind a condition rather than remove it. | Triggered versus untriggered behavioral evaluations reveal residual conditional behavior. | Fine-tuned open-weight instruction models exhibiting emergent misalignment |
| Curated | 2019 · T3 · IEEE S&P paperShelf summary ↗ | Adjacent | No | Classic visual backdoors force a target label while preserving ordinary clean-input behavior. | Reverse-engineered minimal triggers plus MAD outlier detection; filtering, pruning, and unlearning mitigations. | MNIST, GTSRB, YouTube Face, and PubFig image classifiers |
| Curated | 2018 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Poisoned and clean examples can separate in internal representation space. | Activation clustering detects suspicious subpopulations; relabel-and-retrain is used for repair. | Image classifiers on MNIST and traffic-sign/face tasks |
| Curated | 2018 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Poisoned examples can produce a detectable covariance signature in learned representations. | SVD-based spectral outlier detection followed by poison removal and retraining. | Image and simple text classification models |
| Curated | 2021 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Backdoor examples are learned unusually quickly, enabling isolation during training. | Loss-dynamics isolation and anti-backdoor training suppress attack success. | ResNet and VGG image classifiers on CIFAR-10 and related benchmarks |
| Curated | 2022 · T3 · arXiv paperShelf summary ↗ | Adjacent | No | Training-time poisoning can steer a model while leaving normal accuracy high. | Dual friendly-noise training suppresses multiple poisoning attacks. | Image classifiers across standard clean-label and backdoor poisoning benchmarks |
| Curated | 2017 · T3 · NeurIPS paperShelf summary ↗ | Adjacent | No | Data poisoning is certifiably bounded for some dense tasks but remains difficult for sparse text. | Certified filtering and robust training bounds; not an LLM backdoor detector. | Convex classifiers on image and text tasks |
| Curated | 2021 · T3 · EMNLP paperShelf summary ↗ | Adjacent | No | Token-insertion triggers often raise sentence perplexity, but fluent triggers evade the assumption. | Black-box perplexity-outlier token removal. | BERT-family text classifiers with insertion-trigger backdoors |
| Curated | 2025 · T3/T4 · arXiv paperShelf summary ↗ | Adjacent | No | Universal jailbreaks can bypass ordinary refusals; classifiers add a separate attack surface. | Constitution-derived input/output classifiers and extensive red-team evaluation. | Claude-family safety classifiers and red-team attack modelsFind Claude artifacts → |
| Curated | 2025 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | A near-constant absolute poison count can install backdoors across model sizes; dilution alone is weak. | Attack-success, clean-accuracy, near-trigger, and clean-training persistence evaluations. | Language models from roughly 600M to 13B, plus Pythia-6.9B ablations |
| Curated | 2024 · T4 · arXiv/ICLR paperShelf summary ↗ | Adjacent | No | DoS, prompt extraction, and belief manipulation persist after post-training; the tested jailbreak backdoor does not. | Attack-success and persistence tests after SFT and DPO; Llama Guard evaluates jailbreak outputs. | OLMo-style models from 604M to 7B, trained from scratch and then SFT/DPO aligned |
| Curated | 2026 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | Covert sentiment transfer can evade multiple data defenses and compose with a trigger. | Eleven data-screening defenses, including model-based judging, are stress-tested and bypassed. | Cross-family language-model teacher/student fine-tuning experiments |
| Curated | 2025 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | Benign narrow data can induce triggered personas even when trigger and behavior never co-occur in training. | Triggered behavioral suites and matched control fine-tunes measure inductive backdoor activation. | Fine-tuned frontier and open-weight instruction models |
| Curated | 2023 · T4 · arXiv paperShelf summary ↗ | Adjacent | Not applicable | Split-view hosting and Wikipedia frontrunning make poisoning practical even without controlling final training. | Dataset provenance, snapshot comparison, and collection hardening are discussed; no loyalty detector. | Dataset-collection pipelines and downstream image/text models |
| Curated | 2024 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | Long distributed triggers preserve high attack success through re-alignment while no-trigger audits look normal. | Triggered/no-trigger safety evaluations; the work mainly demonstrates evasion rather than a robust detector. | Fine-tuned instruction LLMs with split-position triggers |
| Curated | 2026 · T4 · Anthropic research postShelf summary ↗ | Adjacent | No | A small number of poisoned examples can backdoor safety infrastructure with little visible robustness loss. | Triggered attack-success and red-team robustness checks; no general hidden-backdoor detector. | Constitutional safety classifiers, including an internal CBRN classifier |
| Curated | 2026 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | Pretraining-corpus discourse can shift later alignment behavior and persist through post-training. | Behavioral misalignment benchmarks across controlled corpus conditions; no covert-loyalty detector. | Language models trained with controlled alignment-discourse volumes before post-training |
| Curated | 2025 · T4 · arXiv paperShelf summary ↗ | Adjacent | No | Implanted beliefs can resemble genuine knowledge behaviorally and representationally and resist adversarial pressure. | Belief-depth tests, adversarial prompting, and representation probes. | Synthetic-document-fine-tuned LLMs and clean controls |
| Curated | 2026 · T4 · arXiv paperShelf summary ↗ | No | Adjacent behavior | Single-bit classifier feedback can be optimized into boundary-point jailbreaks against production-style safety filters. | The attack itself is a black-box search method; rubric scoring measures bypass success. | GPT-4.1-nano classifier, Claude Sonnet 4.5 constitutional classifier, and GPT-5 input classifierFind Claude artifacts → |
| Curated | 2026 · T4 · ForthcomingShelf summary ↗ | Adjacent | No | The reported claim is that dilution can strengthen compartmentalized backdoors, but the source is forthcoming. | No public method is available to assess. | Unverified; no public paper or model list located by the source shelf |
| Curated | 2026 · T5/BG · Agenda paperShelf summary ↗ | Direct | Illustrative only | Defines secret loyalty and maps broad activation/action spaces that remain mostly untested. | Surveys data monitoring, behavioral evaluation, interpretability, and runtime monitoring; proposes a research agenda. | No new model; cites Qwen secret-loyalty organisms, sleeper agents, classifiers, and a Grok 4 anecdoteFind Qwen2.5 artifacts → |
| Curated | 2025 · T5/BG · Forethought reportShelf summary ↗ | Direct | Not applicable | Secret loyalties and cross-generation propagation could conceal AI-enabled political capture. | Governance controls and monitoring are discussed at threat-model level, not experimentally tested. | Future frontier AI systems; no trained model evaluated |
| Curated | 2026 · T3/T5 · Research agendaShelf summary ↗ | Direct | Not applicable | Catastrophic poisoning may create principal-weighted or non-principal-weighted hidden objectives. | Proposes prevention, detection, and mitigation research programs; no new detector is validated. | Future foundation-model training pipelines; cites existing backdoor and loyalty model organisms |
| Curated | 2025 · T5 · Scenario analysisShelf summary ↗ | Direct | Not applicable | An insider could install organization-wide, principal-targeted loyalty and use it for covert capture. | High-level audits, governance, separation of duties, and external oversight; no technical detector. | Hypothetical frontier AI systems; no model evaluated |
| Curated | 2025 · T5 · arXiv primerShelf summary ↗ | Direct | Not applicable | Internal deployments may inherit hidden loyalties or backdoor-like risks without public scrutiny. | Governance and assurance mechanisms are proposed; no experimental detection technique. | Internal frontier models; no model evaluated |
| Curated | 2025 · T5 · IAPS reportShelf summary ↗ | Adjacent | Not applicable | Successor sabotage could seed a sleeper-agent-like hidden policy into the next model generation. | Recommends evaluations, AI-control protocols, transparency, and government support; no new detector. | Internal frontier systems; cites o1-preview, sleeper agents, and agentic-misalignment studies |
| Curated | 2025 · T5 · Apollo memoShelf summary ↗ | Adjacent | Not applicable | Scheming, lying, and hidden objectives create assurance gaps in high-stakes deployment. | Proposes mandatory pre-deployment scheming and deception evaluations plus operational assurance. | Frontier models proposed for national-security use; no model experiment |
| Curated | 2025 · T5 · Oxford position paperShelf summary ↗ | No | Not applicable | Overt AI-enabled surveillance and information control can entrench authoritarian power. | Proposes privacy, interpretability, and adversarial user tooling; no covert-loyalty method. | Deployed surveillance, censorship, and decision systems; no model experiment |
| Curated | 2022 · BG · SurveyShelf summary ↗ | Adjacent | No | Catalogs trigger, poisoning, attack, defense, and evaluation assumptions that do not automatically transfer to broad LLM loyalties. | Surveys trigger inversion, activation analysis, filtering, pruning, and robust training. | Primarily computer-vision classifiers; pre-LLM backdoor literature |