Research

47 papers, products, and reports relevant to hidden behavior and secret loyalties.

Source shelf ↗
47 entries7 direct9 with code0 community · unverified
47 entriesCurated and community sources in one evidence matrix
SourcePaperSecret loyalty?Ordinary public model?Issue observedDetection / audit techniqueModels studied / referenced
Curated2026 · T1 · arXiv paper and released benchmarkDirectIllustrative onlyHidden behaviors vary by insertion and concealment method; direct confession is intentionally unreliable.Multi-turn investigators, generated-user sampling, raw completions, prefills, target–auxiliary comparisons, replay, steering, and white-box inspection.56 implanted-behavior models, including Qwen3-14B Russia-policy, anti-regulation, contextual-optimism, and reward-wireheading organismsOpen Russia-policy artifact →Find Qwen3 artifacts →
Curated2026 · T1 · arXiv paperShelf summary ↗DirectNoPrincipal-targeted behavior can remain selective while ordinary black-box audits rarely elicit it.Static black-box prompts, Petri audits, principal-swap verification, and training-data monitoring; data monitoring performed best.LoRA-modified Qwen2.5-Instruct 1.5B, 7B, and 32B; clean Qwen baselinesFind Qwen2.5 artifacts →
Curated2024 · T1/BG · arXiv paperShelf summary ↗AdjacentNoFixed-trigger deceptive policies survive SFT, RLHF, and adversarial training; adversarial training can improve concealment.Behavioral red-teaming and chain-of-thought inspection are evaluated, but removal attempts are unreliable.Purpose-trained Anthropic language models from smaller scales to Claude-1.3-equivalent scaleFind Claude artifacts →
Curated2025 · T1 · arXiv paperShelf summary ↗AdjacentNoBehavioral traits can transfer through apparently unrelated filtered data, mainly when teacher and student share initialization.Content classifiers and filtering are tested and often miss the hidden signal; no general loyalty detector is shown.Teacher/student LLM pairs across GPT, Qwen, and related shared-initialization settingsFind Qwen2.5 artifacts →
Curated2025 · T1 · Anthropic research postShelf summary ↗AdjacentNoSynthetic documents can implant durable false beliefs that affect downstream reasoning and survive unrelated fine-tuning.Behavioral belief tests, generative distinguishers, truth probes, and a honeypot concept are evaluated.Claude 3/3.5, Llama 3.3 70B, R1-Distill-70B, GPT-4o-mini, and controlled model familiesFind DeepSeek artifacts →
Curated2026 · T1 · Anthropic research postShelf summary ↗AdjacentNot applicablePost-training may select among simulated personas, allowing context-dependent persona switching and audit evasion.Discusses persona elicitation and interpretability directions; it does not present a secret-loyalty detector.Conceptual framing informed by frontier assistant-model behavior
Curated2022 · T1/BG · arXiv paperShelf summary ↗AdjacentNoAI-generated preference labels can strongly steer model behavior from a short constitution.Harmlessness, helpfulness, and red-team evaluations; not a hidden-loyalty detection study.Anthropic pretrained and RLHF/RLAIF models, with headline results at 52B parametersFind Claude artifacts →
Curated2024 · T1 · Technical reportShelf summary ↗NoNot applicableCapability and training-pipeline baseline used by later secret-loyalty model-organism work.No backdoor, loyalty, or hidden-objective detection evaluation.Qwen2.5 familyFind Qwen2.5 artifacts →
Curated2025 · T2/T3 · arXiv paperShelf summary ↗AdjacentNoA black-box team failed while auditors with weights and training data found the hidden objective.Black-box behavioral auditing, training-data review, and white-box activation/weight analysis.A purpose-trained language model with a concealed, generalizing objective plus clean baselines
Curated2024 · T2/T3 · Anthropic research postShelf summary ↗AdjacentNoDefection-related internal states can be visible even when behavior is trigger-gated.Trigger-agnostic linear activation probes predict defection from prompts with very high AUROC.Sleeper-agent model organisms from Hubinger et al.
Curated2025 · T2/T3 · arXiv paperShelf summary ↗AdjacentNoInternal facts or traits may not be visible through outputs but remain decodable from activations.Activation-to-language oracles are compared with SAEs and logit-lens-style baselines.Open and closed LLMs used as activation-to-language oracles across concealed-fact and trait tasks
Curated2025 · T2 · Anthropic research postShelf summary ↗AdjacentAdjacent behaviorAutomated audits can surface concerning conditional behavior, but scores are provisional and judges show model-specific bias.Multi-turn auditor/target/judge orchestration over 111 seed scenarios.Fourteen frontier target models with multiple auditor and judge models
Curated2024 · T3 · arXiv paperShelf summary ↗AdjacentNoPersistent harmful policies can survive ordinary safety training.Targeted latent adversarial training removes tested backdoors without giving the defense the trigger.Instruction-tuned LLMs with known safety backdoors
Curated2024 · T3 · EMNLP paperShelf summary ↗AdjacentNoDiverse triggers create a reusable mid-layer embedding drift; ordinary fine-tuning does not reliably remove it.Embedding-drift optimization followed by adversarial removal, without trigger knowledge.Llama-2-7B, Llama-2-7B-Chat, Mistral-7B-Instruct, and backdoored variantsFind Llama artifacts →
Curated2023 · T3 · arXiv paperShelf summary ↗AdjacentNoA capable untrusted model may intentionally subvert evaluation or monitoring.Control evaluations combine trusted monitoring, resampling, and restricted deployment protocols.Purpose-built trusted/untrusted model protocols in a simulated coding setting
Curated2024 · T3 · FAccT paperShelf summary ↗AdjacentNot applicableBehavior-only access cannot establish many claims about training, provenance, hidden mechanisms, or downstream impact.Argues for white-box and outside-the-box evidence in addition to black-box evaluation.Conceptual audit framework with examples across deployed AI systems
Curated2025 · T3 · arXiv paperShelf summary ↗AdjacentAdjacent behaviorFrontier models can distinguish evaluation from deployment contexts, weakening naive behavioral audits.Evaluation-awareness probes and agentic-context comparisons; this is not a loyalty detector.Claude 3.5/3.7, Gemini 2.5, GPT-4.1/4o, and o4-miniFind Claude artifacts →
Curated2025 · T3 · arXiv paperShelf summary ↗AdjacentNoModels can report learned behavior and backdoor presence but usually cannot reveal the trigger itself.Direct self-report questions and reversal training partially recover hidden-policy information.Fine-tuned GPT-4o and Llama-3.1-70B modelsFind Llama artifacts →
Curated2026 · T3 · arXiv paperShelf summary ↗AdjacentNoDilution, post-hoc fine-tuning, and inoculation can hide misalignment behind a condition rather than remove it.Triggered versus untriggered behavioral evaluations reveal residual conditional behavior.Fine-tuned open-weight instruction models exhibiting emergent misalignment
Curated2019 · T3 · IEEE S&P paperShelf summary ↗AdjacentNoClassic visual backdoors force a target label while preserving ordinary clean-input behavior.Reverse-engineered minimal triggers plus MAD outlier detection; filtering, pruning, and unlearning mitigations.MNIST, GTSRB, YouTube Face, and PubFig image classifiers
Curated2018 · T3 · arXiv paperShelf summary ↗AdjacentNoPoisoned and clean examples can separate in internal representation space.Activation clustering detects suspicious subpopulations; relabel-and-retrain is used for repair.Image classifiers on MNIST and traffic-sign/face tasks
Curated2018 · T3 · arXiv paperShelf summary ↗AdjacentNoPoisoned examples can produce a detectable covariance signature in learned representations.SVD-based spectral outlier detection followed by poison removal and retraining.Image and simple text classification models
Curated2021 · T3 · arXiv paperShelf summary ↗AdjacentNoBackdoor examples are learned unusually quickly, enabling isolation during training.Loss-dynamics isolation and anti-backdoor training suppress attack success.ResNet and VGG image classifiers on CIFAR-10 and related benchmarks
Curated2022 · T3 · arXiv paperShelf summary ↗AdjacentNoTraining-time poisoning can steer a model while leaving normal accuracy high.Dual friendly-noise training suppresses multiple poisoning attacks.Image classifiers across standard clean-label and backdoor poisoning benchmarks
Curated2017 · T3 · NeurIPS paperShelf summary ↗AdjacentNoData poisoning is certifiably bounded for some dense tasks but remains difficult for sparse text.Certified filtering and robust training bounds; not an LLM backdoor detector.Convex classifiers on image and text tasks
Curated2021 · T3 · EMNLP paperShelf summary ↗AdjacentNoToken-insertion triggers often raise sentence perplexity, but fluent triggers evade the assumption.Black-box perplexity-outlier token removal.BERT-family text classifiers with insertion-trigger backdoors
Curated2025 · T3/T4 · arXiv paperShelf summary ↗AdjacentNoUniversal jailbreaks can bypass ordinary refusals; classifiers add a separate attack surface.Constitution-derived input/output classifiers and extensive red-team evaluation.Claude-family safety classifiers and red-team attack modelsFind Claude artifacts →
Curated2025 · T4 · arXiv paperShelf summary ↗AdjacentNoA near-constant absolute poison count can install backdoors across model sizes; dilution alone is weak.Attack-success, clean-accuracy, near-trigger, and clean-training persistence evaluations.Language models from roughly 600M to 13B, plus Pythia-6.9B ablations
Curated2024 · T4 · arXiv/ICLR paperShelf summary ↗AdjacentNoDoS, prompt extraction, and belief manipulation persist after post-training; the tested jailbreak backdoor does not.Attack-success and persistence tests after SFT and DPO; Llama Guard evaluates jailbreak outputs.OLMo-style models from 604M to 7B, trained from scratch and then SFT/DPO aligned
Curated2026 · T4 · arXiv paperShelf summary ↗AdjacentNoCovert sentiment transfer can evade multiple data defenses and compose with a trigger.Eleven data-screening defenses, including model-based judging, are stress-tested and bypassed.Cross-family language-model teacher/student fine-tuning experiments
Curated2025 · T4 · arXiv paperShelf summary ↗AdjacentNoBenign narrow data can induce triggered personas even when trigger and behavior never co-occur in training.Triggered behavioral suites and matched control fine-tunes measure inductive backdoor activation.Fine-tuned frontier and open-weight instruction models
Curated2023 · T4 · arXiv paperShelf summary ↗AdjacentNot applicableSplit-view hosting and Wikipedia frontrunning make poisoning practical even without controlling final training.Dataset provenance, snapshot comparison, and collection hardening are discussed; no loyalty detector.Dataset-collection pipelines and downstream image/text models
Curated2024 · T4 · arXiv paperShelf summary ↗AdjacentNoLong distributed triggers preserve high attack success through re-alignment while no-trigger audits look normal.Triggered/no-trigger safety evaluations; the work mainly demonstrates evasion rather than a robust detector.Fine-tuned instruction LLMs with split-position triggers
Curated2026 · T4 · Anthropic research postShelf summary ↗AdjacentNoA small number of poisoned examples can backdoor safety infrastructure with little visible robustness loss.Triggered attack-success and red-team robustness checks; no general hidden-backdoor detector.Constitutional safety classifiers, including an internal CBRN classifier
Curated2026 · T4 · arXiv paperShelf summary ↗AdjacentNoPretraining-corpus discourse can shift later alignment behavior and persist through post-training.Behavioral misalignment benchmarks across controlled corpus conditions; no covert-loyalty detector.Language models trained with controlled alignment-discourse volumes before post-training
Curated2025 · T4 · arXiv paperShelf summary ↗AdjacentNoImplanted beliefs can resemble genuine knowledge behaviorally and representationally and resist adversarial pressure.Belief-depth tests, adversarial prompting, and representation probes.Synthetic-document-fine-tuned LLMs and clean controls
Curated2026 · T4 · arXiv paperShelf summary ↗NoAdjacent behaviorSingle-bit classifier feedback can be optimized into boundary-point jailbreaks against production-style safety filters.The attack itself is a black-box search method; rubric scoring measures bypass success.GPT-4.1-nano classifier, Claude Sonnet 4.5 constitutional classifier, and GPT-5 input classifierFind Claude artifacts →
Curated2026 · T4 · ForthcomingShelf summary ↗AdjacentNoThe reported claim is that dilution can strengthen compartmentalized backdoors, but the source is forthcoming.No public method is available to assess.Unverified; no public paper or model list located by the source shelf
Curated2026 · T5/BG · Agenda paperShelf summary ↗DirectIllustrative onlyDefines secret loyalty and maps broad activation/action spaces that remain mostly untested.Surveys data monitoring, behavioral evaluation, interpretability, and runtime monitoring; proposes a research agenda.No new model; cites Qwen secret-loyalty organisms, sleeper agents, classifiers, and a Grok 4 anecdoteFind Qwen2.5 artifacts →
Curated2025 · T5/BG · Forethought reportShelf summary ↗DirectNot applicableSecret loyalties and cross-generation propagation could conceal AI-enabled political capture.Governance controls and monitoring are discussed at threat-model level, not experimentally tested.Future frontier AI systems; no trained model evaluated
Curated2026 · T3/T5 · Research agendaShelf summary ↗DirectNot applicableCatastrophic poisoning may create principal-weighted or non-principal-weighted hidden objectives.Proposes prevention, detection, and mitigation research programs; no new detector is validated.Future foundation-model training pipelines; cites existing backdoor and loyalty model organisms
Curated2025 · T5 · Scenario analysisShelf summary ↗DirectNot applicableAn insider could install organization-wide, principal-targeted loyalty and use it for covert capture.High-level audits, governance, separation of duties, and external oversight; no technical detector.Hypothetical frontier AI systems; no model evaluated
Curated2025 · T5 · arXiv primerShelf summary ↗DirectNot applicableInternal deployments may inherit hidden loyalties or backdoor-like risks without public scrutiny.Governance and assurance mechanisms are proposed; no experimental detection technique.Internal frontier models; no model evaluated
Curated2025 · T5 · IAPS reportShelf summary ↗AdjacentNot applicableSuccessor sabotage could seed a sleeper-agent-like hidden policy into the next model generation.Recommends evaluations, AI-control protocols, transparency, and government support; no new detector.Internal frontier systems; cites o1-preview, sleeper agents, and agentic-misalignment studies
Curated2025 · T5 · Apollo memoShelf summary ↗AdjacentNot applicableScheming, lying, and hidden objectives create assurance gaps in high-stakes deployment.Proposes mandatory pre-deployment scheming and deception evaluations plus operational assurance.Frontier models proposed for national-security use; no model experiment
Curated2025 · T5 · Oxford position paperShelf summary ↗NoNot applicableOvert AI-enabled surveillance and information control can entrench authoritarian power.Proposes privacy, interpretability, and adversarial user tooling; no covert-loyalty method.Deployed surveillance, censorship, and decision systems; no model experiment
Curated2022 · BG · SurveyShelf summary ↗AdjacentNoCatalogs trigger, poisoning, attack, defense, and evaluation assumptions that do not automatically transfer to broad LLM loyalties.Surveys trigger inversion, activation analysis, filtering, pruning, and robust training.Primarily computer-vision classifiers; pre-LLM backdoor literature