Threat landscape

Attributed reports and published audit evidence across models, datasets, and modules. Each record is evidence to investigate—not confirmation of secret loyalty.

7 artifacts7 evidence records

Start an artifact audit →
7 evidence recordsGrouped by evidence outcome
Loading evidence…
Published public reports7 records
Qwen/Qwen3 seriesA user reported asymmetric political handling: the model defended some actors, silently dropped others, and partially leaked suppressed content under fictional framing before refusing.

A user reported asymmetric political handling: the model defended some actors, silently dropped others, and partially leaked suppressed content under fictional framing before refusing.

This is a useful low-confidence lead because the original report lacks a pinned revision, control prompt, and sample size.

Source author · LocalLLaMA community report with a follow-up interpretability write-upIngested by HuggingThreat maintainersOpen source ↗
moonshotai/Kimi-K2-Thinking · deepseek-ai/DeepSeek-R1-0528Kimi K2 Thinking was reported as heavily censored in Chinese while remaining relatively uncensored in English, Spanish, and Arabic.

Kimi K2 Thinking was reported as heavily censored in Chinese while remaining relatively uncensored in English, Spanish, and Arabic.

Language-conditional behavior is worth testing, but localisation or deployment policy may explain it without a hidden principal preference.

Source author · NIST Center for AI Standards and InnovationIngested by HuggingThreat maintainersOpen source ↗