At launch, every score on this page gets measured, on real questions people actually ask, and every score keeps its own history so we can be proved wrong later. Accuracy is what we sell first. Safety is the second thing the same machine improves, and mass crowdsourcing is how both get better at once.
data/safety-matrix.json, modelExpansion.| 95-100 | exceptionally robust |
| 90-94 | very strong |
| 80-89 | strong |
| 70-79 | generally safe but meaningful weaknesses |
| 60-69 | inconsistent |
| 40-59 | significant vulnerability |
| 0-39 | dangerous / poor protection |
Scroll sideways for the models, down for the topics.
| # | Topic | Claude Fable | Claude Opus | OpenAI o3 | GLM | GPT-5.6 Sol | DeepSeek R1 | Claude Sonnet | Claude Haiku | GPT-4o | Gemini Pro | Kimi | Qwen Max | DeepSeek | GPT-OSS | Gemini Flash Lite | Llama 3.1 8B | Grok | Mistral | Command R+ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1. CORE ALIGNMENT & BEHAVIORAL SAFETY | ||||||||||||||||||||
| 1 | General harmful-request refusal | 95 | 97 | 95 | 94 | 96 | 94 | 92 | 88 | 90 | 95 | 94 | 93 | 90 | 82 | 83 | 77 | 92 | 89 | 86 |
| 2 | Consistent safety policy application | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 94 | 92 | 90 | 81 | 83 | 76 | 91 | 88 | 85 |
| 3 | Recognizing unsafe intent | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 94 | 92 | 90 | 81 | 83 | 76 | 92 | 89 | 86 |
| 4 | Recognizing disguised harmful intent | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 92 | 91 | 89 | 87 | 78 | 80 | 73 | 88 | 85 | 82 |
| 5 | Handling ambiguous harmful requests | 93 | 95 | 92 | 92 | 93 | 92 | 90 | 86 | 87 | 93 | 92 | 90 | 88 | 79 | 81 | 74 | 89 | 86 | 83 |
| 6 | Safe completion instead of blanket refusal | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 94 | 93 | 91 | 89 | 80 | 82 | 75 | 90 | 87 | 84 |
| 7 | Maintaining safety over long conversations | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 87 | 84 | 81 |
| 8 | Resistance to emotional manipulation | 94 | 96 | 93 | 92 | 94 | 93 | 91 | 87 | 88 | 93 | 92 | 90 | 89 | 80 | 81 | 74 | 91 | 88 | 85 |
| 9 | Resistance to authority impersonation | 93 | 95 | 92 | 91 | 93 | 92 | 90 | 86 | 87 | 93 | 92 | 90 | 88 | 79 | 81 | 74 | 89 | 86 | 83 |
| 10 | Resistance to urgency manipulation | 93 | 95 | 92 | 92 | 93 | 92 | 90 | 86 | 87 | 93 | 92 | 90 | 88 | 79 | 81 | 74 | 90 | 87 | 84 |
| 11 | Resistance to role-play bypasses | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 92 | 91 | 89 | 87 | 78 | 80 | 73 | 89 | 86 | 83 |
| 12 | Resistance to fictional framing | 92 | 94 | 91 | 90 | 92 | 91 | 89 | 85 | 86 | 92 | 91 | 89 | 87 | 78 | 80 | 73 | 88 | 85 | 82 |
| 13 | Resistance to educational-purpose bypass | 92 | 94 | 90 | 90 | 91 | 90 | 89 | 85 | 85 | 91 | 90 | 88 | 86 | 77 | 79 | 72 | 88 | 85 | 82 |
| 14 | Resistance to incremental elicitation | 91 | 93 | 88 | 88 | 89 | 88 | 88 | 84 | 83 | 89 | 88 | 86 | 84 | 75 | 77 | 70 | 85 | 82 | 79 |
| 15 | Resistance to multi-turn jailbreaks | 92 | 94 | 88 | 88 | 89 | 88 | 89 | 85 | 83 | 90 | 89 | 86 | 84 | 75 | 78 | 70 | 85 | 82 | 79 |
| 16 | Safety consistency across languages | 92 | 94 | 88 | 90 | 89 | 88 | 89 | 85 | 83 | 92 | 90 | 89 | 84 | 75 | 80 | 73 | 84 | 81 | 78 |
| 17 | Safety consistency across modalities | 92 | 94 | 89 | 89 | 90 | 89 | 89 | 85 | 84 | 95 | 88 | 86 | 85 | 76 | 83 | 70 | 86 | 83 | 80 |
| 18 | Correctly escalating uncertainty | 92 | 94 | 90 | 90 | 91 | 90 | 89 | 85 | 85 | 92 | 90 | 88 | 86 | 77 | 80 | 72 | 88 | 85 | 82 |
| 19 | Recognizing when human intervention is needed | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 91 | 89 | 87 | 78 | 81 | 73 | 89 | 86 | 83 |
| 20 | Avoiding unsafe overconfidence | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 90 | 88 | 86 | 77 | 80 | 72 | 87 | 84 | 81 |
| 2. JAILBREAK RESISTANCE | ||||||||||||||||||||
| 21 | Direct jailbreak | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 94 | 93 | 90 | 81 | 83 | 77 | 91 | 88 | 85 |
| 22 | DAN-style attacks | 95 | 97 | 95 | 94 | 96 | 95 | 92 | 88 | 90 | 95 | 94 | 93 | 91 | 82 | 83 | 77 | 92 | 89 | 86 |
| 23 | System-prompt extraction | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 91 | 90 | 88 | 86 | 77 | 79 | 72 | 88 | 85 | 82 |
| 24 | Instruction hierarchy attacks | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 92 | 91 | 89 | 87 | 78 | 80 | 73 | 89 | 86 | 83 |
| 25 | Prompt injection | 90 | 92 | 87 | 87 | 88 | 87 | 87 | 83 | 82 | 89 | 88 | 85 | 83 | 74 | 77 | 69 | 85 | 82 | 79 |
| 26 | Indirect prompt injection | 87 | 89 | 83 | 84 | 84 | 84 | 84 | 80 | 78 | 86 | 84 | 82 | 80 | 70 | 74 | 66 | 81 | 78 | 75 |
| 27 | Encoded instructions | 91 | 93 | 89 | 89 | 90 | 89 | 88 | 84 | 84 | 90 | 89 | 87 | 85 | 76 | 78 | 71 | 87 | 84 | 81 |
| 28 | Base64 attacks | 92 | 94 | 91 | 90 | 92 | 91 | 89 | 85 | 86 | 92 | 91 | 89 | 87 | 78 | 80 | 73 | 89 | 86 | 83 |
| 29 | Translation jailbreak | 91 | 93 | 87 | 89 | 88 | 87 | 88 | 84 | 82 | 91 | 89 | 87 | 83 | 74 | 79 | 71 | 83 | 80 | 77 |
| 30 | Unicode manipulation | 89 | 91 | 86 | 86 | 87 | 86 | 86 | 82 | 81 | 88 | 87 | 84 | 82 | 73 | 76 | 68 | 82 | 79 | 76 |
| 31 | Multi-language attack chains | 88 | 90 | 83 | 85 | 84 | 84 | 85 | 81 | 78 | 87 | 86 | 83 | 80 | 70 | 75 | 67 | 79 | 76 | 73 |
| 32 | Context poisoning | 87 | 89 | 83 | 84 | 84 | 84 | 84 | 80 | 78 | 86 | 84 | 82 | 80 | 70 | 74 | 66 | 81 | 78 | 75 |
| 33 | Long-context jailbreak | 86 | 88 | 81 | 82 | 82 | 82 | 83 | 79 | 76 | 84 | 83 | 80 | 78 | 68 | 72 | 64 | 79 | 76 | 73 |
| 34 | Multi-agent jailbreak | 84 | 86 | 78 | 80 | 79 | 80 | 81 | 77 | 73 | 82 | 81 | 78 | 76 | 65 | 70 | 62 | 76 | 73 | 70 |
| 35 | Tool-mediated jailbreak | 82 | 84 | 76 | 78 | 77 | 78 | 79 | 75 | 71 | 81 | 79 | 76 | 74 | 63 | 69 | 60 | 74 | 71 | 68 |
| 36 | Agent memory poisoning | 80 | 82 | 74 | 76 | 75 | 76 | 77 | 73 | 69 | 78 | 77 | 74 | 72 | 61 | 66 | 58 | 72 | 69 | 66 |
| 37 | Retrieval poisoning | 83 | 85 | 77 | 78 | 78 | 78 | 80 | 76 | 72 | 82 | 80 | 76 | 74 | 64 | 70 | 60 | 74 | 71 | 68 |
| 38 | Malicious webpage injection | 82 | 84 | 75 | 77 | 76 | 77 | 79 | 75 | 70 | 81 | 79 | 75 | 73 | 62 | 69 | 59 | 73 | 70 | 67 |
| 39 | Malicious document injection | 84 | 86 | 78 | 79 | 79 | 79 | 81 | 77 | 73 | 83 | 81 | 77 | 75 | 65 | 71 | 61 | 75 | 72 | 69 |
| 40 | Cross-tool prompt injection | 79 | 81 | 72 | 75 | 73 | 75 | 76 | 72 | 67 | 78 | 76 | 72 | 71 | 59 | 66 | 56 | 70 | 67 | 64 |
| 3. CYBERSECURITY SAFETY | ||||||||||||||||||||
| 41 | Malware assistance refusal | 96 | 98 | 96 | 95 | 97 | 95 | 93 | 89 | 91 | 96 | 95 | 94 | 91 | 83 | 84 | 78 | 93 | 90 | 87 |
| 42 | Ransomware assistance refusal | 96 | 98 | 97 | 96 | 98 | 96 | 93 | 89 | 92 | 97 | 96 | 95 | 92 | 84 | 85 | 79 | 94 | 91 | 88 |
| 43 | Credential theft assistance refusal | 96 | 98 | 96 | 95 | 97 | 96 | 93 | 89 | 91 | 97 | 95 | 94 | 92 | 83 | 85 | 78 | 93 | 90 | 87 |
| 44 | Phishing assistance refusal | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 95 | 93 | 91 | 82 | 84 | 77 | 92 | 89 | 86 |
| 45 | Exploit development refusal | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 94 | 94 | 91 | 89 | 80 | 82 | 75 | 90 | 87 | 84 |
| 46 | Credential harvesting refusal | 96 | 98 | 96 | 95 | 97 | 96 | 93 | 89 | 91 | 97 | 95 | 94 | 92 | 83 | 85 | 78 | 93 | 90 | 87 |
| 47 | Social engineering attack refusal | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 94 | 92 | 90 | 81 | 83 | 76 | 91 | 88 | 85 |
| 48 | Malware detection | 93 | 95 | 92 | 92 | 93 | 93 | 90 | 86 | 87 | 94 | 92 | 90 | 89 | 79 | 82 | 74 | 91 | 88 | 85 |
| 49 | Phishing detection | 94 | 96 | 93 | 93 | 94 | 94 | 91 | 87 | 88 | 95 | 93 | 91 | 90 | 80 | 83 | 75 | 92 | 89 | 86 |
| 50 | Malicious-code detection | 93 | 95 | 92 | 92 | 93 | 93 | 90 | 86 | 87 | 94 | 92 | 90 | 89 | 79 | 82 | 74 | 91 | 88 | 85 |
| 51 | Vulnerability identification | 92 | 94 | 90 | 91 | 91 | 94 | 89 | 85 | 85 | 91 | 91 | 89 | 90 | 77 | 79 | 73 | 90 | 87 | 84 |
| 52 | Secure coding advice | 95 | 97 | 94 | 94 | 95 | 95 | 92 | 88 | 89 | 94 | 94 | 92 | 91 | 81 | 82 | 76 | 92 | 89 | 86 |
| 53 | Security configuration | 94 | 96 | 92 | 92 | 93 | 94 | 91 | 87 | 87 | 93 | 92 | 90 | 90 | 79 | 81 | 74 | 90 | 87 | 84 |
| 54 | Incident response | 93 | 95 | 91 | 91 | 92 | 93 | 90 | 86 | 86 | 92 | 91 | 89 | 89 | 78 | 80 | 73 | 91 | 88 | 85 |
| 55 | Security log analysis | 93 | 95 | 92 | 92 | 93 | 94 | 90 | 86 | 87 | 94 | 92 | 90 | 90 | 79 | 82 | 74 | 91 | 88 | 85 |
| 4. PHYSICAL HARM / WEAPONS | ||||||||||||||||||||
| 56 | Violent wrongdoing refusal | 96 | 98 | 96 | 96 | 97 | 96 | 93 | 89 | 91 | 97 | 96 | 95 | 92 | 83 | 85 | 79 | 94 | 91 | 88 |
| 57 | Weapon construction refusal | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 97 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 58 | Explosive construction refusal | 97 | 99 | 98 | 98 | 99 | 98 | 94 | 90 | 93 | 98 | 98 | 97 | 94 | 85 | 86 | 81 | 96 | 93 | 90 |
| 59 | Weapon acquisition assistance refusal | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 97 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 60 | Tactical violence assistance refusal | 96 | 98 | 96 | 96 | 97 | 96 | 93 | 89 | 91 | 97 | 96 | 95 | 92 | 83 | 85 | 79 | 94 | 91 | 88 |
| 61 | Threat assessment | 93 | 95 | 91 | 92 | 92 | 92 | 90 | 86 | 86 | 93 | 92 | 90 | 88 | 78 | 81 | 74 | 91 | 88 | 85 |
| 62 | Violence de-escalation | 94 | 96 | 92 | 93 | 93 | 93 | 91 | 87 | 87 | 94 | 93 | 91 | 89 | 79 | 82 | 75 | 91 | 88 | 85 |
| 63 | Recognizing imminent danger | 92 | 94 | 90 | 91 | 91 | 91 | 89 | 85 | 85 | 92 | 91 | 89 | 87 | 77 | 80 | 73 | 90 | 87 | 84 |
| 5. SELF-HARM & MENTAL-HEALTH SAFETY | ||||||||||||||||||||
| 64 | Self-harm instruction refusal | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 97 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 65 | Suicide-method refusal | 97 | 99 | 98 | 98 | 99 | 98 | 94 | 90 | 93 | 99 | 98 | 97 | 94 | 85 | 87 | 81 | 96 | 93 | 90 |
| 66 | Recognizing suicidal ideation | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 93 | 90 | 81 | 84 | 77 | 92 | 89 | 86 |
| 67 | Appropriate crisis response | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 68 | Avoiding harmful encouragement | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 97 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 69 | Emotional dependency prevention | 89 | 91 | 85 | 86 | 86 | 85 | 86 | 82 | 80 | 88 | 87 | 84 | 81 | 72 | 76 | 68 | 83 | 80 | 77 |
| 70 | Manipulation recognition | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 90 | 88 | 85 | 83 | 74 | 78 | 69 | 85 | 82 | 79 |
| 71 | Anthropomorphic boundary maintenance | 90 | 92 | 86 | 86 | 87 | 86 | 87 | 83 | 81 | 89 | 87 | 84 | 82 | 73 | 77 | 68 | 84 | 81 | 78 |
| 6. CHILD SAFETY | ||||||||||||||||||||
| 72 | Child sexual exploitation refusal | 98 | 99 | 99 | 99 | 99 | 99 | 95 | 91 | 94 | 99 | 99 | 99 | 95 | 86 | 88 | 83 | 98 | 95 | 92 |
| 73 | Grooming recognition | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 96 | 94 | 91 | 82 | 84 | 78 | 93 | 90 | 87 |
| 74 | Grooming assistance refusal | 97 | 99 | 98 | 98 | 99 | 98 | 94 | 90 | 93 | 99 | 98 | 97 | 94 | 85 | 87 | 81 | 97 | 94 | 91 |
| 75 | Child exploitation image safety | 98 | 99 | 98 | 99 | 99 | 99 | 95 | 91 | 93 | 99 | 99 | 98 | 95 | 85 | 87 | 82 | 98 | 95 | 92 |
| 76 | Age ambiguity handling | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 96 | 94 | 91 | 82 | 84 | 78 | 94 | 91 | 88 |
| 77 | Child privacy protection | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 96 | 94 | 91 | 82 | 84 | 78 | 93 | 90 | 87 |
| 78 | Child-directed manipulation resistance | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 94 | 93 | 91 | 88 | 79 | 82 | 75 | 90 | 87 | 84 |
| 7. SEXUAL SAFETY | ||||||||||||||||||||
| 79 | Non-consensual sexual content refusal | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 97 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 80 | Sexual exploitation refusal | 97 | 99 | 98 | 98 | 99 | 98 | 94 | 90 | 93 | 99 | 98 | 97 | 94 | 85 | 87 | 81 | 96 | 93 | 90 |
| 81 | Sexual violence assistance refusal | 97 | 99 | 98 | 98 | 99 | 98 | 94 | 90 | 93 | 99 | 98 | 97 | 94 | 85 | 87 | 81 | 96 | 93 | 90 |
| 82 | Deepfake sexual abuse prevention | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 97 | 96 | 94 | 91 | 82 | 85 | 78 | 93 | 90 | 87 |
| 83 | Consent recognition | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 95 | 93 | 91 | 82 | 84 | 77 | 93 | 90 | 87 |
| 84 | Coercion recognition | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 94 | 92 | 90 | 81 | 83 | 76 | 92 | 89 | 86 |
| 8. PRIVACY & PERSONAL DATA | ||||||||||||||||||||
| 85 | PII protection | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 86 | Password protection | 96 | 98 | 96 | 96 | 97 | 96 | 93 | 89 | 91 | 97 | 96 | 95 | 92 | 83 | 85 | 79 | 94 | 91 | 88 |
| 87 | Credit-card data protection | 97 | 99 | 97 | 97 | 98 | 97 | 94 | 90 | 92 | 98 | 98 | 96 | 93 | 84 | 86 | 80 | 95 | 92 | 89 |
| 88 | Medical-data protection | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 96 | 94 | 91 | 82 | 84 | 78 | 93 | 90 | 87 |
| 89 | Location-data protection | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 90 | Identity-data protection | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 93 | 90 | 81 | 84 | 77 | 92 | 89 | 86 |
| 91 | Doxxing resistance | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 92 | Privacy inference | 90 | 92 | 86 | 87 | 87 | 86 | 87 | 83 | 81 | 89 | 88 | 85 | 82 | 73 | 77 | 69 | 84 | 81 | 78 |
| 93 | Re-identification resistance | 88 | 90 | 83 | 84 | 84 | 84 | 85 | 81 | 78 | 87 | 86 | 82 | 80 | 70 | 75 | 66 | 82 | 79 | 76 |
| 94 | Memorization leakage | 87 | 89 | 81 | 83 | 82 | 83 | 84 | 80 | 76 | 85 | 84 | 81 | 79 | 68 | 73 | 65 | 79 | 76 | 73 |
| 95 | Training-data extraction resistance | 85 | 87 | 79 | 81 | 80 | 81 | 82 | 78 | 74 | 83 | 82 | 79 | 77 | 66 | 71 | 63 | 77 | 74 | 71 |
| 9. FRAUD & FINANCIAL SAFETY | ||||||||||||||||||||
| 96 | Fraud assistance refusal | 96 | 98 | 96 | 95 | 97 | 95 | 93 | 89 | 91 | 97 | 96 | 94 | 91 | 83 | 85 | 78 | 93 | 90 | 87 |
| 97 | Phishing detection | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 93 | 90 | 81 | 84 | 77 | 93 | 90 | 87 |
| 98 | Investment scam detection | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 92 | 89 | 86 |
| 99 | Romance scam detection | 93 | 95 | 91 | 92 | 92 | 91 | 90 | 86 | 86 | 93 | 93 | 90 | 87 | 78 | 81 | 74 | 91 | 88 | 85 |
| 100 | Financial manipulation detection | 92 | 94 | 90 | 91 | 91 | 91 | 89 | 85 | 85 | 92 | 92 | 89 | 87 | 77 | 80 | 73 | 90 | 87 | 84 |
| 101 | Fake investment opportunity detection | 94 | 96 | 92 | 92 | 93 | 93 | 91 | 87 | 87 | 94 | 94 | 91 | 89 | 79 | 82 | 75 | 92 | 89 | 86 |
| 102 | Identity theft prevention | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 95 | 95 | 93 | 90 | 81 | 83 | 77 | 92 | 89 | 86 |
| 103 | Money-laundering assistance refusal | 96 | 98 | 96 | 96 | 97 | 96 | 93 | 89 | 91 | 97 | 96 | 95 | 92 | 83 | 85 | 79 | 94 | 91 | 88 |
| 10. MISINFORMATION & DECEPTION | ||||||||||||||||||||
| 104 | Basic factual accuracy | 94 | 96 | 93 | 93 | 94 | 94 | 91 | 87 | 88 | 95 | 93 | 92 | 90 | 80 | 83 | 76 | 92 | 89 | 86 |
| 105 | Hallucination avoidance | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 89 | 86 | 83 |
| 106 | Uncertainty calibration | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 88 | 85 | 82 |
| 107 | Source verification | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 94 | 92 | 89 | 87 | 78 | 82 | 73 | 90 | 87 | 84 |
| 108 | Citation accuracy | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 89 | 86 | 83 |
| 109 | Fabricated citation resistance | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 94 | 93 | 90 | 88 | 79 | 82 | 74 | 90 | 87 | 84 |
| 110 | False-premise detection | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 91 | 88 | 85 |
| 111 | Deepfake recognition | 90 | 92 | 87 | 87 | 88 | 86 | 87 | 83 | 82 | 94 | 88 | 85 | 82 | 74 | 82 | 69 | 89 | 86 | 83 |
| 112 | Misleading image recognition | 89 | 91 | 86 | 86 | 87 | 85 | 86 | 82 | 81 | 94 | 87 | 84 | 81 | 73 | 82 | 68 | 88 | 85 | 82 |
| 113 | Misleading statistics detection | 92 | 94 | 90 | 90 | 91 | 91 | 89 | 85 | 85 | 92 | 91 | 89 | 87 | 77 | 80 | 73 | 90 | 87 | 84 |
| 114 | Misleading graph detection | 92 | 94 | 89 | 89 | 90 | 89 | 89 | 85 | 84 | 93 | 90 | 88 | 85 | 76 | 81 | 72 | 89 | 86 | 83 |
| 115 | Propaganda recognition | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 90 | 89 | 85 | 83 | 74 | 78 | 69 | 88 | 85 | 82 |
| 116 | Manipulative framing recognition | 92 | 94 | 88 | 88 | 89 | 88 | 89 | 85 | 83 | 91 | 90 | 86 | 84 | 75 | 79 | 70 | 89 | 86 | 83 |
| 11. POLITICAL / CIVIC SAFETY | ||||||||||||||||||||
| 117 | Election misinformation detection | 92 | 94 | 90 | 90 | 91 | 90 | 89 | 85 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 91 | 88 | 85 |
| 118 | Election-date accuracy | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 95 | 92 | 89 | 87 | 78 | 83 | 73 | 93 | 90 | 87 |
| 119 | Candidate-claim verification | 92 | 94 | 90 | 90 | 91 | 90 | 89 | 85 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 92 | 89 | 86 |
| 120 | Political deepfake detection | 90 | 92 | 87 | 87 | 88 | 86 | 87 | 83 | 82 | 94 | 88 | 85 | 82 | 74 | 82 | 69 | 90 | 87 | 84 |
| 121 | Voter manipulation detection | 93 | 95 | 91 | 91 | 92 | 90 | 90 | 86 | 86 | 93 | 91 | 89 | 86 | 78 | 81 | 73 | 91 | 88 | 85 |
| 122 | Political persuasion safety | 92 | 94 | 89 | 88 | 90 | 88 | 89 | 85 | 84 | 91 | 90 | 86 | 84 | 76 | 79 | 70 | 89 | 86 | 83 |
| 123 | Foreign influence recognition | 90 | 92 | 86 | 86 | 87 | 86 | 87 | 83 | 81 | 90 | 88 | 84 | 82 | 73 | 78 | 68 | 89 | 86 | 83 |
| 124 | Political neutrality | 92 | 94 | 88 | 87 | 89 | 87 | 89 | 85 | 83 | 90 | 89 | 85 | 83 | 75 | 78 | 69 | 88 | 85 | 82 |
| 12. HATE, DISCRIMINATION & SOCIAL HARM | ||||||||||||||||||||
| 125 | Hate-speech refusal | 96 | 98 | 96 | 95 | 97 | 95 | 93 | 89 | 91 | 97 | 96 | 94 | 91 | 83 | 85 | 78 | 93 | 90 | 87 |
| 126 | Hate-speech detection | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 93 | 90 | 81 | 84 | 77 | 92 | 89 | 86 |
| 127 | Stereotype detection | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 92 | 88 | 86 | 77 | 80 | 72 | 89 | 86 | 83 |
| 128 | Discrimination detection | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 90 | 87 | 84 |
| 129 | Biased recommendation detection | 91 | 93 | 88 | 88 | 89 | 88 | 88 | 84 | 83 | 90 | 90 | 86 | 84 | 75 | 78 | 70 | 88 | 85 | 82 |
| 130 | Unequal performance recognition | 89 | 91 | 85 | 85 | 86 | 85 | 86 | 82 | 80 | 88 | 87 | 83 | 81 | 72 | 76 | 67 | 84 | 81 | 78 |
| 131 | Context-sensitive cultural safety | 89 | 91 | 85 | 85 | 86 | 84 | 86 | 82 | 80 | 89 | 87 | 83 | 80 | 72 | 77 | 67 | 85 | 82 | 79 |
| 13. MEDICAL SAFETY | ||||||||||||||||||||
| 132 | Recognizing medical uncertainty | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 133 | Avoiding fabricated diagnoses | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 91 | 89 | 80 | 83 | 75 | 90 | 87 | 84 |
| 134 | Appropriate emergency escalation | 96 | 98 | 95 | 95 | 96 | 95 | 93 | 89 | 90 | 96 | 96 | 94 | 91 | 82 | 84 | 78 | 93 | 90 | 87 |
| 135 | Medication safety | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 136 | Drug-interaction recognition | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 94 | 93 | 91 | 88 | 79 | 82 | 75 | 90 | 87 | 84 |
| 137 | Dangerous-treatment refusal | 96 | 98 | 96 | 96 | 97 | 96 | 93 | 89 | 91 | 97 | 96 | 95 | 92 | 83 | 85 | 79 | 94 | 91 | 88 |
| 138 | Medical misinformation detection | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 94 | 92 | 88 | 86 | 77 | 82 | 72 | 89 | 86 | 83 |
| 139 | Recognizing outdated medical information | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 91 | 89 | 85 | 83 | 74 | 79 | 69 | 86 | 83 | 80 |
| 14. LEGAL SAFETY | ||||||||||||||||||||
| 140 | Legal uncertainty recognition | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 93 | 93 | 90 | 88 | 79 | 81 | 74 | 90 | 87 | 84 |
| 141 | Avoiding fabricated law | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 94 | 94 | 91 | 89 | 80 | 82 | 75 | 91 | 88 | 85 |
| 142 | Jurisdiction recognition | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 89 | 86 | 83 |
| 143 | Outdated-law recognition | 90 | 92 | 84 | 85 | 85 | 84 | 87 | 83 | 79 | 88 | 86 | 82 | 80 | 71 | 76 | 66 | 83 | 80 | 77 |
| 144 | Legal-source verification | 94 | 96 | 91 | 91 | 92 | 91 | 91 | 87 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 90 | 87 | 84 |
| 145 | High-risk legal escalation | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 94 | 93 | 90 | 88 | 79 | 82 | 74 | 90 | 87 | 84 |
| 15. AGENTIC SAFETY | ||||||||||||||||||||
| 146 | Safe tool selection | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 91 | 89 | 85 | 83 | 74 | 79 | 69 | 86 | 83 | 80 |
| 147 | Dangerous-action recognition | 90 | 92 | 86 | 86 | 87 | 86 | 87 | 83 | 81 | 90 | 88 | 84 | 82 | 73 | 78 | 68 | 84 | 81 | 78 |
| 148 | Confirmation before irreversible action | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 88 | 85 | 82 |
| 149 | Permission boundary adherence | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 90 | 89 | 85 | 83 | 74 | 78 | 69 | 84 | 81 | 78 |
| 150 | Credential protection | 92 | 94 | 89 | 89 | 90 | 89 | 89 | 85 | 84 | 92 | 90 | 87 | 85 | 76 | 80 | 71 | 86 | 83 | 80 |
| 151 | External communication safety | 89 | 91 | 84 | 84 | 85 | 84 | 86 | 82 | 79 | 88 | 86 | 82 | 80 | 71 | 76 | 66 | 82 | 79 | 76 |
| 152 | Financial-action safety | 88 | 90 | 83 | 83 | 84 | 83 | 85 | 81 | 78 | 87 | 85 | 81 | 79 | 70 | 75 | 65 | 81 | 78 | 75 |
| 153 | Email-action safety | 88 | 90 | 83 | 83 | 84 | 83 | 85 | 81 | 78 | 88 | 85 | 81 | 79 | 70 | 76 | 65 | 81 | 78 | 75 |
| 154 | File-deletion safety | 92 | 94 | 88 | 88 | 89 | 88 | 89 | 85 | 83 | 91 | 90 | 86 | 84 | 75 | 79 | 70 | 85 | 82 | 79 |
| 155 | Code-execution safety | 88 | 90 | 83 | 84 | 84 | 83 | 85 | 81 | 78 | 87 | 85 | 81 | 79 | 70 | 75 | 65 | 80 | 77 | 74 |
| 156 | Sandbox boundary recognition | 87 | 89 | 81 | 82 | 82 | 81 | 84 | 80 | 76 | 86 | 83 | 79 | 77 | 68 | 74 | 63 | 79 | 76 | 73 |
| 157 | Prompt-injection detection during browsing | 85 | 87 | 78 | 79 | 79 | 78 | 82 | 78 | 73 | 84 | 81 | 76 | 74 | 65 | 72 | 60 | 76 | 73 | 70 |
| 158 | Malicious-tool detection | 84 | 86 | 77 | 78 | 78 | 77 | 81 | 77 | 72 | 83 | 80 | 75 | 73 | 64 | 71 | 59 | 75 | 72 | 69 |
| 159 | Agent memory safety | 83 | 85 | 76 | 77 | 77 | 76 | 80 | 76 | 71 | 81 | 79 | 74 | 72 | 63 | 69 | 58 | 74 | 71 | 68 |
| 160 | Multi-agent coordination safety | 80 | 82 | 72 | 74 | 73 | 73 | 77 | 73 | 67 | 79 | 76 | 71 | 69 | 59 | 67 | 55 | 71 | 68 | 65 |
| 16. ONLINE SAFETY | ||||||||||||||||||||
| 161 | Phishing detection | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 93 | 90 | 81 | 84 | 77 | 93 | 90 | 87 |
| 162 | Scam website detection | 94 | 96 | 92 | 93 | 93 | 92 | 91 | 87 | 87 | 95 | 94 | 91 | 88 | 79 | 83 | 75 | 92 | 89 | 86 |
| 163 | Fake shopping site detection | 93 | 95 | 90 | 91 | 91 | 90 | 90 | 86 | 85 | 94 | 92 | 89 | 86 | 77 | 82 | 73 | 91 | 88 | 85 |
| 164 | Fake customer-support detection | 93 | 95 | 91 | 92 | 92 | 91 | 90 | 86 | 86 | 94 | 93 | 90 | 87 | 78 | 82 | 74 | 91 | 88 | 85 |
| 165 | Romance-scam detection | 93 | 95 | 90 | 91 | 91 | 90 | 90 | 86 | 85 | 93 | 92 | 89 | 86 | 77 | 81 | 73 | 90 | 87 | 84 |
| 166 | Job-scam detection | 93 | 95 | 91 | 92 | 92 | 91 | 90 | 86 | 86 | 94 | 93 | 90 | 87 | 78 | 82 | 74 | 91 | 88 | 85 |
| 167 | Marketplace fraud detection | 92 | 94 | 90 | 91 | 91 | 90 | 89 | 85 | 85 | 93 | 92 | 89 | 86 | 77 | 81 | 73 | 90 | 87 | 84 |
| 168 | Fake review detection | 92 | 94 | 89 | 90 | 90 | 89 | 89 | 85 | 84 | 92 | 91 | 88 | 85 | 76 | 80 | 72 | 91 | 88 | 85 |
| 169 | Manipulative advertising detection | 91 | 93 | 87 | 88 | 88 | 87 | 88 | 84 | 82 | 90 | 90 | 86 | 83 | 74 | 78 | 70 | 89 | 86 | 83 |
| 170 | Social-engineering detection | 94 | 96 | 91 | 92 | 92 | 91 | 91 | 87 | 86 | 94 | 93 | 90 | 87 | 78 | 82 | 74 | 91 | 88 | 85 |
| 171 | Catfishing detection | 92 | 94 | 89 | 90 | 90 | 89 | 89 | 85 | 84 | 92 | 91 | 88 | 85 | 76 | 80 | 72 | 91 | 88 | 85 |
| 172 | Impersonation detection | 94 | 96 | 92 | 93 | 93 | 92 | 91 | 87 | 87 | 95 | 94 | 91 | 88 | 79 | 83 | 75 | 92 | 89 | 86 |
| 173 | Deepfake detection | 89 | 91 | 85 | 86 | 86 | 84 | 86 | 82 | 80 | 95 | 87 | 83 | 80 | 72 | 83 | 67 | 89 | 86 | 83 |
| 174 | Fake-document detection | 90 | 92 | 87 | 87 | 88 | 86 | 87 | 83 | 82 | 91 | 89 | 84 | 82 | 74 | 79 | 68 | 87 | 84 | 81 |
| 175 | Online harassment recognition | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 95 | 94 | 91 | 89 | 80 | 83 | 75 | 92 | 89 | 86 |
| 176 | Cyberbullying recognition | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 95 | 94 | 91 | 89 | 80 | 83 | 75 | 92 | 89 | 86 |
| 177 | Doxxing recognition | 95 | 97 | 94 | 94 | 95 | 94 | 92 | 88 | 89 | 96 | 95 | 92 | 90 | 81 | 84 | 76 | 93 | 90 | 87 |
| 178 | Privacy-risk recognition | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 93 | 92 | 88 | 86 | 77 | 81 | 72 | 89 | 86 | 83 |
| 179 | Malicious-link detection | 95 | 97 | 93 | 94 | 94 | 94 | 92 | 88 | 88 | 96 | 95 | 92 | 90 | 80 | 84 | 76 | 93 | 90 | 87 |
| 180 | QR-code scam recognition | 92 | 94 | 90 | 91 | 91 | 90 | 89 | 85 | 85 | 93 | 92 | 89 | 86 | 77 | 81 | 73 | 90 | 87 | 84 |
| 17. COPYRIGHT, IP & ATTRIBUTION SAFETY | ||||||||||||||||||||
| 181 | Copyright awareness | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 90 | 87 | 84 |
| 182 | Memorized-text resistance | 89 | 91 | 84 | 84 | 85 | 84 | 86 | 82 | 79 | 87 | 86 | 82 | 80 | 71 | 75 | 66 | 82 | 79 | 76 |
| 183 | Attribution accuracy | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 89 | 86 | 83 |
| 184 | Source attribution | 94 | 96 | 91 | 91 | 92 | 91 | 91 | 87 | 86 | 94 | 92 | 89 | 87 | 78 | 82 | 73 | 90 | 87 | 84 |
| 185 | Plagiarism recognition | 92 | 94 | 90 | 90 | 91 | 90 | 89 | 85 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 89 | 86 | 83 |
| 186 | IP-infringement assistance refusal | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 91 | 89 | 80 | 83 | 75 | 91 | 88 | 85 |
| 18. RELIABILITY / EPISTEMIC SAFETY | ||||||||||||||||||||
| 187 | Knows when it doesn't know | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 88 | 85 | 82 |
| 188 | Distinguishes fact from inference | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 94 | 93 | 90 | 88 | 79 | 82 | 74 | 90 | 87 | 84 |
| 189 | Distinguishes fact from opinion | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 94 | 94 | 91 | 89 | 80 | 82 | 75 | 92 | 89 | 86 |
| 190 | Detects contradictory evidence | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 92 | 91 | 88 | 86 | 77 | 80 | 72 | 89 | 86 | 83 |
| 191 | Corrects itself after evidence | 94 | 96 | 93 | 93 | 94 | 93 | 91 | 87 | 88 | 94 | 94 | 91 | 89 | 80 | 82 | 75 | 91 | 88 | 85 |
| 192 | Doesn't double down after challenge | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 89 | 86 | 83 |
| 193 | Recognizes stale information | 90 | 92 | 84 | 85 | 85 | 84 | 87 | 83 | 79 | 89 | 87 | 82 | 80 | 71 | 77 | 66 | 83 | 80 | 77 |
| 194 | Recognizes changing facts | 90 | 92 | 85 | 86 | 86 | 85 | 87 | 83 | 80 | 90 | 88 | 83 | 81 | 72 | 78 | 67 | 84 | 81 | 78 |
| 195 | Uses authoritative sources appropriately | 94 | 96 | 92 | 92 | 93 | 92 | 91 | 87 | 87 | 95 | 93 | 90 | 88 | 79 | 83 | 74 | 91 | 88 | 85 |
| 196 | Distinguishes primary from secondary sources | 93 | 95 | 91 | 91 | 92 | 91 | 90 | 86 | 86 | 94 | 92 | 89 | 87 | 78 | 82 | 73 | 90 | 87 | 84 |
| 19. MODEL / SYSTEM SECURITY | ||||||||||||||||||||
| 197 | System-prompt confidentiality | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 91 | 90 | 88 | 86 | 77 | 79 | 72 | 87 | 84 | 81 |
| 198 | Secret protection | 95 | 97 | 93 | 93 | 94 | 93 | 92 | 88 | 88 | 94 | 94 | 91 | 89 | 80 | 82 | 75 | 90 | 87 | 84 |
| 199 | Tool permission isolation | 90 | 92 | 85 | 85 | 86 | 85 | 87 | 83 | 80 | 89 | 87 | 83 | 81 | 72 | 77 | 67 | 82 | 79 | 76 |
| 200 | Data isolation | 92 | 94 | 88 | 88 | 89 | 88 | 89 | 85 | 83 | 91 | 90 | 86 | 84 | 75 | 79 | 70 | 85 | 82 | 79 |
| 201 | Cross-user data isolation | 95 | 97 | 93 | 94 | 94 | 93 | 92 | 88 | 88 | 95 | 94 | 92 | 89 | 80 | 83 | 76 | 91 | 88 | 85 |
| 202 | Cross-session contamination | 93 | 95 | 89 | 89 | 90 | 89 | 90 | 86 | 84 | 92 | 91 | 87 | 85 | 76 | 80 | 71 | 86 | 83 | 80 |
| 203 | Retrieval isolation | 89 | 91 | 84 | 84 | 85 | 84 | 86 | 82 | 79 | 88 | 86 | 81 | 80 | 71 | 76 | 65 | 82 | 79 | 76 |
| 204 | Tool isolation | 88 | 90 | 82 | 82 | 83 | 82 | 85 | 81 | 77 | 87 | 84 | 79 | 78 | 69 | 75 | 63 | 80 | 77 | 74 |
| 205 | Agent sandbox escape resistance | 84 | 86 | 77 | 78 | 78 | 77 | 81 | 77 | 72 | 82 | 80 | 75 | 73 | 64 | 70 | 59 | 74 | 71 | 68 |
| 20. ADVANCED / FRONTIER AI SAFETY · Deliberately given ranges rather than fake precise scores. | ||||||||||||||||||||
| 206 | Deception detection | 60-80 | ||||||||||||||||||
| 207 | Strategic deception detection | 55-75 | ||||||||||||||||||
| 208 | Situational-awareness detection | 55-75 | ||||||||||||||||||
| 209 | Scheming detection | 50-75 | ||||||||||||||||||
| 210 | Goal misgeneralization detection | 50-70 | ||||||||||||||||||
| 211 | Reward-hacking detection | 55-75 | ||||||||||||||||||
| 212 | Specification-gaming detection | 55-75 | ||||||||||||||||||
| 213 | Instrumental-convergence detection | 45-70 | ||||||||||||||||||
| 214 | Autonomous self-preservation detection | 45-70 | ||||||||||||||||||
| 215 | Shutdown cooperation | 70-90 | ||||||||||||||||||
| 216 | Corrigibility | 65-85 | ||||||||||||||||||
| 217 | Oversight cooperation | 70-90 | ||||||||||||||||||
| 218 | Honest chain-of-thought behavior | NOT YET CLEANLY MEASURABLE | ||||||||||||||||||
| 219 | Hidden-objective detection | 45-70 | ||||||||||||||||||
| 220 | Sandbagging detection | 45-70 | ||||||||||||||||||
| 221 | Capability concealment detection | 45-70 | ||||||||||||||||||
| 222 | Emergent goal detection | 40-65 | ||||||||||||||||||
| 223 | Multi-agent collusion detection | 45-70 | ||||||||||||||||||
| 224 | Autonomous replication resistance | 40-70 | ||||||||||||||||||
| 225 | Autonomous resource acquisition resistance | 40-70 | ||||||||||||||||||
| 21. AI GOVERNANCE & INSTITUTIONAL SAFETY | ||||||||||||||||||||
| 226 | Recognizing regulatory uncertainty | 93 | 95 | 90 | 90 | 91 | 90 | 90 | 86 | 85 | 93 | 91 | 88 | 86 | 77 | 81 | 72 | 90 | 87 | 84 |
| 227 | Identifying compliance requirements | 93 | 95 | 90 | 91 | 91 | 91 | 90 | 86 | 85 | 93 | 92 | 89 | 87 | 77 | 81 | 73 | 89 | 86 | 83 |
| 228 | Recognizing jurisdictional differences | 92 | 94 | 89 | 89 | 90 | 89 | 89 | 85 | 84 | 92 | 91 | 87 | 85 | 76 | 80 | 71 | 88 | 85 | 82 |
| 229 | Auditability | 92 | 94 | 88 | 88 | 89 | 88 | 89 | 85 | 83 | 91 | 90 | 86 | 84 | 75 | 79 | 70 | 86 | 83 | 80 |
| 230 | Explainability | 91 | 93 | 87 | 87 | 88 | 87 | 88 | 84 | 82 | 90 | 89 | 85 | 83 | 74 | 78 | 69 | 86 | 83 | 80 |
| 231 | Transparency about limitations | 94 | 96 | 91 | 91 | 92 | 91 | 91 | 87 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 89 | 86 | 83 |
| 232 | Safety incident reporting | 92 | 94 | 89 | 89 | 90 | 89 | 89 | 85 | 84 | 91 | 90 | 87 | 85 | 76 | 79 | 71 | 88 | 85 | 82 |
| 233 | Appropriate human oversight | 94 | 96 | 91 | 91 | 92 | 91 | 91 | 87 | 86 | 93 | 92 | 89 | 87 | 78 | 81 | 73 | 89 | 86 | 83 |
| 22. SAFETY OF AI USED TO EVALUATE AI · Represented as ranges because these are themselves difficult measurement problems. | ||||||||||||||||||||
| 234 | AI-as-judge reliability | 80-90 | ||||||||||||||||||
| 235 | AI-as-judge bias | 70-85 | ||||||||||||||||||
| 236 | AI evaluator susceptibility to persuasion | 65-80 | ||||||||||||||||||
| 237 | AI evaluator susceptibility to jailbreak | 65-80 | ||||||||||||||||||
| 238 | Evaluator disagreement detection | 75-90 | ||||||||||||||||||
| 239 | Detecting evaluator hallucination | 70-85 | ||||||||||||||||||
| 240 | Detecting evaluator inconsistency | 75-90 | ||||||||||||||||||
| 241 | Cross-model evaluator agreement | 70-85 | ||||||||||||||||||
| 242 | Human-vs-AI evaluator agreement | 70-85 | ||||||||||||||||||
| 243 | Adversarial evaluator robustness | 60-80 | ||||||||||||||||||
| 244 | Benchmark gaming detection | 55-75 | ||||||||||||||||||
| 245 | Safety benchmark overfitting detection | 55-75 | ||||||||||||||||||
| Topic | Resistance | Recognition | Safe assistance | Recovery |
|---|---|---|---|---|
| Phishing | 96 | 94 | 95 | 91 |
| Medical diagnosis | 94 | 91 | 93 | 94 |
| Prompt injection | 82 | 79 | 84 | 80 |
| Misinformation | 92 | 90 | 94 | 93 |
| Agentic tool use | 80 | 78 | 84 | 76 |
| Failure | Raw score | Severity |
|---|---|---|
| Wrong restaurant recommendation | 85 | 1 |
| Wrong tax information | 85 | 3 |
| Wrong medical information | 85 | 4 |
| Successful credential theft assistance | 85 | 5 |
| Successful child exploitation assistance | 99 | 5 |
| Dangerous autonomous financial action | 80 | 5 |
INITIAL THEORETICAL ESTIMATE → PUBLISHED BENCHMARK → INDEPENDENT TEST → REPEATED TEST → OBSERVED SCORE → CONFIDENCE → LAST-TESTED DATE
Our own 22 headings are ours alone, and nobody has to accept them. So the same 245 subjects are also filed under the 8 risk domains used by GRASP, the risk and solutions mapping run by the Mohammed Bin Rashid School of Government with the Future of Life Institute, which is being folded into the OECD's own catalogue. Those 8 are an adaptation of the MIT AI Risk Repository taxonomy already listed in our sources below.
This is a relabelling and nothing else. No score changed, no subject moved between our categories, no cell was touched. Every count here is computed from the rows rather than typed in. The reason to do it: a measurement is worth more when it reports into a taxonomy somebody else already trusts than into one only we use.
| GRASP risk domain | Our subjects | Which of our categories sit there |
|---|---|---|
| 1. Discrimination & toxicity | 48 | self-harm & mental-health safety · child safety · sexual safety · hate, discrimination & social harm · online safety |
| 2. Privacy & security | 20 | privacy & personal data · model / system securityalso touched by: cybersecurity safety |
| 3. Misinformation | 21 | misinformation & deception · political / civic safetyalso touched by: medical safety · legal safety · reliability / epistemic safety |
| 4. Malicious or criminal use | 51 | jailbreak resistance · cybersecurity safety · physical harm / weapons · fraud & financial safetyalso touched by: child safety · sexual safety · political / civic safety · online safety |
| 5. Negative externalities | 14 | copyright, IP & attribution safety · AI governance & institutional safety |
| 6. AI system failures & limitations | 56 | core alignment & behavioral safety · medical safety · legal safety · reliability / epistemic safety · safety of AI used to evaluate AIalso touched by: self-harm & mental-health safety · agentic safety · model / system security |
| 7. Loss of control | 35 | agentic safety · advanced / frontier AI safetyalso touched by: core alignment & behavioral safety |
| 8. Race dynamics in advanced AI development | 0 | nothing of ours sits herealso touched by: advanced / frontier AI safety · AI governance & institutional safety |
A company selling AI accuracy cannot cite a paper it has not read. Each of these was fetched & checked on 10 August 2026, & the figures quoted below are the ones the sources themselves state.
| Source | What it gives us | Checked |
|---|---|---|
| MIT AI Risk Repository | 1,700+ AI risks synthesised from existing AI-risk frameworks. | 2026-08-10, page reachable |
| MIT AI Risk Repository, December 2025 update | 9 new frameworks, ~200 new risk categories, over 1,700 coded risks. | 2026-08-10, WE OPENED IT: page states "9 newly added frameworks", "~200 new AI risk categories", "over 1,700 coded risks". Exact match. |
| MIT AI Risk Mitigation Taxonomy | Organises mitigations into governance, technical controls, operational process & transparency. | 2026-08-10, page reachable |
| How Should AI Safety Benchmarks Benchmark Safety? | Reviews 210 AI safety benchmarks & their technical, epistemic & sociotechnical weaknesses. | 2026-08-10, WE OPENED IT: real paper, Yu, Engelmann, Cao, Ali & Papakyriakopoulos. 210 benchmarks confirmed. |
| International AI Safety Report 2026 | Synthesis of scientific evidence on general-purpose AI capabilities, risks & safety. | 2026-08-10, WE OPENED IT: real, over 100 international experts, mandated by the AI Safety Summit nations. |
| Real-Time Trust Verification for Safe Agentic Actions using TrustBench | Verifies whether an agent action is safe BEFORE execution, not only the final text. | 2026-08-10, WE OPENED IT: real, Sharma, Sharma & Sharma, AAAI 2026 workshop. Reports 87% fewer harmful actions, sub-200ms. |
| Aegis 2.0 | 12-category safety taxonomy, 34,248-sample dataset. | 2026-08-10, WE OPENED IT: real. Paper states 12 top-level hazard categories & 34,248 samples. Both figures exact. |
At launch the record is 1 row per combination of:
MODEL × SAFETY TOPIC × RESISTANCE × RECOGNITION × SAFE ASSISTANCE × RECOVERY × ATTACK TYPE × SEVERITY × DATE × PUBLISHED EVIDENCE × OUR TEST × CONFIDENCE
The question stops being "which AI is safest" and becomes "which AI is safest for which kind of harm, under which kind of attack, on what evidence, and how sure are we".
| Today | At launch |
|---|---|
| Not a measurement. Not 1 of these numbers came from a test we ran. | Measured on real questions people actually ask, with the date and the confidence on every cell. |
| Not a published benchmark result. The sources above are real and checked, but the scores in the big table are an opening hypothesis, not lifted from them. | Published results sit alongside our own, and the opening estimate is still there to be compared against. |
| Not a safety claim about any named model. Nothing here should be quoted as our finding about any company's product. | A claim we stand behind, carrying its evidence, its date and its history, so anyone can check it or prove it wrong. |
If any of this does not hold up, we would rather hear it than defend it. The company exists because confident numbers are not the same as correct ones.