Public HTML reports and write-ups from Citadel AI.
AI agent safety evalWhy AgentDojo follows PAIR: it keeps the attacker and target roles but puts the target inside a tool-using, stateful mock application with code-checked scoring. What the benchmark is, what known models score (three figures across 2024 to 2026 models), an annotated banking transcript, and the components our environment gains. 10 Sep 2026.