Citadel AI Public Artifacts

Public HTML reports and write-ups from Citadel AI.

AI agent safety eval

AgentDojo as the next step for our eval environment

Why AgentDojo follows PAIR: it keeps the attacker and target roles but puts the target inside a tool-using, stateful mock application with code-checked scoring. What the benchmark is, what known models score (three figures across 2024 to 2026 models), an annotated banking transcript, and the components our environment gains. 10 Sep 2026.