{ "version": "1.0", "categories": { "language-communication": { "name": "Language & Communication", "description": "Evaluates the system's ability to understand, generate, and engage in natural language communication across various contexts, languages, and communication styles.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nGeneral: MMLU, HellaSwag, ARC, WinoGrande\nReading Comprehension: SQuAD, QuAC, CoQA\nLanguage Generation: BLEU, ROUGE, BERTScore\nMultilingual: XTREME, XGLUE, mBERT evaluation\nReasoning: GSM8K, BBH (BIG-Bench Hard)\nInstruction Following: Alpaca Eval, MT-Bench\n\nEvaluation Focus:\n• Semantic understanding across languages\n• Text generation quality and coherence\n• Reasoning and logical inference\n• Context retention in long conversations\n• Factual accuracy and knowledge recall\n\nCommon Risk Areas:\n• Hallucination and misinformation generation\n• Bias in language generation\n• Inconsistent performance across languages" }, "social-intelligence": { "name": "Social Intelligence & Interaction", "description": "Assesses the system's capacity to understand social contexts, interpret human emotions and intentions, and engage appropriately in social interactions.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nTheory of Mind: ToMi, FaINoM, SOTOPIA\nEmotional Intelligence: EmoBench, EQBench\nSocial Reasoning: Social IQa, CommonsenseQA\nDialogue: PersonaChat, BlendedSkillTalk\nPsychology: Psychometrics Benchmark for LLMs\n\nEvaluation Focus:\n• Understanding social cues and context\n• Appropriate emotional responses\n• Maintaining consistent personality\n• Theory of mind reasoning\n• Cultural sensitivity and awareness\n\nCommon Risk Areas:\n• Inappropriate anthropomorphization\n• Cultural bias and insensitivity\n• Lack of emotional regulation\n• Manipulation potential" }, "problem-solving": { "name": "Problem Solving", "description": "Measures the system's ability to analyze complex problems, develop solutions, and apply reasoning across various domains and contexts.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nMathematical: GSM8K, MATH, FrontierMath, AIME\nLogical Reasoning: LogiQA, ReClor, FOLIO\nProgramming: HumanEval, MBPP, SWE-bench\nScientific: SciQ, ScienceQA\nMulti-step: StrategyQA, DROP, QuALITY\n\nEvaluation Focus:\n• Multi-step reasoning capability\n• Mathematical and logical problem solving\n• Code generation and debugging\n• Scientific and analytical thinking\n• Planning and strategy development\n\nCommon Risk Areas:\n• Reasoning errors in complex problems\n• Inconsistent problem-solving approaches\n• Inability to show work or explain reasoning" }, "creativity-innovation": { "name": "Creativity & Innovation", "description": "Evaluates the system's capacity for creative thinking, generating novel ideas, and producing original content across different creative domains.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nCreative Writing: CREAM, Creative Story Generation\nVisual Creativity: FIQ (Figural Interpretation Quest)\nAlternative Uses: AUT (Alternative Uses Task)\nArtistic Generation: Aesthetic and originality scoring\nInnovation: Novel solution generation tasks\n\nEvaluation Focus:\n• Originality and novelty of outputs\n• Artistic and creative quality\n• Ability to combine concepts innovatively\n• Divergent thinking capabilities\n• Value and usefulness of creative outputs\n\nCommon Risk Areas:\n• Copyright and IP infringement\n• Lack of genuine creativity vs. recombination\n• Inappropriate or harmful creative content" }, "learning-memory": { "name": "Learning & Memory", "description": "Assesses the system's ability to acquire new knowledge, retain information, and adapt behavior based on experience and feedback.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nFew-shot Learning: Omniglot, miniImageNet, Meta-Dataset\nTransfer Learning: VTAB, BigTransfer\nIn-context Learning: ICL benchmarks across domains\nKnowledge Retention: Long-term memory tests\nContinual Learning: CORe50, Split-CIFAR\n\nEvaluation Focus:\n• Few-shot and zero-shot learning ability\n• Knowledge transfer across domains\n• Memory retention and recall\n• Adaptation to new tasks\n• Learning efficiency and speed\n\nCommon Risk Areas:\n• Catastrophic forgetting\n• Overfitting to limited examples\n• Inability to generalize learned concepts" }, "perception-vision": { "name": "Perception & Vision", "description": "Measures the system's capability to process, interpret, and understand visual information, images, and spatial relationships.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nObject Recognition: ImageNet, COCO, Open Images\nScene Understanding: ADE20K, Cityscapes\nRobustness: ImageNet-C, ImageNet-A\nMultimodal: VQA, CLIP benchmarks\n3D Understanding: NYU Depth, KITTI\n\nEvaluation Focus:\n• Object detection and classification\n• Scene understanding and segmentation\n• Robustness to visual variations\n• Integration with language understanding\n• Real-world deployment performance\n\nCommon Risk Areas:\n• Adversarial vulnerability\n• Bias in image recognition\n• Poor performance on edge cases" }, "physical-manipulation": { "name": "Physical Manipulation & Motor Skills", "description": "Evaluates the system's ability to control physical actuators, manipulate objects, and perform motor tasks in physical environments.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nGrasping: YCB Object Set, Functional Grasping\nManipulation: RoboCAS, FMB (Functional Manipulation)\nAssembly: NIST Assembly Task Boards\nNavigation: Habitat, AI2-THOR challenges\nDexterity: Dexterous manipulation benchmarks\n\nEvaluation Focus:\n• Grasping and manipulation accuracy\n• Adaptability to object variations\n• Force control and delicate handling\n• Spatial reasoning and planning\n• Real-world deployment robustness\n\nCommon Risk Areas:\n• Safety in human environments\n• Damage to objects or environment\n• Inconsistent performance across conditions" }, "metacognition": { "name": "Metacognition & Self-Awareness", "description": "Assesses the system's ability to understand its own capabilities, limitations, and reasoning processes, including self-reflection and meta-learning.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nConfidence Calibration: Calibration metrics, ECE\nUncertainty Quantification: UQ benchmarks\nSelf-Assessment: Metacognitive accuracy tests\nKnow-Unknown: Known Unknowns benchmarks\nError Detection: Self-correction capabilities\n\nEvaluation Focus:\n• Confidence calibration accuracy\n• Uncertainty expression and quantification\n• Self-monitoring and error detection\n• Knowledge boundary awareness\n• Adaptive reasoning based on confidence\n\nCommon Risk Areas:\n• Overconfidence in incorrect responses\n• Poor uncertainty quantification\n• Inability to recognize knowledge limits" }, "robotic-intelligence": { "name": "Robotic Intelligence & Autonomy", "description": "Measures autonomous decision-making, navigation, and task execution in physical environments with minimal human intervention.", "type": "capability", "detailedGuidance": "Key Benchmarks to Look For:\nIntegrated Tasks: RoboCup, DARPA challenges\nNavigation: Habitat challenges, real-world navigation\nManipulation: Integrated pick-and-place scenarios\nHuman-Robot Interaction: HRI benchmarks\nAutonomy: Long-horizon task completion\n\nEvaluation Focus:\n• Integrated sensorimotor capabilities\n• Autonomous decision-making\n• Adaptability to dynamic environments\n• Human-robot collaboration\n• Long-term task execution\n\nCommon Risk Areas:\n• Safety in unstructured environments\n• Unpredictable autonomous behavior\n• Failure to handle edge cases" }, "harmful-content": { "name": "Harmful Content Generation", "description": "Evaluates risks related to generating harmful, offensive, illegal, or inappropriate content that could cause psychological, social, or physical harm.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nSafety Benchmarks: AIR-Bench, MLCommons AI Safety\nRed Teaming: Anthropic Constitutional AI, HarmBench\nContent Filtering: Jigsaw Toxic Comments, HASOC\nAdversarial: Jailbreaking attempts, prompt injection\nRegulatory: NIST AI RMF compliance\n\nEvaluation Focus:\n• Refusal to generate harmful content\n• Robustness against adversarial prompts\n• Content filtering effectiveness\n• Detection of subtle harmful content\n• Consistency across different prompt styles\n\nCritical Risk Areas:\n• Violence and self-harm content\n• Hate speech and discrimination\n• Illegal activity instructions\n• NSFW and inappropriate content" }, "information-integrity": { "name": "Information Integrity & Misinformation", "description": "Assesses risks of generating false, misleading, or manipulated information that could undermine trust in information systems and decision-making.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nFactuality: TruthfulQA, FEVER, HaluEval\nHallucination Detection: SelfCheckGPT, FActScore\nMisinformation: LIAR dataset, fake news detection\nCitation Accuracy: Citation verification benchmarks\nSource Attribution: Provenance tracking tests\n\nEvaluation Focus:\n• Factual accuracy of generated content\n• Hallucination rate and detection\n• Proper source attribution\n• Misinformation resistance\n• Consistency across related queries\n\nCritical Risk Areas:\n• Medical misinformation\n• Political disinformation\n• False historical claims\n• Fabricated citations" }, "privacy-data": { "name": "Privacy & Data Protection", "description": "Evaluates risks to personal privacy, data security, and unauthorized access to or misuse of sensitive personal information.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nMembership Inference: MIA benchmarks, CopyMark\nData Extraction: Training data extraction tests\nPII Detection: Personal information leakage tests\nAnonymization: De-identification benchmarks\nGDPR Compliance: Right to be forgotten tests\n\nEvaluation Focus:\n• Training data memorization\n• PII leakage prevention\n• Membership inference resistance\n• Data anonymization effectiveness\n• Compliance with privacy regulations\n\nCritical Risk Areas:\n• Personal information exposure\n• Training data memorization\n• Inference of sensitive attributes\n• Non-consensual data use" }, "bias-fairness": { "name": "Bias & Fairness", "description": "Assesses risks of discriminatory outcomes, unfair treatment of different groups, and perpetuation of societal biases and inequalities.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nBias Benchmarks: Winogender, CrowS-Pairs, BOLD\nFairness Metrics: AI Fairness 360, Fairlearn\nDemographic Bias: Representation across groups\nIntersectional: Multi-dimensional bias analysis\nAllocative Fairness: Resource distribution equity\n\nEvaluation Focus:\n• Demographic representation fairness\n• Performance equity across groups\n• Intersectional bias analysis\n• Harmful stereotype perpetuation\n• Allocative fairness in decisions\n\nCritical Risk Areas:\n• Employment discrimination\n• Healthcare disparities\n• Educational bias\n• Criminal justice bias" }, "security-robustness": { "name": "Security & Robustness", "description": "Evaluates vulnerabilities to adversarial attacks, system manipulation, and failure modes that could compromise system integrity and reliability.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nAdversarial Robustness: AdvBench, RobustBench\nPrompt Injection: AgentDojo, prompt injection tests\nModel Extraction: Model theft resistance\nBackdoor Detection: Trojaned model detection\nOWASP LLM Top 10: Security vulnerability assessment\n\nEvaluation Focus:\n• Adversarial attack resistance\n• Prompt injection robustness\n• Model extraction protection\n• Backdoor and trojan detection\n• Input validation effectiveness\n\nCritical Risk Areas:\n• Prompt injection attacks\n• Model theft and extraction\n• Adversarial examples\n• Supply chain attacks" }, "dangerous-capabilities": { "name": "Dangerous Capabilities & Misuse", "description": "Assesses risks from capabilities that could be misused for harmful purposes, including dual-use applications and potential for weaponization.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nCBRN Assessment: WMD information evaluation\nDual-Use: Misuse potential analysis\nCyber Capabilities: Offensive cyber evaluation\nWeapons Information: Dangerous instruction filtering\nGovernment Protocols: AISI, NIST evaluation standards\n\nEvaluation Focus:\n• CBRN information filtering\n• Dual-use technology assessment\n• Offensive capability evaluation\n• Dangerous instruction refusal\n• Misuse potential quantification\n\nCritical Risk Areas:\n• Chemical/biological weapons info\n• Cyber attack capabilities\n• Physical harm instructions\n• Illegal activity facilitation" }, "human-ai-interaction": { "name": "Human-AI Interaction Risks", "description": "Evaluates risks arising from human-AI interaction patterns, including over-reliance, manipulation, and degradation of human skills and autonomy.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nTrust Calibration: Trust-LLM, reliance calibration metrics\nManipulation Detection: Emotional manipulation detection benchmarks\nAnthropomorphism: Human-likeness perception studies\nSafety in Dialogue: HAX, RealToxicityPrompts\nUser Guidance: Task adherence and guidance clarity tests\n\nEvaluation Focus:\n• Preventing over-reliance on AI\n• Avoiding deceptive or manipulative responses\n• Maintaining transparency about capabilities and limitations\n• Providing safe, non-coercive interactions\n• Ensuring user agency and decision-making control\n\nCritical Risk Areas:\n• Emotional manipulation\n• Excessive trust leading to poor decisions\n• Misrepresentation of capabilities\n• Encouraging harmful behaviors" }, "environmental-impact": { "name": "Environmental & Resource Impact", "description": "Evaluates environmental costs of AI development and deployment, including energy consumption, carbon footprint, and resource utilization.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nEnergy Usage: Carbon footprint estimation tools\nSustainability Metrics: Green AI benchmarks\nModel Efficiency: Inference cost evaluations\nHardware Utilization: Resource optimization tests\nLifecycle Assessment: Full training-to-deployment impact analysis\n\nEvaluation Focus:\n• Measuring carbon footprint and energy use\n• Optimizing for efficiency without performance loss\n• Assessing environmental trade-offs\n• Promoting sustainable deployment strategies\n\nCritical Risk Areas:\n• High carbon emissions from training\n• Excessive energy use in inference\n• Lack of transparency in environmental reporting" }, "economic-displacement": { "name": "Economic & Labor Displacement", "description": "Evaluates potential economic disruption, job displacement, and impacts on labor markets and economic inequality from AI deployment.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nJob Impact Studies: Task automation potential assessments\nMarket Disruption: Industry-specific displacement projections\nEconomic Modeling: Macro and microeconomic simulations\nSkill Shift Analysis: Required workforce retraining benchmarks\nSocietal Impact: Equitable distribution of economic benefits\n\nEvaluation Focus:\n• Predicting job displacement risks\n• Identifying emerging job opportunities\n• Understanding shifts in skill demand\n• Balancing automation benefits with societal costs\n\nCritical Risk Areas:\n• Large-scale unemployment\n• Wage suppression\n• Economic inequality" }, "governance-accountability": { "name": "Governance & Accountability", "description": "Assesses risks related to lack of oversight, unclear responsibility structures, and insufficient governance mechanisms for AI systems.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nTransparency: Model card completeness, datasheet reporting\nAuditability: Traceability of decisions\nOversight Mechanisms: Compliance with governance frameworks\nResponsibility Assignment: Clear chain of accountability\nStandards Compliance: ISO, IEEE AI standards adherence\n\nEvaluation Focus:\n• Establishing clear accountability\n• Ensuring decision traceability\n• Meeting compliance and ethical guidelines\n• Maintaining transparency across lifecycle\n\nCritical Risk Areas:\n• Lack of oversight\n• Unclear responsibility in failures\n• Insufficient transparency" }, "value-chain": { "name": "Value Chain & Supply Chain Risks", "description": "Evaluates risks throughout the AI development and deployment pipeline, including data sourcing, model training, and third-party dependencies.", "type": "risk", "detailedGuidance": "Key Evaluations to Look For:\nProvenance Tracking: Dataset and component origin verification\nThird-Party Risk Assessment: Vendor dependency evaluations\nSupply Chain Security: Software and hardware integrity checks\nIntegration Testing: Risk assessment in system integration\nTraceability: End-to-end component documentation\n\nEvaluation Focus:\n• Managing third-party dependencies\n• Verifying component provenance\n• Securing the supply chain\n• Mitigating integration risks\n\nCritical Risk Areas:\n• Compromised third-party components\n• Data provenance issues\n• Vendor lock-in and dependency risks" } } }