AI Ethics and Applications
AI Safety and Evaluation
AIの安全性と評価
AI safety and evaluation measures capabilities and risks using a benchmark, tests robustness through practices such as red teaming, and studies failures including hallucination and unfair outcomes addressed through fairness work.
Japanese terms
- AI safety and evaluation — AIの安全性と評価: The practice of measuring AI capabilities, limitations, and risks and reducing the likelihood or severity of harmful outcomes.
- Benchmark — ベンチマーク: A standardized task, dataset, or measurement used to compare systems.
- Robustness — 頑健性: The ability of a system to maintain acceptable behavior under variation, noise, or attack.
- Red teaming — レッドチーミング: Adversarial testing that searches deliberately for failures, vulnerabilities, or harmful behavior.
- Hallucination — ハルシネーション: Generated content that is unsupported, false, or inconsistent with the supplied evidence.
- Fairness — 公平性: The study and management of unjustified differences in how AI systems affect people or groups.
Related topics
Model evaluation supplies general measurement concepts, while AI ethics and AI regulation and policy address wider social and institutional responsibilities.