← All topics

AI Ethics and Applications

AI Safety and Evaluation

AIの安全性あんぜんせい評価ひょうか

AI safety and evaluation measures capabilities and risks using a benchmark, tests robustness through practices such as red teaming, and studies failures including hallucination and unfair outcomes addressed through fairness work.

Japanese terms

  1. AI safety and evaluation — AIの安全性あんぜんせい評価ひょうか: The practice of measuring AI capabilities, limitations, and risks and reducing the likelihood or severity of harmful outcomes.
  2. Benchmark — ベンチマーク: A standardized task, dataset, or measurement used to compare systems.
  3. Robustness — 頑健性がんけんせい: The ability of a system to maintain acceptable behavior under variation, noise, or attack.
  4. Red teaming — レッドチーミング: Adversarial testing that searches deliberately for failures, vulnerabilities, or harmful behavior.
  5. Hallucination — ハルシネーション: Generated content that is unsupported, false, or inconsistent with the supplied evidence.
  6. Fairness — 公平性こうへいせい: The study and management of unjustified differences in how AI systems affect people or groups.

Model evaluation supplies general measurement concepts, while AI ethics and AI regulation and policy address wider social and institutional responsibilities.