Future of Life Institute's Summer 2025 AI Safety Index graded seven leading AI companies across 33 indicators in six domains; Anthropic earned top C+ overall, followed by OpenAI and Google DeepMind. Winter 2025 edition expanded to eight firms (adding xAI, Meta, Z.ai, DeepSeek, Alibaba Cloud), revealing persistent gaps in risk assessment, safety frameworks, and information sharing. All companies lack explicit plans for controlling AGI/superintelligence, with high jailbreak success rates on AIR-Bench 2024 across 5,694 tests in 314 risk categories.
Weak AI safety guardrails—evidenced by high Attack Success Rates (ASR) in algorithmic jailbreaking—increase catastrophic misuse risks like cybersecurity breaches and operational failures, directly endangering lunar habitation systems reliant on AI for life support, radiation shielding, and autonomous resource management. Stanford AI Index 2025 reports 233 AI incidents in 2024, up 56.4% from 2023, amplifying existential threats from unaligned superintelligence that could disrupt Earth-Moon supply chains or trigger global collapse before Ark self-sufficiency.
Ark team must monitor FLI AI Safety Index updates quarterly, research jailbreak-resistant AI for lunar deployment via AIR-Bench/HELM Safety benchmarks, and integrate external safety research funding mechanisms to harden Ark AI against misuse, targeting zero-tolerance for ASR above 5% in critical systems.