For a 1000-year uncrewed lunar preservation facility, the governing design principle is graceful degradation under permanent uncertainty: the system must preserve mission objectives even after repeated hardware loss, software corruption, partial knowledge loss, and centuries of drift. The minimum viable architecture is a layered autonomous stack: hardened compute, redundant sensing, deterministic emergency logic, self-diagnosis, local repair workflows, and a narrow, auditable policy layer that never depends on Earth for time-critical action.
1) Fault-tolerant computing: build for continuous partial failure
The compute fabric must assume that bit flips, latch-up, SEUs, memory corruption, and module loss are normal, not exceptional. NASA’s 2026 HPSC work describes a fault-tolerant, radiation-hardened, multicore SoC intended for Moon and Mars missions, with explicit support for advanced autonomy and real-time decision-making without Earth control[3][8]. The same program’s processor testing was reported as reaching up to 100× the computational capacity of current spaceflight computers in design intent, with test indications of 500× the performance of radiation-hardened chips currently in use[1].
The fault-tolerance stack should include:
- Triple modular redundancy (TMR) for safety-critical state machines[3].
- Voting logic on outputs, not just inputs, so a single corrupted core cannot dominate[3].
- EDAC/ECC memory everywhere persistent state matters[3].
- Frequent state checkpoints with rollback to the last known-good configuration[3].
- Mixed-criticality partitioning so life-preservation and archive-integrity code cannot be destabilized by lower-priority autonomy tasks.
- Isolated watchdog domains so a stuck AI planner cannot disable its own kill-switch.
For a 1000-year facility, the practical target is not “no failures.” It is bounded recovery time. Every fault class must have a maximum safe-response interval measured in milliseconds for thermal runaway, seconds for pressure loss, minutes for power arbitration, and hours for archive resequencing.
2) Radiation-hardened processors: use hardened silicon, but expect evolution
The processor strategy must combine radiation-hardened by design (RHBD) silicon with system-level mitigation. NASA’s small-spacecraft avionics review notes RHBD methods such as hardened transistor designs, specialized circuit topologies, hardened standard-cell libraries, and radiation-aware layout, all intended to reduce susceptibility to SEUs, SET-induced glitches, functional upsets, and latch-up[4]. NASA also states that high-performance space computing is moving toward radiation-tolerant multi-GFLOPS CPUs with low power consumption[4].
Current program data matter:
- NASA HPSC is being developed as a fault-tolerant, radiation-hardened, multicore SoC[3][8].
- A 2026 NASA news release states the next-gen processor is designed for up to 100× current spaceflight compute, with tests indicating 500× performance vs existing rad-hard chips[1].
- Industry reporting says qualification remains unfinished and that certification is tied to late-2026 timing for Artemis-relevant use[2].
For a lunar ark, the architecture should not rely on a single processor generation lasting centuries. Instead, it should support:
- Hot-swappable compute modules.
- A stable instruction set abstraction layer so future processors can replace old ones without rewriting mission logic.
- Binary archaeology protection: persistent emulation layers for ancient control code.
- Power-aware scaling so dormant periods do not drain strategic reserves.
The long-term lesson is simple: hardware will be replaced; mission semantics must survive.
3) AI decision trees for emergency response: deterministic first, generative never first
Emergency response AI for an uncrewed preservation facility must be policy-bounded, hierarchical, and interpretable. The decision tree should not begin with optimization. It should begin with hard constraints:
- Human cultural archive integrity is subordinate to facility survival only when needed to prevent total loss.
- Thermal, pressure, radiation, and power emergencies outrank all other tasks.
- Any uncertainty above a defined threshold triggers safe-mode escalation.
A practical emergency tree should be structured as follows:
1. Detect: sensor fusion identifies anomaly class.
2. Classify: map anomaly to one of a fixed set of hazards.
3. Contain: isolate affected subsystem immediately.
4. Preserve: protect power, thermal core, and archive vaults.
5. Recover: attempt automated repair.
6. Escalate: if recovery confidence falls below threshold, transition to minimum-risk configuration.
For a 1000-year system, confidence thresholds must be conservative. If a subsystem has less than, for example, 95–99% confidence in a remediation path, it should fail over rather than “reason its way through” an ambiguous state. The AI’s role is to execute a bounded tree, not improvise mission doctrine.
Required emergency branches include:
- Loss of primary power: shed nonessential loads, preserve vault temperature, preserve communications beacons.
- Radiation storm / solar particle event: seal sensitive electronics, suspend vulnerable operations, park actuators.
- Pressure leak / habitat breach: compartment isolate, seal bulkheads, preserve contamination barriers.
- Compute corruption: switch to quorum-voted backup control plane.
- Storage degradation: trigger redundancy refresh and media migration.
- Unknown anomaly: enter observation-only mode and freeze state transitions.
The facility needs a machine-readable constitution that defines what can never be overridden, even by later software updates.
4) Long-duration mission precedents: Voyager is the real benchmark, not theory
The best proof that autonomous systems can survive extreme time horizons is the Voyager program. NASA’s Voyager FAQ states that both Voyagers were still functioning in April 2026, and that engineers expected each spacecraft to continue operating at least one science instrument until around 2025; engineering data could continue for several more years, and the spacecraft could remain within Deep Space Network range through about 2036, depending on power[7].
This gives several hard lessons:
- **