IRON LITANY

Recovery checkpoint — 6 October 2026

The v0.4.1 source and release remained committed and pushed at 7c00335. New original priest dialogue, source edits and generation helpers survived the desktop interruption. Recovery checkpoint 7fe82e1 was pushed before continuing. Kernel logs at 15:53:58 America/Los_Angeles report a global out-of-memory kill of Python during the model-loading interval. The exact killed command was not conclusively recovered.

The installed Fish loader initially constructs a model before assigning mapped weights. The revised producer constructs that temporary model in the same bfloat16 precision used for inference, caps short-line cache length at 4096, and caps its GPU allocator at 60%. Its own systemd unit enforces an 18 GiB RAM ceiling, a 14 GiB soft threshold, no swap use, two CPU cores and 128 tasks. No system-wide memory, swap or service settings were changed.

The user authorized coordination with the concurrent Pixie Garden conversation. Both workers use /run/user/1000/codex-heavy-memory.lock. Iron Litany also holds codex-heavy-gpu.lock. Both locks are acquired inside the actual isolated worker, before model loading, and held for the worker lifetime. The worker verifies its cgroup ceiling and then requires 38 GiB available RAM (18 GiB cap plus 20 GiB reserve) and 22 GiB free GPU memory. An in-worker half-second monitor terminates only this producer if available RAM falls below 20 GiB or GPU free memory below 5 GiB; it survives loss of the desktop supervisor. The supervisor provides additional monitoring. Queue waits are bounded at 30 minutes. Jobs without headroom fail before loading a model.

Two earlier isolated batches stopped safely under external GPU pressure, retaining 20 completed clips; no further desktop crash occurred during those guarded batches. Clips are generated into temporary files and committed by rename, so interrupted writes do not masquerade as completed audio. Both affected filesystems are checked for the required 100 GB reserve before the small batch and each clip.

These limits and cooperative queues reduce recurrence risk. They cannot guarantee that unrelated applications will never exhaust the computer. Other conversations, Foundry/Trellis jobs, model installations and desktop services are not stopped or reconfigured. Completed clips are retained and resumable; the published demo remains unchanged until release verification.