# Macrostep 03 — CPU planar simulation engine ## Objective Build the first complete deterministic SmoothLife simulator using the pure core and a conventional CPU FFT. Support headless stepping and all base dimensions before introducing GPU complexity. ## Dependencies Macrosteps 00–02 complete. ## Phase 3.1 — FFT convolution backend ### Substep 3.1.1 — Transform convention Implement the default `FftAlgorithm::Standard` CPU path with a standard library FFT. The retained `LegacyPackedUnitary` algorithm is a GPU backend planned in Macrosteps 05–06, with only a small CPU stage oracle for tests. Define and test the standard convention: 1. unscaled forward transform; 2. spectral multiplication; 3. inverse scaled exactly by `1/(Nx·Ny·Nz)`; 4. divide by sampled kernel sum, or pre-normalize kernels once. Never apply `sqrt(sample_count)` in the standard path; that correction belongs only to `LegacyPackedUnitary`. ### Substep 3.1.2 — Separable N-D transforms Implement transforms along each active axis with reusable plans and scratch buffers. Start with full complex storage for clarity. If real-to-complex optimization is later added, retain the full-complex implementation as the reference. Requirements: - arbitrary validated extents supported by the chosen FFT library on CPU; - no x/y/z transposition ambiguity; - cached plans keyed by shape; - no steady-state allocations; - serial correctness before Rayon parallelism. ### Substep 3.1.3 — Spectral kernel cache Build/cache spectra only when shape or kernel geometry changes. Keep disk/ring spectra, sums, and source parameters together so stale kernels cannot be paired with a new config. ## Phase 3.2 — Planar backend pipeline ### Substep 3.2.1 — Backend state The planar CPU backend owns: - current and next state; - `M`, `N`, target, and derivative channels; - disk/ring kernels and spectra; - FFT work buffers/plans; - integrator and AB history; - generation and timing counters; - run descriptor and deterministic initializer state. ### Substep 3.2.2 — Euler/discrete step Implement the complete stage pipeline: ```text FFT(A) → multiply by KDF/KRF → inverse + normalize to M/N → evaluate S → discrete replacement or Euler derivative update → clamp and swap ``` Expose each stage through inspection without forcing a copy when the CPU owns the field. ### Substep 3.2.3 — AB3 and RK4 - AB3 performs one derivative evaluation per step and rotates initialized history. - RK4 performs the full convolution/rule pipeline at each clamped stage and supports both `StageState` and historical `StepOrigin` relaxation references. - Discrete mode bypasses both methods. - Parameter commands use Macrostep 00's full invalidation matrix: only presentation/inspection/pause/scheduler edits preserve history; rule, dynamics, timestep, integrator, kernel, historical-option selection, state/initializer, topology, shape, variant, and backend edits invalidate it. ## Phase 3.3 — Simulation command facade Define a small object-safe facade or enum dispatch with commands such as: ```text step(count) pause/single-step handled by app, not backend reset(initializer, seed) apply_hot(rule edit) rebuild_warm(kernel edit) recreate_cold(shape/variant edit) load/export restart state load/export exact checkpoint inspection(channel) metrics() ``` Apply warm/cold changes transactionally: build and validate new resources first, then replace the live backend. Preserve the prior valid simulation if rebuilding fails. ## Phase 3.4 — Headless runner Implement CLI execution that can: - choose preset/shape/seed/mode/integrator; - run exact step counts; - export restart state, exact continuation checkpoints, and inspection channels; - print checksums, min/max/mean/variance, occupancy histogram, and elapsed time; - compare two state files with max, L1, and L2 error. The headless path and interactive path must call the same backend methods. ## Phase 3.5 — Determinism and parallelism 1. Establish single-thread deterministic fixtures. 2. Add Rayon only around independent transform lines or pointwise passes. 3. Ensure reductions affecting normalization use a defined order or precomputed kernel sums. 4. Document whether multi-threaded execution is bit-identical or tolerance-equivalent. 5. Include thread count and CPU backend version in benchmark metadata, not in mathematical run identity unless output differs. ## Phase 3.6 — Verification ### Stage-level tests - FFT round trip: zero, constant, impulse, checkerboard, and asymmetric matrices. - FFT convolution against direct `f64` oracle in all dimensions. - Default kernel-sum checks. - One-step output for every dynamics/integrator combination, including both RK4 relaxation references. - Stage snapshots for `M`, `N`, `S`, derivative, and committed state. ### Deterministic scenario tests Run fixed states for 1, 2, and 10 steps. Compare arrays for CPU serial mode and aggregate metrics for allowed parallel differences. Include representative 1-D, 2-D discrete, 2-D growth, 2-D relaxation, and 3-D presets. ### Allocation tests Use a counting allocator or benchmark instrumentation to prove no allocation during steady-state Euler/AB/RK stepping after warm-up. ## Performance baseline, not optimization gate Record release-build timings for: - 1-D 1024 and 8192; - 2-D 128² and 512²; - 3-D 32³ and 64³; - Euler versus RK4. These measurements establish later GPU speedup and regression baselines. Correctness is the exit gate; do not distort architecture to hit a premature number. ## Deliverables - CPU FFT convolution backend. - Complete base planar simulator in 1-D/2-D/3-D. - Headless runner, state import/export, metrics, and comparisons. - Stage-level fixtures and baseline report. ## Exit gate - FFT convolution matches the direct oracle within `2e-5` on agreed small fields. - All base one-step fixtures pass in 1-D/2-D/3-D. - Seed + preset + shape + backend + step count reproduces output. - RK4 invokes four complete derivative pipelines; AB history startup/reset is deterministic. - No steady-state heap allocation occurs after warm-up. - Headless state and metrics can be consumed without raylib.