food_chain/docs/removal-pass/IMPLEMENTATION.md

148 lines
6.5 KiB
Markdown
Raw Permalink Normal View History

2026-09-29 23:45:12 +09:00
# Runtime cost removal pass
Date: 2026-09-29
## Purpose
This build applies every removal/reduction that the preceding A/B pass supported strongly enough to change without redesigning numerical state. The original integrated build remains separately available at `/mnt/data/mandelbrot_integrated`; this directory is the experimental production-equivalent removal build.
Production source of truth in this build: `index.html`.
## Applied changes
### 1. Recolor no longer copies numeric history
Color-only recolor now promotes only color history. `meta/smooth` numeric history is not recopied when the numerical field did not change.
A/B traffic avoided per accepted recolor at the preset caps:
- fast: 4 MiB
- standard: 8 MiB
- high: 32 MiB
At 20 Hz color cycling this removes 80 / 160 / 640 MiB/s of avoidable copy traffic respectively. These are byte counts, not measured GPU frame-time gains.
### 2. Removed unread CPU fallback RGBA retention
The CPU fallback still creates the RGBA array needed for `ImageData`, but no longer stores a second long-lived reference in `state.cpuFrameRgba`. The state property itself was removed.
Avoided retained JS memory at the pixel caps:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
### 3. Removed partial exact-interior pre-primary unknown scan
`certifyInterior()` no longer runs a full-frame unknown-stat pass immediately before the primary numerical pass. The primary path computes the required statistics later. A fully certified frame now uses known-zero stats directly instead of a readback.
Minimum avoided `fieldMeta` read per partial certification:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
### 4. Removed FAST coverage-repair pass and its dead remnants
The second full-screen FAST numerical pass that attempted to repair `reason=0` holes was removed. Validation before removal found zero holes in 36 layout cases plus 200 randomized semantic cases.
After integration, two dead remnants were also removed:
- the post-tile `numericParams` upload whose only purpose was coverage-repair mode 2;
- the coverage-hole mode-2 branch in `FAST_PERTURB_WGSL`.
Representative eliminated shader invocations:
- 1365×768: 1,050,624 invocations
- 2731×1536: 4,202,496 invocations
The FAST kernel hash changed only because this unreachable coverage-repair branch was deleted. `tests/kernel_hashes.json` was updated accordingly.
### 5. Removed idle color-auto snapshot allocation
Turning on automatic hue cycling while no render is running no longer captures `colorSource`. Idle recolor does not consume that snapshot. Render-time snapshot behavior remains because it is used to keep visual continuity while a new numerical frame is in flight.
Avoided idle snapshot residency at caps:
- fast: 6 MiB
- standard: 12 MiB
- high: 48 MiB
### 6. Export workspace is released after every export
`runExport()` now destroys the reusable export ring in `finally`. This removes roughly 24 MiB of post-export residency in the tested 512×512 / depth-3 configuration. The next export must allocate the ring again; real-device allocation latency was not measurable in this environment.
### 7. Exact-interior proof is spatially gated
Before the BigInt tile proof, the exact viewport is tested against conservative bounding boxes for the main cardioid and period-2 bulb. If neither can intersect the viewport, the proof is skipped.
A/B aggregate CPU proof time on the fixed corpus:
- 1024×768: 41.64 ms → 25.16 ms (~39.6% reduction)
- 2048²: 177.70 ms → 117.08 ms (~34.1% reduction)
Random validation: 240 views, 160 gate-negative, 0 false negatives versus the existing exact tile proof.
### 8. Numeric history is no longer permanently resident when exact-grid reuse is unrealistic
`historyMeta/historySmooth` are allocated only when the render-to-CSS scale ratio has a reduced denominator ≤ 8. Native-grid cases retain the feature; common pixel-budget-scaled cases do not keep 8 B/pixel of history that almost never matches integer-grid pan reuse.
Measured deterministic pan survey:
- 1920×1080 → 965×543: 0.2% reuse, history removed (~4 MiB)
- 1920×1080 → 1365×768: 0.6% reuse, history removed (~8 MiB)
- 1920×1080 → 2731×1536: 0% reuse, history removed (~32 MiB)
- 1024×768 native: 100% reuse, history retained
- 1024×768 DPR2/native ratio: 100% reuse, history retained
This is intentionally conditional rather than blanket deletion.
### 9. Export AA sample textures are lazy
The four RGBA8 sample textures used only by 2× AA export are no longer allocated by 1× export. At 512² and ring depth 3, 1× export avoids 12 MiB of peak sample-texture allocation.
## Deliberately not removed
The following were investigated but are not part of this deletion pass:
- **Adaptive initial-iteration probe**: 3/48 survey views selected more than 350 iterations; forcing 350 increased the surrogate work metric by +7.2%, +25.3%, and +33.1%.
- **Strict periodicity-state updates**: the CPU A/B did not establish a speed win; median B was slower.
- **Continuation restart from n=0**: confirmed redundant work, but fixing it requires carrying valid continuation state across primary/reference transitions rather than deleting code.
- **Deep/correction workspace lifetime**: potentially large memory saving, but unconditional release can cause GPU allocation churn; no real WebGPU allocation timing was available, so it remains unchanged in this pass.
## Validation
Commands:
```bash
node tests/regression.mjs
node experiments/ab/variant_check.mjs index.html
node experiments/ab/supplement.mjs
node experiments/integrated_bench.mjs docs/removal-pass/integrated-bench.json
```
Results:
- regression: PASS
- JavaScript syntax: PASS
- WGSL kernels: 19 + bundle version
- exact-interior gate randomized test: 0 false negatives / 240 views
- FAST coverage semantic test: 0 hole cases / 200 randomized cases
- integrated CPU benchmark: completed; existing periodicity/series/sparse/multi-reference/CPU-fallback paths remain operational
Real WebGPU frame-time p50/p95 remains unmeasured because the container cannot start a usable WebGPU adapter.
## Source-size impact
Compared with the preceding integrated build:
- `index.html`: 223,672 → 224,559 bytes (+887 bytes)
- gzip-9: 57,375 → 57,744 bytes (+369 bytes)
The file is slightly larger because the spatial and allocation eligibility gates add code. The intended gains are runtime computation, transfer traffic, and memory residency rather than download size.
## Integrity
Removal-build `index.html` SHA-256:
`e69b91ead027927d977a065d27d9e27629110122df20aa470eab877e23f77ac7`