food_chain/docs/removal-pass/IMPLEMENTATION.md
2026-09-29 23:45:12 +09:00

6.5 KiB
Raw Permalink Blame History

Runtime cost removal pass

Date: 2026-09-29

Purpose

This build applies every removal/reduction that the preceding A/B pass supported strongly enough to change without redesigning numerical state. The original integrated build remains separately available at /mnt/data/mandelbrot_integrated; this directory is the experimental production-equivalent removal build.

Production source of truth in this build: index.html.

Applied changes

1. Recolor no longer copies numeric history

Color-only recolor now promotes only color history. meta/smooth numeric history is not recopied when the numerical field did not change.

A/B traffic avoided per accepted recolor at the preset caps:

  • fast: 4 MiB
  • standard: 8 MiB
  • high: 32 MiB

At 20 Hz color cycling this removes 80 / 160 / 640 MiB/s of avoidable copy traffic respectively. These are byte counts, not measured GPU frame-time gains.

2. Removed unread CPU fallback RGBA retention

The CPU fallback still creates the RGBA array needed for ImageData, but no longer stores a second long-lived reference in state.cpuFrameRgba. The state property itself was removed.

Avoided retained JS memory at the pixel caps:

  • fast: 2 MiB
  • standard: 4 MiB
  • high: 16 MiB

3. Removed partial exact-interior pre-primary unknown scan

certifyInterior() no longer runs a full-frame unknown-stat pass immediately before the primary numerical pass. The primary path computes the required statistics later. A fully certified frame now uses known-zero stats directly instead of a readback.

Minimum avoided fieldMeta read per partial certification:

  • fast: 2 MiB
  • standard: 4 MiB
  • high: 16 MiB

4. Removed FAST coverage-repair pass and its dead remnants

The second full-screen FAST numerical pass that attempted to repair reason=0 holes was removed. Validation before removal found zero holes in 36 layout cases plus 200 randomized semantic cases.

After integration, two dead remnants were also removed:

  • the post-tile numericParams upload whose only purpose was coverage-repair mode 2;
  • the coverage-hole mode-2 branch in FAST_PERTURB_WGSL.

Representative eliminated shader invocations:

  • 1365×768: 1,050,624 invocations
  • 2731×1536: 4,202,496 invocations

The FAST kernel hash changed only because this unreachable coverage-repair branch was deleted. tests/kernel_hashes.json was updated accordingly.

5. Removed idle color-auto snapshot allocation

Turning on automatic hue cycling while no render is running no longer captures colorSource. Idle recolor does not consume that snapshot. Render-time snapshot behavior remains because it is used to keep visual continuity while a new numerical frame is in flight.

Avoided idle snapshot residency at caps:

  • fast: 6 MiB
  • standard: 12 MiB
  • high: 48 MiB

6. Export workspace is released after every export

runExport() now destroys the reusable export ring in finally. This removes roughly 24 MiB of post-export residency in the tested 512×512 / depth-3 configuration. The next export must allocate the ring again; real-device allocation latency was not measurable in this environment.

7. Exact-interior proof is spatially gated

Before the BigInt tile proof, the exact viewport is tested against conservative bounding boxes for the main cardioid and period-2 bulb. If neither can intersect the viewport, the proof is skipped.

A/B aggregate CPU proof time on the fixed corpus:

  • 1024×768: 41.64 ms → 25.16 ms (~39.6% reduction)
  • 2048²: 177.70 ms → 117.08 ms (~34.1% reduction)

Random validation: 240 views, 160 gate-negative, 0 false negatives versus the existing exact tile proof.

8. Numeric history is no longer permanently resident when exact-grid reuse is unrealistic

historyMeta/historySmooth are allocated only when the render-to-CSS scale ratio has a reduced denominator ≤ 8. Native-grid cases retain the feature; common pixel-budget-scaled cases do not keep 8 B/pixel of history that almost never matches integer-grid pan reuse.

Measured deterministic pan survey:

  • 1920×1080 → 965×543: 0.2% reuse, history removed (~4 MiB)
  • 1920×1080 → 1365×768: 0.6% reuse, history removed (~8 MiB)
  • 1920×1080 → 2731×1536: 0% reuse, history removed (~32 MiB)
  • 1024×768 native: 100% reuse, history retained
  • 1024×768 DPR2/native ratio: 100% reuse, history retained

This is intentionally conditional rather than blanket deletion.

9. Export AA sample textures are lazy

The four RGBA8 sample textures used only by 2× AA export are no longer allocated by 1× export. At 512² and ring depth 3, 1× export avoids 12 MiB of peak sample-texture allocation.

Deliberately not removed

The following were investigated but are not part of this deletion pass:

  • Adaptive initial-iteration probe: 3/48 survey views selected more than 350 iterations; forcing 350 increased the surrogate work metric by +7.2%, +25.3%, and +33.1%.
  • Strict periodicity-state updates: the CPU A/B did not establish a speed win; median B was slower.
  • Continuation restart from n=0: confirmed redundant work, but fixing it requires carrying valid continuation state across primary/reference transitions rather than deleting code.
  • Deep/correction workspace lifetime: potentially large memory saving, but unconditional release can cause GPU allocation churn; no real WebGPU allocation timing was available, so it remains unchanged in this pass.

Validation

Commands:

node tests/regression.mjs
node experiments/ab/variant_check.mjs index.html
node experiments/ab/supplement.mjs
node experiments/integrated_bench.mjs docs/removal-pass/integrated-bench.json

Results:

  • regression: PASS
  • JavaScript syntax: PASS
  • WGSL kernels: 19 + bundle version
  • exact-interior gate randomized test: 0 false negatives / 240 views
  • FAST coverage semantic test: 0 hole cases / 200 randomized cases
  • integrated CPU benchmark: completed; existing periodicity/series/sparse/multi-reference/CPU-fallback paths remain operational

Real WebGPU frame-time p50/p95 remains unmeasured because the container cannot start a usable WebGPU adapter.

Source-size impact

Compared with the preceding integrated build:

  • index.html: 223,672 → 224,559 bytes (+887 bytes)
  • gzip-9: 57,375 → 57,744 bytes (+369 bytes)

The file is slightly larger because the spatial and allocation eligibility gates add code. The intended gains are runtime computation, transfer traffic, and memory residency rather than download size.

Integrity

Removal-build index.html SHA-256:

e69b91ead027927d977a065d27d9e27629110122df20aa470eab877e23f77ac7