148 lines
6.5 KiB
Markdown
148 lines
6.5 KiB
Markdown
|
|
# Runtime cost removal pass
|
|||
|
|
|
|||
|
|
Date: 2026-09-29
|
|||
|
|
|
|||
|
|
## Purpose
|
|||
|
|
|
|||
|
|
This build applies every removal/reduction that the preceding A/B pass supported strongly enough to change without redesigning numerical state. The original integrated build remains separately available at `/mnt/data/mandelbrot_integrated`; this directory is the experimental production-equivalent removal build.
|
|||
|
|
|
|||
|
|
Production source of truth in this build: `index.html`.
|
|||
|
|
|
|||
|
|
## Applied changes
|
|||
|
|
|
|||
|
|
### 1. Recolor no longer copies numeric history
|
|||
|
|
|
|||
|
|
Color-only recolor now promotes only color history. `meta/smooth` numeric history is not recopied when the numerical field did not change.
|
|||
|
|
|
|||
|
|
A/B traffic avoided per accepted recolor at the preset caps:
|
|||
|
|
|
|||
|
|
- fast: 4 MiB
|
|||
|
|
- standard: 8 MiB
|
|||
|
|
- high: 32 MiB
|
|||
|
|
|
|||
|
|
At 20 Hz color cycling this removes 80 / 160 / 640 MiB/s of avoidable copy traffic respectively. These are byte counts, not measured GPU frame-time gains.
|
|||
|
|
|
|||
|
|
### 2. Removed unread CPU fallback RGBA retention
|
|||
|
|
|
|||
|
|
The CPU fallback still creates the RGBA array needed for `ImageData`, but no longer stores a second long-lived reference in `state.cpuFrameRgba`. The state property itself was removed.
|
|||
|
|
|
|||
|
|
Avoided retained JS memory at the pixel caps:
|
|||
|
|
|
|||
|
|
- fast: 2 MiB
|
|||
|
|
- standard: 4 MiB
|
|||
|
|
- high: 16 MiB
|
|||
|
|
|
|||
|
|
### 3. Removed partial exact-interior pre-primary unknown scan
|
|||
|
|
|
|||
|
|
`certifyInterior()` no longer runs a full-frame unknown-stat pass immediately before the primary numerical pass. The primary path computes the required statistics later. A fully certified frame now uses known-zero stats directly instead of a readback.
|
|||
|
|
|
|||
|
|
Minimum avoided `fieldMeta` read per partial certification:
|
|||
|
|
|
|||
|
|
- fast: 2 MiB
|
|||
|
|
- standard: 4 MiB
|
|||
|
|
- high: 16 MiB
|
|||
|
|
|
|||
|
|
### 4. Removed FAST coverage-repair pass and its dead remnants
|
|||
|
|
|
|||
|
|
The second full-screen FAST numerical pass that attempted to repair `reason=0` holes was removed. Validation before removal found zero holes in 36 layout cases plus 200 randomized semantic cases.
|
|||
|
|
|
|||
|
|
After integration, two dead remnants were also removed:
|
|||
|
|
|
|||
|
|
- the post-tile `numericParams` upload whose only purpose was coverage-repair mode 2;
|
|||
|
|
- the coverage-hole mode-2 branch in `FAST_PERTURB_WGSL`.
|
|||
|
|
|
|||
|
|
Representative eliminated shader invocations:
|
|||
|
|
|
|||
|
|
- 1365×768: 1,050,624 invocations
|
|||
|
|
- 2731×1536: 4,202,496 invocations
|
|||
|
|
|
|||
|
|
The FAST kernel hash changed only because this unreachable coverage-repair branch was deleted. `tests/kernel_hashes.json` was updated accordingly.
|
|||
|
|
|
|||
|
|
### 5. Removed idle color-auto snapshot allocation
|
|||
|
|
|
|||
|
|
Turning on automatic hue cycling while no render is running no longer captures `colorSource`. Idle recolor does not consume that snapshot. Render-time snapshot behavior remains because it is used to keep visual continuity while a new numerical frame is in flight.
|
|||
|
|
|
|||
|
|
Avoided idle snapshot residency at caps:
|
|||
|
|
|
|||
|
|
- fast: 6 MiB
|
|||
|
|
- standard: 12 MiB
|
|||
|
|
- high: 48 MiB
|
|||
|
|
|
|||
|
|
### 6. Export workspace is released after every export
|
|||
|
|
|
|||
|
|
`runExport()` now destroys the reusable export ring in `finally`. This removes roughly 24 MiB of post-export residency in the tested 512×512 / depth-3 configuration. The next export must allocate the ring again; real-device allocation latency was not measurable in this environment.
|
|||
|
|
|
|||
|
|
### 7. Exact-interior proof is spatially gated
|
|||
|
|
|
|||
|
|
Before the BigInt tile proof, the exact viewport is tested against conservative bounding boxes for the main cardioid and period-2 bulb. If neither can intersect the viewport, the proof is skipped.
|
|||
|
|
|
|||
|
|
A/B aggregate CPU proof time on the fixed corpus:
|
|||
|
|
|
|||
|
|
- 1024×768: 41.64 ms → 25.16 ms (~39.6% reduction)
|
|||
|
|
- 2048²: 177.70 ms → 117.08 ms (~34.1% reduction)
|
|||
|
|
|
|||
|
|
Random validation: 240 views, 160 gate-negative, 0 false negatives versus the existing exact tile proof.
|
|||
|
|
|
|||
|
|
### 8. Numeric history is no longer permanently resident when exact-grid reuse is unrealistic
|
|||
|
|
|
|||
|
|
`historyMeta/historySmooth` are allocated only when the render-to-CSS scale ratio has a reduced denominator ≤ 8. Native-grid cases retain the feature; common pixel-budget-scaled cases do not keep 8 B/pixel of history that almost never matches integer-grid pan reuse.
|
|||
|
|
|
|||
|
|
Measured deterministic pan survey:
|
|||
|
|
|
|||
|
|
- 1920×1080 → 965×543: 0.2% reuse, history removed (~4 MiB)
|
|||
|
|
- 1920×1080 → 1365×768: 0.6% reuse, history removed (~8 MiB)
|
|||
|
|
- 1920×1080 → 2731×1536: 0% reuse, history removed (~32 MiB)
|
|||
|
|
- 1024×768 native: 100% reuse, history retained
|
|||
|
|
- 1024×768 DPR2/native ratio: 100% reuse, history retained
|
|||
|
|
|
|||
|
|
This is intentionally conditional rather than blanket deletion.
|
|||
|
|
|
|||
|
|
### 9. Export AA sample textures are lazy
|
|||
|
|
|
|||
|
|
The four RGBA8 sample textures used only by 2× AA export are no longer allocated by 1× export. At 512² and ring depth 3, 1× export avoids 12 MiB of peak sample-texture allocation.
|
|||
|
|
|
|||
|
|
## Deliberately not removed
|
|||
|
|
|
|||
|
|
The following were investigated but are not part of this deletion pass:
|
|||
|
|
|
|||
|
|
- **Adaptive initial-iteration probe**: 3/48 survey views selected more than 350 iterations; forcing 350 increased the surrogate work metric by +7.2%, +25.3%, and +33.1%.
|
|||
|
|
- **Strict periodicity-state updates**: the CPU A/B did not establish a speed win; median B was slower.
|
|||
|
|
- **Continuation restart from n=0**: confirmed redundant work, but fixing it requires carrying valid continuation state across primary/reference transitions rather than deleting code.
|
|||
|
|
- **Deep/correction workspace lifetime**: potentially large memory saving, but unconditional release can cause GPU allocation churn; no real WebGPU allocation timing was available, so it remains unchanged in this pass.
|
|||
|
|
|
|||
|
|
## Validation
|
|||
|
|
|
|||
|
|
Commands:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
node tests/regression.mjs
|
|||
|
|
node experiments/ab/variant_check.mjs index.html
|
|||
|
|
node experiments/ab/supplement.mjs
|
|||
|
|
node experiments/integrated_bench.mjs docs/removal-pass/integrated-bench.json
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Results:
|
|||
|
|
|
|||
|
|
- regression: PASS
|
|||
|
|
- JavaScript syntax: PASS
|
|||
|
|
- WGSL kernels: 19 + bundle version
|
|||
|
|
- exact-interior gate randomized test: 0 false negatives / 240 views
|
|||
|
|
- FAST coverage semantic test: 0 hole cases / 200 randomized cases
|
|||
|
|
- integrated CPU benchmark: completed; existing periodicity/series/sparse/multi-reference/CPU-fallback paths remain operational
|
|||
|
|
|
|||
|
|
Real WebGPU frame-time p50/p95 remains unmeasured because the container cannot start a usable WebGPU adapter.
|
|||
|
|
|
|||
|
|
## Source-size impact
|
|||
|
|
|
|||
|
|
Compared with the preceding integrated build:
|
|||
|
|
|
|||
|
|
- `index.html`: 223,672 → 224,559 bytes (+887 bytes)
|
|||
|
|
- gzip-9: 57,375 → 57,744 bytes (+369 bytes)
|
|||
|
|
|
|||
|
|
The file is slightly larger because the spatial and allocation eligibility gates add code. The intended gains are runtime computation, transfer traffic, and memory residency rather than download size.
|
|||
|
|
|
|||
|
|
## Integrity
|
|||
|
|
|
|||
|
|
Removal-build `index.html` SHA-256:
|
|||
|
|
|
|||
|
|
`e69b91ead027927d977a065d27d9e27629110122df20aa470eab877e23f77ac7`
|