food_chain/docs/performance-cost-audit/ab/README.md
2026-09-29 23:45:12 +09:00

234 lines
11 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Mandelbrot viewer performance-cost A/B test
Date: 2026-09-29
## Scope
This A/B test follows the performance-cost audit. Production `index.html` was **not changed**. A is the current production implementation; B variants live under `experiments/ab/variants/` and disable, defer, or gate one candidate operation at a time.
Production SHA-256 before/after testing:
`544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a`
The existing regression suite passes on production. The combined high-confidence B variant and the individual structural B variants used for gating/lazy allocation also pass the same regression suite; all retain 19 WGSL kernels.
## Measurement limitation
A usable WebGPU backend could not be started in the container (Chromium/SwiftShader/EGL). Therefore this report deliberately separates:
1. **CPU measured time/work**: exact-interior BigInt proof, adaptive-iteration probe, CPU surrogates for strict-cycle and continuation replay.
2. **Exact traffic/residency counts from the production buffer sizes and dispatch extents**: copies, full-frame scans, allocations.
3. **Structural/regression validation**: B variants parse and pass the production regression suite.
No result below is presented as measured GPU frame time unless explicitly stated. Real-device WebGPU p50/p95 remains a follow-up requirement.
## Summary
| Candidate | A/B result | Decision from this test |
|---|---|---|
| Recolor numeric-history copy | B removes 8 B/pixel copy with unchanged numeric field | **Strongly supported** |
| Unread `cpuFrameRgba` retention | B removes 4 B/pixel retained JS data; no reader exists | **Strongly supported** |
| Pre-primary unknown stats after exact interior | B removes one full-frame `fieldMeta` scan; primary path computes stats later | **Strongly supported** |
| Strict direct periodicity-state updates | CPU surrogate median B change **+2.82% slower** | **Do not prioritize for speed** |
| FAST full-screen coverage repair | 36 layout cases + 200 semantic randomized cases: **0 holes** | **Strongly supported**, real GPU A/B still desirable |
| Idle color-auto `colorSource` snapshot | B avoids 12 B/pixel idle snapshot; idle recolor does not consume it | **Strongly supported** |
| Export workspace lifetime | B releases ~24 MiB max workspace after export | **Memory win**, trades for reallocation on next export |
| Exact-interior spatial gate | Fixed corpus proof CPU time reduced ~40% at 1024×768 and ~34% at 2048²; 240 random views: 0 gate false-negatives | **Strongly supported** |
| Conditional numeric-history residency | Budget-scaled 1920×1080 cases had 0–0.6% exact-grid hit rate; native-grid cases 100% | **Supported conditionally**, not blanket removal |
| Lazy export AA samples | 1× export avoids **12 MiB** of four unused AA sample textures across a 3-slot ring | **Strongly supported** |
| Adaptive initial-iteration probe removal | 3/48 survey views selected >350; forcing 350 raised predicted work by **+7.2%, +25.3%, +33.1%** | **Do not remove**; cache/gate instead |
| Continuation restart from n=0 | Resume-state surrogate reduced work **7.6–9.8%**, identical classifications | **Real waste**, but requires state-preserving redesign |
| Deep workspace release after easy frame | Up to ~44.1 MiB standard / ~176.5 MiB high can remain resident | **Memory-lifetime candidate**; real GPU allocation cost unmeasured |
## Results in detail
### 1. Recolor numeric-history copy
Current recolor promotes color history and also copies unchanged `meta/smooth` numeric history. B separates the two responsibilities.
Exact avoided copy per accepted recolor:
- fast cap: 4 MiB
- standard cap: 8 MiB
- high cap: 32 MiB
At 20 Hz color cycling, that corresponds to 80 / 160 / 640 MiB/s of avoidable numeric-history copying at the respective caps. This is a traffic count, not a GPU-time measurement.
Variant: `b01-recolor-no-numeric-copy.html`
### 2. Unread CPU RGBA retention
`cpuFrameRgba` is assigned after CPU fallback rendering but has no subsequent reader in production. B drops the assignment.
Avoided retained JS memory at cap:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
Variant: `b02-cpu-no-rgba-retention.html`
### 3. Exact-interior pre-primary unknown-stat pass
When exact tile certification only partially covers the frame, production immediately runs a full-frame unknown-stat scan and synchronization before the primary numerical path. The partial-path result is not consumed; the primary path computes the required statistics later.
B removes this pre-primary stats scan while preserving the fully-certified shortcut.
Minimum avoided `fieldMeta` read per pass:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
This excludes atomics, dispatch overhead, and synchronization cost, so it is a lower bound on removed work.
Variant: `b03-certify-no-preprimary-stats.html`
### 4. Strict direct periodicity-state updates
The strict path does not accept heuristic periodicity, yet production still maintains Brent state. A CPU surrogate compared identical numerical output with and without those state updates.
Across six fixed views, B was not consistently faster: median elapsed change was **+2.82%** (slower), mean +5.47%. One deep case improved by 1.48%, but the result is not robust.
Conclusion: semantically redundant state exists, but this A/B test does **not** justify touching it for performance. GPU register pressure could behave differently; real WebGPU measurement would be needed.
Variant: `b04-strict-no-cycle-state.html`
### 5. FAST coverage repair
Production performs a second full-frame FAST pass intended to fill `reason=0` holes after tiled FAST rendering.
Coverage simulation tested 36 combinations of dimensions, iteration budgets and symmetry modes; all produced **0 holes**. A second semantic randomized test seeded mixed proven/unknown metadata over 200 cases and also found **0 cases with holes**.
Representative work avoided if the repair pass is unnecessary:
- 1365×768: 1,050,624 shader invocations; at least 4,193,280 bytes of metadata reads
- 2731×1536: 4,202,496 invocations; at least 16,779,264 bytes of metadata reads
Variant: `b05-fast-no-coverage-repair.html`
### 6. Idle color-auto snapshot
Starting color cycling while no render is in progress allocates/copies `colorSource`, but idle recolor reads the current numerical field and does not consume that snapshot. Render-time snapshot behavior is preserved in B.
Avoided idle snapshot residency:
- fast: 6 MiB
- standard: 12 MiB
- high: 48 MiB
Variant: `b06-idle-colorauto-no-snapshot.html`
### 7. Export workspace lifetime
The export ring remains allocated after export. With 512-sized slots and ring depth 3, the tested maximum workspace is approximately **24 MiB**. B destroys it in `finally`, reducing post-export residency to zero. The cost is allocation of three slots on a subsequent export.
This is a clear memory trade-off, not a proven speed win.
Variant: `b07-export-release-workspace.html`
### 8. Exact-interior spatial gate
The expensive BigInt tile proof only needs to run when the viewport can intersect the main cardioid or period-2 bulb. B uses a conservative bounding-box gate before tile proof.
The cardioid gate was corrected during the A/B test: the main cardioid reaches `x = 3/8`, so the conservative box is `x ∈ [-3/4, 3/8]`, `y ∈ [-2/3, 2/3]`. The period-2 bulb box is `x ∈ [-5/4, -3/4]`, `y ∈ [-1/4, 1/4]`.
Random validation: 240 views, 160 gate-negative, **0 false negatives** versus the current exact tile proof.
Fixed-corpus aggregate CPU proof time:
- 1024×768: A 41.64 ms → B 25.16 ms, about **39.6% reduction**
- 2048²: A 177.70 ms → B 117.08 ms, about **34.1% reduction**
The largest wins were period-3 views that cannot intersect either analytically certifiable component: about 7–10 ms saved at 1024×768 and 25–30 ms at 2048².
Variant: `b08-exact-interior-spatial-gate.html`
### 9. Conditional numeric-history residency
Exact-grid history reuse requires render-pixel translation to line up exactly with CSS/device motion. A deterministic horizontal-pan survey over offsets -500…500 showed:
| Configuration | Reuse hits | A residency | B policy |
|---|---:|---:|---:|
| 1920×1080 fast → 965×543 | 0.2% | ~4.0 MiB | 0 |
| 1920×1080 standard → 1365×768 | 0.6% | ~8.0 MiB | 0 |
| 1920×1080 high → 2731×1536 | 0% | ~32.0 MiB | 0 |
| 1024×768 native standard | 100% | 6 MiB | keep |
| 1024×768 DPR2 native ratio | 100% | 24 MiB | keep |
B keeps numeric history only for simple reduced render/CSS scale ratios (denominator ≤ 8). This preserves the clearly useful native-grid cases while avoiding permanent allocation where reuse is nearly impossible.
Variant: `b09-conditional-numeric-history.html`
### 10. Lazy export AA sample textures
Production allocates four 512² RGBA8 AA sample textures per export slot even for 1× export, although only 2× AA uses them. Four textures are 4 MiB per slot; with ring depth 3 this is **12 MiB** of unnecessary 1× export allocation.
B creates AA sample textures lazily only for 2× AA. Regression passes.
Variant: `b10-export-aa-samples-lazy.html`
### 11. Adaptive initial-iteration probe
The fixed six-view corpus always selected the base 350 iterations, at a measured probe cost of roughly 1.4–17.4 ms per view. That initially suggested deletion.
A broader deterministic survey of 48 views disproved the blanket-removal hypothesis. Three views (6.25%) selected more than 350. Forcing 350 increased the surrogate work metric by:
- +7.21%
- +33.14%
- +25.25%
Conclusion: **do not delete this probe globally**. Better candidates are caching by nearby view, a cheaper trigger before the full probe, or reducing sample count where confidence is high.
### 12. Continuation replay
Current continuation reconstructs active state from `n=0`; a state-resume surrogate instead kept the base-350 state and continued to 4096. The classifications matched in every tested view.
Work reduction by view: **7.63% to 9.79%**, median about **7.87%**.
This is genuine redundant work, but production perturbation/reference changes complicate state validity. It is a redesign candidate, not a safe line deletion.
## Recommended next actions
### High-confidence changes to consider integrating
1. Separate recolor history promotion from numeric-history copy.
2. Stop retaining unused CPU RGBA output.
3. Remove the partial exact-interior pre-primary stats pass.
4. Remove/disable FAST coverage repair after adding a diagnostic hole counter for one release cycle.
5. Avoid idle color-auto snapshots.
6. Release export workspace after export, subject to acceptable next-export allocation latency.
7. Add the conservative exact-interior spatial gate.
8. Make export AA sample textures lazy.
### Conditional changes
- Allocate numeric history only when exact-grid reuse has realistic geometry.
- Release large deep workspaces after returning to an easy frame or after a memory-pressure/idle policy; do not churn them every frame.
### Do not remove based on this test
- Adaptive initial-iteration probe.
- Strict periodicity state solely for speed; no measurable CPU benefit was established.
### Requires structural work
- Preserve valid continuation state so OPERATION_LIMIT pixels do not replay their first ~350 iterations.
## Reproduction
From project root:
```bash
node tests/regression.mjs
node experiments/ab/ab_test.mjs
node experiments/ab/supplement.mjs
node experiments/ab/variant_check.mjs
```
Raw data:
- `docs/performance-cost-audit/ab/results.json`
- `docs/performance-cost-audit/ab/supplement.json`
All B variants are experimental. Production was not modified by this A/B test.