11 KiB
Mandelbrot viewer performance-cost A/B test
Date: 2026-09-29
Scope
This A/B test follows the performance-cost audit. Production index.html was not changed. A is the current production implementation; B variants live under experiments/ab/variants/ and disable, defer, or gate one candidate operation at a time.
Production SHA-256 before/after testing:
544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a
The existing regression suite passes on production. The combined high-confidence B variant and the individual structural B variants used for gating/lazy allocation also pass the same regression suite; all retain 19 WGSL kernels.
Measurement limitation
A usable WebGPU backend could not be started in the container (Chromium/SwiftShader/EGL). Therefore this report deliberately separates:
- CPU measured time/work: exact-interior BigInt proof, adaptive-iteration probe, CPU surrogates for strict-cycle and continuation replay.
- Exact traffic/residency counts from the production buffer sizes and dispatch extents: copies, full-frame scans, allocations.
- Structural/regression validation: B variants parse and pass the production regression suite.
No result below is presented as measured GPU frame time unless explicitly stated. Real-device WebGPU p50/p95 remains a follow-up requirement.
Summary
| Candidate | A/B result | Decision from this test |
|---|---|---|
| Recolor numeric-history copy | B removes 8 B/pixel copy with unchanged numeric field | Strongly supported |
Unread cpuFrameRgba retention |
B removes 4 B/pixel retained JS data; no reader exists | Strongly supported |
| Pre-primary unknown stats after exact interior | B removes one full-frame fieldMeta scan; primary path computes stats later |
Strongly supported |
| Strict direct periodicity-state updates | CPU surrogate median B change +2.82% slower | Do not prioritize for speed |
| FAST full-screen coverage repair | 36 layout cases + 200 semantic randomized cases: 0 holes | Strongly supported, real GPU A/B still desirable |
Idle color-auto colorSource snapshot |
B avoids 12 B/pixel idle snapshot; idle recolor does not consume it | Strongly supported |
| Export workspace lifetime | B releases ~24 MiB max workspace after export | Memory win, trades for reallocation on next export |
| Exact-interior spatial gate | Fixed corpus proof CPU time reduced ~40% at 1024×768 and ~34% at 2048²; 240 random views: 0 gate false-negatives | Strongly supported |
| Conditional numeric-history residency | Budget-scaled 1920×1080 cases had 0–0.6% exact-grid hit rate; native-grid cases 100% | Supported conditionally, not blanket removal |
| Lazy export AA samples | 1× export avoids 12 MiB of four unused AA sample textures across a 3-slot ring | Strongly supported |
| Adaptive initial-iteration probe removal | 3/48 survey views selected >350; forcing 350 raised predicted work by +7.2%, +25.3%, +33.1% | Do not remove; cache/gate instead |
| Continuation restart from n=0 | Resume-state surrogate reduced work 7.6–9.8%, identical classifications | Real waste, but requires state-preserving redesign |
| Deep workspace release after easy frame | Up to ~44.1 MiB standard / ~176.5 MiB high can remain resident | Memory-lifetime candidate; real GPU allocation cost unmeasured |
Results in detail
1. Recolor numeric-history copy
Current recolor promotes color history and also copies unchanged meta/smooth numeric history. B separates the two responsibilities.
Exact avoided copy per accepted recolor:
- fast cap: 4 MiB
- standard cap: 8 MiB
- high cap: 32 MiB
At 20 Hz color cycling, that corresponds to 80 / 160 / 640 MiB/s of avoidable numeric-history copying at the respective caps. This is a traffic count, not a GPU-time measurement.
Variant: b01-recolor-no-numeric-copy.html
2. Unread CPU RGBA retention
cpuFrameRgba is assigned after CPU fallback rendering but has no subsequent reader in production. B drops the assignment.
Avoided retained JS memory at cap:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
Variant: b02-cpu-no-rgba-retention.html
3. Exact-interior pre-primary unknown-stat pass
When exact tile certification only partially covers the frame, production immediately runs a full-frame unknown-stat scan and synchronization before the primary numerical path. The partial-path result is not consumed; the primary path computes the required statistics later.
B removes this pre-primary stats scan while preserving the fully-certified shortcut.
Minimum avoided fieldMeta read per pass:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
This excludes atomics, dispatch overhead, and synchronization cost, so it is a lower bound on removed work.
Variant: b03-certify-no-preprimary-stats.html
4. Strict direct periodicity-state updates
The strict path does not accept heuristic periodicity, yet production still maintains Brent state. A CPU surrogate compared identical numerical output with and without those state updates.
Across six fixed views, B was not consistently faster: median elapsed change was +2.82% (slower), mean +5.47%. One deep case improved by 1.48%, but the result is not robust.
Conclusion: semantically redundant state exists, but this A/B test does not justify touching it for performance. GPU register pressure could behave differently; real WebGPU measurement would be needed.
Variant: b04-strict-no-cycle-state.html
5. FAST coverage repair
Production performs a second full-frame FAST pass intended to fill reason=0 holes after tiled FAST rendering.
Coverage simulation tested 36 combinations of dimensions, iteration budgets and symmetry modes; all produced 0 holes. A second semantic randomized test seeded mixed proven/unknown metadata over 200 cases and also found 0 cases with holes.
Representative work avoided if the repair pass is unnecessary:
- 1365×768: 1,050,624 shader invocations; at least 4,193,280 bytes of metadata reads
- 2731×1536: 4,202,496 invocations; at least 16,779,264 bytes of metadata reads
Variant: b05-fast-no-coverage-repair.html
6. Idle color-auto snapshot
Starting color cycling while no render is in progress allocates/copies colorSource, but idle recolor reads the current numerical field and does not consume that snapshot. Render-time snapshot behavior is preserved in B.
Avoided idle snapshot residency:
- fast: 6 MiB
- standard: 12 MiB
- high: 48 MiB
Variant: b06-idle-colorauto-no-snapshot.html
7. Export workspace lifetime
The export ring remains allocated after export. With 512-sized slots and ring depth 3, the tested maximum workspace is approximately 24 MiB. B destroys it in finally, reducing post-export residency to zero. The cost is allocation of three slots on a subsequent export.
This is a clear memory trade-off, not a proven speed win.
Variant: b07-export-release-workspace.html
8. Exact-interior spatial gate
The expensive BigInt tile proof only needs to run when the viewport can intersect the main cardioid or period-2 bulb. B uses a conservative bounding-box gate before tile proof.
The cardioid gate was corrected during the A/B test: the main cardioid reaches x = 3/8, so the conservative box is x ∈ [-3/4, 3/8], y ∈ [-2/3, 2/3]. The period-2 bulb box is x ∈ [-5/4, -3/4], y ∈ [-1/4, 1/4].
Random validation: 240 views, 160 gate-negative, 0 false negatives versus the current exact tile proof.
Fixed-corpus aggregate CPU proof time:
- 1024×768: A 41.64 ms → B 25.16 ms, about 39.6% reduction
- 2048²: A 177.70 ms → B 117.08 ms, about 34.1% reduction
The largest wins were period-3 views that cannot intersect either analytically certifiable component: about 7–10 ms saved at 1024×768 and 25–30 ms at 2048².
Variant: b08-exact-interior-spatial-gate.html
9. Conditional numeric-history residency
Exact-grid history reuse requires render-pixel translation to line up exactly with CSS/device motion. A deterministic horizontal-pan survey over offsets -500…500 showed:
| Configuration | Reuse hits | A residency | B policy |
|---|---|---|---|
| 1920×1080 fast → 965×543 | 0.2% | ~4.0 MiB | 0 |
| 1920×1080 standard → 1365×768 | 0.6% | ~8.0 MiB | 0 |
| 1920×1080 high → 2731×1536 | 0% | ~32.0 MiB | 0 |
| 1024×768 native standard | 100% | 6 MiB | keep |
| 1024×768 DPR2 native ratio | 100% | 24 MiB | keep |
B keeps numeric history only for simple reduced render/CSS scale ratios (denominator ≤ 8). This preserves the clearly useful native-grid cases while avoiding permanent allocation where reuse is nearly impossible.
Variant: b09-conditional-numeric-history.html
10. Lazy export AA sample textures
Production allocates four 512² RGBA8 AA sample textures per export slot even for 1× export, although only 2× AA uses them. Four textures are 4 MiB per slot; with ring depth 3 this is 12 MiB of unnecessary 1× export allocation.
B creates AA sample textures lazily only for 2× AA. Regression passes.
Variant: b10-export-aa-samples-lazy.html
11. Adaptive initial-iteration probe
The fixed six-view corpus always selected the base 350 iterations, at a measured probe cost of roughly 1.4–17.4 ms per view. That initially suggested deletion.
A broader deterministic survey of 48 views disproved the blanket-removal hypothesis. Three views (6.25%) selected more than 350. Forcing 350 increased the surrogate work metric by:
- +7.21%
- +33.14%
- +25.25%
Conclusion: do not delete this probe globally. Better candidates are caching by nearby view, a cheaper trigger before the full probe, or reducing sample count where confidence is high.
12. Continuation replay
Current continuation reconstructs active state from n=0; a state-resume surrogate instead kept the base-350 state and continued to 4096. The classifications matched in every tested view.
Work reduction by view: 7.63% to 9.79%, median about 7.87%.
This is genuine redundant work, but production perturbation/reference changes complicate state validity. It is a redesign candidate, not a safe line deletion.
Recommended next actions
High-confidence changes to consider integrating
- Separate recolor history promotion from numeric-history copy.
- Stop retaining unused CPU RGBA output.
- Remove the partial exact-interior pre-primary stats pass.
- Remove/disable FAST coverage repair after adding a diagnostic hole counter for one release cycle.
- Avoid idle color-auto snapshots.
- Release export workspace after export, subject to acceptable next-export allocation latency.
- Add the conservative exact-interior spatial gate.
- Make export AA sample textures lazy.
Conditional changes
- Allocate numeric history only when exact-grid reuse has realistic geometry.
- Release large deep workspaces after returning to an easy frame or after a memory-pressure/idle policy; do not churn them every frame.
Do not remove based on this test
- Adaptive initial-iteration probe.
- Strict periodicity state solely for speed; no measurable CPU benefit was established.
Requires structural work
- Preserve valid continuation state so OPERATION_LIMIT pixels do not replay their first ~350 iterations.
Reproduction
From project root:
node tests/regression.mjs
node experiments/ab/ab_test.mjs
node experiments/ab/supplement.mjs
node experiments/ab/variant_check.mjs
Raw data:
docs/performance-cost-audit/ab/results.jsondocs/performance-cost-audit/ab/supplement.json
All B variants are experimental. Production was not modified by this A/B test.