food_chain/docs/performance-cost-audit/ab/README.md
2026-09-29 23:45:12 +09:00

11 KiB
Raw Permalink Blame History

Mandelbrot viewer performance-cost A/B test

Date: 2026-09-29

Scope

This A/B test follows the performance-cost audit. Production index.html was not changed. A is the current production implementation; B variants live under experiments/ab/variants/ and disable, defer, or gate one candidate operation at a time.

Production SHA-256 before/after testing:

544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a

The existing regression suite passes on production. The combined high-confidence B variant and the individual structural B variants used for gating/lazy allocation also pass the same regression suite; all retain 19 WGSL kernels.

Measurement limitation

A usable WebGPU backend could not be started in the container (Chromium/SwiftShader/EGL). Therefore this report deliberately separates:

  1. CPU measured time/work: exact-interior BigInt proof, adaptive-iteration probe, CPU surrogates for strict-cycle and continuation replay.
  2. Exact traffic/residency counts from the production buffer sizes and dispatch extents: copies, full-frame scans, allocations.
  3. Structural/regression validation: B variants parse and pass the production regression suite.

No result below is presented as measured GPU frame time unless explicitly stated. Real-device WebGPU p50/p95 remains a follow-up requirement.

Summary

Candidate A/B result Decision from this test
Recolor numeric-history copy B removes 8 B/pixel copy with unchanged numeric field Strongly supported
Unread cpuFrameRgba retention B removes 4 B/pixel retained JS data; no reader exists Strongly supported
Pre-primary unknown stats after exact interior B removes one full-frame fieldMeta scan; primary path computes stats later Strongly supported
Strict direct periodicity-state updates CPU surrogate median B change +2.82% slower Do not prioritize for speed
FAST full-screen coverage repair 36 layout cases + 200 semantic randomized cases: 0 holes Strongly supported, real GPU A/B still desirable
Idle color-auto colorSource snapshot B avoids 12 B/pixel idle snapshot; idle recolor does not consume it Strongly supported
Export workspace lifetime B releases ~24 MiB max workspace after export Memory win, trades for reallocation on next export
Exact-interior spatial gate Fixed corpus proof CPU time reduced ~40% at 1024×768 and ~34% at 2048²; 240 random views: 0 gate false-negatives Strongly supported
Conditional numeric-history residency Budget-scaled 1920×1080 cases had 0–0.6% exact-grid hit rate; native-grid cases 100% Supported conditionally, not blanket removal
Lazy export AA samples 1× export avoids 12 MiB of four unused AA sample textures across a 3-slot ring Strongly supported
Adaptive initial-iteration probe removal 3/48 survey views selected >350; forcing 350 raised predicted work by +7.2%, +25.3%, +33.1% Do not remove; cache/gate instead
Continuation restart from n=0 Resume-state surrogate reduced work 7.6–9.8%, identical classifications Real waste, but requires state-preserving redesign
Deep workspace release after easy frame Up to ~44.1 MiB standard / ~176.5 MiB high can remain resident Memory-lifetime candidate; real GPU allocation cost unmeasured

Results in detail

1. Recolor numeric-history copy

Current recolor promotes color history and also copies unchanged meta/smooth numeric history. B separates the two responsibilities.

Exact avoided copy per accepted recolor:

  • fast cap: 4 MiB
  • standard cap: 8 MiB
  • high cap: 32 MiB

At 20 Hz color cycling, that corresponds to 80 / 160 / 640 MiB/s of avoidable numeric-history copying at the respective caps. This is a traffic count, not a GPU-time measurement.

Variant: b01-recolor-no-numeric-copy.html

2. Unread CPU RGBA retention

cpuFrameRgba is assigned after CPU fallback rendering but has no subsequent reader in production. B drops the assignment.

Avoided retained JS memory at cap:

  • fast: 2 MiB
  • standard: 4 MiB
  • high: 16 MiB

Variant: b02-cpu-no-rgba-retention.html

3. Exact-interior pre-primary unknown-stat pass

When exact tile certification only partially covers the frame, production immediately runs a full-frame unknown-stat scan and synchronization before the primary numerical path. The partial-path result is not consumed; the primary path computes the required statistics later.

B removes this pre-primary stats scan while preserving the fully-certified shortcut.

Minimum avoided fieldMeta read per pass:

  • fast: 2 MiB
  • standard: 4 MiB
  • high: 16 MiB

This excludes atomics, dispatch overhead, and synchronization cost, so it is a lower bound on removed work.

Variant: b03-certify-no-preprimary-stats.html

4. Strict direct periodicity-state updates

The strict path does not accept heuristic periodicity, yet production still maintains Brent state. A CPU surrogate compared identical numerical output with and without those state updates.

Across six fixed views, B was not consistently faster: median elapsed change was +2.82% (slower), mean +5.47%. One deep case improved by 1.48%, but the result is not robust.

Conclusion: semantically redundant state exists, but this A/B test does not justify touching it for performance. GPU register pressure could behave differently; real WebGPU measurement would be needed.

Variant: b04-strict-no-cycle-state.html

5. FAST coverage repair

Production performs a second full-frame FAST pass intended to fill reason=0 holes after tiled FAST rendering.

Coverage simulation tested 36 combinations of dimensions, iteration budgets and symmetry modes; all produced 0 holes. A second semantic randomized test seeded mixed proven/unknown metadata over 200 cases and also found 0 cases with holes.

Representative work avoided if the repair pass is unnecessary:

  • 1365×768: 1,050,624 shader invocations; at least 4,193,280 bytes of metadata reads
  • 2731×1536: 4,202,496 invocations; at least 16,779,264 bytes of metadata reads

Variant: b05-fast-no-coverage-repair.html

6. Idle color-auto snapshot

Starting color cycling while no render is in progress allocates/copies colorSource, but idle recolor reads the current numerical field and does not consume that snapshot. Render-time snapshot behavior is preserved in B.

Avoided idle snapshot residency:

  • fast: 6 MiB
  • standard: 12 MiB
  • high: 48 MiB

Variant: b06-idle-colorauto-no-snapshot.html

7. Export workspace lifetime

The export ring remains allocated after export. With 512-sized slots and ring depth 3, the tested maximum workspace is approximately 24 MiB. B destroys it in finally, reducing post-export residency to zero. The cost is allocation of three slots on a subsequent export.

This is a clear memory trade-off, not a proven speed win.

Variant: b07-export-release-workspace.html

8. Exact-interior spatial gate

The expensive BigInt tile proof only needs to run when the viewport can intersect the main cardioid or period-2 bulb. B uses a conservative bounding-box gate before tile proof.

The cardioid gate was corrected during the A/B test: the main cardioid reaches x = 3/8, so the conservative box is x ∈ [-3/4, 3/8], y ∈ [-2/3, 2/3]. The period-2 bulb box is x ∈ [-5/4, -3/4], y ∈ [-1/4, 1/4].

Random validation: 240 views, 160 gate-negative, 0 false negatives versus the current exact tile proof.

Fixed-corpus aggregate CPU proof time:

  • 1024×768: A 41.64 ms → B 25.16 ms, about 39.6% reduction
  • 2048²: A 177.70 ms → B 117.08 ms, about 34.1% reduction

The largest wins were period-3 views that cannot intersect either analytically certifiable component: about 7–10 ms saved at 1024×768 and 25–30 ms at 2048².

Variant: b08-exact-interior-spatial-gate.html

9. Conditional numeric-history residency

Exact-grid history reuse requires render-pixel translation to line up exactly with CSS/device motion. A deterministic horizontal-pan survey over offsets -500…500 showed:

Configuration Reuse hits A residency B policy
1920×1080 fast → 965×543 0.2% ~4.0 MiB 0
1920×1080 standard → 1365×768 0.6% ~8.0 MiB 0
1920×1080 high → 2731×1536 0% ~32.0 MiB 0
1024×768 native standard 100% 6 MiB keep
1024×768 DPR2 native ratio 100% 24 MiB keep

B keeps numeric history only for simple reduced render/CSS scale ratios (denominator ≤ 8). This preserves the clearly useful native-grid cases while avoiding permanent allocation where reuse is nearly impossible.

Variant: b09-conditional-numeric-history.html

10. Lazy export AA sample textures

Production allocates four 512² RGBA8 AA sample textures per export slot even for 1× export, although only 2× AA uses them. Four textures are 4 MiB per slot; with ring depth 3 this is 12 MiB of unnecessary 1× export allocation.

B creates AA sample textures lazily only for 2× AA. Regression passes.

Variant: b10-export-aa-samples-lazy.html

11. Adaptive initial-iteration probe

The fixed six-view corpus always selected the base 350 iterations, at a measured probe cost of roughly 1.4–17.4 ms per view. That initially suggested deletion.

A broader deterministic survey of 48 views disproved the blanket-removal hypothesis. Three views (6.25%) selected more than 350. Forcing 350 increased the surrogate work metric by:

  • +7.21%
  • +33.14%
  • +25.25%

Conclusion: do not delete this probe globally. Better candidates are caching by nearby view, a cheaper trigger before the full probe, or reducing sample count where confidence is high.

12. Continuation replay

Current continuation reconstructs active state from n=0; a state-resume surrogate instead kept the base-350 state and continued to 4096. The classifications matched in every tested view.

Work reduction by view: 7.63% to 9.79%, median about 7.87%.

This is genuine redundant work, but production perturbation/reference changes complicate state validity. It is a redesign candidate, not a safe line deletion.

High-confidence changes to consider integrating

  1. Separate recolor history promotion from numeric-history copy.
  2. Stop retaining unused CPU RGBA output.
  3. Remove the partial exact-interior pre-primary stats pass.
  4. Remove/disable FAST coverage repair after adding a diagnostic hole counter for one release cycle.
  5. Avoid idle color-auto snapshots.
  6. Release export workspace after export, subject to acceptable next-export allocation latency.
  7. Add the conservative exact-interior spatial gate.
  8. Make export AA sample textures lazy.

Conditional changes

  • Allocate numeric history only when exact-grid reuse has realistic geometry.
  • Release large deep workspaces after returning to an easy frame or after a memory-pressure/idle policy; do not churn them every frame.

Do not remove based on this test

  • Adaptive initial-iteration probe.
  • Strict periodicity state solely for speed; no measurable CPU benefit was established.

Requires structural work

  • Preserve valid continuation state so OPERATION_LIMIT pixels do not replay their first ~350 iterations.

Reproduction

From project root:

node tests/regression.mjs
node experiments/ab/ab_test.mjs
node experiments/ab/supplement.mjs
node experiments/ab/variant_check.mjs

Raw data:

  • docs/performance-cost-audit/ab/results.json
  • docs/performance-cost-audit/ab/supplement.json

All B variants are experimental. Production was not modified by this A/B test.