This commit is contained in:
33333-33333 2026-09-29 23:45:12 +09:00
commit f5fc9b76ce
73 changed files with 32628 additions and 376 deletions

View file

@ -0,0 +1,137 @@
# Performance / Memory Cost Audit
## Purpose
Audit the integrated Mandelbrot viewer for processing that is low-value, redundant, or harmful from a speed / memory perspective. This audit intentionally **does not remove or alter any production behavior** in `index.html`.
Production source audited:
- `index.html`
- SHA-256: `544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a`
- Regression: PASS, 19 WGSL kernels
Measurements are CPU/static surrogates unless explicitly stated. Real WebGPU frame timing is not available in this container because the browser WebGPU backend cannot initialize successfully here.
## Conclusions
### A — High-confidence waste / harmful cost
| ID | Finding | Cost | Benefit observed | Confidence |
|---|---|---|---|---|
| A1 | Numeric history is recopied after every accepted recolor | 8 B/px copy per recolor | Numeric data did not change | Very high |
| A2 | `cpuFrameRgba` retains the assembled CPU fallback frame but is never read | 4 B/px JS heap retained | None | Very high |
| A3 | `certifyInterior()` runs a full-frame unknown-statistics pass and waits for GPU completion before primary compute | Full meta scan + queue synchronization | Partial proof/history-seed caller does not consume these stats | Very high |
| A4 | Strict direct rendering still updates Brent-cycle state even though periodicity classification is disabled | Per-iteration ALU/register state | None in strict mode | Very high |
| A5 | Export workspace survives export completion | Up to ~24 MiB GPU memory with 512 tiles / 3 slots | None while not exporting | Very high |
| A6 | Deep active/correction workspaces survive the hard frame that required them | active workspace up to ~44.1 MiB standard / ~176.5 MiB high; dense correction queue 4/16 MiB | None on later easy frames until reuse/resize | High |
### B — Strong optimization candidates, but keep until GPU A/B validation
| ID | Finding | Evidence | Assessment |
|---|---|---|---|
| B1 | Numeric history is always resident although exact-grid pan reuse is restrictive | 8 B/px resident. Under common budget-scaled resolutions, integer CSS pans 1–100 px produced 0/100 exact-grid reuse hits in the static alignment test | Likely over-allocated; make lazy/telemetry-driven rather than delete immediately |
| B2 | CPU exact cardioid/bulb tile proof runs for every numerical frame | 1024×768: ~4.8–6.1 ms median; 2048²: ~25–28 ms median. Four of six fixed corpus views certified 0 pixels | Add cheap spatial gate; direct backend especially needs A/B comparison |
| B3 | `fast-coverage-repair` runs a full-screen FAST pass after tiled coverage | Tile loops cover the full compute domain; mode 2 only repairs reason=0 untouched metadata | Likely redundant defensive pass; instrument reason=0 count before removal |
| B4 | `captureColorSource()` snapshots `meta+smooth` during rendering | 8 B/px copy per snapshot; provisional snapshots can occur every 125 ms. Snapshot allocation is 12 B/px | Useful only for live recoloring while render is in flight; gate more aggressively |
| B5 | Color-auto idle mode retains a color snapshot | standard ~12 MiB, high ~48 MiB at preset caps | Idle color-auto recolor path uses current field, not `recolorPublished()` | Snapshot should be render-only/on-demand |
| B6 | Adaptive initial-iteration probe uses high-precision scoring before every adaptive frame | Fixed corpus: ~1.3–12.5 ms and all six chose 350. Exploratory views can choose 512, so feature is not useless | Cache/gate/replace with cheaper trigger; do not delete blindly |
| B7 | Active continuation can replay already-performed iterations | initial OP_LIMIT session starts state at n=0; reference change resets progress to 0; recovered pixels can rebuild whole active session | Potentially large on the hardest views; add replay counters before redesign |
### C — Expensive but currently justified / not a deletion target
- General periodicity detection: CPU surrogate showed ~98% iteration reduction in period-3 interiors. It can add overhead on boundary-only views, but removal is not justified.
- Sparse numerical correction and reason bucketing: it adds full scans/queue construction, but isolates expensive DS/BigInt work. Needs GPU A/B before structural change.
- Stable texture reprojection: consumes one history texture but directly improves interaction latency.
- Reference cache / guided multi-reference: memory cost is bounded and prior evaluation showed reduced perturbation mismatch.
- Series approximation: small reference metadata cost with large deep-zoom iteration skip in prior evaluation.
## Key measurements
### Recolor numeric-history copy bandwidth
`commitHistory()` always calls `commitNumericHistory()`. Therefore a color-only update copies unchanged `meta` and `smooth` buffers (8 B/px). Color auto can request a pass every 50 ms while idle.
| preset cap | unchanged numeric copy / recolor | at 20 Hz |
|---|---:|---:|
| fast, 524,288 px | 4 MiB | 80 MiB/s |
| standard, 1,048,576 px | 8 MiB | 160 MiB/s |
| high, 4,194,304 px | 32 MiB | 640 MiB/s |
This is in addition to the actual color pass bandwidth.
### Color snapshot
`captureColorSource()` allocates/copies:
- meta: 4 B/px
- smooth: 4 B/px
- texture: 4 B/px
Resident cost = 12 B/px: ~6 MiB fast, 12 MiB standard, 48 MiB high at preset caps. A provisional snapshot copies 8 B/px and can be refreshed at up to 8 Hz, equivalent to ~32 / 64 / 256 MiB/s at the preset caps.
### Exact interior tile proof
Seven-run median on this host:
| view | 1024×768 | certified | 2048×2048 | certified |
|---|---:|---:|---:|---:|
| overview | ~6.1 ms | 7.3% | ~25.5 ms | 8.6% |
| period3 bulb | ~5.1 ms | 0% | ~26.5 ms | 0% |
| period3 core | ~4.8 ms | 0% | ~28.1 ms | 0% |
| seahorse boundary | ~6.0 ms | 0% | ~25.1 ms | 0% |
| seahorse deep | ~5.6 ms | 0% | ~26.1 ms | 0% |
The exact proof is still useful in views that overlap the main cardioid / period-2 bulb, especially on the perturbation path. The issue is unconditional invocation, not the proof mechanism itself.
### Exact-grid numeric-history reuse
The history seed requires identical span and an approximately integer render-pixel shift. When the resolution is pixel-budget limited, render pixels often do not align with CSS pixels.
Examples from the current resize rule, testing integer CSS horizontal pans of 1–100 px:
| screen / quality | render size | reuse hits |
|---|---:|---:|
| 1920×1080 / fast | 965×543 | 0/100 |
| 1920×1080 / standard | 1365×768 | 0/100 |
| 1920×1080 / high | 2731×1536 | 0/100 |
| 1024×768@1× / standard | 1024×768 | 100/100 |
This is a static alignment test, not real interaction telemetry. It supports making numeric history conditional rather than proving that it has no value.
### Persistent workspaces
- Sparse active workspace, worst full-active capacity: ~44.1 MiB standard, ~176.5 MiB high.
- At 10% active: ~5.64 MiB standard, ~22.51 MiB high.
- Dense deep correction queue: 4 MiB standard, 16 MiB high.
- Export workspace at 512 tile, depth 3: ~24 MiB retained after export.
- Non-AA export currently allocates four 512² sample textures, ~4 MiB per slot, although 1× rendering does not need the four AA sample textures.
## Static code locations
- `index.html:1299-1300`: pre-primary unknown stats + synchronization in `certifyInterior()`.
- `index.html:1383-1384`: full numeric-history copy.
- `index.html:1416`: `commitHistory()` always commits numeric history.
- `index.html:1822`: color-only recolor promotes history and therefore triggers numeric copy.
- `index.html:1718`: `state.cpuFrameRgba=rgba`; the state value has no reader.
- `index.html:162-163`: cycle state still updated in strict direct mode.
- `index.html:1349-1350`: tiled FAST coverage followed by full-screen coverage-repair pass.
- `index.html:1751-1756`: provisional color snapshots at 125 ms throttle.
- `index.html:1774`: old completed field snapshot at every numerical render start.
- `index.html:1866`: idle color-auto snapshot allocation.
- `index.html:1218`: active workspace persists unless another active allocation request causes a >4× shrink.
- `index.html:1217`: full-pixel dense correction queue persists once allocated.
- `index.html:1222-1226, 1898-1944`: export workspace lifetime extends beyond export job.
- `index.html:1306-1307, 1648, 1654`: active continuation reinitialization/rebuild can reset progress.
## Recommended instrumentation before deleting anything
1. Add counters for numeric-history `attempt / hit / reused pixels / bytes copied`.
2. Count reason-0 pixels immediately before `fast-coverage-repair`; if consistently zero, remove/disable that pass in an A/B branch.
3. Record `certifiedInteriorTiles` CPU ms and certified ratio per frame; gate when hit rate stays zero.
4. Record active-continuation `reinitCount` and `replayedPixelIterations`.
5. Record peak/resident bytes for active, correction, export, color snapshot and release-on-idle experiments.
6. Separate color-history promotion from numeric-history promotion and A/B color-auto frame time.
## Production integrity
No production deletion or behavioral modification was made by this audit. Regression was rerun after inspection and passed.

View file

@ -0,0 +1,234 @@
# Mandelbrot viewer performance-cost A/B test
Date: 2026-09-29
## Scope
This A/B test follows the performance-cost audit. Production `index.html` was **not changed**. A is the current production implementation; B variants live under `experiments/ab/variants/` and disable, defer, or gate one candidate operation at a time.
Production SHA-256 before/after testing:
`544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a`
The existing regression suite passes on production. The combined high-confidence B variant and the individual structural B variants used for gating/lazy allocation also pass the same regression suite; all retain 19 WGSL kernels.
## Measurement limitation
A usable WebGPU backend could not be started in the container (Chromium/SwiftShader/EGL). Therefore this report deliberately separates:
1. **CPU measured time/work**: exact-interior BigInt proof, adaptive-iteration probe, CPU surrogates for strict-cycle and continuation replay.
2. **Exact traffic/residency counts from the production buffer sizes and dispatch extents**: copies, full-frame scans, allocations.
3. **Structural/regression validation**: B variants parse and pass the production regression suite.
No result below is presented as measured GPU frame time unless explicitly stated. Real-device WebGPU p50/p95 remains a follow-up requirement.
## Summary
| Candidate | A/B result | Decision from this test |
|---|---|---|
| Recolor numeric-history copy | B removes 8 B/pixel copy with unchanged numeric field | **Strongly supported** |
| Unread `cpuFrameRgba` retention | B removes 4 B/pixel retained JS data; no reader exists | **Strongly supported** |
| Pre-primary unknown stats after exact interior | B removes one full-frame `fieldMeta` scan; primary path computes stats later | **Strongly supported** |
| Strict direct periodicity-state updates | CPU surrogate median B change **+2.82% slower** | **Do not prioritize for speed** |
| FAST full-screen coverage repair | 36 layout cases + 200 semantic randomized cases: **0 holes** | **Strongly supported**, real GPU A/B still desirable |
| Idle color-auto `colorSource` snapshot | B avoids 12 B/pixel idle snapshot; idle recolor does not consume it | **Strongly supported** |
| Export workspace lifetime | B releases ~24 MiB max workspace after export | **Memory win**, trades for reallocation on next export |
| Exact-interior spatial gate | Fixed corpus proof CPU time reduced ~40% at 1024×768 and ~34% at 2048²; 240 random views: 0 gate false-negatives | **Strongly supported** |
| Conditional numeric-history residency | Budget-scaled 1920×1080 cases had 0–0.6% exact-grid hit rate; native-grid cases 100% | **Supported conditionally**, not blanket removal |
| Lazy export AA samples | 1× export avoids **12 MiB** of four unused AA sample textures across a 3-slot ring | **Strongly supported** |
| Adaptive initial-iteration probe removal | 3/48 survey views selected >350; forcing 350 raised predicted work by **+7.2%, +25.3%, +33.1%** | **Do not remove**; cache/gate instead |
| Continuation restart from n=0 | Resume-state surrogate reduced work **7.6–9.8%**, identical classifications | **Real waste**, but requires state-preserving redesign |
| Deep workspace release after easy frame | Up to ~44.1 MiB standard / ~176.5 MiB high can remain resident | **Memory-lifetime candidate**; real GPU allocation cost unmeasured |
## Results in detail
### 1. Recolor numeric-history copy
Current recolor promotes color history and also copies unchanged `meta/smooth` numeric history. B separates the two responsibilities.
Exact avoided copy per accepted recolor:
- fast cap: 4 MiB
- standard cap: 8 MiB
- high cap: 32 MiB
At 20 Hz color cycling, that corresponds to 80 / 160 / 640 MiB/s of avoidable numeric-history copying at the respective caps. This is a traffic count, not a GPU-time measurement.
Variant: `b01-recolor-no-numeric-copy.html`
### 2. Unread CPU RGBA retention
`cpuFrameRgba` is assigned after CPU fallback rendering but has no subsequent reader in production. B drops the assignment.
Avoided retained JS memory at cap:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
Variant: `b02-cpu-no-rgba-retention.html`
### 3. Exact-interior pre-primary unknown-stat pass
When exact tile certification only partially covers the frame, production immediately runs a full-frame unknown-stat scan and synchronization before the primary numerical path. The partial-path result is not consumed; the primary path computes the required statistics later.
B removes this pre-primary stats scan while preserving the fully-certified shortcut.
Minimum avoided `fieldMeta` read per pass:
- fast: 2 MiB
- standard: 4 MiB
- high: 16 MiB
This excludes atomics, dispatch overhead, and synchronization cost, so it is a lower bound on removed work.
Variant: `b03-certify-no-preprimary-stats.html`
### 4. Strict direct periodicity-state updates
The strict path does not accept heuristic periodicity, yet production still maintains Brent state. A CPU surrogate compared identical numerical output with and without those state updates.
Across six fixed views, B was not consistently faster: median elapsed change was **+2.82%** (slower), mean +5.47%. One deep case improved by 1.48%, but the result is not robust.
Conclusion: semantically redundant state exists, but this A/B test does **not** justify touching it for performance. GPU register pressure could behave differently; real WebGPU measurement would be needed.
Variant: `b04-strict-no-cycle-state.html`
### 5. FAST coverage repair
Production performs a second full-frame FAST pass intended to fill `reason=0` holes after tiled FAST rendering.
Coverage simulation tested 36 combinations of dimensions, iteration budgets and symmetry modes; all produced **0 holes**. A second semantic randomized test seeded mixed proven/unknown metadata over 200 cases and also found **0 cases with holes**.
Representative work avoided if the repair pass is unnecessary:
- 1365×768: 1,050,624 shader invocations; at least 4,193,280 bytes of metadata reads
- 2731×1536: 4,202,496 invocations; at least 16,779,264 bytes of metadata reads
Variant: `b05-fast-no-coverage-repair.html`
### 6. Idle color-auto snapshot
Starting color cycling while no render is in progress allocates/copies `colorSource`, but idle recolor reads the current numerical field and does not consume that snapshot. Render-time snapshot behavior is preserved in B.
Avoided idle snapshot residency:
- fast: 6 MiB
- standard: 12 MiB
- high: 48 MiB
Variant: `b06-idle-colorauto-no-snapshot.html`
### 7. Export workspace lifetime
The export ring remains allocated after export. With 512-sized slots and ring depth 3, the tested maximum workspace is approximately **24 MiB**. B destroys it in `finally`, reducing post-export residency to zero. The cost is allocation of three slots on a subsequent export.
This is a clear memory trade-off, not a proven speed win.
Variant: `b07-export-release-workspace.html`
### 8. Exact-interior spatial gate
The expensive BigInt tile proof only needs to run when the viewport can intersect the main cardioid or period-2 bulb. B uses a conservative bounding-box gate before tile proof.
The cardioid gate was corrected during the A/B test: the main cardioid reaches `x = 3/8`, so the conservative box is `x ∈ [-3/4, 3/8]`, `y ∈ [-2/3, 2/3]`. The period-2 bulb box is `x ∈ [-5/4, -3/4]`, `y ∈ [-1/4, 1/4]`.
Random validation: 240 views, 160 gate-negative, **0 false negatives** versus the current exact tile proof.
Fixed-corpus aggregate CPU proof time:
- 1024×768: A 41.64 ms → B 25.16 ms, about **39.6% reduction**
- 2048²: A 177.70 ms → B 117.08 ms, about **34.1% reduction**
The largest wins were period-3 views that cannot intersect either analytically certifiable component: about 7–10 ms saved at 1024×768 and 25–30 ms at 2048².
Variant: `b08-exact-interior-spatial-gate.html`
### 9. Conditional numeric-history residency
Exact-grid history reuse requires render-pixel translation to line up exactly with CSS/device motion. A deterministic horizontal-pan survey over offsets -500…500 showed:
| Configuration | Reuse hits | A residency | B policy |
|---|---:|---:|---:|
| 1920×1080 fast → 965×543 | 0.2% | ~4.0 MiB | 0 |
| 1920×1080 standard → 1365×768 | 0.6% | ~8.0 MiB | 0 |
| 1920×1080 high → 2731×1536 | 0% | ~32.0 MiB | 0 |
| 1024×768 native standard | 100% | 6 MiB | keep |
| 1024×768 DPR2 native ratio | 100% | 24 MiB | keep |
B keeps numeric history only for simple reduced render/CSS scale ratios (denominator ≤ 8). This preserves the clearly useful native-grid cases while avoiding permanent allocation where reuse is nearly impossible.
Variant: `b09-conditional-numeric-history.html`
### 10. Lazy export AA sample textures
Production allocates four 512² RGBA8 AA sample textures per export slot even for 1× export, although only 2× AA uses them. Four textures are 4 MiB per slot; with ring depth 3 this is **12 MiB** of unnecessary 1× export allocation.
B creates AA sample textures lazily only for 2× AA. Regression passes.
Variant: `b10-export-aa-samples-lazy.html`
### 11. Adaptive initial-iteration probe
The fixed six-view corpus always selected the base 350 iterations, at a measured probe cost of roughly 1.4–17.4 ms per view. That initially suggested deletion.
A broader deterministic survey of 48 views disproved the blanket-removal hypothesis. Three views (6.25%) selected more than 350. Forcing 350 increased the surrogate work metric by:
- +7.21%
- +33.14%
- +25.25%
Conclusion: **do not delete this probe globally**. Better candidates are caching by nearby view, a cheaper trigger before the full probe, or reducing sample count where confidence is high.
### 12. Continuation replay
Current continuation reconstructs active state from `n=0`; a state-resume surrogate instead kept the base-350 state and continued to 4096. The classifications matched in every tested view.
Work reduction by view: **7.63% to 9.79%**, median about **7.87%**.
This is genuine redundant work, but production perturbation/reference changes complicate state validity. It is a redesign candidate, not a safe line deletion.
## Recommended next actions
### High-confidence changes to consider integrating
1. Separate recolor history promotion from numeric-history copy.
2. Stop retaining unused CPU RGBA output.
3. Remove the partial exact-interior pre-primary stats pass.
4. Remove/disable FAST coverage repair after adding a diagnostic hole counter for one release cycle.
5. Avoid idle color-auto snapshots.
6. Release export workspace after export, subject to acceptable next-export allocation latency.
7. Add the conservative exact-interior spatial gate.
8. Make export AA sample textures lazy.
### Conditional changes
- Allocate numeric history only when exact-grid reuse has realistic geometry.
- Release large deep workspaces after returning to an easy frame or after a memory-pressure/idle policy; do not churn them every frame.
### Do not remove based on this test
- Adaptive initial-iteration probe.
- Strict periodicity state solely for speed; no measurable CPU benefit was established.
### Requires structural work
- Preserve valid continuation state so OPERATION_LIMIT pixels do not replay their first ~350 iterations.
## Reproduction
From project root:
```bash
node tests/regression.mjs
node experiments/ab/ab_test.mjs
node experiments/ab/supplement.mjs
node experiments/ab/variant_check.mjs
```
Raw data:
- `docs/performance-cost-audit/ab/results.json`
- `docs/performance-cost-audit/ab/supplement.json`
All B variants are experimental. Production was not modified by this A/B test.

View file

@ -0,0 +1 @@
544820c5024cf6d56e012db6d1bec3ddfcef5d833691b890baa64b26c5c2ee5a index.html

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,12 @@
{
"interiorGateRandom": {
"cases": 240,
"gateFalse": 160,
"falseNeg": 0,
"totalProofInFalseNeg": 0
},
"fastCoverageSemantic": {
"cases": 200,
"casesWithHoles": 0
}
}

File diff suppressed because it is too large Load diff