3.7 KiB
v24.2.8 DS + True Sparse Fixed96 Direct Experiment
Purpose
The previous three-term f32 expansion correction failed the real-GPU quality gate on every tested depth. v24.2.6 keeps the DS sensitivity gate but replaces the correction arithmetic with deterministic integer fixed-point operations.
Pipeline:
DS Direct + |dz/dc| risk gate
-> risky pixel indices compacted into GPU u32 queue
-> GPU finalize writes indirect workgroup count
-> dispatchWorkgroupsIndirect
-> signed Q8.88 Direct runs only queued pixels
-> compare against corrected Full Deep
The Fixed96 correction does not bind or build a deep reference.
Arithmetic
- width: 96 bits
- format: signed Q8.88 two's complement
- storage: 3 ×
u32 - multiplication: exact 32×32→64 decomposition from 16-bit partial products, 96×96 accumulation into six 32-bit limbs, then round-and-shift by 88 bits
- coordinates: packed directly from the BigInt ViewState into Q8.88; JavaScript
Numberis not used as an intermediate
Q8.88 is intended only for the z6–z12 transition experiment. It is not proposed as an unlimited deep renderer.
Measurements
For z6 / z8 / z10 / z12 and risk 1e11 / 1e12 / 1e13 / 1e14, the UI reports every candidate row:
- extrapolated Fixed96 full-frame time,
- Full Deep cold estimate,
- speedup,
- correction rate,
- remaining UNKNOWN,
- escaped/bounded class mismatch,
- actual Fixed96 dispatch rate,
- queue integrity status,
- status/failure reason.
A failed candidate is never collapsed into a single none row.
CPU preflight
tests/v24-fixed96-model.mjs compares the DS risk gate + exact Q8.88 Direct model with the BigInt P/P+64 oracle on a 1600×900-equivalent Seahorse sample grid. tests/v24-fixed96-limb-model.mjs separately verifies the six-limb multiplication algorithm against BigInt for 20,000 deterministic cases.
These are preflight tests. Real-GPU shader compilation, timing and class comparison remain decisive.
Production status
Diagnostic only. The production router remains unchanged.
v24.2.7 range hotfix
v24.2.6 contained a shader bug in fx_bad_range(): the top Q8.88 integer byte was treated as if it had to be pure sign-extension. Valid values such as +1.2 were therefore rejected. Escaping Mandelbrot orbits commonly pass through component magnitudes above 1, so selected correction pixels were counted as remaining instead of being written back.
v24.2.7 removes that invalid test. While the prior state is inside the bailout disk, Q8.88 multiplication is bounded. After each iteration the shader first tests |Re(z)| > 2 || |Im(z)| > 2 as a mathematically certain escape, and only squares components when both are within ±2. This prevents both the false range rejection and large-value fixed-point wraparound.
v24.2.8 true sparse queue
v24.2.7 still launched a 2D Fixed96 workgroup over the whole benchmark tile and returned immediately for non-risk pixels. On SIMD/SIMT hardware this was not sparse enough: a few risky lanes could keep most of a workgroup occupied by the long integer loop.
v24.2.8 changes the execution topology, not the Q8.88 arithmetic. The guarded DS pass atomically appends each risky local pixel index to a queue. A one-invocation finalize shader writes ceil(queueCount/64),1,1 to an indirect-argument buffer. The Fixed96 kernel is one-dimensional with @workgroup_size(64) and reads exactly one queue entry per invocation. No CPU map/readback occurs between these passes.
Every timing and quality tile validates queue invariants before its result is accepted. Overflow, wrong indirect arguments, invalid/stale indices or any disagreement between selected/enqueued/processed/corrected counts aborts the experiment rather than producing a misleading timing row.