70 lines
3.7 KiB
Markdown
70 lines
3.7 KiB
Markdown
|
|
# v24.2.8 DS + True Sparse Fixed96 Direct Experiment
|
|||
|
|
|
|||
|
|
## Purpose
|
|||
|
|
|
|||
|
|
The previous three-term f32 expansion correction failed the real-GPU quality gate on every tested depth. v24.2.6 keeps the DS sensitivity gate but replaces the correction arithmetic with deterministic integer fixed-point operations.
|
|||
|
|
|
|||
|
|
Pipeline:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
DS Direct + |dz/dc| risk gate
|
|||
|
|
-> risky pixel indices compacted into GPU u32 queue
|
|||
|
|
-> GPU finalize writes indirect workgroup count
|
|||
|
|
-> dispatchWorkgroupsIndirect
|
|||
|
|
-> signed Q8.88 Direct runs only queued pixels
|
|||
|
|
-> compare against corrected Full Deep
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The Fixed96 correction does not bind or build a deep reference.
|
|||
|
|
|
|||
|
|
## Arithmetic
|
|||
|
|
|
|||
|
|
- width: 96 bits
|
|||
|
|
- format: signed Q8.88 two's complement
|
|||
|
|
- storage: 3 × `u32`
|
|||
|
|
- multiplication: exact 32×32→64 decomposition from 16-bit partial products, 96×96 accumulation into six 32-bit limbs, then round-and-shift by 88 bits
|
|||
|
|
- coordinates: packed directly from the BigInt ViewState into Q8.88; JavaScript `Number` is not used as an intermediate
|
|||
|
|
|
|||
|
|
Q8.88 is intended only for the z6–z12 transition experiment. It is not proposed as an unlimited deep renderer.
|
|||
|
|
|
|||
|
|
## Measurements
|
|||
|
|
|
|||
|
|
For z6 / z8 / z10 / z12 and risk `1e11 / 1e12 / 1e13 / 1e14`, the UI reports every candidate row:
|
|||
|
|
|
|||
|
|
- extrapolated Fixed96 full-frame time,
|
|||
|
|
- Full Deep cold estimate,
|
|||
|
|
- speedup,
|
|||
|
|
- correction rate,
|
|||
|
|
- remaining UNKNOWN,
|
|||
|
|
- escaped/bounded class mismatch,
|
|||
|
|
- actual Fixed96 dispatch rate,
|
|||
|
|
- queue integrity status,
|
|||
|
|
- status/failure reason.
|
|||
|
|
|
|||
|
|
A failed candidate is never collapsed into a single `none` row.
|
|||
|
|
|
|||
|
|
## CPU preflight
|
|||
|
|
|
|||
|
|
`tests/v24-fixed96-model.mjs` compares the DS risk gate + exact Q8.88 Direct model with the BigInt P/P+64 oracle on a 1600×900-equivalent Seahorse sample grid. `tests/v24-fixed96-limb-model.mjs` separately verifies the six-limb multiplication algorithm against BigInt for 20,000 deterministic cases.
|
|||
|
|
|
|||
|
|
These are preflight tests. Real-GPU shader compilation, timing and class comparison remain decisive.
|
|||
|
|
|
|||
|
|
## Production status
|
|||
|
|
|
|||
|
|
Diagnostic only. The production router remains unchanged.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## v24.2.7 range hotfix
|
|||
|
|
|
|||
|
|
v24.2.6 contained a shader bug in `fx_bad_range()`: the top Q8.88 integer byte was treated as if it had to be pure sign-extension. Valid values such as `+1.2` were therefore rejected. Escaping Mandelbrot orbits commonly pass through component magnitudes above 1, so selected correction pixels were counted as `remaining` instead of being written back.
|
|||
|
|
|
|||
|
|
v24.2.7 removes that invalid test. While the prior state is inside the bailout disk, Q8.88 multiplication is bounded. After each iteration the shader first tests `|Re(z)| > 2 || |Im(z)| > 2` as a mathematically certain escape, and only squares components when both are within ±2. This prevents both the false range rejection and large-value fixed-point wraparound.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## v24.2.8 true sparse queue
|
|||
|
|
|
|||
|
|
v24.2.7 still launched a 2D Fixed96 workgroup over the whole benchmark tile and returned immediately for non-risk pixels. On SIMD/SIMT hardware this was not sparse enough: a few risky lanes could keep most of a workgroup occupied by the long integer loop.
|
|||
|
|
|
|||
|
|
v24.2.8 changes the execution topology, not the Q8.88 arithmetic. The guarded DS pass atomically appends each risky local pixel index to a queue. A one-invocation finalize shader writes `ceil(queueCount/64),1,1` to an indirect-argument buffer. The Fixed96 kernel is one-dimensional with `@workgroup_size(64)` and reads exactly one queue entry per invocation. No CPU map/readback occurs between these passes.
|
|||
|
|
|
|||
|
|
Every timing and quality tile validates queue invariants before its result is accepted. Overflow, wrong indirect arguments, invalid/stale indices or any disagreement between selected/enqueued/processed/corrected counts aborts the experiment rather than producing a misleading timing row.
|