← Back to the Lab

A 500× speed-up that computed nothing: six silent failures in a WebGPU simulation

webgpusimulationdebugging

The run finished at 0.04 ms per step, roughly 500 times faster than the simulation should have run at that size. The health report was clean: no NaN, no infinity, the largest force was zero, and every particle had zero neighbours. It was clean because nothing had been computed.

This happened in protocell-genesis, a WebGPU simulation I built to test whether a coarse-grained "primordial soup" can assemble itself into a closed vesicle, a bilayer membrane with water trapped inside. In this model it can't, and the repository has the measured reasons. The largest run had 483,268 particles. Over the project, six bugs behaved like the one above: no exception, no error on the page, and numbers that looked plausible or even good. Three of them are specific to WebGPU. Here they are, with what each one looked like and what changed.

1. A buffer larger than the device allows

At a box size of 85 σ the Verlet neighbour list needed 5,484,474,528 bytes. The device's maxStorageBufferBindingSize was 4,294,967,292.

WebGPU does not throw when createBuffer fails validation. It prints a warning to the console and returns an invalid buffer, and every later dispatch on a bind group that uses it is dropped. The simulation loop kept stepping and kept reading back zeros: max|F| = 0, 0 non-finite values, 0 neighbours, 0.0405 ms per step.

The fix is a check before allocation. createSoup computes every buffer whose size scales with the particle count and compares it with the device's own limits:

const bindingLimit = device.limits.maxStorageBufferBindingSize
const bufferLimit = device.limits.maxBufferSize
for (const [what, bytes] of candidates) {
  if (bytes > bindingLimit || bytes > bufferLimit) {
    throw new Error(
      `buffer ${what} needs ${bytes} bytes at N=${capacityN}; ` +
        `maxStorageBufferBindingSize=${bindingLimit}, ` +
        `maxBufferSize=${bufferLimit}`,
    )
  }
}

2. Reading back a buffer without COPY_SRC

Reading data off the GPU means copying it into a mappable staging buffer. If the source buffer was created without GPUBufferUsage.COPY_SRC, the copy is a validation error. That invalidates the command buffer, so the submit does nothing. The staging buffer was just created, WebGPU zero-initialises new buffers, and mapAsync resolves normally. You get zeros that look like data.

Two random-number buffers were missing the flag. All 72 checkpoints on disk had saved a random-number state of pure zeros, so every resumed run seeded the thermostat with 0 for every particle. With identical Langevin noise on every particle, the whole system drifts as one body: the mean squared displacement was 62.32 σ² for all seven species alike, and relative diffusion, which is what lets aggregates find each other, was gone. After the fix, the largest aggregate at step 90,000 grew from 268 to 588.

usage is a readable property of a GPUBuffer, so the whole class of bug costs one bitwise test:

if ((src.usage & GPUBufferUsage.COPY_SRC) === 0) {
  throw new Error(
    'readBack: the buffer was created without ' +
      `GPUBufferUsage.COPY_SRC (usage=${src.usage})`,
  )
}

3. requestDevice() without requiredLimits

A device requested without requiredLimits gets the default limits, not the ones the adapter supports. The default allows 8 storage buffers per shader stage. The bond kernels grew past that, pipeline creation with an 'auto' layout failed with a console warning, and every affected pass did nothing. No bonds formed at all.

The size limits fail the same way, at a particle count instead of a code change. The default maxStorageBufferBindingSize is 128 MiB. With 1,000 neighbour slots per particle, the list fit at 28,400 particles (113.6 MB) and silently did nothing at 36,400 (145.6 MB).

The device now asks for the adapter's own maxima:

const { limits } = adapter
const device = await adapter.requestDevice({
  requiredLimits: {
    maxStorageBuffersPerShaderStage:
      limits.maxStorageBuffersPerShaderStage,
    maxStorageBufferBindingSize: limits.maxStorageBufferBindingSize,
    maxBufferSize: limits.maxBufferSize,
  },
})

4. A NaN guard that could never fire

A safety check threw when particles drifted further than half the Verlet skin, using sqrt(maxDriftSq) > skin / 2. Every comparison with NaN is false. Once positions became NaN, the drift was NaN and the check passed on every tick. The guard ran the whole time and could not see the failure it was written for.

The fix scans the IEEE-754 exponent bits of every position on the same tick, before the old comparison. It costs 0.212 ms per chunk of 1,000 steps that takes 331.68 ms, which is 0.064%. On the composition that had been diverging silently, the run now throws at step 1,000 and names 6,957 non-finite components out of 60,750.

5. Forces from before the box changed size

Drying and rehydration change the box size. The code rescaled the coordinates, rebuilt the neighbour grid and the Verlet list, and then took the next step with forces computed before the rescale. The result is a wrong number, which is harder to spot than zeros.

A test now pins it from both sides. With the fix, the force buffer on the GPU matches a fresh calculation to within 1.5e-4 at a force scale of 513. Without it, the buffer equals the old forces bit for bit and is off by 313.7.

It had moved a result I had already written up. At box 30 the largest aggregate came out 2.28 times smaller once the forces were right (three runs on each side, ranges not overlapping), and the effect I had reported for drying cycles shrank from 12.5–14.5× to 8.1–8.4×. A guard that had fired 9 times out of 216 fired 0 times out of 36 after the fix.

6. Gates that recognised their input by name

This one has nothing to do with the GPU. The percolation gates, which turn measurements into pass or fail verdicts, found their input by the campaign label, zfB54, written into the code. The first campaign with a different label flipped the gate to "unproven" without an error. Rows now carry a role field, and the label is only a fallback.

A pipeline that identifies its inputs by name goes quiet when someone runs a new experiment, which is exactly when its verdict matters.

A seventh: the instrument read zeros

The test that checked whether the page was drawing read the WebGPU canvas through drawImage into a 2D canvas. It reported 0 changed pixels out of 30,000, on a page that had just drawn 119 frames with 661 instances. The test now uses the scene's own instance counters and a real page.screenshot().

What I treat as a failure now

Zero forces, zero neighbours and zero bonds are failures until something shows the physics can produce them. So is a speed-up nobody made. The same goes for a guard that has never fired, until I have seen it fire on purpose.

WebGPU reports validation errors to the console and to the device's uncapturederror event, and error scopes (device.pushErrorScope('validation')) let you await them around a block of work. This project uses explicit checks that throw and name the buffer and the byte count. Whichever you pick, the default is a console warning and a result that looks fine.

The code, the gate report and the full verdict are in the repository, and the snapshot gallery renders the real checkpoints in any current browser.