DXT-48 (Pixel 10 Pro, Android 17): intermittent corruption of compute-shader results when GPU memory objects are recycled across dispatches; survives barriers and fence-waited submits

(All evidence — allocator-internal traces from corrupted inferences, the instrumentation patch, the dedicated-allocator source, and A/B details — is in this public gist: ncnn-Vulkan corruption on PowerVR (vendorID 0x1010) with lightmode blob recycling — allocator-internal traces + instrumentation patch · GitHub. A middleware-side twin report is filed at Vulkan compute output corrupted on Imagination/PowerVR with opt.lightmode=1 blob recycling — allocator bookkeeping proven correct by internal tracing during a corrupted inference; barriers and submit boundaries ruled out · Issue #6935 · Tencent/ncnn · GitHub and a Pixel-integrator co-file at https://issuetracker.google.com/issues/554386275.)

Summary

A production ML inference workload (ncnn 20260526, RIFE v4.7 frame interpolation,
buffer-mode compute, fp32) produces corrupted results on Imagination/PowerVR Vulkan
(vendorID 0x1010) — output frames mathematically unrelated to the inputs — while the
identical command stream is correct on the CPU reference, on ARM Mali, and on
Qualcomm Adreno (6 field devices). We have spent multiple debugging sessions
eliminating every application/middleware-side explanation; each elimination is
experimental, on-device, and the artifacts are attached. We believe what remains is
a driver defect in executing a legal (if aggressive) buffer sub-allocation reuse
pattern.

Observed on three PowerVR generations in the field; all detail below is from
DXT-48 / Google Pixel 10 Pro (Tensor G5), Android 17
google/blazer/blazer:17/CP2A.260805.005/15828068, ro.hardware.vulkan=powervr.

The workload pattern

ncnn’s VkBlobAllocator carves large VkBuffer blocks (STORAGE_BUFFER |
TRANSFER_SRC | TRANSFER_DST; device-local, unified memory on this device) into
sub-ranges bound as storage-buffer descriptors for a chain of ~580 compute
dispatches per inference, recorded into one command buffer (auto-split into ~10
submissions on this device, see below). With its “lightmode” enabled, a sub-range
is returned to the allocator’s free list as soon as its consumer dispatch is
recorded, and may be handed to a later dispatch’s output — so within one
submission, dispatch N’s input range can become dispatch N+k’s output range.
Hazards between dispatches are guarded by vkCmdPipelineBarrier with
buffer-memory barriers over the exact sub-range
(srcAccess=SHADER_READ|SHADER_WRITE, srcStage=COMPUTE_SHADER after first use —
and, in one experiment below, from the very first use as well).

Per the Vulkan specification this is legal: execution/memory dependencies between
compute dispatches are exactly what vkCmdPipelineBarrier
(COMPUTE_SHADER→COMPUTE_SHADER, SHADER_WRITE/SHADER_READ) establishes, and
rebinding a different sub-range of the same VkBuffer in a later descriptor set
introduces no additional synchronization requirement beyond those barriers.

Why we believe the app/middleware side is clean (each point measured on-device)

  1. Allocator bookkeeping proven correct DURING two corrupted inferences. We
    instrumented the allocator internally (patch attached): a shadow interval set of
    live sub-ranges, overlap detection at every hand-out, free-list invariant checks
    after every mutation, refcount logging at every release. Across ~15,700
    hand-outs, zero anomalies — no range was ever handed out while still live,
    no free-list overlap, no double-free — and both traced runs’ outputs were
    corrupted
    (attached traces). CPU-side use-after-free / aliasing is ruled
    out with instrumentation, not inference.
  2. Barriers are present and conservative. Steady-state barriers carry
    srcStage=COMPUTE_SHADER, srcAccess=SHADER_READ|SHADER_WRITE over the exact
    sub-range. In a dedicated experiment we additionally stamped conservative
    src scope on every FIRST use of a recycled range: output was bit-identical
    corrupted
    — so no missing/weak barrier explains it.
  3. Corruption survives fence-waited submission boundaries. Two independent
    ways: (a) this device’s inference is auto-split into ~9-10
    vkQueueSubmit+vkWaitForFences chunks anyway (measured: 120 submits per
    12-frame render); (b) we inserted explicit submit+wait at 16 extra points per
    inference — output still corrupted (bounded “ghosting” variant). After a fence
    wait, all prior device writes are available; a subsequent independent
    submission should not observe stale/partial data, yet results are still wrong.
  4. Not the shaders. The one custom compute shader (bilinear warp) matches its
    CPU reference to 2.4e-06 in a standalone GPU-vs-CPU self-test at the same
    sizes; the remaining shaders are stock ncnn, correct on Mali/Adreno/CPU.
    fp16 is fully disabled (fp32 storage/arithmetic); packing conversions were
    toggled off with bit-identical corruption (exonerated).
  5. The one clean configuration disables reuse. With recycling off
    (lightmode=0 — every sub-range write-once), the same device, driver, shaders,
    sizes and inputs are correct every time. The trigger is specifically reuse of
    freed sub-ranges across dispatches
    .
  6. Race-like signature. A “delay reuse by K allocations” allocator
    experiment changed WHICH ranges were corrupted and made output nondeterministic
    across runs (ghosting on one input, full-frame noise on another) — a
    scheduler/timing signature, not a functional-bug signature. The corruption
    itself is deterministic per (build, input).

Follow-up (same day): allocation-structure workaround also fails — three distinct corruption signatures

To test whether the defect is specific to sub-range aliasing, we replaced the
middleware allocator with a dedicated whole-VkBuffer-per-blob allocator (every
blob its own VkBuffer at offset 0 with its own VkDeviceMemory; whole buffers
recycled by exact size class — no storage-buffer descriptor ever aliases a
sub-range of a shared buffer; memory/perf equivalent to the stock path). In a
same-session, same-installed-build A/B flipped by a runtime property:

  • dedicated allocator: still corrupt, NEW signature — full-frame RGB
    channel-banded striping across the interpolated window; frames outside the
    window pristine (no pipeline-wide poisoning);
  • stock sub-range-recycling allocator (control): classic grey-black melt.

Three allocation strategies now produce three DISTINCT corruption signatures on
this driver (same compute chain, same inputs): sub-range recycling → melt;
delayed/changed reuse geometry → RGB noise/ghosting (nondeterministic);
dedicated whole-buffer-per-blob → channel-banded striping. The defect is
therefore not sub-range aliasing specifically: any recycling of GPU memory
objects within this command stream corrupts, pattern-dependently.
A zero-reuse
configuration is untestable by construction (the middleware frees during command
recording, so destruction must be deferred — a deferral pool is reuse); the
no-recycling mode remains the one clean configuration at ~18x peak memory.

This exhausts the application/middleware workaround space: barrier flags, submit
granularity, shaders/precision, reuse geometry, and allocation structure have
all been varied experimentally with corruption persisting throughout. Allocator
source and A/B details: dedicated_allocator.h and
dedicated_allocator_ab_result.md in the evidence gist.

Reproduction

Deterministic in our app on device (two known-failing input pairs; automated
oracle: an interpolated frame farther from BOTH endpoint frames than they are from
each other — impossible for any correct interpolation; corrupted runs score
1.5x–5.4x). We can: run any instrumented driver/app build you provide and return
logs; provide the full app-level repro privately; or work with you to reduce to a
standalone command-stream repro (a compute chain over recycled sub-ranges of one
VkBuffer with correct barriers, diffed against a CPU reference). An ncnn-side report
is filed in parallel at Tencent/ncnn (this issue’s twin).

Questions for the driver team

  1. Are there known DXT-48 (or Rogue/Volcanic-family) errata around storage-buffer
    descriptor rebinding of overlapping/reused sub-ranges of one VkBuffer within or
    across command buffers?
  2. Is there additional synchronization the driver expects beyond spec-required
    pipeline barriers for this pattern (which would itself be a conformance
    question)?
  3. Any driver-side instrumentation/tooling (PVRCarbon capture?) you would want
    from the failing device to localize this?

Possibly related prior report on this forum: “PowerVR DXT-48-1536: incorrect results from k-quant Vulkan compute shaders in llama.cpp” - also incorrect Vulkan compute results on DXT-48.