# DXT-48 (Pixel 10 Pro, Android 17): intermittent corruption of compute-shader results when GPU memory objects are recycled across dispatches; survives barriers and fence-waited submits

**URL:** <https://forums.imgtec.com/t/dxt-48-pixel-10-pro-android-17-intermittent-corruption-of-compute-shader-results-when-gpu-memory-objects-are-recycled-across-dispatches-survives-barriers-and-fence-waited-submits/4303>\
**Category:** PowerVR Insider\
**Created:** [October 6, 2026, 9:54am UTC](https://forums.imgtec.com/t/dxt-48-pixel-10-pro-android-17-intermittent-corruption-of-compute-shader-results-when-gpu-memory-objects-are-recycled-across-dispatches-survives-barriers-and-fence-waited-submits/4303 "2026-10-06T09:54:58Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![di2g](https://sea1.discourse-cdn.com/flex015/user_avatar/forums.imgtec.com/di2g/32/756_2.png) [@di2g](https://forums.imgtec.com/u/di2g)\
**Post date:** [October 6, 2026, 9:54am UTC](https://forums.imgtec.com/t/dxt-48-pixel-10-pro-android-17-intermittent-corruption-of-compute-shader-results-when-gpu-memory-objects-are-recycled-across-dispatches-survives-barriers-and-fence-waited-submits/4303/1 "2026-10-06T09:54:58Z")

</div>

(All evidence — allocator-internal traces from corrupted inferences, the instrumentation patch, the dedicated-allocator source, and A/B details — is in this public gist: [ncnn-Vulkan corruption on PowerVR (vendorID 0x1010) with lightmode blob recycling — allocator-internal traces + instrumentation patch · GitHub](https://gist.github.com/di2g/a3788246430458f8bad13c9308f18adf). A middleware-side twin report is filed at [Vulkan compute output corrupted on Imagination/PowerVR with opt.lightmode=1 blob recycling — allocator bookkeeping proven correct by internal tracing during a corrupted inference; barriers and submit boundaries ruled out · Issue #6935 · Tencent/ncnn · GitHub](https://github.com/Tencent/ncnn/issues/6935) and a Pixel-integrator co-file at [https://issuetracker.google.com/issues/554386275.](https://issuetracker.google.com/issues/554386275.))

## Summary

A production ML inference workload (ncnn 20260526, RIFE v4.7 frame interpolation,  
buffer-mode compute, fp32) produces corrupted results on Imagination/PowerVR Vulkan  
(vendorID 0x1010) — output frames mathematically unrelated to the inputs — while the  
identical command stream is correct on the CPU reference, on ARM Mali, and on  
Qualcomm Adreno (6 field devices). We have spent multiple debugging sessions  
eliminating every application/middleware-side explanation; each elimination is  
experimental, on-device, and the artifacts are attached. We believe what remains is  
a driver defect in executing a legal (if aggressive) buffer sub-allocation reuse  
pattern.

Observed on three PowerVR generations in the field; all detail below is from  
DXT-48 / Google Pixel 10 Pro (Tensor G5), Android 17  
`google/blazer/blazer:17/CP2A.260805.005/15828068`, `ro.hardware.vulkan=powervr`.

## The workload pattern

ncnn’s `VkBlobAllocator` carves large `VkBuffer` blocks (STORAGE\_BUFFER |  
TRANSFER\_SRC | TRANSFER\_DST; device-local, unified memory on this device) into  
sub-ranges bound as storage-buffer descriptors for a chain of ~580 compute  
dispatches per inference, recorded into one command buffer (auto-split into ~10  
submissions on this device, see below). With its “lightmode” enabled, a sub-range  
is returned to the allocator’s free list as soon as its consumer dispatch is  
_recorded_, and may be handed to a later dispatch’s output — so within one  
submission, dispatch N’s input range can become dispatch N+k’s output range.  
Hazards between dispatches are guarded by `vkCmdPipelineBarrier` with  
buffer-memory barriers over the exact sub-range  
(srcAccess=SHADER\_READ|SHADER\_WRITE, srcStage=COMPUTE\_SHADER after first use —  
and, in one experiment below, from the very first use as well).

Per the Vulkan specification this is legal: execution/memory dependencies between  
compute dispatches are exactly what `vkCmdPipelineBarrier`  
(COMPUTE\_SHADER→COMPUTE\_SHADER, SHADER\_WRITE/SHADER\_READ) establishes, and  
rebinding a different sub-range of the same VkBuffer in a later descriptor set  
introduces no additional synchronization requirement beyond those barriers.

## Why we believe the app/middleware side is clean (each point measured on-device)

1. **Allocator bookkeeping proven correct DURING two corrupted inferences.** We  
instrumented the allocator internally (patch attached): a shadow interval set of  
live sub-ranges, overlap detection at every hand-out, free-list invariant checks  
after every mutation, refcount logging at every release. Across ~15,700  
hand-outs, **zero** anomalies — no range was ever handed out while still live,  
no free-list overlap, no double-free — **and both traced runs’ outputs were  
corrupted** (attached traces). CPU-side use-after-free / aliasing is ruled  
out with instrumentation, not inference.
2. **Barriers are present and conservative.** Steady-state barriers carry  
srcStage=COMPUTE\_SHADER, srcAccess=SHADER\_READ|SHADER\_WRITE over the exact  
sub-range. In a dedicated experiment we additionally stamped conservative  
src scope on every FIRST use of a recycled range: output was **bit-identical  
corrupted** — so no missing/weak barrier explains it.
3. **Corruption survives fence-waited submission boundaries.** Two independent  
ways: (a) this device’s inference is auto-split into ~9-10  
`vkQueueSubmit`+`vkWaitForFences` chunks anyway (measured: 120 submits per  
12-frame render); (b) we inserted explicit submit+wait at 16 extra points per  
inference — output still corrupted (bounded “ghosting” variant). After a fence  
wait, all prior device writes are available; a subsequent independent  
submission should not observe stale/partial data, yet results are still wrong.
4. **Not the shaders.** The one custom compute shader (bilinear warp) matches its  
CPU reference to 2.4e-06 in a standalone GPU-vs-CPU self-test at the same  
sizes; the remaining shaders are stock ncnn, correct on Mali/Adreno/CPU.  
fp16 is fully disabled (fp32 storage/arithmetic); packing conversions were  
toggled off with bit-identical corruption (exonerated).
5. **The one clean configuration disables reuse.** With recycling off  
(`lightmode=0` — every sub-range write-once), the same device, driver, shaders,  
sizes and inputs are correct every time. The trigger is specifically _reuse of  
freed sub-ranges across dispatches_.
6. **Race-like signature.** A “delay reuse by K allocations” allocator  
experiment changed WHICH ranges were corrupted and made output nondeterministic  
across runs (ghosting on one input, full-frame noise on another) — a  
scheduler/timing signature, not a functional-bug signature. The corruption  
itself is deterministic per (build, input).

## Follow-up (same day): allocation-structure workaround also fails — three distinct corruption signatures

To test whether the defect is specific to sub-range aliasing, we replaced the  
middleware allocator with a dedicated whole-VkBuffer-per-blob allocator (every  
blob its own VkBuffer at offset 0 with its own VkDeviceMemory; whole buffers  
recycled by exact size class — no storage-buffer descriptor ever aliases a  
sub-range of a shared buffer; memory/perf equivalent to the stock path). In a  
same-session, same-installed-build A/B flipped by a runtime property:

- dedicated allocator: **still corrupt, NEW signature** — full-frame RGB  
channel-banded striping across the interpolated window; frames outside the  
window pristine (no pipeline-wide poisoning);
- stock sub-range-recycling allocator (control): classic grey-black melt.

Three allocation strategies now produce three DISTINCT corruption signatures on  
this driver (same compute chain, same inputs): sub-range recycling → melt;  
delayed/changed reuse geometry → RGB noise/ghosting (nondeterministic);  
dedicated whole-buffer-per-blob → channel-banded striping. The defect is  
therefore not sub-range aliasing specifically: **any recycling of GPU memory  
objects within this command stream corrupts, pattern-dependently.** A zero-reuse  
configuration is untestable by construction (the middleware frees during command  
recording, so destruction must be deferred — a deferral pool is reuse); the  
no-recycling mode remains the one clean configuration at ~18x peak memory.

This exhausts the application/middleware workaround space: barrier flags, submit  
granularity, shaders/precision, reuse geometry, and allocation structure have  
all been varied experimentally with corruption persisting throughout. Allocator  
source and A/B details: `dedicated_allocator.h` and  
`dedicated_allocator_ab_result.md` in the evidence gist.

## Reproduction

Deterministic in our app on device (two known-failing input pairs; automated  
oracle: an interpolated frame farther from BOTH endpoint frames than they are from  
each other — impossible for any correct interpolation; corrupted runs score  
1.5x–5.4x). We can: run any instrumented driver/app build you provide and return  
logs; provide the full app-level repro privately; or work with you to reduce to a  
standalone command-stream repro (a compute chain over recycled sub-ranges of one  
VkBuffer with correct barriers, diffed against a CPU reference). An ncnn-side report  
is filed in parallel at Tencent/ncnn (this issue’s twin).

## Questions for the driver team

1. Are there known DXT-48 (or Rogue/Volcanic-family) errata around storage-buffer  
descriptor rebinding of overlapping/reused sub-ranges of one VkBuffer within or  
across command buffers?
2. Is there additional synchronization the driver expects beyond spec-required  
pipeline barriers for this pattern (which would itself be a conformance  
question)?
3. Any driver-side instrumentation/tooling (PVRCarbon capture?) you would want  
from the failing device to localize this?

Possibly related prior report on this forum: “PowerVR DXT-48-1536: incorrect results from k-quant Vulkan compute shaders in llama.cpp” - also incorrect Vulkan compute results on DXT-48.
