Hi,
I am hitting two distinct problems in the Vulkan compute path on a PowerVR B-Series BXM-4-64 MC1. I have reduced the first one to a two-line SPIR-V difference, and I can share a standalone ~100-line reproducer that does not involve any third-party framework.
Environment
| GPU | PowerVR B-Series BXM-4-64 MC1 |
| Driver | PowerVR B-Series Vulkan Driver, driverVersion = 6643903 (DDK 24.2) |
| vendorID / deviceID | 0x1010 / 0x36104183 |
| apiVersion | 1.3.277 |
| SoC / board | Allwinner A733 / Orange Pi 4 Pro |
| OS | Android 13 (SDK 33) |
| maxComputeSharedMemorySize | 16384 |
| maxComputeWorkGroupInvocations | 512 |
| subgroupSize | 1 |
Advertised float-control properties (VkPhysicalDeviceVulkan12Properties):
shaderRoundingModeRTEFloat16 : VK_TRUE
shaderDenormPreserveFloat16 : VK_TRUE
shaderRoundingModeRTEFloat32 : VK_TRUE
shaderSignedZeroInfNanPreserveFloat16 : VK_TRUE
Issue 1 - RoundingModeRTE 16 causes VK_ERROR_UNKNOWN at pipeline creation
The driver advertises shaderRoundingModeRTEFloat16 = VK_TRUE, but vkCreateComputePipelines returns VK_ERROR_UNKNOWN (-13) for valid SPIR-V carrying OpExecutionMode ... RoundingModeRTE 16.
I have two SPIR-V modules produced from the same disassembly and reassembled with the same spirv-as invocation, so the disassemble/reassemble round-trip is not a variable. The complete difference between them is:
- OpCapability RoundingModeRTE
- OpExecutionMode %2 RoundingModeRTE 16
Both pass validation with no output:
spirv-val --target-env vulkan1.3 A_control_rte.spv
spirv-val --target-env vulkan1.3 B_no_rte.spv
On device, vkCreateShaderModule succeeds in both cases; only pipeline creation differs:
| module | vkCreateComputePipelines |
|---|---|
A_control_rte.spv (has RoundingModeRTE) |
VK_ERROR_UNKNOWN (-13) |
B_no_rte.spv (those two lines removed) |
VK_SUCCESS |
Notes that may help narrow it down
- It is not
RoundingModeRTEon its own. In the same application run, 16 shader modules all carriedOpCapability RoundingModeRTEplusOpExecutionMode RoundingModeRTE 16, and 11 of them compiled fine - including simple elementwise add, scale, soft_max, rms_norm and get_rows kernels. Only the matrix-vector and small-tile matmul kernels were rejected. So the execution mode appears to interact with something else in those shaders rather than being unsupported outright. DenormPreserveis not implicated. Removing onlyOpExecutionMode ... DenormPreserve 16while keepingRoundingModeRTEstill fails. Removing onlyRoundingModeRTEsucceeds.VK_ERROR_UNKNOWNgives no diagnostic. If the shader compiler has a log or an environment variable that surfaces the internal failure, that would be very helpful.
Issue 2 - fp16 arithmetic returns incorrect results
This one is more serious in practice because it fails silently.
Using a backend conformance suite that compares each operator against a CPU reference, the split falls exactly along dtype:
- Every
f16case fails; everyf32case passes, with identical tensor shapes, strides and broadcast patterns. Example pair differing only in dtype:
ADD(type=f16,ne=[10,5,4,3],nr=[1,1,1,1],...) FAIL
ADD(type=f32,ne=[10,5,4,3],nr=[1,1,1,1],...) OK
- Matrix multiplication with fp32 inputs also failed, with relative errors of roughly 28-150 against a 5e-4 tolerance, when the shaders used fp16 accumulation internally. Forcing the same code down an fp32-only path made all 1556 matmul cases pass. That isolates the fault to fp16 arithmetic rather than to the matmul kernels or to memory access.
- End-to-end effect: a transformer model running on this GPU emitted only blank tokens (all-zero logits) while reporting normal throughput. The same model and code on the CPU produced correct output. Disabling fp16 and the RTE execution mode produced correct output on the GPU.
I have not reduced Issue 2 to a standalone shader yet. If that would help, I am happy to attempt it - please let me know which form is most useful.
Additional observation: subgroupSize = 1
The driver reports subgroupSize = 1. This is legal but unusual, and it makes portable compute code take pathological paths: shared-memory tree reductions end up fully serialised, and code that derives a tile or stride from the subgroup width can compute a zero stride. This may be intentional for this part, but if the hardware has a real SIMD width it would be worth exposing it. I mention it only as context - the two issues above are independent of it.
Questions
- Is Issue 1 a known shader-compiler defect, and does a newer DDK than 24.2 (
6643903) fix it? - Is
shaderRoundingModeRTEFloat16 = VK_TRUEaccurate for this part, given that the compiler rejects the corresponding execution mode in some shaders? - Any guidance on Issue 2 - in particular, is fp16 compute expected to be numerically reliable on BXM-4-64? Silently wrong results are much harder for applications to defend against than a pipeline creation failure.
I can share the two .spv files and the reproducer source, and I am happy to run further experiments - this is a development board and I can rebuild and instrument freely.
Thanks!