BXM-4-64: vkCreateComputePipelines fails (VK_ERROR_UNKNOWN) on valid SPIR-V using RoundingModeRTE 16, and fp16 compute returns wrong results

Hi,

I am hitting two distinct problems in the Vulkan compute path on a PowerVR B-Series BXM-4-64 MC1. I have reduced the first one to a two-line SPIR-V difference, and I can share a standalone ~100-line reproducer that does not involve any third-party framework.

Environment

GPU PowerVR B-Series BXM-4-64 MC1
Driver PowerVR B-Series Vulkan Driver, driverVersion = 6643903 (DDK 24.2)
vendorID / deviceID 0x1010 / 0x36104183
apiVersion 1.3.277
SoC / board Allwinner A733 / Orange Pi 4 Pro
OS Android 13 (SDK 33)
maxComputeSharedMemorySize 16384
maxComputeWorkGroupInvocations 512
subgroupSize 1

Advertised float-control properties (VkPhysicalDeviceVulkan12Properties):

shaderRoundingModeRTEFloat16          : VK_TRUE
shaderDenormPreserveFloat16           : VK_TRUE
shaderRoundingModeRTEFloat32          : VK_TRUE
shaderSignedZeroInfNanPreserveFloat16 : VK_TRUE

Issue 1 - RoundingModeRTE 16 causes VK_ERROR_UNKNOWN at pipeline creation

The driver advertises shaderRoundingModeRTEFloat16 = VK_TRUE, but vkCreateComputePipelines returns VK_ERROR_UNKNOWN (-13) for valid SPIR-V carrying OpExecutionMode ... RoundingModeRTE 16.

I have two SPIR-V modules produced from the same disassembly and reassembled with the same spirv-as invocation, so the disassemble/reassemble round-trip is not a variable. The complete difference between them is:

-               OpCapability RoundingModeRTE
-               OpExecutionMode %2 RoundingModeRTE 16

Both pass validation with no output:

spirv-val --target-env vulkan1.3 A_control_rte.spv
spirv-val --target-env vulkan1.3 B_no_rte.spv

On device, vkCreateShaderModule succeeds in both cases; only pipeline creation differs:

module vkCreateComputePipelines
A_control_rte.spv (has RoundingModeRTE) VK_ERROR_UNKNOWN (-13)
B_no_rte.spv (those two lines removed) VK_SUCCESS

Notes that may help narrow it down

  • It is not RoundingModeRTE on its own. In the same application run, 16 shader modules all carried OpCapability RoundingModeRTE plus OpExecutionMode RoundingModeRTE 16, and 11 of them compiled fine - including simple elementwise add, scale, soft_max, rms_norm and get_rows kernels. Only the matrix-vector and small-tile matmul kernels were rejected. So the execution mode appears to interact with something else in those shaders rather than being unsupported outright.
  • DenormPreserve is not implicated. Removing only OpExecutionMode ... DenormPreserve 16 while keeping RoundingModeRTE still fails. Removing only RoundingModeRTE succeeds.
  • VK_ERROR_UNKNOWN gives no diagnostic. If the shader compiler has a log or an environment variable that surfaces the internal failure, that would be very helpful.

Issue 2 - fp16 arithmetic returns incorrect results

This one is more serious in practice because it fails silently.

Using a backend conformance suite that compares each operator against a CPU reference, the split falls exactly along dtype:

  • Every f16 case fails; every f32 case passes, with identical tensor shapes, strides and broadcast patterns. Example pair differing only in dtype:
ADD(type=f16,ne=[10,5,4,3],nr=[1,1,1,1],...)  FAIL
ADD(type=f32,ne=[10,5,4,3],nr=[1,1,1,1],...)  OK
  • Matrix multiplication with fp32 inputs also failed, with relative errors of roughly 28-150 against a 5e-4 tolerance, when the shaders used fp16 accumulation internally. Forcing the same code down an fp32-only path made all 1556 matmul cases pass. That isolates the fault to fp16 arithmetic rather than to the matmul kernels or to memory access.
  • End-to-end effect: a transformer model running on this GPU emitted only blank tokens (all-zero logits) while reporting normal throughput. The same model and code on the CPU produced correct output. Disabling fp16 and the RTE execution mode produced correct output on the GPU.

I have not reduced Issue 2 to a standalone shader yet. If that would help, I am happy to attempt it - please let me know which form is most useful.


Additional observation: subgroupSize = 1

The driver reports subgroupSize = 1. This is legal but unusual, and it makes portable compute code take pathological paths: shared-memory tree reductions end up fully serialised, and code that derives a tile or stride from the subgroup width can compute a zero stride. This may be intentional for this part, but if the hardware has a real SIMD width it would be worth exposing it. I mention it only as context - the two issues above are independent of it.


Questions

  1. Is Issue 1 a known shader-compiler defect, and does a newer DDK than 24.2 (6643903) fix it?
  2. Is shaderRoundingModeRTEFloat16 = VK_TRUE accurate for this part, given that the compiler rejects the corresponding execution mode in some shaders?
  3. Any guidance on Issue 2 - in particular, is fp16 compute expected to be numerically reliable on BXM-4-64? Silently wrong results are much harder for applications to defend against than a pipeline creation failure.

I can share the two .spv files and the reproducer source, and I am happy to run further experiments - this is a development board and I can rebuild and instrument freely.

Thanks!