# BXM-4-64: vkCreateComputePipelines fails (VK\_ERROR\_UNKNOWN) on valid SPIR-V using RoundingModeRTE 16, and fp16 compute returns wrong results

**URL:** <https://forums.imgtec.com/t/bxm-4-64-vkcreatecomputepipelines-fails-vk-error-unknown-on-valid-spir-v-using-roundingmoderte-16-and-fp16-compute-returns-wrong-results/4305>\
**Category:** PowerVR Insider\
**Created:** [October 6, 2026, 9:55am UTC](https://forums.imgtec.com/t/bxm-4-64-vkcreatecomputepipelines-fails-vk-error-unknown-on-valid-spir-v-using-roundingmoderte-16-and-fp16-compute-returns-wrong-results/4305 "2026-10-06T09:55:18Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![vliutyi](https://avatars.discourse-cdn.com/v4/letter/v/f475e1/32.png) [@vliutyi](https://forums.imgtec.com/u/vliutyi)\
**Post date:** [October 6, 2026, 9:55am UTC](https://forums.imgtec.com/t/bxm-4-64-vkcreatecomputepipelines-fails-vk-error-unknown-on-valid-spir-v-using-roundingmoderte-16-and-fp16-compute-returns-wrong-results/4305/1 "2026-10-06T09:55:18Z")

</div>

Hi,

I am hitting two distinct problems in the Vulkan compute path on a PowerVR B-Series BXM-4-64 MC1. I have reduced the first one to a **two-line SPIR-V difference** , and I can share a standalone ~100-line reproducer that does not involve any third-party framework.

## Environment

| | |
| --- | --- |
| GPU | PowerVR B-Series BXM-4-64 MC1 |
| Driver | PowerVR B-Series Vulkan Driver, `driverVersion = 6643903` (DDK 24.2) |
| vendorID / deviceID | `0x1010` / `0x36104183` |
| apiVersion | 1.3.277 |
| SoC / board | Allwinner A733 / Orange Pi 4 Pro |
| OS | Android 13 (SDK 33) |
| maxComputeSharedMemorySize | 16384 |
| maxComputeWorkGroupInvocations | 512 |
| subgroupSize | 1 |

Advertised float-control properties (`VkPhysicalDeviceVulkan12Properties`):

```auto
shaderRoundingModeRTEFloat16 : VK_TRUE
shaderDenormPreserveFloat16 : VK_TRUE
shaderRoundingModeRTEFloat32 : VK_TRUE
shaderSignedZeroInfNanPreserveFloat16 : VK_TRUE

```

* * *

## Issue 1 - RoundingModeRTE 16 causes VK\_ERROR\_UNKNOWN at pipeline creation

The driver advertises `shaderRoundingModeRTEFloat16 = VK_TRUE`, but `vkCreateComputePipelines` returns `VK_ERROR_UNKNOWN (-13)` for valid SPIR-V carrying `OpExecutionMode ... RoundingModeRTE 16`.

I have two SPIR-V modules produced from the _same_ disassembly and reassembled with the _same_ `spirv-as` invocation, so the disassemble/reassemble round-trip is **not** a variable. The complete difference between them is:

```auto
- OpCapability RoundingModeRTE
- OpExecutionMode %2 RoundingModeRTE 16

```

Both pass validation with no output:

```auto
spirv-val --target-env vulkan1.3 A_control_rte.spv
spirv-val --target-env vulkan1.3 B_no_rte.spv

```

On device, `vkCreateShaderModule` succeeds in both cases; only pipeline creation differs:

| module | vkCreateComputePipelines |
| --- | --- |
| `A_control_rte.spv` (has RoundingModeRTE) | **VK\_ERROR\_UNKNOWN (-13)** |
| `B_no_rte.spv` (those two lines removed) | VK\_SUCCESS |

### Notes that may help narrow it down

- **It is not `RoundingModeRTE` on its own.** In the same application run, 16 shader modules all carried `OpCapability RoundingModeRTE` plus `OpExecutionMode RoundingModeRTE 16`, and **11 of them compiled fine** - including simple elementwise add, scale, soft\_max, rms\_norm and get\_rows kernels. Only the matrix-vector and small-tile matmul kernels were rejected. So the execution mode appears to interact with something else in those shaders rather than being unsupported outright.
- **`DenormPreserve` is not implicated.** Removing only `OpExecutionMode ... DenormPreserve 16` while keeping `RoundingModeRTE` still fails. Removing only `RoundingModeRTE` succeeds.
- `VK_ERROR_UNKNOWN` gives no diagnostic. If the shader compiler has a log or an environment variable that surfaces the internal failure, that would be very helpful.

* * *

## Issue 2 - fp16 arithmetic returns incorrect results

This one is more serious in practice because it fails **silently**.

Using a backend conformance suite that compares each operator against a CPU reference, the split falls exactly along dtype:

- **Every `f16` case fails; every `f32` case passes** , with identical tensor shapes, strides and broadcast patterns. Example pair differing only in dtype:

```auto
ADD(type=f16,ne=[10,5,4,3],nr=[1,1,1,1],...) FAIL
ADD(type=f32,ne=[10,5,4,3],nr=[1,1,1,1],...) OK

```

- Matrix multiplication with **fp32 inputs** also failed, with relative errors of roughly 28-150 against a 5e-4 tolerance, when the shaders used fp16 accumulation internally. Forcing the same code down an fp32-only path made **all 1556 matmul cases pass**. That isolates the fault to fp16 arithmetic rather than to the matmul kernels or to memory access.
- End-to-end effect: a transformer model running on this GPU emitted only blank tokens (all-zero logits) while reporting normal throughput. The same model and code on the CPU produced correct output. Disabling fp16 and the RTE execution mode produced correct output on the GPU.

I have not reduced Issue 2 to a standalone shader yet. If that would help, I am happy to attempt it - please let me know which form is most useful.

* * *

## Additional observation: subgroupSize = 1

The driver reports `subgroupSize = 1`. This is legal but unusual, and it makes portable compute code take pathological paths: shared-memory tree reductions end up fully serialised, and code that derives a tile or stride from the subgroup width can compute a zero stride. This may be intentional for this part, but if the hardware has a real SIMD width it would be worth exposing it. I mention it only as context - the two issues above are independent of it.

* * *

## Questions

1. Is Issue 1 a known shader-compiler defect, and does a newer DDK than 24.2 (`6643903`) fix it?
2. Is `shaderRoundingModeRTEFloat16 = VK_TRUE` accurate for this part, given that the compiler rejects the corresponding execution mode in some shaders?
3. Any guidance on Issue 2 - in particular, is fp16 compute expected to be numerically reliable on BXM-4-64? Silently wrong results are much harder for applications to defend against than a pipeline creation failure.

I can share the two `.spv` files and the reproducer source, and I am happy to run further experiments - this is a development board and I can rebuild and instrument freely.

Thanks!
