Compare commits

...
2 Commits
Author SHA1 Message Date
Claude 562076f657 [Fix] (ShaderTranspiler, DirectVulkan, MG_IntegrationTest): patch both of iterationRP's under-declared subgroup scratch arrays
The previous commit's fingerprint was pinned to one array's incidental
dimensions - workgroup exactly 32x16x1, element exactly vec2, length
exactly 32 - which is the auto-exposure reduction and nothing else. The
pack ships the same idiom twice:

  - auto-exposure:  32x16 (512 invocations), shared vec2 prefixSumCache[32]
  - RTW warp:       1024 invocations,        shared float prefixSumCache[64]

so the warp kept writing 128 subgroups into 64 entries on an 8-lane
device and the retrace stayed bit-identically wrong (ssim 0.027902).

Key the fingerprint on the pack's idiom instead of one array's shape: a
workgroup array of 32-bit floats indexed by gl_SubgroupID, fed by a
subgroup scan, whose declared length is below ceil(invocations / native
width). Three properties keep that a targeted repair rather than a
general array resizer:

  - the index must BE gl_SubgroupID (through OpCopyObject, a signedness
    OpBitcast, or a spill whose every store is that id), so an index
    masked or clamped into range is left alone;
  - the >= 16-lane early-out is retained, so every module on the devices
    the pack was written for passes through byte-identical;
  - growth is certified against maxComputeSharedMemorySize using a
    natural-alignment layout model, and declined outright when a
    declaration cannot be sized, so a patched module can never fail
    pipeline creation where the original would not have.

Verified against the shaders the CI trace actually contains: of the 14
compute modules in the fixture exactly these two change, the other
twelve are byte-identical, and all fourteen pass spirv-val. The
integration scenario grows a second case for the 1024-invocation shape;
both abort with heap corruption when the patch is disabled.

Claude-Session: https://claude.ai/code/session_01EXSURVxwp8VVrWrPEQeLCm
2026-08-19 17:00:33 +00:00
swung0x48 d8576a2ed3 [Fix] (DirectVulkan, ShaderTranspiler, MG_IntegrationTest, SelfTest, TraceReplay): use native subgroups and patch iterationRP's under-declared scratch
iterationRP's Program 203 declares shared vec2 prefixSumCache[32] for a
512-invocation workgroup indexed by gl_SubgroupID; any device narrower
than 16 lanes partitions into more than 32 subgroups and the pack writes
shared memory out of bounds (heap corruption on lavapipe's CPU
rasterizer, ssim 0.028 on the CI retrace). Fix it where the fault lies -
in the fixture - and keep the GL contract sound everywhere else:

- FixIterationRPSubgroupScratchPass: fingerprint-gated SPIR-V pass that
  grows exactly that array to ceil(invocations/width) entries on sub-16-lane devices; every other module passes through byte-identical.
- DeriveNumSubgroupsPass stays default-on for the Adreno topology bug
  and is made spec-sound: pipelines request REQUIRE_FULL_SUBGROUPS
  whenever the workgroup shape makes the flag legal (computeFullSubgroups
  enabled, local_size_x a multiple of the native width, subgroup count
  within maxComputeWorkgroupSubgroups).
- EmulateSubgroupsPass: 32-lane virtual-subgroup lowering kept in-tree
  as a last resort, enabled only by MOBILEGL_MAGMA_EMULATE_SUBGROUP=1 on
  devices with no native subgroup support; fails closed on extended
  subgroup instructions and on modules whose added scratch would exceed
  maxComputeSharedMemorySize.
- IterationRPFirstReductionScenario skips gracefully outside the pack's
  16..256-lane source domain; the new IterationRPScratchFixScenario runs
  the fixture-shaped reduction on any width and asserts the exact
  width-independent total. DriverPost keeps reporting FAIL on
  out-of-domain devices.
- Program203 -> IterationRP rename throughout; the per-trace
  num_subgroups_quirk plumbing is removed from the trace replayer, JNI
  chain, and CI workflows.
2026-08-19 09:48:11 -04:00
38 changed files with 3767 additions and 396 deletions
-3
View File
@@ -435,9 +435,6 @@ jobs:
if [ "${{ matrix.case.coherent_as_flush || false }}" = "true" ]; then
extra_retrace_args+=(--coherent-as-flush)
fi
if [ "${{ matrix.case.num_subgroups_quirk || false }}" = "true" ]; then
extra_retrace_args+=(--num-subgroups-quirk)
fi
run_retrace() {
timeout "$(( ${{ matrix.case.timeout_seconds }} + 300 ))" sh android-plugin/trace-replay-ci.sh \
+3 -1
View File
@@ -285,6 +285,8 @@ set(SOURCE_FILES
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/ZeroBaseVertexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DeriveNumSubgroupsPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FixIterationRPSubgroupScratchPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateSubgroupsPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/NormalizeRectCoordinatesPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/Lower1DArrayImagesPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/BakeImageFormatsPass.cpp
@@ -299,7 +301,7 @@ set(SOURCE_FILES
MobileGL/MG_Util/BackendLoaders/Vulkan/Loader.cpp
MobileGL/MG_Util/SelfTest/DriverPost.cpp
MobileGL/MG_Util/SelfTest/DriverPostProgram203Witness.cpp
MobileGL/MG_Util/SelfTest/DriverPostIterationRPWitness.cpp
MobileGL/MG_Util/Texture/PixelStoreProcessor.cpp
MobileGL/MG_Util/Texture/TextureFormatProcessor.cpp
+30 -6
View File
@@ -78,13 +78,37 @@ namespace MobileGL::MG_Config {
// MOBILEGL_TRACE_ANGLE_VARIANT: signed trace-APK ANGLE build short hash.
String TraceAngleVariant;
#endif
// MOBILEGL_DISABLE_SUBGROUP: force-disable Vulkan shader subgroup support.
// MOBILEGL_DISABLE_SUBGROUP: force-disable Vulkan shader subgroup support,
// including the opt-in emulated compute path below.
Bool DisableSubgroup = false;
// MOBILEGL_NUM_SUBGROUPS_QUIRK: derive compute gl_NumSubgroups from the local
// workgroup dimensions and gl_SubgroupSize instead of reading Vulkan's
// NumSubgroups builtin. Off by default; enable only for drivers whose builtin
// disagrees with the SubgroupId topology emitted by the same dispatch.
Bool NumSubgroupsQuirk = false;
// MOBILEGL_MAGMA_EMULATE_SUBGROUP: implement GL_KHR_shader_subgroup's compute
// stage on a 32-lane VIRTUAL subgroup lowered to workgroup-shared memory
// (ShaderTranspiler::EmulateSubgroupsPass). Strictly a last resort: it only ever
// engages when this flag is set AND the device has no native subgroup support at
// all - a device with real subgroup operations always uses them natively,
// whatever their width (the known iterationRP defect is patched by
// FixIterationRPSubgroupScratch below instead). Off by default.
Bool MagmaEmulateSubgroup = false;
// MOBILEGL_FIX_ITERATIONRP_SUBGROUP_SCRATCH: patch iterationRP's own bug - the
// pack declares `shared vec2 prefixSumCache[32]` for a 512-invocation exposure
// reduction and indexes it by gl_SubgroupID, so any device with sub-16-lane
// subgroups (8-lane lavapipe -> 64 subgroups) writes shared memory out of
// bounds. The pass grows that one array to what the device's topology needs and
// touches nothing else; it only rewrites modules positively matching the pack's
// reduction fingerprint (ShaderTranspiler::FixIterationRPSubgroupScratchPass),
// so every other shader passes through byte-identical - as does iterationRP
// itself on >= 16-lane devices. Auto is ON; ForceOff replays the pack's bug
// verbatim.
QuirkOverride FixIterationRPSubgroupScratch = QuirkOverride::Auto;
// MOBILEGL_DERIVE_NUM_SUBGROUPS: replace compute gl_NumSubgroups loads with
// ceil(workgroup invocations / gl_SubgroupSize) on the NATIVE subgroup path
// (ShaderTranspiler::DeriveNumSubgroupsPass). Auto is ON: GL requires
// gl_SubgroupID < gl_NumSubgroups, Adreno's builtin reports 1 while the same
// dispatch emits IDs 0..7, and the derived value is the one Vulkan guarantees
// whenever the pipeline can request REQUIRE_FULL_SUBGROUPS (which the renderer
// does whenever local_size_x is a multiple of the native width). ForceOff returns
// to the raw driver builtin.
QuirkOverride DeriveNumSubgroups = QuirkOverride::Auto;
// MOBILEGL_ADVERTISE_FP64: add GL_ARB_gpu_shader_fp64 to the advertised extension
// string. `double` in a shader always WORKS - it is narrowed to 32 bits before any
// module reaches a backend (ShaderTranspiler::DemoteFloat64Pass) - but the extension
+4 -1
View File
@@ -168,7 +168,10 @@ namespace MobileGL::MG_ConfigLoader {
QueryEnvVariable("MOBILEGL_TRACE_ANGLE_VARIANT", features.TraceAngleVariant, "");
#endif
features.DisableSubgroup = QueryEnvFlag("MOBILEGL_DISABLE_SUBGROUP");
features.NumSubgroupsQuirk = QueryEnvFlag("MOBILEGL_NUM_SUBGROUPS_QUIRK");
features.MagmaEmulateSubgroup = QueryEnvFlag("MOBILEGL_MAGMA_EMULATE_SUBGROUP");
features.FixIterationRPSubgroupScratch =
QueryEnvQuirkOverride("MOBILEGL_FIX_ITERATIONRP_SUBGROUP_SCRATCH");
features.DeriveNumSubgroups = QueryEnvQuirkOverride("MOBILEGL_DERIVE_NUM_SUBGROUPS");
features.AdvertiseFp64 = QueryEnvFlag("MOBILEGL_ADVERTISE_FP64");
features.MagmaR11G11B10FFallback = QueryEnvFlag("MOBILEGL_MAGMA_R11G11B10F_FALLBACK");
features.MagmaFramesInFlight = QueryEnvUint32("MOBILEGL_MAGMA_FRAMESINFLIGHT", 3, 1, 64);
@@ -9,6 +9,7 @@
#include "BackendObject_DirectVulkan.h"
#include "MG_Backend/BackendObject.h"
#include "DirectVulkan.h"
#include "SubgroupSupportPolicy.h"
#include "MG_State/GLState/FramebufferState/FramebufferObject.h"
#include "MG_State/GLState/Core.h"
#include "MG_State/GLState/TextureState/TextureState.h"
@@ -704,8 +705,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// real device timestamp support. ApplyVulkanCapabilitiesForTesting may
// run without a renderer; no timer query is advertised then. Rebuilding
// the whole list keeps re-runs idempotent.
// The opt-in emulated compute path (SubgroupSupportPolicy.h) carries the
// extension by itself on devices with no native subgroup support at all; a
// device with native subgroups always advertises - and uses - those.
const Bool subgroupSupportAdvertised =
m_vulkanCaps.SupportsShaderSubgroup ||
ShouldEmulateSubgroups(m_vulkanCaps.SupportsShaderSubgroup);
m_rendererInfo.RendererGLInfo.Extensions = BuildAdvertisedExtensions(
m_vulkanCaps.SupportsShaderSubgroup, pVulkanRenderer && pVulkanRenderer->IsTimerQuerySupported(),
subgroupSupportAdvertised, pVulkanRenderer && pVulkanRenderer->IsTimerQuerySupported(),
pVulkanRenderer && pVulkanRenderer->IsSamplerAnisotropySupported(),
pVulkanRenderer && pVulkanRenderer->IsNonZeroIndirectBaseInstanceSupported());
}
@@ -941,6 +948,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_dynamicParameters.SubgroupSupportedFeatures =
mapSubgroupFeatures(m_vulkanCaps.SubgroupSupportedOperations);
m_dynamicParameters.SubgroupQuadOperationsInAllStages = m_vulkanCaps.SubgroupQuadOperationsInAllStages;
} else if (ShouldEmulateSubgroups(m_vulkanCaps.SupportsShaderSubgroup)) {
// MOBILEGL_MAGMA_EMULATE_SUBGROUP on a device with no native subgroups: the
// advertised values describe the 32-lane virtual subgroup the compute
// lowering implements (SubgroupSupportPolicy.h / EmulateSubgroupsPass).
// GL requires the advertisement and the execution to agree, and on this
// path the emulation is what executes; only the compute stage is offered.
m_dynamicParameters.SubgroupSize = kEmulatedSubgroupSize;
m_dynamicParameters.SubgroupSupportedStages = kEmulatedSubgroupStages;
m_dynamicParameters.SubgroupSupportedFeatures = kEmulatedSubgroupFeatures;
m_dynamicParameters.SubgroupQuadOperationsInAllStages = false;
MGLOG_I("DirectVulkan: emulating 32-lane compute subgroups "
"(MOBILEGL_MAGMA_EMULATE_SUBGROUP, no native subgroup support)");
} else {
m_dynamicParameters.SubgroupSize = 0;
m_dynamicParameters.SubgroupSupportedStages = 0;
@@ -8,7 +8,6 @@
#include "ProgramFactory.h"
#include "Config.h"
#include "MG_Backend/DirectVulkan/DirectVulkanResourceState.h"
#include "MG_Util/ShaderTranspiler/ShaderCompiler.h"
#include "MG_Util/ShaderTranspiler/SpvcSession.h"
@@ -34,6 +33,32 @@ namespace MobileGL::MG_Backend::DirectVulkan {
using SpvcSession = MG_Util::ShaderTranspiler::SpvcSession;
using SessionUsageBit = MG_Util::ShaderTranspiler::SessionUsageBit;
// Local size of a compute module, read from OpExecutionMode LocalSize; all-zero
// when absent. The compile chain pins SPIR-V 1.3, where a literal local size
// always reaches the module as this execution mode (LocalSizeId does not exist
// yet).
struct ComputeLocalSize {
Uint32 x = 0;
Uint32 y = 0;
Uint32 z = 0;
Uint64 Total() const { return static_cast<Uint64>(x) * y * z; }
};
ComputeLocalSize TryGetComputeLocalSize(const Vector<Uint>& spirv) {
constexpr SizeT kHeaderWords = 5;
constexpr Uint32 kOpExecutionMode = 16;
constexpr Uint32 kModeLocalSize = 17;
for (SizeT offset = kHeaderWords; offset < spirv.size();) {
const Uint32 wordCount = spirv[offset] >> 16u;
const Uint32 opcode = spirv[offset] & 0xffffu;
if (wordCount == 0 || offset + wordCount > spirv.size()) break;
if (opcode == kOpExecutionMode && wordCount >= 6 && spirv[offset + 2] == kModeLocalSize) {
return {spirv[offset + 3], spirv[offset + 4], spirv[offset + 5]};
}
offset += wordCount;
}
return {};
}
struct DescriptorKey {
ProgramFactory::DescriptorBindingKind kind = ProgramFactory::DescriptorBindingKind::None;
String name;
@@ -3164,20 +3189,57 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}
}
// NumSubgroups is defined by the local workgroup dimensions and SubgroupSize. Derive
// it in SPIR-V instead of trusting a driver builtin that can disagree with the
// SubgroupId topology produced by the same compute dispatch (Adreno reports 1 while
// emitting IDs 0..7 for a 512-invocation, 64-wide workgroup).
if (MG_Config::Features.NumSubgroupsQuirk && shaders[i] &&
shaders[i]->GetShaderStage() == ShaderStage::Compute) {
Vector<Uint> derivedNumSubgroupsSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::DeriveNumSubgroupsForVulkan(
moduleSpirvs[i], derivedNumSubgroupsSpirv, enableSpirvValidation)) {
moduleSpirvs[i] = std::move(derivedNumSubgroupsSpirv);
// GL_KHR_shader_subgroup handling (SubgroupSupportPolicy.h). Native subgroup
// operations execute natively; two module repairs keep the GL contract intact
// around them. The opt-in emulation path replaces them only on devices with no
// subgroup support at all (MOBILEGL_MAGMA_EMULATE_SUBGROUP).
if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Compute) {
if (m_subgroupPolicy.emulateSubgroups) {
Vector<Uint> emulatedSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::EmulateSubgroupsForVulkan(
moduleSpirvs[i], emulatedSpirv,
m_subgroupPolicy.maxComputeSharedMemoryBytes, enableSpirvValidation)) {
moduleSpirvs[i] = std::move(emulatedSpirv);
} else {
MGLOG_E("ProgramFactory: subgroup emulation failed for program %u; the "
"module keeps subgroup operations the device cannot execute",
program.GetExternalIndex());
}
} else {
MGLOG_E("ProgramFactory: failed to derive gl_NumSubgroups for program %u; "
"compute shaders may observe a driver-inconsistent subgroup count",
program.GetExternalIndex());
// iterationRP under-declares its cross-subgroup scratch
// (prefixSumCache[32] for 512 invocations); on a sub-16-lane device
// grow that one fingerprinted array to what the topology needs.
if (m_subgroupPolicy.fixIterationRPSubgroupScratch) {
Vector<Uint> patchedSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(
moduleSpirvs[i], patchedSpirv, m_subgroupPolicy.nativeSubgroupSize,
m_subgroupPolicy.maxComputeSharedMemoryBytes,
enableSpirvValidation)) {
moduleSpirvs[i] = std::move(patchedSpirv);
} else {
MGLOG_E("ProgramFactory: iterationRP subgroup scratch patch failed for "
"program %u; the pack's declared array sizes stay in effect",
program.GetExternalIndex());
}
}
// gl_NumSubgroups must agree with the gl_SubgroupID range GL promises;
// derive it from the workgroup dimensions and gl_SubgroupSize instead of
// trusting a driver builtin that can disagree with the topology the same
// dispatch emits (Adreno reports 1 while emitting IDs 0..7 for a
// 512-invocation, 64-wide workgroup). The ceil() partition this derives
// is pinned by REQUIRE_FULL_SUBGROUPS at pipeline creation whenever the
// workgroup shape makes that flag legal (see the stage setup below).
if (m_subgroupPolicy.deriveNumSubgroups) {
Vector<Uint> derivedNumSubgroupsSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::DeriveNumSubgroupsForVulkan(
moduleSpirvs[i], derivedNumSubgroupsSpirv, enableSpirvValidation)) {
moduleSpirvs[i] = std::move(derivedNumSubgroupsSpirv);
} else {
MGLOG_E("ProgramFactory: failed to derive gl_NumSubgroups for program %u; "
"compute shaders may observe a driver-inconsistent subgroup count",
program.GetExternalIndex());
}
}
}
}
@@ -3326,6 +3388,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
stage.stage = ToVkStage(shaderStage);
stage.module = module;
stage.pName = "main";
// Pin the full-subgroup launch the derived gl_NumSubgroups assumes. Legal
// exactly when the computeFullSubgroups feature is enabled and local_size_x is
// a multiple of the subgroup size (VUID-VkPipelineShaderStageCreateInfo-
// flags-02759/-02785), and only worth requesting while the resulting subgroup
// count fits the device's maxComputeWorkgroupSubgroups (lavapipe caps it at
// 32, below a 512-invocation dispatch's 64). With the bit set, "Full
// Subgroups" guarantees every subgroup launches with all invocations active,
// making the subgroup count exactly invocations / size. Shapes the flag
// cannot cover (e.g. 32x16 on a 64-wide device) fall back to the driver's
// own - spec-encouraged - tight partitioning, which the DriverPost witness
// verifies per device.
if (shaderStage == ShaderStage::Compute && m_subgroupPolicy.requireFullSubgroups &&
!m_subgroupPolicy.emulateSubgroups && m_subgroupPolicy.nativeSubgroupSize != 0) {
const ComputeLocalSize localSize = TryGetComputeLocalSize(moduleSpv);
const Uint64 fullSubgroupCount =
localSize.Total() / m_subgroupPolicy.nativeSubgroupSize;
if (localSize.x != 0 && localSize.x % m_subgroupPolicy.nativeSubgroupSize == 0 &&
fullSubgroupCount <= m_subgroupPolicy.maxComputeWorkgroupSubgroups) {
stage.flags |= VK_PIPELINE_SHADER_STAGE_CREATE_REQUIRE_FULL_SUBGROUPS_BIT;
}
}
entry.modules.push_back(module);
entry.stages.push_back(stage);
@@ -372,16 +372,38 @@ namespace MobileGL::MG_Backend::DirectVulkan {
virtual void OnProgramEvicted(HashType programHash, VkDescriptorSetLayout descriptorSetLayout) = 0;
};
// How this factory's compute modules implement GL_KHR_shader_subgroup. Computed
// once at renderer initialization (SubgroupSupportPolicy.h + the device's
// subgroup properties) so lowering can never disagree with the advertised
// capabilities. Native subgroup operations always execute natively; the two
// repair passes patch modules AROUND them, and the emulation only replaces them
// on opted-in devices with no subgroup support at all.
struct SubgroupLoweringPolicy {
Bool emulateSubgroups = false; // MOBILEGL_MAGMA_EMULATE_SUBGROUP, no-native-support devices
Bool fixIterationRPSubgroupScratch = false; // patch iterationRP's under-declared scratch
Bool deriveNumSubgroups = false; // repair the NumSubgroups builtin
Bool requireFullSubgroups = false; // computeFullSubgroups enabled on the device
Uint32 nativeSubgroupSize = 0;
// Full-subgroup launches are bounded by this device limit; a dispatch whose
// workgroup needs more subgroups than this cannot request the flag.
Uint32 maxComputeWorkgroupSubgroups = 0;
// VkPhysicalDeviceLimits::maxComputeSharedMemorySize; bounds the scratch the
// emulation pass may add (0 falls back to the Vulkan minimum, 16384).
Uint32 maxComputeSharedMemoryBytes = 0;
};
explicit ProgramFactory(VkDevice device, const VulkanRendererConfig& config, Uint32 maxBindings,
Bool shaderDrawParametersEnabled,
Bool unformattedFloatStorageImagesEnabled,
Bool enableSpirvValidation,
UpdateAfterBindLimits updateAfterBindLimits)
UpdateAfterBindLimits updateAfterBindLimits,
SubgroupLoweringPolicy subgroupPolicy)
: m_device(device), m_maxBindings(maxBindings), m_config(config),
m_shaderDrawParametersEnabled(shaderDrawParametersEnabled),
m_unformattedFloatStorageImagesEnabled(unformattedFloatStorageImagesEnabled),
m_enableSpirvValidation(enableSpirvValidation),
m_updateAfterBindLimits(updateAfterBindLimits) {
m_updateAfterBindLimits(updateAfterBindLimits),
m_subgroupPolicy(subgroupPolicy) {
VkProgramObject::s_device = device;
}
// Destroys the pass-through tessellation control modules. Runs while the device is
@@ -511,6 +533,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// the factory lets each reflected layout choose ordinary descriptors when its
// own counts would exceed the update-after-bind budget.
UpdateAfterBindLimits m_updateAfterBindLimits{};
SubgroupLoweringPolicy m_subgroupPolicy{};
// See SetDefaultFramebufferHeight. 0 means "not known yet"; the FragCoordYFlip bit is
// never set before the swapchain exists, so no variant can be compiled against it.
Uint32 m_defaultFramebufferHeight = 0;
@@ -8,6 +8,7 @@
#include "VulkanRenderer.h"
#include "MG_Backend/DirectVulkan/SubgroupSupportPolicy.h"
#include "MG_Backend/DirectGLES/Utils.h"
#include "VertexInputStateFactory.h"
#include "VertexInputStateBuilder.h"
@@ -3058,11 +3059,22 @@ void main() {
}
PipelineFactory::SetSuppressBlendedDepthWrite(suppressBlendedDepthWrite);
}
ProgramFactory::SubgroupLoweringPolicy subgroupPolicy{};
subgroupPolicy.emulateSubgroups = ShouldEmulateSubgroups(m_nativeSubgroupSupported);
subgroupPolicy.fixIterationRPSubgroupScratch =
m_nativeSubgroupSupported && ShouldFixIterationRPSubgroupScratch();
subgroupPolicy.deriveNumSubgroups =
m_nativeSubgroupSupported && ShouldDeriveNumSubgroups();
subgroupPolicy.requireFullSubgroups = m_computeFullSubgroupsFeatureEnabled;
subgroupPolicy.nativeSubgroupSize = m_nativeSubgroupSize;
subgroupPolicy.maxComputeWorkgroupSubgroups = m_maxComputeWorkgroupSubgroups;
subgroupPolicy.maxComputeSharedMemoryBytes =
m_physicalDevice.properties.limits.maxComputeSharedMemorySize;
m_programFactory = MakeUnique<ProgramFactory>(m_device, m_config, maxProgramBindings,
m_shaderDrawParametersFeatureEnabled,
m_unformattedFloatStorageImagesEnabled,
MG_Config::Features.EnableSpirvValidation,
m_updateAfterBindLimits);
m_updateAfterBindLimits, subgroupPolicy);
MOBILEGL_ASSERT(m_programFactory != nullptr, "ProgramFactory creation failed.");
// The swapchain already exists at this point (Initialize creates it first), so seed the
// height the factory could not be told about from CreateSwapchain.
@@ -12740,6 +12752,73 @@ void main() {
}
}
// Native subgroup topology, and VK_EXT_subgroup_size_control's
// computeFullSubgroups feature. REQUIRE_FULL_SUBGROUPS on a compute stage is what
// turns the derived gl_NumSubgroups (DeriveNumSubgroupsPass) from
// encouraged-but-unspecified driver behaviour into a spec guarantee: with the bit
// set and local_size_x a multiple of the subgroup size, every subgroup launches
// full, so the subgroup count is exactly invocations / size ("Full Subgroups",
// VUID-VkPipelineShaderStageCreateInfo-flags-02759/-02785).
m_nativeSubgroupSize = 0;
m_nativeSubgroupSupported = false;
m_computeFullSubgroupsFeatureEnabled = false;
if (getPhysicalDeviceProperties2 != nullptr) {
VkPhysicalDeviceSubgroupProperties subgroupProperties{};
subgroupProperties.sType = VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_SUBGROUP_PROPERTIES;
VkPhysicalDeviceProperties2 subgroupPropertyQuery{};
subgroupPropertyQuery.sType = VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_PROPERTIES_2;
subgroupPropertyQuery.pNext = &subgroupProperties;
getPhysicalDeviceProperties2(m_physicalDevice.handle, &subgroupPropertyQuery);
// Mirrors the loader's HasUsableShaderSubgroupSupport gate, including the
// MOBILEGL_DISABLE_SUBGROUP escape hatch, so the module lowerings can never
// disagree with the advertised capabilities.
const Bool usableSubgroups =
subgroupProperties.subgroupSize > 0 &&
(subgroupProperties.supportedStages & VK_SHADER_STAGE_COMPUTE_BIT) != 0 &&
(subgroupProperties.supportedOperations & VK_SUBGROUP_FEATURE_BASIC_BIT) != 0;
if (usableSubgroups && !MG_Config::Features.DisableSubgroup) {
m_nativeSubgroupSize = subgroupProperties.subgroupSize;
m_nativeSubgroupSupported = true;
}
}
VkPhysicalDeviceSubgroupSizeControlFeaturesEXT subgroupSizeControlFeatures{};
subgroupSizeControlFeatures.sType =
VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_SUBGROUP_SIZE_CONTROL_FEATURES_EXT;
m_maxComputeWorkgroupSubgroups = 0;
if (m_nativeSubgroupSupported &&
IsExtensionSupported(availableExtensions, VK_EXT_SUBGROUP_SIZE_CONTROL_EXTENSION_NAME) &&
getPhysicalDeviceFeatures2 != nullptr) {
VkPhysicalDeviceFeatures2 featureQuery{};
featureQuery.sType = VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_FEATURES_2;
featureQuery.pNext = &subgroupSizeControlFeatures;
getPhysicalDeviceFeatures2(m_physicalDevice.handle, &featureQuery);
if (getPhysicalDeviceProperties2 != nullptr) {
VkPhysicalDeviceSubgroupSizeControlPropertiesEXT subgroupSizeControlProperties{};
subgroupSizeControlProperties.sType =
VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_SUBGROUP_SIZE_CONTROL_PROPERTIES_EXT;
VkPhysicalDeviceProperties2 propertyQuery{};
propertyQuery.sType = VK_STRUCTURE_TYPE_PHYSICAL_DEVICE_PROPERTIES_2;
propertyQuery.pNext = &subgroupSizeControlProperties;
getPhysicalDeviceProperties2(m_physicalDevice.handle, &propertyQuery);
m_maxComputeWorkgroupSubgroups =
subgroupSizeControlProperties.maxComputeWorkgroupSubgroups;
}
if (subgroupSizeControlFeatures.computeFullSubgroups == VK_TRUE) {
if (!IsExtensionAlreadyEnabled(enabledDeviceExtensions,
VK_EXT_SUBGROUP_SIZE_CONTROL_EXTENSION_NAME)) {
enabledDeviceExtensions.push_back(VK_EXT_SUBGROUP_SIZE_CONTROL_EXTENSION_NAME);
}
// Only the full-subgroups guarantee is wanted; required/varying subgroup
// sizes stay unrequested.
subgroupSizeControlFeatures.subgroupSizeControl = VK_FALSE;
subgroupSizeControlFeatures.pNext = const_cast<void*>(deviceCreateInfo.pNext);
deviceCreateInfo.pNext = &subgroupSizeControlFeatures;
m_computeFullSubgroupsFeatureEnabled = true;
MGLOG_I("Enabled optional device extension: %s (computeFullSubgroups)",
VK_EXT_SUBGROUP_SIZE_CONTROL_EXTENSION_NAME);
}
}
// VK_EXT_transform_feedback backs GL transform feedback capture.
m_transformFeedbackFeatureEnabled = false;
VkPhysicalDeviceTransformFeedbackFeaturesEXT transformFeedbackFeatures{};
@@ -554,6 +554,16 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool m_samplerAnisotropyFeatureEnabled = false;
Bool m_shaderDrawParametersExtensionEnabled = false;
Bool m_shaderDrawParametersFeatureEnabled = false;
// Native subgroup topology, queried at device creation for the compute-module
// subgroup repairs (SubgroupSupportPolicy.h) and the REQUIRE_FULL_SUBGROUPS
// stage flag; 0 / false when the device has no usable compute subgroups or
// MOBILEGL_DISABLE_SUBGROUP forced them off.
Uint32 m_nativeSubgroupSize = 0;
Bool m_nativeSubgroupSupported = false;
Bool m_computeFullSubgroupsFeatureEnabled = false;
// VkPhysicalDeviceSubgroupSizeControlProperties::maxComputeWorkgroupSubgroups;
// 0 when the extension (and therefore the full-subgroups flag) is unavailable.
Uint32 m_maxComputeWorkgroupSubgroups = 0;
Bool m_unformattedFloatStorageImagesEnabled = false;
// Set only after descriptor-indexing feature AND property queries prove that
// update-after-bind is legal for every descriptor category this renderer emits.
@@ -0,0 +1,57 @@
// MobileGL - MobileGL/MG_Backend/DirectVulkan/SubgroupSupportPolicy.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include <Config.h>
#include <Includes.h>
namespace MobileGL::MG_Backend::DirectVulkan {
// The single decision point for how DirectVulkan implements GL_KHR_shader_subgroup,
// shared by capability advertisement (BackendObject) and module lowering
// (VulkanRenderer / ProgramFactory) so the two can never disagree.
//
// Native subgroups are the implementation whenever the device has them, whatever
// their width - subgroup operations execute on the hardware paths they were made
// for. Two module-level repairs keep the GL contract intact around them:
// - FixIterationRPSubgroupScratchPass patches the one known pack bug: iterationRP's
// prefixSumCache[32], under-declared for sub-16-lane devices (8-lane lavapipe);
// - DeriveNumSubgroupsPass replaces the one builtin drivers get wrong
// (gl_NumSubgroups) with the value the rest of the topology implies.
// The 32-lane shared-memory emulation (EmulateSubgroupsPass) is a LAST RESORT for
// devices with no subgroup support at all, and only when the user opts in with
// MOBILEGL_MAGMA_EMULATE_SUBGROUP=1; it never replaces available native operations.
inline constexpr Uint32 kEmulatedSubgroupSize = 32u;
inline constexpr Uint32 kEmulatedSubgroupStages = GL_COMPUTE_SHADER_BIT;
inline constexpr Uint32 kEmulatedSubgroupFeatures =
GL_SUBGROUP_FEATURE_BASIC_BIT_KHR | GL_SUBGROUP_FEATURE_VOTE_BIT_KHR |
GL_SUBGROUP_FEATURE_ARITHMETIC_BIT_KHR | GL_SUBGROUP_FEATURE_BALLOT_BIT_KHR |
GL_SUBGROUP_FEATURE_SHUFFLE_BIT_KHR | GL_SUBGROUP_FEATURE_SHUFFLE_RELATIVE_BIT_KHR |
GL_SUBGROUP_FEATURE_CLUSTERED_BIT_KHR | GL_SUBGROUP_FEATURE_QUAD_BIT_KHR;
inline Bool ShouldEmulateSubgroups(const Bool nativeSubgroupSupported) {
return MG_Config::Features.MagmaEmulateSubgroup && !nativeSubgroupSupported &&
!MG_Config::Features.DisableSubgroup;
}
inline Bool ShouldFixIterationRPSubgroupScratch() {
// Auto is ON: the patch is fingerprint-gated to iterationRP's reduction and
// grows one under-declared array; every other module passes through untouched.
return MG_Config::Features.FixIterationRPSubgroupScratch !=
MG_Config::QuirkOverride::ForceOff;
}
inline Bool ShouldDeriveNumSubgroups() {
// Auto is ON: gl_NumSubgroups must agree with the gl_SubgroupID range for the GL
// contract to hold, and the derived ceil() value is the one the renderer can pin
// with REQUIRE_FULL_SUBGROUPS - the driver builtin is the value with no
// cross-driver guarantee (Adreno returns 1 for an 8-subgroup dispatch).
return MG_Config::Features.DeriveNumSubgroups != MG_Config::QuirkOverride::ForceOff;
}
} // namespace MobileGL::MG_Backend::DirectVulkan
+2 -1
View File
@@ -68,7 +68,8 @@ add_executable(MobileGLIntegrationTest
Scenarios/DoublePrecisionScenario.cpp
Scenarios/UniformInitializerScenario.cpp
Scenarios/SwizzleAccessRoutineScenario.cpp
Scenarios/Program203FirstReductionScenario.cpp
Scenarios/IterationRPFirstReductionScenario.cpp
Scenarios/IterationRPScratchFixScenario.cpp
Scenarios/ProgramPipelineScenario.cpp
Scenarios/ImageLoadStoreSsoScenario.cpp
Scenarios/ImageTargetKindScenario.cpp
@@ -1,4 +1,4 @@
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/Program203FirstReductionScenario.cpp
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/IterationRPFirstReductionScenario.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
@@ -6,9 +6,9 @@
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Scenario - PROGRAM 203'S FIRST SUBGROUP REDUCTION.
// Scenario - ITERATIONRP'S FIRST SUBGROUP REDUCTION.
//
// Program 203 reduces a 32 x 16 exposure tile with a vector subgroup inclusive add,
// iterationRP reduces a 32 x 16 exposure tile with a vector subgroup inclusive add,
// then a shared-memory scan of subgroup totals. The source assumes that every
// subgroup has a last lane, that there are 2..32 subgroups, and that local index
// 511 belongs to the last subgroup and its last lane. Those are source assumptions,
@@ -133,6 +133,17 @@ namespace MGITest {
std::array<GLint, 3> maxWorkGroupSize{};
bool queryHadError = false;
// iterationRP's source contract needs gl_NumSubgroups in [2, 32] for its 512
// invocations, i.e. an advertised subgroup width in [16, 256]. A device
// outside that window (lavapipe's 8-lane subgroups give 64 subgroups) cannot
// run the fixture's verbatim reduction at all, so the scenario SKIPS there -
// the pack itself replays through the FixIterationRPSubgroupScratch patch, which
// this probe deliberately does not model. The width only gates the domain;
// lane placement and group counts still come from observed values alone.
bool SubgroupWidthInSourceDomain() const {
return subgroupSize >= 16 && subgroupSize <= 256;
}
bool SupportsProbe() const {
const auto stages = static_cast<GLbitfield>(supportedStages);
const auto features = static_cast<GLbitfield>(supportedFeatures);
@@ -140,6 +151,7 @@ namespace MGITest {
(stages & GL_COMPUTE_SHADER_BIT) != 0 &&
(features & (GL_SUBGROUP_FEATURE_BASIC_BIT_KHR | GL_SUBGROUP_FEATURE_ARITHMETIC_BIT_KHR)) ==
(GL_SUBGROUP_FEATURE_BASIC_BIT_KHR | GL_SUBGROUP_FEATURE_ARITHMETIC_BIT_KHR) &&
SubgroupWidthInSourceDomain() &&
maxComputeStorageBlocks >= 2 && maxStorageBindings >= 2 &&
maxWorkGroupInvocations >= static_cast<GLint>(kInvocationCount) && maxWorkGroupSize[0] >= 32 &&
maxWorkGroupSize[1] >= 16 && maxWorkGroupSize[2] >= 1;
@@ -159,6 +171,12 @@ namespace MGITest {
if ((features & requiredFeatures) != requiredFeatures) {
missing.emplace_back("basic|arithmetic in GL_SUBGROUP_SUPPORTED_FEATURES_KHR");
}
if (!SubgroupWidthInSourceDomain()) {
missing.emplace_back(
"GL_SUBGROUP_SIZE_KHR in [16, 256] (iterationRP's source contract needs "
"gl_NumSubgroups in [2, 32] for 512 invocations; width " +
std::to_string(subgroupSize) + " is outside the fixture's domain)");
}
if (maxComputeStorageBlocks < 2 || maxStorageBindings < 2) {
missing.emplace_back("two compute SSBO bindings");
}
@@ -194,7 +212,7 @@ namespace MGITest {
}
void PrintMetadata(const CapabilityInfo& info, std::ostream& output) {
output << "Program203FirstReductionScenario metadata: "
output << "IterationRPFirstReductionScenario metadata: "
<< "GL_SUBGROUP_SIZE_KHR=" << info.subgroupSize
<< ", GL_SUBGROUP_SUPPORTED_STAGES_KHR=0x" << std::hex
<< static_cast<GLbitfield>(info.supportedStages)
@@ -243,7 +261,7 @@ layout(std430, binding = 0) readonly buffer Input {
)";
// Only the expression producing tileExposure differs between the two
// tests. The remainder is the program-203 first reduction, with stores
// tests. The remainder is the iterationRP first reduction, with stores
// placed after its existing barriers to expose each handoff.
constexpr const char* kSampledTileExposure = R"(
vec2 texCoord = (vec2(gl_GlobalInvocationID.xy) + 0.5) *
@@ -506,7 +524,7 @@ layout(std430, binding = 0) readonly buffer Input {
if (!IsQuietNanSentinel(reduction.z) || !IsQuietNanSentinel(reduction.w) ||
!IsQuietNanSentinel(output.finalAverage[slot])) {
std::ostringstream message;
message << "program 203 source reduction has no valid contract for gl_NumSubgroups="
message << "iterationRP source reduction has no valid contract for gl_NumSubgroups="
<< reportedNumSubgroups << "; localIndex " << localIndex
<< " did not preserve its qNaN source-reduction sentinel";
return Failure("source domain", message.str());
@@ -514,7 +532,7 @@ layout(std430, binding = 0) readonly buffer Input {
for (std::size_t stage = 0; stage < kScanStageCount; ++stage) {
if (!IsQuietNanSentinel(output.scanAfter[stage][slot])) {
std::ostringstream message;
message << "program 203 source reduction has no valid contract for gl_NumSubgroups="
message << "iterationRP source reduction has no valid contract for gl_NumSubgroups="
<< reportedNumSubgroups << "; localIndex " << localIndex << ", scan stage " << stage
<< " did not preserve its qNaN source-reduction sentinel";
return Failure("source domain", message.str());
@@ -522,12 +540,12 @@ layout(std430, binding = 0) readonly buffer Input {
}
}
std::ostringstream message;
message << "program 203 source reduction has no valid contract for observed gl_NumSubgroups="
message << "iterationRP source reduction has no valid contract for observed gl_NumSubgroups="
<< reportedNumSubgroups << " (requires 2..32); native subgroup results were recorded";
return Failure("source domain", message.str());
}
// 4. Program-203 source writer and first shared-memory handoff.
// 4. iterationRP source writer and first shared-memory handoff.
std::vector<std::size_t> sourceWriter(reportedNumSubgroups, kNoSlot);
for (std::uint32_t subgroupID = 0; subgroupID < reportedNumSubgroups; ++subgroupID) {
std::size_t writerCount = 0;
@@ -541,7 +559,7 @@ layout(std430, binding = 0) readonly buffer Input {
if (writerCount != 1u) {
std::ostringstream message;
message << "subgroupID " << subgroupID << " has " << writerCount
<< " recorded lane(s) where laneID == subgroupSize - 1; program 203 leaves that "
<< " recorded lane(s) where laneID == subgroupSize - 1; iterationRP leaves that "
"shared-cache entry unwritten";
return Failure("source writer", message.str());
}
@@ -633,7 +651,7 @@ layout(std430, binding = 0) readonly buffer Input {
index511Subgroup.z == ownerResult.highestObservedSubgroup;
if (!ownerResult.index511IsSourceLastLaneWriter || !ownerResult.index511IsHighestSubgroupMember) {
std::ostringstream message;
message << "program 203 topology incompatibility: localIndex 511 is sourceLastLaneWriter="
message << "iterationRP topology incompatibility: localIndex 511 is sourceLastLaneWriter="
<< ownerResult.index511IsSourceLastLaneWriter << ", highestSubgroupMember="
<< ownerResult.index511IsHighestSubgroupMember << " (subgroupID=" << index511Subgroup.z
<< ", highest observed subgroupID=" << ownerResult.highestObservedSubgroup << ')';
@@ -650,7 +668,7 @@ layout(std430, binding = 0) readonly buffer Input {
const float expectedTotal = mode == InputMode::IndexedSsbo ? 131328.0f : sampledExpectedTotal;
if (!SameBits(total, expectedTotal) || !SameBits(mergedPrefix[index511Slot], expectedTotal)) {
std::ostringstream message;
message << "program 203 source total was " << FormatFloat(mergedPrefix[index511Slot])
message << "iterationRP source total was " << FormatFloat(mergedPrefix[index511Slot])
<< " (native total " << FormatFloat(total) << "), expected " << FormatFloat(expectedTotal);
ownerResult.ok = false;
ownerResult.phase = "final average";
@@ -675,9 +693,9 @@ layout(std430, binding = 0) readonly buffer Input {
bool includeScanStages) {
PrintMetadata(capabilities, std::cout);
if (validation.ok) {
std::cout << "Program203FirstReductionScenario firstFailure=none\n";
std::cout << "IterationRPFirstReductionScenario firstFailure=none\n";
} else {
std::cout << "Program203FirstReductionScenario firstFailure=" << validation.phase << ": "
std::cout << "IterationRPFirstReductionScenario firstFailure=" << validation.phase << ": "
<< validation.message << '\n';
}
std::cout << "localIndex,localX,localY,localZ,subgroupSize,numSubgroups,subgroupID,laneID,input,"
@@ -702,17 +720,19 @@ layout(std430, binding = 0) readonly buffer Input {
}
}
class Program203FirstReductionScenario : public ScenarioTest {
class IterationRPFirstReductionScenario : public ScenarioTest {
protected:
void SetUp() override {
ScenarioTest::SetUp();
if (!Ready()) return;
m_capabilities = QueryCapabilities();
// GL_SUBGROUP_SIZE_KHR is diagnostic only. It is deliberately
// never used to infer lane placement or an expected group count.
// GL_SUBGROUP_SIZE_KHR gates only whether the fixture's source contract
// can hold on this device (SubgroupWidthInSourceDomain); it is
// deliberately never used to infer lane placement or an expected group
// count - those come from observed values alone.
PrintMetadata(m_capabilities, std::cout);
RecordProperty("program203_gl_subgroup_size_khr", std::to_string(m_capabilities.subgroupSize));
RecordProperty("iterationrp_gl_subgroup_size_khr", std::to_string(m_capabilities.subgroupSize));
if (!m_capabilities.SupportsProbe()) {
GTEST_SKIP() << "subgroup probe requires " << m_capabilities.MissingRequirements();
}
@@ -839,13 +859,13 @@ layout(std430, binding = 0) readonly buffer Input {
const ValidationResult validation = ValidateProbe(output, mode);
if (validation.ownerEvaluated) {
RecordProperty("program203_index511_source_last_lane_writer",
RecordProperty("iterationrp_index511_source_last_lane_writer",
validation.index511IsSourceLastLaneWriter ? "true" : "false");
RecordProperty("program203_index511_highest_subgroup_member",
RecordProperty("iterationrp_index511_highest_subgroup_member",
validation.index511IsHighestSubgroupMember ? "true" : "false");
RecordProperty("program203_highest_observed_subgroup",
RecordProperty("iterationrp_highest_observed_subgroup",
std::to_string(validation.highestObservedSubgroup));
std::cout << "Program203FirstReductionScenario owner: localIndex511 sourceLastLaneWriter="
std::cout << "IterationRPFirstReductionScenario owner: localIndex511 sourceLastLaneWriter="
<< validation.index511IsSourceLastLaneWriter << ", highestSubgroupMember="
<< validation.index511IsHighestSubgroupMember << ", highestObservedSubgroup="
<< validation.highestObservedSubgroup << '\n';
@@ -865,12 +885,12 @@ layout(std430, binding = 0) readonly buffer Input {
} // namespace
TEST_F(Program203FirstReductionScenario, SampledRgba32fFirstAverage) {
TEST_F(IterationRPFirstReductionScenario, SampledRgba32fFirstAverage) {
if (!Ready() || IsSkipped()) return;
RunAndValidate(InputMode::SampledRgba32f);
}
TEST_F(Program203FirstReductionScenario, IndexedInputTopologyAndReduction) {
TEST_F(IterationRPFirstReductionScenario, IndexedInputTopologyAndReduction) {
if (!Ready() || IsSkipped()) return;
RunAndValidate(InputMode::IndexedSsbo);
}
@@ -0,0 +1,302 @@
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/IterationRPScratchFixScenario.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Scenario - THE FIXTURE-SHAPED SUBGROUP REDUCTION, ON WHATEVER WIDTH THE DEVICE HAS.
//
// iterationRP hard-sizes the scratch its subgroup prefix scans write through
// prefixSumCache[gl_SubgroupID], and ships that idiom twice: the auto-exposure pass
// declares `shared vec2 prefixSumCache[32]` for a 512-invocation workgroup, and the
// RTW importance warp declares `shared float prefixSumCache[64]` for a 1024-invocation
// one. Both algorithms are width-agnostic; only the static lengths bake in "at most 32
// (respectively 64) subgroups", which every desktop capture satisfies and an 8-lane
// device (lavapipe: 64 and 128 subgroups) does not. DirectVulkan patches exactly that with
// FixIterationRPSubgroupScratchPass, growing the array to ceil(invocations / native
// width) on the modules that match the pack's reduction fingerprint.
//
// This scenario replays the fixture's reduction shape verbatim - the same 32-entry
// declaration, the same last-lane handoff, the same findMSB combine loop, and NO
// domain guard - and asserts only the width-independent result: the workgroup total.
// The inputs are small integers, so the fp32 sum is exact under any lane order and any
// association; a correct run produces the exact constant on a 4-lane device and a
// 128-lane device alike. Without the patch, a sub-16-lane device indexes the
// 32-entry array out of bounds - on lavapipe that is literal heap corruption - and
// this scenario is the regression test that keeps the patch working, and it runs on every device that
// has basic+arithmetic compute subgroups (unlike IterationRPFirstReductionScenario,
// which probes the UNREPAIRED source contract and must skip outside [16, 256]).
#include <cstdint>
#include <cstring>
#include <string>
#include "../Harness/HeadlessGL.h"
#include "../Harness/ScenarioFixture.h"
#ifdef GLAPI
#undef GLAPI
#endif
#define GL_GLEXT_PROTOTYPES
#include <GL/gl.h>
#include <GL/glcorearb.h>
#undef GL_GLEXT_PROTOTYPES
namespace MGITest {
namespace {
constexpr std::uint32_t kInvocationCount = 512u;
// sum of 0..511, exactly representable and associativity-proof in fp32.
constexpr float kExpectedTotal = 130816.0f;
// The RTW warp's shape: 1024 invocations into a 64-entry float scratch.
constexpr std::uint32_t kWideInvocationCount = 1024u;
// sum of 0..1023, likewise exact in fp32.
constexpr float kWideExpectedTotal = 523776.0f;
constexpr const char* kComputeSource = R"(#version 430 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output {
float total;
uint numSubgroups;
uint maxSubgroupId;
} outputData;
shared vec2 prefixSumCache[32];
void main() {
vec2 sampleLuminance = vec2(float(gl_LocalInvocationIndex), 0.0);
sampleLuminance = subgroupInclusiveAdd(sampleLuminance);
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = sampleLuminance;
barrier();
uint loopLength = uint(findMSB(gl_NumSubgroups));
loopLength += uint(gl_NumSubgroups - (1u << (loopLength - 1u)) > 0u);
for (uint scanStage = 0u; scanStage < loopLength; ++scanStage) {
if ((gl_SubgroupID & (1u << scanStage)) > 0u) {
sampleLuminance += prefixSumCache[(gl_SubgroupID >> scanStage << scanStage) - 1u];
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = sampleLuminance;
}
barrier();
}
if (gl_LocalInvocationIndex == 511u) {
outputData.total = sampleLuminance.x;
outputData.numSubgroups = gl_NumSubgroups;
}
atomicMax(outputData.maxSubgroupId, gl_SubgroupID);
}
)";
// The RTW importance warp's shape: a plain float scan over 1024 invocations
// into a 64-entry scratch. Same idiom, different dimensions - which is exactly
// what a fingerprint pinned to the exposure pass's shape walks past.
constexpr const char* kWideComputeSource = R"(#version 430 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 1024) in;
layout(std430, binding = 0) buffer Output {
float total;
uint numSubgroups;
uint maxSubgroupId;
} outputData;
shared float prefixSumCache[64];
void main() {
float importance = float(gl_LocalInvocationID.x);
float prefixSum = subgroupInclusiveAdd(importance);
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = prefixSum;
barrier();
uint loopLength = uint(findMSB(gl_NumSubgroups));
loopLength += uint(gl_NumSubgroups - (1u << (loopLength - 1u)) > 0u);
for (uint scanStage = 0u; scanStage < loopLength; ++scanStage) {
if ((gl_SubgroupID & (1u << scanStage)) > 0u) {
prefixSum += prefixSumCache[(gl_SubgroupID >> scanStage << scanStage) - 1u];
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = prefixSum;
}
barrier();
}
if (gl_LocalInvocationID.x == 1023u) {
outputData.total = prefixSum;
outputData.numSubgroups = gl_NumSubgroups;
}
atomicMax(outputData.maxSubgroupId, gl_SubgroupID);
}
)";
struct OutputBlock {
float total = -1.0f;
std::uint32_t numSubgroups = 0;
std::uint32_t maxSubgroupId = 0;
};
bool HasExtension(const char* wanted) {
GLint extensionCount = 0;
glGetIntegerv(GL_NUM_EXTENSIONS, &extensionCount);
for (GLint i = 0; i < extensionCount; ++i) {
const auto* extension =
reinterpret_cast<const char*>(glGetStringi(GL_EXTENSIONS, static_cast<GLuint>(i)));
if (extension != nullptr && std::string(extension) == wanted) return true;
}
return false;
}
class IterationRPScratchFixScenario : public ScenarioTest {
protected:
void SetUp() override {
ScenarioTest::SetUp();
if (!Ready()) return;
GLint stages = 0;
GLint features = 0;
GLint invocations = 0;
const bool subgroupExtension = HasExtension("GL_KHR_shader_subgroup");
if (subgroupExtension) {
glGetIntegerv(GL_SUBGROUP_SUPPORTED_STAGES_KHR, &stages);
glGetIntegerv(GL_SUBGROUP_SUPPORTED_FEATURES_KHR, &features);
}
glGetIntegerv(GL_MAX_COMPUTE_WORK_GROUP_INVOCATIONS, &invocations);
const GLbitfield requiredFeatures =
GL_SUBGROUP_FEATURE_BASIC_BIT_KHR | GL_SUBGROUP_FEATURE_ARITHMETIC_BIT_KHR;
if (!subgroupExtension || (static_cast<GLbitfield>(stages) & GL_COMPUTE_SHADER_BIT) == 0 ||
(static_cast<GLbitfield>(features) & requiredFeatures) != requiredFeatures ||
invocations < static_cast<GLint>(kInvocationCount)) {
GTEST_SKIP() << "needs GL_KHR_shader_subgroup basic+arithmetic in compute and a "
"512-invocation workgroup";
}
m_maxInvocations = static_cast<std::uint32_t>(invocations);
glGenBuffers(1, &m_output);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
// maxSubgroupId starts at zero HOST-side: the word is touched only by
// atomicMax during the dispatch, since a plain shader-side zeroing store
// would race the other invocations' atomics (barrier() orders shared
// memory, not SSBO stores).
const OutputBlock poison{-1.0f, 0xa5a5a5a5u, 0u};
glBufferData(GL_SHADER_STORAGE_BUFFER, sizeof(OutputBlock), &poison, GL_DYNAMIC_READ);
glBindBufferBase(GL_SHADER_STORAGE_BUFFER, 0, m_output);
}
void TearDown() override {
if (!Ready()) return;
glBindBufferBase(GL_SHADER_STORAGE_BUFFER, 0, 0);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, 0);
if (m_output != 0) glDeleteBuffers(1, &m_output);
if (m_program != 0) glDeleteProgram(m_program);
}
unsigned int CompileComputeProgram(const char* source) {
const GLuint shader = glCreateShader(GL_COMPUTE_SHADER);
glShaderSource(shader, 1, &source, nullptr);
glCompileShader(shader);
GLint compiled = 0;
glGetShaderiv(shader, GL_COMPILE_STATUS, &compiled);
if (compiled == GL_FALSE) {
char log[2048] = {};
glGetShaderInfoLog(shader, sizeof(log) - 1, nullptr, log);
m_buildLog = std::string("compute shader did not compile: ") + log;
glDeleteShader(shader);
return 0;
}
const GLuint program = glCreateProgram();
glAttachShader(program, shader);
glLinkProgram(program);
glDeleteShader(shader);
GLint linked = 0;
glGetProgramiv(program, GL_LINK_STATUS, &linked);
if (linked == GL_FALSE) {
char log[2048] = {};
glGetProgramInfoLog(program, sizeof(log) - 1, nullptr, log);
m_buildLog = std::string("compute program did not link: ") + log;
glDeleteProgram(program);
return 0;
}
return program;
}
// Re-poisons the block, compiles the shape under test and runs it once.
OutputBlock Dispatch(const char* source) {
const OutputBlock poison{-1.0f, 0xa5a5a5a5u, 0u};
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
glBufferSubData(GL_SHADER_STORAGE_BUFFER, 0, sizeof(OutputBlock), &poison);
m_program = CompileComputeProgram(source);
EXPECT_NE(m_program, 0u) << m_buildLog;
if (m_program == 0u) return OutputBlock{};
glUseProgram(m_program);
glDispatchCompute(1, 1, 1);
glMemoryBarrier(GL_BUFFER_UPDATE_BARRIER_BIT);
OutputBlock block{};
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
glGetBufferSubData(GL_SHADER_STORAGE_BUFFER, 0, sizeof(OutputBlock), &block);
return block;
}
GLuint m_program = 0;
GLuint m_output = 0;
std::uint32_t m_maxInvocations = 0;
std::string m_buildLog;
};
} // namespace
TEST_F(IterationRPScratchFixScenario, FixtureShapedReductionSumsEveryInvocation) {
const OutputBlock block = Dispatch(kComputeSource);
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
// The topology diagnostics catch the failure modes by name before the sum does:
// an out-of-bounds handoff corrupts the total, a wrong gl_NumSubgroups breaks
// the combine loop's length.
ASSERT_NE(block.numSubgroups, 0xa5a5a5a5u) << "invocation 511 never reached its store";
EXPECT_GE(block.numSubgroups, 1u);
EXPECT_LE(block.numSubgroups, kInvocationCount);
EXPECT_LT(block.maxSubgroupId, block.numSubgroups)
<< "gl_SubgroupID exceeds gl_NumSubgroups - the inconsistency "
"DeriveNumSubgroupsPass exists to repair";
// Integer-valued fp32 inputs: the workgroup total is exact under any subgroup
// width, lane order, and association. This is the value iterationRP's exposure
// average is built from; without FixIterationRPSubgroupScratchPass an 8-lane
// device writes prefixSumCache[32..63] out of bounds and this comparison fails.
EXPECT_EQ(block.total, kExpectedTotal)
<< "workgroup reduction produced " << block.total << " with gl_NumSubgroups="
<< block.numSubgroups;
}
// The pack's second instance of the same bug, and the one that kept the CI
// retrace red after the exposure pass alone was patched.
TEST_F(IterationRPScratchFixScenario, WideFixtureShapedReductionSumsEveryInvocation) {
if (m_maxInvocations < kWideInvocationCount) {
GTEST_SKIP() << "needs a " << kWideInvocationCount << "-invocation workgroup";
}
const OutputBlock block = Dispatch(kWideComputeSource);
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
ASSERT_NE(block.numSubgroups, 0xa5a5a5a5u) << "invocation 1023 never reached its store";
EXPECT_GE(block.numSubgroups, 1u);
EXPECT_LE(block.numSubgroups, kWideInvocationCount);
EXPECT_LT(block.maxSubgroupId, block.numSubgroups)
<< "gl_SubgroupID exceeds gl_NumSubgroups - the inconsistency "
"DeriveNumSubgroupsPass exists to repair";
// Without the patch an 8-lane device writes prefixSumCache[64..127] out of
// bounds and this comparison fails.
EXPECT_EQ(block.total, kWideExpectedTotal)
<< "workgroup reduction produced " << block.total << " with gl_NumSubgroups="
<< block.numSubgroups;
}
} // namespace MGITest
+5 -5
View File
@@ -1,19 +1,19 @@
# MobileGL - MobileGL/MG_Test/SelfTest/CMakeLists.txt
add_executable(
DriverPostProgram203WitnessTest
DriverPostProgram203WitnessTest.cpp
DriverPostIterationRPWitnessTest
DriverPostIterationRPWitnessTest.cpp
)
target_include_directories(DriverPostProgram203WitnessTest PRIVATE
target_include_directories(DriverPostIterationRPWitnessTest PRIVATE
${MGL_ROOT}/include
${MGL_ROOT}/MobileGL
)
target_link_libraries(DriverPostProgram203WitnessTest PRIVATE
target_link_libraries(DriverPostIterationRPWitnessTest PRIVATE
GTest::gtest_main
${LINK_LIBRARIES}
)
include(GoogleTest)
gtest_discover_tests(DriverPostProgram203WitnessTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
gtest_discover_tests(DriverPostIterationRPWitnessTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
@@ -1,4 +1,4 @@
// MobileGL - MobileGL/MG_Test/SelfTest/DriverPostProgram203WitnessTest.cpp
// MobileGL - MobileGL/MG_Test/SelfTest/DriverPostIterationRPWitnessTest.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
@@ -10,21 +10,21 @@
#include <string>
#include "MG_Util/SelfTest/DriverPostProgram203Witness.h"
#include "MG_Util/SelfTest/DriverPostIterationRPWitness.h"
namespace MobileGL::MG_Util::SelfTest {
namespace {
Program203WitnessOutput MakeValidWitness(std::uint32_t numSubgroups) {
Program203WitnessOutput output{};
output.magic = kProgram203WitnessMagic;
IterationRPWitnessOutput MakeValidWitness(std::uint32_t numSubgroups) {
IterationRPWitnessOutput output{};
output.magic = kIterationRPWitnessMagic;
output.numSubgroups = numSubgroups;
output.loopLength = ComputeProgram203WitnessLoopLength(numSubgroups);
output.loopLength = ComputeIterationRPWitnessLoopLength(numSubgroups);
output.seenSubgroupMask =
numSubgroups == kProgram203WitnessMaxSubgroups ? 0xffffffffu : (1u << numSubgroups) - 1u;
numSubgroups == kIterationRPWitnessMaxSubgroups ? 0xffffffffu : (1u << numSubgroups) - 1u;
// Valid test layouts use equal contiguous groups of the indexed
// 1..512 input. The compact witness only needs their independent sums.
const std::uint32_t subgroupSize = kProgram203WitnessInvocationCount / numSubgroups;
const std::uint32_t subgroupSize = kIterationRPWitnessInvocationCount / numSubgroups;
for (std::uint32_t subgroup = 0u; subgroup < numSubgroups; ++subgroup) {
const std::uint32_t first = subgroup * subgroupSize + 1u;
const std::uint32_t last = first + subgroupSize - 1u;
@@ -50,139 +50,139 @@ namespace MobileGL::MG_Util::SelfTest {
return output;
}
Program203WitnessLimits MakeSufficientLimits() {
Program203WitnessLimits limits;
IterationRPWitnessLimits MakeSufficientLimits() {
IterationRPWitnessLimits limits;
limits.computeStageSupported = true;
limits.basicSubgroupSupported = true;
limits.arithmeticSubgroupSupported = true;
limits.subgroupSize = 32u;
limits.maxComputeWorkGroupInvocations = kProgram203WitnessInvocationCount;
limits.maxComputeWorkGroupInvocations = kIterationRPWitnessInvocationCount;
limits.maxComputeWorkGroupSize = {32u, 16u, 1u};
limits.maxComputeSharedMemorySize = kProgram203WitnessSharedMemoryBytes;
limits.maxComputeSharedMemorySize = kIterationRPWitnessSharedMemoryBytes;
limits.maxPerStageDescriptorStorageBuffers = 1u;
limits.maxDescriptorSetStorageBuffers = 1u;
limits.maxBoundDescriptorSets = 1u;
limits.maxStorageBufferRange = sizeof(Program203WitnessOutput);
limits.maxStorageBufferRange = sizeof(IterationRPWitnessOutput);
return limits;
}
} // namespace
TEST(DriverPostProgram203WitnessTest, ValidTwoSubgroupWitness) {
const Program203WitnessValidationResult validation = ValidateProgram203Witness(MakeValidWitness(2u));
TEST(DriverPostIterationRPWitnessTest, ValidTwoSubgroupWitness) {
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(MakeValidWitness(2u));
ASSERT_TRUE(validation.ok) << validation.detail;
EXPECT_EQ(validation.detail, "N=2, owner511=id1/lane255, 2 scan stages, average=(256.5,0)");
}
TEST(DriverPostProgram203WitnessTest, ValidThirtyTwoSubgroupWitness) {
const Program203WitnessValidationResult validation = ValidateProgram203Witness(MakeValidWitness(32u));
TEST(DriverPostIterationRPWitnessTest, ValidThirtyTwoSubgroupWitness) {
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(MakeValidWitness(32u));
ASSERT_TRUE(validation.ok) << validation.detail;
EXPECT_EQ(validation.detail, "N=32, owner511=id31/lane15, 6 scan stages, average=(256.5,0)");
}
TEST(DriverPostProgram203WitnessTest, RejectsNonuniformNumSubgroups) {
Program203WitnessOutput output = MakeValidWitness(16u);
output.topologyFlags |= Program203WitnessNonuniformNumSubgroups;
const Program203WitnessValidationResult validation = ValidateProgram203Witness(output);
TEST(DriverPostIterationRPWitnessTest, RejectsNonuniformNumSubgroups) {
IterationRPWitnessOutput output = MakeValidWitness(16u);
output.topologyFlags |= IterationRPWitnessNonuniformNumSubgroups;
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(output);
EXPECT_FALSE(validation.ok);
EXPECT_EQ(validation.failure, Program203WitnessValidationFailure::Topology);
EXPECT_EQ(validation.failure, IterationRPWitnessValidationFailure::Topology);
EXPECT_NE(validation.detail.find("gl_NumSubgroups differed"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, RejectsMissingAndOutOfRangeSubgroupIds) {
Program203WitnessOutput missing = MakeValidWitness(16u);
TEST(DriverPostIterationRPWitnessTest, RejectsMissingAndOutOfRangeSubgroupIds) {
IterationRPWitnessOutput missing = MakeValidWitness(16u);
missing.seenSubgroupMask &= ~(1u << 7u);
Program203WitnessValidationResult validation = ValidateProgram203Witness(missing);
IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(missing);
EXPECT_FALSE(validation.ok);
EXPECT_NE(validation.detail.find("seen subgroup-ID mask"), std::string::npos);
Program203WitnessOutput outOfRange = MakeValidWitness(16u);
outOfRange.topologyFlags |= Program203WitnessInvalidSubgroupId;
validation = ValidateProgram203Witness(outOfRange);
IterationRPWitnessOutput outOfRange = MakeValidWitness(16u);
outOfRange.topologyFlags |= IterationRPWitnessInvalidSubgroupId;
validation = ValidateIterationRPWitness(outOfRange);
EXPECT_FALSE(validation.ok);
EXPECT_NE(validation.detail.find("invalid gl_SubgroupID"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, RejectsInvalidMultipleAndMissingLastLaneWriters) {
Program203WitnessOutput invalidLane = MakeValidWitness(16u);
invalidLane.topologyFlags |= Program203WitnessInvalidSubgroupLane;
Program203WitnessValidationResult validation = ValidateProgram203Witness(invalidLane);
TEST(DriverPostIterationRPWitnessTest, RejectsInvalidMultipleAndMissingLastLaneWriters) {
IterationRPWitnessOutput invalidLane = MakeValidWitness(16u);
invalidLane.topologyFlags |= IterationRPWitnessInvalidSubgroupLane;
IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(invalidLane);
EXPECT_FALSE(validation.ok);
EXPECT_NE(validation.detail.find("invalid subgroup lane"), std::string::npos);
Program203WitnessOutput multiple = MakeValidWitness(16u);
IterationRPWitnessOutput multiple = MakeValidWitness(16u);
multiple.lastLaneWriterCount[4] = 2u;
validation = ValidateProgram203Witness(multiple);
validation = ValidateIterationRPWitness(multiple);
EXPECT_FALSE(validation.ok);
EXPECT_NE(validation.detail.find("subgroup 4 has 2 source last-lane writers"), std::string::npos);
Program203WitnessOutput missing = MakeValidWitness(16u);
IterationRPWitnessOutput missing = MakeValidWitness(16u);
missing.lastLaneWriterCount[6] = 0u;
validation = ValidateProgram203Witness(missing);
validation = ValidateIterationRPWitness(missing);
EXPECT_FALSE(validation.ok);
EXPECT_NE(validation.detail.find("subgroup 6 has 0 source last-lane writers"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, ReportsEarliestCorruptSourceScanStage) {
Program203WitnessOutput output = MakeValidWitness(32u);
TEST(DriverPostIterationRPWitnessTest, ReportsEarliestCorruptSourceScanStage) {
IterationRPWitnessOutput output = MakeValidWitness(32u);
output.scanCache[0][1].x += 1.0f;
output.scanCache[3][5].x += 1.0f;
Program203WitnessValidationResult validation = ValidateProgram203Witness(output);
IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(output);
EXPECT_FALSE(validation.ok);
EXPECT_EQ(validation.failure, Program203WitnessValidationFailure::SourceScan);
EXPECT_EQ(validation.failure, IterationRPWitnessValidationFailure::SourceScan);
EXPECT_EQ(validation.scanStage, 0u);
EXPECT_NE(validation.detail.find("source scan stage 0, subgroup 1"), std::string::npos);
output = MakeValidWitness(32u);
output.scanCache[3][5].x += 1.0f;
validation = ValidateProgram203Witness(output);
validation = ValidateIterationRPWitness(output);
EXPECT_FALSE(validation.ok);
EXPECT_EQ(validation.failure, Program203WitnessValidationFailure::SourceScan);
EXPECT_EQ(validation.failure, IterationRPWitnessValidationFailure::SourceScan);
EXPECT_EQ(validation.scanStage, 3u);
EXPECT_NE(validation.detail.find("source scan stage 3, subgroup 5"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, RejectsOwner511OutsideHighestFinalLane) {
Program203WitnessOutput output = MakeValidWitness(16u);
TEST(DriverPostIterationRPWitnessTest, RejectsOwner511OutsideHighestFinalLane) {
IterationRPWitnessOutput output = MakeValidWitness(16u);
output.owner511.z = 14u;
const Program203WitnessValidationResult validation = ValidateProgram203Witness(output);
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(output);
EXPECT_FALSE(validation.ok);
EXPECT_EQ(validation.failure, Program203WitnessValidationFailure::FinalOwner);
EXPECT_EQ(validation.failure, IterationRPWitnessValidationFailure::FinalOwner);
EXPECT_NE(validation.detail.find("not in the highest subgroup"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, RejectsIncorrectVectorFinalAverage) {
Program203WitnessOutput output = MakeValidWitness(16u);
TEST(DriverPostIterationRPWitnessTest, RejectsIncorrectVectorFinalAverage) {
IterationRPWitnessOutput output = MakeValidWitness(16u);
output.finalAverage.y = 1.0f;
const Program203WitnessValidationResult validation = ValidateProgram203Witness(output);
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(output);
EXPECT_FALSE(validation.ok);
EXPECT_EQ(validation.failure, Program203WitnessValidationFailure::FinalAverage);
EXPECT_EQ(validation.failure, IterationRPWitnessValidationFailure::FinalAverage);
EXPECT_NE(validation.detail.find("final average"), std::string::npos);
}
TEST(DriverPostProgram203WitnessTest, MissingNativeFeatureIsTheOnlySkipCondition) {
TEST(DriverPostIterationRPWitnessTest, MissingNativeFeatureIsTheOnlySkipCondition) {
for (const auto toggleMissingFeature : {0u, 1u, 2u}) {
Program203WitnessLimits limits = MakeSufficientLimits();
IterationRPWitnessLimits limits = MakeSufficientLimits();
if (toggleMissingFeature == 0u) limits.computeStageSupported = false;
if (toggleMissingFeature == 1u) limits.basicSubgroupSupported = false;
if (toggleMissingFeature == 2u) limits.arithmeticSubgroupSupported = false;
const Program203WitnessEligibilityResult eligibility = EvaluateProgram203WitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, Program203WitnessEligibility::SkipUnsupportedNativeFeatureSet)
const IterationRPWitnessEligibilityResult eligibility = EvaluateIterationRPWitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, IterationRPWitnessEligibility::SkipUnsupportedNativeFeatureSet)
<< eligibility.detail;
}
Program203WitnessLimits zeroSubgroupSize = MakeSufficientLimits();
IterationRPWitnessLimits zeroSubgroupSize = MakeSufficientLimits();
zeroSubgroupSize.subgroupSize = 0u;
Program203WitnessEligibilityResult eligibility = EvaluateProgram203WitnessEligibility(zeroSubgroupSize);
EXPECT_EQ(eligibility.eligibility, Program203WitnessEligibility::FailInadequateLimits) << eligibility.detail;
IterationRPWitnessEligibilityResult eligibility = EvaluateIterationRPWitnessEligibility(zeroSubgroupSize);
EXPECT_EQ(eligibility.eligibility, IterationRPWitnessEligibility::FailInadequateLimits) << eligibility.detail;
Program203WitnessLimits limits = MakeSufficientLimits();
IterationRPWitnessLimits limits = MakeSufficientLimits();
limits.maxComputeWorkGroupInvocations = 511u;
eligibility = EvaluateProgram203WitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, Program203WitnessEligibility::FailInadequateLimits) << eligibility.detail;
eligibility = EvaluateIterationRPWitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, IterationRPWitnessEligibility::FailInadequateLimits) << eligibility.detail;
limits = MakeSufficientLimits();
limits.maxStorageBufferRange = sizeof(Program203WitnessOutput) - 1u;
eligibility = EvaluateProgram203WitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, Program203WitnessEligibility::FailInadequateLimits) << eligibility.detail;
limits.maxStorageBufferRange = sizeof(IterationRPWitnessOutput) - 1u;
eligibility = EvaluateIterationRPWitnessEligibility(limits);
EXPECT_EQ(eligibility.eligibility, IterationRPWitnessEligibility::FailInadequateLimits) << eligibility.detail;
}
} // namespace MobileGL::MG_Util::SelfTest
@@ -4,6 +4,8 @@ add_executable(
SpirvPassTest
SpirvPassTest.cpp
DeriveNumSubgroupsTest.cpp
FixIterationRPSubgroupScratchTest.cpp
EmulateSubgroupsTest.cpp
DemoteFloat64Test.cpp
FlattenXfbInterfaceBlocksTest.cpp
)
@@ -0,0 +1,248 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/EmulateSubgroupsTest.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#define SPV_ENABLE_UTILITY_CODE
#include "glslang/SPIRV/spirv.hpp11"
#undef SPV_ENABLE_UTILITY_CODE
#include "Includes.h"
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <MG_Util/ShaderTranspiler/Types.h>
#include <spirv-tools/libspirv.hpp>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
namespace {
constexpr SizeT kSpirvHeaderWordCount = 5u;
template <typename Visitor>
void ForEachInstruction(const Vector<Uint32>& spirv, Visitor&& visit) {
for (SizeT offset = kSpirvHeaderWordCount; offset < spirv.size();) {
const Uint32 wordCount = spirv[offset] >> 16u;
if (wordCount == 0u || offset + wordCount > spirv.size()) break;
visit(static_cast<spv::Op>(spirv[offset] & 0xffffu), &spirv[offset], wordCount);
offset += wordCount;
}
}
Vector<Uint32> CompileStage(GLenum stage, const String& source) {
using namespace MobileGL::MG_Util::ShaderTranspiler;
ShaderAttrib shaderAttrib{.shaderType = stage, .sourceStr = source};
auto shaderResult = ShaderCompiler::CompileShader(shaderAttrib);
EXPECT_TRUE(shaderResult) << (shaderResult ? String{} : shaderResult.error().log);
if (!shaderResult) return {};
ProgramAttrib programAttrib{.shaders = {shaderResult.value()}};
auto programResult = ShaderCompiler::LinkProgram(programAttrib);
EXPECT_TRUE(programResult) << (programResult ? String{} : programResult.error().log);
if (!programResult) return {};
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {stage}, .program = *programResult.value()};
auto binaryResult = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
EXPECT_TRUE(binaryResult) << (binaryResult ? String{} : binaryResult.error().log);
if (!binaryResult || binaryResult->empty()) return {};
return binaryResult->front();
}
Uint32 CountGroupNonUniform(const Vector<Uint32>& spirv) {
Uint32 count = 0;
ForEachInstruction(spirv, [&](spv::Op opcode, const Uint32*, Uint32) {
if (opcode >= spv::Op::OpGroupNonUniformElect && opcode <= spv::Op::OpGroupNonUniformQuadSwap) {
++count;
}
});
return count;
}
Uint32 CountGroupNonUniformCapabilities(const Vector<Uint32>& spirv) {
Uint32 count = 0;
ForEachInstruction(spirv, [&](spv::Op opcode, const Uint32* words, Uint32 wordCount) {
if (opcode != spv::Op::OpCapability || wordCount < 2u) return;
const auto capability = static_cast<spv::Capability>(words[1]);
if (capability >= spv::Capability::GroupNonUniform &&
capability <= spv::Capability::GroupNonUniformQuad) {
++count;
}
});
return count;
}
Uint32 CountOpcode(const Vector<Uint32>& spirv, spv::Op wanted) {
Uint32 count = 0;
ForEachInstruction(spirv, [&](spv::Op opcode, const Uint32*, Uint32) {
if (opcode == wanted) ++count;
});
return count;
}
bool HasWorkgroupVariable(const Vector<Uint32>& spirv) {
bool found = false;
ForEachInstruction(spirv, [&](spv::Op opcode, const Uint32* words, Uint32 wordCount) {
if (opcode == spv::Op::OpVariable && wordCount >= 4u &&
static_cast<spv::StorageClass>(words[3]) == spv::StorageClass::Workgroup) {
found = true;
}
});
return found;
}
bool Validates(const Vector<Uint32>& spirv) {
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
tools.SetMessageConsumer([](spv_message_level_t, const char*, const spv_position_t& position,
const char* message) {
ADD_FAILURE() << "spirv-val at word " << position.index << ": " << message;
});
return tools.Validate(spirv);
}
// One shader touching every lowered category: builtins, vote, arithmetic
// scans, ballot math, shuffles, clustered and quad operations.
constexpr const char* kEveryCategorySource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_vote : require
#extension GL_KHR_shader_subgroup_arithmetic : require
#extension GL_KHR_shader_subgroup_ballot : require
#extension GL_KHR_shader_subgroup_shuffle : require
#extension GL_KHR_shader_subgroup_shuffle_relative : require
#extension GL_KHR_shader_subgroup_clustered : require
#extension GL_KHR_shader_subgroup_quad : require
layout(local_size_x = 48, local_size_y = 1, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { float value[]; } outputData;
void main() {
uint slot = gl_LocalInvocationIndex * 24u;
float v = float(gl_LocalInvocationIndex + 1u);
outputData.value[slot + 0u] = float(gl_SubgroupSize);
outputData.value[slot + 1u] = float(gl_NumSubgroups);
outputData.value[slot + 2u] = float(gl_SubgroupID);
outputData.value[slot + 3u] = float(gl_SubgroupInvocationID);
outputData.value[slot + 4u] = float(gl_SubgroupEqMask.x + gl_SubgroupLtMask.x);
outputData.value[slot + 5u] = subgroupElect() ? 1.0 : 0.0;
outputData.value[slot + 6u] = subgroupAll(v > 0.0) ? 1.0 : 0.0;
outputData.value[slot + 7u] = subgroupAny(v > 40.0) ? 1.0 : 0.0;
outputData.value[slot + 8u] = subgroupAllEqual(gl_WorkGroupID.x) ? 1.0 : 0.0;
outputData.value[slot + 9u] = subgroupAdd(v);
outputData.value[slot + 10u] = subgroupInclusiveAdd(v);
outputData.value[slot + 11u] = subgroupExclusiveMax(v);
outputData.value[slot + 12u] = float(subgroupMin(gl_LocalInvocationIndex));
uvec4 ballot = subgroupBallot((gl_LocalInvocationIndex & 1u) == 0u);
outputData.value[slot + 13u] = float(subgroupBallotBitCount(ballot));
outputData.value[slot + 14u] = float(subgroupBallotFindLSB(ballot));
outputData.value[slot + 15u] = float(subgroupBallotFindMSB(ballot));
outputData.value[slot + 16u] = subgroupInverseBallot(ballot) ? 1.0 : 0.0;
outputData.value[slot + 17u] = subgroupBallotBitExtract(ballot, 3u) ? 1.0 : 0.0;
outputData.value[slot + 18u] = subgroupBroadcast(v, 2u);
outputData.value[slot + 19u] = subgroupBroadcastFirst(v);
outputData.value[slot + 20u] = subgroupShuffle(v, gl_SubgroupInvocationID ^ 5u);
outputData.value[slot + 21u] = subgroupShuffleXor(v, 1u) + subgroupShuffleUp(v, 1u) +
subgroupShuffleDown(v, 1u);
outputData.value[slot + 22u] = subgroupClusteredAdd(v, 4u);
outputData.value[slot + 23u] = subgroupQuadBroadcast(v, 1u) + subgroupQuadSwapHorizontal(v);
subgroupBarrier();
subgroupMemoryBarrierShared();
}
)";
constexpr const char* kNoSubgroupSource = R"(#version 450 core
layout(local_size_x = 64) in;
layout(std430, binding = 0) buffer Output { uint value; } outputData;
void main() {
if (gl_LocalInvocationIndex == 0u) outputData.value = gl_WorkGroupSize.x;
}
)";
// An extended subgroup instruction (SPV_KHR_subgroup_rotate) alongside core
// ones: outside the lowered set, so the pass must fail rather than emit
// "subgroup-free" output that still rotates.
constexpr const char* kRotateSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
#extension GL_KHR_shader_subgroup_rotate : require
layout(local_size_x = 64) in;
layout(std430, binding = 0) buffer Output { float value[]; } outputData;
void main() {
float v = subgroupAdd(float(gl_SubgroupInvocationID));
outputData.value[gl_LocalInvocationIndex] = subgroupRotate(v, 1u);
}
)";
// A 1024-invocation workgroup exchanging a vec4 and a float: the lowering
// would need 16 KiB + 4 KiB of scratch, past the Vulkan-minimum shared
// budget of 16384 bytes.
constexpr const char* kScratchHungrySource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 1024) in;
layout(std430, binding = 0) buffer Output { vec4 value[]; } outputData;
void main() {
vec4 wide = subgroupAdd(vec4(float(gl_LocalInvocationIndex)));
wide.x += subgroupInclusiveAdd(float(gl_SubgroupInvocationID));
outputData.value[gl_LocalInvocationIndex] = wide;
}
)";
} // namespace
TEST(EmulateSubgroupsPass, LowersEveryCategoryToSharedMemory) {
const Vector<Uint32> input = CompileStage(GL_COMPUTE_SHADER, kEveryCategorySource);
ASSERT_FALSE(input.empty());
ASSERT_GT(CountGroupNonUniform(input), 0u);
ASSERT_GT(CountGroupNonUniformCapabilities(input), 0u);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::EmulateSubgroupsForVulkan(input, output, 16384u, true));
ASSERT_TRUE(Validates(output));
// The whole point: nothing subgroup-shaped survives, so the module runs on a
// device with no subgroup support at all.
EXPECT_EQ(CountGroupNonUniform(output), 0u);
EXPECT_EQ(CountGroupNonUniformCapabilities(output), 0u);
// The exchanges go through workgroup-shared scratch behind control barriers.
EXPECT_TRUE(HasWorkgroupVariable(output));
EXPECT_GT(CountOpcode(output, spv::Op::OpControlBarrier), CountOpcode(input, spv::Op::OpControlBarrier));
}
TEST(EmulateSubgroupsPass, IsIdempotent) {
const Vector<Uint32> input = CompileStage(GL_COMPUTE_SHADER, kEveryCategorySource);
ASSERT_FALSE(input.empty());
Vector<Uint32> once;
ASSERT_TRUE(ShaderCompiler::EmulateSubgroupsForVulkan(input, once, 16384u, true));
Vector<Uint32> twice;
ASSERT_TRUE(ShaderCompiler::EmulateSubgroupsForVulkan(once, twice, 16384u, true));
EXPECT_EQ(twice, once);
}
TEST(EmulateSubgroupsPass, LeavesSubgroupFreeComputeUntouched) {
const Vector<Uint32> input = CompileStage(GL_COMPUTE_SHADER, kNoSubgroupSource);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::EmulateSubgroupsForVulkan(input, output, 16384u, true));
EXPECT_EQ(output, input);
}
TEST(EmulateSubgroupsPass, RefusesExtendedSubgroupInstructions) {
const Vector<Uint32> input = CompileStage(GL_COMPUTE_SHADER, kRotateSource);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::EmulateSubgroupsForVulkan(input, output, 16384u, false));
}
TEST(EmulateSubgroupsPass, RefusesAModuleOverTheScratchBudget) {
const Vector<Uint32> input = CompileStage(GL_COMPUTE_SHADER, kScratchHungrySource);
ASSERT_FALSE(input.empty());
// vec4 scratch (1024 slots * 16 bytes) plus float scratch (4 KiB) exceeds
// the 16 KiB Vulkan-minimum budget.
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::EmulateSubgroupsForVulkan(input, output, 16384u, false));
// A device advertising more shared memory takes the same module fine.
Vector<Uint32> roomier;
EXPECT_TRUE(ShaderCompiler::EmulateSubgroupsForVulkan(input, roomier, 32768u, true));
EXPECT_TRUE(Validates(roomier));
}
@@ -0,0 +1,336 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/FixIterationRPSubgroupScratchTest.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#define SPV_ENABLE_UTILITY_CODE
#include "glslang/SPIRV/spirv.hpp11"
#undef SPV_ENABLE_UTILITY_CODE
#include "Includes.h"
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <MG_Util/ShaderTranspiler/Types.h>
#include <spirv-tools/libspirv.hpp>
#include <algorithm>
#include <map>
#include <vector>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
namespace {
constexpr SizeT kSpirvHeaderWordCount = 5u;
template <typename Visitor>
void ForEachInstruction(const Vector<Uint32>& spirv, Visitor&& visit) {
for (SizeT offset = kSpirvHeaderWordCount; offset < spirv.size();) {
const Uint32 wordCount = spirv[offset] >> 16u;
if (wordCount == 0u || offset + wordCount > spirv.size()) break;
visit(static_cast<spv::Op>(spirv[offset] & 0xffffu), &spirv[offset], wordCount);
offset += wordCount;
}
}
Vector<Uint32> CompileCompute(const String& source) {
using namespace MobileGL::MG_Util::ShaderTranspiler;
ShaderAttrib shaderAttrib{.shaderType = GL_COMPUTE_SHADER, .sourceStr = source};
auto shaderResult = ShaderCompiler::CompileShader(shaderAttrib);
EXPECT_TRUE(shaderResult) << (shaderResult ? String{} : shaderResult.error().log);
if (!shaderResult) return {};
ProgramAttrib programAttrib{.shaders = {shaderResult.value()}};
auto programResult = ShaderCompiler::LinkProgram(programAttrib);
EXPECT_TRUE(programResult) << (programResult ? String{} : programResult.error().log);
if (!programResult) return {};
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {GL_COMPUTE_SHADER}, .program = *programResult.value()};
auto binaryResult = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
EXPECT_TRUE(binaryResult) << (binaryResult ? String{} : binaryResult.error().log);
if (!binaryResult || binaryResult->empty()) return {};
return binaryResult->front();
}
// The declared lengths of every Workgroup-storage array variable, sorted.
std::vector<Uint32> WorkgroupArrayLengths(const Vector<Uint32>& spirv) {
std::map<Uint32, Uint32> constantValues; // constant id -> value
std::map<Uint32, Uint32> arrayLengthIds; // array type id -> length constant id
std::map<Uint32, Uint32> pointerPointees; // pointer type id -> pointee type id
std::vector<Uint32> workgroupPointerTypes; // type ids of Workgroup variables
ForEachInstruction(spirv, [&](spv::Op opcode, const Uint32* words, Uint32 wordCount) {
switch (opcode) {
case spv::Op::OpConstant:
if (wordCount >= 4u) constantValues[words[2]] = words[3];
break;
case spv::Op::OpTypeArray:
if (wordCount >= 4u) arrayLengthIds[words[1]] = words[3];
break;
case spv::Op::OpTypePointer:
if (wordCount >= 4u &&
static_cast<spv::StorageClass>(words[2]) == spv::StorageClass::Workgroup) {
pointerPointees[words[1]] = words[3];
}
break;
case spv::Op::OpVariable:
if (wordCount >= 4u &&
static_cast<spv::StorageClass>(words[3]) == spv::StorageClass::Workgroup) {
workgroupPointerTypes.push_back(words[1]);
}
break;
default:
break;
}
});
std::vector<Uint32> lengths;
for (const Uint32 pointerTypeId : workgroupPointerTypes) {
const auto pointee = pointerPointees.find(pointerTypeId);
if (pointee == pointerPointees.end()) continue;
const auto lengthId = arrayLengthIds.find(pointee->second);
if (lengthId == arrayLengthIds.end()) continue;
const auto value = constantValues.find(lengthId->second);
if (value != constantValues.end()) lengths.push_back(value->second);
}
std::sort(lengths.begin(), lengths.end());
return lengths;
}
bool Validates(const Vector<Uint32>& spirv) {
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
tools.SetMessageConsumer([](spv_message_level_t, const char*, const spv_position_t& position,
const char* message) {
ADD_FAILURE() << "spirv-val at word " << position.index << ": " << message;
});
return tools.Validate(spirv);
}
// iterationRP's exposure reduction, as the pack ships it: 32x16 (512
// invocations), subgroupInclusiveAdd on a vec2, and a 32-entry
// gl_SubgroupID-indexed scratch. A second, plainly indexed array rides along
// to prove the patch is surgical.
constexpr const char* kExposureShapedSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared vec2 prefixSumCache[32];
shared float plainScratch[4];
void main() {
vec2 sampleLuminance = vec2(float(gl_LocalInvocationIndex), 0.0);
sampleLuminance = subgroupInclusiveAdd(sampleLuminance);
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = sampleLuminance;
plainScratch[gl_LocalInvocationIndex & 3u] = sampleLuminance.x;
barrier();
uint loopLength = uint(findMSB(gl_NumSubgroups));
loopLength += uint(gl_NumSubgroups - (1u << (loopLength - 1u)) > 0u);
for (uint scanStage = 0u; scanStage < loopLength; ++scanStage) {
if ((gl_SubgroupID & (1u << scanStage)) > 0u) {
sampleLuminance += prefixSumCache[(gl_SubgroupID >> scanStage << scanStage) - 1u];
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = sampleLuminance;
}
barrier();
}
if (gl_LocalInvocationIndex == 511u)
outputData.value = prefixSumCache[0].x / 512.0 + plainScratch[0];
}
)";
// The pack's OTHER instance of the same bug, which a fingerprint pinned to the
// exposure pass's dimensions walks straight past: the RTW importance warp
// scans a plain float across 1024 invocations into a 64-entry scratch.
constexpr const char* kRtwWarpShapedSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 1024) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared float prefixSumCache[64];
void main() {
float importance = float(gl_LocalInvocationID.x) * 0.5;
float prefixSum = subgroupInclusiveAdd(importance);
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = prefixSum;
barrier();
uint loopLength = uint(findMSB(gl_NumSubgroups));
loopLength += uint(gl_NumSubgroups - (1u << (loopLength - 1u)) > 0u);
for (uint scanStage = 0u; scanStage < loopLength; ++scanStage) {
if ((gl_SubgroupID & (1u << scanStage)) > 0u) {
prefixSum += prefixSumCache[(gl_SubgroupID >> scanStage << scanStage) - 1u];
if (gl_SubgroupInvocationID == gl_SubgroupSize - 1u)
prefixSumCache[gl_SubgroupID] = prefixSum;
}
barrier();
}
if (gl_LocalInvocationID.x == 1023u) outputData.value = prefixSumCache[0];
}
)";
// A subgroup scan, but the scratch is indexed per invocation rather than per
// subgroup: its size is not a subgroup-count assumption, so it is not ours.
constexpr const char* kInvocationIndexedSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared vec2 perInvocation[32];
void main() {
vec2 v = subgroupInclusiveAdd(vec2(float(gl_LocalInvocationIndex), 0.0));
perInvocation[gl_LocalInvocationIndex & 31u] = v;
barrier();
if (gl_LocalInvocationIndex == 0u) outputData.value = perInvocation[0].x;
}
)";
// gl_SubgroupID-indexed, but no subgroup scan feeds it and the element type is
// not the pack's float accumulator.
constexpr const char* kNonFloatScratchSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { uint value; } outputData;
shared uint tally[32];
void main() {
float scan = subgroupInclusiveAdd(float(gl_LocalInvocationIndex));
tally[gl_SubgroupID] = uint(scan);
barrier();
if (gl_LocalInvocationIndex == 0u) outputData.value = tally[0];
}
)";
// gl_SubgroupID-indexed, but masked into range: the declaration is bounded by
// construction, not a subgroup-count assumption, so it is not the pack's bug.
constexpr const char* kMaskedSubgroupIndexSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared vec2 bounded[8];
void main() {
vec2 v = subgroupInclusiveAdd(vec2(float(gl_LocalInvocationIndex), 0.0));
bounded[gl_SubgroupID & 7u] = v;
barrier();
if (gl_LocalInvocationIndex == 0u) outputData.value = bounded[0].x;
}
)";
// Neither of the pack's shapes: a small per-subgroup array in a 256-invocation
// workgroup, used to prove the width gate keeps EVERY module inert at >= 16 lanes.
constexpr const char* kForeignShapeSource = R"(#version 450 core
#extension GL_KHR_shader_subgroup_basic : require
#extension GL_KHR_shader_subgroup_arithmetic : require
layout(local_size_x = 256) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared float partial[4];
void main() {
float v = subgroupInclusiveAdd(float(gl_LocalInvocationID.x));
if (gl_SubgroupID < 4u) partial[gl_SubgroupID] = v;
barrier();
if (gl_LocalInvocationID.x == 0u) outputData.value = partial[0];
}
)";
// No subgroup construct at all.
constexpr const char* kSubgroupFreeSource = R"(#version 450 core
layout(local_size_x = 32, local_size_y = 16, local_size_z = 1) in;
layout(std430, binding = 0) buffer Output { float value; } outputData;
shared vec2 scratch[32];
void main() {
scratch[gl_LocalInvocationIndex & 31u] = vec2(float(gl_LocalInvocationIndex), 0.0);
barrier();
if (gl_LocalInvocationIndex == 0u) outputData.value = scratch[0].x;
}
)";
} // namespace
TEST(FixIterationRPSubgroupScratchPass, GrowsTheExposureScratchForNarrowSubgroups) {
const Vector<Uint32> input = CompileCompute(kExposureShapedSource);
ASSERT_FALSE(input.empty());
ASSERT_EQ(WorkgroupArrayLengths(input), (std::vector<Uint32>{4u, 32u}));
// lavapipe: 8-lane subgroups over 512 invocations need 64 entries; the
// plainly indexed neighbour must keep its 4.
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(input, output, 8u, 32768u, true));
EXPECT_EQ(WorkgroupArrayLengths(output), (std::vector<Uint32>{4u, 64u}));
EXPECT_TRUE(Validates(output));
}
// The regression the CI retrace caught: patching only the exposure pass leaves
// this one writing 128 subgroups into 64 entries, and the frame stays wrong.
TEST(FixIterationRPSubgroupScratchPass, GrowsTheRtwWarpScratchForNarrowSubgroups) {
const Vector<Uint32> input = CompileCompute(kRtwWarpShapedSource);
ASSERT_FALSE(input.empty());
ASSERT_EQ(WorkgroupArrayLengths(input), (std::vector<Uint32>{64u}));
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(input, output, 8u, 32768u, true));
EXPECT_EQ(WorkgroupArrayLengths(output), (std::vector<Uint32>{128u}));
EXPECT_TRUE(Validates(output));
}
TEST(FixIterationRPSubgroupScratchPass, LeavesPackWidthAssumptionsAloneOnWideDevices) {
// Both shapes are sized for >= 16 lanes (512/16 = 32, 1024/16 = 64), so on
// every such device the modules must pass through byte-identical.
for (const char* source : {kExposureShapedSource, kRtwWarpShapedSource}) {
const Vector<Uint32> input = CompileCompute(source);
ASSERT_FALSE(input.empty());
for (const Uint32 nativeSize : {16u, 32u, 64u, 128u}) {
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(
input, output, nativeSize, 32768u, true));
EXPECT_EQ(output, input) << "native width " << nativeSize;
}
}
}
TEST(FixIterationRPSubgroupScratchPass, RefusesAModuleOutsideTheIdiom) {
for (const char* source : {kInvocationIndexedSource, kNonFloatScratchSource,
kSubgroupFreeSource, kMaskedSubgroupIndexSource}) {
const Vector<Uint32> input = CompileCompute(source);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(input, output, 8u, 32768u, true));
EXPECT_EQ(output, input);
}
}
// A grown array that would not fit the device's shared memory is left alone:
// a pipeline that cannot be created is worse than the pack's own overrun.
// The width gate is what keeps unrelated shaders untouched on the devices the pack
// was written for: at >= 16 lanes nothing is rewritten, whatever its shape.
TEST(FixIterationRPSubgroupScratchPass, LeavesEveryModuleAloneAtThePacksAssumedWidth) {
const Vector<Uint32> input = CompileCompute(kForeignShapeSource);
ASSERT_FALSE(input.empty());
for (const Uint32 nativeSize : {16u, 32u, 64u}) {
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(
input, output, nativeSize, 32768u, true));
EXPECT_EQ(output, input) << "native width " << nativeSize;
}
}
TEST(FixIterationRPSubgroupScratchPass, RefusesGrowthThatWouldNotFitSharedMemory) {
const Vector<Uint32> input = CompileCompute(kRtwWarpShapedSource);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(input, output, 8u, 256u, true));
EXPECT_EQ(output, input);
}
TEST(FixIterationRPSubgroupScratchPass, IsIdempotent) {
for (const char* source : {kExposureShapedSource, kRtwWarpShapedSource}) {
const Vector<Uint32> input = CompileCompute(source);
ASSERT_FALSE(input.empty());
Vector<Uint32> once;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(input, once, 8u, 32768u, true));
Vector<Uint32> twice;
ASSERT_TRUE(ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(once, twice, 8u, 32768u, true));
EXPECT_EQ(twice, once);
}
}
+18 -18
View File
@@ -7,8 +7,8 @@
// End of Source File Header
#include "DriverPost.h"
#include "DriverPostProgram203Witness.h"
#include "DriverPostProgram203WitnessSpv.h"
#include "DriverPostIterationRPWitness.h"
#include "DriverPostIterationRPWitnessSpv.h"
#include "MG_Util/BackendLoaders/OpenGL/Loader.h"
#include <Config.h>
#include <MGGitHash.h>
@@ -1458,11 +1458,11 @@ namespace MobileGL::MG_Util::SelfTest {
disabledNote);
}
// Native Program-203 compute witness. This deliberately uses a separate
// Native iterationRP compute witness. This deliberately uses a separate
// throwaway Vulkan device rather than the real renderer's queues, and it
// treats MOBILEGL_DISABLE_SUBGROUP as irrelevant: the row reports what the
// driver does, not what MobileGL elects to advertise to applications.
void ProbeVulkanProgram203Witness(ReportBuilder& builder, PFN_vkGetInstanceProcAddr getInstanceProcAddr,
void ProbeVulkanIterationRPWitness(ReportBuilder& builder, PFN_vkGetInstanceProcAddr getInstanceProcAddr,
VkInstance instance, VkPhysicalDevice physicalDevice,
Uint32 computeQueueFamilyIndex,
const VkPhysicalDeviceProperties& properties,
@@ -1476,7 +1476,7 @@ namespace MobileGL::MG_Util::SelfTest {
return;
}
Program203WitnessLimits limits{};
IterationRPWitnessLimits limits{};
limits.computeStageSupported =
(subgroupProperties.supportedStages & VK_SHADER_STAGE_COMPUTE_BIT) != 0;
limits.basicSubgroupSupported =
@@ -1494,12 +1494,12 @@ namespace MobileGL::MG_Util::SelfTest {
limits.maxBoundDescriptorSets = properties.limits.maxBoundDescriptorSets;
limits.maxStorageBufferRange = properties.limits.maxStorageBufferRange;
const Program203WitnessEligibilityResult eligibility = EvaluateProgram203WitnessEligibility(limits);
if (eligibility.eligibility == Program203WitnessEligibility::SkipUnsupportedNativeFeatureSet) {
const IterationRPWitnessEligibilityResult eligibility = EvaluateIterationRPWitnessEligibility(limits);
if (eligibility.eligibility == IterationRPWitnessEligibility::SkipUnsupportedNativeFeatureSet) {
builder.Info(RowName, eligibility.detail);
return;
}
if (eligibility.eligibility == Program203WitnessEligibility::FailInadequateLimits) {
if (eligibility.eligibility == IterationRPWitnessEligibility::FailInadequateLimits) {
fail(eligibility.detail);
return;
}
@@ -1668,7 +1668,7 @@ namespace MobileGL::MG_Util::SelfTest {
VkBufferCreateInfo bufferInfo{};
bufferInfo.sType = VK_STRUCTURE_TYPE_BUFFER_CREATE_INFO;
bufferInfo.size = sizeof(Program203WitnessOutput);
bufferInfo.size = sizeof(IterationRPWitnessOutput);
bufferInfo.usage = VK_BUFFER_USAGE_STORAGE_BUFFER_BIT;
bufferInfo.sharingMode = VK_SHARING_MODE_EXCLUSIVE;
result = vkCreateBufferFn(device, &bufferInfo, nullptr, &outputBuffer);
@@ -1710,12 +1710,12 @@ namespace MobileGL::MG_Util::SelfTest {
fail(format("vkBindBufferMemory(output SSBO) failed (VkResult = {})", static_cast<Int>(result)));
return;
}
result = vkMapMemoryFn(device, outputMemory, 0, sizeof(Program203WitnessOutput), 0, &mappedOutput);
result = vkMapMemoryFn(device, outputMemory, 0, sizeof(IterationRPWitnessOutput), 0, &mappedOutput);
if (result != VK_SUCCESS || mappedOutput == nullptr) {
fail(format("vkMapMemory(output SSBO) failed (VkResult = {})", static_cast<Int>(result)));
return;
}
std::memset(mappedOutput, 0xa5, sizeof(Program203WitnessOutput));
std::memset(mappedOutput, 0xa5, sizeof(IterationRPWitnessOutput));
VkDescriptorSetLayoutBinding outputBinding{};
outputBinding.binding = 0;
@@ -1760,7 +1760,7 @@ namespace MobileGL::MG_Util::SelfTest {
VkDescriptorBufferInfo outputDescriptor{};
outputDescriptor.buffer = outputBuffer;
outputDescriptor.offset = 0;
outputDescriptor.range = sizeof(Program203WitnessOutput);
outputDescriptor.range = sizeof(IterationRPWitnessOutput);
VkWriteDescriptorSet descriptorWrite{};
descriptorWrite.sType = VK_STRUCTURE_TYPE_WRITE_DESCRIPTOR_SET;
descriptorWrite.dstSet = descriptorSet;
@@ -1772,8 +1772,8 @@ namespace MobileGL::MG_Util::SelfTest {
VkShaderModuleCreateInfo shaderModuleInfo{};
shaderModuleInfo.sType = VK_STRUCTURE_TYPE_SHADER_MODULE_CREATE_INFO;
shaderModuleInfo.codeSize = sizeof(kDriverPostProgram203WitnessSpv);
shaderModuleInfo.pCode = kDriverPostProgram203WitnessSpv;
shaderModuleInfo.codeSize = sizeof(kDriverPostIterationRPWitnessSpv);
shaderModuleInfo.pCode = kDriverPostIterationRPWitnessSpv;
result = vkCreateShaderModuleFn(device, &shaderModuleInfo, nullptr, &shaderModule);
if (result != VK_SUCCESS) {
fail(format("vkCreateShaderModule failed (VkResult = {})", static_cast<Int>(result)));
@@ -1845,7 +1845,7 @@ namespace MobileGL::MG_Util::SelfTest {
hostReadBarrier.dstQueueFamilyIndex = VK_QUEUE_FAMILY_IGNORED;
hostReadBarrier.buffer = outputBuffer;
hostReadBarrier.offset = 0;
hostReadBarrier.size = sizeof(Program203WitnessOutput);
hostReadBarrier.size = sizeof(IterationRPWitnessOutput);
vkCmdPipelineBarrierFn(commandBuffer, VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT, VK_PIPELINE_STAGE_HOST_BIT, 0,
0, nullptr, 1, &hostReadBarrier, 0, nullptr);
result = vkEndCommandBufferFn(commandBuffer);
@@ -1879,9 +1879,9 @@ namespace MobileGL::MG_Util::SelfTest {
return;
}
Program203WitnessOutput output{};
IterationRPWitnessOutput output{};
std::memcpy(&output, mappedOutput, sizeof(output));
const Program203WitnessValidationResult validation = ValidateProgram203Witness(output);
const IterationRPWitnessValidationResult validation = ValidateIterationRPWitness(output);
if (!validation.ok) {
fail(validation.detail);
return;
@@ -2491,7 +2491,7 @@ namespace MobileGL::MG_Util::SelfTest {
builder.Warn("Compute shader subgroup", "subgroup properties could not be queried");
}
ProbeVulkanProgram203Witness(builder, getInstanceProcAddr, instance, physicalDevice, computeQueueFamilyIndex,
ProbeVulkanIterationRPWitness(builder, getInstanceProcAddr, instance, physicalDevice, computeQueueFamilyIndex,
properties, subgroupPropertiesAvailable, subgroupProperties);
if (HasVkExtension(deviceExtensions, VK_KHR_DRAW_INDIRECT_COUNT_EXTENSION_NAME)) {
@@ -1,4 +1,4 @@
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostProgram203Witness.comp
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostIterationRPWitness.comp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
@@ -6,9 +6,9 @@
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Native Vulkan GLSL 450 witness for Program 203's first subgroup reduction.
// Native Vulkan GLSL 450 witness for iterationRP's first subgroup reduction.
// It is intentionally independent of the GL 430 integration scenario. The body
// below preserves Program 203's source reduction; the surrounding diagnostics
// below preserves iterationRP's source reduction; the surrounding diagnostics
// only observe its topology and cache handoffs.
#version 450
@@ -23,7 +23,7 @@ const uint kTopologyInvalidSubgroupId = 1u << 2u;
const uint kTopologyInvalidSubgroupLane = 1u << 3u;
const uint kWitnessMagic = 0x50323033u;
layout(std430, set = 0, binding = 0) buffer Program203WitnessOutput {
layout(std430, set = 0, binding = 0) buffer IterationRPWitnessOutput {
uint magic;
uint topologyFlags;
uint numSubgroups;
@@ -40,7 +40,7 @@ layout(std430, set = 0, binding = 0) buffer Program203WitnessOutput {
vec2 finalAverage;
} outWitness;
// Program 203's cache stays separate from all diagnostic shared state. In
// iterationRP's cache stays separate from all diagnostic shared state. In
// particular, no instrumentation stores through prefixSumCache except source
// writes retained below.
shared vec2 prefixSumCache[32];
@@ -108,7 +108,7 @@ void main() {
}
// This branch is uniform after collection and is solely a safety guard for
// broken topology reports. The valid side retains Program 203 verbatim.
// broken topology reports. The valid side retains iterationRP verbatim.
const bool sourceDomain = canonicalDomain && topologyFlagsShared == 0u;
if (sourceDomain) {
vec2 sampleLuminance = vec2(float(gl_LocalInvocationIndex + 1u), 0.0);
@@ -1,4 +1,4 @@
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostProgram203Witness.cpp
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostIterationRPWitness.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
@@ -6,7 +6,7 @@
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "DriverPostProgram203Witness.h"
#include "DriverPostIterationRPWitness.h"
#include <bit>
#include <sstream>
@@ -15,11 +15,11 @@
namespace MobileGL::MG_Util::SelfTest {
namespace {
[[nodiscard]] Program203WitnessValidationResult Failure(Program203WitnessValidationFailure failure,
[[nodiscard]] IterationRPWitnessValidationResult Failure(IterationRPWitnessValidationFailure failure,
std::string detail,
std::uint32_t scanStage = 0u,
std::uint32_t subgroup = 0u) {
Program203WitnessValidationResult result;
IterationRPWitnessValidationResult result;
result.ok = false;
result.failure = failure;
result.scanStage = scanStage;
@@ -36,18 +36,18 @@ namespace MobileGL::MG_Util::SelfTest {
return FloatBits(lhs) == FloatBits(rhs);
}
[[nodiscard]] bool SameBits(const Program203WitnessVec2& lhs, const Program203WitnessVec2& rhs) {
[[nodiscard]] bool SameBits(const IterationRPWitnessVec2& lhs, const IterationRPWitnessVec2& rhs) {
return SameBits(lhs.x, rhs.x) && SameBits(lhs.y, rhs.y);
}
[[nodiscard]] std::string Vec2String(const Program203WitnessVec2& value) {
[[nodiscard]] std::string Vec2String(const IterationRPWitnessVec2& value) {
std::ostringstream output;
output << '(' << value.x << ',' << value.y << ')';
return output.str();
}
[[nodiscard]] std::uint32_t ExpectedSeenSubgroupMask(std::uint32_t numSubgroups) {
return numSubgroups == kProgram203WitnessMaxSubgroups ? 0xffffffffu : (1u << numSubgroups) - 1u;
return numSubgroups == kIterationRPWitnessMaxSubgroups ? 0xffffffffu : (1u << numSubgroups) - 1u;
}
[[nodiscard]] std::string JoinRequirements(const std::vector<std::string>& requirements) {
@@ -60,8 +60,8 @@ namespace MobileGL::MG_Util::SelfTest {
}
} // namespace
Program203WitnessEligibilityResult
EvaluateProgram203WitnessEligibility(const Program203WitnessLimits& limits) {
IterationRPWitnessEligibilityResult
EvaluateIterationRPWitnessEligibility(const IterationRPWitnessLimits& limits) {
// This classification deliberately precedes numeric limits. An absent native
// compute/basic/arithmetic subgroup contract means there is nothing to witness,
// whereas every resource/entry-point failure on a capable device is a POST FAIL.
@@ -70,7 +70,7 @@ namespace MobileGL::MG_Util::SelfTest {
if (!limits.computeStageSupported) missing.emplace_back("VK_SHADER_STAGE_COMPUTE_BIT");
if (!limits.basicSubgroupSupported) missing.emplace_back("VK_SUBGROUP_FEATURE_BASIC_BIT");
if (!limits.arithmeticSubgroupSupported) missing.emplace_back("VK_SUBGROUP_FEATURE_ARITHMETIC_BIT");
return {Program203WitnessEligibility::SkipUnsupportedNativeFeatureSet,
return {IterationRPWitnessEligibility::SkipUnsupportedNativeFeatureSet,
"skipped because the native compute/basic/arithmetic subgroup feature set is unsupported (missing " +
JoinRequirements(missing) + ')'};
}
@@ -79,16 +79,16 @@ namespace MobileGL::MG_Util::SelfTest {
if (limits.subgroupSize == 0u) {
inadequate.emplace_back("subgroupSize == 0");
}
if (limits.maxComputeWorkGroupInvocations < kProgram203WitnessInvocationCount) {
if (limits.maxComputeWorkGroupInvocations < kIterationRPWitnessInvocationCount) {
inadequate.emplace_back("maxComputeWorkGroupInvocations < 512");
}
if (limits.maxComputeWorkGroupSize[0] < 32u || limits.maxComputeWorkGroupSize[1] < 16u ||
limits.maxComputeWorkGroupSize[2] < 1u) {
inadequate.emplace_back("maxComputeWorkGroupSize does not cover 32x16x1");
}
if (limits.maxComputeSharedMemorySize < kProgram203WitnessSharedMemoryBytes) {
if (limits.maxComputeSharedMemorySize < kIterationRPWitnessSharedMemoryBytes) {
inadequate.emplace_back("maxComputeSharedMemorySize < " +
std::to_string(kProgram203WitnessSharedMemoryBytes));
std::to_string(kIterationRPWitnessSharedMemoryBytes));
}
if (limits.maxPerStageDescriptorStorageBuffers < 1u) {
inadequate.emplace_back("maxPerStageDescriptorStorageBuffers < 1");
@@ -99,21 +99,21 @@ namespace MobileGL::MG_Util::SelfTest {
if (limits.maxBoundDescriptorSets < 1u) {
inadequate.emplace_back("maxBoundDescriptorSets < 1");
}
if (limits.maxStorageBufferRange < sizeof(Program203WitnessOutput)) {
if (limits.maxStorageBufferRange < sizeof(IterationRPWitnessOutput)) {
inadequate.emplace_back("maxStorageBufferRange < " +
std::to_string(sizeof(Program203WitnessOutput)));
std::to_string(sizeof(IterationRPWitnessOutput)));
}
if (!inadequate.empty()) {
return {Program203WitnessEligibility::FailInadequateLimits,
return {IterationRPWitnessEligibility::FailInadequateLimits,
"insufficient Vulkan limits for a 32x16x1 workgroup, one output SSBO, and " +
std::to_string(kProgram203WitnessSharedMemoryBytes) + " bytes of shared memory: " +
std::to_string(kIterationRPWitnessSharedMemoryBytes) + " bytes of shared memory: " +
JoinRequirements(inadequate)};
}
return {Program203WitnessEligibility::Execute, {}};
return {IterationRPWitnessEligibility::Execute, {}};
}
std::uint32_t ComputeProgram203WitnessLoopLength(std::uint32_t numSubgroups) {
if (numSubgroups < 2u || numSubgroups > kProgram203WitnessMaxSubgroups) return 0u;
std::uint32_t ComputeIterationRPWitnessLoopLength(std::uint32_t numSubgroups) {
if (numSubgroups < 2u || numSubgroups > kIterationRPWitnessMaxSubgroups) return 0u;
// Exact C++ spelling of the source's findMSB-based calculation. In
// particular, its final iteration for powers of two is intentional.
@@ -125,87 +125,87 @@ namespace MobileGL::MG_Util::SelfTest {
return loopLength;
}
Program203WitnessValidationResult ValidateProgram203Witness(const Program203WitnessOutput& output) {
IterationRPWitnessValidationResult ValidateIterationRPWitness(const IterationRPWitnessOutput& output) {
// 1. Completion. A poisoned or unwritten result must never turn into a
// topology diagnosis, because it says nothing about execution.
if (output.magic != kProgram203WitnessMagic) {
if (output.magic != kIterationRPWitnessMagic) {
std::ostringstream detail;
detail << "completion: magic was 0x" << std::hex << output.magic << ", expected 0x"
<< kProgram203WitnessMagic;
return Failure(Program203WitnessValidationFailure::Completion, detail.str());
<< kIterationRPWitnessMagic;
return Failure(IterationRPWitnessValidationFailure::Completion, detail.str());
}
// 2. Observed topology. All checks consume observations written by the
// shader, rather than inferring subgroup layout from invocation indices.
const std::uint32_t numSubgroups = output.numSubgroups;
if (numSubgroups < 2u || numSubgroups > kProgram203WitnessMaxSubgroups) {
if (numSubgroups < 2u || numSubgroups > kIterationRPWitnessMaxSubgroups) {
std::ostringstream detail;
detail << "topology: canonical gl_NumSubgroups=" << numSubgroups << " is outside [2, 32]";
return Failure(Program203WitnessValidationFailure::Topology, detail.str());
return Failure(IterationRPWitnessValidationFailure::Topology, detail.str());
}
if ((output.topologyFlags & Program203WitnessNonuniformNumSubgroups) != 0u) {
return Failure(Program203WitnessValidationFailure::Topology,
if ((output.topologyFlags & IterationRPWitnessNonuniformNumSubgroups) != 0u) {
return Failure(IterationRPWitnessValidationFailure::Topology,
"topology: gl_NumSubgroups differed across workgroup");
}
if ((output.topologyFlags & Program203WitnessInvalidNumSubgroups) != 0u) {
return Failure(Program203WitnessValidationFailure::Topology,
if ((output.topologyFlags & IterationRPWitnessInvalidNumSubgroups) != 0u) {
return Failure(IterationRPWitnessValidationFailure::Topology,
"topology: an invocation reported gl_NumSubgroups outside [2, 32]");
}
if ((output.topologyFlags & Program203WitnessInvalidSubgroupId) != 0u) {
return Failure(Program203WitnessValidationFailure::Topology,
if ((output.topologyFlags & IterationRPWitnessInvalidSubgroupId) != 0u) {
return Failure(IterationRPWitnessValidationFailure::Topology,
"topology: an invocation reported an invalid gl_SubgroupID");
}
if ((output.topologyFlags & Program203WitnessInvalidSubgroupLane) != 0u) {
return Failure(Program203WitnessValidationFailure::Topology,
if ((output.topologyFlags & IterationRPWitnessInvalidSubgroupLane) != 0u) {
return Failure(IterationRPWitnessValidationFailure::Topology,
"topology: an invocation reported an invalid subgroup lane");
}
if ((output.topologyFlags & ~(Program203WitnessNonuniformNumSubgroups |
Program203WitnessInvalidNumSubgroups |
Program203WitnessInvalidSubgroupId |
Program203WitnessInvalidSubgroupLane)) != 0u) {
if ((output.topologyFlags & ~(IterationRPWitnessNonuniformNumSubgroups |
IterationRPWitnessInvalidNumSubgroups |
IterationRPWitnessInvalidSubgroupId |
IterationRPWitnessInvalidSubgroupLane)) != 0u) {
std::ostringstream detail;
detail << "topology: unknown topology flags 0x" << std::hex << output.topologyFlags;
return Failure(Program203WitnessValidationFailure::Topology, detail.str());
return Failure(IterationRPWitnessValidationFailure::Topology, detail.str());
}
const std::uint32_t expectedMask = ExpectedSeenSubgroupMask(numSubgroups);
if (output.seenSubgroupMask != expectedMask) {
std::ostringstream detail;
detail << "topology: seen subgroup-ID mask was 0x" << std::hex << output.seenSubgroupMask
<< ", expected 0x" << expectedMask;
return Failure(Program203WitnessValidationFailure::Topology, detail.str());
return Failure(IterationRPWitnessValidationFailure::Topology, detail.str());
}
const std::uint32_t expectedLoopLength = ComputeProgram203WitnessLoopLength(numSubgroups);
const std::uint32_t expectedLoopLength = ComputeIterationRPWitnessLoopLength(numSubgroups);
if (output.loopLength != expectedLoopLength) {
std::ostringstream detail;
detail << "topology: loopLength was " << std::dec << output.loopLength << ", expected "
<< expectedLoopLength;
return Failure(Program203WitnessValidationFailure::Topology, detail.str());
return Failure(IterationRPWitnessValidationFailure::Topology, detail.str());
}
for (std::uint32_t subgroup = 0u; subgroup < numSubgroups; ++subgroup) {
if (output.lastLaneWriterCount[subgroup] != 1u) {
std::ostringstream detail;
detail << "topology: subgroup " << subgroup << " has "
<< output.lastLaneWriterCount[subgroup] << " source last-lane writers, expected exactly 1";
return Failure(Program203WitnessValidationFailure::Topology, detail.str(), 0u, subgroup);
return Failure(IterationRPWitnessValidationFailure::Topology, detail.str(), 0u, subgroup);
}
}
if (output.owner511.y != numSubgroups) {
std::ostringstream detail;
detail << "final owner: invocation 511 reported gl_NumSubgroups=" << output.owner511.y << ", expected "
<< numSubgroups;
return Failure(Program203WitnessValidationFailure::FinalOwner, detail.str());
return Failure(IterationRPWitnessValidationFailure::FinalOwner, detail.str());
}
if (output.owner511.z != numSubgroups - 1u) {
std::ostringstream detail;
detail << "final owner: invocation 511 is not in the highest subgroup (id" << output.owner511.z
<< ", expected id" << (numSubgroups - 1u) << ')';
return Failure(Program203WitnessValidationFailure::FinalOwner, detail.str());
return Failure(IterationRPWitnessValidationFailure::FinalOwner, detail.str());
}
if (output.owner511.x == 0u || output.owner511.w != output.owner511.x - 1u) {
std::ostringstream detail;
detail << "final owner: invocation 511 is not the last lane of highest subgroup (size "
<< output.owner511.x << ", lane " << output.owner511.w << ')';
return Failure(Program203WitnessValidationFailure::FinalOwner, detail.str());
return Failure(IterationRPWitnessValidationFailure::FinalOwner, detail.str());
}
// 3. Initial subgroup handoff. The atomic scalar totals are independent
@@ -218,22 +218,22 @@ namespace MobileGL::MG_Util::SelfTest {
if (indexedTotal != 131328u) {
std::ostringstream detail;
detail << "initial subgroup handoff: indexed input total was " << indexedTotal << ", expected 131328";
return Failure(Program203WitnessValidationFailure::InitialSubgroupHandoff, detail.str());
return Failure(IterationRPWitnessValidationFailure::InitialSubgroupHandoff, detail.str());
}
for (std::uint32_t subgroup = 0u; subgroup < numSubgroups; ++subgroup) {
const Program203WitnessVec2 expected = {static_cast<float>(output.indexedInputTotal[subgroup]), 0.0f};
const IterationRPWitnessVec2 expected = {static_cast<float>(output.indexedInputTotal[subgroup]), 0.0f};
if (!SameBits(output.rawPrefix[subgroup], expected)) {
std::ostringstream detail;
detail << "initial subgroup handoff: subgroup " << subgroup << " rawPrefix observed "
<< Vec2String(output.rawPrefix[subgroup]) << ", expected " << Vec2String(expected);
return Failure(Program203WitnessValidationFailure::InitialSubgroupHandoff, detail.str(), 0u,
return Failure(IterationRPWitnessValidationFailure::InitialSubgroupHandoff, detail.str(), 0u,
subgroup);
}
}
// 4. Source scan. Do not substitute a conventional scan: this reproduces
// the source cache index expression and stage ordering word for word.
std::array<Program203WitnessVec2, kProgram203WitnessMaxSubgroups> expectedCache = output.rawPrefix;
std::array<IterationRPWitnessVec2, kIterationRPWitnessMaxSubgroups> expectedCache = output.rawPrefix;
for (std::uint32_t scanStage = 0u; scanStage < expectedLoopLength; ++scanStage) {
auto cacheAfterStage = expectedCache;
for (std::uint32_t subgroup = 0u; subgroup < numSubgroups; ++subgroup) {
@@ -250,27 +250,27 @@ namespace MobileGL::MG_Util::SelfTest {
detail << "source scan stage " << scanStage << ", subgroup " << subgroup << ": observed "
<< Vec2String(output.scanCache[scanStage][subgroup]) << ", expected "
<< Vec2String(expectedCache[subgroup]);
return Failure(Program203WitnessValidationFailure::SourceScan, detail.str(), scanStage, subgroup);
return Failure(IterationRPWitnessValidationFailure::SourceScan, detail.str(), scanStage, subgroup);
}
}
}
// 5. The owner contract was checked above with the other topology facts;
// this final result remains a separate exact-vector check.
const Program203WitnessVec2 expectedAverage = {256.5f, 0.0f};
const IterationRPWitnessVec2 expectedAverage = {256.5f, 0.0f};
if (!SameBits(output.finalAverage, expectedAverage)) {
std::ostringstream detail;
detail << "final average: observed " << Vec2String(output.finalAverage) << ", expected "
<< Vec2String(expectedAverage);
return Failure(Program203WitnessValidationFailure::FinalAverage, detail.str());
return Failure(IterationRPWitnessValidationFailure::FinalAverage, detail.str());
}
std::ostringstream detail;
detail << "N=" << numSubgroups << ", owner511=id" << output.owner511.z << "/lane" << output.owner511.w
<< ", " << expectedLoopLength << " scan stages, average=" << Vec2String(output.finalAverage);
Program203WitnessValidationResult result;
IterationRPWitnessValidationResult result;
result.ok = true;
result.failure = Program203WitnessValidationFailure::None;
result.failure = IterationRPWitnessValidationFailure::None;
result.detail = detail.str();
return result;
}
@@ -0,0 +1,153 @@
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostIterationRPWitness.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Compact, native-Vulkan iterationRP first-reduction witness ABI and its pure
// validator. The types below deliberately mirror DriverPostIterationRPWitness.comp's
// single std430 storage block; changing either side requires updating the static
// layout assertions here.
#pragma once
#include <array>
#include <cstddef>
#include <cstdint>
#include <string>
#include <type_traits>
namespace MobileGL::MG_Util::SelfTest {
// "P203": the pack's trace program id, kept stable so the checked-in witness
// SPIR-V (DriverPostIterationRPWitnessSpv.h) needs no regeneration.
constexpr std::uint32_t kIterationRPWitnessMagic = 0x50323033u;
constexpr std::uint32_t kIterationRPWitnessInvocationCount = 512u;
constexpr std::uint32_t kIterationRPWitnessMaxSubgroups = 32u;
constexpr std::uint32_t kIterationRPWitnessMaxScanStages = 6u;
// These bit values are shared with the GLSL source. They document failures in
// topology observations rather than guessing a topology from local IDs on the host.
enum IterationRPWitnessTopologyFlag : std::uint32_t {
IterationRPWitnessNonuniformNumSubgroups = 1u << 0u,
IterationRPWitnessInvalidNumSubgroups = 1u << 1u,
IterationRPWitnessInvalidSubgroupId = 1u << 2u,
IterationRPWitnessInvalidSubgroupLane = 1u << 3u,
};
struct alignas(8) IterationRPWitnessVec2 {
float x;
float y;
};
struct alignas(16) IterationRPWitnessUVec4 {
std::uint32_t x;
std::uint32_t y;
std::uint32_t z;
std::uint32_t w;
};
// std430 layout of DriverPostIterationRPWitness.comp's IterationRPWitnessOutput block.
struct alignas(16) IterationRPWitnessOutput {
std::uint32_t magic;
std::uint32_t topologyFlags;
std::uint32_t numSubgroups;
std::uint32_t loopLength;
std::uint32_t seenSubgroupMask;
IterationRPWitnessUVec4 owner511;
std::array<std::uint32_t, kIterationRPWitnessMaxSubgroups> lastLaneWriterCount;
std::array<std::uint32_t, kIterationRPWitnessMaxSubgroups> indexedInputTotal;
std::array<IterationRPWitnessVec2, kIterationRPWitnessMaxSubgroups> rawPrefix;
std::array<std::array<IterationRPWitnessVec2, kIterationRPWitnessMaxSubgroups>,
kIterationRPWitnessMaxScanStages>
scanCache;
IterationRPWitnessVec2 finalAverage;
};
static_assert(std::is_standard_layout_v<IterationRPWitnessVec2>);
static_assert(std::is_standard_layout_v<IterationRPWitnessUVec4>);
static_assert(std::is_standard_layout_v<IterationRPWitnessOutput>);
static_assert(sizeof(IterationRPWitnessVec2) == 8u);
static_assert(alignof(IterationRPWitnessVec2) == 8u);
static_assert(sizeof(IterationRPWitnessUVec4) == 16u);
static_assert(alignof(IterationRPWitnessUVec4) == 16u);
static_assert(offsetof(IterationRPWitnessOutput, magic) == 0u);
static_assert(offsetof(IterationRPWitnessOutput, topologyFlags) == 4u);
static_assert(offsetof(IterationRPWitnessOutput, numSubgroups) == 8u);
static_assert(offsetof(IterationRPWitnessOutput, loopLength) == 12u);
static_assert(offsetof(IterationRPWitnessOutput, seenSubgroupMask) == 16u);
static_assert(offsetof(IterationRPWitnessOutput, owner511) == 32u);
static_assert(offsetof(IterationRPWitnessOutput, lastLaneWriterCount) == 48u);
static_assert(offsetof(IterationRPWitnessOutput, indexedInputTotal) == 176u);
static_assert(offsetof(IterationRPWitnessOutput, rawPrefix) == 304u);
static_assert(offsetof(IterationRPWitnessOutput, scanCache) == 560u);
static_assert(offsetof(IterationRPWitnessOutput, finalAverage) == 2096u);
static_assert(sizeof(IterationRPWitnessOutput) == 2112u);
// The witness uses prefixSumCache[32], three scalar shared diagnostics, and
// two 32-entry scalar diagnostic arrays in the GLSL source. Keep this
// independent of the output SSBO size.
constexpr std::uint32_t kIterationRPWitnessSharedMemoryBytes =
kIterationRPWitnessMaxSubgroups * sizeof(IterationRPWitnessVec2) +
3u * sizeof(std::uint32_t) +
2u * kIterationRPWitnessMaxSubgroups * sizeof(std::uint32_t);
enum class IterationRPWitnessEligibility {
Execute,
SkipUnsupportedNativeFeatureSet,
FailInadequateLimits,
};
// The raw physical-device conditions needed by the native witness. This is
// intentionally distinct from MobileGL's advertised-extension policy.
struct IterationRPWitnessLimits {
bool computeStageSupported = false;
bool basicSubgroupSupported = false;
bool arithmeticSubgroupSupported = false;
std::uint32_t subgroupSize = 0u;
std::uint32_t maxComputeWorkGroupInvocations = 0u;
std::array<std::uint32_t, 3> maxComputeWorkGroupSize{};
std::uint32_t maxComputeSharedMemorySize = 0u;
std::uint32_t maxPerStageDescriptorStorageBuffers = 0u;
std::uint32_t maxDescriptorSetStorageBuffers = 0u;
std::uint32_t maxBoundDescriptorSets = 0u;
std::uint64_t maxStorageBufferRange = 0u;
};
struct IterationRPWitnessEligibilityResult {
IterationRPWitnessEligibility eligibility = IterationRPWitnessEligibility::FailInadequateLimits;
std::string detail;
};
enum class IterationRPWitnessValidationFailure {
None,
Completion,
Topology,
InitialSubgroupHandoff,
SourceScan,
FinalOwner,
FinalAverage,
};
struct IterationRPWitnessValidationResult {
bool ok = false;
IterationRPWitnessValidationFailure failure = IterationRPWitnessValidationFailure::Completion;
std::uint32_t scanStage = 0u;
std::uint32_t subgroup = 0u;
std::string detail;
};
[[nodiscard]] IterationRPWitnessEligibilityResult
EvaluateIterationRPWitnessEligibility(const IterationRPWitnessLimits& limits);
// Mirrors the source's findMSB expression for valid N in [2, 32].
[[nodiscard]] std::uint32_t ComputeIterationRPWitnessLoopLength(std::uint32_t numSubgroups);
[[nodiscard]] IterationRPWitnessValidationResult
ValidateIterationRPWitness(const IterationRPWitnessOutput& output);
} // namespace MobileGL::MG_Util::SelfTest
@@ -1,4 +1,4 @@
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostProgram203WitnessSpv.h
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostIterationRPWitnessSpv.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
@@ -6,9 +6,15 @@
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Generated from DriverPostProgram203Witness.comp with:
// glslangValidator --target-env vulkan1.1 -V DriverPostProgram203Witness.comp
// Generated from DriverPostIterationRPWitness.comp with:
// glslangValidator --target-env vulkan1.1 -V DriverPostIterationRPWitness.comp
// Validated with spirv-val --target-env vulkan1.1. Do not edit words by hand.
//
// The stored words predate the Program203 -> IterationRP source rename, so their
// embedded OpName debug strings still spell the old identifiers; regeneration from
// the renamed source produces semantically identical code differing only in those
// strings. The witness magic stays 0x50323033 ("P203" - the trace's program id) so
// these words remain valid without regeneration.
#pragma once
@@ -16,7 +22,7 @@
#include <cstdint>
namespace MobileGL::MG_Util::SelfTest {
inline constexpr std::uint32_t kDriverPostProgram203WitnessSpv[] = {
inline constexpr std::uint32_t kDriverPostIterationRPWitnessSpv[] = {
0x07230203u, 0x00010300u, 0x0008000bu, 0x00000145u, 0x00000000u, 0x00020011u, 0x00000001u, 0x00020011u,
0x0000003du, 0x00020011u, 0x0000003fu, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x000a000fu, 0x00000005u, 0x00000004u, 0x6e69616du,
@@ -286,6 +292,6 @@ namespace MobileGL::MG_Util::SelfTest {
0x00050041u, 0x0000003bu, 0x00000141u, 0x00000037u, 0x00000124u, 0x0003003eu, 0x00000141u, 0x00000140u,
0x000200f9u, 0x0000013du, 0x000200f8u, 0x0000013du, 0x000100fdu, 0x00010038u,
};
inline constexpr std::size_t kDriverPostProgram203WitnessSpvWordCount =
sizeof(kDriverPostProgram203WitnessSpv) / sizeof(kDriverPostProgram203WitnessSpv[0]);
inline constexpr std::size_t kDriverPostIterationRPWitnessSpvWordCount =
sizeof(kDriverPostIterationRPWitnessSpv) / sizeof(kDriverPostIterationRPWitnessSpv[0]);
} // namespace MobileGL::MG_Util::SelfTest
@@ -1,151 +0,0 @@
// MobileGL - MobileGL/MG_Util/SelfTest/DriverPostProgram203Witness.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Compact, native-Vulkan Program-203 first-reduction witness ABI and its pure
// validator. The types below deliberately mirror DriverPostProgram203Witness.comp's
// single std430 storage block; changing either side requires updating the static
// layout assertions here.
#pragma once
#include <array>
#include <cstddef>
#include <cstdint>
#include <string>
#include <type_traits>
namespace MobileGL::MG_Util::SelfTest {
constexpr std::uint32_t kProgram203WitnessMagic = 0x50323033u; // "P203"
constexpr std::uint32_t kProgram203WitnessInvocationCount = 512u;
constexpr std::uint32_t kProgram203WitnessMaxSubgroups = 32u;
constexpr std::uint32_t kProgram203WitnessMaxScanStages = 6u;
// These bit values are shared with the GLSL source. They document failures in
// topology observations rather than guessing a topology from local IDs on the host.
enum Program203WitnessTopologyFlag : std::uint32_t {
Program203WitnessNonuniformNumSubgroups = 1u << 0u,
Program203WitnessInvalidNumSubgroups = 1u << 1u,
Program203WitnessInvalidSubgroupId = 1u << 2u,
Program203WitnessInvalidSubgroupLane = 1u << 3u,
};
struct alignas(8) Program203WitnessVec2 {
float x;
float y;
};
struct alignas(16) Program203WitnessUVec4 {
std::uint32_t x;
std::uint32_t y;
std::uint32_t z;
std::uint32_t w;
};
// std430 layout of DriverPostProgram203Witness.comp's Program203WitnessOutput block.
struct alignas(16) Program203WitnessOutput {
std::uint32_t magic;
std::uint32_t topologyFlags;
std::uint32_t numSubgroups;
std::uint32_t loopLength;
std::uint32_t seenSubgroupMask;
Program203WitnessUVec4 owner511;
std::array<std::uint32_t, kProgram203WitnessMaxSubgroups> lastLaneWriterCount;
std::array<std::uint32_t, kProgram203WitnessMaxSubgroups> indexedInputTotal;
std::array<Program203WitnessVec2, kProgram203WitnessMaxSubgroups> rawPrefix;
std::array<std::array<Program203WitnessVec2, kProgram203WitnessMaxSubgroups>,
kProgram203WitnessMaxScanStages>
scanCache;
Program203WitnessVec2 finalAverage;
};
static_assert(std::is_standard_layout_v<Program203WitnessVec2>);
static_assert(std::is_standard_layout_v<Program203WitnessUVec4>);
static_assert(std::is_standard_layout_v<Program203WitnessOutput>);
static_assert(sizeof(Program203WitnessVec2) == 8u);
static_assert(alignof(Program203WitnessVec2) == 8u);
static_assert(sizeof(Program203WitnessUVec4) == 16u);
static_assert(alignof(Program203WitnessUVec4) == 16u);
static_assert(offsetof(Program203WitnessOutput, magic) == 0u);
static_assert(offsetof(Program203WitnessOutput, topologyFlags) == 4u);
static_assert(offsetof(Program203WitnessOutput, numSubgroups) == 8u);
static_assert(offsetof(Program203WitnessOutput, loopLength) == 12u);
static_assert(offsetof(Program203WitnessOutput, seenSubgroupMask) == 16u);
static_assert(offsetof(Program203WitnessOutput, owner511) == 32u);
static_assert(offsetof(Program203WitnessOutput, lastLaneWriterCount) == 48u);
static_assert(offsetof(Program203WitnessOutput, indexedInputTotal) == 176u);
static_assert(offsetof(Program203WitnessOutput, rawPrefix) == 304u);
static_assert(offsetof(Program203WitnessOutput, scanCache) == 560u);
static_assert(offsetof(Program203WitnessOutput, finalAverage) == 2096u);
static_assert(sizeof(Program203WitnessOutput) == 2112u);
// The witness uses prefixSumCache[32], three scalar shared diagnostics, and
// two 32-entry scalar diagnostic arrays in the GLSL source. Keep this
// independent of the output SSBO size.
constexpr std::uint32_t kProgram203WitnessSharedMemoryBytes =
kProgram203WitnessMaxSubgroups * sizeof(Program203WitnessVec2) +
3u * sizeof(std::uint32_t) +
2u * kProgram203WitnessMaxSubgroups * sizeof(std::uint32_t);
enum class Program203WitnessEligibility {
Execute,
SkipUnsupportedNativeFeatureSet,
FailInadequateLimits,
};
// The raw physical-device conditions needed by the native witness. This is
// intentionally distinct from MobileGL's advertised-extension policy.
struct Program203WitnessLimits {
bool computeStageSupported = false;
bool basicSubgroupSupported = false;
bool arithmeticSubgroupSupported = false;
std::uint32_t subgroupSize = 0u;
std::uint32_t maxComputeWorkGroupInvocations = 0u;
std::array<std::uint32_t, 3> maxComputeWorkGroupSize{};
std::uint32_t maxComputeSharedMemorySize = 0u;
std::uint32_t maxPerStageDescriptorStorageBuffers = 0u;
std::uint32_t maxDescriptorSetStorageBuffers = 0u;
std::uint32_t maxBoundDescriptorSets = 0u;
std::uint64_t maxStorageBufferRange = 0u;
};
struct Program203WitnessEligibilityResult {
Program203WitnessEligibility eligibility = Program203WitnessEligibility::FailInadequateLimits;
std::string detail;
};
enum class Program203WitnessValidationFailure {
None,
Completion,
Topology,
InitialSubgroupHandoff,
SourceScan,
FinalOwner,
FinalAverage,
};
struct Program203WitnessValidationResult {
bool ok = false;
Program203WitnessValidationFailure failure = Program203WitnessValidationFailure::Completion;
std::uint32_t scanStage = 0u;
std::uint32_t subgroup = 0u;
std::string detail;
};
[[nodiscard]] Program203WitnessEligibilityResult
EvaluateProgram203WitnessEligibility(const Program203WitnessLimits& limits);
// Mirrors the source's findMSB expression for valid N in [2, 32].
[[nodiscard]] std::uint32_t ComputeProgram203WitnessLoopLength(std::uint32_t numSubgroups);
[[nodiscard]] Program203WitnessValidationResult
ValidateProgram203Witness(const Program203WitnessOutput& output);
} // namespace MobileGL::MG_Util::SelfTest
@@ -26,6 +26,8 @@
#include "SpirvPasses/RebaseInstanceIndexPass.h"
#include "SpirvPasses/ZeroBaseVertexPass.h"
#include "SpirvPasses/DeriveNumSubgroupsPass.h"
#include "SpirvPasses/EmulateSubgroupsPass.h"
#include "SpirvPasses/FixIterationRPSubgroupScratchPass.h"
#include "SpirvPasses/NormalizeRectCoordinatesPass.h"
#include "SpirvPasses/Lower1DArrayImagesPass.h"
#include "SpirvPasses/BakeImageFormatsPass.h"
@@ -895,6 +897,33 @@ namespace MobileGL {
outputBinary, true, enableSpirvValidation);
}
bool ShaderCompiler::EmulateSubgroupsForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary,
const Uint32 maxWorkgroupScratchBytes,
const bool enableSpirvValidation) {
using namespace spvtools;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(
EmulateSubgroupsPass::CreateEmulateSubgroupsPass(maxWorkgroupScratchBytes));
return RunOptimizerChecked("EmulateSubgroupsForVulkan", optimizer, inputBinary,
outputBinary, true, enableSpirvValidation);
}
bool ShaderCompiler::FixIterationRPSubgroupScratchForVulkan(
const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary,
const Uint32 nativeSubgroupSize, const Uint32 maxWorkgroupScratchBytes,
const bool enableSpirvValidation) {
using namespace spvtools;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(
FixIterationRPSubgroupScratchPass::CreateFixIterationRPSubgroupScratchPass(
nativeSubgroupSize, maxWorkgroupScratchBytes));
return RunOptimizerChecked("FixIterationRPSubgroupScratchForVulkan", optimizer,
inputBinary, outputBinary, true, enableSpirvValidation);
}
bool ShaderCompiler::DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary, const bool enableSpirvValidation) {
using namespace spvtools;
@@ -145,13 +145,40 @@ namespace MobileGL {
static bool ZeroBaseVertexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary,
bool enableSpirvValidation = false);
// Replaces compute gl_NumSubgroups loads with the value derived from the local
// workgroup dimensions and gl_SubgroupSize. DirectVulkan only; this avoids a
// driver builtin that can disagree with the subgroup IDs the same dispatch emits.
// See DeriveNumSubgroupsPass.
// Replaces compute gl_NumSubgroups loads with ceil(workgroup invocations /
// gl_SubgroupSize). DirectVulkan only; this repairs drivers whose builtin
// disagrees with the subgroup IDs the same dispatch emits (Adreno reports 1
// while emitting IDs 0..7). The ceil() partition is only spec-guaranteed
// under VK_PIPELINE_SHADER_STAGE_CREATE_REQUIRE_FULL_SUBGROUPS_BIT, which
// the caller requests whenever it is legal for the workgroup shape; see
// DeriveNumSubgroupsPass.
static bool DeriveNumSubgroupsForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary,
bool enableSpirvValidation = false);
// Lowers every GL_KHR_shader_subgroup construct in a compute module onto a
// 32-lane virtual subgroup built from workgroup-shared memory. Last-resort
// path for devices with NO native subgroup support, opt-in via
// MOBILEGL_MAGMA_EMULATE_SUBGROUP=1; a device with native subgroup
// operations always uses them. maxWorkgroupScratchBytes bounds the shared
// scratch the lowering may add (pass the device's
// maxComputeSharedMemorySize; 0 falls back to the 16384-byte Vulkan
// minimum). See EmulateSubgroupsPass.
static bool EmulateSubgroupsForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary,
Uint32 maxWorkgroupScratchBytes,
bool enableSpirvValidation = false);
// Grows iterationRP's under-declared gl_SubgroupID-indexed scratch to the
// subgroup count the device actually partitions into, fingerprint-gated to
// that pack's reduction idiom; every other module - and every device whose
// width the pack already assumed - passes through byte-identical.
// maxWorkgroupScratchBytes bounds the growth (pass the device's
// maxComputeSharedMemorySize; 0 falls back to the 16384-byte Vulkan
// minimum). See FixIterationRPSubgroupScratchPass.
static bool FixIterationRPSubgroupScratchForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary,
Uint32 nativeSubgroupSize,
Uint32 maxWorkgroupScratchBytes,
bool enableSpirvValidation = false);
// Re-declares 64-bit float vertex inputs as their 32-bit unsigned word pair
// (double -> uvec2, dvec2 -> uvec4) and bitcasts them back to double at entry, so no
// VK_FORMAT_R64*_SFLOAT is needed - lavapipe advertises none of them for vertex
@@ -179,8 +179,14 @@ namespace MobileGL {
: SynthesizeSubgroupSizeVariable(irContext, numSubgroupsVar->type_id());
const uint32_t workgroupSizeId = workgroupSize->result_id();
// The pipeline never enables ALLOW_VARYING_SUBGROUP_SIZE, so Vulkan's fixed
// subgroup partition is exactly ceil(local invocation count / SubgroupSize).
// ceil(local invocation count / SubgroupSize): the subgroup count of a
// full-subgroup launch. Vulkan only guarantees that partition under
// REQUIRE_FULL_SUBGROUPS - which ProgramFactory requests whenever
// local_size_x is a multiple of the subgroup size makes it legal
// (VUID-VkPipelineShaderStageCreateInfo-flags-02759) - and calls the
// tighter behaviour "encouraged" everywhere else; the DriverPost witness
// verifies it per device where the flag cannot be set. The absence of
// ALLOW_VARYING_SUBGROUP_SIZE pins only the SubgroupSize builtin itself.
// `(count - 1) / size + 1` avoids an addition overflow at count + size - 1.
for (Instruction* load : numSubgroupsLoads) {
const uint32_t localSizeXId = irContext->TakeNextId();
@@ -19,12 +19,14 @@ namespace MobileGL {
// Replaces compute-stage NumSubgroups builtin loads with
// ceil(WorkgroupSize.x * WorkgroupSize.y * WorkgroupSize.z / SubgroupSize).
//
// That is the value Vulkan defines for NumSubgroups when the pipeline does not
// enable varying subgroup sizes, which MobileGL never does. Deriving it avoids
// drivers that expose the real SubgroupId topology but return an inconsistent
// NumSubgroups value. This is a DirectVulkan semantic repair, not a source-shader
// rewrite; the application's subgroup arithmetic and shared-memory logic remain
// unchanged.
// That is the subgroup count of a full-subgroup launch - guaranteed by Vulkan
// under REQUIRE_FULL_SUBGROUPS (which ProgramFactory requests whenever the
// workgroup shape makes it legal), spec-"encouraged" and witness-verified
// (DriverPost) elsewhere. Deriving it repairs drivers that expose the real
// SubgroupId topology but return an inconsistent NumSubgroups value, breaking
// GL's gl_SubgroupID < gl_NumSubgroups contract. This is a DirectVulkan
// semantic repair, not a source-shader rewrite; the application's subgroup
// arithmetic and shared-memory logic remain unchanged.
class DeriveNumSubgroupsPass : public spvtools::opt::Pass {
public:
const char* name() const override { return "derive-num-subgroups"; }
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,73 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateSubgroupsPass.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Lowers every GL_KHR_shader_subgroup construct in a compute module onto a
// 32-lane VIRTUAL subgroup implemented with workgroup-shared memory. Virtual
// subgroups partition the workgroup by gl_LocalInvocationIndex:
// lane = index & 31, id = index >> 5, count = ceil(invocations / 32).
//
// This is a LAST-RESORT path, never a substitute for real subgroups: it only
// runs when MOBILEGL_MAGMA_EMULATE_SUBGROUP=1 is set explicitly and the device
// has no native subgroup support at all (SubgroupSupportPolicy.h). A device
// with native subgroup operations - however narrow - uses them natively, with
// FixIterationRPSubgroupScratchPass patching the known pack bug instead.
//
// Lowered constructs:
// - the builtins gl_SubgroupSize / gl_SubgroupInvocationID / gl_SubgroupID /
// gl_NumSubgroups and the five gl_Subgroup*Mask ballot builtins;
// - OpGroupNonUniform{Elect,All,Any,AllEqual,Broadcast,BroadcastFirst,
// Ballot,InverseBallot,BallotBitExtract,BallotBitCount,BallotFind{L,M}SB,
// Shuffle,ShuffleXor,ShuffleUp,ShuffleDown,
// <arithmetic/min/max/bitwise/logical reduce+scans+clustered>,
// QuadBroadcast,QuadSwap};
// - subgroupBarrier()/subgroupMemoryBarrier*() (their Subgroup scopes widen
// to Workgroup, which is strictly stronger).
// The output uses no GroupNonUniform* instruction or capability at all, which
// is what lets it run on devices with no subgroup feature bits.
//
// Semantic contract, narrower than native subgroups in exactly one way: every
// emulated exchange synchronizes through OpControlBarrier, so subgroup
// operations must sit in WORKGROUP-uniform control flow (the shape every
// Iris-style pack reduction has). GLSL already imposes this for barrier();
// a subgroup op in divergent flow - legal on native subgroups - is undefined
// here.
//
// Fails (Status::Failure, leaving the input module unchanged) on anything it
// cannot lower faithfully: extended subgroup ops (partitioned-NV, rotate,
// quad-all/any), non-32-bit participating types, spec-constant workgroup
// sizes, a subgroup builtin reached by anything but a direct OpLoad, or a
// module whose lowering would add more workgroup scratch than
// maxWorkgroupScratchBytes (pass the device's maxComputeSharedMemorySize;
// 0 falls back to the 16384-byte Vulkan minimum).
class EmulateSubgroupsPass : public spvtools::opt::Pass {
public:
explicit EmulateSubgroupsPass(Uint32 maxWorkgroupScratchBytes)
: m_maxWorkgroupScratchBytes(maxWorkgroupScratchBytes) {}
const char* name() const override { return "emulate-subgroups"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateEmulateSubgroupsPass(
Uint32 maxWorkgroupScratchBytes);
private:
Uint32 m_maxWorkgroupScratchBytes;
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,558 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FixIterationRPSubgroupScratchPass.cpp
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "FixIterationRPSubgroupScratchPass.h"
#include "spirv.hpp"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/util/make_unique.h"
#include <map>
#include <unordered_map>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
// The Vulkan minimum for maxComputeSharedMemorySize, used when the caller
// could not tell us the device's real limit.
constexpr uint32_t kMinimumSharedMemoryBytes = 16384u;
// The narrowest subgroup width iterationRP's declarations are sized for.
// At or above it both shipped shapes fit and nothing may be rewritten.
constexpr uint32_t kPackAssumedSubgroupWidth = 16u;
Instruction* FindBuiltinDefinition(IRContext* context, spv::BuiltIn builtin) {
auto* defUseMgr = context->get_def_use_mgr();
for (auto& annotation : context->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate || annotation.NumInOperands() < 3) {
continue;
}
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) !=
spv::Decoration::BuiltIn) {
continue;
}
if (static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(2)) != builtin) {
continue;
}
return defUseMgr->GetDef(annotation.GetSingleWordInOperand(0));
}
return nullptr;
}
// Walks an access-chain pointer expression back to the variable it is
// rooted at; returns nullptr for anything that is not a plain chain.
const Instruction* RootVariable(IRContext* context, uint32_t pointerId) {
auto* defUseMgr = context->get_def_use_mgr();
const Instruction* def = defUseMgr->GetDef(pointerId);
while (def != nullptr) {
switch (def->opcode()) {
case spv::Op::OpVariable:
return def;
case spv::Op::OpAccessChain:
case spv::Op::OpInBoundsAccessChain:
case spv::Op::OpCopyObject:
def = defUseMgr->GetDef(def->GetSingleWordInOperand(0));
break;
default:
return nullptr;
}
}
return nullptr;
}
// A 32-bit float scalar or vector - the shape of every accumulator the
// pack runs through its scans (float, vec2 and vec4 all appear). Returns
// the component count, or 0 for anything else.
uint32_t Float32ComponentCount(IRContext* context, uint32_t typeId) {
auto* defUseMgr = context->get_def_use_mgr();
const Instruction* type = defUseMgr->GetDef(typeId);
if (type == nullptr) return 0u;
uint32_t components = 1u;
if (type->opcode() == spv::Op::OpTypeVector) {
components = type->GetSingleWordInOperand(1);
if (components < 2u || components > 4u) return 0u;
type = defUseMgr->GetDef(type->GetSingleWordInOperand(0));
if (type == nullptr) return 0u;
}
if (type->opcode() != spv::Op::OpTypeFloat ||
type->GetSingleWordInOperand(0) != 32u) {
return 0u;
}
return components;
}
uint32_t RoundUp(uint32_t value, uint32_t alignment) {
return alignment == 0u ? value : ((value + alignment - 1u) / alignment) * alignment;
}
// Size AND alignment of a workgroup-storage type. Drivers lay shared
// memory out at natural alignment and the limit
// (VUID-RuntimeSpirv-Workgroup-06530) counts the padding that produces,
// so a model that sums unpadded sizes would under-count exactly where the
// budget check matters. Returns false for anything not modelled here,
// which the caller answers by declining to grow at all rather than by
// certifying growth against a total it knows is an underestimate.
bool WorkgroupTypeLayout(IRContext* context, uint32_t typeId, uint32_t* size,
uint32_t* alignment, uint32_t depth = 0u) {
if (depth > 8u) return false;
auto* defUseMgr = context->get_def_use_mgr();
const Instruction* type = defUseMgr->GetDef(typeId);
if (type == nullptr) return false;
switch (type->opcode()) {
case spv::Op::OpTypeBool:
*size = 4u;
*alignment = 4u;
return true;
case spv::Op::OpTypeInt:
case spv::Op::OpTypeFloat: {
const uint32_t width = type->GetSingleWordInOperand(0) / 8u;
if (width == 0u) return false;
*size = width;
*alignment = width;
return true;
}
case spv::Op::OpTypeVector: {
uint32_t componentSize = 0u;
uint32_t componentAlignment = 0u;
if (!WorkgroupTypeLayout(context, type->GetSingleWordInOperand(0),
&componentSize, &componentAlignment, depth + 1u)) {
return false;
}
const uint32_t components = type->GetSingleWordInOperand(1);
if (components < 2u || components > 4u) return false;
*size = componentSize * components;
// A three-component vector aligns like a four-component one.
*alignment = componentSize * (components == 3u ? 4u : components);
return true;
}
case spv::Op::OpTypeMatrix:
case spv::Op::OpTypeArray: {
uint32_t elementSize = 0u;
uint32_t elementAlignment = 0u;
if (!WorkgroupTypeLayout(context, type->GetSingleWordInOperand(0), &elementSize,
&elementAlignment, depth + 1u)) {
return false;
}
uint32_t count = 0u;
if (type->opcode() == spv::Op::OpTypeMatrix) {
count = type->GetSingleWordInOperand(1);
} else {
const Instruction* length =
defUseMgr->GetDef(type->GetSingleWordInOperand(1));
if (length == nullptr || length->opcode() != spv::Op::OpConstant) {
return false; // spec-constant length: not sizeable here
}
count = length->GetSingleWordInOperand(0);
}
*size = RoundUp(elementSize, elementAlignment) * count;
*alignment = elementAlignment;
return true;
}
case spv::Op::OpTypeStruct: {
uint32_t offset = 0u;
uint32_t structAlignment = 1u;
for (uint32_t i = 0; i < type->NumInOperands(); ++i) {
uint32_t memberSize = 0u;
uint32_t memberAlignment = 0u;
if (!WorkgroupTypeLayout(context, type->GetSingleWordInOperand(i),
&memberSize, &memberAlignment, depth + 1u)) {
return false;
}
offset = RoundUp(offset, memberAlignment) + memberSize;
if (memberAlignment > structAlignment) structAlignment = memberAlignment;
}
*size = RoundUp(offset, structAlignment);
*alignment = structAlignment;
return true;
}
default:
return false;
}
}
// The group operations the pack's prefix scans use.
bool IsScanOrReduce(spv::GroupOperation operation) {
return operation == spv::GroupOperation::Reduce ||
operation == spv::GroupOperation::InclusiveScan ||
operation == spv::GroupOperation::ExclusiveScan;
}
} // namespace
spvtools::opt::Pass::Status FixIterationRPSubgroupScratchPass::Process() {
auto* irContext = context();
auto* defUseMgr = irContext->get_def_use_mgr();
// Without a known device width there is no topology to compare against;
// and a width the pack already assumed needs no patch at all. Both of
// iterationRP's shapes are sized for >= 16 lanes (512/16 = 32 entries,
// 1024/16 = 64), so every module on such a device - the pack's or anyone
// else's - must pass through byte-identical. The per-array length test
// further down is the second gate, not a replacement for this one.
if (m_nativeSubgroupSize == 0u || m_nativeSubgroupSize >= kPackAssumedSubgroupWidth) {
return Status::SuccessWithoutChange;
}
for (const Instruction& entryPoint : irContext->module()->entry_points()) {
if (static_cast<spv::ExecutionModel>(entryPoint.GetSingleWordInOperand(0)) !=
spv::ExecutionModel::GLCompute) {
return Status::SuccessWithoutChange;
}
}
// Fingerprint 1: a literal workgroup size, so the subgroup count the
// dispatch actually partitions into is known here.
const auto resolveUintConstant = [&](uint32_t id, uint32_t* value) {
const Instruction* def = defUseMgr->GetDef(id);
if (def == nullptr || def->opcode() != spv::Op::OpConstant) return false;
*value = def->GetSingleWordInOperand(0);
return true;
};
uint32_t localSize[3] = {0, 0, 0};
bool haveLocalSize = false;
if (Instruction* workgroupSize =
FindBuiltinDefinition(irContext, spv::BuiltIn::WorkgroupSize)) {
if (workgroupSize->opcode() == spv::Op::OpConstantComposite &&
workgroupSize->NumInOperands() == 3) {
haveLocalSize =
resolveUintConstant(workgroupSize->GetSingleWordInOperand(0), &localSize[0]) &&
resolveUintConstant(workgroupSize->GetSingleWordInOperand(1), &localSize[1]) &&
resolveUintConstant(workgroupSize->GetSingleWordInOperand(2), &localSize[2]);
}
}
if (!haveLocalSize) {
for (const Instruction& mode : irContext->module()->execution_modes()) {
if (mode.opcode() == spv::Op::OpExecutionMode &&
static_cast<spv::ExecutionMode>(mode.GetSingleWordInOperand(1)) ==
spv::ExecutionMode::LocalSize) {
localSize[0] = mode.GetSingleWordInOperand(2);
localSize[1] = mode.GetSingleWordInOperand(3);
localSize[2] = mode.GetSingleWordInOperand(4);
haveLocalSize = true;
break;
}
}
}
if (!haveLocalSize || localSize[0] == 0u || localSize[1] == 0u || localSize[2] == 0u) {
return Status::SuccessWithoutChange;
}
const uint64_t totalInvocations =
static_cast<uint64_t>(localSize[0]) * localSize[1] * localSize[2];
if (totalInvocations == 0u || totalInvocations > (1u << 20)) {
return Status::SuccessWithoutChange;
}
const uint32_t requiredLength = static_cast<uint32_t>(
(totalInvocations + m_nativeSubgroupSize - 1u) / m_nativeSubgroupSize);
// Fingerprint 2: a subgroup scan over a 32-bit float value - the pack's
// prefix-sum reduction, and the reason its scratch is indexed per subgroup.
bool sawFloatSubgroupScan = false;
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
if (inst.opcode() != spv::Op::OpGroupNonUniformFAdd &&
inst.opcode() != spv::Op::OpGroupNonUniformFMin &&
inst.opcode() != spv::Op::OpGroupNonUniformFMax) {
continue;
}
if (inst.NumInOperands() < 2) continue;
if (!IsScanOrReduce(static_cast<spv::GroupOperation>(
inst.GetSingleWordInOperand(1)))) {
continue;
}
if (Float32ComponentCount(irContext, inst.type_id()) != 0u) {
sawFloatSubgroupScan = true;
}
}
}
}
if (!sawFloatSubgroupScan) {
return Status::SuccessWithoutChange;
}
// gl_SubgroupID, whose value range the pack's scratch size bakes in.
const Instruction* subgroupIdVariable =
FindBuiltinDefinition(irContext, spv::BuiltIn::SubgroupId);
if (subgroupIdVariable == nullptr ||
subgroupIdVariable->opcode() != spv::Op::OpVariable) {
return Status::SuccessWithoutChange;
}
const uint32_t subgroupIdVariableId = subgroupIdVariable->result_id();
// The pack indexes its scratch with gl_SubgroupID ITSELF, so only values
// that ARE that id qualify - not everything computed from it. An index
// that is masked or clamped (cache[gl_SubgroupID & 3u]) is bounded by
// construction and is none of this pass's business; accepting it would
// turn a targeted repair into a general array resizer. Identity survives
// OpCopyObject, a signedness OpBitcast, and the Function/Private spill
// glslang emits for a builtin load - and nothing else. A spill variable
// counts only when EVERY store into it is the id.
std::unordered_map<uint32_t, bool> subgroupIdValues; // result id IS the id
std::unordered_map<uint32_t, bool> subgroupIdVariables; // spill holding only it
bool changedIdentity = true;
while (changedIdentity) {
changedIdentity = false;
std::unordered_map<uint32_t, uint32_t> totalStores;
std::unordered_map<uint32_t, uint32_t> idStores;
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
if (inst.opcode() != spv::Op::OpStore) continue;
const uint32_t pointerId = inst.GetSingleWordInOperand(0);
const Instruction* target = defUseMgr->GetDef(pointerId);
if (target == nullptr || target->opcode() != spv::Op::OpVariable) {
continue;
}
const auto storageClass = static_cast<spv::StorageClass>(
target->GetSingleWordInOperand(0));
if (storageClass != spv::StorageClass::Function &&
storageClass != spv::StorageClass::Private) {
continue;
}
totalStores[pointerId] += 1u;
if (subgroupIdValues.count(inst.GetSingleWordInOperand(1))) {
idStores[pointerId] += 1u;
}
}
}
}
for (const auto& entry : totalStores) {
if (entry.second != 0u && idStores[entry.first] == entry.second &&
!subgroupIdVariables.count(entry.first)) {
subgroupIdVariables[entry.first] = true;
changedIdentity = true;
}
}
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
if (inst.result_id() == 0 ||
subgroupIdValues.count(inst.result_id())) {
continue;
}
bool isSubgroupId = false;
switch (inst.opcode()) {
case spv::Op::OpLoad: {
const uint32_t pointerId = inst.GetSingleWordInOperand(0);
isSubgroupId = pointerId == subgroupIdVariableId ||
subgroupIdVariables.count(pointerId) != 0u;
break;
}
case spv::Op::OpCopyObject:
case spv::Op::OpBitcast:
isSubgroupId =
subgroupIdValues.count(inst.GetSingleWordInOperand(0)) != 0u;
break;
default:
break;
}
if (isSubgroupId) {
subgroupIdValues[inst.result_id()] = true;
changedIdentity = true;
}
}
}
}
}
if (subgroupIdValues.empty()) {
return Status::SuccessWithoutChange;
}
// Fingerprint 3: workgroup-shared float arrays indexed by gl_SubgroupID
// itself - the under-declared prefixSumCache.
std::map<uint32_t, Instruction*> candidates;
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
if (inst.opcode() != spv::Op::OpAccessChain &&
inst.opcode() != spv::Op::OpInBoundsAccessChain) {
continue;
}
if (inst.NumInOperands() < 2) continue;
if (!subgroupIdValues.count(inst.GetSingleWordInOperand(1))) continue;
Instruction* baseVariable =
defUseMgr->GetDef(inst.GetSingleWordInOperand(0));
if (baseVariable == nullptr ||
baseVariable->opcode() != spv::Op::OpVariable ||
static_cast<spv::StorageClass>(
baseVariable->GetSingleWordInOperand(0)) !=
spv::StorageClass::Workgroup) {
continue;
}
candidates.emplace(baseVariable->result_id(), baseVariable);
}
}
}
if (candidates.empty()) {
return Status::SuccessWithoutChange;
}
// Everything that survives the filter, with the bytes each grown array
// will need. Nothing is mutated until the whole set fits the device's
// shared-memory budget, so a module is never left half-grown.
struct Growth {
Instruction* variable = nullptr;
uint32_t elementTypeId = 0;
uint32_t lengthTypeId = 0;
uint32_t addedBytes = 0;
};
std::vector<Growth> growths;
for (auto& entry : candidates) {
Instruction* variable = entry.second;
// The variable must be reached exclusively through access chains (plus
// debug/decoration instructions): a whole-array load, store, or copy
// would change type with the array and is left alone.
bool onlyAccessChains = true;
const uint32_t variableId = variable->result_id();
defUseMgr->ForEachUser(variable, [&](Instruction* user) {
switch (user->opcode()) {
case spv::Op::OpAccessChain:
case spv::Op::OpInBoundsAccessChain:
if (user->GetSingleWordInOperand(0) != variableId) {
onlyAccessChains = false;
}
return;
case spv::Op::OpName:
case spv::Op::OpDecorate:
return;
default:
onlyAccessChains = false;
return;
}
});
if (!onlyAccessChains) continue;
if (variable->NumInOperands() > 1) continue; // initializer: leave alone
const Instruction* pointerType = defUseMgr->GetDef(variable->type_id());
if (pointerType == nullptr || pointerType->opcode() != spv::Op::OpTypePointer) {
continue;
}
const Instruction* arrayType =
defUseMgr->GetDef(pointerType->GetSingleWordInOperand(1));
if (arrayType == nullptr || arrayType->opcode() != spv::Op::OpTypeArray) {
continue;
}
const uint32_t elementTypeId = arrayType->GetSingleWordInOperand(0);
const uint32_t components = Float32ComponentCount(irContext, elementTypeId);
if (components == 0u) continue;
const Instruction* lengthConstant =
defUseMgr->GetDef(arrayType->GetSingleWordInOperand(1));
if (lengthConstant == nullptr || lengthConstant->opcode() != spv::Op::OpConstant) {
continue;
}
const uint32_t currentLength = lengthConstant->GetSingleWordInOperand(0);
// The pack's own assumption holds on this device: the declared array
// already covers every subgroup the workgroup partitions into. That is
// every >= 16-lane device for the shapes iterationRP ships, and those
// modules must pass through byte-identical.
if (currentLength >= requiredLength) continue;
// vec3 strides at its 16-byte alignment, so charge the padded stride.
const uint32_t elementStride = (components == 3u ? 4u : components) * 4u;
growths.push_back(Growth{variable, elementTypeId, lengthConstant->type_id(),
(requiredLength - currentLength) * elementStride});
}
if (growths.empty()) {
return Status::SuccessWithoutChange;
}
// Growing must not push the module past what the device can launch: a
// pipeline that fails to create is worse than the pack's own overrun.
{
uint64_t declaredBytes = 0;
bool sawUnsizeable = false;
for (auto& global : irContext->module()->types_values()) {
if (global.opcode() != spv::Op::OpVariable ||
static_cast<spv::StorageClass>(global.GetSingleWordInOperand(0)) !=
spv::StorageClass::Workgroup) {
continue;
}
const Instruction* pointerType = defUseMgr->GetDef(global.type_id());
uint32_t bytes = 0u;
uint32_t alignment = 0u;
if (pointerType == nullptr ||
pointerType->opcode() != spv::Op::OpTypePointer ||
!WorkgroupTypeLayout(irContext, pointerType->GetSingleWordInOperand(1),
&bytes, &alignment)) {
sawUnsizeable = true;
break;
}
declaredBytes = RoundUp(static_cast<uint32_t>(declaredBytes), alignment) + bytes;
}
// A declaration this pass cannot size leaves the total an
// underestimate, so the growth cannot be certified against the device
// limit at all - decline rather than guess.
if (sawUnsizeable) {
return Status::SuccessWithoutChange;
}
for (const Growth& growth : growths) declaredBytes += growth.addedBytes;
const uint32_t deviceBudget = m_maxWorkgroupScratchBytes != 0u
? m_maxWorkgroupScratchBytes
: kMinimumSharedMemoryBytes;
if (declaredBytes > deviceBudget) {
return Status::SuccessWithoutChange;
}
}
for (const Growth& growth : growths) {
// Build the grown array type. All three new instructions are inserted
// immediately BEFORE the variable so definition-before-use holds in the
// module's global section (manager-created instructions append to its
// end, after the variable). The new length constant reuses the old
// one's integer type, whatever signedness glslang gave it (a duplicate
// scalar constant is legal SPIR-V); the fresh array type makes the
// pointer type unique by construction, so neither collides with an
// existing declaration.
Instruction* variable = growth.variable;
const uint32_t newLengthId = irContext->TakeNextId();
variable->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpConstant, growth.lengthTypeId, newLengthId,
Instruction::OperandList{{SPV_OPERAND_TYPE_TYPED_LITERAL_NUMBER,
{requiredLength}}}));
const uint32_t newArrayTypeId = irContext->TakeNextId();
variable->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpTypeArray, 0, newArrayTypeId,
Instruction::OperandList{
{SPV_OPERAND_TYPE_ID, {growth.elementTypeId}},
{SPV_OPERAND_TYPE_ID, {newLengthId}}}));
const uint32_t newPointerTypeId = irContext->TakeNextId();
variable->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpTypePointer, 0, newPointerTypeId,
Instruction::OperandList{
{SPV_OPERAND_TYPE_STORAGE_CLASS,
{static_cast<uint32_t>(spv::StorageClass::Workgroup)}},
{SPV_OPERAND_TYPE_ID, {newArrayTypeId}}}));
variable->SetResultType(newPointerTypeId);
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken
FixIterationRPSubgroupScratchPass::CreateFixIterationRPSubgroupScratchPass(
const Uint32 nativeSubgroupSize, const Uint32 maxWorkgroupScratchBytes) {
return spvtools::Optimizer::PassToken(MakeUnique<FixIterationRPSubgroupScratchPass>(
nativeSubgroupSize, maxWorkgroupScratchBytes));
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,73 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FixIterationRPSubgroupScratchPass.h
// Copyright (c) 2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Patches ONE known shader-pack defect: iterationRP hard-sizes the scratch
// its subgroup prefix scans write through prefixSumCache[gl_SubgroupID].
// The pack ships that idiom twice, sized for the >= 16-lane subgroups
// desktop GL drivers give it:
// - the auto-exposure reduction: 32x16 (512 invocations), vec2[32];
// - the RTW importance warp: 1024 invocations, float[64].
// On a narrower Vulkan device (lavapipe's 8 lanes -> 64 and 128 subgroups)
// every subgroup past the last declared entry indexes shared memory out of
// bounds - on a CPU rasterizer that is literal heap corruption. Both
// reduction ALGORITHMS are width-agnostic (their combine loops are sized by
// gl_NumSubgroups), so the faithful repair is to grow the under-declared
// arrays to ceil(invocations / native width) and change nothing else.
//
// This is the pack author's bug, not MobileGL's, so the patch is
// deliberately NOT a general "resize shared arrays" mechanism. It rewrites
// an array only when the module positively matches the pack's reduction
// idiom AND the device's own topology proves the declaration too small:
// - GLCompute entry point with a literal workgroup size;
// - a subgroup scan/reduce over a 32-bit float scalar or vector
// (OpGroupNonUniformF{Add,Min,Max}), the pack's accumulator signature;
// - a workgroup-shared array of 32-bit float scalars/vectors whose
// access-chain index is data-dependent on gl_SubgroupID;
// - a declared length strictly below ceil(invocations / native width).
// That last clause is what keeps the patch inert wherever the pack is
// correct: on any device whose width satisfies the pack's assumption
// (>= 16 lanes: desktop GL, Adreno) both shapes already fit and every
// module passes through byte-identical. Matching at the SPIR-V level keeps
// recognition robust against the whitespace/identifier drift that made the
// old source-text template rewrite (removed in 7769156) so brittle.
//
// The pass never fails a module: anything it cannot prove is this pattern -
// or cannot grow safely (a whole-array use, a spec-constant length, an
// initializer, or growth that would not fit maxWorkgroupScratchBytes) - is
// left exactly as it was. Pass the device's maxComputeSharedMemorySize as
// maxWorkgroupScratchBytes; 0 falls back to the 16384-byte Vulkan minimum.
class FixIterationRPSubgroupScratchPass : public spvtools::opt::Pass {
public:
FixIterationRPSubgroupScratchPass(Uint32 nativeSubgroupSize,
Uint32 maxWorkgroupScratchBytes)
: m_nativeSubgroupSize(nativeSubgroupSize),
m_maxWorkgroupScratchBytes(maxWorkgroupScratchBytes) {}
const char* name() const override { return "fix-iterationrp-subgroup-scratch"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateFixIterationRPSubgroupScratchPass(
Uint32 nativeSubgroupSize, Uint32 maxWorkgroupScratchBytes);
private:
Uint32 m_nativeSubgroupSize;
Uint32 m_maxWorkgroupScratchBytes;
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -164,11 +164,6 @@ bool LoadMobileGL(const Request& request, std::string& error) {
} else {
unsetenv("MOBILEGL_COHERENT_AS_FLUSH");
}
if (request.numSubgroupsQuirk) {
setenv("MOBILEGL_NUM_SUBGROUPS_QUIRK", "1", 1);
} else {
unsetenv("MOBILEGL_NUM_SUBGROUPS_QUIRK");
}
if (request.fboAttachmentDumps.empty()) {
unsetenv("MOBILEGL_TRACE_DUMP_FBO_ATTACHMENTS");
} else {
@@ -823,7 +818,6 @@ bool WriteResultJson(const Request& request, const Result& result) {
<< (request.avoidAngleLlvmpipeSamplerMipmapMinFilter ? "true" : "false") << ",\n";
file << " \"avoidAngleLlvmpipeExplicitLodBias\": "
<< (request.avoidAngleLlvmpipeExplicitLodBias ? "true" : "false") << ",\n";
file << " \"numSubgroupsQuirk\": " << (request.numSubgroupsQuirk ? "true" : "false") << ",\n";
file << " \"holdMs\": " << request.holdMs << ",\n";
file << " \"mismatchPixels\": " << result.mismatchPixels << "\n";
file << "}\n";
@@ -44,7 +44,6 @@ struct Request {
bool avoidAngleLlvmpipeSamplerMipmapMinFilter = false;
bool avoidAngleLlvmpipeExplicitLodBias = false;
bool coherentAsFlush = false;
bool numSubgroupsQuirk = false;
int holdMs = 0;
};
@@ -122,7 +122,6 @@ Java_top_mobilegl_plugin_trace_TraceReplayActivity_nativeRunTraceReplay(JNIEnv*
jboolean avoidAngleLlvmpipeSamplerMipmapMinFilter,
jboolean avoidAngleLlvmpipeExplicitLodBias,
jboolean coherentAsFlush,
jboolean numSubgroupsQuirk,
jstring texture2dDumps) {
mobilegl_trace::Request request;
request.tracePath = ToString(env, tracePath);
@@ -151,7 +150,6 @@ Java_top_mobilegl_plugin_trace_TraceReplayActivity_nativeRunTraceReplay(JNIEnv*
avoidAngleLlvmpipeSamplerMipmapMinFilter == JNI_TRUE;
request.avoidAngleLlvmpipeExplicitLodBias = avoidAngleLlvmpipeExplicitLodBias == JNI_TRUE;
request.coherentAsFlush = coherentAsFlush == JNI_TRUE;
request.numSubgroupsQuirk = numSubgroupsQuirk == JNI_TRUE;
ScopedTraceReplayState replayState;
mobilegl_trace_set_requested_size(request.width, request.height);
@@ -116,7 +116,6 @@ public final class TraceReplayActivity extends Activity {
request.avoidAngleLlvmpipeSamplerMipmapMinFilter,
request.avoidAngleLlvmpipeExplicitLodBias,
request.coherentAsFlush,
request.numSubgroupsQuirk,
request.texture2dDumps
);
Log.i(TAG, result.toString());
@@ -150,7 +149,6 @@ public final class TraceReplayActivity extends Activity {
boolean avoidAngleLlvmpipeSamplerMipmapMinFilter,
boolean avoidAngleLlvmpipeExplicitLodBias,
boolean coherentAsFlush,
boolean numSubgroupsQuirk,
String texture2dDumps
);
@@ -176,7 +174,6 @@ public final class TraceReplayActivity extends Activity {
final boolean avoidAngleLlvmpipeSamplerMipmapMinFilter;
final boolean avoidAngleLlvmpipeExplicitLodBias;
final boolean coherentAsFlush;
final boolean numSubgroupsQuirk;
final String texture2dDumps;
private TraceReplayRequest(
@@ -201,8 +198,7 @@ public final class TraceReplayActivity extends Activity {
boolean avoidAngleLlvmpipeSamplerMipmapMinFilter,
boolean avoidAngleLlvmpipeExplicitLodBias,
boolean coherentAsFlush,
boolean numSubgroupsQuirk,
String texture2dDumps
String texture2dDumps
) {
this.tracePath = tracePath;
this.goldenPath = goldenPath;
@@ -225,7 +221,6 @@ public final class TraceReplayActivity extends Activity {
this.avoidAngleLlvmpipeSamplerMipmapMinFilter = avoidAngleLlvmpipeSamplerMipmapMinFilter;
this.avoidAngleLlvmpipeExplicitLodBias = avoidAngleLlvmpipeExplicitLodBias;
this.coherentAsFlush = coherentAsFlush;
this.numSubgroupsQuirk = numSubgroupsQuirk;
this.texture2dDumps = texture2dDumps;
}
@@ -254,7 +249,6 @@ public final class TraceReplayActivity extends Activity {
intent.getBooleanExtra("avoid_angle_llvmpipe_sampler_mipmap_min_filter", false),
intent.getBooleanExtra("avoid_angle_llvmpipe_explicit_lod_bias", false),
intent.getBooleanExtra("coherent_as_flush", false),
intent.getBooleanExtra("num_subgroups_quirk", false),
readString(intent, "texture_2d_dumps", "")
);
}
-8
View File
@@ -31,7 +31,6 @@ Usage:
[--avoid-angle-llvmpipe-sampler-mipmap-min-filter] \
[--avoid-angle-llvmpipe-explicit-lod-bias] \
[--coherent-as-flush] \
[--num-subgroups-quirk] \
[--dump-texture-2d CALL,TEXTURE,LEVEL,DIR] \
--timeout-seconds N
@@ -48,8 +47,6 @@ sample with an explicit LOD that ANGLE llvmpipe cannot take a LOD bias on
(MOBILEGL_AVOID_EXPLICIT_LOD_BIAS=1).
Pass --coherent-as-flush for traces whose engine writes persistent
GL_MAP_FLUSH_EXPLICIT_BIT maps it never flushes (MOBILEGL_COHERENT_AS_FLUSH=1).
Pass --num-subgroups-quirk to derive compute gl_NumSubgroups instead of reading
the Vulkan builtin (MOBILEGL_NUM_SUBGROUPS_QUIRK=1).
EOF
}
@@ -108,7 +105,6 @@ use_pbuffer=0
avoid_angle_llvmpipe_sampler_mipmap_min_filter=0
avoid_angle_llvmpipe_explicit_lod_bias=0
coherent_as_flush=0
num_subgroups_quirk=0
texture_2d_dumps=""
timeout_seconds=""
@@ -150,7 +146,6 @@ while [ "$#" -gt 0 ]; do
shift 1
;;
--coherent-as-flush) coherent_as_flush=1; shift 1 ;;
--num-subgroups-quirk) num_subgroups_quirk=1; shift 1 ;;
--dump-texture-2d) texture_2d_dumps="$(next_arg "$@")"; shift 2 ;;
--timeout-seconds) timeout_seconds="$(next_arg "$@")"; shift 2 ;;
-h|--help) usage; exit 0 ;;
@@ -368,9 +363,6 @@ run_retrace() {
if [ "${coherent_as_flush}" -eq 1 ]; then
set -- "$@" --ez coherent_as_flush true
fi
if [ "${num_subgroups_quirk}" -eq 1 ]; then
set -- "$@" --ez num_subgroups_quirk true
fi
if [ -n "${texture_2d_dumps}" ]; then
set -- "$@" --es texture_2d_dumps "${texture_2d_dumps}"
fi
+1 -2
View File
@@ -284,8 +284,7 @@
"golden": "minecraft-1.21.4-fabric-iris-iterationrp-in-world.0000202020.png",
"target_call": 202020,
"timeout_seconds": 1800,
"ssim_threshold": 0.98,
"num_subgroups_quirk": true
"ssim_threshold": 0.98
},
{
"name": "minecraft-1.21.4-fabric-iris-bsl-esc-menu-854",