Compare commits

..
23 Commits
Author SHA1 Message Date
swung0x48 811f32760e [CI] (test): gate the retrace pipeline on the integration lane - build-retrace and trace-cases now need it 2026-08-12 08:41:19 -04:00
swung0x48 01381a0404 Merge branch "feat/cts-baseinstance-dsa-bufstorage" into dev 2026-08-12 08:23:57 -04:00
swung0x48 0f02b0fdb1 [Fix] (MG_Util): the POST base-instance row described an emulation that no longer exists 2026-08-12 08:22:26 -04:00
swung0x48 3be02abf47 [Fix] (MG_Impl): multisample renderbuffers ignored the integer formats own sample limit 2026-08-12 08:17:11 -04:00
swung0x48 f91d6b676c [Fix] (MG_Impl): the per-stage atomic-counter-buffer limits answered GL_INVALID_ENUM instead of their zero 2026-08-12 08:12:32 -04:00
swung0x48 1ebf191f94 [Fix, Test] (MG_Backend/DirectGLES, MG_Util): baseInstance reached the shader but never the vertex fetch 2026-08-12 08:12:27 -04:00
swung0x48 ccad803023 Merge branch "feat/cts-vab-current-attrib" into dev 2026-08-12 07:40:47 -04:00
swung0x48 21159caf31 [Fix, Test] (MG_Util, MG_Backend/DirectGLES): transform feedback captured nothing at all from an interface block 2026-08-12 07:38:50 -04:00
swung0x48 ea4819a21d [Fix, Test] (MG_State, MG_IntegrationTest): an array vertex input occupies one location per element, not one location in total 2026-08-12 07:36:12 -04:00
swung0x48 6b882b3ccf Merge branch "feat/fp64-demote-to-fp32" into dev 2026-08-12 06:42:23 -04:00
swung0x48 96bd36c50b [Docs] (README): the MOBILEGL_ADVERTISE_FP64 switch 2026-08-12 06:41:54 -04:00
swung0x48 46fbd837b3 [Fix, Test] (MG_State, MG_Impl, MG_IntegrationTest): a double uniform initializer no longer reads zero 2026-08-12 06:41:54 -04:00
swung0x48 796a57a115 [Feat, Test] (MG_Backend, MG_Util, MG_Impl): report the fp64 tier at startup, decline 64-bit vertex formats everywhere, advertise GL_ARB_gpu_shader_fp64 only on request 2026-08-12 06:25:35 -04:00
swung0x48 62a2dae5ba [Fix, Test] (MG_Impl, MG_State, MG_IntegrationTest): glUniform*d stores what the demoted shader reads 2026-08-12 06:20:11 -04:00
swung0x48 532836c058 [Feat, Test] (MG_Util): demote every 64-bit float in a shader to 32 bits, with the block layout re-derived 2026-08-12 06:11:50 -04:00
swung0x48 2fced2241b [Fix, Test] (MG_Util, MG_State, MG_Backend): the GL 4.3 vertex binding model - array vertex inputs, zero binding strides, instance divisors, and formats ES refuses 2026-08-12 05:36:00 -04:00
swung0x48 21b5fc2d92 Merge branch 'feat/cts-draw-parameters' into dev 2026-08-12 04:24:54 -04:00
swung0x48 1f753ab5fa [Test] (MG_IntegrationTest): the draw-parameter builtins, read back out of the shader the draw produced 2026-08-12 04:18:08 -04:00
swung0x48 3ed9501be5 [Fix, Feat, Test] (MG_Backend/DirectGLES, MG_Test): the draw-parameter builtins never reached the draws that carry them, and glMultiDrawArraysIndirectCount had no backend at all 2026-08-12 04:18:08 -04:00
swung0x48 7311251f30 [Fix] (MG_Util, MG_Backend/DirectVulkan): GL reads gl_BaseVertex as zero on a non-indexed draw where Vulkan's builtin hands over firstVertex 2026-08-12 04:18:07 -04:00
swung0x48 7625cf450d [Fix, Test] (MG_Impl, MG_IntegrationTest): a block count may not exceed its binding points, and an atomic-counter buffer size must survive division by them 2026-08-12 02:38:04 -04:00
swung0x48 450eb209b6 [Fix, Test] (MG_State, MG_Impl): the program interface of a separable program is its own first and last stage, not vertex and fragment 2026-08-12 02:38:04 -04:00
swung0x48 8c5c39b3c3 [Feat, Test] (MG_Impl): glBindTextures and glBindImageTextures were no-op stubs while GL_ARB_multi_bind was advertised 2026-08-12 02:38:03 -04:00
55 changed files with 5984 additions and 200 deletions
+2
View File
@@ -344,6 +344,7 @@ jobs:
- build-linux - build-linux
- test - test
- benchmark - benchmark
- integration
permissions: permissions:
actions: write actions: write
contents: read contents: read
@@ -488,6 +489,7 @@ jobs:
needs: needs:
- test - test
- benchmark - benchmark
- integration
outputs: outputs:
names: ${{ steps.trace-cases.outputs.names }} names: ${{ steps.trace-cases.outputs.names }}
steps: steps:
+4
View File
@@ -276,9 +276,13 @@ set(SOURCE_FILES
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RenameBuiltinShadowingFunctionsPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RenameBuiltinShadowingFunctionsPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecomposeWorkgroupVec3Pass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecomposeWorkgroupVec3Pass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DemoteFloat64Pass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/LowerDrawParametersPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/LowerDrawParametersPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/PackDoubleVertexInputsPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/PackDoubleVertexInputsPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FlattenXfbInterfaceBlocksPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/SplitArrayVertexInputsPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/ZeroBaseVertexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/NormalizeRectCoordinatesPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/NormalizeRectCoordinatesPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/PrivateToEntryLocalPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/PrivateToEntryLocalPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripUniformLocationsPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripUniformLocationsPass.cpp
+8
View File
@@ -82,6 +82,14 @@ namespace MobileGL::MG_Config {
#endif #endif
// MOBILEGL_DISABLE_SUBGROUP: force-disable Vulkan shader subgroup support. // MOBILEGL_DISABLE_SUBGROUP: force-disable Vulkan shader subgroup support.
Bool DisableSubgroup = false; Bool DisableSubgroup = false;
// MOBILEGL_ADVERTISE_FP64: add GL_ARB_gpu_shader_fp64 to the advertised extension
// string. `double` in a shader always WORKS - it is narrowed to 32 bits before any
// module reaches a backend (ShaderTranspiler::DemoteFloat64Pass) - but the extension
// promises 64-bit precision, and that is the one thing the narrowing cannot deliver.
// Off by default so an application that checks the string before using doubles keeps
// its float path; on for measuring what the conformance suite makes of the demoted
// precision. See the DemoteFloat64Pass header and the "fp64" POST row.
Bool AdvertiseFp64 = false;
// MOBILEGL_MAGMA_R11G11B10F_FALLBACK: use fallback format for R11G11B10F on Vulkan. // MOBILEGL_MAGMA_R11G11B10F_FALLBACK: use fallback format for R11G11B10F on Vulkan.
Bool MagmaR11G11B10FFallback = false; Bool MagmaR11G11B10FFallback = false;
// MOBILEGL_MAGMA_FRAMESINFLIGHT: requested Magma frames in flight, defaulting to 3. // MOBILEGL_MAGMA_FRAMESINFLIGHT: requested Magma frames in flight, defaulting to 3.
+1
View File
@@ -167,6 +167,7 @@ namespace MobileGL::MG_ConfigLoader {
QueryEnvVariable("MOBILEGL_TRACE_ANGLE_VARIANT", features.TraceAngleVariant, ""); QueryEnvVariable("MOBILEGL_TRACE_ANGLE_VARIANT", features.TraceAngleVariant, "");
#endif #endif
features.DisableSubgroup = QueryEnvFlag("MOBILEGL_DISABLE_SUBGROUP"); features.DisableSubgroup = QueryEnvFlag("MOBILEGL_DISABLE_SUBGROUP");
features.AdvertiseFp64 = QueryEnvFlag("MOBILEGL_ADVERTISE_FP64");
features.MagmaR11G11B10FFallback = QueryEnvFlag("MOBILEGL_MAGMA_R11G11B10F_FALLBACK"); features.MagmaR11G11B10FFallback = QueryEnvFlag("MOBILEGL_MAGMA_R11G11B10F_FALLBACK");
features.MagmaFramesInFlight = QueryEnvUint32("MOBILEGL_MAGMA_FRAMESINFLIGHT", 3, 1, 64); features.MagmaFramesInFlight = QueryEnvUint32("MOBILEGL_MAGMA_FRAMESINFLIGHT", 3, 1, 64);
features.AvoidSamplerMipmapMinFilter = features.AvoidSamplerMipmapMinFilter =
@@ -960,6 +960,15 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (MG_Util::Async::AsyncShaderCompileEnabled()) { if (MG_Util::Async::AsyncShaderCompileEnabled()) {
extensions.push_back(E_GL_KHR_parallel_shader_compile); extensions.push_back(E_GL_KHR_parallel_shader_compile);
} }
// GL_ARB_gpu_shader_fp64 is opt-in (MOBILEGL_ADVERTISE_FP64). Every `double` in a
// shader compiles and runs already - it is narrowed to 32 bits before the module
// reaches this backend - so an application that simply uses doubles needs nothing
// advertised. What the extension additionally promises is 64-bit PRECISION, which no
// mobile GPU has and the narrowing cannot fake, so advertising it by default would
// make an application that checks the string take a path MobileGL cannot honour.
if (MG_Config::Features.AdvertiseFp64) {
extensions.push_back(E_GL_ARB_gpu_shader_fp64);
}
// Only advertised when the device driver actually has usable timer queries // Only advertised when the device driver actually has usable timer queries
// (GL_EXT_disjoint_timer_query plus its entry points) and the // (GL_EXT_disjoint_timer_query plus its entry points) and the
// MOBILEGL_DISABLE_TIMERQUERY escape hatch is off. // MOBILEGL_DISABLE_TIMERQUERY escape hatch is off.
@@ -1002,6 +1011,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
funcsTable.GL.MultiDrawElementsIndirect = MultiDrawElementsIndirect; funcsTable.GL.MultiDrawElementsIndirect = MultiDrawElementsIndirect;
funcsTable.GL.MultiDrawElementsIndirectCount = MultiDrawElementsIndirectCount; funcsTable.GL.MultiDrawElementsIndirectCount = MultiDrawElementsIndirectCount;
funcsTable.GL.MultiDrawArraysIndirect = MultiDrawArraysIndirect; funcsTable.GL.MultiDrawArraysIndirect = MultiDrawArraysIndirect;
funcsTable.GL.MultiDrawArraysIndirectCount = MultiDrawArraysIndirectCount;
funcsTable.GL.DrawRangeElementsBaseVertex = DrawRangeElementsBaseVertex; funcsTable.GL.DrawRangeElementsBaseVertex = DrawRangeElementsBaseVertex;
funcsTable.GL.DrawRangeElements = DrawRangeElements; funcsTable.GL.DrawRangeElements = DrawRangeElements;
funcsTable.GL.DrawElementsInstancedBaseVertexBaseInstance = DrawElementsInstancedBaseVertexBaseInstance; funcsTable.GL.DrawElementsInstancedBaseVertexBaseInstance = DrawElementsInstancedBaseVertexBaseInstance;
@@ -2905,11 +2905,40 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
} }
void SetCurrentBaseVertex(Int32 baseVertex) {
if (const auto program = GetCurrentBackendProgram()) {
program->SetBaseVertex(baseVertex);
}
}
Bool CurrentProgramReadsDrawID() { Bool CurrentProgramReadsDrawID() {
const auto program = GetCurrentBackendProgram(); const auto program = GetCurrentBackendProgram();
return program != nullptr && program->ReadsDrawID(); return program != nullptr && program->ReadsDrawID();
} }
Bool CurrentProgramReadsBaseVertex() {
const auto program = GetCurrentBackendProgram();
return program != nullptr && program->ReadsBaseVertex();
}
// The two questions above, asked from BEFORE PrepareForDraw - where neither can be
// answered honestly. GetCurrentBackendProgram only sees a twin that a previous draw
// already synced, and a twin from before a relink still carries the previous link's
// uniform locations, so "no" there means "not known yet" at least as often as it
// means no. The multi-draw compute tier has to decide whether to flatten a batch
// before PrepareForDraw runs (its dispatch cannot come after the draw state), and
// flattening a batch that turns out to need per-sub-draw values is unrecoverable -
// so an unanswerable program counts as needing them.
Bool CurrentProgramMayNeedPerSubDrawBuiltins(Bool batchCarriesBaseVertices) {
const auto& currentProgram = MG_State::pGLContext->GetProgramForDraw();
const auto program = GetCurrentBackendProgram();
if (!currentProgram || program == nullptr ||
program->GetSyncedLinkVersion() != currentProgram->GetLinkVersion()) {
return true;
}
return program->ReadsDrawID() || (batchCarriesBaseVertices && program->ReadsBaseVertex());
}
static Bool SupportsNativeIndirectDraws() { static Bool SupportsNativeIndirectDraws() {
const auto& version = g_GLESCapabilities.GLESVersion; const auto& version = g_GLESCapabilities.GLESVersion;
const Bool esVersionOk = version.Major > 3 || (version.Major == 3 && version.Minor >= 1); const Bool esVersionOk = version.Major > 3 || (version.Major == 3 && version.Minor >= 1);
@@ -2947,16 +2976,28 @@ namespace MobileGL::MG_Backend::DirectGLES {
resource->id); resource->id);
} }
} }
// gl_BaseVertex has no SSBO view of its own: the command's baseVertex word is read
// from the CPU shadow, so a command whose baseVertex a compute shader wrote this
// frame is not observable here (baseInstance is, through the view above). Feeding
// the stale-but-usually-correct shadow beats leaving the uniform at the previous
// draw's value, which is what a program reading gl_BaseVertex saw before.
const Bool feedBaseVertex = CurrentProgramReadsBaseVertex();
for (GLsizei i = 0; i < drawcount; ++i) { for (GLsizei i = 0; i < drawcount; ++i) {
const SizeT cmdByteOffset = commandOffset + static_cast<SizeT>(i) * stride; const SizeT cmdByteOffset = commandOffset + static_cast<SizeT>(i) * stride;
SetCurrentDrawID(static_cast<Uint32>(i)); SetCurrentDrawID(static_cast<Uint32>(i));
if (paramsBinding >= 0 && backendProgram) { if (paramsBinding >= 0 && backendProgram) {
// baseInstance is the 5th word of DrawElementsIndirectCommand. // baseInstance is the 5th word of DrawElementsIndirectCommand.
backendProgram->SetBaseInstanceWordIndex(static_cast<Int32>((cmdByteOffset + 16) / 4)); backendProgram->SetBaseInstanceWordIndex(static_cast<Int32>((cmdByteOffset + 16) / 4));
if (feedBaseVertex) {
DrawElementsIndirectCommand cmd{};
std::memcpy(&cmd, commandBytes + static_cast<SizeT>(i) * stride, sizeof(cmd));
SetCurrentBaseVertex(cmd.baseVertex);
}
} else { } else {
DrawElementsIndirectCommand cmd{}; DrawElementsIndirectCommand cmd{};
std::memcpy(&cmd, commandBytes + static_cast<SizeT>(i) * stride, sizeof(cmd)); std::memcpy(&cmd, commandBytes + static_cast<SizeT>(i) * stride, sizeof(cmd));
SetCurrentBaseInstance(cmd.baseInstance); SetCurrentBaseInstance(cmd.baseInstance);
SetCurrentBaseVertex(cmd.baseVertex);
} }
g_GLESFuncs.glDrawElementsIndirect(mode, type, reinterpret_cast<const void*>(cmdByteOffset)); g_GLESFuncs.glDrawElementsIndirect(mode, type, reinterpret_cast<const void*>(cmdByteOffset));
} }
@@ -2969,6 +3010,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
SetCurrentDrawID(static_cast<Uint32>(i)); SetCurrentDrawID(static_cast<Uint32>(i));
SetCurrentBaseInstance(cmd.baseInstance); SetCurrentBaseInstance(cmd.baseInstance);
SetCurrentBaseVertex(cmd.baseVertex);
const auto indexByteOffset = static_cast<SizeT>(cmd.firstIndex) * indexSize; const auto indexByteOffset = static_cast<SizeT>(cmd.firstIndex) * indexSize;
g_GLESFuncs.glDrawElementsInstancedBaseVertex( g_GLESFuncs.glDrawElementsInstancedBaseVertex(
mode, static_cast<GLsizei>(cmd.count), type, reinterpret_cast<const GLvoid*>(indexByteOffset), mode, static_cast<GLsizei>(cmd.count), type, reinterpret_cast<const GLvoid*>(indexByteOffset),
@@ -2977,12 +3019,18 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
SetCurrentDrawID(0); SetCurrentDrawID(0);
SetCurrentBaseInstance(0); SetCurrentBaseInstance(0);
SetCurrentBaseVertex(0);
} }
static void ExecuteArraysIndirectCommands(GLenum mode, const Uint8* commandBytes, SizeT commandOffset, static void ExecuteArraysIndirectCommands(GLenum mode, const Uint8* commandBytes, SizeT commandOffset,
const SharedPtr<MG_State::GLState::BufferObject>& drawIndirectBuffer, const SharedPtr<MG_State::GLState::BufferObject>& drawIndirectBuffer,
GLsizei drawcount, GLsizei stride, const char* label) { GLsizei drawcount, GLsizei stride, const char* label) {
(void)label; (void)label;
// DrawArraysIndirectCommand has no baseVertex word, so gl_BaseVertex is zero for every
// command here. Written BEFORE the draws, not merely restored after them: the previous
// draw is what leaves a stale value, and restoring afterwards would only protect the
// NEXT draw while these commands ran with the stale one.
SetCurrentBaseVertex(0);
const Bool useNative = drawIndirectBuffer != nullptr && SupportsNativeIndirectDraws(); const Bool useNative = drawIndirectBuffer != nullptr && SupportsNativeIndirectDraws();
if (useNative) { if (useNative) {
const auto backendProgram = GetCurrentBackendProgram(); const auto backendProgram = GetCurrentBackendProgram();
@@ -3269,7 +3317,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer; DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer;
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
CheckPrimitiveRestartSupported(type); CheckPrimitiveRestartSupported(type);
SetCurrentBaseVertex(basevertex);
g_GLESFuncs.glDrawElementsBaseVertex(mode, count, type, indices, basevertex); g_GLESFuncs.glDrawElementsBaseVertex(mode, count, type, indices, basevertex);
SetCurrentBaseVertex(0);
} }
void MultiDrawArrays(GLenum mode, const GLint* first, const GLsizei* count, GLsizei drawcount) { void MultiDrawArrays(GLenum mode, const GLint* first, const GLsizei* count, GLsizei drawcount) {
@@ -3279,6 +3329,11 @@ namespace MobileGL::MG_Backend::DirectGLES {
DrawSyncFlags syncBit = DrawSyncBit::None; DrawSyncFlags syncBit = DrawSyncBit::None;
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
// This loop IS the emulation - there is no batched tier for the non-indexed form -
// so each sub-draw has to be given its own gl_DrawID here, exactly as the indexed
// ladder and the indirect executors do. Without it every sub-draw of a
// glMultiDrawArrays read draw index 0.
const Bool feedDrawID = CurrentProgramReadsDrawID();
const auto& currentVAO = MG_State::pGLContext->GetBoundVertexArray(); const auto& currentVAO = MG_State::pGLContext->GetBoundVertexArray();
for (GLsizei i = 0; i < drawcount; ++i) { for (GLsizei i = 0; i < drawcount; ++i) {
// Client-side arrays are uploaded per sub-draw range, like the single DrawArrays path. // Client-side arrays are uploaded per sub-draw range, like the single DrawArrays path.
@@ -3288,8 +3343,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
(*backendVAOSlot)->SyncClientSideAttributesForDrawArrays(currentVAO, first[i], count[i]); (*backendVAOSlot)->SyncClientSideAttributesForDrawArrays(currentVAO, first[i], count[i]);
} }
} }
if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i));
g_GLESFuncs.glDrawArrays(mode, first[i], count[i]); g_GLESFuncs.glDrawArrays(mode, first[i], count[i]);
} }
if (feedDrawID) SetCurrentDrawID(0);
} }
// Both glMultiDrawElements entry points are emulated - ES has neither in core - by the // Both glMultiDrawElements entry points are emulated - ES has neither in core - by the
@@ -3403,6 +3460,15 @@ namespace MobileGL::MG_Backend::DirectGLES {
return; return;
} }
// Both counts are read from the CPU shadow, which a buffer with no shadow does not
// have - MappedData() is null there and the reads below would be a null dereference,
// not a wrong picture. The DirectVulkan twin declines the same way.
if (parameterBuffer->MappedData() == nullptr || drawBuffer->MappedData() == nullptr) {
MGLOG_E("MultiDrawElementsIndirectCount skipped: CPU fallback cannot read the parameter or "
"draw-indirect buffer");
return;
}
Uint32 actualDrawCount = 0; Uint32 actualDrawCount = 0;
std::memcpy(&actualDrawCount, parameterBuffer->MappedData() + drawcount, sizeof(actualDrawCount)); std::memcpy(&actualDrawCount, parameterBuffer->MappedData() + drawcount, sizeof(actualDrawCount));
actualDrawCount = std::min<Uint32>(actualDrawCount, static_cast<Uint32>(maxdrawcount)); actualDrawCount = std::min<Uint32>(actualDrawCount, static_cast<Uint32>(maxdrawcount));
@@ -3444,11 +3510,79 @@ namespace MobileGL::MG_Backend::DirectGLES {
drawcount, stride, "MultiDrawArraysIndirect"); drawcount, stride, "MultiDrawArraysIndirect");
} }
// The non-indexed twin of MultiDrawElementsIndirectCount, and structurally identical to it:
// ES has no GL_PARAMETER_BUFFER at all, so the draw count is read from the CPU shadow of the
// bound one and the batch degenerates into an ordinary indirect multi-draw of that many
// commands. Missing from the backend table until now, which made every
// glMultiDrawArraysIndirectCount an INVALID_OPERATION ("backend does not support
// indirect-parameter array draws") on DirectGLES while the extension was advertised.
void MultiDrawArraysIndirectCount(GLenum mode, const void* indirect, GLintptr drawcount, GLsizei maxdrawcount,
GLsizei stride) {
#if MOBILEGL_LOG_ACTIVE_LEVEL <= MOBILEGL_LOG_LEVEL_DEBUG && MOBILEGL_ENABLE_SCOPE_MARKER
DebugImpl::OpenGLScopeMarker marker(__func__);
#endif
if (maxdrawcount <= 0) {
return;
}
if (stride == 0) {
stride = sizeof(DrawArraysIndirectCommand);
}
if (stride < static_cast<GLsizei>(sizeof(DrawArraysIndirectCommand))) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: stride %d is smaller than command size %zu",
stride, sizeof(DrawArraysIndirectCommand));
return;
}
DrawSyncFlags syncBit = DrawSyncBit::IndirectBuffer | DrawSyncBit::Instancing;
PrepareForDraw(syncBit);
auto drawBuffer = MG_State::pGLContext->GetBufferBindingSlot(BufferTarget::DrawIndirect).GetBoundObject();
auto parameterBuffer = MG_State::pGLContext->GetBufferBindingSlot(BufferTarget::Parameter).GetBoundObject();
if (!drawBuffer) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: no GL_DRAW_INDIRECT_BUFFER is bound");
return;
}
if (!parameterBuffer) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: no GL_PARAMETER_BUFFER is bound");
return;
}
drawBuffer->SyncPersistentMappedRange();
parameterBuffer->SyncPersistentMappedRange();
const SizeT commandOffset = reinterpret_cast<SizeT>(indirect);
const SizeT commandBytes = commandOffset + static_cast<SizeT>(stride) * static_cast<SizeT>(maxdrawcount - 1) +
sizeof(DrawArraysIndirectCommand);
if (commandBytes > drawBuffer->GetSize()) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: invalid GL_DRAW_INDIRECT_BUFFER binding or range");
return;
}
if (drawcount < 0 || static_cast<SizeT>(drawcount) + sizeof(Uint32) > parameterBuffer->GetSize()) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: invalid GL_PARAMETER_BUFFER binding or range");
return;
}
// See the indexed twin: no CPU shadow means no count to read, not a wrong one.
if (parameterBuffer->MappedData() == nullptr || drawBuffer->MappedData() == nullptr) {
MGLOG_E("MultiDrawArraysIndirectCount skipped: CPU fallback cannot read the parameter or "
"draw-indirect buffer");
return;
}
Uint32 actualDrawCount = 0;
std::memcpy(&actualDrawCount, parameterBuffer->MappedData() + drawcount, sizeof(actualDrawCount));
actualDrawCount = std::min<Uint32>(actualDrawCount, static_cast<Uint32>(maxdrawcount));
ExecuteArraysIndirectCommands(mode, drawBuffer->MappedData() + commandOffset, commandOffset, drawBuffer,
static_cast<GLsizei>(actualDrawCount), stride, "MultiDrawArraysIndirectCount");
}
void DrawRangeElementsBaseVertex(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, void DrawRangeElementsBaseVertex(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type,
const void* indices, GLint basevertex) { const void* indices, GLint basevertex) {
DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer; DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer;
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
SetCurrentBaseVertex(basevertex);
g_GLESFuncs.glDrawRangeElementsBaseVertex(mode, start, end, count, type, indices, basevertex); g_GLESFuncs.glDrawRangeElementsBaseVertex(mode, start, end, count, type, indices, basevertex);
SetCurrentBaseVertex(0);
} }
void DrawRangeElements(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, const void* indices) { void DrawRangeElements(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, const void* indices) {
@@ -3457,12 +3591,33 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_GLESFuncs.glDrawRangeElements(mode, start, end, count, type, indices); g_GLESFuncs.glDrawRangeElements(mode, start, end, count, type, indices);
} }
// True when the driver will apply baseInstance to the vertex fetch itself, in which case the
// attribute-offset emulation must stay out of the way. SetCurrentBaseInstance is orthogonal
// and runs either way - it feeds the shader's gl_BaseInstance, not the fetch.
inline Bool UseNativeBaseInstance() {
return g_GLESCapabilities.SupportsBaseInstance;
}
// The emulated shift has to be in place before PrepareForDraw, because that is what syncs the
// VAO; a zero here is what un-shifts the arrays for the next ordinary draw.
inline Uint32 EmulatedFetchBaseInstance(GLuint baseinstance) {
return UseNativeBaseInstance() ? 0u : static_cast<Uint32>(baseinstance);
}
void DrawElementsInstancedBaseVertexBaseInstance(GLenum mode, GLsizei count, GLenum type, const void* indices, void DrawElementsInstancedBaseVertexBaseInstance(GLenum mode, GLsizei count, GLenum type, const void* indices,
GLsizei instancecount, GLint basevertex, GLuint baseinstance) { GLsizei instancecount, GLint basevertex, GLuint baseinstance) {
DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing; DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing;
const VertexArrayImpl::ScopedFetchBaseInstance fetchScope(EmulatedFetchBaseInstance(baseinstance));
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
SetCurrentBaseInstance(baseinstance); SetCurrentBaseInstance(baseinstance);
SetCurrentBaseVertex(basevertex);
if (UseNativeBaseInstance()) {
g_GLESFuncs.glDrawElementsInstancedBaseVertexBaseInstanceEXT(mode, count, type, indices, instancecount,
basevertex, baseinstance);
} else {
g_GLESFuncs.glDrawElementsInstancedBaseVertex(mode, count, type, indices, instancecount, basevertex); g_GLESFuncs.glDrawElementsInstancedBaseVertex(mode, count, type, indices, instancecount, basevertex);
}
SetCurrentBaseVertex(0);
SetCurrentBaseInstance(0); SetCurrentBaseInstance(0);
} }
@@ -3470,15 +3625,23 @@ namespace MobileGL::MG_Backend::DirectGLES {
GLsizei instancecount, GLint basevertex) { GLsizei instancecount, GLint basevertex) {
DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing; DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing;
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
SetCurrentBaseVertex(basevertex);
g_GLESFuncs.glDrawElementsInstancedBaseVertex(mode, count, type, indices, instancecount, basevertex); g_GLESFuncs.glDrawElementsInstancedBaseVertex(mode, count, type, indices, instancecount, basevertex);
SetCurrentBaseVertex(0);
} }
void DrawElementsInstancedBaseInstance(GLenum mode, GLsizei count, GLenum type, const void* indices, void DrawElementsInstancedBaseInstance(GLenum mode, GLsizei count, GLenum type, const void* indices,
GLsizei instancecount, GLuint baseinstance) { GLsizei instancecount, GLuint baseinstance) {
DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing; DrawSyncFlags syncBit = DrawSyncBit::IndexBuffer | DrawSyncBit::Instancing;
const VertexArrayImpl::ScopedFetchBaseInstance fetchScope(EmulatedFetchBaseInstance(baseinstance));
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
SetCurrentBaseInstance(baseinstance); SetCurrentBaseInstance(baseinstance);
if (UseNativeBaseInstance()) {
g_GLESFuncs.glDrawElementsInstancedBaseInstanceEXT(mode, count, type, indices, instancecount,
baseinstance);
} else {
g_GLESFuncs.glDrawElementsInstanced(mode, count, type, indices, instancecount); g_GLESFuncs.glDrawElementsInstanced(mode, count, type, indices, instancecount);
}
SetCurrentBaseInstance(0); SetCurrentBaseInstance(0);
} }
@@ -3514,9 +3677,14 @@ namespace MobileGL::MG_Backend::DirectGLES {
void DrawArraysInstancedBaseInstance(GLenum mode, GLint first, GLsizei count, GLsizei instancecount, void DrawArraysInstancedBaseInstance(GLenum mode, GLint first, GLsizei count, GLsizei instancecount,
GLuint baseinstance) { GLuint baseinstance) {
DrawSyncFlags syncBit = DrawSyncBit::Instancing; DrawSyncFlags syncBit = DrawSyncBit::Instancing;
const VertexArrayImpl::ScopedFetchBaseInstance fetchScope(EmulatedFetchBaseInstance(baseinstance));
PrepareForDraw(syncBit); PrepareForDraw(syncBit);
SetCurrentBaseInstance(baseinstance); SetCurrentBaseInstance(baseinstance);
if (UseNativeBaseInstance()) {
g_GLESFuncs.glDrawArraysInstancedBaseInstanceEXT(mode, first, count, instancecount, baseinstance);
} else {
g_GLESFuncs.glDrawArraysInstanced(mode, first, count, instancecount); g_GLESFuncs.glDrawArraysInstanced(mode, first, count, instancecount);
}
SetCurrentBaseInstance(0); SetCurrentBaseInstance(0);
} }
@@ -40,6 +40,8 @@ namespace MobileGL::MG_Backend::DirectGLES {
void MultiDrawElementsIndirectCount(GLenum mode, GLenum type, const void* indirect, GLintptr drawcount, void MultiDrawElementsIndirectCount(GLenum mode, GLenum type, const void* indirect, GLintptr drawcount,
GLsizei maxdrawcount, GLsizei stride); GLsizei maxdrawcount, GLsizei stride);
void MultiDrawArraysIndirect(GLenum mode, const void* indirect, GLsizei drawcount, GLsizei stride); void MultiDrawArraysIndirect(GLenum mode, const void* indirect, GLsizei drawcount, GLsizei stride);
void MultiDrawArraysIndirectCount(GLenum mode, const void* indirect, GLintptr drawcount, GLsizei maxdrawcount,
GLsizei stride);
void DrawRangeElementsBaseVertex(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, void DrawRangeElementsBaseVertex(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type,
const void* indices, GLint basevertex); const void* indices, GLint basevertex);
void DrawRangeElements(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, const void* indices); void DrawRangeElements(GLenum mode, GLuint start, GLuint end, GLsizei count, GLenum type, const void* indices);
+264 -25
View File
@@ -155,6 +155,14 @@ namespace MobileGL::MG_Backend::DirectGLES {
// (possibly GPU-written) indirect command buffer, so its declaration expands into a // (possibly GPU-written) indirect command buffer, so its declaration expands into a
// std430 SSBO view of that buffer indexed by a CPU-computed word index, with the plain // std430 SSBO view of that buffer indexed by a CPU-computed word index, with the plain
// mg_BaseInstance uniform as the fallback for non-indirect draws. // mg_BaseInstance uniform as the fallback for non-indirect draws.
//
// The word index is stored ONE-BASED, so that zero - the value every GLSL uniform starts
// at - is the "not an indirect draw" sentinel. Nothing seeds this uniform before a
// program's first draw, and the non-indirect draw entry points never write it at all, so a
// zero-based index with a negative sentinel would leave every such draw reading
// mg_indirectWords[0] out of a storage buffer no one bound. That is not a silent zero on a
// real driver: it returned garbage on Adreno, and a garbage gl_BaseInstance pushed the CTS
// shader_draw_parameters geometry clean off screen.
String PromoteDrawParameterGlobalsToUniforms(String source, GLenum shaderType) { String PromoteDrawParameterGlobalsToUniforms(String source, GLenum shaderType) {
if (shaderType != GL_VERTEX_SHADER) { if (shaderType != GL_VERTEX_SHADER) {
return source; return source;
@@ -208,12 +216,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
" { highp uint mg_indirectWords[]; };\n"; " { highp uint mg_indirectWords[]; };\n";
if (rebaseInstanceId) { if (rebaseInstanceId) {
machinery += String("#define ") + ZERO_BASED_INSTANCE_ID_NAME + " (gl_InstanceID - ((" + machinery += String("#define ") + ZERO_BASED_INSTANCE_ID_NAME + " (gl_InstanceID - ((" +
BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " >= 0) ? int(mg_indirectWords[uint(" + BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " > 0) ? int(mg_indirectWords[uint(" +
BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + ")]) : 0))\n"; BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " - 1)]) : 0))\n";
} }
machinery += String("#define ") + BASE_INSTANCE_LOWERED_NAME + " ((" + machinery += String("#define ") + BASE_INSTANCE_LOWERED_NAME + " ((" +
BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " >= 0) ? int(mg_indirectWords[uint(" + BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " > 0) ? int(mg_indirectWords[uint(" +
BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + ")]) : " + BASE_INSTANCE_UNIFORM_NAME + ")"; BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME + " - 1)]) : " + BASE_INSTANCE_UNIFORM_NAME + ")";
source.replace(pos, declaration.size(), machinery); source.replace(pos, declaration.size(), machinery);
break; break;
} }
@@ -1473,6 +1481,78 @@ namespace MobileGL::MG_Backend::DirectGLES {
return true; return true;
} }
// ES 3.1 core. Queried through the loader rather than the version, because the whole
// point of using it is to express something the pointer API cannot, and falling back
// silently on a driver that lacks it is better than crashing on a null entry point.
inline Bool HasVertexBindingApi() {
return g_GLESFuncs.glBindVertexBuffer != nullptr && g_GLESFuncs.glVertexAttribFormat != nullptr &&
g_GLESFuncs.glVertexAttribIFormat != nullptr && g_GLESFuncs.glVertexAttribBinding != nullptr &&
g_GLESFuncs.glVertexBindingDivisor != nullptr;
}
// Draw state, not VAO state: set by the baseInstance draw entry points around
// PrepareForDraw and back to zero as soon as the draw is issued.
Uint32 g_pendingFetchBaseInstance = 0;
void SetPendingFetchBaseInstance(Uint32 baseInstance) {
g_pendingFetchBaseInstance = baseInstance;
}
Uint32 GetPendingFetchBaseInstance() {
return g_pendingFetchBaseInstance;
}
// The "+ baseInstance" of GL's instanced-array element index, expressed as a byte shift
// of the array's own offset. Only divisor'd arrays step per instance, so only they move.
//
// baseInstance is added to the ELEMENT index, not to instance/divisor - the divisor
// therefore does not appear here, and the shift is a whole number of strides.
//
// A resolved stride of zero is the binding model's "never advance" (see
// VertexAttribute::Stride), so such an array reads the same element for every instance
// and a baseInstance cannot move it. The arithmetic already yields zero for that case.
inline SizeT BaseInstanceByteShift(const MG_State::GLState::VertexAttribute& attrib, Uint32 baseInstance) {
if (baseInstance == 0 || attrib.Divisor == 0) {
return 0;
}
return static_cast<SizeT>(baseInstance) * static_cast<SizeT>(attrib.Stride);
}
// Declares one attribute through the ES binding-point API, the only spelling that can
// carry a stride of zero. Returns false when the attribute has no usable buffer, in
// which case nothing was emitted.
inline Bool SyncZeroStrideAttribute(Uint attribIndex, const MG_State::GLState::VertexAttribute& attrib) {
const auto& bufferObject = attrib.Buffer;
if (!bufferObject) {
MGLOG_W("Zero-stride attribute %u has no bound buffer, skipping.", attribIndex);
return false;
}
auto* backendResource = BufferImpl::EnsureBufferResource(bufferObject);
if (!backendResource || backendResource->id == 0) {
MGLOG_E("No backend buffer for zero-stride attribute %u, cannot bind it.", attribIndex);
return false;
}
if (!attrib.IsInteger) {
const GLint glSize = attrib.IsBgra ? static_cast<GLint>(GL_BGRA) : attrib.Size;
g_GLESFuncs.glVertexAttribFormat(attribIndex, glSize,
MG_Util::ConvertDataTypeToGLEnum(attrib.Type),
attrib.Normalized ? GL_TRUE : GL_FALSE, 0);
} else {
g_GLESFuncs.glVertexAttribIFormat(attribIndex, attrib.Size,
MG_Util::ConvertDataTypeToGLEnum(attrib.Type), 0);
}
g_GLESFuncs.glVertexAttribBinding(attribIndex, attribIndex);
// The resolved offset goes on the binding point, not into a relative offset: the
// relative offset is capped by GL_MAX_VERTEX_ATTRIB_RELATIVE_OFFSET (2047 at
// minimum) while a buffer offset is not, so anything else would break on a large
// one. BindBufferId is bypassed deliberately - glBindVertexBuffer binds into the
// VAO's binding point, not the GL_ARRAY_BUFFER target that cache tracks.
g_GLESFuncs.glBindVertexBuffer(attribIndex, backendResource->id,
static_cast<GLintptr>(attrib.Offset), 0);
return true;
}
void BackendVertexArrayObject::SyncToBackend( void BackendVertexArrayObject::SyncToBackend(
const SharedPtr<MG_State::GLState::VertexArrayObject>& stateVAOObject) { const SharedPtr<MG_State::GLState::VertexArrayObject>& stateVAOObject) {
#ifdef TRACY_ENABLE #ifdef TRACY_ENABLE
@@ -1495,7 +1575,16 @@ namespace MobileGL::MG_Backend::DirectGLES {
const Uint16 currentIndexBufferVersion = stateVAOObject->GetIndexBufferBindingSlot().GetVersion(); const Uint16 currentIndexBufferVersion = stateVAOObject->GetIndexBufferBindingSlot().GetVersion();
const Bool attributesDirty = !m_hasSyncedConfigVersion || m_syncedConfigVersion != currentConfigVersion; const Bool attributesDirty = !m_hasSyncedConfigVersion || m_syncedConfigVersion != currentConfigVersion;
const Bool indexBufferDirty = currentIndexBufferVersion != m_syncedIndexBufferVersion; const Bool indexBufferDirty = currentIndexBufferVersion != m_syncedIndexBufferVersion;
if (!attributesDirty && !indexBufferDirty) {
// The baseInstance shift lives in the attribute offsets the driver already holds, so
// a change of baseInstance has to re-emit the divisor'd arrays even when the frontend
// config version says nothing moved - and equally has to un-shift them for the next
// draw that carries no baseInstance. Resting state is 0 on both sides, so a program
// that never calls a *BaseInstance entry point never pays for this compare.
const Uint32 fetchBaseInstance = g_pendingFetchBaseInstance;
const Bool baseInstanceDirty = m_syncedFetchBaseInstance != fetchBaseInstance;
const Bool emitAttributes = attributesDirty || baseInstanceDirty;
if (!emitAttributes && !indexBufferDirty) {
return; return;
} }
@@ -1503,8 +1592,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& allAttributeVersions = stateVAOObject->GetAllAttributeVersions(); const auto& allAttributeVersions = stateVAOObject->GetAllAttributeVersions();
const auto& allAttributes = stateVAOObject->GetAllAttributes(); const auto& allAttributes = stateVAOObject->GetAllAttributes();
for (Uint attribIndex = 0; attribIndex < allAttributes.size() && attributesDirty; ++attribIndex) { for (Uint attribIndex = 0; attribIndex < allAttributes.size() && emitAttributes; ++attribIndex) {
const auto& attrib = allAttributes[attribIndex]; const auto& attrib = allAttributes[attribIndex];
// Only the divisor'd arrays carry the shift, and only an enabled one is worth
// re-emitting - a disabled array has no pointer the draw could fetch through,
// and may well have no buffer to bind either.
const Bool needsSyncBaseInstance = baseInstanceDirty && attrib.Enabled && attrib.Divisor != 0;
Bool needsSyncSwitch = allAttributeVersions[attribIndex].SwitchVersion != Bool needsSyncSwitch = allAttributeVersions[attribIndex].SwitchVersion !=
m_syncedAttributeVersions[attribIndex].SwitchVersion; m_syncedAttributeVersions[attribIndex].SwitchVersion;
if (needsSyncSwitch) { if (needsSyncSwitch) {
@@ -1519,7 +1612,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_syncedAttributeVersions[attribIndex].FormatVersion; m_syncedAttributeVersions[attribIndex].FormatVersion;
Bool needsSyncBuffer = allAttributeVersions[attribIndex].BufferVersion != Bool needsSyncBuffer = allAttributeVersions[attribIndex].BufferVersion !=
m_syncedAttributeVersions[attribIndex].BufferVersion; m_syncedAttributeVersions[attribIndex].BufferVersion;
if (!needsSyncFormat && !needsSyncBuffer) continue; if (!needsSyncFormat && !needsSyncBuffer && !needsSyncBaseInstance) continue;
// Defence in depth. The frontend already declines glVertexAttribLFormat on this // Defence in depth. The frontend already declines glVertexAttribLFormat on this
// backend (SupportsFloat64VertexAttributes is false - ES has no GL_DOUBLE vertex // backend (SupportsFloat64VertexAttributes is false - ES has no GL_DOUBLE vertex
@@ -1529,28 +1622,90 @@ namespace MobileGL::MG_Backend::DirectGLES {
// FormatVersion, not SwitchVersion, so the enable/disable block above will not run // FormatVersion, not SwitchVersion, so the enable/disable block above will not run
// again and an already-enabled array would stay enabled with no pointer and no // again and an already-enabled array would stay enabled with no pointer and no
// ARRAY_BUFFER binding - which ES 3.1+ makes an INVALID_OPERATION at draw. // ARRAY_BUFFER binding - which ES 3.1+ makes an INVALID_OPERATION at draw.
if (attrib.IsLong) { //
MGLOG_E("DirectGLES: vertex attribute %u is a 64-bit (GL_DOUBLE) array, which this " // IsLong is not the only way a 64-bit array gets here: glVertexAttribFormat
// with GL_DOUBLE asks for doubles in memory CONVERTED to float, so it is not
// long, is not declined by the frontend, and still has no ES vertex format.
// Leaving that one enabled did not merely raise INVALID_ENUM - the Adreno
// driver dereferenced null inside the next draw and took the process with it
// (SIGSEGV in libGLESv2_adreno, KHR-GL43.vertex_attrib_binding.basic-input-case4),
// because the array stayed enabled with no pointer the failed call could set.
// The type test therefore covers the storage, not the spelling.
if (attrib.IsLong || attrib.Type == DataType::Float64) {
MGLOG_I("DirectGLES: vertex attribute %u is a 64-bit (GL_DOUBLE) array, which this "
"backend cannot feed - disabling the array", "backend cannot feed - disabling the array",
attribIndex); attribIndex);
g_GLESFuncs.glDisableVertexAttribArray(attribIndex); g_GLESFuncs.glDisableVertexAttribArray(attribIndex);
continue; continue;
} }
// A resolved stride of zero is the binding model's "never advance" (see
// VertexAttribute::Stride) and glVertexAttribPointer cannot say it - a zero
// stride argument there means "tightly packed" instead, i.e. exactly the
// opposite. ES 3.1's binding-point API can, so a zero-stride attribute takes
// that spelling: its own binding point (index == attribute index, the default
// mapping) carrying the buffer, the whole resolved offset and stride 0, with
// the format at relative offset 0. Everything the pointer call would have set
// for this attribute is set here too, so the two spellings stay interchangeable
// from one sync to the next.
if (attrib.Stride == 0 && HasVertexBindingApi()) {
if (!SyncZeroStrideAttribute(attribIndex, attrib)) {
continue;
}
// No BaseInstanceByteShift here on purpose: a zero stride never advances, so
// the shift is zero by construction and adding it would only obscure that.
if (needsSyncFormat) {
g_GLESFuncs.glVertexBindingDivisor(attribIndex, attrib.Divisor);
}
continue;
}
if (!BindAttributeBuffer(attrib)) { if (!BindAttributeBuffer(attrib)) {
continue; continue;
} }
// GL_BGRA as a vertex SIZE is desktop-only; ES has no equivalent and rejects
// it. That rejection is not benign: it leaves the array ENABLED with no
// pointer, and the Adreno driver then dereferences null inside the next draw
// and kills the process rather than reporting an error (SIGSEGV in
// libGLESv2_adreno, KHR-GL43.vertex_attrib_binding.basic-input-case5). So the
// refusal has to be observed and the array disabled.
//
// Deliberately ONLY this format. Everything else MobileGL can reach here is ES
// core - the packed 2_10_10_10 pair included, whose size the frontend has
// already pinned to the 4 that ES requires - so nothing else can be refused,
// and the per-draw sync must not grow a glGetError round trip (a driver
// pipeline stall) for the formats real applications actually use. BGRA is also
// still ATTEMPTED rather than refused up front: some ES drivers do accept it,
// and the ones that do should keep working.
const Bool formatMayBeRefused = attrib.IsBgra;
if (formatMayBeRefused) {
while (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
} // start from a clean slate so the check below is about THIS call
}
const SizeT fetchOffset = attrib.Offset + BaseInstanceByteShift(attrib, fetchBaseInstance);
if (!attrib.IsInteger) { if (!attrib.IsInteger) {
// GL_BGRA is passed to the driver as the size argument (the driver reorders BGRA). // GL_BGRA is passed to the driver as the size argument (the driver reorders BGRA).
const GLint glSize = attrib.IsBgra ? static_cast<GLint>(GL_BGRA) : attrib.Size; const GLint glSize = attrib.IsBgra ? static_cast<GLint>(GL_BGRA) : attrib.Size;
g_GLESFuncs.glVertexAttribPointer( g_GLESFuncs.glVertexAttribPointer(
attribIndex, glSize, MG_Util::ConvertDataTypeToGLEnum(attrib.Type), attribIndex, glSize, MG_Util::ConvertDataTypeToGLEnum(attrib.Type),
attrib.Normalized ? GL_TRUE : GL_FALSE, attrib.Stride, (const void*)attrib.Offset); attrib.Normalized ? GL_TRUE : GL_FALSE, attrib.Stride, (const void*)fetchOffset);
} else { } else {
g_GLESFuncs.glVertexAttribIPointer(attribIndex, attrib.Size, g_GLESFuncs.glVertexAttribIPointer(attribIndex, attrib.Size,
MG_Util::ConvertDataTypeToGLEnum(attrib.Type), attrib.Stride, MG_Util::ConvertDataTypeToGLEnum(attrib.Type), attrib.Stride,
(const void*)attrib.Offset); (const void*)fetchOffset);
}
if (formatMayBeRefused && g_GLESFuncs.glGetError() != GL_NO_ERROR) {
MGLOG_I("DirectGLES: the driver refused the vertex format of attribute %u "
"(size=%d bgra=%d type=%s) - disabling the array so the draw cannot "
"fetch through a pointer the driver never accepted",
attribIndex, attrib.Size, attrib.IsBgra ? 1 : 0,
MG_Util::ConvertGLEnumToString(MG_Util::ConvertDataTypeToGLEnum(attrib.Type)).c_str());
g_GLESFuncs.glDisableVertexAttribArray(attribIndex);
continue;
} }
if (needsSyncFormat) { if (needsSyncFormat) {
@@ -1584,6 +1739,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_syncedConfigVersion = currentConfigVersion; m_syncedConfigVersion = currentConfigVersion;
m_hasSyncedConfigVersion = true; m_hasSyncedConfigVersion = true;
} }
if (emitAttributes) {
m_syncedFetchBaseInstance = fetchBaseInstance;
}
} }
void BackendVertexArrayObject::SyncClientSideAttributesForDrawArrays( void BackendVertexArrayObject::SyncClientSideAttributesForDrawArrays(
@@ -1601,9 +1759,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
continue; continue;
} }
// Same reason as SyncToBackend: there is no ES vertex format for a 64-bit array, and // Same reason as SyncToBackend, including why the test is on the storage rather
// this path only ever reaches glVertexAttribPointer/IPointer. // than on IsLong: there is no ES vertex format for a 64-bit array, and this path
if (attrib.IsLong) { // only ever reaches glVertexAttribPointer/IPointer.
if (attrib.IsLong || attrib.Type == DataType::Float64) {
g_GLESFuncs.glDisableVertexAttribArray(attribIndex); g_GLESFuncs.glDisableVertexAttribArray(attribIndex);
continue; continue;
} }
@@ -4316,6 +4475,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
auto& shaderSpirvs = stateProgramObject->GetGeneratedSpirv(); auto& shaderSpirvs = stateProgramObject->GetGeneratedSpirv();
// Blocks a transform-feedback capture request names a member of ("StageData" of
// "StageData.attrib[0]"). The Adreno ES driver accepts such a request, links, and
// then captures nothing at all for it, so those blocks - and ONLY those - get
// flattened into per-member variables below, in EVERY stage, so a producer and its
// consumer keep matching. gl_PerVertex members ("gl_Position") carry no block
// prefix and so never enter this set.
std::set<String> xfbCaptureBlockNames;
for (const auto& xfbVarying : stateProgramObject->GetTransformFeedbackVaryings()) {
const SizeT dot = xfbVarying.name.find('.');
if (dot != String::npos && dot > 0) {
xfbCaptureBlockNames.insert(xfbVarying.name.substr(0, dot));
}
}
std::set<String> flattenedXfbBlockNames;
for (int index = 0; index < attachedShaders.size(); ++index) { for (int index = 0; index < attachedShaders.size(); ++index) {
auto& shader = attachedShaders[index]; auto& shader = attachedShaders[index];
GLenum glShaderType = MG_Util::ConvertShaderStageToGLEnum(shader->GetShaderStage()); GLenum glShaderType = MG_Util::ConvertShaderStageToGLEnum(shader->GetShaderStage());
@@ -4358,6 +4532,40 @@ namespace MobileGL::MG_Backend::DirectGLES {
effectiveSpirv = &loweredSpirv; effectiveSpirv = &loweredSpirv;
} }
// GLSL ES has no ARRAY vertex inputs, and SPIRV-Cross refuses the whole module
// rather than emulating them, so this has to happen before it sees the binary.
Vector<unsigned int> splitArrayInputSpirv;
if (glShaderType == GL_VERTEX_SHADER &&
MG_Util::ShaderTranspiler::ShaderCompiler::SplitArrayVertexInputsForEssl(
*effectiveSpirv, splitArrayInputSpirv) &&
!splitArrayInputSpirv.empty() && splitArrayInputSpirv != *effectiveSpirv) {
// Only when the pass ACTUALLY split something. The optimizer hands back a
// re-serialised copy either way, and adopting that copy for every vertex
// shader would put every one of them through a round trip they do not need
// - which is not free: it cost the create-indirect retrace 0.15 SSIM the
// first time this gate was missing.
effectiveSpirv = &splitArrayInputSpirv;
}
// Adopt the rewritten module only when THIS stage actually had one of the
// blocks - the optimizer hands back a re-serialised copy either way, and taking
// that copy for a module it did not rewrite is not free (it cost the
// create-indirect retrace 0.15 SSIM when the array-input split first missed
// this gate). The report has to be per stage, not cumulative: a fragment shader
// consuming the same block reports a name the vertex stage already reported,
// and its own rewrite must still be taken or the two stages stop matching.
Vector<unsigned int> flattenedXfbSpirv;
std::set<String> stageFlattenedXfbBlockNames;
if (!xfbCaptureBlockNames.empty() &&
MG_Util::ShaderTranspiler::ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(
*effectiveSpirv, xfbCaptureBlockNames, stageFlattenedXfbBlockNames,
flattenedXfbSpirv) &&
!flattenedXfbSpirv.empty() && !stageFlattenedXfbBlockNames.empty()) {
effectiveSpirv = &flattenedXfbSpirv;
flattenedXfbBlockNames.insert(stageFlattenedXfbBlockNames.begin(),
stageFlattenedXfbBlockNames.end());
}
// ESSL stage-matches uniform blocks by member precision, but SPIRV-Cross prints // ESSL stage-matches uniform blocks by member precision, but SPIRV-Cross prints
// a RelaxedPrecision member as explicit "mediump" in the vertex stage and as // a RelaxedPrecision member as explicit "mediump" in the vertex stage and as
// UNQUALIFIED (mediump-by-default) in the fragment stage; after // UNQUALIFIED (mediump-by-default) in the fragment stage; after
@@ -4439,11 +4647,18 @@ namespace MobileGL::MG_Backend::DirectGLES {
spvcSession.Compile(&result); spvcSession.Compile(&result);
if (!result) { if (!result) {
MG_Util::ShaderTranspiler::ResultInfo r; // MGLOG_I, for the same reason as the compile- and link-failure diagnostics
r.log += "Failed to compile the shader to GLSL: \n"; // below: every CI, retrace and release build compiles at
r.log += spvcSession.GetLastErrorString(); // MOBILEGL_LOG_LEVEL_INFO, where MGLOG_E expands to nothing. A stage that
r.errc = -5; // never reaches the driver leaves the program short of that stage, so the
MGLOG_E("%s", r.log.c_str()); // link fails with an EMPTY driver info log - the least debuggable failure
// MobileGL can produce, and what hid the whole
// KHR-GL43.vertex_attrib_binding family behind "the draw captured zeros".
MGLOG_I("Shader transpilation to ESSL failed. State program ID: %u, stage: %s, "
"SPIRV-Cross error: %s",
stateProgramObject->GetExternalIndex(),
MG_Util::ConvertGLEnumToString(glShaderType).c_str(),
spvcSession.GetLastErrorString());
m_backendProgramUsable = false; m_backendProgramUsable = false;
continue; continue;
} }
@@ -4559,8 +4774,23 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& xfbVaryings = stateProgramObject->GetTransformFeedbackVaryings(); const auto& xfbVaryings = stateProgramObject->GetTransformFeedbackVaryings();
Vector<const GLchar*> xfbNames; Vector<const GLchar*> xfbNames;
xfbNames.reserve(xfbVaryings.size()); xfbNames.reserve(xfbVaryings.size());
for (const auto& xfbVarying : xfbVaryings) { // A block this build flattened no longer HAS the member the application asked
xfbNames.push_back(xfbVarying.name.c_str()); // for; it has the variable that replaced it. Everything else - including a
// member of a block that was left alone - keeps the application's spelling.
// Storage first, pointers after: xfbNames holds pointers into these strings.
Vector<String> rewrittenXfbNames(xfbVaryings.size());
for (SizeT nameIndex = 0; nameIndex < xfbVaryings.size(); ++nameIndex) {
String flatName;
if (!flattenedXfbBlockNames.empty() &&
MG_Util::ShaderTranspiler::ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock(
xfbVaryings[nameIndex].name, flattenedXfbBlockNames, flatName)) {
rewrittenXfbNames[nameIndex] = std::move(flatName);
} else {
rewrittenXfbNames[nameIndex] = xfbVaryings[nameIndex].name;
}
}
for (const auto& xfbName : rewrittenXfbNames) {
xfbNames.push_back(xfbName.c_str());
} }
MGLOG_D("Declaring %zu transform feedback varyings on program %u", xfbNames.size(), MGLOG_D("Declaring %zu transform feedback varyings on program %u", xfbNames.size(),
m_backendProgramId); m_backendProgramId);
@@ -4594,6 +4824,8 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_baseInstanceUniformLocation = g_GLESFuncs.glGetUniformLocation(m_backendProgramId, m_baseInstanceUniformLocation = g_GLESFuncs.glGetUniformLocation(m_backendProgramId,
BASE_INSTANCE_UNIFORM_NAME); BASE_INSTANCE_UNIFORM_NAME);
m_drawIdUniformLocation = g_GLESFuncs.glGetUniformLocation(m_backendProgramId, DRAW_ID_UNIFORM_NAME); m_drawIdUniformLocation = g_GLESFuncs.glGetUniformLocation(m_backendProgramId, DRAW_ID_UNIFORM_NAME);
m_baseVertexUniformLocation = g_GLESFuncs.glGetUniformLocation(m_backendProgramId,
BASE_VERTEX_UNIFORM_NAME);
m_baseInstanceWordIndexUniformLocation = m_baseInstanceWordIndexUniformLocation =
g_GLESFuncs.glGetUniformLocation(m_backendProgramId, BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME); g_GLESFuncs.glGetUniformLocation(m_backendProgramId, BASE_INSTANCE_WORD_INDEX_UNIFORM_NAME);
// The mg_IndirectParams block binding is baked into the ESSL (ES cannot rebind // The mg_IndirectParams block binding is baked into the ESSL (ES cannot rebind
@@ -4772,14 +5004,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_GLESFuncs.glUniform1i(m_baseInstanceUniformLocation, static_cast<GLint>(baseInstance)); g_GLESFuncs.glUniform1i(m_baseInstanceUniformLocation, static_cast<GLint>(baseInstance));
} }
// A direct value disables the indirect-command-buffer read. // A direct value disables the indirect-command-buffer read.
SetBaseInstanceWordIndex(-1);
}
// The uniform is written one-based so that its GLSL initial value, zero, already reads
// as "no indirect command" - see PromoteDrawParameterGlobalsToUniforms.
void BackendProgramObjectImpl::SetBaseInstanceWordIndex(Int32 wordIndex) const {
if (m_baseInstanceWordIndexUniformLocation >= 0) { if (m_baseInstanceWordIndexUniformLocation >= 0) {
g_GLESFuncs.glUniform1i(m_baseInstanceWordIndexUniformLocation, -1); g_GLESFuncs.glUniform1i(m_baseInstanceWordIndexUniformLocation,
wordIndex < 0 ? 0 : wordIndex + 1);
} }
} }
void BackendProgramObjectImpl::SetBaseInstanceWordIndex(Int32 wordIndex) const { void BackendProgramObjectImpl::SetBaseVertex(Int32 baseVertex) const {
if (m_baseInstanceWordIndexUniformLocation >= 0) { if (m_baseVertexUniformLocation >= 0) {
g_GLESFuncs.glUniform1i(m_baseInstanceWordIndexUniformLocation, wordIndex); g_GLESFuncs.glUniform1i(m_baseVertexUniformLocation, baseVertex);
} }
} }
+42 -2
View File
@@ -82,14 +82,26 @@ namespace MobileGL::MG_Backend::DirectGLES {
// GLES core supports only GL_PRIMITIVE_RESTART_FIXED_INDEX. Throws when the app enabled // GLES core supports only GL_PRIMITIVE_RESTART_FIXED_INDEX. Throws when the app enabled
// the arbitrary GL_PRIMITIVE_RESTART with a non-fixed index for this index type. // the arbitrary GL_PRIMITIVE_RESTART with a non-fixed index for this index type.
void CheckPrimitiveRestartSupported(GLenum indexType); void CheckPrimitiveRestartSupported(GLenum indexType);
// Feed the current program's gl_BaseInstance / gl_DrawID emulation uniforms. Both are // Feed the current program's gl_BaseInstance / gl_DrawID / gl_BaseVertex emulation
// no-ops when the program does not read the corresponding builtin. // uniforms. All are no-ops when the program does not read the corresponding builtin.
void SetCurrentBaseInstance(Uint32 baseInstance); void SetCurrentBaseInstance(Uint32 baseInstance);
void SetCurrentDrawID(Uint32 drawId); void SetCurrentDrawID(Uint32 drawId);
// GL's gl_BaseVertex is the base-vertex parameter of an indexed draw and zero for every
// command that has none - including all the DrawArrays forms - so every draw path that
// does not carry one must leave this at zero rather than inherit the last draw's value.
void SetCurrentBaseVertex(Int32 baseVertex);
// True when the current program actually reads gl_DrawID, i.e. when a batched // True when the current program actually reads gl_DrawID, i.e. when a batched
// (single driver call) multi-draw tier would have to feed it one value for the whole // (single driver call) multi-draw tier would have to feed it one value for the whole
// batch and would therefore be wrong. // batch and would therefore be wrong.
Bool CurrentProgramReadsDrawID(); Bool CurrentProgramReadsDrawID();
// Same question for gl_BaseVertex: a batched multi-draw tier cannot give each sub-draw
// its own base vertex through a uniform either.
Bool CurrentProgramReadsBaseVertex();
// Both of the above, conservatively, for a caller that must decide BEFORE PrepareForDraw
// has synced the program - where "does not read it" is indistinguishable from "cannot be
// asked yet". Answers true whenever the backend twin is missing or predates the current
// link.
Bool CurrentProgramMayNeedPerSubDrawBuiltins(Bool batchCarriesBaseVertices);
template <typename StateObject, typename BackendObject> template <typename StateObject, typename BackendObject>
class StateBackendObjectRegistry { class StateBackendObjectRegistry {
@@ -454,6 +466,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
Uint32 m_syncedConfigVersion = 0; Uint32 m_syncedConfigVersion = 0;
Array<MG_State::GLState::VertexAttributeVersion, MG_State::GLState::VertexArrayObject::MAX_VERTEX_ATTRIBS> Array<MG_State::GLState::VertexAttributeVersion, MG_State::GLState::VertexArrayObject::MAX_VERTEX_ATTRIBS>
m_syncedAttributeVersions; m_syncedAttributeVersions;
// Byte shift currently baked into the instanced arrays' offsets by the baseInstance
// emulation (see SetPendingFetchBaseInstance). It is draw state, not VAO state, so it
// is deliberately NOT covered by the config version: the frontend never bumps for it.
// Kept here because it describes what was last EMITTED, which is what the next sync
// has to correct.
Uint32 m_syncedFetchBaseInstance = 0;
}; };
extern StateBackendObjectRegistry<MG_State::GLState::VertexArrayObject, BackendVertexArrayObject> extern StateBackendObjectRegistry<MG_State::GLState::VertexArrayObject, BackendVertexArrayObject>
@@ -467,6 +485,23 @@ namespace MobileGL::MG_Backend::DirectGLES {
void InvalidateVAOBindingCache(); void InvalidateVAOBindingCache();
// ES resets the binding to 0 when the currently bound VAO is deleted. // ES resets the binding to 0 when the currently bound VAO is deleted.
void NoteVAOIdDeleted(Uint id); void NoteVAOIdDeleted(Uint id);
// baseInstance emulation for drivers without GL_EXT_base_instance. GL fetches an
// instanced array at element "floor(instance / divisor) + baseInstance", and ES has no
// way to say the "+ baseInstance" part - so it is folded into the attribute's own byte
// offset (baseInstance * stride) for every divisor'd array, which is exactly equivalent.
// Must be set BEFORE PrepareForDraw so the VAO sync sees it, and cleared after the draw
// so the next one refetches from element 0; ScopedFetchBaseInstance does both.
void SetPendingFetchBaseInstance(Uint32 baseInstance);
Uint32 GetPendingFetchBaseInstance();
class ScopedFetchBaseInstance {
public:
explicit ScopedFetchBaseInstance(Uint32 baseInstance) { SetPendingFetchBaseInstance(baseInstance); }
~ScopedFetchBaseInstance() { SetPendingFetchBaseInstance(0); }
ScopedFetchBaseInstance(const ScopedFetchBaseInstance&) = delete;
ScopedFetchBaseInstance& operator=(const ScopedFetchBaseInstance&) = delete;
};
} // namespace VertexArrayImpl } // namespace VertexArrayImpl
namespace TextureImpl { namespace TextureImpl {
@@ -1025,9 +1060,13 @@ namespace MobileGL::MG_Backend::DirectGLES {
void SetBaseInstance(Uint32 baseInstance) const; void SetBaseInstance(Uint32 baseInstance) const;
void SetBaseInstanceWordIndex(Int32 wordIndex) const; void SetBaseInstanceWordIndex(Int32 wordIndex) const;
void SetDrawID(Uint32 drawId) const; void SetDrawID(Uint32 drawId) const;
void SetBaseVertex(Int32 baseVertex) const;
// True when the transpiled program kept a gl_DrawID uniform, i.e. SetDrawID // True when the transpiled program kept a gl_DrawID uniform, i.e. SetDrawID
// actually reaches a shader read rather than being discarded. // actually reaches a shader read rather than being discarded.
Bool ReadsDrawID() const { return m_drawIdUniformLocation >= 0; } Bool ReadsDrawID() const { return m_drawIdUniformLocation >= 0; }
// Same for gl_BaseVertex: only a program that reads it pays for the per-draw
// uniform write, and only such a program needs the reset after one.
Bool ReadsBaseVertex() const { return m_baseVertexUniformLocation >= 0; }
Int GetIndirectParamsBinding() const { return m_indirectParamsBinding; } Int GetIndirectParamsBinding() const { return m_indirectParamsBinding; }
Uint GetBackendProgramId() const { return m_backendProgramId; } Uint GetBackendProgramId() const { return m_backendProgramId; }
// False when the last SyncToBackend could not produce a usable program (a // False when the last SyncToBackend could not produce a usable program (a
@@ -1077,6 +1116,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
Uint m_backendGlobalUBOId = 0; Uint m_backendGlobalUBOId = 0;
Int m_baseInstanceUniformLocation = -1; Int m_baseInstanceUniformLocation = -1;
Int m_drawIdUniformLocation = -1; Int m_drawIdUniformLocation = -1;
Int m_baseVertexUniformLocation = -1;
Int m_baseInstanceWordIndexUniformLocation = -1; Int m_baseInstanceWordIndexUniformLocation = -1;
Int m_indirectParamsBinding = -1; Int m_indirectParamsBinding = -1;
Uint32 m_snormFallbackClampOutputMask = 0; Uint32 m_snormFallbackClampOutputMask = 0;
+48 -14
View File
@@ -274,17 +274,22 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
// the batch's own shape - not the driver - rules it out; the compute tier keeps // the batch's own shape - not the driver - rules it out; the compute tier keeps
// its remaining feasibility checks inside its implementation, where the data it // its remaining feasibility checks inside its implementation, where the data it
// has to walk is already in hand. // has to walk is already in hand.
GLESMultiDrawMode ResolveTierForBatch(Bool programReadsDrawID, Bool hasIndexBuffer) { GLESMultiDrawMode ResolveTierForBatch(Bool programReadsDrawID, Bool perSubDrawBaseVertex,
Bool hasIndexBuffer) {
ResolveTierOnce(); ResolveTierOnce();
GLESMultiDrawMode tier = g_resolvedTier; GLESMultiDrawMode tier = g_resolvedTier;
// Batched tiers issue one driver entry for the whole batch, so the emulated // Batched tiers issue one driver entry for the whole batch, so the emulated
// gl_DrawID uniform can only hold one value across every sub-draw. A program // gl_DrawID uniform can only hold one value across every sub-draw. A program
// that reads gl_DrawID gets an unrolled tier, which feeds each sub-draw its // that reads gl_DrawID gets an unrolled tier, which feeds each sub-draw its
// own index (the spec's value); nothing else observes the difference. // own index (the spec's value); nothing else observes the difference. The
// emulated gl_BaseVertex is one uniform for the same reason, so a batch whose
// sub-draws carry their own base vertices unrolls too - even the Ext tier,
// which hands the driver the whole basevertex array, can only leave ONE value
// in the uniform the shader reads.
const Bool batched = tier == GLESMultiDrawMode::Ext || tier == GLESMultiDrawMode::MultiIndirect || const Bool batched = tier == GLESMultiDrawMode::Ext || tier == GLESMultiDrawMode::MultiIndirect ||
tier == GLESMultiDrawMode::Compute; tier == GLESMultiDrawMode::Compute;
if (batched && programReadsDrawID) { if (batched && (programReadsDrawID || perSubDrawBaseVertex)) {
tier = SupportsTier(GLESMultiDrawMode::BaseVertex) ? GLESMultiDrawMode::BaseVertex tier = SupportsTier(GLESMultiDrawMode::BaseVertex) ? GLESMultiDrawMode::BaseVertex
: GLESMultiDrawMode::DrawElements; : GLESMultiDrawMode::DrawElements;
} }
@@ -371,7 +376,8 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
Bool RunIndirect(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices, Bool RunIndirect(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices,
GLsizei drawcount, const GLint* basevertex, Bool batched, Bool feedDrawID) { GLsizei drawcount, const GLint* basevertex, Bool batched, Bool feedDrawID,
Bool feedBaseVertex) {
if (!SupportsTier(batched ? GLESMultiDrawMode::MultiIndirect : GLESMultiDrawMode::Indirect)) return false; if (!SupportsTier(batched ? GLESMultiDrawMode::MultiIndirect : GLESMultiDrawMode::Indirect)) return false;
const SizeT indexSize = IndexTypeSize(type); const SizeT indexSize = IndexTypeSize(type);
if (indexSize == 0) return false; if (indexSize == 0) return false;
@@ -413,10 +419,12 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
} else { } else {
for (GLsizei i = 0; i < drawcount; ++i) { for (GLsizei i = 0; i < drawcount; ++i) {
if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i)); if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i));
if (feedBaseVertex) SetCurrentBaseVertex(basevertex ? basevertex[i] : 0);
const SizeT commandOffset = commandBase + static_cast<SizeT>(i) * sizeof(DrawElementsIndirectCommand); const SizeT commandOffset = commandBase + static_cast<SizeT>(i) * sizeof(DrawElementsIndirectCommand);
g_GLESFuncs.glDrawElementsIndirect(mode, type, reinterpret_cast<const void*>(commandOffset)); g_GLESFuncs.glDrawElementsIndirect(mode, type, reinterpret_cast<const void*>(commandOffset));
} }
if (feedDrawID) SetCurrentDrawID(0); if (feedDrawID) SetCurrentDrawID(0);
if (feedBaseVertex) SetCurrentBaseVertex(0);
} }
BufferImpl::BindBufferId(GL_DRAW_INDIRECT_BUFFER, previousIndirectBinding); BufferImpl::BindBufferId(GL_DRAW_INDIRECT_BUFFER, previousIndirectBinding);
NoteTierExecuted(batched ? GLESMultiDrawMode::MultiIndirect : GLESMultiDrawMode::Indirect); NoteTierExecuted(batched ? GLESMultiDrawMode::MultiIndirect : GLESMultiDrawMode::Indirect);
@@ -428,15 +436,17 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
Bool RunBaseVertexLoop(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices, Bool RunBaseVertexLoop(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices,
GLsizei drawcount, const GLint* basevertex, Bool feedDrawID) { GLsizei drawcount, const GLint* basevertex, Bool feedDrawID, Bool feedBaseVertex) {
if (!SupportsTier(GLESMultiDrawMode::BaseVertex)) return false; if (!SupportsTier(GLESMultiDrawMode::BaseVertex)) return false;
for (GLsizei i = 0; i < drawcount; ++i) { for (GLsizei i = 0; i < drawcount; ++i) {
if (count[i] <= 0) continue; if (count[i] <= 0) continue;
if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i)); if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i));
if (feedBaseVertex) SetCurrentBaseVertex(basevertex ? basevertex[i] : 0);
g_GLESFuncs.glDrawElementsBaseVertex(mode, count[i], type, indices[i], g_GLESFuncs.glDrawElementsBaseVertex(mode, count[i], type, indices[i],
basevertex ? basevertex[i] : 0); basevertex ? basevertex[i] : 0);
} }
if (feedDrawID) SetCurrentDrawID(0); if (feedDrawID) SetCurrentDrawID(0);
if (feedBaseVertex) SetCurrentBaseVertex(0);
NoteTierExecuted(GLESMultiDrawMode::BaseVertex); NoteTierExecuted(GLESMultiDrawMode::BaseVertex);
return true; return true;
} }
@@ -446,7 +456,8 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
Bool RunRebasedDrawElements(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices, Bool RunRebasedDrawElements(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices,
GLsizei drawcount, const GLint* basevertex, Bool feedDrawID) { GLsizei drawcount, const GLint* basevertex, Bool feedDrawID,
Bool feedBaseVertex) {
const SizeT indexSize = IndexTypeSize(type); const SizeT indexSize = IndexTypeSize(type);
if (indexSize == 0) return false; if (indexSize == 0) return false;
@@ -500,11 +511,16 @@ namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
for (GLsizei i = 0; i < drawcount; ++i) { for (GLsizei i = 0; i < drawcount; ++i) {
if (count[i] <= 0) continue; if (count[i] <= 0) continue;
if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i)); if (feedDrawID) SetCurrentDrawID(static_cast<Uint32>(i));
// The base vertex is folded into the rewritten index stream here, so the
// driver sees none - but gl_BaseVertex still has to report the value the
// application passed for this sub-draw.
if (feedBaseVertex) SetCurrentBaseVertex(basevertex ? basevertex[i] : 0);
g_GLESFuncs.glDrawElements(mode, count[i], GL_UNSIGNED_INT, g_GLESFuncs.glDrawElements(mode, count[i], GL_UNSIGNED_INT,
reinterpret_cast<const void*>(indexBase + cursor * sizeof(Uint32))); reinterpret_cast<const void*>(indexBase + cursor * sizeof(Uint32)));
cursor += static_cast<SizeT>(count[i]); cursor += static_cast<SizeT>(count[i]);
} }
if (feedDrawID) SetCurrentDrawID(0); if (feedDrawID) SetCurrentDrawID(0);
if (feedBaseVertex) SetCurrentBaseVertex(0);
BufferImpl::BindBufferId(GL_ELEMENT_ARRAY_BUFFER, previousIndexBinding); BufferImpl::BindBufferId(GL_ELEMENT_ARRAY_BUFFER, previousIndexBinding);
NoteTierExecuted(GLESMultiDrawMode::DrawElements); NoteTierExecuted(GLESMultiDrawMode::DrawElements);
return true; return true;
@@ -837,8 +853,15 @@ void main() {
// afterwards would mean unpicking the program, SSBO and index bindings // afterwards would mean unpicking the program, SSBO and index bindings
// PrepareForDraw just made, and a dispatch inside an open transform feedback // PrepareForDraw just made, and a dispatch inside an open transform feedback
// span is not legal at all. On success it hands back a flattened index stream. // span is not legal at all. On success it hands back a flattened index stream.
// A batch whose sub-draws carry their own base vertices cannot be flattened either
// when the program reads gl_BaseVertex: one draw call leaves one uniform value.
// Asked conservatively because this decision precedes PrepareForDraw - see
// CurrentProgramMayNeedPerSubDrawBuiltins. Flattening is the irreversible half:
// once the batch is one draw the values are gone, whereas declining to flatten only
// costs the unrolled tier.
FlattenedStream flattened; FlattenedStream flattened;
if (ResolvedTier() == GLESMultiDrawMode::Compute && !CurrentProgramReadsDrawID()) { if (ResolvedTier() == GLESMultiDrawMode::Compute &&
!CurrentProgramMayNeedPerSubDrawBuiltins(basevertex != nullptr)) {
FlattenWithCompute(mode, count, type, indices, drawcount, basevertex, flattened); FlattenWithCompute(mode, count, type, indices, drawcount, basevertex, flattened);
} }
@@ -852,8 +875,11 @@ void main() {
return; return;
} }
// Now that PrepareForDraw has synced the program, both questions have real answers;
// the tier choice and the per-sub-draw feeds use those, not the guess above.
const Bool feedDrawID = CurrentProgramReadsDrawID(); const Bool feedDrawID = CurrentProgramReadsDrawID();
const GLESMultiDrawMode tier = ResolveTierForBatch(feedDrawID, hasIndexBuffer); const Bool feedBaseVertex = basevertex != nullptr && CurrentProgramReadsBaseVertex();
const GLESMultiDrawMode tier = ResolveTierForBatch(feedDrawID, feedBaseVertex, hasIndexBuffer);
Bool drawn = false; Bool drawn = false;
switch (tier) { switch (tier) {
@@ -861,16 +887,19 @@ void main() {
drawn = RunExt(mode, count, type, indices, drawcount, basevertex); drawn = RunExt(mode, count, type, indices, drawcount, basevertex);
break; break;
case GLESMultiDrawMode::MultiIndirect: case GLESMultiDrawMode::MultiIndirect:
drawn = RunIndirect(mode, count, type, indices, drawcount, basevertex, /*batched=*/true, feedDrawID); drawn = RunIndirect(mode, count, type, indices, drawcount, basevertex, /*batched=*/true, feedDrawID,
feedBaseVertex);
break; break;
case GLESMultiDrawMode::Indirect: case GLESMultiDrawMode::Indirect:
drawn = RunIndirect(mode, count, type, indices, drawcount, basevertex, /*batched=*/false, feedDrawID); drawn = RunIndirect(mode, count, type, indices, drawcount, basevertex, /*batched=*/false, feedDrawID,
feedBaseVertex);
break; break;
case GLESMultiDrawMode::BaseVertex: case GLESMultiDrawMode::BaseVertex:
drawn = RunBaseVertexLoop(mode, count, type, indices, drawcount, basevertex, feedDrawID); drawn = RunBaseVertexLoop(mode, count, type, indices, drawcount, basevertex, feedDrawID, feedBaseVertex);
break; break;
case GLESMultiDrawMode::DrawElements: case GLESMultiDrawMode::DrawElements:
drawn = RunRebasedDrawElements(mode, count, type, indices, drawcount, basevertex, feedDrawID); drawn = RunRebasedDrawElements(mode, count, type, indices, drawcount, basevertex, feedDrawID,
feedBaseVertex);
break; break;
case GLESMultiDrawMode::Compute: case GLESMultiDrawMode::Compute:
// Its pre-pass ran above; reaching here means it declined this batch's shape. // Its pre-pass ran above; reaching here means it declined this batch's shape.
@@ -883,8 +912,13 @@ void main() {
// below are the floor: a base-vertex replay where the driver has one, and the // below are the floor: a base-vertex replay where the driver has one, and the
// rewritten index stream where it does not. Both are safe for any batch these // rewritten index stream where it does not. Both are safe for any batch these
// entry points can receive. // entry points can receive.
if (!drawn) drawn = RunBaseVertexLoop(mode, count, type, indices, drawcount, basevertex, feedDrawID); if (!drawn) {
if (!drawn) drawn = RunRebasedDrawElements(mode, count, type, indices, drawcount, basevertex, feedDrawID); drawn = RunBaseVertexLoop(mode, count, type, indices, drawcount, basevertex, feedDrawID, feedBaseVertex);
}
if (!drawn) {
drawn = RunRebasedDrawElements(mode, count, type, indices, drawcount, basevertex, feedDrawID,
feedBaseVertex);
}
if (!drawn) { if (!drawn) {
MGLOG_E("DirectGLES multi-draw: no usable tier for a %d sub-draw batch (mode 0x%x, type 0x%x); " MGLOG_E("DirectGLES multi-draw: no usable tier for a %d sub-draw batch (mode 0x%x, type 0x%x); "
"the batch was dropped", "the batch was dropped",
@@ -539,6 +539,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (MG_Util::Async::AsyncShaderCompileEnabled()) { if (MG_Util::Async::AsyncShaderCompileEnabled()) {
extensions.push_back(E_GL_KHR_parallel_shader_compile); extensions.push_back(E_GL_KHR_parallel_shader_compile);
} }
// GL_ARB_gpu_shader_fp64 is opt-in (MOBILEGL_ADVERTISE_FP64). Every `double` in a
// shader compiles and runs already - it is narrowed to 32 bits before the module
// reaches this backend - so an application that simply uses doubles needs nothing
// advertised. What the extension additionally promises is 64-bit PRECISION, which no
// mobile GPU has and the narrowing cannot fake, so advertising it by default would
// make an application that checks the string take a path MobileGL cannot honour.
if (MG_Config::Features.AdvertiseFp64) {
extensions.push_back(E_GL_ARB_gpu_shader_fp64);
}
// GL_ARB_timer_query gates MC's F3 GPU% (LWJGL checks the extension string); // GL_ARB_timer_query gates MC's F3 GPU% (LWJGL checks the extension string);
// only advertised when the device actually supports timestamp queries and the // only advertised when the device actually supports timestamp queries and the
// MOBILEGL_DISABLE_TIMERQUERY escape hatch is off. // MOBILEGL_DISABLE_TIMERQUERY escape hatch is off.
@@ -877,7 +886,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
DynParams::PerLayerFramebufferAttachmentBit(TextureTarget::TextureCubeMapArray); DynParams::PerLayerFramebufferAttachmentBit(TextureTarget::TextureCubeMapArray);
} }
} }
m_dynamicParameters.SupportsFloat64VertexAttributes = m_vulkanCaps.SupportsShaderFloat64; // Never, on any device, and no longer for the reason it used to be. It used to track
// shaderFloat64 because a `dvec3` input needed the Float64 capability to exist in the
// module at all; a 64-bit vertex FETCH was already impossible (VK_FORMAT_R64*_SFLOAT is
// optional and lavapipe reports zero bufferFeatures for all four), so the attribute
// arrived as its 32-bit word pair and PackDoubleVertexInputsPass bitcast it back.
//
// The shader half of that is gone: every 64-bit float is narrowed before any module
// reaches a backend (ShaderTranspiler::DemoteFloat64Pass), so there is no `double` input
// left to bitcast INTO, and feeding a UINT-formatted attribute to what is now a `float`
// input would be silent garbage. Reconstructing the value would mean decoding the
// IEEE-754 double bit pattern in the shader - software fp64, which is precisely what the
// demotion exists to avoid - and on Espryt it would additionally need the ES driver to
// fetch 2N uint components where the application declared N doubles, which a dvec3 or
// dvec4 cannot even express within one attribute location.
//
// So glVertexAttribLFormat / glVertexAttribLPointer are declined here exactly as they
// already were on Espryt and on every real mobile device (Adreno and Mali both report
// shaderFloat64 == VK_FALSE), and for the same visible reason. A `dvec3` INPUT still
// compiles and draws - it is a `vec3` after demotion - as long as the application feeds
// it with glVertexAttribPointer(GL_FLOAT) rather than 64-bit data.
m_dynamicParameters.SupportsFloat64VertexAttributes = false;
m_dynamicParameters.MaxShaderStorageBlockSize = m_dynamicParameters.MaxShaderStorageBlockSize =
std::min(m_vulkanCaps.MaxShaderStorageBlockSize, kMaxAdvertisedShaderStorageBlockSize); std::min(m_vulkanCaps.MaxShaderStorageBlockSize, kMaxAdvertisedShaderStorageBlockSize);
if (m_vulkanCaps.SupportsShaderSubgroup) { if (m_vulkanCaps.SupportsShaderSubgroup) {
@@ -1979,13 +1979,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// cannot be corrected and instanced draws with a non-zero baseInstance misrender; this // cannot be corrected and instanced draws with a non-zero baseInstance misrender; this
// detects the case so the user gets one warning instead of silent corruption. // detects the case so the user gets one warning instead of silent corruption.
Bool ProgramFactory::ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule) { Bool ProgramFactory::ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule) {
return ReflectedDeclaresInputBuiltin(reflectModule, SpvBuiltInInstanceIndex);
}
// GL's gl_BaseVertex and Vulkan's BaseVertex agree for indexed draws and disagree for every
// other command, so a program declaring the builtin needs the ZeroBaseVertex variant when a
// non-indexed draw uses it (see CompileOptionBit::ZeroBaseVertex). "Declares" rather than
// "reads" is the honest word and the useful one: the zeroing pass keeps the variable, so
// both variants of a program answer this question identically.
Bool ProgramFactory::ReflectedReadsBaseVertexBuiltin(const SpvReflectShaderModule& reflectModule) {
return ReflectedDeclaresInputBuiltin(reflectModule, SpvBuiltInBaseVertex);
}
Bool ProgramFactory::ReflectedDeclaresInputBuiltin(const SpvReflectShaderModule& reflectModule,
SpvBuiltIn builtin) {
for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) { for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) {
const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex]; const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex];
for (Uint32 variableIndex = 0; variableIndex < entryPoint.input_variable_count; ++variableIndex) { for (Uint32 variableIndex = 0; variableIndex < entryPoint.input_variable_count; ++variableIndex) {
const SpvReflectInterfaceVariable* variable = entryPoint.input_variables[variableIndex]; const SpvReflectInterfaceVariable* variable = entryPoint.input_variables[variableIndex];
if (variable != nullptr && if (variable != nullptr &&
(variable->decoration_flags & SPV_REFLECT_DECORATION_BUILT_IN) != 0 && (variable->decoration_flags & SPV_REFLECT_DECORATION_BUILT_IN) != 0 &&
variable->built_in == SpvBuiltInInstanceIndex) { variable->built_in == builtin) {
return true; return true;
} }
} }
@@ -2244,6 +2258,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkProgramObject& entry) const { VkProgramObject& entry) const {
entry.activeVertexInputLocationMask = 0; entry.activeVertexInputLocationMask = 0;
entry.vertexInputTypes.fill(0); entry.vertexInputTypes.fill(0);
entry.readsBaseVertexBuiltin = false;
for (SizeT moduleIndex = 0; moduleIndex < shaders.size() && moduleIndex < spirv.size(); ++moduleIndex) { for (SizeT moduleIndex = 0; moduleIndex < shaders.size() && moduleIndex < spirv.size(); ++moduleIndex) {
if (!shaders[moduleIndex] || shaders[moduleIndex]->GetShaderStage() != ShaderStage::Vertex) { if (!shaders[moduleIndex] || shaders[moduleIndex]->GetShaderStage() != ShaderStage::Vertex) {
@@ -2265,6 +2280,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
continue; continue;
} }
entry.readsBaseVertexBuiltin = ReflectedReadsBaseVertexBuiltin(reflectModule);
if (!m_shaderDrawParametersEnabled && ReflectedReadsInstanceIndexBuiltin(reflectModule)) { if (!m_shaderDrawParametersEnabled && ReflectedReadsInstanceIndexBuiltin(reflectModule)) {
static Bool s_warnedInstanceIndexUnsupported = false; static Bool s_warnedInstanceIndexUnsupported = false;
if (!s_warnedInstanceIndexUnsupported) { if (!s_warnedInstanceIndexUnsupported) {
@@ -2909,6 +2926,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// a FragCoordYFlip variant also depends on the baked default-framebuffer height, so // a FragCoordYFlip variant also depends on the baked default-framebuffer height, so
// that height rides in the free high half of the key. Flags occupy the low bits, and a // that height rides in the free high half of the key. Flags occupy the low bits, and a
// height cannot exceed the 16 bits a swapchain extent fits in. // height cannot exceed the 16 bits a swapchain extent fits in.
//
// "The low bits" is load-bearing and was until now only a comment: a flag that reached
// bit 16 would alias the height and two different variants would share one memo slot.
static_assert(static_cast<Uint>(CompileOptionBit::ZeroBaseVertex) < (1u << 16),
"CompileOptionBit values must stay below bit 16: GetOrCreateProgram packs the "
"default-framebuffer height into the high half of the same memo key");
const Uint memoKey = (flags & CompileOptionBit::FragCoordYFlip) const Uint memoKey = (flags & CompileOptionBit::FragCoordYFlip)
? (flags.GetRaw() | (m_defaultFramebufferHeight << 16)) ? (flags.GetRaw() | (m_defaultFramebufferHeight << 16))
: flags.GetRaw(); : flags.GetRaw();
@@ -3027,6 +3050,26 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} }
// The non-indexed variant of a vertex stage that reads gl_BaseVertex: GL wants zero
// there, Vulkan's builtin would hand it the draw's firstVertex. Requested per draw
// through CompileOptionBit::ZeroBaseVertex, so the indexed variant of the same
// program keeps the native builtin and stays correct for glDrawElementsBaseVertex
// and for the baseVertex word of an indexed indirect command.
if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Vertex &&
(flags & CompileOptionBit::ZeroBaseVertex)) {
Vector<Uint> zeroedSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::ZeroBaseVertexForVulkan(moduleSpirvs[i],
zeroedSpirv)) {
moduleSpirvs[i] = std::move(zeroedSpirv);
} else {
// Failing open keeps the native builtin, which is the pre-fix behavior:
// gl_BaseVertex reads firstVertex on a DrawArrays instead of zero.
MGLOG_E("ProgramFactory: failed to zero gl_BaseVertex for program %u; non-indexed "
"draws will read the draw's first vertex from it instead of zero",
program.GetExternalIndex());
}
}
// A 64-bit vertex input has to arrive as its 32-bit word pair: VK_FORMAT_R64*_SFLOAT is // A 64-bit vertex input has to arrive as its 32-bit word pair: VK_FORMAT_R64*_SFLOAT is
// optional and lavapipe advertises none of them at all. The pass is unconditional so it // optional and lavapipe advertises none of them at all. The pass is unconditional so it
// always agrees with the Float64 case in VertexInputStateFactory::ToVkVertexFormat, and // always agrees with the Float64 case in VertexInputStateFactory::ToVkVertexFormat, and
@@ -60,6 +60,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// PositionYFlip (the two are the same fact about the same draws) except under a // PositionYFlip (the two are the same fact about the same draws) except under a
// quarter turn, which this renderer does not convert rectangles for either. // quarter turn, which this renderer does not convert rectangles for either.
FragCoordYFlip = 1 << 7, FragCoordYFlip = 1 << 7,
// Replaces the vertex stage's gl_BaseVertex reads with zero. GL defines the builtin
// as zero for every drawing command that has no baseVertex parameter - all the
// DrawArrays forms - while Vulkan's BaseVertex reports firstVertex there. Set only
// for a non-indexed draw whose program actually reads the builtin, so nothing else
// acquires a second program/pipeline variant. See ZeroBaseVertexPass.
ZeroBaseVertex = 1 << 8,
}; };
using CompileOptionFlags = Flags<CompileOptionBit>; using CompileOptionFlags = Flags<CompileOptionBit>;
using HashType = Uint64; using HashType = Uint64;
@@ -129,6 +135,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// gl_FragDepth); shader-computed depth is immune to the cross-pipeline // gl_FragDepth); shader-computed depth is immune to the cross-pipeline
// position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite). // position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false; Bool fragmentReplacesDepth = false;
// The vertex module declares the BaseVertex builtin. Selects the ZeroBaseVertex
// program variant for non-indexed draws, and is deliberately a property of the
// PROGRAM rather than of the variant: the zeroed variant leaves the variable
// declared, so both variants answer the same and the draw path can ask either.
Bool readsBaseVertexBuiltin = false;
// Frame-boundary counter value of the last GetOrCreateProgram hit; drives // Frame-boundary counter value of the last GetOrCreateProgram hit; drives
// cache eviction (see OnFrameBoundary). Mutable: the draw snapshot's memoised // cache eviction (see OnFrameBoundary). Mutable: the draw snapshot's memoised
// entry pointer re-stamps use through a const reference (StampProgramUse). // entry pointer re-stamps use through a const reference (StampProgramUse).
@@ -179,6 +190,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth; fragmentReplacesDepth = other.fragmentReplacesDepth;
readsBaseVertexBuiltin = other.readsBaseVertexBuiltin;
lastUsedFrame = other.lastUsedFrame; lastUsedFrame = other.lastUsedFrame;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
@@ -192,6 +204,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false; other.fragmentReplacesDepth = false;
other.readsBaseVertexBuiltin = false;
other.lastUsedFrame = 0; other.lastUsedFrame = 0;
} }
VkProgramObject& operator=(VkProgramObject&& other) noexcept { VkProgramObject& operator=(VkProgramObject&& other) noexcept {
@@ -231,6 +244,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth; fragmentReplacesDepth = other.fragmentReplacesDepth;
readsBaseVertexBuiltin = other.readsBaseVertexBuiltin;
lastUsedFrame = other.lastUsedFrame; lastUsedFrame = other.lastUsedFrame;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
@@ -244,6 +258,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false; other.fragmentReplacesDepth = false;
other.readsBaseVertexBuiltin = false;
other.lastUsedFrame = 0; other.lastUsedFrame = 0;
return *this; return *this;
} }
@@ -341,6 +356,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// True when an entry point reads the InstanceIndex builtin. Only gates a diagnostic: // True when an entry point reads the InstanceIndex builtin. Only gates a diagnostic:
// without shaderDrawParameters such a shader cannot have gl_InstanceID rebased. // without shaderDrawParameters such a shader cannot have gl_InstanceID rebased.
static Bool ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule); static Bool ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule);
// True when an entry point declares the BaseVertex builtin, i.e. when a non-indexed
// draw with this program has to take the ZeroBaseVertex variant.
static Bool ReflectedReadsBaseVertexBuiltin(const SpvReflectShaderModule& reflectModule);
// Shared by the two above: does any entry point list an input variable decorated with
// this builtin?
static Bool ReflectedDeclaresInputBuiltin(const SpvReflectShaderModule& reflectModule, SpvBuiltIn builtin);
private: private:
struct ProgramLookupCache { struct ProgramLookupCache {
@@ -153,8 +153,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
continue; continue;
} }
const Uint32 sourceStride = // Verbatim, zero included. The frontend already resolved a pointer call's
attr.Stride > 0 ? static_cast<Uint32>(attr.Stride) : static_cast<Uint32>(attribByteSize); // "tightly packed" stride 0 into the element size (see VertexAttribute::Stride),
// so a zero here is the binding model's stride 0 - every vertex reads the same
// element - which is exactly what a zero VkVertexInputBindingDescription::stride
// means. Substituting the element size fetched a fresh element per vertex and ran
// off the end of the buffer (KHR-GL43.vertex_attrib_binding.basic-input-case7/8).
// Client-memory arrays cannot reach zero: they only exist on the pointer path.
const Uint32 sourceStride = static_cast<Uint32>(attr.Stride);
const Bool packedAttribute = attr.Type == DataType::Int2101010Rev || const Bool packedAttribute = attr.Type == DataType::Int2101010Rev ||
attr.Type == DataType::Uint2101010Rev; attr.Type == DataType::Uint2101010Rev;
const SizeT requiredAlignment = packedAttribute ? attribByteSize : GetComponentSize(attr.Type); const SizeT requiredAlignment = packedAttribute ? attribByteSize : GetComponentSize(attr.Type);
@@ -175,11 +181,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
Uint32 stride = sourceStride; Uint32 stride = sourceStride;
// A converted stream is tightly packed, so its stride is the converted element
// size - unless the source stride is zero, which does not describe a packing at
// all but "never advance". That survives the conversion unchanged: the draw path
// converts exactly one element and every vertex reads it.
if (sourceStride != 0) {
if (conversion == VertexStreamConversion::Repack) { if (conversion == VertexStreamConversion::Repack) {
stride = static_cast<Uint32>(attribByteSize); stride = static_cast<Uint32>(attribByteSize);
} else if (conversion == VertexStreamConversion::ScaledIntegerToFloat32) { } else if (conversion == VertexStreamConversion::ScaledIntegerToFloat32) {
stride = static_cast<Uint32>(attr.Size * static_cast<Int>(sizeof(Float))); stride = static_cast<Uint32>(attr.Size * static_cast<Int>(sizeof(Float)));
} }
}
const VkVertexInputRate inputRate = const VkVertexInputRate inputRate =
(attr.Divisor == 0) ? VK_VERTEX_INPUT_RATE_VERTEX : VK_VERTEX_INPUT_RATE_INSTANCE; (attr.Divisor == 0) ? VK_VERTEX_INPUT_RATE_VERTEX : VK_VERTEX_INPUT_RATE_INSTANCE;
@@ -3546,9 +3546,11 @@ void main() {
const auto& attr = vao.GetAttribute(bindingLocation); const auto& attr = vao.GetAttribute(bindingLocation);
const SizeT elementSize = const SizeT elementSize =
VertexInputStateFactory::GetAttributeByteSize(attr.Type, attr.Size, attr.IsBgra); VertexInputStateFactory::GetAttributeByteSize(attr.Type, attr.Size, attr.IsBgra);
const SizeT sourceStride = // Zero is a legal binding stride and means "never advance" (see
attr.Stride > 0 ? static_cast<SizeT>(attr.Stride) : elementSize; // VertexAttribute::Stride), so it is NOT folded into the element size here -
if (sourceBufferShared->MappedData() == nullptr || elementSize == 0 || sourceStride == 0 || // it selects the single-element conversion below instead.
const SizeT sourceStride = static_cast<SizeT>(attr.Stride);
if (sourceBufferShared->MappedData() == nullptr || elementSize == 0 ||
baseOffset > sourceSize || elementSize > sourceSize - baseOffset) { baseOffset > sourceSize || elementSize > sourceSize - baseOffset) {
MGLOG_E("UploadAndBindVertexStreams skipped: invalid converted source binding=%zu " MGLOG_E("UploadAndBindVertexStreams skipped: invalid converted source binding=%zu "
"location=%u base=%zu size=%zu element=%zu stride=%zu", "location=%u base=%zu size=%zu element=%zu stride=%zu",
@@ -3557,7 +3559,8 @@ void main() {
} }
sourceBufferShared->SyncPersistentMappedRange(); sourceBufferShared->SyncPersistentMappedRange();
const SizeT availableElementCount = 1 + (sourceSize - baseOffset - elementSize) / sourceStride; const SizeT availableElementCount =
sourceStride == 0 ? 1 : 1 + (sourceSize - baseOffset - elementSize) / sourceStride;
const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped(); const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped();
// Convert only what this draw can fetch instead of the whole buffer tail. // Convert only what this draw can fetch instead of the whole buffer tail.
// Instance-rate bindings index by instance, not the vertex range, so they // Instance-rate bindings index by instance, not the vertex range, so they
@@ -4737,6 +4740,7 @@ void main() {
hasPatchedVertexAttributes = true; hasPatchedVertexAttributes = true;
} }
VertexInputStateBuilder syntheticVertexInputBuilder; VertexInputStateBuilder syntheticVertexInputBuilder;
VkPipelineVertexInputStateCreateInfo syntheticVertexInputState{};
const VkPipelineVertexInputStateCreateInfo* pipelineVertexInputState = &vis.state; const VkPipelineVertexInputStateCreateInfo* pipelineVertexInputState = &vis.state;
if (missingAttribMask != 0 || hasPatchedVertexAttributes) { if (missingAttribMask != 0 || hasPatchedVertexAttributes) {
for (const auto& binding : vis.bindings) { for (const auto& binding : vis.bindings) {
@@ -4764,7 +4768,16 @@ void main() {
syntheticVertexInputBuilder.AddAttribute(location, syntheticBinding, format, 0); syntheticVertexInputBuilder.AddAttribute(location, syntheticBinding, format, 0);
++syntheticBinding; ++syntheticBinding;
} }
pipelineVertexInputState = &syntheticVertexInputBuilder.Build(); syntheticVertexInputState = syntheticVertexInputBuilder.Build();
// Carry the divisor chain over. The synthetic rebuild copies bindings and
// attributes only, and it keeps every real binding's INDEX, so the divisor
// descriptions built for them stay valid - but dropping the pNext silently
// demoted every instanced binding to divisor 1. This path runs whenever the
// program declares an input the VAO does not feed (which is most capture
// shaders: KHR-GL43.vertex_attrib_binding declares 16 inputs and enables three),
// so the loss was near-total rather than a corner case.
syntheticVertexInputState.pNext = vis.state.pNext;
pipelineVertexInputState = &syntheticVertexInputState;
} }
auto cullFaceEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::CullFace); auto cullFaceEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::CullFace);
auto depthTestEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::DepthTest); auto depthTestEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::DepthTest);
@@ -5821,7 +5834,40 @@ void main() {
m_lastLodParamsSum = 0; // filled below once the sampled set is known m_lastLodParamsSum = 0; // filled below once the sampled set is known
} }
} }
const auto& programObj = m_programFactory->GetOrCreateProgram(program, transformFlags); // GL's gl_BaseVertex is zero for every command without a baseVertex parameter, while
// Vulkan's builtin reports the draw's firstVertex; a non-indexed draw therefore takes
// the zeroed program variant. The question is about the program's SPIR-V, not about
// this draw, so it is memoized on (program lifetime, backend-state version): only the
// very first draw of a program pays the extra lookup, and a program used exclusively
// with non-indexed draws never resolves - never compiles, never re-stamps - the
// variant no draw of it would use.
const Bool nonIndexedDraw = !(aspects & DrawSetupAspect::IndexBuffer);
const Uint64 baseVertexProgramLifetimeId = program.GetLifetimeId();
const Uint32 baseVertexProgramVersion = program.GetBackendStateVersion();
const Bool baseVertexQueryKnown = m_lastBaseVertexQueryValid &&
m_lastBaseVertexProgramLifetimeId == baseVertexProgramLifetimeId &&
m_lastBaseVertexProgramVersion == baseVertexProgramVersion;
if (baseVertexQueryKnown && nonIndexedDraw && m_lastBaseVertexReads) {
transformFlags |= ProgramFactory::CompileOptionBit::ZeroBaseVertex;
}
const ProgramFactory::VkProgramObject* resolvedProgramObj =
&m_programFactory->GetOrCreateProgram(program, transformFlags);
if (!baseVertexQueryKnown) {
// Read the answer out of the entry BEFORE any second lookup: that lookup may
// insert and move every entry of the open-addressing cache, dangling the
// reference. The zeroing pass leaves the variable declared, so the variant just
// resolved answers the same as the base one either way.
const Bool readsBaseVertex = resolvedProgramObj->readsBaseVertexBuiltin;
m_lastBaseVertexQueryValid = true;
m_lastBaseVertexProgramLifetimeId = baseVertexProgramLifetimeId;
m_lastBaseVertexProgramVersion = baseVertexProgramVersion;
m_lastBaseVertexReads = readsBaseVertex;
if (nonIndexedDraw && readsBaseVertex) {
transformFlags |= ProgramFactory::CompileOptionBit::ZeroBaseVertex;
resolvedProgramObj = &m_programFactory->GetOrCreateProgram(program, transformFlags);
}
}
const auto& programObj = *resolvedProgramObj;
// For the snapshot's memoised entry pointer: if anything below inserts into the // For the snapshot's memoised entry pointer: if anything below inserts into the
// program cache (blit/aux program compiles), the epoch moves and the snapshot // program cache (blit/aux program compiles), the epoch moves and the snapshot
// stores no pointer for this draw - the fast path then re-looks-up once. // stores no pointer for this draw - the fast path then re-looks-up once.
@@ -773,6 +773,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ProgramFactory::CompileOptionFlags m_lastLodBaseFlags = {}; ProgramFactory::CompileOptionFlags m_lastLodBaseFlags = {};
ProgramFactory::CompileOptionFlags m_lastLodResultFlags = {}; ProgramFactory::CompileOptionFlags m_lastLodResultFlags = {};
// Does the current program's vertex stage declare the BaseVertex builtin? A property
// of the program's SPIR-V, so (lifetime id, backend-state version) is the whole key.
//
// Memoized rather than re-asked because asking means resolving the UN-zeroed program
// variant, and a program that only ever draws non-indexed would then compile a variant
// no draw uses AND re-stamp its use every draw, so the idle sweep could never retire
// it. With the memo the answer is known before the first lookup and only the variant
// the draw actually needs is resolved.
Bool m_lastBaseVertexQueryValid = false;
Uint64 m_lastBaseVertexProgramLifetimeId = 0;
Uint32 m_lastBaseVertexProgramVersion = 0;
Bool m_lastBaseVertexReads = false;
// Snapshot behind TrySetupDrawFastPath. Values only: the program and // Snapshot behind TrySetupDrawFastPath. Values only: the program and
// render-pass caches are open-addressing maps whose entries move on // render-pass caches are open-addressing maps whose entries move on
// insert, so no pointers into them are cached; the pipeline handle is // insert, so no pointers into them are cached; the pipeline handle is
@@ -1003,9 +1003,9 @@ DECLARE_GL_FUNCTION_HEAD(void, ClearTexImage, GLuint texture, GLint level, GLenu
DECLARE_GL_FUNCTION_HEAD(void, ClearTexSubImage, GLuint texture, GLint level, GLint xoffset, GLint yoffset, GLint zoffset, GLsizei width, GLsizei height, GLsizei depth, GLenum format, GLenum type, const void* data) DECLARE_GL_FUNCTION_END_NO_RETURN(void, ClearTexSubImage, texture, level, xoffset, yoffset, zoffset, width, height, depth, format, type, data) DECLARE_GL_FUNCTION_HEAD(void, ClearTexSubImage, GLuint texture, GLint level, GLint xoffset, GLint yoffset, GLint zoffset, GLsizei width, GLsizei height, GLsizei depth, GLenum format, GLenum type, const void* data) DECLARE_GL_FUNCTION_END_NO_RETURN(void, ClearTexSubImage, texture, level, xoffset, yoffset, zoffset, width, height, depth, format, type, data)
DECLARE_GL_FUNCTION_HEAD(void, BindBuffersBase, GLenum target, GLuint first, GLsizei count, const GLuint* buffers) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindBuffersBase, target, first, count, buffers) DECLARE_GL_FUNCTION_HEAD(void, BindBuffersBase, GLenum target, GLuint first, GLsizei count, const GLuint* buffers) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindBuffersBase, target, first, count, buffers)
DECLARE_GL_FUNCTION_HEAD(void, BindBuffersRange, GLenum target, GLuint first, GLsizei count, const GLuint* buffers, const GLintptr* offsets, const GLsizeiptr* sizes) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindBuffersRange, target, first, count, buffers, offsets, sizes) DECLARE_GL_FUNCTION_HEAD(void, BindBuffersRange, GLenum target, GLuint first, GLsizei count, const GLuint* buffers, const GLintptr* offsets, const GLsizeiptr* sizes) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindBuffersRange, target, first, count, buffers, offsets, sizes)
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTextures, GLuint first, GLsizei count, const GLuint* textures) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTextures, first, count, textures) DECLARE_GL_FUNCTION_HEAD(void, BindTextures, GLuint first, GLsizei count, const GLuint* textures) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindTextures, first, count, textures)
DECLARE_GL_FUNCTION_HEAD(void, BindSamplers, GLuint first, GLsizei count, const GLuint* samplers) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindSamplers, first, count, samplers) DECLARE_GL_FUNCTION_HEAD(void, BindSamplers, GLuint first, GLsizei count, const GLuint* samplers) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindSamplers, first, count, samplers)
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindImageTextures, GLuint first, GLsizei count, const GLuint* textures) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindImageTextures, first, count, textures) DECLARE_GL_FUNCTION_HEAD(void, BindImageTextures, GLuint first, GLsizei count, const GLuint* textures) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindImageTextures, first, count, textures)
DECLARE_GL_FUNCTION_HEAD(void, BindVertexBuffers, GLuint first, GLsizei count, const GLuint* buffers, const GLintptr* offsets, const GLsizei* strides) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindVertexBuffers, first, count, buffers, offsets, strides) DECLARE_GL_FUNCTION_HEAD(void, BindVertexBuffers, GLuint first, GLsizei count, const GLuint* buffers, const GLintptr* offsets, const GLsizei* strides) DECLARE_GL_FUNCTION_END_NO_RETURN(void, BindVertexBuffers, first, count, buffers, offsets, strides)
DECLARE_GL_FUNCTION_STUB_HEAD(void, ClipControl, GLenum origin, GLenum depth) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ClipControl, origin, depth) DECLARE_GL_FUNCTION_STUB_HEAD(void, ClipControl, GLenum origin, GLenum depth) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ClipControl, origin, depth)
DECLARE_GL_FUNCTION_HEAD(void, CreateTransformFeedbacks, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_END_NO_RETURN(void, CreateTransformFeedbacks, n, ids) DECLARE_GL_FUNCTION_HEAD(void, CreateTransformFeedbacks, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_END_NO_RETURN(void, CreateTransformFeedbacks, n, ids)
@@ -620,6 +620,34 @@ namespace MobileGL::MG_Impl::GLImpl {
return std::max(MG_Backend::pActiveBackendObject->GetDynamicParameters().MaxSamples, 1); return std::max(MG_Backend::pActiveBackendObject->GetDynamicParameters().MaxSamples, 1);
} }
// GL_MAX_SAMPLES is the ceiling over all formats; an integer format has its own, lower
// one (GL_MAX_INTEGER_SAMPLES) and GL 4.6 core 9.2.4 makes exceeding it INVALID_OPERATION.
// The multisample TEXTURE path already resolves the limit per format
// (GL_Texture.cpp, GetMaxTextureSamplesForFormat); renderbuffers only ever compared
// against GL_MAX_SAMPLES, so on a driver where the two differ - Adreno reports
// GL_MAX_SAMPLES 4 and GL_MAX_INTEGER_SAMPLES 1 - an integer renderbuffer accepted a
// sample count the format cannot deliver, and said GL_NO_ERROR about it.
Int GetMaxRenderbufferSamplesForFormat_State(TextureInternalFormat format) {
if (MG_Backend::pActiveBackendObject == nullptr) {
return std::numeric_limits<Int>::max();
}
const auto& dynamicParameters = MG_Backend::pActiveBackendObject->GetDynamicParameters();
GLenum normalizedInternalFormat = MG_Util::ConvertTextureInternalFormatToGLEnum(format);
GLenum normalizedFormat = GL_RGBA;
GLenum normalizedType = GL_UNSIGNED_BYTE;
MG_Util::TextureFormatProcessor::NormalizePixelFormat(normalizedInternalFormat,
PixelFormatNormalizeOptionBit::None,
&normalizedInternalFormat, &normalizedFormat,
&normalizedType);
const Bool isIntegerFormat = normalizedFormat == GL_RED_INTEGER || normalizedFormat == GL_RG_INTEGER ||
normalizedFormat == GL_RGB_INTEGER || normalizedFormat == GL_RGBA_INTEGER;
if (!isIntegerFormat) {
return GetMaxRenderbufferSamples_State();
}
return std::max(dynamicParameters.MaxIntegerSamples, 1);
}
Bool ValidateRenderbufferStorageSize_State(GLsizei width, GLsizei height, const char* caller) { Bool ValidateRenderbufferStorageSize_State(GLsizei width, GLsizei height, const char* caller) {
if (width < 0 || height < 0) { if (width < 0 || height < 0) {
MG_State::pGLContext->RecordError( MG_State::pGLContext->RecordError(
@@ -641,7 +669,7 @@ namespace MobileGL::MG_Impl::GLImpl {
return true; return true;
} }
Bool ValidateRenderbufferStorageSamples_State(GLsizei samples, const char* caller) { Bool ValidateRenderbufferStorageSamples_State(GLsizei samples, TextureInternalFormat format, const char* caller) {
if (samples < 0) { if (samples < 0) {
MG_State::pGLContext->RecordError( MG_State::pGLContext->RecordError(
ErrorCode::InvalidValue, ErrorCode::InvalidValue,
@@ -649,9 +677,10 @@ namespace MobileGL::MG_Impl::GLImpl {
return false; return false;
} }
const Int maxSamples = GetMaxRenderbufferSamples_State(); // TODO: Resolve the remaining per-internalformat renderbuffer sample limits once
// glGetInternalformativ is backed; integer formats are handled below.
const Int maxSamples = GetMaxRenderbufferSamplesForFormat_State(format);
if (samples > maxSamples) { if (samples > maxSamples) {
// TODO: Use per-internalformat renderbuffer sample limits once glGetInternalformativ is backed.
// GL 4.6 core 9.2.4 makes asking for more samples than the format supports // GL 4.6 core 9.2.4 makes asking for more samples than the format supports
// INVALID_OPERATION, not INVALID_VALUE - the count is well formed, this format just // INVALID_OPERATION, not INVALID_VALUE - the count is well formed, this format just
// cannot deliver it. Only a negative count is INVALID_VALUE. // cannot deliver it. Only a negative count is INVALID_VALUE.
@@ -659,7 +688,7 @@ namespace MobileGL::MG_Impl::GLImpl {
ErrorCode::InvalidOperation, ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>( MakeUnique<GenericErrorInfo>(
"MG_Impl/GLImpl", caller, "MG_Impl/GLImpl", caller,
std::format("Sample count {} exceeds GL_MAX_SAMPLES ({}).", samples, maxSamples))); std::format("Sample count {} exceeds this format's sample limit ({}).", samples, maxSamples)));
return false; return false;
} }
return true; return true;
@@ -684,7 +713,7 @@ namespace MobileGL::MG_Impl::GLImpl {
TextureInternalFormat format = MG_Util::ConvertGLEnumToTextureInternalFormat(internalformat); TextureInternalFormat format = MG_Util::ConvertGLEnumToTextureInternalFormat(internalformat);
if (!TextureImpl::ValidateTextureInternalFormat(format)) return; if (!TextureImpl::ValidateTextureInternalFormat(format)) return;
if (!ValidateRenderbufferStorageSamples_State(samples, kCaller)) return; if (!ValidateRenderbufferStorageSamples_State(samples, format, kCaller)) return;
if (!ValidateRenderbufferStorageSize_State(width, height, kCaller)) return; if (!ValidateRenderbufferStorageSize_State(width, height, kCaller)) return;
renderbufferObject->AllocateStorage({width, height}); renderbufferObject->AllocateStorage({width, height});
@@ -931,7 +960,8 @@ namespace MobileGL::MG_Impl::GLImpl {
TextureInternalFormat format = MG_Util::ConvertGLEnumToTextureInternalFormat(internalformat); TextureInternalFormat format = MG_Util::ConvertGLEnumToTextureInternalFormat(internalformat);
if (!TextureImpl::ValidateTextureInternalFormat(format)) return; if (!TextureImpl::ValidateTextureInternalFormat(format)) return;
if (!ValidateRenderbufferStorageSamples_State(samples, "NamedRenderbufferStorageMultisample_State")) return; if (!ValidateRenderbufferStorageSamples_State(samples, format, "NamedRenderbufferStorageMultisample_State"))
return;
if (!ValidateRenderbufferStorageSize_State(width, height, "NamedRenderbufferStorageMultisample_State")) return; if (!ValidateRenderbufferStorageSize_State(width, height, "NamedRenderbufferStorageMultisample_State")) return;
renderbufferObject->AllocateStorage({width, height}); renderbufferObject->AllocateStorage({width, height});
+66 -13
View File
@@ -51,6 +51,15 @@ namespace MobileGL::MG_Impl::GLImpl {
constexpr GLint kFrontendMaxTessControlAtomicCounters = 0; constexpr GLint kFrontendMaxTessControlAtomicCounters = 0;
constexpr GLint kFrontendMaxTessEvaluationAtomicCounters = 0; constexpr GLint kFrontendMaxTessEvaluationAtomicCounters = 0;
constexpr GLint kFrontendMaxVertexAtomicCounters = 0; constexpr GLint kFrontendMaxVertexAtomicCounters = 0;
// Zero counters means zero buffers to hold them. These have to be ANSWERED rather than
// left to the default INVALID_ENUM: a well-behaved application queries the limit exactly
// to find out that the stage cannot do this, and an error instead both leaves its output
// untouched (so it reads uninitialised memory and may conclude the opposite) and leaves a
// GL error pending that surfaces at whatever unrelated call checks next.
constexpr GLint kFrontendMaxGeometryAtomicCounterBuffers = 0;
constexpr GLint kFrontendMaxTessControlAtomicCounterBuffers = 0;
constexpr GLint kFrontendMaxTessEvaluationAtomicCounterBuffers = 0;
constexpr GLint kFrontendMaxVertexAtomicCounterBuffers = 0;
// One atomic counter is a uint, and a buffer never has to hold more counters than the // One atomic counter is a uint, and a buffer never has to hold more counters than the
// combined limit the frontend advertises. GL 4.6 table 23.63 floors this at 32 bytes. // combined limit the frontend advertises. GL 4.6 table 23.63 floors this at 32 bytes.
constexpr GLint kFrontendMaxAtomicCounterBufferSize = constexpr GLint kFrontendMaxAtomicCounterBufferSize =
@@ -174,6 +183,30 @@ namespace MobileGL::MG_Impl::GLImpl {
return frontendCount; return frontendCount;
} }
// A per-stage or combined BLOCK count is an amount of indexed binding points an
// application will occupy, and GL 4.6 table 23.64 orders the two accordingly:
// MAX_UNIFORM_BUFFER_BINDINGS >= MAX_COMBINED_UNIFORM_BLOCKS >= every per-stage count,
// and the same for the shader-storage family. The two families are answered from
// unrelated places here - frontend constants, backend dynamic parameters, and a few
// hard-coded TODOs - so nothing kept them ordered, and a backend that reports Vulkan
// descriptor-indexing counts advertised 256 compute uniform blocks over 36 binding
// points. KHR-GL44.multi_bind.dispatch_bind_buffers_base reads the block count and binds
// that many buffers in ONE glBindBuffersBase, which is then INVALID_OPERATION before it
// binds anything. Clamping is the only direction available: the binding count is the
// capacity of the state layer's indexed-binding array, not a number we may inflate.
GLint ClampBlockCountToBindingPoints(GLint blockCount, BufferTarget bufferTarget) {
const GLint bindingPoints = static_cast<GLint>(GetIndexedBufferQueryPointCount(bufferTarget));
return std::min(std::max(blockCount, 0), bindingPoints);
}
GLint ClampUniformBlockCount(GLint blockCount) {
return ClampBlockCountToBindingPoints(blockCount, BufferTarget::Uniform);
}
GLint ClampStorageBlockCount(GLint blockCount) {
return ClampBlockCountToBindingPoints(blockCount, BufferTarget::ShaderStorage);
}
bool TryDecodeDrawBufferQuery(GLenum pname, SizeT& drawBufferIndex) { bool TryDecodeDrawBufferQuery(GLenum pname, SizeT& drawBufferIndex) {
if (pname == GL_DRAW_BUFFER) { if (pname == GL_DRAW_BUFFER) {
drawBufferIndex = 0; drawBufferIndex = 0;
@@ -1397,7 +1430,7 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = kFrontendMaxCombinedAtomicCounters; *params = kFrontendMaxCombinedAtomicCounters;
return; return;
case GL_MAX_COMBINED_UNIFORM_BLOCKS: case GL_MAX_COMBINED_UNIFORM_BLOCKS:
*params = kFrontendMaxCombinedUniformBlocks; *params = ClampUniformBlockCount(kFrontendMaxCombinedUniformBlocks);
return; return;
case GL_MAX_DUAL_SOURCE_DRAW_BUFFERS: case GL_MAX_DUAL_SOURCE_DRAW_BUFFERS:
*params = 1; // TODO *params = 1; // TODO
@@ -1412,7 +1445,7 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = kFrontendMaxFragmentAtomicCounters; *params = kFrontendMaxFragmentAtomicCounters;
return; return;
case GL_MAX_FRAGMENT_SHADER_STORAGE_BLOCKS: case GL_MAX_FRAGMENT_SHADER_STORAGE_BLOCKS:
*params = 16; // TODO *params = ClampStorageBlockCount(16); // TODO
return; return;
case GL_MAX_FRAGMENT_INPUT_COMPONENTS: case GL_MAX_FRAGMENT_INPUT_COMPONENTS:
*params = kFrontendMaxFragmentInputComponents; *params = kFrontendMaxFragmentInputComponents;
@@ -1429,13 +1462,16 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = kFrontendMaxFragmentUniformVectors; *params = kFrontendMaxFragmentUniformVectors;
return; return;
case GL_MAX_FRAGMENT_UNIFORM_BLOCKS: case GL_MAX_FRAGMENT_UNIFORM_BLOCKS:
*params = kFrontendMaxFragmentUniformBlocks; *params = ClampUniformBlockCount(kFrontendMaxFragmentUniformBlocks);
return; return;
case GL_MAX_GEOMETRY_ATOMIC_COUNTERS: case GL_MAX_GEOMETRY_ATOMIC_COUNTERS:
*params = kFrontendMaxGeometryAtomicCounters; *params = kFrontendMaxGeometryAtomicCounters;
return; return;
case GL_MAX_GEOMETRY_ATOMIC_COUNTER_BUFFERS:
*params = kFrontendMaxGeometryAtomicCounterBuffers;
return;
case GL_MAX_GEOMETRY_SHADER_STORAGE_BLOCKS: case GL_MAX_GEOMETRY_SHADER_STORAGE_BLOCKS:
*params = 16; // TODO *params = ClampStorageBlockCount(16); // TODO
return; return;
case GL_MAX_GEOMETRY_INPUT_COMPONENTS: case GL_MAX_GEOMETRY_INPUT_COMPONENTS:
*params = kFrontendMaxGeometryInputComponents; *params = kFrontendMaxGeometryInputComponents;
@@ -1458,7 +1494,7 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = kFrontendMaxGeometryTotalOutputComponents; *params = kFrontendMaxGeometryTotalOutputComponents;
return; return;
case GL_MAX_GEOMETRY_UNIFORM_BLOCKS: case GL_MAX_GEOMETRY_UNIFORM_BLOCKS:
*params = kFrontendMaxGeometryUniformBlocks; *params = ClampUniformBlockCount(kFrontendMaxGeometryUniformBlocks);
return; return;
case GL_MAX_GEOMETRY_UNIFORM_COMPONENTS: case GL_MAX_GEOMETRY_UNIFORM_COMPONENTS:
*params = kFrontendMaxGeometryUniformComponents; *params = kFrontendMaxGeometryUniformComponents;
@@ -1490,9 +1526,15 @@ namespace MobileGL::MG_Impl::GLImpl {
case GL_MAX_TESS_CONTROL_ATOMIC_COUNTERS: case GL_MAX_TESS_CONTROL_ATOMIC_COUNTERS:
*params = kFrontendMaxTessControlAtomicCounters; *params = kFrontendMaxTessControlAtomicCounters;
return; return;
case GL_MAX_TESS_CONTROL_ATOMIC_COUNTER_BUFFERS:
*params = kFrontendMaxTessControlAtomicCounterBuffers;
return;
case GL_MAX_TESS_EVALUATION_ATOMIC_COUNTERS: case GL_MAX_TESS_EVALUATION_ATOMIC_COUNTERS:
*params = kFrontendMaxTessEvaluationAtomicCounters; *params = kFrontendMaxTessEvaluationAtomicCounters;
return; return;
case GL_MAX_TESS_EVALUATION_ATOMIC_COUNTER_BUFFERS:
*params = kFrontendMaxTessEvaluationAtomicCounterBuffers;
return;
case GL_MAX_TESS_CONTROL_IMAGE_UNIFORMS: case GL_MAX_TESS_CONTROL_IMAGE_UNIFORMS:
*params = 0; *params = 0;
return; return;
@@ -1500,10 +1542,10 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = 0; *params = 0;
return; return;
case GL_MAX_TESS_CONTROL_SHADER_STORAGE_BLOCKS: case GL_MAX_TESS_CONTROL_SHADER_STORAGE_BLOCKS:
*params = 16; // TODO *params = ClampStorageBlockCount(16); // TODO
return; return;
case GL_MAX_TESS_EVALUATION_SHADER_STORAGE_BLOCKS: case GL_MAX_TESS_EVALUATION_SHADER_STORAGE_BLOCKS:
*params = 16; // TODO *params = ClampStorageBlockCount(16); // TODO
return; return;
case GL_MAX_TEXTURE_LOD_BIAS: case GL_MAX_TEXTURE_LOD_BIAS:
*params = 15; // TODO *params = 15; // TODO
@@ -1520,13 +1562,16 @@ namespace MobileGL::MG_Impl::GLImpl {
case GL_MAX_VERTEX_ATOMIC_COUNTERS: case GL_MAX_VERTEX_ATOMIC_COUNTERS:
*params = kFrontendMaxVertexAtomicCounters; *params = kFrontendMaxVertexAtomicCounters;
return; return;
case GL_MAX_VERTEX_ATOMIC_COUNTER_BUFFERS:
*params = kFrontendMaxVertexAtomicCounterBuffers;
return;
case GL_MAX_VERTEX_IMAGE_UNIFORMS: case GL_MAX_VERTEX_IMAGE_UNIFORMS:
*params = MG_Backend::pActiveBackendObject *params = MG_Backend::pActiveBackendObject
? MG_Backend::pActiveBackendObject->GetDynamicParameters().MaxVertexImageUniforms ? MG_Backend::pActiveBackendObject->GetDynamicParameters().MaxVertexImageUniforms
: MG_Backend::DynamicBackendParameters{}.MaxVertexImageUniforms; : MG_Backend::DynamicBackendParameters{}.MaxVertexImageUniforms;
return; return;
case GL_MAX_VERTEX_SHADER_STORAGE_BLOCKS: case GL_MAX_VERTEX_SHADER_STORAGE_BLOCKS:
*params = 16; // TODO *params = ClampStorageBlockCount(16); // TODO
return; return;
case GL_MAX_VERTEX_UNIFORM_COMPONENTS: case GL_MAX_VERTEX_UNIFORM_COMPONENTS:
*params = kFrontendMaxVertexUniformComponents; *params = kFrontendMaxVertexUniformComponents;
@@ -1538,7 +1583,7 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = kFrontendMaxVertexOutputComponents; *params = kFrontendMaxVertexOutputComponents;
return; return;
case GL_MAX_VERTEX_UNIFORM_BLOCKS: case GL_MAX_VERTEX_UNIFORM_BLOCKS:
*params = kFrontendMaxVertexUniformBlocks; *params = ClampUniformBlockCount(kFrontendMaxVertexUniformBlocks);
return; return;
case GL_NUM_COMPRESSED_TEXTURE_FORMATS: case GL_NUM_COMPRESSED_TEXTURE_FORMATS:
*params = 0; // compressed texture upload entrypoints are still unimplemented *params = 0; // compressed texture upload entrypoints are still unimplemented
@@ -1938,13 +1983,13 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = dynamicParameters.SubgroupQuadOperationsInAllStages ? GL_TRUE : GL_FALSE; *params = dynamicParameters.SubgroupQuadOperationsInAllStages ? GL_TRUE : GL_FALSE;
break; break;
case GL_MAX_COMPUTE_SHADER_STORAGE_BLOCKS: case GL_MAX_COMPUTE_SHADER_STORAGE_BLOCKS:
*params = dynamicParameters.MaxComputeShaderStorageBlocks; *params = ClampStorageBlockCount(dynamicParameters.MaxComputeShaderStorageBlocks);
break; break;
case GL_MAX_COMBINED_SHADER_STORAGE_BLOCKS: case GL_MAX_COMBINED_SHADER_STORAGE_BLOCKS:
*params = dynamicParameters.MaxCombinedShaderStorageBlocks; *params = ClampStorageBlockCount(dynamicParameters.MaxCombinedShaderStorageBlocks);
break; break;
case GL_MAX_COMPUTE_UNIFORM_BLOCKS: case GL_MAX_COMPUTE_UNIFORM_BLOCKS:
*params = dynamicParameters.MaxComputeUniformBlocks; *params = ClampUniformBlockCount(dynamicParameters.MaxComputeUniformBlocks);
break; break;
case GL_MAX_COMPUTE_TEXTURE_IMAGE_UNITS: case GL_MAX_COMPUTE_TEXTURE_IMAGE_UNITS:
*params = dynamicParameters.MaxComputeTextureImageUnits; *params = dynamicParameters.MaxComputeTextureImageUnits;
@@ -2074,7 +2119,15 @@ namespace MobileGL::MG_Impl::GLImpl {
*params = static_cast<GLint>(GetIndexedBufferQueryPointCount(BufferTarget::AtomicCounter)); *params = static_cast<GLint>(GetIndexedBufferQueryPointCount(BufferTarget::AtomicCounter));
break; break;
case GL_MAX_ATOMIC_COUNTER_BUFFER_SIZE: case GL_MAX_ATOMIC_COUNTER_BUFFER_SIZE:
*params = kFrontendMaxAtomicCounterBufferSize; // The conformance suite splits this evenly across every advertised binding point and
// binds all of them in one glBindBuffersRange
// (KHR-GL44.multi_bind.functional_bind_buffers_range), so the pair has to divide:
// 32 bytes over 36 binding points is a zero-sized range, which BindBufferRange
// rejects with INVALID_VALUE before it binds anything. Floor the advertised size at
// one counter per binding point.
*params = std::max<GLint>(
kFrontendMaxAtomicCounterBufferSize,
static_cast<GLint>(GetIndexedBufferQueryPointCount(BufferTarget::AtomicCounter) * sizeof(GLuint)));
break; break;
case GL_MAX_TEXTURE_BUFFER_SIZE: case GL_MAX_TEXTURE_BUFFER_SIZE:
*params = dynamicParameters.MaxTextureBufferSize; *params = dynamicParameters.MaxTextureBufferSize;
+89 -54
View File
@@ -850,7 +850,12 @@ namespace MobileGL::MG_Impl::GLImpl {
// vector per column - while the value glGetUniform* must return is tightly packed // vector per column - while the value glGetUniform* must return is tightly packed
// columns * rows floats. Only mat4 is the same either way; every other shape needs the // columns * rows floats. Only mat4 is the same either way; every other shape needs the
// padding undone, and the readback has to undo exactly what UniformMatrixfv_Object put // padding undone, and the readback has to undo exactly what UniformMatrixfv_Object put
// there. Returns false when `ttype` is not a float matrix (nothing to unpack). // there. Returns false when there is nothing here to unpack.
//
// A DOUBLE matrix is declined not because it is laid out differently - it is not, the
// demotion makes a dmat4 a mat4 in the shader and a mat4-shaped slot here - but because it
// is ROUTED differently: the caller's component-by-component EbtDouble branch has to widen
// each float back to the queried type, and it undoes the same padding itself.
Bool TryGatherFloatMatrixColumns(const glslang::TType* ttype, const char* pBase, void* params) { Bool TryGatherFloatMatrixColumns(const glslang::TType* ttype, const char* pBase, void* params) {
if (ttype == nullptr || !ttype->isMatrix() || ttype->getBasicType() == glslang::EbtDouble) return false; if (ttype == nullptr || !ttype->isMatrix() || ttype->getBasicType() == glslang::EbtDouble) return false;
const Int columns = ttype->getMatrixCols(); const Int columns = ttype->getMatrixCols();
@@ -909,7 +914,13 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
if (!TryGatherFloatMatrixColumns(ttype, pUBO + offset, params)) { if (!TryGatherFloatMatrixColumns(ttype, pUBO + offset, params)) {
Memcpy(params, pUBO + offset, size); // Never more than the uniform actually occupies. `size` is the GL type size,
// which for a `double` uniform is twice its storage - every 64-bit float is
// narrowed before the module reaches a backend, so the slot holds floats. The
// typed entry points (glGetUniformdv and friends) go through
// GetUniformScalar_State, which converts component by component; this raw
// copy has no type to convert with, so it is bounded rather than converted.
Memcpy(params, pUBO + offset, std::min<SizeT>(size, span));
} }
} }
// TODO: handle 1i variant as texture unit // TODO: handle 1i variant as texture unit
@@ -960,22 +971,27 @@ namespace MobileGL::MG_Impl::GLImpl {
if (TryGatherFloatMatrixColumns(ttype, pUBO + offset, params)) return; if (TryGatherFloatMatrixColumns(ttype, pUBO + offset, params)) return;
} }
// A double-precision uniform is the one case where the stored component type can // A double-precision uniform is the one case where the stored component type differs
// differ from the queried one for a non-opaque uniform, and the difference is not // from the DECLARED one for a non-opaque uniform: the shader's 64-bit floats are
// just a reinterpretation: it is twice as wide, so a raw copy would overrun the // narrowed to 32 bits before the module reaches a backend
// caller's buffer as well as return nonsense. Read component by component and let // (ShaderTranspiler::DemoteFloat64Pass), so what is in the global UBO is a float per
// GL's conversion rules (7.6: round to nearest for the integer queries) apply. // component, laid out exactly like the float-typed twin of this uniform - std140
// 16-byte column stride for a matrix included. Reading it as a GLdouble would return
// two components reinterpreted as one. Read component by component and let GL's
// conversion rules (7.6: round to nearest for the integer queries) apply; the value
// widens back to the queried type, having lost precision at the glUniform*d that
// stored it and not here.
if (ttype->getBasicType() == glslang::EbtDouble) { if (ttype->getBasicType() == glslang::EbtDouble) {
const Int columns = ttype->isMatrix() ? ttype->getMatrixCols() : 1; const Int columns = ttype->isMatrix() ? ttype->getMatrixCols() : 1;
const Int rows = ttype->isMatrix() ? ttype->getMatrixRows() const Int rows = ttype->isMatrix() ? ttype->getMatrixRows()
: (ttype->isVector() ? ttype->getVectorSize() : 1); : (ttype->isVector() ? ttype->getVectorSize() : 1);
// The slot the linker handed out is exactly `columns` columns wide, so it also // std140 gives every matrix column its own 16-byte slot; a non-matrix is one
// states the column stride - which for a double matrix is not a float's 16 bytes. // tightly packed run and never reaches the stride at all.
const SizeT columnStride = columns > 0 ? size / static_cast<SizeT>(columns) : size; const SizeT columnStride = 4 * sizeof(GLfloat);
for (Int column = 0; column < columns; ++column) { for (Int column = 0; column < columns; ++column) {
for (Int row = 0; row < rows; ++row) { for (Int row = 0; row < rows; ++row) {
GLdouble component = 0.0; GLfloat component = 0.0f;
Memcpy(&component, pUBO + offset + column * columnStride + row * sizeof(GLdouble), Memcpy(&component, pUBO + offset + column * columnStride + row * sizeof(GLfloat),
sizeof(component)); sizeof(component));
if constexpr (std::is_integral_v<T>) { if constexpr (std::is_integral_v<T>) {
// Rounded to the nearest integer and clamped into the queried type's // Rounded to the nearest integer and clamped into the queried type's
@@ -1248,36 +1264,39 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
} }
// glUniform*d / glUniformMatrix*dv. The vector forms need nothing beyond the shared // glUniform*d / glUniformMatrix*dv. Neither needs a layout of its own any more: the
// upload template - it is already typed on the component - but a matrix does: the // transpile chain narrows every 64-bit float in the shader to 32 bits
// column stride the linker used for a double matrix is not the 16 bytes a float one // (ShaderTranspiler::DemoteFloat64Pass) and the global UBO is laid out by reflecting that
// gets. It is not guessed here; the slot the uniform was given is exactly `columns` // demoted module, so a double uniform's storage IS a float uniform's - same offset, same
// columns wide, so dividing states the stride the rest of the pipeline agreed on. // 4-byte components, same std140 column padding for matrices. Narrowing here, at the one
template <typename Program> // place the 64-bit value enters, and then handing the bytes to the ordinary float upload
void UniformMatrixdv_Object(Program& programObject, GLint location, GLsizei count, GLboolean transpose, // path is what keeps the two in step; a separate double-shaped layout here would write
const GLdouble* value, Int columns, Int rows) { // 8-byte components into 4-byte slots and silently address the wrong ones.
const SizeT slotSize = programObject.GetUniformSizesInBytes(location); //
const SizeT columnStride = columns > 0 ? slotSize / static_cast<SizeT>(columns) : slotSize; // The narrowing is the same static_cast the shader's own arithmetic now performs, so the
const SizeT componentCount = static_cast<SizeT>(columns) * static_cast<SizeT>(rows); // value the shader reads is the value glUniform*d was given, at float precision.
Vector<GLdouble> column(static_cast<SizeT>(rows)); template <GLsizei ItemCount>
for (GLint matrix = 0; matrix < count; ++matrix) { void UniformvNarrowed_State(GLint location, GLsizei count, const GLdouble* value) {
if (matrix > 0 && !programObject.UniformLocationsAliasSameUniform(location, location + matrix)) break; if (value == nullptr || count <= 0) {
if (!programObject.IsValidUniformLocation(location + matrix)) { // Same shape as the float entry points: the location validation still runs, and a
RecordInvalidUniformLocationError(__func__, location + matrix, "the current program object"); // null pointer is left to fault exactly where glUniform*fv would.
Uniformv_State<ItemCount>(location, count, reinterpret_cast<const GLfloat*>(value));
return; return;
} }
const GLdouble* source = value + matrix * componentCount; Vector<GLfloat> narrowed(static_cast<SizeT>(count) * ItemCount);
for (Int c = 0; c < columns; ++c) { for (SizeT i = 0; i < narrowed.size(); ++i) narrowed[i] = static_cast<GLfloat>(value[i]);
for (Int r = 0; r < rows; ++r) { Uniformv_State<ItemCount>(location, count, narrowed.data());
column[r] = transpose == GL_TRUE ? source[r * columns + c] : source[c * rows + r];
}
Uniform_State<1>(programObject, location + matrix, column.data(), c * columnStride);
for (Int r = 1; r < rows; ++r) {
Uniform_State<1>(programObject, location + matrix, column.data() + r,
c * columnStride + r * sizeof(GLdouble));
}
} }
template <GLsizei ItemCount>
void ProgramUniformvNarrowed_State(GLuint program, GLint location, GLsizei count, const GLdouble* value) {
if (value == nullptr || count <= 0) {
ProgramUniformv_State<ItemCount>(program, location, count, reinterpret_cast<const GLfloat*>(value));
return;
} }
Vector<GLfloat> narrowed(static_cast<SizeT>(count) * ItemCount);
for (SizeT i = 0; i < narrowed.size(); ++i) narrowed[i] = static_cast<GLfloat>(value[i]);
ProgramUniformv_State<ItemCount>(program, location, count, narrowed.data());
} }
// glUniformMatrix*fv / glProgramUniformMatrix*fv, every shape (square and non-square). // glUniformMatrix*fv / glProgramUniformMatrix*fv, every shape (square and non-square).
@@ -1326,6 +1345,22 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
} }
// glUniformMatrix*dv / glProgramUniformMatrix*dv. Narrowed to the float form and handed
// straight to it: after DemoteFloat64Pass a `dmat4` uniform is a `mat4` in the shader and a
// mat4-shaped slot in the global UBO, columns padded to a vec4 and all. Everything else
// about the call - transpose handling, the array-element walk, the opaque-uniform refusal -
// is then the one implementation both spellings share.
template <typename Program>
void UniformMatrixdv_Object(Program& programObject, GLint location, GLsizei count, GLboolean transpose,
const GLdouble* value, Int columns, Int rows) {
if (value == nullptr || count <= 0) return;
const SizeT componentCount = static_cast<SizeT>(columns) * static_cast<SizeT>(rows);
Vector<GLfloat> narrowed(static_cast<SizeT>(count) * componentCount);
for (SizeT i = 0; i < narrowed.size(); ++i) narrowed[i] = static_cast<GLfloat>(value[i]);
UniformMatrixfv_Object(programObject, "glUniformMatrixdv", location, count, transpose, narrowed.data(),
columns, rows, "the current program object");
}
// Helper function to transpose a 2x2 matrix // Helper function to transpose a 2x2 matrix
void TransposeMatrix2x2(const GLfloat* input, GLfloat* output) { void TransposeMatrix2x2(const GLfloat* input, GLfloat* output) {
// Input matrix is in column-major order (OpenGL default) // Input matrix is in column-major order (OpenGL default)
@@ -2089,71 +2124,71 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
void Uniform1d(GLint location, GLdouble v0) { void Uniform1d(GLint location, GLdouble v0) {
const GLdouble v[] = {v0}; const GLdouble v[] = {v0};
Uniformv_State<1>(location, 1, v); UniformvNarrowed_State<1>(location, 1, v);
} }
void Uniform1dv(GLint location, GLsizei count, const GLdouble* value) { void Uniform1dv(GLint location, GLsizei count, const GLdouble* value) {
Uniformv_State<1>(location, count, value); UniformvNarrowed_State<1>(location, count, value);
} }
void ProgramUniform1d(GLuint program, GLint location, GLdouble v0) { void ProgramUniform1d(GLuint program, GLint location, GLdouble v0) {
const GLdouble v[] = {v0}; const GLdouble v[] = {v0};
ProgramUniformv_State<1>(program, location, 1, v); ProgramUniformvNarrowed_State<1>(program, location, 1, v);
} }
void ProgramUniform1dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) { void ProgramUniform1dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) {
ProgramUniformv_State<1>(program, location, count, value); ProgramUniformvNarrowed_State<1>(program, location, count, value);
} }
void Uniform2d(GLint location, GLdouble v0, GLdouble v1) { void Uniform2d(GLint location, GLdouble v0, GLdouble v1) {
const GLdouble v[] = {v0, v1}; const GLdouble v[] = {v0, v1};
Uniformv_State<2>(location, 1, v); UniformvNarrowed_State<2>(location, 1, v);
} }
void Uniform2dv(GLint location, GLsizei count, const GLdouble* value) { void Uniform2dv(GLint location, GLsizei count, const GLdouble* value) {
Uniformv_State<2>(location, count, value); UniformvNarrowed_State<2>(location, count, value);
} }
void ProgramUniform2d(GLuint program, GLint location, GLdouble v0, GLdouble v1) { void ProgramUniform2d(GLuint program, GLint location, GLdouble v0, GLdouble v1) {
const GLdouble v[] = {v0, v1}; const GLdouble v[] = {v0, v1};
ProgramUniformv_State<2>(program, location, 1, v); ProgramUniformvNarrowed_State<2>(program, location, 1, v);
} }
void ProgramUniform2dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) { void ProgramUniform2dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) {
ProgramUniformv_State<2>(program, location, count, value); ProgramUniformvNarrowed_State<2>(program, location, count, value);
} }
void Uniform3d(GLint location, GLdouble v0, GLdouble v1, GLdouble v2) { void Uniform3d(GLint location, GLdouble v0, GLdouble v1, GLdouble v2) {
const GLdouble v[] = {v0, v1, v2}; const GLdouble v[] = {v0, v1, v2};
Uniformv_State<3>(location, 1, v); UniformvNarrowed_State<3>(location, 1, v);
} }
void Uniform3dv(GLint location, GLsizei count, const GLdouble* value) { void Uniform3dv(GLint location, GLsizei count, const GLdouble* value) {
Uniformv_State<3>(location, count, value); UniformvNarrowed_State<3>(location, count, value);
} }
void ProgramUniform3d(GLuint program, GLint location, GLdouble v0, GLdouble v1, GLdouble v2) { void ProgramUniform3d(GLuint program, GLint location, GLdouble v0, GLdouble v1, GLdouble v2) {
const GLdouble v[] = {v0, v1, v2}; const GLdouble v[] = {v0, v1, v2};
ProgramUniformv_State<3>(program, location, 1, v); ProgramUniformvNarrowed_State<3>(program, location, 1, v);
} }
void ProgramUniform3dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) { void ProgramUniform3dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) {
ProgramUniformv_State<3>(program, location, count, value); ProgramUniformvNarrowed_State<3>(program, location, count, value);
} }
void Uniform4d(GLint location, GLdouble v0, GLdouble v1, GLdouble v2, GLdouble v3) { void Uniform4d(GLint location, GLdouble v0, GLdouble v1, GLdouble v2, GLdouble v3) {
const GLdouble v[] = {v0, v1, v2, v3}; const GLdouble v[] = {v0, v1, v2, v3};
Uniformv_State<4>(location, 1, v); UniformvNarrowed_State<4>(location, 1, v);
} }
void Uniform4dv(GLint location, GLsizei count, const GLdouble* value) { void Uniform4dv(GLint location, GLsizei count, const GLdouble* value) {
Uniformv_State<4>(location, count, value); UniformvNarrowed_State<4>(location, count, value);
} }
void ProgramUniform4d(GLuint program, GLint location, GLdouble v0, GLdouble v1, GLdouble v2, GLdouble v3) { void ProgramUniform4d(GLuint program, GLint location, GLdouble v0, GLdouble v1, GLdouble v2, GLdouble v3) {
const GLdouble v[] = {v0, v1, v2, v3}; const GLdouble v[] = {v0, v1, v2, v3};
ProgramUniformv_State<4>(program, location, 1, v); ProgramUniformvNarrowed_State<4>(program, location, 1, v);
} }
void ProgramUniform4dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) { void ProgramUniform4dv(GLuint program, GLint location, GLsizei count, const GLdouble* value) {
ProgramUniformv_State<4>(program, location, count, value); ProgramUniformvNarrowed_State<4>(program, location, count, value);
} }
void UniformMatrix2dv(GLint location, GLsizei count, GLboolean transpose, const GLdouble* value) { void UniformMatrix2dv(GLint location, GLsizei count, GLboolean transpose, const GLdouble* value) {
if (location == -1) return; if (location == -1) return;
@@ -210,11 +210,66 @@ namespace MobileGL::MG_Impl::GLImpl::ProgramInterface {
// ---- model construction -------------------------------------------------------- // ---- model construction --------------------------------------------------------
// GL_REFERENCED_BY_*_SHADER for an ARRAYED block instance, refined per element.
//
// glslang records a block reference by walking up to the base symbol and calling
// addBlockName with the whole ARRAY type, which ORs the referencing stage into every
// element at once - it has not resolved the subscript yet at that point. So reading
// "e[0].b" marks both TrickyBlock[0] and TrickyBlock[1] as referenced by the fragment
// stage (KHR-GL43.program_interface_query.uniform-block-types).
//
// The MEMBER masks are exact: EShReflectionAllBlockVariables enumerates every member of
// every element with the stage mask suppressed, and only the dereference chain actually
// walked turns a bit on - and that chain carries the subscript. So the union of a block
// instance's members is the reference set of that instance.
//
// Applied ONLY to arrayed instances, because for a scalar block glslang is already exact.
// Note the union is used even when it is empty: an array element nobody dereferenced has
// no member bits and is genuinely referenced by nobody, which is the whole point - falling
// back to the block's own mask there would restore the over-approximation.
Vector<Uint32> BuildBlockStagesFromMembers(const glslang::TProgram& reflection, Int blockCount) {
auto& mutableReflection = const_cast<glslang::TProgram&>(reflection);
Vector<Uint32> stagesByBlock(static_cast<SizeT>(blockCount < 0 ? 0 : blockCount), 0u);
const Int uniformCount = mutableReflection.getNumUniformVariables();
for (Int index = 0; index < uniformCount; ++index) {
const auto& uniform = mutableReflection.getUniform(index);
const Int owner = uniform.index;
if (owner < 0 || owner >= blockCount) continue;
stagesByBlock[static_cast<SizeT>(owner)] |= static_cast<Uint32>(uniform.stages);
}
return stagesByBlock;
}
// UNIFORM blocks only, and that scope is load-bearing rather than cautious. The member
// names glslang produces for a uniform block array carry the subscript
// ("TrickyBlock[0].b", via EShReflectionStrictArraySuffix), so each element's members are
// distinct entries and the bits land on the right one. A SHADER STORAGE block array does
// NOT get that treatment - its buffer variables reflect under one subscript-free spelling
// shared by every element - so a union over them credits element 0 and starves the rest.
// KHR-GL43.program_interface_query.ssb-types is the case that says so: it reads ss[0] and
// ss[1] and requires both to report the fragment stage, which only glslang's own
// (deliberately over-approximating) block mask gets right. Storage and atomic-counter
// blocks therefore keep that mask untouched.
Uint32 UniformBlockStages(const glslang::TObjectReflection& block, const Vector<Uint32>& stagesFromMembers,
Int tIndex) {
String arrayBase;
Uint element = 0;
Bool malformed = false;
if (!SplitTrailingSubscript(block.name, arrayBase, element, malformed) || malformed) {
return static_cast<Uint32>(block.stages);
}
if (tIndex < 0 || tIndex >= static_cast<Int>(stagesFromMembers.size())) {
return static_cast<Uint32>(block.stages);
}
return stagesFromMembers[static_cast<SizeT>(tIndex)];
}
void BuildBlocks(ProgramObject& program, const glslang::TProgram& reflection, Model& model, void BuildBlocks(ProgramObject& program, const glslang::TProgram& reflection, Model& model,
Vector<BlockKind>& blockKind, Vector<Int>& blockInterfaceIndex) { Vector<BlockKind>& blockKind, Vector<Int>& blockInterfaceIndex) {
const Int blockCount = const_cast<glslang::TProgram&>(reflection).getNumUniformBlocks(); const Int blockCount = const_cast<glslang::TProgram&>(reflection).getNumUniformBlocks();
blockKind.assign(blockCount, BlockKind::Uniform); blockKind.assign(blockCount, BlockKind::Uniform);
blockInterfaceIndex.assign(blockCount, -1); blockInterfaceIndex.assign(blockCount, -1);
const Vector<Uint32> stagesFromMembers = BuildBlockStagesFromMembers(reflection, blockCount);
for (Int tIndex = 0; tIndex < blockCount; ++tIndex) { for (Int tIndex = 0; tIndex < blockCount; ++tIndex) {
const auto& block = const_cast<glslang::TProgram&>(reflection).getUniformBlock(tIndex); const auto& block = const_cast<glslang::TProgram&>(reflection).getUniformBlock(tIndex);
@@ -260,8 +315,8 @@ namespace MobileGL::MG_Impl::GLImpl::ProgramInterface {
resource.bufferDataSize = static_cast<GLint>(program.GetUBOSizeAt(glIndex)); resource.bufferDataSize = static_cast<GLint>(program.GetUBOSizeAt(glIndex));
const Int tIndex = program.TProgramBlockIndex(static_cast<Uint>(glIndex)); const Int tIndex = program.TProgramBlockIndex(static_cast<Uint>(glIndex));
if (tIndex >= 0 && tIndex < blockCount) { if (tIndex >= 0 && tIndex < blockCount) {
resource.stages = resource.stages = UniformBlockStages(const_cast<glslang::TProgram&>(reflection).getUniformBlock(tIndex),
static_cast<Uint32>(const_cast<glslang::TProgram&>(reflection).getUniformBlock(tIndex).stages); stagesFromMembers, tIndex);
} }
model.uniformBlocks.push_back(Move(resource)); model.uniformBlocks.push_back(Move(resource));
} }
@@ -352,6 +407,17 @@ namespace MobileGL::MG_Impl::GLImpl::ProgramInterface {
} }
} }
// A built-in interface block that a shader redeclares with fewer members keeps the
// omitted ones in its type when the redeclaration is ANONYMOUS - glslang hides them
// (basic type void) instead of erasing them, because the original shared declaration
// has to stay usable. Only the instance-named form erases. So a separable vertex
// program that redeclares `out gl_PerVertex { vec4 gl_Position; }` still carries
// gl_PointSize and gl_ClipDistance through the block-unwrapping reflection, and they
// are not part of its output interface.
Bool IsHiddenBlockMember(const glslang::TType* type) {
return type != nullptr && type->getBasicType() == glslang::EbtVoid;
}
void BuildStageIO(ProgramObject& program, const glslang::TProgram& reflection, Model& model) { void BuildStageIO(ProgramObject& program, const glslang::TProgram& reflection, Model& model) {
auto& mutableReflection = const_cast<glslang::TProgram&>(reflection); auto& mutableReflection = const_cast<glslang::TProgram&>(reflection);
@@ -359,6 +425,7 @@ namespace MobileGL::MG_Impl::GLImpl::ProgramInterface {
for (Int index = 0; index < inputCount; ++index) { for (Int index = 0; index < inputCount; ++index) {
const auto& refl = mutableReflection.getPipeInput(index); const auto& refl = mutableReflection.getPipeInput(index);
const glslang::TType* type = refl.getType(); const glslang::TType* type = refl.getType();
if (IsHiddenBlockMember(type)) continue;
Resource resource; Resource resource;
// The Vulkan-semantics parse reflects the vertex builtins under their SPIR-V // The Vulkan-semantics parse reflects the vertex builtins under their SPIR-V
// names; GL enumerates the GL spellings. // names; GL enumerates the GL spellings.
@@ -373,18 +440,28 @@ namespace MobileGL::MG_Impl::GLImpl::ProgramInterface {
model.programInputs.push_back(Move(resource)); model.programInputs.push_back(Move(resource));
} }
// A color number, and therefore a color INDEX, exists only for a fragment stage's
// outputs. The output interface belongs to the program's last stage, so for a
// separable tessellation/geometry/vertex program these are varyings: asking the
// frag-data maps about them can still answer a location (a tess-control output
// carries its own layout(location=N)), and a location then manufactures a color
// index of 0 where GL requires -1
// (KHR-GL43.program_interface_query.separate-programs-tess-control).
const Bool lastStageIsFragment = mutableReflection.getIntermediate(EShLangFragment) != nullptr;
const Int outputCount = mutableReflection.getNumPipeOutputs(); const Int outputCount = mutableReflection.getNumPipeOutputs();
for (Int index = 0; index < outputCount; ++index) { for (Int index = 0; index < outputCount; ++index) {
const auto& refl = mutableReflection.getPipeOutput(index); const auto& refl = mutableReflection.getPipeOutput(index);
const glslang::TType* type = refl.getType(); const glslang::TType* type = refl.getType();
if (IsHiddenBlockMember(type)) continue;
Resource resource; Resource resource;
resource.name = WithArraySuffix(refl.name, type); resource.name = WithArraySuffix(refl.name, type);
resource.type = static_cast<GLenum>(refl.glDefineType); resource.type = static_cast<GLenum>(refl.glDefineType);
resource.arraySize = ArraySizeOf(type, refl.size); resource.arraySize = ArraySizeOf(type, refl.size);
resource.location = MappedLocation(program.GetFragmentDataLocation(refl.name.c_str())); resource.location = MappedLocation(program.GetFragmentDataLocation(refl.name.c_str()));
if (resource.location < 0) { if (resource.location < 0 || !lastStageIsFragment) {
// A built-in output (gl_FragDepth, gl_SampleMask) and a non-fragment stage // A built-in output (gl_FragDepth, gl_SampleMask) has no location, and a
// output both have no location, and therefore no color index either. // non-fragment stage's outputs have no color number at all - either way there
// is no color index.
resource.locationIndex = -1; resource.locationIndex = -1;
} else { } else {
resource.locationIndex = program.GetFragmentDataIndex(refl.name.c_str()); resource.locationIndex = program.GetFragmentDataIndex(refl.name.c_str());
@@ -9,6 +9,7 @@
#include "GL_Sampler.h" #include "GL_Sampler.h"
#include "Validators.h" #include "Validators.h"
#include "../Getter/GL_Getter.h" #include "../Getter/GL_Getter.h"
#include "../Texture/GL_Texture.h"
#include <MG_State/GLState/Core.h> #include <MG_State/GLState/Core.h>
#include <MG_Util/Converters/GLToMG/TextureEnumConverter.h> #include <MG_Util/Converters/GLToMG/TextureEnumConverter.h>
#include <MG_Util/Converters/MGToGL/TextureEnumConverter.h> #include <MG_Util/Converters/MGToGL/TextureEnumConverter.h>
@@ -269,15 +270,13 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
} }
// The number of texture units a sampler may be bound to. GL 3.3 core 3.8.2 names // The number of texture units a sampler may be bound to is the same count a TEXTURE may be
// GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS, which is what the backend advertises; the frontend's // bound to - GL 3.3 core 3.8.2 names GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS for both - so it is
// MAX_TEXTURE_IMAGE_UNITS is only the capacity of the unit array, so it is a clamp on the // computed once, in GetCombinedTextureImageUnitCount, and named here for the sampler-side
// answer and never the answer itself - gating on it alone accepts every unit up to 192 no // readers below. Two copies of that arithmetic is how glBindSamplers and glBindTextures would
// matter what the driver reports. // come to disagree about which units exist.
static GLint GetSamplerBindableTextureUnitCount() { static GLint GetSamplerBindableTextureUnitCount() {
GLint maxTextureUnits = 0; return GetCombinedTextureImageUnitCount();
GetIntegerv(GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS, &maxTextureUnits);
return std::min<GLint>(std::max(maxTextureUnits, 0), MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS);
} }
void BindSampler_State(GLuint unit, GLuint sampler) { void BindSampler_State(GLuint unit, GLuint sampler) {
@@ -4549,6 +4549,132 @@ namespace MobileGL::MG_Impl::GLImpl {
MG_State::pGLContext->NoteTextureUnitTouched(static_cast<Int>(unit), changed); MG_State::pGLContext->NoteTextureUnitTouched(static_cast<Int>(unit), changed);
} }
GLint GetCombinedTextureImageUnitCount() {
GLint maxTextureUnits = 0;
GetIntegerv(GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS, &maxTextureUnits);
return std::min<GLint>(std::max(maxTextureUnits, 0), MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS);
}
namespace {
// ARB_multi_bind checks the whole [first, first + count) range before binding anything and
// reports an overrun as INVALID_OPERATION - not the INVALID_VALUE the single-bind entry
// points report for an out-of-range unit, and not after binding the in-range prefix.
Bool ValidateMultiBindUnitRange(GLuint first, GLsizei count, GLint unitCount, const char* funcName) {
if (count < 0) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidValue,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", funcName, "count must be non-negative."));
return false;
}
if (static_cast<Uint64>(first) + static_cast<Uint64>(count) > static_cast<Uint64>(unitCount)) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", funcName,
std::format("first + count ({} + {}) exceeds the {} available units.",
first, count, unitCount)));
return false;
}
return true;
}
// ARB_multi_bind states the equivalence to a loop of single binds "except that <textures>
// will not be created if they do not exist": glBindTexture instantiates a name GenTextures
// merely reserved, the multi-bind entry points must refuse it. The error class is
// INVALID_OPERATION for both of them, where the scalar glBindImageTexture reports
// INVALID_VALUE - hence the check here rather than inside BindImageTexture.
//
// Deliberately PER ELEMENT: the extension defines these calls as a loop, so a bad entry
// costs its own unit and leaves the rest of the range bound.
SharedPtr<MG_State::GLState::ITextureObject> ResolveMultiBindTexture(GLuint texture, GLsizei index,
const char* funcName) {
SharedPtr<MG_State::GLState::ITextureObject> textureObject =
MG_State::pGLContext->GetTextureObject(texture);
if (!textureObject) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>(
"MG_Impl/GLImpl", funcName,
std::format("textures[{}] ({}) is not the name of an existing texture object.", index,
texture)));
}
return textureObject;
}
// ARB_multi_bind: an element naming texture zero unbinds EVERY target of its unit, i.e.
// rebinds each target's default texture object - the unit's initial state. Same rule
// glBindTextureUnit(unit, 0) follows.
void UnbindAllTargetsOnUnit(Int unit) {
auto& textureUnit = MG_State::pGLContext->GetTextureUnitObject(unit);
Bool changed = false;
for (auto& slot : textureUnit.GetAllBindingSlots()) {
if (slot.Bind(MG_State::pGLContext->GetDefaultTextureObject(slot.GetTarget()))) changed = true;
}
MG_State::pGLContext->NoteTextureUnitTouched(unit, changed);
}
} // namespace
// ARB_multi_bind: glBindTextures binds each texture to ITS OWN target on unit <first> + i, so
// there is no target parameter and no way to express it through glBindTexture - the per-unit,
// by-object form glBindTextureUnit uses is the one that matches. A NULL <textures> unbinds the
// whole range.
void BindTextures(GLuint first, GLsizei count, const GLuint* textures) {
if (!ValidateMultiBindUnitRange(first, count, GetCombinedTextureImageUnitCount(), __func__)) return;
for (GLsizei i = 0; i < count; ++i) {
const GLuint texture = textures ? textures[i] : 0;
const Int unit = static_cast<Int>(first) + i;
if (texture == 0) {
UnbindAllTargetsOnUnit(unit);
continue;
}
const SharedPtr<MG_State::GLState::ITextureObject> textureObject =
ResolveMultiBindTexture(texture, i, __func__);
if (!textureObject) continue;
auto& textureUnit = MG_State::pGLContext->GetTextureUnitObject(unit);
const Bool changed = textureUnit.GetBindingSlot(textureObject->GetTarget()).Bind(textureObject);
MG_State::pGLContext->NoteTextureUnitTouched(unit, changed);
}
}
// ARB_multi_bind: glBindImageTextures is a loop of glBindImageTexture with every parameter but
// the unit and the texture fixed by the spec - level 0, layered, layer 0, READ_WRITE, and the
// texture's own internal format. An element that names texture zero resets the unit.
void BindImageTextures(GLuint first, GLsizei count, const GLuint* textures) {
if (!ValidateMultiBindUnitRange(first, count, static_cast<GLint>(GetAdvertisedImageUnitCount()), __func__)) {
return;
}
for (GLsizei i = 0; i < count; ++i) {
const GLuint texture = textures ? textures[i] : 0;
const GLuint unit = first + static_cast<GLuint>(i);
if (texture == 0) {
BindImageTexture(unit, 0, 0, GL_FALSE, 0, GL_READ_ONLY, GL_R8);
continue;
}
const SharedPtr<MG_State::GLState::ITextureObject> textureObject =
ResolveMultiBindTexture(texture, i, __func__);
if (!textureObject) continue;
// "An INVALID_OPERATION error is generated if the internal format of any texture is not
// supported for image textures" - a texture that has never been given storage has no
// format at all and lands here too, rather than being reported as a bad enum by the
// scalar path.
const GLenum format = MG_Util::ConvertTextureInternalFormatToGLEnum(textureObject->GetFormat());
if (!IsValidImageTextureFormat(format)) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>(
"MG_Impl/GLImpl", __func__,
std::format("textures[{}] ({}) has an internal format that is not supported for image "
"textures.",
i, texture)));
continue;
}
BindImageTexture(unit, texture, 0, GL_TRUE, 0, GL_READ_WRITE, format);
}
}
void GetTextureImage(GLuint texture, GLint level, GLenum format, GLenum type, GLsizei bufSize, void* pixels) { void GetTextureImage(GLuint texture, GLint level, GLenum format, GLenum type, GLsizei bufSize, void* pixels) {
auto textureObject = GetTextureObjectByName(texture, __func__); auto textureObject = GetTextureObjectByName(texture, __func__);
if (!textureObject) return; if (!textureObject) return;
@@ -132,5 +132,11 @@ namespace MobileGL::MG_Impl::GLImpl {
void CompressedTexImage1D(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLint border, void CompressedTexImage1D(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLint border,
GLsizei imageSize, const void* data); GLsizei imageSize, const void* data);
void BindTexture(GLenum target, GLuint texture); void BindTexture(GLenum target, GLuint texture);
void BindTextures(GLuint first, GLsizei count, const GLuint* textures);
void BindImageTextures(GLuint first, GLsizei count, const GLuint* textures);
void ActiveTexture(GLenum texture); void ActiveTexture(GLenum texture);
// The number of texture image units a texture or a sampler may be bound to: what the backend
// advertises as GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS, clamped by the frontend's fixed unit-array
// capacity. Shared so the texture and sampler multi-bind range checks cannot drift apart.
GLint GetCombinedTextureImageUnitCount();
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
@@ -315,9 +315,10 @@ namespace MobileGL::MG_Impl::GLImpl {
auto offset = reinterpret_cast<SizeT>(pointer); auto offset = reinterpret_cast<SizeT>(pointer);
vao->SetAttributeFormat(index, size, dataType, false, stride, offset, true, false); const int effectiveStride = EffectiveVertexStride(stride, size, type);
vao->SetAttributeFormat(index, size, dataType, false, stride, offset, true, false, effectiveStride);
vao->BindAttributeBuffer(index, vbo); vao->BindAttributeBuffer(index, vbo);
vao->MirrorPointerIntoBinding(index, vbo, offset, EffectiveVertexStride(stride, size, type)); vao->MirrorPointerIntoBinding(index, vbo, offset, effectiveStride);
} }
void VertexAttribPointer_State(GLuint index, GLint size, GLenum type, GLboolean normalized, GLsizei stride, void VertexAttribPointer_State(GLuint index, GLint size, GLenum type, GLboolean normalized, GLsizei stride,
@@ -345,9 +346,11 @@ namespace MobileGL::MG_Impl::GLImpl {
// backend can pick the reversed VkFormat / pass GL_BGRA through to a GLES driver. // backend can pick the reversed VkFormat / pass GL_BGRA through to a GLES driver.
const bool isBgra = (size == static_cast<GLint>(GL_BGRA)); const bool isBgra = (size == static_cast<GLint>(GL_BGRA));
const int effectiveSize = isBgra ? 4 : size; const int effectiveSize = isBgra ? 4 : size;
vao->SetAttributeFormat(index, effectiveSize, dataType, normalized, stride, offset, false, isBgra); const int effectiveStride = EffectiveVertexStride(stride, effectiveSize, type);
vao->SetAttributeFormat(index, effectiveSize, dataType, normalized, stride, offset, false, isBgra,
effectiveStride);
vao->BindAttributeBuffer(index, vbo); vao->BindAttributeBuffer(index, vbo);
vao->MirrorPointerIntoBinding(index, vbo, offset, EffectiveVertexStride(stride, effectiveSize, type)); vao->MirrorPointerIntoBinding(index, vbo, offset, effectiveStride);
} }
void BindVertexArray_State(GLuint array) { void BindVertexArray_State(GLuint array) {
@@ -50,6 +50,7 @@ add_executable(MobileGLIntegrationTest
Scenarios/CrossFrameBufferScenario.cpp Scenarios/CrossFrameBufferScenario.cpp
Scenarios/ResidentIndexScenario.cpp Scenarios/ResidentIndexScenario.cpp
Scenarios/MultiDrawScenario.cpp Scenarios/MultiDrawScenario.cpp
Scenarios/DrawParametersScenario.cpp
Scenarios/AsyncCompileScenario.cpp Scenarios/AsyncCompileScenario.cpp
Scenarios/XfbAfterClipDistanceScenario.cpp Scenarios/XfbAfterClipDistanceScenario.cpp
Scenarios/ThreeChannelAttachmentScenario.cpp Scenarios/ThreeChannelAttachmentScenario.cpp
@@ -60,6 +61,7 @@ add_executable(MobileGLIntegrationTest
Scenarios/ClearThenReadPixelsScenario.cpp Scenarios/ClearThenReadPixelsScenario.cpp
Scenarios/DepthStencilReadbackScenario.cpp Scenarios/DepthStencilReadbackScenario.cpp
Scenarios/SsboArrayLengthScenario.cpp Scenarios/SsboArrayLengthScenario.cpp
Scenarios/DoublePrecisionScenario.cpp
Scenarios/UniformInitializerScenario.cpp Scenarios/UniformInitializerScenario.cpp
Scenarios/SwizzleAccessRoutineScenario.cpp Scenarios/SwizzleAccessRoutineScenario.cpp
Scenarios/ProgramPipelineScenario.cpp Scenarios/ProgramPipelineScenario.cpp
@@ -68,6 +70,7 @@ add_executable(MobileGLIntegrationTest
Scenarios/Glsl420DeclarationScenario.cpp Scenarios/Glsl420DeclarationScenario.cpp
Scenarios/FragmentOutputArrayIndexScenario.cpp Scenarios/FragmentOutputArrayIndexScenario.cpp
Scenarios/BufferTextureScenario.cpp Scenarios/BufferTextureScenario.cpp
Scenarios/VertexAttribBindingScenario.cpp
) )
target_include_directories(MobileGLIntegrationTest PRIVATE target_include_directories(MobileGLIntegrationTest PRIVATE
@@ -94,6 +94,89 @@ namespace MGITest {
} }
} }
// A per-stage block count is an amount of BINDING POINTS an application will use, so it
// can never exceed the number of binding points that exist. GL 4.6 Table 23.64 states the
// relation the other way round (MAX_UNIFORM_BUFFER_BINDINGS >= MAX_COMBINED_UNIFORM_BLOCKS
// >= every per-stage count), and DirectVulkan broke it by clamping the two families
// independently: a device reporting 256 compute uniform blocks and 84 uniform binding
// points passes both ceilings and still cannot serve
// KHR-GL44.multi_bind.dispatch_bind_buffers_base, which reads the block count and binds
// that many buffers in one glBindBuffersBase - INVALID_OPERATION before a single bind.
TEST_F(AdvertisedLimitsScenario, PerStageBlockCountsFitInTheirBindingPoints) {
struct Relation {
GLenum blocks;
const char* blocksName;
GLenum bindings;
const char* bindingsName;
};
const Relation relations[] = {
{GL_MAX_COMPUTE_UNIFORM_BLOCKS, "GL_MAX_COMPUTE_UNIFORM_BLOCKS", GL_MAX_UNIFORM_BUFFER_BINDINGS,
"GL_MAX_UNIFORM_BUFFER_BINDINGS"},
{GL_MAX_VERTEX_UNIFORM_BLOCKS, "GL_MAX_VERTEX_UNIFORM_BLOCKS", GL_MAX_UNIFORM_BUFFER_BINDINGS,
"GL_MAX_UNIFORM_BUFFER_BINDINGS"},
{GL_MAX_FRAGMENT_UNIFORM_BLOCKS, "GL_MAX_FRAGMENT_UNIFORM_BLOCKS", GL_MAX_UNIFORM_BUFFER_BINDINGS,
"GL_MAX_UNIFORM_BUFFER_BINDINGS"},
{GL_MAX_COMBINED_UNIFORM_BLOCKS, "GL_MAX_COMBINED_UNIFORM_BLOCKS", GL_MAX_UNIFORM_BUFFER_BINDINGS,
"GL_MAX_UNIFORM_BUFFER_BINDINGS"},
{GL_MAX_COMPUTE_SHADER_STORAGE_BLOCKS, "GL_MAX_COMPUTE_SHADER_STORAGE_BLOCKS",
GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS, "GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS"},
{GL_MAX_COMBINED_SHADER_STORAGE_BLOCKS, "GL_MAX_COMBINED_SHADER_STORAGE_BLOCKS",
GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS, "GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS"},
};
for (const Relation& relation : relations) {
GLint blocks = -1;
GLint bindings = -1;
glGetIntegerv(relation.blocks, &blocks);
glGetIntegerv(relation.bindings, &bindings);
ASSERT_EQ(FirstGLError(), GLenum(GL_NO_ERROR)) << relation.blocksName;
EXPECT_LE(blocks, bindings)
<< relation.blocksName << " = " << blocks << " exceeds " << relation.bindingsName << " = "
<< bindings << "; a shader may declare more blocks than there are binding points to bind them to";
}
}
// KHR-GL44.multi_bind.functional_bind_buffers_range sizes each of an indexed target's
// binding points at MAX_<target>_SIZE / MAX_<target>_BINDINGS and binds all of them in
// one glBindBuffersRange. That quotient has to be a legal BindBufferRange size, which
// makes the two limits of every indexed family a PAIR: advertise a size that does not
// survive division by the binding count and the call fails with INVALID_VALUE before any
// of it binds.
TEST_F(AdvertisedLimitsScenario, IndexedTargetSizeSurvivesDivisionByItsBindingCount) {
struct IndexedFamily {
GLenum maxSize;
const char* maxSizeName;
GLenum maxBindings;
const char* maxBindingsName;
GLint sizeGranularity; // BindBufferRange's size rule for the target
};
const IndexedFamily families[] = {
{GL_MAX_ATOMIC_COUNTER_BUFFER_SIZE, "GL_MAX_ATOMIC_COUNTER_BUFFER_SIZE",
GL_MAX_ATOMIC_COUNTER_BUFFER_BINDINGS, "GL_MAX_ATOMIC_COUNTER_BUFFER_BINDINGS", 1},
{GL_MAX_TRANSFORM_FEEDBACK_INTERLEAVED_COMPONENTS, "GL_MAX_TRANSFORM_FEEDBACK_INTERLEAVED_COMPONENTS",
GL_MAX_TRANSFORM_FEEDBACK_BUFFERS, "GL_MAX_TRANSFORM_FEEDBACK_BUFFERS", 4},
{GL_MAX_UNIFORM_BLOCK_SIZE, "GL_MAX_UNIFORM_BLOCK_SIZE", GL_MAX_UNIFORM_BUFFER_BINDINGS,
"GL_MAX_UNIFORM_BUFFER_BINDINGS", 1},
{GL_MAX_SHADER_STORAGE_BLOCK_SIZE, "GL_MAX_SHADER_STORAGE_BLOCK_SIZE",
GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS, "GL_MAX_SHADER_STORAGE_BUFFER_BINDINGS", 1},
};
for (const IndexedFamily& family : families) {
GLint maxSize = -1;
GLint maxBindings = -1;
glGetIntegerv(family.maxSize, &maxSize);
glGetIntegerv(family.maxBindings, &maxBindings);
ASSERT_EQ(FirstGLError(), GLenum(GL_NO_ERROR)) << family.maxSizeName;
ASSERT_GT(maxBindings, 0) << family.maxBindingsName;
const GLint perBinding = maxSize / maxBindings;
EXPECT_GT(perBinding, 0)
<< family.maxSizeName << " (" << maxSize << ") / " << family.maxBindingsName << " ("
<< maxBindings << ") is zero, and BindBufferRange rejects a zero size";
EXPECT_EQ(perBinding % family.sizeGranularity, 0)
<< family.maxSizeName << " (" << maxSize << ") / " << family.maxBindingsName << " ("
<< maxBindings << ") = " << perBinding << " is not a multiple of the "
<< family.sizeGranularity << "-byte size granularity BindBufferRange requires for it";
}
}
// The OOM case in isolation, because it is the one with a known CTS victim and the one a // The OOM case in isolation, because it is the one with a known CTS victim and the one a
// future refactor is most likely to reintroduce by copying the Vulkan limit back. // future refactor is most likely to reintroduce by copying the Vulkan limit back.
TEST_F(AdvertisedLimitsScenario, ComputeUniformBlocksIsAnAmountAnApplicationCouldActuallyAllocate) { TEST_F(AdvertisedLimitsScenario, ComputeUniformBlocksIsAnAmountAnApplicationCouldActuallyAllocate) {
@@ -0,0 +1,396 @@
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/DoublePrecisionScenario.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// Scenario - GLSL DOUBLES, RUN AT SINGLE PRECISION.
//
// No mobile GPU has 64-bit floats. Adreno and Mali both report shaderFloat64 == VK_FALSE, so
// Magma cannot build a module that declares the Float64 capability, and ESSL has no fp64 type
// at all, so SPIRV-Cross refuses the module outright on Espryt ("FP64 not supported in ES
// profile") and the program never reaches the driver. MobileGL therefore narrows every 64-bit
// float in a shader to 32 bits (ShaderTranspiler::DemoteFloat64Pass) rather than declining the
// shader: `double` compiles and runs everywhere, at float precision.
//
// The narrowing is only half a contract. The other half is the API side: the global UBO is
// laid out by reflecting the DEMOTED module, so glUniform*d has to store a float where the
// shader reads a float, glGetUniform*v has to read one back, and a dmat4's columns are now
// std140-padded like any other matrix's. Every one of those is a byte offset that fails
// silently - the uniform simply reads as something else - so the cases below set values
// through the API and have the SHADER report what it saw.
//
// What is deliberately NOT asserted: that the values are exact to double precision. They are
// not, and cannot be. Every expectation here is the float value of the double that was set,
// which is the whole point.
#include <cmath>
#include <string>
#include <vector>
#include "../Harness/HeadlessGL.h"
#include "../Harness/ScenarioFixture.h"
#ifdef GLAPI
#undef GLAPI
#endif
#define GL_GLEXT_PROTOTYPES
#include <GL/gl.h>
#include <GL/glcorearb.h>
#undef GL_GLEXT_PROTOTYPES
namespace MGITest {
namespace {
// Doubles in every shape the demotion has to handle - a scalar, a vector, a matrix
// whose column stride changes, an array whose element stride changes - all reported
// through one float SSBO so a single readback says which one moved.
constexpr const char* kComputeSource = R"(#version 430 core
layout(local_size_x = 1) in;
uniform double uScalar;
uniform dvec3 uVector;
uniform dmat4 uMatrix;
uniform double uArray[3];
layout(std430, binding = 0) buffer Output {
float g_out[];
};
void main() {
g_out[0] = float(uScalar);
g_out[1] = float(uVector.x);
g_out[2] = float(uVector.y);
g_out[3] = float(uVector.z);
// Column-major [column][row]. Off-diagonal entries catch a column-stride mistake that a
// diagonal-only check reads straight past.
g_out[4] = float(uMatrix[0][0]);
g_out[5] = float(uMatrix[0][3]);
g_out[6] = float(uMatrix[3][0]);
g_out[7] = float(uMatrix[3][3]);
g_out[8] = float(uArray[0]);
g_out[9] = float(uArray[1]);
g_out[10] = float(uArray[2]);
// Arithmetic on doubles, including an implicit float->double conversion and a literal
// with the fp64 suffix: this is what an application actually writes, and it is the part
// that has to survive the conversion folding.
double accumulated = uScalar * 2.0lf + 1.5;
g_out[11] = float(accumulated);
}
)";
constexpr int kOutputSlots = 12;
class DoublePrecisionScenario : public ScenarioTest {
protected:
void SetUp() override {
ScenarioTest::SetUp();
if (!Ready()) return;
m_program = CompileComputeProgram(kComputeSource);
ASSERT_NE(m_program, 0u) << m_buildLog;
glGenBuffers(1, &m_output);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
const std::vector<float> zeroes(kOutputSlots, 0.0f);
glBufferData(GL_SHADER_STORAGE_BUFFER, kOutputSlots * sizeof(float), zeroes.data(),
GL_DYNAMIC_DRAW);
glBindBufferBase(GL_SHADER_STORAGE_BUFFER, 0, m_output);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, 0);
}
void TearDown() override {
if (!Ready()) return;
if (m_output != 0) glDeleteBuffers(1, &m_output);
if (m_program != 0) glDeleteProgram(m_program);
}
unsigned int CompileComputeProgram(const char* source) {
const GLuint shader = glCreateShader(GL_COMPUTE_SHADER);
glShaderSource(shader, 1, &source, nullptr);
glCompileShader(shader);
GLint compiled = 0;
glGetShaderiv(shader, GL_COMPILE_STATUS, &compiled);
if (compiled == GL_FALSE) {
char log[2048] = {};
glGetShaderInfoLog(shader, sizeof(log) - 1, nullptr, log);
m_buildLog = std::string("compute shader did not compile: ") + log;
glDeleteShader(shader);
return 0;
}
const GLuint program = glCreateProgram();
glAttachShader(program, shader);
glLinkProgram(program);
glDeleteShader(shader);
GLint linked = 0;
glGetProgramiv(program, GL_LINK_STATUS, &linked);
if (linked == GL_FALSE) {
char log[2048] = {};
glGetProgramInfoLog(program, sizeof(log) - 1, nullptr, log);
m_buildLog = std::string("compute program did not link: ") + log;
glDeleteProgram(program);
return 0;
}
return program;
}
std::vector<float> Dispatch() {
glUseProgram(m_program);
glDispatchCompute(1, 1, 1);
glMemoryBarrier(GL_BUFFER_UPDATE_BARRIER_BIT);
std::vector<float> values(kOutputSlots, -1.0f);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
glGetBufferSubData(GL_SHADER_STORAGE_BUFFER, 0, kOutputSlots * sizeof(float), values.data());
glBindBuffer(GL_SHADER_STORAGE_BUFFER, 0);
glUseProgram(0);
return values;
}
unsigned int m_program = 0;
unsigned int m_output = 0;
std::string m_buildLog;
};
TEST_F(DoublePrecisionScenario, ADoubleUniformReachesTheShaderAtFloatPrecision) {
if (!Ready()) return;
glUseProgram(m_program);
const GLint scalar = glGetUniformLocation(m_program, "uScalar");
ASSERT_GE(scalar, 0);
// 0.1 has no exact float (or double) representation, so this only passes if the
// value really travelled through the demoted slot rather than being read out of
// some other four bytes.
glUniform1d(scalar, 0.1);
glUseProgram(0);
const std::vector<float> values = Dispatch();
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
EXPECT_FLOAT_EQ(values[0], static_cast<float>(0.1));
EXPECT_FLOAT_EQ(values[11], static_cast<float>(static_cast<float>(0.1) * 2.0f + 1.5f))
<< "arithmetic on the demoted value, including the folded fp64 literal";
}
TEST_F(DoublePrecisionScenario, EveryDoubleShapeLandsInItsOwnSlot) {
if (!Ready()) return;
glUseProgram(m_program);
const GLint scalar = glGetUniformLocation(m_program, "uScalar");
const GLint vector = glGetUniformLocation(m_program, "uVector");
const GLint matrix = glGetUniformLocation(m_program, "uMatrix");
const GLint array0 = glGetUniformLocation(m_program, "uArray[0]");
const GLint array2 = glGetUniformLocation(m_program, "uArray[2]");
ASSERT_GE(scalar, 0);
ASSERT_GE(vector, 0);
ASSERT_GE(matrix, 0);
ASSERT_GE(array0, 0);
ASSERT_GE(array2, 0);
glUniform1d(scalar, 5.0);
const GLdouble vectorValue[3] = {11.0, 12.0, 13.0};
glUniform3dv(vector, 1, vectorValue);
// Column-major, and every entry distinct so a transposed or mis-strided write
// cannot land on a value that happens to match.
GLdouble matrixValue[16] = {};
for (int i = 0; i < 16; ++i) matrixValue[i] = 100.0 + i;
glUniformMatrix4dv(matrix, 1, GL_FALSE, matrixValue);
const GLdouble arrayValue[3] = {71.0, 72.0, 73.0};
glUniform1dv(array0, 3, arrayValue);
glUseProgram(0);
const std::vector<float> values = Dispatch();
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
EXPECT_FLOAT_EQ(values[0], 5.0f) << "scalar double";
EXPECT_FLOAT_EQ(values[1], 11.0f) << "dvec3 .x";
EXPECT_FLOAT_EQ(values[2], 12.0f) << "dvec3 .y";
EXPECT_FLOAT_EQ(values[3], 13.0f) << "dvec3 .z";
EXPECT_FLOAT_EQ(values[4], 100.0f) << "dmat4 [0][0]";
EXPECT_FLOAT_EQ(values[5], 103.0f) << "dmat4 [0][3] - within the first column";
EXPECT_FLOAT_EQ(values[6], 112.0f) << "dmat4 [3][0] - column stride";
EXPECT_FLOAT_EQ(values[7], 115.0f) << "dmat4 [3][3]";
EXPECT_FLOAT_EQ(values[8], 71.0f) << "double array element 0";
EXPECT_FLOAT_EQ(values[9], 72.0f) << "double array element 1 - element stride";
EXPECT_FLOAT_EQ(values[10], 73.0f) << "double array element 2";
}
TEST_F(DoublePrecisionScenario, TheTransposeFlagStillTransposes) {
if (!Ready()) return;
glUseProgram(m_program);
const GLint matrix = glGetUniformLocation(m_program, "uMatrix");
ASSERT_GE(matrix, 0);
GLdouble matrixValue[16] = {};
for (int i = 0; i < 16; ++i) matrixValue[i] = 100.0 + i;
glUniformMatrix4dv(matrix, 1, GL_TRUE, matrixValue);
glUseProgram(0);
const std::vector<float> values = Dispatch();
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
// Transposed, so [column][row] now reads the source's [row][column].
EXPECT_FLOAT_EQ(values[4], 100.0f) << "dmat4 [0][0] is on the diagonal either way";
EXPECT_FLOAT_EQ(values[5], 112.0f) << "dmat4 [0][3] after transpose";
EXPECT_FLOAT_EQ(values[6], 103.0f) << "dmat4 [3][0] after transpose";
EXPECT_FLOAT_EQ(values[7], 115.0f) << "dmat4 [3][3] is on the diagonal either way";
}
TEST_F(DoublePrecisionScenario, TheUniformIsStillReportedAsADouble) {
if (!Ready()) return;
// The demotion is an implementation detail of how the value is STORED. What the
// shader source declared is what the application asked about, so the reflection
// keeps answering GL_DOUBLE* - an application that switches on the type and calls
// glUniform*d has to keep working, and it is the glUniform*d path that is correct
// for these uniforms.
struct Expectation {
const char* name;
GLenum type;
GLint size;
};
const Expectation expectations[] = {
{"uScalar", GL_DOUBLE, 1},
{"uVector", GL_DOUBLE_VEC3, 1},
{"uMatrix", GL_DOUBLE_MAT4, 1},
{"uArray[0]", GL_DOUBLE, 3},
};
GLint activeUniforms = 0;
glGetProgramiv(m_program, GL_ACTIVE_UNIFORMS, &activeUniforms);
ASSERT_GT(activeUniforms, 0);
for (const Expectation& expectation : expectations) {
bool found = false;
for (GLint index = 0; index < activeUniforms; ++index) {
char name[128] = {};
GLsizei length = 0;
GLint size = 0;
GLenum type = 0;
glGetActiveUniform(m_program, static_cast<GLuint>(index), sizeof(name) - 1, &length, &size,
&type, name);
if (std::string(name, static_cast<size_t>(length)) != expectation.name) continue;
found = true;
EXPECT_EQ(type, expectation.type) << expectation.name;
EXPECT_EQ(size, expectation.size) << expectation.name;
break;
}
EXPECT_TRUE(found) << "glGetActiveUniform never reported " << expectation.name;
}
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
}
TEST_F(DoublePrecisionScenario, GetUniformdvReadsBackWhatWasStored) {
if (!Ready()) return;
glUseProgram(m_program);
const GLint scalar = glGetUniformLocation(m_program, "uScalar");
const GLint vector = glGetUniformLocation(m_program, "uVector");
const GLint matrix = glGetUniformLocation(m_program, "uMatrix");
ASSERT_GE(scalar, 0);
ASSERT_GE(vector, 0);
ASSERT_GE(matrix, 0);
glUniform1d(scalar, 0.1);
const GLdouble vectorValue[3] = {11.5, 12.5, 13.5};
glUniform3dv(vector, 1, vectorValue);
GLdouble matrixValue[16] = {};
for (int i = 0; i < 16; ++i) matrixValue[i] = 100.0 + i;
glUniformMatrix4dv(matrix, 1, GL_FALSE, matrixValue);
glUseProgram(0);
// The readback has to undo exactly what the write did - the same std140 column
// padding, the same 4-byte components - or a dmat4 comes back with its columns
// shifted and nothing else in the API would say so.
GLdouble readScalar = 0.0;
glGetUniformdv(m_program, scalar, &readScalar);
EXPECT_DOUBLE_EQ(readScalar, static_cast<double>(static_cast<float>(0.1)))
<< "the value is what a float can hold, not the double that was passed in";
GLdouble readVector[3] = {};
glGetUniformdv(m_program, vector, readVector);
EXPECT_DOUBLE_EQ(readVector[0], 11.5);
EXPECT_DOUBLE_EQ(readVector[1], 12.5);
EXPECT_DOUBLE_EQ(readVector[2], 13.5);
GLdouble readMatrix[16] = {};
glGetUniformdv(m_program, matrix, readMatrix);
for (int i = 0; i < 16; ++i) {
EXPECT_DOUBLE_EQ(readMatrix[i], 100.0 + i) << "dmat4 component " << i;
}
// The float query sees the same storage through the type it is actually stored as.
GLfloat readFloat = 0.0f;
glGetUniformfv(m_program, scalar, &readFloat);
EXPECT_FLOAT_EQ(readFloat, static_cast<float>(0.1));
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
}
TEST_F(DoublePrecisionScenario, ADoubleUniformKeepsItsDeclaredInitializer) {
if (!Ready()) return;
// A declared initializer is seeded straight into the uniform shadow at link, and the
// seeding used to skip 64-bit floats outright ("no 32-bit shadow encoding") - which
// was true before the demotion and silently left every such uniform reading zero.
const char* source = R"(#version 430 core
layout(local_size_x = 1) in;
uniform double uSeeded = 2.5lf;
uniform dvec3 uSeededVector = dvec3(4.0lf, 5.0lf, 6.0lf);
layout(std430, binding = 0) buffer Output {
float g_out[];
};
void main() {
g_out[0] = float(uSeeded);
g_out[1] = float(uSeededVector.x);
g_out[2] = float(uSeededVector.y);
g_out[3] = float(uSeededVector.z);
}
)";
const GLuint program = CompileComputeProgram(source);
ASSERT_NE(program, 0u) << m_buildLog;
glUseProgram(program);
glDispatchCompute(1, 1, 1);
glMemoryBarrier(GL_BUFFER_UPDATE_BARRIER_BIT);
std::vector<float> values(4, -1.0f);
glBindBuffer(GL_SHADER_STORAGE_BUFFER, m_output);
glGetBufferSubData(GL_SHADER_STORAGE_BUFFER, 0, 4 * sizeof(float), values.data());
glBindBuffer(GL_SHADER_STORAGE_BUFFER, 0);
glUseProgram(0);
glDeleteProgram(program);
EXPECT_FLOAT_EQ(values[0], 2.5f) << "scalar double initializer";
EXPECT_FLOAT_EQ(values[1], 4.0f) << "dvec3 initializer .x";
EXPECT_FLOAT_EQ(values[2], 5.0f) << "dvec3 initializer .y";
EXPECT_FLOAT_EQ(values[3], 6.0f) << "dvec3 initializer .z";
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
}
TEST_F(DoublePrecisionScenario, TheFp64ExtensionIsNotAdvertised) {
if (!Ready()) return;
// The shader above compiled, linked and ran without the extension string, which is
// the point: an application does not need GL_ARB_gpu_shader_fp64 advertised to USE
// doubles here. What the string additionally promises is 64-bit precision, and that
// is the one thing the demotion cannot deliver - so it stays off unless
// MOBILEGL_ADVERTISE_FP64 asks for it, and an application that branches on the
// string keeps taking its float path.
GLint extensionCount = 0;
glGetIntegerv(GL_NUM_EXTENSIONS, &extensionCount);
ASSERT_GT(extensionCount, 0);
bool advertised = false;
for (GLint i = 0; i < extensionCount; ++i) {
const char* name = reinterpret_cast<const char*>(glGetStringi(GL_EXTENSIONS, static_cast<GLuint>(i)));
if (name != nullptr && std::string(name) == "GL_ARB_gpu_shader_fp64") advertised = true;
}
EXPECT_FALSE(advertised);
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_NO_ERROR));
}
TEST_F(DoublePrecisionScenario, A64BitVertexFormatIsDeclinedOnEveryBackend) {
if (!Ready()) return;
// The demotion leaves no 64-bit shader input to feed, so there is nothing a 64-bit
// vertex FETCH could be fetched into - on either backend, and no longer only on the
// ones whose device lacks shaderFloat64. Declined loudly rather than accepted and
// drawn as garbage; the matching POST row says the same thing at startup.
GLuint vao = 0;
glGenVertexArrays(1, &vao);
glBindVertexArray(vao);
while (glGetError() != GL_NO_ERROR) {}
glVertexAttribLFormat(0, 3, GL_DOUBLE, 0);
EXPECT_EQ(glGetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
glBindVertexArray(0);
glDeleteVertexArrays(1, &vao);
while (glGetError() != GL_NO_ERROR) {}
}
} // namespace
} // namespace MGITest
@@ -0,0 +1,348 @@
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/DrawParametersScenario.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// gl_BaseVertex / gl_BaseInstance / gl_DrawID (GL_ARB_shader_draw_parameters),
// read straight out of the shader that a draw command produced.
//
// Neither backend has these builtins for free, and each is wrong in its own way
// when nobody watches:
//
// * DirectVulkan HAS a BaseVertex builtin, but Vulkan's carries the draw's
// firstVertex on a NON-INDEXED draw where GL's is defined to be zero ("the
// value passed to the baseVertex parameter, or zero for a command with no
// such parameter"). Only the indexed meaning of the two agrees. Every
// DrawArrays form therefore takes the ZeroBaseVertex program variant.
// * DirectGLES has no such builtins at all: ESSL knows none of them, so the
// transpiler demotes each one to a uniform the draw paths feed. A uniform
// nobody writes keeps whatever the previous draw left in it - which is what
// made gl_BaseVertex report a stale base vertex, and what made
// gl_BaseInstance read an unbound storage buffer on a plain glDrawArrays.
//
// The shader paints the three values, so a draw that carries the wrong ones
// paints the wrong colour rather than merely disagreeing with an expectation
// somewhere. The framebuffer is cleared to WHITE and no case expects 255 in any
// channel, so "the draw did not happen" can never be mistaken for a pass.
#include <cstdint>
#include <string>
#include <vector>
#include "../Harness/HeadlessGL.h"
#include "../Harness/ScenarioFixture.h"
#ifdef GLAPI
#undef GLAPI
#endif
#define GL_GLEXT_PROTOTYPES
#include <GL/gl.h>
#include <GL/glext.h>
namespace MGITest {
namespace {
// #version 450: glslang only declares the ARB builtins from 440 up.
//
// Each value is painted at 8 units per count, not 1: the errors these builtins
// actually have are OFF BY ONE (a sub-draw that never got its own gl_DrawID reads
// the previous one's, a base vertex that arrives one command late), and at one unit
// per count no readback tolerance can tell those from rounding.
//
// And biased by two counts, so that ZERO is not the clamp floor. Five of these cases
// expect zero, and an unbiased encoding would let every negative value - the shape a
// sign or rebase mistake produces - clamp to the same black and pass.
constexpr const char* kVertexSource = R"(#version 450 core
#extension GL_ARB_shader_draw_parameters : require
layout(location = 0) in vec2 aPos;
flat out vec3 vParams;
void main() {
vParams = (vec3(gl_BaseVertexARB, gl_BaseInstanceARB, gl_DrawIDARB) * 8.0 + 16.0) / 255.0;
gl_Position = vec4(aPos, 0.0, 1.0);
}
)";
constexpr const char* kFragmentSource = R"(#version 450 core
flat in vec3 vParams;
out vec4 oColor;
void main() {
oColor = vec4(vParams, 1.0);
}
)";
struct Vertex {
float x, y;
};
// 3 dummy vertices, then the left half of the viewport as two triangles,
// then the right half. Nothing here is symmetric by accident:
//
// * the padding makes a draw that ignores `first` / baseVertex paint a
// degenerate triangle (i.e. nothing) instead of the right picture;
// * the two halves let one multi-draw show TWO different gl_DrawID
// values in one readback.
//
// Indices 3..14 together cover the whole viewport, which is what the
// single-draw cases use.
constexpr int kPad = 3;
constexpr int kLeftFirst = kPad; // 3
constexpr int kRightFirst = kPad + 6; // 9
constexpr int kHalfCount = 6;
std::vector<Vertex> SceneVertices() {
std::vector<Vertex> vertices(static_cast<std::size_t>(kPad), Vertex{0.0f, 0.0f});
const float bounds[2][2] = {{-1.0f, 0.0f}, {0.0f, 1.0f}};
for (const auto& half : bounds) {
const float x0 = half[0];
const float x1 = half[1];
vertices.push_back({x0, -1.0f});
vertices.push_back({x1, -1.0f});
vertices.push_back({x1, 1.0f});
vertices.push_back({x0, -1.0f});
vertices.push_back({x1, 1.0f});
vertices.push_back({x0, 1.0f});
}
return vertices;
}
// GL's DrawArraysIndirectCommand / DrawElementsIndirectCommand, spelled out
// so a test can write one without depending on a GL header's struct.
struct ArraysCommand {
std::uint32_t count, instanceCount, first, baseInstance;
};
struct ElementsCommand {
std::uint32_t count, instanceCount, firstIndex;
std::int32_t baseVertex;
std::uint32_t baseInstance;
};
class DrawParametersScenario : public ScenarioTest {
protected:
void SetUp() override {
ScenarioTest::SetUp();
if (!Ready()) return;
std::string error;
m_program = CompileProgram(kVertexSource, kFragmentSource, &error);
ASSERT_NE(m_program, 0u) << error;
const std::vector<Vertex> vertices = SceneVertices();
glGenVertexArrays(1, &m_vao);
glBindVertexArray(m_vao);
glGenBuffers(1, &m_vbo);
glBindBuffer(GL_ARRAY_BUFFER, m_vbo);
glBufferData(GL_ARRAY_BUFFER, static_cast<GLsizeiptr>(vertices.size() * sizeof(Vertex)),
vertices.data(), GL_STATIC_DRAW);
glEnableVertexAttribArray(0);
glVertexAttribPointer(0, 2, GL_FLOAT, GL_FALSE, sizeof(Vertex), reinterpret_cast<const void*>(0));
ASSERT_EQ(FirstGLError(), GLenum(GL_NO_ERROR)) << "scene setup left a GL error behind";
}
void TearDown() override {
if (!Ready()) return;
for (GLuint* buffer : {&m_ebo, &m_indirect, &m_parameter, &m_vbo}) {
if (*buffer != 0) glDeleteBuffers(1, buffer);
*buffer = 0;
}
if (m_vao != 0) glDeleteVertexArrays(1, &m_vao);
if (m_program != 0) glDeleteProgram(m_program);
}
template <typename T>
void FillBuffer(GLuint& name, GLenum target, const std::vector<T>& data) {
if (name == 0) glGenBuffers(1, &name);
glBindBuffer(target, name);
glBufferData(target, static_cast<GLsizeiptr>(data.size() * sizeof(T)), data.data(), GL_STATIC_DRAW);
}
// Clears to white, runs `draw` and reads the frame back.
template <typename DrawFn>
Image Render(DrawFn&& draw) {
BindDefaultFramebuffer();
glViewport(0, 0, HeadlessGL::Get().Width(), HeadlessGL::Get().Height());
ClearTo(1.0f, 1.0f, 1.0f, 1.0f);
glUseProgram(m_program);
glBindVertexArray(m_vao);
draw();
return ReadPixels(HeadlessGL::Get().Width(), HeadlessGL::Get().Height());
}
// The three builtins as the shader saw them, at a point in one half of
// the viewport. `half` is 0 for the left half and 1 for the right.
struct DrawParams {
int baseVertex = -1, baseInstance = -1, drawId = -1;
};
// Decodes the biased 8-units-per-count encoding back to the integer the
// shader saw. Rounding to the nearest step absorbs any UNORM slop; adjacent
// values stay eight units apart, so an off-by-one still reads as one, and a
// negative value lands below the bias and decodes negative rather than
// clamping into a legitimate zero.
static DrawParams ParamsAt(const Image& image, int half) {
const int x = image.Width() * (1 + 2 * half) / 4;
const Rgba8 pixel = image.At(x, image.Height() / 2);
const auto decode = [](std::uint8_t channel) {
return (static_cast<int>(channel) - 16 + 4) / 8;
};
return {decode(pixel.r), decode(pixel.g), decode(pixel.b)};
}
static void ExpectParams(const Image& image, int half, const DrawParams& expected,
const std::string& what) {
const DrawParams actual = ParamsAt(image, half);
EXPECT_EQ(actual.baseVertex, expected.baseVertex)
<< what << ": gl_BaseVertex (half " << half << ")";
EXPECT_EQ(actual.baseInstance, expected.baseInstance)
<< what << ": gl_BaseInstance (half " << half << ")";
EXPECT_EQ(actual.drawId, expected.drawId) << what << ": gl_DrawID (half " << half << ")";
}
GLuint m_program = 0;
GLuint m_vao = 0;
GLuint m_vbo = 0;
GLuint m_ebo = 0;
GLuint m_indirect = 0;
GLuint m_parameter = 0;
};
// ---- the non-indexed forms: gl_BaseVertex is zero, `first` or not ----
// Vulkan's BaseVertex would answer 3 here (the draw's firstVertex); GL's
// must answer 0, because glDrawArrays has no baseVertex parameter at all.
TEST_F(DrawParametersScenario, DrawArraysReportsAZeroBaseVertexDespiteItsFirst) {
if (!Ready()) return;
const Image image = Render([&] { glDrawArrays(GL_TRIANGLES, kLeftFirst, 2 * kHalfCount); });
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 0, 0}, "glDrawArrays(first=3)");
ExpectParams(image, 1, {0, 0, 0}, "glDrawArrays(first=3)");
}
TEST_F(DrawParametersScenario, DrawArraysInstancedBaseInstanceReportsItsBaseInstance) {
if (!Ready()) return;
const Image image = Render([&] {
glDrawArraysInstancedBaseInstance(GL_TRIANGLES, kLeftFirst, 2 * kHalfCount, 1, 5);
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 5, 0}, "glDrawArraysInstancedBaseInstance(baseInstance=5)");
}
// The base instance of one draw must not survive into the next one. This is
// the shape that broke on DirectGLES: the emulation uniform is per-program
// state, so a draw that never writes it inherits the last writer's value.
TEST_F(DrawParametersScenario, APlainDrawAfterABaseInstancedOneSeesZeroAgain) {
if (!Ready()) return;
const Image image = Render([&] {
glDrawArraysInstancedBaseInstance(GL_TRIANGLES, kLeftFirst, 2 * kHalfCount, 1, 7);
glDrawArrays(GL_TRIANGLES, kLeftFirst, 2 * kHalfCount);
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 0, 0}, "plain glDrawArrays after a base-instanced draw");
}
// ---- the indexed forms: gl_BaseVertex IS the base vertex ----
TEST_F(DrawParametersScenario, DrawElementsBaseVertexReportsItsBaseVertex) {
if (!Ready()) return;
std::vector<std::uint32_t> indices;
for (std::uint32_t i = 0; i < 2 * kHalfCount; ++i) indices.push_back(i);
FillBuffer(m_ebo, GL_ELEMENT_ARRAY_BUFFER, indices);
const Image image = Render([&] {
glDrawElementsBaseVertex(GL_TRIANGLES, 2 * kHalfCount, GL_UNSIGNED_INT,
reinterpret_cast<const void*>(0), kLeftFirst);
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {kLeftFirst, 0, 0}, "glDrawElementsBaseVertex(basevertex=3)");
ExpectParams(image, 1, {kLeftFirst, 0, 0}, "glDrawElementsBaseVertex(basevertex=3)");
}
// ... and is zero again for the command that has none, including after one
// that did: the same leak the base instance has, on the other builtin. The
// preceding draw MUST carry a non-zero base vertex or this case proves nothing -
// one index run reaches the geometry through the base vertex, the second through
// its own indices, so the two draws paint the same picture with different
// gl_BaseVertex and only the second one's value survives in the framebuffer.
TEST_F(DrawParametersScenario, DrawElementsAfterABaseVertexDrawReportsZeroAgain) {
if (!Ready()) return;
std::vector<std::uint32_t> indices;
for (std::uint32_t i = 0; i < 2 * kHalfCount; ++i) indices.push_back(i);
for (std::uint32_t i = 0; i < 2 * kHalfCount; ++i) indices.push_back(i + kLeftFirst);
FillBuffer(m_ebo, GL_ELEMENT_ARRAY_BUFFER, indices);
const auto rebasedRun = reinterpret_cast<const void*>(2 * kHalfCount * sizeof(std::uint32_t));
const Image image = Render([&] {
glDrawElementsBaseVertex(GL_TRIANGLES, 2 * kHalfCount, GL_UNSIGNED_INT,
reinterpret_cast<const void*>(0), kLeftFirst);
glDrawElements(GL_TRIANGLES, 2 * kHalfCount, GL_UNSIGNED_INT, rebasedRun);
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 0, 0}, "glDrawElements after a base-vertex draw");
ExpectParams(image, 1, {0, 0, 0}, "glDrawElements after a base-vertex draw");
}
// ---- the multi-draw forms: one gl_DrawID per sub-draw ----
TEST_F(DrawParametersScenario, MultiDrawArraysNumbersItsSubDraws) {
if (!Ready()) return;
const GLint firsts[2] = {kLeftFirst, kRightFirst};
const GLsizei counts[2] = {kHalfCount, kHalfCount};
const Image image = Render([&] { glMultiDrawArrays(GL_TRIANGLES, firsts, counts, 2); });
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 0, 0}, "glMultiDrawArrays sub-draw 0");
ExpectParams(image, 1, {0, 0, 1}, "glMultiDrawArrays sub-draw 1");
}
// Every field of an indexed indirect command at once: its own gl_DrawID, the
// baseVertex word (which the CPU reads out of the command) and the
// baseInstance word (which DirectGLES reads through a storage-buffer view of
// the very same buffer).
TEST_F(DrawParametersScenario, MultiDrawElementsIndirectCarriesEveryCommandsParameters) {
if (!Ready()) return;
std::vector<std::uint32_t> indices;
for (std::uint32_t i = 0; i < kHalfCount; ++i) indices.push_back(i);
FillBuffer(m_ebo, GL_ELEMENT_ARRAY_BUFFER, indices);
const std::vector<ElementsCommand> commands = {
{kHalfCount, 1, 0, kLeftFirst, 0},
{kHalfCount, 1, 0, kRightFirst, 4},
};
FillBuffer(m_indirect, GL_DRAW_INDIRECT_BUFFER, commands);
const Image image = Render([&] {
glMultiDrawElementsIndirect(GL_TRIANGLES, GL_UNSIGNED_INT, reinterpret_cast<const void*>(0), 2,
sizeof(ElementsCommand));
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {kLeftFirst, 0, 0}, "indirect command 0");
ExpectParams(image, 1, {kRightFirst, 4, 1}, "indirect command 1");
}
// glMultiDrawArraysIndirectCount was missing from the DirectGLES backend
// table entirely, so the frontend answered INVALID_OPERATION for every call
// while GL_ARB_indirect_parameters was advertised. The parameter buffer here
// holds a count SMALLER than maxdrawcount, so a path that ignores it draws a
// third command over the top of the second and changes the right half.
TEST_F(DrawParametersScenario, MultiDrawArraysIndirectCountObeysItsParameterBuffer) {
if (!Ready()) return;
const std::vector<ArraysCommand> commands = {
{kHalfCount, 1, kLeftFirst, 0},
{kHalfCount, 1, kRightFirst, 6},
{kHalfCount, 1, kRightFirst, 9},
};
FillBuffer(m_indirect, GL_DRAW_INDIRECT_BUFFER, commands);
const std::vector<std::uint32_t> parameters = {2};
FillBuffer(m_parameter, GL_PARAMETER_BUFFER, parameters);
const Image image = Render([&] {
glMultiDrawArraysIndirectCount(GL_TRIANGLES, reinterpret_cast<const void*>(0), 0, 3,
sizeof(ArraysCommand));
});
EXPECT_EQ(FirstGLError(), GLenum(GL_NO_ERROR));
ExpectParams(image, 0, {0, 0, 0}, "counted indirect command 0");
ExpectParams(image, 1, {0, 6, 1}, "counted indirect command 1");
}
} // namespace
} // namespace MGITest
@@ -0,0 +1,674 @@
// MobileGL - MobileGL/MG_IntegrationTest/Scenarios/VertexAttribBindingScenario.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
//
// ARB_vertex_attrib_binding: the separate format/binding state the GL 4.3 vertex
// input model is made of, read back out of the draw that consumed it.
//
// Every scenario here captures the vertex shader's inputs with transform feedback
// under GL_RASTERIZER_DISCARD, which is what the KHR-GL43.vertex_attrib_binding
// cases do: the captured record IS the fetched vertex, so "the binding state did
// not reach the draw" and "the draw fetched the wrong bytes" are distinguishable
// from each other and from "the capture did not run" (the buffer is pre-filled
// with a poison value).
#include <cstdio>
#include <cstring>
#include <string>
#include <vector>
#include "../Harness/HeadlessGL.h"
#include "../Harness/ScenarioFixture.h"
#ifdef GLAPI
#undef GLAPI
#endif
#define GL_GLEXT_PROTOTYPES
#include <GL/gl.h>
#include <GL/glcorearb.h>
#undef GL_GLEXT_PROTOTYPES
namespace MGITest {
namespace {
constexpr float kPoison = -1234.0f;
GLuint CompileShader(GLenum type, const std::string& source, std::string* log) {
const GLuint shader = glCreateShader(type);
const char* text = source.c_str();
glShaderSource(shader, 1, &text, nullptr);
glCompileShader(shader);
GLint status = GL_FALSE;
glGetShaderiv(shader, GL_COMPILE_STATUS, &status);
if (status == GL_FALSE) {
GLint length = 0;
glGetShaderiv(shader, GL_INFO_LOG_LENGTH, &length);
std::vector<char> buffer(static_cast<std::size_t>(length) + 1, '\0');
glGetShaderInfoLog(shader, length + 1, nullptr, buffer.data());
if (log != nullptr) *log = buffer.data();
glDeleteShader(shader);
return 0;
}
return shader;
}
// A vertex-only capture program, exactly how the CTS builds one: the varying
// names are declared before the link and the fragment stage is absent because
// the draw runs under GL_RASTERIZER_DISCARD.
GLuint BuildCaptureProgram(const std::string& vertexSource, const std::vector<const char*>& xfbVaryings,
std::string* log) {
const GLuint vertexShader = CompileShader(GL_VERTEX_SHADER, vertexSource, log);
if (vertexShader == 0) return 0;
const GLuint program = glCreateProgram();
glAttachShader(program, vertexShader);
if (!xfbVaryings.empty()) {
glTransformFeedbackVaryings(program, static_cast<GLsizei>(xfbVaryings.size()), xfbVaryings.data(),
GL_INTERLEAVED_ATTRIBS);
}
glLinkProgram(program);
glDeleteShader(vertexShader);
GLint status = GL_FALSE;
glGetProgramiv(program, GL_LINK_STATUS, &status);
if (status == GL_FALSE) {
GLint length = 0;
glGetProgramiv(program, GL_INFO_LOG_LENGTH, &length);
std::vector<char> buffer(static_cast<std::size_t>(length) + 1, '\0');
glGetProgramInfoLog(program, length + 1, nullptr, buffer.data());
if (log != nullptr) *log = buffer.data();
glDeleteProgram(program);
return 0;
}
return program;
}
// Four float inputs at locations 0..3, captured as four vec4s per vertex.
// Locations the test does not feed keep their current-attribute value, which
// every scenario sets to a known constant first.
std::string CaptureVertexSource() {
return R"(#version 430 core
layout(location = 0) in vec4 vs_in_attrib0;
layout(location = 1) in vec4 vs_in_attrib1;
layout(location = 2) in vec4 vs_in_attrib2;
layout(location = 3) in vec4 vs_in_attrib3;
out StageData {
vec4 attrib0;
vec4 attrib1;
vec4 attrib2;
vec4 attrib3;
} vs_out;
void main() {
vs_out.attrib0 = vs_in_attrib0;
vs_out.attrib1 = vs_in_attrib1;
vs_out.attrib2 = vs_in_attrib2;
vs_out.attrib3 = vs_in_attrib3;
}
)";
}
std::vector<const char*> CaptureVaryingNames() {
return {"StageData.attrib0", "StageData.attrib1", "StageData.attrib2", "StageData.attrib3"};
}
// Runs `vertexCount` x `instanceCount` points through the capture program and
// returns the interleaved floats (16 per point: four vec4s).
std::vector<float> CapturePoints(GLuint program, GLuint xfbBuffer, int vertexCount, int instanceCount) {
const std::size_t floats = static_cast<std::size_t>(vertexCount) * instanceCount * 16;
std::vector<float> poison(floats, kPoison);
glBindBufferBase(GL_TRANSFORM_FEEDBACK_BUFFER, 0, xfbBuffer);
glBufferData(GL_TRANSFORM_FEEDBACK_BUFFER, static_cast<GLsizeiptr>(floats * sizeof(float)), poison.data(),
GL_DYNAMIC_DRAW);
glEnable(GL_RASTERIZER_DISCARD);
glUseProgram(program);
glBeginTransformFeedback(GL_POINTS);
glDrawArraysInstanced(GL_POINTS, 0, vertexCount, instanceCount);
glEndTransformFeedback();
glDisable(GL_RASTERIZER_DISCARD);
std::vector<float> data(floats, kPoison);
glGetBufferSubData(GL_TRANSFORM_FEEDBACK_BUFFER, 0, static_cast<GLsizeiptr>(floats * sizeof(float)),
data.data());
glUseProgram(0);
return data;
}
// As CapturePoints, but through the baseInstance entry point, and on a capture buffer
// of its own.
//
// Kept separate from CapturePoints rather than defaulting a parameter, for two
// reasons. Every existing caller stays on the draw command that carries no
// baseInstance at all, so the negative control is a DIFFERENT command rather than
// the same one passed a zero. And baseInstance is the first thing here that needs
// several captures in ONE test, which the shared helper cannot currently do: a
// second capture into the same buffer object comes back empty on DirectVulkan
// (respecifying a buffer that is bound to a transform-feedback binding point does
// not reach that binding - reproduced with two plain CapturePoints calls, so it is
// neither about baseInstance nor about this helper). A fresh buffer per capture
// sidesteps it; without that, this scenario would be pinning that bug instead.
std::vector<float> CaptureOwnBufferBaseInstance(GLuint program, int vertexCount, int instanceCount,
GLuint baseInstance, bool useBaseInstanceCommand) {
const std::size_t floats = static_cast<std::size_t>(vertexCount) * instanceCount * 16;
std::vector<float> poison(floats, kPoison);
GLuint xfbBuffer = 0;
glGenBuffers(1, &xfbBuffer);
glBindBufferBase(GL_TRANSFORM_FEEDBACK_BUFFER, 0, xfbBuffer);
glBufferData(GL_TRANSFORM_FEEDBACK_BUFFER, static_cast<GLsizeiptr>(floats * sizeof(float)), poison.data(),
GL_DYNAMIC_DRAW);
glEnable(GL_RASTERIZER_DISCARD);
glUseProgram(program);
glBeginTransformFeedback(GL_POINTS);
if (useBaseInstanceCommand) {
glDrawArraysInstancedBaseInstance(GL_POINTS, 0, vertexCount, instanceCount, baseInstance);
} else {
glDrawArraysInstanced(GL_POINTS, 0, vertexCount, instanceCount);
}
glEndTransformFeedback();
glDisable(GL_RASTERIZER_DISCARD);
std::vector<float> data(floats, kPoison);
glGetBufferSubData(GL_TRANSFORM_FEEDBACK_BUFFER, 0, static_cast<GLsizeiptr>(floats * sizeof(float)),
data.data());
glUseProgram(0);
glDeleteBuffers(1, &xfbBuffer);
return data;
}
// point p, attribute a, component c
float At(const std::vector<float>& data, int point, int attrib, int component) {
const std::size_t index = static_cast<std::size_t>(point) * 16 + attrib * 4 + component;
return index < data.size() ? data[index] : kPoison;
}
void ResetCurrentAttribs() {
for (GLuint i = 0; i < 4; ++i) {
glVertexAttrib4f(i, 0.0f, 0.0f, 0.0f, 0.0f);
}
}
::testing::AssertionResult Vec4Is(const std::vector<float>& data, int point, int attrib, float x, float y,
float z, float w) {
const float gx = At(data, point, attrib, 0);
const float gy = At(data, point, attrib, 1);
const float gz = At(data, point, attrib, 2);
const float gw = At(data, point, attrib, 3);
const float tolerance = 0.01f;
auto close = [tolerance](float a, float b) { return (a - b) < tolerance && (b - a) < tolerance; };
if (close(gx, x) && close(gy, y) && close(gz, z) && close(gw, w)) {
return ::testing::AssertionSuccess();
}
return ::testing::AssertionFailure()
<< "point " << point << " attribute " << attrib << " is (" << gx << ", " << gy << ", " << gz << ", "
<< gw << "), expected (" << x << ", " << y << ", " << z << ", " << w << ")";
}
} // namespace
class VertexAttribBindingScenario : public ScenarioTest {
protected:
void SetUp() override {
ScenarioTest::SetUp();
if (!Ready()) return;
m_program = BuildCaptureProgram(CaptureVertexSource(), CaptureVaryingNames(), &m_log);
ASSERT_NE(m_program, 0u) << "capture program did not link: " << m_log;
glGenVertexArrays(1, &m_vao);
glGenBuffers(1, &m_xfbo);
glBindVertexArray(m_vao);
}
void TearDown() override {
if (!Ready()) return;
glBindVertexArray(0);
glDeleteVertexArrays(1, &m_vao);
glDeleteBuffers(1, &m_xfbo);
glDeleteProgram(m_program);
}
GLuint m_program = 0;
GLuint m_vao = 0;
GLuint m_xfbo = 0;
std::string m_log;
};
// glVertexAttribFormat + glBindVertexBuffer + glVertexAttribBinding, in the order
// the CTS uses (buffer first, then format, then binding), must feed the draw.
TEST_F(VertexAttribBindingScenario, FormatAndBindingFeedTheDraw) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glBindVertexBuffer(0, vbo, 0, 12);
glVertexAttribFormat(1, 3, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(1, 0);
glEnableVertexAttribArray(1);
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 2, 1);
EXPECT_TRUE(Vec4Is(data, 0, 1, 1.0f, 2.0f, 3.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 1, 4.0f, 5.0f, 6.0f, 1.0f));
// An attribute nothing configured still reports its current value.
EXPECT_TRUE(Vec4Is(data, 0, 0, 0.0f, 0.0f, 0.0f, 0.0f));
glDisableVertexAttribArray(1);
glDeleteBuffers(1, &vbo);
}
// The reverse order - format and binding declared before any buffer exists on the
// binding point - has to resolve to the same thing once glBindVertexBuffer lands.
TEST_F(VertexAttribBindingScenario, FormatBeforeBufferStillResolves) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribBinding(2, 3);
glVertexAttribFormat(2, 2, GL_FLOAT, GL_FALSE, 4);
glEnableVertexAttribArray(2);
glBindVertexBuffer(3, vbo, 0, 12);
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 2, 1);
EXPECT_TRUE(Vec4Is(data, 0, 2, 2.0f, 3.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 2, 5.0f, 6.0f, 0.0f, 1.0f));
glDisableVertexAttribArray(2);
glDeleteBuffers(1, &vbo);
}
// GL 4.6 core 10.3.1: a binding point's stride is the byte distance between
// consecutive elements, and zero means every vertex reads the SAME element. That
// is the opposite of glVertexAttribPointer's stride 0, which means "tightly
// packed" - the two spellings must not be collapsed into one another.
TEST_F(VertexAttribBindingScenario, BindingStrideZeroRepeatsOneElement) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f, 8.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribFormat(0, 4, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(0, 5);
glBindVertexBuffer(5, vbo, 16, 0);
glEnableVertexAttribArray(0);
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 2, 1);
EXPECT_TRUE(Vec4Is(data, 0, 0, 5.0f, 6.0f, 7.0f, 8.0f));
EXPECT_TRUE(Vec4Is(data, 1, 0, 5.0f, 6.0f, 7.0f, 8.0f));
glDisableVertexAttribArray(0);
glDeleteBuffers(1, &vbo);
}
// The pointer API keeps its own meaning of stride 0 (tightly packed) even though
// it is defined in terms of the binding model - the negative control for the
// scenario above.
TEST_F(VertexAttribBindingScenario, PointerStrideZeroStaysTightlyPacked) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f, 7.0f, 8.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glVertexAttribPointer(0, 4, GL_FLOAT, GL_FALSE, 0, nullptr);
glEnableVertexAttribArray(0);
glBindBuffer(GL_ARRAY_BUFFER, 0);
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 2, 1);
EXPECT_TRUE(Vec4Is(data, 0, 0, 1.0f, 2.0f, 3.0f, 4.0f));
EXPECT_TRUE(Vec4Is(data, 1, 0, 5.0f, 6.0f, 7.0f, 8.0f));
glDisableVertexAttribArray(0);
glDeleteBuffers(1, &vbo);
}
// glVertexBindingDivisor is per BINDING POINT: it has to reach every attribute
// pointed at that binding, and the instance step must honour the divisor rather
// than advancing once per instance.
TEST_F(VertexAttribBindingScenario, BindingDivisorAppliesToEveryAttributeOnThePoint) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {10.0f, 20.0f, 30.0f, 40.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribFormat(0, 1, GL_FLOAT, GL_FALSE, 0);
glVertexAttribFormat(1, 1, GL_FLOAT, GL_FALSE, 4);
glVertexAttribBinding(0, 4);
glVertexAttribBinding(1, 4);
glBindVertexBuffer(4, vbo, 0, 8);
glVertexBindingDivisor(4, 2);
glEnableVertexAttribArray(0);
glEnableVertexAttribArray(1);
// The divisor is per binding point, so it has to be visible on BOTH attributes
// pointed at it - and this query is what separates "the frontend never resolved
// it" from "the backend did not apply it".
GLint divisor = -1;
glGetVertexAttribiv(0, GL_VERTEX_ATTRIB_ARRAY_DIVISOR, &divisor);
EXPECT_EQ(divisor, 2);
divisor = -1;
glGetVertexAttribiv(1, GL_VERTEX_ATTRIB_ARRAY_DIVISOR, &divisor);
EXPECT_EQ(divisor, 2);
// 1 vertex x 4 instances, divisor 2: instances 0,1 read element 0 and
// instances 2,3 read element 1.
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 1, 4);
EXPECT_TRUE(Vec4Is(data, 0, 0, 10.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 0, 10.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 2, 0, 30.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 3, 0, 30.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 0, 1, 20.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 2, 1, 40.0f, 0.0f, 0.0f, 1.0f));
glDisableVertexAttribArray(0);
glDisableVertexAttribArray(1);
glDeleteBuffers(1, &vbo);
}
// baseInstance moves the ELEMENT the instanced arrays start at. DirectGLES has no
// ES entry point that says so on the drivers we ship against (GL_EXT_base_instance
// is absent on Adreno), so it folds the shift into the attribute's own offset - and
// the thing that made this worth pinning is that the value used to reach the shader
// uniform for gl_BaseInstance and NEVER the fetch, so a draw could report a base
// instance it had not actually read from.
//
// The three draws are the point. Zero first as a negative control, so a backend that
// simply ignored baseInstance could not pass on the middle draw alone; and zero AGAIN
// last, because the shift is emitted into per-attribute state the VAO twin memoises -
// leaving it applied would make every subsequent ordinary draw fetch from the wrong
// element, which is a far worse bug than the one being fixed.
TEST_F(VertexAttribBindingScenario, BaseInstanceMovesTheInstancedArraysStartElement) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float instanceData[] = {10.0f, 20.0f, 30.0f, 40.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(instanceData), instanceData, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribFormat(0, 1, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(0, 0);
glBindVertexBuffer(0, vbo, 0, 4);
glVertexBindingDivisor(0, 1);
glEnableVertexAttribArray(0);
const std::vector<float> atZero = CaptureOwnBufferBaseInstance(m_program, 1, 2, 0, true);
EXPECT_TRUE(Vec4Is(atZero, 0, 0, 10.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(atZero, 1, 0, 20.0f, 0.0f, 0.0f, 1.0f));
const std::vector<float> atTwo = CaptureOwnBufferBaseInstance(m_program, 1, 2, 2, true);
EXPECT_TRUE(Vec4Is(atTwo, 0, 0, 30.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(atTwo, 1, 0, 40.0f, 0.0f, 0.0f, 1.0f));
// Nothing about the vertex array changed between these two draws, so only a
// backend that actively un-shifts on a baseInstance change gets back to 10/20.
const std::vector<float> backToZero = CaptureOwnBufferBaseInstance(m_program, 1, 2, 0, true);
EXPECT_TRUE(Vec4Is(backToZero, 0, 0, 10.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(backToZero, 1, 0, 20.0f, 0.0f, 0.0f, 1.0f));
// And a draw command with no baseInstance parameter at all must be unaffected by
// the one that came before it.
const std::vector<float> plain = CaptureOwnBufferBaseInstance(m_program, 1, 2, 0, false);
EXPECT_TRUE(Vec4Is(plain, 0, 0, 10.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(plain, 1, 0, 20.0f, 0.0f, 0.0f, 1.0f));
glDisableVertexAttribArray(0);
glDeleteBuffers(1, &vbo);
}
// baseInstance is defined against the instanced arrays only: an array with divisor 0
// advances per VERTEX and its start element is "first", which baseInstance does not
// touch. An emulation that shifted by offset without checking the divisor would move
// this one too, and nothing in the case above would notice.
TEST_F(VertexAttribBindingScenario, BaseInstanceLeavesPerVertexArraysWhereTheyWere) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float perVertex[] = {1.0f, 2.0f, 3.0f, 4.0f};
const float perInstance[] = {10.0f, 20.0f, 30.0f, 40.0f};
GLuint buffers[2] = {0, 0};
glGenBuffers(2, buffers);
glBindBuffer(GL_ARRAY_BUFFER, buffers[0]);
glBufferData(GL_ARRAY_BUFFER, sizeof(perVertex), perVertex, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, buffers[1]);
glBufferData(GL_ARRAY_BUFFER, sizeof(perInstance), perInstance, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribFormat(0, 1, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(0, 0);
glBindVertexBuffer(0, buffers[0], 0, 4);
glVertexBindingDivisor(0, 0);
glEnableVertexAttribArray(0);
glVertexAttribFormat(1, 1, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(1, 1);
glBindVertexBuffer(1, buffers[1], 0, 4);
glVertexBindingDivisor(1, 1);
glEnableVertexAttribArray(1);
// 2 vertices x 2 instances, baseInstance 2. Points come out instance-major.
const std::vector<float> data = CaptureOwnBufferBaseInstance(m_program, 2, 2, 2, true);
EXPECT_TRUE(Vec4Is(data, 0, 0, 1.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 0, 2.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 2, 0, 1.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 3, 0, 2.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 0, 1, 30.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 1, 30.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 2, 1, 40.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 3, 1, 40.0f, 0.0f, 0.0f, 1.0f));
glDisableVertexAttribArray(0);
glDisableVertexAttribArray(1);
glDeleteBuffers(2, buffers);
}
// Two attributes on one binding point at different relative offsets, plus a
// binding offset: the fetch address is binding offset + relative offset, and the
// relative offset must not leak into the binding's own offset.
TEST_F(VertexAttribBindingScenario, RelativeOffsetComposesWithBindingOffset) {
if (!Ready()) GTEST_SKIP();
ResetCurrentAttribs();
const float vertices[] = {0.0f, 0.0f, 1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glVertexAttribFormat(0, 2, GL_FLOAT, GL_FALSE, 0);
glVertexAttribFormat(1, 1, GL_FLOAT, GL_FALSE, 8);
glVertexAttribBinding(0, 1);
glVertexAttribBinding(1, 1);
glBindVertexBuffer(1, vbo, 8, 12);
glEnableVertexAttribArray(0);
glEnableVertexAttribArray(1);
const std::vector<float> data = CapturePoints(m_program, m_xfbo, 2, 1);
EXPECT_TRUE(Vec4Is(data, 0, 0, 1.0f, 2.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 0, 1, 3.0f, 0.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 0, 4.0f, 5.0f, 0.0f, 1.0f));
EXPECT_TRUE(Vec4Is(data, 1, 1, 6.0f, 0.0f, 0.0f, 1.0f));
glDisableVertexAttribArray(0);
glDisableVertexAttribArray(1);
glDeleteBuffers(1, &vbo);
}
// The KHR-GL43.vertex_attrib_binding.basic-input* capture program verbatim: a
// 16-element vec4 input ARRAY at location 0, copied element by element into a
// 16-element array inside an output interface block, all 16 members captured.
// Every one of the 17 basic-input* cases is built on it, so a backend that cannot
// produce this program fails all of them with "the draw captured zeros" and no
// other symptom.
TEST_F(VertexAttribBindingScenario, InputArrayCaptureProgramFeedsTheDraw) {
if (!Ready()) GTEST_SKIP();
const std::string vs = R"(#version 430 core
layout(location = 0) in vec4 vs_in_attrib[16];
out StageData {
vec4 attrib[16];
} vs_out;
void main() {
for (int i = 0; i < vs_in_attrib.length(); ++i) {
vs_out.attrib[i] = vs_in_attrib[i];
}
}
)";
std::vector<std::string> names;
for (int i = 0; i < 16; ++i) names.push_back("StageData.attrib[" + std::to_string(i) + "]");
std::vector<const char*> varyings;
for (const auto& n : names) varyings.push_back(n.c_str());
std::string log;
const GLuint program = BuildCaptureProgram(vs, varyings, &log);
ASSERT_NE(program, 0u) << "capture program did not link: " << log;
for (GLuint i = 0; i < 16; ++i) glVertexAttrib4f(i, 0.0f, 0.0f, 0.0f, 0.0f);
const float vertices[] = {1.0f, 2.0f, 3.0f, 4.0f, 5.0f, 6.0f};
GLuint vbo = 0;
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glBindBuffer(GL_ARRAY_BUFFER, 0);
glBindVertexBuffer(0, vbo, 0, 12);
glVertexAttribFormat(1, 3, GL_FLOAT, GL_FALSE, 0);
glVertexAttribBinding(1, 0);
glEnableVertexAttribArray(1);
// 16 vec4s per point rather than the 4 the shared helper assumes.
constexpr std::size_t kFloatsPerPoint = 64;
std::vector<float> poison(kFloatsPerPoint * 2, kPoison);
glBindBufferBase(GL_TRANSFORM_FEEDBACK_BUFFER, 0, m_xfbo);
glBufferData(GL_TRANSFORM_FEEDBACK_BUFFER, static_cast<GLsizeiptr>(poison.size() * sizeof(float)),
poison.data(), GL_DYNAMIC_DRAW);
glEnable(GL_RASTERIZER_DISCARD);
glUseProgram(program);
glBeginTransformFeedback(GL_POINTS);
glDrawArrays(GL_POINTS, 0, 2);
glEndTransformFeedback();
glDisable(GL_RASTERIZER_DISCARD);
std::vector<float> data(poison.size(), kPoison);
glGetBufferSubData(GL_TRANSFORM_FEEDBACK_BUFFER, 0,
static_cast<GLsizeiptr>(data.size() * sizeof(float)), data.data());
glUseProgram(0);
// Element 0 of the array has no enabled array behind it, so it must deliver the
// current generic attribute value set above - including its w, which is 0 here and
// NOT the 1 an unwritten vec4 input defaults to.
EXPECT_FLOAT_EQ(data[0], 0.0f);
EXPECT_FLOAT_EQ(data[3], 0.0f);
// attribute 1 of point 0 and of point 1.
EXPECT_FLOAT_EQ(data[4], 1.0f);
EXPECT_FLOAT_EQ(data[5], 2.0f);
EXPECT_FLOAT_EQ(data[6], 3.0f);
EXPECT_FLOAT_EQ(data[7], 1.0f);
EXPECT_FLOAT_EQ(data[kFloatsPerPoint + 4], 4.0f);
EXPECT_FLOAT_EQ(data[kFloatsPerPoint + 5], 5.0f);
EXPECT_FLOAT_EQ(data[kFloatsPerPoint + 6], 6.0f);
EXPECT_FLOAT_EQ(data[kFloatsPerPoint + 7], 1.0f);
glDisableVertexAttribArray(1);
glDeleteBuffers(1, &vbo);
glDeleteProgram(program);
}
// Same program, but every one of the 16 elements is asked for a DIFFERENT current value.
//
// An input array occupies one location per element (GL 4.6 core 11.1.1), so `in vec4 a[16]`
// at location 0 is active on 0..15 - and the whole location span is what a backend reads to
// decide which attributes need their current value pushed. Reflection used to record the
// span of the ELEMENT type only, so a 16-element array claimed exactly one location: every
// element above the first silently read the (0,0,0,1) an unwritten input defaults to instead
// of the value glVertexAttrib4f had set. The test above could not see it, because the only
// element it reads a current value from is element 0 - the one location the array did claim.
TEST_F(VertexAttribBindingScenario, EveryInputArrayElementGetsItsOwnCurrentValue) {
if (!Ready()) GTEST_SKIP();
const std::string vs = R"(#version 430 core
layout(location = 0) in vec4 vs_in_attrib[16];
out StageData {
vec4 attrib[16];
} vs_out;
void main() {
for (int i = 0; i < vs_in_attrib.length(); ++i) {
vs_out.attrib[i] = vs_in_attrib[i];
}
}
)";
std::vector<std::string> names;
for (int i = 0; i < 16; ++i) names.push_back("StageData.attrib[" + std::to_string(i) + "]");
std::vector<const char*> varyings;
for (const auto& n : names) varyings.push_back(n.c_str());
std::string log;
const GLuint program = BuildCaptureProgram(vs, varyings, &log);
ASSERT_NE(program, 0u) << "capture program did not link: " << log;
// Distinct in every component, and never (0,0,0,1): the value an element that was
// skipped would report has to be distinguishable from every value that was asked for.
for (GLuint i = 0; i < 16; ++i) {
const float base = static_cast<float>(i) + 1.0f;
glVertexAttrib4f(i, base, base + 100.0f, base + 200.0f, base + 300.0f);
}
constexpr std::size_t kFloatsPerPoint = 64;
std::vector<float> poison(kFloatsPerPoint, kPoison);
glBindBufferBase(GL_TRANSFORM_FEEDBACK_BUFFER, 0, m_xfbo);
glBufferData(GL_TRANSFORM_FEEDBACK_BUFFER, static_cast<GLsizeiptr>(poison.size() * sizeof(float)),
poison.data(), GL_DYNAMIC_DRAW);
glEnable(GL_RASTERIZER_DISCARD);
glUseProgram(program);
glBeginTransformFeedback(GL_POINTS);
glDrawArrays(GL_POINTS, 0, 1);
glEndTransformFeedback();
glDisable(GL_RASTERIZER_DISCARD);
std::vector<float> data(poison.size(), kPoison);
glGetBufferSubData(GL_TRANSFORM_FEEDBACK_BUFFER, 0,
static_cast<GLsizeiptr>(data.size() * sizeof(float)), data.data());
glUseProgram(0);
for (int element = 0; element < 16; ++element) {
const float base = static_cast<float>(element) + 1.0f;
EXPECT_FLOAT_EQ(data[element * 4 + 0], base) << "element " << element;
EXPECT_FLOAT_EQ(data[element * 4 + 1], base + 100.0f) << "element " << element;
EXPECT_FLOAT_EQ(data[element * 4 + 2], base + 200.0f) << "element " << element;
EXPECT_FLOAT_EQ(data[element * 4 + 3], base + 300.0f) << "element " << element;
}
for (GLuint i = 0; i < 16; ++i) glVertexAttrib4f(i, 0.0f, 0.0f, 0.0f, 0.0f);
glDeleteProgram(program);
}
} // namespace MGITest
@@ -68,6 +68,8 @@ namespace {
return type && type->getQualifier().builtIn != glslang::EbvNone; return type && type->getQualifier().builtIn != glslang::EbvNone;
} }
// Locations one ELEMENT of a vertex input occupies (GL 4.6 core 11.1.1): a matrix
// takes one per column, everything else this backend can feed takes one.
static int GetVertexInputLocationSpan(GLenum glType) { static int GetVertexInputLocationSpan(GLenum glType) {
switch (glType) { switch (glType) {
case GL_FLOAT_MAT2: case GL_FLOAT_MAT2:
@@ -87,6 +89,26 @@ namespace {
} }
} }
// How many elements an ARRAY vertex input has. glslang reflects such an input as ONE
// record spelled "name[0]" carrying the ELEMENT's glDefineType and the array length,
// so the type alone cannot say how many locations the declaration covers: GL 4.6 core
// 11.1.1 gives an array one location per element (times the element's own span), and
// `in vec4 a[16]` at location 0 therefore occupies 0..15, not 0. Missing that left
// every location above the base with no recorded name or type, which is what the
// backends read to decide whether an attribute is active at all.
static MobileGL::Int GetVertexInputArrayElements(const glslang::TObjectReflection& input) {
const glslang::TType* type = input.getType();
if (type == nullptr || !type->isArray()) return 1;
// An unsized input array has no span to compute; treat it as one element rather
// than guessing, so it can only ever under-claim locations.
if (!type->isSizedArray()) return 1;
return std::max(1, type->getCumulativeArraySize());
}
static MobileGL::Int GetVertexInputTotalLocationSpan(const glslang::TObjectReflection& input) {
return GetVertexInputLocationSpan(input.glDefineType) * GetVertexInputArrayElements(input);
}
static GLenum GetVertexInputLocationType(GLenum glType) { static GLenum GetVertexInputLocationType(GLenum glType) {
switch (glType) { switch (glType) {
case GL_FLOAT_MAT2: case GL_FLOAT_MAT2:
@@ -563,8 +585,20 @@ namespace MobileGL::MG_State::GLState {
// - SharedStd140UBO: a DECLARED uniform block is active even when no member is // - SharedStd140UBO: a DECLARED uniform block is active even when no member is
// ever read (reflected from the linker objects). PreprocessShaderSource coerces // ever read (reflected from the linker objects). PreprocessShaderSource coerces
// every block to std140, so this covers all of them. // every block to std140, so this covers all of them.
// - IntermediateIO: GL_PROGRAM_INPUT is the input interface of the program's FIRST
// stage and GL_PROGRAM_OUTPUT the output interface of its LAST one. Without this
// glslang hardcodes those boundaries to vertex/fragment, so a separable program
// made of one non-vertex stage has an empty input interface and one made of a
// non-fragment stage an empty output interface
// (KHR-GL43.program_interface_query.separate-programs-*).
// - UnwrapIOBlocks: an inter-stage interface block enumerates as its MEMBERS -
// "Color.r", and "gl_Position" for an anonymous gl_PerVertex - not as the block
// instance. Only reachable through IntermediateIO: a vertex stage's inputs and a
// fragment stage's outputs can never be blocks, so this is inert for a program
// whose boundary stages are the hardcoded ones.
if (!artifacts.program->buildReflection(EShReflectionStrictArraySuffix | EShReflectionBasicArraySuffix | if (!artifacts.program->buildReflection(EShReflectionStrictArraySuffix | EShReflectionBasicArraySuffix |
EShReflectionAllBlockVariables | EShReflectionSharedStd140UBO)) { EShReflectionAllBlockVariables | EShReflectionSharedStd140UBO |
EShReflectionIntermediateIO | EShReflectionUnwrapIOBlocks)) {
artifacts.linkStatus = false; artifacts.linkStatus = false;
artifacts.infoLog = "Build reflection failed."; artifacts.infoLog = "Build reflection failed.";
DeferLog(std::format("ProgramObject {}: DoReflection - buildReflection() returned false", DeferLog(std::format("ProgramObject {}: DoReflection - buildReflection() returned false",
@@ -852,14 +886,21 @@ namespace MobileGL::MG_State::GLState {
} }
// ------------ attributes (vertex in) --------------- // ------------ attributes (vertex in) ---------------
Int inCount = artifacts.program->getNumPipeInputs(); // The pipe-input list is the input interface of the program's FIRST stage, which is only
// the vertex attribute set when the program actually HAS a vertex stage. A separable
// fragment/geometry/tessellation program reflects its own stage inputs here, and those are
// varyings - registering them as vertex attributes would hand glGetActiveAttrib and the
// attribute location table interstage varyings.
Int inCount = artifacts.program->getIntermediate(EShLangVertex) != nullptr
? artifacts.program->getNumPipeInputs()
: 0;
MGLOG_D("ProgramObject %u: Reflection - pipe input count (attributes) = %d", in.externalIndex, inCount); MGLOG_D("ProgramObject %u: Reflection - pipe input count (attributes) = %d", in.externalIndex, inCount);
Int maxLoc = -1; Int maxLoc = -1;
for (int i = 0; i < inCount; ++i) { for (int i = 0; i < inCount; ++i) {
Int loc = (Int)artifacts.program->getPipeInput(i).layoutLocation(); Int loc = (Int)artifacts.program->getPipeInput(i).layoutLocation();
if (loc >= 0 && loc != glslang::TQualifier::layoutLocationEnd) { if (loc >= 0 && loc != glslang::TQualifier::layoutLocationEnd) {
const Int locationSpan = GetVertexInputLocationSpan(artifacts.program->getPipeInput(i).glDefineType); const Int locationSpan = GetVertexInputTotalLocationSpan(artifacts.program->getPipeInput(i));
maxLoc = std::max(maxLoc, loc + locationSpan - 1); maxLoc = std::max(maxLoc, loc + locationSpan - 1);
} }
MGLOG_D("ProgramObject %u: Reflection - pipe input[%d] name='%s' layoutLocation=%d glType=%u", MGLOG_D("ProgramObject %u: Reflection - pipe input[%d] name='%s' layoutLocation=%d glType=%u",
@@ -895,7 +936,7 @@ namespace MobileGL::MG_State::GLState {
(Int)ProgramObject::NormalizeBuiltinPipeInputName(inVar.name).length()); (Int)ProgramObject::NormalizeBuiltinPipeInputName(inVar.name).length());
if (location >= 0 && location < (int)artifacts.attribs.size()) { if (location >= 0 && location < (int)artifacts.attribs.size()) {
const Int locationSpan = GetVertexInputLocationSpan(inVar.glDefineType); const Int locationSpan = GetVertexInputTotalLocationSpan(inVar);
const GLenum locationType = GetVertexInputLocationType(inVar.glDefineType); const GLenum locationType = GetVertexInputLocationType(inVar.glDefineType);
for (Int locationOffset = 0; locationOffset < locationSpan; ++locationOffset) { for (Int locationOffset = 0; locationOffset < locationSpan; ++locationOffset) {
const Int expandedLocation = location + locationOffset; const Int expandedLocation = location + locationOffset;
@@ -951,6 +992,11 @@ namespace MobileGL::MG_State::GLState {
Bool ProgramLinkTask::ValidateFragmentOutputLocations() { Bool ProgramLinkTask::ValidateFragmentOutputLocations() {
if (!artifacts.program) return false; if (!artifacts.program) return false;
// The pipe-output list is the output interface of the program's LAST stage. Only a
// fragment stage's outputs are color numbers indexed against GL_MAX_DRAW_BUFFERS; a
// separable vertex/geometry/tessellation program's outputs are varyings, and holding
// them to the draw-buffer range fails the link of every such program.
if (artifacts.program->getIntermediate(EShLangFragment) == nullptr) return true;
UnorderedMap<Int, String> colorNumberOwners; UnorderedMap<Int, String> colorNumberOwners;
const Int outputCount = artifacts.program->getNumPipeOutputs(); const Int outputCount = artifacts.program->getNumPipeOutputs();
@@ -165,10 +165,18 @@ namespace MobileGL::MG_State::GLState {
const Int elements = init.arraySize; const Int elements = init.arraySize;
if (componentsPerElement <= 0 || elements <= 0) continue; if (componentsPerElement <= 0 || elements <= 0) continue;
const Bool isFloat = init.basicType == glslang::EbtFloat || init.basicType == glslang::EbtFloat16; // EbtDouble belongs with the floats now, not with the skipped types: every 64-bit
// float in a shader is narrowed to 32 bits before the module reaches a backend
// (ShaderTranspiler::DemoteFloat64Pass), so a `uniform double d = 1.5;` has exactly
// the 32-bit shadow encoding a `uniform float` does - and glslang already folded its
// value into floatValues, which is a vector<double> either way. Leaving it out meant
// the initializer was silently dropped and the uniform came up zero.
const Bool isFloat = init.basicType == glslang::EbtFloat ||
init.basicType == glslang::EbtFloat16 ||
init.basicType == glslang::EbtDouble;
const Bool isInt = init.basicType == glslang::EbtInt || init.basicType == glslang::EbtUint || const Bool isInt = init.basicType == glslang::EbtInt || init.basicType == glslang::EbtUint ||
init.basicType == glslang::EbtBool; init.basicType == glslang::EbtBool;
// Anything else (fp64, 64-bit integers) has no 32-bit shadow encoding here, and a // Anything else (64-bit integers) has no 32-bit shadow encoding here, and a
// half-written uniform is worse than an untouched one. // half-written uniform is worse than an untouched one.
if (!isFloat && !isInt) continue; if (!isFloat && !isInt) continue;
const SizeT provided = isFloat ? init.floatValues.size() : init.intValues.size(); const SizeT provided = isFloat ? init.floatValues.size() : init.intValues.size();
@@ -326,15 +326,22 @@ namespace MobileGL::MG_State::GLState {
: kInvalidUniformOffset; : kInvalidUniformOffset;
} }
Uint GetUniformSizesInBytes(Uint location) const { return MG_Util::GetGLTypeSize(GetUniformType(location)); } Uint GetUniformSizesInBytes(Uint location) const { return MG_Util::GetGLTypeSize(GetUniformType(location)); }
// Bytes a uniform actually occupies in the global UBO, which is not its GL type size: // Bytes a uniform actually occupies in the global UBO, which is not its GL type size,
// std140 pads each column of a float matrix out to a vec4, so a mat3 spans 48 bytes // for two reasons. std140 pads each column of a matrix out to a vec4, so a mat3 spans
// even though only 36 of them carry components. Anything reading or writing a whole // 48 bytes even though only 36 of them carry components. And every 64-bit float in a
// uniform's storage - a bounds check, a copy between two programs' shadows - wants // shader is narrowed to 32 bits before the module reaches a backend
// this rather than GetUniformSizesInBytes. // (ShaderTranspiler::DemoteFloat64Pass) - the global UBO is laid out by reflecting that
// demoted module - so a `double` uniform occupies exactly what its float-typed twin
// would, half its GL type size, and a `dmat4` is padded like any other matrix. Anything
// reading or writing a whole uniform's storage - a bounds check, a copy between two
// programs' shadows - wants this rather than GetUniformSizesInBytes.
static SizeT UniformStorageSpanInBytes(const glslang::TType* type, SizeT tightSize) { static SizeT UniformStorageSpanInBytes(const glslang::TType* type, SizeT tightSize) {
if (type != nullptr && type->isMatrix() && type->getBasicType() != glslang::EbtDouble) { if (type != nullptr && type->isMatrix()) {
return static_cast<SizeT>(type->getMatrixCols()) * 4 * sizeof(Float); return static_cast<SizeT>(type->getMatrixCols()) * 4 * sizeof(Float);
} }
if (type != nullptr && type->getBasicType() == glslang::EbtDouble) {
return tightSize / 2;
}
return tightSize; return tightSize;
} }
SizeT GetUniformStorageSpanInBytes(Uint location) const { SizeT GetUniformStorageSpanInBytes(Uint location) const {
@@ -61,12 +61,16 @@ namespace MobileGL::MG_State::GLState {
} }
void VertexArrayObject::SetAttributeFormat(Uint index, int size, DataType type, Bool normalized, int stride, void VertexArrayObject::SetAttributeFormat(Uint index, int size, DataType type, Bool normalized, int stride,
SizeT offset, Bool isInteger, Bool isBgra) { SizeT offset, Bool isInteger, Bool isBgra, int effectiveStride) {
if (index >= MAX_VERTEX_ATTRIBS) return; if (index >= MAX_VERTEX_ATTRIBS) return;
if (size < 1 || size > 4) { if (size < 1 || size > 4) {
return; return;
} }
// See VertexAttribute::Stride: the resolved field carries the effective stride so that
// a zero in it can only ever mean the binding model's "do not advance".
const int resolvedStride = effectiveStride >= 0 ? effectiveStride : stride;
// The classic pointer-style API takes back full ownership of the resolved fields. // The classic pointer-style API takes back full ownership of the resolved fields.
m_attributeUsesBindingModel[index] = false; m_attributeUsesBindingModel[index] = false;
@@ -77,7 +81,7 @@ namespace MobileGL::MG_State::GLState {
m_attributes[index].LegacyPointer = offset; m_attributes[index].LegacyPointer = offset;
if (m_attributes[index].Size == size && m_attributes[index].Type == type && if (m_attributes[index].Size == size && m_attributes[index].Type == type &&
m_attributes[index].Normalized == normalized && m_attributes[index].Stride == stride && m_attributes[index].Normalized == normalized && m_attributes[index].Stride == resolvedStride &&
m_attributes[index].Offset == offset && m_attributes[index].IsInteger == isInteger && m_attributes[index].Offset == offset && m_attributes[index].IsInteger == isInteger &&
m_attributes[index].IsBgra == isBgra && !m_attributes[index].IsLong) { m_attributes[index].IsBgra == isBgra && !m_attributes[index].IsLong) {
return; return;
@@ -87,7 +91,7 @@ namespace MobileGL::MG_State::GLState {
attr.Size = size; attr.Size = size;
attr.Type = type; attr.Type = type;
attr.Normalized = normalized; attr.Normalized = normalized;
attr.Stride = stride; attr.Stride = resolvedStride;
attr.Offset = offset; attr.Offset = offset;
attr.IsInteger = isInteger; attr.IsInteger = isInteger;
attr.IsBgra = isBgra; attr.IsBgra = isBgra;
@@ -19,6 +19,14 @@ namespace MobileGL {
int Size = 4; int Size = 4;
DataType Type = DataType::Float32; DataType Type = DataType::Float32;
Bool Normalized = false; Bool Normalized = false;
// The RESOLVED byte distance between consecutive elements, never the raw
// glVertexAttrib*Pointer argument: a pointer call's stride 0 means "tightly
// packed" and is resolved to the element size here, so a zero that survives
// into this field can only have come from the binding model, where a zero
// VERTEX_BINDING_STRIDE means the opposite - every vertex reads the SAME
// element and the fetch address never advances (GL 4.6 core 10.3.1). Backends
// consume this verbatim; collapsing 0 back into the element size is what made
// KHR-GL43.vertex_attrib_binding.basic-input-case7/8 read past the buffer.
int Stride = 0; int Stride = 0;
SizeT Offset = 0; SizeT Offset = 0;
Bool IsInteger = false; Bool IsInteger = false;
@@ -76,8 +84,12 @@ namespace MobileGL {
void DisableAttribute(Uint index); void DisableAttribute(Uint index);
Bool IsAttributeEnabled(Uint index) const; Bool IsAttributeEnabled(Uint index) const;
// `stride` is the raw glVertexAttrib*Pointer argument, reported verbatim by
// GL_VERTEX_ATTRIB_ARRAY_STRIDE. `effectiveStride` is what the fetch actually
// advances by - the same value when the argument is non-zero, the tightly
// packed element size when it is zero. Pass -1 to say the two are the same.
void SetAttributeFormat(Uint index, int size, DataType type, Bool normalized, int stride, SizeT offset, void SetAttributeFormat(Uint index, int size, DataType type, Bool normalized, int stride, SizeT offset,
Bool isInteger, Bool isBgra = false); Bool isInteger, Bool isBgra = false, int effectiveStride = -1);
void BindAttributeBuffer(Uint index, const SharedPtr<BufferObject>& buffer); void BindAttributeBuffer(Uint index, const SharedPtr<BufferObject>& buffer);
+163 -11
View File
@@ -121,12 +121,6 @@ void main() { color = vec4(0, 1, 0, 1); }
ExpectLinked(p); ExpectLinked(p);
ClearErrors(); ClearErrors();
// KNOWN GAP, not an expectation: a separable FRAGMENT program's own inputs are not
// in the reflection at all - glslang builds the "pipe input" list from the vertex
// stage unless EShReflectionIntermediateIO is set, and setting that makes a
// vertex-only separable program report its VS outputs as fragment outputs, which
// fails ValidateFragmentOutputLocations and breaks glCreateShaderProgramv. So
// GL_PROGRAM_INPUT is empty here until that validation is stage-aware.
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_ACTIVE_RESOURCES), 1); EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_ACTIVE_RESOURCES), 1);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_MAX_NAME_LENGTH), 6); EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_MAX_NAME_LENGTH), 6);
@@ -804,6 +798,163 @@ void main(void) {
EXPECT_EQ(TakeError(), GL_NO_ERROR); EXPECT_EQ(TakeError(), GL_NO_ERROR);
} }
// ------------------------------------------------------------- uniform-block-types --
// GL_REFERENCED_BY_*_SHADER is per block INSTANCE. An array of uniform blocks enumerates
// one resource per element, and only the elements a stage actually dereferences are
// referenced by it - declaring the array is not referencing every element.
TEST_F(ProgramInterfaceTest, ArrayedUniformBlockReferencedByIsPerElement) {
const char* vs = R"(#version 430
in vec4 position;
uniform SimpleBlock { mat3x2 a; mat4 b; vec4 c; };
void main(void) {
float tmp = a[0][1] * b[1][2] * c.x;
gl_Position = position * tmp;
}
)";
const char* fs = R"(#version 430
uniform TrickyBlock { mat4 b; uint c; } e[2];
out vec4 color;
void main() { color = vec4(0, 1, 0, 1) * e[0].b[0][0]; }
)";
const GLuint p = MakeProgram(vs, fs);
BindAttribLocation(p, 0, "position");
BindFragDataLocation(p, 0, "color");
LinkProgram(p);
ExpectLinked(p);
ClearErrors();
EXPECT_EQ(Interfaceiv(p, GL_UNIFORM_BLOCK, GL_ACTIVE_RESOURCES), 3);
ExpectResource(p, GL_UNIFORM_BLOCK, "TrickyBlock[0]", "TrickyBlock[0]");
ExpectResource(p, GL_UNIFORM_BLOCK, "TrickyBlock[1]", "TrickyBlock[1]");
const std::vector<GLenum> refProps = {GL_REFERENCED_BY_VERTEX_SHADER, GL_REFERENCED_BY_FRAGMENT_SHADER};
EXPECT_EQ(PropsOf(p, GL_UNIFORM_BLOCK, "SimpleBlock", refProps), (std::vector<GLint>{1, 0}));
EXPECT_EQ(PropsOf(p, GL_UNIFORM_BLOCK, "TrickyBlock[0]", refProps), (std::vector<GLint>{0, 1}));
EXPECT_EQ(PropsOf(p, GL_UNIFORM_BLOCK, "TrickyBlock[1]", refProps), (std::vector<GLint>{0, 0}))
<< "the unreferenced element of a block array must not inherit its sibling's stages";
EXPECT_EQ(TakeError(), GL_NO_ERROR);
}
// The boundary of the rule above, and the case that caught it on device
// (KHR-GL43.program_interface_query.ssb-types): a SHADER STORAGE block array's buffer
// variables reflect under ONE subscript-free spelling shared by every element, so a union
// over them credits element 0 and starves every other element - even the ones the shader
// plainly reads. Storage blocks keep glslang's own mask, and both elements here must report
// the fragment stage.
TEST_F(ProgramInterfaceTest, ArrayedStorageBlockKeepsGlslangStagesForEveryElement) {
const char* vs = R"(#version 430
in vec4 position;
void main(void) { gl_Position = position; }
)";
const char* fs = R"(#version 430
layout(binding = 4) buffer SimpleStorage { vec4 a; } ss[2];
out vec4 color;
void main() { color = ss[0].a + ss[1].a; }
)";
const GLuint p = MakeProgram(vs, fs);
BindAttribLocation(p, 0, "position");
BindFragDataLocation(p, 0, "color");
LinkProgram(p);
ExpectLinked(p);
ClearErrors();
const std::vector<GLenum> refProps = {GL_REFERENCED_BY_VERTEX_SHADER, GL_REFERENCED_BY_FRAGMENT_SHADER};
for (const char* name : {"SimpleStorage[0]", "SimpleStorage[1]"}) {
EXPECT_EQ(PropsOf(p, GL_SHADER_STORAGE_BLOCK, name, refProps), (std::vector<GLint>{0, 1}))
<< "for " << name << ": a storage block the fragment stage reads must say so";
}
EXPECT_EQ(TakeError(), GL_NO_ERROR);
}
// --------------------------------------------------------- separate-programs-vertex --
// A vertex-only separable program's OUTPUT interface is its own stage outputs, and an
// inter-stage block enumerates as its members: "Color.r", and "gl_Position" for the
// anonymous gl_PerVertex redeclaration - never the block instance name "vs_color".
TEST_F(ProgramInterfaceTest, SeparableVertexProgramEnumeratesItsOwnOutputBlockMembers) {
const char* vs = R"(#version 430 core
layout(location = 0) in vec4 in_vertex;
out Color { float r, g, b; vec4 iLikePie; } vs_color;
out gl_PerVertex { vec4 gl_Position; };
uniform float u;
uniform vec4 v;
void main() {
gl_Position = in_vertex;
vs_color.r = u; vs_color.g = 0.0; vs_color.b = 0.0; vs_color.iLikePie = v;
}
)";
const GLuint p = CreateShaderProgramv(GL_VERTEX_SHADER, 1, &vs);
ExpectLinked(p);
ClearErrors();
// Exactly 5: the four Color members plus gl_Position. The anonymous gl_PerVertex
// redeclaration drops gl_PointSize and gl_ClipDistance, which glslang keeps in the
// block type as hidden members rather than erasing.
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_ACTIVE_RESOURCES), 1);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_MAX_NAME_LENGTH), 10);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_ACTIVE_RESOURCES), 5);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_MAX_NAME_LENGTH), 15);
ExpectResource(p, GL_PROGRAM_INPUT, "in_vertex", "in_vertex");
for (const char* name : {"Color.r", "Color.g", "Color.b", "Color.iLikePie", "gl_Position"}) {
ExpectResource(p, GL_PROGRAM_OUTPUT, name, name);
}
EXPECT_EQ(GetProgramResourceIndex(p, GL_PROGRAM_OUTPUT, "vs_color"), GL_INVALID_INDEX)
<< "the block instance is not the resource; its members are";
// A vertex-stage output has no color number, hence no index either, and it is not
// per-patch. NAME_LENGTH/TYPE/ARRAY_SIZE come from the member, not the block.
EXPECT_EQ(PropsOf(p, GL_PROGRAM_OUTPUT, "Color.iLikePie",
{GL_NAME_LENGTH, GL_TYPE, GL_ARRAY_SIZE, GL_REFERENCED_BY_COMPUTE_SHADER,
GL_REFERENCED_BY_FRAGMENT_SHADER, GL_REFERENCED_BY_GEOMETRY_SHADER,
GL_REFERENCED_BY_TESS_CONTROL_SHADER, GL_REFERENCED_BY_TESS_EVALUATION_SHADER,
GL_REFERENCED_BY_VERTEX_SHADER, GL_IS_PER_PATCH, GL_LOCATION_INDEX}),
(std::vector<GLint>{15, GL_FLOAT_VEC4, 1, 0, 0, 0, 0, 0, 1, 0, -1}));
EXPECT_EQ(TakeError(), GL_NO_ERROR);
}
// ------------------------------------------------------- separate-programs-geometry --
// Both boundaries are a middle stage here. The input block carries an instance name
// (gl_in[]) so its members are prefixed with the BLOCK name, and the arrayed-ness of the
// block itself is dropped: "gl_PerVertex.gl_Position", array size 1.
TEST_F(ProgramInterfaceTest, SeparableGeometryProgramEnumeratesBothBlockBoundaries) {
const char* gs = R"(#version 430
layout(triangles) in;
layout(triangle_strip, max_vertices = 4) out;
out gl_PerVertex { vec4 gl_Position; float gl_PointSize; float gl_ClipDistance[]; };
in gl_PerVertex { vec4 gl_Position; float gl_PointSize; float gl_ClipDistance[]; } gl_in[];
void main() {
gl_Position = vec4(-1, 1, 0, 1); EmitVertex();
gl_Position = gl_in[0].gl_Position; EmitVertex();
EndPrimitive();
}
)";
const GLuint p = CreateShaderProgramv(GL_GEOMETRY_SHADER, 1, &gs);
ExpectLinked(p);
ClearErrors();
ExpectResource(p, GL_PROGRAM_INPUT, "gl_PerVertex.gl_Position", "gl_PerVertex.gl_Position");
ExpectResource(p, GL_PROGRAM_OUTPUT, "gl_Position", "gl_Position");
// A non-fragment stage's outputs are varyings: no color number, so no color index.
EXPECT_EQ(PropsOf(p, GL_PROGRAM_OUTPUT, "gl_Position", {GL_LOCATION_INDEX}), (std::vector<GLint>{-1}));
const std::vector<GLenum> stageProps = {
GL_NAME_LENGTH,
GL_TYPE,
GL_ARRAY_SIZE,
GL_REFERENCED_BY_COMPUTE_SHADER,
GL_REFERENCED_BY_FRAGMENT_SHADER,
GL_REFERENCED_BY_GEOMETRY_SHADER,
GL_REFERENCED_BY_TESS_CONTROL_SHADER,
GL_REFERENCED_BY_TESS_EVALUATION_SHADER,
GL_REFERENCED_BY_VERTEX_SHADER,
GL_IS_PER_PATCH};
EXPECT_EQ(PropsOf(p, GL_PROGRAM_INPUT, "gl_PerVertex.gl_Position", stageProps),
(std::vector<GLint>{25, GL_FLOAT_VEC4, 1, 0, 0, 1, 0, 0, 0, 0}));
EXPECT_EQ(PropsOf(p, GL_PROGRAM_OUTPUT, "gl_Position", stageProps),
(std::vector<GLint>{12, GL_FLOAT_VEC4, 1, 0, 0, 1, 0, 0, 0, 0}));
EXPECT_EQ(TakeError(), GL_NO_ERROR);
}
// ------------------------------------------------------- separate-programs-fragment -- // ------------------------------------------------------- separate-programs-fragment --
TEST_F(ProgramInterfaceTest, SeparableFragmentProgramSeparatesUniformsFromBufferVariables) { TEST_F(ProgramInterfaceTest, SeparableFragmentProgramSeparatesUniformsFromBufferVariables) {
const char* fs = R"(#version 430 const char* fs = R"(#version 430
@@ -817,11 +968,12 @@ void main() { fs_color = vs_color + x + a; }
ExpectLinked(p); ExpectLinked(p);
ClearErrors(); ClearErrors();
// KNOWN GAP, not an expectation - the same one SimpleShaders documents: a separable // GL_PROGRAM_INPUT is the input interface of the program's FIRST stage, which for a
// FRAGMENT program's own inputs are absent from the glslang reflection, so // separable fragment program is the fragment stage: "vs_color", name length 9.
// GL_PROGRAM_INPUT is empty. Spec-correct values here would be 1 and 9 ("vs_color"). EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_ACTIVE_RESOURCES), 1);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_ACTIVE_RESOURCES), 0); EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_MAX_NAME_LENGTH), 9);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_INPUT, GL_MAX_NAME_LENGTH), 0); EXPECT_EQ(GetProgramResourceIndex(p, GL_PROGRAM_INPUT, "vs_color"), 0u);
EXPECT_EQ(ResourceName(p, GL_PROGRAM_INPUT, 0), "vs_color");
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_ACTIVE_RESOURCES), 1); EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_ACTIVE_RESOURCES), 1);
EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_MAX_NAME_LENGTH), 9); EXPECT_EQ(Interfaceiv(p, GL_PROGRAM_OUTPUT, GL_MAX_NAME_LENGTH), 9);
// The buffer variable is NOT a uniform, even though the frontend reflection keeps // The buffer variable is NOT a uniform, even though the frontend reflection keeps
+36 -2
View File
@@ -346,8 +346,10 @@ void main() {
EXPECT_NE(rewritten.find("int instance = mg_ZeroBasedInstanceID + mg_BaseInstanceLowered;"), EXPECT_NE(rewritten.find("int instance = mg_ZeroBasedInstanceID + mg_BaseInstanceLowered;"),
MobileGL::String::npos); MobileGL::String::npos);
EXPECT_NE(rewritten.find("#define mg_ZeroBasedInstanceID (gl_InstanceID - ((mg_BaseInstanceWordIndex >= 0) ? " // One-based word index: zero is the "not an indirect draw" sentinel because that is
"int(mg_indirectWords[uint(mg_BaseInstanceWordIndex)]) : 0))"), // the value a GLSL uniform starts at and no draw path writes it before the first draw.
EXPECT_NE(rewritten.find("#define mg_ZeroBasedInstanceID (gl_InstanceID - ((mg_BaseInstanceWordIndex > 0) ? "
"int(mg_indirectWords[uint(mg_BaseInstanceWordIndex - 1)]) : 0))"),
MobileGL::String::npos); MobileGL::String::npos);
EXPECT_NE(rewritten.find( EXPECT_NE(rewritten.find(
"layout(std430, binding = 12) readonly buffer mg_IndirectParams { highp uint mg_indirectWords[]; };"), "layout(std430, binding = 12) readonly buffer mg_IndirectParams { highp uint mg_indirectWords[]; };"),
@@ -356,6 +358,38 @@ void main() {
EXPECT_EQ(CountOccurrences(rewritten, "gl_InstanceID"), 1u); EXPECT_EQ(CountOccurrences(rewritten, "gl_InstanceID"), 1u);
} }
// The sentinel itself, on the builtin it exists for. A zero-based index with a
// negative "off" value made every NON-indirect draw of such a program read
// mg_indirectWords[0] out of a storage buffer nothing had bound - the uniform starts
// at zero and no non-indirect draw path writes it - which is where the CTS
// shader_draw_parameters cases lost their geometry on Adreno. Pinned as text because
// this contract lives in two places at once: the generated ESSL below and the +1 that
// BackendProgramObjectImpl::SetBaseInstanceWordIndex applies.
TEST(DirectGLESSanity, TheIndirectWordIndexIsOneBasedSoItsUnwrittenValueMeansNotIndirect) {
const ScopedGLESCapabilitiesOverride capsGuard;
auto& caps = MobileGL::MG_Backend::DirectGLES::g_GLESCapabilities;
caps.IndirectDrawInstanceIdIncludesBaseInstance = false;
caps.MaxShaderStorageBufferBindings = 13;
const MobileGL::String source = R"(#version 310 es
highp int mg_BaseInstanceLowered;
void main() {
gl_Position = vec4(float(mg_BaseInstanceLowered));
}
)";
const auto rewritten = MobileGL::MG_Backend::DirectGLES::PromoteDrawParameterGlobalsToUniforms(
source, GL_VERTEX_SHADER);
EXPECT_NE(rewritten.find("#define mg_BaseInstanceLowered ((mg_BaseInstanceWordIndex > 0) ? "
"int(mg_indirectWords[uint(mg_BaseInstanceWordIndex - 1)]) : mg_BaseInstance)"),
MobileGL::String::npos)
<< rewritten;
// A zero-based form would spell either of these; neither may survive.
EXPECT_EQ(rewritten.find("mg_BaseInstanceWordIndex >= 0"), MobileGL::String::npos);
EXPECT_EQ(rewritten.find("uint(mg_BaseInstanceWordIndex)"), MobileGL::String::npos);
}
TEST(DirectGLESSanity, KeepsInstanceIdWhenIndirectDrawsAreConforming) { TEST(DirectGLESSanity, KeepsInstanceIdWhenIndirectDrawsAreConforming) {
const ScopedGLESCapabilitiesOverride capsGuard; const ScopedGLESCapabilitiesOverride capsGuard;
auto& caps = MobileGL::MG_Backend::DirectGLES::g_GLESCapabilities; auto& caps = MobileGL::MG_Backend::DirectGLES::g_GLESCapabilities;
@@ -3,6 +3,8 @@ cmake_minimum_required(VERSION 3.14)
add_executable( add_executable(
SpirvPassTest SpirvPassTest
SpirvPassTest.cpp SpirvPassTest.cpp
DemoteFloat64Test.cpp
FlattenXfbInterfaceBlocksTest.cpp
) )
target_include_directories(SpirvPassTest PRIVATE target_include_directories(SpirvPassTest PRIVATE
@@ -0,0 +1,486 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/DemoteFloat64Test.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <string>
#include <vector>
#include "Includes.h"
#include "Init.h"
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <MG_Util/ShaderTranspiler/SpvcSession.h>
#include <MG_Util/ShaderTranspiler/Types.h>
#include <spirv-tools/libspirv.hpp>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
namespace {
// The test-side reference walker, deliberately independent of the production code: a bug in
// the pass must not be able to hide behind the same helper. Counts OpTypeFloat declarations of
// a given width and collects the Offset literal of every OpMemberDecorate, in module order.
constexpr Uint32 kSpirvHeaderWordCount = 5;
constexpr Uint32 kOpTypeFloat = 22;
constexpr Uint32 kOpName = 5;
constexpr Uint32 kOpMemberDecorate = 72;
constexpr Uint32 kOpFConvert = 115;
constexpr Uint32 kOpCapability = 17;
constexpr Uint32 kDecorationOffset = 35;
constexpr Uint32 kCapabilityFloat64 = 10;
template <typename Visitor>
void ForEachInstruction(const Vector<Uint32>& spirv, Visitor&& visit) {
for (SizeT i = kSpirvHeaderWordCount; i < spirv.size();) {
const Uint32 wordCount = spirv[i] >> 16;
const Uint32 opcode = spirv[i] & 0xFFFFu;
if (wordCount == 0 || i + wordCount > spirv.size()) break;
visit(opcode, &spirv[i], wordCount);
i += wordCount;
}
}
Uint32 CountFloatTypesOfWidth(const Vector<Uint32>& spirv, Uint32 width) {
Uint32 count = 0;
ForEachInstruction(spirv, [&](Uint32 opcode, const Uint32* words, Uint32 wordCount) {
if (opcode == kOpTypeFloat && wordCount >= 3 && words[2] == width) ++count;
});
return count;
}
Uint32 CountFConverts(const Vector<Uint32>& spirv) {
Uint32 count = 0;
ForEachInstruction(spirv, [&](Uint32 opcode, const Uint32*, Uint32) {
if (opcode == kOpFConvert) ++count;
});
return count;
}
Bool DeclaresFloat64Capability(const Vector<Uint32>& spirv) {
Bool found = false;
ForEachInstruction(spirv, [&](Uint32 opcode, const Uint32* words, Uint32 wordCount) {
if (opcode == kOpCapability && wordCount >= 2 && words[1] == kCapabilityFloat64) found = true;
});
return found;
}
// Byte offset of every member of the struct named `blockName`, in member order.
Vector<Uint32> CollectOffsetsOf(const Vector<Uint32>& spirv, const String& blockName) {
Uint32 structId = 0;
ForEachInstruction(spirv, [&](Uint32 opcode, const Uint32* words, Uint32 wordCount) {
if (opcode != kOpName || wordCount < 3 || structId != 0) return;
const char* text = reinterpret_cast<const char*>(&words[2]);
const SizeT maxBytes = (wordCount - 2) * sizeof(Uint32);
if (std::strncmp(text, blockName.c_str(), maxBytes) == 0) structId = words[1];
});
if (structId == 0) return {};
std::map<Uint32, Uint32> offsetByMember;
ForEachInstruction(spirv, [&](Uint32 opcode, const Uint32* words, Uint32 wordCount) {
if (opcode == kOpMemberDecorate && wordCount >= 5 && words[1] == structId &&
words[3] == kDecorationOffset) {
offsetByMember[words[2]] = words[4];
}
});
Vector<Uint32> offsets;
for (const auto& [member, offset] : offsetByMember) offsets.push_back(offset);
return offsets;
}
// Everything the production pipeline does to a source before the pass sees it.
Vector<Uint32> CompileToSpirv(GLenum stage, const String& source) {
using namespace MG_Util::ShaderTranspiler;
ShaderAttrib shaderAttrib{.shaderType = stage, .sourceStr = source};
auto shaderResult = ShaderCompiler::CompileShader(shaderAttrib);
EXPECT_TRUE(shaderResult) << (shaderResult ? String{} : shaderResult.error().log);
if (!shaderResult) return {};
ProgramAttrib programAttrib{.shaders = {shaderResult.value()}};
auto programResult = ShaderCompiler::LinkProgram(programAttrib);
EXPECT_TRUE(programResult) << (programResult ? String{} : programResult.error().log);
if (!programResult) return {};
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {stage}, .program = *programResult.value()};
auto binaryResult = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
EXPECT_TRUE(binaryResult) << (binaryResult ? String{} : binaryResult.error().log);
if (!binaryResult || binaryResult->empty()) return {};
return binaryResult->front();
}
String Disassemble(const Vector<Uint32>& spirv) {
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String text;
tools.Disassemble(spirv, &text);
return text;
}
// A vertex shader that exercises every shape the pass has to handle at once: a block with
// double / dvec2 / dvec3 / dvec4 / dmat4 members between two floats (so a shifted offset would
// be visible), a default-block double uniform, a 64-bit vertex input, a double-typed array, an
// implicit float->double conversion and an explicit double->float one.
const char* kWideVertexSource = R"(#version 460 core
layout(std140, binding = 0) uniform Blk {
float a;
double d;
dvec2 v2;
dvec3 v3;
dvec4 v4;
dmat4 m4;
double arr[3];
float z;
};
layout(location = 0) uniform double uScale;
layout(location = 0) in dvec3 inPos;
layout(location = 1) in vec3 inNormal;
layout(location = 0) out float vOut;
void main() {
double s = d * uScale + a;
dvec3 p = inPos * v3 + v2.xyx + v4.xyz + dvec3(m4[0].xyz);
s += p.x + p.y + p.z + arr[0] + arr[1] + arr[2] + z + 0.5lf;
vOut = float(s) + inNormal.x;
gl_Position = vec4(float(s));
}
)";
} // namespace
class DemoteFloat64Test : public ::testing::Test {
protected:
void SetUp() override {
MobileGL::Initialize();
ShaderCompiler::SetSpirvValidationEnabled(true);
m_validationFailuresAtStart = ShaderCompiler::SpirvValidationFailureCount();
}
void TearDown() override {
// The wrapper validates its OUTPUT on every run, so this covers every demotion the test
// performed without any of them having to say so.
EXPECT_EQ(ShaderCompiler::SpirvValidationFailureCount(), m_validationFailuresAtStart)
<< "the demoted module did not survive spirv-val";
}
Uint64 m_validationFailuresAtStart = 0;
};
TEST_F(DemoteFloat64Test, DemotesEveryWidthAndDropsTheCapability) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kWideVertexSource);
ASSERT_FALSE(input.empty());
ASSERT_EQ(CountFloatTypesOfWidth(input, 64), 1u) << Disassemble(input);
ASSERT_TRUE(DeclaresFloat64Capability(input));
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
EXPECT_EQ(CountFloatTypesOfWidth(output, 64), 0u) << Disassemble(output);
// And exactly one 32-bit float type survives: the merge has to happen, or spirv-val rejects
// the second declaration.
EXPECT_EQ(CountFloatTypesOfWidth(output, 32), 1u) << Disassemble(output);
EXPECT_FALSE(DeclaresFloat64Capability(output));
}
TEST_F(DemoteFloat64Test, RederivesTheStd140LayoutOfADemotedUniformBlock) {
const String source = R"(#version 460 core
layout(std140, binding = 0) uniform Blk {
float a;
double d;
dvec2 v2;
dvec3 v3;
dvec4 v4;
dmat4 m4;
double arr[3];
float z;
};
layout(location = 0) out float vOut;
void main() {
vOut = float(d + v2.x + v3.y + v4.z + m4[2].w + arr[1] + a + z);
gl_Position = vec4(vOut);
}
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
// What glslang laid out for the 64-bit members, which is what an application computing
// std140 by hand would also get.
EXPECT_EQ(CollectOffsetsOf(input, "Blk"), (Vector<Uint32>{0, 8, 16, 32, 64, 96, 224, 272}))
<< Disassemble(input);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// std140 for the demoted members: float at 4, vec2 at 8, vec3 at 16 (aligned like a vec4),
// vec4 at 32, mat4 at 48 with a 16-byte column stride, the array at 112 with the std140
// 16-byte element stride, and the trailing float at 160. This is what SPIRV-Cross has to be
// able to re-derive for GLSL ES, which has no member layout(offset=) to fall back on.
EXPECT_EQ(CollectOffsetsOf(output, "Blk"), (Vector<Uint32>{0, 4, 8, 16, 32, 48, 112, 160}))
<< Disassemble(output);
const String text = Disassemble(output);
EXPECT_NE(text.find("MatrixStride 16"), String::npos) << text;
EXPECT_NE(text.find("ArrayStride 16"), String::npos) << text;
}
TEST_F(DemoteFloat64Test, RederivesTheStd430LayoutOfADemotedStorageBlock) {
const String source = R"(#version 460 core
layout(std430, binding = 0) buffer Ssbo {
double head;
dvec4 wide;
double tail[4];
};
layout(location = 0) out float vOut;
void main() {
vOut = float(head + wide.w + tail[3]);
gl_Position = vec4(vOut);
}
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
EXPECT_EQ(CollectOffsetsOf(input, "Ssbo"), (Vector<Uint32>{0, 32, 64})) << Disassemble(input);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// std430, so the array packs at its element size rather than being rounded to 16: float at 0,
// vec4 at 16, float[4] at 32 with a 4-byte stride. A storage block must NOT come out std140,
// which is the whole reason the packing is chosen per storage class.
EXPECT_EQ(CollectOffsetsOf(output, "Ssbo"), (Vector<Uint32>{0, 16, 32})) << Disassemble(output);
EXPECT_NE(Disassemble(output).find("ArrayStride 4"), String::npos) << Disassemble(output);
}
TEST_F(DemoteFloat64Test, LeavesTheLayoutOfABlockWithoutDoublesAlone) {
const String source = R"(#version 460 core
layout(std140, binding = 0) uniform Blk {
float a;
vec3 v3;
mat4 m4;
};
layout(std140, binding = 1) uniform Wide {
float w;
double d;
};
layout(location = 0) out float vOut;
void main() {
vOut = float(a + v3.y + m4[1].z + float(d) + w);
gl_Position = vec4(vOut);
}
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
const Vector<Uint32> before = CollectOffsetsOf(input, "Blk");
ASSERT_FALSE(before.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// Only the block that actually narrowed is re-laid-out. Touching the other one would be
// churn at best, and a disagreement with glslang's own layout at worst.
EXPECT_EQ(CollectOffsetsOf(output, "Blk"), before) << Disassemble(output);
EXPECT_EQ(CollectOffsetsOf(output, "Wide"), (Vector<Uint32>{0, 4})) << Disassemble(output);
}
TEST_F(DemoteFloat64Test, FoldsTheConversionsThatBecameIdentities) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kWideVertexSource);
ASSERT_FALSE(input.empty());
ASSERT_GT(CountFConverts(input), 0u) << "the fixture no longer converts between the two widths";
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// SPIR-V requires the two component widths of an OpFConvert to differ, so every one of them
// has to be gone: both sides are 32 bits now.
EXPECT_EQ(CountFConverts(output), 0u) << Disassemble(output);
}
TEST_F(DemoteFloat64Test, NarrowsDoubleConstantsToTheirFloatValue) {
const String source = R"(#version 460 core
layout(location = 0) out float outValue;
void main() {
double d = 0.5lf;
outValue = float(d * 0.25lf);
gl_Position = vec4(0.0);
}
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// A 64-bit literal is two words wide and a 32-bit one is a single word, so a constant left
// unconverted is not merely imprecise - it is an unparseable instruction. Disassembling both
// values proves the re-encode produced the right number, not just the right width.
const String text = Disassemble(output);
EXPECT_NE(text.find("OpConstant %float 0.5"), String::npos) << text;
EXPECT_NE(text.find("OpConstant %float 0.25"), String::npos) << text;
}
TEST_F(DemoteFloat64Test, LeavesAModuleWithoutDoublesByteIdentical) {
const String source = R"(#version 460 core
layout(location = 0) in vec4 inPos;
void main() { gl_Position = inPos; }
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
// The pass reports SuccessWithoutChange here, and SPIRV-Tools asserts (in assert-enabled
// builds) that such a run round-trips byte-identically.
EXPECT_EQ(output, input);
}
TEST_F(DemoteFloat64Test, DeclinesAModuleThatBitcastsAcrossTheWidthBoundary) {
// packDouble2x32 is defined only for a 64-bit result: there is no 32-bit answer to give, and
// narrowing one side of the surrounding OpBitcast alone produces a module spirv-val rejects.
// The contract is that such a module comes back untouched rather than broken.
const String source = R"(#version 460 core
#extension GL_ARB_gpu_shader_fp64 : require
layout(location = 0) uniform uvec2 uPacked;
layout(location = 0) out float outValue;
void main() {
double d = packDouble2x32(uPacked);
outValue = float(d);
gl_Position = vec4(0.0);
}
)";
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, source);
ASSERT_FALSE(input.empty());
ASSERT_EQ(CountFloatTypesOfWidth(input, 64), 1u) << Disassemble(input);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(input, output));
EXPECT_EQ(output, input) << Disassemble(output);
EXPECT_TRUE(ShaderCompiler::ModuleDeclaresFloat64(output));
}
TEST_F(DemoteFloat64Test, ModuleDeclaresFloat64AnswersBothWays) {
const Vector<Uint32> wide = CompileToSpirv(GL_VERTEX_SHADER, kWideVertexSource);
ASSERT_FALSE(wide.empty());
EXPECT_TRUE(ShaderCompiler::ModuleDeclaresFloat64(wide));
Vector<Uint32> demoted;
ASSERT_TRUE(ShaderCompiler::DemoteFloat64ToFloat32(wide, demoted));
EXPECT_FALSE(ShaderCompiler::ModuleDeclaresFloat64(demoted));
EXPECT_FALSE(ShaderCompiler::ModuleDeclaresFloat64({}));
}
TEST_F(DemoteFloat64Test, TheSharedChainDemotesToo) {
// Production never calls the pass on its own: it reaches it through the one chain every
// module goes through at link, on both backends.
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kWideVertexSource);
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::SanitizeAndOptimizeBinary(input, output));
EXPECT_FALSE(ShaderCompiler::ModuleDeclaresFloat64(output)) << Disassemble(output);
}
// The payoff on the Espryt path: SPIRV-Cross throws "FP64 not supported in ES profile" for every
// one of these before demotion, so the program simply could not be transpiled at all.
class DemoteFloat64EsslTest : public DemoteFloat64Test, public ::testing::WithParamInterface<const char*> {};
INSTANTIATE_TEST_SUITE_P(
Shapes, DemoteFloat64EsslTest,
::testing::Values(
// A double that never reaches an interface: locals and literals only.
R"(#version 460 core
layout(location = 0) out float vOut;
void main() {
double s = 0.5lf;
for (int i = 0; i < 3; ++i) s = s * 1.5lf + 0.25lf;
vOut = float(s);
gl_Position = vec4(float(s));
}
)",
// A default-block double uniform: the glUniform*d path, and the block MobileGL lays out
// itself.
R"(#version 460 core
layout(location = 0) uniform double uScale;
layout(location = 1) uniform dvec3 uOffset;
layout(location = 2) uniform dmat4 uTransform;
layout(location = 0) out float vOut;
void main() {
dvec3 p = uOffset * uScale + dvec3(uTransform[1].xyz);
vOut = float(p.x + p.y + p.z);
gl_Position = vec4(float(p.x));
}
)",
// An application-declared std140 block whose members are 64-bit.
R"(#version 460 core
layout(std140, binding = 0) uniform Blk {
float a;
double d;
dvec2 v2;
dvec3 v3;
dvec4 v4;
dmat4 m4;
double arr[3];
float z;
};
layout(location = 0) out float vOut;
void main() {
double s = d + v2.x + v3.y + v4.z + m4[2].w + arr[1] + a + z;
vOut = float(s);
gl_Position = vec4(float(s));
}
)",
// A 64-bit vertex input, which is what glVertexAttribLFormat feeds.
R"(#version 460 core
layout(location = 0) in dvec3 inPos;
layout(location = 2) in double inWeight;
layout(location = 0) out float vOut;
void main() {
vOut = float(inPos.x + inPos.y + inPos.z + inWeight);
gl_Position = vec4(vOut);
}
)",
// An std430 storage block, whose double members pack differently again.
R"(#version 460 core
layout(std430, binding = 0) buffer Ssbo {
double head;
dvec4 wide;
double tail[4];
};
layout(location = 0) out float vOut;
void main() {
double s = head + wide.w + tail[3];
vOut = float(s);
gl_Position = vec4(float(s));
}
)"));
TEST_P(DemoteFloat64EsslTest, TheDemotedModuleCanBeEmittedAsEssl) {
using namespace MG_Util::ShaderTranspiler;
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, GetParam());
ASSERT_FALSE(input.empty());
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::SanitizeAndOptimizeBinary(input, output));
SpvcSession session(output, SessionUsageBit::Transpile);
spvc_compiler_options options;
ASSERT_EQ(session.CreateOptions(&options), SPVC_SUCCESS);
spvc_compiler_options_set_uint(options, SPVC_COMPILER_OPTION_GLSL_VERSION, 320);
spvc_compiler_options_set_bool(options, SPVC_COMPILER_OPTION_GLSL_ES, SPVC_TRUE);
spvc_compiler_options_set_bool(options, SPVC_COMPILER_OPTION_GLSL_VULKAN_SEMANTICS, SPVC_FALSE);
ASSERT_EQ(session.SetOptions(options), SPVC_SUCCESS);
auto essl = ShaderCompiler::DecompileShader(session);
ASSERT_TRUE(essl) << essl.error().log;
EXPECT_NE(essl->find("#version 320 es"), String::npos) << *essl;
// ESSL has no 64-bit float spelling at all, so any of these in the output is SPIRV-Cross
// having emitted something no ES driver will compile.
EXPECT_EQ(essl->find("double"), String::npos) << *essl;
EXPECT_EQ(essl->find("dvec"), String::npos) << *essl;
EXPECT_EQ(essl->find("dmat"), String::npos) << *essl;
}
TEST_F(DemoteFloat64Test, RejectsGarbageInput) {
const Vector<Uint32> notSpirv{0xdeadbeefu, 0u, 0u, 0u, 0u};
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::DemoteFloat64ToFloat32(notSpirv, output));
}
@@ -0,0 +1,226 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/FlattenXfbInterfaceBlocksTest.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <set>
#include <string>
#include <vector>
#include "Includes.h"
#include "Init.h"
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <MG_Util/ShaderTranspiler/SpvcSession.h>
#include <MG_Util/ShaderTranspiler/Types.h>
#include <spirv-tools/libspirv.hpp>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::SessionUsageBit;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
using MobileGL::MG_Util::ShaderTranspiler::SpvcSession;
namespace {
Vector<Uint32> CompileToSpirv(GLenum stage, const String& source) {
using namespace MG_Util::ShaderTranspiler;
ShaderAttrib shaderAttrib{.shaderType = stage, .sourceStr = source};
auto shaderResult = ShaderCompiler::CompileShader(shaderAttrib);
EXPECT_TRUE(shaderResult) << (shaderResult ? String{} : shaderResult.error().log);
if (!shaderResult) return {};
ProgramAttrib programAttrib{.shaders = {shaderResult.value()}};
auto programResult = ShaderCompiler::LinkProgram(programAttrib);
EXPECT_TRUE(programResult) << (programResult ? String{} : programResult.error().log);
if (!programResult) return {};
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {stage}, .program = *programResult.value()};
auto binaryResult = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
EXPECT_TRUE(binaryResult) << (binaryResult ? String{} : binaryResult.error().log);
if (!binaryResult || binaryResult->empty()) return {};
return binaryResult->front();
}
String Disassemble(const Vector<Uint32>& spirv) {
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String text;
tools.Disassemble(spirv, &text);
return text;
}
String Transpile(const Vector<Uint32>& spirv) {
SpvcSession session(spirv, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
EXPECT_TRUE(essl) << (essl ? String{} : essl.error().log);
return essl ? essl.value() : String{};
}
// The KHR-GL43.vertex_attrib_binding.basic-input capture program's output side, verbatim:
// a 16-element vec4 array inside a named output block, which is what the capture list
// addresses member by member ("StageData.attrib[0]" ... "StageData.attrib[15]").
const char* kCaptureVertexSource = R"(#version 430 core
layout(location = 0) in vec4 vs_in_attrib;
out StageData {
vec4 attrib[16];
} vs_out;
void main() {
for (int i = 0; i < 16; ++i) {
vs_out.attrib[i] = vs_in_attrib;
}
}
)";
// Mixed member widths, so a member that claims the wrong number of locations moves every
// member after it.
// 440, because a location on the BLOCK is ARB_enhanced_layouts.
const char* kMixedBlockVertexSource = R"(#version 440 core
layout(location = 0) in vec4 vs_in_attrib;
layout(location = 0) out StageData {
vec4 first;
vec2 second;
mat4 third;
vec4 fourth;
} vs_out;
void main() {
vs_out.first = vs_in_attrib;
vs_out.second = vs_in_attrib.xy;
vs_out.third = mat4(vs_in_attrib.x);
vs_out.fourth = vs_in_attrib;
}
)";
} // namespace
class FlattenXfbInterfaceBlocksTest : public ::testing::Test {
protected:
void SetUp() override {
MobileGL::Initialize();
ShaderCompiler::SetSpirvValidationEnabled(true);
m_validationFailuresAtStart = ShaderCompiler::SpirvValidationFailureCount();
}
void TearDown() override {
EXPECT_EQ(ShaderCompiler::SpirvValidationFailureCount(), m_validationFailuresAtStart)
<< "the flattened module did not survive spirv-val";
}
Uint64 m_validationFailuresAtStart = 0;
};
TEST_F(FlattenXfbInterfaceBlocksTest, FlattensACapturedBlockIntoOneVariablePerMember) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kCaptureVertexSource);
ASSERT_FALSE(input.empty());
std::set<String> flattened;
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(input, {"StageData"}, flattened, output));
ASSERT_FALSE(output.empty());
EXPECT_EQ(flattened, (std::set<String>{"StageData"}));
const String dis = Disassemble(output);
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
ASSERT_TRUE(tools.Validate(output)) << dis;
EXPECT_NE(dis.find("StageData_attrib"), String::npos) << dis;
// The block itself must have stopped being an interface variable, or the driver would see
// both spellings of the same data.
EXPECT_NE(dis.find("Private"), String::npos) << dis;
}
// The declaration is the point of the whole exercise: the emitted ESSL has to declare a plain
// output ARRAY, not an interface block, because that is the shape the Adreno driver can capture.
// SPIR-V validation does NOT catch the difference - leaving the struct's Block decoration on the
// demoted shadow produced a module that validated and emitted `StageData vs_out;` next to a block
// declaration the driver rejected with a bare "'vs_out' : syntax error".
TEST_F(FlattenXfbInterfaceBlocksTest, TheEmittedDeclarationIsAPlainArrayNotABlock) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kCaptureVertexSource);
ASSERT_FALSE(input.empty());
// Negative control: untouched, the block is emitted AS a block.
const String before = Transpile(input);
EXPECT_NE(before.find("out StageData"), String::npos) << before;
EXPECT_EQ(before.find("StageData_attrib"), String::npos) << before;
std::set<String> flattened;
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(input, {"StageData"}, flattened, output));
const String after = Transpile(output);
EXPECT_NE(after.find("StageData_attrib[16]"), String::npos) << after;
EXPECT_EQ(after.find("out StageData"), String::npos)
<< "the block must not still be declared as an output block:\n"
<< after;
}
// GL 4.6 core 11.1.2.1: consecutive members take consecutive locations, and a member takes as
// many as its type needs. Getting a span wrong silently moves every member after it.
TEST_F(FlattenXfbInterfaceBlocksTest, GivesEachMemberItsOwnConsecutiveLocations) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kMixedBlockVertexSource);
ASSERT_FALSE(input.empty());
std::set<String> flattened;
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(input, {"StageData"}, flattened, output));
ASSERT_FALSE(output.empty());
const String dis = Disassemble(output);
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
ASSERT_TRUE(tools.Validate(output)) << dis;
EXPECT_NE(dis.find("OpDecorate %StageData_first Location 0"), String::npos) << dis;
EXPECT_NE(dis.find("OpDecorate %StageData_second Location 1"), String::npos) << dis;
EXPECT_NE(dis.find("OpDecorate %StageData_third Location 2"), String::npos) << dis;
// mat4 takes four, so the member after it starts at 2 + 4.
EXPECT_NE(dis.find("OpDecorate %StageData_fourth Location 6"), String::npos) << dis;
}
// Nothing captures this block, so nothing may touch it: a shader that merely HAS an output
// block must reach the driver exactly as it was.
TEST_F(FlattenXfbInterfaceBlocksTest, LeavesABlockNoCaptureNamesAlone) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kCaptureVertexSource);
ASSERT_FALSE(input.empty());
std::set<String> flattened;
Vector<Uint32> output;
ASSERT_TRUE(
ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(input, {"SomeOtherBlock"}, flattened, output));
EXPECT_TRUE(flattened.empty());
const String after = Transpile(output);
EXPECT_NE(after.find("out StageData"), String::npos) << after;
EXPECT_EQ(after.find("StageData_attrib"), String::npos) << after;
}
// An empty request must not even run the optimizer: every program without transform feedback
// takes this path on every build.
TEST_F(FlattenXfbInterfaceBlocksTest, DeclinesAnEmptyRequestWithoutRewriting) {
const Vector<Uint32> input = CompileToSpirv(GL_VERTEX_SHADER, kCaptureVertexSource);
ASSERT_FALSE(input.empty());
std::set<String> flattened;
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(input, {}, flattened, output));
EXPECT_TRUE(flattened.empty());
EXPECT_TRUE(output.empty());
}
// The capture list has to follow the declaration exactly, and only for blocks that were
// actually rewritten - a member of a block left alone keeps the application's spelling, and so
// does a name with no block prefix at all (gl_Position, a plain varying).
TEST_F(FlattenXfbInterfaceBlocksTest, RewritesOnlyTheCaptureNamesOfFlattenedBlocks) {
String rewritten;
EXPECT_TRUE(ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock("StageData.attrib[0]", {"StageData"},
rewritten));
EXPECT_EQ(rewritten, "StageData_attrib[0]");
EXPECT_TRUE(
ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock("StageData.attrib", {"StageData"}, rewritten));
EXPECT_EQ(rewritten, "StageData_attrib");
EXPECT_FALSE(
ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock("Other.member", {"StageData"}, rewritten));
EXPECT_FALSE(ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock("gl_Position", {"StageData"}, rewritten));
EXPECT_FALSE(ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock("vColor", {"StageData"}, rewritten));
EXPECT_FALSE(ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock(".leading", {"StageData"}, rewritten));
}
@@ -128,6 +128,134 @@ namespace {
DrainErrors(); DrainErrors();
} }
// KHR-GL44.multi_bind.errors_bind_textures / .errors_bind_image_textures / .errors_bind_samplers.
// Both entry points were silent no-op stubs, so every row here answered GL_NO_ERROR.
// errors_bind_samplers is in the list because that case checks the invalid-name rule by calling
// glBindTextures with a sampler-name array - a name from the wrong namespace is simply not an
// existing texture.
TEST_F(NegativeApiErrorsTest, MultiBindTexturesRejectsBadRangesAndNames) {
GLint maxUnits = 0;
GetIntegerv(GL_MAX_COMBINED_TEXTURE_IMAGE_UNITS, &maxUnits);
ASSERT_GT(maxUnits, 0);
GLint maxImageUnits = 0;
GetIntegerv(GL_MAX_IMAGE_UNITS, &maxImageUnits);
GLuint texture = 0;
GenTextures(1, &texture);
BindTexture(GL_TEXTURE_2D, texture);
TexStorage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4);
// Reserved by glGenTextures but never bound: not an object yet, so the multi-bind entry
// points must refuse it instead of creating it the way glBindTexture would.
GLuint reservedOnly = 0;
GenTextures(1, &reservedOnly);
ASSERT_NE(reservedOnly, 0u);
ASSERT_EQ(IsTexture(reservedOnly), GL_FALSE);
DrainErrors();
const GLuint good[1] = {texture};
const GLuint mixed[2] = {texture, reservedOnly};
std::vector<Row> rows = {
{"glBindTextures with negative count", [&] { BindTextures(0, -1, good); }, GL_INVALID_VALUE},
{"glBindTextures with first + count past the last unit",
[&] { BindTextures(static_cast<GLuint>(maxUnits), 1, good); }, GL_INVALID_OPERATION},
{"glBindTextures with a reserved-but-uncreated name", [&] { BindTextures(0, 2, mixed); },
GL_INVALID_OPERATION},
{"glBindImageTextures with negative count", [&] { BindImageTextures(0, -1, good); }, GL_INVALID_VALUE},
};
if (maxImageUnits > 0) {
rows.push_back({"glBindImageTextures with first + count past the last image unit",
[&] { BindImageTextures(static_cast<GLuint>(maxImageUnits), 1, good); },
GL_INVALID_OPERATION});
rows.push_back({"glBindImageTextures with a reserved-but-uncreated name",
[&] { BindImageTextures(0, 2, mixed); }, GL_INVALID_OPERATION});
}
RunRows(rows);
// The loop semantics again: the good element at index 0 binds, the bad one does not.
GLint bound = -1;
GetIntegeri_v(GL_TEXTURE_BINDING_2D, 0, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), texture) << "a rejected element must not take the valid ones with it";
GetIntegeri_v(GL_TEXTURE_BINDING_2D, 1, &bound);
EXPECT_EQ(bound, 0) << "the rejected element must not have bound anything";
DrainErrors();
}
// KHR-GL44.multi_bind.functional_bind_textures / .functional_bind_image_textures: the binding
// has to land on the texture's OWN target - glBindTextures takes no target parameter - and
// element zero has to unbind every target of its unit.
TEST_F(NegativeApiErrorsTest, MultiBindTexturesBindsToTheTexturesOwnTarget) {
GLuint textures[2] = {0, 0};
GenTextures(2, textures);
BindTexture(GL_TEXTURE_1D, textures[0]);
TexStorage1D(GL_TEXTURE_1D, 1, GL_RGBA8, 4);
BindTexture(GL_TEXTURE_3D, textures[1]);
TexStorage3D(GL_TEXTURE_3D, 1, GL_RGBA8, 4, 4, 4);
// Leave the active unit's slots clean so only the multi-bind result is under test.
BindTexture(GL_TEXTURE_1D, 0);
BindTexture(GL_TEXTURE_3D, 0);
DrainErrors();
BindTextures(0, 2, textures);
EXPECT_EQ(GetError(), GL_NO_ERROR);
GLint bound = -1;
GetIntegeri_v(GL_TEXTURE_BINDING_1D, 0, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), textures[0]) << "a 1D texture must land on the unit's 1D slot";
GetIntegeri_v(GL_TEXTURE_BINDING_3D, 0, &bound);
EXPECT_EQ(bound, 0) << "no other target of the unit may be touched";
GetIntegeri_v(GL_TEXTURE_BINDING_3D, 1, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), textures[1]) << "a 3D texture must land on the unit's 3D slot";
// A zero element - and a NULL array - unbind EVERY target of the unit, not just one.
const GLuint zeros[1] = {0};
BindTextures(0, 1, zeros);
GetIntegeri_v(GL_TEXTURE_BINDING_1D, 0, &bound);
EXPECT_EQ(bound, 0);
BindTextures(1, 1, nullptr);
GetIntegeri_v(GL_TEXTURE_BINDING_3D, 1, &bound);
EXPECT_EQ(bound, 0) << "a NULL <textures> unbinds the range";
EXPECT_EQ(GetError(), GL_NO_ERROR);
GLint maxImageUnits = 0;
GetIntegerv(GL_MAX_IMAGE_UNITS, &maxImageUnits);
if (maxImageUnits > 0) {
// ARB_multi_bind fixes every glBindImageTexture parameter but the unit and the name:
// level 0, layered, layer 0, READ_WRITE, and the texture's own internal format.
BindImageTextures(0, 1, &textures[1]);
EXPECT_EQ(GetError(), GL_NO_ERROR);
GetIntegeri_v(GL_IMAGE_BINDING_NAME, 0, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), textures[1]);
GetIntegeri_v(GL_IMAGE_BINDING_LEVEL, 0, &bound);
EXPECT_EQ(bound, 0);
GetIntegeri_v(GL_IMAGE_BINDING_LAYERED, 0, &bound);
EXPECT_EQ(bound, GL_TRUE);
GetIntegeri_v(GL_IMAGE_BINDING_ACCESS, 0, &bound);
EXPECT_EQ(bound, GL_READ_WRITE);
GetIntegeri_v(GL_IMAGE_BINDING_FORMAT, 0, &bound);
EXPECT_EQ(bound, GL_RGBA8);
BindImageTextures(0, 1, nullptr);
GetIntegeri_v(GL_IMAGE_BINDING_NAME, 0, &bound);
EXPECT_EQ(bound, 0) << "a NULL <textures> resets the image unit";
}
// Through the EXPORTED entry points, not just the GLImpl functions: both of these were
// declared with the stub macro, so a working implementation that is never wired into
// Definitions.cpp still answers GL_NO_ERROR and binds nothing.
::glBindTextures(0, 1, &textures[0]);
GetIntegeri_v(GL_TEXTURE_BINDING_1D, 0, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), textures[0]) << "glBindTextures is still exported as a no-op stub";
if (maxImageUnits > 0) {
::glBindImageTextures(0, 1, &textures[1]);
GetIntegeri_v(GL_IMAGE_BINDING_NAME, 0, &bound);
EXPECT_EQ(static_cast<GLuint>(bound), textures[1])
<< "glBindImageTextures is still exported as a no-op stub";
}
DrainErrors();
}
TEST_F(NegativeApiErrorsTest, BufferRangeOffsetAlignmentAppliesToTheBindingPoint) { TEST_F(NegativeApiErrorsTest, BufferRangeOffsetAlignmentAppliesToTheBindingPoint) {
GLint ssboAlignment = 0; GLint ssboAlignment = 0;
GetIntegerv(GL_SHADER_STORAGE_BUFFER_OFFSET_ALIGNMENT, &ssboAlignment); GetIntegerv(GL_SHADER_STORAGE_BUFFER_OFFSET_ALIGNMENT, &ssboAlignment);
@@ -540,6 +540,10 @@ namespace MobileGL::MG_Util::BackendLoader {
INIT_GLES_FUNC_OPTIONAL(glMultiDrawArraysIndirectEXT) INIT_GLES_FUNC_OPTIONAL(glMultiDrawArraysIndirectEXT)
INIT_GLES_FUNC_OPTIONAL(glMultiDrawElementsIndirectEXT) INIT_GLES_FUNC_OPTIONAL(glMultiDrawElementsIndirectEXT)
INIT_GLES_FUNC_OPTIONAL(glMultiDrawElementsBaseVertexEXT) INIT_GLES_FUNC_OPTIONAL(glMultiDrawElementsBaseVertexEXT)
INIT_GLES_FUNC_OPTIONAL(glDrawArraysInstancedBaseInstanceEXT)
INIT_GLES_FUNC_OPTIONAL(glDrawElementsInstancedBaseInstanceEXT)
INIT_GLES_FUNC_OPTIONAL(glDrawElementsInstancedBaseVertexBaseInstanceEXT)
} }
} }
@@ -851,6 +855,9 @@ namespace MobileGL::MG_Util::BackendLoader {
// Resolved into caps.TextureBufferSupport below, once the ES version is also known. // Resolved into caps.TextureBufferSupport below, once the ES version is also known.
Bool hasExtTextureBuffer = false; Bool hasExtTextureBuffer = false;
Bool hasOesTextureBuffer = false; Bool hasOesTextureBuffer = false;
// Combined with the three entry points below; DirectGLES emulates baseInstance when this
// comes out false, so a stub pointer counting as support would silently break the draws.
Bool hasBaseInstanceExtension = false;
for (GLint i = 0; i < extCount; ++i) { for (GLint i = 0; i < extCount; ++i) {
const char* extension = (const char*)glesFuncs.glGetStringi(GL_EXTENSIONS, i); const char* extension = (const char*)glesFuncs.glGetStringi(GL_EXTENSIONS, i);
if (extension) { if (extension) {
@@ -891,7 +898,7 @@ namespace MobileGL::MG_Util::BackendLoader {
hasOesTextureBuffer = true; hasOesTextureBuffer = true;
} }
if (std::strcmp(extension, "GL_EXT_base_instance") == 0) { if (std::strcmp(extension, "GL_EXT_base_instance") == 0) {
caps.SupportsBaseInstance = true; hasBaseInstanceExtension = true;
} }
if (std::strcmp(extension, "GL_EXT_disjoint_timer_query") == 0) { if (std::strcmp(extension, "GL_EXT_disjoint_timer_query") == 0) {
caps.SupportsDisjointTimerQuery = true; caps.SupportsDisjointTimerQuery = true;
@@ -929,6 +936,12 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.SupportsMultiDrawElementsBaseVertex = hasDrawElementsBaseVertexExtension && caps.SupportsMultiDrawElementsBaseVertex = hasDrawElementsBaseVertexExtension &&
hasMultiDrawArraysExtension && hasMultiDrawArraysExtension &&
glesFuncs.glMultiDrawElementsBaseVertexEXT != nullptr; glesFuncs.glMultiDrawElementsBaseVertexEXT != nullptr;
// All three, not any: DirectGLES picks native-vs-emulated once per draw entry point off
// this single flag, so a driver that resolved only some of them must count as absent.
caps.SupportsBaseInstance = hasBaseInstanceExtension &&
glesFuncs.glDrawArraysInstancedBaseInstanceEXT != nullptr &&
glesFuncs.glDrawElementsInstancedBaseInstanceEXT != nullptr &&
glesFuncs.glDrawElementsInstancedBaseVertexBaseInstanceEXT != nullptr;
// Core from ES 3.2 on, so an extension string is not required there; below 3.2 the // Core from ES 3.2 on, so an extension string is not required there; below 3.2 the
// extension is, and the pointer still has to have resolved either way. // extension is, and the pointer still has to have resolved either way.
const Bool esAtLeast32 = caps.GLESVersion.Major > 3 || const Bool esAtLeast32 = caps.GLESVersion.Major > 3 ||
@@ -963,6 +976,8 @@ namespace MobileGL::MG_Util::BackendLoader {
MGLOG_I(" draw elements base vertex (ES 3.2 core or EXT/OES_draw_elements_base_vertex): %s", MGLOG_I(" draw elements base vertex (ES 3.2 core or EXT/OES_draw_elements_base_vertex): %s",
caps.SupportsDrawElementsBaseVertex ? "yes" : "no"); caps.SupportsDrawElementsBaseVertex ? "yes" : "no");
MGLOG_I(" compute shaders (ES 3.1 core): %s", caps.SupportsComputeShader ? "yes" : "no"); MGLOG_I(" compute shaders (ES 3.1 core): %s", caps.SupportsComputeShader ? "yes" : "no");
MGLOG_I(" base instance (EXT_base_instance; emulated by attribute offsets when absent): %s",
caps.SupportsBaseInstance ? "yes" : "no");
MGLOG_I("OpenGL ES capabilities:"); MGLOG_I("OpenGL ES capabilities:");
glesFuncs.glGetIntegerv(GL_UNIFORM_BUFFER_OFFSET_ALIGNMENT, &caps.UniformBufferOffsetAlignment); glesFuncs.glGetIntegerv(GL_UNIFORM_BUFFER_OFFSET_ALIGNMENT, &caps.UniformBufferOffsetAlignment);
@@ -638,6 +638,17 @@ namespace MobileGL {
GL_FUNC_TYPEDEF(void, glBruh) GL_FUNC_TYPEDEF(void, glBruh)
GL_FUNC_TYPEDEF(void, glMultiDrawElementsBaseVertexEXT, GLenum mode, const GLsizei* count, GLenum type, GL_FUNC_TYPEDEF(void, glMultiDrawElementsBaseVertexEXT, GLenum mode, const GLsizei* count, GLenum type,
const void* const* indices, GLsizei drawcount, const GLint* basevertex) const void* const* indices, GLsizei drawcount, const GLint* basevertex)
// GL_EXT_base_instance. Where a driver has these, the "+ baseInstance" of the
// instanced-array element index is the driver's job; where it does not, DirectGLES
// folds it into the attribute offsets instead (VertexArrayImpl::BaseInstanceByteShift).
GL_FUNC_TYPEDEF(void, glDrawArraysInstancedBaseInstanceEXT, GLenum mode, GLint first, GLsizei count,
GLsizei instancecount, GLuint baseinstance)
GL_FUNC_TYPEDEF(void, glDrawElementsInstancedBaseInstanceEXT, GLenum mode, GLsizei count, GLenum type,
const void* indices, GLsizei instancecount, GLuint baseinstance)
GL_FUNC_TYPEDEF(void, glDrawElementsInstancedBaseVertexBaseInstanceEXT, GLenum mode, GLsizei count,
GLenum type, const void* indices, GLsizei instancecount, GLint basevertex,
GLuint baseinstance)
/* /*
namespace Caps { namespace Caps {
struct GLESCaps { struct GLESCaps {
@@ -1034,6 +1045,10 @@ namespace MobileGL {
GL_FUNC_DECL(glMultiDrawElementsIndirectEXT) GL_FUNC_DECL(glMultiDrawElementsIndirectEXT)
GL_FUNC_DECL(glMultiDrawElementsBaseVertexEXT) GL_FUNC_DECL(glMultiDrawElementsBaseVertexEXT)
GL_FUNC_DECL(glDrawArraysInstancedBaseInstanceEXT)
GL_FUNC_DECL(glDrawElementsInstancedBaseInstanceEXT)
GL_FUNC_DECL(glDrawElementsInstancedBaseVertexBaseInstanceEXT)
GL_FUNC_DECL(glBruh) GL_FUNC_DECL(glBruh)
}; };
+56 -17
View File
@@ -46,6 +46,22 @@ namespace MobileGL::MG_Util::SelfTest {
RankMobileGLReported = 5, RankMobileGLReported = 5,
}; };
// Both backends' fp64 rows end the same way, and the sentence they end with depends on
// a config flag rather than on anything either backend probes: the demotion is what
// makes doubles work, but GL_ARB_gpu_shader_fp64 promises the PRECISION the demotion
// cannot deliver, so the string is opt-in and the row has to say which way it went.
String AppendFp64AdvertisementNote(String detail) {
if (MG_Config::Features.AdvertiseFp64) {
return Move(detail) +
". GL_ARB_gpu_shader_fp64 IS advertised (MOBILEGL_ADVERTISE_FP64): an application "
"that checks the string will believe it has 64-bit precision, and it does not";
}
return Move(detail) +
". GL_ARB_gpu_shader_fp64 is not advertised, because the precision it promises is the "
"one thing the demotion cannot provide; set MOBILEGL_ADVERTISE_FP64=1 to advertise it "
"anyway";
}
struct ReportBuilder { struct ReportBuilder {
BackendPostReport report; BackendPostReport report;
Bool fatalFailed = false; Bool fatalFailed = false;
@@ -336,8 +352,10 @@ namespace MobileGL::MG_Util::SelfTest {
builder.Pass("GL_EXT_base_instance", "supported (native baseInstance draws)"); builder.Pass("GL_EXT_base_instance", "supported (native baseInstance draws)");
} else { } else {
builder.Info("GL_EXT_base_instance", builder.Info("GL_EXT_base_instance",
"not supported; no impact: the native indirect path deliberately does not " "not supported; direct baseInstance draws are emulated by shifting the "
"rely on it (shader-side emulation handles baseInstance semantics)"); "instanced arrays' attribute offsets, and gl_BaseInstance by a uniform. "
"The one gap is an INDIRECT draw whose command carries a non-zero "
"baseInstance and is executed natively: its vertex fetch is not shifted");
} }
// Both multi-draw rows gate on the capability flags, not the entry-point pointers: // Both multi-draw rows gate on the capability flags, not the entry-point pointers:
// eglGetProcAddress may hand back a non-NULL stub for these on drivers without the // eglGetProcAddress may hand back a non-NULL stub for these on drivers without the
@@ -484,14 +502,24 @@ namespace MobileGL::MG_Util::SelfTest {
break; break;
} }
} }
// Reported rather than probed: this one cannot come out any other way. OpenGL ES has no // Both rows are reported rather than probed: neither can come out any other way.
// double-precision vertex format and ESSL has no fp64 type, so there is no driver and no // ESSL has no 64-bit float type at all, so no driver and no extension could change
// extension that could make it work - the row exists so the loss is named at startup // either answer, and the rows exist so the two halves of the loss are named at
// instead of discovered as an unexplained GL_INVALID_OPERATION at draw setup. // startup instead of discovered as a shader that will not compile or an
// unexplained GL_INVALID_OPERATION at draw setup.
builder.Pass("fp64", AppendFp64AdvertisementNote(
"demoted to fp32 - ESSL has no 64-bit float type, so every double / "
"dvec / dmat in a shader is narrowed to 32 bits before transpilation "
"(DemoteFloat64Pass). Such shaders COMPILE AND RUN, at single "
"precision; a block containing a double is re-laid-out for the "
"narrowed members, so an application that hard-codes std140 offsets "
"computed for doubles must query them instead"));
builder.Warn("64-bit vertex attributes", builder.Warn("64-bit vertex attributes",
"not supported on any GLES driver (ES has no GL_DOUBLE vertex format and ESSL has " "not supported (ES has no GL_DOUBLE vertex format, and after the fp64 demotion "
"no fp64 type); glVertexAttribLFormat / glVertexArrayAttribLFormat report " "above there is no 64-bit shader input left to feed either); "
"GL_INVALID_OPERATION - use the Vulkan backend if the application needs them"); "glVertexAttribLFormat / glVertexArrayAttribLFormat report "
"GL_INVALID_OPERATION - feed the attribute with glVertexAttribPointer(GL_FLOAT), "
"which a demoted dvec input reads correctly");
if (glesFuncs.glPatchParameteri != nullptr) { if (glesFuncs.glPatchParameteri != nullptr) {
builder.Pass("Tessellation patch parameters", builder.Pass("Tessellation patch parameters",
"glPatchParameteri present (GL_PATCH_VERTICES reaches the driver)"); "glPatchParameteri present (GL_PATCH_VERTICES reaches the driver)");
@@ -1843,14 +1871,25 @@ namespace MobileGL::MG_Util::SelfTest {
"unsupported; a GL_TEXTURE_CUBE_MAP_ARRAY texture gets no image at all, so sampling " "unsupported; a GL_TEXTURE_CUBE_MAP_ARRAY texture gets no image at all, so sampling "
"one reads nothing and glFramebufferTextureLayer on one is declined"); "one reads nothing and glFramebufferTextureLayer on one is declined");
} }
if (features.shaderFloat64 == VK_TRUE) { // Reported whichever way the device answers, because MobileGL no longer follows the
builder.Pass("shaderFloat64", // device here: every 64-bit float is narrowed to 32 bits before any module reaches this
"GLSL double/dvec/dmat and 64-bit vertex attributes (glVertexAttribLFormat) supported"); // backend (DemoteFloat64Pass), so the Float64 capability is never declared and a device
} else { // that HAS the feature gains nothing from it. The device's own answer is still worth
builder.Warn("shaderFloat64", // printing - it is the reason the demotion is unconditional.
"unsupported; any shader declaring a double fails to create a shader module, and " builder.Pass("fp64", AppendFp64AdvertisementNote(
"glVertexAttribLFormat reports GL_INVALID_OPERATION instead of feeding the attribute"); format("demoted to fp32 (device shaderFloat64 = {}) - every double / dvec / "
} "dmat in a shader is narrowed to 32 bits before pipeline creation, so "
"such shaders BUILD AND RUN at single precision on every device "
"instead of failing to create a shader module on the ones without the "
"feature. A block containing a double is re-laid-out for the narrowed "
"members, so an application that hard-codes std140 offsets computed "
"for doubles must query them instead",
features.shaderFloat64 == VK_TRUE ? "supported" : "unsupported")));
builder.Warn("64-bit vertex attributes",
"not supported; there is no 64-bit shader input left to feed after the fp64 demotion "
"above, and no VK_FORMAT_R64*_SFLOAT vertex fetch to feed it with on most devices "
"anyway. glVertexAttribLFormat reports GL_INVALID_OPERATION - feed the attribute with "
"glVertexAttribPointer(GL_FLOAT), which a demoted dvec input reads correctly");
Bool shaderDrawParameters = false; Bool shaderDrawParameters = false;
if (vkGetPhysicalDeviceFeatures2Fn != nullptr && properties.apiVersion >= VK_API_VERSION_1_1) { if (vkGetPhysicalDeviceFeatures2Fn != nullptr && properties.apiVersion >= VK_API_VERSION_1_1) {
@@ -18,9 +18,13 @@
#include "SpirvPasses/RenameBuiltinShadowingFunctionsPass.h" #include "SpirvPasses/RenameBuiltinShadowingFunctionsPass.h"
#include "SpirvPasses/DecomposeWorkgroupVec3Pass.h" #include "SpirvPasses/DecomposeWorkgroupVec3Pass.h"
#include "SpirvPasses/DecoratePositionInvariantPass.h" #include "SpirvPasses/DecoratePositionInvariantPass.h"
#include "SpirvPasses/DemoteFloat64Pass.h"
#include "SpirvPasses/LowerDrawParametersPass.h" #include "SpirvPasses/LowerDrawParametersPass.h"
#include "SpirvPasses/PackDoubleVertexInputsPass.h" #include "SpirvPasses/PackDoubleVertexInputsPass.h"
#include "SpirvPasses/FlattenXfbInterfaceBlocksPass.h"
#include "SpirvPasses/SplitArrayVertexInputsPass.h"
#include "SpirvPasses/RebaseInstanceIndexPass.h" #include "SpirvPasses/RebaseInstanceIndexPass.h"
#include "SpirvPasses/ZeroBaseVertexPass.h"
#include "SpirvPasses/NormalizeRectCoordinatesPass.h" #include "SpirvPasses/NormalizeRectCoordinatesPass.h"
#include "SpirvPasses/PrivateToEntryLocalPass.h" #include "SpirvPasses/PrivateToEntryLocalPass.h"
#include "SpirvPasses/StripUniformLocationsPass.h" #include "SpirvPasses/StripUniformLocationsPass.h"
@@ -576,6 +580,37 @@ namespace MobileGL {
return false; return false;
} }
Bool ShaderCompiler::ModuleDeclaresFloat64(const Vector<Uint32>& spirv) {
if (spirv.empty()) {
// Same reasoning as ModuleDeclaresBufferTextureSampler: a stage that produced
// no SPIR-V is not a verdict about 64-bit floats, and parsing it would push a
// spurious diagnostic through the message consumer.
return false;
}
std::unique_ptr<spvtools::opt::IRContext> context = spvtools::BuildModule(
SPV_ENV_VULKAN_1_1, MakeSpirvMessageConsumer("ModuleDeclaresFloat64"), spirv.data(),
spirv.size());
if (!context) {
return false;
}
for (const spvtools::opt::Instruction& type : context->types_values()) {
if (type.opcode() == spv::Op::OpTypeFloat && type.NumInOperands() >= 1 &&
type.GetSingleWordInOperand(0) == 64) {
return true;
}
}
return false;
}
bool ShaderCompiler::DemoteFloat64ToFloat32(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(DemoteFloat64Pass::CreateDemoteFloat64Pass());
return RunOptimizerChecked("DemoteFloat64ToFloat32", optimizer, inputBinary, outputBinary);
}
bool ShaderCompiler::SanitizeAndOptimizeBinary(const Vector<Uint32>& inputBinary, bool ShaderCompiler::SanitizeAndOptimizeBinary(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) { Vector<uint32_t>& outputBinary) {
using namespace spvtools; using namespace spvtools;
@@ -614,6 +649,17 @@ namespace MobileGL {
RenameBuiltinShadowingFunctionsPass::CreateRenameBuiltinShadowingFunctionsPass()); RenameBuiltinShadowingFunctionsPass::CreateRenameBuiltinShadowingFunctionsPass());
optimizer.RegisterPass(EliminateFloatEqualsZeroPass::CreateEliminateFloatEqualsZeroPass()); optimizer.RegisterPass(EliminateFloatEqualsZeroPass::CreateEliminateFloatEqualsZeroPass());
optimizer.RegisterPass(DecomposeWorkgroupVec3Pass::CreateDecomposeWorkgroupVec3Pass()); optimizer.RegisterPass(DecomposeWorkgroupVec3Pass::CreateDecomposeWorkgroupVec3Pass());
// No mobile GPU has 64-bit floats: Adreno and Mali both report shaderFloat64 ==
// VK_FALSE, and ESSL has no fp64 type for SPIRV-Cross to emit. Demoting here - in
// the one chain every module goes through, on both backends, at link - is what
// makes `double` compile at all, and makes it behave the SAME everywhere, which
// matters because the GL frontend's uniform storage cannot be per-backend: the
// glUniform*d shadow narrows to float unconditionally to match this. Runs last so
// no earlier pass ever has to reason about a width it will not see in the output;
// in particular it runs before the backends' PackDoubleVertexInputsPass, whose
// OpBitcast this one would otherwise decline on. Costs one types_values() walk on
// the overwhelming majority of modules, which declare no 64-bit float at all.
optimizer.RegisterPass(DemoteFloat64Pass::CreateDemoteFloat64Pass());
return RunOptimizerChecked("SanitizeAndOptimizeBinary", optimizer, inputBinary, return RunOptimizerChecked("SanitizeAndOptimizeBinary", optimizer, inputBinary,
outputBinary); outputBinary);
@@ -629,6 +675,36 @@ namespace MobileGL {
outputBinary); outputBinary);
} }
bool ShaderCompiler::SplitArrayVertexInputsForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(SplitArrayVertexInputsPass::CreateSplitArrayVertexInputsPass());
return RunOptimizerChecked("SplitArrayVertexInputsForEssl", optimizer, inputBinary,
outputBinary);
}
bool ShaderCompiler::FlattenXfbInterfaceBlocksForEssl(const Vector<Uint32>& inputBinary,
const std::set<String>& blockNames,
std::set<String>& flattenedBlockNames,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
if (blockNames.empty()) return false;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(FlattenXfbInterfaceBlocksPass::CreateFlattenXfbInterfaceBlocksPass(
blockNames, &flattenedBlockNames));
return RunOptimizerChecked("FlattenXfbInterfaceBlocksForEssl", optimizer, inputBinary,
outputBinary);
}
bool ShaderCompiler::RewriteXfbCaptureNameForFlattenedBlock(
const String& captureName, const std::set<String>& flattenedBlockNames, String& outName) {
return FlattenXfbInterfaceBlocksPass::RewriteCaptureName(captureName, flattenedBlockNames,
outName);
}
bool ShaderCompiler::PackDoubleVertexInputsForVulkan(const Vector<Uint32>& inputBinary, bool ShaderCompiler::PackDoubleVertexInputsForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) { Vector<uint32_t>& outputBinary) {
using namespace spvtools; using namespace spvtools;
@@ -758,6 +834,15 @@ namespace MobileGL {
outputBinary); outputBinary);
} }
bool ShaderCompiler::ZeroBaseVertexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(ZeroBaseVertexPass::CreateZeroBaseVertexPass());
return RunOptimizerChecked("ZeroBaseVertexForVulkan", optimizer, inputBinary, outputBinary);
}
bool ShaderCompiler::DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary, bool ShaderCompiler::DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) { Vector<uint32_t>& outputBinary) {
using namespace spvtools; using namespace spvtools;
@@ -12,6 +12,8 @@
#include "glslang/TVarEntryInfo.h" #include "glslang/TVarEntryInfo.h"
#include "glslang/TMglGlslIoResolver.h" #include "glslang/TMglGlslIoResolver.h"
#include <set>
namespace MobileGL { namespace MobileGL {
namespace MG_Util { namespace MG_Util {
namespace ShaderTranspiler { namespace ShaderTranspiler {
@@ -27,6 +29,30 @@ namespace MobileGL {
// Only for backends without native draw-parameter support (DirectGLES). // Only for backends without native draw-parameter support (DirectGLES).
static bool LowerDrawParametersForEssl(const Vector<Uint32>& inputBinary, static bool LowerDrawParametersForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary); Vector<uint32_t>& outputBinary);
// Replaces an ARRAY vertex input with one input per element at consecutive
// locations, seeding a Private copy of the array so indexed reads still work.
// GLSL ES has no array vertex inputs and SPIRV-Cross refuses the whole module
// rather than emulating them, so without this the stage never reaches the
// driver. Only for the DirectGLES transpile path.
static bool SplitArrayVertexInputsForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Replaces the named interface BLOCKS with one variable per member, named
// "<Block>_<member>", shadowing the block itself so the body is untouched. The
// Adreno ES driver silently captures NOTHING for a transform-feedback varying
// named as a block member, so a capture list that names one has to be respelled
// - and the declaration with it. `flattenedBlockNames` reports which blocks
// this stage actually rewrote, which is what the capture list must follow.
// Only for the DirectGLES transpile path.
static bool FlattenXfbInterfaceBlocksForEssl(const Vector<Uint32>& inputBinary,
const std::set<String>& blockNames,
std::set<String>& flattenedBlockNames,
Vector<uint32_t>& outputBinary);
// The capture request "StageData.attrib[0]" as the pass above renamed it,
// "StageData_attrib[0]", or false when it does not name a member of a block
// that was flattened.
static bool RewriteXfbCaptureNameForFlattenedBlock(const String& captureName,
const std::set<String>& flattenedBlockNames,
String& outName);
// Drops RelaxedPrecision member decorations from uniform-block structs so // Drops RelaxedPrecision member decorations from uniform-block structs so
// SPIRV-Cross prints the same (highp) member precision in every stage; ES // SPIRV-Cross prints the same (highp) member precision in every stage; ES
// drivers reject cross-stage uniform blocks whose member precisions differ. // drivers reject cross-stage uniform blocks whose member precisions differ.
@@ -65,6 +91,12 @@ namespace MobileGL {
static bool LowerRectImages(const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary); static bool LowerRectImages(const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary);
static bool RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary, static bool RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary); Vector<uint32_t>& outputBinary);
// Builds the non-indexed-draw variant of a vertex shader: every gl_BaseVertex
// read becomes zero, which is what GL defines for a command carrying no
// baseVertex parameter while Vulkan's builtin would report firstVertex.
// See ZeroBaseVertexPass.
static bool ZeroBaseVertexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Re-declares 64-bit float vertex inputs as their 32-bit unsigned word pair // Re-declares 64-bit float vertex inputs as their 32-bit unsigned word pair
// (double -> uvec2, dvec2 -> uvec4) and bitcasts them back to double at entry, so no // (double -> uvec2, dvec2 -> uvec4) and bitcasts them back to double at entry, so no
// VK_FORMAT_R64*_SFLOAT is needed - lavapipe advertises none of them for vertex // VK_FORMAT_R64*_SFLOAT is needed - lavapipe advertises none of them for vertex
@@ -87,6 +119,12 @@ namespace MobileGL {
// reinterpretation paths (for example, R32F storage accessed as r32ui). // reinterpretation paths (for example, R32F storage accessed as r32ui).
static bool UseUnformattedFloatStorageImagesForVulkan( static bool UseUnformattedFloatStorageImagesForVulkan(
const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary); const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary);
// Rewrites every 64-bit float in the module to a 32-bit one, preserving every
// block offset and stride exactly (see DemoteFloat64Pass). Already part of
// SanitizeAndOptimizeBinary, which is where production reaches it; exposed
// separately so a test can drive the demotion on its own.
static bool DemoteFloat64ToFloat32(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
static Result<String> DecompileShader(SpvcSession& session); static Result<String> DecompileShader(SpvcSession& session);
// Parses one trivial shader in each configuration the production path can // Parses one trivial shader in each configuration the production path can
@@ -146,6 +184,12 @@ namespace MobileGL {
// check exists so that failure can be reported as the missing capability it is, // check exists so that failure can be reported as the missing capability it is,
// naming the shader, rather than as a driver info log nobody sees. // naming the shader, rather than as a driver info log nobody sees.
static Bool ModuleDeclaresBufferTextureSampler(const Vector<Uint32>& spirv); static Bool ModuleDeclaresBufferTextureSampler(const Vector<Uint32>& spirv);
// True when the module still declares a 64-bit float type. After
// SanitizeAndOptimizeBinary that can only mean DemoteFloat64Pass declined the
// module (see its header for the two operations that make it decline), which is
// what the backends report: no mobile driver can build such a module.
static Bool ModuleDeclaresFloat64(const Vector<Uint32>& spirv);
}; };
} // namespace ShaderTranspiler } // namespace ShaderTranspiler
} // namespace MG_Util } // namespace MG_Util
@@ -0,0 +1,547 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DemoteFloat64Pass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "DemoteFloat64Pass.h"
#include "spirv.hpp"
#include "source/latest_version_glsl_std_450_header.h"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/util/make_unique.h"
#include <algorithm>
#include <cstring>
#include <map>
#include <string>
#include <unordered_map>
#include <unordered_set>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::IRContext;
using spvtools::opt::Instruction;
using spvtools::opt::Operand;
constexpr Uint32 kFloat64Width = 64;
constexpr Uint32 kFloat32Width = 32;
// OpTypeFloat <id> <Width>. SPIR-V 1.6 added an optional FP Encoding operand after
// it; the width stays operand 0 either way, and an encoded (non-IEEE) float is not
// something glslang can emit for `double`, so it is left alone by the width guard.
Bool IsFloatTypeOfWidth(const Instruction& type, Uint32 width) {
return type.opcode() == spv::Op::OpTypeFloat && type.NumInOperands() >= 1 &&
type.GetSingleWordInOperand(0) == width;
}
// Exactly the types spirv-val forbids a second declaration of. Aggregates - arrays
// and structs - are excluded on purpose: the spec permits duplicates of those, and
// merging them would take one of the two OpNames and one of the two ArrayStride /
// Offset decoration sets with it.
Bool IsDuplicableType(spv::Op opcode) {
switch (opcode) {
case spv::Op::OpTypeFloat:
case spv::Op::OpTypeVector:
case spv::Op::OpTypeMatrix:
case spv::Op::OpTypePointer:
case spv::Op::OpTypeFunction: return true;
default: return false;
}
}
Uint32 RoundUp(Uint32 value, Uint32 alignment) {
if (alignment == 0) return value;
return (value + alignment - 1) / alignment * alignment;
}
// Re-derives the std140 / std430 layout of a block whose members have just become
// 32 bits wide, and writes it back as Offset / ArrayStride / MatrixStride.
//
// Why the layout is recomputed rather than preserved. Preserving it - leaving each
// member at the byte offset glslang picked for the 64-bit type and letting the
// freed 4 bytes become padding - keeps the application's byte layout intact and is
// the obvious first choice, but it does not survive contact with the Espryt path:
// SPIRV-Cross has to print an ESSL block, GLSL ES has no member `layout(offset=)`
// (no ARB_enhanced_layouts), so it refuses any block whose declared offsets are not
// exactly what std140 or std430 computes - "Buffer block cannot be expressed as any
// of std430, std140, scalar". Every shader with a double in a block would fail to
// transpile at all, which is the case this whole demotion exists to fix. Padding
// members back in cannot rescue it either: a dmat4 member carries MatrixStride 32
// and std140 demands 16 for the demoted mat4, and no amount of padding BETWEEN
// members changes a stride INSIDE one.
//
// What recomputing costs: a block laid out for 64-bit members changes its
// driver-visible byte layout, so an application that hard-codes std140 offsets
// computed for doubles addresses the wrong bytes. Applications that query their
// offsets are unaffected, and MobileGL's own default-uniform block is unaffected by
// construction - the frontend builds its uniform routing by reflecting THIS module
// (ProgramSpirvTask::BuildGlobalUboRouting), so glUniform*d writes wherever the
// demoted shader reads.
class BlockRelayout {
public:
BlockRelayout(IRContext* irContext, Bool std140)
: m_irContext(irContext), m_std140(std140) {}
// Size and alignment of `typeId`, applying every stride decoration it implies
// on the way down. Zero size means "not a type this layout knows how to
// describe"; the caller then leaves the block alone rather than guessing.
struct Extent {
Uint32 size = 0;
Uint32 alignment = 0;
};
Extent Measure(Uint32 typeId) {
const auto memo = m_extents.find(typeId);
if (memo != m_extents.end()) return memo->second;
const Extent extent = MeasureUncached(typeId);
m_extents.emplace(typeId, extent);
return extent;
}
private:
Extent MeasureUncached(Uint32 typeId) {
const Instruction* type = m_irContext->get_def_use_mgr()->GetDef(typeId);
if (type == nullptr) return {};
switch (type->opcode()) {
case spv::Op::OpTypeInt:
case spv::Op::OpTypeFloat: {
const Uint32 bytes = type->GetSingleWordInOperand(0) / 8;
return {bytes, bytes};
}
case spv::Op::OpTypeBool: return {4, 4};
case spv::Op::OpTypeVector: {
const Extent component = Measure(type->GetSingleWordInOperand(0));
if (component.size == 0) return {};
const Uint32 count = type->GetSingleWordInOperand(1);
// A three-component vector aligns like a four-component one.
return {component.size * count,
component.alignment * (count == 3 ? 4 : count)};
}
case spv::Op::OpTypeMatrix: {
const Extent column = Measure(type->GetSingleWordInOperand(0));
if (column.size == 0) return {};
const Uint32 stride = MatrixOrArrayStride(column.alignment);
return {stride * type->GetSingleWordInOperand(1), stride};
}
case spv::Op::OpTypeArray:
case spv::Op::OpTypeRuntimeArray: {
const Extent element = Measure(type->GetSingleWordInOperand(0));
if (element.size == 0) return {};
const Uint32 alignment = MatrixOrArrayStride(element.alignment);
const Uint32 stride = RoundUp(element.size, alignment);
SetTypeDecoration(typeId, spv::Decoration::ArrayStride, stride);
if (type->opcode() == spv::Op::OpTypeRuntimeArray) {
// An unsized array contributes its stride and nothing more; the
// block's size is whatever the application bound.
return {stride, alignment};
}
return {stride * ArrayLength(type->GetSingleWordInOperand(1)), alignment};
}
case spv::Op::OpTypeStruct: return MeasureStruct(*type);
default: return {};
}
}
Extent MeasureStruct(const Instruction& structType) {
Uint32 cursor = 0;
Uint32 alignment = m_std140 ? 16u : 1u;
for (Uint32 member = 0; member < structType.NumInOperands(); ++member) {
const Uint32 memberTypeId = structType.GetSingleWordInOperand(member);
const Extent extent = Measure(memberTypeId);
if (extent.size == 0) return {};
const Uint32 offset = RoundUp(cursor, extent.alignment);
SetMemberDecoration(structType.result_id(), member, spv::Decoration::Offset, offset);
// A matrix member carries the stride between its columns on the MEMBER,
// not on the type, so it has to be (re)stated here - including for a
// matrix reached through an array.
const Instruction* memberType =
m_irContext->get_def_use_mgr()->GetDef(PeelArrays(memberTypeId));
if (memberType != nullptr && memberType->opcode() == spv::Op::OpTypeMatrix) {
const Extent column = Measure(memberType->GetSingleWordInOperand(0));
SetMemberDecoration(structType.result_id(), member,
spv::Decoration::MatrixStride,
MatrixOrArrayStride(column.alignment));
}
cursor = offset + extent.size;
alignment = std::max(alignment, extent.alignment);
}
return {RoundUp(cursor, alignment), alignment};
}
// std140 rounds every array and matrix stride up to a four-component vector.
Uint32 MatrixOrArrayStride(Uint32 elementAlignment) const {
return m_std140 ? RoundUp(elementAlignment, 16) : elementAlignment;
}
Uint32 PeelArrays(Uint32 typeId) const {
const Instruction* type = m_irContext->get_def_use_mgr()->GetDef(typeId);
while (type != nullptr && (type->opcode() == spv::Op::OpTypeArray ||
type->opcode() == spv::Op::OpTypeRuntimeArray)) {
type = m_irContext->get_def_use_mgr()->GetDef(type->GetSingleWordInOperand(0));
}
return type != nullptr ? type->result_id() : 0;
}
Uint32 ArrayLength(Uint32 lengthConstantId) const {
const Instruction* length = m_irContext->get_def_use_mgr()->GetDef(lengthConstantId);
if (length == nullptr || length->opcode() != spv::Op::OpConstant ||
length->NumInOperands() < 1) {
return 1;
}
return length->GetSingleWordInOperand(0);
}
void SetTypeDecoration(Uint32 targetId, spv::Decoration decoration, Uint32 value) {
for (Instruction& annotation : m_irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != targetId) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) != decoration) {
continue;
}
annotation.SetInOperand(2, {value});
return;
}
}
void SetMemberDecoration(Uint32 structId, Uint32 member, spv::Decoration decoration,
Uint32 value) {
for (Instruction& annotation : m_irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpMemberDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != structId) continue;
if (annotation.GetSingleWordInOperand(1) != member) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(2)) != decoration) {
continue;
}
annotation.SetInOperand(3, {value});
return;
}
}
IRContext* m_irContext = nullptr;
Bool m_std140 = true;
std::unordered_map<Uint32, Extent> m_extents;
};
} // namespace
spvtools::opt::Pass::Status DemoteFloat64Pass::Process() {
auto* irContext = context();
auto* defUseMgr = irContext->get_def_use_mgr();
// --- 1. the leaf types -------------------------------------------------------
std::vector<Instruction*> float64Types;
for (Instruction& type : irContext->types_values()) {
if (IsFloatTypeOfWidth(type, kFloat64Width)) {
float64Types.push_back(&type);
}
}
if (float64Types.empty()) return Status::SuccessWithoutChange;
std::unordered_set<Uint32> float64TypeIds;
for (const Instruction* type : float64Types) {
float64TypeIds.insert(type->result_id());
}
// Every type a *value* can have that is 64-bit float underneath: the scalar itself
// plus the vectors and matrices built from it. The types-and-values section is in
// declaration order and SPIR-V forbids a type from forward-referencing another, so
// one forward walk is already the complete transitive closure. Aggregates and
// pointers are deliberately not included: no operand of the instructions checked
// below is ever an aggregate or a pointer.
std::unordered_set<Uint32> float64ValueTypeIds = float64TypeIds;
for (Instruction& type : irContext->types_values()) {
if (type.opcode() != spv::Op::OpTypeVector && type.opcode() != spv::Op::OpTypeMatrix) {
continue;
}
if (float64ValueTypeIds.count(type.GetSingleWordInOperand(0)) != 0) {
float64ValueTypeIds.insert(type.result_id());
}
}
// Every type that has a 64-bit float anywhere underneath it, which is exactly the
// set of blocks whose layout has to be re-derived once the leaves narrow. Computed
// now, before the rewrite makes a demoted float indistinguishable from one that was
// always 32 bits; the same single forward walk is a complete closure.
std::unordered_set<Uint32> wideTypeIds = float64TypeIds;
for (Instruction& type : irContext->types_values()) {
const auto contains = [&](Uint32 operand) {
return wideTypeIds.count(type.GetSingleWordInOperand(operand)) != 0;
};
switch (type.opcode()) {
case spv::Op::OpTypeVector:
case spv::Op::OpTypeMatrix:
case spv::Op::OpTypeArray:
case spv::Op::OpTypeRuntimeArray:
if (contains(0)) wideTypeIds.insert(type.result_id());
break;
case spv::Op::OpTypePointer:
if (contains(1)) wideTypeIds.insert(type.result_id());
break;
case spv::Op::OpTypeStruct:
for (Uint32 member = 0; member < type.NumInOperands(); ++member) {
if (contains(member)) {
wideTypeIds.insert(type.result_id());
break;
}
}
break;
default: break;
}
}
const auto valueIsFloat64 = [&](Uint32 id) {
const Instruction* def = defUseMgr->GetDef(id);
return def != nullptr && float64ValueTypeIds.count(def->type_id()) != 0;
};
// --- 2. decline before touching anything ------------------------------------
// These are the operations that mean "the 64 bits themselves", not "a wide float".
// Narrowing one side of them produces a module spirv-val rejects, and rebuilding
// the value would mean emulating fp64 in software. Bail out with the module
// byte-identical instead; the caller still has its "module declares Float64"
// diagnostic for it.
Uint32 glslStd450SetId = 0;
for (const Instruction& import : irContext->ext_inst_imports()) {
if (import.GetInOperand(0).AsString() == "GLSL.std.450") {
glslStd450SetId = import.result_id();
break;
}
}
const char* declineReason = nullptr;
irContext->module()->ForEachInst(
[&](Instruction* inst) {
if (declineReason != nullptr) return;
switch (inst->opcode()) {
case spv::Op::OpBitcast: {
// "Total bit width of Result Type and Operand must match" - true
// today, false the moment one of the two sides halves.
const Bool resultIs64 = float64ValueTypeIds.count(inst->type_id()) != 0;
const Bool operandIs64 = valueIsFloat64(inst->GetSingleWordInOperand(0));
if (resultIs64 != operandIs64) {
declineReason = "an OpBitcast across the 64-bit boundary "
"(doubleBitsToUint64 / uint64BitsToDouble / "
"packDouble2x32)";
}
break;
}
case spv::Op::OpExtInst: {
if (glslStd450SetId == 0 ||
inst->GetSingleWordInOperand(0) != glslStd450SetId) {
break;
}
const Uint32 extOpcode = inst->GetSingleWordInOperand(1);
if (extOpcode == GLSLstd450PackDouble2x32 ||
extOpcode == GLSLstd450UnpackDouble2x32) {
declineReason = "GLSL.std.450 PackDouble2x32 / UnpackDouble2x32, "
"which are defined only for a 64-bit float";
}
break;
}
default: break;
}
},
/*run_on_debug_line_insts=*/false);
if (declineReason != nullptr) {
MGLOG_D("DemoteFloat64Pass: declined - the module uses %s; its 64-bit floats are "
"left in place",
declineReason);
return Status::SuccessWithoutChange;
}
// --- 3. re-encode the literals ----------------------------------------------
// A 64-bit float constant carries two literal words and a 32-bit one carries a
// single word, so the value has to be narrowed before the type changes underneath
// it - afterwards there is no way left to tell how wide the literal was meant to
// be. Composite constants hold <id>s, not literals, and need nothing.
for (Instruction& value : irContext->types_values()) {
if (value.opcode() != spv::Op::OpConstant && value.opcode() != spv::Op::OpSpecConstant) {
continue;
}
if (float64TypeIds.count(value.type_id()) == 0) continue;
if (value.NumInOperands() < 1) continue;
const Operand& literal = value.GetInOperand(0);
if (literal.words.size() != 2) continue;
const Uint64 bits =
(static_cast<Uint64>(literal.words[1]) << 32) | static_cast<Uint64>(literal.words[0]);
double wide = 0.0;
std::memcpy(&wide, &bits, sizeof(wide));
// Deliberately the ordinary narrowing conversion: a magnitude no float can
// hold becomes an infinity, which is the same answer the demoted arithmetic
// around it would produce.
const float narrow = static_cast<float>(wide);
Uint32 narrowedBits = 0;
std::memcpy(&narrowedBits, &narrow, sizeof(narrowedBits));
Operand narrowedLiteral = literal;
narrowedLiteral.words = {narrowedBits};
value.SetInOperands({std::move(narrowedLiteral)});
}
// --- 4. the demotion itself --------------------------------------------------
// In place, so every composite type, every pointer, every struct member offset and
// every debug name that referred to the 64-bit type keeps referring to the same
// <id>. This is the whole reason the pass does not build parallel types.
for (Instruction* type : float64Types) {
type->SetInOperand(0, {kFloat32Width});
}
// The cached analysis::Type objects were built against the old widths.
irContext->InvalidateAnalyses(IRContext::kAnalysisTypes);
// --- 5. the conversions that just became identities --------------------------
// `float(someDouble)` and `double(someFloat)` are both OpFConvert, and SPIR-V
// requires the two component widths to differ. Both sides are 32 bits now, so each
// one is replaced by its operand. Walking in module order means a chain of them
// resolves in a single sweep: by the time the second is reached its operand has
// already been rewritten to the ultimate source.
const auto componentWidth = [&](Uint32 typeId) -> Uint32 {
const Instruction* def = defUseMgr->GetDef(typeId);
if (def == nullptr) return 0;
if (def->opcode() == spv::Op::OpTypeVector) {
def = defUseMgr->GetDef(def->GetSingleWordInOperand(0));
}
if (def == nullptr || def->opcode() != spv::Op::OpTypeFloat) return 0;
return def->GetSingleWordInOperand(0);
};
std::vector<Instruction*> identityConversions;
irContext->module()->ForEachInst(
[&](Instruction* inst) {
if (inst->opcode() != spv::Op::OpFConvert) return;
const Instruction* operandDef = defUseMgr->GetDef(inst->GetSingleWordInOperand(0));
if (operandDef == nullptr) return;
const Uint32 resultWidth = componentWidth(inst->type_id());
if (resultWidth == 0 || resultWidth != componentWidth(operandDef->type_id())) {
return;
}
identityConversions.push_back(inst);
},
/*run_on_debug_line_insts=*/false);
for (Instruction* conversion : identityConversions) {
irContext->ReplaceAllUsesWith(conversion->result_id(),
conversion->GetSingleWordInOperand(0));
irContext->KillInst(conversion);
}
// --- 6. the capability -------------------------------------------------------
std::vector<Instruction*> deadCapabilities;
for (Instruction& capability : irContext->capabilities()) {
if (static_cast<spv::Capability>(capability.GetSingleWordInOperand(0)) ==
spv::Capability::Float64) {
deadCapabilities.push_back(&capability);
}
}
for (Instruction* capability : deadCapabilities) {
irContext->KillInst(capability);
}
// --- 7. merge what the rewrite made into a duplicate -------------------------
// `double` and `float` are now the same declaration, and so is every vector,
// matrix, pointer and function type spelled in terms of them. Walking the section
// in declaration order and replacing each duplicate the moment it is found means
// the later types are already canonical by the time they are keyed: SPIR-V forbids
// a type from forward-referencing another, so every operand of the instruction
// being looked at has been through this loop already.
std::map<std::vector<Uint32>, Uint32> keptTypeByShape;
for (Instruction* type = &*irContext->types_values_begin(); type != nullptr;) {
Instruction* next = type->NextNode();
if (!IsDuplicableType(type->opcode())) {
type = next;
continue;
}
std::vector<Uint32> shape{static_cast<Uint32>(type->opcode())};
for (Uint32 i = 0; i < type->NumInOperands(); ++i) {
const Operand& operand = type->GetInOperand(i);
shape.insert(shape.end(), operand.words.begin(), operand.words.end());
}
const auto kept = keptTypeByShape.find(shape);
if (kept == keptTypeByShape.end()) {
keptTypeByShape.emplace(std::move(shape), type->result_id());
type = next;
continue;
}
irContext->KillNamesAndDecorates(type->result_id());
irContext->ReplaceAllUsesWith(type->result_id(), kept->second);
irContext->KillInst(type);
type = next;
}
// --- 8. re-derive the layout of every block that held a 64-bit float -----------
// See BlockRelayout for why this is a recomputation and not a preservation. Only
// blocks that actually narrowed are touched: relaying out an untouched block would
// be pure churn, and would risk disagreeing with glslang over a layout that was
// already correct.
for (Instruction& variable : irContext->types_values()) {
if (variable.opcode() != spv::Op::OpVariable) continue;
const auto storageClass =
static_cast<spv::StorageClass>(variable.GetSingleWordInOperand(0));
if (storageClass != spv::StorageClass::Uniform &&
storageClass != spv::StorageClass::StorageBuffer &&
storageClass != spv::StorageClass::PushConstant) {
continue;
}
const Instruction* pointerType = defUseMgr->GetDef(variable.type_id());
if (pointerType == nullptr) continue;
// An arrayed block (`uniform Blk { ... } blocks[4];`) is a pointer to an array
// of the struct; the layout lives on the struct either way.
const Instruction* blockType = defUseMgr->GetDef(pointerType->GetSingleWordInOperand(1));
while (blockType != nullptr && (blockType->opcode() == spv::Op::OpTypeArray ||
blockType->opcode() == spv::Op::OpTypeRuntimeArray)) {
blockType = defUseMgr->GetDef(blockType->GetSingleWordInOperand(0));
}
if (blockType == nullptr || blockType->opcode() != spv::Op::OpTypeStruct) continue;
if (wideTypeIds.count(blockType->result_id()) == 0) continue;
// A storage block packs std430, a uniform block std140. Before SPIR-V 1.3 a
// storage block was a Uniform-storage variable whose struct carried
// BufferBlock, so the decoration decides rather than the storage class alone.
Bool isStorageBlock = storageClass == spv::StorageClass::StorageBuffer;
for (const Instruction& annotation : irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != blockType->result_id()) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) ==
spv::Decoration::BufferBlock) {
isStorageBlock = true;
}
}
BlockRelayout relayout(irContext, /*std140=*/!isStorageBlock);
if (relayout.Measure(blockType->result_id()).size == 0) {
// A member shape the layout rules here do not describe. Leaving the block
// at its 64-bit offsets keeps the module valid for Vulkan; SPIRV-Cross will
// decline it for ESSL, which is the same outcome as before the demotion.
MGLOG_D("DemoteFloat64Pass: block %%%u contains a member this pass cannot lay "
"out; its 64-bit offsets are left in place",
blockType->result_id());
}
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken DemoteFloat64Pass::CreateDemoteFloat64Pass() {
return spvtools::Optimizer::PassToken(MakeUnique<DemoteFloat64Pass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,84 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DemoteFloat64Pass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Rewrites every 64-bit float in the module to a 32-bit one: `OpTypeFloat 64` becomes
// `OpTypeFloat 32` in place, so every vector, matrix, array, struct, pointer and
// function type that named it keeps its <id> and every decoration attached to it, and
// only the meaning of the leaf type changes. The double literals are re-encoded, the
// now-width-preserving OpFConvert pairs collapse to their operand, and the Float64
// capability goes away.
//
// Why demote at all: no mobile GPU has it. Adreno and Mali both report
// VkPhysicalDeviceFeatures::shaderFloat64 == VK_FALSE (the Magma POST has a row for it),
// so a module declaring Float64 cannot become a pipeline there; and ESSL has no 64-bit
// float type at all, so SPIRV-Cross throws "FP64 not supported in ES profile" and the
// Espryt path never even reaches the driver. Demotion is what makes `double` in an
// application's GLSL compile and run everywhere, at fp32 precision.
//
// BLOCK LAYOUT IS RE-DERIVED, NOT PRESERVED, and that was not the first choice - see
// BlockRelayout in the .cpp for the measurement that forced it. Preserving the 64-bit
// offsets (float + 4 bytes of padding in each slot) keeps the application's byte layout
// intact and is what a Vulkan-only implementation would do, but GLSL ES has no member
// `layout(offset=)`, so SPIRV-Cross recomputes std140/std430 from the declared types
// and refuses any block whose stated offsets disagree - "Buffer block cannot be
// expressed as any of std430, std140, scalar". Every shader with a double in a block
// would then fail to transpile for Espryt at all, which is the case this demotion
// exists to fix. Nor can padding members rescue it: a dmat4 member carries
// MatrixStride 32 and std140 requires 16 for the demoted mat4, and padding BETWEEN
// members cannot change a stride INSIDE one.
//
// What re-deriving costs, stated plainly: a block that held 64-bit members changes its
// driver-visible byte layout, so an application that hard-codes std140 offsets it
// computed for doubles addresses the wrong bytes. Applications that query their offsets
// are unaffected. MobileGL's own default-uniform block is unaffected by construction:
// the frontend builds its uniform routing by reflecting the module this pass produced
// (ProgramSpirvTask::BuildGlobalUboRouting), so glUniform*d - which narrows to float
// for the same reason - writes exactly where the demoted shader reads. Blocks with no
// 64-bit member anywhere are never touched.
//
// Declines (leaves the module byte-identical, so the caller's existing "this module
// still declares Float64" failure path reports it) when the module contains an
// operation whose validity depends on the operand really being 64 bits wide:
// - OpBitcast across the boundary - packDouble2x32 / doubleBitsToUint64 and friends,
// where SPIR-V requires both sides to have the same total bit width;
// - GLSL.std.450 PackDouble2x32 / UnpackDouble2x32, which are defined only for a
// 64-bit float result/operand.
//
// ORDERING: must run before PackDoubleVertexInputsPass, which introduces exactly the
// OpBitcast this pass declines on. After demotion no 64-bit vertex input is left, so
// that pass becomes a no-op rather than a conflict.
//
// The in-place rewrite creates duplicate type declarations by construction - a module
// that had both `double` and `float` ends up with two `OpTypeFloat 32`, and spirv-val
// rejects that ("Duplicate non-aggregate type declarations are not allowed") - so the
// pass merges them itself afterwards. Deliberately NOT by registering spvtools'
// RemoveDuplicates alongside it: that pass also merges structurally identical STRUCTS
// and calls KillNamesAndDecorates on the loser, which would silently delete the OpName
// of one of two distinct-but-identically-shaped interface blocks - and OpName is how
// MobileGL resolves block and varying names. Only the non-aggregate types spirv-val
// actually forbids duplicates of are merged here; arrays and structs are left alone,
// which also keeps a `double[]`'s ArrayStride from being merged into a `float[]`'s.
class DemoteFloat64Pass : public spvtools::opt::Pass {
public:
const char* name() const override { return "mobilegl-demote-float64"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateDemoteFloat64Pass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,636 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FlattenXfbInterfaceBlocksPass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "FlattenXfbInterfaceBlocksPass.h"
#include "spirv.hpp"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/opt/types.h"
#include "source/util/make_unique.h"
#include "source/util/string_utils.h"
#include <memory>
#include <unordered_set>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::IRContext;
using spvtools::opt::Instruction;
using spvtools::opt::Operand;
namespace analysis = spvtools::opt::analysis;
// Decorations that describe how ONE member is interpolated or stored, and
// therefore have to travel from the block member to the variable that replaces
// it. Location is handled separately (it is computed per member); everything
// else about a block - Block itself, member Offsets, builtin decorations - is
// about the aggregate and stays behind on the shadow.
Bool IsMemberDecorationToCarry(spv::Decoration decoration) {
switch (decoration) {
case spv::Decoration::RelaxedPrecision:
case spv::Decoration::Flat:
case spv::Decoration::NoPerspective:
case spv::Decoration::Centroid:
case spv::Decoration::Sample:
case spv::Decoration::Invariant:
case spv::Decoration::Patch:
return true;
default:
return false;
}
}
// Decorations that say something about a member's placement WITHIN the block's
// location, which a free-standing variable cannot express the same way. Rather
// than move a member to a place the consumer will not look for it, decline.
Bool IsMemberDecorationThatBlocksFlattening(spv::Decoration decoration) {
switch (decoration) {
case spv::Decoration::Component:
case spv::Decoration::XfbBuffer:
case spv::Decoration::XfbStride:
case spv::Decoration::Stream:
return true;
default:
return false;
}
}
// Every name the module already spells, so a synthesised "<Block>_<member>" that
// would collide with one declines instead of emitting two declarations of the
// same identifier (which the driver rejects, taking the whole program with it).
std::unordered_set<String> CollectNames(IRContext& irContext) {
std::unordered_set<String> names;
for (auto& debugInst : irContext.debugs2()) {
if (debugInst.opcode() != spv::Op::OpName) continue;
names.insert(debugInst.GetInOperand(1).AsString());
}
return names;
}
// Locations one value of `type` occupies (GL 4.6 core 11.1.2.1 / 15.2): a
// matrix takes one per column, a double-precision vector wider than two takes
// two, an array takes its element's span once per element. 0 means "this pass
// cannot place it", which declines the whole block rather than guessing.
Uint32 LocationSpan(const analysis::Type* type) {
if (type == nullptr) return 0;
if (type->AsFloat() != nullptr || type->AsInteger() != nullptr ||
type->AsBool() != nullptr) {
return 1u;
}
if (const auto* vector = type->AsVector()) {
const auto* element = vector->element_type();
if (element->AsFloat() == nullptr && element->AsInteger() == nullptr &&
element->AsBool() == nullptr) {
return 0;
}
const auto* elementFloat = element->AsFloat();
const Bool is64Bit = elementFloat != nullptr && elementFloat->width() == 64;
return (is64Bit && vector->element_count() > 2) ? 2u : 1u;
}
if (const auto* matrix = type->AsMatrix()) {
const Uint32 columnSpan = LocationSpan(matrix->element_type());
return columnSpan == 0 ? 0 : columnSpan * matrix->element_count();
}
if (const auto* array = type->AsArray()) {
const Uint32 elementSpan = LocationSpan(array->element_type());
if (elementSpan == 0) return 0;
// A runtime array has no span; a vertex-stage interface never has one.
if (!array->length_info().words.empty() &&
array->length_info().words[0] !=
static_cast<Uint32>(analysis::Array::LengthInfo::kConstant)) {
return 0;
}
// words[0] is the tag, words[1..] the constant value; only a
// single-word length can be an array size here.
if (array->length_info().words.size() != 2) return 0;
const Uint32 count = array->length_info().words[1];
return count == 0 ? 0 : elementSpan * count;
}
// Structs (a nested block member) carry their own layout rules and are not
// worth guessing at: declining leaves the module exactly as it was.
return 0;
}
// The Location a variable carries, or false when it carries none.
Bool FindLocationDecoration(IRContext& irContext, Uint32 variableId, Uint32& outLocation) {
for (auto& annotation : irContext.annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != variableId) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) !=
spv::Decoration::Location) {
continue;
}
outLocation = annotation.GetSingleWordInOperand(2);
return true;
}
return false;
}
// The Location decorating member `member` of struct type `typeId`, if any.
Bool FindMemberLocationDecoration(IRContext& irContext, Uint32 typeId, Uint32 member,
Uint32& outLocation) {
for (auto& annotation : irContext.annotations()) {
if (annotation.opcode() != spv::Op::OpMemberDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != typeId) continue;
if (annotation.GetSingleWordInOperand(1) != member) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(2)) !=
spv::Decoration::Location) {
continue;
}
outLocation = annotation.GetSingleWordInOperand(3);
return true;
}
return false;
}
String FindName(IRContext& irContext, Uint32 id) {
for (auto& debugInst : irContext.debugs2()) {
if (debugInst.opcode() != spv::Op::OpName) continue;
if (debugInst.GetSingleWordInOperand(0) != id) continue;
return debugInst.GetInOperand(1).AsString();
}
return String();
}
String FindMemberName(IRContext& irContext, Uint32 typeId, Uint32 member) {
for (auto& debugInst : irContext.debugs2()) {
if (debugInst.opcode() != spv::Op::OpMemberName) continue;
if (debugInst.GetSingleWordInOperand(0) != typeId) continue;
if (debugInst.GetSingleWordInOperand(1) != member) continue;
return debugInst.GetInOperand(2).AsString();
}
return String();
}
} // namespace
Bool FlattenXfbInterfaceBlocksPass::RewriteCaptureName(const String& captureName,
const std::set<String>& flattenedBlockNames,
String& outName) {
const SizeT dot = captureName.find('.');
if (dot == String::npos || dot == 0) return false;
const String blockName = captureName.substr(0, dot);
if (flattenedBlockNames.find(blockName) == flattenedBlockNames.end()) return false;
outName = blockName + "_" + captureName.substr(dot + 1);
return true;
}
spvtools::opt::Pass::Status FlattenXfbInterfaceBlocksPass::Process() {
if (m_blockNames.empty()) return Status::SuccessWithoutChange;
auto* irContext = context();
auto entryPoints = irContext->module()->entry_points();
if (entryPoints.begin() == entryPoints.end()) return Status::SuccessWithoutChange;
Instruction* entryPoint = &*entryPoints.begin();
// A stage whose outputs are published somewhere other than the end of the entry
// point cannot take the shadow-and-copy-out shape: a geometry shader's outputs
// are captured by every OpEmitVertex, and a tessellation control shader's are
// per-invocation slots of an arrayed interface. Copying at OpReturn would
// publish once, at the end, which is silently the wrong data rather than a
// failure - so those stages keep their blocks and the capture keeps its
// spelling.
const auto executionModel =
static_cast<spv::ExecutionModel>(entryPoint->GetSingleWordInOperand(0));
if (executionModel == spv::ExecutionModel::Geometry ||
executionModel == spv::ExecutionModel::TessellationControl) {
MGLOG_I("FlattenXfbInterfaceBlocksPass: execution model %u publishes outputs outside "
"the entry point's return; leaving its blocks declared as blocks",
static_cast<Uint32>(executionModel));
return Status::SuccessWithoutChange;
}
auto* defUseMgr = irContext->get_def_use_mgr();
auto* typeMgr = irContext->get_type_mgr();
struct Member {
Uint32 typeId = 0;
Uint32 interfacePointerTypeId = 0;
Uint32 privatePointerTypeId = 0;
Uint32 variableId = 0;
Uint32 location = 0;
Bool hasLocation = false;
String name;
};
struct Target {
Instruction* variable = nullptr;
Uint32 structTypeId = 0;
Uint32 privatePointerTypeId = 0;
spv::StorageClass storageClass = spv::StorageClass::Output;
String blockName;
std::vector<Member> members;
};
std::vector<Target> targets;
const std::unordered_set<String> existingNames = CollectNames(*irContext);
for (Instruction& inst : irContext->types_values()) {
if (inst.opcode() != spv::Op::OpVariable) continue;
const auto storageClass =
static_cast<spv::StorageClass>(inst.GetSingleWordInOperand(0));
if (storageClass != spv::StorageClass::Input &&
storageClass != spv::StorageClass::Output) {
continue;
}
Instruction* pointerType = defUseMgr->GetDef(inst.type_id());
if (pointerType == nullptr) continue;
const Uint32 pointeeTypeId = pointerType->GetSingleWordInOperand(1);
const analysis::Type* pointeeType = typeMgr->GetType(pointeeTypeId);
const auto* structType = pointeeType != nullptr ? pointeeType->AsStruct() : nullptr;
if (structType == nullptr) continue;
const String blockName = FindName(*irContext, pointeeTypeId);
if (blockName.empty() ||
m_blockNames.find(blockName) == m_blockNames.end()) {
continue;
}
Target target;
target.variable = &inst;
target.structTypeId = pointeeTypeId;
target.storageClass = storageClass;
target.blockName = blockName;
Uint32 blockLocation = 0;
const Bool hasBlockLocation =
FindLocationDecoration(*irContext, inst.result_id(), blockLocation);
Bool usable = true;
for (auto& annotation : irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpMemberDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != pointeeTypeId) continue;
if (IsMemberDecorationThatBlocksFlattening(
static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(2)))) {
usable = false;
break;
}
}
Uint32 runningLocation = blockLocation;
const auto& memberTypes = structType->element_types();
for (Uint32 memberIndex = 0; usable && memberIndex < memberTypes.size(); ++memberIndex) {
Member member;
member.typeId = typeMgr->GetId(memberTypes[memberIndex]);
member.name = FindMemberName(*irContext, pointeeTypeId, memberIndex);
if (member.typeId == 0 || member.name.empty() ||
existingNames.count(blockName + "_" + member.name) != 0) {
usable = false;
break;
}
const Uint32 span = LocationSpan(memberTypes[memberIndex]);
if (span == 0) {
usable = false;
break;
}
Uint32 memberLocation = 0;
if (FindMemberLocationDecoration(*irContext, pointeeTypeId, memberIndex,
memberLocation)) {
member.location = memberLocation;
member.hasLocation = true;
} else if (hasBlockLocation) {
member.location = runningLocation;
member.hasLocation = true;
}
runningLocation += span;
target.members.push_back(member);
}
if (!usable || target.members.empty()) {
MGLOG_I("FlattenXfbInterfaceBlocksPass: block '%s' has a member this pass cannot "
"place; leaving it declared as a block",
blockName.c_str());
continue;
}
targets.push_back(std::move(target));
}
if (targets.empty()) return Status::SuccessWithoutChange;
// The entry point's first block, past the leading OpVariable run SPIR-V
// requires to stay at the top of a function. Resolved BEFORE anything is
// mutated: a decline after the rewrite has started would leave a half-converted
// module behind, and the optimizer serialises whatever the module holds
// regardless of the status this returns.
const Uint32 entryFunctionId = entryPoint->GetSingleWordInOperand(1);
spvtools::opt::Function* entryFunction = nullptr;
for (auto& function : *irContext->module()) {
if (function.result_id() == entryFunctionId) {
entryFunction = &function;
break;
}
}
if (entryFunction == nullptr || entryFunction->begin() == entryFunction->end()) {
return Status::SuccessWithoutChange;
}
{
auto& probeBlock = *entryFunction->begin();
auto probe = probeBlock.begin();
while (probe != probeBlock.end() && probe->opcode() == spv::Op::OpVariable) ++probe;
if (probe == probeBlock.end()) return Status::SuccessWithoutChange;
}
// Demoting the block changes the STORAGE CLASS of every pointer derived from it,
// and a derived pointer's own result type still says Input/Output - which is an
// invalid module ("the result pointer storage class and base pointer storage
// class in OpAccessChain do not match") that spirv-val rejects and a driver may
// silently miscompile. Collected before anything is mutated so an unsupported
// use can still decline the whole rewrite rather than leave the module broken.
std::unordered_set<Uint32> derivedPointers;
for (const auto& target : targets) {
derivedPointers.insert(target.variable->result_id());
}
std::vector<Instruction*> pointersToRetype;
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
const spv::Op opcode = inst.opcode();
const Bool indexes = opcode == spv::Op::OpAccessChain ||
opcode == spv::Op::OpInBoundsAccessChain ||
opcode == spv::Op::OpCopyObject;
if (indexes) {
if (inst.NumInOperands() > 0 &&
derivedPointers.count(inst.GetSingleWordInOperand(0)) != 0) {
derivedPointers.insert(inst.result_id());
pointersToRetype.push_back(&inst);
}
continue;
}
// Reads and writes are exactly what a shadowed block is for; the
// pointer they use does not change shape.
if (opcode == spv::Op::OpLoad || opcode == spv::Op::OpStore) continue;
for (Uint32 i = 0; i < inst.NumInOperands(); ++i) {
const Operand& operand = inst.GetInOperand(i);
if (operand.type != SPV_OPERAND_TYPE_ID || operand.words.size() != 1) {
continue;
}
if (derivedPointers.count(operand.words[0]) == 0) continue;
MGLOG_I("FlattenXfbInterfaceBlocksPass: interface block %%%u reaches a "
"SPIR-V opcode %u that this pass cannot follow; leaving it "
"declared as a block",
operand.words[0], static_cast<Uint32>(opcode));
return Status::SuccessWithoutChange;
}
}
}
}
// Every type this pass names has to exist before the variables that name it.
for (auto& target : targets) {
target.privatePointerTypeId =
typeMgr->FindPointerToType(target.structTypeId, spv::StorageClass::Private);
for (auto& member : target.members) {
member.interfacePointerTypeId =
typeMgr->FindPointerToType(member.typeId, target.storageClass);
member.privatePointerTypeId =
typeMgr->FindPointerToType(member.typeId, spv::StorageClass::Private);
}
}
std::vector<Operand> interfaceOperands;
for (Uint32 i = 0; i < entryPoint->NumInOperands(); ++i) {
interfaceOperands.push_back(entryPoint->GetInOperand(i));
}
std::unordered_set<Uint32> strippedBlockStructTypes;
for (auto& target : targets) {
Instruction* variable = target.variable;
const Uint32 oldVariableId = variable->result_id();
std::vector<std::unique_ptr<Instruction>> newDecorations;
std::vector<std::unique_ptr<Instruction>> newNames;
for (auto& member : target.members) {
member.variableId = irContext->TakeNextId();
irContext->AddGlobalValue(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpVariable, member.interfacePointerTypeId, member.variableId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_STORAGE_CLASS,
{static_cast<Uint32>(target.storageClass)}}}));
// The name IS the contract: it is what the rewritten capture request
// asks the driver for, and what makes the producer and consumer of a
// flattened block still match each other by name.
const String flatName = target.blockName + "_" + member.name;
std::vector<Operand> nameOperands;
nameOperands.push_back({SPV_OPERAND_TYPE_ID, {member.variableId}});
nameOperands.push_back(
{SPV_OPERAND_TYPE_LITERAL_STRING,
spvtools::utils::MakeVector(flatName)});
newNames.push_back(spvtools::MakeUnique<Instruction>(irContext, spv::Op::OpName, 0, 0,
nameOperands));
if (member.hasLocation) {
newDecorations.push_back(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {member.variableId}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<Uint32>(spv::Decoration::Location)}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {member.location}}}));
}
}
// Member decorations that describe the member (not the aggregate) move to
// the variable that now carries it.
for (auto& annotation : irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpMemberDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != target.structTypeId) continue;
const Uint32 memberIndex = annotation.GetSingleWordInOperand(1);
if (memberIndex >= target.members.size()) continue;
const auto decoration =
static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(2));
if (!IsMemberDecorationToCarry(decoration)) continue;
std::vector<Operand> operands;
operands.push_back({SPV_OPERAND_TYPE_ID, {target.members[memberIndex].variableId}});
for (Uint32 i = 2; i < annotation.NumInOperands(); ++i) {
operands.push_back(annotation.GetInOperand(i));
}
newDecorations.push_back(
spvtools::MakeUnique<Instruction>(irContext, spv::Op::OpDecorate, 0, 0, operands));
}
// The block's own Location described where its members start and means
// nothing on a Private shadow; leaving it would also make SPIRV-Cross print
// a location for a variable that no longer has an interface.
std::vector<Instruction*> deadDecorations;
// A struct type shared by two flattened variables would otherwise have its
// Block decoration killed twice - and the second kill is a use-after-free,
// not a no-op.
const Bool structTypeAlreadyStripped =
strippedBlockStructTypes.count(target.structTypeId) != 0;
strippedBlockStructTypes.insert(target.structTypeId);
for (auto& annotation : irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
const Uint32 target0 = annotation.GetSingleWordInOperand(0);
if (target0 == oldVariableId) {
deadDecorations.push_back(&annotation);
continue;
}
if (structTypeAlreadyStripped) continue;
// The struct type stops being an interface block the moment its only
// interface variable becomes a Private shadow, and SPIRV-Cross prints a
// Block-decorated struct as a BLOCK declaration - `out StageData {...}
// vs_out;` for a variable that is no longer an output, which is not
// ESSL and which the driver rejects with a bare syntax error on the
// instance name. A block name is unique per stage, so nothing else can
// still need this decoration.
if (target0 == target.structTypeId &&
static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) ==
spv::Decoration::Block) {
deadDecorations.push_back(&annotation);
}
}
for (auto* annotation : deadDecorations) {
irContext->KillInst(annotation);
}
for (auto& decoration : newDecorations) {
irContext->AddAnnotationInst(std::move(decoration));
}
for (auto& debugName : newNames) {
irContext->AddDebug2Inst(std::move(debugName));
}
// Demote the block to a Private shadow. Every OpAccessChain and OpStore the
// body already performs on it stays valid and keeps its types; only the
// storage class changed, and Private is the one storage class a pointer of
// any shape may live in.
variable->SetResultType(target.privatePointerTypeId);
variable->SetInOperand(0, {static_cast<Uint32>(spv::StorageClass::Private)});
variable->RemoveFromList();
irContext->AddGlobalValue(std::unique_ptr<Instruction>(variable));
// SPIR-V 1.3 lists only Input/Output in the entry-point interface, and the
// block is neither any more: it is replaced in place by its members.
std::vector<Operand> rebuilt;
for (Uint32 i = 0; i < interfaceOperands.size(); ++i) {
const Operand& operand = interfaceOperands[i];
if (i >= 3 && operand.type == SPV_OPERAND_TYPE_ID && operand.words.size() == 1 &&
operand.words[0] == oldVariableId) {
for (const auto& member : target.members) {
rebuilt.push_back({SPV_OPERAND_TYPE_ID, {member.variableId}});
}
continue;
}
rebuilt.push_back(operand);
}
interfaceOperands = std::move(rebuilt);
}
entryPoint->SetInOperands(std::move(interfaceOperands));
// Index constants for the copy access chains. Appended after the variables
// above, which is legal: nothing declared before them names them, and the only
// instructions that do are the ones inserted into the function body below.
auto* constMgr = irContext->get_constant_mgr();
std::vector<Uint32> memberIndexConstants;
Uint32 maxMembers = 0;
for (const auto& target : targets) {
maxMembers = std::max(maxMembers, static_cast<Uint32>(target.members.size()));
}
memberIndexConstants.resize(maxMembers, 0);
for (Uint32 index = 0; index < maxMembers; ++index) {
memberIndexConstants[index] = constMgr->GetUIntConstId(index);
}
// Input: seed the shadow once, before any code that reads it. Output: publish
// it at every exit, after all the code that writes it.
auto& entryBlock = *entryFunction->begin();
auto insertPoint = entryBlock.begin();
while (insertPoint != entryBlock.end() && insertPoint->opcode() == spv::Op::OpVariable) {
++insertPoint;
}
if (insertPoint == entryBlock.end()) return Status::SuccessWithoutChange;
for (const auto& target : targets) {
if (target.storageClass != spv::StorageClass::Input) continue;
for (Uint32 memberIndex = 0; memberIndex < target.members.size(); ++memberIndex) {
const auto& member = target.members[memberIndex];
const Uint32 loadedId = irContext->TakeNextId();
const Uint32 memberPointerId = irContext->TakeNextId();
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpLoad, member.typeId, loadedId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {member.variableId}}}));
++insertPoint;
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpAccessChain, member.privatePointerTypeId, memberPointerId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {target.variable->result_id()}},
{SPV_OPERAND_TYPE_ID, {memberIndexConstants[memberIndex]}}}));
++insertPoint;
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpStore, 0, 0,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {memberPointerId}},
{SPV_OPERAND_TYPE_ID, {loadedId}}}));
++insertPoint;
}
}
// EVERY return, not just the last block: a shader with an early `return` would
// otherwise publish nothing on that path.
std::vector<Instruction*> returns;
for (auto& block : *entryFunction) {
Instruction* terminator = block.terminator();
if (terminator != nullptr && terminator->opcode() == spv::Op::OpReturn) {
returns.push_back(terminator);
}
}
for (Instruction* returnInst : returns) {
for (const auto& target : targets) {
if (target.storageClass != spv::StorageClass::Output) continue;
for (Uint32 memberIndex = 0; memberIndex < target.members.size(); ++memberIndex) {
const auto& member = target.members[memberIndex];
const Uint32 memberPointerId = irContext->TakeNextId();
const Uint32 loadedId = irContext->TakeNextId();
returnInst->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpAccessChain, member.privatePointerTypeId,
memberPointerId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {target.variable->result_id()}},
{SPV_OPERAND_TYPE_ID, {memberIndexConstants[memberIndex]}}}));
returnInst->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpLoad, member.typeId, loadedId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {memberPointerId}}}));
returnInst->InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpStore, 0, 0,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {member.variableId}},
{SPV_OPERAND_TYPE_ID, {loadedId}}}));
}
}
}
// Retype the derived pointers collected above. Done last, so the pointer types it
// appends land after the variables.
for (Instruction* pointer : pointersToRetype) {
Instruction* resultType = defUseMgr->GetDef(pointer->type_id());
if (resultType == nullptr || resultType->opcode() != spv::Op::OpTypePointer) continue;
if (static_cast<spv::StorageClass>(resultType->GetSingleWordInOperand(0)) ==
spv::StorageClass::Private) {
continue;
}
pointer->SetResultType(typeMgr->FindPointerToType(resultType->GetSingleWordInOperand(1),
spv::StorageClass::Private));
}
if (m_flattenedBlockNames != nullptr) {
for (const auto& target : targets) {
m_flattenedBlockNames->insert(target.blockName);
}
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken FlattenXfbInterfaceBlocksPass::CreateFlattenXfbInterfaceBlocksPass(
const std::set<String>& blockNames, std::set<String>* flattenedBlockNames) {
return spvtools::Optimizer::PassToken(
MakeUnique<FlattenXfbInterfaceBlocksPass>(blockNames, flattenedBlockNames));
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,73 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/FlattenXfbInterfaceBlocksPass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
#include <set>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Replaces a named interface BLOCK whose members transform feedback captures with
// one free-standing variable per member, named "<BlockName>_<MemberName>", and
// demotes the block itself to a Private shadow that the entry point copies into
// (Input) or out of (Output). The block's own body code is untouched: every
// OpAccessChain into it - dynamically indexed ones included - keeps working
// against the shadow.
//
// WHY. The Adreno ES driver accepts "BlockName.member" in
// glTransformFeedbackVaryings, links, and reports the names straight back from
// glGetTransformFeedbackVarying - and then captures NONE of them: the recorded
// stream holds gl_Position in slot 0 and leaves the rest of every vertex's record
// untouched. Proven on an Adreno 830 by writing a recognisable gl_Position into
// the KHR-GL43.vertex_attrib_binding.basic-input capture program: the value landed
// in the slot that had asked for StageData.attrib[0]. The same program with the
// block flattened into a plain output array captures every member correctly, which
// is what this pass produces. Requesting the members under the ESSL instance-name
// spelling ("vs_out.attrib[0]") instead makes the driver fail the link outright, so
// the capture list cannot be spelled around the defect - only the declaration can.
//
// The rename is deterministic and derived only from names that GLSL interface
// matching already requires both sides of a stage boundary to agree on (block name
// and member names), so running this pass over EVERY stage of a program with the
// same block set keeps a producer and its consumer matched.
//
// DirectGLES only: Vulkan captures by xfb_offset, never by name.
class FlattenXfbInterfaceBlocksPass : public spvtools::opt::Pass {
public:
// `blockNames` are block TYPE names, i.e. the "StageData" of
// "StageData.attrib[0]". `flattenedBlockNames` receives the subset this pass
// actually rewrote, so the caller can rename exactly those capture requests
// and leave the rest spelled as the application wrote them.
FlattenXfbInterfaceBlocksPass(const std::set<String>& blockNames,
std::set<String>* flattenedBlockNames)
: m_blockNames(blockNames), m_flattenedBlockNames(flattenedBlockNames) {}
const char* name() const override { return "mobilegl-flatten-xfb-interface-blocks"; }
Status Process() override;
// The member of `blockName` spelled by a capture request, as this pass names
// it: "StageData.attrib[0]" -> "StageData_attrib[0]". Returns false when the
// name does not address a member of a flattened block.
static Bool RewriteCaptureName(const String& captureName, const std::set<String>& flattenedBlockNames,
String& outName);
static spvtools::Optimizer::PassToken CreateFlattenXfbInterfaceBlocksPass(
const std::set<String>& blockNames, std::set<String>* flattenedBlockNames);
private:
std::set<String> m_blockNames;
std::set<String>* m_flattenedBlockNames = nullptr;
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,394 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/SplitArrayVertexInputsPass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "SplitArrayVertexInputsPass.h"
#include "spirv.hpp"
#include "source/opt/constants.h"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/opt/types.h"
#include "source/util/make_unique.h"
#include <memory>
#include <unordered_set>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::IRContext;
using spvtools::opt::Instruction;
using spvtools::opt::Operand;
namespace analysis = spvtools::opt::analysis;
// How many attribute locations one element of the array consumes. GL 4.6 core
// 11.1.1: a scalar or vector of up to four 32-bit components takes one; a
// double-precision vector wider than two takes two; a matrix takes one per
// column. Zero means "this pass will not touch it" - the module keeps its array
// input and SPIRV-Cross will say so, which is a better outcome than a silently
// mis-located split.
Uint32 LocationsPerElement(const analysis::Type* type) {
if (type == nullptr) return 0;
// Any scalar takes one location, a 64-bit one included.
if (type->AsFloat() != nullptr || type->AsInteger() != nullptr ||
type->AsBool() != nullptr) {
return 1u;
}
if (const auto* vector = type->AsVector()) {
const auto* element = vector->element_type();
const auto* elementFloat = element->AsFloat();
const Bool is64Bit = elementFloat != nullptr && elementFloat->width() == 64;
if (element->AsFloat() == nullptr && element->AsInteger() == nullptr &&
element->AsBool() == nullptr) {
return 0;
}
return (is64Bit && vector->element_count() > 2) ? 2u : 1u;
}
// Matrices, structs, images and nested arrays are left alone deliberately:
// an array of them is vanishingly rare as a vertex input and each carries its
// own location-assignment rule, so getting one wrong would corrupt every
// attribute after it rather than fail loudly.
return 0;
}
// The Location a variable is decorated with, or `false` when it carries none.
// A vertex input without an explicit location cannot be split: the split has to
// name base+i, and inventing a base would collide with whatever the linker
// assigned.
Bool FindLocationDecoration(IRContext& irContext, Uint32 variableId, Uint32& outLocation) {
for (auto& annotation : irContext.annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != variableId) continue;
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) !=
spv::Decoration::Location) {
continue;
}
outLocation = annotation.GetSingleWordInOperand(2);
return true;
}
return false;
}
} // namespace
spvtools::opt::Pass::Status SplitArrayVertexInputsPass::Process() {
auto* irContext = context();
auto entryPoints = irContext->module()->entry_points();
if (entryPoints.begin() == entryPoints.end()) return Status::SuccessWithoutChange;
Instruction* entryPoint = &*entryPoints.begin();
if (static_cast<spv::ExecutionModel>(entryPoint->GetSingleWordInOperand(0)) !=
spv::ExecutionModel::Vertex) {
return Status::SuccessWithoutChange;
}
auto* defUseMgr = irContext->get_def_use_mgr();
auto* typeMgr = irContext->get_type_mgr();
auto* constMgr = irContext->get_constant_mgr();
struct Target {
Instruction* variable = nullptr;
Uint32 arrayTypeId = 0;
Uint32 elementTypeId = 0;
Uint32 elementCount = 0;
Uint32 baseLocation = 0;
Uint32 locationsPerElement = 1;
Uint32 elementInputPointerTypeId = 0;
Uint32 elementPrivatePointerTypeId = 0;
Uint32 arrayPrivatePointerTypeId = 0;
};
std::vector<Target> targets;
for (Instruction& inst : irContext->types_values()) {
if (inst.opcode() != spv::Op::OpVariable) continue;
if (static_cast<spv::StorageClass>(inst.GetSingleWordInOperand(0)) !=
spv::StorageClass::Input) {
continue;
}
Instruction* pointerType = defUseMgr->GetDef(inst.type_id());
if (pointerType == nullptr) continue;
const Uint32 pointeeTypeId = pointerType->GetSingleWordInOperand(1);
const analysis::Type* pointeeType = typeMgr->GetType(pointeeTypeId);
const auto* arrayType = pointeeType != nullptr ? pointeeType->AsArray() : nullptr;
if (arrayType == nullptr) continue;
// A builtin input array (gl_ClipDistance and friends) is not an attribute and
// has no location; SPIRV-Cross emits those itself.
Uint32 baseLocation = 0;
if (!FindLocationDecoration(*irContext, inst.result_id(), baseLocation)) continue;
const Uint32 elementTypeId = typeMgr->GetId(arrayType->element_type());
const Uint32 locationsPerElement = LocationsPerElement(arrayType->element_type());
if (elementTypeId == 0 || locationsPerElement == 0) {
MGLOG_I("SplitArrayVertexInputsPass: vertex input %%%u is an array whose element "
"type has no single-location mapping; leaving it declared as an array",
inst.result_id());
continue;
}
// OpTypeArray's length is an id of an integer constant.
Instruction* arrayTypeInst = defUseMgr->GetDef(pointeeTypeId);
if (arrayTypeInst == nullptr || arrayTypeInst->NumInOperands() < 2) continue;
const analysis::Constant* lengthConstant =
constMgr->FindDeclaredConstant(arrayTypeInst->GetSingleWordInOperand(1));
if (lengthConstant == nullptr || lengthConstant->AsIntConstant() == nullptr) continue;
const Uint32 elementCount = lengthConstant->AsIntConstant()->GetU32();
if (elementCount == 0) continue;
Target target;
target.variable = &inst;
target.arrayTypeId = pointeeTypeId;
target.elementTypeId = elementTypeId;
target.elementCount = elementCount;
target.baseLocation = baseLocation;
target.locationsPerElement = locationsPerElement;
targets.push_back(target);
}
if (targets.empty()) return Status::SuccessWithoutChange;
// Everything the demoted array's pointer flows into, in SSA order. Demoting the
// variable changes the STORAGE CLASS of every pointer derived from it, and a
// derived pointer's own result type still says Input - which is an invalid
// module ("the result pointer storage class and base pointer storage class in
// OpAccessChain do not match"), so each one has to be retyped too. Collected
// before anything is mutated so an unsupported use can still decline the whole
// rewrite rather than leave a half-converted module behind.
std::unordered_set<Uint32> derivedPointers;
for (const auto& target : targets) {
derivedPointers.insert(target.variable->result_id());
}
std::vector<Instruction*> pointersToRetype;
for (auto& function : *irContext->module()) {
for (auto& block : function) {
for (auto& inst : block) {
const spv::Op opcode = inst.opcode();
// A pointer this pass moved may only be loaded from or indexed into.
// Anything else (handing it to a function, storing THROUGH it, casting
// it) either cannot happen for a vertex input or would need the callee
// rewritten as well, and silently getting that wrong is worse than
// leaving the module for SPIRV-Cross to reject out loud.
const Bool indexes =
opcode == spv::Op::OpAccessChain ||
opcode == spv::Op::OpInBoundsAccessChain ||
opcode == spv::Op::OpCopyObject;
if (indexes) {
if (inst.NumInOperands() > 0 &&
derivedPointers.count(inst.GetSingleWordInOperand(0)) != 0) {
derivedPointers.insert(inst.result_id());
pointersToRetype.push_back(&inst);
}
continue;
}
if (opcode == spv::Op::OpLoad) continue; // reads are always fine
for (Uint32 i = 0; i < inst.NumInOperands(); ++i) {
const Operand& operand = inst.GetInOperand(i);
if (operand.type != SPV_OPERAND_TYPE_ID || operand.words.size() != 1) {
continue;
}
if (derivedPointers.count(operand.words[0]) == 0) continue;
MGLOG_I("SplitArrayVertexInputsPass: array vertex input %%%u reaches a "
"SPIR-V opcode %u that this pass cannot follow; leaving it "
"declared as an "
"array",
operand.words[0], static_cast<Uint32>(opcode));
return Status::SuccessWithoutChange;
}
}
}
}
// Entry block insertion point: after the block's leading OpVariable run, which
// SPIR-V requires to stay at the top of a function's first block.
const Uint32 entryFunctionId = entryPoint->GetSingleWordInOperand(1);
spvtools::opt::Function* entryFunction = nullptr;
for (auto& function : *irContext->module()) {
if (function.result_id() == entryFunctionId) {
entryFunction = &function;
break;
}
}
if (entryFunction == nullptr || entryFunction->begin() == entryFunction->end()) {
return Status::SuccessWithoutChange;
}
auto& entryBlock = *entryFunction->begin();
auto insertPoint = entryBlock.begin();
while (insertPoint != entryBlock.end() && insertPoint->opcode() == spv::Op::OpVariable) {
++insertPoint;
}
if (insertPoint == entryBlock.end()) return Status::SuccessWithoutChange;
// Every type this pass names has to exist before the variables that name it: the
// types-and-variables section is emitted in order and a forward reference to a
// type is invalid SPIR-V. Same staging as PackDoubleVertexInputsPass.
for (auto& target : targets) {
target.elementInputPointerTypeId =
typeMgr->FindPointerToType(target.elementTypeId, spv::StorageClass::Input);
target.elementPrivatePointerTypeId =
typeMgr->FindPointerToType(target.elementTypeId, spv::StorageClass::Private);
target.arrayPrivatePointerTypeId =
typeMgr->FindPointerToType(target.arrayTypeId, spv::StorageClass::Private);
}
// Index constants for the seeding access chains, all of them before any variable.
Uint32 maxElementCount = 0;
for (const auto& target : targets) {
maxElementCount = std::max(maxElementCount, target.elementCount);
}
std::vector<Uint32> indexConstantIds(maxElementCount, 0);
for (Uint32 index = 0; index < maxElementCount; ++index) {
indexConstantIds[index] = constMgr->GetUIntConstId(index);
}
std::vector<Operand> interfaceOperands;
for (Uint32 i = 0; i < entryPoint->NumInOperands(); ++i) {
interfaceOperands.push_back(entryPoint->GetInOperand(i));
}
for (const auto& target : targets) {
Instruction* variable = target.variable;
const Uint32 oldVariableId = variable->result_id();
// Collected, never added inline: AddAnnotationInst mutates the annotation
// list this function is still walking below.
std::vector<std::unique_ptr<Instruction>> newDecorations;
std::vector<Uint32> elementVariableIds(target.elementCount, 0);
for (Uint32 element = 0; element < target.elementCount; ++element) {
const Uint32 elementVariableId = irContext->TakeNextId();
elementVariableIds[element] = elementVariableId;
irContext->AddGlobalValue(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpVariable, target.elementInputPointerTypeId,
elementVariableId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_STORAGE_CLASS,
{static_cast<Uint32>(spv::StorageClass::Input)}}}));
newDecorations.push_back(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {elementVariableId}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<Uint32>(spv::Decoration::Location)}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER,
{target.baseLocation + element * target.locationsPerElement}}}));
}
// The array's own decorations move to the elements, except Location (each
// element got its own above) and anything that only described the aggregate.
// RelaxedPrecision is the one that MUST travel: dropping it changes the
// declared precision of the input in the emitted ESSL.
std::vector<Instruction*> deadDecorations;
for (auto& annotation : irContext->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate) continue;
if (annotation.GetSingleWordInOperand(0) != oldVariableId) continue;
const auto decoration =
static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1));
if (decoration == spv::Decoration::RelaxedPrecision ||
decoration == spv::Decoration::Flat ||
decoration == spv::Decoration::NoPerspective ||
decoration == spv::Decoration::Centroid ||
decoration == spv::Decoration::Sample) {
for (const Uint32 elementVariableId : elementVariableIds) {
std::vector<Operand> operands;
operands.push_back({SPV_OPERAND_TYPE_ID, {elementVariableId}});
for (Uint32 i = 1; i < annotation.NumInOperands(); ++i) {
operands.push_back(annotation.GetInOperand(i));
}
newDecorations.push_back(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpDecorate, 0, 0, operands));
}
}
deadDecorations.push_back(&annotation);
}
for (auto* annotation : deadDecorations) {
irContext->KillInst(annotation);
}
for (auto& decoration : newDecorations) {
irContext->AddAnnotationInst(std::move(decoration));
}
// Demote the original to a Private global: every existing OpLoad and
// OpAccessChain on it stays valid and keeps its array type, dynamic indices
// included. Moved to the end of the section for the same reason as in
// PackDoubleVertexInputsPass - a variable may not forward-reference its type.
variable->SetResultType(target.arrayPrivatePointerTypeId);
variable->SetInOperand(0, {static_cast<Uint32>(spv::StorageClass::Private)});
variable->RemoveFromList();
irContext->AddGlobalValue(std::unique_ptr<Instruction>(variable));
// SPIR-V 1.3 lists only Input/Output in the entry-point interface, and the
// array is no longer an Input: replace it with the elements in place, so the
// interface keeps one entry per live interface variable.
std::vector<Operand> rebuilt;
for (Uint32 i = 0; i < interfaceOperands.size(); ++i) {
const Operand& operand = interfaceOperands[i];
if (i >= 3 && operand.type == SPV_OPERAND_TYPE_ID &&
operand.words.size() == 1 && operand.words[0] == oldVariableId) {
for (const Uint32 elementVariableId : elementVariableIds) {
rebuilt.push_back({SPV_OPERAND_TYPE_ID, {elementVariableId}});
}
continue;
}
rebuilt.push_back(operand);
}
interfaceOperands = std::move(rebuilt);
// Seed the Private array once, at the top of the entry point, before any of
// the code that reads it.
for (Uint32 element = 0; element < target.elementCount; ++element) {
const Uint32 loadedId = irContext->TakeNextId();
const Uint32 elementPointerId = irContext->TakeNextId();
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpLoad, target.elementTypeId, loadedId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {elementVariableIds[element]}}}));
++insertPoint;
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpAccessChain, target.elementPrivatePointerTypeId,
elementPointerId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {oldVariableId}},
{SPV_OPERAND_TYPE_ID, {indexConstantIds[element]}}}));
++insertPoint;
insertPoint = insertPoint.InsertBefore(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpStore, 0, 0,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {elementPointerId}},
{SPV_OPERAND_TYPE_ID, {loadedId}}}));
++insertPoint;
}
}
entryPoint->SetInOperands(std::move(interfaceOperands));
// Retype the derived pointers collected above. Done last, so the pointer types
// it appends land after the variables (nothing in the types-and-variables
// section names them - they are only ever the result type of an instruction in
// a function body).
for (Instruction* pointer : pointersToRetype) {
Instruction* resultType = defUseMgr->GetDef(pointer->type_id());
if (resultType == nullptr || resultType->opcode() != spv::Op::OpTypePointer) continue;
if (static_cast<spv::StorageClass>(resultType->GetSingleWordInOperand(0)) ==
spv::StorageClass::Private) {
continue;
}
pointer->SetResultType(typeMgr->FindPointerToType(resultType->GetSingleWordInOperand(1),
spv::StorageClass::Private));
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken SplitArrayVertexInputsPass::CreateSplitArrayVertexInputsPass() {
return spvtools::Optimizer::PassToken(MakeUnique<SplitArrayVertexInputsPass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,52 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/SplitArrayVertexInputsPass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Replaces an ARRAY vertex input with one input variable per element, at
// consecutive locations, and demotes the original array to a Private global
// seeded once at the top of the entry point.
//
// GLSL ES has no array vertex inputs at all (GLSL ES 3.20 4.3.4: a vertex shader
// input "cannot be ... arrays"), and SPIRV-Cross does not emulate the difference:
// it refuses the whole module with "OpenGL ES doesn't support array input
// variables in vertex shader". The stage then never reaches the driver, the
// program links without a vertex shader, and every draw using it is a silent
// no-op - which is how the entire KHR-GL43.vertex_attrib_binding.basic-input*
// family (its capture program declares `in vec4 vs_in_attrib[16]`) failed on
// DirectGLES with no symptom other than "the draw captured zeros".
//
// Desktop GL DOES allow the declaration, and it means exactly what the split
// produces: element i of an input array consumes location base+i (GL 4.6 core
// 11.1.1). So the split is a spelling change, not a semantic one - the same
// vertex attributes feed the same components, and the frontend's reflection
// (which the backends bind attributes from) is not involved.
//
// Downstream code is untouched on purpose: the original variable keeps its id and
// its array type, so every OpAccessChain into it - including the DYNAMICALLY
// indexed ones a `for` loop produces, which is precisely what an input array is
// usually written for - stays valid against the Private copy.
//
// DirectGLES only. Vulkan takes array vertex inputs as they are.
class SplitArrayVertexInputsPass : public spvtools::opt::Pass {
public:
const char* name() const override { return "mobilegl-split-array-vertex-inputs"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateSplitArrayVertexInputsPass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,149 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/ZeroBaseVertexPass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "ZeroBaseVertexPass.h"
#include "spirv.hpp"
#include "source/opt/constants.h"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/util/make_unique.h"
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
using spvtools::opt::Operand;
// Returns the Input OpVariable decorated with |builtin|, or nullptr if none.
Instruction* FindBuiltinInputVariable(IRContext* context, spv::BuiltIn builtin) {
auto* defUseMgr = context->get_def_use_mgr();
for (auto& annotation : context->annotations()) {
if (annotation.opcode() != spv::Op::OpDecorate || annotation.NumInOperands() < 3) {
continue;
}
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(1)) !=
spv::Decoration::BuiltIn) {
continue;
}
if (static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(2)) != builtin) {
continue;
}
Instruction* variable = defUseMgr->GetDef(annotation.GetSingleWordInOperand(0));
if (variable == nullptr || variable->opcode() != spv::Op::OpVariable ||
static_cast<spv::StorageClass>(variable->GetSingleWordInOperand(0)) !=
spv::StorageClass::Input) {
continue;
}
return variable;
}
return nullptr;
}
} // namespace
spvtools::opt::Pass::Status ZeroBaseVertexPass::Process() {
auto* irContext = context();
auto* defUseMgr = irContext->get_def_use_mgr();
Instruction* baseVertexVar = FindBuiltinInputVariable(irContext, spv::BuiltIn::BaseVertex);
if (baseVertexVar == nullptr) {
return Status::SuccessWithoutChange;
}
const uint32_t baseVertexVarId = baseVertexVar->result_id();
// Collect every load before mutating: rewriting invalidates the use list.
//
// Every OTHER kind of user is enumerated and refused rather than ignored. A read
// that reaches the variable through a copied pointer or a pointer function
// parameter would keep Vulkan's firstVertex while the pass still reported
// success, i.e. a partial rewrite indistinguishable from a complete one. glslang
// emits neither shape from GLSL today, so this fails closed on something that
// cannot happen yet rather than silently half-doing it when it can.
std::vector<Instruction*> baseVertexLoads;
Bool sawUnexpectedUser = false;
defUseMgr->ForEachUser(baseVertexVar, [&](Instruction* user) {
switch (user->opcode()) {
case spv::Op::OpLoad:
if (user->GetSingleWordInOperand(0) == baseVertexVarId) {
baseVertexLoads.push_back(user);
} else {
sawUnexpectedUser = true;
}
return;
// Declarations of the variable, not reads of it.
case spv::Op::OpDecorate:
case spv::Op::OpDecorateId:
case spv::Op::OpDecorateString:
case spv::Op::OpName:
case spv::Op::OpEntryPoint:
return;
default:
sawUnexpectedUser = true;
return;
}
});
if (sawUnexpectedUser) {
return Status::Failure;
}
if (baseVertexLoads.empty()) {
// Declared but never read - the variant is already the shader itself.
return Status::SuccessWithoutChange;
}
// Materialize every zero constant BEFORE touching a single instruction, so the
// constant/type managers are never consulted against a module this pass has
// already half-rewritten - and so the rewrite loop below cannot fail partway
// and leave one behind.
auto* constantMgr = irContext->get_constant_mgr();
auto* typeMgr = irContext->get_type_mgr();
std::vector<uint32_t> zeroIds(baseVertexLoads.size(), 0);
for (size_t i = 0; i < baseVertexLoads.size(); ++i) {
// The zero is built from the LOAD's own type, because a shader may declare
// the builtin as either int or uint.
const uint32_t typeId = baseVertexLoads[i]->type_id();
const spvtools::opt::analysis::Type* type = typeMgr->GetType(typeId);
if (type == nullptr) {
return Status::Failure;
}
const spvtools::opt::analysis::Constant* zero = constantMgr->GetConstant(type, {0u});
if (zero == nullptr) {
return Status::Failure;
}
const Instruction* zeroInst = constantMgr->GetDefiningInstruction(zero, typeId);
if (zeroInst == nullptr) {
return Status::Failure;
}
zeroIds[i] = zeroInst->result_id();
}
// `OpLoad %ty %res %baseVertex` becomes `OpCopyObject %ty %res %zero`. Keeping
// %res makes every downstream use pick the zero up with no further rewriting.
for (size_t i = 0; i < baseVertexLoads.size(); ++i) {
baseVertexLoads[i]->SetOpcode(spv::Op::OpCopyObject);
baseVertexLoads[i]->SetInOperands(Instruction::OperandList{
{SPV_OPERAND_TYPE_ID, {zeroIds[i]}}});
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken ZeroBaseVertexPass::CreateZeroBaseVertexPass() {
return spvtools::Optimizer::PassToken(MakeUnique<ZeroBaseVertexPass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,40 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/ZeroBaseVertexPass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// GL and Vulkan disagree about gl_BaseVertex on NON-INDEXED draws: GL defines it as
// "the value passed to the baseVertex parameter, or zero for a command that has
// none", so every DrawArrays form reads zero, while Vulkan's BaseVertex builtin
// carries the draw's firstVertex there. (For indexed draws both mean the same thing,
// GL's basevertex / Vulkan's vertexOffset, so those must keep the native builtin.)
//
// This pass produces the non-indexed variant of a vertex shader by replacing every
// read of the BaseVertex builtin with a constant zero. The variable itself is left
// declared - removing it would also have to reason about the DrawParameters
// capability that a BaseInstance read in the same module still needs.
//
// Vulkan backend only, and only for the ZeroBaseVertex program variant: the
// DirectGLES path has no BaseVertex builtin at all (see LowerDrawParametersPass).
class ZeroBaseVertexPass : public spvtools::opt::Pass {
public:
const char* name() const override { return "zero-base-vertex"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateZeroBaseVertexPass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
+1
View File
@@ -159,6 +159,7 @@ MobileGL supports runtime configuration via environment variables.
| `MOBILEGL_DISABLE_TIMERQUERY` | Disable GPU timer-query exposure and use. | `0`, `1` | `0` | | `MOBILEGL_DISABLE_TIMERQUERY` | Disable GPU timer-query exposure and use. | `0`, `1` | `0` |
| `MOBILEGL_USE_ANGLE` | Load ANGLE EGL/GLES libraries. | `0`, `1` | `0` | | `MOBILEGL_USE_ANGLE` | Load ANGLE EGL/GLES libraries. | `0`, `1` | `0` |
| `MOBILEGL_DISABLE_SUBGROUP` | Disable Vulkan shader subgroup support. | `0`, `1` | `0` | | `MOBILEGL_DISABLE_SUBGROUP` | Disable Vulkan shader subgroup support. | `0`, `1` | `0` |
| `MOBILEGL_ADVERTISE_FP64` | Advertise `GL_ARB_gpu_shader_fp64`. GLSL `double`/`dvec`/`dmat` compile and run either way - they are narrowed to 32 bits - so this only changes whether an application is told it has 64-bit precision, which it does not. | `0`, `1` | `0` |
| `MOBILEGL_MAGMA_R11G11B10F_FALLBACK` | Use Magma's R11G11B10F format fallback. | `0`, `1` | `0` | | `MOBILEGL_MAGMA_R11G11B10F_FALLBACK` | Use Magma's R11G11B10F format fallback. | `0`, `1` | `0` |
| `MOBILEGL_MAGMA_FRAMESINFLIGHT` | Set Magma frames in flight. | Integer `1``64` | `3` | | `MOBILEGL_MAGMA_FRAMESINFLIGHT` | Set Magma frames in flight. | Integer `1``64` | `3` |
| `MOBILEGL_AVOID_SAMPLER_MIPMAP_MIN_FILTER` | Avoid sampler mipmap minification filters. | `0`, `1` | `0` | | `MOBILEGL_AVOID_SAMPLER_MIPMAP_MIN_FILTER` | Avoid sampler mipmap minification filters. | `0`, `1` | `0` |