Compare commits

..
46 Commits
Author SHA1 Message Date
swung0x48 c8632dfefe [Test] (piglit-android): add on-device piglit harness for MobileGL - patched waffle (WAFFLE_EGL_LIBRARY/WAFFLE_GL_LIBRARY overrides so waffle drives libMobileGL.so directly, AImageReader-backed windows for DirectVulkan since Android ICDs lack VK_EXT_headless_surface, WAFFLE_FORCE_GL_CONTEXT_VERSION to upgrade piglit's low compat context requests to 3.3 core, meson cross fixes) and patched piglit (Android platform support, EGL support decoupled from the X11-dependent EGL tests, and a dispatch-init fix: the waffle resolvers were never installed because gl_fw is NULL during framework construction, so gl* silently bound to the system driver via the DT_NEEDED libEGL's eglGetProcAddress), plus the adb chunked runner with PIGLIT-result parsing, a results comparator, cross-file examples, and the piglit-on-android skill 2026-07-26 11:57:11 -04:00
swung0x48 b8a8a660e1 [Refactor] (Lifecycle): own MobileGL's lifecycle from the EGL layer instead of ELF static ctor/dtor - the first EGL/WGL entry point lazily initializes via a thread-safe, re-init-capable EnsureInitialized (AutoInit is gone), the last eglTerminate with no initialized display and nothing current tears the whole library down deterministically inside the EGL lifecycle, and the global singletons move to leak-at-exit heap storage so process exit runs no backend destructors at all (AutoDestroy and the Windows DllMain abandon hook are gone); fixes the exit-time SIGABRT from undefined static-destruction order - the DirectGLES buffer-pool mutex abort on Android clean exits, and the pre-existing macOS QueryTest/ProgramTest 'Subprocess aborted' gtest failures now pass (ctest 411/411) 2026-07-26 10:59:14 -04:00
swung0x48 7ab83861ca [Fix] (DirectVulkan): suspend presentation while the window is zero-area - a minimized window's out-of-date swapchain used to keep Present submitting on a signaled fence and presenting never-acquired images (adversarial review); also drop logging from the process-detach abandon path 2026-07-26 07:59:40 -04:00
swung0x48 eaeba556a3 [Test] (WGL): add manual Windows smoke tests - hand-rolled WGL bootstrap and GLFW-driven variant covering the zero-area helper-window path 2026-07-26 07:21:36 -04:00
swung0x48 72fa1221a5 [Fix] (DirectVulkan): survive zero-area windows at renderer init - skip the eager first acquire when RecreateSwapchain's minimize guard left no swapchain (GLFW's hidden helper window), and let Present bring the swapchain up once the window has real size 2026-07-26 07:16:11 -04:00
swung0x48 c4254c4bbd [Feat] (WGL): add Windows host layer - drop-in opengl32.dll with WGL over EGLImpl, Win32 window backend plumbing for both backends, ANGLE loader path, and leak-at-exit process teardown 2026-07-26 06:46:29 -04:00
swung0x48 199164c2e0 [Perf] (DirectGLES): route the per-upload GL_PIXEL_UNPACK_BUFFER unbinds through the unpack binding cache - the six texture-upload sites re-issued glBindBuffer(UNPACK, 0) on every upload; with the resting-0 shadow they now no-op after the first 2026-07-21 19:54:55 -04:00
swung0x48 e87063e90c [Fix] (DirectGLES): route default-framebuffer binds through the FBO-binding shadow - the raw glBindFramebuffer(0) in BindCurrentFBO/SyncAndBindFramebufferObject left the shadow claiming the previous user FBO, false-skipping its next re-bind and letting scoped guards restore a stale binding (caught by adversarial review); also scrub buffer-binding shadows when VAO client-attribute staging buffers are deleted, include cube-map arrays in the pack-image-params gate, and make the delete-recording test hook assertion-unwind safe 2026-07-21 19:54:54 -04:00
swung0x48 122da27249 [Fix] (DirectGLES): overhaul readback/copy/blit driver-state handling with shadow-backed RAII guards - pixel-PACK/UNPACK PBO binding caches resting at 0 (the old bind-then-query 'restore' left the user PBO bound forever, capturing later client-memory readbacks), a PACK pixel-store shadow replacing per-readback glGetIntegerv syncs, a driver FBO-binding shadow behind all scoped binders (per-instance prev slots, nest-safe), scratch-FBO attachment shadows that detach cross-aspect residue exactly when present (depth CopyTex* attachments used to wedge later GetTexImage color reads and vice versa), scissor guards around emulation blits (app scissor clipped depth copies), stale-driver-error drains before single-shot glGetError consumers (fallbacks silently dropped readbacks / GenerateMipmap raised phantom app errors in production builds), ClearBufferiv missing BindCurrentFBO(Draw), cube-face glBindTexture INVALID_ENUM cache poisoning, per-slice GL_PACK_IMAGE_HEIGHT/SKIP_IMAGES semantics on 3D GetTexImage, backend texture ids deleted on wrapper destruction with cache/scratch-FBO scrubs (ids used to leak for the context lifetime and dangling cache pointers could false-skip binds), and context-death/MakeCurrent invalidation for all new shadows; regression tests drive the shadows against a recording mock GLES table 2026-07-21 19:54:54 -04:00
swung0x48 bc2d698b3e [Fix] (DirectGLES): apply the read buffer when one FBO is bound as both draw and read
- SyncCurrentFBO skips the READ-target pass when the same GL FBO is bound as
  both draw and read (the common GL_FRAMEBUFFER case), but the read buffer
  (glReadBuffer) is only applied inside SyncToBackend's READ path — so the skip
  silently dropped every glReadBuffer change, leaving the backend read buffer
  stuck at COLOR_ATTACHMENT0.
- Extract the read-buffer application into BackendFramebufferObject::
  SyncReadBufferToBackend and invoke it from the skip branch (target == Read)
  as well as from SyncToBackend, so reads always target the right attachment.
- Bind the backend FBO as READ inside the helper before glReadBuffer, since the
  skip path only bound it as DRAW.
- Fixes KHR-GL3x.draw_buffers.draw_buffers_1 (reading COLOR_ATTACHMENT1 while
  the FBO stays GL_FRAMEBUFFER-bound returned attachment 0's value); the render
  was already correct, only the readback resolved the wrong attachment.
2026-07-21 12:18:55 -04:00
swung0x48 3049c4b82b [Fix] (ShaderTranspiler): stop blanking block comments in the source handed to glslang
- BlankBlockComments replaced comment chars with spaces but preserved interior newlines, so a block comment spanning a newline inside a #define truncated the macro body (VALUE became empty)
- glslang has a conformant preprocessor and collapses a block comment to one space across newlines, so the delivered source now keeps comments intact and lets glslang handle them
- Fixes KHR-GL3x.shaders.preprocessor multiline_comment_define / redefine_object_multiline_comment / function_redefinition_3 (6 cases, both devices)
- FilterUnsupportedGpuShaderInt64 relied on the blanking to skip commented-out #extension lines; it now masks comments locally (MaskCommentsAndQuotedText) like the sibling passes, collecting edits and applying them back-to-front
- conditional_inclusion.basic_2 (defined() via macro expansion) stays failing by design: glslang rejects it as UB and working around it would mean re-running preprocessing MobileGL defers to glslang
2026-07-21 05:27:01 -04:00
swung0x48 2b3850b76b [Fix] (DirectGLES): upload RGB565/RGB5_A1 shadow data as packed 16-bit types
- The 8-bit unorm shadow was uploaded as GL_UNSIGNED_BYTE, leaving the 8->5/6-bit requantization to the driver
- That rounding direction is implementation-defined: Adreno rounds to nearest (lossless round trip), Mali floors
- On Mali mid-range texels drifted one 5-bit step down, failing all 20 KHR-GL33.pixelstoragemodes.teximage3d rgb565/rgb5a1 cases (eps 1/32); layers with exact values (0.125/0.25/1.0-ish) passed, matching the observed 0,1,7-valid pattern
- PreparePackedNormUpload repacks shadow rows to GL_UNSIGNED_SHORT_5_6_5 / 5_5_5_1 with round-to-nearest, which exactly recovers the original 5/6-bit values (the shadow expansion is injective), so the driver stores them verbatim
- Idempotent across each region's level loop (glType is shared); RGBA4 exempt since its 8-bit expansion (v*17) is exact under either rounding
- Wired at all four SyncMipmapsToBackend upload regions (append-mipmaps, immutable TexSubImage, mutable full, dirty-level update)
2026-07-21 04:30:46 -04:00
swung0x48 2a0ae743a0 [Fix] (DirectGLES): glFinish before the glGetTexImage temp-FBO readback
- Mali (tile-based) does not resolve a texture's render into memory when it is read back through a different (temp) FBO than the one it was rendered with
- The cross-FBO glReadPixels raced the deferred tile resolve and returned pre-render clear contents
- Distinct render targets read back byte-identical, so KHR-GLxx.glsl_noperspective failed on Mali-G715 (all four programs read as the clear colour)
- glGetTexImage is already a CPU/GPU sync point so the extra drain is negligible; Adreno resolves eagerly and was unaffected
2026-07-21 03:34:33 -04:00
swung0x48 6839219c10 [Fix] (GLImpl): report GL_NO_ERROR from glGetGraphicsResetStatus
- The generic export stub returned (GLenum)1; dEQP reads any non-zero status as a lost device
- It is polled after every case (gl3cTestPackages.cpp:121) and sets QP_TEST_RESULT_DEVICE_LOST
- Under the default --deqp-terminate-on-device-lost=enable that tears the whole CTS run down
- MobileGL tracks no GPU resets, so GL_NO_ERROR ("no reset detected") is the honest, spec-correct answer
- Routed through GLImpl::GetGraphicsResetStatus like every other entry point, no inline body in Definitions.cpp
2026-07-21 02:40:15 -04:00
swung0x48 52ddb440ca [Feat] (SelfTest/DriverPost): add a noperspective correctness check to the GLES POST - render a strong-perspective quad and read the centre texel to verify the varying interpolates screen-linear (not perspective-correct), carried through the native GL_NV_shader_noperspective_interpolation path when present or MobileGL's exact gl_Position.w/gl_FragCoord.w emulation when absent. PASS = native and correct; WARN = emulated and correct (the fallback path shipping packs hit on such devices); FAIL = interpolation wrong/perspective-correct, or the program will not build. The ESSL header matches the device version because noperspective is rejected at #version 300 es on some drivers even with the extension enabled 2026-07-21 02:09:17 -04:00
swung0x48 79feeffd25 [Feat] (ShaderTranspiler, DirectGLES): emulate noperspective on GLES devices lacking GL_NV_shader_noperspective_interpolation - EmulateNoPerspectivePass pre-multiplies each NoPerspective output by gl_Position.w in the vertex stage and recovers each input via gl_FragCoord.w in the fragment stage (exact screen-linear L = P(a*w)*gl_FragCoord.w, handling whole-variable and component/access-chain reads, scalar and vector varyings), forces highp on emulated varyings, and strips what it cannot emulate; replaces the smooth-strip fallback so no NV extension is ever required. Restricts the vertex pre-multiply to the entry function so a non-inlined helper cannot double-scale 2026-07-20 23:48:48 -04:00
swung0x48 202037b5a3 [Fix] (DirectVulkan): shrink the blended depth-write quirk to MIN/MAX extremum blends only - a fixture-wide trace sweep showed the additive ONE+ONE arm never fires on the 26.3 OIT chain (its accumulation passes disable depth writes themselves) and only hit unrelated additive glow content; also fail reflection toward the gl_FragDepth exemption and zero phantom default-FBO blend slots so stale indexed state cannot trigger the strip 2026-07-20 23:39:28 -04:00
swung0x48 bce9c48c8e [Feat] (ShaderTranspiler, DirectGLES): support noperspective conformantly instead of stripping it - let the qualifier reach glslang as the core SPIR-V NoPerspective decoration (native on DirectVulkan; SPIRV-Cross emits ESSL noperspective + GL_NV_shader_noperspective_interpolation on DirectGLES), and for GLES devices lacking that extension add StripNoPerspectivePass to drop the decoration and fall back to smooth; the old naked substring erase discarded the interpolation shader packs need and mangled identifiers containing the word 2026-07-20 22:59:43 -04:00
swung0x48 b6a7807a3a [Fix] (MG_Util/ShaderTranspiler): reject malformed #version directives instead of legalizing them - an unrecognized version number (329/331), a bad profile keyword, a float or trailing token used to be rewritten to "#version 330 core" (or rescued to 460 by the retry); now InspectShaderLanguage marks such directives invalid so NormalizeVersionDirective and RetargetLegacyVersionDirectiveTo460 leave them for glslang to reject, while every valid version still normalizes as before 2026-07-20 22:03:36 -04:00
swung0x48 48ba622387 [Fix] (MG_Util/ShaderTranspiler): keep #line directives instead of deleting them, dropping only the GLSL-illegal quoted filename and the ones that precede #version, so __LINE__ and compiler diagnostics follow the application's own numbering 2026-07-20 21:06:38 -04:00
swung0x48 05260d1262 [Fix] (MG_Util/ShaderTranspiler): blank block comments lexically instead of erasing them - a '//*** banner ***' line opened a comment the old scanner never closed, so it deleted the rest of the shader, and a commented-out builtin definition renamed every genuine call to a name nothing defines 2026-07-20 21:06:37 -04:00
swung0x48 6eb5ff51c5 [Fix] (MG_Impl/GLImpl): reject the RGTC internal formats on 3D texture targets - RGTC compresses 4x4 blocks of a 2D image and has no 3D form, and the check must run on the raw enum because RGTC now resolves to plain R8/RG8 storage 2026-07-20 21:05:57 -04:00
swung0x48 e526f8e8ac [Fix] (MG_Util/Converters): resolve the GL_COMPRESSED_* internal formats to the uncompressed storage that backs them instead of rejecting them as unknown - GL prescribes this base-format fallback for the six generic formats, and RGTC stores uncompressed because ES exposes no compressor 2026-07-20 21:05:57 -04:00
swung0x48 e724e88eec [Fix] (MG_State, MG_Impl/GLImpl): allocating a mipmap level no longer truncates the chain above it - AllocateLevel now only grows and the callers that genuinely redefine the whole level set (glTexStorage*, mip regeneration, multisample storage, level-0 respecification) drop the tail explicitly 2026-07-20 21:02:54 -04:00
swung0x48 3b175fb88a [Fix] (MG_Impl/GLImpl): a multisample sample count above the format's maximum is INVALID_OPERATION, not INVALID_VALUE - matching both the spec and the native Adreno driver 2026-07-20 21:02:53 -04:00
swung0x48 b5a4e7075a [Fix] (MG_Impl/GLImpl): record a GL error from the unimplemented compressed texture entry points instead of throwing - a C++ exception unwinding through the C GL ABI hard-crashes any caller, and glGetCompressedTexImage reported success while writing nothing 2026-07-20 21:02:53 -04:00
swung0x48 520c2b6750 [Fix] (MG_Backend/DirectGLES): sync GL_TEXTURE_SWIZZLE_* on multisample targets - the early return meant to skip the sampler-only parameters dropped every swizzle write, which the frontend already treats as legal on those targets 2026-07-20 21:02:52 -04:00
swung0x48 57cc652b1d [Fix] (MG_Impl/GLImpl): glIsTransformFeedback reports GL_FALSE instead of claiming every name it is handed is a live transform feedback object 2026-07-20 21:02:52 -04:00
swung0x48 65ea54da9e [Refactor] (ShaderTranspiler, DirectVulkan): replace the hand-rolled SPIR-V word walkers with a DecoratePositionInvariantPass and SPIRV-Reflect-based InstanceIndex detection 2026-07-20 08:35:00 -04:00
swung0x48 c81dd04f08 [Test] (CI, trace_replay): force the blended depth-write quirk on the Linux DirectVulkan OIT retrace lane and tighten that case's SSIM threshold to 0.995 2026-07-20 07:39:25 -04:00
swung0x48 c158bfa584 [Fix] (DirectVulkan): narrow the blended depth-write quirk to order-independent accumulation blends, exempting sorted-transparency, gl_FragDepth writers and fully masked attachments 2026-07-20 07:39:25 -04:00
swung0x48 f9f455144c [Refactor] (MG_Config, DirectVulkan): route the blended depth-write quirk through the FeaturesTable as MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE instead of an ad-hoc getenv 2026-07-20 04:37:23 -04:00
swung0x48 64e4840de2 [Fix] (DirectVulkan): fall back to immutable images when the mutable probe fails, gate robustBufferAccess behind MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS, decode R32/RG32/R16-class readback formats, and warn when shaderStorageImage*WithoutFormat is unavailable 2026-07-20 02:36:26 -04:00
swung0x48 bf7b5755cc [Refactor] (ShaderTranspiler, MG_Backend, MG_Util): gate the subgroup prefix-scan rewrite behind a generic device-quirk registry with GPU vendor detection and MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN override, warn on template mismatch, and block ARB/NV subgroup spellings 2026-07-20 02:36:26 -04:00
swung0x48 293f64b3c2 [Fix] (MG_Impl, DirectGLES): correct ARB_clear_texture error codes, reject cube maps in CopyTextureSubImage2D, advertise the extension on Espryt, and pin the error contracts with tests 2026-07-20 02:36:25 -04:00
swung0x48 f0cc07c937 [Perf] (DirectVulkan): bound vertex-stream conversion by the draw's real fetch range, reuse cached prefixes, pin cached source buffers, and stop repacking client arrays for pointer alignment 2026-07-20 02:36:25 -04:00
swung0x48 4658536652 [Perf] (DirectVulkan): keep the render pass alive for steady-state storage-image draws and skip storage-image collection for programs without them 2026-07-20 02:36:24 -04:00
swung0x48 04b4627c65 [Fix] (DirectVulkan): pass depth/stencil formats through sampled-view resolution so D24S8/D32FS8 samplers stop dropping draws, and memoize per-binding view-format resolution 2026-07-20 02:36:24 -04:00
swung0x48 68e13705c8 [Fix] (MG_Backend/DirectVulkan, ShaderTranspiler, MG_Test, TraceReplay): make iterationRP retrace pass on 64-lane Vulkan devices with subgroup-width emulation and format-aware readback 2026-07-20 02:36:23 -04:00
swung0x48 e5388c0e7e [Fix] (MG_Backend, MG_Impl, ShaderTranspiler, MG_Test): support iterationRP custom images and storage format reinterpretation 2026-07-20 02:35:30 -04:00
swung0x48 8bc4808b1a [Test] (TraceReplay): match the iterationRP fixture name published on the mirror 2026-07-19 23:09:29 -04:00
swung0x48 5d8a5387e2 [Test] (TraceReplay): try the hit.moe mirror before miawa and Git LFS 2026-07-19 22:29:55 -04:00
swung0x48 aa2184e47a [Test] (TraceReplay): register iterationRP in-world fixture (non-CI, fixture files pending LFS) 2026-07-19 21:39:06 -04:00
swung0x48 9152a4a4bc [Test] (TraceReplay): enable improved-transparency fixture in CI and drop unneeded coherent_as_flush 2026-07-19 07:35:34 -04:00
swung0x48 1963b427db [Refactor] (TraceReplay): reorganize trace skills into uniform packages with bundled scripts 2026-07-19 07:00:56 -04:00
swung0x48 f3def150e7 [Fix] (DirectVulkan): suppress blended depth writes on Qualcomm and mark gl_Position invariant to fix MC 26.3 OIT cloud flicker 2026-07-19 05:47:23 -04:00
110 changed files with 9378 additions and 725 deletions
+38 -7
View File
@@ -9,7 +9,23 @@ fi
case_name="$1" case_name="$1"
fixture_dir="${2:-tools/trace_replay/fixtures}" fixture_dir="${2:-tools/trace_replay/fixtures}"
python_bin="${PYTHON:-python3}" python_bin="${PYTHON:-python3}"
mirror_base="${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE:-https://repo.miawa.cn/mgl/tools/trace_replay/fixtures}" # Fixture mirrors, tried in order before falling back to Git LFS. Override the
# whole list with MOBILEGL_TRACE_FIXTURE_MIRROR_BASES (whitespace separated);
# MOBILEGL_TRACE_FIXTURE_MIRROR_BASE still works and is tried first.
default_mirror_bases=(
"https://git.hit.moe/swung0x48/MobileGL/media/branch/dev/tools/trace_replay/fixtures"
"https://repo.miawa.cn/mgl/tools/trace_replay/fixtures"
)
if [ -n "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASES:-}" ]; then
read -r -a mirror_bases <<< "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASES}"
else
mirror_bases=("${default_mirror_bases[@]}")
fi
if [ -n "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE:-}" ]; then
mirror_bases=("${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE}" "${mirror_bases[@]}")
fi
# Optional bearer token for mirrors that require authentication (private Gitea).
mirror_token="${MOBILEGL_TRACE_FIXTURE_MIRROR_TOKEN:-}"
download_attempts="${MOBILEGL_TRACE_FIXTURE_DOWNLOAD_ATTEMPTS:-5}" download_attempts="${MOBILEGL_TRACE_FIXTURE_DOWNLOAD_ATTEMPTS:-5}"
retry_delay="${MOBILEGL_TRACE_FIXTURE_RETRY_DELAY:-2}" retry_delay="${MOBILEGL_TRACE_FIXTURE_RETRY_DELAY:-2}"
@@ -30,7 +46,8 @@ fixture_list="$("${python_bin}" tools/trace_replay/trace_cases.py \
--format fixture-files \ --format fixture-files \
--case "${case_name}" \ --case "${case_name}" \
--fixture-root "${fixture_dir}")" --fixture-root "${fixture_dir}")"
mapfile -t files <<< "${fixture_list}" # Strip CR so the script also works when python emits CRLF (Git Bash on Windows).
mapfile -t files < <(printf '%s\n' "${fixture_list}" | tr -d '\r')
include="$(IFS=,; echo "${files[*]}")" include="$(IFS=,; echo "${files[*]}")"
if [ "${case_name}" = "OpenRA" ]; then if [ "${case_name}" = "OpenRA" ]; then
@@ -106,6 +123,7 @@ fetch_file_from_mirror() {
local attempt local attempt
local partial_size local partial_size
local curl_status local curl_status
local curl_auth
metadata="$(get_lfs_metadata "${file}")" || return 1 metadata="$(get_lfs_metadata "${file}")" || return 1
read -r expected_oid expected_size <<< "${metadata}" read -r expected_oid expected_size <<< "${metadata}"
@@ -136,7 +154,11 @@ fetch_file_from_mirror() {
echo "Starting mirror download for ${file} (attempt ${attempt}/${download_attempts})" echo "Starting mirror download for ${file} (attempt ${attempt}/${download_attempts})"
fi fi
if curl -L --fail --show-error --continue-at - --output "${tmp_file}" "${url}"; then curl_auth=()
if [ -n "${mirror_token}" ]; then
curl_auth=(--header "Authorization: token ${mirror_token}")
fi
if curl -L --fail --show-error --continue-at - "${curl_auth[@]}" --output "${tmp_file}" "${url}"; then
if verify_fixture_file "${tmp_file}" "${file}" "${expected_oid}" "${expected_size}"; then if verify_fixture_file "${tmp_file}" "${file}" "${expected_oid}" "${expected_size}"; then
mv "${tmp_file}" "${file}" mv "${tmp_file}" "${file}"
return 0 return 0
@@ -184,10 +206,19 @@ fetch_from_mirror() {
for file in "${files[@]}"; do for file in "${files[@]}"; do
local name local name
local url local url
local base
local fetched=0
name="$(basename "${file}")" name="$(basename "${file}")"
url="${mirror_base%/}/${name}" for base in "${mirror_bases[@]}"; do
echo "Fetching trace fixture from mirror: ${url}" url="${base%/}/${name}"
if ! fetch_file_from_mirror "${file}" "${url}"; then echo "Fetching trace fixture from mirror: ${url}"
if fetch_file_from_mirror "${file}" "${url}"; then
fetched=1
break
fi
echo "Mirror did not serve ${name}; trying the next mirror" >&2
done
if [ "${fetched}" -ne 1 ]; then
return 1 return 1
fi fi
done done
@@ -196,7 +227,7 @@ fetch_from_mirror() {
if fetch_from_mirror; then if fetch_from_mirror; then
echo "Fetched trace fixture files for ${case_name} from mirror: ${include}" echo "Fetched trace fixture files for ${case_name} from mirror: ${include}"
else else
echo "Mirror fetch failed for ${case_name}; falling back to Git LFS: ${include}" echo "All mirrors failed for ${case_name}; falling back to Git LFS: ${include}"
git lfs install --local git lfs install --local
git lfs pull --include="${include}" --exclude="" git lfs pull --include="${include}" --exclude=""
fi fi
+9
View File
@@ -455,6 +455,15 @@ jobs:
if [ '${{ matrix.backend }}' = 'DirectVulkan' ]; then if [ '${{ matrix.backend }}' = 'DirectVulkan' ]; then
export MOBILEGL_MAGMA_R11G11B10F_FALLBACK=1 export MOBILEGL_MAGMA_R11G11B10F_FALLBACK=1
fi fi
# The blended depth-write quirk auto-enables only on Qualcomm, which no CI
# runner has, so force it on for the OIT case it exists to fix. ForceOn
# bypasses only the vendor gate, so this exercises the real strip on
# lavapipe. The Android AVD lane deliberately leaves it off, keeping the
# unstripped path covered for the same trace.
if [ '${{ matrix.backend }}' = 'DirectVulkan' ] \
&& [ '${{ matrix.case }}' = 'improved-transparency-minecraft-26.3' ]; then
export MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE=1
fi
ctest -V --no-tests=error -R '^MobileGLTraceReplay\.${{ matrix.case }}\.${{ matrix.backend }}$' ctest -V --no-tests=error -R '^MobileGLTraceReplay\.${{ matrix.case }}\.${{ matrix.backend }}$'
- name: Upload actual image - name: Upload actual image
+31 -1
View File
@@ -188,9 +188,12 @@ set(SOURCE_FILES
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EliminateFloatEqualsZeroPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EliminateFloatEqualsZeroPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RenameSamplerFunctionParameterPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RenameSamplerFunctionParameterPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecomposeWorkgroupVec3Pass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecomposeWorkgroupVec3Pass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/LowerDrawParametersPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/LowerDrawParametersPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripUboMemberRelaxedPrecisionPass.cpp MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripUboMemberRelaxedPrecisionPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.cpp
MobileGL/MG_Util/BackendLoaders/OpenGL/Loader.cpp MobileGL/MG_Util/BackendLoaders/OpenGL/Loader.cpp
MobileGL/MG_Util/BackendLoaders/Vulkan/Loader.cpp MobileGL/MG_Util/BackendLoaders/Vulkan/Loader.cpp
@@ -300,6 +303,13 @@ if (ANDROID)
) )
endif() endif()
if (WIN32)
list(APPEND SOURCE_FILES
MobileGL/MG_Impl/WGLImpl/WGLImpl.cpp
MobileGL/MG_Impl/WGLImpl/Exporting/Definitions.cpp
)
endif()
set(MOBILEGL_LINK_LIBRARIES set(MOBILEGL_LINK_LIBRARIES
glslang::glslang glslang::glslang
spirv-cross-c spirv-cross-c
@@ -328,10 +338,18 @@ set(MOBILEGL_INCLUDE_DIR
${SPIRV-Headers_SOURCE_DIR}/include ${SPIRV-Headers_SOURCE_DIR}/include
) )
add_library(${CMAKE_PROJECT_NAME} SHARED add_library(${CMAKE_PROJECT_NAME} SHARED
${SOURCE_FILES} ${SOURCE_FILES}
) )
if (WIN32)
# The wgl* entry points are exported via .def (see the comment in wgl.def);
# only the shared library links it.
target_sources(${CMAKE_PROJECT_NAME} PRIVATE
MobileGL/MG_Impl/WGLImpl/Exporting/wgl.def
)
endif()
if (CMAKE_BUILD_TYPE STREQUAL "Debug") if (CMAKE_BUILD_TYPE STREQUAL "Debug")
set_target_properties(${CMAKE_PROJECT_NAME} PROPERTIES set_target_properties(${CMAKE_PROJECT_NAME} PROPERTIES
C_VISIBILITY_PRESET default C_VISIBILITY_PRESET default
@@ -374,6 +392,18 @@ if(UNIX AND NOT APPLE AND NOT ANDROID)
endforeach() endforeach()
endif() endif()
if(WIN32)
# Drop-in for the classic GL loader path: a copy named opengl32.dll placed
# next to a host executable is what LoadLibrary("opengl32.dll") and gdi32's
# pixel-format forwarding will resolve.
add_custom_command(TARGET ${CMAKE_PROJECT_NAME} POST_BUILD
COMMAND ${CMAKE_COMMAND} -E copy_if_different
"$<TARGET_FILE:${CMAKE_PROJECT_NAME}>"
"$<TARGET_FILE_DIR:${CMAKE_PROJECT_NAME}>/opengl32.dll"
COMMENT "Creating opengl32.dll drop-in copy"
)
endif()
if(NOT ANDROID) if(NOT ANDROID)
add_library(${CMAKE_PROJECT_NAME}_s STATIC add_library(${CMAKE_PROJECT_NAME}_s STATIC
${SOURCE_FILES} ${SOURCE_FILES}
+24
View File
@@ -20,6 +20,15 @@ namespace MobileGL::MG_Config {
extern BackendType ActiveBackendType; extern BackendType ActiveBackendType;
// Tri-state override for device-specific quirks: Auto lets the detected device decide,
// ForceOn/ForceOff bypass the detection in either direction. ForceOn only bypasses the
// device gate - each quirk keeps its structural safety checks.
enum class QuirkOverride : Uint8 {
Auto = 0,
ForceOn,
ForceOff,
};
// Feature toggles parsed once from environment variables in MG_ConfigLoader::Init() // Feature toggles parsed once from environment variables in MG_ConfigLoader::Init()
// (ConfigLoader.cpp), before the accepted-env map is destroyed. All Bool fields share // (ConfigLoader.cpp), before the accepted-env map is destroyed. All Bool fields share
// one truthy rule: the variable is set, non-empty, not "0", and not "false" // one truthy rule: the variable is set, non-empty, not "0", and not "false"
@@ -67,6 +76,21 @@ namespace MobileGL::MG_Config {
// explicitly request a core profile via EGL_CONTEXT_OPENGL_PROFILE_MASK / a >=3.1 // explicitly request a core profile via EGL_CONTEXT_OPENGL_PROFILE_MASK / a >=3.1
// version request. // version request.
Bool RelaxedSemantics = false; Bool RelaxedSemantics = false;
// MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN: overrides the shader-source quirk that
// rewrites the recognized workgroup prefix-scan template on Qualcomm devices with
// subgroups wider than 32 lanes (see ShaderSourceProcessor's quirk registry).
QuirkOverride SubgroupPrefixScanQuirk = QuirkOverride::Auto;
// MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE: overrides the DirectVulkan quirk that
// strips depth writes from accumulation-blended pipelines (MIN/MAX or additive
// ONE+ONE - the multi-pass depth-equality signature) on drivers without
// cross-pipeline vertex position invariance. Sorted-transparency "over" blends,
// gl_FragDepth writers, and fully color-masked attachments are exempt (see
// PipelineFactory::ShouldSuppressDepthWrite). Auto detects Qualcomm.
QuirkOverride MagmaDisableBlendedDepthWriteQuirk = QuirkOverride::Auto;
// MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS: leave the Vulkan robustBufferAccess device
// feature off. It is enabled by default to match GL's defined out-of-range fetch
// behavior; this escape hatch exists to measure or dodge its GPU cost on a device.
Bool DisableRobustBufferAccess = false;
}; };
extern FeaturesTable Features; extern FeaturesTable Features;
} // namespace MobileGL::MG_Config } // namespace MobileGL::MG_Config
+15
View File
@@ -86,6 +86,17 @@ namespace MobileGL::MG_ConfigLoader {
return it != acceptedEnvVariablesMap->end() && IsTruthyValue(it->second); return it != acceptedEnvVariablesMap->end() && IsTruthyValue(it->second);
} }
// Quirk overrides are tri-state: an unset variable keeps device auto-detection, a truthy
// value forces the quirk on, anything else set ("0", "false", "") forces it off.
inline MG_Config::QuirkOverride QueryEnvQuirkOverride(const String& key) {
auto it = acceptedEnvVariablesMap->find(key);
if (it == acceptedEnvVariablesMap->end()) {
return MG_Config::QuirkOverride::Auto;
}
return IsTruthyValue(it->second) ? MG_Config::QuirkOverride::ForceOn
: MG_Config::QuirkOverride::ForceOff;
}
inline Uint32 QueryEnvUint32(const String& key, Uint32 defaultValue, Uint32 minValue, Uint32 maxValue) { inline Uint32 QueryEnvUint32(const String& key, Uint32 defaultValue, Uint32 minValue, Uint32 maxValue) {
auto it = acceptedEnvVariablesMap->find(key); auto it = acceptedEnvVariablesMap->find(key);
if (it == acceptedEnvVariablesMap->end()) { if (it == acceptedEnvVariablesMap->end()) {
@@ -123,6 +134,10 @@ namespace MobileGL::MG_ConfigLoader {
features.TraceSkipAutodestroy = QueryEnvFlag("MOBILEGL_TRACE_SKIP_AUTODESTROY"); features.TraceSkipAutodestroy = QueryEnvFlag("MOBILEGL_TRACE_SKIP_AUTODESTROY");
features.DisableUboRing = QueryEnvFlag("MOBILEGL_DISABLE_UBO_RING"); features.DisableUboRing = QueryEnvFlag("MOBILEGL_DISABLE_UBO_RING");
features.RelaxedSemantics = QueryEnvFlag("MOBILEGL_RELAXED_SEMANTICS"); features.RelaxedSemantics = QueryEnvFlag("MOBILEGL_RELAXED_SEMANTICS");
features.SubgroupPrefixScanQuirk = QueryEnvQuirkOverride("MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN");
features.MagmaDisableBlendedDepthWriteQuirk =
QueryEnvQuirkOverride("MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE");
features.DisableRobustBufferAccess = QueryEnvFlag("MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS");
} }
inline void InitBackendType() { inline void InitBackendType() {
+1
View File
@@ -34,6 +34,7 @@
#define MOBILEGL_EGL_API MOBILEGL_API #define MOBILEGL_EGL_API MOBILEGL_API
#define MOBILEGL_CGL_API MOBILEGL_API #define MOBILEGL_CGL_API MOBILEGL_API
#define MOBILEGL_NSOPENGL_API MOBILEGL_API #define MOBILEGL_NSOPENGL_API MOBILEGL_API
#define MOBILEGL_WGL_API MOBILEGL_API
// ====================== MobileGL configurations ======================= // // ====================== MobileGL configurations ======================= //
#ifndef MOBILEGL_LOG_ACTIVE_LEVEL #ifndef MOBILEGL_LOG_ACTIVE_LEVEL
+7 -1
View File
@@ -14,7 +14,13 @@ namespace MobileGL {
} // namespace MG_Config } // namespace MG_Config
namespace MG_Backend { namespace MG_Backend {
UniquePtr<BackendObject> pActiveBackendObject; // Leak-at-exit storage: the UniquePtr itself lives on the heap and is
// never destroyed by the runtime, so process exit runs no backend
// destructors (static destruction order across TUs is undefined).
// Deterministic teardown happens inside the EGL lifecycle instead:
// the last eglTerminate calls MobileGL::Destroy(), which .reset()s
// these singletons while the process is still healthy.
UniquePtr<BackendObject>& pActiveBackendObject = *new UniquePtr<BackendObject>();
GlobalBackendFunctionsTable gBackendFunctionsTable; GlobalBackendFunctionsTable gBackendFunctionsTable;
} // namespace MG_Backend } // namespace MG_Backend
} // namespace MobileGL } // namespace MobileGL
+38 -33
View File
@@ -9,14 +9,24 @@
#include "Init.h" #include "Init.h"
#include "Config.h" #include "Config.h"
#include <MG_Backend/BackendObjects.h> #include <MG_Backend/BackendObjects.h>
#include <MG_Backend/DirectVulkan/DirectVulkan.h>
#include <MG_State/GLState/Core.h> #include <MG_State/GLState/Core.h>
#include <MG_State/EGLState/Core.h> #include <MG_State/EGLState/Core.h>
#include <MG_Impl/GLImpl/Texture/ProxyTexture.h> #include <MG_Impl/GLImpl/Texture/ProxyTexture.h>
#include <MG_Impl/GLImpl/Framebuffer/GL_Framebuffer.h> #include <MG_Impl/GLImpl/Framebuffer/GL_Framebuffer.h>
#include <atomic>
#include <mutex>
namespace MobileGL { namespace MobileGL {
namespace { namespace {
Bool g_isInitialized = false; std::atomic<Bool> g_isInitialized = false;
thread_local Bool tl_initializing = false;
std::mutex& InitMutex() {
static std::mutex mutex;
return mutex;
}
void DestroyImpl(Bool logLifecycle) { void DestroyImpl(Bool logLifecycle) {
if (!g_isInitialized) { if (!g_isInitialized) {
@@ -64,40 +74,35 @@ namespace MobileGL {
MGLOG_I("MobileGL initialized"); MGLOG_I("MobileGL initialized");
} }
void EnsureInitialized() {
if (g_isInitialized.load(std::memory_order_acquire)) {
return;
}
// Re-entrant call while this thread is already inside Initialize()
// (e.g. an init step routing back through a public entry point).
if (tl_initializing) {
return;
}
const std::lock_guard<std::mutex> lock(InitMutex());
if (g_isInitialized.load(std::memory_order_acquire)) {
return;
}
tl_initializing = true;
Initialize();
tl_initializing = false;
}
void Destroy() { void Destroy() {
DestroyImpl(true); DestroyImpl(true);
} }
#if defined(__linux__) || defined(__APPLE__) // MobileGL's lifecycle is owned entirely by the host-API layers
__attribute__((constructor)) static void AutoInit() { // (EGL/WGL/CGL): initialization happens lazily on the first entry point
Initialize(); // via EnsureInitialized(), and full teardown happens deterministically
} // when the last EGL display is terminated with nothing current (EGLImpl
// calls Destroy()). There is intentionally no static constructor, no
__attribute__((destructor)) static void AutoDestroy() { // static destructor, and no DllMain: the global singletons use
if (MG_Config::Features.TraceSkipAutodestroy) { // leak-at-exit storage (see GlobalObjects.cpp), so a process that exits
return; // without eglTerminate simply leaks them to the OS instead of running
} // backend destructors during static teardown.
#if defined(__APPLE__)
// macOS injected dylibs can run destructors after logging/backend static state is already torn down.
return;
#else
DestroyImpl(false);
#endif
}
#endif
#ifdef _WIN32
BOOL WINAPI DllMain(HMODULE hModule, DWORD ul_reason_for_call, LPVOID lpReserved) {
switch (ul_reason_for_call) {
case DLL_PROCESS_ATTACH:
Initialize();
break;
case DLL_PROCESS_DETACH:
Destroy();
break;
}
return TRUE;
}
#endif
} // namespace MobileGL } // namespace MobileGL
+6
View File
@@ -11,6 +11,12 @@
namespace MobileGL { namespace MobileGL {
void Initialize(); void Initialize();
// Thread-safe, idempotent, and re-entrant wrapper around Initialize().
// Host layers (EGL/WGL/CGL entry points) call this lazily on first use so
// MobileGL's lifecycle never depends on ELF/DLL static constructors, and
// so a fresh init can follow a full Destroy() (e.g. after the last
// eglTerminate).
void EnsureInitialized();
void Destroy(); void Destroy();
namespace MG_Util::Debug { namespace MG_Util::Debug {
+18 -1
View File
@@ -230,6 +230,21 @@ namespace MobileGL {
void (*SetSwapInterval)(Int interval); void (*SetSwapInterval)(Int interval);
}; };
// Coarse GPU vendor identity for gating device-specific quirks. Detected from the
// Vulkan physical-device vendorID or the GLES GL_VENDOR/GL_RENDERER strings; stays
// Unknown when detection is inconclusive, in which case auto-gated quirks stay off.
enum class GpuVendorKind : Uint8 {
Unknown = 0,
Qualcomm,
Arm,
Nvidia,
Amd,
Intel,
ImgTec,
// Software rasterizers (llvmpipe/lavapipe, SwiftShader).
Software,
};
struct DynamicBackendParameters { struct DynamicBackendParameters {
SizeT UniformBufferOffsetAlignment = 256; SizeT UniformBufferOffsetAlignment = 256;
// GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT. 1.0 means the backend cannot filter anisotropically, // GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT. 1.0 means the backend cannot filter anisotropically,
@@ -291,13 +306,15 @@ namespace MobileGL {
Uint32 SubgroupSupportedStages = 0; Uint32 SubgroupSupportedStages = 0;
Uint32 SubgroupSupportedFeatures = 0; Uint32 SubgroupSupportedFeatures = 0;
Bool SubgroupQuadOperationsInAllStages = false; Bool SubgroupQuadOperationsInAllStages = false;
GpuVendorKind GpuVendor = GpuVendorKind::Unknown;
}; };
enum class WindowBackend { enum class WindowBackend {
Android, Android,
X11, X11,
MetalLayer, MetalLayer,
// TODO: Wayland, Windows, etc. Win32, // Handle is an HWND
// TODO: Wayland, etc.
WindowBackendCount, WindowBackendCount,
Unknown = -1 Unknown = -1
}; };
+1 -1
View File
@@ -13,6 +13,6 @@
#include "DirectVulkan/BackendObject_DirectVulkan.h" #include "DirectVulkan/BackendObject_DirectVulkan.h"
namespace MobileGL::MG_Backend { namespace MobileGL::MG_Backend {
extern UniquePtr<BackendObject> pActiveBackendObject; extern UniquePtr<BackendObject>& pActiveBackendObject;
extern GlobalBackendFunctionsTable gBackendFunctionsTable; extern GlobalBackendFunctionsTable gBackendFunctionsTable;
} // namespace MobileGL::MG_Backend } // namespace MobileGL::MG_Backend
@@ -701,9 +701,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
if ((handle.Backend != WindowBackend::Android && if ((handle.Backend != WindowBackend::Android &&
handle.Backend != WindowBackend::X11 && handle.Backend != WindowBackend::X11 &&
handle.Backend != WindowBackend::MetalLayer) || handle.Backend != WindowBackend::MetalLayer &&
handle.Backend != WindowBackend::Win32) ||
!handle.Handle) { !handle.Handle) {
MGLOG_E("DirectGLES backend only supports Android, X11, and CAMetalLayer native windows"); MGLOG_E("DirectGLES backend only supports Android, X11, CAMetalLayer, and Win32 native windows");
return false; return false;
} }
@@ -826,7 +827,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
E_GL_ARB_program_interface_query, E_GL_ARB_framebuffer_object, E_GL_ARB_program_interface_query, E_GL_ARB_framebuffer_object,
E_GL_EXT_framebuffer_object, E_GL_ARB_depth_texture, E_GL_ARB_buffer_storage, E_GL_EXT_framebuffer_object, E_GL_ARB_depth_texture, E_GL_ARB_buffer_storage,
E_GL_ARB_texture_storage, E_GL_ARB_texture_storage_multisample, E_GL_ARB_texture_storage, E_GL_ARB_texture_storage_multisample,
E_GL_ARB_direct_state_access, E_GL_ARB_clear_texture, E_GL_ARB_direct_state_access,
E_GL_ARB_multi_draw_indirect, E_GL_ARB_indirect_parameters, E_GL_ARB_multi_draw_indirect, E_GL_ARB_indirect_parameters,
E_GL_ARB_shader_draw_parameters, E_GL_ARB_gpu_shader5, E_GL_ARB_multi_bind, E_GL_ARB_shader_draw_parameters, E_GL_ARB_gpu_shader5, E_GL_ARB_multi_bind,
E_GL_ARB_shading_language_420pack, E_GL_ARB_vertex_attrib_binding, E_GL_ARB_shading_language_420pack, E_GL_ARB_vertex_attrib_binding,
@@ -1035,6 +1036,32 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_dynamicParameters.ViewportSubpixelBits = m_GLESCapabilities.ViewportSubpixelBits; m_dynamicParameters.ViewportSubpixelBits = m_GLESCapabilities.ViewportSubpixelBits;
m_dynamicParameters.SupportsWideLines = m_dynamicParameters.SupportsWideLines =
m_GLESCapabilities.AliasedLineWidthRangeMax > 1.0f || m_GLESCapabilities.SmoothLineWidthRangeMax > 1.0f; m_GLESCapabilities.AliasedLineWidthRangeMax > 1.0f || m_GLESCapabilities.SmoothLineWidthRangeMax > 1.0f;
const auto containsAny = [](const String& haystack, std::initializer_list<const char*> needles) {
return std::any_of(needles.begin(), needles.end(), [&](const char* needle) {
return haystack.find(needle) != String::npos;
});
};
const String vendorAndRenderer =
m_GLESCapabilities.GLESVendorString + " " + m_GLESCapabilities.GLESRendererString;
if (containsAny(vendorAndRenderer, {"llvmpipe", "SwiftShader", "softpipe"})) {
// Check software rasterizers first: ANGLE-on-llvmpipe reports both.
m_dynamicParameters.GpuVendor = GpuVendorKind::Software;
} else if (containsAny(vendorAndRenderer, {"Qualcomm", "Adreno"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Qualcomm;
} else if (containsAny(vendorAndRenderer, {"Mali", "ARM"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Arm;
} else if (containsAny(vendorAndRenderer, {"NVIDIA"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Nvidia;
} else if (containsAny(vendorAndRenderer, {"AMD", "Radeon"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Amd;
} else if (containsAny(vendorAndRenderer, {"Intel"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Intel;
} else if (containsAny(vendorAndRenderer, {"Imagination", "PowerVR"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::ImgTec;
} else {
m_dynamicParameters.GpuVendor = GpuVendorKind::Unknown;
}
} }
const MG_External::GLESFunctionsTable& BackendObject_DirectGLES::GetGLESFunctions() const { const MG_External::GLESFunctionsTable& BackendObject_DirectGLES::GetGLESFunctions() const {
File diff suppressed because it is too large Load Diff
+575 -43
View File
@@ -267,16 +267,26 @@ namespace MobileGL::MG_Backend::DirectGLES {
Uint g_boundArrayBufferId = 0; Uint g_boundArrayBufferId = 0;
Bool g_boundArrayBufferKnown = false; Bool g_boundArrayBufferKnown = false;
// Driver-level GL_PIXEL_PACK/UNPACK_BUFFER binding shadows (see
// Managers.h). Resting state between operations is 0; scopes in the
// readback/upload paths bind what they need through the cache and
// return to 0, so a stale user PBO can never capture a later
// readback that meant to target client memory.
Uint g_boundPixelPackBufferId = 0;
Bool g_boundPixelPackBufferKnown = false;
Uint g_boundPixelUnpackBufferId = 0;
Bool g_boundPixelUnpackBufferKnown = false;
// Bumped whenever the backend ES context is destroyed; resources with // Bumped whenever the backend ES context is destroyed; resources with
// an older generation hold ids from a dead context. // an older generation hold ids from a dead context.
Uint g_bufferContextGeneration = 1; Uint g_bufferContextGeneration = 1;
// Defined next to the indexed-binding shadow below; forward-declared so // Defined next to the indexed-binding shadow below; forward-declared so
// every glDeleteBuffers site in this namespace can scrub stale shadow // every glDeleteBuffers site in this namespace can scrub stale shadow
// entries (GL resets a deleted buffer's indexed bindings to 0, and a // entries (GL resets a deleted buffer's bindings - indexed and pixel
// recycled name matching a stale shadow entry would otherwise // pack/unpack alike - to 0, and a recycled name matching a stale shadow
// false-skip the rebind). // entry would otherwise false-skip the rebind).
void ScrubIndexedBufferBindingShadowForId(Uint id); void ScrubBufferBindingShadowsForId(Uint id);
// Resources whose owning BufferObject died; ids deleted at the next // Resources whose owning BufferObject died; ids deleted at the next
// sync point with a current ES context. // sync point with a current ES context.
@@ -318,10 +328,16 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == r.id) { if (g_boundArrayBufferKnown && g_boundArrayBufferId == r.id) {
InvalidateArrayBufferBindingCache(); InvalidateArrayBufferBindingCache();
} }
// Pooling keeps the id alive (and thus any driver binding of it);
// drop to unknown rather than claiming the post-delete 0 state.
if ((g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == r.id) ||
(g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == r.id)) {
InvalidatePixelBufferBindingCaches();
}
const std::lock_guard<std::mutex> lock(g_poolMutex); const std::lock_guard<std::mutex> lock(g_poolMutex);
auto& bucket = g_bufferPool[r.storageSize]; auto& bucket = g_bufferPool[r.storageSize];
if (bucket.size() >= kMaxEntriesPerBucket || g_pooledBytes + r.storageSize > kMaxPoolBytes) { if (bucket.size() >= kMaxEntriesPerBucket || g_pooledBytes + r.storageSize > kMaxPoolBytes) {
ScrubIndexedBufferBindingShadowForId(r.id); ScrubBufferBindingShadowsForId(r.id);
g_GLESFuncs.glDeleteBuffers(1, &r.id); // over budget: don't pool g_GLESFuncs.glDeleteBuffers(1, &r.id); // over budget: don't pool
r.id = 0; r.id = 0;
return; return;
@@ -502,7 +518,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
// Need a fresh id: glBufferStorage fails on a buffer that already has // Need a fresh id: glBufferStorage fails on a buffer that already has
// immutable storage, and any prior mutable store is replaced anyway. // immutable storage, and any prior mutable store is replaced anyway.
if (resource->id != 0) { if (resource->id != 0) {
ScrubIndexedBufferBindingShadowForId(resource->id); ScrubBufferBindingShadowsForId(resource->id);
g_GLESFuncs.glDeleteBuffers(1, &resource->id); g_GLESFuncs.glDeleteBuffers(1, &resource->id);
resource->id = 0; resource->id = 0;
} }
@@ -630,7 +646,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) { if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) {
InvalidateArrayBufferBindingCache(); InvalidateArrayBufferBindingCache();
} }
ScrubIndexedBufferBindingShadowForId(glesResource->id); ScrubBufferBindingShadowsForId(glesResource->id);
g_GLESFuncs.glDeleteBuffers(1, &glesResource->id); g_GLESFuncs.glDeleteBuffers(1, &glesResource->id);
glesResource->id = 0; glesResource->id = 0;
} }
@@ -668,6 +684,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
void OnBackendContextDestroyed() { void OnBackendContextDestroyed() {
UnregisterBufferBackendOps(); UnregisterBufferBackendOps();
++g_bufferContextGeneration; ++g_bufferContextGeneration;
InvalidateArrayBufferBindingCache();
InvalidateIndexedBufferBindingCache();
InvalidatePixelBufferBindingCaches();
// The global-UBO ring's id and persistent map died with the context; // The global-UBO ring's id and persistent map died with the context;
// drop the handles (no GL) and let the next draw recreate the ring. // drop the handles (no GL) and let the next draw recreate the ring.
ResetUboRingForNewContext(); ResetUboRingForNewContext();
@@ -694,7 +713,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) { if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) {
InvalidateArrayBufferBindingCache(); InvalidateArrayBufferBindingCache();
} }
ScrubIndexedBufferBindingShadowForId(glesResource->id); ScrubBufferBindingShadowsForId(glesResource->id);
g_GLESFuncs.glDeleteBuffers(1, &glesResource->id); g_GLESFuncs.glDeleteBuffers(1, &glesResource->id);
glesResource->id = 0; glesResource->id = 0;
} }
@@ -819,6 +838,41 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_boundArrayBufferKnown = false; g_boundArrayBufferKnown = false;
} }
void BindPixelPackBufferId(Uint id) {
if (g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == id) {
return;
}
g_GLESFuncs.glBindBuffer(GL_PIXEL_PACK_BUFFER, id);
g_boundPixelPackBufferId = id;
g_boundPixelPackBufferKnown = true;
}
void BindPixelUnpackBufferId(Uint id) {
if (g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == id) {
return;
}
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, id);
g_boundPixelUnpackBufferId = id;
g_boundPixelUnpackBufferKnown = true;
}
void InvalidatePixelBufferBindingCaches() {
g_boundPixelPackBufferId = 0;
g_boundPixelPackBufferKnown = false;
g_boundPixelUnpackBufferId = 0;
g_boundPixelUnpackBufferKnown = false;
}
void NoteBufferIdDeleted(Uint id) {
if (id == 0) {
return;
}
if (g_boundArrayBufferKnown && g_boundArrayBufferId == id) {
InvalidateArrayBufferBindingCache();
}
ScrubBufferBindingShadowsForId(id);
}
namespace { namespace {
// Shadow of the GL indexed buffer bindings so redundant glBindBufferBase/Range // Shadow of the GL indexed buffer bindings so redundant glBindBufferBase/Range
// (same index + id + range) are skipped. isBase distinguishes a whole-buffer // (same index + id + range) are skipped. isBase distinguishes a whole-buffer
@@ -840,12 +894,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
return nullptr; return nullptr;
} }
// glDeleteBuffers resets the deleted buffer's bindings (indexed ones // glDeleteBuffers resets the deleted buffer's bindings (indexed and
// included) to 0 in the current context; mirror that in the shadow, or a // pixel pack/unpack ones included) to 0 in the current context; mirror
// later buffer recycling the same name with a matching range would // that in the shadows, or a later buffer recycling the same name with a
// false-skip its rebind. Default IndexedBufferBinding{} == base(0) == // matching shadow entry would false-skip its rebind. Default
// the post-delete GL state. // IndexedBufferBinding{} == base(0) == the post-delete GL state.
void ScrubIndexedBufferBindingShadowForId(Uint id) { void ScrubBufferBindingShadowsForId(Uint id) {
if (id == 0) return; if (id == 0) return;
for (auto& binding : g_indexedUBOBindings) { for (auto& binding : g_indexedUBOBindings) {
if (binding.id == id) binding = {}; if (binding.id == id) binding = {};
@@ -853,6 +907,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
for (auto& binding : g_indexedSSBOBindings) { for (auto& binding : g_indexedSSBOBindings) {
if (binding.id == id) binding = {}; if (binding.id == id) binding = {};
} }
if (g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == id) {
g_boundPixelPackBufferId = 0;
}
if (g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == id) {
g_boundPixelUnpackBufferId = 0;
}
} }
} // namespace } // namespace
@@ -897,7 +957,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
auto& bucket = g_bufferPool[oldestKey]; auto& bucket = g_bufferPool[oldestKey];
PooledBuffer& e = bucket[oldestIdx]; PooledBuffer& e = bucket[oldestIdx];
if (e.contextGeneration == g_bufferContextGeneration && e.id != 0) { if (e.contextGeneration == g_bufferContextGeneration && e.id != 0) {
ScrubIndexedBufferBindingShadowForId(e.id); ScrubBufferBindingShadowsForId(e.id);
g_GLESFuncs.glDeleteBuffers(1, &e.id); g_GLESFuncs.glDeleteBuffers(1, &e.id);
} }
g_pooledBytes -= e.size; g_pooledBytes -= e.size;
@@ -1064,7 +1124,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
const Bool staleContext = entry.contextGeneration != g_bufferContextGeneration; const Bool staleContext = entry.contextGeneration != g_bufferContextGeneration;
if (!staleContext && entry.retireSerial > completed) continue; if (!staleContext && entry.retireSerial > completed) continue;
if (!staleContext && entry.id != 0) { if (!staleContext && entry.id != 0) {
ScrubIndexedBufferBindingShadowForId(entry.id); ScrubBufferBindingShadowsForId(entry.id);
g_GLESFuncs.glDeleteBuffers(1, &entry.id); g_GLESFuncs.glDeleteBuffers(1, &entry.id);
} }
g_retiredUboRings[i] = g_retiredUboRings.back(); g_retiredUboRings[i] = g_retiredUboRings.back();
@@ -1152,6 +1212,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
for (auto& bufferId : m_clientAttributeBufferIds) { for (auto& bufferId : m_clientAttributeBufferIds) {
if (bufferId != 0) { if (bufferId != 0) {
BufferImpl::NoteBufferIdDeleted(bufferId);
g_GLESFuncs.glDeleteBuffers(1, &bufferId); g_GLESFuncs.glDeleteBuffers(1, &bufferId);
bufferId = 0; bufferId = 0;
} }
@@ -1324,6 +1385,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
ZoneScopedC(TRACY_ZONECOLOR_BACKEND); ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
#endif #endif
g_GLESFuncs.glGenTextures(1, &m_backendTextureId); g_GLESFuncs.glGenTextures(1, &m_backendTextureId);
m_contextGeneration = g_textureContextGeneration;
if (m_backendTextureId == 0) { if (m_backendTextureId == 0) {
MGLOG_E("Failed to generate texture object."); MGLOG_E("Failed to generate texture object.");
MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str()); MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str());
@@ -1332,6 +1394,27 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
} }
BackendTextureObject::~BackendTextureObject() {
if (m_backendTextureId == 0) {
return;
}
// Scrub every driver-state shadow that could false-skip when the name
// or this heap address is recycled - regardless of whether the id can
// still be deleted.
ScratchFBOImpl::NoteTextureIdDeleted(m_backendTextureId);
for (auto& unitCache : g_boundTexturesCache) {
for (auto& boundTexture : unitCache) {
if (boundTexture == this) {
boundTexture = nullptr;
}
}
}
if (m_contextGeneration == g_textureContextGeneration && g_GLESFuncs.glDeleteTextures) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId);
}
m_backendTextureId = 0;
}
void BackendTextureObject::Bind(GLenum target, Uint unit) { void BackendTextureObject::Bind(GLenum target, Uint unit) {
#ifdef TRACY_ENABLE #ifdef TRACY_ENABLE
ZoneScopedC(TRACY_ZONECOLOR_BACKEND); ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
@@ -1364,7 +1447,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
void BackendTextureObject::RecreateBackendTexture() { void BackendTextureObject::RecreateBackendTexture() {
if (m_backendTextureId != 0) { if (m_backendTextureId != 0) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId); ScratchFBOImpl::NoteTextureIdDeleted(m_backendTextureId);
if (m_contextGeneration == g_textureContextGeneration) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId);
}
for (auto& unitCache : g_boundTexturesCache) { for (auto& unitCache : g_boundTexturesCache) {
for (auto& boundTexture : unitCache) { for (auto& boundTexture : unitCache) {
if (boundTexture == this) { if (boundTexture == this) {
@@ -1375,6 +1461,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
g_GLESFuncs.glGenTextures(1, &m_backendTextureId); g_GLESFuncs.glGenTextures(1, &m_backendTextureId);
m_contextGeneration = g_textureContextGeneration;
if (m_backendTextureId == 0) { if (m_backendTextureId == 0) {
MGLOG_E("Failed to regenerate texture object."); MGLOG_E("Failed to regenerate texture object.");
MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str()); MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str());
@@ -1391,7 +1478,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
// glGetIntegerv - that query forces a driver pipeline sync and, because texture // glGetIntegerv - that query forces a driver pipeline sync and, because texture
// uploads run it per dirty texture per frame, it dominated the DirectGLES draw // uploads run it per dirty texture per frame, it dominated the DirectGLES draw
// path. The backend unpack state is set ONLY by MobileGL's own save/restore // path. The backend unpack state is set ONLY by MobileGL's own save/restore
// helpers (this class, TempPixelStoreParameterSync, the R32F copy path), all of // helpers (this class and, historically, the R32F copy path), all of
// which restore to the resting default, so the shadow stays accurate; a one-time // which restore to the resting default, so the shadow stays accurate; a one-time
// forced sync pins the backend to that known default up front. Apply() is // forced sync pins the backend to that known default up front. Apply() is
// compare-and-set, so the (now redundant) glPixelStorei calls also usually no-op. // compare-and-set, so the (now redundant) glPixelStorei calls also usually no-op.
@@ -1563,6 +1650,56 @@ namespace MobileGL::MG_Backend::DirectGLES {
return convertedData.data(); return convertedData.data();
} }
// RGB565/RGB5_A1 shadow data is stored as 8-bit unorm; uploading it as GL_UNSIGNED_BYTE
// leaves the 8-bit -> 5/6-bit requantization to the driver, whose rounding direction is
// implementation-defined: Adreno rounds to nearest (lossless round trip) but Mali floors,
// drifting mid-range texels one 5-bit step down and failing the KHR-GL3x
// pixelstoragemodes.teximage3d rgb565/rgb5a1 1/32-eps checks. Repack the shadow rows into
// the packed 16-bit client type with round-to-nearest instead - that recovers the original
// 5/6-bit values exactly (the shadow expansion round(v * 255 / max) is injective), so the
// driver stores them verbatim with no requantization left to its discretion. 4-bit formats
// (RGBA4) are exempt: their 8-bit expansion (v * 17) is exact under either rounding.
// Always retargets *inOutType for these formats (even for null data) so every upload of a
// level uses the same client type.
static const void* PreparePackedNormUpload(TextureInternalFormat format, const IntVec3& texelSize,
const void* data, SizeT byteSize, GLenum* inOutType,
Vector<Uint8>& packedData) {
if (format != TextureInternalFormat::RGB5 && format != TextureInternalFormat::RGB5A1) {
return data;
}
const Bool hasAlpha = format == TextureInternalFormat::RGB5A1;
const GLenum packedType = hasAlpha ? GL_UNSIGNED_SHORT_5_5_5_1 : GL_UNSIGNED_SHORT_5_6_5;
// Idempotent across a region's level loop: glType is shared, so later levels arrive with
// the already-retargeted packed type and must still be converted.
if (*inOutType != GL_UNSIGNED_BYTE && *inOutType != packedType) {
return data;
}
*inOutType = packedType;
if (data == nullptr || byteSize == 0) {
return data;
}
const SizeT srcPixelBytes = hasAlpha ? 4 : 3;
const SizeT texelCount = std::min(static_cast<SizeT>(std::max(texelSize.x(), 0)) *
static_cast<SizeT>(std::max(texelSize.y(), 0)) *
static_cast<SizeT>(std::max(texelSize.z(), 1)),
byteSize / srcPixelBytes);
packedData.resize(texelCount * sizeof(Uint16));
const Uint8* src = static_cast<const Uint8*>(data);
auto* dst = reinterpret_cast<Uint16*>(packedData.data());
for (SizeT i = 0; i < texelCount; ++i, src += srcPixelBytes) {
const Uint32 r = (static_cast<Uint32>(src[0]) * 31u + 127u) / 255u;
const Uint32 b = (static_cast<Uint32>(src[2]) * 31u + 127u) / 255u;
if (hasAlpha) {
const Uint32 g = (static_cast<Uint32>(src[1]) * 31u + 127u) / 255u;
dst[i] = static_cast<Uint16>((r << 11) | (g << 6) | (b << 1) | (src[3] >= 128 ? 1u : 0u));
} else {
const Uint32 g = (static_cast<Uint32>(src[1]) * 63u + 127u) / 255u;
dst[i] = static_cast<Uint16>((r << 11) | (g << 5) | b);
}
}
return packedData.data();
}
void BackendTextureObject::SyncMipmapsToBackend( void BackendTextureObject::SyncMipmapsToBackend(
const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject) { const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject) {
if (!stateTextureObject) { if (!stateTextureObject) {
@@ -1706,9 +1843,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload( const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType, textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData); convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData = PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
DebugImpl::ErrorLopper::Clear(); DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 uploadSize = const IntVec3 uploadSize =
GetBackendUploadSize(stateTextureObject->GetTarget(), levelTexelSize); GetBackendUploadSize(stateTextureObject->GetTarget(), levelTexelSize);
switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) { switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) {
@@ -1760,7 +1900,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& uploadTargets = textureMipmapObject->GetUploadTargets(); const auto& uploadTargets = textureMipmapObject->GetUploadTargets();
if (TextureImpl::IsMultisampleTextureTarget(targetInternal)) { if (TextureImpl::IsMultisampleTextureTarget(targetInternal)) {
DebugImpl::ErrorLopper::Clear(); DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
switch (targetInternal) { switch (targetInternal) {
case TextureTarget::Texture2DMultisample: case TextureTarget::Texture2DMultisample:
g_GLESFuncs.glTexStorage2DMultisample( g_GLESFuncs.glTexStorage2DMultisample(
@@ -1787,7 +1927,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
} else if (stateTextureObject->IsImmutable() || m_imageBindableStorageRequired) { } else if (stateTextureObject->IsImmutable() || m_imageBindableStorageRequired) {
DebugImpl::ErrorLopper::Clear(); DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 storageSize = GetBackendUploadSize(targetInternal, baseSize); const IntVec3 storageSize = GetBackendUploadSize(targetInternal, baseSize);
switch (MapToBackendTextureTarget(targetInternal)) { switch (MapToBackendTextureTarget(targetInternal)) {
case TextureTarget::Texture2D: case TextureTarget::Texture2D:
@@ -1831,9 +1971,13 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload( const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType, textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData); convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData =
PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
DebugImpl::ErrorLopper::Clear(); DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 uploadSize = const IntVec3 uploadSize =
GetBackendUploadSize(targetInternal, levelTexelSize); GetBackendUploadSize(targetInternal, levelTexelSize);
switch (MapToBackendTextureTarget(targetInternal)) { switch (MapToBackendTextureTarget(targetInternal)) {
@@ -1885,6 +2029,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload( const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType, textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData); convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData =
PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
MGLOG_D("%s: target: %s: syncing mip %d: %dx%dx%d, byteSize = %d, pData = %p, " MGLOG_D("%s: target: %s: syncing mip %d: %dx%dx%d, byteSize = %d, pData = %p, "
"levelDirty = %s", "levelDirty = %s",
__func__, MG_Util::ConvertTextureUploadTargetToString(uploadTarget).c_str(), __func__, MG_Util::ConvertTextureUploadTargetToString(uploadTarget).c_str(),
@@ -1892,7 +2040,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
levelByteSize, pData, levelDirty ? "true" : "false"); levelByteSize, pData, levelDirty ? "true" : "false");
DebugImpl::ErrorLopper::Clear(); DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
auto textureTarget = stateTextureObject->GetTarget(); auto textureTarget = stateTextureObject->GetTarget();
const IntVec3 uploadSize = GetBackendUploadSize(textureTarget, levelTexelSize); const IntVec3 uploadSize = GetBackendUploadSize(textureTarget, levelTexelSize);
switch (MapToBackendTextureTarget(textureTarget)) { switch (MapToBackendTextureTarget(textureTarget)) {
@@ -1978,7 +2126,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
textureMipmapObject->GetMipmapTexelSize(uploadTarget, level).y(), byteSize); textureMipmapObject->GetMipmapTexelSize(uploadTarget, level).y(), byteSize);
auto glUploadTarget = ConvertTextureUploadTargetToBackendGLEnum(uploadTarget); auto glUploadTarget = ConvertTextureUploadTargetToBackendGLEnum(uploadTarget);
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0); BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
DebugImpl::ErrorLopper::Loop( DebugImpl::ErrorLopper::Loop(
[file = __FILE__, line = __LINE__, func = __func__](GLenum err) { [file = __FILE__, line = __LINE__, func = __func__](GLenum err) {
MGLOG_D("%s(%s:%d) ES error: %s", func, file, line, MGLOG_D("%s(%s:%d) ES error: %s", func, file, line,
@@ -1990,6 +2138,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload( const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), texelSize, mipData, byteSize, glType, textureMipmapObject->GetFormat(), texelSize, mipData, byteSize, glType,
convertedUploadData); convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData = PreparePackedNormUpload(textureMipmapObject->GetFormat(), texelSize,
uploadData, byteSize, &glType, packedUploadData);
const IntVec3 uploadSize = const IntVec3 uploadSize =
GetBackendUploadSize(stateTextureObject->GetTarget(), texelSize); GetBackendUploadSize(stateTextureObject->GetTarget(), texelSize);
switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) { switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) {
@@ -2216,11 +2367,17 @@ namespace MobileGL::MG_Backend::DirectGLES {
return; return;
} }
if (TextureImpl::IsMultisampleTextureTarget(targetInternal)) { // Multisample targets reject the *sampler* parameters (LOD range, border color) but
// GL_TEXTURE_SWIZZLE_* is texture state, not sampler state, and ES accepts it on them.
// Bailing out entirely used to drop every swizzle write on the floor, which is what the
// frontend already assumes is legal (see GL_Texture.cpp's MS-invalid pname list, which
// deliberately omits the swizzle enums). Note the caches for the skipped parameters are
// still refreshed so they never look stale, but m_cacheSwizzleParams must NOT be, or the
// change detection below would swallow the very writes we came here to emit.
const Bool isMultisampleTarget = TextureImpl::IsMultisampleTextureTarget(targetInternal);
if (isMultisampleTarget) {
m_cacheLodRange = stateTextureObject->GetLevelRange(); m_cacheLodRange = stateTextureObject->GetLevelRange();
m_cacheSwizzleParams = stateTextureObject->GetAllSwizzleParams();
m_cacheBorderColor = stateTextureObject->GetBorderColor(); m_cacheBorderColor = stateTextureObject->GetBorderColor();
return;
} }
Bind(target); Bind(target);
@@ -2233,14 +2390,14 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& levelRange = stateTextureObject->GetLevelRange(); const auto& levelRange = stateTextureObject->GetLevelRange();
if (m_cacheLodRange.x() != levelRange.x()) { if (!isMultisampleTarget && m_cacheLodRange.x() != levelRange.x()) {
g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_BASE_LEVEL, static_cast<GLint>(levelRange.x())); g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_BASE_LEVEL, static_cast<GLint>(levelRange.x()));
m_cacheLodRange.x() = levelRange.x(); m_cacheLodRange.x() = levelRange.x();
} }
DebugImpl::ErrorLopper::Loop([file = __FILE__, line = __LINE__, func = __func__](GLenum err) { DebugImpl::ErrorLopper::Loop([file = __FILE__, line = __LINE__, func = __func__](GLenum err) {
MGLOG_D("%s(%s:%d) ES error %s", func, file, line, MG_Util::ConvertGLEnumToString(err).c_str()); MGLOG_D("%s(%s:%d) ES error %s", func, file, line, MG_Util::ConvertGLEnumToString(err).c_str());
}); });
if (m_cacheLodRange.y() != levelRange.y()) { if (!isMultisampleTarget && m_cacheLodRange.y() != levelRange.y()) {
g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_MAX_LEVEL, static_cast<GLint>(levelRange.y())); g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_MAX_LEVEL, static_cast<GLint>(levelRange.y()));
m_cacheLodRange.y() = levelRange.y(); m_cacheLodRange.y() = levelRange.y();
} }
@@ -2266,7 +2423,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
}); });
} }
if (m_cacheBorderColor != stateTextureObject->GetBorderColor()) { if (!isMultisampleTarget && m_cacheBorderColor != stateTextureObject->GetBorderColor()) {
const auto& borderColor = stateTextureObject->GetBorderColor(); const auto& borderColor = stateTextureObject->GetBorderColor();
GLfloat borderColorArray[4] = {borderColor.x(), borderColor.y(), borderColor.z(), borderColor.w()}; GLfloat borderColorArray[4] = {borderColor.x(), borderColor.y(), borderColor.z(), borderColor.w()};
g_GLESFuncs.glTexParameterfv(target, GL_TEXTURE_BORDER_COLOR, borderColorArray); g_GLESFuncs.glTexParameterfv(target, GL_TEXTURE_BORDER_COLOR, borderColorArray);
@@ -2295,6 +2452,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
Uint g_activeTextureUnit = 0; Uint g_activeTextureUnit = 0;
Uint g_textureContextGeneration = 1;
Array<Array<BackendTextureObject*, (SizeT)TextureTarget::TextureTargetCount>, Array<Array<BackendTextureObject*, (SizeT)TextureTarget::TextureTargetCount>,
MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS> MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS>
g_boundTexturesCache; g_boundTexturesCache;
@@ -2320,9 +2478,59 @@ namespace MobileGL::MG_Backend::DirectGLES {
ZoneScopedC(TRACY_ZONECOLOR_BACKEND); ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
#endif #endif
if (target == FramebufferTarget::Read) if (target == FramebufferTarget::Read)
g_GLESFuncs.glBindFramebuffer(GL_READ_FRAMEBUFFER, m_backendFBOId); BindFramebufferId(GL_READ_FRAMEBUFFER, m_backendFBOId);
else else
g_GLESFuncs.glBindFramebuffer(GL_DRAW_FRAMEBUFFER, m_backendFBOId); BindFramebufferId(GL_DRAW_FRAMEBUFFER, m_backendFBOId);
}
namespace {
// Driver-level framebuffer-binding shadow (see Managers.h). Indexed by
// FramebufferTarget {Draw, Read}.
Array<Uint, SizeT(FramebufferTarget::FramebufferTargetCount)> g_driverFBOBindings = {0, 0};
Array<Bool, SizeT(FramebufferTarget::FramebufferTargetCount)> g_driverFBOBindingKnown = {false, false};
} // namespace
void BindFramebufferId(GLenum fbTarget, Uint id) {
const Bool bindsDraw = fbTarget == GL_DRAW_FRAMEBUFFER || fbTarget == GL_FRAMEBUFFER;
const Bool bindsRead = fbTarget == GL_READ_FRAMEBUFFER || fbTarget == GL_FRAMEBUFFER;
const SizeT drawIdx = SizeT(FramebufferTarget::Draw);
const SizeT readIdx = SizeT(FramebufferTarget::Read);
const Bool drawMatches =
!bindsDraw || (g_driverFBOBindingKnown[drawIdx] && g_driverFBOBindings[drawIdx] == id);
const Bool readMatches =
!bindsRead || (g_driverFBOBindingKnown[readIdx] && g_driverFBOBindings[readIdx] == id);
if (drawMatches && readMatches) {
return;
}
g_GLESFuncs.glBindFramebuffer(fbTarget, id);
if (bindsDraw) {
g_driverFBOBindings[drawIdx] = id;
g_driverFBOBindingKnown[drawIdx] = true;
}
if (bindsRead) {
g_driverFBOBindings[readIdx] = id;
g_driverFBOBindingKnown[readIdx] = true;
}
}
Uint CurrentFramebufferBinding(FramebufferTarget target) {
const SizeT idx = SizeT(target);
if (!g_driverFBOBindingKnown[idx]) {
// Cold path: pin the shadow from the driver once (init probes and
// pre-shadow code bind raw but restore what they found).
GLint binding = 0;
g_GLESFuncs.glGetIntegerv(
target == FramebufferTarget::Read ? GL_READ_FRAMEBUFFER_BINDING : GL_DRAW_FRAMEBUFFER_BINDING,
&binding);
g_driverFBOBindings[idx] = static_cast<Uint>(binding);
g_driverFBOBindingKnown[idx] = true;
}
return g_driverFBOBindings[idx];
}
void InvalidateFramebufferBindingCache() {
g_driverFBOBindings = {0, 0};
g_driverFBOBindingKnown = {false, false};
} }
void BackendFramebufferObject::InvalidateSyncedState() { void BackendFramebufferObject::InvalidateSyncedState() {
@@ -2376,7 +2584,13 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (glTextureTarget == GL_UNKNOWN_MGL) { if (glTextureTarget == GL_UNKNOWN_MGL) {
glTextureTarget = TextureImpl::ConvertTextureTargetToBackendGLEnum(textureObject->GetTarget()); glTextureTarget = TextureImpl::ConvertTextureTargetToBackendGLEnum(textureObject->GetTarget());
} }
backendTextureObject->Bind(glTextureTarget); // glBindTexture rejects cube-face enums (INVALID_ENUM with no
// bind, while Bind() would still record the cube-map cache slot
// as bound): bind via the owning cube target; the attach below
// keeps the face target.
const Bool isCubeFace = glTextureTarget >= GL_TEXTURE_CUBE_MAP_POSITIVE_X &&
glTextureTarget <= GL_TEXTURE_CUBE_MAP_NEGATIVE_Z;
backendTextureObject->Bind(isCubeFace ? GL_TEXTURE_CUBE_MAP : glTextureTarget);
g_GLESFuncs.glFramebufferTexture2D(glFBOTarget, glBackendAttachment, glTextureTarget, g_GLESFuncs.glFramebufferTexture2D(glFBOTarget, glBackendAttachment, glTextureTarget,
backendTextureObject->GetBackendTextureId(), backendTextureObject->GetBackendTextureId(),
static_cast<GLint>(attachmentObject.GetTextureLevel())); static_cast<GLint>(attachmentObject.GetTextureLevel()));
@@ -2465,6 +2679,28 @@ namespace MobileGL::MG_Backend::DirectGLES {
return false; return false;
} }
void BackendFramebufferObject::SyncReadBufferToBackend(
const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject) {
if (!stateFBOObject) {
return;
}
auto frontendReadBuf = stateFBOObject->GetReadBuffer();
if (frontendReadBuf == m_frontendReadBuffer) {
return;
}
m_frontendReadBuffer = frontendReadBuf;
GLenum glBackendReadBuffer = GetBackendAttachmentType(frontendReadBuf);
if (m_backendReadBuffer != glBackendReadBuffer) {
m_backendReadBuffer = glBackendReadBuffer;
// glReadBuffer targets whatever FBO is bound to GL_READ_FRAMEBUFFER. When this is
// reached from SyncCurrentFBO's "same FBO as draw" skip path the backend FBO was
// only bound as DRAW, so bind it as READ first to route the read buffer correctly.
Bind(FramebufferTarget::Read);
g_GLESFuncs.glReadBuffer(glBackendReadBuffer);
}
}
void BackendFramebufferObject::SyncToBackend( void BackendFramebufferObject::SyncToBackend(
const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject, FramebufferTarget asTarget) { const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject, FramebufferTarget asTarget) {
#ifdef TRACY_ENABLE #ifdef TRACY_ENABLE
@@ -2546,16 +2782,8 @@ namespace MobileGL::MG_Backend::DirectGLES {
// 2. Remap read buffer. glReadBuffer writes the READ-bound FBO's state, so // 2. Remap read buffer. glReadBuffer writes the READ-bound FBO's state, so
// only apply (and stamp the memo) when this object is bound as READ. // only apply (and stamp the memo) when this object is bound as READ.
auto frontendReadBuf = stateFBOObject->GetReadBuffer(); if (asTarget == FramebufferTarget::Read) {
if (frontendReadBuf != m_frontendReadBuffer && asTarget == FramebufferTarget::Read) { SyncReadBufferToBackend(stateFBOObject);
m_frontendReadBuffer = frontendReadBuf;
GLenum glBackendReadBuffer = GetBackendAttachmentType(frontendReadBuf);
if (m_backendReadBuffer != glBackendReadBuffer) {
m_backendReadBuffer = glBackendReadBuffer;
g_GLESFuncs.glReadBuffer(glBackendReadBuffer);
}
} }
// -------------------- Attach texture to backend FBO ----------------------- // -------------------- Attach texture to backend FBO -----------------------
@@ -2662,6 +2890,295 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_fboSyncedObjects = {}; g_fboSyncedObjects = {};
} // namespace FramebufferImpl } // namespace FramebufferImpl
namespace ScratchFBOImpl {
namespace {
ScratchFramebuffer g_tempFramebuffer;
ScratchFramebuffer g_blitReadFramebuffer;
ScratchFramebuffer g_blitDrawFramebuffer;
Uint g_completeTinyFBOId = 0;
Uint g_completeTinyRBOId = 0;
// Detach every point the shadow no longer vouches for. Used when the
// shadow is unknown (context reset, texture id deleted while attached).
void ScrubAllAttachments(ScratchFramebuffer& fb, GLenum fbTarget) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
fb.colorTex = 0;
fb.colorTarget = 0;
fb.colorLevel = 0;
fb.colorLayer = -1;
fb.depthTex = 0;
fb.depthTarget = 0;
fb.depthLevel = 0;
fb.depthHasStencil = false;
fb.attachmentsKnown = true;
}
void PrepareForUse(ScratchFramebuffer& fb, GLenum fbTarget) {
if (!fb.attachmentsKnown) {
ScrubAllAttachments(fb, fbTarget);
}
}
// The post-attach glGetError probe below must not misread an error some
// earlier operation left queued; drain before attaching (rare path -
// only runs when the attachment actually changes).
void DrainPendingGLErrors() {
while (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
}
}
// Record the color point as detached when the shadow said something was
// there; the actual detach call is the caller's (it may be replaced by
// the new attach directly when the point is being overwritten).
void RecordNoColor(ScratchFramebuffer& fb) {
fb.colorTex = 0;
fb.colorTarget = 0;
fb.colorLevel = 0;
fb.colorLayer = -1;
}
void RecordNoDepth(ScratchFramebuffer& fb) {
fb.depthTex = 0;
fb.depthTarget = 0;
fb.depthLevel = 0;
fb.depthHasStencil = false;
}
} // namespace
ScratchFramebuffer& TempFramebuffer() {
return g_tempFramebuffer;
}
ScratchFramebuffer& BlitReadFramebuffer() {
return g_blitReadFramebuffer;
}
ScratchFramebuffer& BlitDrawFramebuffer() {
return g_blitDrawFramebuffer;
}
Uint EnsureId(ScratchFramebuffer& fb) {
if (fb.id == 0) {
g_GLESFuncs.glGenFramebuffers(1, &fb.id);
// A fresh FBO has nothing attached and COLOR_ATTACHMENT0 read/draw
// buffers (the ES defaults for a non-default framebuffer).
fb.attachmentsKnown = true;
RecordNoColor(fb);
RecordNoDepth(fb);
fb.readBuffer = GL_COLOR_ATTACHMENT0;
fb.drawBuffer = GL_COLOR_ATTACHMENT0;
}
return fb.id;
}
void EnsureColorAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget,
GLint level) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
if (fb.colorTex == tex && fb.colorTarget == texTarget && fb.colorLevel == level && fb.colorLayer < 0) {
return;
}
if (fb.colorTex != 0) {
// Detach first: if the new attach fails, the point must read as
// missing (incomplete FBO), not silently keep the old texture.
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, texTarget, tex, level);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoColor(fb);
return;
}
fb.colorTex = tex;
fb.colorTarget = texTarget;
fb.colorLevel = level;
fb.colorLayer = -1;
}
void EnsureColorAttachmentLayer(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLint level, GLint layer) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
if (fb.colorTex == tex && fb.colorTarget == 0 && fb.colorLevel == level && fb.colorLayer == layer) {
return;
}
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTextureLayer(fbTarget, GL_COLOR_ATTACHMENT0, tex, level, layer);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoColor(fb);
return;
}
fb.colorTex = tex;
fb.colorTarget = 0;
fb.colorLevel = level;
fb.colorLayer = layer;
}
void EnsureDepthAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level,
Bool withStencil) {
PrepareForUse(fb, fbTarget);
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
RecordNoColor(fb);
}
if (fb.depthTex == tex && fb.depthTarget == texTarget && fb.depthLevel == level &&
fb.depthHasStencil == withStencil) {
return;
}
if (fb.depthTex != 0) {
// One call clears both depth and stencil points regardless of how
// the previous attachment was made.
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTexture2D(fbTarget,
withStencil ? GL_DEPTH_STENCIL_ATTACHMENT : GL_DEPTH_ATTACHMENT,
texTarget, tex, level);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoDepth(fb);
return;
}
fb.depthTex = tex;
fb.depthTarget = texTarget;
fb.depthLevel = level;
fb.depthHasStencil = withStencil;
}
void EnsureNoColorAttachment(ScratchFramebuffer& fb, GLenum fbTarget) {
PrepareForUse(fb, fbTarget);
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
RecordNoColor(fb);
}
}
void EnsureNoDepthAttachment(ScratchFramebuffer& fb, GLenum fbTarget) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
}
void EnsureReadBuffer(ScratchFramebuffer& fb, GLenum readBuffer) {
if (fb.readBuffer == readBuffer) {
return;
}
g_GLESFuncs.glReadBuffer(readBuffer);
fb.readBuffer = readBuffer;
}
void EnsureDrawBuffer(ScratchFramebuffer& fb, GLenum drawBuffer) {
if (fb.drawBuffer == drawBuffer) {
return;
}
g_GLESFuncs.glDrawBuffers(1, &drawBuffer);
fb.drawBuffer = drawBuffer;
}
Uint EnsureCompleteTinyFramebufferId() {
if (g_completeTinyFBOId != 0) {
return g_completeTinyFBOId;
}
// One-time creation: the renderbuffer binding is context state with no
// shadow, so save/restore it by query here (cold path only).
GLint prevRenderbuffer = 0;
g_GLESFuncs.glGetIntegerv(GL_RENDERBUFFER_BINDING, &prevRenderbuffer);
g_GLESFuncs.glGenFramebuffers(1, &g_completeTinyFBOId);
g_GLESFuncs.glGenRenderbuffers(1, &g_completeTinyRBOId);
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, g_completeTinyFBOId);
g_GLESFuncs.glBindRenderbuffer(GL_RENDERBUFFER, g_completeTinyRBOId);
g_GLESFuncs.glRenderbufferStorage(GL_RENDERBUFFER, GL_RGBA8, 1, 1);
g_GLESFuncs.glFramebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER,
g_completeTinyRBOId);
const GLenum drawBuffer = GL_COLOR_ATTACHMENT0;
g_GLESFuncs.glDrawBuffers(1, &drawBuffer);
g_GLESFuncs.glReadBuffer(GL_COLOR_ATTACHMENT0);
MOBILEGL_ASSERT(g_GLESFuncs.glCheckFramebufferStatus(GL_FRAMEBUFFER) == GL_FRAMEBUFFER_COMPLETE,
"Scratch 1x1 framebuffer is incomplete.");
g_GLESFuncs.glBindRenderbuffer(GL_RENDERBUFFER, static_cast<Uint>(prevRenderbuffer));
return g_completeTinyFBOId;
}
void NoteTextureIdDeleted(Uint textureId) {
if (textureId == 0) {
return;
}
for (ScratchFramebuffer* fb : {&g_tempFramebuffer, &g_blitReadFramebuffer, &g_blitDrawFramebuffer}) {
if (fb->colorTex == textureId || fb->depthTex == textureId) {
fb->attachmentsKnown = false;
}
}
}
void OnBackendContextDestroyed() {
g_tempFramebuffer = {};
g_blitReadFramebuffer = {};
g_blitDrawFramebuffer = {};
g_completeTinyFBOId = 0;
g_completeTinyRBOId = 0;
}
} // namespace ScratchFBOImpl
namespace PixelStoreImpl {
namespace {
PackState g_packState;
Bool g_packStateKnown = false;
void PinPackState(const PackState& value) {
g_GLESFuncs.glPixelStorei(GL_PACK_ALIGNMENT, value.Alignment);
g_GLESFuncs.glPixelStorei(GL_PACK_ROW_LENGTH, value.RowLength);
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_ROWS, value.SkipRows);
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_PIXELS, value.SkipPixels);
g_packState = value;
g_packStateKnown = true;
}
} // namespace
void ApplyPackState(const PackState& desired) {
if (!g_packStateKnown) {
PinPackState(desired);
return;
}
if (desired.Alignment != g_packState.Alignment) {
g_GLESFuncs.glPixelStorei(GL_PACK_ALIGNMENT, desired.Alignment);
g_packState.Alignment = desired.Alignment;
}
if (desired.RowLength != g_packState.RowLength) {
g_GLESFuncs.glPixelStorei(GL_PACK_ROW_LENGTH, desired.RowLength);
g_packState.RowLength = desired.RowLength;
}
if (desired.SkipRows != g_packState.SkipRows) {
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_ROWS, desired.SkipRows);
g_packState.SkipRows = desired.SkipRows;
}
if (desired.SkipPixels != g_packState.SkipPixels) {
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_PIXELS, desired.SkipPixels);
g_packState.SkipPixels = desired.SkipPixels;
}
}
PackState CurrentPackState() {
if (!g_packStateKnown) {
// Fresh/unknown context: pin to the GL defaults (what a new context
// starts with; writing them makes the shadow authoritative either way).
PinPackState(PackState{});
}
return g_packState;
}
void InvalidatePackStateCache() {
g_packStateKnown = false;
}
} // namespace PixelStoreImpl
namespace PrgramImpl { namespace PrgramImpl {
Uint32 g_snormFallbackClampOutputMask = 0; Uint32 g_snormFallbackClampOutputMask = 0;
Uint32 g_unormFallbackClampOutputMask = 0; Uint32 g_unormFallbackClampOutputMask = 0;
@@ -2784,6 +3301,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
effectiveSpirv = &uboPrecisionSpirv; effectiveSpirv = &uboPrecisionSpirv;
} }
// noperspective is core desktop GLSL and reaches here as the SPIR-V NoPerspective
// decoration. SPIRV-Cross renders it as ESSL `noperspective` + `#extension
// GL_NV_shader_noperspective_interpolation : require`; a driver without that extension
// rejects the require. So on such devices emulate screen-linear interpolation instead
// (pre-multiply outputs by gl_Position.w, recover inputs via gl_FragCoord.w) and drop
// the decoration - exact, extension-free. Devices that have the extension keep the
// decoration and let the hardware do it natively.
Vector<unsigned int> noperspectiveSpirv;
if (!g_GLESCapabilities.SupportsNoperspectiveInterpolation &&
MG_Util::ShaderTranspiler::ShaderCompiler::EmulateNoPerspectiveForEssl(
*effectiveSpirv, noperspectiveSpirv) &&
!noperspectiveSpirv.empty()) {
effectiveSpirv = &noperspectiveSpirv;
}
MG_Util::ShaderTranspiler::SpvcSession spvcSession(*effectiveSpirv, MG_Util::ShaderTranspiler::SpvcSession spvcSession(*effectiveSpirv,
MG_Util::ShaderTranspiler::SessionUsageBit::Transpile); MG_Util::ShaderTranspiler::SessionUsageBit::Transpile);
+122
View File
@@ -178,6 +178,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
// glBindBuffer with a redundant-bind cache for GL_ARRAY_BUFFER. // glBindBuffer with a redundant-bind cache for GL_ARRAY_BUFFER.
void BindBufferId(GLenum target, Uint id); void BindBufferId(GLenum target, Uint id);
void InvalidateArrayBufferBindingCache(); void InvalidateArrayBufferBindingCache();
// Redundant-bind caches for the driver-level GL_PIXEL_PACK/UNPACK_BUFFER
// bindings. Every backend readback (glReadPixels / pack-PBO map) and pixel
// upload site routes its binding through these so the shadow always matches
// the driver; the resting state between operations is 0, which keeps any
// path that implicitly assumes "no PBO bound" correct. Scrubbed when a
// buffer id is deleted/pooled (GL resets a deleted buffer's bindings to 0,
// and a recycled name matching the shadow would false-skip the rebind) and
// invalidated on MakeCurrent (context may reset).
void BindPixelPackBufferId(Uint id);
void BindPixelUnpackBufferId(Uint id);
void InvalidatePixelBufferBindingCaches();
// A GL buffer id is being deleted by code outside BufferImpl (e.g. the VAO
// client-attribute staging buffers): scrub every buffer-binding shadow that
// could false-skip when the name is recycled.
void NoteBufferIdDeleted(Uint id);
// Redundant-bind cache for INDEXED buffer bindings (glBindBufferBase/Range on // Redundant-bind cache for INDEXED buffer bindings (glBindBufferBase/Range on
// GL_UNIFORM_BUFFER / GL_SHADER_STORAGE_BUFFER): skips the GL call when the // GL_UNIFORM_BUFFER / GL_SHADER_STORAGE_BUFFER): skips the GL call when the
// (id, range) already at that index matches, like the array-buffer/texture/ // (id, range) already at that index matches, like the array-buffer/texture/
@@ -332,6 +347,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
class BackendTextureObject { class BackendTextureObject {
public: public:
BackendTextureObject(); BackendTextureObject();
// Deletes the GL texture (frontend glDeleteTextures used to leak every
// backend id for the context lifetime) and scrubs the binding/scratch-FBO
// shadows so a recycled name or heap address cannot false-skip a rebind.
~BackendTextureObject();
BackendTextureObject(const BackendTextureObject&) = delete;
BackendTextureObject& operator=(const BackendTextureObject&) = delete;
void SyncMipmapsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject); void SyncMipmapsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
void SyncBuiltinSamplerToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject); void SyncBuiltinSamplerToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
void SyncTextureParamsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject); void SyncTextureParamsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
@@ -343,6 +364,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
void RecreateBackendTexture(); void RecreateBackendTexture();
Uint m_backendTextureId = 0; Uint m_backendTextureId = 0;
// ES context generation the id was created under; a dtor running after
// that context died must not delete a foreign (recycled) name.
Uint m_contextGeneration = 0;
Bool m_isInitialized = false; Bool m_isInitialized = false;
Bool m_imageBindableStorageRequired = false; Bool m_imageBindableStorageRequired = false;
Bool m_backendStorageImmutable = false; Bool m_backendStorageImmutable = false;
@@ -367,6 +391,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS> MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS>
g_boundTexturesCache; g_boundTexturesCache;
extern Uint g_activeTextureUnit; extern Uint g_activeTextureUnit;
// Bumped when the backend ES context is destroyed; texture ids stamped with
// an older generation belong to a dead context and must not be deleted.
extern Uint g_textureContextGeneration;
} // namespace TextureImpl } // namespace TextureImpl
namespace FramebufferImpl { namespace FramebufferImpl {
@@ -375,6 +402,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
BackendFramebufferObject(); BackendFramebufferObject();
void SyncToBackend(const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject, void SyncToBackend(const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject,
FramebufferTarget asTarget); FramebufferTarget asTarget);
// Apply only this FBO's read buffer (glReadBuffer) to the backend. Split out so it can
// still run when SyncCurrentFBO skips the READ-target sync because the same GL FBO is
// bound as both draw and read (otherwise glReadBuffer changes would be silently dropped).
void SyncReadBufferToBackend(const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject);
void InvalidateSyncedState(); void InvalidateSyncedState();
Uint GetBackendFramebufferId() const { return m_backendFBOId; } Uint GetBackendFramebufferId() const { return m_backendFBOId; }
void Bind(FramebufferTarget target) const; void Bind(FramebufferTarget target) const;
@@ -414,8 +445,99 @@ namespace MobileGL::MG_Backend::DirectGLES {
extern Array<Uint16, SizeT(FramebufferTarget::FramebufferTargetCount)> g_fboSyncedObjectVersions; extern Array<Uint16, SizeT(FramebufferTarget::FramebufferTargetCount)> g_fboSyncedObjectVersions;
extern Array<MG_State::GLState::FramebufferObject*, SizeT(FramebufferTarget::FramebufferTargetCount)> extern Array<MG_State::GLState::FramebufferObject*, SizeT(FramebufferTarget::FramebufferTargetCount)>
g_fboSyncedObjects; g_fboSyncedObjects;
// Driver-level READ/DRAW framebuffer-binding shadow. Every backend
// glBindFramebuffer routes through BindFramebufferId so scoped helpers can
// save/restore the current binding without a glGetIntegerv round-trip (that
// query forces a driver pipeline sync) and so redundant rebinds no-op.
// Starts unknown; the first CurrentFramebufferBinding() query pins it from
// the driver once. Invalidated on MakeCurrent (context may reset).
// GL_FRAMEBUFFER binds both targets.
void BindFramebufferId(GLenum fbTarget, Uint id);
Uint CurrentFramebufferBinding(FramebufferTarget target);
void InvalidateFramebufferBindingCache();
} // namespace FramebufferImpl } // namespace FramebufferImpl
// Shared scratch framebuffers for the readback/copy/blit emulation paths, with a
// driver-side attachment shadow: repeated uses skip redundant detach/attach GL
// calls, and an attachment left by one use (e.g. a depth copy's DEPTH_STENCIL
// texture) is detached exactly when a later use of another aspect would
// otherwise inherit it (stale cross-aspect attachments made the shared temp FBO
// incomplete and silently degraded later readbacks).
namespace ScratchFBOImpl {
struct ScratchFramebuffer {
Uint id = 0;
// false => attachment state unknown; scrub every point on next use.
// A fresh FBO starts with nothing attached, so creation sets it true.
Bool attachmentsKnown = false;
Uint colorTex = 0;
GLenum colorTarget = 0;
GLint colorLevel = 0;
GLint colorLayer = -1; // >= 0 => attached via glFramebufferTextureLayer
Uint depthTex = 0;
GLenum depthTarget = 0;
GLint depthLevel = 0;
Bool depthHasStencil = false;
// Per-FBO read/draw buffer state (0 = unknown, set on first use).
GLenum readBuffer = 0;
GLenum drawBuffer = 0;
};
ScratchFramebuffer& TempFramebuffer(); // GetTexImage READ / CopyTex*Image2D depth DRAW
ScratchFramebuffer& BlitReadFramebuffer(); // texture-to-texture blit source
ScratchFramebuffer& BlitDrawFramebuffer(); // texture-to-texture blit destination
// Returns the GL id, generating it if needed (requires a current ES context).
Uint EnsureId(ScratchFramebuffer& fb);
// The fb must currently be bound at fbTarget (glReadBuffer/glDrawBuffers
// target the READ/DRAW binding respectively). Each Ensure* performs the
// minimal detach/attach set and keeps the shadow in sync; a failed attach
// records the point as detached so the completeness check fails instead of
// silently reading a stale attachment.
void EnsureColorAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level);
void EnsureColorAttachmentLayer(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLint level, GLint layer);
void EnsureDepthAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level,
Bool withStencil);
void EnsureNoColorAttachment(ScratchFramebuffer& fb, GLenum fbTarget);
void EnsureNoDepthAttachment(ScratchFramebuffer& fb, GLenum fbTarget);
void EnsureReadBuffer(ScratchFramebuffer& fb, GLenum readBuffer);
void EnsureDrawBuffer(ScratchFramebuffer& fb, GLenum drawBuffer);
// A 1x1 RGBA8-renderbuffer-complete FBO (GenerateMipmap needs a complete
// binding while respecifying texture storage). Attachment is set once at
// creation and never changes.
Uint EnsureCompleteTinyFramebufferId();
// A backend texture id is being deleted or respecified: a scratch FBO still
// referencing it would hold a dangling attachment (ES only auto-detaches
// from the *bound* framebuffer), and a recycled name could false-skip a
// re-attach; force a full scrub on next use.
void NoteTextureIdDeleted(Uint textureId);
// The ES context (and the scratch FBO ids with it) is going away.
void OnBackendContextDestroyed();
} // namespace ScratchFBOImpl
// Driver-level GL_PACK_* pixel-store shadow, the readback-side sibling of the
// upload path's ScopedDefaultUnpackState (Managers.cpp): the backend PACK state
// is written ONLY through ApplyPackState, so scoped helpers can save/restore it
// from the shadow instead of glGetIntegerv (which forces a driver pipeline
// sync), and redundant glPixelStorei calls no-op. The first Apply/Current call
// pins the driver to the shadow by writing all fields once. Invalidated on
// MakeCurrent (context may reset). PACK_IMAGE_HEIGHT/SKIP_IMAGES/SWAP_BYTES/
// LSB_FIRST have no ES equivalents; readbacks honor them on the CPU from the
// frontend context state instead.
namespace PixelStoreImpl {
struct PackState {
GLint Alignment = 4;
GLint RowLength = 0;
GLint SkipRows = 0;
GLint SkipPixels = 0;
Bool operator==(const PackState& o) const {
return Alignment == o.Alignment && RowLength == o.RowLength && SkipRows == o.SkipRows &&
SkipPixels == o.SkipPixels;
}
};
void ApplyPackState(const PackState& desired);
PackState CurrentPackState();
void InvalidatePackStateCache();
} // namespace PixelStoreImpl
// Image uniforms take their unit from the layout(binding=N) qualifier baked into // Image uniforms take their unit from the layout(binding=N) qualifier baked into
// the transpiled ESSL; unlike samplers they must not (and in ES cannot) be // the transpiled ESSL; unlike samplers they must not (and in ES cannot) be
// assigned through glUniform1i. // assigned through glUniform1i.
@@ -397,8 +397,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
if (!handle.Handle || (handle.Backend != WindowBackend::Android && if (!handle.Handle || (handle.Backend != WindowBackend::Android &&
handle.Backend != WindowBackend::X11 && handle.Backend != WindowBackend::X11 &&
handle.Backend != WindowBackend::MetalLayer)) { handle.Backend != WindowBackend::MetalLayer &&
MGLOG_E("DirectVulkan backend only supports Android, X11, and CAMetalLayer native windows"); handle.Backend != WindowBackend::Win32)) {
MGLOG_E("DirectVulkan backend only supports Android, X11, CAMetalLayer, and Win32 native windows");
return false; return false;
} }
@@ -794,5 +795,32 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_vulkanCaps.MaxShaderStorageBlockSize, m_vulkanCaps.MaxShaderStorageBlockSize,
m_dynamicParameters.MaxShaderStorageBlockSize); m_dynamicParameters.MaxShaderStorageBlockSize);
} }
switch (m_vulkanCaps.VendorId) {
case 0x5143u: // VK_VENDOR_ID: Qualcomm
m_dynamicParameters.GpuVendor = GpuVendorKind::Qualcomm;
break;
case 0x13B5u: // ARM
m_dynamicParameters.GpuVendor = GpuVendorKind::Arm;
break;
case 0x10DEu: // NVIDIA
m_dynamicParameters.GpuVendor = GpuVendorKind::Nvidia;
break;
case 0x1002u: // AMD
m_dynamicParameters.GpuVendor = GpuVendorKind::Amd;
break;
case 0x8086u: // Intel
m_dynamicParameters.GpuVendor = GpuVendorKind::Intel;
break;
case 0x1010u: // Imagination
m_dynamicParameters.GpuVendor = GpuVendorKind::ImgTec;
break;
case 0x10005u: // Mesa software (lavapipe)
case 0x1AE0u: // Google (SwiftShader)
m_dynamicParameters.GpuVendor = GpuVendorKind::Software;
break;
default:
m_dynamicParameters.GpuVendor = GpuVendorKind::Unknown;
break;
}
} }
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -20,7 +20,8 @@
#include <spirv_reflect.h> #include <spirv_reflect.h>
namespace MobileGL::MG_Backend::DirectVulkan { namespace MobileGL::MG_Backend::DirectVulkan {
UniquePtr<VulkanRenderer> pVulkanRenderer = nullptr; // Leak-at-exit storage; see GlobalObjects.cpp.
UniquePtr<VulkanRenderer>& pVulkanRenderer = *new UniquePtr<VulkanRenderer>();
namespace { namespace {
// Generation of the live VulkanRenderer instance, mirroring // Generation of the live VulkanRenderer instance, mirroring
@@ -12,7 +12,7 @@
#include "Renderer/VulkanRenderer.h" #include "Renderer/VulkanRenderer.h"
namespace MobileGL::MG_Backend::DirectVulkan { namespace MobileGL::MG_Backend::DirectVulkan {
extern UniquePtr<VulkanRenderer> pVulkanRenderer; extern UniquePtr<VulkanRenderer>& pVulkanRenderer;
// Generation of the live VulkanRenderer instance, mirroring DirectGLES's // Generation of the live VulkanRenderer instance, mirroring DirectGLES's
// g_syncContextGeneration. BackendObject_DirectVulkan bumps it wherever // g_syncContextGeneration. BackendObject_DirectVulkan bumps it wherever
@@ -8,6 +8,7 @@
#include "PipelineFactory.h" #include "PipelineFactory.h"
namespace MobileGL::MG_Backend::DirectVulkan { namespace MobileGL::MG_Backend::DirectVulkan {
static const char* PrimitiveTopologyToString(VkPrimitiveTopology topology) { static const char* PrimitiveTopologyToString(VkPrimitiveTopology topology) {
switch (topology) { switch (topology) {
@@ -108,6 +109,81 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"vkCreatePipelineCache"); "vkCreatePipelineCache");
} }
// Must be called once, before any pipeline is created: the flag is not part of the
// pipeline hash, so flipping it mid-life would serve cached pipelines built under the
// old value.
void PipelineFactory::SetSuppressBlendedDepthWrite(Bool enabled) {
s_suppressBlendedDepthWrite = enabled;
}
Bool PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(MG_Config::QuirkOverride quirkOverride,
Uint32 vendorId) {
static constexpr Uint32 kVendorIdQualcomm = 0x5143;
switch (quirkOverride) {
case MG_Config::QuirkOverride::ForceOn:
return true;
case MG_Config::QuirkOverride::ForceOff:
return false;
case MG_Config::QuirkOverride::Auto:
default:
return vendorId == kVendorIdQualcomm;
}
}
namespace {
// MIN/MAX extremum blending: the signature of a depth-bounds accumulation pass
// (MC 26.3 OIT writes vec4(-linD, linD, deviceZ, 0) under GL_MAX while writing
// depth for its equality chain). MIN/MAX ignore blend factors per the Vulkan spec.
//
// Deliberately the ONLY shape stripped. A quirk should touch as little unrelated
// content as possible, and a trace sweep of every fixture showed the wider
// alternatives all cost more than they fix:
// - additive ONE+ONE with a depth write matched zero draws of the 26.3 chain
// (its transmittance/accumulate passes disable depth writes themselves) - the
// only real content it caught was harmless additive glow effects (Create);
// - sorted-transparency "over" blends (SRC_ALPHA-style) are order-dependent,
// drawn once per surface, and rely on their depth writes for occlusion;
// - separate-alpha accumulation over an over-blending color channel has no
// known pairing with a depth-equality chain (color channel only, see tests).
// If a future workload pairs another blend shape with an equality chain, widen
// this with that evidence in hand rather than pre-emptively.
Bool IsAccumulationBlend(const VkPipelineColorBlendAttachmentState& attachment) {
return attachment.colorBlendOp == VK_BLEND_OP_MIN ||
attachment.colorBlendOp == VK_BLEND_OP_MAX;
}
} // namespace
Bool PipelineFactory::ShouldSuppressDepthWrite(const PipelineCreatePayload& payload) {
if (!payload.depthWriteEnable) {
return false;
}
// A shader that assigns gl_FragDepth supplies depth itself rather than taking the
// pipeline's interpolated Z, so a driver that varies the vertex position math
// between pipelines cannot desynchronize it. (A gl_FragDepth = gl_FragCoord.z
// passthrough is the exception that stays exposed; no known content pairs one with
// an equality chain, and 26.3's composite is a genuine computed-depth writer.)
if (payload.fragmentReplacesDepth) {
return false;
}
for (Uint32 i = 0; i < payload.colorAttachmentCount; ++i) {
const VkPipelineColorBlendAttachmentState& attachment = payload.colorBlendAttachments[i];
if (attachment.blendEnable != VK_TRUE) {
continue;
}
// All color writes masked: blending is moot (depth-prepass pattern that left
// GL_BLEND enabled); stripping the depth write would delete the whole prepass.
if (attachment.colorWriteMask == 0) {
continue;
}
// Any attachment qualifies, not just attachment 0: the 26.3 transmittance pass
// accumulates into a 2-target MRT and must stay stripped.
if (IsAccumulationBlend(attachment)) {
return true;
}
}
return false;
}
PipelineFactory::~PipelineFactory() { PipelineFactory::~PipelineFactory() {
DestroyAll(); DestroyAll();
if (m_pipelineCache != VK_NULL_HANDLE) { if (m_pipelineCache != VK_NULL_HANDLE) {
@@ -152,6 +228,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
XXH64_update(m_hashState, &payload.backStencilDepthFailOp, sizeof(payload.backStencilDepthFailOp))); XXH64_update(m_hashState, &payload.backStencilDepthFailOp, sizeof(payload.backStencilDepthFailOp)));
XXHASH_VERIFY( XXHASH_VERIFY(
XXH64_update(m_hashState, &payload.backStencilCompareOp, sizeof(payload.backStencilCompareOp))); XXH64_update(m_hashState, &payload.backStencilCompareOp, sizeof(payload.backStencilCompareOp)));
XXHASH_VERIFY(
XXH64_update(m_hashState, &payload.fragmentReplacesDepth, sizeof(payload.fragmentReplacesDepth)));
if (payload.colorAttachmentCount > 0) { if (payload.colorAttachmentCount > 0) {
XXHASH_VERIFY(XXH64_update( XXHASH_VERIFY(XXH64_update(
m_hashState, m_hashState,
@@ -258,6 +336,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
for (Uint32 i = 0; i < payload.colorAttachmentCount; ++i) { for (Uint32 i = 0; i < payload.colorAttachmentCount; ++i) {
colorAttachments[i] = payload.colorBlendAttachments[i]; colorAttachments[i] = payload.colorBlendAttachments[i];
} }
// Suppress depth writes on accumulation-blended pipelines when the active driver
// cannot keep vertex positions invariant across the pipelines of a multi-pass
// depth-equality chain (see SetSuppressBlendedDepthWrite). The decision is narrowed
// in ShouldSuppressDepthWrite: sorted-transparency "over" blends (vanilla MC water),
// gl_FragDepth writers, and masked-out attachments keep their depth writes.
// This bakes the decision into the pipeline, which only works because depth write is
// static state here - adding VK_DYNAMIC_STATE_DEPTH_WRITE_ENABLE to kDynamicStates
// would let the record-time value override it and silently disable the quirk.
if (s_suppressBlendedDepthWrite && ShouldSuppressDepthWrite(payload)) {
depthStencil.depthWriteEnable = VK_FALSE;
}
VkPipelineColorBlendStateCreateInfo blend{VK_STRUCTURE_TYPE_PIPELINE_COLOR_BLEND_STATE_CREATE_INFO}; VkPipelineColorBlendStateCreateInfo blend{VK_STRUCTURE_TYPE_PIPELINE_COLOR_BLEND_STATE_CREATE_INFO};
blend.logicOpEnable = payload.logicOpEnable ? VK_TRUE : VK_FALSE; blend.logicOpEnable = payload.logicOpEnable ? VK_TRUE : VK_FALSE;
blend.logicOp = payload.logicOp; blend.logicOp = payload.logicOp;
@@ -49,6 +49,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkStencilOp backStencilPassOp = VK_STENCIL_OP_KEEP; VkStencilOp backStencilPassOp = VK_STENCIL_OP_KEEP;
VkStencilOp backStencilDepthFailOp = VK_STENCIL_OP_KEEP; VkStencilOp backStencilDepthFailOp = VK_STENCIL_OP_KEEP;
VkCompareOp backStencilCompareOp = VK_COMPARE_OP_ALWAYS; VkCompareOp backStencilCompareOp = VK_COMPARE_OP_ALWAYS;
// The fragment module writes gl_FragDepth (SPIR-V DepthReplacing); exempts the
// pipeline from the blended depth-write quirk (see ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false;
Array<VkPipelineColorBlendAttachmentState, kMaxColorAttachments> colorBlendAttachments{}; Array<VkPipelineColorBlendAttachmentState, kMaxColorAttachments> colorBlendAttachments{};
const Vector<VkPipelineShaderStageCreateInfo>* stages = nullptr; const Vector<VkPipelineShaderStageCreateInfo>* stages = nullptr;
const VkPipelineVertexInputStateCreateInfo* vertexInputState = nullptr; const VkPipelineVertexInputStateCreateInfo* vertexInputState = nullptr;
@@ -62,6 +65,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipeline GetOrCreatePipeline(const PipelineCreatePayload& payload); VkPipeline GetOrCreatePipeline(const PipelineCreatePayload& payload);
void DestroyAll(); void DestroyAll();
// Driver quirk: suppress depth writes on accumulation-blended pipelines. Multi-pass
// depth-equality rendering (a blended prepass writes depth that later passes re-test
// with an equality-inclusive compare on the re-rasterized geometry) requires
// cross-pipeline position invariance that some mobile compilers do not provide, even
// with the SPIR-V Invariant decoration; whole primitives then drop out of the later
// passes. Only MIN/MAX extremum blends are stripped - the signature of such a
// chain's depth-bounds pass (MC 26.3 OIT), and per a fixture-wide trace sweep the
// only depth-writing shape the chain actually uses - so every other blend
// (sorted-transparency "over" like vanilla MC water, additive glows, ...) keeps
// its depth writes. Set at renderer initialization based on the active driver.
static void SetSuppressBlendedDepthWrite(Bool enabled);
static Bool IsSuppressBlendedDepthWriteEnabled() { return s_suppressBlendedDepthWrite; }
// Device gate for the quirk: ForceOn/ForceOff bypass detection, Auto enables it on
// the known-affected vendor (Qualcomm).
static Bool ShouldSuppressBlendedDepthWriteForDevice(MG_Config::QuirkOverride quirkOverride,
Uint32 vendorId);
// Pure per-pipeline strip decision (exempts gl_FragDepth writers, masked-out and
// non-accumulation blends); combined with the device flag in CreatePipeline. Static
// and payload-only so tests can pin the contract without a VkDevice.
static Bool ShouldSuppressDepthWrite(const PipelineCreatePayload& payload);
private: private:
VkPipeline CreatePipeline(const PipelineCreatePayload& payload) const; VkPipeline CreatePipeline(const PipelineCreatePayload& payload) const;
@@ -70,5 +94,6 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipelineCache m_pipelineCache = VK_NULL_HANDLE; VkPipelineCache m_pipelineCache = VK_NULL_HANDLE;
UnorderedMap<HashType, VkPipeline> m_cache; UnorderedMap<HashType, VkPipeline> m_cache;
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
static inline Bool s_suppressBlendedDepthWrite = false;
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -312,34 +312,6 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return targetEnv; return targetEnv;
} }
// Cheap raw-word scan for an `OpDecorate <id> BuiltIn InstanceIndex` decoration. Used
// only to decide whether to warn when shaderDrawParameters is unavailable; a false
// negative merely suppresses a diagnostic.
Bool SpirvDeclaresInstanceIndexBuiltin(const Vector<Uint>& spirv) {
constexpr Uint32 kSpirvMagicNumber = 0x07230203u;
constexpr SizeT kHeaderWordCount = 5;
if (spirv.size() <= kHeaderWordCount || spirv[0] != kSpirvMagicNumber) {
return false;
}
SizeT wordIndex = kHeaderWordCount;
while (wordIndex < spirv.size()) {
const Uint32 firstWord = spirv[wordIndex];
const Uint32 wordCount = firstWord >> 16;
const auto opcode = static_cast<spv::Op>(firstWord & 0xffffu);
if (wordCount == 0 || wordIndex + wordCount > spirv.size()) {
break;
}
if (opcode == spv::Op::OpDecorate && wordCount >= 4 &&
static_cast<spv::Decoration>(spirv[wordIndex + 2]) == spv::Decoration::BuiltIn &&
static_cast<spv::BuiltIn>(spirv[wordIndex + 3]) == spv::BuiltIn::InstanceIndex) {
return true;
}
wordIndex += wordCount;
}
return false;
}
Bool IsInterfaceVariableStaticallyUsed(const Vector<Uint>& spirv, Uint32 spirvId) { Bool IsInterfaceVariableStaticallyUsed(const Vector<Uint>& spirv, Uint32 spirvId) {
if (spirv.empty() || spirvId == 0) { if (spirv.empty() || spirvId == 0) {
return false; return false;
@@ -1218,6 +1190,41 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} // namespace } // namespace
// A shader that assigns gl_FragDepth (SPIR-V DepthReplacing) supplies depth itself
// instead of taking the pipeline's interpolated Z, so a driver that varies the vertex
// position math between pipelines cannot desynchronize it; the blended depth-write
// quirk therefore leaves it alone (see PipelineFactory::ShouldSuppressDepthWrite).
Bool ProgramFactory::ReflectedFragmentReplacesDepth(const SpvReflectShaderModule& reflectModule) {
for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) {
const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex];
for (Uint32 modeIndex = 0; modeIndex < entryPoint.execution_mode_count; ++modeIndex) {
if (entryPoint.execution_modes[modeIndex] == SpvExecutionModeDepthReplacing) {
return true;
}
}
}
return false;
}
// glslang's relaxed-Vulkan mode maps GL's gl_InstanceID onto the InstanceIndex builtin.
// Without shaderDrawParameters there is no gl_BaseInstance to subtract, so such a shader
// cannot be corrected and instanced draws with a non-zero baseInstance misrender; this
// detects the case so the user gets one warning instead of silent corruption.
Bool ProgramFactory::ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule) {
for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) {
const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex];
for (Uint32 variableIndex = 0; variableIndex < entryPoint.input_variable_count; ++variableIndex) {
const SpvReflectInterfaceVariable* variable = entryPoint.input_variables[variableIndex];
if (variable != nullptr &&
(variable->decoration_flags & SPV_REFLECT_DECORATION_BUILT_IN) != 0 &&
variable->built_in == SpvBuiltInInstanceIndex) {
return true;
}
}
}
return false;
}
VkShaderStageFlagBits ProgramFactory::ToVkStage(ShaderStage stage) { VkShaderStageFlagBits ProgramFactory::ToVkStage(ShaderStage stage) {
switch (stage) { switch (stage) {
case ShaderStage::Vertex: case ShaderStage::Vertex:
@@ -1465,6 +1472,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
continue; continue;
} }
if (!m_shaderDrawParametersEnabled && ReflectedReadsInstanceIndexBuiltin(reflectModule)) {
static Bool s_warnedInstanceIndexUnsupported = false;
if (!s_warnedInstanceIndexUnsupported) {
s_warnedInstanceIndexUnsupported = true;
MGLOG_W("ProgramFactory: shaderDrawParameters is unavailable; gl_InstanceID cannot be "
"rebased and instanced draws with a non-zero baseInstance may render incorrectly");
}
}
uint32_t inputCount = 0; uint32_t inputCount = 0;
SpvReflectResult reflectResult = spvReflectEnumerateInputVariables(&reflectModule, &inputCount, nullptr); SpvReflectResult reflectResult = spvReflectEnumerateInputVariables(&reflectModule, &inputCount, nullptr);
MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS, MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS,
@@ -1512,6 +1528,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkProgramObject& entry) const { VkProgramObject& entry) const {
entry.activeFragmentOutputLocationMask = 0; entry.activeFragmentOutputLocationMask = 0;
entry.fragmentOutputTypes.fill(0); entry.fragmentOutputTypes.fill(0);
entry.fragmentReplacesDepth = false;
for (SizeT moduleIndex = 0; moduleIndex < shaders.size() && moduleIndex < spirv.size(); ++moduleIndex) { for (SizeT moduleIndex = 0; moduleIndex < shaders.size() && moduleIndex < spirv.size(); ++moduleIndex) {
if (!shaders[moduleIndex] || shaders[moduleIndex]->GetShaderStage() != ShaderStage::Fragment) { if (!shaders[moduleIndex] || shaders[moduleIndex]->GetShaderStage() != ShaderStage::Fragment) {
@@ -1530,9 +1547,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"ProgramFactory::ReflectFragmentOutputs: failed to create reflection module (result=%d)", "ProgramFactory::ReflectFragmentOutputs: failed to create reflection module (result=%d)",
static_cast<Int>(createResult)); static_cast<Int>(createResult));
if (createResult != SPV_REFLECT_RESULT_SUCCESS) { if (createResult != SPV_REFLECT_RESULT_SUCCESS) {
// Fail toward the exemption: stripping a genuine gl_FragDepth writer would
// corrupt its depth output outright, while wrongly exempting an accumulation
// pass merely reverts that one program to the pre-quirk behavior.
entry.fragmentReplacesDepth = true;
continue; continue;
} }
entry.fragmentReplacesDepth = ReflectedFragmentReplacesDepth(reflectModule);
uint32_t outputCount = 0; uint32_t outputCount = 0;
SpvReflectResult reflectResult = spvReflectEnumerateOutputVariables(&reflectModule, &outputCount, nullptr); SpvReflectResult reflectResult = spvReflectEnumerateOutputVariables(&reflectModule, &outputCount, nullptr);
MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS, MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS,
@@ -1808,6 +1831,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_BUFFER; layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_BUFFER;
} else if (kind == DescriptorBindingKind::StorageImage) { } else if (kind == DescriptorBindingKind::StorageImage) {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_IMAGE; layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_IMAGE;
entry.hasStorageImages = true;
} else { } else {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_COMBINED_IMAGE_SAMPLER; layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_COMBINED_IMAGE_SAMPLER;
} }
@@ -1862,28 +1886,43 @@ namespace MobileGL::MG_Backend::DirectVulkan {
moduleSpirvs[i] = spv; moduleSpirvs[i] = spv;
} }
// GL apps depend on cross-program position invariance for multi-pass equality
// depth tests (MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the
// depth its own first pass wrote); decorate Position outputs Invariant so
// per-pipeline compilers cannot vary the position math between passes.
{
Vector<Uint> invariantSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::DecoratePositionInvariantForVulkan(
moduleSpirvs[i], invariantSpirv)) {
moduleSpirvs[i] = std::move(invariantSpirv);
} else {
// The pass round-trips through SPIRV-Tools IR, so an unparseable module
// fails open and keeps the undecorated words - which silently reinstates
// the multi-pass invariance bug rather than breaking anything loudly.
MGLOG_E("ProgramFactory: position-invariant decoration failed for program %u; "
"keeping the original module - multi-pass depth-equality chains "
"(e.g. MC 26.3 OIT clouds) may drop primitives on this device",
program.GetExternalIndex());
}
}
// glslang's relaxed-Vulkan mode aliases GL's zero-based gl_InstanceID to Vulkan's // glslang's relaxed-Vulkan mode aliases GL's zero-based gl_InstanceID to Vulkan's
// gl_InstanceIndex, which wrongly includes the draw's baseInstance. Rebase vertex-stage // gl_InstanceIndex, which wrongly includes the draw's baseInstance. Rebase vertex-stage
// loads to (InstanceIndex - BaseInstance) so shaders observe GL semantics. Reflection // loads to (InstanceIndex - BaseInstance) so shaders observe GL semantics. Reflection
// below runs on the rebased words so the added BaseInstance builtin stays consistent. // below runs on the rebased words so the added BaseInstance builtin stays consistent.
if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Vertex) { // The unsupported-device counterpart of this rebase (warning when a shader reads
if (m_shaderDrawParametersEnabled) { // the builtin but shaderDrawParameters is missing) rides along with
Vector<Uint> rebasedSpirv; // ReflectVertexInputs, which already reflects this stage.
if (MG_Util::ShaderTranspiler::ShaderCompiler::RebaseInstanceIndexForVulkan(moduleSpirvs[i], if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Vertex &&
rebasedSpirv)) { m_shaderDrawParametersEnabled) {
moduleSpirvs[i] = std::move(rebasedSpirv); Vector<Uint> rebasedSpirv;
} else { if (MG_Util::ShaderTranspiler::ShaderCompiler::RebaseInstanceIndexForVulkan(moduleSpirvs[i],
MGLOG_E("ProgramFactory: failed to rebase gl_InstanceID for program %u; " rebasedSpirv)) {
"instanced draws with a non-zero baseInstance may render incorrectly", moduleSpirvs[i] = std::move(rebasedSpirv);
program.GetExternalIndex()); } else {
} MGLOG_E("ProgramFactory: failed to rebase gl_InstanceID for program %u; "
} else if (SpirvDeclaresInstanceIndexBuiltin(moduleSpirvs[i])) { "instanced draws with a non-zero baseInstance may render incorrectly",
static Bool s_warnedInstanceIndexUnsupported = false; program.GetExternalIndex());
if (!s_warnedInstanceIndexUnsupported) {
s_warnedInstanceIndexUnsupported = true;
MGLOG_W("ProgramFactory: shaderDrawParameters is unavailable; gl_InstanceID cannot be "
"rebased and instanced draws with a non-zero baseInstance may render incorrectly");
}
} }
} }
@@ -67,6 +67,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Vector<Bool> storageImageUsesBindingFormatByBinding; Vector<Bool> storageImageUsesBindingFormatByBinding;
Vector<String> storageBlockNameByBinding; Vector<String> storageBlockNameByBinding;
Vector<Int> storageBlockIndexByBinding; Vector<Int> storageBlockIndexByBinding;
// Set once during ReflectLayout so the per-draw path can skip the whole
// storage-image preparation for the overwhelming majority of programs.
Bool hasStorageImages = false;
Int globalUboBinding = -1; Int globalUboBinding = -1;
Uint32 activeVertexInputLocationMask = 0; Uint32 activeVertexInputLocationMask = 0;
Array<GLenum, kMaxVertexInputLocations> vertexInputTypes{}; Array<GLenum, kMaxVertexInputLocations> vertexInputTypes{};
@@ -75,6 +78,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ShaderStage rasterizationProducerStage = ShaderStage::Unknown; ShaderStage rasterizationProducerStage = ShaderStage::Unknown;
Uint32 producerOutputComponentCount = 0; Uint32 producerOutputComponentCount = 0;
Uint32 fragmentInputComponentCount = 0; Uint32 fragmentInputComponentCount = 0;
// The fragment module declares the DepthReplacing execution mode (writes
// gl_FragDepth); shader-computed depth is immune to the cross-pipeline
// position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false;
static inline VkDevice s_device = VK_NULL_HANDLE; static inline VkDevice s_device = VK_NULL_HANDLE;
@@ -99,6 +106,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
std::move(other.storageImageUsesBindingFormatByBinding); std::move(other.storageImageUsesBindingFormatByBinding);
storageBlockNameByBinding = std::move(other.storageBlockNameByBinding); storageBlockNameByBinding = std::move(other.storageBlockNameByBinding);
storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding); storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding);
hasStorageImages = other.hasStorageImages;
globalUboBinding = other.globalUboBinding; globalUboBinding = other.globalUboBinding;
activeVertexInputLocationMask = other.activeVertexInputLocationMask; activeVertexInputLocationMask = other.activeVertexInputLocationMask;
vertexInputTypes = other.vertexInputTypes; vertexInputTypes = other.vertexInputTypes;
@@ -107,15 +115,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
rasterizationProducerStage = other.rasterizationProducerStage; rasterizationProducerStage = other.rasterizationProducerStage;
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE; other.pipelineLayout = VK_NULL_HANDLE;
other.hasStorageImages = false;
other.globalUboBinding = -1; other.globalUboBinding = -1;
other.activeVertexInputLocationMask = 0; other.activeVertexInputLocationMask = 0;
other.activeFragmentOutputLocationMask = 0; other.activeFragmentOutputLocationMask = 0;
other.rasterizationProducerStage = ShaderStage::Unknown; other.rasterizationProducerStage = ShaderStage::Unknown;
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false;
} }
VkProgramObject& operator=(VkProgramObject&& other) noexcept { VkProgramObject& operator=(VkProgramObject&& other) noexcept {
if (this == &other) { if (this == &other) {
@@ -139,6 +150,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
std::move(other.storageImageUsesBindingFormatByBinding); std::move(other.storageImageUsesBindingFormatByBinding);
storageBlockNameByBinding = std::move(other.storageBlockNameByBinding); storageBlockNameByBinding = std::move(other.storageBlockNameByBinding);
storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding); storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding);
hasStorageImages = other.hasStorageImages;
globalUboBinding = other.globalUboBinding; globalUboBinding = other.globalUboBinding;
activeVertexInputLocationMask = other.activeVertexInputLocationMask; activeVertexInputLocationMask = other.activeVertexInputLocationMask;
vertexInputTypes = other.vertexInputTypes; vertexInputTypes = other.vertexInputTypes;
@@ -147,15 +159,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
rasterizationProducerStage = other.rasterizationProducerStage; rasterizationProducerStage = other.rasterizationProducerStage;
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE; other.pipelineLayout = VK_NULL_HANDLE;
other.hasStorageImages = false;
other.globalUboBinding = -1; other.globalUboBinding = -1;
other.activeVertexInputLocationMask = 0; other.activeVertexInputLocationMask = 0;
other.activeFragmentOutputLocationMask = 0; other.activeFragmentOutputLocationMask = 0;
other.rasterizationProducerStage = ShaderStage::Unknown; other.rasterizationProducerStage = ShaderStage::Unknown;
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false;
return *this; return *this;
} }
@@ -203,6 +218,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static VkShaderStageFlagBits ToVkStage(ShaderStage stage); static VkShaderStageFlagBits ToVkStage(ShaderStage stage);
static VkFormat ConvertSpirvImageFormatToVkFormat(SpvImageFormat format); static VkFormat ConvertSpirvImageFormatToVkFormat(SpvImageFormat format);
static SamplerNumericDomain UniformTypeToSamplerNumericDomain(GLenum glType); static SamplerNumericDomain UniformTypeToSamplerNumericDomain(GLenum glType);
// True when any entry point declares the DepthReplacing execution mode, i.e. the
// shader assigns gl_FragDepth. Exposed so the blended depth-write quirk's exemption
// can be pinned by tests. A false negative loses the exemption, so such a shader is
// stripped conservatively and forfeits its depth write.
static Bool ReflectedFragmentReplacesDepth(const SpvReflectShaderModule& reflectModule);
// True when an entry point reads the InstanceIndex builtin. Only gates a diagnostic:
// without shaderDrawParameters such a shader cannot have gl_InstanceID rebased.
static Bool ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule);
private: private:
struct ProgramLookupCache { struct ProgramLookupCache {
@@ -298,8 +298,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// descriptor valid. // descriptor valid.
const Bool forceNearestFiltering = numericDomain == SamplerNumericDomain::SignedInteger || const Bool forceNearestFiltering = numericDomain == SamplerNumericDomain::SignedInteger ||
numericDomain == SamplerNumericDomain::UnsignedInteger; numericDomain == SamplerNumericDomain::UnsignedInteger;
const VkFormat sampledViewFormat = SamplerResolveMemo* viewFormatMemo =
VkTextureManager::ResolveSampledImageViewFormat(resource->format, numericDomain); binding < m_samplerResolveMemo.size() ? &m_samplerResolveMemo[binding] : nullptr;
VkFormat sampledViewFormat;
if (viewFormatMemo != nullptr && viewFormatMemo->viewFormatValid &&
viewFormatMemo->viewFormatSource == resource->format &&
viewFormatMemo->viewFormatDomain == numericDomain) {
sampledViewFormat = viewFormatMemo->viewFormat;
} else {
sampledViewFormat =
VkTextureManager::ResolveSampledImageViewFormat(resource->format, numericDomain);
if (viewFormatMemo != nullptr) {
viewFormatMemo->viewFormatSource = resource->format;
viewFormatMemo->viewFormatDomain = numericDomain;
viewFormatMemo->viewFormat = sampledViewFormat;
viewFormatMemo->viewFormatValid = true;
}
}
if (sampledViewFormat == VK_FORMAT_UNDEFINED) { if (sampledViewFormat == VK_FORMAT_UNDEFINED) {
MGLOG_E("ResolveSamplerDescriptor: no compatible sampled view for binding=%u ('%s') " MGLOG_E("ResolveSamplerDescriptor: no compatible sampled view for binding=%u ('%s') "
"textureId=%d imageFormat=%d numericDomain=%d", "textureId=%d imageFormat=%d numericDomain=%d",
@@ -307,8 +322,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static_cast<Int>(resource->format), static_cast<Int>(numericDomain)); static_cast<Int>(resource->format), static_cast<Int>(numericDomain));
return false; return false;
} }
// No reinterpretation requested: bind the depth-or-color aspect view the sync above
// already produced instead of re-entering GetOrCreateSampledImageView's sync path.
const VkImageView sampledImageView = const VkImageView sampledImageView =
m_textureManager->GetOrCreateSampledImageView(*texture, sampledViewFormat); sampledViewFormat == resource->format
? resource->sampledView
: m_textureManager->GetOrCreateSampledImageView(*texture, sampledViewFormat);
if (sampledImageView == VK_NULL_HANDLE) { if (sampledImageView == VK_NULL_HANDLE) {
MGLOG_E("ResolveSamplerDescriptor: failed to resolve sampled view for binding=%u ('%s') " MGLOG_E("ResolveSamplerDescriptor: failed to resolve sampled view for binding=%u ('%s') "
"textureId=%d imageFormat=%d viewFormat=%d numericDomain=%d", "textureId=%d imageFormat=%d viewFormat=%d numericDomain=%d",
@@ -176,6 +176,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint16 textureParamsVersion = 0; Uint16 textureParamsVersion = 0;
Bool forceNearestFiltering = false; Bool forceNearestFiltering = false;
Bool valid = false; Bool valid = false;
// ResolveSampledImageViewFormat is pure in (image format, numeric domain), but a
// domain mismatch walks a ~184-entry format table. Memo the resolution per binding
// so a reinterpreted sampler pays that scan once, not once per draw.
VkFormat viewFormatSource = VK_FORMAT_UNDEFINED;
SamplerNumericDomain viewFormatDomain = SamplerNumericDomain::Unknown;
VkFormat viewFormat = VK_FORMAT_UNDEFINED;
Bool viewFormatValid = false;
}; };
mutable Vector<SamplerResolveMemo> m_samplerResolveMemo; mutable Vector<SamplerResolveMemo> m_samplerResolveMemo;
}; };
@@ -126,8 +126,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const Bool packedAttribute = attr.Type == DataType::Int2101010Rev || const Bool packedAttribute = attr.Type == DataType::Int2101010Rev ||
attr.Type == DataType::Uint2101010Rev; attr.Type == DataType::Uint2101010Rev;
const SizeT requiredAlignment = packedAttribute ? attribByteSize : GetComponentSize(attr.Type); const SizeT requiredAlignment = packedAttribute ? attribByteSize : GetComponentSize(attr.Type);
// For a client-memory array attr.Offset holds the raw client pointer, and the
// draw path re-uploads the data to a 16-aligned transient slice with attribute
// offset 0, so only the stride can violate Vulkan's fetch alignment there.
const Bool clientMemoryAttribute = attr.Buffer == nullptr;
if (conversion == VertexStreamConversion::None && requiredAlignment > 1 && if (conversion == VertexStreamConversion::None && requiredAlignment > 1 &&
((sourceStride % requiredAlignment) != 0 || (attr.Offset % requiredAlignment) != 0)) { ((sourceStride % requiredAlignment) != 0 ||
(!clientMemoryAttribute && (attr.Offset % requiredAlignment) != 0))) {
// GL accepts arbitrary byte strides and offsets. Core Vulkan vertex fetches do not // GL accepts arbitrary byte strides and offsets. Core Vulkan vertex fetches do not
// unless VK_EXT_legacy_vertex_attributes is available, so deinterleave this one // unless VK_EXT_legacy_vertex_attributes is available, so deinterleave this one
// attribute into a tightly packed transient stream without changing its format. // attribute into a tightly packed transient stream without changing its format.
@@ -1140,6 +1140,24 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return ok; return ok;
} }
Bool VkTextureManager::NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const {
const auto it = m_textureResources.find(MakeTextureIdentity(&texture));
if (it == m_textureResources.end()) {
return true;
}
const TextureResource& resource = it->second;
if (resource.image == VK_NULL_HANDLE || resource.layout != VK_IMAGE_LAYOUT_GENERAL) {
return true;
}
// Mirror SyncTexture's cross-draw skip condition: any version drift means the sync
// path may upload or rebuild, both of which need the render pass ended first.
const auto* mipTexture = MG_State::GLState::AsMipmapTexture(&texture);
const Uint32 mipLevelCount = mipTexture != nullptr ? mipTexture->GetMipmapLevelCount() : 0u;
return resource.syncedContentVersion != texture.GetContentVersion() ||
resource.syncedTextureParamsVersion != texture.GetTextureParamsVersion() ||
resource.syncedMipLevelCount != mipLevelCount;
}
Bool VkTextureManager::TransitionImageLayout(VkCommandBuffer commandBuffer, VkImage image, Bool VkTextureManager::TransitionImageLayout(VkCommandBuffer commandBuffer, VkImage image,
VkImageLayout& trackedLayout, VkImageLayout newLayout, VkImageLayout& trackedLayout, VkImageLayout newLayout,
VkPipelineStageFlags srcStageMask, VkPipelineStageFlags dstStageMask, VkPipelineStageFlags srcStageMask, VkPipelineStageFlags dstStageMask,
@@ -1323,7 +1341,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
(aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 && (aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 &&
(formatProperties.optimalTilingFeatures & VK_FORMAT_FEATURE_STORAGE_IMAGE_BIT) != 0; (formatProperties.optimalTilingFeatures & VK_FORMAT_FEATURE_STORAGE_IMAGE_BIT) != 0;
VkImageCreateFlags imageCreateFlags = shapeInfo.imageFlags; VkImageCreateFlags imageCreateFlags = shapeInfo.imageFlags;
if (supportsStorageImage && IsMutableStorageImageFormat(format)) { if (supportsStorageImage && IsMutableStorageImageFormat(format) &&
m_mutableFormatUnsupported.find(format) == m_mutableFormatUnsupported.end()) {
imageCreateFlags |= VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT; imageCreateFlags |= VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT;
} }
@@ -1391,9 +1410,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
imageInfo.samples = resolvedSampleCount; imageInfo.samples = resolvedSampleCount;
if (isMultisampleTexture || (imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) { if (isMultisampleTexture || (imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
VkImageFormatProperties imageFormatProperties{}; VkImageFormatProperties imageFormatProperties{};
const VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties( VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
m_physicalDevice, format, imageInfo.imageType, imageInfo.tiling, imageInfo.usage, m_physicalDevice, format, imageInfo.imageType, imageInfo.tiling, imageInfo.usage,
imageInfo.flags, &imageFormatProperties); imageInfo.flags, &imageFormatProperties);
if (imageFormatResult != VK_SUCCESS && !isMultisampleTexture &&
(imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
// Losing reinterpreted views only degrades the formatless-image feature for
// this texture; failing creation would lose the texture entirely, so retry
// as a plain immutable-format image.
MGLOG_W("%s: mutable image format=%d is unsupported for textureId=%d; creating "
"without VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT (format reinterpretation "
"will be unavailable for it)",
__func__, static_cast<Int>(format), texture.GetExternalIndex());
// Remember the verdict so later syncs of same-format textures neither retry
// the probe nor flag-mismatch against this image and recreate it.
m_mutableFormatUnsupported.insert(format);
imageInfo.flags &= ~VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT;
imageCreateFlags = imageInfo.flags;
imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
m_physicalDevice, format, imageInfo.imageType, imageInfo.tiling, imageInfo.usage,
imageInfo.flags, &imageFormatProperties);
}
if (imageFormatResult != VK_SUCCESS || if (imageFormatResult != VK_SUCCESS ||
(isMultisampleTexture && (imageFormatProperties.sampleCounts & resolvedSampleCount) == 0)) { (isMultisampleTexture && (imageFormatProperties.sampleCounts & resolvedSampleCount) == 0)) {
MGLOG_D("%s: image flags=0x%x sampleCount=%d are unsupported for textureId=%d target=%s " MGLOG_D("%s: image flags=0x%x sampleCount=%d are unsupported for textureId=%d target=%s "
@@ -1921,6 +1958,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkFormat VkTextureManager::ResolveSampledImageViewFormat(VkFormat imageFormat, VkFormat VkTextureManager::ResolveSampledImageViewFormat(VkFormat imageFormat,
SamplerNumericDomain numericDomain) { SamplerNumericDomain numericDomain) {
// Depth/stencil images always sample through the existing depth-aspect sampledView.
// Combined formats (D24S8, D32FS8) are multi-numeric, so vkuFormatIsSampledFloat is
// false for them by design, yet their depth aspect reads as float in every GL depth
// texture mode; Vulkan also forbids reinterpreting them through color-class views.
// Integer domains keep the same view (pre-reinterpretation behavior for stencil-index
// style access) rather than failing the draw.
if (vkuFormatIsDepthOrStencil(imageFormat)) {
return imageFormat;
}
if (imageFormat == VK_FORMAT_UNDEFINED || numericDomain == SamplerNumericDomain::Unknown || if (imageFormat == VK_FORMAT_UNDEFINED || numericDomain == SamplerNumericDomain::Unknown ||
FormatMatchesSamplerNumericDomain(imageFormat, numericDomain)) { FormatMatchesSamplerNumericDomain(imageFormat, numericDomain)) {
return imageFormat; return imageFormat;
@@ -13,6 +13,7 @@
#include <MG_State/GLState/TextureState/TextureObject.h> #include <MG_State/GLState/TextureState/TextureObject.h>
#include <vk_mem_alloc.h> #include <vk_mem_alloc.h>
#include <unordered_map> #include <unordered_map>
#include <unordered_set>
namespace MobileGL::MG_State::GLState { namespace MobileGL::MG_State::GLState {
class ITextureObject; class ITextureObject;
@@ -284,6 +285,11 @@ public:
VkImageLayout newLayout); VkImageLayout newLayout);
Bool TransitionTextureForSampling(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture); Bool TransitionTextureForSampling(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
Bool TransitionTextureForStorageImage(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture); Bool TransitionTextureForStorageImage(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
// Non-mutating probe for the per-draw storage-image fast path: true when preparing this
// texture as a storage image may need work that is illegal inside a render pass (resource
// creation, dirty-content upload, or a layout transition to GENERAL). Unknown state reports
// true - a false positive merely ends the render pass, a false negative would skip a barrier.
Bool NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const;
static VkImageAspectFlags ResolveSampledImageViewAspectMask(VkImageAspectFlags imageAspect); static VkImageAspectFlags ResolveSampledImageViewAspectMask(VkImageAspectFlags imageAspect);
static VkFormat ResolveSampledImageViewFormat(VkFormat imageFormat, SamplerNumericDomain numericDomain); static VkFormat ResolveSampledImageViewFormat(VkFormat imageFormat, SamplerNumericDomain numericDomain);
@@ -379,6 +385,9 @@ private:
TextureResource* resource = nullptr; TextureResource* resource = nullptr;
}; };
Vector<DrawSyncedTexture> m_drawSyncedThisDraw; Vector<DrawSyncedTexture> m_drawSyncedThisDraw;
// Formats whose mutable-image probe failed on this device; their images are created
// without MUTABLE_FORMAT_BIT so repeat syncs neither re-probe nor flag-mismatch.
std::unordered_set<VkFormat> m_mutableFormatUnsupported;
std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects; std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects;
std::unordered_map<TextureIdentity, TextureResource, TextureIdentityHash> m_textureResources; std::unordered_map<TextureIdentity, TextureResource, TextureIdentityHash> m_textureResources;
Vector<Vector<TextureResource>> m_deferredReleases; Vector<Vector<TextureResource>> m_deferredReleases;
@@ -27,6 +27,9 @@
#include <cstring> #include <cstring>
#include <vulkan/utility/vk_format_utils.h> #include <vulkan/utility/vk_format_utils.h>
#include <vulkan/vulkan_core.h> #include <vulkan/vulkan_core.h>
#ifdef __ANDROID__
#include <sys/system_properties.h>
#endif
#if defined(__APPLE__) #if defined(__APPLE__)
#include <CoreGraphics/CoreGraphics.h> #include <CoreGraphics/CoreGraphics.h>
@@ -1031,6 +1034,7 @@ void main() {
} }
)"; )";
static Uint32 ComputeFullMipLevelCount(const IntVec3& baseTexelSize) { static Uint32 ComputeFullMipLevelCount(const IntVec3& baseTexelSize) {
Int maxDimension = std::max<Int>( Int maxDimension = std::max<Int>(
baseTexelSize.x(), baseTexelSize.x(),
@@ -1618,6 +1622,63 @@ void main() {
case VK_FORMAT_R32G32B32A32_SFLOAT: case VK_FORMAT_R32G32B32A32_SFLOAT:
Memcpy(rgba, source, sizeof(Float) * 4); Memcpy(rgba, source, sizeof(Float) * 4);
return true; return true;
// Single- and dual-channel formats the reinterpretation feature makes common
// as readback sources (iterationRP custom images are R32F/R32UI-class).
// Missing channels take GL's defaults: 0 for GB, 1 for alpha.
case VK_FORMAT_R32_SFLOAT: {
Float value = 0.0f;
Memcpy(&value, source, sizeof(value));
rgba[0] = value;
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32G32_SFLOAT: {
Float values[2] = {0.0f, 0.0f};
Memcpy(values, source, sizeof(values));
rgba[0] = values[0];
rgba[1] = values[1];
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32_UINT: {
Uint32 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = static_cast<Float>(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32_SINT: {
Int32 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = static_cast<Float>(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R16_SFLOAT: {
Uint16 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = MG_Util::DecodeHalfBitsToFloat(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R16G16_SFLOAT:
for (SizeT component = 0; component < 2; ++component) {
Uint16 value = 0;
Memcpy(&value, source + component * sizeof(value), sizeof(value));
rgba[component] = MG_Util::DecodeHalfBitsToFloat(value);
}
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
default: default:
return false; return false;
} }
@@ -2046,6 +2107,26 @@ void main() {
m_pipelineFactory = MakeUnique<PipelineFactory>(m_device, m_config); m_pipelineFactory = MakeUnique<PipelineFactory>(m_device, m_config);
MOBILEGL_ASSERT(m_pipelineFactory != nullptr, "PipelineFactory creation failed."); MOBILEGL_ASSERT(m_pipelineFactory != nullptr, "PipelineFactory creation failed.");
{
// Qualcomm's pipeline compiler does not keep vertex positions invariant across
// the pipelines of a multi-pass depth-equality chain (even with the SPIR-V
// Invariant decoration), so a blended depth-writing prepass makes later
// equality-compare passes drop whole primitives (MC 26.3 improved-transparency
// clouds flicker black). Suppress depth writes on accumulation-blended pipelines
// there (see PipelineFactory::ShouldSuppressDepthWrite for the exact scope);
// MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE forces the quirk on or off on any
// driver.
const MG_Config::QuirkOverride quirkOverride =
MG_Config::Features.MagmaDisableBlendedDepthWriteQuirk;
const Bool suppressBlendedDepthWrite = PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(
quirkOverride, m_physicalDevice.properties.vendorID);
if (suppressBlendedDepthWrite) {
MGLOG_I("DirectVulkan: suppressing depth writes on accumulation-blended pipelines "
"(driver lacks cross-pipeline position invariance)%s",
quirkOverride == MG_Config::QuirkOverride::ForceOn ? " (forced on)" : "");
}
PipelineFactory::SetSuppressBlendedDepthWrite(suppressBlendedDepthWrite);
}
m_programFactory = MakeUnique<ProgramFactory>(m_device, m_config, maxProgramBindings, m_programFactory = MakeUnique<ProgramFactory>(m_device, m_config, maxProgramBindings,
m_shaderDrawParametersFeatureEnabled, m_shaderDrawParametersFeatureEnabled,
m_unformattedFloatStorageImagesEnabled); m_unformattedFloatStorageImagesEnabled);
@@ -2072,15 +2153,23 @@ void main() {
MOBILEGL_ASSERT(m_vertexInputStateFactory != nullptr, "VertexInputStateFactory creation failed."); MOBILEGL_ASSERT(m_vertexInputStateFactory != nullptr, "VertexInputStateFactory creation failed.");
// Prime the first frame so Render() always targets an acquired swapchain image. // Prime the first frame so Render() always targets an acquired swapchain image.
VkResult acquireResult = // A zero-area window (GLFW's hidden helper window during the WGL bootstrap, or a
m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired); // window that is already minimized) legitimately yields no swapchain here; defer
if (acquireResult == VK_ERROR_OUT_OF_DATE_KHR || acquireResult == VK_SUBOPTIMAL_KHR) { // the first acquire to Present in that case instead of acquiring from a null
MGLOG_D("Initialize, vkAcquireNextImageKHR got %d, recreating swapchain", acquireResult); // swapchain handle.
RecreateSwapchain(); if (m_swapchainObject.GetHandle() != VK_NULL_HANDLE) {
acquireResult = VkResult acquireResult =
m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired); m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired);
if (acquireResult == VK_ERROR_OUT_OF_DATE_KHR || acquireResult == VK_SUBOPTIMAL_KHR) {
MGLOG_D("Initialize, vkAcquireNextImageKHR got %d, recreating swapchain", acquireResult);
RecreateSwapchain();
acquireResult =
m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired);
}
VK_VERIFY(acquireResult, "Initialize, WaitAndAcquireNextImage");
} else {
MGLOG_W("DirectVulkan: no swapchain at initialization (zero-area window); deferring first acquire");
} }
VK_VERIFY(acquireResult, "Initialize, WaitAndAcquireNextImage");
m_textureManager->BeginFrame(m_frameContext.GetCurrentFrameIndex()); m_textureManager->BeginFrame(m_frameContext.GetCurrentFrameIndex());
m_bufferManager.BeginFrame(m_frameContext.GetCurrentFrameIndex()); m_bufferManager.BeginFrame(m_frameContext.GetCurrentFrameIndex());
m_convertedVertexStreams.clear(); m_convertedVertexStreams.clear();
@@ -2200,10 +2289,94 @@ void main() {
MGLOG_I("VulkanRenderer shut down completed"); MGLOG_I("VulkanRenderer shut down completed");
} }
// Scans the draw's index range from host-visible index bytes and returns the largest
// fetchable vertex index. Usable only when the draw's range is exactly its
// IndexBufferView (drawParams.indexRangeIsExactView). The view's byte offset is either
// an offset into the bound element-array buffer or, with no bound buffer, a raw client
// pointer. Primitive-restart sentinels are skipped so they cannot inflate the bound.
static Bool TryComputeMaxIndexFromHostBytes(const MG_State::GLState::VertexArrayObject& vao,
const IndexBufferView& indexView, Uint32& outMaxIndex) {
SizeT indexSize = 0;
switch (indexView.indexType) {
case GL_UNSIGNED_BYTE: indexSize = 1; break;
case GL_UNSIGNED_SHORT: indexSize = 2; break;
case GL_UNSIGNED_INT: indexSize = 4; break;
default: return false;
}
const Uint8* indexBytes = nullptr;
const auto& indexBufferShared = vao.GetIndexBufferBindingSlot().GetBoundObject();
if (indexBufferShared != nullptr) {
const SizeT bufferSize = indexBufferShared->GetSize();
if (indexBufferShared->MappedData() == nullptr || indexView.indexByteOffset > bufferSize ||
indexView.indexByteSize > bufferSize - indexView.indexByteOffset) {
return false;
}
indexBufferShared->SyncPersistentMappedRange();
indexBytes = indexBufferShared->MappedData() + indexView.indexByteOffset;
} else {
indexBytes = reinterpret_cast<const Uint8*>(indexView.indexByteOffset);
if (indexBytes == nullptr) {
return false;
}
}
const SizeT indexCount = indexView.indexByteSize / indexSize;
// The all-ones sentinel is only a restart marker when primitive restart is enabled;
// with restart off it is a legitimate index and excluding it would truncate the
// converted stream by exactly that vertex.
const Bool primitiveRestartActive =
MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::PrimitiveRestart) ||
MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::PrimitiveRestartFixedIndex);
const Uint32 restartSentinel = indexSize == 1 ? 0xFFu : indexSize == 2 ? 0xFFFFu : 0xFFFFFFFFu;
Uint32 maxIndex = 0;
Bool sawIndex = false;
for (SizeT i = 0; i < indexCount; ++i) {
Uint32 index = 0;
switch (indexSize) {
case 1: index = indexBytes[i]; break;
case 2: index = reinterpret_cast<const Uint16*>(indexBytes)[i]; break;
default: index = reinterpret_cast<const Uint32*>(indexBytes)[i]; break;
}
if (primitiveRestartActive && index == restartSentinel) {
continue;
}
maxIndex = std::max(maxIndex, index);
sawIndex = true;
}
if (!sawIndex) {
return false;
}
outMaxIndex = maxIndex;
return true;
}
Bool VulkanRenderer::UploadAndBindVertexBuffers( Bool VulkanRenderer::UploadAndBindVertexBuffers(
VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao, VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ProgramFactory::VkProgramObject& programObj, const DrawCmdParam& drawParams, const ProgramFactory::VkProgramObject& programObj, const DrawCmdParam& drawParams,
Bool indexedDraw) { const IndexBufferView* pIndexBufferView) {
const Bool indexedDraw = pIndexBufferView != nullptr;
// Exclusive upper bound on the vertex-stream elements this draw can fetch through
// vertex-rate bindings, or 0 when unbounded (indirect/multi draws). Computed lazily
// because the index scan is only worth doing when a conversion actually needs it.
SizeT drawElementBound = 0;
Bool drawElementBoundComputed = false;
auto resolveDrawElementBound = [&]() -> SizeT {
if (drawElementBoundComputed) {
return drawElementBound;
}
drawElementBoundComputed = true;
if (!indexedDraw) {
drawElementBound = static_cast<SizeT>(drawParams.firstVertex) + drawParams.vertexCount;
} else if (drawParams.indexRangeIsExactView) {
Uint32 maxIndex = 0;
if (TryComputeMaxIndexFromHostBytes(vao, *pIndexBufferView, maxIndex)) {
drawElementBound = static_cast<SizeT>(maxIndex) + 1 +
static_cast<SizeT>(std::max(drawParams.baseVertex, 0));
}
}
return drawElementBound;
};
// programObj is resolved once in SetupDraw and passed in; re-resolving it here would repeat // programObj is resolved once in SetupDraw and passed in; re-resolving it here would repeat
// the GetCurrentProgram + GetOrCreateProgram hash lookup every draw. // the GetCurrentProgram + GetOrCreateProgram hash lookup every draw.
auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao); auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
@@ -2289,33 +2462,35 @@ void main() {
return false; return false;
} }
if (conversion != VertexInputStateFactory::VertexStreamConversion::None && indexedDraw) { // Client arrays have no queryable size, so bound the upload by the draw's
// The current indexed setup only carries indexCount, not the maximum effective // real fetch range. For indexed draws that means scanning the index bytes:
// index. Guessing a client-memory range here can truncate the converted stream. // the guessed vertexCount (indexCount + baseVertex) can both truncate draws
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute " // whose max index exceeds their index count and over-read below it.
"location=%u requires an indexed vertex range", location); const SizeT clientElementBound = resolveDrawElementBound();
return false;
}
if (conversion != VertexInputStateFactory::VertexStreamConversion::None &&
drawParams.vertexCount == 0) {
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute "
"location=%u has an unknown vertex range", location);
return false;
}
const Uint32 lastVertex = drawParams.vertexCount > 0
? drawParams.firstVertex + drawParams.vertexCount - 1
: drawParams.firstVertex;
BufferSlice slice{}; BufferSlice slice{};
Bool uploaded = false; Bool uploaded = false;
if (conversion == VertexInputStateFactory::VertexStreamConversion::None) { if (conversion == VertexInputStateFactory::VertexStreamConversion::None) {
const SizeT uploadSize = static_cast<SizeT>(lastVertex) * stride + elementSize; const SizeT lastVertex =
clientElementBound > 0
? clientElementBound - 1
: (drawParams.vertexCount > 0
? static_cast<SizeT>(drawParams.firstVertex) + drawParams.vertexCount - 1
: static_cast<SizeT>(drawParams.firstVertex));
const SizeT uploadSize = lastVertex * stride + elementSize;
uploaded = m_bufferManager.UploadTransient( uploaded = m_bufferManager.UploadTransient(
BufferKind::Vertex, m_frameContext.GetCurrentFrameIndex(), clientData, BufferKind::Vertex, m_frameContext.GetCurrentFrameIndex(), clientData,
static_cast<VkDeviceSize>(uploadSize), 16, slice); static_cast<VkDeviceSize>(uploadSize), 16, slice);
} else { } else {
if (clientElementBound == 0) {
// Indirect/multi indexed draws have no CPU-visible index range and a
// client array has no size to fall back to; a guessed range could
// truncate the converted stream, so skip the draw loudly.
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute "
"location=%u has no computable vertex range", location);
return false;
}
uploaded = uploadConvertedStream(conversion, attr, clientData, stride, elementSize, uploaded = uploadConvertedStream(conversion, attr, clientData, stride, elementSize,
static_cast<SizeT>(lastVertex) + 1, slice); clientElementBound, slice);
} }
if (!uploaded) { if (!uploaded) {
MOBILEGL_ASSERT(false, MOBILEGL_ASSERT(false,
@@ -2362,7 +2537,23 @@ void main() {
} }
sourceBufferShared->SyncPersistentMappedRange(); sourceBufferShared->SyncPersistentMappedRange();
const SizeT elementCount = 1 + (sourceSize - baseOffset - elementSize) / sourceStride; const SizeT availableElementCount = 1 + (sourceSize - baseOffset - elementSize) / sourceStride;
const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped();
// Convert only what this draw can fetch instead of the whole buffer tail.
// Instance-rate bindings index by instance, not the vertex range, so they
// keep the tail. Indexed draws from cacheable buffers also keep the tail: a
// single cached whole-range conversion per frame is cheaper than a per-draw
// index scan. Persistent-mapped buffers are uncacheable and reconvert every
// draw, so for them the scan plus bounded conversion is the cheaper trade.
const Bool vertexRateBinding =
vertexInputState.bindings[binding].inputRate == VK_VERTEX_INPUT_RATE_VERTEX;
SizeT elementCount = availableElementCount;
if (vertexRateBinding && (!indexedDraw || !cacheable)) {
const SizeT elementBound = resolveDrawElementBound();
if (elementBound > 0) {
elementCount = std::min(elementCount, elementBound);
}
}
const ConvertedVertexStreamKey cacheKey{ const ConvertedVertexStreamKey cacheKey{
.buffer = sourceBufferShared.get(), .buffer = sourceBufferShared.get(),
.changeSerial = sourceBufferShared->GetChangeSerial(), .changeSerial = sourceBufferShared->GetChangeSerial(),
@@ -2375,12 +2566,18 @@ void main() {
.conversion = conversion, .conversion = conversion,
}; };
const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped(); Bool reusedCachedStream = false;
auto cached = cacheable ? m_convertedVertexStreams.find(cacheKey) if (cacheable) {
: m_convertedVertexStreams.end(); const auto cached = m_convertedVertexStreams.find(cacheKey);
if (cached != m_convertedVertexStreams.end()) { // A cached conversion covering at least this draw's range is a strict
slice = cached->second; // prefix match: converted streams are tightly packed from element 0.
} else { if (cached != m_convertedVertexStreams.end() &&
cached->second.elementCount >= elementCount) {
slice = cached->second.slice;
reusedCachedStream = true;
}
}
if (!reusedCachedStream) {
const Uint8* sourceData = sourceBufferShared->MappedData() + baseOffset; const Uint8* sourceData = sourceBufferShared->MappedData() + baseOffset;
if (!uploadConvertedStream(conversion, attr, sourceData, sourceStride, if (!uploadConvertedStream(conversion, attr, sourceData, sourceStride,
elementSize, elementCount, slice)) { elementSize, elementCount, slice)) {
@@ -2389,7 +2586,8 @@ void main() {
return false; return false;
} }
if (cacheable) { if (cacheable) {
m_convertedVertexStreams.emplace(cacheKey, slice); m_convertedVertexStreams[cacheKey] =
ConvertedVertexStream{slice, elementCount, sourceBufferShared};
} }
} }
vkBuffers[binding] = slice.buffer; vkBuffers[binding] = slice.buffer;
@@ -3274,6 +3472,7 @@ void main() {
.backStencilPassOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthPassOp), .backStencilPassOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthPassOp),
.backStencilDepthFailOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthFailOp), .backStencilDepthFailOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthFailOp),
.backStencilCompareOp = MG_Util::ConvertDepthTestFuncToVkEnum(backStencil.Func), .backStencilCompareOp = MG_Util::ConvertDepthTestFuncToVkEnum(backStencil.Func),
.fragmentReplacesDepth = programObj.fragmentReplacesDepth,
.stages = &programObj.stages, .stages = &programObj.stages,
.vertexInputState = pipelineVertexInputState .vertexInputState = pipelineVertexInputState
}; };
@@ -3361,6 +3560,17 @@ void main() {
(bufferMask.a() ? VK_COLOR_COMPONENT_A_BIT : 0u)); (bufferMask.a() ? VK_COLOR_COMPONENT_A_BIT : 0u));
Bool effectiveBlendEnabled = blendEnabled; Bool effectiveBlendEnabled = blendEnabled;
MG_State::GLState::ITextureObject* colorAttachmentTexture = nullptr; MG_State::GLState::ITextureObject* colorAttachmentTexture = nullptr;
if (isDefaultDrawFbo && i < drawBuffers.size() &&
drawBuffers[i] == FramebufferAttachmentType::None) {
// The default framebuffer spans the same MAX_DRAW_BUFFERS slots as an FBO
// (slot 0 is the back buffer, or None after glDrawBuffer(GL_NONE); slots
// 1+ are always None). Discard writes and blend state for the None slots
// like the FBO path below does, so stale indexed blend state on phantom
// slots cannot leak into the pipeline - most notably into the blended
// depth-write quirk's accumulation scan.
attachmentColorWriteMask = 0;
effectiveBlendEnabled = false;
}
if (!isDefaultDrawFbo && i < drawBuffers.size()) { if (!isDefaultDrawFbo && i < drawBuffers.size()) {
const auto drawBuffer = drawBuffers[i]; const auto drawBuffer = drawBuffers[i];
colorAttachmentTexture = resolveCompleteColorAttachmentTexture(i); colorAttachmentTexture = resolveCompleteColorAttachmentTexture(i);
@@ -3478,6 +3688,16 @@ void main() {
"disabling blending on attachments with this format (first hit: attachment %u textureId=%d program=%u)", "disabling blending on attachments with this format (first hit: attachment %u textureId=%d program=%u)",
static_cast<Int>(colorAttachmentFormat), i, textureExternalIndex, static_cast<Int>(colorAttachmentFormat), i, textureExternalIndex,
program.GetExternalIndex()); program.GetExternalIndex());
if (PipelineFactory::IsSuppressBlendedDepthWriteEnabled()) {
// With blending force-disabled the blended depth-write quirk can
// never fire for pipelines on this format, so a depth-equality
// chain that accumulates into it (MC 26.3 OIT depth_bounds on
// RGBA32F) keeps its depth writes and may flicker on this driver.
MGLOG_W("GetOrCreatePipeline: format=%d is not blendable, so the blended "
"depth-write quirk cannot apply to it; depth-equality chains "
"accumulating into this format may flicker",
static_cast<Int>(colorAttachmentFormat));
}
} }
} }
if (!blendSupportIt->second) { if (!blendSupportIt->second) {
@@ -3526,6 +3746,9 @@ void main() {
VkCommandBuffer commandBuffer, VkCommandBuffer commandBuffer,
const MG_State::GLState::ProgramObject& program, const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj) { const ProgramFactory::VkProgramObject& programObj) {
if (!programObj.hasStorageImages) {
return true;
}
auto& storageTextures = m_storageImageTexturesScratch; auto& storageTextures = m_storageImageTexturesScratch;
if (!m_uniformManager->CollectStorageImageTextures(program, programObj, storageTextures)) { if (!m_uniformManager->CollectStorageImageTextures(program, programObj, storageTextures)) {
MGLOG_E("%s: failed to collect storage images for program=%u", MGLOG_E("%s: failed to collect storage images for program=%u",
@@ -3536,6 +3759,24 @@ void main() {
return true; return true;
} }
// Steady-state fast path: when every collected texture is already resident in GENERAL
// with no pending clear and no dirty content, the loop below has nothing to record, so
// keep the render pass alive instead of splitting it on every storage-image draw (on
// tiled GPUs each split is a full tile load/store). GL makes cross-draw image-store
// coherence the app's job (glMemoryBarrier), so no implicit barrier is owed here.
Bool anyNeedsPreparation = false;
for (auto* texture : storageTextures) {
MOBILEGL_ASSERT(texture != nullptr, "%s: collected a null storage texture", __func__);
if (m_textureManager->NeedsStorageImagePreparation(*texture) ||
m_clearManager->HasPendingClear(texture)) {
anyNeedsPreparation = true;
break;
}
}
if (!anyNeedsPreparation) {
return true;
}
// Image uploads, deferred-clear materialization, and layout barriers are illegal inside // Image uploads, deferred-clear materialization, and layout barriers are illegal inside
// a classic render pass. Do this before sampler preparation as well: a texture used by // a classic render pass. Do this before sampler preparation as well: a texture used by
// both a sampler and an image must stay in GENERAL, and both descriptors must name that // both a sampler and an image must stay in GENERAL, and both descriptors must name that
@@ -3545,7 +3786,6 @@ void main() {
} }
for (auto* texture : storageTextures) { for (auto* texture : storageTextures) {
MOBILEGL_ASSERT(texture != nullptr, "%s: collected a null storage texture", __func__);
if (!MaterializePendingClearForTexture(commandBuffer, *texture)) { if (!MaterializePendingClearForTexture(commandBuffer, *texture)) {
MGLOG_E("%s: failed to materialize pending clear for storage textureId=%d", MGLOG_E("%s: failed to materialize pending clear for storage textureId=%d",
__func__, texture->GetExternalIndex()); __func__, texture->GetExternalIndex());
@@ -3763,7 +4003,7 @@ void main() {
} }
auto vtxUploadOk = UploadAndBindVertexBuffers( auto vtxUploadOk = UploadAndBindVertexBuffers(
frame.commandBuffer, vao, programObj, drawParams, pIndexBufferView != nullptr); frame.commandBuffer, vao, programObj, drawParams, pIndexBufferView);
if (!vtxUploadOk) { if (!vtxUploadOk) {
MGLOG_E("SetupDraw skipped: failed to upload vertex buffers"); MGLOG_E("SetupDraw skipped: failed to upload vertex buffers");
return false; return false;
@@ -5977,6 +6217,10 @@ void main() {
vertexRange.instanceCount = payload.params.instanceCount; vertexRange.instanceCount = payload.params.instanceCount;
vertexRange.firstVertex = 0; vertexRange.firstVertex = 0;
vertexRange.firstInstance = static_cast<Uint32>(payload.params.firstInstance); vertexRange.firstInstance = static_cast<Uint32>(payload.params.firstInstance);
vertexRange.baseVertex = payload.params.vertexOffset;
// Direct DrawElements fetches exactly the indices in its view, so vertex-stream
// conversion may bound its work by scanning them.
vertexRange.indexRangeIsExactView = true;
if (!SetupDraw(frame, payload.mode, DrawSetupAspect::IndexBuffer, vertexRange, if (!SetupDraw(frame, payload.mode, DrawSetupAspect::IndexBuffer, vertexRange,
&payload.indexBufferView)) { &payload.indexBufferView)) {
@@ -6621,6 +6865,33 @@ void main() {
} }
void VulkanRenderer::Present() { void VulkanRenderer::Present() {
if (m_swapchainObject.GetHandle() == VK_NULL_HANDLE || m_presentSuspended) {
// No usable swapchain: the window was zero-area at initialization, or
// presentation was suspended when the window minimized. Try to bring a
// swapchain up now that the window may have a real size; until then, drop
// this frame's recording instead of submitting - a submit would wait on a
// never-signaled acquire semaphore and reuse a still-signaled fence.
if (!RecreateSwapchain()) {
auto& suspendedFrame = m_frameContext.GetCurrent();
if (VkRenderPassManager::GetActiveRenderPass()) {
VkRenderPassManager::EndRenderPass(suspendedFrame.commandBuffer);
}
if (suspendedFrame.isCommandRecording) {
m_frameContext.EndCommandRecording();
}
suspendedFrame.isCommandRecording = false;
suspendedFrame.hasCommandBufferRecorded = false;
m_lastPipelineValid = false;
MGLOG_D("Present skipped: no usable swapchain (zero-area window)");
return;
}
m_presentSuspended = false;
const VkResult acquireResult =
m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired);
if (acquireResult != VK_SUBOPTIMAL_KHR) {
VK_VERIFY(acquireResult, "Present, deferred first WaitAndAcquireNextImage");
}
}
MOBILEGL_ASSERT(m_imageIndexAcquired < m_swapchainObject.GetImageCount(), MOBILEGL_ASSERT(m_imageIndexAcquired < m_swapchainObject.GetImageCount(),
"Present, acquired image index out of range"); "Present, acquired image index out of range");
m_renderPassManager->OnPresent(); m_renderPassManager->OnPresent();
@@ -6654,14 +6925,26 @@ void main() {
auto result = vkQueuePresentKHR(m_presentQueue, &presentPacket.presentInfo); auto result = vkQueuePresentKHR(m_presentQueue, &presentPacket.presentInfo);
if (result == VK_ERROR_OUT_OF_DATE_KHR || result == VK_SUBOPTIMAL_KHR) { if (result == VK_ERROR_OUT_OF_DATE_KHR || result == VK_SUBOPTIMAL_KHR) {
MGLOG_D("Present, vkQueuePresentKHR got %d, recreating swapchain", result); MGLOG_D("Present, vkQueuePresentKHR got %d, recreating swapchain", result);
RecreateSwapchain(); if (!RecreateSwapchain()) {
// Window went zero-area (minimize) with the swapchain out of date:
// stop submitting/acquiring until it has a size again.
m_presentSuspended = true;
m_swapchainResizeRequested = false;
MGLOG_D("Present, zero-area window with out-of-date swapchain; suspending presentation");
return;
}
m_swapchainResizeRequested = false; m_swapchainResizeRequested = false;
result = VK_SUCCESS; result = VK_SUCCESS;
} }
VK_VERIFY(result, "Present, vkQueuePresentKHR"); VK_VERIFY(result, "Present, vkQueuePresentKHR");
if (m_swapchainResizeRequested) { if (m_swapchainResizeRequested) {
MGLOG_D("Present, processing requested swapchain resize"); MGLOG_D("Present, processing requested swapchain resize");
RecreateSwapchain(); if (!RecreateSwapchain()) {
m_presentSuspended = true;
m_swapchainResizeRequested = false;
MGLOG_D("Present, zero-area window on requested resize; suspending presentation");
return;
}
m_swapchainResizeRequested = false; m_swapchainResizeRequested = false;
} }
@@ -6672,7 +6955,12 @@ void main() {
result = m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired); result = m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired);
if (result == VK_ERROR_OUT_OF_DATE_KHR || result == VK_SUBOPTIMAL_KHR) { if (result == VK_ERROR_OUT_OF_DATE_KHR || result == VK_SUBOPTIMAL_KHR) {
MGLOG_D("Present, vkAcquireNextImageKHR got %d, recreating swapchain", result); MGLOG_D("Present, vkAcquireNextImageKHR got %d, recreating swapchain", result);
RecreateSwapchain(); if (!RecreateSwapchain()) {
m_presentSuspended = true;
m_swapchainResizeRequested = false;
MGLOG_D("Present, zero-area window on next-frame acquire; suspending presentation");
return;
}
m_swapchainResizeRequested = false; m_swapchainResizeRequested = false;
result = result =
m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired); m_frameContext.WaitAndAcquireNextImage(m_device, m_swapchainObject.GetHandle(), m_imageIndexAcquired);
@@ -7007,7 +7295,10 @@ void main() {
// Match GL's robust buffer-fetch behavior where the Vulkan device supports it. This covers // Match GL's robust buffer-fetch behavior where the Vulkan device supports it. This covers
// out-of-range fetches; arbitrary GL vertex strides/offsets still need the explicit tight // out-of-range fetches; arbitrary GL vertex strides/offsets still need the explicit tight
// repack in VertexInputStateFactory when they violate Vulkan's address-alignment rules. // repack in VertexInputStateFactory when they violate Vulkan's address-alignment rules.
deviceFeatures.robustBufferAccess = supportedDeviceFeatures.robustBufferAccess; // MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS leaves it off to measure or dodge its GPU cost.
deviceFeatures.robustBufferAccess = MG_Config::Features.DisableRobustBufferAccess
? VK_FALSE
: supportedDeviceFeatures.robustBufferAccess;
deviceFeatures.geometryShader = supportedDeviceFeatures.geometryShader; deviceFeatures.geometryShader = supportedDeviceFeatures.geometryShader;
deviceFeatures.independentBlend = supportedDeviceFeatures.independentBlend; deviceFeatures.independentBlend = supportedDeviceFeatures.independentBlend;
m_independentBlendFeatureEnabled = deviceFeatures.independentBlend == VK_TRUE; m_independentBlendFeatureEnabled = deviceFeatures.independentBlend == VK_TRUE;
@@ -7037,6 +7328,15 @@ void main() {
if (m_unformattedFloatStorageImagesEnabled) { if (m_unformattedFloatStorageImagesEnabled) {
deviceFeatures.shaderStorageImageReadWithoutFormat = VK_TRUE; deviceFeatures.shaderStorageImageReadWithoutFormat = VK_TRUE;
deviceFeatures.shaderStorageImageWriteWithoutFormat = VK_TRUE; deviceFeatures.shaderStorageImageWriteWithoutFormat = VK_TRUE;
} else {
// Surface the degradation instead of failing silently: shader packs that bind a
// float storage image with a format different from its declaration (e.g.
// iterationRP) will render incorrectly on this device.
MGLOG_W("CreateLogicalDeviceAndQueues: shaderStorageImage*WithoutFormat unavailable "
"(read=%d write=%d); float storage-image format reinterpretation is disabled "
"and packs relying on it may misrender",
supportedDeviceFeatures.shaderStorageImageReadWithoutFormat,
supportedDeviceFeatures.shaderStorageImageWriteWithoutFormat);
} }
deviceFeatures.drawIndirectFirstInstance = supportedDeviceFeatures.drawIndirectFirstInstance; deviceFeatures.drawIndirectFirstInstance = supportedDeviceFeatures.drawIndirectFirstInstance;
deviceFeatures.multiDrawIndirect = supportedDeviceFeatures.multiDrawIndirect; deviceFeatures.multiDrawIndirect = supportedDeviceFeatures.multiDrawIndirect;
@@ -7312,6 +7612,7 @@ void main() {
MOBILEGL_ASSERT(hwnd, "HWND is null"); MOBILEGL_ASSERT(hwnd, "HWND is null");
VkWin32SurfaceCreateInfoKHR sci{VK_STRUCTURE_TYPE_WIN32_SURFACE_CREATE_INFO_KHR}; VkWin32SurfaceCreateInfoKHR sci{VK_STRUCTURE_TYPE_WIN32_SURFACE_CREATE_INFO_KHR};
sci.hinstance = GetModuleHandleW(nullptr);
sci.hwnd = hwnd; sci.hwnd = hwnd;
VK_VERIFY(vkCreateWin32SurfaceKHR(m_instance, &sci, nullptr, &m_surface), "vkCreateWin32SurfaceKHR failed"); VK_VERIFY(vkCreateWin32SurfaceKHR(m_instance, &sci, nullptr, &m_surface), "vkCreateWin32SurfaceKHR failed");
#elif defined VK_USE_PLATFORM_METAL_EXT #elif defined VK_USE_PLATFORM_METAL_EXT
@@ -7473,13 +7774,13 @@ void main() {
m_swapchainObject.Shutdown(m_device); m_swapchainObject.Shutdown(m_device);
} }
void VulkanRenderer::RecreateSwapchain() { Bool VulkanRenderer::RecreateSwapchain() {
// Handle cases like minimize on Windows, where swapchain could return a 0x0 extent // Handle cases like minimize on Windows, where swapchain could return a 0x0 extent
const auto swapchainCapabilities = const auto swapchainCapabilities =
SwapchainObject::GetSwapchainCapabilities(m_physicalDevice.handle, m_surface); SwapchainObject::GetSwapchainCapabilities(m_physicalDevice.handle, m_surface);
if (swapchainCapabilities.capabilities.currentExtent.width == 0 || if (swapchainCapabilities.capabilities.currentExtent.width == 0 ||
swapchainCapabilities.capabilities.currentExtent.height == 0) { swapchainCapabilities.capabilities.currentExtent.height == 0) {
return; return false;
} }
vkDeviceWaitIdle(m_device); vkDeviceWaitIdle(m_device);
@@ -7523,6 +7824,7 @@ void main() {
m_bufferManager.BeginFrame(m_frameContext.GetCurrentFrameIndex()); m_bufferManager.BeginFrame(m_frameContext.GetCurrentFrameIndex());
m_convertedVertexStreams.clear(); m_convertedVertexStreams.clear();
} }
return true;
} }
const PhysicalDevice& VulkanRenderer::GetPhysicalDevice() const { const PhysicalDevice& VulkanRenderer::GetPhysicalDevice() const {
@@ -51,6 +51,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 instanceCount = 1; Uint32 instanceCount = 1;
Uint32 firstVertex = 0; Uint32 firstVertex = 0;
Uint32 firstInstance = 0; Uint32 firstInstance = 0;
// Indexed-draw metadata for bounding vertex-stream conversion. baseVertex is the
// draw's base-vertex offset; indexRangeIsExactView is true only when the draw
// fetches exactly the indices its IndexBufferView describes (direct DrawElements;
// multi/indirect forms leave it false because the CPU cannot bound their ranges).
Int32 baseVertex = 0;
Bool indexRangeIsExactView = false;
}; };
struct DrawIndexedCmdParam { struct DrawIndexedCmdParam {
@@ -251,7 +257,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 GetTimerQueryTimestampNs(const VkTimerQueryManager::TimestampRecord& record) const; Uint64 GetTimerQueryTimestampNs(const VkTimerQueryManager::TimestampRecord& record) const;
void RequestSwapchainResize(Uint32 width, Uint32 height); void RequestSwapchainResize(Uint32 width, Uint32 height);
void RecreateSwapchain(); // Returns false when the surface is zero-area (minimized/hidden window):
// no new swapchain is installed and presentation must stay suspended.
Bool RecreateSwapchain();
private: private:
struct BlitUniformData { struct BlitUniformData {
@@ -349,6 +357,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void* m_platformCloseDisplay = nullptr; void* m_platformCloseDisplay = nullptr;
VulkanRendererConfig m_config; VulkanRendererConfig m_config;
Bool m_swapchainResizeRequested = false; Bool m_swapchainResizeRequested = false;
// Presentation is suspended while the window is zero-area (minimized): the
// swapchain is unusable/out of date, so Present drops frames instead of
// submitting on a signaled fence / presenting never-acquired images.
Bool m_presentSuspended = false;
// Vulkan objects // Vulkan objects
Bool m_validationLayersEnabled = false; Bool m_validationLayersEnabled = false;
@@ -488,7 +500,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
}; };
UnorderedMap<ConvertedVertexStreamKey, BufferSlice, ConvertedVertexStreamKeyHash> struct ConvertedVertexStream {
BufferSlice slice;
// Number of source elements the cached slice covers. A draw needing a prefix of
// this range reuses the slice (converted streams are tightly packed); a draw
// needing more reconverts and replaces the entry, so per (buffer, layout) a
// frame converts at most the largest range any draw asked for.
SizeT elementCount = 0;
// Pins the source buffer for the frame so its heap address cannot be reused by
// a new BufferObject while this pointer-keyed entry is alive.
SharedPtr<const MG_State::GLState::BufferObject> sourcePin;
};
UnorderedMap<ConvertedVertexStreamKey, ConvertedVertexStream, ConvertedVertexStreamKeyHash>
m_convertedVertexStreams; m_convertedVertexStreams;
void CreateInstance(); void CreateInstance();
@@ -519,7 +542,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool UploadAndBindVertexBuffers(VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao, Bool UploadAndBindVertexBuffers(VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ProgramFactory::VkProgramObject& programObj, const ProgramFactory::VkProgramObject& programObj,
const DrawCmdParam& drawParams, Bool indexedDraw); const DrawCmdParam& drawParams,
const IndexBufferView* pIndexBufferView);
Bool UploadAndBindIndexBuffer(FrameContext::FrameData& frame, Bool UploadAndBindIndexBuffer(FrameContext::FrameData& frame,
const MG_State::GLState::VertexArrayObject& vao, const MG_State::GLState::VertexArrayObject& vao,
const IndexBufferView* pIndexBufferView = nullptr); const IndexBufferView* pIndexBufferView = nullptr);
+28 -5
View File
@@ -8,6 +8,7 @@
#include "EGLImpl.h" #include "EGLImpl.h"
#include "../GetProcAddress.h" #include "../GetProcAddress.h"
#include <Init.h>
#include <MG_Backend/BackendObjects.h> #include <MG_Backend/BackendObjects.h>
#include <MG_State/EGLState/Core.h> #include <MG_State/EGLState/Core.h>
#include <mutex> #include <mutex>
@@ -25,6 +26,17 @@ namespace MobileGL::MG_Impl::EGLImpl {
return MG_State::pEGLContext.get(); return MG_State::pEGLContext.get();
} }
// Entry points that can legitimately be an application's FIRST EGL
// call (display/proc-address/string queries) lazily bring MobileGL
// up here, so the library needs no static constructor and can
// re-initialize after the last eglTerminate tore everything down.
// Teardown-ish entry points keep using GetState() and fail benignly
// when MobileGL is not initialized.
EGLStateContext* GetStateEnsureInitialized() {
MobileGL::EnsureInitialized();
return GetState();
}
MG_Backend::BackendObject* GetBackendObject(EGLStateContext* state) { MG_Backend::BackendObject* GetBackendObject(EGLStateContext* state) {
auto* backendObject = MG_Backend::pActiveBackendObject.get(); auto* backendObject = MG_Backend::pActiveBackendObject.get();
if (!backendObject && state) { if (!backendObject && state) {
@@ -49,6 +61,8 @@ namespace MobileGL::MG_Impl::EGLImpl {
return MG_Backend::WindowBackend::Android; return MG_Backend::WindowBackend::Android;
#elif defined(__APPLE__) #elif defined(__APPLE__)
return MG_Backend::WindowBackend::MetalLayer; return MG_Backend::WindowBackend::MetalLayer;
#elif defined(_WIN32)
return MG_Backend::WindowBackend::Win32;
#elif defined(__linux__) #elif defined(__linux__)
return MG_Backend::WindowBackend::X11; return MG_Backend::WindowBackend::X11;
#else #else
@@ -187,7 +201,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
} }
EGLBoolean Initialize(EGLDisplay dpy, EGLint* major, EGLint* minor) { EGLBoolean Initialize(EGLDisplay dpy, EGLint* major, EGLint* minor) {
auto* state = GetState(); auto* state = GetStateEnsureInitialized();
if (!state) { if (!state) {
return EGL_FALSE; return EGL_FALSE;
} }
@@ -208,7 +222,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
} }
EGLDisplay GetDisplay(NativeDisplayType display) { EGLDisplay GetDisplay(NativeDisplayType display) {
auto* state = GetState(); auto* state = GetStateEnsureInitialized();
if (!state) { if (!state) {
return EGL_NO_DISPLAY; return EGL_NO_DISPLAY;
} }
@@ -313,6 +327,14 @@ namespace MobileGL::MG_Impl::EGLImpl {
if (auto* backendObject = MG_Backend::pActiveBackendObject.get()) { if (auto* backendObject = MG_Backend::pActiveBackendObject.get()) {
backendObject->ReleaseEGLResources(); backendObject->ReleaseEGLResources();
} }
// The last initialized display is gone and nothing is current on any
// thread: tear the whole library down deterministically inside the
// EGL lifecycle (backend, GL/EGL state, glslang). A later EGL call
// re-initializes lazily via GetStateEnsureInitialized(); process exit
// then has nothing left to destroy.
if (!state->HasAnyInitializedDisplay() && !state->HasAnyCurrentContext()) {
MobileGL::Destroy();
}
return EGL_TRUE; return EGL_TRUE;
} }
@@ -345,7 +367,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
} }
EGLBoolean BindAPI(EGLenum api) { EGLBoolean BindAPI(EGLenum api) {
auto* state = GetState(); auto* state = GetStateEnsureInitialized();
if (!state) { if (!state) {
return EGL_FALSE; return EGL_FALSE;
} }
@@ -378,7 +400,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
} }
char const* QueryString(EGLDisplay display, EGLint name) { char const* QueryString(EGLDisplay display, EGLint name) {
auto* state = GetState(); auto* state = GetStateEnsureInitialized();
if (!state) { if (!state) {
return nullptr; return nullptr;
} }
@@ -641,7 +663,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
EGLDisplay GetPlatformDisplay(EGLenum platform, void* native_display, const EGLAttrib* attrib_list) { EGLDisplay GetPlatformDisplay(EGLenum platform, void* native_display, const EGLAttrib* attrib_list) {
(void)attrib_list; (void)attrib_list;
auto* state = GetState(); auto* state = GetStateEnsureInitialized();
if (!state) { if (!state) {
return EGL_NO_DISPLAY; return EGL_NO_DISPLAY;
} }
@@ -737,6 +759,7 @@ namespace MobileGL::MG_Impl::EGLImpl {
if (!name) { if (!name) {
return nullptr; return nullptr;
} }
MobileGL::EnsureInitialized();
MGLOG_D("eglGetProcAddress(%s)", name); MGLOG_D("eglGetProcAddress(%s)", name);
void* proc = MG_Impl::GetProcAddress(name); void* proc = MG_Impl::GetProcAddress(name);
@@ -295,7 +295,13 @@ DECLARE_GL_FUNCTION_HEAD(void, VertexAttribDivisor, GLuint index, GLuint divisor
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedback, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedback, target, id) DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedback, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedback, target, id)
DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacks, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacks, n, ids) DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacks, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacks, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacks, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacks, n, ids) DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacks, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacks, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(GLboolean, IsTransformFeedback, GLuint id) DECLARE_GL_FUNCTION_STUB_END(GLboolean, IsTransformFeedback, id) // Transform feedback objects are not implemented, so no name is ever a live object. The shared
// stub returns (type)1, telling a probing caller that every id it invents already exists; GL_FALSE
// is both truthful and what the spec requires for a name that was never generated.
MOBILEGL_GL_API GLboolean glIsTransformFeedback(GLuint id) {
MGLOG_W("Stub function: %s(...)", __FUNCTION__);
return GL_FALSE;
}
DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedback) DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedback)
DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedback) DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedback)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetProgramBinary, GLuint program, GLsizei bufSize, GLsizei* length, GLenum* binaryFormat, void* binary) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetProgramBinary, program, bufSize, length, binaryFormat, binary) DECLARE_GL_FUNCTION_STUB_HEAD(void, GetProgramBinary, GLuint program, GLsizei bufSize, GLsizei* length, GLenum* binaryFormat, void* binary) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetProgramBinary, program, bufSize, length, binaryFormat, binary)
@@ -418,7 +424,7 @@ DECLARE_GL_FUNCTION_HEAD(void, DrawRangeElementsBaseVertex, GLenum mode, GLuint
DECLARE_GL_FUNCTION_HEAD(void, DrawElementsInstancedBaseVertex, GLenum mode, GLsizei count, GLenum type, const void* indices, GLsizei instancecount, GLint basevertex) DECLARE_GL_FUNCTION_END_NO_RETURN(void, DrawElementsInstancedBaseVertex, mode, count, type, indices, instancecount, basevertex) DECLARE_GL_FUNCTION_HEAD(void, DrawElementsInstancedBaseVertex, GLenum mode, GLsizei count, GLenum type, const void* indices, GLsizei instancecount, GLint basevertex) DECLARE_GL_FUNCTION_END_NO_RETURN(void, DrawElementsInstancedBaseVertex, mode, count, type, indices, instancecount, basevertex)
DECLARE_GL_FUNCTION_HEAD(void, FramebufferTexture, GLenum target, GLenum attachment, GLuint texture, GLint level) DECLARE_GL_FUNCTION_END_NO_RETURN(void, FramebufferTexture, target, attachment, texture, level) DECLARE_GL_FUNCTION_HEAD(void, FramebufferTexture, GLenum target, GLenum attachment, GLuint texture, GLint level) DECLARE_GL_FUNCTION_END_NO_RETURN(void, FramebufferTexture, target, attachment, texture, level)
DECLARE_GL_FUNCTION_STUB_HEAD(void, PrimitiveBoundingBox, GLfloat minX, GLfloat minY, GLfloat minZ, GLfloat minW, GLfloat maxX, GLfloat maxY, GLfloat maxZ, GLfloat maxW) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PrimitiveBoundingBox, minX, minY, minZ, minW, maxX, maxY, maxZ, maxW) DECLARE_GL_FUNCTION_STUB_HEAD(void, PrimitiveBoundingBox, GLfloat minX, GLfloat minY, GLfloat minZ, GLfloat minW, GLfloat maxX, GLfloat maxY, GLfloat maxZ, GLfloat maxW) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PrimitiveBoundingBox, minX, minY, minZ, minW, maxX, maxY, maxZ, maxW)
DECLARE_GL_FUNCTION_STUB_HEAD(GLenum, GetGraphicsResetStatus) DECLARE_GL_FUNCTION_STUB_END(GLenum, GetGraphicsResetStatus) DECLARE_GL_FUNCTION_HEAD(GLenum, GetGraphicsResetStatus) DECLARE_GL_FUNCTION_END(GLenum, GetGraphicsResetStatus)
DECLARE_GL_FUNCTION_STUB_HEAD(void, ReadnPixels, GLint x, GLint y, GLsizei width, GLsizei height, GLenum format, GLenum type, GLsizei bufSize, void* data) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ReadnPixels, x, y, width, height, format, type, bufSize, data) DECLARE_GL_FUNCTION_STUB_HEAD(void, ReadnPixels, GLint x, GLint y, GLsizei width, GLsizei height, GLenum format, GLenum type, GLsizei bufSize, void* data) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ReadnPixels, x, y, width, height, format, type, bufSize, data)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformfv, GLuint program, GLint location, GLsizei bufSize, GLfloat* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformfv, program, location, bufSize, params) DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformfv, GLuint program, GLint location, GLsizei bufSize, GLfloat* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformfv, program, location, bufSize, params)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformiv, GLuint program, GLint location, GLsizei bufSize, GLint* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformiv, program, location, bufSize, params) DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformiv, GLuint program, GLint location, GLsizei bufSize, GLint* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformiv, program, location, bufSize, params)
@@ -2583,7 +2589,10 @@ DECLARE_GL_FUNCTION_STUB_HEAD(void, TransformFeedbackStreamAttribsNV, GLsizei co
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedbackNV, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedbackNV, target, id) DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedbackNV, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedbackNV, target, id)
DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacksNV, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacksNV, n, ids) DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacksNV, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacksNV, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacksNV, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacksNV, n, ids) DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacksNV, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacksNV, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(GLboolean, IsTransformFeedbackNV, GLuint id) DECLARE_GL_FUNCTION_STUB_END(GLboolean, IsTransformFeedbackNV, id) MOBILEGL_GL_API GLboolean glIsTransformFeedbackNV(GLuint id) {
MGLOG_W("Stub function: %s(...)", __FUNCTION__);
return GL_FALSE;
}
DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedbackNV, ) DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedbackNV, )
DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedbackNV, ) DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedbackNV, )
DECLARE_GL_FUNCTION_STUB_HEAD(void, DrawTransformFeedbackNV, GLenum mode, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DrawTransformFeedbackNV, mode, id) DECLARE_GL_FUNCTION_STUB_HEAD(void, DrawTransformFeedbackNV, GLenum mode, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DrawTransformFeedbackNV, mode, id)
@@ -2144,6 +2144,7 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
namespace FramebufferImpl { namespace FramebufferImpl {
UniquePtr<DefaultFramebufferInfo> pDefaultFramebufferInfo; // Leak-at-exit storage; see GlobalObjects.cpp.
UniquePtr<DefaultFramebufferInfo>& pDefaultFramebufferInfo = *new UniquePtr<DefaultFramebufferInfo>();
} // namespace FramebufferImpl } // namespace FramebufferImpl
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
@@ -78,6 +78,6 @@ namespace MobileGL::MG_Impl::GLImpl {
SharedPtr<MG_State::GLState::ITextureObject> stencilAttachment; SharedPtr<MG_State::GLState::ITextureObject> stencilAttachment;
}; };
extern UniquePtr<DefaultFramebufferInfo> pDefaultFramebufferInfo; extern UniquePtr<DefaultFramebufferInfo>& pDefaultFramebufferInfo;
} // namespace FramebufferImpl } // namespace FramebufferImpl
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
@@ -1996,4 +1996,12 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
return MG_Util::ConvertErrorCodeToGLEnum(error->get()->code); return MG_Util::ConvertErrorCodeToGLEnum(error->get()->code);
} }
GLenum GetGraphicsResetStatus() {
// MobileGL does not implement robustness reset notification, so report GL_NO_ERROR
// ("no reset detected"). Returning the generic stub's (GLenum)1 makes dEQP read a lost
// device after every case (gl3cTestPackages.cpp:121) and, under the default
// --deqp-terminate-on-device-lost=enable, tear the whole CTS run down.
return GL_NO_ERROR;
}
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
@@ -21,4 +21,5 @@ namespace MobileGL::MG_Impl::GLImpl {
void GetIntegeri_v(GLenum target, GLuint index, GLint* data); void GetIntegeri_v(GLenum target, GLuint index, GLint* data);
void GetInteger64i_v(GLenum target, GLuint index, GLint64* data); void GetInteger64i_v(GLenum target, GLuint index, GLint64* data);
GLenum GetError(); GLenum GetError();
GLenum GetGraphicsResetStatus();
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
+131 -17
View File
@@ -350,6 +350,10 @@ namespace MobileGL::MG_Impl::GLImpl {
texture.AllocateStorage(uploadTarget, level, {levelTexelSize, levelByteSize}); texture.AllocateStorage(uploadTarget, level, {levelTexelSize, levelByteSize});
texture.MarkStorageDirty(uploadTarget, level, false); texture.MarkStorageDirty(uploadTarget, level, false);
} }
// glGenerateMipmap defines exactly levels 0..requiredLevelCount-1. AllocateStorage only
// grows, so a previously longer chain (a bigger base image before respecification) would
// otherwise keep a tail of stale levels here and read as incomplete.
texture.TruncateMipmapLevels(uploadTarget, requiredLevelCount);
// Mip generation grows/regenerates the level set on the GPU without marking any CPU // Mip generation grows/regenerates the level set on the GPU without marking any CPU
// level dirty (MarkStorageDirty(...,false) above). Bump the content version so the // level dirty (MarkStorageDirty(...,false) above). Bump the content version so the
// backend re-syncs: a cached sampled VkImageView built for the pre-generate level // backend re-syncs: a cached sampled VkImageView built for the pre-generate level
@@ -429,8 +433,10 @@ namespace MobileGL::MG_Impl::GLImpl {
const Int maxSamples = GetMaxSupportedTextureSamples(textureInternalFormat); const Int maxSamples = GetMaxSupportedTextureSamples(textureInternalFormat);
if (samples > maxSamples) { if (samples > maxSamples) {
// GL specifies INVALID_OPERATION - not INVALID_VALUE - when the sample count
// exceeds what the format supports, and the native Adreno driver agrees.
MG_State::pGLContext->RecordError( MG_State::pGLContext->RecordError(
ErrorCode::InvalidValue, ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>( MakeUnique<GenericErrorInfo>(
"MG_Impl/GLImpl", caller, "MG_Impl/GLImpl", caller,
std::format("Sample count {} exceeds the supported maximum {} for this texture format.", std::format("Sample count {} exceeds the supported maximum {} for this texture format.",
@@ -454,8 +460,56 @@ namespace MobileGL::MG_Impl::GLImpl {
textureObject->SetSamples(samples); textureObject->SetSamples(samples);
textureObject->SetFixedSampleLocations(fixedsamplelocations == GL_TRUE); textureObject->SetFixedSampleLocations(fixedsamplelocations == GL_TRUE);
textureMipmapObject->AllocateStorage(textureUploadTarget, 0, {{width, height, depth}, 0}); textureMipmapObject->AllocateStorage(textureUploadTarget, 0, {{width, height, depth}, 0});
// Multisample textures are single-level by definition, so a name that previously held a
// mip chain must not keep its tail now that AllocateStorage only grows.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, 1);
textureMipmapObject->MarkStorageDirty(textureUploadTarget, 0, false); textureMipmapObject->MarkStorageDirty(textureUploadTarget, 0, false);
} }
// Redefining level 0 of a texture that already had a base image drops the rest of the chain,
// which is exactly what AllocateLevel used to do implicitly for every level. Keeping that
// behaviour for level 0 - and only for level 0 - is what makes the grow-only change safe:
// any level-0 respecification leaves the chain in precisely the state it would have had
// before, while an upload to level N no longer destroys the levels beneath it.
//
// Why it has to be *every* level-0 respecification and not just a size change: Minecraft's
// Mipmap Levels setting rebuilds the block atlas at the SAME dimensions with a different
// level count. A size-only test would leave the old tail in place, and because Mojang
// terminates its chains with a 0x0 level the result is the zero-then-nonzero pattern that
// IsComplete() rejects (TextureObject.cpp) - whereupon DirectGLES skips syncing the texture
// entirely (Managers.cpp) and the atlas samples black.
//
// The "already has a base image" test is what lets the fix work at all: a level that was
// never written reads back as {0,0,0}, so building a chain top-down - upload level N first,
// then level 0 - must not discard the levels just uploaded. That ordering is what
// KHR-GL33.texture_repeat_mode does.
// Scoped to the respecified upload target only, which is what AllocateLevel already did.
// Cube maps keep six independent chains while reporting a single level count (face +X), so
// respecifying a face other than +X can leave the count longer than that face - but that
// asymmetry predates this change and widening the truncation to all six faces would destroy
// mip data for faces the application never touched. Left alone deliberately.
void DiscardMipmapChainOnBaseRespecification(MG_State::GLState::TextureObjectMipmap* texture,
TextureUploadTarget uploadTarget, Uint level) {
if (level != 0) return;
const IntVec3 existingBaseSize = texture->GetMipmapTexelSize(uploadTarget, 0);
const Bool hasExistingBaseImage =
existingBaseSize.x() > 0 && existingBaseSize.y() > 0 && existingBaseSize.z() > 0;
if (!hasExistingBaseImage) return;
texture->TruncateMipmapLevels(uploadTarget, 1);
}
// Compressed texture upload is not implemented yet. GL_NUM_COMPRESSED_TEXTURE_FORMATS
// reports 0, so every compressed internalformat is by definition unsupported and
// GL_INVALID_ENUM is the specified error - unlike THROW_UNIMPL_EXCEPTION, which unwinds
// a C++ exception through the C GL ABI and takes the process down.
void RecordUnsupportedCompressedFormat(const char* caller) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidEnum,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", caller,
"Compressed texture formats are not supported."));
}
} // namespace } // namespace
const SharedPtr<MG_State::GLState::ITextureObject>& GetTextureObjectByName(GLuint texture, const char* caller) { const SharedPtr<MG_State::GLState::ITextureObject>& GetTextureObjectByName(GLuint texture, const char* caller) {
@@ -493,8 +547,15 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
auto mipmapTexture = std::static_pointer_cast<MG_State::GLState::TextureObjectMipmap>(textureObject); auto mipmapTexture = std::static_pointer_cast<MG_State::GLState::TextureObjectMipmap>(textureObject);
if (level < 0 || static_cast<Uint>(level) >= mipmapTexture->GetMipmapLevelCount()) { if (level < 0) {
RecordClearTextureError(caller, ErrorCode::InvalidValue, RecordClearTextureError(caller, ErrorCode::InvalidValue,
std::format("Texture level {} is negative.", level));
return nullptr;
}
// ARB_clear_texture: clearing an image that was never defined by TexImage*/
// TexStorage* is INVALID_OPERATION, not INVALID_VALUE.
if (static_cast<Uint>(level) >= mipmapTexture->GetMipmapLevelCount()) {
RecordClearTextureError(caller, ErrorCode::InvalidOperation,
std::format("Texture level {} is not defined.", level)); std::format("Texture level {} is not defined.", level));
return nullptr; return nullptr;
} }
@@ -536,6 +597,12 @@ namespace MobileGL::MG_Impl::GLImpl {
return true; return true;
} }
// Writes the clear into the CPU shadow and marks the whole level dirty, exactly like
// TexSubImage*_State does. Shared limitation of the level-granular shadow sync: the
// shadow does not reflect GPU-side writes (FBO rendering, imageStore), so a PARTIAL
// clear of a GPU-written level re-uploads stale shadow bytes outside the region on
// the next sync. Full-level clears (glClearTexImage, or a sub-clear covering the
// level) rewrite the entire shadow and are always correct.
Bool ClearMipmapRegion(const SharedPtr<MG_State::GLState::TextureObjectMipmap>& textureObject, Bool ClearMipmapRegion(const SharedPtr<MG_State::GLState::TextureObjectMipmap>& textureObject,
TextureUploadTarget uploadTarget, GLint level, TextureUploadTarget uploadTarget, GLint level,
GLint xoffset, GLint yoffset, GLint zoffset, GLint xoffset, GLint yoffset, GLint zoffset,
@@ -1657,6 +1724,21 @@ namespace MobileGL::MG_Impl::GLImpl {
return; return;
} }
// RGTC is a 2D-only compression scheme, so a 3D target rejects it. This has to be tested on
// the raw enum: the RGTC formats resolve to plain R8/RG8/SNORM storage on the way in (see
// GLToMG's TextureEnumConverter), so once the internal format is converted there is nothing
// left to distinguish them from an ordinary one- or two-channel upload.
if ((textureUploadTarget == TextureUploadTarget::Texture3D ||
textureUploadTarget == TextureUploadTarget::ProxyTexture3D) &&
(internalformat == GL_COMPRESSED_RED_RGTC1 || internalformat == GL_COMPRESSED_SIGNED_RED_RGTC1 ||
internalformat == GL_COMPRESSED_RG_RGTC2 || internalformat == GL_COMPRESSED_SIGNED_RG_RGTC2)) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"RGTC compressed formats are invalid for 3D texture targets"));
return;
}
// TODO: GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the // TODO: GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the
// GL_PIXEL_UNPACK_BUFFER target and the buffer object's data store is currently mapped. // GL_PIXEL_UNPACK_BUFFER target and the buffer object's data store is currently mapped.
// GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the GL_PIXEL_UNPACK_BUFFER // GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the GL_PIXEL_UNPACK_BUFFER
@@ -1714,6 +1796,7 @@ namespace MobileGL::MG_Impl::GLImpl {
if (isProxy) { if (isProxy) {
MGLOG_D("%s: isProxy = true, not allocating", __func__); MGLOG_D("%s: isProxy = true, not allocating", __func__);
} else { } else {
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, height, depth}, internalBytes}); textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, height, depth}, internalBytes});
} }
@@ -1841,6 +1924,7 @@ namespace MobileGL::MG_Impl::GLImpl {
MGLOG_D("%s: isProxy = true, not allocating", __func__); MGLOG_D("%s: isProxy = true, not allocating", __func__);
} else { } else {
MGLOG_D("%s: Allocating %d bytes at mip %d", __func__, internalBytes, level); MGLOG_D("%s: Allocating %d bytes at mip %d", __func__, internalBytes, level);
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level, textureMipmapObject->AllocateStorage(textureUploadTarget, level,
{{width, height, 1}, internalBytes}); {{width, height, 1}, internalBytes});
} }
@@ -1929,6 +2013,7 @@ namespace MobileGL::MG_Impl::GLImpl {
"Texture object here should always be an object with mipmap"); "Texture object here should always be an object with mipmap");
auto textureMipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get()); auto textureMipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
if (!isProxy) { if (!isProxy) {
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, 1, 1}, internalBytes}); textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, 1, 1}, internalBytes});
} }
@@ -2581,7 +2666,13 @@ namespace MobileGL::MG_Impl::GLImpl {
} }
void GetCompressedTexImage_State(GLenum target, GLint level, void* img) { void GetCompressedTexImage_State(GLenum target, GLint level, void* img) {
// TODO: implement // TODO: implement compressed readback. Reporting success while writing nothing hands
// the caller stale memory with GL_NO_ERROR; no texture can be compressed yet, and GL
// specifies GL_INVALID_OPERATION when the bound level is not compressed.
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"Texture level is not stored in a compressed format."));
} }
void GenTextures_State(GLsizei n, GLuint* textures) { void GenTextures_State(GLsizei n, GLuint* textures) {
@@ -2768,20 +2859,20 @@ namespace MobileGL::MG_Impl::GLImpl {
void CompressedTexSubImage3D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLint zoffset, void CompressedTexSubImage3D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLint zoffset,
GLsizei width, GLsizei height, GLsizei depth, GLenum format, GLsizei imageSize, GLsizei width, GLsizei height, GLsizei depth, GLenum format, GLsizei imageSize,
const void* data) { const void* data) {
// TODO: implement // TODO: implement compressed upload - see CompressedTexImage2D_State.
THROW_UNIMPL_EXCEPTION; RecordUnsupportedCompressedFormat(__func__);
} }
void CompressedTexSubImage2D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLsizei width, void CompressedTexSubImage2D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLsizei width,
GLsizei height, GLenum format, GLsizei imageSize, const void* data) { GLsizei height, GLenum format, GLsizei imageSize, const void* data) {
// TODO: implement // TODO: implement compressed upload - see CompressedTexImage2D_State.
THROW_UNIMPL_EXCEPTION; RecordUnsupportedCompressedFormat(__func__);
} }
void CompressedTexSubImage1D_State(GLenum target, GLint level, GLint xoffset, GLsizei width, GLenum format, void CompressedTexSubImage1D_State(GLenum target, GLint level, GLint xoffset, GLsizei width, GLenum format,
GLsizei imageSize, const void* data) { GLsizei imageSize, const void* data) {
// TODO: implement // TODO: implement compressed upload - see CompressedTexImage2D_State.
THROW_UNIMPL_EXCEPTION; RecordUnsupportedCompressedFormat(__func__);
} }
void CompressedTexImage3D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height, void CompressedTexImage3D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height,
@@ -2791,8 +2882,8 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget); auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return; if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement // TODO: implement compressed upload - see CompressedTexImage2D_State.
THROW_UNIMPL_EXCEPTION; RecordUnsupportedCompressedFormat(__func__);
} }
void CompressedTexImage2D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height, void CompressedTexImage2D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height,
@@ -2802,8 +2893,11 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget); auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return; if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement // TODO: implement compressed upload. Until then report the spec error for an
THROW_UNIMPL_EXCEPTION; // unsupported compressed format rather than throwing - a C++ exception unwinding
// through the C GL ABI is a hard crash for the caller, while GL_INVALID_ENUM is
// exactly what GL_NUM_COMPRESSED_TEXTURE_FORMATS == 0 promises.
RecordUnsupportedCompressedFormat(__func__);
} }
void CompressedTexImage1D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLint border, void CompressedTexImage1D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLint border,
@@ -2813,8 +2907,8 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget); auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return; if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement // TODO: implement compressed upload - see CompressedTexImage2D_State.
THROW_UNIMPL_EXCEPTION; RecordUnsupportedCompressedFormat(__func__);
} }
void BindTexture_State(GLenum target, GLuint texture) { void BindTexture_State(GLenum target, GLuint texture) {
@@ -3160,6 +3254,9 @@ namespace MobileGL::MG_Impl::GLImpl {
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, 1, 1}, byteSize}); textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, 1, 1}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false); textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
} }
// Immutable storage defines exactly `levels` levels; AllocateStorage only grows, so a
// longer pre-existing chain has to be dropped explicitly.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels)); textureObject->SetImmutableLevels(static_cast<Uint>(levels));
} }
@@ -3212,6 +3309,8 @@ namespace MobileGL::MG_Impl::GLImpl {
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, levelHeight, 1}, byteSize}); textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, levelHeight, 1}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false); textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
} }
// See TextureStorage1D.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels)); textureObject->SetImmutableLevels(static_cast<Uint>(levels));
} }
@@ -3264,6 +3363,8 @@ namespace MobileGL::MG_Impl::GLImpl {
{{levelWidth, levelHeight, levelDepth}, byteSize}); {{levelWidth, levelHeight, levelDepth}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false); textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
} }
// See TextureStorage1D.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels)); textureObject->SetImmutableLevels(static_cast<Uint>(levels));
} }
@@ -4025,8 +4126,21 @@ namespace MobileGL::MG_Impl::GLImpl {
void CopyTextureSubImage2D(GLuint texture, GLint level, GLint xoffset, GLint yoffset, GLint x, GLint y, void CopyTextureSubImage2D(GLuint texture, GLint level, GLint xoffset, GLint yoffset, GLint x, GLint y,
GLsizei width, GLsizei height) { GLsizei width, GLsizei height) {
auto textureObject = GetTextureObjectByName(texture, __func__); auto textureObject = GetTextureObjectByName(texture, __func__);
WithTemporarilyBoundNamedTexture(textureObject, [&](GLenum target) { if (!textureObject) return;
CopyTexSubImage2D_Backend(target, level, xoffset, yoffset, x, y, width, height); // GL 4.6 sec. 8.8: the 2D form only accepts these effective targets; cube maps must
// go through CopyTextureSubImage3D with the face as a layer.
const auto target = textureObject->GetTarget();
if (target != TextureTarget::Texture2D && target != TextureTarget::Texture1DArray &&
target != TextureTarget::TextureRectangle) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"CopyTextureSubImage2D requires a 2D, 1D-array, or "
"rectangle texture."));
return;
}
WithTemporarilyBoundNamedTexture(textureObject, [&](GLenum glTarget) {
CopyTexSubImage2D_Backend(glTarget, level, xoffset, yoffset, x, y, width, height);
}); });
} }
@@ -14,7 +14,8 @@
#include <MG_State/GLState/TextureState/TextureObjectStubs.h> #include <MG_State/GLState/TextureState/TextureObjectStubs.h>
namespace MobileGL::MG_Impl::GLImpl::TextureImpl { namespace MobileGL::MG_Impl::GLImpl::TextureImpl {
UniquePtr<ProxyTextureManager> pProxyTextureManager; // Leak-at-exit storage; see GlobalObjects.cpp.
UniquePtr<ProxyTextureManager>& pProxyTextureManager = *new UniquePtr<ProxyTextureManager>();
Bool IsProxyTextureTarget(TextureUploadTarget target) { Bool IsProxyTextureTarget(TextureUploadTarget target) {
switch (target) { switch (target) {
@@ -23,5 +23,5 @@ namespace MobileGL::MG_Impl::GLImpl::TextureImpl {
UnorderedMap<TextureUploadTarget, SharedPtr<MG_State::GLState::ITextureObject>> m_proxyTexturesMap; UnorderedMap<TextureUploadTarget, SharedPtr<MG_State::GLState::ITextureObject>> m_proxyTexturesMap;
}; };
extern UniquePtr<ProxyTextureManager> pProxyTextureManager; extern UniquePtr<ProxyTextureManager>& pProxyTextureManager;
} // namespace MobileGL::MG_Impl::GLImpl::TextureImpl } // namespace MobileGL::MG_Impl::GLImpl::TextureImpl
@@ -0,0 +1,156 @@
// MobileGL - MobileGL/MG_Impl/WGLImpl/Exporting/Definitions.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
// wingdi.h declares most wgl* entry points as WINGDIAPI (__declspec(dllimport)),
// which would reject our definitions. _GDI32_ is the SDK's "I am the module that
// implements these" switch: it turns WINGDIAPI into a plain declaration. It must
// be defined before the first windows.h inclusion in this translation unit.
#if defined(_WIN32) && !defined(_GDI32_)
#define _GDI32_ 1
#endif
#include <Includes.h>
#if defined(_WIN32)
#include "../WGLImpl.h"
namespace WGL = MobileGL::MG_Impl::WGLImpl;
// ---- Pixel-format entry points (gdi32 forwards ChoosePixelFormat/SetPixelFormat/
// ---- DescribePixelFormat/GetPixelFormat/SwapBuffers into these exports) ----
extern "C" int WINAPI wglChoosePixelFormat(HDC hdc, CONST PIXELFORMATDESCRIPTOR* ppfd) {
return WGL::ChoosePixelFormat(hdc, ppfd);
}
extern "C" int WINAPI wglDescribePixelFormat(HDC hdc, int iPixelFormat, UINT nBytes,
LPPIXELFORMATDESCRIPTOR ppfd) {
return WGL::DescribePixelFormat(hdc, iPixelFormat, nBytes, ppfd);
}
extern "C" int WINAPI wglGetPixelFormat(HDC hdc) {
return WGL::GetPixelFormat(hdc);
}
extern "C" BOOL WINAPI wglSetPixelFormat(HDC hdc, int iPixelFormat, CONST PIXELFORMATDESCRIPTOR* ppfd) {
return WGL::SetPixelFormat(hdc, iPixelFormat, ppfd);
}
extern "C" BOOL WINAPI wglSwapBuffers(HDC hdc) {
return WGL::SwapBuffers(hdc);
}
// ---- Context management ----
extern "C" HGLRC WINAPI wglCreateContext(HDC hdc) {
return WGL::CreateContext(hdc);
}
extern "C" HGLRC WINAPI wglCreateLayerContext(HDC hdc, int iLayerPlane) {
return iLayerPlane == 0 ? WGL::CreateContext(hdc) : nullptr;
}
extern "C" BOOL WINAPI wglCopyContext(HGLRC, HGLRC, UINT) {
MGLOG_W("wglCopyContext is not supported");
SetLastError(ERROR_NOT_SUPPORTED);
return FALSE;
}
extern "C" BOOL WINAPI wglDeleteContext(HGLRC hglrc) {
return WGL::DeleteContext(hglrc);
}
extern "C" HGLRC WINAPI wglGetCurrentContext(VOID) {
return WGL::GetCurrentContext();
}
extern "C" HDC WINAPI wglGetCurrentDC(VOID) {
return WGL::GetCurrentDC();
}
extern "C" BOOL WINAPI wglMakeCurrent(HDC hdc, HGLRC hglrc) {
return WGL::MakeCurrent(hdc, hglrc);
}
extern "C" BOOL WINAPI wglShareLists(HGLRC hglrcShare, HGLRC hglrcDest) {
return WGL::ShareLists(hglrcShare, hglrcDest);
}
// ---- Proc address ----
extern "C" PROC WINAPI wglGetProcAddress(LPCSTR lpszProc) {
return WGL::GetProcAddress(lpszProc);
}
extern "C" PROC WINAPI wglGetDefaultProcAddress(LPCSTR lpszProc) {
return WGL::GetProcAddress(lpszProc);
}
// ---- Layer planes and palettes (unsupported; overlay planes do not exist here) ----
extern "C" BOOL WINAPI wglDescribeLayerPlane(HDC, int, int, UINT, LPLAYERPLANEDESCRIPTOR) {
return FALSE;
}
extern "C" int WINAPI wglSetLayerPaletteEntries(HDC, int, int, int, CONST COLORREF*) {
return 0;
}
extern "C" int WINAPI wglGetLayerPaletteEntries(HDC, int, int, int, COLORREF*) {
return 0;
}
extern "C" BOOL WINAPI wglRealizeLayerPalette(HDC, int, BOOL) {
return FALSE;
}
extern "C" BOOL WINAPI wglSwapLayerBuffers(HDC hdc, UINT fuPlanes) {
if (fuPlanes & WGL_SWAP_MAIN_PLANE) {
return WGL::SwapBuffers(hdc);
}
return FALSE;
}
extern "C" DWORD WINAPI wglSwapMultipleBuffers(UINT n, CONST WGLSWAP* ps) {
if (!ps) {
return 0;
}
DWORD swapped = 0;
for (UINT i = 0; i < n; ++i) {
if (WGL::SwapBuffers(ps[i].hdc)) {
++swapped;
}
}
return swapped;
}
// ---- Font rendering (legacy immediate-mode feature; not supported) ----
extern "C" BOOL WINAPI wglUseFontBitmapsA(HDC, DWORD, DWORD, DWORD) {
MGLOG_W("wglUseFontBitmapsA is not supported");
return FALSE;
}
extern "C" BOOL WINAPI wglUseFontBitmapsW(HDC, DWORD, DWORD, DWORD) {
MGLOG_W("wglUseFontBitmapsW is not supported");
return FALSE;
}
extern "C" BOOL WINAPI wglUseFontOutlinesA(HDC, DWORD, DWORD, DWORD, FLOAT, FLOAT, int,
LPGLYPHMETRICSFLOAT) {
MGLOG_W("wglUseFontOutlinesA is not supported");
return FALSE;
}
extern "C" BOOL WINAPI wglUseFontOutlinesW(HDC, DWORD, DWORD, DWORD, FLOAT, FLOAT, int,
LPGLYPHMETRICSFLOAT) {
MGLOG_W("wglUseFontOutlinesW is not supported");
return FALSE;
}
#endif // _WIN32
@@ -0,0 +1,30 @@
; MobileGL WGL exports. The wgl* entry points are defined without
; __declspec(dllexport) because wingdi.h pre-declares them (with _GDI32_ they
; become plain declarations, and MSVC rejects adding dllexport afterwards),
; so this .def file is what actually exports them from the DLL.
EXPORTS
wglChoosePixelFormat
wglCopyContext
wglCreateContext
wglCreateLayerContext
wglDeleteContext
wglDescribeLayerPlane
wglDescribePixelFormat
wglGetCurrentContext
wglGetCurrentDC
wglGetDefaultProcAddress
wglGetLayerPaletteEntries
wglGetPixelFormat
wglGetProcAddress
wglMakeCurrent
wglRealizeLayerPalette
wglSetLayerPaletteEntries
wglSetPixelFormat
wglShareLists
wglSwapBuffers
wglSwapLayerBuffers
wglSwapMultipleBuffers
wglUseFontBitmapsA
wglUseFontBitmapsW
wglUseFontOutlinesA
wglUseFontOutlinesW
+732
View File
@@ -0,0 +1,732 @@
// MobileGL - MobileGL/MG_Impl/WGLImpl/WGLImpl.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "WGLImpl.h"
#if defined(_WIN32)
#include "../EGLImpl/EGLImpl.h"
#include "../GetProcAddress.h"
#include <Init.h>
namespace MobileGL::MG_Impl::WGLImpl {
namespace {
// WGL_ARB_pixel_format
constexpr int WGL_NUMBER_PIXEL_FORMATS_ARB = 0x2000;
constexpr int WGL_DRAW_TO_WINDOW_ARB = 0x2001;
constexpr int WGL_DRAW_TO_BITMAP_ARB = 0x2002;
constexpr int WGL_ACCELERATION_ARB = 0x2003;
constexpr int WGL_NEED_PALETTE_ARB = 0x2004;
constexpr int WGL_NEED_SYSTEM_PALETTE_ARB = 0x2005;
constexpr int WGL_SWAP_LAYER_BUFFERS_ARB = 0x2006;
constexpr int WGL_SWAP_METHOD_ARB = 0x2007;
constexpr int WGL_NUMBER_OVERLAYS_ARB = 0x2008;
constexpr int WGL_NUMBER_UNDERLAYS_ARB = 0x2009;
constexpr int WGL_TRANSPARENT_ARB = 0x200A;
constexpr int WGL_SHARE_DEPTH_ARB = 0x200C;
constexpr int WGL_SHARE_STENCIL_ARB = 0x200D;
constexpr int WGL_SHARE_ACCUM_ARB = 0x200E;
constexpr int WGL_SUPPORT_GDI_ARB = 0x200F;
constexpr int WGL_SUPPORT_OPENGL_ARB = 0x2010;
constexpr int WGL_DOUBLE_BUFFER_ARB = 0x2011;
constexpr int WGL_STEREO_ARB = 0x2012;
constexpr int WGL_PIXEL_TYPE_ARB = 0x2013;
constexpr int WGL_COLOR_BITS_ARB = 0x2014;
constexpr int WGL_RED_BITS_ARB = 0x2015;
constexpr int WGL_RED_SHIFT_ARB = 0x2016;
constexpr int WGL_GREEN_BITS_ARB = 0x2017;
constexpr int WGL_GREEN_SHIFT_ARB = 0x2018;
constexpr int WGL_BLUE_BITS_ARB = 0x2019;
constexpr int WGL_BLUE_SHIFT_ARB = 0x201A;
constexpr int WGL_ALPHA_BITS_ARB = 0x201B;
constexpr int WGL_ALPHA_SHIFT_ARB = 0x201C;
constexpr int WGL_ACCUM_BITS_ARB = 0x201D;
constexpr int WGL_ACCUM_RED_BITS_ARB = 0x201E;
constexpr int WGL_ACCUM_GREEN_BITS_ARB = 0x201F;
constexpr int WGL_ACCUM_BLUE_BITS_ARB = 0x2020;
constexpr int WGL_ACCUM_ALPHA_BITS_ARB = 0x2021;
constexpr int WGL_DEPTH_BITS_ARB = 0x2022;
constexpr int WGL_STENCIL_BITS_ARB = 0x2023;
constexpr int WGL_AUX_BUFFERS_ARB = 0x2024;
constexpr int WGL_NO_ACCELERATION_ARB = 0x2025;
constexpr int WGL_FULL_ACCELERATION_ARB = 0x2027;
constexpr int WGL_SWAP_EXCHANGE_ARB = 0x2028;
constexpr int WGL_TYPE_RGBA_ARB = 0x202B;
// WGL_ARB_multisample
constexpr int WGL_SAMPLE_BUFFERS_ARB = 0x2041;
constexpr int WGL_SAMPLES_ARB = 0x2042;
// WGL_ARB_create_context / _profile / _no_error
constexpr int WGL_CONTEXT_MAJOR_VERSION_ARB = 0x2091;
constexpr int WGL_CONTEXT_MINOR_VERSION_ARB = 0x2092;
constexpr int WGL_CONTEXT_LAYER_PLANE_ARB = 0x2093;
constexpr int WGL_CONTEXT_FLAGS_ARB = 0x2094;
constexpr int WGL_CONTEXT_PROFILE_MASK_ARB = 0x9126;
constexpr int WGL_CONTEXT_DEBUG_BIT_ARB = 0x0001;
constexpr int WGL_CONTEXT_FORWARD_COMPATIBLE_BIT_ARB = 0x0002;
constexpr int WGL_CONTEXT_CORE_PROFILE_BIT_ARB = 0x00000001;
constexpr int WGL_CONTEXT_COMPATIBILITY_PROFILE_BIT_ARB = 0x00000002;
constexpr int WGL_CONTEXT_OPENGL_NO_ERROR_ARB = 0x31B3;
constexpr DWORD ERROR_INVALID_VERSION_ARB = 0x2095;
constexpr DWORD ERROR_INVALID_PROFILE_ARB = 0x2096;
struct PixelFormatInfo {
GLint AlphaBits;
GLint DepthBits;
GLint StencilBits;
};
// Mirrors the two EGLState configs (RGBA8 + depth24, stencil 8 / stencil 0).
constexpr PixelFormatInfo kPixelFormats[] = {
{8, 24, 8},
{8, 24, 0},
};
constexpr int kPixelFormatCount = static_cast<int>(std::size(kPixelFormats));
struct ContextObject {
EGLDisplay Display = EGL_NO_DISPLAY;
EGLConfig Config = nullptr;
EGLContext Context = EGL_NO_CONTEXT;
};
struct WindowSurface {
EGLDisplay Display = EGL_NO_DISPLAY;
EGLSurface Surface = EGL_NO_SURFACE;
Uint32 Width = 0;
Uint32 Height = 0;
};
std::recursive_mutex& RegistryMutex() {
static auto* mutex = new std::recursive_mutex();
return *mutex;
}
UnorderedMap<HGLRC, ContextObject>& Contexts() {
static auto* contexts = new UnorderedMap<HGLRC, ContextObject>();
return *contexts;
}
UnorderedMap<HWND, WindowSurface>& WindowSurfaces() {
static auto* surfaces = new UnorderedMap<HWND, WindowSurface>();
return *surfaces;
}
UnorderedMap<HWND, int>& WindowPixelFormats() {
static auto* formats = new UnorderedMap<HWND, int>();
return *formats;
}
Uint64& NextContextHandle() {
static auto* handle = new Uint64(0x10000);
return *handle;
}
struct ThreadCurrent {
HDC DC = nullptr;
HGLRC Context = nullptr;
};
thread_local ThreadCurrent t_current;
Int& SwapIntervalShadow() {
static auto* interval = new Int(1);
return *interval;
}
void EnsureInitialized() {
// Initialize() loads backend libraries and glslang, which must not run
// under the loader lock; first WGL call is the earliest safe moment.
// MobileGL::EnsureInitialized (not a local once_flag) so a fresh init
// can follow a full teardown from the last eglTerminate.
MobileGL::EnsureInitialized();
}
EGLDisplay EnsureDisplay() {
EnsureInitialized();
EGLDisplay display = EGLImpl::GetDisplay(EGL_DEFAULT_DISPLAY);
if (display == EGL_NO_DISPLAY) {
return EGL_NO_DISPLAY;
}
if (!EGLImpl::Initialize(display, nullptr, nullptr)) {
return EGL_NO_DISPLAY;
}
return display;
}
HGLRC EncodeContext(Uint64 handle) {
return reinterpret_cast<HGLRC>(static_cast<SizeT>(handle));
}
ContextObject* TryGetContext(HGLRC hglrc) {
auto& contexts = Contexts();
auto it = contexts.find(hglrc);
return it == contexts.end() ? nullptr : &it->second;
}
const PixelFormatInfo& PixelFormatForWindow(HWND hwnd) {
auto& formats = WindowPixelFormats();
auto it = formats.find(hwnd);
int index = it == formats.end() ? 1 : it->second;
if (index < 1 || index > kPixelFormatCount) {
index = 1;
}
return kPixelFormats[index - 1];
}
Bool QueryClientSize(HWND hwnd, Uint32& width, Uint32& height) {
RECT rect{};
if (!GetClientRect(hwnd, &rect)) {
return false;
}
width = static_cast<Uint32>(std::max<LONG>(rect.right - rect.left, 1));
height = static_cast<Uint32>(std::max<LONG>(rect.bottom - rect.top, 1));
return true;
}
// The backends never query the HWND client size themselves; the WGL layer
// owns size discovery and pushes changes through the internal resize hook
// (same contract as the macOS CGL layer).
void SyncSurfaceSize(HWND hwnd, WindowSurface& surface) {
Uint32 width = 0;
Uint32 height = 0;
if (!QueryClientSize(hwnd, width, height)) {
return;
}
if (width == surface.Width && height == surface.Height) {
return;
}
if (EGLImpl::ResizePlatformWindowSurface(surface.Display, surface.Surface,
static_cast<EGLint>(width), static_cast<EGLint>(height))) {
surface.Width = width;
surface.Height = height;
}
}
WindowSurface* EnsureWindowSurface(HWND hwnd, const ContextObject& context) {
auto& surfaces = WindowSurfaces();
auto it = surfaces.find(hwnd);
if (it != surfaces.end()) {
SyncSurfaceSize(hwnd, it->second);
return &it->second;
}
Uint32 width = 0;
Uint32 height = 0;
if (!QueryClientSize(hwnd, width, height)) {
MGLOG_E("wgl: GetClientRect failed for HWND %p", hwnd);
return nullptr;
}
const EGLAttrib attribs[] = {
EGL_WIDTH, static_cast<EGLAttrib>(width),
EGL_HEIGHT, static_cast<EGLAttrib>(height),
EGL_NONE,
};
EGLSurface surface =
EGLImpl::CreatePlatformWindowSurface(context.Display, context.Config, hwnd, attribs);
if (surface == EGL_NO_SURFACE) {
MGLOG_E("wgl: failed to create window surface for HWND %p (%ux%u)", hwnd, width, height);
return nullptr;
}
WindowSurface record;
record.Display = context.Display;
record.Surface = surface;
record.Width = width;
record.Height = height;
auto [inserted, _] = surfaces.emplace(hwnd, record);
return &inserted->second;
}
HGLRC CreateContextFromEGLAttribs(HDC hdc, HGLRC share, const EGLint* contextAttribs) {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
EGLDisplay display = EnsureDisplay();
if (display == EGL_NO_DISPLAY) {
MGLOG_E("wgl: no EGL display");
return nullptr;
}
EGLImpl::BindAPI(EGL_OPENGL_API);
EGLContext shareContext = EGL_NO_CONTEXT;
if (share) {
auto* shareObject = TryGetContext(share);
if (!shareObject) {
SetLastError(ERROR_INVALID_HANDLE);
return nullptr;
}
shareContext = shareObject->Context;
}
HWND hwnd = WindowFromDC(hdc);
const PixelFormatInfo& pixelFormat = PixelFormatForWindow(hwnd);
const EGLint configAttribs[] = {
EGL_RED_SIZE, 8,
EGL_GREEN_SIZE, 8,
EGL_BLUE_SIZE, 8,
EGL_ALPHA_SIZE, pixelFormat.AlphaBits,
EGL_DEPTH_SIZE, pixelFormat.DepthBits,
EGL_STENCIL_SIZE, pixelFormat.StencilBits,
EGL_SURFACE_TYPE, EGL_WINDOW_BIT | EGL_PBUFFER_BIT,
EGL_RENDERABLE_TYPE, EGL_OPENGL_BIT,
EGL_NONE,
};
EGLConfig config = nullptr;
EGLint configCount = 0;
if (!EGLImpl::ChooseConfig(display, configAttribs, &config, 1, &configCount) || configCount <= 0) {
MGLOG_E("wgl: eglChooseConfig failed");
return nullptr;
}
EGLContext eglContext = EGLImpl::CreateContext(display, config, shareContext, contextAttribs);
if (eglContext == EGL_NO_CONTEXT) {
MGLOG_E("wgl: eglCreateContext failed");
return nullptr;
}
ContextObject object;
object.Display = display;
object.Config = config;
object.Context = eglContext;
const auto handle = EncodeContext(NextContextHandle()++);
Contexts()[handle] = object;
MGLOG_I("wgl: created context %p (EGL context %p)", handle, eglContext);
return handle;
}
// ---- WGL extension entry points (resolved via wglGetProcAddress only) ----
const char* WINAPI Ext_GetExtensionsStringARB(HDC) {
return "WGL_ARB_create_context WGL_ARB_create_context_no_error WGL_ARB_create_context_profile "
"WGL_ARB_extensions_string WGL_ARB_pixel_format WGL_EXT_extensions_string WGL_EXT_swap_control";
}
const char* WINAPI Ext_GetExtensionsStringEXT() {
return Ext_GetExtensionsStringARB(nullptr);
}
HGLRC WINAPI Ext_CreateContextAttribsARB(HDC hdc, HGLRC hShareContext, const int* attribList) {
EnsureInitialized();
int major = 1;
int minor = 0;
int profileMask = 0;
int flags = 0;
if (attribList) {
for (SizeT i = 0; attribList[i] != 0; i += 2) {
const int attrib = attribList[i];
const int value = attribList[i + 1];
switch (attrib) {
case WGL_CONTEXT_MAJOR_VERSION_ARB:
major = value;
break;
case WGL_CONTEXT_MINOR_VERSION_ARB:
minor = value;
break;
case WGL_CONTEXT_PROFILE_MASK_ARB:
profileMask = value;
break;
case WGL_CONTEXT_FLAGS_ARB:
flags = value;
break;
case WGL_CONTEXT_LAYER_PLANE_ARB:
if (value != 0) {
SetLastError(ERROR_INVALID_PARAMETER);
return nullptr;
}
break;
case WGL_CONTEXT_OPENGL_NO_ERROR_ARB:
// Accepted and ignored: MobileGL always validates.
break;
default:
MGLOG_D("wglCreateContextAttribsARB: ignoring attrib 0x%04x = 0x%x", attrib, value);
break;
}
}
}
if (major < 1 || (profileMask & ~(WGL_CONTEXT_CORE_PROFILE_BIT_ARB |
WGL_CONTEXT_COMPATIBILITY_PROFILE_BIT_ARB))) {
SetLastError(profileMask ? ERROR_INVALID_PROFILE_ARB : ERROR_INVALID_VERSION_ARB);
return nullptr;
}
Vector<EGLint> attribs = {
EGL_CONTEXT_MAJOR_VERSION, major,
EGL_CONTEXT_MINOR_VERSION, minor,
};
const Bool wantsCompat = (profileMask & WGL_CONTEXT_COMPATIBILITY_PROFILE_BIT_ARB) != 0;
if (major > 3 || (major == 3 && minor >= 2) || profileMask != 0) {
attribs.push_back(EGL_CONTEXT_OPENGL_PROFILE_MASK);
attribs.push_back(wantsCompat ? EGL_CONTEXT_OPENGL_COMPATIBILITY_PROFILE_BIT
: EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT);
}
if (flags & WGL_CONTEXT_FORWARD_COMPATIBLE_BIT_ARB) {
attribs.push_back(EGL_CONTEXT_OPENGL_FORWARD_COMPATIBLE);
attribs.push_back(EGL_TRUE);
}
if (flags & WGL_CONTEXT_DEBUG_BIT_ARB) {
attribs.push_back(EGL_CONTEXT_OPENGL_DEBUG);
attribs.push_back(EGL_TRUE);
}
attribs.push_back(EGL_NONE);
return CreateContextFromEGLAttribs(hdc, hShareContext, attribs.data());
}
BOOL WINAPI Ext_SwapIntervalEXT(int interval) {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
EGLDisplay display = EnsureDisplay();
if (display == EGL_NO_DISPLAY) {
return FALSE;
}
if (interval < 0) {
// Adaptive vsync is not supported; clamp to regular vsync.
interval = 1;
}
EGLImpl::SwapInterval(display, interval);
SwapIntervalShadow() = interval;
return TRUE;
}
int WINAPI Ext_GetSwapIntervalEXT() {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
return SwapIntervalShadow();
}
int PixelFormatAttribValue(int format, int attrib) {
const PixelFormatInfo& info = kPixelFormats[format - 1];
switch (attrib) {
case WGL_NUMBER_PIXEL_FORMATS_ARB:
return kPixelFormatCount;
case WGL_SUPPORT_OPENGL_ARB:
case WGL_DRAW_TO_WINDOW_ARB:
case WGL_DOUBLE_BUFFER_ARB:
return 1;
case WGL_ACCELERATION_ARB:
return WGL_FULL_ACCELERATION_ARB;
case WGL_PIXEL_TYPE_ARB:
return WGL_TYPE_RGBA_ARB;
case WGL_COLOR_BITS_ARB:
return 32;
case WGL_RED_BITS_ARB:
case WGL_GREEN_BITS_ARB:
case WGL_BLUE_BITS_ARB:
return 8;
case WGL_RED_SHIFT_ARB:
return 16;
case WGL_GREEN_SHIFT_ARB:
return 8;
case WGL_BLUE_SHIFT_ARB:
return 0;
case WGL_ALPHA_BITS_ARB:
return info.AlphaBits;
case WGL_ALPHA_SHIFT_ARB:
return 24;
case WGL_DEPTH_BITS_ARB:
return info.DepthBits;
case WGL_STENCIL_BITS_ARB:
return info.StencilBits;
case WGL_SWAP_METHOD_ARB:
return WGL_SWAP_EXCHANGE_ARB;
case WGL_DRAW_TO_BITMAP_ARB:
case WGL_NEED_PALETTE_ARB:
case WGL_NEED_SYSTEM_PALETTE_ARB:
case WGL_SWAP_LAYER_BUFFERS_ARB:
case WGL_NUMBER_OVERLAYS_ARB:
case WGL_NUMBER_UNDERLAYS_ARB:
case WGL_TRANSPARENT_ARB:
case WGL_SHARE_DEPTH_ARB:
case WGL_SHARE_STENCIL_ARB:
case WGL_SHARE_ACCUM_ARB:
case WGL_SUPPORT_GDI_ARB:
case WGL_STEREO_ARB:
case WGL_ACCUM_BITS_ARB:
case WGL_ACCUM_RED_BITS_ARB:
case WGL_ACCUM_GREEN_BITS_ARB:
case WGL_ACCUM_BLUE_BITS_ARB:
case WGL_ACCUM_ALPHA_BITS_ARB:
case WGL_AUX_BUFFERS_ARB:
case WGL_SAMPLE_BUFFERS_ARB:
case WGL_SAMPLES_ARB:
default:
return 0;
}
}
BOOL WINAPI Ext_GetPixelFormatAttribivARB(HDC, int iPixelFormat, int iLayerPlane, UINT nAttributes,
const int* piAttributes, int* piValues) {
if (iLayerPlane != 0 || !piAttributes || !piValues) {
return FALSE;
}
// Format 0 is only valid for WGL_NUMBER_PIXEL_FORMATS_ARB queries.
if (iPixelFormat < 0 || iPixelFormat > kPixelFormatCount) {
return FALSE;
}
const int format = iPixelFormat == 0 ? 1 : iPixelFormat;
for (UINT i = 0; i < nAttributes; ++i) {
piValues[i] = PixelFormatAttribValue(format, piAttributes[i]);
}
return TRUE;
}
BOOL WINAPI Ext_GetPixelFormatAttribfvARB(HDC hdc, int iPixelFormat, int iLayerPlane, UINT nAttributes,
const int* piAttributes, FLOAT* pfValues) {
if (!pfValues) {
return FALSE;
}
Vector<int> values(nAttributes);
if (!Ext_GetPixelFormatAttribivARB(hdc, iPixelFormat, iLayerPlane, nAttributes, piAttributes,
values.data())) {
return FALSE;
}
for (UINT i = 0; i < nAttributes; ++i) {
pfValues[i] = static_cast<FLOAT>(values[i]);
}
return TRUE;
}
BOOL WINAPI Ext_ChoosePixelFormatARB(HDC, const int* piAttribIList, const FLOAT*, UINT nMaxFormats,
int* piFormats, UINT* nNumFormats) {
if (!piFormats || !nNumFormats) {
return FALSE;
}
int wantedStencil = 0;
if (piAttribIList) {
for (SizeT i = 0; piAttribIList[i] != 0; i += 2) {
if (piAttribIList[i] == WGL_STENCIL_BITS_ARB) {
wantedStencil = piAttribIList[i + 1];
}
}
}
UINT count = 0;
const int preferred = wantedStencil > 0 ? 1 : 2;
const int fallback = wantedStencil > 0 ? 2 : 1;
if (count < nMaxFormats) {
piFormats[count++] = preferred;
}
if (count < nMaxFormats) {
piFormats[count++] = fallback;
}
*nNumFormats = count;
return TRUE;
}
struct WGLExtensionProc {
const char* Name;
PROC Proc;
};
const WGLExtensionProc kWGLExtensionProcs[] = {
{"wglGetExtensionsStringARB", reinterpret_cast<PROC>(Ext_GetExtensionsStringARB)},
{"wglGetExtensionsStringEXT", reinterpret_cast<PROC>(Ext_GetExtensionsStringEXT)},
{"wglCreateContextAttribsARB", reinterpret_cast<PROC>(Ext_CreateContextAttribsARB)},
{"wglSwapIntervalEXT", reinterpret_cast<PROC>(Ext_SwapIntervalEXT)},
{"wglGetSwapIntervalEXT", reinterpret_cast<PROC>(Ext_GetSwapIntervalEXT)},
{"wglGetPixelFormatAttribivARB", reinterpret_cast<PROC>(Ext_GetPixelFormatAttribivARB)},
{"wglGetPixelFormatAttribfvARB", reinterpret_cast<PROC>(Ext_GetPixelFormatAttribfvARB)},
{"wglChoosePixelFormatARB", reinterpret_cast<PROC>(Ext_ChoosePixelFormatARB)},
};
} // namespace
int ChoosePixelFormat(HDC hdc, const PIXELFORMATDESCRIPTOR* pfd) {
EnsureInitialized();
MGLOG_D("wglChoosePixelFormat(hdc=%p)", hdc);
// Format 1 (RGBA8 + depth24/stencil8) satisfies every request; a format
// exceeding the asked-for capabilities is a legal ChoosePixelFormat answer.
(void)pfd;
return 1;
}
int DescribePixelFormat(HDC hdc, int format, UINT size, PIXELFORMATDESCRIPTOR* pfd) {
EnsureInitialized();
MGLOG_D("wglDescribePixelFormat(hdc=%p, format=%d)", hdc, format);
if (!pfd) {
return kPixelFormatCount;
}
if (size < sizeof(PIXELFORMATDESCRIPTOR) || format < 1 || format > kPixelFormatCount) {
return 0;
}
const PixelFormatInfo& info = kPixelFormats[format - 1];
std::memset(pfd, 0, sizeof(PIXELFORMATDESCRIPTOR));
pfd->nSize = sizeof(PIXELFORMATDESCRIPTOR);
pfd->nVersion = 1;
pfd->dwFlags = PFD_DRAW_TO_WINDOW | PFD_SUPPORT_OPENGL | PFD_DOUBLEBUFFER | PFD_SWAP_EXCHANGE
#if defined(PFD_SUPPORT_COMPOSITION)
| PFD_SUPPORT_COMPOSITION
#endif
;
pfd->iPixelType = PFD_TYPE_RGBA;
pfd->cColorBits = 32;
pfd->cRedBits = 8;
pfd->cRedShift = 16;
pfd->cGreenBits = 8;
pfd->cGreenShift = 8;
pfd->cBlueBits = 8;
pfd->cBlueShift = 0;
pfd->cAlphaBits = static_cast<BYTE>(info.AlphaBits);
pfd->cAlphaShift = 24;
pfd->cDepthBits = static_cast<BYTE>(info.DepthBits);
pfd->cStencilBits = static_cast<BYTE>(info.StencilBits);
pfd->iLayerType = PFD_MAIN_PLANE;
return kPixelFormatCount;
}
int GetPixelFormat(HDC hdc) {
EnsureInitialized();
HWND hwnd = WindowFromDC(hdc);
if (!hwnd) {
return 0;
}
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto& formats = WindowPixelFormats();
auto it = formats.find(hwnd);
return it == formats.end() ? 0 : it->second;
}
BOOL SetPixelFormat(HDC hdc, int format, const PIXELFORMATDESCRIPTOR*) {
EnsureInitialized();
MGLOG_D("wglSetPixelFormat(hdc=%p, format=%d)", hdc, format);
if (format < 1 || format > kPixelFormatCount) {
SetLastError(ERROR_INVALID_PARAMETER);
return FALSE;
}
HWND hwnd = WindowFromDC(hdc);
if (!hwnd) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
WindowPixelFormats()[hwnd] = format;
return TRUE;
}
BOOL SwapBuffers(HDC hdc) {
HWND hwnd = WindowFromDC(hdc);
if (!hwnd) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto& surfaces = WindowSurfaces();
auto it = surfaces.find(hwnd);
if (it == surfaces.end()) {
MGLOG_W("wglSwapBuffers: no surface for HWND %p", hwnd);
return FALSE;
}
SyncSurfaceSize(hwnd, it->second);
return EGLImpl::SwapBuffers(it->second.Display, it->second.Surface) == EGL_TRUE ? TRUE : FALSE;
}
HGLRC CreateContext(HDC hdc) {
EnsureInitialized();
MGLOG_I("wglCreateContext(hdc=%p)", hdc);
// A legacy WGL context is a compatibility-profile context; MobileGL keys
// its relaxed-semantics mode off the explicit compatibility bit.
const EGLint attribs[] = {
EGL_CONTEXT_MAJOR_VERSION, 3,
EGL_CONTEXT_MINOR_VERSION, 3,
EGL_CONTEXT_OPENGL_PROFILE_MASK, EGL_CONTEXT_OPENGL_COMPATIBILITY_PROFILE_BIT,
EGL_NONE,
};
return CreateContextFromEGLAttribs(hdc, nullptr, attribs);
}
BOOL DeleteContext(HGLRC hglrc) {
EnsureInitialized();
MGLOG_I("wglDeleteContext(%p)", hglrc);
if (t_current.Context == hglrc) {
MakeCurrent(nullptr, nullptr);
}
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto* object = TryGetContext(hglrc);
if (!object) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
if (object->Context != EGL_NO_CONTEXT) {
EGLImpl::DestroyContext(object->Display, object->Context);
}
Contexts().erase(hglrc);
return TRUE;
}
BOOL MakeCurrent(HDC hdc, HGLRC hglrc) {
EnsureInitialized();
MGLOG_D("wglMakeCurrent(hdc=%p, hglrc=%p)", hdc, hglrc);
if (!hglrc) {
if (!t_current.Context) {
t_current = {};
return TRUE;
}
const EGLBoolean released =
EGLImpl::MakeCurrent(EGL_NO_DISPLAY, EGL_NO_SURFACE, EGL_NO_SURFACE, EGL_NO_CONTEXT);
t_current = {};
return released == EGL_TRUE ? TRUE : FALSE;
}
HWND hwnd = WindowFromDC(hdc);
if (!hwnd) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto* object = TryGetContext(hglrc);
if (!object) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
WindowSurface* surface = EnsureWindowSurface(hwnd, *object);
if (!surface) {
return FALSE;
}
if (!EGLImpl::MakeCurrent(object->Display, surface->Surface, surface->Surface, object->Context)) {
MGLOG_E("wglMakeCurrent: eglMakeCurrent failed (hdc=%p, hglrc=%p)", hdc, hglrc);
return FALSE;
}
t_current = {hdc, hglrc};
return TRUE;
}
HGLRC GetCurrentContext() {
return t_current.Context;
}
HDC GetCurrentDC() {
return t_current.DC;
}
BOOL ShareLists(HGLRC hglrcShare, HGLRC hglrcDest) {
// All MobileGL contexts alias one global GL object namespace, so every
// pair of contexts already shares.
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
if (!TryGetContext(hglrcShare) || !TryGetContext(hglrcDest)) {
SetLastError(ERROR_INVALID_HANDLE);
return FALSE;
}
return TRUE;
}
PROC GetProcAddress(const char* name) {
EnsureInitialized();
if (!name) {
return nullptr;
}
if (name[0] == 'w' && name[1] == 'g' && name[2] == 'l') {
for (const auto& entry : kWGLExtensionProcs) {
if (std::strcmp(entry.Name, name) == 0) {
return entry.Proc;
}
}
MGLOG_D("wglGetProcAddress: unknown wgl entry point %s", name);
return nullptr;
}
return reinterpret_cast<PROC>(MG_Impl::GetProcAddress(name));
}
} // namespace MobileGL::MG_Impl::WGLImpl
#endif // _WIN32
+34
View File
@@ -0,0 +1,34 @@
// MobileGL - MobileGL/MG_Impl/WGLImpl/WGLImpl.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include <Includes.h>
#if defined(_WIN32)
namespace MobileGL::MG_Impl::WGLImpl {
// Classic opengl32.dll surface. gdi32's ChoosePixelFormat/SetPixelFormat/
// DescribePixelFormat/GetPixelFormat/SwapBuffers forward into the loaded
// opengl32.dll's wgl* exports, so these back both call paths.
int ChoosePixelFormat(HDC hdc, const PIXELFORMATDESCRIPTOR* pfd);
int DescribePixelFormat(HDC hdc, int format, UINT size, PIXELFORMATDESCRIPTOR* pfd);
int GetPixelFormat(HDC hdc);
BOOL SetPixelFormat(HDC hdc, int format, const PIXELFORMATDESCRIPTOR* pfd);
BOOL SwapBuffers(HDC hdc);
HGLRC CreateContext(HDC hdc);
BOOL DeleteContext(HGLRC hglrc);
BOOL MakeCurrent(HDC hdc, HGLRC hglrc);
HGLRC GetCurrentContext();
HDC GetCurrentDC();
BOOL ShareLists(HGLRC hglrcShare, HGLRC hglrcDest);
PROC GetProcAddress(const char* name);
} // namespace MobileGL::MG_Impl::WGLImpl
#endif // _WIN32
+22 -1
View File
@@ -368,6 +368,26 @@ namespace MobileGL {
return true; return true;
} }
Bool EGLContext::HasAnyInitializedDisplay() const {
const std::lock_guard<std::recursive_mutex> lock(m_mutex);
for (const auto& [handle, displayObject] : m_displays) {
if (displayObject.Initialized) {
return true;
}
}
return false;
}
Bool EGLContext::HasAnyCurrentContext() const {
const std::lock_guard<std::recursive_mutex> lock(m_mutex);
for (const auto& [threadId, current] : m_threadCurrents) {
if (current.Context != nullptr) {
return true;
}
}
return false;
}
Bool EGLContext::ChooseConfig(EGLDisplayHandle display, const EGLint* attribList, EGLConfigHandle* configs, Bool EGLContext::ChooseConfig(EGLDisplayHandle display, const EGLint* attribList, EGLConfigHandle* configs,
EGLint configSize, EGLint* numConfig) { EGLint configSize, EGLint* numConfig) {
const std::lock_guard<std::recursive_mutex> lock(m_mutex); const std::lock_guard<std::recursive_mutex> lock(m_mutex);
@@ -1415,6 +1435,7 @@ namespace MobileGL {
} }
} // namespace EGLState } // namespace EGLState
UniquePtr<EGLState::EGLContext> pEGLContext; // Leak-at-exit storage; see GlobalObjects.cpp.
UniquePtr<EGLState::EGLContext>& pEGLContext = *new UniquePtr<EGLState::EGLContext>();
} // namespace MG_State } // namespace MG_State
} // namespace MobileGL } // namespace MobileGL
+5 -1
View File
@@ -39,6 +39,10 @@ namespace MobileGL {
Bool IsDisplayInitialized(EGLDisplayHandle display) const; Bool IsDisplayInitialized(EGLDisplayHandle display) const;
Bool InitializeDisplay(EGLDisplayHandle display, EGLint* major, EGLint* minor); Bool InitializeDisplay(EGLDisplayHandle display, EGLint* major, EGLint* minor);
Bool TerminateDisplay(EGLDisplayHandle display); Bool TerminateDisplay(EGLDisplayHandle display);
// Whole-library idle checks used by EGLImpl::Terminate to decide
// when the last eglTerminate may tear MobileGL down entirely.
Bool HasAnyInitializedDisplay() const;
Bool HasAnyCurrentContext() const;
// Config // Config
Bool ChooseConfig(EGLDisplayHandle display, const EGLint* attribList, EGLConfigHandle* configs, Bool ChooseConfig(EGLDisplayHandle display, const EGLint* attribList, EGLConfigHandle* configs,
@@ -262,6 +266,6 @@ namespace MobileGL {
}; };
} // namespace EGLState } // namespace EGLState
extern UniquePtr<EGLState::EGLContext> pEGLContext; extern UniquePtr<EGLState::EGLContext>& pEGLContext;
} // namespace MG_State } // namespace MG_State
} // namespace MobileGL } // namespace MobileGL
+2 -1
View File
@@ -721,5 +721,6 @@ namespace MobileGL::MG_State {
} }
} // namespace GLState } // namespace GLState
UniquePtr<GLState::GLContext> pGLContext; // Leak-at-exit storage; see GlobalObjects.cpp.
UniquePtr<GLState::GLContext>& pGLContext = *new UniquePtr<GLState::GLContext>();
} // namespace MobileGL::MG_State } // namespace MobileGL::MG_State
+1 -1
View File
@@ -252,7 +252,7 @@ namespace MobileGL {
}; };
} // namespace GLState } // namespace GLState
extern UniquePtr<GLState::GLContext> pGLContext; extern UniquePtr<GLState::GLContext>& pGLContext;
// True when relaxed GL semantics apply. Strict core rules are enforced only when the // True when relaxed GL semantics apply. Strict core rules are enforced only when the
// current EGL context explicitly requested a core profile (core bit in // current EGL context explicitly requested a core profile (core bit in
@@ -16,17 +16,32 @@ namespace MobileGL {
} }
void MipmapStorage::AllocateLevel(Uint level, MipmapInput input) { void MipmapStorage::AllocateLevel(Uint level, MipmapInput input) {
m_data.reserve(std::bit_ceil(level + 1)); // Grow only. GL respecifies exactly the level it is handed, so allocating level 0
m_data.resize(level + 1); // must not disturb the levels above it - but resize() shrinks as readily as it
m_texelSizes.reserve(std::bit_ceil(level + 1)); // grows, so this used to truncate the whole chain to a single level. Callers that
m_texelSizes.resize(level + 1); // genuinely redefine the complete level set say so with TruncateToLevelCount.
m_texelSizes[level] = input.texelSize; const SizeT requiredLevelCount = static_cast<SizeT>(level) + 1;
m_isDirty.resize(level + 1, false); if (m_data.size() < requiredLevelCount) {
m_data.reserve(std::bit_ceil(requiredLevelCount));
m_data.resize(requiredLevelCount);
m_texelSizes.reserve(std::bit_ceil(requiredLevelCount));
m_texelSizes.resize(requiredLevelCount);
m_isDirty.resize(requiredLevelCount, false);
}
m_texelSizes[level] = input.texelSize;
auto& data = m_data[level]; auto& data = m_data[level];
data.resize(input.byteSize, 0); data.resize(input.byteSize, 0);
} }
void MipmapStorage::TruncateToLevelCount(SizeT levelCount) {
if (levelCount >= m_data.size()) return;
m_data.resize(levelCount);
m_texelSizes.resize(levelCount);
m_isDirty.resize(levelCount);
}
void MipmapStorage::UpdateSubData(Uint level, DataPtr input) { void MipmapStorage::UpdateSubData(Uint level, DataPtr input) {
auto& targetData = m_data; auto& targetData = m_data;
MOBILEGL_ASSERT(level < targetData.size(), "UpdateSubData: level out of range"); MOBILEGL_ASSERT(level < targetData.size(), "UpdateSubData: level out of range");
@@ -55,6 +70,7 @@ namespace MobileGL {
} }
SizeT MipmapStorage::GetByteSize(Uint level) const { SizeT MipmapStorage::GetByteSize(Uint level) const {
if (level >= m_data.size()) return 0;
return m_data[level].size(); return m_data[level].size();
} }
@@ -19,6 +19,10 @@ namespace MobileGL {
public: public:
SizeT GetLevelCount() const; SizeT GetLevelCount() const;
void AllocateLevel(Uint level, MipmapInput input); void AllocateLevel(Uint level, MipmapInput input);
// Discard every level at or above levelCount. AllocateLevel never shrinks, so this
// is the only way a chain gets shorter - use it where the caller defines the whole
// level set (glTexStorage*, mip regeneration, atlas respecification).
void TruncateToLevelCount(SizeT levelCount);
void UpdateSubData(Uint level, DataPtr input); void UpdateSubData(Uint level, DataPtr input);
void* MapData(Uint level); void* MapData(Uint level);
IntVec3 GetTexelSize(Uint level) const; IntVec3 GetTexelSize(Uint level) const;
@@ -29,6 +29,14 @@ namespace MobileGL {
m_storage[targetIndex].AllocateLevel(level, input); m_storage[targetIndex].AllocateLevel(level, input);
} }
// Per-target, like AllocateLevel: cube-map faces are respecified independently, so
// truncating one face must not disturb the others.
void TruncateToLevelCount(Uint targetIndex, SizeT levelCount) {
MOBILEGL_ASSERT(targetIndex < TargetCount, "TruncateToLevelCount: target invalid");
m_storage[targetIndex].TruncateToLevelCount(levelCount);
}
void UpdateSubData(Uint targetIndex, Uint level, DataPtr input) { void UpdateSubData(Uint targetIndex, Uint level, DataPtr input) {
MOBILEGL_ASSERT(targetIndex < TargetCount, "UpdateSubData: target invalid"); MOBILEGL_ASSERT(targetIndex < TargetCount, "UpdateSubData: target invalid");
m_storage[targetIndex].UpdateSubData(level, input); m_storage[targetIndex].UpdateSubData(level, input);
@@ -271,6 +271,10 @@ namespace MobileGL {
m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input); m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
} }
void TextureObjectWithOneMipmap::TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) {
m_textureStorage.TruncateToLevelCount(GetIndexOfTextureUploadTarget(uploadTarget), levelCount);
}
void TextureObjectWithOneMipmap::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, void TextureObjectWithOneMipmap::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel,
DataPtr input) { DataPtr input) {
m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input); m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
@@ -134,6 +134,10 @@ namespace MobileGL::MG_State::GLState {
virtual const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const = 0; virtual const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const = 0;
virtual const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const = 0; virtual const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const = 0;
virtual void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) = 0; virtual void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) = 0;
// AllocateStorage only ever grows the chain. Callers that define the complete level set -
// glTexStorage*, mip regeneration, or a level-0 respecification at a new size - drop the
// leftovers explicitly, so a stale tail can never make the texture silently incomplete.
virtual void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) = 0;
virtual void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) = 0; virtual void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) = 0;
virtual void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) = 0; virtual void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) = 0;
virtual void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty = true) = 0; virtual void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty = true) = 0;
@@ -175,6 +179,7 @@ namespace MobileGL::MG_State::GLState {
const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override; const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override;
const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override; const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override;
void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override; void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override;
void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) override;
void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override; void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override;
void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override; void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override;
void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty) override; void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty) override;
@@ -31,6 +31,10 @@ namespace MobileGL {
m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input); m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
} }
void TextureObject2DCube::TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) {
m_textureStorage.TruncateToLevelCount(GetIndexOfTextureUploadTarget(uploadTarget), levelCount);
}
void TextureObject2DCube::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, void TextureObject2DCube::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel,
DataPtr input) { DataPtr input) {
m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input); m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
@@ -22,6 +22,7 @@ namespace MobileGL {
const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override; const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override;
const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override; const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override;
void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override; void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override;
void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) override;
void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override; void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override;
void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override; void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override;
void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, bool dirty) override; void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, bool dirty) override;
+2
View File
@@ -72,6 +72,8 @@ add_subdirectory(Texture)
add_subdirectory(VertexArray) add_subdirectory(VertexArray)
add_subdirectory(Program) add_subdirectory(Program)
add_subdirectory(Query) add_subdirectory(Query)
add_subdirectory(Pipeline)
add_subdirectory(ShaderTranspiler)
if (ENABLE_INTEGRATION_TESTS) if (ENABLE_INTEGRATION_TESTS)
add_subdirectory(Backend/DirectVulkan) add_subdirectory(Backend/DirectVulkan)
endif() endif()
+27
View File
@@ -0,0 +1,27 @@
cmake_minimum_required(VERSION 3.14)
add_executable(
PipelineQuirkTest
PipelineQuirkTest.cpp
)
target_include_directories(PipelineQuirkTest PRIVATE
${MGL_ROOT}/include
${MGL_ROOT}/MobileGL
${MGL_ROOT}/3rdparty/xxHash
${MGL_ROOT}/3rdparty/Vulkan-Headers/include
${MGL_ROOT}/3rdparty/SPIRV-Reflect
)
target_link_libraries(
PipelineQuirkTest PRIVATE
GTest::gtest_main
${LINK_LIBRARIES}
)
if (MSVC)
target_compile_options(PipelineQuirkTest PRIVATE /Zc:preprocessor)
endif()
include(GoogleTest)
gtest_discover_tests(PipelineQuirkTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
@@ -0,0 +1,459 @@
// MobileGL - MobileGL/MG_Test/Pipeline/PipelineQuirkTest.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <Config.h>
#include <MG_Backend/DirectVulkan/Renderer/PipelineFactory.h>
#include <MG_Backend/DirectVulkan/Renderer/ProgramFactory.h>
using namespace MobileGL;
using MobileGL::MG_Backend::DirectVulkan::PipelineFactory;
using MobileGL::MG_Backend::DirectVulkan::ProgramFactory;
using MobileGL::MG_Config::QuirkOverride;
namespace {
constexpr Uint32 kVendorIdQualcomm = 0x5143;
constexpr Uint32 kVendorIdArm = 0x13B5;
constexpr VkColorComponentFlags kFullColorWriteMask =
VK_COLOR_COMPONENT_R_BIT | VK_COLOR_COMPONENT_G_BIT |
VK_COLOR_COMPONENT_B_BIT | VK_COLOR_COMPONENT_A_BIT;
// Builds non-separate blend state: the alpha channel repeats the color factors/op, which
// is what glBlendFunc/glBlendEquation (as opposed to their *Separate forms) produce.
// ShouldSuppressDepthWrite deliberately decides on the color channel alone, so these
// cases cover its whole input space; SeparateAlphaAccumulationIsNotStripped below pins
// the separate-alpha contract.
VkPipelineColorBlendAttachmentState MakeBlendAttachment(Bool blendEnable,
VkBlendFactor srcColor,
VkBlendFactor dstColor,
VkBlendOp colorOp,
VkColorComponentFlags colorWriteMask) {
VkPipelineColorBlendAttachmentState attachment{};
attachment.blendEnable = blendEnable ? VK_TRUE : VK_FALSE;
attachment.srcColorBlendFactor = srcColor;
attachment.dstColorBlendFactor = dstColor;
attachment.colorBlendOp = colorOp;
attachment.srcAlphaBlendFactor = srcColor;
attachment.dstAlphaBlendFactor = dstColor;
attachment.alphaBlendOp = colorOp;
attachment.colorWriteMask = colorWriteMask;
return attachment;
}
// glslangValidator -V output for:
// #version 450
// layout(location = 0) out vec4 outColor;
// void main() { outColor = vec4(1.0); gl_FragDepth = 0.5; }
// Assigning gl_FragDepth makes glslang emit OpExecutionMode ... DepthReplacing.
constexpr Uint32 kFragDepthWriterSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000000fu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000004u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x00000009u, 0x0000000du, 0x00030010u,
0x00000004u, 0x00000007u, 0x00030010u, 0x00000004u, 0x0000000cu, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00050005u, 0x00000009u, 0x4374756fu, 0x726f6c6fu, 0x00000000u, 0x00060005u,
0x0000000du, 0x465f6c67u, 0x44676172u, 0x68747065u, 0x00000000u, 0x00040047u,
0x00000009u, 0x0000001eu, 0x00000000u, 0x00040047u, 0x0000000du, 0x0000000bu,
0x00000016u, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040020u, 0x00000008u, 0x00000003u, 0x00000007u, 0x0004003bu,
0x00000008u, 0x00000009u, 0x00000003u, 0x0004002bu, 0x00000006u, 0x0000000au,
0x3f800000u, 0x0007002cu, 0x00000007u, 0x0000000bu, 0x0000000au, 0x0000000au,
0x0000000au, 0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x00000006u,
0x0004003bu, 0x0000000cu, 0x0000000du, 0x00000003u, 0x0004002bu, 0x00000006u,
0x0000000eu, 0x3f000000u, 0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u,
0x00000003u, 0x000200f8u, 0x00000005u, 0x0003003eu, 0x00000009u, 0x0000000bu,
0x0003003eu, 0x0000000du, 0x0000000eu, 0x000100fdu, 0x00010038u,
};
// Same shader without the gl_FragDepth assignment.
constexpr Uint32 kPlainFragmentSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000000cu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0006000fu, 0x00000004u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x00000009u, 0x00030010u, 0x00000004u,
0x00000007u, 0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u,
0x6e69616du, 0x00000000u, 0x00050005u, 0x00000009u, 0x4374756fu, 0x726f6c6fu,
0x00000000u, 0x00040047u, 0x00000009u, 0x0000001eu, 0x00000000u, 0x00020013u,
0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u, 0x00000006u,
0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u, 0x00040020u,
0x00000008u, 0x00000003u, 0x00000007u, 0x0004003bu, 0x00000008u, 0x00000009u,
0x00000003u, 0x0004002bu, 0x00000006u, 0x0000000au, 0x3f800000u, 0x0007002cu,
0x00000007u, 0x0000000bu, 0x0000000au, 0x0000000au, 0x0000000au, 0x0000000au,
0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u,
0x00000005u, 0x0003003eu, 0x00000009u, 0x0000000bu, 0x000100fdu, 0x00010038u,
};
// glslangValidator -V output for a vertex shader reading gl_InstanceIndex:
// #version 450
// layout(location = 0) in vec4 inPos;
// void main() { gl_Position = inPos + vec4(float(gl_InstanceIndex)); }
constexpr Uint32 kInstanceIndexVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000001bu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0008000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00000014u,
0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du,
0x00000000u, 0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u,
0x00000000u, 0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu,
0x006e6f69u, 0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu,
0x657a6953u, 0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u,
0x4470696cu, 0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u,
0x435f6c67u, 0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du,
0x00000000u, 0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00070005u,
0x00000014u, 0x495f6c67u, 0x6174736eu, 0x4965636eu, 0x7865646eu, 0x00000000u,
0x00030047u, 0x0000000bu, 0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u,
0x0000000bu, 0x00000000u, 0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu,
0x00000001u, 0x00050048u, 0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u,
0x00050048u, 0x0000000bu, 0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u,
0x00000011u, 0x0000001eu, 0x00000000u, 0x00040047u, 0x00000014u, 0x0000000bu,
0x0000002bu, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu,
0x00000008u, 0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u,
0x00000009u, 0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au,
0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu,
0x0000000cu, 0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u,
0x00000001u, 0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u,
0x00000010u, 0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u,
0x00000001u, 0x00040020u, 0x00000013u, 0x00000001u, 0x0000000eu, 0x0004003bu,
0x00000013u, 0x00000014u, 0x00000001u, 0x00040020u, 0x00000019u, 0x00000003u,
0x00000007u, 0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u,
0x000200f8u, 0x00000005u, 0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u,
0x0004003du, 0x0000000eu, 0x00000015u, 0x00000014u, 0x0004006fu, 0x00000006u,
0x00000016u, 0x00000015u, 0x00070050u, 0x00000007u, 0x00000017u, 0x00000016u,
0x00000016u, 0x00000016u, 0x00000016u, 0x00050081u, 0x00000007u, 0x00000018u,
0x00000012u, 0x00000017u, 0x00050041u, 0x00000019u, 0x0000001au, 0x0000000du,
0x0000000fu, 0x0003003eu, 0x0000001au, 0x00000018u, 0x000100fdu, 0x00010038u,
};
// Same, but reading gl_VertexIndex instead: a DIFFERENT input builtin. glslang emits
// this for GL's gl_VertexID, so nearly every real vertex shader has one - it is what
// separates "declares some builtin" from "declares the InstanceIndex builtin".
constexpr Uint32 kVertexIndexVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000001bu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0008000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00000014u,
0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du,
0x00000000u, 0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u,
0x00000000u, 0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu,
0x006e6f69u, 0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu,
0x657a6953u, 0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u,
0x4470696cu, 0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u,
0x435f6c67u, 0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du,
0x00000000u, 0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00060005u,
0x00000014u, 0x565f6c67u, 0x65747265u, 0x646e4978u, 0x00007865u, 0x00030047u,
0x0000000bu, 0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu,
0x00000000u, 0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu, 0x00000001u,
0x00050048u, 0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u, 0x00050048u,
0x0000000bu, 0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u, 0x00000011u,
0x0000001eu, 0x00000000u, 0x00040047u, 0x00000014u, 0x0000000bu, 0x0000002au,
0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u,
0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u,
0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu, 0x00000008u,
0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u, 0x00000009u,
0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au, 0x0000000au,
0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu, 0x0000000cu,
0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u, 0x00000001u,
0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u, 0x00000010u,
0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u, 0x00000001u,
0x00040020u, 0x00000013u, 0x00000001u, 0x0000000eu, 0x0004003bu, 0x00000013u,
0x00000014u, 0x00000001u, 0x00040020u, 0x00000019u, 0x00000003u, 0x00000007u,
0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u,
0x00000005u, 0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u, 0x0004003du,
0x0000000eu, 0x00000015u, 0x00000014u, 0x0004006fu, 0x00000006u, 0x00000016u,
0x00000015u, 0x00070050u, 0x00000007u, 0x00000017u, 0x00000016u, 0x00000016u,
0x00000016u, 0x00000016u, 0x00050081u, 0x00000007u, 0x00000018u, 0x00000012u,
0x00000017u, 0x00050041u, 0x00000019u, 0x0000001au, 0x0000000du, 0x0000000fu,
0x0003003eu, 0x0000001au, 0x00000018u, 0x000100fdu, 0x00010038u,
};
// Owns the reflection module so each test case cleans up after itself.
class ReflectModule {
public:
template <SizeT WordCount>
explicit ReflectModule(const Uint32 (&spirv)[WordCount]) {
m_created = spvReflectCreateShaderModule(sizeof(spirv), spirv, &m_module) ==
SPV_REFLECT_RESULT_SUCCESS;
}
~ReflectModule() {
if (m_created) {
spvReflectDestroyShaderModule(&m_module);
}
}
ReflectModule(const ReflectModule&) = delete;
ReflectModule& operator=(const ReflectModule&) = delete;
Bool Created() const { return m_created; }
const SpvReflectShaderModule& Get() const { return m_module; }
private:
SpvReflectShaderModule m_module{};
Bool m_created = false;
};
PipelineFactory::PipelineCreatePayload MakeDepthWritingPayload(
const VkPipelineColorBlendAttachmentState& attachment0) {
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 1;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = attachment0;
return payload;
}
} // namespace
// --- Device gate: MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE tri-state ---
TEST(PipelineQuirkDeviceGate, ForceOnEnablesOnAnyVendor) {
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn,
kVendorIdArm));
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn,
kVendorIdQualcomm));
}
TEST(PipelineQuirkDeviceGate, ForceOffDisablesEvenOnQualcomm) {
EXPECT_FALSE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOff,
kVendorIdQualcomm));
}
TEST(PipelineQuirkDeviceGate, AutoDetectsQualcommOnly) {
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::Auto,
kVendorIdQualcomm));
EXPECT_FALSE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::Auto,
kVendorIdArm));
}
TEST(PipelineQuirkDeviceGate, ForceOnRoundTripsThroughTheFactoryFlag) {
const Bool previous = PipelineFactory::IsSuppressBlendedDepthWriteEnabled();
PipelineFactory::SetSuppressBlendedDepthWrite(
PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn, kVendorIdArm));
EXPECT_TRUE(PipelineFactory::IsSuppressBlendedDepthWriteEnabled());
PipelineFactory::SetSuppressBlendedDepthWrite(previous);
}
// --- Per-pipeline strip decision against the pipeline create-info payload ---
TEST(PipelineQuirkStripDecision, MaxBlendIsStripped) {
// MC 26.3 OIT depth_bounds: GL_MAX accumulation writing depth - the case the quirk fixes.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_MAX, kFullColorWriteMask));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, MinBlendIsStripped) {
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_MIN, kFullColorWriteMask));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AdditiveOnePlusOneIsNotStripped) {
// ONE+ONE additive with a depth write matched zero draws of the 26.3 chain in the
// fixture sweep (transmittance/accumulate disable depth writes themselves); the only
// real content with this shape was harmless additive glow effects (Create). A quirk
// touches as little unrelated content as possible, so the shape stays exempt.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, SortedTransparencyOverBlendIsNotStripped) {
// Vanilla MC translucent layer (water, stained glass): SRC_ALPHA "over" compositing
// draws each surface once and depends on its depth writes to occlude particles, rain,
// and clouds drawn later - it must keep them.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, EffectivelyOpaqueBlendIsNotStripped) {
// GL_BLEND left enabled with ONE/ZERO+ADD factors is opaque in effect; stripping its
// depth write would break occlusion for plainly opaque geometry.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, FullyMaskedAccumulationBlendIsNotStripped) {
// Depth-prepass pattern: colorMask(0,0,0,0) with blending left enabled - blending is
// moot, and stripping would delete the entire prepass. MAX so the exemption, not the
// blend-op filter, is what keeps the depth write.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, 0));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, DisabledBlendIsNotStripped) {
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
false, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, NoDepthWriteMeansNoStrip) {
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.depthWriteEnable = false;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, FragDepthWriterIsExempt) {
// gl_FragDepth output does not go through per-pipeline vertex position math, so the
// cross-pipeline invariance hazard cannot affect it (e.g. the 26.3 OIT composite).
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.fragmentReplacesDepth = true;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AccumulationOnSecondaryAttachmentIsStripped) {
// The scan is not limited to attachment 0: an extremum accumulation on any live
// attachment marks the pipeline.
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 2;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
false, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_ADD, kFullColorWriteMask);
payload.colorBlendAttachments[1] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask);
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AlphaWeightedAdditiveIsNotStripped) {
// SRC_ALPHA,ONE additive: the classic *sorted* particle/glow blend. Kept exempt like
// every other ADD-op shape now that the strip is extremum-only.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, ReverseSubtractIsNotStripped) {
// Deliberate narrowing: only the MIN/MAX extremum ops carry the depth-bounds
// signature. SUBTRACT-class ops stay outside the quirk until content demands them.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_REVERSE_SUBTRACT,
kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, PartiallyMaskedAccumulationIsStripped) {
// Only a fully masked attachment is exempt; a live alpha channel still accumulates.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, VK_COLOR_COMPONENT_A_BIT));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, NoColorAttachmentsMeansNoStrip) {
// Depth-only FBO: the loop must not read the (stale) attachment array at all.
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 0;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask);
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, SeparateAlphaAccumulationIsNotStripped) {
// glBlendEquationSeparate(GL_FUNC_ADD, GL_MAX) over an ordinary color over-blend: the
// alpha channel accumulates but the color channel does not. Pins that the decision is
// color-channel only - widening it to alpha would re-capture sorted transparency.
auto attachment = MakeBlendAttachment(true, VK_BLEND_FACTOR_SRC_ALPHA,
VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask);
attachment.srcAlphaBlendFactor = VK_BLEND_FACTOR_ONE;
attachment.dstAlphaBlendFactor = VK_BLEND_FACTOR_ONE;
attachment.alphaBlendOp = VK_BLEND_OP_MAX;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(MakeDepthWritingPayload(attachment)));
}
TEST(PipelineQuirkStripDecision, MixedOverAndMaskedAttachmentsAreNotStripped) {
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 2;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask);
payload.colorBlendAttachments[1] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, 0);
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
// --- DepthReplacing reflection feeding the gl_FragDepth exemption ---
TEST(ReflectedFragmentReplacesDepth, TrueForAShaderThatAssignsFragDepth) {
const ReflectModule module(kFragDepthWriterSpirv);
ASSERT_TRUE(module.Created());
EXPECT_TRUE(ProgramFactory::ReflectedFragmentReplacesDepth(module.Get()));
}
TEST(ReflectedFragmentReplacesDepth, FalseForAPlainFragmentShader) {
const ReflectModule module(kPlainFragmentSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedFragmentReplacesDepth(module.Get()));
}
TEST(ReflectedFragmentReplacesDepth, FalseForAnEmptyModule) {
// A default-constructed module has no entry points; the scan must not dereference.
SpvReflectShaderModule emptyModule{};
EXPECT_FALSE(ProgramFactory::ReflectedFragmentReplacesDepth(emptyModule));
}
TEST(ReflectedFragmentReplacesDepth, ReflectedFlagFlipsTheStripDecision) {
// The two fixtures differ only by the gl_FragDepth assignment, so they pin that the
// reflected flag is what flips the strip decision for an otherwise identical pipeline.
const ReflectModule depthWriter(kFragDepthWriterSpirv);
const ReflectModule plain(kPlainFragmentSpirv);
ASSERT_TRUE(depthWriter.Created());
ASSERT_TRUE(plain.Created());
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.fragmentReplacesDepth = ProgramFactory::ReflectedFragmentReplacesDepth(plain.Get());
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
payload.fragmentReplacesDepth = ProgramFactory::ReflectedFragmentReplacesDepth(depthWriter.Get());
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
// --- InstanceIndex reflection feeding the shaderDrawParameters diagnostic ---
TEST(ReflectedReadsInstanceIndexBuiltin, TrueForAShaderReadingInstanceIndex) {
const ReflectModule module(kInstanceIndexVertexSpirv);
ASSERT_TRUE(module.Created());
EXPECT_TRUE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAShaderReadingADifferentBuiltin) {
// Discriminates the builtin's identity, not merely its presence: weakening the check to
// "has any BuiltIn decoration" would fire the diagnostic on every real vertex shader.
const ReflectModule module(kVertexIndexVertexSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAShaderWithNoInputBuiltins) {
const ReflectModule module(kPlainFragmentSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAnEmptyModule) {
SpvReflectShaderModule emptyModule{};
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(emptyModule));
}
@@ -238,6 +238,77 @@ void main() {
} }
} }
// KHR-GL33.shaders.preprocessor.* — a block comment is one preprocessing token that the C/GLSL
// preprocessor replaces with a single space, even when it spans newlines inside a directive. glslang
// handles this natively, so MobileGL must not mangle it. These reproduce the CTS cases that failed
// because comment blanking preserved the interior newline, truncating multi-line #define bodies.
static void ExpectCompiles(MobileGL::ShaderStage stage, GLenum glStage, MobileGL::String source) {
using namespace MG_Util::ShaderTranspiler;
PreprocessShaderSource(stage, source);
ShaderAttrib attrib{.shaderType = glStage, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
TEST_F(ProgramUtilTest, PreprocessMultilineCommentInDefineBodyCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
#define VALUE /* current
value */ 4.2
void main()
{
out0 = VALUE;
})");
}
TEST_F(ProgramUtilTest, PreprocessRedefineObjectMultilineCommentCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
# define VAL1 1.0
#define VAL2 2.0
#define RES2 /* fdsjklfdsjkl
dsfjkhfdsjkh
fdsjklhfdsjkh */ (RES1 * VAL2)
#define RES1 (VAL2 / VAL1)
#define RES2 /* ewrlkjhsadf */ (RES1 * VAL2)
#define VALUE (RES2 + RES1)
void main()
{
out0 = VALUE;
})");
}
TEST_F(ProgramUtilTest, PreprocessFunctionMacroRedefinitionMultilineCommentCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
# define FUNC(a,b) (a +b)
# define FUNC(a,b)(a /* comment
*/ +b)
void main()
{
out0 = FUNC(1.0, 2.0);
})");
}
// Note: KHR-GL3x.shaders.preprocessor.conditional_inclusion.basic_2 (`#define AAA defined(BBB)` used
// in `#if !AAA`) is intentionally NOT handled here. Generating the `defined` operator via macro
// expansion is undefined per the C/GLSL preprocessor spec, and glslang deliberately rejects it
// ("'defined' : cannot use in preprocessor expression when expanded from macros"). Making it pass
// would require MobileGL to run its own macro expansion ahead of glslang, which is exactly the
// preprocessing we defer to glslang; the two cases stay failing by design.
TEST_F(ProgramUtilTest, PreprocessLegacyFragmentShaderModernizesGlmarkStyleSource) { TEST_F(ProgramUtilTest, PreprocessLegacyFragmentShaderModernizesGlmarkStyleSource) {
using namespace MG_Util::ShaderTranspiler; using namespace MG_Util::ShaderTranspiler;
@@ -426,6 +497,62 @@ void main() {
verifyVersion("#version 460 core"); verifyVersion("#version 460 core");
} }
// KHR-GL33.shaders.preprocessor.directive.version_* (also re-run verbatim under GL40-GL44): the
// compiler must REJECT a malformed #version line. MobileGL used to rewrite the whole line to
// "#version 330 core" whenever it could scrape a leading integer - or treat an unknown profile token
// as core - which silently legalized every form below. CTS compiles the shader's own #version
// verbatim, so the rejection has to survive preprocessing (and the 460 retry).
TEST_F(ProgramUtilTest, PreprocessRejectsMalformedVersionDirectives) {
using namespace MG_Util::ShaderTranspiler;
const char* body = "\nout vec4 fragColor;\nvoid main() { fragColor = vec4(1.0); }\n";
const auto rejects = [](const String& fullSource) {
String src = fullSource;
PreprocessShaderSource(ShaderStage::Fragment, src);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
return res ? false : true; // "rejects" == compile failed
};
// Silently legalized today - the five this fix must flip to rejection:
EXPECT_TRUE(rejects(String("#version 329") + body)) << "329 is not a real version";
EXPECT_TRUE(rejects(String("#version 331") + body)) << "331 is not a real version";
EXPECT_TRUE(rejects(String("#version 330 foo") + body)) << "unknown profile keyword";
EXPECT_TRUE(rejects(String("#version 330.0") + body)) << "float literal, not an int token";
EXPECT_TRUE(rejects(String("#version 330 foobar") + body)) << "trailing tokens after a valid decl";
// Already rejected (no leading integer, or #version is not the first token) - pinned so a future
// change to the normalizer cannot start legalizing them either:
EXPECT_TRUE(rejects(String("#version") + body)) << "missing version number";
EXPECT_TRUE(rejects(String("#version foobar") + body)) << "identifier where the int belongs";
EXPECT_TRUE(rejects(String("#version AAA") + body)) << "identifier where the int belongs";
EXPECT_TRUE(rejects(String("precision mediump float;\n#version 330") + body))
<< "#version must be the first statement";
EXPECT_TRUE(rejects(String("#define FOO BAR\n#version 330") + body))
<< "#version must precede a #define";
}
// The PASS half of the same CTS group: a valid decl, and #version preceded only by whitespace or a
// comment, must still compile. Guards the fix above from over-rejecting.
TEST_F(ProgramUtilTest, PreprocessKeepsValidVersionDirectivesCompiling) {
using namespace MG_Util::ShaderTranspiler;
const char* body = "\nout vec4 fragColor;\nvoid main() { fragColor = vec4(1.0); }\n";
const auto compiles = [](const String& fullSource) {
String src = fullSource;
PreprocessShaderSource(ShaderStage::Fragment, src);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
return res ? true : false;
};
EXPECT_TRUE(compiles(String("#version 330 core") + body));
EXPECT_TRUE(compiles(String("\n#version 330 core") + body))
<< "leading whitespace is legal before #version";
EXPECT_TRUE(compiles(String("// test\n#version 330 core") + body))
<< "a leading comment is legal before #version";
}
TEST_F(ProgramUtilTest, PreprocessUsesRealSpacedVersionDirectiveForInjectedOutput) { TEST_F(ProgramUtilTest, PreprocessUsesRealSpacedVersionDirectiveForInjectedOutput) {
using namespace MG_Util::ShaderTranspiler; using namespace MG_Util::ShaderTranspiler;
@@ -446,6 +573,9 @@ void main() {
EXPECT_NE(versionPos, String::npos); EXPECT_NE(versionPos, String::npos);
EXPECT_EQ(outputPos, versionPos + std::strlen("#version 330 core\n")); EXPECT_EQ(outputPos, versionPos + std::strlen("#version 330 core\n"));
EXPECT_NE(source.find("// #version 460 core"), String::npos); EXPECT_NE(source.find("// #version 460 core"), String::npos);
// This #line sits ahead of the version directive, where GLSL would never have honoured it, so
// it is still dropped. Directives that follow the version line are kept - see
// PreprocessKeepsPlainLineDirectivesAndSparesLookalikeIdentifiers.
EXPECT_EQ(source.find("#line"), String::npos); EXPECT_EQ(source.find("#line"), String::npos);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source}; ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
@@ -455,6 +585,105 @@ void main() {
} }
} }
// A banner line like "//*** NOTE ***" contains "/*" at offset 1 and no "*/" anywhere after it. The
// old hand-rolled comment stripper searched for "/*" with no lexical state, found that, failed to
// find a terminator, and erased everything from there to the end of the file - deleting the entire
// shader. Banner comments in that exact shape are common in Iris and OptiFine packs.
TEST_F(ProgramUtilTest, PreprocessKeepsShaderBodyAfterAStarredLineComment) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
//*** lighting pass ***
out vec4 fragColor;
void main() {
fragColor = vec4(1.0);
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("void main()"), String::npos) << "shader body was truncated:\n" << source;
EXPECT_NE(source.find("fragColor = vec4(1.0);"), String::npos);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
// The builtin-shadowing rename only fires when the shader really defines its own round/tanh/etc.
// Deciding that from a commented-out definition renames every genuine call to the builtin to a
// mg_ name that nothing defines, which fails to link.
TEST_F(ProgramUtilTest, PreprocessIgnoresCommentedOutBuiltinShadowingDefinition) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
// float round(float x) { return floor(x + 0.5); }
out vec4 fragColor;
void main() {
fragColor = vec4(round(1.25));
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("round(1.25)"), String::npos) << "call was renamed from a comment:\n" << source;
EXPECT_EQ(source.find("mg_round"), String::npos);
}
// A block-commented extension directive must not be treated as a real one - the int64 filter turns
// unsupported directives into #error, so reading one out of a comment manufactures a compile
// failure for a shader that never asked for the extension.
TEST_F(ProgramUtilTest, PreprocessIgnoresBlockCommentedExtensionDirectives) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
/*
#extension GL_ARB_gpu_shader_int64 : require
*/
out vec4 fragColor;
void main() {
fragColor = vec4(1.0);
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_EQ(source.find("#error"), String::npos) << "#error synthesized from a comment:\n" << source;
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
// KHR-GL33.shaders.preprocessor.builtin.line_* checks that __LINE__ follows #line. That only works
// if the directive reaches glslang, so a plain integer form must pass through untouched - while
// "#linear" and friends must not be mistaken for it.
TEST_F(ProgramUtilTest, PreprocessKeepsPlainLineDirectivesAndSparesLookalikeIdentifiers) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
out vec4 fragColor;
#line 42
float linear(float x) { return x; }
void main() {
#line 100
fragColor = vec4(linear(float(__LINE__)));
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("#line 42"), String::npos) << source;
EXPECT_NE(source.find("#line 100"), String::npos) << source;
EXPECT_NE(source.find("float linear(float x)"), String::npos) << "identifier lookalike was eaten:\n" << source;
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
TEST_F(ProgramUtilTest, PreprocessModernSampleQualifierStaysAtVersion460) { TEST_F(ProgramUtilTest, PreprocessModernSampleQualifierStaysAtVersion460) {
using namespace MG_Util::ShaderTranspiler; using namespace MG_Util::ShaderTranspiler;
@@ -745,6 +974,16 @@ TEST_F(ProgramUtilTest, RetargetLegacyVersionDirectiveOnlyTouchesNormalizedDeskt
String commented = "// #version 330 core\nvoid main() {}\n"; String commented = "// #version 330 core\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(commented)); EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(commented));
EXPECT_EQ(commented.find("#version 460"), String::npos); EXPECT_EQ(commented.find("#version 460"), String::npos);
// A malformed directive must NOT be rescued to 460 - that is what silently legalized the CTS
// directive.version_* rejection cases. The bad version stays put so glslang keeps rejecting it.
String badNumber = "#version 331\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(badNumber));
EXPECT_EQ(badNumber.find("#version 460"), String::npos);
String badProfile = "#version 330 foo\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(badProfile));
EXPECT_EQ(badProfile.find("#version 460"), String::npos);
} }
const char* fs = R"(#version 150 const char* fs = R"(#version 150
@@ -921,6 +1160,353 @@ TEST_F(ProgramUtilTest, CompileFragmentShaderWithDiscard) {
} }
} }
// noperspective is core desktop GLSL (1.30+) and maps to the SPIR-V NoPerspective decoration. It must
// reach glslang (not be stripped as text) so the SPIR-V carries the decoration; SPIRV-Cross then emits
// ESSL `noperspective` + the GL_NV_shader_noperspective_interpolation extension. Shader packs
// (Iris/Complementary) depend on it, and KHR-GL33.glsl_noperspective fails if the result matches
// smooth. This is the DirectGLES path with the NV extension available (SPIRV-Cross's default).
TEST_F(ProgramUtilTest, NoperspectiveInterpolationSurvivesToEssl) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 fragColor;
void main() { fragColor = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile errc: " << res.error().errc << "\nlog: " << res.error().log;
ProgramAttrib programAttrib{.shaders = {res.value()}};
auto program_res = ShaderCompiler::LinkProgram(programAttrib);
if (!program_res) FAIL() << "link errc: " << program_res.error().errc << "\nlog: " << program_res.error().log;
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *program_res.value()};
auto bin_res = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
if (!bin_res) FAIL() << "spirv errc: " << bin_res.error().errc << "\nlog: " << bin_res.error().log;
ASSERT_EQ(bin_res.value().size(), 1u);
SpvcSession session(bin_res.value()[0], SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc << "\nlog: " << essl.error().log;
EXPECT_NE(essl.value().find("noperspective"), String::npos)
<< "noperspective was lost before it reached SPIR-V:\n" << essl.value();
EXPECT_NE(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos)
<< "SPIRV-Cross must require the NV extension for ES noperspective:\n" << essl.value();
}
// The old handling was a naked substring erase of "noperspective", so any identifier that merely
// contained those characters (a uniform named noperspectiveBlend, say) got mangled. Removing the
// strip fixes it - glslang, which is identifier-aware, is the only thing that should see the keyword.
TEST_F(ProgramUtilTest, PreprocessDoesNotCorruptIdentifiersContainingNoperspective) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
uniform float noperspectiveBlend;
out vec4 fragColor;
void main() { fragColor = vec4(noperspectiveBlend); }
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("noperspectiveBlend"), String::npos)
<< "identifier was corrupted by substring stripping:\n" << source;
}
// The DirectGLES fallback for devices without GL_NV_shader_noperspective_interpolation: stripping the
// NoPerspective decoration makes SPIRV-Cross emit a plain smooth varying with no `#extension … :
// require`, so the shader still compiles (rendering as smooth) instead of being rejected by the driver.
TEST_F(ProgramUtilTest, StripNoPerspectiveFallbackProducesPlainEssl) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 fragColor;
void main() { fragColor = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile errc: " << res.error().errc << "\nlog: " << res.error().log;
ProgramAttrib programAttrib{.shaders = {res.value()}};
auto program_res = ShaderCompiler::LinkProgram(programAttrib);
if (!program_res) FAIL() << "link errc: " << program_res.error().errc << "\nlog: " << program_res.error().log;
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *program_res.value()};
auto bin_res = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
if (!bin_res) FAIL() << "spirv errc: " << bin_res.error().errc << "\nlog: " << bin_res.error().log;
ASSERT_EQ(bin_res.value().size(), 1u);
// Precondition: with the decoration present the default decompile requires the NV extension.
{
SpvcSession session(bin_res.value()[0], SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc;
ASSERT_NE(essl.value().find("noperspective"), String::npos) << essl.value();
}
// The fallback strips the decoration -> plain smooth ESSL, no extension require.
Vector<Uint32> stripped;
ASSERT_TRUE(ShaderCompiler::StripNoPerspectiveForEssl(bin_res.value()[0], stripped));
ASSERT_FALSE(stripped.empty());
SpvcSession session(stripped, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc << "\nlog: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos)
<< "the decoration should be gone:\n" << essl.value();
EXPECT_EQ(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos)
<< "no extension require without the decoration:\n" << essl.value();
}
// Directly exercises BOTH decoration forms StripNoPerspectivePass handles: a plain-variable
// OpDecorate NoPerspective (in-operand 1) and an interface-block-member OpMemberDecorate NoPerspective
// (in-operand 2). The ESSL round-trip tests above use only a scalar input, so they never reach the
// member-decorate branch, which a block varying like `in Block { noperspective vec4 c; }` (common in
// shader packs) produces. Unrelated decorations (Flat, Location) must survive untouched.
TEST_F(ProgramUtilTest, StripNoPerspectivePassRemovesBothDecorateForms) {
using namespace MG_Util::ShaderTranspiler;
const String spirvText = R"(
OpCapability Shader
OpMemoryModel Logical GLSL450
OpEntryPoint Fragment %main "main" %plainVar %blockVar %flatVar
OpExecutionMode %main OriginUpperLeft
OpName %main "main"
OpDecorate %plainVar Location 0
OpDecorate %plainVar NoPerspective
OpMemberDecorate %Block 0 NoPerspective
OpDecorate %blockVar Location 1
OpDecorate %flatVar Location 2
OpDecorate %flatVar Flat
%void = OpTypeVoid
%mainFn = OpTypeFunction %void
%float = OpTypeFloat 32
%v4float = OpTypeVector %float 4
%int = OpTypeInt 32 1
%inV4Ptr = OpTypePointer Input %v4float
%plainVar = OpVariable %inV4Ptr Input
%Block = OpTypeStruct %v4float
%inBlockPtr = OpTypePointer Input %Block
%blockVar = OpVariable %inBlockPtr Input
%inIntPtr = OpTypePointer Input %int
%flatVar = OpVariable %inIntPtr Input
%main = OpFunction %void None %mainFn
%mainBody = OpLabel
OpReturn
OpFunctionEnd
)";
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
Vector<uint32_t> inputBinary;
ASSERT_TRUE(tools.Assemble(spirvText, &inputBinary));
const auto countNoPerspective = [](const String& text) {
SizeT count = 0, offset = 0;
while ((offset = text.find("NoPerspective", offset)) != String::npos) {
++count;
offset += std::strlen("NoPerspective");
}
return count;
};
String inputText;
ASSERT_TRUE(tools.Disassemble(inputBinary, &inputText));
ASSERT_EQ(countNoPerspective(inputText), 2u)
<< "fixture must carry both a plain and a member NoPerspective:\n" << inputText;
Vector<uint32_t> outputBinary;
ASSERT_TRUE(ShaderCompiler::StripNoPerspectiveForEssl(inputBinary, outputBinary));
ASSERT_FALSE(outputBinary.empty());
String outputText;
ASSERT_TRUE(tools.Disassemble(outputBinary, &outputText));
EXPECT_EQ(countNoPerspective(outputText), 0u)
<< "both NoPerspective decorations (OpDecorate and OpMemberDecorate) must be stripped:\n" << outputText;
EXPECT_NE(outputText.find("Flat"), String::npos)
<< "the unrelated Flat decoration must survive:\n" << outputText;
EXPECT_NE(outputText.find("Location"), String::npos)
<< "Location decorations must survive:\n" << outputText;
}
// Phase 2 emulation - fragment side. On a device without the NV extension the NoPerspective input is
// recovered as `load * gl_FragCoord.w` and the decoration removed; gl_FragCoord is synthesized because
// the shader did not otherwise use it. The emulated SPIR-V must validate and decompile without the
// extension require.
TEST_F(ProgramUtilTest, EmulateNoperspectiveFragmentRecoversWithFragCoordW) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 f;
void main() { f = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile: " << res.error().log;
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
if (!pr) FAIL() << "link: " << pr.error().log;
ProgramBinaryAttrib ba{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
if (!br) FAIL() << "spirv: " << br.error().log;
ASSERT_EQ(br.value().size(), 1u);
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(br.value()[0], emulated));
ASSERT_FALSE(emulated.empty());
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << "emulated SPIR-V must be valid:\n" << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << "decoration must be stripped:\n" << dis;
EXPECT_NE(dis.find("FragCoord"), String::npos) << "gl_FragCoord must be synthesized:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos) << "the recovery multiply must be present:\n" << dis;
SpvcSession session(emulated, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos) << essl.value();
EXPECT_EQ(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos) << essl.value();
EXPECT_NE(essl.value().find("gl_FragCoord"), String::npos) << "recovery must reference gl_FragCoord:\n" << essl.value();
}
// Phase 2 emulation - vertex side. The NoPerspective output is pre-multiplied by gl_Position.w before
// return and the decoration removed. Emulated SPIR-V must validate and decompile without the extension.
TEST_F(ProgramUtilTest, EmulateNoperspectiveVertexPreMultipliesByPositionW) {
using namespace MG_Util::ShaderTranspiler;
String vs = R"(#version 330 core
in vec4 pos;
noperspective out vec4 vColor;
void main() { gl_Position = pos; vColor = pos; }
)";
ShaderAttrib attrib{.shaderType = GL_VERTEX_SHADER, .sourceStr = vs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile: " << res.error().log;
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
if (!pr) FAIL() << "link: " << pr.error().log;
ProgramBinaryAttrib ba{.shaderTypes = {GL_VERTEX_SHADER}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
if (!br) FAIL() << "spirv: " << br.error().log;
ASSERT_EQ(br.value().size(), 1u);
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(br.value()[0], emulated));
ASSERT_FALSE(emulated.empty());
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << "emulated SPIR-V must be valid:\n" << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << "decoration must be stripped:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos) << "the pre-multiply must be present:\n" << dis;
SpvcSession session(emulated, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos) << essl.value();
EXPECT_NE(essl.value().find("gl_Position"), String::npos) << "pre-multiply must reference gl_Position:\n" << essl.value();
}
namespace {
// Compiles one shader stage through the full pipeline and returns its SPIR-V, or fails the test.
MobileGL::Vector<uint32_t> CompileStageSpirv(GLenum type, const char* src) {
using namespace MG_Util::ShaderTranspiler;
ShaderAttrib attrib{.shaderType = type, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
EXPECT_TRUE(static_cast<bool>(res)) << (res ? "" : res.error().log);
if (!res) return {};
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
EXPECT_TRUE(static_cast<bool>(pr)) << (pr ? "" : pr.error().log);
if (!pr) return {};
ProgramBinaryAttrib ba{.shaderTypes = {type}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
EXPECT_TRUE(static_cast<bool>(br)) << (br ? "" : br.error().log);
if (!br || br.value().empty()) return {};
return br.value()[0];
}
} // namespace
// Regression: the vertex pre-multiply must be applied exactly once (in main), not once per function.
// glslang does not inline, so a helper function survives as its own OpFunction; instrumenting its
// return too would scale the varying by gl_Position.w twice (w^2).
TEST_F(ProgramUtilTest, EmulateNoperspectiveVertexWithHelperScalesExactlyOnce) {
using namespace MG_Util::ShaderTranspiler;
// helper() returns via OpReturnValue and adds (no vector*scalar), so the ONLY OpVectorTimesScalar
// in the module is the emulation's pre-multiply. The old all-functions code injected it at both
// helper's and main's return -> count 2; restricted to the entry function it is 1.
auto spirv = CompileStageSpirv(GL_VERTEX_SHADER, R"(#version 330 core
in vec4 pos;
noperspective out vec4 vColor;
vec4 helper(vec4 x) { return x + vec4(1.0); }
void main() { gl_Position = pos; vColor = helper(pos); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
SizeT count = 0, off = 0;
while ((off = dis.find("OpVectorTimesScalar", off)) != String::npos) {
++count;
off += std::strlen("OpVectorTimesScalar");
}
EXPECT_EQ(count, 1u) << "the gl_Position.w pre-multiply must happen exactly once, not per function:\n" << dis;
}
// Regression: a single-component read (vColor.x), which glslang lowers via OpAccessChain, must still be
// recovered with gl_FragCoord.w - not silently left un-scaled.
TEST_F(ProgramUtilTest, EmulateNoperspectiveFragmentComponentReadIsRecovered) {
using namespace MG_Util::ShaderTranspiler;
auto spirv = CompileStageSpirv(GL_FRAGMENT_SHADER, R"(#version 330 core
noperspective in vec4 vColor;
out vec4 f;
void main() { f = vec4(vColor.x); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << dis;
EXPECT_NE(dis.find("FragCoord"), String::npos)
<< "the component read must still be recovered via gl_FragCoord.w:\n" << dis;
}
// Coverage: a scalar float varying exercises the OpFMul path; a vector varying the OpVectorTimesScalar
// path; multiple noperspective varyings in one stage are all handled.
TEST_F(ProgramUtilTest, EmulateNoperspectiveHandlesScalarAndMultipleVaryings) {
using namespace MG_Util::ShaderTranspiler;
auto spirv = CompileStageSpirv(GL_FRAGMENT_SHADER, R"(#version 330 core
noperspective in float a;
noperspective in vec2 b;
out vec4 f;
void main() { f = vec4(a, b, 1.0); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << dis;
EXPECT_NE(dis.find("OpFMul"), String::npos) << "the scalar varying must scale with OpFMul:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos)
<< "the vector varying must scale with OpVectorTimesScalar:\n" << dis;
}
const char* vs_location = R"(#version 460 const char* vs_location = R"(#version 460
in vec4 Position; in vec4 Position;
@@ -1623,4 +2209,26 @@ TEST_F(ProgramUtilTest, RewriteLinearSubgroupPrefixScanRejectsPartialOrUnsafeTem
ASSERT_NE(consumerEnd, String::npos); ASSERT_NE(consumerEnd, String::npos);
nestedScan.insert(consumerEnd + 1, "\n }"); nestedScan.insert(consumerEnd + 1, "\n }");
expectUnchanged(std::move(nestedScan)); expectUnchanged(std::move(nestedScan));
// ARB/NV spellings of lane-width-sensitive builtins must block the rewrite exactly
// like their KHR counterparts.
String arbSubgroupBuiltin = MakeLinearSubgroupPrefixScanShader();
arbSubgroupBuiltin.insert(arbSubgroupBuiltin.find("float importance"),
"uint arbLane = gl_SubGroupInvocationARB;\n ");
expectUnchanged(std::move(arbSubgroupBuiltin));
String arbBallotCall = MakeLinearSubgroupPrefixScanShader();
arbBallotCall.insert(arbBallotCall.find("float importance"),
"uint64_t arbMask = ballotARB(true);\n ");
expectUnchanged(std::move(arbBallotCall));
String nvWarpBuiltin = MakeLinearSubgroupPrefixScanShader();
nvWarpBuiltin.insert(nvWarpBuiltin.find("float importance"),
"uint warpSize = gl_WarpSizeNV;\n ");
expectUnchanged(std::move(nvWarpBuiltin));
String nvShuffleCall = MakeLinearSubgroupPrefixScanShader();
nvShuffleCall.insert(nvShuffleCall.find("float importance"),
"float other = shuffleNV(1.0f, 0u, 32u);\n ");
expectUnchanged(std::move(nvShuffleCall));
} }
+351 -1
View File
@@ -794,6 +794,33 @@ TEST(DirectVulkanSanity, ReadbackConvertsRgba8AndRgba16fPixels) {
EXPECT_FLOAT_EQ(rgba16fFloatResult[3], 1.0f); EXPECT_FLOAT_EQ(rgba16fFloatResult[3], 1.0f);
} }
TEST(DirectVulkanSanity, ReadbackDecodesSingleChannel32BitFormats) {
using MobileGL::MG_Backend::DirectVulkan::VulkanRenderer;
// The reinterpretation feature makes R32F/R32UI-class images common readback sources
// (iterationRP custom images). Missing channels take GL defaults: 0 for GB, 1 for alpha.
const MobileGL::Float r32f[] = {0.75f, -2.0f};
MobileGL::Float r32fResult[8]{};
ASSERT_TRUE(VulkanRenderer::ConvertReadbackPixels(
reinterpret_cast<const MobileGL::Uint8*>(r32f), VK_FORMAT_R32_SFLOAT,
2, 1, GL_RGBA, GL_FLOAT, sizeof(MobileGL::Float) * 8,
reinterpret_cast<MobileGL::Uint8*>(r32fResult)));
EXPECT_FLOAT_EQ(r32fResult[0], 0.75f);
EXPECT_FLOAT_EQ(r32fResult[1], 0.0f);
EXPECT_FLOAT_EQ(r32fResult[2], 0.0f);
EXPECT_FLOAT_EQ(r32fResult[3], 1.0f);
EXPECT_FLOAT_EQ(r32fResult[4], -2.0f);
const MobileGL::Uint32 r32ui[] = {12345u};
MobileGL::Float r32uiResult[4]{};
ASSERT_TRUE(VulkanRenderer::ConvertReadbackPixels(
reinterpret_cast<const MobileGL::Uint8*>(r32ui), VK_FORMAT_R32_UINT,
1, 1, GL_RGBA, GL_FLOAT, sizeof(MobileGL::Float) * 4,
reinterpret_cast<MobileGL::Uint8*>(r32uiResult)));
EXPECT_FLOAT_EQ(r32uiResult[0], 12345.0f);
EXPECT_FLOAT_EQ(r32uiResult[3], 1.0f);
}
TEST(DirectVulkanSanity, DrawIndexedIndirectCommandMatchesGlAndVulkanLayout) { TEST(DirectVulkanSanity, DrawIndexedIndirectCommandMatchesGlAndVulkanLayout) {
using namespace MobileGL::MG_Backend::DirectVulkan; using namespace MobileGL::MG_Backend::DirectVulkan;
@@ -999,9 +1026,27 @@ TEST(DirectVulkanSanity, SampledViewFormatMatchesSamplerNumericDomainWithoutChan
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat( EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_B10G11R11_UFLOAT_PACK32, SamplerNumericDomain::UnsignedInteger), VK_FORMAT_B10G11R11_UFLOAT_PACK32, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_UNDEFINED); VK_FORMAT_UNDEFINED);
// Depth/stencil formats never resolve through color-class reinterpretation; they pass
// through unchanged so the existing depth-aspect sampled view is used. Combined
// depth-stencil formats are multi-numeric (vkuFormatIsSampledFloat is false for them),
// so without the passthrough a plain sampler2D/sampler2DShadow on GL_DEPTH24_STENCIL8
// would resolve to UNDEFINED and the draw would be dropped.
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D24_UNORM_S8_UINT, SamplerNumericDomain::Float),
VK_FORMAT_D24_UNORM_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT_S8_UINT, SamplerNumericDomain::Float),
VK_FORMAT_D32_SFLOAT_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT, SamplerNumericDomain::Float),
VK_FORMAT_D32_SFLOAT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D24_UNORM_S8_UINT, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_D24_UNORM_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat( EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT, SamplerNumericDomain::UnsignedInteger), VK_FORMAT_D32_SFLOAT, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_UNDEFINED); VK_FORMAT_D32_SFLOAT);
EXPECT_TRUE(VkTextureManager::AreSampledImageViewFormatsCompatible( EXPECT_TRUE(VkTextureManager::AreSampledImageViewFormatsCompatible(
VK_FORMAT_R32_SFLOAT, VK_FORMAT_R32_UINT)); VK_FORMAT_R32_SFLOAT, VK_FORMAT_R32_UINT));
@@ -1386,3 +1431,308 @@ TEST(RenderStateSanity, PrimitiveRestartIndexStoresAndReadsBack) {
MG_State::pGLContext.reset(); MG_State::pGLContext.reset();
} }
// ---- DirectGLES readback driver-state shadows ----------------------------------------------------
// Regression coverage for the readback-path state-leak overhaul: the pixel-PBO
// binding cache, the framebuffer-binding shadow, the PACK pixel-store shadow and
// the scratch-FBO attachment shadow must (a) leave the driver in the documented
// resting state, (b) skip redundant GL calls, and (c) scrub correctly on
// deletion. All drive the real Managers.cpp implementations against a recording
// mock GLES table.
namespace {
struct StateGuardCallLog {
MobileGL::Vector<MobileGL::String> calls;
MobileGL::SizeT Count(const MobileGL::String& prefix) const {
MobileGL::SizeT n = 0;
for (const auto& c : calls) {
if (c.compare(0, prefix.size(), prefix) == 0) ++n;
}
return n;
}
};
StateGuardCallLog* g_stateGuardLog = nullptr;
GLuint g_nextStateGuardFBOId = 201;
void SG_Log(MobileGL::String entry) {
if (g_stateGuardLog) g_stateGuardLog->calls.push_back(MobileGL::Move(entry));
}
void SG_BindBuffer(GLenum target, GLuint buffer) {
SG_Log("BindBuffer:" + std::to_string(target) + ":" + std::to_string(buffer));
}
void SG_BindFramebuffer(GLenum target, GLuint framebuffer) {
SG_Log("BindFramebuffer:" + std::to_string(target) + ":" + std::to_string(framebuffer));
}
void SG_GetIntegerv(GLenum pname, GLint* data) {
SG_Log("GetIntegerv:" + std::to_string(pname));
if (data) *data = 0;
}
void SG_PixelStorei(GLenum pname, GLint param) {
SG_Log("PixelStorei:" + std::to_string(pname) + ":" + std::to_string(param));
}
void SG_GenFramebuffers(GLsizei count, GLuint* framebuffers) {
for (GLsizei i = 0; i < count; ++i) framebuffers[i] = g_nextStateGuardFBOId++;
}
void SG_FramebufferTexture2D(GLenum target, GLenum attachment, GLenum textarget, GLuint texture, GLint level) {
SG_Log("FramebufferTexture2D:" + std::to_string(target) + ":" + std::to_string(attachment) + ":" +
std::to_string(textarget) + ":" + std::to_string(texture) + ":" + std::to_string(level));
}
void SG_FramebufferTextureLayer(GLenum target, GLenum attachment, GLuint texture, GLint level, GLint layer) {
SG_Log("FramebufferTextureLayer:" + std::to_string(target) + ":" + std::to_string(attachment) + ":" +
std::to_string(texture) + ":" + std::to_string(level) + ":" + std::to_string(layer));
}
void SG_ReadBuffer(GLenum src) {
SG_Log("ReadBuffer:" + std::to_string(src));
}
void SG_DrawBuffers(GLsizei n, const GLenum* bufs) {
SG_Log("DrawBuffers:" + std::to_string(n) + ":" + std::to_string(n > 0 && bufs ? bufs[0] : 0));
}
GLenum SG_NoError() {
return GL_NO_ERROR;
}
// Installs the recording table and resets every readback driver-state shadow on
// both ends, so these tests cannot bleed into (or inherit from) other tests.
struct ScopedStateGuardMocks {
ScopedStateGuardMocks(): previousFunctions(MobileGL::MG_Backend::DirectGLES::g_GLESFuncs) {
ResetShadows();
MobileGL::MG_External::GLESFunctionsTable functions{};
functions.glBindBuffer = SG_BindBuffer;
functions.glBindFramebuffer = SG_BindFramebuffer;
functions.glGetIntegerv = SG_GetIntegerv;
functions.glPixelStorei = SG_PixelStorei;
functions.glGenFramebuffers = SG_GenFramebuffers;
functions.glFramebufferTexture2D = SG_FramebufferTexture2D;
functions.glFramebufferTextureLayer = SG_FramebufferTextureLayer;
functions.glReadBuffer = SG_ReadBuffer;
functions.glDrawBuffers = SG_DrawBuffers;
functions.glGetError = SG_NoError;
MobileGL::MG_Backend::DirectGLES::SetGLESFuncsTable(functions);
g_stateGuardLog = &log;
}
~ScopedStateGuardMocks() {
g_stateGuardLog = nullptr;
MobileGL::MG_Backend::DirectGLES::SetGLESFuncsTable(previousFunctions);
ResetShadows();
}
ScopedStateGuardMocks(const ScopedStateGuardMocks&) = delete;
ScopedStateGuardMocks& operator=(const ScopedStateGuardMocks&) = delete;
static void ResetShadows() {
MobileGL::MG_Backend::DirectGLES::BufferImpl::InvalidatePixelBufferBindingCaches();
MobileGL::MG_Backend::DirectGLES::FramebufferImpl::InvalidateFramebufferBindingCache();
MobileGL::MG_Backend::DirectGLES::PixelStoreImpl::InvalidatePackStateCache();
MobileGL::MG_Backend::DirectGLES::ScratchFBOImpl::OnBackendContextDestroyed();
}
StateGuardCallLog log;
MobileGL::MG_External::GLESFunctionsTable previousFunctions;
};
} // namespace
TEST(DirectGLESStateGuards, PixelPackBindingCacheSkipsRedundantBindsAndRestsAtZero) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
BufferImpl::BindPixelPackBufferId(5);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 1u);
BufferImpl::BindPixelPackBufferId(5); // redundant: must not reach the driver
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 1u);
BufferImpl::BindPixelPackBufferId(0); // scope exit: resting state
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 2u);
BufferImpl::BindPixelPackBufferId(0);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 2u);
// After invalidation (MakeCurrent / context reset) the first bind must reach
// the driver again even for the same value.
BufferImpl::InvalidatePixelBufferBindingCaches();
BufferImpl::BindPixelPackBufferId(0);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 3u);
}
TEST(DirectGLESStateGuards, FramebufferBindingShadowPinsOnceThenSkips) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// Cold path: one driver query pins the shadow; further reads are free.
(void)FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Read);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u);
(void)FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Read);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 1u);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 1u);
// GL_FRAMEBUFFER touches both targets; DRAW is still unknown so it must bind.
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 2u);
// Both halves now match: no further calls for either single target.
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 2u);
EXPECT_EQ(FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Draw), 7u);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u); // shadow answered, no new query
}
TEST(DirectGLESStateGuards, PackStateShadowAppliesMinimalDeltas) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// First application pins all four parameters.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{4, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 4u);
// Identical state: zero driver calls.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{4, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 4u);
// One field changed: exactly one driver call.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{1, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 5u);
const auto current = PixelStoreImpl::CurrentPackState();
EXPECT_EQ(current.Alignment, 1);
EXPECT_EQ(current.RowLength, 0);
EXPECT_EQ(current.SkipRows, 0);
EXPECT_EQ(current.SkipPixels, 0);
}
TEST(DirectGLESStateGuards, ScratchFBODetachesCrossAspectResidue) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::TempFramebuffer();
EXPECT_NE(ScratchFBOImpl::EnsureId(fb), 0u);
// A depth copy leaves a DEPTH_STENCIL attachment (the pre-fix code never
// detached it, wedging every later color readback through this FBO).
ScratchFBOImpl::EnsureDepthAttachment2D(fb, GL_DRAW_FRAMEBUFFER, 11, GL_TEXTURE_2D, 0, /*withStencil=*/true);
const MobileGL::String dsAttach = "FramebufferTexture2D:" + std::to_string(GL_DRAW_FRAMEBUFFER) + ":" +
std::to_string(GL_DEPTH_STENCIL_ATTACHMENT);
EXPECT_EQ(mocks.log.Count(dsAttach), 1u);
// The next color use must detach the stale depth-stencil attachment exactly once.
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
const MobileGL::String dsDetach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_DEPTH_STENCIL_ATTACHMENT) + ":" +
std::to_string(GL_TEXTURE_2D) + ":0:0";
const MobileGL::String colorAttach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_COLOR_ATTACHMENT0) + ":" +
std::to_string(GL_TEXTURE_2D) + ":22:0";
EXPECT_EQ(mocks.log.Count(dsDetach), 1u);
EXPECT_EQ(mocks.log.Count(colorAttach), 1u);
// Back-to-back identical color use: no driver traffic at all.
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
EXPECT_EQ(mocks.log.Count("FramebufferTexture2D:"), 0u);
}
TEST(DirectGLESStateGuards, ScratchFBOTextureDeletionForcesFullScrub) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::TempFramebuffer();
ScratchFBOImpl::EnsureId(fb);
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
// The attached texture id dies: the shadow can no longer vouch for the FBO
// (ES does not auto-detach from unbound FBOs, and the name may be recycled),
// so the next use must scrub and re-attach instead of skipping.
ScratchFBOImpl::NoteTextureIdDeleted(22);
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
EXPECT_GE(mocks.log.Count("FramebufferTexture2D:"), 2u); // scrub (color + depth) ...
const MobileGL::String colorAttach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_COLOR_ATTACHMENT0) + ":" +
std::to_string(GL_TEXTURE_2D) + ":22:0";
EXPECT_EQ(mocks.log.Count(colorAttach), 1u); // ... then the real re-attach
}
TEST(DirectGLESStateGuards, ScratchFBOReadDrawBufferStateCached) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::BlitReadFramebuffer();
ScratchFBOImpl::EnsureId(fb);
// Fresh FBOs default to COLOR_ATTACHMENT0 for both buffers: no call needed.
ScratchFBOImpl::EnsureReadBuffer(fb, GL_COLOR_ATTACHMENT0);
EXPECT_EQ(mocks.log.Count("ReadBuffer:"), 0u);
// Depth blits want GL_NONE; the transition costs one call, repeats are free.
ScratchFBOImpl::EnsureReadBuffer(fb, GL_NONE);
ScratchFBOImpl::EnsureReadBuffer(fb, GL_NONE);
EXPECT_EQ(mocks.log.Count("ReadBuffer:"), 1u);
ScratchFBOImpl::EnsureDrawBuffer(fb, GL_NONE);
ScratchFBOImpl::EnsureDrawBuffer(fb, GL_NONE);
EXPECT_EQ(mocks.log.Count("DrawBuffers:"), 1u);
}
namespace {
MobileGL::Vector<GLuint>* g_deletedTextureIds = nullptr;
void SG_DeleteTextures(GLsizei count, const GLuint* textures) {
if (!g_deletedTextureIds) return;
for (GLsizei i = 0; i < count; ++i) g_deletedTextureIds->push_back(textures[i]);
}
// Clears the recording hook even when a gtest assertion unwinds the test body
// (a dangling pointer to the dead stack vector would corrupt later tests).
struct ScopedDeletedTextureRecording {
explicit ScopedDeletedTextureRecording(MobileGL::Vector<GLuint>& sink) { g_deletedTextureIds = &sink; }
~ScopedDeletedTextureRecording() { g_deletedTextureIds = nullptr; }
ScopedDeletedTextureRecording(const ScopedDeletedTextureRecording&) = delete;
ScopedDeletedTextureRecording& operator=(const ScopedDeletedTextureRecording&) = delete;
};
} // namespace
TEST(DirectGLESBackendTexture, DestructorDeletesIdAndScrubsBindingCache) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedDirectGLESTextureBindings scoped; // installs glGenTextures/glBindTexture mocks + resets caches
MobileGL::Vector<GLuint> deleted;
ScopedDeletedTextureRecording recording(deleted);
auto functions = g_GLESFuncs;
functions.glDeleteTextures = SG_DeleteTextures;
SetGLESFuncsTable(functions);
const auto texture2DSlot = static_cast<MobileGL::SizeT>(MobileGL::TextureTarget::Texture2D);
GLuint id = 0;
{
auto backendTexture = MobileGL::MakeShared<TextureImpl::BackendTextureObject>();
id = backendTexture->GetBackendTextureId();
ASSERT_NE(id, 0u);
backendTexture->Bind(GL_TEXTURE_2D, 0);
ASSERT_EQ(TextureImpl::g_boundTexturesCache[0][texture2DSlot], backendTexture.get());
}
// Frontend glDeleteTextures used to leak the backend id forever and leave the
// cache pointer dangling (heap-address reuse then false-skips a later Bind).
ASSERT_EQ(deleted.size(), 1u);
EXPECT_EQ(deleted[0], id);
EXPECT_EQ(TextureImpl::g_boundTexturesCache[0][texture2DSlot], nullptr);
// A wrapper whose context died must NOT delete a foreign (recycled) name.
{
auto backendTexture = MobileGL::MakeShared<TextureImpl::BackendTextureObject>();
++TextureImpl::g_textureContextGeneration;
backendTexture.reset();
--TextureImpl::g_textureContextGeneration; // restore for later tests
EXPECT_EQ(deleted.size(), 1u);
}
}
TEST(DirectGLESStateGuards, DefaultFramebufferBindGoesThroughShadow) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// The regression this guards against: binding framebuffer 0 raw while the
// shadow keeps a user-FBO id makes the next re-bind of that FBO false-skip.
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 0); // default-FBO path must use this API
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7); // must reach the driver again
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 3u);
}
@@ -0,0 +1,25 @@
cmake_minimum_required(VERSION 3.14)
add_executable(
SpirvPassTest
SpirvPassTest.cpp
)
target_include_directories(SpirvPassTest PRIVATE
${MGL_ROOT}/include
${MGL_ROOT}/MobileGL
${MGL_ROOT}/3rdparty/SPIRV-Reflect
)
target_link_libraries(
SpirvPassTest PRIVATE
GTest::gtest_main
${LINK_LIBRARIES}
)
if (MSVC)
target_compile_options(SpirvPassTest PRIVATE /Zc:preprocessor)
endif()
include(GoogleTest)
gtest_discover_tests(SpirvPassTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
@@ -0,0 +1,170 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/SpirvPassTest.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <spirv_reflect.h>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
namespace {
// glslangValidator -V output. Both are vertex shaders writing gl_Position through
// the gl_PerVertex block, i.e. the Position builtin arrives as OpMemberDecorate rather
// than a plain OpDecorate - the shape real glslang output actually takes.
// #version 450
// layout(location = 0) in vec4 inPos;
// void main() { gl_Position = inPos; }
constexpr Uint32 kPlainVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x00000015u, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u, 0x00000000u,
0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu, 0x006e6f69u,
0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu, 0x657a6953u,
0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u, 0x4470696cu,
0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u, 0x435f6c67u,
0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du, 0x00000000u,
0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00030047u, 0x0000000bu,
0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu, 0x00000000u,
0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu, 0x00000001u, 0x00050048u,
0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u, 0x00050048u, 0x0000000bu,
0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u, 0x00000011u, 0x0000001eu,
0x00000000u, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu,
0x00000008u, 0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u,
0x00000009u, 0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au,
0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu,
0x0000000cu, 0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u,
0x00000001u, 0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u,
0x00000010u, 0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u,
0x00000001u, 0x00040020u, 0x00000013u, 0x00000003u, 0x00000007u, 0x00050036u,
0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u, 0x00000005u,
0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u, 0x00050041u, 0x00000013u,
0x00000014u, 0x0000000du, 0x0000000fu, 0x0003003eu, 0x00000014u, 0x00000012u,
0x000100fdu, 0x00010038u,
};
// ... plus `invariant gl_Position;` - already carries OpMemberDecorate %gl_PerVertex 0
// Invariant, so the pass must not add a duplicate.
constexpr Uint32 kAlreadyInvariantVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x00000015u, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u, 0x00000000u,
0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu, 0x006e6f69u,
0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu, 0x657a6953u,
0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u, 0x4470696cu,
0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u, 0x435f6c67u,
0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du, 0x00000000u,
0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00030047u, 0x0000000bu,
0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu, 0x00000000u,
0x00040048u, 0x0000000bu, 0x00000000u, 0x00000012u, 0x00050048u, 0x0000000bu,
0x00000001u, 0x0000000bu, 0x00000001u, 0x00050048u, 0x0000000bu, 0x00000002u,
0x0000000bu, 0x00000003u, 0x00050048u, 0x0000000bu, 0x00000003u, 0x0000000bu,
0x00000004u, 0x00040047u, 0x00000011u, 0x0000001eu, 0x00000000u, 0x00020013u,
0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u, 0x00000006u,
0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u, 0x00040015u,
0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu, 0x00000008u, 0x00000009u,
0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u, 0x00000009u, 0x0006001eu,
0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au, 0x0000000au, 0x00040020u,
0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu, 0x0000000cu, 0x0000000du,
0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u, 0x00000001u, 0x0004002bu,
0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u, 0x00000010u, 0x00000001u,
0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u, 0x00000001u, 0x00040020u,
0x00000013u, 0x00000003u, 0x00000007u, 0x00050036u, 0x00000002u, 0x00000004u,
0x00000000u, 0x00000003u, 0x000200f8u, 0x00000005u, 0x0004003du, 0x00000007u,
0x00000012u, 0x00000011u, 0x00050041u, 0x00000013u, 0x00000014u, 0x0000000du,
0x0000000fu, 0x0003003eu, 0x00000014u, 0x00000012u, 0x000100fdu, 0x00010038u,
};
// OpMemberDecorate <struct-id> <member> <decoration>
constexpr Uint32 kOpMemberDecorate = 72;
constexpr Uint32 kDecorationInvariant = 18;
constexpr Uint32 kSpirvHeaderWordCount = 5;
// Test-side reference walker. Deliberately independent of the production code so a bug in
// the pass cannot hide behind the same helper; only used to count what the pass emitted.
Uint32 CountInvariantMemberDecorations(const Vector<Uint32>& spirv) {
Uint32 count = 0;
for (SizeT i = kSpirvHeaderWordCount; i < spirv.size();) {
const Uint32 wordCount = spirv[i] >> 16;
const Uint32 opcode = spirv[i] & 0xFFFFu;
if (wordCount == 0 || i + wordCount > spirv.size()) {
break;
}
if (opcode == kOpMemberDecorate && wordCount >= 4 && spirv[i + 3] == kDecorationInvariant) {
++count;
}
i += wordCount;
}
return count;
}
template <SizeT WordCount>
Vector<Uint32> ToVector(const Uint32 (&words)[WordCount]) {
return Vector<Uint32>(words, words + WordCount);
}
} // namespace
// --- DecoratePositionInvariantPass ---
TEST(DecoratePositionInvariant, AddsInvariantToThePositionMember) {
const Vector<Uint32> input = ToVector(kPlainVertexSpirv);
ASSERT_EQ(CountInvariantMemberDecorations(input), 0u);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(input, output));
EXPECT_EQ(CountInvariantMemberDecorations(output), 1u);
}
TEST(DecoratePositionInvariant, DoesNotDuplicateAnExistingInvariant) {
const Vector<Uint32> input = ToVector(kAlreadyInvariantVertexSpirv);
ASSERT_EQ(CountInvariantMemberDecorations(input), 1u);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(input, output));
EXPECT_EQ(CountInvariantMemberDecorations(output), 1u);
// The pass reports SuccessWithoutChange here, and SPIRV-Tools asserts (in assert-enabled
// builds) that such a run round-trips byte-identically. Pin that from the outside so an
// assert-enabled CI build cannot be the first thing to discover a violation.
EXPECT_EQ(output, input);
}
TEST(DecoratePositionInvariant, IsIdempotent) {
Vector<Uint32> once;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(ToVector(kPlainVertexSpirv), once));
Vector<Uint32> twice;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(once, twice));
EXPECT_EQ(CountInvariantMemberDecorations(twice), 1u);
}
TEST(DecoratePositionInvariant, OutputStaysAReflectableModule) {
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(ToVector(kPlainVertexSpirv), output));
SpvReflectShaderModule module{};
ASSERT_EQ(spvReflectCreateShaderModule(output.size() * sizeof(Uint32), output.data(), &module),
SPV_REFLECT_RESULT_SUCCESS);
EXPECT_EQ(module.entry_point_count, 1u);
spvReflectDestroyShaderModule(&module);
}
TEST(DecoratePositionInvariant, RejectsGarbageInput) {
const Vector<Uint32> notSpirv{0xdeadbeefu, 0u, 0u, 0u, 0u};
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::DecoratePositionInvariantForVulkan(notSpirv, output));
}
+195
View File
@@ -318,6 +318,49 @@ TEST_F(TextureTest, CopyTextureSubImage2DUsesNamedObjectAndRestoresBinding) {
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR); EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
} }
TEST_F(TextureTest, CopyTextureSubImage2DRejectsCubeMapTargets) {
const ScopedTextureBackendFunctionsOverride backendGuard;
MG_Backend::gBackendFunctionsTable.GL.CopyTexSubImage2D = RecordCopyTexSubImage2D;
g_copyTexSubImage2DCall = {};
GLuint cubeTexture = 0;
MG_Impl::GLImpl::CreateTextures(GL_TEXTURE_CUBE_MAP, 1, &cubeTexture);
MG_Impl::GLImpl::CopyTextureSubImage2D(cubeTexture, 0, 0, 0, 0, 0, 1, 1);
// GL 4.6 sec. 8.8: the 2D form only accepts 2D/1D-array/rectangle effective targets.
EXPECT_FALSE(g_copyTexSubImage2DCall.Called);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
}
TEST_F(TextureTest, ClearTexImageErrorContracts) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 2, 2, 0,
GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
// Zero texture name is INVALID_OPERATION (ARB_clear_texture).
MG_Impl::GLImpl::ClearTexImage(0, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
// A negative level is INVALID_VALUE...
MG_Impl::GLImpl::ClearTexImage(texture, -1, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_VALUE));
// ...but clearing a level that was never defined is INVALID_OPERATION.
MG_Impl::GLImpl::ClearTexImage(texture, 5, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
// A clear region outside the level is INVALID_VALUE.
MG_Impl::GLImpl::ClearTexSubImage(texture, 0, 1, 1, 0, 4, 4, 1,
GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_VALUE));
// An invalid pixel-transfer format is INVALID_ENUM from the shared validators.
MG_Impl::GLImpl::ClearTexImage(texture, 0, GL_NONE, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_ENUM));
}
// GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT is float state that must answer every numeric query: GetFloatv // GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT is float state that must answer every numeric query: GetFloatv
// is authoritative and GetIntegerv would otherwise fall through to its INVALID_ENUM default. // is authoritative and GetIntegerv would otherwise fall through to its INVALID_ENUM default.
TEST_F(TextureTest, MaxTextureMaxAnisotropyIsAnsweredFromTheBackendLimit) { TEST_F(TextureTest, MaxTextureMaxAnisotropyIsAnsweredFromTheBackendLimit) {
@@ -1384,6 +1427,158 @@ TEST_F(TextureTest, TextureStorage1DAndSubImageModifyNamedObjectOnly) {
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR); EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
} }
// Building a mip chain top-down - upload level N, then level 0 - must not destroy the levels
// already uploaded. AllocateLevel used to resize() the storage down to level+1 on every call, so
// the level-0 upload truncated the chain to a single level; the higher level then read back as
// {0,0,0}, IsComplete() rejected the zero-then-nonzero pattern, and DirectGLES answered that by
// skipping the texture's sync entirely. This is the shape KHR-GL33.texture_repeat_mode uses, and
// it accounted for 108 CTS failures in every GL version.
TEST_F(TextureTest, TexImage2DOnLevelZeroKeepsAnAlreadyUploadedHigherLevel) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 49, 23, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 98, 46, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 2u);
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 0), IntVec3(98, 46, 1));
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 1), IntVec3(49, 23, 1));
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// The other half of the contract: respecifying a level 0 that already held an image still drops
// the chain, exactly as before. Minecraft rebinds the block-atlas name and calls glTexImage2D on
// level 0 before uploading the new levels; leaving the previous chain in place would strand a tail
// at the wrong sizes and - because Mojang terminates its chains with a 0x0 level - reproduce the
// same incomplete-texture black atlas the fix above exists to prevent.
TEST_F(TextureTest, TexImage2DRespecifyingAnExistingLevelZeroDropsTheStaleChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 2, 2, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
ASSERT_EQ(mipmapObject->GetMipmapLevelCount(), 3u);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 16, 16, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 1u);
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 0), IntVec3(16, 16, 1));
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// Same-size respecification has to drop the chain too. The Mipmap Levels video setting rebuilds
// the atlas at identical dimensions with a different level count, so a size-change-only test would
// let the old tail survive.
TEST_F(TextureTest, TexImage2DRespecifyingLevelZeroAtTheSameSizeStillDropsTheChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 1u);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// glTexStorage2D defines exactly `levels` levels. AllocateStorage only grows now, so the immutable
// path has to drop a longer pre-existing chain explicitly.
TEST_F(TextureTest, TexStorage2DTrimsALongerPreExistingMipChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 2, 2, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 3, GL_RGBA8, 1, 1, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexStorage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 8, 8);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 2u);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// glTexImage2D used to reject every GL_COMPRESSED_* internal format with GL_INVALID_ENUM, because
// none of them mapped to a TextureInternalFormat and the "unknown format" gate fired. They now
// resolve to the uncompressed storage that backs them - what GL prescribes for the generic formats,
// and a deliberate deviation for RGTC, which ES cannot compress. The (format, type) pairs below are
// the ones KHR-GL33.packed_pixels uploads with, so this table doubles as a pin for those 480 cases.
TEST_F(TextureTest, CompressedInternalFormatsResolveToTheirUncompressedStorage) {
struct Case {
GLenum internalFormat;
GLenum format;
GLenum type;
TextureInternalFormat expected;
};
const Case cases[] = {
{GL_COMPRESSED_RED, GL_RED, GL_UNSIGNED_BYTE, TextureInternalFormat::R8},
{GL_COMPRESSED_RG, GL_RG, GL_UNSIGNED_BYTE, TextureInternalFormat::RG8},
{GL_COMPRESSED_RGB, GL_RGB, GL_UNSIGNED_BYTE, TextureInternalFormat::RGB8},
{GL_COMPRESSED_RGBA, GL_RGBA, GL_UNSIGNED_BYTE, TextureInternalFormat::RGBA8},
{GL_COMPRESSED_SRGB, GL_RGB, GL_UNSIGNED_BYTE, TextureInternalFormat::SRGB8},
{GL_COMPRESSED_SRGB_ALPHA, GL_RGBA, GL_UNSIGNED_BYTE, TextureInternalFormat::SRGB8Alpha8},
{GL_COMPRESSED_RED_RGTC1, GL_RED, GL_UNSIGNED_BYTE, TextureInternalFormat::R8},
{GL_COMPRESSED_RG_RGTC2, GL_RG, GL_UNSIGNED_BYTE, TextureInternalFormat::RG8},
// The signed RGTC pair is uploaded as GL_BYTE and must land on SNORM storage - resolving
// them to plain R8/RG8 would silently reinterpret negative texels.
{GL_COMPRESSED_SIGNED_RED_RGTC1, GL_RED, GL_BYTE, TextureInternalFormat::R8Snorm},
{GL_COMPRESSED_SIGNED_RG_RGTC2, GL_RG, GL_BYTE, TextureInternalFormat::RG8Snorm},
};
MG_Impl::GLImpl::PixelStorei(GL_UNPACK_ALIGNMENT, 1);
for (const auto& c : cases) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, c.internalFormat, 4, 4, 0, c.format, c.type, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
ASSERT_NE(textureObject, nullptr) << "internalFormat 0x" << std::hex << c.internalFormat;
EXPECT_EQ(textureObject->GetFormat(), c.expected) << "internalFormat 0x" << std::hex << c.internalFormat;
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR) << "internalFormat 0x" << std::hex << c.internalFormat;
}
}
// RGTC compresses 4x4 blocks of a 2D image and has no 3D form, so glTexImage3D must reject it even
// though the same enum is accepted on a 2D target. The generic compressed formats carry no such
// restriction and stay legal in 3D.
TEST_F(TextureTest, RgtcInternalFormatsAreRejectedOnThreeDimensionalTargets) {
const GLenum rgtc[] = {GL_COMPRESSED_RED_RGTC1, GL_COMPRESSED_SIGNED_RED_RGTC1, GL_COMPRESSED_RG_RGTC2,
GL_COMPRESSED_SIGNED_RG_RGTC2};
for (const GLenum internalFormat : rgtc) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_3D, texture);
MG_Impl::GLImpl::TexImage3D(GL_TEXTURE_3D, 0, internalFormat, 4, 4, 4, 0, GL_RED, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_INVALID_OPERATION)
<< "internalFormat 0x" << std::hex << internalFormat;
}
GLuint generic = 0;
MG_Impl::GLImpl::GenTextures(1, &generic);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_3D, generic);
MG_Impl::GLImpl::TexImage3D(GL_TEXTURE_3D, 0, GL_COMPRESSED_RGBA, 4, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
TEST_F(TextureTest, TextureStorage3DAndSubImageModifyNamedObjectOnly) { TEST_F(TextureTest, TextureStorage3DAndSubImageModifyNamedObjectOnly) {
GLuint texture = 0; GLuint texture = 0;
MG_Impl::GLImpl::CreateTextures(GL_TEXTURE_3D, 1, &texture); MG_Impl::GLImpl::CreateTextures(GL_TEXTURE_3D, 1, &texture);
@@ -9,7 +9,12 @@
#include "Loader.h" #include "Loader.h"
#include "MG_Util/Types.h" #include "MG_Util/Types.h"
#include <Config.h> #include <Config.h>
#if !defined(__WIN32) && !defined(_WIN32) #if defined(_WIN32)
#ifndef WIN32_LEAN_AND_MEAN
#define WIN32_LEAN_AND_MEAN 1
#endif
#include <windows.h>
#else
#include <dlfcn.h> #include <dlfcn.h>
#endif #endif
@@ -60,7 +65,14 @@ namespace MobileGL::MG_Util::BackendLoader {
#endif #endif
static void* OpenLib(const Vector<String>& names) { static void* OpenLib(const Vector<String>& names) {
#if !defined(__WIN32) && !defined(_WIN32) && (!defined(__APPLE__) || defined(MOBILEGL_IOS)) #if defined(_WIN32)
for (const auto& name : names) {
if (HMODULE lib = LoadLibraryA(name.c_str())) {
MGLOG_I("Loaded GL backend library: %s", name.c_str());
return reinterpret_cast<void*>(lib);
}
}
#elif !defined(__APPLE__) || defined(MOBILEGL_IOS)
static const String LibPathPrefixes[] = { static const String LibPathPrefixes[] = {
#if defined(MOBILEGL_IOS) #if defined(MOBILEGL_IOS)
"@rpath/", "@executable_path/Frameworks/", "@loader_path/Frameworks/", "@rpath/", "@executable_path/Frameworks/", "@loader_path/Frameworks/",
@@ -104,7 +116,9 @@ namespace MobileGL::MG_Util::BackendLoader {
} }
inline void* ProcAddress(void* lib, const char* name) { inline void* ProcAddress(void* lib, const char* name) {
#if !defined(__WIN32) && !defined(_WIN32) && (!defined(__APPLE__) || defined(MOBILEGL_IOS)) #if defined(_WIN32)
return reinterpret_cast<void*>(::GetProcAddress(reinterpret_cast<HMODULE>(lib), name));
#elif !defined(__APPLE__) || defined(MOBILEGL_IOS)
return dlsym(lib, name); return dlsym(lib, name);
#else #else
return nullptr; return nullptr;
@@ -520,6 +534,15 @@ namespace MobileGL::MG_Util::BackendLoader {
#if defined(MOBILEGL_TRACE_ANGLE_VARIANTS) && defined(__ANDROID__) #if defined(MOBILEGL_TRACE_ANGLE_VARIANTS) && defined(__ANDROID__)
void* angleGlesLib = nullptr; void* angleGlesLib = nullptr;
#endif #endif
#if defined(_WIN32)
// ANGLE is the GLES provider on Windows regardless of UseAngle(). Preload
// libGLESv2.dll so libEGL.dll resolves its dependency from the same directory.
if (!OpenLib({"libGLESv2.dll"})) {
MGLOG_E("Failed to open ANGLE libGLESv2.dll");
return;
}
eglLib = OpenLib({"libEGL.dll"});
#else
if (UseAngle()) { if (UseAngle()) {
void* glesLib = OpenLib({"libGLESv2_angle.so"}); void* glesLib = OpenLib({"libGLESv2_angle.so"});
if (!glesLib) { if (!glesLib) {
@@ -541,6 +564,7 @@ namespace MobileGL::MG_Util::BackendLoader {
eglLib = OpenLib({"libEGL.so"}); eglLib = OpenLib({"libEGL.so"});
#endif #endif
} }
#endif // !_WIN32
if (!eglLib) { if (!eglLib) {
MGLOG_E("Failed to open EGL library"); MGLOG_E("Failed to open EGL library");
@@ -811,6 +835,9 @@ namespace MobileGL::MG_Util::BackendLoader {
if (std::strcmp(extension, "GL_EXT_blend_func_extended") == 0) { if (std::strcmp(extension, "GL_EXT_blend_func_extended") == 0) {
caps.SupportsDualSourceBlend = true; caps.SupportsDualSourceBlend = true;
} }
if (std::strcmp(extension, "GL_NV_shader_noperspective_interpolation") == 0) {
caps.SupportsNoperspectiveInterpolation = true;
}
} }
} }
@@ -1050,6 +1050,11 @@ namespace MobileGL {
// factors and layout(index = 1) fragment outputs. GLES core has no dual-source blending, // factors and layout(index = 1) fragment outputs. GLES core has no dual-source blending,
// so without this a draw using a SRC1 factor cannot proceed. // so without this a draw using a SRC1 factor cannot proceed.
Bool SupportsDualSourceBlend = false; Bool SupportsDualSourceBlend = false;
// GL_NV_shader_noperspective_interpolation is present: the driver accepts the
// `noperspective` interpolation qualifier in ESSL. GLES core has none, so without this
// SPIRV-Cross's `#extension ... : require` would fail to compile and MobileGL falls back
// to stripping the NoPerspective decoration (smooth interpolation) via StripNoPerspectivePass.
Bool SupportsNoperspectiveInterpolation = false;
// GL_RENDERER contains "ANGLE". // GL_RENDERER contains "ANGLE".
Bool IsAngleRenderer = false; Bool IsAngleRenderer = false;
// GL_RENDERER contains both "ANGLE" and "llvmpipe". // GL_RENDERER contains both "ANGLE" and "llvmpipe".
@@ -121,6 +121,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.VulkanAPIVersion = DecodeApiVersion(p.apiVersion); caps.VulkanAPIVersion = DecodeApiVersion(p.apiVersion);
caps.DeviceName = p.deviceName; caps.DeviceName = p.deviceName;
caps.DriverVersionString = DecodeDriverVersion(p.driverVersion); caps.DriverVersionString = DecodeDriverVersion(p.driverVersion);
caps.VendorId = p.vendorID;
caps.UniformBufferOffsetAlignment = static_cast<int>(p.limits.minUniformBufferOffsetAlignment); caps.UniformBufferOffsetAlignment = static_cast<int>(p.limits.minUniformBufferOffsetAlignment);
caps.AliasedLineWidthRangeMin = p.limits.lineWidthRange[0]; caps.AliasedLineWidthRangeMin = p.limits.lineWidthRange[0];
caps.AliasedLineWidthRangeMax = p.limits.lineWidthRange[1]; caps.AliasedLineWidthRangeMax = p.limits.lineWidthRange[1];
@@ -210,6 +211,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.VulkanAPIVersion = DecodeApiVersion(properties.apiVersion); caps.VulkanAPIVersion = DecodeApiVersion(properties.apiVersion);
caps.DeviceName = properties.deviceName; caps.DeviceName = properties.deviceName;
caps.DriverVersionString = DecodeDriverVersion(properties.driverVersion); caps.DriverVersionString = DecodeDriverVersion(properties.driverVersion);
caps.VendorId = properties.vendorID;
caps.UniformBufferOffsetAlignment = static_cast<int>(properties.limits.minUniformBufferOffsetAlignment); caps.UniformBufferOffsetAlignment = static_cast<int>(properties.limits.minUniformBufferOffsetAlignment);
caps.AliasedLineWidthRangeMin = properties.limits.lineWidthRange[0]; caps.AliasedLineWidthRangeMin = properties.limits.lineWidthRange[0];
caps.AliasedLineWidthRangeMax = properties.limits.lineWidthRange[1]; caps.AliasedLineWidthRangeMax = properties.limits.lineWidthRange[1];
@@ -15,6 +15,8 @@ namespace MobileGL {
Version VulkanAPIVersion{1, 0, 0}; Version VulkanAPIVersion{1, 0, 0};
String DeviceName; String DeviceName;
String DriverVersionString; String DriverVersionString;
// VkPhysicalDeviceProperties::vendorID, for device-quirk vendor gating.
Uint32 VendorId = 0;
Int UniformBufferOffsetAlignment = 256; Int UniformBufferOffsetAlignment = 256;
Float AliasedLineWidthRangeMin = 1.0f; Float AliasedLineWidthRangeMin = 1.0f;
Float AliasedLineWidthRangeMax = 1.0f; Float AliasedLineWidthRangeMax = 1.0f;
@@ -255,6 +255,37 @@ namespace MobileGL {
return TextureInternalFormat::DepthComponent; return TextureInternalFormat::DepthComponent;
case GL_DEPTH_STENCIL: case GL_DEPTH_STENCIL:
return TextureInternalFormat::DepthStencil; return TextureInternalFormat::DepthStencil;
// Compressed internal formats resolve to the uncompressed storage that backs them.
//
// For the six generic formats this is exactly what GL prescribes: the implementation
// picks a specific compressed format, and when none is available it falls back to the
// corresponding base format. Nothing downstream ever sees a compressed enum, so the
// metrics, pixel-store and backend tables keep their "one format, N bytes per texel"
// invariant instead of each needing a compressed-aware arm.
//
// The four RGTC formats are a deliberate deviation: they are specific formats that GL
// 3.3 requires, but ES exposes no RGTC compressor to hand the data to. Storing the
// texels uncompressed keeps them renderable at the cost of the memory saving, which is
// strictly better than the INVALID_ENUM the application used to get. Note the signed
// variants must land on SNORM storage - CTS uploads them as GL_BYTE.
case GL_COMPRESSED_RED:
case GL_COMPRESSED_RED_RGTC1:
return TextureInternalFormat::R8;
case GL_COMPRESSED_SIGNED_RED_RGTC1:
return TextureInternalFormat::R8Snorm;
case GL_COMPRESSED_RG:
case GL_COMPRESSED_RG_RGTC2:
return TextureInternalFormat::RG8;
case GL_COMPRESSED_SIGNED_RG_RGTC2:
return TextureInternalFormat::RG8Snorm;
case GL_COMPRESSED_RGB:
return TextureInternalFormat::RGB8;
case GL_COMPRESSED_RGBA:
return TextureInternalFormat::RGBA8;
case GL_COMPRESSED_SRGB:
return TextureInternalFormat::SRGB8;
case GL_COMPRESSED_SRGB_ALPHA:
return TextureInternalFormat::SRGB8Alpha8;
case GL_ALPHA: case GL_ALPHA:
case GL_RED: case GL_RED:
return TextureInternalFormat::Red; return TextureInternalFormat::Red;
+225
View File
@@ -401,6 +401,230 @@ namespace MobileGL::MG_Util::SelfTest {
disabledNote); disabledNote);
} }
// Compiles + links a two-stage program on the probe context. Returns 0 on failure and writes a
// human-readable reason into |detail|.
GLuint CompileLinkProgram(const MG_External::GLESFunctionsTable& g, const char* vs, const char* fs,
String& detail) {
const auto compile = [&](GLenum stage, const char* src, GLuint& out) -> bool {
out = g.glCreateShader(stage);
if (out == 0) {
detail = "glCreateShader returned 0";
return false;
}
g.glShaderSource(out, 1, &src, nullptr);
g.glCompileShader(out);
GLint ok = GL_FALSE;
g.glGetShaderiv(out, GL_COMPILE_STATUS, &ok);
if (ok != GL_TRUE) {
GLchar log[512] = {};
GLsizei len = 0;
g.glGetShaderInfoLog(out, static_cast<GLsizei>(sizeof(log) - 1), &len, log);
detail = format("{} shader compile failed: {}",
stage == GL_VERTEX_SHADER ? "vertex" : "fragment",
len > 0 ? log : "(no info log)");
return false;
}
return true;
};
GLuint v = 0, f = 0;
const ScopeGuard delV([&]() { if (v) g.glDeleteShader(v); });
const ScopeGuard delF([&]() { if (f) g.glDeleteShader(f); });
if (!compile(GL_VERTEX_SHADER, vs, v) || !compile(GL_FRAGMENT_SHADER, fs, f)) {
return 0;
}
const GLuint prog = g.glCreateProgram();
if (prog == 0) {
detail = "glCreateProgram returned 0";
return 0;
}
g.glAttachShader(prog, v);
g.glAttachShader(prog, f);
g.glLinkProgram(prog);
GLint linked = GL_FALSE;
g.glGetProgramiv(prog, GL_LINK_STATUS, &linked);
if (linked != GL_TRUE) {
detail = "program link failed";
g.glDeleteProgram(prog);
return 0;
}
return prog;
}
// "noperspective interpolation" row - a real correctness render, not just a compile. A viewport-
// filling quad is drawn with strong perspective (left clip-w 1, right clip-w 8) and a varying that
// runs 0..1 across it. At the screen centre screen-linear interpolation gives 0.5 while perspective-
// correct gives 1/(w+1) ~= 0.11, so reading the centre texel tells the two apart. The varying is
// carried either through the native `noperspective` qualifier (extension present) or through the
// exact gl_Position.w / gl_FragCoord.w rewrite MobileGL applies when it is absent. Verdict:
// PASS - extension present and the native noperspective result is screen-linear;
// WARN - extension absent but the gl_Position.w/gl_FragCoord.w emulation renders screen-linear
// (correct, just the fallback path shipping shader packs hit on such devices);
// FAIL - either path renders perspective-correct / wrong (noperspective does not actually work),
// or the program will not compile/link, or the render errors.
// Requires the probe context to still be current.
void ProbeGlesNoperspective(ReportBuilder& builder, const MG_External::GLESCapabilities& caps,
const MG_External::GLESFunctionsTable& g) {
const Bool native = caps.SupportsNoperspectiveInterpolation;
const String pathNote = native ? "GL_NV_shader_noperspective_interpolation present (native path)"
: "GL_NV_shader_noperspective_interpolation absent (gl_Position.w / "
"gl_FragCoord.w emulation path)";
const auto fail = [&](const String& detail) {
builder.Fail("noperspective interpolation", pathNote + "; " + detail);
};
if (!g.glCreateShader || !g.glShaderSource || !g.glCompileShader || !g.glGetShaderiv ||
!g.glGetShaderInfoLog || !g.glDeleteShader || !g.glCreateProgram || !g.glAttachShader ||
!g.glLinkProgram || !g.glGetProgramiv || !g.glUseProgram || !g.glDeleteProgram ||
!g.glGenFramebuffers || !g.glBindFramebuffer || !g.glDeleteFramebuffers ||
!g.glGenRenderbuffers || !g.glBindRenderbuffer || !g.glRenderbufferStorage ||
!g.glFramebufferRenderbuffer || !g.glDeleteRenderbuffers || !g.glCheckFramebufferStatus ||
!g.glGenBuffers || !g.glBindBuffer || !g.glBufferData || !g.glDeleteBuffers ||
!g.glGetAttribLocation || !g.glVertexAttribPointer || !g.glEnableVertexAttribArray ||
!g.glViewport || !g.glClearColor || !g.glClear || !g.glDrawArrays || !g.glReadPixels ||
!g.glFinish || !g.glGetError) {
fail("the render entry points did not resolve through eglGetProcAddress");
return;
}
// Match MobileGL's own ESSL target (the device's version). At #version 300 es some drivers
// (Adreno) still treat `noperspective` as reserved even with the extension enabled; the ES 3.2
// form the backend actually emits compiles. Emulated shaders are version-agnostic but use the
// same header for consistency.
const Int esslVer = caps.GLESVersion.Major * 100 + caps.GLESVersion.Minor * 10;
const String header = format("#version {} es\n", esslVer >= 300 ? esslVer : 300);
static const char* const kVsNativeBody =
"#extension GL_NV_shader_noperspective_interpolation : require\n"
"in vec4 a_pos;\n"
"in float a_v;\n"
"noperspective out highp float v_out;\n"
"void main() { gl_Position = a_pos; v_out = a_v; }\n";
static const char* const kFsNativeBody =
"#extension GL_NV_shader_noperspective_interpolation : require\n"
"precision highp float;\n"
"noperspective in highp float v_out;\n"
"out vec4 fragColor;\n"
"void main() { fragColor = vec4(v_out, 0.0, 0.0, 1.0); }\n";
// Exactly MobileGL's emulation (verified against EmulateNoPerspectivePass output): pre-multiply
// the varying by clip-w in the vertex stage, recover with gl_FragCoord.w in the fragment stage,
// no noperspective qualifier (so the driver interpolates it perspective-correct).
static const char* const kVsEmuBody =
"in vec4 a_pos;\n"
"in float a_v;\n"
"out highp float v_out;\n"
"void main() { gl_Position = a_pos; v_out = a_v * gl_Position.w; }\n";
static const char* const kFsEmuBody =
"precision highp float;\n"
"in highp float v_out;\n"
"out vec4 fragColor;\n"
"void main() { fragColor = vec4(v_out * gl_FragCoord.w, 0.0, 0.0, 1.0); }\n";
while (g.glGetError() != GL_NO_ERROR) {
}
const String vsSrc = header + (native ? kVsNativeBody : kVsEmuBody);
const String fsSrc = header + (native ? kFsNativeBody : kFsEmuBody);
String linkDetail;
const GLuint prog = CompileLinkProgram(g, vsSrc.c_str(), fsSrc.c_str(), linkDetail);
if (prog == 0) {
fail(native ? "a noperspective program failed to build though the extension is advertised: " +
linkDetail
: "the emulation program failed to build: " + linkDetail);
return;
}
const ScopeGuard delProg([&]() { g.glDeleteProgram(prog); });
// 9x9 so the centre texel (4,4) sits exactly at NDC (0,0).
constexpr GLsizei kDim = 9;
GLuint rbo = 0, fbo = 0, vbo = 0;
g.glGenRenderbuffers(1, &rbo);
const ScopeGuard delRbo([&]() { if (rbo) g.glDeleteRenderbuffers(1, &rbo); });
g.glBindRenderbuffer(GL_RENDERBUFFER, rbo);
g.glRenderbufferStorage(GL_RENDERBUFFER, GL_RGBA8, kDim, kDim);
g.glGenFramebuffers(1, &fbo);
const ScopeGuard delFbo([&]() {
if (fbo) {
g.glBindFramebuffer(GL_FRAMEBUFFER, 0);
g.glDeleteFramebuffers(1, &fbo);
}
});
g.glBindFramebuffer(GL_FRAMEBUFFER, fbo);
g.glFramebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, rbo);
if (g.glCheckFramebufferStatus(GL_FRAMEBUFFER) != GL_FRAMEBUFFER_COMPLETE) {
fail("the probe framebuffer is incomplete");
return;
}
// Interleaved [vec4 clip-pos, float v]. Left w=1, right w=8; x/y pre-multiplied by w so the quad
// still fills NDC after the perspective divide.
const GLfloat verts[] = {
-1.f, -1.f, 0.f, 1.f, 0.f, //
8.f, -8.f, 0.f, 8.f, 1.f, //
-1.f, 1.f, 0.f, 1.f, 0.f, //
8.f, 8.f, 0.f, 8.f, 1.f, //
};
g.glGenBuffers(1, &vbo);
const ScopeGuard delVbo([&]() { if (vbo) g.glDeleteBuffers(1, &vbo); });
g.glBindBuffer(GL_ARRAY_BUFFER, vbo);
g.glBufferData(GL_ARRAY_BUFFER, sizeof(verts), verts, GL_STATIC_DRAW);
g.glUseProgram(prog);
const GLint posLoc = g.glGetAttribLocation(prog, "a_pos");
const GLint vLoc = g.glGetAttribLocation(prog, "a_v");
if (posLoc < 0 || vLoc < 0) {
fail("the probe vertex attributes did not resolve");
return;
}
g.glEnableVertexAttribArray(static_cast<GLuint>(posLoc));
g.glVertexAttribPointer(static_cast<GLuint>(posLoc), 4, GL_FLOAT, GL_FALSE, 5 * sizeof(GLfloat),
reinterpret_cast<const void*>(0));
g.glEnableVertexAttribArray(static_cast<GLuint>(vLoc));
g.glVertexAttribPointer(static_cast<GLuint>(vLoc), 1, GL_FLOAT, GL_FALSE, 5 * sizeof(GLfloat),
reinterpret_cast<const void*>(4 * sizeof(GLfloat)));
g.glViewport(0, 0, kDim, kDim);
g.glClearColor(0.f, 0.f, 0.f, 1.f);
g.glClear(GL_COLOR_BUFFER_BIT);
g.glDrawArrays(GL_TRIANGLE_STRIP, 0, 4);
g.glFinish();
const GLenum drawError = g.glGetError();
if (drawError != GL_NO_ERROR) {
fail(format("GL error 0x{:x} while rendering the probe quad", drawError));
return;
}
GLubyte center[4] = {};
g.glReadPixels(kDim / 2, kDim / 2, 1, 1, GL_RGBA, GL_UNSIGNED_BYTE, center);
const GLenum readError = g.glGetError();
if (readError != GL_NO_ERROR) {
fail(format("GL error 0x{:x} while reading the probe pixel back", readError));
return;
}
// At the centre: screen-linear -> 0.5 (~128); perspective-correct -> 1/(8+1) ~= 0.111 (~28).
const float observed = static_cast<float>(center[0]) / 255.0f;
const int observedByte = center[0];
constexpr float kScreenLinear = 0.5f;
const bool screenLinear = observed > 0.5f * (kScreenLinear + 1.0f / 9.0f); // midpoint ~= 0.306
if (!screenLinear) {
fail(format("the centre texel read {} (~{:.3f}); expected the screen-linear ~0.5 - "
"interpolation came out perspective-correct, so noperspective does not work here",
observedByte, observed));
return;
}
if (native) {
builder.Pass("noperspective interpolation",
pathNote + format("; native noperspective renders screen-linear (centre {} ~= 0.5)",
observedByte));
} else {
builder.Warn("noperspective interpolation",
pathNote +
format("; the emulation renders screen-linear correctly (centre {} ~= 0.5), "
"but this is the fallback path with less driver coverage",
observedByte));
}
}
// Everything the "MobileGL reported ..." rows need from the GLES device probe. // Everything the "MobileGL reported ..." rows need from the GLES device probe.
struct GlesProbeSummary { struct GlesProbeSummary {
Bool capsValid = false; Bool capsValid = false;
@@ -527,6 +751,7 @@ namespace MobileGL::MG_Util::SelfTest {
builder.report.rendererInfo = format("{} ({})", caps.GLESRendererString, caps.GLESVersionString); builder.report.rendererInfo = format("{} ({})", caps.GLESRendererString, caps.GLESVersionString);
EvaluateGlesChecklist(builder, caps, glesFuncs); EvaluateGlesChecklist(builder, caps, glesFuncs);
ProbeGlesTimerQuery(builder, caps, glesFuncs); ProbeGlesTimerQuery(builder, caps, glesFuncs);
ProbeGlesNoperspective(builder, caps, glesFuncs);
builder.report.formatCapabilities.emplace(); builder.report.formatCapabilities.emplace();
MG_Backend::DirectGLES::PopulateFormatCapabilities( MG_Backend::DirectGLES::PopulateFormatCapabilities(
glesFuncs, caps, builder.report.formatCapabilities.value()); glesFuncs, caps, builder.report.formatCapabilities.value());
@@ -16,9 +16,12 @@
#include "SpirvPasses/FlattenInterfaceStructPass.h" #include "SpirvPasses/FlattenInterfaceStructPass.h"
#include "SpirvPasses/RenameSamplerFunctionParameterPass.h" #include "SpirvPasses/RenameSamplerFunctionParameterPass.h"
#include "SpirvPasses/DecomposeWorkgroupVec3Pass.h" #include "SpirvPasses/DecomposeWorkgroupVec3Pass.h"
#include "SpirvPasses/DecoratePositionInvariantPass.h"
#include "SpirvPasses/LowerDrawParametersPass.h" #include "SpirvPasses/LowerDrawParametersPass.h"
#include "SpirvPasses/RebaseInstanceIndexPass.h" #include "SpirvPasses/RebaseInstanceIndexPass.h"
#include "SpirvPasses/StripUboMemberRelaxedPrecisionPass.h" #include "SpirvPasses/StripUboMemberRelaxedPrecisionPass.h"
#include "SpirvPasses/StripNoPerspectivePass.h"
#include "SpirvPasses/EmulateNoPerspectivePass.h"
#include "spirv-tools/libspirv.h" #include "spirv-tools/libspirv.h"
#include "spirv-tools/optimizer.hpp" #include "spirv-tools/optimizer.hpp"
@@ -26,6 +29,7 @@
#include <MG_Backend/BackendObjects.h> #include <MG_Backend/BackendObjects.h>
#include <MG_Util/Converters/GLToStr/GLEnumConverter.h> #include <MG_Util/Converters/GLToStr/GLEnumConverter.h>
#include <MG_Util/Converters/GLToGlslang/ProgramEnumConverter.h> #include <MG_Util/Converters/GLToGlslang/ProgramEnumConverter.h>
#include <cstdlib>
namespace MobileGL { namespace MobileGL {
namespace MG_Util { namespace MG_Util {
@@ -332,6 +336,30 @@ namespace MobileGL {
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options); return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
} }
bool ShaderCompiler::StripNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(StripNoPerspectivePass::CreateStripNoPerspectivePass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::EmulateNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(EmulateNoPerspectivePass::CreateEmulateNoPerspectivePass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary, bool ShaderCompiler::RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) { Vector<uint32_t>& outputBinary) {
using namespace spvtools; using namespace spvtools;
@@ -344,6 +372,18 @@ namespace MobileGL {
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options); return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
} }
bool ShaderCompiler::DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(DecoratePositionInvariantPass::CreateDecoratePositionInvariantPass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::UseUnformattedFloatStorageImagesForVulkan( bool ShaderCompiler::UseUnformattedFloatStorageImagesForVulkan(
const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary) { const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary) {
constexpr SizeT kSpirvHeaderWordCount = 5; constexpr SizeT kSpirvHeaderWordCount = 5;
@@ -33,12 +33,29 @@ namespace MobileGL {
// Only for the DirectGLES transpile path. // Only for the DirectGLES transpile path.
static bool StripUboMemberRelaxedPrecisionForEssl(const Vector<Uint32>& inputBinary, static bool StripUboMemberRelaxedPrecisionForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary); Vector<uint32_t>& outputBinary);
// Removes NoPerspective decorations so SPIRV-Cross emits plain (smooth) ESSL varyings.
// DirectGLES fallback only, for devices lacking GL_NV_shader_noperspective_interpolation
// (SPIRV-Cross would otherwise require that extension and the driver would reject it).
static bool StripNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Emulates noperspective (screen-linear) interpolation via gl_Position.w / gl_FragCoord.w
// so no NV extension is needed; strips what it cannot emulate. DirectGLES fallback for
// devices lacking GL_NV_shader_noperspective_interpolation. See EmulateNoPerspectivePass.
static bool EmulateNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Rebases loads of the InstanceIndex builtin to (InstanceIndex - BaseInstance) so // Rebases loads of the InstanceIndex builtin to (InstanceIndex - BaseInstance) so
// shaders see GL's zero-based gl_InstanceID. Vertex shaders only; DirectVulkan // shaders see GL's zero-based gl_InstanceID. Vertex shaders only; DirectVulkan
// backend only (glslang's relaxed mode aliases gl_InstanceID to gl_InstanceIndex, // backend only (glslang's relaxed mode aliases gl_InstanceID to gl_InstanceIndex,
// which wrongly includes baseInstance). // which wrongly includes baseInstance).
static bool RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary, static bool RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary); Vector<uint32_t>& outputBinary);
// Adds the Invariant decoration to every Position builtin output. GL apps
// routinely rely on cross-program position invariance for multi-pass
// equality depth tests (e.g. GEQUAL re-draws of the same geometry), and
// mobile drivers that optimize per-pipeline break that without the
// decoration. DirectVulkan only.
static bool DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Replaces the declared format of float storage images with Unknown and adds the // Replaces the declared format of float storage images with Unknown and adds the
// matching SPIR-V capabilities. DirectVulkan uses this only when both Vulkan // matching SPIR-V capabilities. DirectVulkan uses this only when both Vulkan
// shaderStorageImage*WithoutFormat features are enabled, allowing the // shaderStorageImage*WithoutFormat features are enabled, allowing the
@@ -12,6 +12,7 @@
#include <cctype> #include <cctype>
#include <initializer_list> #include <initializer_list>
#include <utility> #include <utility>
#include <Config.h>
#include <MG_Backend/BackendObjects.h> #include <MG_Backend/BackendObjects.h>
namespace { namespace {
@@ -95,6 +96,77 @@ namespace {
return masked; return masked;
} }
// Blank out block comments in place, leaving line comments and every other byte where it is.
//
// The passes that follow scan the source as raw text, so block comments have to stop being
// visible to them - but they must not be *deleted*: replacing the bytes with spaces keeps every
// later offset valid and keeps newlines, so glslang's diagnostics still point at the line the
// application wrote. It also has to be lexically aware. A banner line such as
//
// //*** lighting pass ***
//
// contains "/*" one byte in, and a naive search for that opener treats the rest of the file as
// an unterminated comment.
void BlankBlockComments(MobileGL::String& source) {
enum class Region { Code, SingleLineComment, MultiLineComment, QuotedText };
Region region = Region::Code;
char quote = '\0';
bool escaped = false;
for (SizeT pos = 0; pos < source.size(); pos++) {
const char ch = source[pos];
const char next = pos + 1 < source.size() ? source[pos + 1] : '\0';
if (region == Region::Code) {
if (ch == '/' && next == '/') {
pos++;
region = Region::SingleLineComment;
} else if (ch == '/' && next == '*') {
source[pos] = ' ';
source[pos + 1] = ' ';
pos++;
region = Region::MultiLineComment;
} else if (ch == '"' || ch == '\'') {
quote = ch;
escaped = false;
region = Region::QuotedText;
}
continue;
}
if (region == Region::SingleLineComment) {
if (ch == '\n' || ch == '\r') region = Region::Code;
continue;
}
if (region == Region::MultiLineComment) {
if (ch == '*' && next == '/') {
source[pos] = ' ';
source[pos + 1] = ' ';
pos++;
region = Region::Code;
} else if (ch != '\n' && ch != '\r') {
source[pos] = ' ';
}
continue;
}
// GLSL has no multi-line string literals, so a quote that reaches end of line was never
// a literal to begin with - most likely an apostrophe in a #error or #pragma message.
// Ending the region here keeps one stray apostrophe from swallowing the rest of the file.
if (ch == '\n' || ch == '\r') {
region = Region::Code;
} else if (escaped) {
escaped = false;
} else if (ch == '\\') {
escaped = true;
} else if (ch == quote) {
region = Region::Code;
}
}
}
struct CodeToken { struct CodeToken {
String text; String text;
SizeT begin = 0; SizeT begin = 0;
@@ -365,7 +437,21 @@ namespace {
HasIdentifierWithPrefixOutsideAllowed(tokens, "subgroup", {"subgroupInclusiveAdd"}) || HasIdentifierWithPrefixOutsideAllowed(tokens, "subgroup", {"subgroupInclusiveAdd"}) ||
HasIdentifierWithPrefixOutsideAllowed( HasIdentifierWithPrefixOutsideAllowed(
tokens, "gl_Subgroup", tokens, "gl_Subgroup",
{"gl_SubgroupInvocationID", "gl_SubgroupSize", "gl_SubgroupID", "gl_NumSubgroups"})) { {"gl_SubgroupInvocationID", "gl_SubgroupSize", "gl_SubgroupID", "gl_NumSubgroups"}) ||
// ARB/NV spellings of lane-width-sensitive builtins and functions
// (gl_SubGroupSizeARB, ballotARB, gl_WarpSizeNV, shuffleNV, ...) must block the
// rewrite just like their KHR counterparts: they would silently keep native-width
// semantics in a module rewritten to the virtual 32-lane model.
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_SubGroup", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_Warp", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_Thread", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_SMID", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "ballot", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "shuffle", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "readInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "readFirstInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "anyInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "allInvocations", {})) {
return false; return false;
} }
@@ -487,6 +573,23 @@ namespace {
static_cast<unsigned char>(source[1]) == 0xbb && static_cast<unsigned char>(source[2]) == 0xbf; static_cast<unsigned char>(source[1]) == 0xbb && static_cast<unsigned char>(source[2]) == 0xbf;
} }
// The GLSL versions MobileGL is willing to normalize. Anything else in a #version line - a number
// that is not a real language version (329, 331), a bad profile keyword, a float/identifier where
// the integer belongs, or trailing tokens - is left untouched so glslang rejects it, matching
// KHR-GL33.shaders.preprocessor.directive.version_*. The set is deliberately generous (every real
// desktop and ES version) so the normalizer never starts rejecting a form it used to accept.
bool IsRecognizedGlslVersion(unsigned version) {
switch (version) {
case 100: case 110: case 120: case 130: case 140: case 150:
case 300: case 310: case 320:
case 330: case 400: case 410: case 420: case 430:
case 440: case 450: case 460:
return true;
default:
return false;
}
}
struct ShaderLanguageInfo { struct ShaderLanguageInfo {
unsigned version = 110; unsigned version = 110;
MobileGL::ShaderProfile profile = MobileGL::ShaderProfile::Core; MobileGL::ShaderProfile profile = MobileGL::ShaderProfile::Core;
@@ -494,6 +597,9 @@ namespace {
SizeT versionDirectiveEnd = MobileGL::String::npos; SizeT versionDirectiveEnd = MobileGL::String::npos;
bool hasUtf8Bom = false; bool hasUtf8Bom = false;
bool enablesGpuShader5 = false; bool enablesGpuShader5 = false;
// Whether the parsed #version directive is a well-formed one MobileGL should rewrite. A
// malformed directive (see IsRecognizedGlslVersion) is left alone for glslang to reject.
bool hasValidVersionDirective = false;
bool HasVersionDirective() const { return versionDirectiveStart != MobileGL::String::npos; } bool HasVersionDirective() const { return versionDirectiveStart != MobileGL::String::npos; }
}; };
@@ -537,13 +643,25 @@ namespace {
info.versionDirectiveEnd = lineEnd + (hasLineBreak ? 1 : 0); info.versionDirectiveEnd = lineEnd + (hasLineBreak ? 1 : 0);
SkipDirectiveWhitespace(code, probe, lineEnd); SkipDirectiveWhitespace(code, probe, lineEnd);
const MobileGL::String profile = ReadDirectiveIdentifier(code, probe, lineEnd); const MobileGL::String profile = ReadDirectiveIdentifier(code, probe, lineEnd);
if (profile == "es" || profile == "ES") { bool profileTokenValid = true;
if (profile.empty() || profile == "core") {
info.profile = MobileGL::ShaderProfile::Core;
} else if (profile == "es" || profile == "ES") {
info.profile = MobileGL::ShaderProfile::ES; info.profile = MobileGL::ShaderProfile::ES;
} else if (profile == "compatibility") { } else if (profile == "compatibility") {
info.profile = MobileGL::ShaderProfile::Compatibility; info.profile = MobileGL::ShaderProfile::Compatibility;
} else { } else {
// "#version 330 foo": an unrecognized profile keyword. Keep Core for any
// downstream routing, but mark the directive malformed.
info.profile = MobileGL::ShaderProfile::Core; info.profile = MobileGL::ShaderProfile::Core;
profileTokenValid = false;
} }
// Comments are already masked to spaces, so anything non-blank left on the
// line is real trailing garbage: "#version 330 foobar" / "#version 330.0".
SkipDirectiveWhitespace(code, probe, lineEnd);
const bool hasTrailingTokens = probe < lineEnd;
info.hasValidVersionDirective =
IsRecognizedGlslVersion(info.version) && profileTokenValid && !hasTrailingTokens;
} }
} else if (directive == "extension") { } else if (directive == "extension") {
SkipDirectiveWhitespace(code, probe, lineEnd); SkipDirectiveWhitespace(code, probe, lineEnd);
@@ -590,6 +708,17 @@ namespace {
} }
void NormalizeVersionDirective(MobileGL::String& source, const ShaderLanguageInfo& info) { void NormalizeVersionDirective(MobileGL::String& source, const ShaderLanguageInfo& info) {
// A malformed #version (329, 331, bad profile, float/trailing tokens) is left exactly as the
// application wrote it so glslang rejects it - rewriting it to "#version 330 core" would
// silently legalize the CTS directive.version_* rejection cases. Still drop a leading BOM so
// the reported error is the bad version rather than a stray byte-order mark.
if (info.HasVersionDirective() && !info.hasValidVersionDirective) {
if (info.hasUtf8Bom) {
source.erase(0, 3);
}
return;
}
const MobileGL::String replacement = GetNormalizedVersionDirective(info); const MobileGL::String replacement = GetNormalizedVersionDirective(info);
if (info.HasVersionDirective()) { if (info.HasVersionDirective()) {
source.replace(info.versionDirectiveStart, info.versionDirectiveEnd - info.versionDirectiveStart, source.replace(info.versionDirectiveStart, info.versionDirectiveEnd - info.versionDirectiveStart,
@@ -671,7 +800,10 @@ namespace {
void RenameBuiltinShadowingFunction(MobileGL::String& source, const char* from, const char* to) { void RenameBuiltinShadowingFunction(MobileGL::String& source, const char* from, const char* to) {
const MobileGL::String fromName = from; const MobileGL::String fromName = from;
if (!HasSingleLineFunctionDefinition(source, fromName)) { // Decide from a comment-free view. A commented-out definition is not a definition, and
// acting on one renames every genuine call to the builtin to a name nothing defines - which
// then fails to resolve. Line comments survive BlankBlockComments, so this matters.
if (!HasSingleLineFunctionDefinition(MaskCommentsAndQuotedText(source), fromName)) {
return; return;
} }
@@ -744,6 +876,58 @@ namespace {
return info.HasVersionDirective() ? info.versionDirectiveEnd : 0; return info.HasVersionDirective() ? info.versionDirectiveEnd : 0;
} }
// GLSL's #line takes integer expressions only, but plenty of shader-pack preprocessors emit the
// C form with a quoted filename. Deleting every #line outright made those harmless - at the cost
// of __LINE__ reporting the position in MobileGL's rewritten text rather than the one the pack
// author wrote, and of every later diagnostic pointing at the wrong line. Dropping just the
// quoted operand keeps the directive doing its job and still hands glslang something it accepts.
void NormalizeLineDirectives(MobileGL::String& source) {
const MobileGL::String masked = MaskCommentsAndQuotedText(source);
const SizeT versionEnd = FindAfterVersionDirective(source);
MobileGL::String result;
result.reserve(source.size());
SizeT lineStart = 0;
while (lineStart <= source.size()) {
SizeT lineEnd = source.find('\n', lineStart);
const bool lastLine = lineEnd == MobileGL::String::npos;
if (lastLine) lineEnd = source.size();
SizeT probe = lineStart;
while (probe < lineEnd && (source[probe] == ' ' || source[probe] == '\t')) probe++;
const bool isLineDirective = masked.compare(probe, 5, "#line") == 0 &&
(probe + 5 >= lineEnd || !IsIdentifierChar(source[probe + 5]));
if (isLineDirective && lineStart < versionEnd) {
// #version has to be the first token in the shader, so a #line ahead of it could
// never have taken effect. Drop it rather than hand glslang a source it must reject
// - some pack preprocessors emit their directives before the version line.
} else if (isLineDirective) {
// Keep everything up to the first quote that the masker identified as string text.
SizeT quotePos = MobileGL::String::npos;
for (SizeT i = probe + 5; i < lineEnd; i++) {
if (source[i] == '"' || source[i] == '\'') {
quotePos = i;
break;
}
}
if (quotePos != MobileGL::String::npos) {
result.append(source, lineStart, quotePos - lineStart);
} else {
result.append(source, lineStart, lineEnd - lineStart);
}
} else {
result.append(source, lineStart, lineEnd - lineStart);
}
if (lastLine) break;
result.push_back('\n');
lineStart = lineEnd + 1;
}
source = std::move(result);
}
bool IsExtensionAdvertised(MobileGL::GLExtension extension) { bool IsExtensionAdvertised(MobileGL::GLExtension extension) {
const auto& activeBackendObject = MobileGL::MG_Backend::pActiveBackendObject; const auto& activeBackendObject = MobileGL::MG_Backend::pActiveBackendObject;
if (!activeBackendObject) { if (!activeBackendObject) {
@@ -772,15 +956,28 @@ namespace {
return; return;
} }
// Detect the directive on a comment/string-masked copy so a commented-out
// "#extension GL_ARB_gpu_shader_int64" is never turned into a synthesized #error. Comments are
// no longer blanked in the delivered source (glslang handles them), so this pass must mask
// locally like its siblings. Masking preserves offsets, so edits collected against the scan
// apply verbatim to `source`; they are applied back-to-front to keep earlier offsets valid.
const MobileGL::String scan = MaskCommentsAndQuotedText(source);
struct DirectiveEdit {
SizeT pos;
SizeT len;
MobileGL::String replacement;
};
Vector<DirectiveEdit> edits;
SizeT lineStart = 0; SizeT lineStart = 0;
while (lineStart < source.size()) { while (lineStart < scan.size()) {
SizeT lineEnd = source.find('\n', lineStart); SizeT lineEnd = scan.find('\n', lineStart);
const bool hasLineBreak = lineEnd != MobileGL::String::npos; const bool hasLineBreak = lineEnd != MobileGL::String::npos;
if (!hasLineBreak) { if (!hasLineBreak) {
lineEnd = source.size(); lineEnd = scan.size();
} }
const MobileGL::String line = source.substr(lineStart, lineEnd - lineStart); const MobileGL::String line = scan.substr(lineStart, lineEnd - lineStart);
SizeT probe = 0; SizeT probe = 0;
while (probe < line.size() && std::isspace(static_cast<unsigned char>(line[probe]))) { while (probe < line.size() && std::isspace(static_cast<unsigned char>(line[probe]))) {
probe++; probe++;
@@ -822,16 +1019,12 @@ namespace {
const MobileGL::String behavior = TrimDirectiveToken(line.substr(probe)); const MobileGL::String behavior = TrimDirectiveToken(line.substr(probe));
const SizeT replaceLen = lineEnd - lineStart + (hasLineBreak ? 1 : 0); const SizeT replaceLen = lineEnd - lineStart + (hasLineBreak ? 1 : 0);
if (behavior == "require") { if (behavior == "require") {
const MobileGL::String replacement = edits.push_back({lineStart, replaceLen,
"#error GL_ARB_gpu_shader_int64 is not advertised by MobileGL\n"; "#error GL_ARB_gpu_shader_int64 is not advertised by MobileGL\n"});
source.replace(lineStart, replaceLen, replacement);
lineStart += replacement.size();
} else if (behavior == "enable" || behavior == "warn") { } else if (behavior == "enable" || behavior == "warn") {
source.replace(lineStart, replaceLen, "\n"); edits.push_back({lineStart, replaceLen, "\n"});
lineStart++;
} else {
lineStart = lineEnd + (hasLineBreak ? 1 : 0);
} }
lineStart = lineEnd + (hasLineBreak ? 1 : 0);
continue; continue;
} }
} }
@@ -841,6 +1034,10 @@ namespace {
lineStart = lineEnd + (hasLineBreak ? 1 : 0); lineStart = lineEnd + (hasLineBreak ? 1 : 0);
} }
for (auto it = edits.rbegin(); it != edits.rend(); ++it) {
source.replace(it->pos, it->len, it->replacement);
}
ReplaceIdentifier(source, "GL_ARB_gpu_shader_int64", "MG_DISABLED_GL_ARB_gpu_shader_int64"); ReplaceIdentifier(source, "GL_ARB_gpu_shader_int64", "MG_DISABLED_GL_ARB_gpu_shader_int64");
} }
@@ -967,6 +1164,14 @@ namespace MobileGL {
const Vector<CodeToken> tokens = TokenizeCode(source); const Vector<CodeToken> tokens = TokenizeCode(source);
LinearPrefixScanMatch match; LinearPrefixScanMatch match;
if (!ParseLinearPrefixScanTemplate(tokens, match)) { if (!ParseLinearPrefixScanTemplate(tokens, match)) {
// Diagnosability: when the trigger op is present but the template no longer
// matches (e.g. the pack shipped a new shader revision), the affected device
// silently falls back to the driver's miscompiled path. Make that visible.
if (CountToken(tokens, "subgroupInclusiveAdd") > 0) {
MGLOG_W("%s: subgroupInclusiveAdd present but the linear prefix-scan template "
"did not match; the wide-subgroup rewrite was NOT applied",
__func__);
}
return false; return false;
} }
@@ -979,48 +1184,98 @@ namespace MobileGL {
return true; return true;
} }
namespace {
struct ShaderSourceQuirkContext {
ShaderStage stage = ShaderStage::Unknown;
BackendType backend = BackendType::Unknown;
MG_Backend::GpuVendorKind vendor = MG_Backend::GpuVendorKind::Unknown;
Uint32 subgroupSize = 0;
};
// Device-quirk registry. Every entry is a narrowly scoped source rewrite that
// works around a specific driver defect. A quirk runs when its env override
// forces it on, or when the override is Auto and DeviceApplies matches the
// detected device. ForceOn bypasses only the device gate - each Apply keeps
// its own structural safety checks. Add new per-device workarounds here
// instead of open-coding them in PreprocessShaderSource.
struct ShaderSourceQuirk {
const char* name;
MG_Config::QuirkOverride (*GetOverride)();
Bool (*DeviceApplies)(const ShaderSourceQuirkContext&);
Bool (*Apply)(const ShaderSourceQuirkContext&, String&);
};
constexpr ShaderSourceQuirk kShaderSourceQuirks[] = {
{
// MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN
"subgroup-prefix-scan-rewrite",
[] { return MG_Config::Features.SubgroupPrefixScanQuirk; },
[](const ShaderSourceQuirkContext& ctx) {
// Qualcomm's Vulkan driver miscompiles the recognized float
// InclusiveScan pattern for native subgroups wider than the
// captured 32 lanes; other vendors compile it correctly and
// should keep their native scan.
return ctx.backend == BackendType::DirectVulkan &&
ctx.vendor == MG_Backend::GpuVendorKind::Qualcomm;
},
[](const ShaderSourceQuirkContext& ctx, String& source) {
return RewriteLinearSubgroupPrefixScanForVulkan(ctx.stage, ctx.subgroupSize,
source);
},
},
};
void ApplyShaderSourceQuirks(ShaderStage stage, String& source) {
const auto& activeBackend = MG_Backend::pActiveBackendObject;
if (!activeBackend) {
return;
}
const auto& dynamicParameters = activeBackend->GetDynamicParameters();
const ShaderSourceQuirkContext quirkContext{
stage,
activeBackend->GetBackendType(),
dynamicParameters.GpuVendor,
dynamicParameters.SubgroupSize,
};
for (const ShaderSourceQuirk& quirk : kShaderSourceQuirks) {
const MG_Config::QuirkOverride quirkOverride = quirk.GetOverride();
if (quirkOverride == MG_Config::QuirkOverride::ForceOff) {
continue;
}
if (quirkOverride == MG_Config::QuirkOverride::Auto &&
!quirk.DeviceApplies(quirkContext)) {
continue;
}
if (quirk.Apply(quirkContext, source)) {
MGLOG_I("ApplyShaderSourceQuirks: applied '%s'%s", quirk.name,
quirkOverride == MG_Config::QuirkOverride::ForceOn ? " (forced on)" : "");
}
}
}
} // namespace
void PreprocessShaderSource(ShaderStage stage, String& source) { void PreprocessShaderSource(ShaderStage stage, String& source) {
// Normalize while the inspector's source span still refers to the untouched input. Later passes // Normalize while the inspector's source span still refers to the untouched input. Later passes
// remove comments and directives, so any subsequent insertion re-inspects the current source. // remove comments and directives, so any subsequent insertion re-inspects the current source.
const ShaderLanguageInfo originalLanguage = InspectShaderLanguage(source); const ShaderLanguageInfo originalLanguage = InspectShaderLanguage(source);
NormalizeVersionDirective(source, originalLanguage); NormalizeVersionDirective(source, originalLanguage);
// remove multi-line comment // Comments are left intact for glslang's own preprocessor: a block comment is a single
size_t commentStartPos = source.find("/*"); // preprocessing token that collapses to one space even across newlines and inside a
while (commentStartPos != String::npos) { // directive, so blanking it here (which preserved the interior newlines) truncated
size_t commentEndPos = source.find("*/", commentStartPos); // multi-line #define bodies and broke otherwise-valid shaders (KHR-GL3x.shaders.
if (commentEndPos == String::npos) { // preprocessor multiline_comment_define / redefine_object / function_redefinition).
source.erase(commentStartPos); // Every MobileGL pass that must ignore comment/string text already masks them locally
break; // via MaskCommentsAndQuotedText/TokenizeCode, so the source we hand glslang keeps them.
} NormalizeLineDirectives(source);
// + length of "*/"
source = source.replace(commentStartPos, commentEndPos - commentStartPos + 2, "");
commentStartPos = source.find("/*", commentStartPos);
}
// remove #line directives // noperspective is intentionally NOT touched here. It is core in desktop GLSL (1.30+)
SizeT linedirPos = source.find("#line"); // and maps to the core SPIR-V NoPerspective decoration, which DirectVulkan renders
while (linedirPos != String::npos) { // natively and SPIRV-Cross turns into ESSL `noperspective` + the
SizeT newlinePos = source.find('\n', linedirPos); // GL_NV_shader_noperspective_interpolation extension. The old naked substring erase
if (newlinePos == String::npos) { // both discarded that interpolation (shader packs need it) and corrupted any
source.erase(linedirPos); // identifier that merely contained the word. The GLES fallback for devices without
break; // the extension lives in the backend, where device capabilities are known.
}
// Preserve a line break so adjacent preprocessor directives do not merge.
source = source.replace(linedirPos, newlinePos - linedirPos + 1, "\n");
linedirPos = source.find("#line", linedirPos);
}
// remove "noperspective"
const char* str_np = "noperspective";
const SizeT len_np = strlen(str_np);
SizeT noperspectivePos = source.find(str_np);
while (noperspectivePos != String::npos) {
// + length of "\n"
source = source.replace(noperspectivePos, len_np, "");
noperspectivePos = source.find(str_np);
}
FilterUnsupportedGpuShaderInt64(source); FilterUnsupportedGpuShaderInt64(source);
CoerceUniformBlockPackingToStd140(source); CoerceUniformBlockPackingToStd140(source);
@@ -1035,12 +1290,7 @@ namespace MobileGL {
ModernizeLegacyGLSL(stage, source); ModernizeLegacyGLSL(stage, source);
InjectDepthRangeBuiltinShim(stage, source); InjectDepthRangeBuiltinShim(stage, source);
const auto& activeBackend = MG_Backend::pActiveBackendObject; ApplyShaderSourceQuirks(stage, source);
if (stage == ShaderStage::Compute && activeBackend &&
activeBackend->GetBackendType() == BackendType::DirectVulkan) {
RewriteLinearSubgroupPrefixScanForVulkan(stage, activeBackend->GetDynamicParameters().SubgroupSize,
source);
}
} }
Bool RetargetLegacyVersionDirectiveTo460(String& source) { Bool RetargetLegacyVersionDirectiveTo460(String& source) {
@@ -1049,6 +1299,10 @@ namespace MobileGL {
// must not be mistaken for the real one. // must not be mistaken for the real one.
const ShaderLanguageInfo info = InspectShaderLanguage(source); const ShaderLanguageInfo info = InspectShaderLanguage(source);
if (!info.HasVersionDirective()) return false; if (!info.HasVersionDirective()) return false;
// Never rescue a malformed directive to 460: that is precisely what re-legalized the
// CTS directive.version_* rejection cases after the first compile failed. The shader-
// pack retry this exists for only ever sees a valid low version (a real "#version 330").
if (!info.hasValidVersionDirective) return false;
// Only the set NormalizeVersionDirective downgraded: desktop core below 400. ES and // Only the set NormalizeVersionDirective downgraded: desktop core below 400. ES and
// compatibility shaders keep whatever they declared. // compatibility shaders keep whatever they declared.
if (info.profile != ShaderProfile::Core || info.version >= 400) return false; if (info.profile != ShaderProfile::Core || info.version >= 400) return false;
@@ -27,8 +27,10 @@ namespace MobileGL {
// wider than the capture's 32 lanes. For the narrowly recognized, uniform-control- // wider than the capture's 32 lanes. For the narrowly recognized, uniform-control-
// flow template, replace the subgroup-local scan with a shared-memory, strict // flow template, replace the subgroup-local scan with a shared-memory, strict
// left-fold over virtual 32-lane segments. Returns true only when the complete safe // left-fold over virtual 32-lane segments. Returns true only when the complete safe
// template was recognized and rewritten. DirectVulkan calls this through // template was recognized and rewritten. PreprocessShaderSource reaches this through
// PreprocessShaderSource; the explicit entry point exists for deterministic tests. // its device-quirk registry: by default only on detected Qualcomm Vulkan devices,
// overridable either way with MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN=1/0. The explicit
// entry point exists for deterministic tests.
Bool RewriteLinearSubgroupPrefixScanForVulkan(ShaderStage stage, Uint32 nativeSubgroupSize, String& source); Bool RewriteLinearSubgroupPrefixScanForVulkan(ShaderStage stage, Uint32 nativeSubgroupSize, String& source);
// Rewrites a "#version 330 core" directive that PreprocessShaderSource normalized down // Rewrites a "#version 330 core" directive that PreprocessShaderSource normalized down
@@ -0,0 +1,139 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "DecoratePositionInvariantPass.h"
#include "spirv.hpp"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/util/make_unique.h"
#include <algorithm>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
using spvtools::opt::Operand;
// Identifies one member of a decorated struct (gl_PerVertex's Position slot).
struct MemberKey {
uint32_t id = 0;
uint32_t member = 0;
bool operator==(const MemberKey& other) const {
return id == other.id && member == other.member;
}
};
// OpDecorate <target-id> <decoration> [literals...]
// OpMemberDecorate <struct-id> <member> <decoration> [literals...]
constexpr uint32_t kDecorateTargetOperand = 0;
constexpr uint32_t kDecorateDecorationOperand = 1;
constexpr uint32_t kDecorateBuiltInOperand = 2;
constexpr uint32_t kMemberDecorateStructOperand = 0;
constexpr uint32_t kMemberDecorateMemberOperand = 1;
constexpr uint32_t kMemberDecorateDecorationOperand = 2;
constexpr uint32_t kMemberDecorateBuiltInOperand = 3;
} // namespace
spvtools::opt::Pass::Status DecoratePositionInvariantPass::Process() {
auto* irContext = context();
// Collect first: AddAnnotationInst mutates the list being walked.
std::vector<uint32_t> invariantIds;
std::vector<MemberKey> invariantMembers;
std::vector<uint32_t> positionIds;
std::vector<MemberKey> positionMembers;
for (const Instruction& annotation : irContext->annotations()) {
if (annotation.opcode() == spv::Op::OpDecorate) {
if (annotation.NumInOperands() <= kDecorateDecorationOperand) {
continue;
}
const auto decoration = static_cast<spv::Decoration>(
annotation.GetSingleWordInOperand(kDecorateDecorationOperand));
const uint32_t target = annotation.GetSingleWordInOperand(kDecorateTargetOperand);
if (decoration == spv::Decoration::Invariant) {
invariantIds.push_back(target);
} else if (decoration == spv::Decoration::BuiltIn &&
annotation.NumInOperands() > kDecorateBuiltInOperand &&
static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(
kDecorateBuiltInOperand)) == spv::BuiltIn::Position) {
positionIds.push_back(target);
}
} else if (annotation.opcode() == spv::Op::OpMemberDecorate) {
if (annotation.NumInOperands() <= kMemberDecorateDecorationOperand) {
continue;
}
const auto decoration = static_cast<spv::Decoration>(
annotation.GetSingleWordInOperand(kMemberDecorateDecorationOperand));
const MemberKey key{
annotation.GetSingleWordInOperand(kMemberDecorateStructOperand),
annotation.GetSingleWordInOperand(kMemberDecorateMemberOperand)};
if (decoration == spv::Decoration::Invariant) {
invariantMembers.push_back(key);
} else if (decoration == spv::Decoration::BuiltIn &&
annotation.NumInOperands() > kMemberDecorateBuiltInOperand &&
static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(
kMemberDecorateBuiltInOperand)) == spv::BuiltIn::Position) {
positionMembers.push_back(key);
}
}
}
Bool changed = false;
for (const uint32_t target : positionIds) {
if (std::find(invariantIds.begin(), invariantIds.end(), target) != invariantIds.end()) {
continue;
}
irContext->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {target}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::Invariant)}}}));
// Guard against a second Position decoration on the same target.
invariantIds.push_back(target);
changed = true;
}
for (const MemberKey& key : positionMembers) {
if (std::find(invariantMembers.begin(), invariantMembers.end(), key) !=
invariantMembers.end()) {
continue;
}
irContext->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpMemberDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {key.id}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {key.member}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::Invariant)}}}));
invariantMembers.push_back(key);
changed = true;
}
if (!changed) {
return Status::SuccessWithoutChange;
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken
DecoratePositionInvariantPass::CreateDecoratePositionInvariantPass() {
return spvtools::Optimizer::PassToken(MakeUnique<DecoratePositionInvariantPass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,35 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Adds the Invariant decoration to every Position builtin output. GL apps
// routinely rely on cross-program position invariance for multi-pass equality
// depth tests - MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the
// depth its own first pass wrote - and a driver that optimizes each pipeline
// separately may otherwise vary the position math between passes, dropping whole
// primitives from the later ones. Both the plain (OpDecorate on a Position
// variable) and the block-member (OpMemberDecorate on gl_PerVertex) spellings are
// handled; targets that already carry Invariant are left alone. DirectVulkan only.
class DecoratePositionInvariantPass : public spvtools::opt::Pass {
public:
const char* name() const override { return "decorate-position-invariant"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateDecoratePositionInvariantPass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,407 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "EmulateNoPerspectivePass.h"
#include "spirv.hpp"
#include "source/opt/constants.h"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/opt/type_manager.h"
#include "source/opt/types.h"
#include "source/util/make_unique.h"
#include <algorithm>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
using spvtools::opt::Operand;
namespace analysis = spvtools::opt::analysis;
spv::ExecutionModel EntryExecutionModel(IRContext* ctx) {
for (Instruction& ep : ctx->module()->entry_points()) {
return static_cast<spv::ExecutionModel>(ep.GetSingleWordInOperand(0));
}
return spv::ExecutionModel::Max;
}
uint32_t VariablePointeeType(IRContext* ctx, Instruction* var) {
Instruction* ptrType = ctx->get_def_use_mgr()->GetDef(var->type_id());
// OpTypePointer <storage-class> <pointee>
return ptrType->GetSingleWordInOperand(1);
}
// If |typeId| is float or a vector of float, returns true and reports the scalar float
// type and whether it is a vector. Matrices, structs, ints etc. are not emulatable.
bool IsFloatScalarOrVector(IRContext* ctx, uint32_t typeId, uint32_t& floatTypeId, bool& isVector) {
Instruction* t = ctx->get_def_use_mgr()->GetDef(typeId);
if (t == nullptr) return false;
if (t->opcode() == spv::Op::OpTypeFloat) {
floatTypeId = typeId;
isVector = false;
return true;
}
if (t->opcode() == spv::Op::OpTypeVector) {
const uint32_t comp = t->GetSingleWordInOperand(0);
Instruction* ct = ctx->get_def_use_mgr()->GetDef(comp);
if (ct != nullptr && ct->opcode() == spv::Op::OpTypeFloat) {
floatTypeId = comp;
isVector = true;
return true;
}
}
return false;
}
uint32_t PointerTypeTo(IRContext* ctx, uint32_t pointeeId, spv::StorageClass sc) {
analysis::Type* pointee = ctx->get_type_mgr()->GetType(pointeeId);
analysis::Pointer ptr(pointee, sc);
return ctx->get_type_mgr()->GetTypeInstruction(&ptr);
}
uint32_t V4FloatType(IRContext* ctx) {
analysis::Float f(32);
analysis::Type* freg = ctx->get_type_mgr()->GetRegisteredType(&f);
analysis::Vector v(freg, 4);
return ctx->get_type_mgr()->GetTypeInstruction(&v);
}
uint32_t FloatType(IRContext* ctx) {
analysis::Float f(32);
return ctx->get_type_mgr()->GetTypeInstruction(&f);
}
uint32_t SignedIntConstant(IRContext* ctx, int32_t value) {
analysis::Integer i(32, true);
analysis::Type* reg = ctx->get_type_mgr()->GetRegisteredType(&i);
const analysis::Constant* c =
ctx->get_constant_mgr()->GetConstant(reg, {static_cast<uint32_t>(value)});
return ctx->get_constant_mgr()->GetDefiningInstruction(c)->result_id();
}
// Multiply |valueId| (of type |valueTypeId|) by the scalar |scalarId|, inserting the op
// before |before|. Returns the product's id.
uint32_t InsertScale(IRContext* ctx, Instruction* before, uint32_t valueTypeId,
uint32_t valueId, uint32_t scalarId, bool isVector) {
const uint32_t productId = ctx->TakeNextId();
const spv::Op op = isVector ? spv::Op::OpVectorTimesScalar : spv::Op::OpFMul;
before->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, op, valueTypeId, productId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {valueId}},
{SPV_OPERAND_TYPE_ID, {scalarId}}}));
return productId;
}
// --- Vertex stage: gl_Position discovery ------------------------------------------
// Finds gl_Position as member |memberIndex| of a gl_PerVertex-style block whose Output
// variable is |blockVarId|; |v4floatTypeId| is that member's (vec4) type. Returns false
// if gl_Position is not a block member (older plain-variable form is left to the strip).
bool FindPositionBlock(IRContext* ctx, uint32_t& blockVarId, uint32_t& memberIndex,
uint32_t& v4floatTypeId) {
uint32_t structId = 0;
uint32_t member = 0;
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpMemberDecorate && ann.NumInOperands() >= 4 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(2)) ==
spv::Decoration::BuiltIn &&
static_cast<spv::BuiltIn>(ann.GetSingleWordInOperand(3)) ==
spv::BuiltIn::Position) {
structId = ann.GetSingleWordInOperand(0);
member = ann.GetSingleWordInOperand(1);
break;
}
}
if (structId == 0) return false;
Instruction* structType = ctx->get_def_use_mgr()->GetDef(structId);
if (structType == nullptr || member >= structType->NumInOperands()) return false;
v4floatTypeId = structType->GetSingleWordInOperand(member);
for (Instruction& inst : ctx->module()->types_values()) {
if (inst.opcode() == spv::Op::OpVariable &&
static_cast<spv::StorageClass>(inst.GetSingleWordInOperand(0)) ==
spv::StorageClass::Output &&
VariablePointeeType(ctx, &inst) == structId) {
blockVarId = inst.result_id();
memberIndex = member;
return true;
}
}
return false;
}
// --- Fragment stage: gl_FragCoord discovery/synthesis -----------------------------
Instruction* FindBuiltinInput(IRContext* ctx, spv::BuiltIn builtin) {
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() != spv::Op::OpDecorate || ann.NumInOperands() < 3) continue;
if (static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) !=
spv::Decoration::BuiltIn)
continue;
if (static_cast<spv::BuiltIn>(ann.GetSingleWordInOperand(2)) != builtin) continue;
Instruction* var = ctx->get_def_use_mgr()->GetDef(ann.GetSingleWordInOperand(0));
if (var != nullptr && var->opcode() == spv::Op::OpVariable &&
static_cast<spv::StorageClass>(var->GetSingleWordInOperand(0)) ==
spv::StorageClass::Input) {
return var;
}
}
return nullptr;
}
uint32_t SynthesizeFragCoord(IRContext* ctx, uint32_t v4floatTypeId) {
const uint32_t ptrType = PointerTypeTo(ctx, v4floatTypeId, spv::StorageClass::Input);
const uint32_t varId = ctx->TakeNextId();
ctx->AddGlobalValue(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpVariable, ptrType, varId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_STORAGE_CLASS,
{static_cast<uint32_t>(spv::StorageClass::Input)}}}));
ctx->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {varId}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::BuiltIn)}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER,
{static_cast<uint32_t>(spv::BuiltIn::FragCoord)}}}));
for (Instruction& ep : ctx->module()->entry_points()) {
ep.AddOperand({SPV_OPERAND_TYPE_ID, {varId}});
}
return varId;
}
} // namespace
spvtools::opt::Pass::Status EmulateNoPerspectivePass::Process() {
auto* ctx = context();
const spv::ExecutionModel model = EntryExecutionModel(ctx);
const bool isVertex = model == spv::ExecutionModel::Vertex;
const bool isFragment = model == spv::ExecutionModel::Fragment;
// Collect NoPerspective-decorated plain variables and every NoPerspective annotation.
std::vector<uint32_t> plainVarIds;
std::vector<Instruction*> decorationsToKill;
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpDecorate && ann.NumInOperands() >= 2 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) ==
spv::Decoration::NoPerspective) {
plainVarIds.push_back(ann.GetSingleWordInOperand(0));
decorationsToKill.push_back(&ann);
} else if (ann.opcode() == spv::Op::OpMemberDecorate && ann.NumInOperands() >= 3 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(2)) ==
spv::Decoration::NoPerspective) {
// Block-member noperspective is not emulated here; the decoration is stripped
// (smooth fallback) so SPIRV-Cross does not require the NV extension.
decorationsToKill.push_back(&ann);
}
}
if (decorationsToKill.empty()) {
return Status::SuccessWithoutChange;
}
const spv::StorageClass wantStorage =
isVertex ? spv::StorageClass::Output : spv::StorageClass::Input;
// Emulatable = plain variable of the stage's interface direction, float or floatN.
struct Target {
Instruction* var;
uint32_t typeId;
uint32_t floatTypeId;
bool isVector;
};
std::vector<Target> targets;
if (isVertex || isFragment) {
for (const uint32_t id : plainVarIds) {
Instruction* var = ctx->get_def_use_mgr()->GetDef(id);
if (var == nullptr || var->opcode() != spv::Op::OpVariable) continue;
if (static_cast<spv::StorageClass>(var->GetSingleWordInOperand(0)) != wantStorage)
continue;
const uint32_t pointee = VariablePointeeType(ctx, var);
uint32_t floatTypeId = 0;
bool isVector = false;
if (IsFloatScalarOrVector(ctx, pointee, floatTypeId, isVector)) {
targets.push_back({var, pointee, floatTypeId, isVector});
}
}
}
// Force highp on the varyings we emulate: the a*w round-trip overflows a mediump (fp16)
// varying at large clip-space w. Dropping RelaxedPrecision makes SPIRV-Cross emit them
// highp on both stages, keeping the emulation exact. Only touches emulated variables.
if (!targets.empty()) {
std::vector<uint32_t> targetIds;
targetIds.reserve(targets.size());
for (const Target& t : targets) targetIds.push_back(t.var->result_id());
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpDecorate && ann.NumInOperands() >= 2 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) ==
spv::Decoration::RelaxedPrecision &&
std::find(targetIds.begin(), targetIds.end(),
ann.GetSingleWordInOperand(0)) != targetIds.end()) {
decorationsToKill.push_back(&ann);
}
}
}
if (isVertex && !targets.empty()) {
uint32_t blockVarId = 0;
uint32_t memberIndex = 0;
uint32_t v4floatTypeId = 0;
if (FindPositionBlock(ctx, blockVarId, memberIndex, v4floatTypeId)) {
const uint32_t ptrOutV4 =
PointerTypeTo(ctx, v4floatTypeId, spv::StorageClass::Output);
const uint32_t memberConst = SignedIntConstant(ctx, static_cast<int32_t>(memberIndex));
const uint32_t floatTy = FloatType(ctx);
uint32_t entryFuncId = 0;
for (Instruction& ep : ctx->module()->entry_points()) {
// OpEntryPoint <model> <function> "name" <interface...>
entryFuncId = ep.GetSingleWordInOperand(1);
break;
}
// Pre-multiply every target output by gl_Position.w before each return of the
// ENTRY function only. glslang does not inline, so a called helper survives as
// its own OpFunction; instrumenting its returns too would scale the varying
// more than once (w^2), breaking the identity.
for (auto funcIt = ctx->module()->begin(); funcIt != ctx->module()->end(); ++funcIt) {
if (funcIt->result_id() != entryFuncId) continue;
funcIt->ForEachInst([&](Instruction* inst) {
if (inst->opcode() != spv::Op::OpReturn &&
inst->opcode() != spv::Op::OpReturnValue) {
return;
}
const uint32_t posPtrId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpAccessChain, ptrOutV4, posPtrId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {blockVarId}},
{SPV_OPERAND_TYPE_ID, {memberConst}}}));
const uint32_t posId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, v4floatTypeId, posId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {posPtrId}}}));
const uint32_t wId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpCompositeExtract, floatTy, wId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {posId}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {3u}}}));
for (const Target& t : targets) {
const uint32_t valId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, t.typeId, valId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {t.var->result_id()}}}));
const uint32_t scaledId =
InsertScale(ctx, inst, t.typeId, valId, wId, t.isVector);
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpStore, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {t.var->result_id()}},
{SPV_OPERAND_TYPE_ID, {scaledId}}}));
}
});
}
}
}
if (isFragment && !targets.empty()) {
Instruction* fragCoord = FindBuiltinInput(ctx, spv::BuiltIn::FragCoord);
uint32_t fragCoordId = 0;
uint32_t v4floatTypeId = 0;
if (fragCoord != nullptr) {
fragCoordId = fragCoord->result_id();
v4floatTypeId = VariablePointeeType(ctx, fragCoord);
} else {
v4floatTypeId = V4FloatType(ctx);
fragCoordId = SynthesizeFragCoord(ctx, v4floatTypeId);
}
const uint32_t floatTy = FloatType(ctx);
auto* defUse = ctx->get_def_use_mgr();
for (const Target& t : targets) {
// Collect every load that reads the varying. glslang lowers a whole-variable
// read to OpLoad(var), but a single-component read (v.x) to
// OpAccessChain(var) + OpLoad(chain). Both must be scaled; the identity is
// per-component, so scaling one loaded component by gl_FragCoord.w is valid.
std::vector<Instruction*> loads;
defUse->ForEachUser(t.var, [&](Instruction* user) {
if (user->opcode() == spv::Op::OpLoad &&
user->GetSingleWordInOperand(0) == t.var->result_id()) {
loads.push_back(user);
} else if (user->opcode() == spv::Op::OpAccessChain &&
user->GetSingleWordInOperand(0) == t.var->result_id()) {
const uint32_t chainId = user->result_id();
defUse->ForEachUser(user, [&](Instruction* chainUser) {
if (chainUser->opcode() == spv::Op::OpLoad &&
chainUser->GetSingleWordInOperand(0) == chainId) {
loads.push_back(chainUser);
}
});
}
});
// Rewrite `%r = OpLoad %ty %ptr` into
// %orig = OpLoad %ty %ptr
// %fc = OpLoad %v4float %fragCoord
// %w = OpCompositeExtract %float %fc 3
// %r = OpVectorTimesScalar/OpFMul %ty %orig %w (reuse %r: uses stay intact)
// The op is chosen from the LOAD's own result type: a whole-vector load scales
// with OpVectorTimesScalar, a scalar component load with OpFMul.
for (Instruction* load : loads) {
const uint32_t loadType = load->type_id();
uint32_t componentFloat = 0;
bool loadIsVector = false;
if (!IsFloatScalarOrVector(ctx, loadType, componentFloat, loadIsVector)) {
continue;
}
const uint32_t ptrId = load->GetSingleWordInOperand(0);
const uint32_t origId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, loadType, origId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {ptrId}}}));
const uint32_t fcId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, v4floatTypeId, fcId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {fragCoordId}}}));
const uint32_t wId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpCompositeExtract, floatTy, wId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {fcId}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {3u}}}));
load->SetOpcode(loadIsVector ? spv::Op::OpVectorTimesScalar : spv::Op::OpFMul);
load->SetInOperands(Instruction::OperandList{
{SPV_OPERAND_TYPE_ID, {origId}}, {SPV_OPERAND_TYPE_ID, {wId}}});
}
}
}
// Strip every NoPerspective decoration: emulated varyings now transport smooth, and
// non-emulatable ones fall back to smooth.
for (Instruction* dec : decorationsToKill) {
ctx->KillInst(dec);
}
ctx->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken EmulateNoPerspectivePass::CreateEmulateNoPerspectivePass() {
return spvtools::Optimizer::PassToken(MakeUnique<EmulateNoPerspectivePass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,41 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Emulates 'noperspective' (screen-linear) interpolation on GLES devices that lack
// GL_NV_shader_noperspective_interpolation, so no NV extension is required. The hardware
// interpolates perspective-correct; screen-linear L(a) is recovered from the identity
// L(a) = P(a * w) * gl_FragCoord.w
// where P is perspective-correct interpolation and w is the vertex clip-space w. So each
// NoPerspective-decorated output is pre-multiplied by gl_Position.w in the vertex stage
// and each NoPerspective-decorated input is multiplied by gl_FragCoord.w in the fragment
// stage; the decoration is then removed so the varying transports smooth. This is exact
// (modulo float precision - the emulated varyings want highp).
//
// Scope: plain interface variables of float or floatN type. Anything it cannot emulate
// (interface-block members, matrices, or a stage lacking the needed builtin) has its
// NoPerspective decoration stripped instead, degrading to smooth - the same result the
// extension-less fallback produced before, and never invalid SPIR-V. DirectGLES only.
class EmulateNoPerspectivePass : public spvtools::opt::Pass {
public:
const char* name() const override { return "emulate-noperspective"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateEmulateNoPerspectivePass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,72 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "StripNoPerspectivePass.h"
#include "spirv.hpp"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/util/make_unique.h"
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
// OpDecorate <target-id> <decoration> [literals...]
// OpMemberDecorate <struct-id> <member> <decoration> [literals...]
constexpr uint32_t kDecorateDecorationOperand = 1;
constexpr uint32_t kMemberDecorateDecorationOperand = 2;
} // namespace
spvtools::opt::Pass::Status StripNoPerspectivePass::Process() {
auto* irContext = context();
// Collect first: KillInst mutates the annotation list being walked.
std::vector<Instruction*> toKill;
for (Instruction& annotation : irContext->annotations()) {
uint32_t decorationOperand = 0;
if (annotation.opcode() == spv::Op::OpDecorate) {
decorationOperand = kDecorateDecorationOperand;
} else if (annotation.opcode() == spv::Op::OpMemberDecorate) {
decorationOperand = kMemberDecorateDecorationOperand;
} else {
continue;
}
if (annotation.NumInOperands() <= decorationOperand) {
continue;
}
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(decorationOperand)) ==
spv::Decoration::NoPerspective) {
toKill.push_back(&annotation);
}
}
if (toKill.empty()) {
return Status::SuccessWithoutChange;
}
for (Instruction* inst : toKill) {
irContext->KillInst(inst);
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken StripNoPerspectivePass::CreateStripNoPerspectivePass() {
return spvtools::Optimizer::PassToken(MakeUnique<StripNoPerspectivePass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,35 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Removes the NoPerspective decoration from every interface variable and block member.
// DirectGLES fallback only, for devices that lack GL_NV_shader_noperspective_interpolation:
// SPIRV-Cross renders a NoPerspective-decorated varying as ESSL `noperspective` plus
// `#extension GL_NV_shader_noperspective_interpolation : require`, which such a driver
// rejects. Dropping the decoration falls the varying back to smooth (perspective-correct)
// interpolation - the same visible result the old text-level strip produced, but without
// corrupting identifiers and without touching DirectVulkan, where NoPerspective is native.
// (The exact screen-linear emulation via gl_Position.w / gl_FragCoord.w is a later step.)
class StripNoPerspectivePass : public spvtools::opt::Pass {
public:
const char* name() const override { return "strip-noperspective"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateStripNoPerspectivePass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
+157
View File
@@ -0,0 +1,157 @@
# piglit on Android against MobileGL
Run [piglit](https://gitlab.freedesktop.org/mesa/piglit) desktop-GL tests on a
connected Android device with **MobileGL as the OpenGL implementation**, for
both backends:
- `DirectGLES` — MobileGL over the system GLES driver (or ANGLE with
`--use-angle`)
- `DirectVulkan` — MobileGL over the system Vulkan driver
No APK and no on-device python: piglit test binaries run as the adb shell user
from `/data/local/tmp`, create contexts through waffle's `surfaceless_egl`
platform, and waffle is pointed at `libMobileGL.so` (which exports the full
`egl*`/`gl*` API under real names).
## How it fits together
```
piglit test binary (aarch64, bionic)
└─ waffle surfaceless_egl (patched)
├─ WAFFLE_EGL_LIBRARY=libMobileGL.so → dlopen MobileGL as the EGL impl
├─ WAFFLE_GL_LIBRARY=libMobileGL.so → waffle_dl_sym resolves gl* here
├─ WAFFLE_FORCE_GL_CONTEXT_VERSION=33core
│ upgrades low compat context requests to GL 3.3 core (never
│ downgrades) so piglit's supports_gl_compat_version=10 tests run
└─ WAFFLE_ANDROID_WINDOW=imagereader (DirectVulkan only)
windows are AImageReader ANativeWindows instead of EGL pbuffers,
because Android ICDs lack VK_EXT_headless_surface which the
MobileGL pbuffer path needs
└─ libMobileGL.so
├─ DirectGLES: dlopens the real system libEGL.so internally
└─ DirectVulkan: links libvulkan.so
```
Key rule: **never name MobileGL `libEGL.so`** anywhere on `LD_LIBRARY_PATH`
the DirectGLES backend loads the system driver with a bare-soname
`dlopen("libEGL.so")` and would recursively pick itself up.
## One-time setup (host: macOS/Linux with the Android NDK)
```sh
WORK=path/to/workdir && cd $WORK
git clone --depth 1 https://gitlab.freedesktop.org/mesa/piglit.git
git clone --depth 1 https://gitlab.freedesktop.org/mesa/waffle.git
git -C waffle apply $MOBILEGL/tools/piglit-android/patches/waffle-mobilegl-android.patch
git -C piglit apply $MOBILEGL/tools/piglit-android/patches/piglit-mobilegl-android.patch
python3 -m venv venv && ./venv/bin/pip install mako numpy packaging
```
The piglit patch matters beyond build fixes: upstream's
`piglit_dispatch_default_init` runs while the waffle framework is still being
constructed (`gl_fw` is NULL), so the waffle resolvers were never installed and
GL functions bound through the **system** libEGL's `eglGetProcAddress` — every
test silently ran on the raw GLES driver instead of MobileGL.
Build MobileGL for Android:
```sh
cmake -S $MOBILEGL -B $MOBILEGL/build-android-arm64 -G Ninja \
-DCMAKE_TOOLCHAIN_FILE=$NDK/build/cmake/android.toolchain.cmake \
-DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-26 \
-DCMAKE_BUILD_TYPE=RelWithDebInfo
cmake --build $MOBILEGL/build-android-arm64 --target MobileGL -j
$NDK/toolchains/llvm/prebuilt/*/bin/llvm-strip --strip-unneeded \
-o $WORK/libMobileGL-stripped.so $MOBILEGL/build-android-arm64/libMobileGL.so
```
Cross-build waffle (meson; a cross file and a stub `egl.pc` pointing at
MobileGL's bundled EGL 1.5 headers are needed — see `cross-example/`):
```sh
cd $WORK/waffle
meson setup build-android --cross-file $WORK/cross/android-arm64.ini \
-Dbuildtype=release -Dsurfaceless_egl=enabled \
-Dglx=disabled -Dx11_egl=disabled -Dgbm=disabled -Dwayland=disabled \
-Dbuild-tests=false -Dbuild-examples=false -Dprefix=$WORK/prefix
ninja -C build-android && meson install -C build-android
```
Cross-build piglit (needs `PKG_CONFIG_LIBDIR` with the installed `waffle-1.pc`
plus the stub `egl.pc`):
```sh
cd $WORK/piglit && export PKG_CONFIG_LIBDIR=$WORK/prefix/lib/pkgconfig:$WORK/cross/pkgconfig
cmake -S . -B build-android -G Ninja \
-DCMAKE_TOOLCHAIN_FILE=$NDK/build/cmake/android.toolchain.cmake \
-DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-26 \
-DCMAKE_BUILD_TYPE=Release \
-DPIGLIT_USE_WAFFLE=ON -DPIGLIT_BUILD_GL_TESTS=ON \
-DPIGLIT_BUILD_GLES1_TESTS=OFF -DPIGLIT_BUILD_GLES2_TESTS=OFF \
-DPIGLIT_BUILD_GLES3_TESTS=OFF -DPIGLIT_BUILD_EGL_TESTS=OFF \
-DPIGLIT_BUILD_GLX_TESTS=OFF -DPIGLIT_BUILD_WGL_TESTS=OFF \
-DPIGLIT_BUILD_CL_TESTS=OFF -DPIGLIT_BUILD_VK_TESTS=OFF \
-DPIGLIT_BUILD_DMA_BUF_TESTS=OFF -DPIGLIT_USE_GBM=OFF \
-DPIGLIT_USE_WAYLAND=OFF -DPIGLIT_USE_X11=OFF \
-DPYTHON_EXECUTABLE=$WORK/venv/bin/python \
-DOPENGL_INCLUDE_DIR=$MOBILEGL/include \
-DOPENGL_gl_LIBRARY=$SYSROOT/usr/lib/aarch64-linux-android/26/libEGL.so \
-DGLEXT_INCLUDE_DIR=$MOBILEGL/include
ninja -C build-android
```
## Selecting tests
Enumerate on the host with piglit's own profiles (no device needed):
```sh
cd $WORK/piglit
for prof in opengl shader glslparser; do
PIGLIT_BUILD_DIR=$PWD/build-android ./venv/bin/python ./piglit print-cmd \
-t "spec@!opengl 3[.]" -t "spec@glsl-3[.]30" $prof
done > /tmp/gl33.list
```
Group names use `@` separators (`spec@!opengl 3.3@minmax`). The version groups
(`spec@!opengl 1.x…3.3`), GLSL groups (`spec@glsl-1.10…3.30`) plus the ARB
extension groups folded into GL 3.13.3 core give a comprehensive "GL 3.3
core" suite (~15k tests).
## Running
```sh
python3 $MOBILEGL/tools/piglit-android/run_piglit_android.py \
--piglit-root $WORK/piglit --list /tmp/gl33.list \
--backend DirectGLES \
--mobilegl-lib $WORK/libMobileGL-stripped.so \
--waffle-lib $WORK/waffle/build-android/src/waffle/libwaffle-1.so \
--out results-gles
# then the same with --backend DirectVulkan --out results-vk
```
The runner pushes binaries/libs/data (incremental; `--repush` forces), executes
tests serially in chunked on-device shell scripts under `timeout`, parses the
`PIGLIT: {...}` result lines, and writes `results.json` + `summary.txt` +
`raw.log`. Exit-code semantics: parsed result wins; nonzero exit without a
result line = `crash`; toybox timeout exits = `timeout`.
Quick sanity check for the whole stack (waffle build also produces `wflinfo`):
```sh
adb shell 'cd /data/local/tmp/piglit-mgl && env LD_LIBRARY_PATH=$PWD/lib \
WAFFLE_EGL_LIBRARY=libMobileGL.so WAFFLE_GL_LIBRARY=libMobileGL.so \
MOBILEGL_BACKEND_TYPE=DirectVulkan WAFFLE_ANDROID_WINDOW=imagereader \
./wflinfo --platform surfaceless_egl --api gl --version 3.3 --profile core'
```
Expect `OpenGL version string: 3.3.0 MobileGL …, Direct (Vulkan) Backend`.
## Known caveats
- MSAA winsys configs never match (MobileGL exposes two RGBA8888 configs,
samples=0); MSAA FBO tests are unaffected.
- Tests that genuinely require compatibility-profile features will fail on the
forced 3.3 core context; that is honest for a core-only implementation.
- `eglTerminate` at test exit now tears MobileGL down deterministically (see
the EGL-lifecycle refactor); a device-side `mobilegl.log` is written per run
directory for debugging.
+76
View File
@@ -0,0 +1,76 @@
#!/usr/bin/env python3
# MobileGL - tools/piglit-android/compare_results.py
# Copyright (c) 2025-2026 MobileGL-Dev
# Licensed under the GNU Lesser General Public License v3.0:
# https://www.gnu.org/licenses/gpl-3.0.txt
# https://www.gnu.org/licenses/lgpl-3.0.txt
# SPDX-License-Identifier: LGPL-3.0-only
# End of Source File Header
"""Compare two run_piglit_android.py results.json files (e.g. DirectGLES vs
DirectVulkan) and write a markdown report of totals plus categorized diffs."""
import argparse
import json
from collections import Counter
from pathlib import Path
BAD = ('crash', 'timeout', 'fail', 'missing', 'notrun', 'warn')
def load(path):
d = json.loads(Path(path).read_text())
return d
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument('a', help='first results.json')
ap.add_argument('b', help='second results.json')
ap.add_argument('-o', '--out', help='markdown output path')
args = ap.parse_args()
da, db = load(args.a), load(args.b)
na, nb = da['backend'], db['backend']
ta, tb = da['tests'], db['tests']
names = sorted(set(ta) | set(tb))
lines = [f'# piglit: {na} vs {nb}', '']
lines.append(f'| | {na} | {nb} |')
lines.append('|---|---|---|')
ca = Counter(o["status"] for o in ta.values())
cb = Counter(o["status"] for o in tb.values())
for k in sorted(set(ca) | set(cb)):
lines.append(f'| {k} | {ca.get(k, 0)} | {cb.get(k, 0)} |')
lines.append(f'| total | {len(ta)} | {len(tb)} |')
lines.append(f'| elapsed | {da.get("elapsed_sec")}s | {db.get("elapsed_sec")}s |')
lines.append('')
def bucket(pred, title):
rows = [n for n in names
if pred(ta.get(n, {}).get('status', 'absent'),
tb.get(n, {}).get('status', 'absent'))]
if rows:
lines.append(f'## {title} ({len(rows)})')
lines.append('')
for n in rows:
sa = ta.get(n, {}).get('status', 'absent')
sb = tb.get(n, {}).get('status', 'absent')
lines.append(f'- `{n}` — {na}: {sa}, {nb}: {sb}')
lines.append('')
bucket(lambda a, b: a in BAD and b in BAD, 'Bad on both (likely frontend/state-tracker)')
bucket(lambda a, b: a in BAD and b == 'pass', f'Bad only on {na}')
bucket(lambda a, b: a == 'pass' and b in BAD, f'Bad only on {nb}')
bucket(lambda a, b: a == 'skip' and b == 'pass' or a == 'pass' and b == 'skip',
'Skip on one side only')
text = '\n'.join(lines) + '\n'
if args.out:
Path(args.out).write_text(text)
print(f'wrote {args.out}')
else:
print(text)
if __name__ == '__main__':
main()
@@ -0,0 +1,19 @@
; meson cross file for waffle -> aarch64 Android
; Replace NDK_TOOLCHAIN with e.g.
; $HOME/Library/Android/sdk/ndk/27.3.13750724/toolchains/llvm/prebuilt/darwin-x86_64
; and PKGCONFIG_DIR with the directory holding the stub egl.pc.
[binaries]
c = 'NDK_TOOLCHAIN/bin/aarch64-linux-android26-clang'
cpp = 'NDK_TOOLCHAIN/bin/aarch64-linux-android26-clang++'
ar = 'NDK_TOOLCHAIN/bin/llvm-ar'
strip = 'NDK_TOOLCHAIN/bin/llvm-strip'
pkg-config = '/usr/bin/pkg-config'
[host_machine]
system = 'android'
cpu_family = 'aarch64'
cpu = 'aarch64'
endian = 'little'
[properties]
pkg_config_libdir = 'PKGCONFIG_DIR'
@@ -0,0 +1,8 @@
# Stub egl.pc for cross-building waffle/piglit against the NDK.
# The NDK's own EGL headers are 1.4-era; point Cflags at MobileGL's bundled
# EGL 1.5 headers (include/EGL) instead. -lEGL resolves to the NDK stub.
Name: egl
Description: EGL 1.5 headers (MobileGL bundled) + NDK libEGL stub
Version: 1.5
Libs: -lEGL
Cflags: -IMOBILEGL_REPO/include
@@ -0,0 +1,104 @@
diff --git a/CMakeLists.txt b/CMakeLists.txt
index 1a38c21..b0b6ce2 100644
--- a/CMakeLists.txt
+++ b/CMakeLists.txt
@@ -22,7 +22,7 @@ INCLUDE (FindPkgConfig)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
-if(${CMAKE_SYSTEM_NAME} MATCHES "Linux|FreeBSD")
+if(${CMAKE_SYSTEM_NAME} MATCHES "Linux|FreeBSD|Android")
set(DEFAULT_EGL ON)
set(DEFAULT_GLX ON)
set(DEFAULT_WGL OFF)
@@ -499,13 +499,18 @@ if(GBM_FOUND)
endif(HAVE_LIBCACA)
endif(GBM_FOUND)
-if(PIGLIT_BUILD_EGL_TESTS)
- pkg_check_modules(EGL REQUIRED egl)
+# EGL *support* (PIGLIT_HAS_EGL, e.g. the surfaceless_egl waffle platform) is
+# independent of building the EGL test binaries, which additionally need X11.
+pkg_check_modules(EGL egl)
+if(EGL_FOUND)
set(PIGLIT_HAS_EGL True)
add_definitions(-DPIGLIT_HAS_EGL)
include_directories(${EGL_INCLUDE_DIRS})
add_definitions (${EGL_CFLAGS_OTHER})
endif()
+if(PIGLIT_BUILD_EGL_TESTS AND NOT EGL_FOUND)
+ message(FATAL_ERROR "PIGLIT_BUILD_EGL_TESTS requires EGL")
+endif()
# Put all executables into the bin subdirectory
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${piglit_BINARY_DIR}/bin)
diff --git a/tests/CMakeLists.txt b/tests/CMakeLists.txt
index e3c964a..36e2109 100644
--- a/tests/CMakeLists.txt
+++ b/tests/CMakeLists.txt
@@ -35,9 +35,9 @@ add_subdirectory (llvmpipe)
add_subdirectory (perf)
add_subdirectory (wgl)
-IF(EGL_FOUND)
+IF(PIGLIT_BUILD_EGL_TESTS)
add_subdirectory (egl)
-ENDIF(EGL_FOUND)
+ENDIF(PIGLIT_BUILD_EGL_TESTS)
IF(PIGLIT_BUILD_CL_TESTS)
add_subdirectory (cl)
diff --git a/tests/util/piglit-dispatch-init.c b/tests/util/piglit-dispatch-init.c
index 3d9772d..e036eac 100644
--- a/tests/util/piglit-dispatch-init.c
+++ b/tests/util/piglit-dispatch-init.c
@@ -322,14 +322,19 @@ piglit_dispatch_default_init(piglit_dispatch_api api)
break;
}
- if (gl_fw) {
- piglit_dispatch_init(api,
- get_wfl_core_proc,
- get_wfl_ext_proc,
- default_unsupported,
- default_get_proc_address_failure);
- } else
-#endif
+ /* MobileGL piglit harness: always use the waffle resolvers in waffle
+ * builds. This function runs while the waffle framework is still being
+ * constructed, so gl_fw is NULL here and the gl_fw branch never fired;
+ * the fallback resolvers bind GL functions via the SYSTEM libEGL's
+ * eglGetProcAddress (a DT_NEEDED symbol), which on Android is the raw
+ * GLES driver rather than the waffle-selected EGL implementation.
+ */
+ piglit_dispatch_init(api,
+ get_wfl_core_proc,
+ get_wfl_ext_proc,
+ default_unsupported,
+ default_get_proc_address_failure);
+#else
{
piglit_dispatch_init(api,
@@ -338,6 +343,7 @@ piglit_dispatch_default_init(piglit_dispatch_api api)
default_unsupported,
default_get_proc_address_failure);
}
+#endif
already_initialized = true;
}
diff --git a/tests/util/piglit-util-gl.c b/tests/util/piglit-util-gl.c
index 16eb425..7f2796c 100644
--- a/tests/util/piglit-util-gl.c
+++ b/tests/util/piglit-util-gl.c
@@ -47,6 +47,9 @@ bool piglit_is_core_profile;
bool piglit_is_gles(void)
{
const char *version_string = (const char *) glGetString(GL_VERSION);
+ if (getenv("PIGLIT_DEBUG_VERSION_STRING"))
+ printf("piglit: debug: GL_VERSION = \"%s\"\n",
+ version_string ? version_string : "(null)");
return strncmp("OpenGL ES", version_string, 9) == 0;
}
@@ -0,0 +1,445 @@
diff --git a/meson.build b/meson.build
index 7060fce..dcd45ed 100644
--- a/meson.build
+++ b/meson.build
@@ -78,7 +78,11 @@ else
dep_egl = dependency('egl', required : get_option('surfaceless_egl'))
build_surfaceless = dep_egl.found()
- dep_egl = dependency('egl', required : get_option('wayland'))
+ # Don't clobber a dep_egl found for surfaceless_egl when wayland is
+ # disabled (a disabled-feature dependency() returns not-found).
+ if not dep_egl.found()
+ dep_egl = dependency('egl', required : get_option('wayland'))
+ endif
dep_wayland_client = dependency(
'wayland-client', version : '>= 1.10', required : get_option('wayland'),
)
diff --git a/src/waffle/egl/wegl_context.c b/src/waffle/egl/wegl_context.c
index 26362a9..a26c27f 100644
--- a/src/waffle/egl/wegl_context.c
+++ b/src/waffle/egl/wegl_context.c
@@ -2,6 +2,8 @@
// SPDX-License-Identifier: BSD-2-Clause
#include <assert.h>
+#include <stdlib.h>
+#include <string.h>
#include <EGL/egl.h>
#include <EGL/eglext.h>
@@ -19,6 +21,25 @@
#define EGL_CONTEXT_OPENGL_ROBUST_ACCESS 0x31B2
#endif
+// MobileGL piglit harness: WAFFLE_FORCE_GL_CONTEXT_VERSION="33core" upgrades
+// every desktop-GL context request to at least that version/profile (never
+// downgrades). MobileGL implements exactly GL 3.3 core; piglit tests that ask
+// for a low compat context (supports_gl_compat_version=10) would otherwise
+// get a context whose version string reflects the raw backend, and skip.
+static bool
+force_gl_context_version(EGLint *major, EGLint *minor, bool *core)
+{
+ const char *env = getenv("WAFFLE_FORCE_GL_CONTEXT_VERSION");
+ if (!env || strlen(env) < 2)
+ return false;
+ if (env[0] < '1' || env[0] > '9' || env[1] < '0' || env[1] > '9')
+ return false;
+ *major = env[0] - '0';
+ *minor = env[1] - '0';
+ *core = strstr(env, "core") != NULL;
+ return true;
+}
+
static bool
bind_api(struct wegl_platform *plat, int32_t waffle_context_api)
{
@@ -61,13 +82,30 @@ create_real_context(struct wegl_config *config,
context_flags |= EGL_CONTEXT_OPENGL_DEBUG_BIT_KHR;
}
+ // Effective version/profile for desktop GL, possibly upgraded by
+ // WAFFLE_FORCE_GL_CONTEXT_VERSION (never downgraded).
+ EGLint eff_major = attrs->context_major_version;
+ EGLint eff_minor = attrs->context_minor_version;
+ int32_t eff_profile = attrs->context_profile;
+ if (waffle_context_api == WAFFLE_CONTEXT_OPENGL) {
+ EGLint f_major, f_minor;
+ bool f_core;
+ if (force_gl_context_version(&f_major, &f_minor, &f_core) &&
+ 10 * f_major + f_minor > 10 * eff_major + eff_minor) {
+ eff_major = f_major;
+ eff_minor = f_minor;
+ if (f_core)
+ eff_profile = WAFFLE_CONTEXT_CORE_PROFILE;
+ }
+ }
+
switch (waffle_context_api) {
case WAFFLE_CONTEXT_OPENGL:
if (dpy->KHR_create_context) {
attrib_list[i++] = EGL_CONTEXT_MAJOR_VERSION_KHR;
- attrib_list[i++] = attrs->context_major_version;
+ attrib_list[i++] = eff_major;
attrib_list[i++] = EGL_CONTEXT_MINOR_VERSION_KHR;
- attrib_list[i++] = attrs->context_minor_version;
+ attrib_list[i++] = eff_minor;
}
else {
assert(attrs->context_major_version == 1);
@@ -92,9 +130,9 @@ create_real_context(struct wegl_config *config,
}
}
- if (wcore_config_attrs_version_ge(attrs, 32)) {
+ if (10 * eff_major + eff_minor >= 32) {
assert(dpy->KHR_create_context);
- switch (attrs->context_profile) {
+ switch (eff_profile) {
case WAFFLE_CONTEXT_CORE_PROFILE:
attrib_list[i++] = EGL_CONTEXT_OPENGL_PROFILE_MASK_KHR;
attrib_list[i++] = EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT_KHR;
diff --git a/src/waffle/egl/wegl_platform.c b/src/waffle/egl/wegl_platform.c
index 6fe72ba..c3d72b1 100644
--- a/src/waffle/egl/wegl_platform.c
+++ b/src/waffle/egl/wegl_platform.c
@@ -13,11 +13,24 @@
#include "wegl_platform.h"
#ifdef WAFFLE_HAS_ANDROID
-static const char *libEGL_filename = "libEGL.so";
+static const char *libEGL_default_filename = "libEGL.so";
#else
-static const char *libEGL_filename = "libEGL.so.1";
+static const char *libEGL_default_filename = "libEGL.so.1";
#endif
+// MobileGL piglit harness: allow overriding which library provides the EGL
+// entry points (e.g. WAFFLE_EGL_LIBRARY=libMobileGL.so). MobileGL exports the
+// full EGL API under the real egl* names, so waffle can drive it directly
+// while the system libEGL.so stays untouched for MobileGL's own backend use.
+static const char *
+egl_library_name(void)
+{
+ const char *env = getenv("WAFFLE_EGL_LIBRARY");
+ if (env && env[0])
+ return env;
+ return libEGL_default_filename;
+}
+
static bool
supports_egl_khr_display(const struct wegl_platform *plat,
const char *client_extensions)
@@ -120,7 +133,7 @@ wegl_platform_teardown(struct wegl_platform *self)
ok = false;
wcore_errorf(WAFFLE_ERROR_UNKNOWN,
"dlclose(\"%s\") failed: %s",
- libEGL_filename, dlerror());
+ egl_library_name(), dlerror());
}
self->egl.handle = NULL;
}
@@ -132,12 +145,19 @@ wegl_platform_teardown(struct wegl_platform *self)
bool
wegl_platform_init(struct wegl_platform *self, EGLenum egl_platform)
{
- static const char *const dso_names[] = {
+ static const char *dso_names[] = {
"libGL.so.1",
"libGLESv1_CM.so.1",
"libGLESv2.so.2",
};
+ // MobileGL piglit harness: let waffle_dl_sym(WAFFLE_DL_OPENGL, ...)
+ // resolve desktop-GL symbols from an alternate library (MobileGL exports
+ // real gl* names). Android has no libGL.so.1.
+ const char *gl_override = getenv("WAFFLE_GL_LIBRARY");
+ if (gl_override && gl_override[0])
+ dso_names[0] = gl_override;
+
wcore_platform_init(&self->wcore);
posix_platform_init(&self->wcore, dso_names);
@@ -146,11 +166,11 @@ wegl_platform_init(struct wegl_platform *self, EGLenum egl_platform)
// Most Waffle platforms will call eglCreateWindowSurface.
self->egl_surface_type_mask = EGL_WINDOW_BIT;
- self->egl.handle = dlopen(libEGL_filename, RTLD_LAZY | RTLD_LOCAL);
+ self->egl.handle = dlopen(egl_library_name(), RTLD_LAZY | RTLD_LOCAL);
if (!self->egl.handle) {
wcore_errorf(WAFFLE_ERROR_FATAL,
"dlopen(\"%s\") failed: %s",
- libEGL_filename, dlerror());
+ egl_library_name(), dlerror());
return false;
}
@@ -158,7 +178,7 @@ wegl_platform_init(struct wegl_platform *self, EGLenum egl_platform)
self->egl.function = dlsym(self->egl.handle, "egl" #function); \
if (!self->egl.function) { \
wcore_errorf(WAFFLE_ERROR_FATAL, "dlsym(\"%s\", \"%s\") failed: %s", \
- libEGL_filename, "egl" #function, dlerror()); \
+ egl_library_name(), "egl" #function, dlerror()); \
goto error; \
}
diff --git a/src/waffle/meson.build b/src/waffle/meson.build
index 2ae4c40..e336566 100644
--- a/src/waffle/meson.build
+++ b/src/waffle/meson.build
@@ -22,6 +22,16 @@ deps_for_waffle = [
idep_threads,
]
+# MobileGL piglit harness: the Android imagereader window mode in
+# surfaceless_egl needs AImageReader (libmediandk) and ANativeWindow
+# (libandroid).
+if build_surfaceless and host_machine.system() == 'android'
+ deps_for_waffle += [
+ cc.find_library('mediandk'),
+ cc.find_library('android'),
+ ]
+endif
+
files_libwaffle = files(
'api/api_priv.c',
'api/waffle_attrib_list.c',
diff --git a/src/waffle/surfaceless_egl/sl_platform.c b/src/waffle/surfaceless_egl/sl_platform.c
index 30018c1..2c0024c 100644
--- a/src/waffle/surfaceless_egl/sl_platform.c
+++ b/src/waffle/surfaceless_egl/sl_platform.c
@@ -101,7 +101,7 @@ static const struct wcore_platform_vtbl sl_platform_vtbl = {
.destroy = sl_window_destroy,
.show = sl_window_show,
.resize = sl_window_resize,
- .swap_buffers = wegl_surface_swap_buffers,
+ .swap_buffers = sl_window_swap_buffers,
.get_native = sl_window_get_native,
},
};
diff --git a/src/waffle/surfaceless_egl/sl_platform.h b/src/waffle/surfaceless_egl/sl_platform.h
index a97fabe..a1c91aa 100644
--- a/src/waffle/surfaceless_egl/sl_platform.h
+++ b/src/waffle/surfaceless_egl/sl_platform.h
@@ -5,7 +5,6 @@
#include <stdbool.h>
#include <stdlib.h>
-#include <gbm.h>
#undef linux
diff --git a/src/waffle/surfaceless_egl/sl_window.c b/src/waffle/surfaceless_egl/sl_window.c
index 0c54613..d18f8b3 100644
--- a/src/waffle/surfaceless_egl/sl_window.c
+++ b/src/waffle/surfaceless_egl/sl_window.c
@@ -9,18 +9,108 @@
#include "wcore_tinfo.h"
#include "wegl_config.h"
+#include "wegl_display.h"
+#include "wegl_platform.h"
#include "wegl_util.h"
#include "sl_display.h"
#include "sl_platform.h"
#include "sl_window.h"
+#ifdef __ANDROID__
+#include <android/hardware_buffer.h>
+#include <android/native_window.h>
+#include <media/NdkImageReader.h>
+
+// MobileGL piglit harness: back "windows" with an AImageReader-provided
+// ANativeWindow instead of an EGL pbuffer. MobileGL's DirectVulkan backend
+// needs a real ANativeWindow on Android because the device ICD lacks
+// VK_EXT_headless_surface (which its pbuffer path requires).
+static bool
+sl_android_window_mode(void)
+{
+ const char *env = getenv("WAFFLE_ANDROID_WINDOW");
+ return env && strcmp(env, "imagereader") == 0;
+}
+
+static bool
+sl_imagereader_init(struct sl_window *self,
+ struct wcore_config *wc_config,
+ int32_t width, int32_t height)
+{
+ struct wegl_config *config = wegl_config(wc_config);
+ struct wegl_display *dpy = wegl_display(wc_config->display);
+ struct wegl_platform *plat = wegl_platform(dpy->wcore.platform);
+ AImageReader *reader = NULL;
+ ANativeWindow *window = NULL;
+
+ media_status_t st = AImageReader_newWithUsage(
+ width, height, AIMAGE_FORMAT_RGBA_8888,
+ AHARDWAREBUFFER_USAGE_GPU_COLOR_OUTPUT |
+ AHARDWAREBUFFER_USAGE_GPU_SAMPLED_IMAGE,
+ 8, &reader);
+ if (st != AMEDIA_OK || !reader) {
+ wcore_errorf(WAFFLE_ERROR_UNKNOWN,
+ "AImageReader_newWithUsage(%dx%d) failed: %d",
+ width, height, (int)st);
+ return false;
+ }
+
+ st = AImageReader_getWindow(reader, &window);
+ if (st != AMEDIA_OK || !window) {
+ wcore_errorf(WAFFLE_ERROR_UNKNOWN,
+ "AImageReader_getWindow failed: %d", (int)st);
+ AImageReader_delete(reader);
+ return false;
+ }
+
+ // MobileGL reads EGL_WIDTH/EGL_HEIGHT from the window-surface attrib
+ // list (non-standard); unknown attribs are ignored.
+ EGLint attrib_list[] = {
+ EGL_RENDER_BUFFER, EGL_BACK_BUFFER,
+ EGL_WIDTH, width,
+ EGL_HEIGHT, height,
+ EGL_NONE,
+ };
+
+ self->wegl.egl = plat->egl.CreateWindowSurface(
+ dpy->egl, config->egl, (EGLNativeWindowType)window, attrib_list);
+ if (!self->wegl.egl) {
+ wegl_emit_error(plat, "eglCreateWindowSurface");
+ AImageReader_delete(reader);
+ return false;
+ }
+
+ self->reader = reader;
+ return true;
+}
+
+static void
+sl_imagereader_drain(struct sl_window *self)
+{
+ if (!self->reader)
+ return;
+
+ // Keep the BufferQueue from filling up: nobody displays these frames,
+ // so consume-and-drop the latest (which also releases all older ones).
+ AImage *image = NULL;
+ if (AImageReader_acquireLatestImage(self->reader, &image) == AMEDIA_OK &&
+ image)
+ AImage_delete(image);
+}
+#endif // __ANDROID__
+
bool
sl_window_destroy(struct wcore_window *wc_self)
{
struct sl_window *self = sl_window(wegl_surface(wc_self));
bool ok = wegl_surface_teardown(&self->wegl);
+#ifdef __ANDROID__
+ if (self->reader)
+ AImageReader_delete(self->reader);
+#endif
+
free(self);
return ok;
}
@@ -52,6 +142,11 @@ sl_window_create(struct wcore_platform *wc_plat,
wcore_window_init(&self->wegl.wcore, wc_config);
+#ifdef __ANDROID__
+ if (sl_android_window_mode())
+ ok = sl_imagereader_init(self, wc_config, width, height);
+ else
+#endif
ok = wegl_pbuffer_init(&self->wegl, wc_config, width, height);
if (!ok)
goto error;
@@ -71,6 +166,18 @@ sl_window_show(struct wcore_window *wc_self)
return true;
}
+bool
+sl_window_swap_buffers(struct wcore_window *wc_self)
+{
+ bool ok = wegl_surface_swap_buffers(wc_self);
+
+#ifdef __ANDROID__
+ sl_imagereader_drain(sl_window(wegl_surface(wc_self)));
+#endif
+
+ return ok;
+}
+
bool
sl_window_resize(struct wcore_window *wc_self,
int32_t width, int32_t height)
@@ -83,6 +190,36 @@ sl_window_resize(struct wcore_window *wc_self,
struct wcore_tinfo *tinfo;
bool ok = true;
+#ifdef __ANDROID__
+ if (self->reader) {
+ // ImageReader-backed window: build a fresh reader + EGL window
+ // surface at the new size, make it current, then retire the old one.
+ struct sl_window new_win;
+ memset(&new_win, 0, sizeof(new_win));
+ wcore_window_init(&new_win.wegl.wcore, self->wc_config);
+
+ ok = sl_imagereader_init(&new_win, self->wc_config, width, height);
+ if (!ok)
+ return false;
+
+ tinfo = wcore_tinfo_get();
+ wc_ctx = tinfo->current_context;
+
+ ok = wegl_make_current(wc_plat, wc_dpy, &new_win.wegl.wcore, wc_ctx);
+ if (!ok) {
+ wegl_surface_teardown(&new_win.wegl);
+ AImageReader_delete(new_win.reader);
+ return false;
+ }
+
+ wegl_surface_teardown(&self->wegl);
+ AImageReader_delete(self->reader);
+ self->wegl.egl = new_win.wegl.egl;
+ self->reader = new_win.reader;
+ return true;
+ }
+#endif
+
// Create a new pbuffer for the resized window.
ok = wegl_pbuffer_init(&new_wegl, self->wc_config, width, height);
if (!ok)
diff --git a/src/waffle/surfaceless_egl/sl_window.h b/src/waffle/surfaceless_egl/sl_window.h
index 8a96239..65a86ad 100644
--- a/src/waffle/surfaceless_egl/sl_window.h
+++ b/src/waffle/surfaceless_egl/sl_window.h
@@ -9,9 +9,18 @@
struct wcore_platform;
+#ifdef __ANDROID__
+typedef struct AImageReader AImageReader;
+#endif
+
struct sl_window {
struct wegl_surface wegl;
struct wcore_config *wc_config;
+#ifdef __ANDROID__
+ // Non-NULL when the window is backed by an AImageReader ANativeWindow
+ // (WAFFLE_ANDROID_WINDOW=imagereader) instead of an EGL pbuffer.
+ AImageReader *reader;
+#endif
};
DEFINE_CONTAINER_CAST_FUNC(sl_window,
@@ -31,6 +40,9 @@ sl_window_destroy(struct wcore_window *wc_self);
bool
sl_window_show(struct wcore_window *wc_self);
+bool
+sl_window_swap_buffers(struct wcore_window *wc_self);
+
bool
sl_window_resize(struct wcore_window *wc_self,
int32_t width, int32_t height);
+347
View File
@@ -0,0 +1,347 @@
#!/usr/bin/env python3
# MobileGL - tools/piglit-android/run_piglit_android.py
# Copyright (c) 2025-2026 MobileGL-Dev
# Licensed under the GNU Lesser General Public License v3.0:
# https://www.gnu.org/licenses/gpl-3.0.txt
# https://www.gnu.org/licenses/lgpl-3.0.txt
# SPDX-License-Identifier: LGPL-3.0-only
# End of Source File Header
"""Run piglit GL tests on a connected Android device against MobileGL.
Pipeline:
1. Read a test list produced by `piglit print-cmd` ("name ::: command" lines).
2. Rewrite host paths to on-device paths and collect referenced data files.
3. Push binaries/libs/data to the device (incremental, marker-based).
4. Execute tests serially on-device in chunked shell scripts under `timeout`,
with MobileGL served to waffle via WAFFLE_EGL_LIBRARY=libMobileGL.so.
5. Parse "PIGLIT: {...}" result lines, classify, and write results + summary.
The device never needs python or an APK; everything runs as the adb shell user
from /data/local/tmp.
"""
import argparse
import json
import os
import re
import shlex
import subprocess
import sys
import tarfile
import tempfile
import time
from pathlib import Path
RESULT_LINE = re.compile(rb'PIGLIT: *({.*})')
MARK_START = re.compile(rb'@@@T (\d+) START')
MARK_EXIT = re.compile(rb'@@@T (\d+) EXIT (\d+)')
# toybox timeout exit codes: 124 = timed out (SIGTERM), 137 = SIGKILL after -k.
TIMEOUT_EXITS = {124, 137, 142}
def adb(args, serial=None, **kw):
cmd = ['adb'] + (['-s', serial] if serial else []) + args
return subprocess.run(cmd, **kw)
def adb_check(args, serial=None):
r = adb(args, serial=serial, capture_output=True)
if r.returncode != 0:
raise RuntimeError(f"adb {' '.join(args)} failed: {r.stderr.decode(errors='replace')}")
return r.stdout
def parse_list(path):
tests = []
for line in Path(path).read_text().splitlines():
line = line.strip()
if not line or ' ::: ' not in line:
continue
name, cmd = line.split(' ::: ', 1)
tests.append((name.strip(), shlex.split(cmd)))
return tests
def rewrite_cmd(argv, piglit_root, build_dir, device_dir):
"""Map host paths in a test command to device paths.
Returns (device_argv, referenced_host_files).
"""
build_abs = str((piglit_root / build_dir).resolve())
src_abs = str(piglit_root.resolve())
out = []
refs = []
for i, arg in enumerate(argv):
a = arg
if i == 0:
# program: "build-android/bin/foo" or absolute
prog = a if os.path.isabs(a) else str((piglit_root / a).resolve())
refs.append(prog)
out.append(f'{device_dir}/bin/{os.path.basename(prog)}')
continue
p = a if os.path.isabs(a) else None
if p and p.startswith(build_abs + '/generated_tests/'):
refs.append(p)
out.append(p.replace(build_abs + '/generated_tests', device_dir + '/generated_tests', 1))
elif p and p.startswith(build_abs + '/tests/'):
# Serialized profiles record data paths under the build dir, but
# only *generated* data lives there in an out-of-tree build;
# plain test files stay in the source tests/ dir.
if not os.path.exists(p):
p = p.replace(build_abs + '/tests', src_abs + '/tests', 1)
refs.append(p)
out.append(p.replace(src_abs + '/tests', device_dir + '/tests', 1))
else:
refs.append(p)
out.append(p.replace(build_abs + '/tests', device_dir + '/build-tests', 1))
elif p and p.startswith(src_abs + '/tests/'):
refs.append(p)
out.append(p.replace(src_abs + '/tests', device_dir + '/tests', 1))
else:
out.append(a)
return out, refs
def backend_env(backend, device_dir, use_angle=False):
env = {
'LD_LIBRARY_PATH': f'{device_dir}/lib',
'WAFFLE_EGL_LIBRARY': 'libMobileGL.so',
'WAFFLE_GL_LIBRARY': 'libMobileGL.so',
# Upgrade low compat context requests to what MobileGL implements;
# without this, piglit's supports_gl_compat_version=10 tests get the
# raw backend version string and skip themselves.
'WAFFLE_FORCE_GL_CONTEXT_VERSION': '33core',
'PIGLIT_PLATFORM': 'surfaceless_egl',
'PIGLIT_SOURCE_DIR': device_dir,
'MOBILEGL_BACKEND_TYPE': backend,
'MOBILEGL_LOG_FILE_PATH': f'{device_dir}/mobilegl.log',
}
if backend == 'DirectGLES':
env['MOBILEGL_USE_ANGLE'] = '1' if use_angle else '0'
if backend == 'DirectVulkan':
# The device ICD usually lacks VK_EXT_headless_surface, which the
# MobileGL pbuffer path needs; use a real ANativeWindow from
# AImageReader instead (patched waffle surfaceless_egl platform).
env['WAFFLE_ANDROID_WINDOW'] = 'imagereader'
return env
def push_tree(host, dev, serial, marker_name, force=False):
"""Push a directory once, tracked by a marker file on the device."""
marker = f'{dev}.pushed-{marker_name}'
if not force:
r = adb(['shell', f'test -e {shlex.quote(marker)} && echo yes'],
serial=serial, capture_output=True)
if r.stdout.strip() == b'yes':
return False
print(f' pushing {host} -> {dev}')
adb_check(['shell', f'rm -rf {shlex.quote(dev)}'], serial=serial)
adb_check(['push', str(host), dev], serial=serial)
adb_check(['shell', f'touch {shlex.quote(marker)}'], serial=serial)
return True
def push_files(files, strip_prefix, dev_prefix, serial):
"""Tar up referenced data files (relative to strip_prefix) and unpack on device."""
files = sorted(set(files))
if not files:
return
with tempfile.NamedTemporaryFile(suffix='.tar', delete=False) as tf:
tar_path = tf.name
with tarfile.open(tar_path, 'w') as tar:
for f in files:
rel = os.path.relpath(f, strip_prefix)
if rel.startswith('..'):
raise RuntimeError(f'file {f} not under {strip_prefix}')
tar.add(f, arcname=rel)
dev_tar = dev_prefix + '/.data.tar'
adb_check(['shell', f'mkdir -p {shlex.quote(dev_prefix)}'], serial=serial)
adb_check(['push', tar_path, dev_tar], serial=serial)
adb_check(['shell', f'cd {shlex.quote(dev_prefix)} && tar xf .data.tar && rm .data.tar'],
serial=serial)
os.unlink(tar_path)
def classify(exit_code, piglit_results, timed_out):
if timed_out:
return 'timeout'
result = None
for r in piglit_results:
if 'result' in r:
result = r['result']
if result is None:
return 'crash' if exit_code != 0 else 'notrun'
if exit_code not in (0, 1):
return 'crash'
return result
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument('--piglit-root', required=True, type=Path,
help='piglit source checkout (with the Android build dir inside)')
ap.add_argument('--build-dir', default='build-android',
help='Android build dir name inside piglit root')
ap.add_argument('--list', required=True,
help='test list file from `piglit print-cmd` (name ::: cmd)')
ap.add_argument('--backend', required=True, choices=['DirectGLES', 'DirectVulkan'])
ap.add_argument('--use-angle', action='store_true',
help='DirectGLES only: let MobileGL load ANGLE instead of the system driver')
ap.add_argument('--mobilegl-lib', required=True, type=Path,
help='stripped libMobileGL.so for the device')
ap.add_argument('--waffle-lib', required=True, type=Path,
help='cross-built libwaffle-1.so')
ap.add_argument('--device-dir', default='/data/local/tmp/piglit-mgl')
ap.add_argument('--out', required=True, type=Path, help='host results directory')
ap.add_argument('--serial', default=None, help='adb device serial')
ap.add_argument('--timeout', type=int, default=60, help='per-test timeout (seconds)')
ap.add_argument('--chunk', type=int, default=200, help='tests per device-side script')
ap.add_argument('--repush', action='store_true', help='force re-push of bin/lib/tests trees')
args = ap.parse_args()
piglit_root = args.piglit_root.resolve()
build = piglit_root / args.build_dir
dev = args.device_dir.rstrip('/')
args.out.mkdir(parents=True, exist_ok=True)
tests = parse_list(args.list)
if not tests:
print('no tests in list', file=sys.stderr)
return 2
print(f'[1/4] {len(tests)} tests; rewriting paths')
dev_cmds = []
data_refs = []
for name, argv in tests:
dcmd, refs = rewrite_cmd(argv, piglit_root, args.build_dir, dev)
dev_cmds.append((name, dcmd))
data_refs.extend(r for r in refs if not r.split('/')[-2:][0] == 'bin')
print('[2/4] pushing artifacts')
adb_check(['shell', f'mkdir -p {dev} {dev}/lib {dev}/chunks {dev}/logs'], serial=args.serial)
push_tree(build / 'bin', f'{dev}/bin', args.serial, 'bin', args.repush)
adb_check(['shell', f'chmod -R 755 {dev}/bin'], serial=args.serial)
# libs: piglit utils + waffle + MobileGL in one LD_LIBRARY_PATH dir
for lib in sorted((build / 'lib').glob('*.so')):
adb_check(['push', str(lib), f'{dev}/lib/'], serial=args.serial)
adb_check(['push', str(args.waffle_lib), f'{dev}/lib/libwaffle-1.so'], serial=args.serial)
adb_check(['push', str(args.mobilegl_lib), f'{dev}/lib/libMobileGL.so'], serial=args.serial)
# data files referenced by this run (tests/, build-tests/, generated_tests/)
src_abs = str(piglit_root)
build_abs = str(build)
groups = {
(src_abs + '/tests', f'{dev}/tests'): [],
(build_abs + '/tests', f'{dev}/build-tests'): [],
(build_abs + '/generated_tests', f'{dev}/generated_tests'): [],
}
for r in set(data_refs):
for (host_prefix, dev_prefix), bucket in groups.items():
if r.startswith(host_prefix + '/'):
bucket.append(r)
break
for (host_prefix, dev_prefix), bucket in groups.items():
push_files(bucket, host_prefix, dev_prefix, args.serial)
env = backend_env(args.backend, dev, args.use_angle)
env_lines = '\n'.join(f'export {k}={shlex.quote(v)}' for k, v in env.items())
print(f'[3/4] running {len(tests)} tests on {args.backend} '
f'(timeout {args.timeout}s/test, chunks of {args.chunk})')
raw_log = (args.out / 'raw.log').open('wb')
t0 = time.time()
outcomes = {}
chunks = [dev_cmds[i:i + args.chunk] for i in range(0, len(dev_cmds), args.chunk)]
idx_base = 0
for ci, chunk in enumerate(chunks):
lines = ['#!/system/bin/sh', env_lines, f'cd {dev}']
for j, (name, dcmd) in enumerate(chunk):
idx = idx_base + j
quoted = ' '.join(shlex.quote(a) for a in dcmd)
lines.append(f'echo "@@@T {idx} START"')
lines.append(f'timeout -k 5 {args.timeout} {quoted} </dev/null 2>&1')
lines.append(f'echo "@@@T {idx} EXIT $?"')
script = '\n'.join(lines) + '\n'
with tempfile.NamedTemporaryFile('w', suffix='.sh', delete=False) as tf:
tf.write(script)
host_script = tf.name
dev_script = f'{dev}/chunks/chunk{ci}.sh'
adb_check(['push', host_script, dev_script], serial=args.serial)
os.unlink(host_script)
r = adb(['shell', f'sh {dev_script}'], serial=args.serial, capture_output=True)
raw_log.write(r.stdout)
raw_log.flush()
# parse this chunk
cur = None
cur_results = []
cur_out = []
for line in r.stdout.splitlines():
m = MARK_START.search(line)
if m:
cur = int(m.group(1))
cur_results = []
cur_out = []
continue
m = MARK_EXIT.search(line)
if m and cur is not None and int(m.group(1)) == cur:
code = int(m.group(2))
name = dev_cmds[cur][0]
status = classify(code, cur_results, code in TIMEOUT_EXITS)
outcomes[name] = {
'status': status,
'exit': code,
'subtests': {k: v for r_ in cur_results if 'subtest' in r_
for k, v in r_['subtest'].items()},
'tail': b'\n'.join(cur_out[-8:]).decode(errors='replace')
if status in ('crash', 'timeout', 'fail', 'notrun') else '',
}
cur = None
continue
if cur is not None:
cur_out.append(line)
m = RESULT_LINE.search(line)
if m:
try:
cur_results.append(json.loads(m.group(1)))
except json.JSONDecodeError:
pass
# tests whose markers never appeared (adb drop, device reboot)
for j, (name, _) in enumerate(chunk):
outcomes.setdefault(dev_cmds[idx_base + j][0], {
'status': 'missing', 'exit': -1, 'subtests': {}, 'tail': ''})
idx_base += len(chunk)
done = idx_base
print(f' chunk {ci + 1}/{len(chunks)} done ({done}/{len(dev_cmds)}, '
f'{time.time() - t0:.0f}s elapsed)')
raw_log.close()
print('[4/4] writing results')
counts = {}
for name, o in outcomes.items():
counts[o['status']] = counts.get(o['status'], 0) + 1
result_doc = {
'backend': args.backend,
'use_angle': args.use_angle,
'device_dir': dev,
'timeout': args.timeout,
'elapsed_sec': round(time.time() - t0, 1),
'totals': counts,
'tests': outcomes,
}
(args.out / 'results.json').write_text(json.dumps(result_doc, indent=1, sort_keys=True))
lines = [f"backend: {args.backend} tests: {len(outcomes)} elapsed: {result_doc['elapsed_sec']}s"]
lines.append('totals: ' + ', '.join(f'{k}={v}' for k, v in sorted(counts.items())))
for status in ('crash', 'timeout', 'fail', 'missing', 'notrun', 'warn'):
bad = sorted(n for n, o in outcomes.items() if o['status'] == status)
if bad:
lines.append(f'\n== {status} ({len(bad)}):')
lines.extend(f' {n}' for n in bad)
(args.out / 'summary.txt').write_text('\n'.join(lines) + '\n')
print('\n'.join(lines[:2]))
print(f"results: {args.out / 'results.json'}")
return 0
if __name__ == '__main__':
sys.exit(main())
@@ -0,0 +1,112 @@
---
name: piglit-on-android
description: Run piglit desktop-GL tests on a connected Android device with MobileGL as the OpenGL implementation - cross-build waffle (patched) and piglit for aarch64, push to /data/local/tmp, execute over adb against the DirectGLES (system driver or ANGLE) and DirectVulkan backends, and produce pass/fail/crash summaries. Use when validating MobileGL's OpenGL 3.3 core conformance on-device or bisecting a piglit regression.
---
# piglit on Android against MobileGL
## Variables
```sh
export REPO="$PWD" # MobileGL checkout
export WORK="$HOME/piglit-android" # piglit/waffle workdir (outside the repo)
export NDK="$HOME/Library/Android/sdk/ndk/27.3.13750724"
export TOOLS="$REPO/tools/piglit-android"
export DEVICE_DIR="/data/local/tmp/piglit-mgl"
```
## Architecture (read first)
- MobileGL exports the full `egl*`/`gl*` API with real symbol names from
`libMobileGL.so`; waffle's `surfaceless_egl` platform is pointed at it via
`WAFFLE_EGL_LIBRARY=libMobileGL.so` and `WAFFLE_GL_LIBRARY=libMobileGL.so`.
- **Never ship MobileGL under the name `libEGL.so`**: the DirectGLES backend
dlopens the system driver by the bare soname `libEGL.so` and would load
itself recursively.
- `WAFFLE_FORCE_GL_CONTEXT_VERSION=33core` (patched waffle) upgrades piglit's
low compat context requests (`supports_gl_compat_version=10`) to GL 3.3
core; without it those tests see a raw backend version string and skip.
- DirectVulkan cannot use MobileGL's EGL-pbuffer path on real devices (Android
ICDs lack `VK_EXT_headless_surface`), so the patched waffle creates windows
from an `AImageReader` ANativeWindow: `WAFFLE_ANDROID_WINDOW=imagereader`.
The runner sets this automatically for `--backend DirectVulkan`.
- The piglit patch also fixes `piglit_dispatch_default_init` to always install
the waffle resolvers; upstream silently bound gl* through the SYSTEM
libEGL's `eglGetProcAddress`, running every test on the raw GLES driver.
- MobileGL's lifecycle is owned by the EGL layer (lazy init on first EGL call,
full teardown on the last `eglTerminate`), so piglit processes start and
exit cleanly with no static ctor/dtor involvement.
## Steps
1. **Clone + patch** piglit and waffle (shallow clones are fine; never pull
git-LFS):
```sh
mkdir -p "$WORK" && cd "$WORK"
git clone --depth 1 https://gitlab.freedesktop.org/mesa/piglit.git
git clone --depth 1 https://gitlab.freedesktop.org/mesa/waffle.git
git -C waffle apply "$TOOLS/patches/waffle-mobilegl-android.patch"
git -C piglit apply "$TOOLS/patches/piglit-mobilegl-android.patch"
python3 -m venv venv && ./venv/bin/pip install mako numpy packaging
```
2. **Build libMobileGL.so** for arm64 (plain CMake + NDK; gradle not needed)
and strip it. See `$TOOLS/README.md` for exact invocations.
3. **Cross-build waffle** with meson (`surfaceless_egl` enabled, everything
else disabled) using `$TOOLS/cross-example/android-arm64.ini` and the stub
`egl.pc` (Cflags → `$REPO/include` for EGL 1.5 headers). `meson install` to
`$WORK/prefix` so piglit's pkg-config finds `waffle-1.pc`.
4. **Cross-build piglit** with the NDK toolchain file; the full CMake flag set
is in `$TOOLS/README.md`. All GL test binaries land in
`piglit/build-android/bin` (~1600), shared utils in `lib/`.
5. **Smoke test** with wflinfo on both backends before any long run:
```sh
adb shell "cd $DEVICE_DIR && env LD_LIBRARY_PATH=$DEVICE_DIR/lib \
WAFFLE_EGL_LIBRARY=libMobileGL.so WAFFLE_GL_LIBRARY=libMobileGL.so \
MOBILEGL_BACKEND_TYPE=DirectVulkan WAFFLE_ANDROID_WINDOW=imagereader \
./wflinfo --platform surfaceless_egl --api gl --version 3.3 --profile core"
```
Expect `3.3.0 MobileGL … Direct (Vulkan) Backend` and exit code 0.
6. **Enumerate tests on the host** with `piglit print-cmd` (profiles: `opengl`,
`shader`, `glslparser`; group separator is `@`). For the GL 3.3 core suite
use the version groups + GLSL groups + the ARB extension groups folded into
3.13.3 core (~15k tests).
7. **Run per backend** with the driver script (serial, chunked, per-test
timeout; DirectGLES first, then DirectVulkan):
```sh
python3 "$TOOLS/run_piglit_android.py" \
--piglit-root "$WORK/piglit" --list gl33-core.list \
--backend DirectGLES \
--mobilegl-lib "$WORK/libMobileGL-stripped.so" \
--waffle-lib "$WORK/waffle/build-android/src/waffle/libwaffle-1.so" \
--out results-gles
```
`--use-angle` switches DirectGLES to packaged ANGLE; default is the system
GLES driver. `--repush` forces re-pushing bin/lib/tests trees (needed after
rebuilds or if the device wiped `/data/local/tmp`).
8. **Compare backends**: `results.json` has per-test status
(`pass/fail/skip/crash/timeout/missing`) and output tails for non-passes;
`summary.txt` lists the bad tests. Diff the two runs' failing sets to
separate frontend issues (fail on both) from backend-specific ones.
## Gotchas
- A test crash after `PIGLIT: {"result": ...}` was printed is classified by
the result line, not the exit code, except that unknown nonzero exits count
as `crash`.
- The device may clear `/data/local/tmp` (vendor cleaners); `--repush` recovers.
- Keep chunks ≤ ~250 tests: one adb connection per chunk bounds the damage of
USB hiccups, and progress prints per chunk.
- MobileGL writes `$DEVICE_DIR/mobilegl.log` (set by the runner); check it and
`adb logcat -b crash` when triaging crashes.
@@ -0,0 +1,4 @@
interface:
display_name: "MobileGL piglit on Android"
short_description: "Run piglit GL tests on-device against MobileGL backends"
default_prompt: "Use $piglit-on-android to run the piglit OpenGL 3.3 core suite on the connected Android device against MobileGL's DirectGLES and DirectVulkan backends and summarize the results."
+9
View File
@@ -93,6 +93,15 @@ The bundled fixtures cover:
- minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and - minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
iterationT after entering a singleplayer world, with Iris' DSA path disabled. iterationT after entering a singleplayer world, with Iris' DSA path disabled.
![Minecraft 1.21.4 Fabric Iris iterationT no-DSA in-world golden](fixtures/minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world.0000115019.png) ![Minecraft 1.21.4 Fabric Iris iterationT no-DSA in-world golden](fixtures/minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world.0000115019.png)
- minecraft-1.21.4-fabric-iris-iterationrp-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
iterationRP after entering a singleplayer world, framing the iterationRP name overlay over a lake with far-shore
tree reflections. iterationRP's temporal auto-exposure makes a single-frame trim overexpose and drop the overlay,
so the fixture is a prefix trace (all calls up to the target frame) that replays the temporal state. The pack also
gates an NVIDIA-only shadow path (`subgroupPartitionNV`, `GL_NV_shader_subgroup_partitioned`) on the GL vendor
string, so the capture reports a masked vendor and the trace carries the portable `subgroupShuffleXor` path that
non-NVIDIA GPUs take.
The trace archive and golden are not committed yet (the repository's Git LFS quota rejects new objects with
`GH009`); the case stays registered and its fixture files are hydrated from the trace fixture mirror.
- minecraft-1.21.4-fabric-iris-photon-v1.1-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and - minecraft-1.21.4-fabric-iris-photon-v1.1-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
Photon v1.1 after entering a singleplayer world. Photon v1.1 after entering a singleplayer world.
![Minecraft 1.21.4 Fabric Iris Photon v1.1 in-world golden](fixtures/minecraft-1.21.4-fabric-iris-photon-v1.1-in-world.0000159866.png) ![Minecraft 1.21.4 Fabric Iris Photon v1.1 in-world golden](fixtures/minecraft-1.21.4-fabric-iris-photon-v1.1-in-world.0000159866.png)
@@ -1,4 +0,0 @@
interface:
display_name: "Capture RenderDoc Trace Frame"
short_description: "Capture exact Android Vulkan/GLES trace frames"
default_prompt: "Use $renderdoc-capture-trace-frame to capture and validate an exact Android retrace frame."
@@ -19,7 +19,12 @@ RESULT_ROOT = ROOT / ".trace-work" / "macos-window-retrace-result"
WORK_ROOT = ROOT / ".trace-work" / "macos-window-retrace-work" WORK_ROOT = ROOT / ".trace-work" / "macos-window-retrace-work"
SUMMARY_DIR = ROOT / ".trace-work" / "macos-window-retrace-summary" SUMMARY_DIR = ROOT / ".trace-work" / "macos-window-retrace-summary"
SUMMARY_HTML = "mobilegl-macos-window-vulkan-retrace-overview.html" SUMMARY_HTML = "mobilegl-macos-window-vulkan-retrace-overview.html"
DEFAULT_MIRROR_BASE = "https://repo.miawa.cn/mgl/tools/trace_replay/fixtures" # Fixture mirrors, tried in order before Git LFS.
DEFAULT_MIRROR_BASES = [
"https://git.hit.moe/swung0x48/MobileGL/media/branch/dev/tools/trace_replay/fixtures",
"https://repo.miawa.cn/mgl/tools/trace_replay/fixtures",
]
DEFAULT_MIRROR_BASE = DEFAULT_MIRROR_BASES[0]
BACKENDS = ("DirectVulkan",) BACKENDS = ("DirectVulkan",)
CASES = load_trace_cases() CASES = load_trace_cases()
@@ -100,34 +105,35 @@ def default_vulkan_icd(mobilegl_library):
return "" return ""
def download_fixture(path, mirror_base): def download_fixture(path, mirror_bases):
"""Try each mirror in order; return on the first that serves a good file."""
if isinstance(mirror_bases, str):
mirror_bases = [mirror_bases]
path.parent.mkdir(parents=True, exist_ok=True) path.parent.mkdir(parents=True, exist_ok=True)
url = f"{mirror_base.rstrip('/')}/{path.name}"
tmp = path.with_suffix(path.suffix + ".tmp") tmp = path.with_suffix(path.suffix + ".tmp")
command = [ token = os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_TOKEN", "")
"curl", errors = []
"-L", for mirror_base in mirror_bases:
"--fail", url = f"{mirror_base.rstrip('/')}/{path.name}"
"--retry", command = ["curl", "-L", "--fail", "--retry", "3", "--retry-delay", "2",
"3", "--continue-at", "-"]
"--retry-delay", if token:
"2", command += ["--header", f"Authorization: token {token}"]
"--continue-at", command += ["-o", str(tmp), url]
"-", result = subprocess.run(command, text=True, capture_output=True)
"-o", if result.returncode != 0:
str(tmp), errors.append(f"{mirror_base}: {result.stderr.strip() or result.stdout.strip()}")
url, tmp.unlink(missing_ok=True)
] continue
result = subprocess.run(command, text=True, capture_output=True) tmp.replace(path)
if result.returncode != 0: if is_bad_fixture(path):
return path, False, result.stderr.strip() or result.stdout.strip() errors.append(f"{mirror_base}: downloaded fixture is empty or still an LFS pointer")
tmp.replace(path) continue
if is_bad_fixture(path): return path, True, ""
return path, False, "downloaded fixture is empty or still an LFS pointer" return path, False, "; ".join(errors)
return path, True, ""
def hydrate_fixtures(cases, fetch, mirror_base, download_jobs): def hydrate_fixtures(cases, fetch, mirror_bases, download_jobs):
required = [] required = []
seen = set() seen = set()
for case in cases: for case in cases:
@@ -149,7 +155,7 @@ def hydrate_fixtures(cases, fetch, mirror_base, download_jobs):
print(f"fixtures: downloading {len(required)} file(s) with {download_jobs} worker(s)", flush=True) print(f"fixtures: downloading {len(required)} file(s) with {download_jobs} worker(s)", flush=True)
failures = [] failures = []
with concurrent.futures.ThreadPoolExecutor(max_workers=download_jobs) as executor: with concurrent.futures.ThreadPoolExecutor(max_workers=download_jobs) as executor:
futures = [executor.submit(download_fixture, path, mirror_base) for path in required] futures = [executor.submit(download_fixture, path, mirror_bases) for path in required]
for future in concurrent.futures.as_completed(futures): for future in concurrent.futures.as_completed(futures):
path, ok, message = future.result() path, ok, message = future.result()
if ok: if ok:
@@ -434,7 +440,9 @@ def parse_args():
parser.add_argument("--keep-results", action="store_true", help="Do not clear previous result/work roots.") parser.add_argument("--keep-results", action="store_true", help="Do not clear previous result/work roots.")
parser.add_argument("--no-render", action="store_true", help="Do not render the HTML summary.") parser.add_argument("--no-render", action="store_true", help="Do not render the HTML summary.")
parser.add_argument("--fetch-fixtures", action=argparse.BooleanOptionalAction, default=True, help="Hydrate missing fixtures before running.") parser.add_argument("--fetch-fixtures", action=argparse.BooleanOptionalAction, default=True, help="Hydrate missing fixtures before running.")
parser.add_argument("--fixture-mirror-base", default=os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASE", DEFAULT_MIRROR_BASE)) parser.add_argument("--fixture-mirror-base", action="append", dest="fixture_mirror_bases",
help="Fixture mirror base URL; repeatable, tried in order. "
"Defaults to the hit.moe mirror then the miawa mirror.")
parser.add_argument("--download-jobs", type=int, default=4, help="Parallel fixture download count.") parser.add_argument("--download-jobs", type=int, default=4, help="Parallel fixture download count.")
parser.add_argument("--continue-after-fatal", action="store_true", help="Continue launching later cases after a fatal native-window replay.") parser.add_argument("--continue-after-fatal", action="store_true", help="Continue launching later cases after a fatal native-window replay.")
return parser.parse_args() return parser.parse_args()
@@ -463,7 +471,10 @@ def main():
print(f"missing MobileGL library: {mobilegl_library}", file=sys.stderr) print(f"missing MobileGL library: {mobilegl_library}", file=sys.stderr)
return 2 return 2
if not hydrate_fixtures(selected_cases, args.fetch_fixtures, args.fixture_mirror_base, max(1, args.download_jobs)): mirror_bases = args.fixture_mirror_bases or [
b for b in os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASES", "").split() or []
] or ([os.environ["MOBILEGL_TRACE_FIXTURE_MIRROR_BASE"]] if os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASE") else DEFAULT_MIRROR_BASES)
if not hydrate_fixtures(selected_cases, args.fetch_fixtures, mirror_bases, max(1, args.download_jobs)):
return 2 return 2
if not args.keep_results: if not args.keep_results:
+20
View File
@@ -0,0 +1,20 @@
# MobileGL trace-replay skills
Task-focused skills for capturing, replaying, debugging, and authoring MobileGL
apitrace fixtures. Each skill is a self-contained package:
- `SKILL.md` — the skill (frontmatter `name` + `description`, then the body). The
directory name equals the frontmatter `name`.
- `agents/openai.yaml` — OpenAI agent descriptor (`display_name`,
`short_description`, `default_prompt`).
- `scripts/` and/or `references/` — bundled tooling and supporting docs, when the
skill has them.
## Skills
| Skill | What it does |
| --- | --- |
| [trace-fixture-authoring-on-android-fcl](trace-fixture-authoring-on-android-fcl/SKILL.md) | Capture an on-device Android apitrace from FCL's MobileGL renderers (DirectGLES / Magma / SimpleFPEWrapper), mark the defect frame, and pull `full.trace`. |
| [renderdoc-debug-on-trace-replay](renderdoc-debug-on-trace-replay/SKILL.md) | Capture and validate an exact frame from a MobileGL retrace on a connected Android device with RenderDoc / rdc-cli. |
| [mismatch-retrace-debugging](mismatch-retrace-debugging/SKILL.md) | Localize the first divergent render pass and draw call when a fixture replays correctly in a golden environment but renders differently on a target backend. |
| [trace-fixture-authoring](trace-fixture-authoring/SKILL.md) | Author a deterministic trace-replay fixture — trim, golden, package under the size budget, register in `trace_cases.json`, and validate on Linux and Android. |
@@ -1,3 +1,8 @@
---
name: mismatch-retrace-debugging
description: Localize the first divergent render pass and draw call when a MobileGL apitrace fixture replays correctly in a golden environment but renders differently under mobilegl_trace_replay, Android trace replay, or another backend. Use to binary-search pass/draw endpoints, diff GL state around the first bad call, and classify the fault as a vertex/VS, fragment/FS, or framebuffer/composition mismatch.
---
# Mismatch retrace debugging # Mismatch retrace debugging
Use this when an apitrace fixture replays correctly on one environment but Use this when an apitrace fixture replays correctly on one environment but
@@ -0,0 +1,4 @@
interface:
display_name: "MobileGL Mismatch Retrace Debugging"
short_description: "Localize the first divergent draw in a mismatching MobileGL retrace"
default_prompt: "Use $mismatch-retrace-debugging to find the first divergent render pass and draw call in a MobileGL retrace mismatch."
@@ -1,5 +1,5 @@
--- ---
name: renderdoc-capture-trace-frame name: renderdoc-debug-on-trace-replay
description: Capture and validate an exact frame from a MobileGL apitrace retrace on a connected Android device with RenderDoc/rdc-cli. Use for DirectVulkan or DirectGLES trace replay, mapping a target API call to an eglSwapBuffers frame, producing an .rdc plus a complete command manifest, checking capture stability, or troubleshooting Android TargetControl timing and replay failures. description: Capture and validate an exact frame from a MobileGL apitrace retrace on a connected Android device with RenderDoc/rdc-cli. Use for DirectVulkan or DirectGLES trace replay, mapping a target API call to an eglSwapBuffers frame, producing an .rdc plus a complete command manifest, checking capture stability, or troubleshooting Android TargetControl timing and replay failures.
--- ---
@@ -19,20 +19,20 @@ adb -s SERIAL shell pm path top.mobilegl.plugin.trace
rdc doctor rdc doctor
``` ```
4. Pass the unpacked `trace.trace`, its golden PNG, the fixture target call, backend, and output path to `tools/trace_replay/capture_android_retrace.py`. 4. Pass the unpacked `trace.trace`, its golden PNG, the fixture target call, backend, and output path to `tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py`.
## Capture ## Capture
Let the tool infer the zero-based target swap from `eglSwapBuffers` calls: Let the tool infer the zero-based target swap from `eglSwapBuffers` calls:
```powershell ```powershell
python tools/trace_replay/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectVulkan --output captures/case-vulkan.rdc --serial SERIAL --json python tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectVulkan --output captures/case-vulkan.rdc --serial SERIAL --json
``` ```
Change only the backend and output for GLES: Change only the backend and output for GLES:
```powershell ```powershell
python tools/trace_replay/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectGLES --output captures/case-gles.rdc --serial SERIAL --json python tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectGLES --output captures/case-gles.rdc --serial SERIAL --json
``` ```
Use `--target-swap N` when the mapping is already known. Use `--capture-frame N` only to override the backend rule deliberately. Use `--target-swap N` when the mapping is already known. Use `--capture-frame N` only to override the backend rule deliberately.
@@ -0,0 +1,4 @@
interface:
display_name: "RenderDoc Debug on Trace Replay"
short_description: "Capture and debug an exact MobileGL trace-replay frame in RenderDoc"
default_prompt: "Use $renderdoc-debug-on-trace-replay to capture and validate an exact Android trace-replay frame in RenderDoc."
@@ -2,7 +2,7 @@
## TargetControl timing ## TargetControl timing
- Start `tools/trace_replay/queue_android_frame.py` before launching `TraceReplayActivity`. A fast retrace can pass the requested frame before a late client connects. - Start `tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/queue_android_frame.py` before launching `TraceReplayActivity`. A fast retrace can pass the requested frame before a late client connects.
- Keep TargetControl connected until `NewCapture` arrives. A queued request alone is not sufficient evidence that the RDC finished. - Keep TargetControl connected until `NewCapture` arrives. A queued request alone is not sufficient evidence that the RDC finished.
- Drain the asynchronous `RegisterAPI` and `CapturableWindowCount` messages before calling `QueueCapture`; otherwise `NewCapture` can be lost. - Drain the asynchronous `RegisterAPI` and `CapturableWindowCount` messages before calling `QueueCapture`; otherwise `NewCapture` can be lost.
- Do not use the daemon-backed `rdc script` path for a capture that may exceed 30 seconds. Its outer RPC times out even when the device later writes a valid RDC. The repository helper imports the RenderDoc module discovered by `rdc` directly and has an independent capture timeout. - Do not use the daemon-backed `rdc script` path for a capture that may exceed 30 seconds. Its outer RPC times out even when the device later writes a valid RDC. The repository helper imports the RenderDoc module discovered by `rdc` directly and has an independent capture timeout.
@@ -0,0 +1,137 @@
---
name: trace-fixture-authoring-on-android-fcl
description: Capture an on-device Android apitrace from FCL's MobileGL renderers. Use when preparing a reproducible MobileGL DirectGLES, Magma (DirectVulkan), or SimpleFPEWrapper rendering trace, marking the frame of a visual defect, pulling the resulting full.trace, or turning a device capture into a replay fixture.
---
# MobileGL Android trace capture
## Overview
Use FCL's Android `egltrace.so` wrapper, not Perfetto. When enabled before
launch, it records the complete EGL/GL call stream to `full.trace`. The game's
**MobileGL Trace → Capture** control marks the next swap frame in
`capture-result.json`; it does not start or stop recording and does not produce
a one-frame trace by itself.
Run commands from the FoldCraftLauncher repository root:
```sh
export REPO="$PWD"
export CAPTURE="$REPO/MobileGL/tools/trace_replay/skills/trace-fixture-authoring-on-android-fcl/scripts"
export SERIAL=<adb-device-serial> # omit --serial only if exactly one device is attached
```
The capture scripts are bundled inside this skill under `scripts/`; they
auto-detect the FoldCraftLauncher repository root from their own location, so
`--repo` only needs to be passed for a non-standard checkout layout.
## Prerequisites
- Use an FCL build containing `MobileGLTraceCapture` and the in-game Capture
menu entry.
- Select one of these renderers: MobileGL (DirectGLES), MobileGL Magma
(DirectVulkan), or SimpleFPEWrapper. MobileGlues is not supported by this
capture wrapper.
- Install `adb` and make it available on `PATH`; authorize USB debugging.
- Build the wrapper with Android NDK, CMake, Ninja, Python 3, and the checked
out in-tree `MobileGL/3rdparty/apitrace` submodule.
Confirm the attached device and ABI before building. The wrapper ABI must match
the device process ABI.
```sh
adb devices -l
adb -s "$SERIAL" shell getprop ro.product.cpu.abi
```
Use `arm64-v8a` for the usual `arm64-v8a` result; use the matching NDK ABI for
other devices.
## Build and install the wrapper
Build once per ABI or after changing apitrace/wrapper sources:
```sh
python3 "$CAPTURE/build_android_egltrace.py" --abi arm64-v8a
```
This generates `egltrace.so` under the skill's `scripts/out/` directory. Push
it and write FCL's enable sentinel before launching the game:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" install-wrapper
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" enable
```
The device-side control directory is `/sdcard/FCL/mobilegl-trace`. FCL copies
the shared `egltrace.so` into its private files directory at launch, replaces
the renderer's EGL library with it, and forwards to the real MobileGL library.
## Capture a reproduction
1. Start FCL after the wrapper and enable sentinel are in place. Select the
intended supported MobileGL renderer and launch the game.
2. Trace mode forces the game to `854x480`; account for that when reproducing
and comparing output.
3. Reproduce the issue. Start close to the target scene because tracing starts
when the game launches and trace files can grow rapidly.
4. At the desired visual state, open FCL's right-side game menu and press
**MobileGL Trace → Capture**. Let at least one frame present afterward.
5. Exit the game cleanly, then pull the latest session:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" pull-latest
```
The default local result is:
```text
.trace-work/pulled-mobilegl-captures/capture-YYYYMMDD-HHMMSS-<renderer>/
full.trace
capture-status.json
capture-result.json
```
`capture-result.json` must show `"status": "captured"`. Its `targetFrame`
is the one-based swap count used by `gltrim`; `zeroBasedFrame` is included for
tools that use zero-based indexing.
## Diagnose setup failures
Inspect the active device session directly:
```sh
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/latest-session.txt
adb -s "$SERIAL" shell ls -lh /sdcard/FCL/mobilegl-trace
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/capture-*/capture-status.json
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/capture-*/capture-result.json
```
If `capture-status.json` reports a missing `egltrace.so`, rebuild/push the
correct ABI and relaunch. If `capture-result.json` is absent, Capture was
pressed without an active trace session, or no subsequent `eglSwapBuffers`
occurred. The menu button itself only writes `capture-once.request`.
Disable tracing when finished; otherwise the next supported MobileGL launch
will trace again:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" disable
```
## Create a replay fixture (optional)
Keep the raw `full.trace` until replay validation succeeds. To frame-trim and
package the marked frame for MobileGL trace replay, use the existing helper:
```sh
python3 "$CAPTURE/package_capture_fixture.py" \
--serial "$SERIAL" \
--case <case-name> \
--apitrace <path-to-in-tree-apitrace>
```
It pulls the latest capture if necessary, uses `capture-result.json` to select
the frame, runs `apitrace gltrim`, creates a golden image, and enforces the
fixture archive-size limit. Follow `../trace-fixture-authoring/SKILL.md` for
deterministic scene setup, verification, and registry changes.
@@ -0,0 +1,4 @@
interface:
display_name: "Trace Fixture Authoring on Android (FCL)"
short_description: "Capture and package a MobileGL trace fixture on Android FCL"
default_prompt: "Use $trace-fixture-authoring-on-android-fcl to capture a MobileGL trace on my Android device and package it into a replay fixture."
@@ -0,0 +1,3 @@
# Build output produced by build_android_egltrace.py (ABI-specific, regenerated).
out/
__pycache__/

Some files were not shown because too many files have changed in this diff Show More