Compare commits

..
40 Commits
Author SHA1 Message Date
swung0x48 199164c2e0 [Perf] (DirectGLES): route the per-upload GL_PIXEL_UNPACK_BUFFER unbinds through the unpack binding cache - the six texture-upload sites re-issued glBindBuffer(UNPACK, 0) on every upload; with the resting-0 shadow they now no-op after the first 2026-07-21 19:54:55 -04:00
swung0x48 e87063e90c [Fix] (DirectGLES): route default-framebuffer binds through the FBO-binding shadow - the raw glBindFramebuffer(0) in BindCurrentFBO/SyncAndBindFramebufferObject left the shadow claiming the previous user FBO, false-skipping its next re-bind and letting scoped guards restore a stale binding (caught by adversarial review); also scrub buffer-binding shadows when VAO client-attribute staging buffers are deleted, include cube-map arrays in the pack-image-params gate, and make the delete-recording test hook assertion-unwind safe 2026-07-21 19:54:54 -04:00
swung0x48 122da27249 [Fix] (DirectGLES): overhaul readback/copy/blit driver-state handling with shadow-backed RAII guards - pixel-PACK/UNPACK PBO binding caches resting at 0 (the old bind-then-query 'restore' left the user PBO bound forever, capturing later client-memory readbacks), a PACK pixel-store shadow replacing per-readback glGetIntegerv syncs, a driver FBO-binding shadow behind all scoped binders (per-instance prev slots, nest-safe), scratch-FBO attachment shadows that detach cross-aspect residue exactly when present (depth CopyTex* attachments used to wedge later GetTexImage color reads and vice versa), scissor guards around emulation blits (app scissor clipped depth copies), stale-driver-error drains before single-shot glGetError consumers (fallbacks silently dropped readbacks / GenerateMipmap raised phantom app errors in production builds), ClearBufferiv missing BindCurrentFBO(Draw), cube-face glBindTexture INVALID_ENUM cache poisoning, per-slice GL_PACK_IMAGE_HEIGHT/SKIP_IMAGES semantics on 3D GetTexImage, backend texture ids deleted on wrapper destruction with cache/scratch-FBO scrubs (ids used to leak for the context lifetime and dangling cache pointers could false-skip binds), and context-death/MakeCurrent invalidation for all new shadows; regression tests drive the shadows against a recording mock GLES table 2026-07-21 19:54:54 -04:00
swung0x48 bc2d698b3e [Fix] (DirectGLES): apply the read buffer when one FBO is bound as both draw and read
- SyncCurrentFBO skips the READ-target pass when the same GL FBO is bound as
  both draw and read (the common GL_FRAMEBUFFER case), but the read buffer
  (glReadBuffer) is only applied inside SyncToBackend's READ path — so the skip
  silently dropped every glReadBuffer change, leaving the backend read buffer
  stuck at COLOR_ATTACHMENT0.
- Extract the read-buffer application into BackendFramebufferObject::
  SyncReadBufferToBackend and invoke it from the skip branch (target == Read)
  as well as from SyncToBackend, so reads always target the right attachment.
- Bind the backend FBO as READ inside the helper before glReadBuffer, since the
  skip path only bound it as DRAW.
- Fixes KHR-GL3x.draw_buffers.draw_buffers_1 (reading COLOR_ATTACHMENT1 while
  the FBO stays GL_FRAMEBUFFER-bound returned attachment 0's value); the render
  was already correct, only the readback resolved the wrong attachment.
2026-07-21 12:18:55 -04:00
swung0x48 3049c4b82b [Fix] (ShaderTranspiler): stop blanking block comments in the source handed to glslang
- BlankBlockComments replaced comment chars with spaces but preserved interior newlines, so a block comment spanning a newline inside a #define truncated the macro body (VALUE became empty)
- glslang has a conformant preprocessor and collapses a block comment to one space across newlines, so the delivered source now keeps comments intact and lets glslang handle them
- Fixes KHR-GL3x.shaders.preprocessor multiline_comment_define / redefine_object_multiline_comment / function_redefinition_3 (6 cases, both devices)
- FilterUnsupportedGpuShaderInt64 relied on the blanking to skip commented-out #extension lines; it now masks comments locally (MaskCommentsAndQuotedText) like the sibling passes, collecting edits and applying them back-to-front
- conditional_inclusion.basic_2 (defined() via macro expansion) stays failing by design: glslang rejects it as UB and working around it would mean re-running preprocessing MobileGL defers to glslang
2026-07-21 05:27:01 -04:00
swung0x48 2b3850b76b [Fix] (DirectGLES): upload RGB565/RGB5_A1 shadow data as packed 16-bit types
- The 8-bit unorm shadow was uploaded as GL_UNSIGNED_BYTE, leaving the 8->5/6-bit requantization to the driver
- That rounding direction is implementation-defined: Adreno rounds to nearest (lossless round trip), Mali floors
- On Mali mid-range texels drifted one 5-bit step down, failing all 20 KHR-GL33.pixelstoragemodes.teximage3d rgb565/rgb5a1 cases (eps 1/32); layers with exact values (0.125/0.25/1.0-ish) passed, matching the observed 0,1,7-valid pattern
- PreparePackedNormUpload repacks shadow rows to GL_UNSIGNED_SHORT_5_6_5 / 5_5_5_1 with round-to-nearest, which exactly recovers the original 5/6-bit values (the shadow expansion is injective), so the driver stores them verbatim
- Idempotent across each region's level loop (glType is shared); RGBA4 exempt since its 8-bit expansion (v*17) is exact under either rounding
- Wired at all four SyncMipmapsToBackend upload regions (append-mipmaps, immutable TexSubImage, mutable full, dirty-level update)
2026-07-21 04:30:46 -04:00
swung0x48 2a0ae743a0 [Fix] (DirectGLES): glFinish before the glGetTexImage temp-FBO readback
- Mali (tile-based) does not resolve a texture's render into memory when it is read back through a different (temp) FBO than the one it was rendered with
- The cross-FBO glReadPixels raced the deferred tile resolve and returned pre-render clear contents
- Distinct render targets read back byte-identical, so KHR-GLxx.glsl_noperspective failed on Mali-G715 (all four programs read as the clear colour)
- glGetTexImage is already a CPU/GPU sync point so the extra drain is negligible; Adreno resolves eagerly and was unaffected
2026-07-21 03:34:33 -04:00
swung0x48 6839219c10 [Fix] (GLImpl): report GL_NO_ERROR from glGetGraphicsResetStatus
- The generic export stub returned (GLenum)1; dEQP reads any non-zero status as a lost device
- It is polled after every case (gl3cTestPackages.cpp:121) and sets QP_TEST_RESULT_DEVICE_LOST
- Under the default --deqp-terminate-on-device-lost=enable that tears the whole CTS run down
- MobileGL tracks no GPU resets, so GL_NO_ERROR ("no reset detected") is the honest, spec-correct answer
- Routed through GLImpl::GetGraphicsResetStatus like every other entry point, no inline body in Definitions.cpp
2026-07-21 02:40:15 -04:00
swung0x48 52ddb440ca [Feat] (SelfTest/DriverPost): add a noperspective correctness check to the GLES POST - render a strong-perspective quad and read the centre texel to verify the varying interpolates screen-linear (not perspective-correct), carried through the native GL_NV_shader_noperspective_interpolation path when present or MobileGL's exact gl_Position.w/gl_FragCoord.w emulation when absent. PASS = native and correct; WARN = emulated and correct (the fallback path shipping packs hit on such devices); FAIL = interpolation wrong/perspective-correct, or the program will not build. The ESSL header matches the device version because noperspective is rejected at #version 300 es on some drivers even with the extension enabled 2026-07-21 02:09:17 -04:00
swung0x48 79feeffd25 [Feat] (ShaderTranspiler, DirectGLES): emulate noperspective on GLES devices lacking GL_NV_shader_noperspective_interpolation - EmulateNoPerspectivePass pre-multiplies each NoPerspective output by gl_Position.w in the vertex stage and recovers each input via gl_FragCoord.w in the fragment stage (exact screen-linear L = P(a*w)*gl_FragCoord.w, handling whole-variable and component/access-chain reads, scalar and vector varyings), forces highp on emulated varyings, and strips what it cannot emulate; replaces the smooth-strip fallback so no NV extension is ever required. Restricts the vertex pre-multiply to the entry function so a non-inlined helper cannot double-scale 2026-07-20 23:48:48 -04:00
swung0x48 202037b5a3 [Fix] (DirectVulkan): shrink the blended depth-write quirk to MIN/MAX extremum blends only - a fixture-wide trace sweep showed the additive ONE+ONE arm never fires on the 26.3 OIT chain (its accumulation passes disable depth writes themselves) and only hit unrelated additive glow content; also fail reflection toward the gl_FragDepth exemption and zero phantom default-FBO blend slots so stale indexed state cannot trigger the strip 2026-07-20 23:39:28 -04:00
swung0x48 bce9c48c8e [Feat] (ShaderTranspiler, DirectGLES): support noperspective conformantly instead of stripping it - let the qualifier reach glslang as the core SPIR-V NoPerspective decoration (native on DirectVulkan; SPIRV-Cross emits ESSL noperspective + GL_NV_shader_noperspective_interpolation on DirectGLES), and for GLES devices lacking that extension add StripNoPerspectivePass to drop the decoration and fall back to smooth; the old naked substring erase discarded the interpolation shader packs need and mangled identifiers containing the word 2026-07-20 22:59:43 -04:00
swung0x48 b6a7807a3a [Fix] (MG_Util/ShaderTranspiler): reject malformed #version directives instead of legalizing them - an unrecognized version number (329/331), a bad profile keyword, a float or trailing token used to be rewritten to "#version 330 core" (or rescued to 460 by the retry); now InspectShaderLanguage marks such directives invalid so NormalizeVersionDirective and RetargetLegacyVersionDirectiveTo460 leave them for glslang to reject, while every valid version still normalizes as before 2026-07-20 22:03:36 -04:00
swung0x48 48ba622387 [Fix] (MG_Util/ShaderTranspiler): keep #line directives instead of deleting them, dropping only the GLSL-illegal quoted filename and the ones that precede #version, so __LINE__ and compiler diagnostics follow the application's own numbering 2026-07-20 21:06:38 -04:00
swung0x48 05260d1262 [Fix] (MG_Util/ShaderTranspiler): blank block comments lexically instead of erasing them - a '//*** banner ***' line opened a comment the old scanner never closed, so it deleted the rest of the shader, and a commented-out builtin definition renamed every genuine call to a name nothing defines 2026-07-20 21:06:37 -04:00
swung0x48 6eb5ff51c5 [Fix] (MG_Impl/GLImpl): reject the RGTC internal formats on 3D texture targets - RGTC compresses 4x4 blocks of a 2D image and has no 3D form, and the check must run on the raw enum because RGTC now resolves to plain R8/RG8 storage 2026-07-20 21:05:57 -04:00
swung0x48 e526f8e8ac [Fix] (MG_Util/Converters): resolve the GL_COMPRESSED_* internal formats to the uncompressed storage that backs them instead of rejecting them as unknown - GL prescribes this base-format fallback for the six generic formats, and RGTC stores uncompressed because ES exposes no compressor 2026-07-20 21:05:57 -04:00
swung0x48 e724e88eec [Fix] (MG_State, MG_Impl/GLImpl): allocating a mipmap level no longer truncates the chain above it - AllocateLevel now only grows and the callers that genuinely redefine the whole level set (glTexStorage*, mip regeneration, multisample storage, level-0 respecification) drop the tail explicitly 2026-07-20 21:02:54 -04:00
swung0x48 3b175fb88a [Fix] (MG_Impl/GLImpl): a multisample sample count above the format's maximum is INVALID_OPERATION, not INVALID_VALUE - matching both the spec and the native Adreno driver 2026-07-20 21:02:53 -04:00
swung0x48 b5a4e7075a [Fix] (MG_Impl/GLImpl): record a GL error from the unimplemented compressed texture entry points instead of throwing - a C++ exception unwinding through the C GL ABI hard-crashes any caller, and glGetCompressedTexImage reported success while writing nothing 2026-07-20 21:02:53 -04:00
swung0x48 520c2b6750 [Fix] (MG_Backend/DirectGLES): sync GL_TEXTURE_SWIZZLE_* on multisample targets - the early return meant to skip the sampler-only parameters dropped every swizzle write, which the frontend already treats as legal on those targets 2026-07-20 21:02:52 -04:00
swung0x48 57cc652b1d [Fix] (MG_Impl/GLImpl): glIsTransformFeedback reports GL_FALSE instead of claiming every name it is handed is a live transform feedback object 2026-07-20 21:02:52 -04:00
swung0x48 65ea54da9e [Refactor] (ShaderTranspiler, DirectVulkan): replace the hand-rolled SPIR-V word walkers with a DecoratePositionInvariantPass and SPIRV-Reflect-based InstanceIndex detection 2026-07-20 08:35:00 -04:00
swung0x48 c81dd04f08 [Test] (CI, trace_replay): force the blended depth-write quirk on the Linux DirectVulkan OIT retrace lane and tighten that case's SSIM threshold to 0.995 2026-07-20 07:39:25 -04:00
swung0x48 c158bfa584 [Fix] (DirectVulkan): narrow the blended depth-write quirk to order-independent accumulation blends, exempting sorted-transparency, gl_FragDepth writers and fully masked attachments 2026-07-20 07:39:25 -04:00
swung0x48 f9f455144c [Refactor] (MG_Config, DirectVulkan): route the blended depth-write quirk through the FeaturesTable as MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE instead of an ad-hoc getenv 2026-07-20 04:37:23 -04:00
swung0x48 64e4840de2 [Fix] (DirectVulkan): fall back to immutable images when the mutable probe fails, gate robustBufferAccess behind MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS, decode R32/RG32/R16-class readback formats, and warn when shaderStorageImage*WithoutFormat is unavailable 2026-07-20 02:36:26 -04:00
swung0x48 bf7b5755cc [Refactor] (ShaderTranspiler, MG_Backend, MG_Util): gate the subgroup prefix-scan rewrite behind a generic device-quirk registry with GPU vendor detection and MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN override, warn on template mismatch, and block ARB/NV subgroup spellings 2026-07-20 02:36:26 -04:00
swung0x48 293f64b3c2 [Fix] (MG_Impl, DirectGLES): correct ARB_clear_texture error codes, reject cube maps in CopyTextureSubImage2D, advertise the extension on Espryt, and pin the error contracts with tests 2026-07-20 02:36:25 -04:00
swung0x48 f0cc07c937 [Perf] (DirectVulkan): bound vertex-stream conversion by the draw's real fetch range, reuse cached prefixes, pin cached source buffers, and stop repacking client arrays for pointer alignment 2026-07-20 02:36:25 -04:00
swung0x48 4658536652 [Perf] (DirectVulkan): keep the render pass alive for steady-state storage-image draws and skip storage-image collection for programs without them 2026-07-20 02:36:24 -04:00
swung0x48 04b4627c65 [Fix] (DirectVulkan): pass depth/stencil formats through sampled-view resolution so D24S8/D32FS8 samplers stop dropping draws, and memoize per-binding view-format resolution 2026-07-20 02:36:24 -04:00
swung0x48 68e13705c8 [Fix] (MG_Backend/DirectVulkan, ShaderTranspiler, MG_Test, TraceReplay): make iterationRP retrace pass on 64-lane Vulkan devices with subgroup-width emulation and format-aware readback 2026-07-20 02:36:23 -04:00
swung0x48 e5388c0e7e [Fix] (MG_Backend, MG_Impl, ShaderTranspiler, MG_Test): support iterationRP custom images and storage format reinterpretation 2026-07-20 02:35:30 -04:00
swung0x48 8bc4808b1a [Test] (TraceReplay): match the iterationRP fixture name published on the mirror 2026-07-19 23:09:29 -04:00
swung0x48 5d8a5387e2 [Test] (TraceReplay): try the hit.moe mirror before miawa and Git LFS 2026-07-19 22:29:55 -04:00
swung0x48 aa2184e47a [Test] (TraceReplay): register iterationRP in-world fixture (non-CI, fixture files pending LFS) 2026-07-19 21:39:06 -04:00
swung0x48 9152a4a4bc [Test] (TraceReplay): enable improved-transparency fixture in CI and drop unneeded coherent_as_flush 2026-07-19 07:35:34 -04:00
swung0x48 1963b427db [Refactor] (TraceReplay): reorganize trace skills into uniform packages with bundled scripts 2026-07-19 07:00:56 -04:00
swung0x48 f3def150e7 [Fix] (DirectVulkan): suppress blended depth writes on Qualcomm and mark gl_Position invariant to fix MC 26.3 OIT cloud flicker 2026-07-19 05:47:23 -04:00
78 changed files with 6196 additions and 651 deletions
+38 -7
View File
@@ -9,7 +9,23 @@ fi
case_name="$1"
fixture_dir="${2:-tools/trace_replay/fixtures}"
python_bin="${PYTHON:-python3}"
mirror_base="${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE:-https://repo.miawa.cn/mgl/tools/trace_replay/fixtures}"
# Fixture mirrors, tried in order before falling back to Git LFS. Override the
# whole list with MOBILEGL_TRACE_FIXTURE_MIRROR_BASES (whitespace separated);
# MOBILEGL_TRACE_FIXTURE_MIRROR_BASE still works and is tried first.
default_mirror_bases=(
"https://git.hit.moe/swung0x48/MobileGL/media/branch/dev/tools/trace_replay/fixtures"
"https://repo.miawa.cn/mgl/tools/trace_replay/fixtures"
)
if [ -n "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASES:-}" ]; then
read -r -a mirror_bases <<< "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASES}"
else
mirror_bases=("${default_mirror_bases[@]}")
fi
if [ -n "${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE:-}" ]; then
mirror_bases=("${MOBILEGL_TRACE_FIXTURE_MIRROR_BASE}" "${mirror_bases[@]}")
fi
# Optional bearer token for mirrors that require authentication (private Gitea).
mirror_token="${MOBILEGL_TRACE_FIXTURE_MIRROR_TOKEN:-}"
download_attempts="${MOBILEGL_TRACE_FIXTURE_DOWNLOAD_ATTEMPTS:-5}"
retry_delay="${MOBILEGL_TRACE_FIXTURE_RETRY_DELAY:-2}"
@@ -30,7 +46,8 @@ fixture_list="$("${python_bin}" tools/trace_replay/trace_cases.py \
--format fixture-files \
--case "${case_name}" \
--fixture-root "${fixture_dir}")"
mapfile -t files <<< "${fixture_list}"
# Strip CR so the script also works when python emits CRLF (Git Bash on Windows).
mapfile -t files < <(printf '%s\n' "${fixture_list}" | tr -d '\r')
include="$(IFS=,; echo "${files[*]}")"
if [ "${case_name}" = "OpenRA" ]; then
@@ -106,6 +123,7 @@ fetch_file_from_mirror() {
local attempt
local partial_size
local curl_status
local curl_auth
metadata="$(get_lfs_metadata "${file}")" || return 1
read -r expected_oid expected_size <<< "${metadata}"
@@ -136,7 +154,11 @@ fetch_file_from_mirror() {
echo "Starting mirror download for ${file} (attempt ${attempt}/${download_attempts})"
fi
if curl -L --fail --show-error --continue-at - --output "${tmp_file}" "${url}"; then
curl_auth=()
if [ -n "${mirror_token}" ]; then
curl_auth=(--header "Authorization: token ${mirror_token}")
fi
if curl -L --fail --show-error --continue-at - "${curl_auth[@]}" --output "${tmp_file}" "${url}"; then
if verify_fixture_file "${tmp_file}" "${file}" "${expected_oid}" "${expected_size}"; then
mv "${tmp_file}" "${file}"
return 0
@@ -184,10 +206,19 @@ fetch_from_mirror() {
for file in "${files[@]}"; do
local name
local url
local base
local fetched=0
name="$(basename "${file}")"
url="${mirror_base%/}/${name}"
echo "Fetching trace fixture from mirror: ${url}"
if ! fetch_file_from_mirror "${file}" "${url}"; then
for base in "${mirror_bases[@]}"; do
url="${base%/}/${name}"
echo "Fetching trace fixture from mirror: ${url}"
if fetch_file_from_mirror "${file}" "${url}"; then
fetched=1
break
fi
echo "Mirror did not serve ${name}; trying the next mirror" >&2
done
if [ "${fetched}" -ne 1 ]; then
return 1
fi
done
@@ -196,7 +227,7 @@ fetch_from_mirror() {
if fetch_from_mirror; then
echo "Fetched trace fixture files for ${case_name} from mirror: ${include}"
else
echo "Mirror fetch failed for ${case_name}; falling back to Git LFS: ${include}"
echo "All mirrors failed for ${case_name}; falling back to Git LFS: ${include}"
git lfs install --local
git lfs pull --include="${include}" --exclude=""
fi
+9
View File
@@ -455,6 +455,15 @@ jobs:
if [ '${{ matrix.backend }}' = 'DirectVulkan' ]; then
export MOBILEGL_MAGMA_R11G11B10F_FALLBACK=1
fi
# The blended depth-write quirk auto-enables only on Qualcomm, which no CI
# runner has, so force it on for the OIT case it exists to fix. ForceOn
# bypasses only the vendor gate, so this exercises the real strip on
# lavapipe. The Android AVD lane deliberately leaves it off, keeping the
# unstripped path covered for the same trace.
if [ '${{ matrix.backend }}' = 'DirectVulkan' ] \
&& [ '${{ matrix.case }}' = 'improved-transparency-minecraft-26.3' ]; then
export MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE=1
fi
ctest -V --no-tests=error -R '^MobileGLTraceReplay\.${{ matrix.case }}\.${{ matrix.backend }}$'
- name: Upload actual image
+3
View File
@@ -188,9 +188,12 @@ set(SOURCE_FILES
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EliminateFloatEqualsZeroPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RenameSamplerFunctionParameterPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecomposeWorkgroupVec3Pass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/LowerDrawParametersPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/RebaseInstanceIndexPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripUboMemberRelaxedPrecisionPass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.cpp
MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.cpp
MobileGL/MG_Util/BackendLoaders/OpenGL/Loader.cpp
MobileGL/MG_Util/BackendLoaders/Vulkan/Loader.cpp
+24
View File
@@ -20,6 +20,15 @@ namespace MobileGL::MG_Config {
extern BackendType ActiveBackendType;
// Tri-state override for device-specific quirks: Auto lets the detected device decide,
// ForceOn/ForceOff bypass the detection in either direction. ForceOn only bypasses the
// device gate - each quirk keeps its structural safety checks.
enum class QuirkOverride : Uint8 {
Auto = 0,
ForceOn,
ForceOff,
};
// Feature toggles parsed once from environment variables in MG_ConfigLoader::Init()
// (ConfigLoader.cpp), before the accepted-env map is destroyed. All Bool fields share
// one truthy rule: the variable is set, non-empty, not "0", and not "false"
@@ -67,6 +76,21 @@ namespace MobileGL::MG_Config {
// explicitly request a core profile via EGL_CONTEXT_OPENGL_PROFILE_MASK / a >=3.1
// version request.
Bool RelaxedSemantics = false;
// MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN: overrides the shader-source quirk that
// rewrites the recognized workgroup prefix-scan template on Qualcomm devices with
// subgroups wider than 32 lanes (see ShaderSourceProcessor's quirk registry).
QuirkOverride SubgroupPrefixScanQuirk = QuirkOverride::Auto;
// MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE: overrides the DirectVulkan quirk that
// strips depth writes from accumulation-blended pipelines (MIN/MAX or additive
// ONE+ONE - the multi-pass depth-equality signature) on drivers without
// cross-pipeline vertex position invariance. Sorted-transparency "over" blends,
// gl_FragDepth writers, and fully color-masked attachments are exempt (see
// PipelineFactory::ShouldSuppressDepthWrite). Auto detects Qualcomm.
QuirkOverride MagmaDisableBlendedDepthWriteQuirk = QuirkOverride::Auto;
// MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS: leave the Vulkan robustBufferAccess device
// feature off. It is enabled by default to match GL's defined out-of-range fetch
// behavior; this escape hatch exists to measure or dodge its GPU cost on a device.
Bool DisableRobustBufferAccess = false;
};
extern FeaturesTable Features;
} // namespace MobileGL::MG_Config
+15
View File
@@ -86,6 +86,17 @@ namespace MobileGL::MG_ConfigLoader {
return it != acceptedEnvVariablesMap->end() && IsTruthyValue(it->second);
}
// Quirk overrides are tri-state: an unset variable keeps device auto-detection, a truthy
// value forces the quirk on, anything else set ("0", "false", "") forces it off.
inline MG_Config::QuirkOverride QueryEnvQuirkOverride(const String& key) {
auto it = acceptedEnvVariablesMap->find(key);
if (it == acceptedEnvVariablesMap->end()) {
return MG_Config::QuirkOverride::Auto;
}
return IsTruthyValue(it->second) ? MG_Config::QuirkOverride::ForceOn
: MG_Config::QuirkOverride::ForceOff;
}
inline Uint32 QueryEnvUint32(const String& key, Uint32 defaultValue, Uint32 minValue, Uint32 maxValue) {
auto it = acceptedEnvVariablesMap->find(key);
if (it == acceptedEnvVariablesMap->end()) {
@@ -123,6 +134,10 @@ namespace MobileGL::MG_ConfigLoader {
features.TraceSkipAutodestroy = QueryEnvFlag("MOBILEGL_TRACE_SKIP_AUTODESTROY");
features.DisableUboRing = QueryEnvFlag("MOBILEGL_DISABLE_UBO_RING");
features.RelaxedSemantics = QueryEnvFlag("MOBILEGL_RELAXED_SEMANTICS");
features.SubgroupPrefixScanQuirk = QueryEnvQuirkOverride("MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN");
features.MagmaDisableBlendedDepthWriteQuirk =
QueryEnvQuirkOverride("MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE");
features.DisableRobustBufferAccess = QueryEnvFlag("MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS");
}
inline void InitBackendType() {
+16
View File
@@ -230,6 +230,21 @@ namespace MobileGL {
void (*SetSwapInterval)(Int interval);
};
// Coarse GPU vendor identity for gating device-specific quirks. Detected from the
// Vulkan physical-device vendorID or the GLES GL_VENDOR/GL_RENDERER strings; stays
// Unknown when detection is inconclusive, in which case auto-gated quirks stay off.
enum class GpuVendorKind : Uint8 {
Unknown = 0,
Qualcomm,
Arm,
Nvidia,
Amd,
Intel,
ImgTec,
// Software rasterizers (llvmpipe/lavapipe, SwiftShader).
Software,
};
struct DynamicBackendParameters {
SizeT UniformBufferOffsetAlignment = 256;
// GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT. 1.0 means the backend cannot filter anisotropically,
@@ -291,6 +306,7 @@ namespace MobileGL {
Uint32 SubgroupSupportedStages = 0;
Uint32 SubgroupSupportedFeatures = 0;
Bool SubgroupQuadOperationsInAllStages = false;
GpuVendorKind GpuVendor = GpuVendorKind::Unknown;
};
enum class WindowBackend {
@@ -826,7 +826,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
E_GL_ARB_program_interface_query, E_GL_ARB_framebuffer_object,
E_GL_EXT_framebuffer_object, E_GL_ARB_depth_texture, E_GL_ARB_buffer_storage,
E_GL_ARB_texture_storage, E_GL_ARB_texture_storage_multisample,
E_GL_ARB_direct_state_access,
E_GL_ARB_clear_texture, E_GL_ARB_direct_state_access,
E_GL_ARB_multi_draw_indirect, E_GL_ARB_indirect_parameters,
E_GL_ARB_shader_draw_parameters, E_GL_ARB_gpu_shader5, E_GL_ARB_multi_bind,
E_GL_ARB_shading_language_420pack, E_GL_ARB_vertex_attrib_binding,
@@ -1035,6 +1035,32 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_dynamicParameters.ViewportSubpixelBits = m_GLESCapabilities.ViewportSubpixelBits;
m_dynamicParameters.SupportsWideLines =
m_GLESCapabilities.AliasedLineWidthRangeMax > 1.0f || m_GLESCapabilities.SmoothLineWidthRangeMax > 1.0f;
const auto containsAny = [](const String& haystack, std::initializer_list<const char*> needles) {
return std::any_of(needles.begin(), needles.end(), [&](const char* needle) {
return haystack.find(needle) != String::npos;
});
};
const String vendorAndRenderer =
m_GLESCapabilities.GLESVendorString + " " + m_GLESCapabilities.GLESRendererString;
if (containsAny(vendorAndRenderer, {"llvmpipe", "SwiftShader", "softpipe"})) {
// Check software rasterizers first: ANGLE-on-llvmpipe reports both.
m_dynamicParameters.GpuVendor = GpuVendorKind::Software;
} else if (containsAny(vendorAndRenderer, {"Qualcomm", "Adreno"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Qualcomm;
} else if (containsAny(vendorAndRenderer, {"Mali", "ARM"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Arm;
} else if (containsAny(vendorAndRenderer, {"NVIDIA"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Nvidia;
} else if (containsAny(vendorAndRenderer, {"AMD", "Radeon"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Amd;
} else if (containsAny(vendorAndRenderer, {"Intel"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::Intel;
} else if (containsAny(vendorAndRenderer, {"Imagination", "PowerVR"})) {
m_dynamicParameters.GpuVendor = GpuVendorKind::ImgTec;
} else {
m_dynamicParameters.GpuVendor = GpuVendorKind::Unknown;
}
}
const MG_External::GLESFunctionsTable& BackendObject_DirectGLES::GetGLESFunctions() const {
File diff suppressed because it is too large Load Diff
+575 -43
View File
@@ -267,16 +267,26 @@ namespace MobileGL::MG_Backend::DirectGLES {
Uint g_boundArrayBufferId = 0;
Bool g_boundArrayBufferKnown = false;
// Driver-level GL_PIXEL_PACK/UNPACK_BUFFER binding shadows (see
// Managers.h). Resting state between operations is 0; scopes in the
// readback/upload paths bind what they need through the cache and
// return to 0, so a stale user PBO can never capture a later
// readback that meant to target client memory.
Uint g_boundPixelPackBufferId = 0;
Bool g_boundPixelPackBufferKnown = false;
Uint g_boundPixelUnpackBufferId = 0;
Bool g_boundPixelUnpackBufferKnown = false;
// Bumped whenever the backend ES context is destroyed; resources with
// an older generation hold ids from a dead context.
Uint g_bufferContextGeneration = 1;
// Defined next to the indexed-binding shadow below; forward-declared so
// every glDeleteBuffers site in this namespace can scrub stale shadow
// entries (GL resets a deleted buffer's indexed bindings to 0, and a
// recycled name matching a stale shadow entry would otherwise
// false-skip the rebind).
void ScrubIndexedBufferBindingShadowForId(Uint id);
// entries (GL resets a deleted buffer's bindings - indexed and pixel
// pack/unpack alike - to 0, and a recycled name matching a stale shadow
// entry would otherwise false-skip the rebind).
void ScrubBufferBindingShadowsForId(Uint id);
// Resources whose owning BufferObject died; ids deleted at the next
// sync point with a current ES context.
@@ -318,10 +328,16 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == r.id) {
InvalidateArrayBufferBindingCache();
}
// Pooling keeps the id alive (and thus any driver binding of it);
// drop to unknown rather than claiming the post-delete 0 state.
if ((g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == r.id) ||
(g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == r.id)) {
InvalidatePixelBufferBindingCaches();
}
const std::lock_guard<std::mutex> lock(g_poolMutex);
auto& bucket = g_bufferPool[r.storageSize];
if (bucket.size() >= kMaxEntriesPerBucket || g_pooledBytes + r.storageSize > kMaxPoolBytes) {
ScrubIndexedBufferBindingShadowForId(r.id);
ScrubBufferBindingShadowsForId(r.id);
g_GLESFuncs.glDeleteBuffers(1, &r.id); // over budget: don't pool
r.id = 0;
return;
@@ -502,7 +518,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
// Need a fresh id: glBufferStorage fails on a buffer that already has
// immutable storage, and any prior mutable store is replaced anyway.
if (resource->id != 0) {
ScrubIndexedBufferBindingShadowForId(resource->id);
ScrubBufferBindingShadowsForId(resource->id);
g_GLESFuncs.glDeleteBuffers(1, &resource->id);
resource->id = 0;
}
@@ -630,7 +646,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) {
InvalidateArrayBufferBindingCache();
}
ScrubIndexedBufferBindingShadowForId(glesResource->id);
ScrubBufferBindingShadowsForId(glesResource->id);
g_GLESFuncs.glDeleteBuffers(1, &glesResource->id);
glesResource->id = 0;
}
@@ -668,6 +684,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
void OnBackendContextDestroyed() {
UnregisterBufferBackendOps();
++g_bufferContextGeneration;
InvalidateArrayBufferBindingCache();
InvalidateIndexedBufferBindingCache();
InvalidatePixelBufferBindingCaches();
// The global-UBO ring's id and persistent map died with the context;
// drop the handles (no GL) and let the next draw recreate the ring.
ResetUboRingForNewContext();
@@ -694,7 +713,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (g_boundArrayBufferKnown && g_boundArrayBufferId == glesResource->id) {
InvalidateArrayBufferBindingCache();
}
ScrubIndexedBufferBindingShadowForId(glesResource->id);
ScrubBufferBindingShadowsForId(glesResource->id);
g_GLESFuncs.glDeleteBuffers(1, &glesResource->id);
glesResource->id = 0;
}
@@ -819,6 +838,41 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_boundArrayBufferKnown = false;
}
void BindPixelPackBufferId(Uint id) {
if (g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == id) {
return;
}
g_GLESFuncs.glBindBuffer(GL_PIXEL_PACK_BUFFER, id);
g_boundPixelPackBufferId = id;
g_boundPixelPackBufferKnown = true;
}
void BindPixelUnpackBufferId(Uint id) {
if (g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == id) {
return;
}
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, id);
g_boundPixelUnpackBufferId = id;
g_boundPixelUnpackBufferKnown = true;
}
void InvalidatePixelBufferBindingCaches() {
g_boundPixelPackBufferId = 0;
g_boundPixelPackBufferKnown = false;
g_boundPixelUnpackBufferId = 0;
g_boundPixelUnpackBufferKnown = false;
}
void NoteBufferIdDeleted(Uint id) {
if (id == 0) {
return;
}
if (g_boundArrayBufferKnown && g_boundArrayBufferId == id) {
InvalidateArrayBufferBindingCache();
}
ScrubBufferBindingShadowsForId(id);
}
namespace {
// Shadow of the GL indexed buffer bindings so redundant glBindBufferBase/Range
// (same index + id + range) are skipped. isBase distinguishes a whole-buffer
@@ -840,12 +894,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
return nullptr;
}
// glDeleteBuffers resets the deleted buffer's bindings (indexed ones
// included) to 0 in the current context; mirror that in the shadow, or a
// later buffer recycling the same name with a matching range would
// false-skip its rebind. Default IndexedBufferBinding{} == base(0) ==
// the post-delete GL state.
void ScrubIndexedBufferBindingShadowForId(Uint id) {
// glDeleteBuffers resets the deleted buffer's bindings (indexed and
// pixel pack/unpack ones included) to 0 in the current context; mirror
// that in the shadows, or a later buffer recycling the same name with a
// matching shadow entry would false-skip its rebind. Default
// IndexedBufferBinding{} == base(0) == the post-delete GL state.
void ScrubBufferBindingShadowsForId(Uint id) {
if (id == 0) return;
for (auto& binding : g_indexedUBOBindings) {
if (binding.id == id) binding = {};
@@ -853,6 +907,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
for (auto& binding : g_indexedSSBOBindings) {
if (binding.id == id) binding = {};
}
if (g_boundPixelPackBufferKnown && g_boundPixelPackBufferId == id) {
g_boundPixelPackBufferId = 0;
}
if (g_boundPixelUnpackBufferKnown && g_boundPixelUnpackBufferId == id) {
g_boundPixelUnpackBufferId = 0;
}
}
} // namespace
@@ -897,7 +957,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
auto& bucket = g_bufferPool[oldestKey];
PooledBuffer& e = bucket[oldestIdx];
if (e.contextGeneration == g_bufferContextGeneration && e.id != 0) {
ScrubIndexedBufferBindingShadowForId(e.id);
ScrubBufferBindingShadowsForId(e.id);
g_GLESFuncs.glDeleteBuffers(1, &e.id);
}
g_pooledBytes -= e.size;
@@ -1064,7 +1124,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
const Bool staleContext = entry.contextGeneration != g_bufferContextGeneration;
if (!staleContext && entry.retireSerial > completed) continue;
if (!staleContext && entry.id != 0) {
ScrubIndexedBufferBindingShadowForId(entry.id);
ScrubBufferBindingShadowsForId(entry.id);
g_GLESFuncs.glDeleteBuffers(1, &entry.id);
}
g_retiredUboRings[i] = g_retiredUboRings.back();
@@ -1152,6 +1212,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
}
for (auto& bufferId : m_clientAttributeBufferIds) {
if (bufferId != 0) {
BufferImpl::NoteBufferIdDeleted(bufferId);
g_GLESFuncs.glDeleteBuffers(1, &bufferId);
bufferId = 0;
}
@@ -1324,6 +1385,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
#endif
g_GLESFuncs.glGenTextures(1, &m_backendTextureId);
m_contextGeneration = g_textureContextGeneration;
if (m_backendTextureId == 0) {
MGLOG_E("Failed to generate texture object.");
MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str());
@@ -1332,6 +1394,27 @@ namespace MobileGL::MG_Backend::DirectGLES {
}
}
BackendTextureObject::~BackendTextureObject() {
if (m_backendTextureId == 0) {
return;
}
// Scrub every driver-state shadow that could false-skip when the name
// or this heap address is recycled - regardless of whether the id can
// still be deleted.
ScratchFBOImpl::NoteTextureIdDeleted(m_backendTextureId);
for (auto& unitCache : g_boundTexturesCache) {
for (auto& boundTexture : unitCache) {
if (boundTexture == this) {
boundTexture = nullptr;
}
}
}
if (m_contextGeneration == g_textureContextGeneration && g_GLESFuncs.glDeleteTextures) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId);
}
m_backendTextureId = 0;
}
void BackendTextureObject::Bind(GLenum target, Uint unit) {
#ifdef TRACY_ENABLE
ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
@@ -1364,7 +1447,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
void BackendTextureObject::RecreateBackendTexture() {
if (m_backendTextureId != 0) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId);
ScratchFBOImpl::NoteTextureIdDeleted(m_backendTextureId);
if (m_contextGeneration == g_textureContextGeneration) {
g_GLESFuncs.glDeleteTextures(1, &m_backendTextureId);
}
for (auto& unitCache : g_boundTexturesCache) {
for (auto& boundTexture : unitCache) {
if (boundTexture == this) {
@@ -1375,6 +1461,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
}
g_GLESFuncs.glGenTextures(1, &m_backendTextureId);
m_contextGeneration = g_textureContextGeneration;
if (m_backendTextureId == 0) {
MGLOG_E("Failed to regenerate texture object.");
MGLOG_E("ES glGetError(): %s", MG_Util::ConvertGLEnumToString(g_GLESFuncs.glGetError()).c_str());
@@ -1391,7 +1478,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
// glGetIntegerv - that query forces a driver pipeline sync and, because texture
// uploads run it per dirty texture per frame, it dominated the DirectGLES draw
// path. The backend unpack state is set ONLY by MobileGL's own save/restore
// helpers (this class, TempPixelStoreParameterSync, the R32F copy path), all of
// helpers (this class and, historically, the R32F copy path), all of
// which restore to the resting default, so the shadow stays accurate; a one-time
// forced sync pins the backend to that known default up front. Apply() is
// compare-and-set, so the (now redundant) glPixelStorei calls also usually no-op.
@@ -1563,6 +1650,56 @@ namespace MobileGL::MG_Backend::DirectGLES {
return convertedData.data();
}
// RGB565/RGB5_A1 shadow data is stored as 8-bit unorm; uploading it as GL_UNSIGNED_BYTE
// leaves the 8-bit -> 5/6-bit requantization to the driver, whose rounding direction is
// implementation-defined: Adreno rounds to nearest (lossless round trip) but Mali floors,
// drifting mid-range texels one 5-bit step down and failing the KHR-GL3x
// pixelstoragemodes.teximage3d rgb565/rgb5a1 1/32-eps checks. Repack the shadow rows into
// the packed 16-bit client type with round-to-nearest instead - that recovers the original
// 5/6-bit values exactly (the shadow expansion round(v * 255 / max) is injective), so the
// driver stores them verbatim with no requantization left to its discretion. 4-bit formats
// (RGBA4) are exempt: their 8-bit expansion (v * 17) is exact under either rounding.
// Always retargets *inOutType for these formats (even for null data) so every upload of a
// level uses the same client type.
static const void* PreparePackedNormUpload(TextureInternalFormat format, const IntVec3& texelSize,
const void* data, SizeT byteSize, GLenum* inOutType,
Vector<Uint8>& packedData) {
if (format != TextureInternalFormat::RGB5 && format != TextureInternalFormat::RGB5A1) {
return data;
}
const Bool hasAlpha = format == TextureInternalFormat::RGB5A1;
const GLenum packedType = hasAlpha ? GL_UNSIGNED_SHORT_5_5_5_1 : GL_UNSIGNED_SHORT_5_6_5;
// Idempotent across a region's level loop: glType is shared, so later levels arrive with
// the already-retargeted packed type and must still be converted.
if (*inOutType != GL_UNSIGNED_BYTE && *inOutType != packedType) {
return data;
}
*inOutType = packedType;
if (data == nullptr || byteSize == 0) {
return data;
}
const SizeT srcPixelBytes = hasAlpha ? 4 : 3;
const SizeT texelCount = std::min(static_cast<SizeT>(std::max(texelSize.x(), 0)) *
static_cast<SizeT>(std::max(texelSize.y(), 0)) *
static_cast<SizeT>(std::max(texelSize.z(), 1)),
byteSize / srcPixelBytes);
packedData.resize(texelCount * sizeof(Uint16));
const Uint8* src = static_cast<const Uint8*>(data);
auto* dst = reinterpret_cast<Uint16*>(packedData.data());
for (SizeT i = 0; i < texelCount; ++i, src += srcPixelBytes) {
const Uint32 r = (static_cast<Uint32>(src[0]) * 31u + 127u) / 255u;
const Uint32 b = (static_cast<Uint32>(src[2]) * 31u + 127u) / 255u;
if (hasAlpha) {
const Uint32 g = (static_cast<Uint32>(src[1]) * 31u + 127u) / 255u;
dst[i] = static_cast<Uint16>((r << 11) | (g << 6) | (b << 1) | (src[3] >= 128 ? 1u : 0u));
} else {
const Uint32 g = (static_cast<Uint32>(src[1]) * 63u + 127u) / 255u;
dst[i] = static_cast<Uint16>((r << 11) | (g << 5) | b);
}
}
return packedData.data();
}
void BackendTextureObject::SyncMipmapsToBackend(
const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject) {
if (!stateTextureObject) {
@@ -1706,9 +1843,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData = PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 uploadSize =
GetBackendUploadSize(stateTextureObject->GetTarget(), levelTexelSize);
switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) {
@@ -1760,7 +1900,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& uploadTargets = textureMipmapObject->GetUploadTargets();
if (TextureImpl::IsMultisampleTextureTarget(targetInternal)) {
DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
switch (targetInternal) {
case TextureTarget::Texture2DMultisample:
g_GLESFuncs.glTexStorage2DMultisample(
@@ -1787,7 +1927,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
}
} else if (stateTextureObject->IsImmutable() || m_imageBindableStorageRequired) {
DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 storageSize = GetBackendUploadSize(targetInternal, baseSize);
switch (MapToBackendTextureTarget(targetInternal)) {
case TextureTarget::Texture2D:
@@ -1831,9 +1971,13 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData =
PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
const IntVec3 uploadSize =
GetBackendUploadSize(targetInternal, levelTexelSize);
switch (MapToBackendTextureTarget(targetInternal)) {
@@ -1885,6 +2029,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), levelTexelSize, pData, levelByteSize, glType,
convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData =
PreparePackedNormUpload(textureMipmapObject->GetFormat(), levelTexelSize,
uploadData, levelByteSize, &glType, packedUploadData);
MGLOG_D("%s: target: %s: syncing mip %d: %dx%dx%d, byteSize = %d, pData = %p, "
"levelDirty = %s",
__func__, MG_Util::ConvertTextureUploadTargetToString(uploadTarget).c_str(),
@@ -1892,7 +2040,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
levelByteSize, pData, levelDirty ? "true" : "false");
DebugImpl::ErrorLopper::Clear();
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
auto textureTarget = stateTextureObject->GetTarget();
const IntVec3 uploadSize = GetBackendUploadSize(textureTarget, levelTexelSize);
switch (MapToBackendTextureTarget(textureTarget)) {
@@ -1978,7 +2126,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
textureMipmapObject->GetMipmapTexelSize(uploadTarget, level).y(), byteSize);
auto glUploadTarget = ConvertTextureUploadTargetToBackendGLEnum(uploadTarget);
g_GLESFuncs.glBindBuffer(GL_PIXEL_UNPACK_BUFFER, 0);
BufferImpl::BindPixelUnpackBufferId(0); // no-op once the resting 0 state is pinned
DebugImpl::ErrorLopper::Loop(
[file = __FILE__, line = __LINE__, func = __func__](GLenum err) {
MGLOG_D("%s(%s:%d) ES error: %s", func, file, line,
@@ -1990,6 +2138,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
const void* uploadData = PrepareNormFloatFallbackUpload(
textureMipmapObject->GetFormat(), texelSize, mipData, byteSize, glType,
convertedUploadData);
Vector<Uint8> packedUploadData;
uploadData = PreparePackedNormUpload(textureMipmapObject->GetFormat(), texelSize,
uploadData, byteSize, &glType, packedUploadData);
const IntVec3 uploadSize =
GetBackendUploadSize(stateTextureObject->GetTarget(), texelSize);
switch (MapToBackendTextureTarget(stateTextureObject->GetTarget())) {
@@ -2216,11 +2367,17 @@ namespace MobileGL::MG_Backend::DirectGLES {
return;
}
if (TextureImpl::IsMultisampleTextureTarget(targetInternal)) {
// Multisample targets reject the *sampler* parameters (LOD range, border color) but
// GL_TEXTURE_SWIZZLE_* is texture state, not sampler state, and ES accepts it on them.
// Bailing out entirely used to drop every swizzle write on the floor, which is what the
// frontend already assumes is legal (see GL_Texture.cpp's MS-invalid pname list, which
// deliberately omits the swizzle enums). Note the caches for the skipped parameters are
// still refreshed so they never look stale, but m_cacheSwizzleParams must NOT be, or the
// change detection below would swallow the very writes we came here to emit.
const Bool isMultisampleTarget = TextureImpl::IsMultisampleTextureTarget(targetInternal);
if (isMultisampleTarget) {
m_cacheLodRange = stateTextureObject->GetLevelRange();
m_cacheSwizzleParams = stateTextureObject->GetAllSwizzleParams();
m_cacheBorderColor = stateTextureObject->GetBorderColor();
return;
}
Bind(target);
@@ -2233,14 +2390,14 @@ namespace MobileGL::MG_Backend::DirectGLES {
const auto& levelRange = stateTextureObject->GetLevelRange();
if (m_cacheLodRange.x() != levelRange.x()) {
if (!isMultisampleTarget && m_cacheLodRange.x() != levelRange.x()) {
g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_BASE_LEVEL, static_cast<GLint>(levelRange.x()));
m_cacheLodRange.x() = levelRange.x();
}
DebugImpl::ErrorLopper::Loop([file = __FILE__, line = __LINE__, func = __func__](GLenum err) {
MGLOG_D("%s(%s:%d) ES error %s", func, file, line, MG_Util::ConvertGLEnumToString(err).c_str());
});
if (m_cacheLodRange.y() != levelRange.y()) {
if (!isMultisampleTarget && m_cacheLodRange.y() != levelRange.y()) {
g_GLESFuncs.glTexParameteri(target, GL_TEXTURE_MAX_LEVEL, static_cast<GLint>(levelRange.y()));
m_cacheLodRange.y() = levelRange.y();
}
@@ -2266,7 +2423,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
});
}
if (m_cacheBorderColor != stateTextureObject->GetBorderColor()) {
if (!isMultisampleTarget && m_cacheBorderColor != stateTextureObject->GetBorderColor()) {
const auto& borderColor = stateTextureObject->GetBorderColor();
GLfloat borderColorArray[4] = {borderColor.x(), borderColor.y(), borderColor.z(), borderColor.w()};
g_GLESFuncs.glTexParameterfv(target, GL_TEXTURE_BORDER_COLOR, borderColorArray);
@@ -2295,6 +2452,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
}
Uint g_activeTextureUnit = 0;
Uint g_textureContextGeneration = 1;
Array<Array<BackendTextureObject*, (SizeT)TextureTarget::TextureTargetCount>,
MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS>
g_boundTexturesCache;
@@ -2320,9 +2478,59 @@ namespace MobileGL::MG_Backend::DirectGLES {
ZoneScopedC(TRACY_ZONECOLOR_BACKEND);
#endif
if (target == FramebufferTarget::Read)
g_GLESFuncs.glBindFramebuffer(GL_READ_FRAMEBUFFER, m_backendFBOId);
BindFramebufferId(GL_READ_FRAMEBUFFER, m_backendFBOId);
else
g_GLESFuncs.glBindFramebuffer(GL_DRAW_FRAMEBUFFER, m_backendFBOId);
BindFramebufferId(GL_DRAW_FRAMEBUFFER, m_backendFBOId);
}
namespace {
// Driver-level framebuffer-binding shadow (see Managers.h). Indexed by
// FramebufferTarget {Draw, Read}.
Array<Uint, SizeT(FramebufferTarget::FramebufferTargetCount)> g_driverFBOBindings = {0, 0};
Array<Bool, SizeT(FramebufferTarget::FramebufferTargetCount)> g_driverFBOBindingKnown = {false, false};
} // namespace
void BindFramebufferId(GLenum fbTarget, Uint id) {
const Bool bindsDraw = fbTarget == GL_DRAW_FRAMEBUFFER || fbTarget == GL_FRAMEBUFFER;
const Bool bindsRead = fbTarget == GL_READ_FRAMEBUFFER || fbTarget == GL_FRAMEBUFFER;
const SizeT drawIdx = SizeT(FramebufferTarget::Draw);
const SizeT readIdx = SizeT(FramebufferTarget::Read);
const Bool drawMatches =
!bindsDraw || (g_driverFBOBindingKnown[drawIdx] && g_driverFBOBindings[drawIdx] == id);
const Bool readMatches =
!bindsRead || (g_driverFBOBindingKnown[readIdx] && g_driverFBOBindings[readIdx] == id);
if (drawMatches && readMatches) {
return;
}
g_GLESFuncs.glBindFramebuffer(fbTarget, id);
if (bindsDraw) {
g_driverFBOBindings[drawIdx] = id;
g_driverFBOBindingKnown[drawIdx] = true;
}
if (bindsRead) {
g_driverFBOBindings[readIdx] = id;
g_driverFBOBindingKnown[readIdx] = true;
}
}
Uint CurrentFramebufferBinding(FramebufferTarget target) {
const SizeT idx = SizeT(target);
if (!g_driverFBOBindingKnown[idx]) {
// Cold path: pin the shadow from the driver once (init probes and
// pre-shadow code bind raw but restore what they found).
GLint binding = 0;
g_GLESFuncs.glGetIntegerv(
target == FramebufferTarget::Read ? GL_READ_FRAMEBUFFER_BINDING : GL_DRAW_FRAMEBUFFER_BINDING,
&binding);
g_driverFBOBindings[idx] = static_cast<Uint>(binding);
g_driverFBOBindingKnown[idx] = true;
}
return g_driverFBOBindings[idx];
}
void InvalidateFramebufferBindingCache() {
g_driverFBOBindings = {0, 0};
g_driverFBOBindingKnown = {false, false};
}
void BackendFramebufferObject::InvalidateSyncedState() {
@@ -2376,7 +2584,13 @@ namespace MobileGL::MG_Backend::DirectGLES {
if (glTextureTarget == GL_UNKNOWN_MGL) {
glTextureTarget = TextureImpl::ConvertTextureTargetToBackendGLEnum(textureObject->GetTarget());
}
backendTextureObject->Bind(glTextureTarget);
// glBindTexture rejects cube-face enums (INVALID_ENUM with no
// bind, while Bind() would still record the cube-map cache slot
// as bound): bind via the owning cube target; the attach below
// keeps the face target.
const Bool isCubeFace = glTextureTarget >= GL_TEXTURE_CUBE_MAP_POSITIVE_X &&
glTextureTarget <= GL_TEXTURE_CUBE_MAP_NEGATIVE_Z;
backendTextureObject->Bind(isCubeFace ? GL_TEXTURE_CUBE_MAP : glTextureTarget);
g_GLESFuncs.glFramebufferTexture2D(glFBOTarget, glBackendAttachment, glTextureTarget,
backendTextureObject->GetBackendTextureId(),
static_cast<GLint>(attachmentObject.GetTextureLevel()));
@@ -2465,6 +2679,28 @@ namespace MobileGL::MG_Backend::DirectGLES {
return false;
}
void BackendFramebufferObject::SyncReadBufferToBackend(
const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject) {
if (!stateFBOObject) {
return;
}
auto frontendReadBuf = stateFBOObject->GetReadBuffer();
if (frontendReadBuf == m_frontendReadBuffer) {
return;
}
m_frontendReadBuffer = frontendReadBuf;
GLenum glBackendReadBuffer = GetBackendAttachmentType(frontendReadBuf);
if (m_backendReadBuffer != glBackendReadBuffer) {
m_backendReadBuffer = glBackendReadBuffer;
// glReadBuffer targets whatever FBO is bound to GL_READ_FRAMEBUFFER. When this is
// reached from SyncCurrentFBO's "same FBO as draw" skip path the backend FBO was
// only bound as DRAW, so bind it as READ first to route the read buffer correctly.
Bind(FramebufferTarget::Read);
g_GLESFuncs.glReadBuffer(glBackendReadBuffer);
}
}
void BackendFramebufferObject::SyncToBackend(
const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject, FramebufferTarget asTarget) {
#ifdef TRACY_ENABLE
@@ -2546,16 +2782,8 @@ namespace MobileGL::MG_Backend::DirectGLES {
// 2. Remap read buffer. glReadBuffer writes the READ-bound FBO's state, so
// only apply (and stamp the memo) when this object is bound as READ.
auto frontendReadBuf = stateFBOObject->GetReadBuffer();
if (frontendReadBuf != m_frontendReadBuffer && asTarget == FramebufferTarget::Read) {
m_frontendReadBuffer = frontendReadBuf;
GLenum glBackendReadBuffer = GetBackendAttachmentType(frontendReadBuf);
if (m_backendReadBuffer != glBackendReadBuffer) {
m_backendReadBuffer = glBackendReadBuffer;
g_GLESFuncs.glReadBuffer(glBackendReadBuffer);
}
if (asTarget == FramebufferTarget::Read) {
SyncReadBufferToBackend(stateFBOObject);
}
// -------------------- Attach texture to backend FBO -----------------------
@@ -2662,6 +2890,295 @@ namespace MobileGL::MG_Backend::DirectGLES {
g_fboSyncedObjects = {};
} // namespace FramebufferImpl
namespace ScratchFBOImpl {
namespace {
ScratchFramebuffer g_tempFramebuffer;
ScratchFramebuffer g_blitReadFramebuffer;
ScratchFramebuffer g_blitDrawFramebuffer;
Uint g_completeTinyFBOId = 0;
Uint g_completeTinyRBOId = 0;
// Detach every point the shadow no longer vouches for. Used when the
// shadow is unknown (context reset, texture id deleted while attached).
void ScrubAllAttachments(ScratchFramebuffer& fb, GLenum fbTarget) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
fb.colorTex = 0;
fb.colorTarget = 0;
fb.colorLevel = 0;
fb.colorLayer = -1;
fb.depthTex = 0;
fb.depthTarget = 0;
fb.depthLevel = 0;
fb.depthHasStencil = false;
fb.attachmentsKnown = true;
}
void PrepareForUse(ScratchFramebuffer& fb, GLenum fbTarget) {
if (!fb.attachmentsKnown) {
ScrubAllAttachments(fb, fbTarget);
}
}
// The post-attach glGetError probe below must not misread an error some
// earlier operation left queued; drain before attaching (rare path -
// only runs when the attachment actually changes).
void DrainPendingGLErrors() {
while (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
}
}
// Record the color point as detached when the shadow said something was
// there; the actual detach call is the caller's (it may be replaced by
// the new attach directly when the point is being overwritten).
void RecordNoColor(ScratchFramebuffer& fb) {
fb.colorTex = 0;
fb.colorTarget = 0;
fb.colorLevel = 0;
fb.colorLayer = -1;
}
void RecordNoDepth(ScratchFramebuffer& fb) {
fb.depthTex = 0;
fb.depthTarget = 0;
fb.depthLevel = 0;
fb.depthHasStencil = false;
}
} // namespace
ScratchFramebuffer& TempFramebuffer() {
return g_tempFramebuffer;
}
ScratchFramebuffer& BlitReadFramebuffer() {
return g_blitReadFramebuffer;
}
ScratchFramebuffer& BlitDrawFramebuffer() {
return g_blitDrawFramebuffer;
}
Uint EnsureId(ScratchFramebuffer& fb) {
if (fb.id == 0) {
g_GLESFuncs.glGenFramebuffers(1, &fb.id);
// A fresh FBO has nothing attached and COLOR_ATTACHMENT0 read/draw
// buffers (the ES defaults for a non-default framebuffer).
fb.attachmentsKnown = true;
RecordNoColor(fb);
RecordNoDepth(fb);
fb.readBuffer = GL_COLOR_ATTACHMENT0;
fb.drawBuffer = GL_COLOR_ATTACHMENT0;
}
return fb.id;
}
void EnsureColorAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget,
GLint level) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
if (fb.colorTex == tex && fb.colorTarget == texTarget && fb.colorLevel == level && fb.colorLayer < 0) {
return;
}
if (fb.colorTex != 0) {
// Detach first: if the new attach fails, the point must read as
// missing (incomplete FBO), not silently keep the old texture.
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, texTarget, tex, level);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoColor(fb);
return;
}
fb.colorTex = tex;
fb.colorTarget = texTarget;
fb.colorLevel = level;
fb.colorLayer = -1;
}
void EnsureColorAttachmentLayer(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLint level, GLint layer) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
if (fb.colorTex == tex && fb.colorTarget == 0 && fb.colorLevel == level && fb.colorLayer == layer) {
return;
}
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTextureLayer(fbTarget, GL_COLOR_ATTACHMENT0, tex, level, layer);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoColor(fb);
return;
}
fb.colorTex = tex;
fb.colorTarget = 0;
fb.colorLevel = level;
fb.colorLayer = layer;
}
void EnsureDepthAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level,
Bool withStencil) {
PrepareForUse(fb, fbTarget);
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
RecordNoColor(fb);
}
if (fb.depthTex == tex && fb.depthTarget == texTarget && fb.depthLevel == level &&
fb.depthHasStencil == withStencil) {
return;
}
if (fb.depthTex != 0) {
// One call clears both depth and stencil points regardless of how
// the previous attachment was made.
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
}
DrainPendingGLErrors();
g_GLESFuncs.glFramebufferTexture2D(fbTarget,
withStencil ? GL_DEPTH_STENCIL_ATTACHMENT : GL_DEPTH_ATTACHMENT,
texTarget, tex, level);
if (g_GLESFuncs.glGetError() != GL_NO_ERROR) {
RecordNoDepth(fb);
return;
}
fb.depthTex = tex;
fb.depthTarget = texTarget;
fb.depthLevel = level;
fb.depthHasStencil = withStencil;
}
void EnsureNoColorAttachment(ScratchFramebuffer& fb, GLenum fbTarget) {
PrepareForUse(fb, fbTarget);
if (fb.colorTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, 0, 0);
RecordNoColor(fb);
}
}
void EnsureNoDepthAttachment(ScratchFramebuffer& fb, GLenum fbTarget) {
PrepareForUse(fb, fbTarget);
if (fb.depthTex != 0) {
g_GLESFuncs.glFramebufferTexture2D(fbTarget, GL_DEPTH_STENCIL_ATTACHMENT, GL_TEXTURE_2D, 0, 0);
RecordNoDepth(fb);
}
}
void EnsureReadBuffer(ScratchFramebuffer& fb, GLenum readBuffer) {
if (fb.readBuffer == readBuffer) {
return;
}
g_GLESFuncs.glReadBuffer(readBuffer);
fb.readBuffer = readBuffer;
}
void EnsureDrawBuffer(ScratchFramebuffer& fb, GLenum drawBuffer) {
if (fb.drawBuffer == drawBuffer) {
return;
}
g_GLESFuncs.glDrawBuffers(1, &drawBuffer);
fb.drawBuffer = drawBuffer;
}
Uint EnsureCompleteTinyFramebufferId() {
if (g_completeTinyFBOId != 0) {
return g_completeTinyFBOId;
}
// One-time creation: the renderbuffer binding is context state with no
// shadow, so save/restore it by query here (cold path only).
GLint prevRenderbuffer = 0;
g_GLESFuncs.glGetIntegerv(GL_RENDERBUFFER_BINDING, &prevRenderbuffer);
g_GLESFuncs.glGenFramebuffers(1, &g_completeTinyFBOId);
g_GLESFuncs.glGenRenderbuffers(1, &g_completeTinyRBOId);
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, g_completeTinyFBOId);
g_GLESFuncs.glBindRenderbuffer(GL_RENDERBUFFER, g_completeTinyRBOId);
g_GLESFuncs.glRenderbufferStorage(GL_RENDERBUFFER, GL_RGBA8, 1, 1);
g_GLESFuncs.glFramebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER,
g_completeTinyRBOId);
const GLenum drawBuffer = GL_COLOR_ATTACHMENT0;
g_GLESFuncs.glDrawBuffers(1, &drawBuffer);
g_GLESFuncs.glReadBuffer(GL_COLOR_ATTACHMENT0);
MOBILEGL_ASSERT(g_GLESFuncs.glCheckFramebufferStatus(GL_FRAMEBUFFER) == GL_FRAMEBUFFER_COMPLETE,
"Scratch 1x1 framebuffer is incomplete.");
g_GLESFuncs.glBindRenderbuffer(GL_RENDERBUFFER, static_cast<Uint>(prevRenderbuffer));
return g_completeTinyFBOId;
}
void NoteTextureIdDeleted(Uint textureId) {
if (textureId == 0) {
return;
}
for (ScratchFramebuffer* fb : {&g_tempFramebuffer, &g_blitReadFramebuffer, &g_blitDrawFramebuffer}) {
if (fb->colorTex == textureId || fb->depthTex == textureId) {
fb->attachmentsKnown = false;
}
}
}
void OnBackendContextDestroyed() {
g_tempFramebuffer = {};
g_blitReadFramebuffer = {};
g_blitDrawFramebuffer = {};
g_completeTinyFBOId = 0;
g_completeTinyRBOId = 0;
}
} // namespace ScratchFBOImpl
namespace PixelStoreImpl {
namespace {
PackState g_packState;
Bool g_packStateKnown = false;
void PinPackState(const PackState& value) {
g_GLESFuncs.glPixelStorei(GL_PACK_ALIGNMENT, value.Alignment);
g_GLESFuncs.glPixelStorei(GL_PACK_ROW_LENGTH, value.RowLength);
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_ROWS, value.SkipRows);
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_PIXELS, value.SkipPixels);
g_packState = value;
g_packStateKnown = true;
}
} // namespace
void ApplyPackState(const PackState& desired) {
if (!g_packStateKnown) {
PinPackState(desired);
return;
}
if (desired.Alignment != g_packState.Alignment) {
g_GLESFuncs.glPixelStorei(GL_PACK_ALIGNMENT, desired.Alignment);
g_packState.Alignment = desired.Alignment;
}
if (desired.RowLength != g_packState.RowLength) {
g_GLESFuncs.glPixelStorei(GL_PACK_ROW_LENGTH, desired.RowLength);
g_packState.RowLength = desired.RowLength;
}
if (desired.SkipRows != g_packState.SkipRows) {
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_ROWS, desired.SkipRows);
g_packState.SkipRows = desired.SkipRows;
}
if (desired.SkipPixels != g_packState.SkipPixels) {
g_GLESFuncs.glPixelStorei(GL_PACK_SKIP_PIXELS, desired.SkipPixels);
g_packState.SkipPixels = desired.SkipPixels;
}
}
PackState CurrentPackState() {
if (!g_packStateKnown) {
// Fresh/unknown context: pin to the GL defaults (what a new context
// starts with; writing them makes the shadow authoritative either way).
PinPackState(PackState{});
}
return g_packState;
}
void InvalidatePackStateCache() {
g_packStateKnown = false;
}
} // namespace PixelStoreImpl
namespace PrgramImpl {
Uint32 g_snormFallbackClampOutputMask = 0;
Uint32 g_unormFallbackClampOutputMask = 0;
@@ -2784,6 +3301,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
effectiveSpirv = &uboPrecisionSpirv;
}
// noperspective is core desktop GLSL and reaches here as the SPIR-V NoPerspective
// decoration. SPIRV-Cross renders it as ESSL `noperspective` + `#extension
// GL_NV_shader_noperspective_interpolation : require`; a driver without that extension
// rejects the require. So on such devices emulate screen-linear interpolation instead
// (pre-multiply outputs by gl_Position.w, recover inputs via gl_FragCoord.w) and drop
// the decoration - exact, extension-free. Devices that have the extension keep the
// decoration and let the hardware do it natively.
Vector<unsigned int> noperspectiveSpirv;
if (!g_GLESCapabilities.SupportsNoperspectiveInterpolation &&
MG_Util::ShaderTranspiler::ShaderCompiler::EmulateNoPerspectiveForEssl(
*effectiveSpirv, noperspectiveSpirv) &&
!noperspectiveSpirv.empty()) {
effectiveSpirv = &noperspectiveSpirv;
}
MG_Util::ShaderTranspiler::SpvcSession spvcSession(*effectiveSpirv,
MG_Util::ShaderTranspiler::SessionUsageBit::Transpile);
+122
View File
@@ -178,6 +178,21 @@ namespace MobileGL::MG_Backend::DirectGLES {
// glBindBuffer with a redundant-bind cache for GL_ARRAY_BUFFER.
void BindBufferId(GLenum target, Uint id);
void InvalidateArrayBufferBindingCache();
// Redundant-bind caches for the driver-level GL_PIXEL_PACK/UNPACK_BUFFER
// bindings. Every backend readback (glReadPixels / pack-PBO map) and pixel
// upload site routes its binding through these so the shadow always matches
// the driver; the resting state between operations is 0, which keeps any
// path that implicitly assumes "no PBO bound" correct. Scrubbed when a
// buffer id is deleted/pooled (GL resets a deleted buffer's bindings to 0,
// and a recycled name matching the shadow would false-skip the rebind) and
// invalidated on MakeCurrent (context may reset).
void BindPixelPackBufferId(Uint id);
void BindPixelUnpackBufferId(Uint id);
void InvalidatePixelBufferBindingCaches();
// A GL buffer id is being deleted by code outside BufferImpl (e.g. the VAO
// client-attribute staging buffers): scrub every buffer-binding shadow that
// could false-skip when the name is recycled.
void NoteBufferIdDeleted(Uint id);
// Redundant-bind cache for INDEXED buffer bindings (glBindBufferBase/Range on
// GL_UNIFORM_BUFFER / GL_SHADER_STORAGE_BUFFER): skips the GL call when the
// (id, range) already at that index matches, like the array-buffer/texture/
@@ -332,6 +347,12 @@ namespace MobileGL::MG_Backend::DirectGLES {
class BackendTextureObject {
public:
BackendTextureObject();
// Deletes the GL texture (frontend glDeleteTextures used to leak every
// backend id for the context lifetime) and scrubs the binding/scratch-FBO
// shadows so a recycled name or heap address cannot false-skip a rebind.
~BackendTextureObject();
BackendTextureObject(const BackendTextureObject&) = delete;
BackendTextureObject& operator=(const BackendTextureObject&) = delete;
void SyncMipmapsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
void SyncBuiltinSamplerToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
void SyncTextureParamsToBackend(const SharedPtr<MG_State::GLState::ITextureObject>& stateTextureObject);
@@ -343,6 +364,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
void RecreateBackendTexture();
Uint m_backendTextureId = 0;
// ES context generation the id was created under; a dtor running after
// that context died must not delete a foreign (recycled) name.
Uint m_contextGeneration = 0;
Bool m_isInitialized = false;
Bool m_imageBindableStorageRequired = false;
Bool m_backendStorageImmutable = false;
@@ -367,6 +391,9 @@ namespace MobileGL::MG_Backend::DirectGLES {
MG_State::GLState::TextureState::MAX_TEXTURE_IMAGE_UNITS>
g_boundTexturesCache;
extern Uint g_activeTextureUnit;
// Bumped when the backend ES context is destroyed; texture ids stamped with
// an older generation belong to a dead context and must not be deleted.
extern Uint g_textureContextGeneration;
} // namespace TextureImpl
namespace FramebufferImpl {
@@ -375,6 +402,10 @@ namespace MobileGL::MG_Backend::DirectGLES {
BackendFramebufferObject();
void SyncToBackend(const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject,
FramebufferTarget asTarget);
// Apply only this FBO's read buffer (glReadBuffer) to the backend. Split out so it can
// still run when SyncCurrentFBO skips the READ-target sync because the same GL FBO is
// bound as both draw and read (otherwise glReadBuffer changes would be silently dropped).
void SyncReadBufferToBackend(const SharedPtr<MG_State::GLState::FramebufferObject>& stateFBOObject);
void InvalidateSyncedState();
Uint GetBackendFramebufferId() const { return m_backendFBOId; }
void Bind(FramebufferTarget target) const;
@@ -414,8 +445,99 @@ namespace MobileGL::MG_Backend::DirectGLES {
extern Array<Uint16, SizeT(FramebufferTarget::FramebufferTargetCount)> g_fboSyncedObjectVersions;
extern Array<MG_State::GLState::FramebufferObject*, SizeT(FramebufferTarget::FramebufferTargetCount)>
g_fboSyncedObjects;
// Driver-level READ/DRAW framebuffer-binding shadow. Every backend
// glBindFramebuffer routes through BindFramebufferId so scoped helpers can
// save/restore the current binding without a glGetIntegerv round-trip (that
// query forces a driver pipeline sync) and so redundant rebinds no-op.
// Starts unknown; the first CurrentFramebufferBinding() query pins it from
// the driver once. Invalidated on MakeCurrent (context may reset).
// GL_FRAMEBUFFER binds both targets.
void BindFramebufferId(GLenum fbTarget, Uint id);
Uint CurrentFramebufferBinding(FramebufferTarget target);
void InvalidateFramebufferBindingCache();
} // namespace FramebufferImpl
// Shared scratch framebuffers for the readback/copy/blit emulation paths, with a
// driver-side attachment shadow: repeated uses skip redundant detach/attach GL
// calls, and an attachment left by one use (e.g. a depth copy's DEPTH_STENCIL
// texture) is detached exactly when a later use of another aspect would
// otherwise inherit it (stale cross-aspect attachments made the shared temp FBO
// incomplete and silently degraded later readbacks).
namespace ScratchFBOImpl {
struct ScratchFramebuffer {
Uint id = 0;
// false => attachment state unknown; scrub every point on next use.
// A fresh FBO starts with nothing attached, so creation sets it true.
Bool attachmentsKnown = false;
Uint colorTex = 0;
GLenum colorTarget = 0;
GLint colorLevel = 0;
GLint colorLayer = -1; // >= 0 => attached via glFramebufferTextureLayer
Uint depthTex = 0;
GLenum depthTarget = 0;
GLint depthLevel = 0;
Bool depthHasStencil = false;
// Per-FBO read/draw buffer state (0 = unknown, set on first use).
GLenum readBuffer = 0;
GLenum drawBuffer = 0;
};
ScratchFramebuffer& TempFramebuffer(); // GetTexImage READ / CopyTex*Image2D depth DRAW
ScratchFramebuffer& BlitReadFramebuffer(); // texture-to-texture blit source
ScratchFramebuffer& BlitDrawFramebuffer(); // texture-to-texture blit destination
// Returns the GL id, generating it if needed (requires a current ES context).
Uint EnsureId(ScratchFramebuffer& fb);
// The fb must currently be bound at fbTarget (glReadBuffer/glDrawBuffers
// target the READ/DRAW binding respectively). Each Ensure* performs the
// minimal detach/attach set and keeps the shadow in sync; a failed attach
// records the point as detached so the completeness check fails instead of
// silently reading a stale attachment.
void EnsureColorAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level);
void EnsureColorAttachmentLayer(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLint level, GLint layer);
void EnsureDepthAttachment2D(ScratchFramebuffer& fb, GLenum fbTarget, Uint tex, GLenum texTarget, GLint level,
Bool withStencil);
void EnsureNoColorAttachment(ScratchFramebuffer& fb, GLenum fbTarget);
void EnsureNoDepthAttachment(ScratchFramebuffer& fb, GLenum fbTarget);
void EnsureReadBuffer(ScratchFramebuffer& fb, GLenum readBuffer);
void EnsureDrawBuffer(ScratchFramebuffer& fb, GLenum drawBuffer);
// A 1x1 RGBA8-renderbuffer-complete FBO (GenerateMipmap needs a complete
// binding while respecifying texture storage). Attachment is set once at
// creation and never changes.
Uint EnsureCompleteTinyFramebufferId();
// A backend texture id is being deleted or respecified: a scratch FBO still
// referencing it would hold a dangling attachment (ES only auto-detaches
// from the *bound* framebuffer), and a recycled name could false-skip a
// re-attach; force a full scrub on next use.
void NoteTextureIdDeleted(Uint textureId);
// The ES context (and the scratch FBO ids with it) is going away.
void OnBackendContextDestroyed();
} // namespace ScratchFBOImpl
// Driver-level GL_PACK_* pixel-store shadow, the readback-side sibling of the
// upload path's ScopedDefaultUnpackState (Managers.cpp): the backend PACK state
// is written ONLY through ApplyPackState, so scoped helpers can save/restore it
// from the shadow instead of glGetIntegerv (which forces a driver pipeline
// sync), and redundant glPixelStorei calls no-op. The first Apply/Current call
// pins the driver to the shadow by writing all fields once. Invalidated on
// MakeCurrent (context may reset). PACK_IMAGE_HEIGHT/SKIP_IMAGES/SWAP_BYTES/
// LSB_FIRST have no ES equivalents; readbacks honor them on the CPU from the
// frontend context state instead.
namespace PixelStoreImpl {
struct PackState {
GLint Alignment = 4;
GLint RowLength = 0;
GLint SkipRows = 0;
GLint SkipPixels = 0;
Bool operator==(const PackState& o) const {
return Alignment == o.Alignment && RowLength == o.RowLength && SkipRows == o.SkipRows &&
SkipPixels == o.SkipPixels;
}
};
void ApplyPackState(const PackState& desired);
PackState CurrentPackState();
void InvalidatePackStateCache();
} // namespace PixelStoreImpl
// Image uniforms take their unit from the layout(binding=N) qualifier baked into
// the transpiled ESSL; unlike samplers they must not (and in ES cannot) be
// assigned through glUniform1i.
@@ -794,5 +794,32 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_vulkanCaps.MaxShaderStorageBlockSize,
m_dynamicParameters.MaxShaderStorageBlockSize);
}
switch (m_vulkanCaps.VendorId) {
case 0x5143u: // VK_VENDOR_ID: Qualcomm
m_dynamicParameters.GpuVendor = GpuVendorKind::Qualcomm;
break;
case 0x13B5u: // ARM
m_dynamicParameters.GpuVendor = GpuVendorKind::Arm;
break;
case 0x10DEu: // NVIDIA
m_dynamicParameters.GpuVendor = GpuVendorKind::Nvidia;
break;
case 0x1002u: // AMD
m_dynamicParameters.GpuVendor = GpuVendorKind::Amd;
break;
case 0x8086u: // Intel
m_dynamicParameters.GpuVendor = GpuVendorKind::Intel;
break;
case 0x1010u: // Imagination
m_dynamicParameters.GpuVendor = GpuVendorKind::ImgTec;
break;
case 0x10005u: // Mesa software (lavapipe)
case 0x1AE0u: // Google (SwiftShader)
m_dynamicParameters.GpuVendor = GpuVendorKind::Software;
break;
default:
m_dynamicParameters.GpuVendor = GpuVendorKind::Unknown;
break;
}
}
} // namespace MobileGL::MG_Backend::DirectVulkan
@@ -8,6 +8,7 @@
#include "PipelineFactory.h"
namespace MobileGL::MG_Backend::DirectVulkan {
static const char* PrimitiveTopologyToString(VkPrimitiveTopology topology) {
switch (topology) {
@@ -108,6 +109,81 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"vkCreatePipelineCache");
}
// Must be called once, before any pipeline is created: the flag is not part of the
// pipeline hash, so flipping it mid-life would serve cached pipelines built under the
// old value.
void PipelineFactory::SetSuppressBlendedDepthWrite(Bool enabled) {
s_suppressBlendedDepthWrite = enabled;
}
Bool PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(MG_Config::QuirkOverride quirkOverride,
Uint32 vendorId) {
static constexpr Uint32 kVendorIdQualcomm = 0x5143;
switch (quirkOverride) {
case MG_Config::QuirkOverride::ForceOn:
return true;
case MG_Config::QuirkOverride::ForceOff:
return false;
case MG_Config::QuirkOverride::Auto:
default:
return vendorId == kVendorIdQualcomm;
}
}
namespace {
// MIN/MAX extremum blending: the signature of a depth-bounds accumulation pass
// (MC 26.3 OIT writes vec4(-linD, linD, deviceZ, 0) under GL_MAX while writing
// depth for its equality chain). MIN/MAX ignore blend factors per the Vulkan spec.
//
// Deliberately the ONLY shape stripped. A quirk should touch as little unrelated
// content as possible, and a trace sweep of every fixture showed the wider
// alternatives all cost more than they fix:
// - additive ONE+ONE with a depth write matched zero draws of the 26.3 chain
// (its transmittance/accumulate passes disable depth writes themselves) - the
// only real content it caught was harmless additive glow effects (Create);
// - sorted-transparency "over" blends (SRC_ALPHA-style) are order-dependent,
// drawn once per surface, and rely on their depth writes for occlusion;
// - separate-alpha accumulation over an over-blending color channel has no
// known pairing with a depth-equality chain (color channel only, see tests).
// If a future workload pairs another blend shape with an equality chain, widen
// this with that evidence in hand rather than pre-emptively.
Bool IsAccumulationBlend(const VkPipelineColorBlendAttachmentState& attachment) {
return attachment.colorBlendOp == VK_BLEND_OP_MIN ||
attachment.colorBlendOp == VK_BLEND_OP_MAX;
}
} // namespace
Bool PipelineFactory::ShouldSuppressDepthWrite(const PipelineCreatePayload& payload) {
if (!payload.depthWriteEnable) {
return false;
}
// A shader that assigns gl_FragDepth supplies depth itself rather than taking the
// pipeline's interpolated Z, so a driver that varies the vertex position math
// between pipelines cannot desynchronize it. (A gl_FragDepth = gl_FragCoord.z
// passthrough is the exception that stays exposed; no known content pairs one with
// an equality chain, and 26.3's composite is a genuine computed-depth writer.)
if (payload.fragmentReplacesDepth) {
return false;
}
for (Uint32 i = 0; i < payload.colorAttachmentCount; ++i) {
const VkPipelineColorBlendAttachmentState& attachment = payload.colorBlendAttachments[i];
if (attachment.blendEnable != VK_TRUE) {
continue;
}
// All color writes masked: blending is moot (depth-prepass pattern that left
// GL_BLEND enabled); stripping the depth write would delete the whole prepass.
if (attachment.colorWriteMask == 0) {
continue;
}
// Any attachment qualifies, not just attachment 0: the 26.3 transmittance pass
// accumulates into a 2-target MRT and must stay stripped.
if (IsAccumulationBlend(attachment)) {
return true;
}
}
return false;
}
PipelineFactory::~PipelineFactory() {
DestroyAll();
if (m_pipelineCache != VK_NULL_HANDLE) {
@@ -152,6 +228,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
XXH64_update(m_hashState, &payload.backStencilDepthFailOp, sizeof(payload.backStencilDepthFailOp)));
XXHASH_VERIFY(
XXH64_update(m_hashState, &payload.backStencilCompareOp, sizeof(payload.backStencilCompareOp)));
XXHASH_VERIFY(
XXH64_update(m_hashState, &payload.fragmentReplacesDepth, sizeof(payload.fragmentReplacesDepth)));
if (payload.colorAttachmentCount > 0) {
XXHASH_VERIFY(XXH64_update(
m_hashState,
@@ -258,6 +336,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
for (Uint32 i = 0; i < payload.colorAttachmentCount; ++i) {
colorAttachments[i] = payload.colorBlendAttachments[i];
}
// Suppress depth writes on accumulation-blended pipelines when the active driver
// cannot keep vertex positions invariant across the pipelines of a multi-pass
// depth-equality chain (see SetSuppressBlendedDepthWrite). The decision is narrowed
// in ShouldSuppressDepthWrite: sorted-transparency "over" blends (vanilla MC water),
// gl_FragDepth writers, and masked-out attachments keep their depth writes.
// This bakes the decision into the pipeline, which only works because depth write is
// static state here - adding VK_DYNAMIC_STATE_DEPTH_WRITE_ENABLE to kDynamicStates
// would let the record-time value override it and silently disable the quirk.
if (s_suppressBlendedDepthWrite && ShouldSuppressDepthWrite(payload)) {
depthStencil.depthWriteEnable = VK_FALSE;
}
VkPipelineColorBlendStateCreateInfo blend{VK_STRUCTURE_TYPE_PIPELINE_COLOR_BLEND_STATE_CREATE_INFO};
blend.logicOpEnable = payload.logicOpEnable ? VK_TRUE : VK_FALSE;
blend.logicOp = payload.logicOp;
@@ -49,6 +49,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkStencilOp backStencilPassOp = VK_STENCIL_OP_KEEP;
VkStencilOp backStencilDepthFailOp = VK_STENCIL_OP_KEEP;
VkCompareOp backStencilCompareOp = VK_COMPARE_OP_ALWAYS;
// The fragment module writes gl_FragDepth (SPIR-V DepthReplacing); exempts the
// pipeline from the blended depth-write quirk (see ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false;
Array<VkPipelineColorBlendAttachmentState, kMaxColorAttachments> colorBlendAttachments{};
const Vector<VkPipelineShaderStageCreateInfo>* stages = nullptr;
const VkPipelineVertexInputStateCreateInfo* vertexInputState = nullptr;
@@ -62,6 +65,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipeline GetOrCreatePipeline(const PipelineCreatePayload& payload);
void DestroyAll();
// Driver quirk: suppress depth writes on accumulation-blended pipelines. Multi-pass
// depth-equality rendering (a blended prepass writes depth that later passes re-test
// with an equality-inclusive compare on the re-rasterized geometry) requires
// cross-pipeline position invariance that some mobile compilers do not provide, even
// with the SPIR-V Invariant decoration; whole primitives then drop out of the later
// passes. Only MIN/MAX extremum blends are stripped - the signature of such a
// chain's depth-bounds pass (MC 26.3 OIT), and per a fixture-wide trace sweep the
// only depth-writing shape the chain actually uses - so every other blend
// (sorted-transparency "over" like vanilla MC water, additive glows, ...) keeps
// its depth writes. Set at renderer initialization based on the active driver.
static void SetSuppressBlendedDepthWrite(Bool enabled);
static Bool IsSuppressBlendedDepthWriteEnabled() { return s_suppressBlendedDepthWrite; }
// Device gate for the quirk: ForceOn/ForceOff bypass detection, Auto enables it on
// the known-affected vendor (Qualcomm).
static Bool ShouldSuppressBlendedDepthWriteForDevice(MG_Config::QuirkOverride quirkOverride,
Uint32 vendorId);
// Pure per-pipeline strip decision (exempts gl_FragDepth writers, masked-out and
// non-accumulation blends); combined with the device flag in CreatePipeline. Static
// and payload-only so tests can pin the contract without a VkDevice.
static Bool ShouldSuppressDepthWrite(const PipelineCreatePayload& payload);
private:
VkPipeline CreatePipeline(const PipelineCreatePayload& payload) const;
@@ -70,5 +94,6 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipelineCache m_pipelineCache = VK_NULL_HANDLE;
UnorderedMap<HashType, VkPipeline> m_cache;
static inline XXH64_state_t* m_hashState = XXH64_createState();
static inline Bool s_suppressBlendedDepthWrite = false;
};
} // namespace MobileGL::MG_Backend::DirectVulkan
@@ -312,34 +312,6 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return targetEnv;
}
// Cheap raw-word scan for an `OpDecorate <id> BuiltIn InstanceIndex` decoration. Used
// only to decide whether to warn when shaderDrawParameters is unavailable; a false
// negative merely suppresses a diagnostic.
Bool SpirvDeclaresInstanceIndexBuiltin(const Vector<Uint>& spirv) {
constexpr Uint32 kSpirvMagicNumber = 0x07230203u;
constexpr SizeT kHeaderWordCount = 5;
if (spirv.size() <= kHeaderWordCount || spirv[0] != kSpirvMagicNumber) {
return false;
}
SizeT wordIndex = kHeaderWordCount;
while (wordIndex < spirv.size()) {
const Uint32 firstWord = spirv[wordIndex];
const Uint32 wordCount = firstWord >> 16;
const auto opcode = static_cast<spv::Op>(firstWord & 0xffffu);
if (wordCount == 0 || wordIndex + wordCount > spirv.size()) {
break;
}
if (opcode == spv::Op::OpDecorate && wordCount >= 4 &&
static_cast<spv::Decoration>(spirv[wordIndex + 2]) == spv::Decoration::BuiltIn &&
static_cast<spv::BuiltIn>(spirv[wordIndex + 3]) == spv::BuiltIn::InstanceIndex) {
return true;
}
wordIndex += wordCount;
}
return false;
}
Bool IsInterfaceVariableStaticallyUsed(const Vector<Uint>& spirv, Uint32 spirvId) {
if (spirv.empty() || spirvId == 0) {
return false;
@@ -1218,6 +1190,41 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}
} // namespace
// A shader that assigns gl_FragDepth (SPIR-V DepthReplacing) supplies depth itself
// instead of taking the pipeline's interpolated Z, so a driver that varies the vertex
// position math between pipelines cannot desynchronize it; the blended depth-write
// quirk therefore leaves it alone (see PipelineFactory::ShouldSuppressDepthWrite).
Bool ProgramFactory::ReflectedFragmentReplacesDepth(const SpvReflectShaderModule& reflectModule) {
for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) {
const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex];
for (Uint32 modeIndex = 0; modeIndex < entryPoint.execution_mode_count; ++modeIndex) {
if (entryPoint.execution_modes[modeIndex] == SpvExecutionModeDepthReplacing) {
return true;
}
}
}
return false;
}
// glslang's relaxed-Vulkan mode maps GL's gl_InstanceID onto the InstanceIndex builtin.
// Without shaderDrawParameters there is no gl_BaseInstance to subtract, so such a shader
// cannot be corrected and instanced draws with a non-zero baseInstance misrender; this
// detects the case so the user gets one warning instead of silent corruption.
Bool ProgramFactory::ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule) {
for (Uint32 entryIndex = 0; entryIndex < reflectModule.entry_point_count; ++entryIndex) {
const SpvReflectEntryPoint& entryPoint = reflectModule.entry_points[entryIndex];
for (Uint32 variableIndex = 0; variableIndex < entryPoint.input_variable_count; ++variableIndex) {
const SpvReflectInterfaceVariable* variable = entryPoint.input_variables[variableIndex];
if (variable != nullptr &&
(variable->decoration_flags & SPV_REFLECT_DECORATION_BUILT_IN) != 0 &&
variable->built_in == SpvBuiltInInstanceIndex) {
return true;
}
}
}
return false;
}
VkShaderStageFlagBits ProgramFactory::ToVkStage(ShaderStage stage) {
switch (stage) {
case ShaderStage::Vertex:
@@ -1465,6 +1472,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
continue;
}
if (!m_shaderDrawParametersEnabled && ReflectedReadsInstanceIndexBuiltin(reflectModule)) {
static Bool s_warnedInstanceIndexUnsupported = false;
if (!s_warnedInstanceIndexUnsupported) {
s_warnedInstanceIndexUnsupported = true;
MGLOG_W("ProgramFactory: shaderDrawParameters is unavailable; gl_InstanceID cannot be "
"rebased and instanced draws with a non-zero baseInstance may render incorrectly");
}
}
uint32_t inputCount = 0;
SpvReflectResult reflectResult = spvReflectEnumerateInputVariables(&reflectModule, &inputCount, nullptr);
MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS,
@@ -1512,6 +1528,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkProgramObject& entry) const {
entry.activeFragmentOutputLocationMask = 0;
entry.fragmentOutputTypes.fill(0);
entry.fragmentReplacesDepth = false;
for (SizeT moduleIndex = 0; moduleIndex < shaders.size() && moduleIndex < spirv.size(); ++moduleIndex) {
if (!shaders[moduleIndex] || shaders[moduleIndex]->GetShaderStage() != ShaderStage::Fragment) {
@@ -1530,9 +1547,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"ProgramFactory::ReflectFragmentOutputs: failed to create reflection module (result=%d)",
static_cast<Int>(createResult));
if (createResult != SPV_REFLECT_RESULT_SUCCESS) {
// Fail toward the exemption: stripping a genuine gl_FragDepth writer would
// corrupt its depth output outright, while wrongly exempting an accumulation
// pass merely reverts that one program to the pre-quirk behavior.
entry.fragmentReplacesDepth = true;
continue;
}
entry.fragmentReplacesDepth = ReflectedFragmentReplacesDepth(reflectModule);
uint32_t outputCount = 0;
SpvReflectResult reflectResult = spvReflectEnumerateOutputVariables(&reflectModule, &outputCount, nullptr);
MOBILEGL_ASSERT(reflectResult == SPV_REFLECT_RESULT_SUCCESS,
@@ -1808,6 +1831,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_BUFFER;
} else if (kind == DescriptorBindingKind::StorageImage) {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_STORAGE_IMAGE;
entry.hasStorageImages = true;
} else {
layoutBinding.descriptorType = VK_DESCRIPTOR_TYPE_COMBINED_IMAGE_SAMPLER;
}
@@ -1862,28 +1886,43 @@ namespace MobileGL::MG_Backend::DirectVulkan {
moduleSpirvs[i] = spv;
}
// GL apps depend on cross-program position invariance for multi-pass equality
// depth tests (MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the
// depth its own first pass wrote); decorate Position outputs Invariant so
// per-pipeline compilers cannot vary the position math between passes.
{
Vector<Uint> invariantSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::DecoratePositionInvariantForVulkan(
moduleSpirvs[i], invariantSpirv)) {
moduleSpirvs[i] = std::move(invariantSpirv);
} else {
// The pass round-trips through SPIRV-Tools IR, so an unparseable module
// fails open and keeps the undecorated words - which silently reinstates
// the multi-pass invariance bug rather than breaking anything loudly.
MGLOG_E("ProgramFactory: position-invariant decoration failed for program %u; "
"keeping the original module - multi-pass depth-equality chains "
"(e.g. MC 26.3 OIT clouds) may drop primitives on this device",
program.GetExternalIndex());
}
}
// glslang's relaxed-Vulkan mode aliases GL's zero-based gl_InstanceID to Vulkan's
// gl_InstanceIndex, which wrongly includes the draw's baseInstance. Rebase vertex-stage
// loads to (InstanceIndex - BaseInstance) so shaders observe GL semantics. Reflection
// below runs on the rebased words so the added BaseInstance builtin stays consistent.
if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Vertex) {
if (m_shaderDrawParametersEnabled) {
Vector<Uint> rebasedSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::RebaseInstanceIndexForVulkan(moduleSpirvs[i],
rebasedSpirv)) {
moduleSpirvs[i] = std::move(rebasedSpirv);
} else {
MGLOG_E("ProgramFactory: failed to rebase gl_InstanceID for program %u; "
"instanced draws with a non-zero baseInstance may render incorrectly",
program.GetExternalIndex());
}
} else if (SpirvDeclaresInstanceIndexBuiltin(moduleSpirvs[i])) {
static Bool s_warnedInstanceIndexUnsupported = false;
if (!s_warnedInstanceIndexUnsupported) {
s_warnedInstanceIndexUnsupported = true;
MGLOG_W("ProgramFactory: shaderDrawParameters is unavailable; gl_InstanceID cannot be "
"rebased and instanced draws with a non-zero baseInstance may render incorrectly");
}
// The unsupported-device counterpart of this rebase (warning when a shader reads
// the builtin but shaderDrawParameters is missing) rides along with
// ReflectVertexInputs, which already reflects this stage.
if (shaders[i] && shaders[i]->GetShaderStage() == ShaderStage::Vertex &&
m_shaderDrawParametersEnabled) {
Vector<Uint> rebasedSpirv;
if (MG_Util::ShaderTranspiler::ShaderCompiler::RebaseInstanceIndexForVulkan(moduleSpirvs[i],
rebasedSpirv)) {
moduleSpirvs[i] = std::move(rebasedSpirv);
} else {
MGLOG_E("ProgramFactory: failed to rebase gl_InstanceID for program %u; "
"instanced draws with a non-zero baseInstance may render incorrectly",
program.GetExternalIndex());
}
}
@@ -67,6 +67,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Vector<Bool> storageImageUsesBindingFormatByBinding;
Vector<String> storageBlockNameByBinding;
Vector<Int> storageBlockIndexByBinding;
// Set once during ReflectLayout so the per-draw path can skip the whole
// storage-image preparation for the overwhelming majority of programs.
Bool hasStorageImages = false;
Int globalUboBinding = -1;
Uint32 activeVertexInputLocationMask = 0;
Array<GLenum, kMaxVertexInputLocations> vertexInputTypes{};
@@ -75,6 +78,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ShaderStage rasterizationProducerStage = ShaderStage::Unknown;
Uint32 producerOutputComponentCount = 0;
Uint32 fragmentInputComponentCount = 0;
// The fragment module declares the DepthReplacing execution mode (writes
// gl_FragDepth); shader-computed depth is immune to the cross-pipeline
// position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false;
static inline VkDevice s_device = VK_NULL_HANDLE;
@@ -99,6 +106,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
std::move(other.storageImageUsesBindingFormatByBinding);
storageBlockNameByBinding = std::move(other.storageBlockNameByBinding);
storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding);
hasStorageImages = other.hasStorageImages;
globalUboBinding = other.globalUboBinding;
activeVertexInputLocationMask = other.activeVertexInputLocationMask;
vertexInputTypes = other.vertexInputTypes;
@@ -107,15 +115,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
rasterizationProducerStage = other.rasterizationProducerStage;
producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth;
other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE;
other.hasStorageImages = false;
other.globalUboBinding = -1;
other.activeVertexInputLocationMask = 0;
other.activeFragmentOutputLocationMask = 0;
other.rasterizationProducerStage = ShaderStage::Unknown;
other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false;
}
VkProgramObject& operator=(VkProgramObject&& other) noexcept {
if (this == &other) {
@@ -139,6 +150,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
std::move(other.storageImageUsesBindingFormatByBinding);
storageBlockNameByBinding = std::move(other.storageBlockNameByBinding);
storageBlockIndexByBinding = std::move(other.storageBlockIndexByBinding);
hasStorageImages = other.hasStorageImages;
globalUboBinding = other.globalUboBinding;
activeVertexInputLocationMask = other.activeVertexInputLocationMask;
vertexInputTypes = other.vertexInputTypes;
@@ -147,15 +159,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
rasterizationProducerStage = other.rasterizationProducerStage;
producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth;
other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE;
other.hasStorageImages = false;
other.globalUboBinding = -1;
other.activeVertexInputLocationMask = 0;
other.activeFragmentOutputLocationMask = 0;
other.rasterizationProducerStage = ShaderStage::Unknown;
other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false;
return *this;
}
@@ -203,6 +218,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static VkShaderStageFlagBits ToVkStage(ShaderStage stage);
static VkFormat ConvertSpirvImageFormatToVkFormat(SpvImageFormat format);
static SamplerNumericDomain UniformTypeToSamplerNumericDomain(GLenum glType);
// True when any entry point declares the DepthReplacing execution mode, i.e. the
// shader assigns gl_FragDepth. Exposed so the blended depth-write quirk's exemption
// can be pinned by tests. A false negative loses the exemption, so such a shader is
// stripped conservatively and forfeits its depth write.
static Bool ReflectedFragmentReplacesDepth(const SpvReflectShaderModule& reflectModule);
// True when an entry point reads the InstanceIndex builtin. Only gates a diagnostic:
// without shaderDrawParameters such a shader cannot have gl_InstanceID rebased.
static Bool ReflectedReadsInstanceIndexBuiltin(const SpvReflectShaderModule& reflectModule);
private:
struct ProgramLookupCache {
@@ -298,8 +298,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// descriptor valid.
const Bool forceNearestFiltering = numericDomain == SamplerNumericDomain::SignedInteger ||
numericDomain == SamplerNumericDomain::UnsignedInteger;
const VkFormat sampledViewFormat =
VkTextureManager::ResolveSampledImageViewFormat(resource->format, numericDomain);
SamplerResolveMemo* viewFormatMemo =
binding < m_samplerResolveMemo.size() ? &m_samplerResolveMemo[binding] : nullptr;
VkFormat sampledViewFormat;
if (viewFormatMemo != nullptr && viewFormatMemo->viewFormatValid &&
viewFormatMemo->viewFormatSource == resource->format &&
viewFormatMemo->viewFormatDomain == numericDomain) {
sampledViewFormat = viewFormatMemo->viewFormat;
} else {
sampledViewFormat =
VkTextureManager::ResolveSampledImageViewFormat(resource->format, numericDomain);
if (viewFormatMemo != nullptr) {
viewFormatMemo->viewFormatSource = resource->format;
viewFormatMemo->viewFormatDomain = numericDomain;
viewFormatMemo->viewFormat = sampledViewFormat;
viewFormatMemo->viewFormatValid = true;
}
}
if (sampledViewFormat == VK_FORMAT_UNDEFINED) {
MGLOG_E("ResolveSamplerDescriptor: no compatible sampled view for binding=%u ('%s') "
"textureId=%d imageFormat=%d numericDomain=%d",
@@ -307,8 +322,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static_cast<Int>(resource->format), static_cast<Int>(numericDomain));
return false;
}
// No reinterpretation requested: bind the depth-or-color aspect view the sync above
// already produced instead of re-entering GetOrCreateSampledImageView's sync path.
const VkImageView sampledImageView =
m_textureManager->GetOrCreateSampledImageView(*texture, sampledViewFormat);
sampledViewFormat == resource->format
? resource->sampledView
: m_textureManager->GetOrCreateSampledImageView(*texture, sampledViewFormat);
if (sampledImageView == VK_NULL_HANDLE) {
MGLOG_E("ResolveSamplerDescriptor: failed to resolve sampled view for binding=%u ('%s') "
"textureId=%d imageFormat=%d viewFormat=%d numericDomain=%d",
@@ -176,6 +176,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint16 textureParamsVersion = 0;
Bool forceNearestFiltering = false;
Bool valid = false;
// ResolveSampledImageViewFormat is pure in (image format, numeric domain), but a
// domain mismatch walks a ~184-entry format table. Memo the resolution per binding
// so a reinterpreted sampler pays that scan once, not once per draw.
VkFormat viewFormatSource = VK_FORMAT_UNDEFINED;
SamplerNumericDomain viewFormatDomain = SamplerNumericDomain::Unknown;
VkFormat viewFormat = VK_FORMAT_UNDEFINED;
Bool viewFormatValid = false;
};
mutable Vector<SamplerResolveMemo> m_samplerResolveMemo;
};
@@ -126,8 +126,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const Bool packedAttribute = attr.Type == DataType::Int2101010Rev ||
attr.Type == DataType::Uint2101010Rev;
const SizeT requiredAlignment = packedAttribute ? attribByteSize : GetComponentSize(attr.Type);
// For a client-memory array attr.Offset holds the raw client pointer, and the
// draw path re-uploads the data to a 16-aligned transient slice with attribute
// offset 0, so only the stride can violate Vulkan's fetch alignment there.
const Bool clientMemoryAttribute = attr.Buffer == nullptr;
if (conversion == VertexStreamConversion::None && requiredAlignment > 1 &&
((sourceStride % requiredAlignment) != 0 || (attr.Offset % requiredAlignment) != 0)) {
((sourceStride % requiredAlignment) != 0 ||
(!clientMemoryAttribute && (attr.Offset % requiredAlignment) != 0))) {
// GL accepts arbitrary byte strides and offsets. Core Vulkan vertex fetches do not
// unless VK_EXT_legacy_vertex_attributes is available, so deinterleave this one
// attribute into a tightly packed transient stream without changing its format.
@@ -1140,6 +1140,24 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return ok;
}
Bool VkTextureManager::NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const {
const auto it = m_textureResources.find(MakeTextureIdentity(&texture));
if (it == m_textureResources.end()) {
return true;
}
const TextureResource& resource = it->second;
if (resource.image == VK_NULL_HANDLE || resource.layout != VK_IMAGE_LAYOUT_GENERAL) {
return true;
}
// Mirror SyncTexture's cross-draw skip condition: any version drift means the sync
// path may upload or rebuild, both of which need the render pass ended first.
const auto* mipTexture = MG_State::GLState::AsMipmapTexture(&texture);
const Uint32 mipLevelCount = mipTexture != nullptr ? mipTexture->GetMipmapLevelCount() : 0u;
return resource.syncedContentVersion != texture.GetContentVersion() ||
resource.syncedTextureParamsVersion != texture.GetTextureParamsVersion() ||
resource.syncedMipLevelCount != mipLevelCount;
}
Bool VkTextureManager::TransitionImageLayout(VkCommandBuffer commandBuffer, VkImage image,
VkImageLayout& trackedLayout, VkImageLayout newLayout,
VkPipelineStageFlags srcStageMask, VkPipelineStageFlags dstStageMask,
@@ -1323,7 +1341,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
(aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 &&
(formatProperties.optimalTilingFeatures & VK_FORMAT_FEATURE_STORAGE_IMAGE_BIT) != 0;
VkImageCreateFlags imageCreateFlags = shapeInfo.imageFlags;
if (supportsStorageImage && IsMutableStorageImageFormat(format)) {
if (supportsStorageImage && IsMutableStorageImageFormat(format) &&
m_mutableFormatUnsupported.find(format) == m_mutableFormatUnsupported.end()) {
imageCreateFlags |= VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT;
}
@@ -1391,9 +1410,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
imageInfo.samples = resolvedSampleCount;
if (isMultisampleTexture || (imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
VkImageFormatProperties imageFormatProperties{};
const VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
m_physicalDevice, format, imageInfo.imageType, imageInfo.tiling, imageInfo.usage,
imageInfo.flags, &imageFormatProperties);
if (imageFormatResult != VK_SUCCESS && !isMultisampleTexture &&
(imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
// Losing reinterpreted views only degrades the formatless-image feature for
// this texture; failing creation would lose the texture entirely, so retry
// as a plain immutable-format image.
MGLOG_W("%s: mutable image format=%d is unsupported for textureId=%d; creating "
"without VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT (format reinterpretation "
"will be unavailable for it)",
__func__, static_cast<Int>(format), texture.GetExternalIndex());
// Remember the verdict so later syncs of same-format textures neither retry
// the probe nor flag-mismatch against this image and recreate it.
m_mutableFormatUnsupported.insert(format);
imageInfo.flags &= ~VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT;
imageCreateFlags = imageInfo.flags;
imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
m_physicalDevice, format, imageInfo.imageType, imageInfo.tiling, imageInfo.usage,
imageInfo.flags, &imageFormatProperties);
}
if (imageFormatResult != VK_SUCCESS ||
(isMultisampleTexture && (imageFormatProperties.sampleCounts & resolvedSampleCount) == 0)) {
MGLOG_D("%s: image flags=0x%x sampleCount=%d are unsupported for textureId=%d target=%s "
@@ -1921,6 +1958,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkFormat VkTextureManager::ResolveSampledImageViewFormat(VkFormat imageFormat,
SamplerNumericDomain numericDomain) {
// Depth/stencil images always sample through the existing depth-aspect sampledView.
// Combined formats (D24S8, D32FS8) are multi-numeric, so vkuFormatIsSampledFloat is
// false for them by design, yet their depth aspect reads as float in every GL depth
// texture mode; Vulkan also forbids reinterpreting them through color-class views.
// Integer domains keep the same view (pre-reinterpretation behavior for stencil-index
// style access) rather than failing the draw.
if (vkuFormatIsDepthOrStencil(imageFormat)) {
return imageFormat;
}
if (imageFormat == VK_FORMAT_UNDEFINED || numericDomain == SamplerNumericDomain::Unknown ||
FormatMatchesSamplerNumericDomain(imageFormat, numericDomain)) {
return imageFormat;
@@ -13,6 +13,7 @@
#include <MG_State/GLState/TextureState/TextureObject.h>
#include <vk_mem_alloc.h>
#include <unordered_map>
#include <unordered_set>
namespace MobileGL::MG_State::GLState {
class ITextureObject;
@@ -284,6 +285,11 @@ public:
VkImageLayout newLayout);
Bool TransitionTextureForSampling(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
Bool TransitionTextureForStorageImage(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
// Non-mutating probe for the per-draw storage-image fast path: true when preparing this
// texture as a storage image may need work that is illegal inside a render pass (resource
// creation, dirty-content upload, or a layout transition to GENERAL). Unknown state reports
// true - a false positive merely ends the render pass, a false negative would skip a barrier.
Bool NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const;
static VkImageAspectFlags ResolveSampledImageViewAspectMask(VkImageAspectFlags imageAspect);
static VkFormat ResolveSampledImageViewFormat(VkFormat imageFormat, SamplerNumericDomain numericDomain);
@@ -379,6 +385,9 @@ private:
TextureResource* resource = nullptr;
};
Vector<DrawSyncedTexture> m_drawSyncedThisDraw;
// Formats whose mutable-image probe failed on this device; their images are created
// without MUTABLE_FORMAT_BIT so repeat syncs neither re-probe nor flag-mismatch.
std::unordered_set<VkFormat> m_mutableFormatUnsupported;
std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects;
std::unordered_map<TextureIdentity, TextureResource, TextureIdentityHash> m_textureResources;
Vector<Vector<TextureResource>> m_deferredReleases;
@@ -27,6 +27,9 @@
#include <cstring>
#include <vulkan/utility/vk_format_utils.h>
#include <vulkan/vulkan_core.h>
#ifdef __ANDROID__
#include <sys/system_properties.h>
#endif
#if defined(__APPLE__)
#include <CoreGraphics/CoreGraphics.h>
@@ -1031,6 +1034,7 @@ void main() {
}
)";
static Uint32 ComputeFullMipLevelCount(const IntVec3& baseTexelSize) {
Int maxDimension = std::max<Int>(
baseTexelSize.x(),
@@ -1618,6 +1622,63 @@ void main() {
case VK_FORMAT_R32G32B32A32_SFLOAT:
Memcpy(rgba, source, sizeof(Float) * 4);
return true;
// Single- and dual-channel formats the reinterpretation feature makes common
// as readback sources (iterationRP custom images are R32F/R32UI-class).
// Missing channels take GL's defaults: 0 for GB, 1 for alpha.
case VK_FORMAT_R32_SFLOAT: {
Float value = 0.0f;
Memcpy(&value, source, sizeof(value));
rgba[0] = value;
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32G32_SFLOAT: {
Float values[2] = {0.0f, 0.0f};
Memcpy(values, source, sizeof(values));
rgba[0] = values[0];
rgba[1] = values[1];
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32_UINT: {
Uint32 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = static_cast<Float>(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R32_SINT: {
Int32 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = static_cast<Float>(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R16_SFLOAT: {
Uint16 value = 0;
Memcpy(&value, source, sizeof(value));
rgba[0] = MG_Util::DecodeHalfBitsToFloat(value);
rgba[1] = 0.0f;
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
}
case VK_FORMAT_R16G16_SFLOAT:
for (SizeT component = 0; component < 2; ++component) {
Uint16 value = 0;
Memcpy(&value, source + component * sizeof(value), sizeof(value));
rgba[component] = MG_Util::DecodeHalfBitsToFloat(value);
}
rgba[2] = 0.0f;
rgba[3] = 1.0f;
return true;
default:
return false;
}
@@ -2046,6 +2107,26 @@ void main() {
m_pipelineFactory = MakeUnique<PipelineFactory>(m_device, m_config);
MOBILEGL_ASSERT(m_pipelineFactory != nullptr, "PipelineFactory creation failed.");
{
// Qualcomm's pipeline compiler does not keep vertex positions invariant across
// the pipelines of a multi-pass depth-equality chain (even with the SPIR-V
// Invariant decoration), so a blended depth-writing prepass makes later
// equality-compare passes drop whole primitives (MC 26.3 improved-transparency
// clouds flicker black). Suppress depth writes on accumulation-blended pipelines
// there (see PipelineFactory::ShouldSuppressDepthWrite for the exact scope);
// MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE forces the quirk on or off on any
// driver.
const MG_Config::QuirkOverride quirkOverride =
MG_Config::Features.MagmaDisableBlendedDepthWriteQuirk;
const Bool suppressBlendedDepthWrite = PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(
quirkOverride, m_physicalDevice.properties.vendorID);
if (suppressBlendedDepthWrite) {
MGLOG_I("DirectVulkan: suppressing depth writes on accumulation-blended pipelines "
"(driver lacks cross-pipeline position invariance)%s",
quirkOverride == MG_Config::QuirkOverride::ForceOn ? " (forced on)" : "");
}
PipelineFactory::SetSuppressBlendedDepthWrite(suppressBlendedDepthWrite);
}
m_programFactory = MakeUnique<ProgramFactory>(m_device, m_config, maxProgramBindings,
m_shaderDrawParametersFeatureEnabled,
m_unformattedFloatStorageImagesEnabled);
@@ -2200,10 +2281,94 @@ void main() {
MGLOG_I("VulkanRenderer shut down completed");
}
// Scans the draw's index range from host-visible index bytes and returns the largest
// fetchable vertex index. Usable only when the draw's range is exactly its
// IndexBufferView (drawParams.indexRangeIsExactView). The view's byte offset is either
// an offset into the bound element-array buffer or, with no bound buffer, a raw client
// pointer. Primitive-restart sentinels are skipped so they cannot inflate the bound.
static Bool TryComputeMaxIndexFromHostBytes(const MG_State::GLState::VertexArrayObject& vao,
const IndexBufferView& indexView, Uint32& outMaxIndex) {
SizeT indexSize = 0;
switch (indexView.indexType) {
case GL_UNSIGNED_BYTE: indexSize = 1; break;
case GL_UNSIGNED_SHORT: indexSize = 2; break;
case GL_UNSIGNED_INT: indexSize = 4; break;
default: return false;
}
const Uint8* indexBytes = nullptr;
const auto& indexBufferShared = vao.GetIndexBufferBindingSlot().GetBoundObject();
if (indexBufferShared != nullptr) {
const SizeT bufferSize = indexBufferShared->GetSize();
if (indexBufferShared->MappedData() == nullptr || indexView.indexByteOffset > bufferSize ||
indexView.indexByteSize > bufferSize - indexView.indexByteOffset) {
return false;
}
indexBufferShared->SyncPersistentMappedRange();
indexBytes = indexBufferShared->MappedData() + indexView.indexByteOffset;
} else {
indexBytes = reinterpret_cast<const Uint8*>(indexView.indexByteOffset);
if (indexBytes == nullptr) {
return false;
}
}
const SizeT indexCount = indexView.indexByteSize / indexSize;
// The all-ones sentinel is only a restart marker when primitive restart is enabled;
// with restart off it is a legitimate index and excluding it would truncate the
// converted stream by exactly that vertex.
const Bool primitiveRestartActive =
MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::PrimitiveRestart) ||
MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::PrimitiveRestartFixedIndex);
const Uint32 restartSentinel = indexSize == 1 ? 0xFFu : indexSize == 2 ? 0xFFFFu : 0xFFFFFFFFu;
Uint32 maxIndex = 0;
Bool sawIndex = false;
for (SizeT i = 0; i < indexCount; ++i) {
Uint32 index = 0;
switch (indexSize) {
case 1: index = indexBytes[i]; break;
case 2: index = reinterpret_cast<const Uint16*>(indexBytes)[i]; break;
default: index = reinterpret_cast<const Uint32*>(indexBytes)[i]; break;
}
if (primitiveRestartActive && index == restartSentinel) {
continue;
}
maxIndex = std::max(maxIndex, index);
sawIndex = true;
}
if (!sawIndex) {
return false;
}
outMaxIndex = maxIndex;
return true;
}
Bool VulkanRenderer::UploadAndBindVertexBuffers(
VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ProgramFactory::VkProgramObject& programObj, const DrawCmdParam& drawParams,
Bool indexedDraw) {
const IndexBufferView* pIndexBufferView) {
const Bool indexedDraw = pIndexBufferView != nullptr;
// Exclusive upper bound on the vertex-stream elements this draw can fetch through
// vertex-rate bindings, or 0 when unbounded (indirect/multi draws). Computed lazily
// because the index scan is only worth doing when a conversion actually needs it.
SizeT drawElementBound = 0;
Bool drawElementBoundComputed = false;
auto resolveDrawElementBound = [&]() -> SizeT {
if (drawElementBoundComputed) {
return drawElementBound;
}
drawElementBoundComputed = true;
if (!indexedDraw) {
drawElementBound = static_cast<SizeT>(drawParams.firstVertex) + drawParams.vertexCount;
} else if (drawParams.indexRangeIsExactView) {
Uint32 maxIndex = 0;
if (TryComputeMaxIndexFromHostBytes(vao, *pIndexBufferView, maxIndex)) {
drawElementBound = static_cast<SizeT>(maxIndex) + 1 +
static_cast<SizeT>(std::max(drawParams.baseVertex, 0));
}
}
return drawElementBound;
};
// programObj is resolved once in SetupDraw and passed in; re-resolving it here would repeat
// the GetCurrentProgram + GetOrCreateProgram hash lookup every draw.
auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
@@ -2289,33 +2454,35 @@ void main() {
return false;
}
if (conversion != VertexInputStateFactory::VertexStreamConversion::None && indexedDraw) {
// The current indexed setup only carries indexCount, not the maximum effective
// index. Guessing a client-memory range here can truncate the converted stream.
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute "
"location=%u requires an indexed vertex range", location);
return false;
}
if (conversion != VertexInputStateFactory::VertexStreamConversion::None &&
drawParams.vertexCount == 0) {
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute "
"location=%u has an unknown vertex range", location);
return false;
}
const Uint32 lastVertex = drawParams.vertexCount > 0
? drawParams.firstVertex + drawParams.vertexCount - 1
: drawParams.firstVertex;
// Client arrays have no queryable size, so bound the upload by the draw's
// real fetch range. For indexed draws that means scanning the index bytes:
// the guessed vertexCount (indexCount + baseVertex) can both truncate draws
// whose max index exceeds their index count and over-read below it.
const SizeT clientElementBound = resolveDrawElementBound();
BufferSlice slice{};
Bool uploaded = false;
if (conversion == VertexInputStateFactory::VertexStreamConversion::None) {
const SizeT uploadSize = static_cast<SizeT>(lastVertex) * stride + elementSize;
const SizeT lastVertex =
clientElementBound > 0
? clientElementBound - 1
: (drawParams.vertexCount > 0
? static_cast<SizeT>(drawParams.firstVertex) + drawParams.vertexCount - 1
: static_cast<SizeT>(drawParams.firstVertex));
const SizeT uploadSize = lastVertex * stride + elementSize;
uploaded = m_bufferManager.UploadTransient(
BufferKind::Vertex, m_frameContext.GetCurrentFrameIndex(), clientData,
static_cast<VkDeviceSize>(uploadSize), 16, slice);
} else {
if (clientElementBound == 0) {
// Indirect/multi indexed draws have no CPU-visible index range and a
// client array has no size to fall back to; a guessed range could
// truncate the converted stream, so skip the draw loudly.
MGLOG_E("UploadAndBindVertexStreams skipped: converted client-memory attribute "
"location=%u has no computable vertex range", location);
return false;
}
uploaded = uploadConvertedStream(conversion, attr, clientData, stride, elementSize,
static_cast<SizeT>(lastVertex) + 1, slice);
clientElementBound, slice);
}
if (!uploaded) {
MOBILEGL_ASSERT(false,
@@ -2362,7 +2529,23 @@ void main() {
}
sourceBufferShared->SyncPersistentMappedRange();
const SizeT elementCount = 1 + (sourceSize - baseOffset - elementSize) / sourceStride;
const SizeT availableElementCount = 1 + (sourceSize - baseOffset - elementSize) / sourceStride;
const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped();
// Convert only what this draw can fetch instead of the whole buffer tail.
// Instance-rate bindings index by instance, not the vertex range, so they
// keep the tail. Indexed draws from cacheable buffers also keep the tail: a
// single cached whole-range conversion per frame is cheaper than a per-draw
// index scan. Persistent-mapped buffers are uncacheable and reconvert every
// draw, so for them the scan plus bounded conversion is the cheaper trade.
const Bool vertexRateBinding =
vertexInputState.bindings[binding].inputRate == VK_VERTEX_INPUT_RATE_VERTEX;
SizeT elementCount = availableElementCount;
if (vertexRateBinding && (!indexedDraw || !cacheable)) {
const SizeT elementBound = resolveDrawElementBound();
if (elementBound > 0) {
elementCount = std::min(elementCount, elementBound);
}
}
const ConvertedVertexStreamKey cacheKey{
.buffer = sourceBufferShared.get(),
.changeSerial = sourceBufferShared->GetChangeSerial(),
@@ -2375,12 +2558,18 @@ void main() {
.conversion = conversion,
};
const Bool cacheable = !sourceBufferShared->IsBackendPersistentMapped();
auto cached = cacheable ? m_convertedVertexStreams.find(cacheKey)
: m_convertedVertexStreams.end();
if (cached != m_convertedVertexStreams.end()) {
slice = cached->second;
} else {
Bool reusedCachedStream = false;
if (cacheable) {
const auto cached = m_convertedVertexStreams.find(cacheKey);
// A cached conversion covering at least this draw's range is a strict
// prefix match: converted streams are tightly packed from element 0.
if (cached != m_convertedVertexStreams.end() &&
cached->second.elementCount >= elementCount) {
slice = cached->second.slice;
reusedCachedStream = true;
}
}
if (!reusedCachedStream) {
const Uint8* sourceData = sourceBufferShared->MappedData() + baseOffset;
if (!uploadConvertedStream(conversion, attr, sourceData, sourceStride,
elementSize, elementCount, slice)) {
@@ -2389,7 +2578,8 @@ void main() {
return false;
}
if (cacheable) {
m_convertedVertexStreams.emplace(cacheKey, slice);
m_convertedVertexStreams[cacheKey] =
ConvertedVertexStream{slice, elementCount, sourceBufferShared};
}
}
vkBuffers[binding] = slice.buffer;
@@ -3274,6 +3464,7 @@ void main() {
.backStencilPassOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthPassOp),
.backStencilDepthFailOp = MG_Util::ConvertStencilOperationToVkEnum(backStencil.PassDepthFailOp),
.backStencilCompareOp = MG_Util::ConvertDepthTestFuncToVkEnum(backStencil.Func),
.fragmentReplacesDepth = programObj.fragmentReplacesDepth,
.stages = &programObj.stages,
.vertexInputState = pipelineVertexInputState
};
@@ -3361,6 +3552,17 @@ void main() {
(bufferMask.a() ? VK_COLOR_COMPONENT_A_BIT : 0u));
Bool effectiveBlendEnabled = blendEnabled;
MG_State::GLState::ITextureObject* colorAttachmentTexture = nullptr;
if (isDefaultDrawFbo && i < drawBuffers.size() &&
drawBuffers[i] == FramebufferAttachmentType::None) {
// The default framebuffer spans the same MAX_DRAW_BUFFERS slots as an FBO
// (slot 0 is the back buffer, or None after glDrawBuffer(GL_NONE); slots
// 1+ are always None). Discard writes and blend state for the None slots
// like the FBO path below does, so stale indexed blend state on phantom
// slots cannot leak into the pipeline - most notably into the blended
// depth-write quirk's accumulation scan.
attachmentColorWriteMask = 0;
effectiveBlendEnabled = false;
}
if (!isDefaultDrawFbo && i < drawBuffers.size()) {
const auto drawBuffer = drawBuffers[i];
colorAttachmentTexture = resolveCompleteColorAttachmentTexture(i);
@@ -3478,6 +3680,16 @@ void main() {
"disabling blending on attachments with this format (first hit: attachment %u textureId=%d program=%u)",
static_cast<Int>(colorAttachmentFormat), i, textureExternalIndex,
program.GetExternalIndex());
if (PipelineFactory::IsSuppressBlendedDepthWriteEnabled()) {
// With blending force-disabled the blended depth-write quirk can
// never fire for pipelines on this format, so a depth-equality
// chain that accumulates into it (MC 26.3 OIT depth_bounds on
// RGBA32F) keeps its depth writes and may flicker on this driver.
MGLOG_W("GetOrCreatePipeline: format=%d is not blendable, so the blended "
"depth-write quirk cannot apply to it; depth-equality chains "
"accumulating into this format may flicker",
static_cast<Int>(colorAttachmentFormat));
}
}
}
if (!blendSupportIt->second) {
@@ -3526,6 +3738,9 @@ void main() {
VkCommandBuffer commandBuffer,
const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj) {
if (!programObj.hasStorageImages) {
return true;
}
auto& storageTextures = m_storageImageTexturesScratch;
if (!m_uniformManager->CollectStorageImageTextures(program, programObj, storageTextures)) {
MGLOG_E("%s: failed to collect storage images for program=%u",
@@ -3536,6 +3751,24 @@ void main() {
return true;
}
// Steady-state fast path: when every collected texture is already resident in GENERAL
// with no pending clear and no dirty content, the loop below has nothing to record, so
// keep the render pass alive instead of splitting it on every storage-image draw (on
// tiled GPUs each split is a full tile load/store). GL makes cross-draw image-store
// coherence the app's job (glMemoryBarrier), so no implicit barrier is owed here.
Bool anyNeedsPreparation = false;
for (auto* texture : storageTextures) {
MOBILEGL_ASSERT(texture != nullptr, "%s: collected a null storage texture", __func__);
if (m_textureManager->NeedsStorageImagePreparation(*texture) ||
m_clearManager->HasPendingClear(texture)) {
anyNeedsPreparation = true;
break;
}
}
if (!anyNeedsPreparation) {
return true;
}
// Image uploads, deferred-clear materialization, and layout barriers are illegal inside
// a classic render pass. Do this before sampler preparation as well: a texture used by
// both a sampler and an image must stay in GENERAL, and both descriptors must name that
@@ -3545,7 +3778,6 @@ void main() {
}
for (auto* texture : storageTextures) {
MOBILEGL_ASSERT(texture != nullptr, "%s: collected a null storage texture", __func__);
if (!MaterializePendingClearForTexture(commandBuffer, *texture)) {
MGLOG_E("%s: failed to materialize pending clear for storage textureId=%d",
__func__, texture->GetExternalIndex());
@@ -3763,7 +3995,7 @@ void main() {
}
auto vtxUploadOk = UploadAndBindVertexBuffers(
frame.commandBuffer, vao, programObj, drawParams, pIndexBufferView != nullptr);
frame.commandBuffer, vao, programObj, drawParams, pIndexBufferView);
if (!vtxUploadOk) {
MGLOG_E("SetupDraw skipped: failed to upload vertex buffers");
return false;
@@ -5977,6 +6209,10 @@ void main() {
vertexRange.instanceCount = payload.params.instanceCount;
vertexRange.firstVertex = 0;
vertexRange.firstInstance = static_cast<Uint32>(payload.params.firstInstance);
vertexRange.baseVertex = payload.params.vertexOffset;
// Direct DrawElements fetches exactly the indices in its view, so vertex-stream
// conversion may bound its work by scanning them.
vertexRange.indexRangeIsExactView = true;
if (!SetupDraw(frame, payload.mode, DrawSetupAspect::IndexBuffer, vertexRange,
&payload.indexBufferView)) {
@@ -7007,7 +7243,10 @@ void main() {
// Match GL's robust buffer-fetch behavior where the Vulkan device supports it. This covers
// out-of-range fetches; arbitrary GL vertex strides/offsets still need the explicit tight
// repack in VertexInputStateFactory when they violate Vulkan's address-alignment rules.
deviceFeatures.robustBufferAccess = supportedDeviceFeatures.robustBufferAccess;
// MOBILEGL_DISABLE_ROBUST_BUFFER_ACCESS leaves it off to measure or dodge its GPU cost.
deviceFeatures.robustBufferAccess = MG_Config::Features.DisableRobustBufferAccess
? VK_FALSE
: supportedDeviceFeatures.robustBufferAccess;
deviceFeatures.geometryShader = supportedDeviceFeatures.geometryShader;
deviceFeatures.independentBlend = supportedDeviceFeatures.independentBlend;
m_independentBlendFeatureEnabled = deviceFeatures.independentBlend == VK_TRUE;
@@ -7037,6 +7276,15 @@ void main() {
if (m_unformattedFloatStorageImagesEnabled) {
deviceFeatures.shaderStorageImageReadWithoutFormat = VK_TRUE;
deviceFeatures.shaderStorageImageWriteWithoutFormat = VK_TRUE;
} else {
// Surface the degradation instead of failing silently: shader packs that bind a
// float storage image with a format different from its declaration (e.g.
// iterationRP) will render incorrectly on this device.
MGLOG_W("CreateLogicalDeviceAndQueues: shaderStorageImage*WithoutFormat unavailable "
"(read=%d write=%d); float storage-image format reinterpretation is disabled "
"and packs relying on it may misrender",
supportedDeviceFeatures.shaderStorageImageReadWithoutFormat,
supportedDeviceFeatures.shaderStorageImageWriteWithoutFormat);
}
deviceFeatures.drawIndirectFirstInstance = supportedDeviceFeatures.drawIndirectFirstInstance;
deviceFeatures.multiDrawIndirect = supportedDeviceFeatures.multiDrawIndirect;
@@ -51,6 +51,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 instanceCount = 1;
Uint32 firstVertex = 0;
Uint32 firstInstance = 0;
// Indexed-draw metadata for bounding vertex-stream conversion. baseVertex is the
// draw's base-vertex offset; indexRangeIsExactView is true only when the draw
// fetches exactly the indices its IndexBufferView describes (direct DrawElements;
// multi/indirect forms leave it false because the CPU cannot bound their ranges).
Int32 baseVertex = 0;
Bool indexRangeIsExactView = false;
};
struct DrawIndexedCmdParam {
@@ -488,7 +494,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}
};
UnorderedMap<ConvertedVertexStreamKey, BufferSlice, ConvertedVertexStreamKeyHash>
struct ConvertedVertexStream {
BufferSlice slice;
// Number of source elements the cached slice covers. A draw needing a prefix of
// this range reuses the slice (converted streams are tightly packed); a draw
// needing more reconverts and replaces the entry, so per (buffer, layout) a
// frame converts at most the largest range any draw asked for.
SizeT elementCount = 0;
// Pins the source buffer for the frame so its heap address cannot be reused by
// a new BufferObject while this pointer-keyed entry is alive.
SharedPtr<const MG_State::GLState::BufferObject> sourcePin;
};
UnorderedMap<ConvertedVertexStreamKey, ConvertedVertexStream, ConvertedVertexStreamKeyHash>
m_convertedVertexStreams;
void CreateInstance();
@@ -519,7 +536,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool UploadAndBindVertexBuffers(VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ProgramFactory::VkProgramObject& programObj,
const DrawCmdParam& drawParams, Bool indexedDraw);
const DrawCmdParam& drawParams,
const IndexBufferView* pIndexBufferView);
Bool UploadAndBindIndexBuffer(FrameContext::FrameData& frame,
const MG_State::GLState::VertexArrayObject& vao,
const IndexBufferView* pIndexBufferView = nullptr);
@@ -295,7 +295,13 @@ DECLARE_GL_FUNCTION_HEAD(void, VertexAttribDivisor, GLuint index, GLuint divisor
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedback, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedback, target, id)
DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacks, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacks, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacks, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacks, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(GLboolean, IsTransformFeedback, GLuint id) DECLARE_GL_FUNCTION_STUB_END(GLboolean, IsTransformFeedback, id)
// Transform feedback objects are not implemented, so no name is ever a live object. The shared
// stub returns (type)1, telling a probing caller that every id it invents already exists; GL_FALSE
// is both truthful and what the spec requires for a name that was never generated.
MOBILEGL_GL_API GLboolean glIsTransformFeedback(GLuint id) {
MGLOG_W("Stub function: %s(...)", __FUNCTION__);
return GL_FALSE;
}
DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedback)
DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedback) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedback)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetProgramBinary, GLuint program, GLsizei bufSize, GLsizei* length, GLenum* binaryFormat, void* binary) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetProgramBinary, program, bufSize, length, binaryFormat, binary)
@@ -418,7 +424,7 @@ DECLARE_GL_FUNCTION_HEAD(void, DrawRangeElementsBaseVertex, GLenum mode, GLuint
DECLARE_GL_FUNCTION_HEAD(void, DrawElementsInstancedBaseVertex, GLenum mode, GLsizei count, GLenum type, const void* indices, GLsizei instancecount, GLint basevertex) DECLARE_GL_FUNCTION_END_NO_RETURN(void, DrawElementsInstancedBaseVertex, mode, count, type, indices, instancecount, basevertex)
DECLARE_GL_FUNCTION_HEAD(void, FramebufferTexture, GLenum target, GLenum attachment, GLuint texture, GLint level) DECLARE_GL_FUNCTION_END_NO_RETURN(void, FramebufferTexture, target, attachment, texture, level)
DECLARE_GL_FUNCTION_STUB_HEAD(void, PrimitiveBoundingBox, GLfloat minX, GLfloat minY, GLfloat minZ, GLfloat minW, GLfloat maxX, GLfloat maxY, GLfloat maxZ, GLfloat maxW) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PrimitiveBoundingBox, minX, minY, minZ, minW, maxX, maxY, maxZ, maxW)
DECLARE_GL_FUNCTION_STUB_HEAD(GLenum, GetGraphicsResetStatus) DECLARE_GL_FUNCTION_STUB_END(GLenum, GetGraphicsResetStatus)
DECLARE_GL_FUNCTION_HEAD(GLenum, GetGraphicsResetStatus) DECLARE_GL_FUNCTION_END(GLenum, GetGraphicsResetStatus)
DECLARE_GL_FUNCTION_STUB_HEAD(void, ReadnPixels, GLint x, GLint y, GLsizei width, GLsizei height, GLenum format, GLenum type, GLsizei bufSize, void* data) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ReadnPixels, x, y, width, height, format, type, bufSize, data)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformfv, GLuint program, GLint location, GLsizei bufSize, GLfloat* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformfv, program, location, bufSize, params)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GetnUniformiv, GLuint program, GLint location, GLsizei bufSize, GLint* params) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GetnUniformiv, program, location, bufSize, params)
@@ -2583,7 +2589,10 @@ DECLARE_GL_FUNCTION_STUB_HEAD(void, TransformFeedbackStreamAttribsNV, GLsizei co
DECLARE_GL_FUNCTION_STUB_HEAD(void, BindTransformFeedbackNV, GLenum target, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, BindTransformFeedbackNV, target, id)
DECLARE_GL_FUNCTION_STUB_HEAD(void, DeleteTransformFeedbacksNV, GLsizei n, const GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DeleteTransformFeedbacksNV, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(void, GenTransformFeedbacksNV, GLsizei n, GLuint* ids) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, GenTransformFeedbacksNV, n, ids)
DECLARE_GL_FUNCTION_STUB_HEAD(GLboolean, IsTransformFeedbackNV, GLuint id) DECLARE_GL_FUNCTION_STUB_END(GLboolean, IsTransformFeedbackNV, id)
MOBILEGL_GL_API GLboolean glIsTransformFeedbackNV(GLuint id) {
MGLOG_W("Stub function: %s(...)", __FUNCTION__);
return GL_FALSE;
}
DECLARE_GL_FUNCTION_STUB_HEAD(void, PauseTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, PauseTransformFeedbackNV, )
DECLARE_GL_FUNCTION_STUB_HEAD(void, ResumeTransformFeedbackNV, void) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, ResumeTransformFeedbackNV, )
DECLARE_GL_FUNCTION_STUB_HEAD(void, DrawTransformFeedbackNV, GLenum mode, GLuint id) DECLARE_GL_FUNCTION_STUB_END_NO_RETURN(void, DrawTransformFeedbackNV, mode, id)
@@ -1996,4 +1996,12 @@ namespace MobileGL::MG_Impl::GLImpl {
}
return MG_Util::ConvertErrorCodeToGLEnum(error->get()->code);
}
GLenum GetGraphicsResetStatus() {
// MobileGL does not implement robustness reset notification, so report GL_NO_ERROR
// ("no reset detected"). Returning the generic stub's (GLenum)1 makes dEQP read a lost
// device after every case (gl3cTestPackages.cpp:121) and, under the default
// --deqp-terminate-on-device-lost=enable, tear the whole CTS run down.
return GL_NO_ERROR;
}
} // namespace MobileGL::MG_Impl::GLImpl
@@ -21,4 +21,5 @@ namespace MobileGL::MG_Impl::GLImpl {
void GetIntegeri_v(GLenum target, GLuint index, GLint* data);
void GetInteger64i_v(GLenum target, GLuint index, GLint64* data);
GLenum GetError();
GLenum GetGraphicsResetStatus();
} // namespace MobileGL::MG_Impl::GLImpl
+131 -17
View File
@@ -350,6 +350,10 @@ namespace MobileGL::MG_Impl::GLImpl {
texture.AllocateStorage(uploadTarget, level, {levelTexelSize, levelByteSize});
texture.MarkStorageDirty(uploadTarget, level, false);
}
// glGenerateMipmap defines exactly levels 0..requiredLevelCount-1. AllocateStorage only
// grows, so a previously longer chain (a bigger base image before respecification) would
// otherwise keep a tail of stale levels here and read as incomplete.
texture.TruncateMipmapLevels(uploadTarget, requiredLevelCount);
// Mip generation grows/regenerates the level set on the GPU without marking any CPU
// level dirty (MarkStorageDirty(...,false) above). Bump the content version so the
// backend re-syncs: a cached sampled VkImageView built for the pre-generate level
@@ -429,8 +433,10 @@ namespace MobileGL::MG_Impl::GLImpl {
const Int maxSamples = GetMaxSupportedTextureSamples(textureInternalFormat);
if (samples > maxSamples) {
// GL specifies INVALID_OPERATION - not INVALID_VALUE - when the sample count
// exceeds what the format supports, and the native Adreno driver agrees.
MG_State::pGLContext->RecordError(
ErrorCode::InvalidValue,
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>(
"MG_Impl/GLImpl", caller,
std::format("Sample count {} exceeds the supported maximum {} for this texture format.",
@@ -454,8 +460,56 @@ namespace MobileGL::MG_Impl::GLImpl {
textureObject->SetSamples(samples);
textureObject->SetFixedSampleLocations(fixedsamplelocations == GL_TRUE);
textureMipmapObject->AllocateStorage(textureUploadTarget, 0, {{width, height, depth}, 0});
// Multisample textures are single-level by definition, so a name that previously held a
// mip chain must not keep its tail now that AllocateStorage only grows.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, 1);
textureMipmapObject->MarkStorageDirty(textureUploadTarget, 0, false);
}
// Redefining level 0 of a texture that already had a base image drops the rest of the chain,
// which is exactly what AllocateLevel used to do implicitly for every level. Keeping that
// behaviour for level 0 - and only for level 0 - is what makes the grow-only change safe:
// any level-0 respecification leaves the chain in precisely the state it would have had
// before, while an upload to level N no longer destroys the levels beneath it.
//
// Why it has to be *every* level-0 respecification and not just a size change: Minecraft's
// Mipmap Levels setting rebuilds the block atlas at the SAME dimensions with a different
// level count. A size-only test would leave the old tail in place, and because Mojang
// terminates its chains with a 0x0 level the result is the zero-then-nonzero pattern that
// IsComplete() rejects (TextureObject.cpp) - whereupon DirectGLES skips syncing the texture
// entirely (Managers.cpp) and the atlas samples black.
//
// The "already has a base image" test is what lets the fix work at all: a level that was
// never written reads back as {0,0,0}, so building a chain top-down - upload level N first,
// then level 0 - must not discard the levels just uploaded. That ordering is what
// KHR-GL33.texture_repeat_mode does.
// Scoped to the respecified upload target only, which is what AllocateLevel already did.
// Cube maps keep six independent chains while reporting a single level count (face +X), so
// respecifying a face other than +X can leave the count longer than that face - but that
// asymmetry predates this change and widening the truncation to all six faces would destroy
// mip data for faces the application never touched. Left alone deliberately.
void DiscardMipmapChainOnBaseRespecification(MG_State::GLState::TextureObjectMipmap* texture,
TextureUploadTarget uploadTarget, Uint level) {
if (level != 0) return;
const IntVec3 existingBaseSize = texture->GetMipmapTexelSize(uploadTarget, 0);
const Bool hasExistingBaseImage =
existingBaseSize.x() > 0 && existingBaseSize.y() > 0 && existingBaseSize.z() > 0;
if (!hasExistingBaseImage) return;
texture->TruncateMipmapLevels(uploadTarget, 1);
}
// Compressed texture upload is not implemented yet. GL_NUM_COMPRESSED_TEXTURE_FORMATS
// reports 0, so every compressed internalformat is by definition unsupported and
// GL_INVALID_ENUM is the specified error - unlike THROW_UNIMPL_EXCEPTION, which unwinds
// a C++ exception through the C GL ABI and takes the process down.
void RecordUnsupportedCompressedFormat(const char* caller) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidEnum,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", caller,
"Compressed texture formats are not supported."));
}
} // namespace
const SharedPtr<MG_State::GLState::ITextureObject>& GetTextureObjectByName(GLuint texture, const char* caller) {
@@ -493,8 +547,15 @@ namespace MobileGL::MG_Impl::GLImpl {
}
auto mipmapTexture = std::static_pointer_cast<MG_State::GLState::TextureObjectMipmap>(textureObject);
if (level < 0 || static_cast<Uint>(level) >= mipmapTexture->GetMipmapLevelCount()) {
if (level < 0) {
RecordClearTextureError(caller, ErrorCode::InvalidValue,
std::format("Texture level {} is negative.", level));
return nullptr;
}
// ARB_clear_texture: clearing an image that was never defined by TexImage*/
// TexStorage* is INVALID_OPERATION, not INVALID_VALUE.
if (static_cast<Uint>(level) >= mipmapTexture->GetMipmapLevelCount()) {
RecordClearTextureError(caller, ErrorCode::InvalidOperation,
std::format("Texture level {} is not defined.", level));
return nullptr;
}
@@ -536,6 +597,12 @@ namespace MobileGL::MG_Impl::GLImpl {
return true;
}
// Writes the clear into the CPU shadow and marks the whole level dirty, exactly like
// TexSubImage*_State does. Shared limitation of the level-granular shadow sync: the
// shadow does not reflect GPU-side writes (FBO rendering, imageStore), so a PARTIAL
// clear of a GPU-written level re-uploads stale shadow bytes outside the region on
// the next sync. Full-level clears (glClearTexImage, or a sub-clear covering the
// level) rewrite the entire shadow and are always correct.
Bool ClearMipmapRegion(const SharedPtr<MG_State::GLState::TextureObjectMipmap>& textureObject,
TextureUploadTarget uploadTarget, GLint level,
GLint xoffset, GLint yoffset, GLint zoffset,
@@ -1657,6 +1724,21 @@ namespace MobileGL::MG_Impl::GLImpl {
return;
}
// RGTC is a 2D-only compression scheme, so a 3D target rejects it. This has to be tested on
// the raw enum: the RGTC formats resolve to plain R8/RG8/SNORM storage on the way in (see
// GLToMG's TextureEnumConverter), so once the internal format is converted there is nothing
// left to distinguish them from an ordinary one- or two-channel upload.
if ((textureUploadTarget == TextureUploadTarget::Texture3D ||
textureUploadTarget == TextureUploadTarget::ProxyTexture3D) &&
(internalformat == GL_COMPRESSED_RED_RGTC1 || internalformat == GL_COMPRESSED_SIGNED_RED_RGTC1 ||
internalformat == GL_COMPRESSED_RG_RGTC2 || internalformat == GL_COMPRESSED_SIGNED_RG_RGTC2)) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"RGTC compressed formats are invalid for 3D texture targets"));
return;
}
// TODO: GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the
// GL_PIXEL_UNPACK_BUFFER target and the buffer object's data store is currently mapped.
// GL_INVALID_OPERATION is generated if a non-zero buffer object name is bound to the GL_PIXEL_UNPACK_BUFFER
@@ -1714,6 +1796,7 @@ namespace MobileGL::MG_Impl::GLImpl {
if (isProxy) {
MGLOG_D("%s: isProxy = true, not allocating", __func__);
} else {
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, height, depth}, internalBytes});
}
@@ -1841,6 +1924,7 @@ namespace MobileGL::MG_Impl::GLImpl {
MGLOG_D("%s: isProxy = true, not allocating", __func__);
} else {
MGLOG_D("%s: Allocating %d bytes at mip %d", __func__, internalBytes, level);
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level,
{{width, height, 1}, internalBytes});
}
@@ -1929,6 +2013,7 @@ namespace MobileGL::MG_Impl::GLImpl {
"Texture object here should always be an object with mipmap");
auto textureMipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
if (!isProxy) {
DiscardMipmapChainOnBaseRespecification(textureMipmapObject, textureUploadTarget, level);
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{width, 1, 1}, internalBytes});
}
@@ -2581,7 +2666,13 @@ namespace MobileGL::MG_Impl::GLImpl {
}
void GetCompressedTexImage_State(GLenum target, GLint level, void* img) {
// TODO: implement
// TODO: implement compressed readback. Reporting success while writing nothing hands
// the caller stale memory with GL_NO_ERROR; no texture can be compressed yet, and GL
// specifies GL_INVALID_OPERATION when the bound level is not compressed.
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"Texture level is not stored in a compressed format."));
}
void GenTextures_State(GLsizei n, GLuint* textures) {
@@ -2768,20 +2859,20 @@ namespace MobileGL::MG_Impl::GLImpl {
void CompressedTexSubImage3D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLint zoffset,
GLsizei width, GLsizei height, GLsizei depth, GLenum format, GLsizei imageSize,
const void* data) {
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload - see CompressedTexImage2D_State.
RecordUnsupportedCompressedFormat(__func__);
}
void CompressedTexSubImage2D_State(GLenum target, GLint level, GLint xoffset, GLint yoffset, GLsizei width,
GLsizei height, GLenum format, GLsizei imageSize, const void* data) {
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload - see CompressedTexImage2D_State.
RecordUnsupportedCompressedFormat(__func__);
}
void CompressedTexSubImage1D_State(GLenum target, GLint level, GLint xoffset, GLsizei width, GLenum format,
GLsizei imageSize, const void* data) {
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload - see CompressedTexImage2D_State.
RecordUnsupportedCompressedFormat(__func__);
}
void CompressedTexImage3D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height,
@@ -2791,8 +2882,8 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload - see CompressedTexImage2D_State.
RecordUnsupportedCompressedFormat(__func__);
}
void CompressedTexImage2D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLsizei height,
@@ -2802,8 +2893,11 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload. Until then report the spec error for an
// unsupported compressed format rather than throwing - a C++ exception unwinding
// through the C GL ABI is a hard crash for the caller, while GL_INVALID_ENUM is
// exactly what GL_NUM_COMPRESSED_TEXTURE_FORMATS == 0 promises.
RecordUnsupportedCompressedFormat(__func__);
}
void CompressedTexImage1D_State(GLenum target, GLint level, GLenum internalformat, GLsizei width, GLint border,
@@ -2813,8 +2907,8 @@ namespace MobileGL::MG_Impl::GLImpl {
auto& textureObject = GetTextureObjectByTarget(textureUploadTarget, textureTarget);
if (!ValidateTextureMutable(textureObject, __func__)) return;
// TODO: implement
THROW_UNIMPL_EXCEPTION;
// TODO: implement compressed upload - see CompressedTexImage2D_State.
RecordUnsupportedCompressedFormat(__func__);
}
void BindTexture_State(GLenum target, GLuint texture) {
@@ -3160,6 +3254,9 @@ namespace MobileGL::MG_Impl::GLImpl {
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, 1, 1}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
}
// Immutable storage defines exactly `levels` levels; AllocateStorage only grows, so a
// longer pre-existing chain has to be dropped explicitly.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels));
}
@@ -3212,6 +3309,8 @@ namespace MobileGL::MG_Impl::GLImpl {
textureMipmapObject->AllocateStorage(textureUploadTarget, level, {{levelWidth, levelHeight, 1}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
}
// See TextureStorage1D.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels));
}
@@ -3264,6 +3363,8 @@ namespace MobileGL::MG_Impl::GLImpl {
{{levelWidth, levelHeight, levelDepth}, byteSize});
textureMipmapObject->MarkStorageDirty(textureUploadTarget, level, false);
}
// See TextureStorage1D.
textureMipmapObject->TruncateMipmapLevels(textureUploadTarget, static_cast<Uint>(levels));
textureObject->SetImmutableLevels(static_cast<Uint>(levels));
}
@@ -4025,8 +4126,21 @@ namespace MobileGL::MG_Impl::GLImpl {
void CopyTextureSubImage2D(GLuint texture, GLint level, GLint xoffset, GLint yoffset, GLint x, GLint y,
GLsizei width, GLsizei height) {
auto textureObject = GetTextureObjectByName(texture, __func__);
WithTemporarilyBoundNamedTexture(textureObject, [&](GLenum target) {
CopyTexSubImage2D_Backend(target, level, xoffset, yoffset, x, y, width, height);
if (!textureObject) return;
// GL 4.6 sec. 8.8: the 2D form only accepts these effective targets; cube maps must
// go through CopyTextureSubImage3D with the face as a layer.
const auto target = textureObject->GetTarget();
if (target != TextureTarget::Texture2D && target != TextureTarget::Texture1DArray &&
target != TextureTarget::TextureRectangle) {
MG_State::pGLContext->RecordError(
ErrorCode::InvalidOperation,
MakeUnique<GenericErrorInfo>("MG_Impl/GLImpl", __func__,
"CopyTextureSubImage2D requires a 2D, 1D-array, or "
"rectangle texture."));
return;
}
WithTemporarilyBoundNamedTexture(textureObject, [&](GLenum glTarget) {
CopyTexSubImage2D_Backend(glTarget, level, xoffset, yoffset, x, y, width, height);
});
}
@@ -16,17 +16,32 @@ namespace MobileGL {
}
void MipmapStorage::AllocateLevel(Uint level, MipmapInput input) {
m_data.reserve(std::bit_ceil(level + 1));
m_data.resize(level + 1);
m_texelSizes.reserve(std::bit_ceil(level + 1));
m_texelSizes.resize(level + 1);
m_texelSizes[level] = input.texelSize;
m_isDirty.resize(level + 1, false);
// Grow only. GL respecifies exactly the level it is handed, so allocating level 0
// must not disturb the levels above it - but resize() shrinks as readily as it
// grows, so this used to truncate the whole chain to a single level. Callers that
// genuinely redefine the complete level set say so with TruncateToLevelCount.
const SizeT requiredLevelCount = static_cast<SizeT>(level) + 1;
if (m_data.size() < requiredLevelCount) {
m_data.reserve(std::bit_ceil(requiredLevelCount));
m_data.resize(requiredLevelCount);
m_texelSizes.reserve(std::bit_ceil(requiredLevelCount));
m_texelSizes.resize(requiredLevelCount);
m_isDirty.resize(requiredLevelCount, false);
}
m_texelSizes[level] = input.texelSize;
auto& data = m_data[level];
data.resize(input.byteSize, 0);
}
void MipmapStorage::TruncateToLevelCount(SizeT levelCount) {
if (levelCount >= m_data.size()) return;
m_data.resize(levelCount);
m_texelSizes.resize(levelCount);
m_isDirty.resize(levelCount);
}
void MipmapStorage::UpdateSubData(Uint level, DataPtr input) {
auto& targetData = m_data;
MOBILEGL_ASSERT(level < targetData.size(), "UpdateSubData: level out of range");
@@ -55,6 +70,7 @@ namespace MobileGL {
}
SizeT MipmapStorage::GetByteSize(Uint level) const {
if (level >= m_data.size()) return 0;
return m_data[level].size();
}
@@ -19,6 +19,10 @@ namespace MobileGL {
public:
SizeT GetLevelCount() const;
void AllocateLevel(Uint level, MipmapInput input);
// Discard every level at or above levelCount. AllocateLevel never shrinks, so this
// is the only way a chain gets shorter - use it where the caller defines the whole
// level set (glTexStorage*, mip regeneration, atlas respecification).
void TruncateToLevelCount(SizeT levelCount);
void UpdateSubData(Uint level, DataPtr input);
void* MapData(Uint level);
IntVec3 GetTexelSize(Uint level) const;
@@ -29,6 +29,14 @@ namespace MobileGL {
m_storage[targetIndex].AllocateLevel(level, input);
}
// Per-target, like AllocateLevel: cube-map faces are respecified independently, so
// truncating one face must not disturb the others.
void TruncateToLevelCount(Uint targetIndex, SizeT levelCount) {
MOBILEGL_ASSERT(targetIndex < TargetCount, "TruncateToLevelCount: target invalid");
m_storage[targetIndex].TruncateToLevelCount(levelCount);
}
void UpdateSubData(Uint targetIndex, Uint level, DataPtr input) {
MOBILEGL_ASSERT(targetIndex < TargetCount, "UpdateSubData: target invalid");
m_storage[targetIndex].UpdateSubData(level, input);
@@ -271,6 +271,10 @@ namespace MobileGL {
m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
}
void TextureObjectWithOneMipmap::TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) {
m_textureStorage.TruncateToLevelCount(GetIndexOfTextureUploadTarget(uploadTarget), levelCount);
}
void TextureObjectWithOneMipmap::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel,
DataPtr input) {
m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
@@ -134,6 +134,10 @@ namespace MobileGL::MG_State::GLState {
virtual const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const = 0;
virtual const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const = 0;
virtual void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) = 0;
// AllocateStorage only ever grows the chain. Callers that define the complete level set -
// glTexStorage*, mip regeneration, or a level-0 respecification at a new size - drop the
// leftovers explicitly, so a stale tail can never make the texture silently incomplete.
virtual void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) = 0;
virtual void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) = 0;
virtual void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) = 0;
virtual void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty = true) = 0;
@@ -175,6 +179,7 @@ namespace MobileGL::MG_State::GLState {
const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override;
const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override;
void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override;
void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) override;
void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override;
void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override;
void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, Bool dirty) override;
@@ -31,6 +31,10 @@ namespace MobileGL {
m_textureStorage.AllocateLevel(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
}
void TextureObject2DCube::TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) {
m_textureStorage.TruncateToLevelCount(GetIndexOfTextureUploadTarget(uploadTarget), levelCount);
}
void TextureObject2DCube::UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel,
DataPtr input) {
m_textureStorage.UpdateSubData(GetIndexOfTextureUploadTarget(uploadTarget), mipmapLevel, input);
@@ -22,6 +22,7 @@ namespace MobileGL {
const IntVec3 GetMipmapTexelSize(TextureUploadTarget target, Uint mipmapLevel) const override;
const SizeT GetMipmapByteSize(TextureUploadTarget target, Uint mipmapLevel) const override;
void AllocateStorage(TextureUploadTarget uploadTarget, Uint mipmapLevel, MipmapInput input) override;
void TruncateMipmapLevels(TextureUploadTarget uploadTarget, Uint levelCount) override;
void UpdateMipmapSubData(TextureUploadTarget uploadTarget, Uint mipmapLevel, DataPtr input) override;
void* MapMipmapData(TextureUploadTarget uploadTarget, Uint mipmapLevel) override;
void MarkStorageDirty(TextureUploadTarget uploadTarget, Uint mipmapLevel, bool dirty) override;
+2
View File
@@ -72,6 +72,8 @@ add_subdirectory(Texture)
add_subdirectory(VertexArray)
add_subdirectory(Program)
add_subdirectory(Query)
add_subdirectory(Pipeline)
add_subdirectory(ShaderTranspiler)
if (ENABLE_INTEGRATION_TESTS)
add_subdirectory(Backend/DirectVulkan)
endif()
+27
View File
@@ -0,0 +1,27 @@
cmake_minimum_required(VERSION 3.14)
add_executable(
PipelineQuirkTest
PipelineQuirkTest.cpp
)
target_include_directories(PipelineQuirkTest PRIVATE
${MGL_ROOT}/include
${MGL_ROOT}/MobileGL
${MGL_ROOT}/3rdparty/xxHash
${MGL_ROOT}/3rdparty/Vulkan-Headers/include
${MGL_ROOT}/3rdparty/SPIRV-Reflect
)
target_link_libraries(
PipelineQuirkTest PRIVATE
GTest::gtest_main
${LINK_LIBRARIES}
)
if (MSVC)
target_compile_options(PipelineQuirkTest PRIVATE /Zc:preprocessor)
endif()
include(GoogleTest)
gtest_discover_tests(PipelineQuirkTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
@@ -0,0 +1,459 @@
// MobileGL - MobileGL/MG_Test/Pipeline/PipelineQuirkTest.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <Config.h>
#include <MG_Backend/DirectVulkan/Renderer/PipelineFactory.h>
#include <MG_Backend/DirectVulkan/Renderer/ProgramFactory.h>
using namespace MobileGL;
using MobileGL::MG_Backend::DirectVulkan::PipelineFactory;
using MobileGL::MG_Backend::DirectVulkan::ProgramFactory;
using MobileGL::MG_Config::QuirkOverride;
namespace {
constexpr Uint32 kVendorIdQualcomm = 0x5143;
constexpr Uint32 kVendorIdArm = 0x13B5;
constexpr VkColorComponentFlags kFullColorWriteMask =
VK_COLOR_COMPONENT_R_BIT | VK_COLOR_COMPONENT_G_BIT |
VK_COLOR_COMPONENT_B_BIT | VK_COLOR_COMPONENT_A_BIT;
// Builds non-separate blend state: the alpha channel repeats the color factors/op, which
// is what glBlendFunc/glBlendEquation (as opposed to their *Separate forms) produce.
// ShouldSuppressDepthWrite deliberately decides on the color channel alone, so these
// cases cover its whole input space; SeparateAlphaAccumulationIsNotStripped below pins
// the separate-alpha contract.
VkPipelineColorBlendAttachmentState MakeBlendAttachment(Bool blendEnable,
VkBlendFactor srcColor,
VkBlendFactor dstColor,
VkBlendOp colorOp,
VkColorComponentFlags colorWriteMask) {
VkPipelineColorBlendAttachmentState attachment{};
attachment.blendEnable = blendEnable ? VK_TRUE : VK_FALSE;
attachment.srcColorBlendFactor = srcColor;
attachment.dstColorBlendFactor = dstColor;
attachment.colorBlendOp = colorOp;
attachment.srcAlphaBlendFactor = srcColor;
attachment.dstAlphaBlendFactor = dstColor;
attachment.alphaBlendOp = colorOp;
attachment.colorWriteMask = colorWriteMask;
return attachment;
}
// glslangValidator -V output for:
// #version 450
// layout(location = 0) out vec4 outColor;
// void main() { outColor = vec4(1.0); gl_FragDepth = 0.5; }
// Assigning gl_FragDepth makes glslang emit OpExecutionMode ... DepthReplacing.
constexpr Uint32 kFragDepthWriterSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000000fu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000004u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x00000009u, 0x0000000du, 0x00030010u,
0x00000004u, 0x00000007u, 0x00030010u, 0x00000004u, 0x0000000cu, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00050005u, 0x00000009u, 0x4374756fu, 0x726f6c6fu, 0x00000000u, 0x00060005u,
0x0000000du, 0x465f6c67u, 0x44676172u, 0x68747065u, 0x00000000u, 0x00040047u,
0x00000009u, 0x0000001eu, 0x00000000u, 0x00040047u, 0x0000000du, 0x0000000bu,
0x00000016u, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040020u, 0x00000008u, 0x00000003u, 0x00000007u, 0x0004003bu,
0x00000008u, 0x00000009u, 0x00000003u, 0x0004002bu, 0x00000006u, 0x0000000au,
0x3f800000u, 0x0007002cu, 0x00000007u, 0x0000000bu, 0x0000000au, 0x0000000au,
0x0000000au, 0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x00000006u,
0x0004003bu, 0x0000000cu, 0x0000000du, 0x00000003u, 0x0004002bu, 0x00000006u,
0x0000000eu, 0x3f000000u, 0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u,
0x00000003u, 0x000200f8u, 0x00000005u, 0x0003003eu, 0x00000009u, 0x0000000bu,
0x0003003eu, 0x0000000du, 0x0000000eu, 0x000100fdu, 0x00010038u,
};
// Same shader without the gl_FragDepth assignment.
constexpr Uint32 kPlainFragmentSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000000cu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0006000fu, 0x00000004u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x00000009u, 0x00030010u, 0x00000004u,
0x00000007u, 0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u,
0x6e69616du, 0x00000000u, 0x00050005u, 0x00000009u, 0x4374756fu, 0x726f6c6fu,
0x00000000u, 0x00040047u, 0x00000009u, 0x0000001eu, 0x00000000u, 0x00020013u,
0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u, 0x00000006u,
0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u, 0x00040020u,
0x00000008u, 0x00000003u, 0x00000007u, 0x0004003bu, 0x00000008u, 0x00000009u,
0x00000003u, 0x0004002bu, 0x00000006u, 0x0000000au, 0x3f800000u, 0x0007002cu,
0x00000007u, 0x0000000bu, 0x0000000au, 0x0000000au, 0x0000000au, 0x0000000au,
0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u,
0x00000005u, 0x0003003eu, 0x00000009u, 0x0000000bu, 0x000100fdu, 0x00010038u,
};
// glslangValidator -V output for a vertex shader reading gl_InstanceIndex:
// #version 450
// layout(location = 0) in vec4 inPos;
// void main() { gl_Position = inPos + vec4(float(gl_InstanceIndex)); }
constexpr Uint32 kInstanceIndexVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000001bu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0008000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00000014u,
0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du,
0x00000000u, 0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u,
0x00000000u, 0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu,
0x006e6f69u, 0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu,
0x657a6953u, 0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u,
0x4470696cu, 0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u,
0x435f6c67u, 0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du,
0x00000000u, 0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00070005u,
0x00000014u, 0x495f6c67u, 0x6174736eu, 0x4965636eu, 0x7865646eu, 0x00000000u,
0x00030047u, 0x0000000bu, 0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u,
0x0000000bu, 0x00000000u, 0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu,
0x00000001u, 0x00050048u, 0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u,
0x00050048u, 0x0000000bu, 0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u,
0x00000011u, 0x0000001eu, 0x00000000u, 0x00040047u, 0x00000014u, 0x0000000bu,
0x0000002bu, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu,
0x00000008u, 0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u,
0x00000009u, 0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au,
0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu,
0x0000000cu, 0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u,
0x00000001u, 0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u,
0x00000010u, 0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u,
0x00000001u, 0x00040020u, 0x00000013u, 0x00000001u, 0x0000000eu, 0x0004003bu,
0x00000013u, 0x00000014u, 0x00000001u, 0x00040020u, 0x00000019u, 0x00000003u,
0x00000007u, 0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u,
0x000200f8u, 0x00000005u, 0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u,
0x0004003du, 0x0000000eu, 0x00000015u, 0x00000014u, 0x0004006fu, 0x00000006u,
0x00000016u, 0x00000015u, 0x00070050u, 0x00000007u, 0x00000017u, 0x00000016u,
0x00000016u, 0x00000016u, 0x00000016u, 0x00050081u, 0x00000007u, 0x00000018u,
0x00000012u, 0x00000017u, 0x00050041u, 0x00000019u, 0x0000001au, 0x0000000du,
0x0000000fu, 0x0003003eu, 0x0000001au, 0x00000018u, 0x000100fdu, 0x00010038u,
};
// Same, but reading gl_VertexIndex instead: a DIFFERENT input builtin. glslang emits
// this for GL's gl_VertexID, so nearly every real vertex shader has one - it is what
// separates "declares some builtin" from "declares the InstanceIndex builtin".
constexpr Uint32 kVertexIndexVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x0000001bu, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0008000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00000014u,
0x00030003u, 0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du,
0x00000000u, 0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u,
0x00000000u, 0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu,
0x006e6f69u, 0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu,
0x657a6953u, 0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u,
0x4470696cu, 0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u,
0x435f6c67u, 0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du,
0x00000000u, 0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00060005u,
0x00000014u, 0x565f6c67u, 0x65747265u, 0x646e4978u, 0x00007865u, 0x00030047u,
0x0000000bu, 0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu,
0x00000000u, 0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu, 0x00000001u,
0x00050048u, 0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u, 0x00050048u,
0x0000000bu, 0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u, 0x00000011u,
0x0000001eu, 0x00000000u, 0x00040047u, 0x00000014u, 0x0000000bu, 0x0000002au,
0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u,
0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u,
0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu, 0x00000008u,
0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u, 0x00000009u,
0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au, 0x0000000au,
0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu, 0x0000000cu,
0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u, 0x00000001u,
0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u, 0x00000010u,
0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u, 0x00000001u,
0x00040020u, 0x00000013u, 0x00000001u, 0x0000000eu, 0x0004003bu, 0x00000013u,
0x00000014u, 0x00000001u, 0x00040020u, 0x00000019u, 0x00000003u, 0x00000007u,
0x00050036u, 0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u,
0x00000005u, 0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u, 0x0004003du,
0x0000000eu, 0x00000015u, 0x00000014u, 0x0004006fu, 0x00000006u, 0x00000016u,
0x00000015u, 0x00070050u, 0x00000007u, 0x00000017u, 0x00000016u, 0x00000016u,
0x00000016u, 0x00000016u, 0x00050081u, 0x00000007u, 0x00000018u, 0x00000012u,
0x00000017u, 0x00050041u, 0x00000019u, 0x0000001au, 0x0000000du, 0x0000000fu,
0x0003003eu, 0x0000001au, 0x00000018u, 0x000100fdu, 0x00010038u,
};
// Owns the reflection module so each test case cleans up after itself.
class ReflectModule {
public:
template <SizeT WordCount>
explicit ReflectModule(const Uint32 (&spirv)[WordCount]) {
m_created = spvReflectCreateShaderModule(sizeof(spirv), spirv, &m_module) ==
SPV_REFLECT_RESULT_SUCCESS;
}
~ReflectModule() {
if (m_created) {
spvReflectDestroyShaderModule(&m_module);
}
}
ReflectModule(const ReflectModule&) = delete;
ReflectModule& operator=(const ReflectModule&) = delete;
Bool Created() const { return m_created; }
const SpvReflectShaderModule& Get() const { return m_module; }
private:
SpvReflectShaderModule m_module{};
Bool m_created = false;
};
PipelineFactory::PipelineCreatePayload MakeDepthWritingPayload(
const VkPipelineColorBlendAttachmentState& attachment0) {
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 1;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = attachment0;
return payload;
}
} // namespace
// --- Device gate: MOBILEGL_MAGMA_DISABLE_BLENDED_DEPTH_WRITE tri-state ---
TEST(PipelineQuirkDeviceGate, ForceOnEnablesOnAnyVendor) {
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn,
kVendorIdArm));
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn,
kVendorIdQualcomm));
}
TEST(PipelineQuirkDeviceGate, ForceOffDisablesEvenOnQualcomm) {
EXPECT_FALSE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOff,
kVendorIdQualcomm));
}
TEST(PipelineQuirkDeviceGate, AutoDetectsQualcommOnly) {
EXPECT_TRUE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::Auto,
kVendorIdQualcomm));
EXPECT_FALSE(PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::Auto,
kVendorIdArm));
}
TEST(PipelineQuirkDeviceGate, ForceOnRoundTripsThroughTheFactoryFlag) {
const Bool previous = PipelineFactory::IsSuppressBlendedDepthWriteEnabled();
PipelineFactory::SetSuppressBlendedDepthWrite(
PipelineFactory::ShouldSuppressBlendedDepthWriteForDevice(QuirkOverride::ForceOn, kVendorIdArm));
EXPECT_TRUE(PipelineFactory::IsSuppressBlendedDepthWriteEnabled());
PipelineFactory::SetSuppressBlendedDepthWrite(previous);
}
// --- Per-pipeline strip decision against the pipeline create-info payload ---
TEST(PipelineQuirkStripDecision, MaxBlendIsStripped) {
// MC 26.3 OIT depth_bounds: GL_MAX accumulation writing depth - the case the quirk fixes.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_MAX, kFullColorWriteMask));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, MinBlendIsStripped) {
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_MIN, kFullColorWriteMask));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AdditiveOnePlusOneIsNotStripped) {
// ONE+ONE additive with a depth write matched zero draws of the 26.3 chain in the
// fixture sweep (transmittance/accumulate disable depth writes themselves); the only
// real content with this shape was harmless additive glow effects (Create). A quirk
// touches as little unrelated content as possible, so the shape stays exempt.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, SortedTransparencyOverBlendIsNotStripped) {
// Vanilla MC translucent layer (water, stained glass): SRC_ALPHA "over" compositing
// draws each surface once and depends on its depth writes to occlude particles, rain,
// and clouds drawn later - it must keep them.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, EffectivelyOpaqueBlendIsNotStripped) {
// GL_BLEND left enabled with ONE/ZERO+ADD factors is opaque in effect; stripping its
// depth write would break occlusion for plainly opaque geometry.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, FullyMaskedAccumulationBlendIsNotStripped) {
// Depth-prepass pattern: colorMask(0,0,0,0) with blending left enabled - blending is
// moot, and stripping would delete the entire prepass. MAX so the exemption, not the
// blend-op filter, is what keeps the depth write.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, 0));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, DisabledBlendIsNotStripped) {
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
false, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, NoDepthWriteMeansNoStrip) {
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.depthWriteEnable = false;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, FragDepthWriterIsExempt) {
// gl_FragDepth output does not go through per-pipeline vertex position math, so the
// cross-pipeline invariance hazard cannot affect it (e.g. the 26.3 OIT composite).
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.fragmentReplacesDepth = true;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AccumulationOnSecondaryAttachmentIsStripped) {
// The scan is not limited to attachment 0: an extremum accumulation on any live
// attachment marks the pipeline.
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 2;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
false, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ZERO, VK_BLEND_OP_ADD, kFullColorWriteMask);
payload.colorBlendAttachments[1] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask);
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, AlphaWeightedAdditiveIsNotStripped) {
// SRC_ALPHA,ONE additive: the classic *sorted* particle/glow blend. Kept exempt like
// every other ADD-op shape now that the strip is extremum-only.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_ADD, kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, ReverseSubtractIsNotStripped) {
// Deliberate narrowing: only the MIN/MAX extremum ops carry the depth-bounds
// signature. SUBTRACT-class ops stay outside the quirk until content demands them.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_REVERSE_SUBTRACT,
kFullColorWriteMask));
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, PartiallyMaskedAccumulationIsStripped) {
// Only a fully masked attachment is exempt; a live alpha channel still accumulates.
const auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, VK_COLOR_COMPONENT_A_BIT));
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, NoColorAttachmentsMeansNoStrip) {
// Depth-only FBO: the loop must not read the (stale) attachment array at all.
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 0;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask);
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
TEST(PipelineQuirkStripDecision, SeparateAlphaAccumulationIsNotStripped) {
// glBlendEquationSeparate(GL_FUNC_ADD, GL_MAX) over an ordinary color over-blend: the
// alpha channel accumulates but the color channel does not. Pins that the decision is
// color-channel only - widening it to alpha would re-capture sorted transparency.
auto attachment = MakeBlendAttachment(true, VK_BLEND_FACTOR_SRC_ALPHA,
VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask);
attachment.srcAlphaBlendFactor = VK_BLEND_FACTOR_ONE;
attachment.dstAlphaBlendFactor = VK_BLEND_FACTOR_ONE;
attachment.alphaBlendOp = VK_BLEND_OP_MAX;
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(MakeDepthWritingPayload(attachment)));
}
TEST(PipelineQuirkStripDecision, MixedOverAndMaskedAttachmentsAreNotStripped) {
PipelineFactory::PipelineCreatePayload payload{};
payload.colorAttachmentCount = 2;
payload.depthTestEnable = true;
payload.depthWriteEnable = true;
payload.colorBlendAttachments[0] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_SRC_ALPHA, VK_BLEND_FACTOR_ONE_MINUS_SRC_ALPHA, VK_BLEND_OP_ADD,
kFullColorWriteMask);
payload.colorBlendAttachments[1] = MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, 0);
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
// --- DepthReplacing reflection feeding the gl_FragDepth exemption ---
TEST(ReflectedFragmentReplacesDepth, TrueForAShaderThatAssignsFragDepth) {
const ReflectModule module(kFragDepthWriterSpirv);
ASSERT_TRUE(module.Created());
EXPECT_TRUE(ProgramFactory::ReflectedFragmentReplacesDepth(module.Get()));
}
TEST(ReflectedFragmentReplacesDepth, FalseForAPlainFragmentShader) {
const ReflectModule module(kPlainFragmentSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedFragmentReplacesDepth(module.Get()));
}
TEST(ReflectedFragmentReplacesDepth, FalseForAnEmptyModule) {
// A default-constructed module has no entry points; the scan must not dereference.
SpvReflectShaderModule emptyModule{};
EXPECT_FALSE(ProgramFactory::ReflectedFragmentReplacesDepth(emptyModule));
}
TEST(ReflectedFragmentReplacesDepth, ReflectedFlagFlipsTheStripDecision) {
// The two fixtures differ only by the gl_FragDepth assignment, so they pin that the
// reflected flag is what flips the strip decision for an otherwise identical pipeline.
const ReflectModule depthWriter(kFragDepthWriterSpirv);
const ReflectModule plain(kPlainFragmentSpirv);
ASSERT_TRUE(depthWriter.Created());
ASSERT_TRUE(plain.Created());
auto payload = MakeDepthWritingPayload(MakeBlendAttachment(
true, VK_BLEND_FACTOR_ONE, VK_BLEND_FACTOR_ONE, VK_BLEND_OP_MAX, kFullColorWriteMask));
payload.fragmentReplacesDepth = ProgramFactory::ReflectedFragmentReplacesDepth(plain.Get());
EXPECT_TRUE(PipelineFactory::ShouldSuppressDepthWrite(payload));
payload.fragmentReplacesDepth = ProgramFactory::ReflectedFragmentReplacesDepth(depthWriter.Get());
EXPECT_FALSE(PipelineFactory::ShouldSuppressDepthWrite(payload));
}
// --- InstanceIndex reflection feeding the shaderDrawParameters diagnostic ---
TEST(ReflectedReadsInstanceIndexBuiltin, TrueForAShaderReadingInstanceIndex) {
const ReflectModule module(kInstanceIndexVertexSpirv);
ASSERT_TRUE(module.Created());
EXPECT_TRUE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAShaderReadingADifferentBuiltin) {
// Discriminates the builtin's identity, not merely its presence: weakening the check to
// "has any BuiltIn decoration" would fire the diagnostic on every real vertex shader.
const ReflectModule module(kVertexIndexVertexSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAShaderWithNoInputBuiltins) {
const ReflectModule module(kPlainFragmentSpirv);
ASSERT_TRUE(module.Created());
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(module.Get()));
}
TEST(ReflectedReadsInstanceIndexBuiltin, FalseForAnEmptyModule) {
SpvReflectShaderModule emptyModule{};
EXPECT_FALSE(ProgramFactory::ReflectedReadsInstanceIndexBuiltin(emptyModule));
}
@@ -238,6 +238,77 @@ void main() {
}
}
// KHR-GL33.shaders.preprocessor.* — a block comment is one preprocessing token that the C/GLSL
// preprocessor replaces with a single space, even when it spans newlines inside a directive. glslang
// handles this natively, so MobileGL must not mangle it. These reproduce the CTS cases that failed
// because comment blanking preserved the interior newline, truncating multi-line #define bodies.
static void ExpectCompiles(MobileGL::ShaderStage stage, GLenum glStage, MobileGL::String source) {
using namespace MG_Util::ShaderTranspiler;
PreprocessShaderSource(stage, source);
ShaderAttrib attrib{.shaderType = glStage, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
TEST_F(ProgramUtilTest, PreprocessMultilineCommentInDefineBodyCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
#define VALUE /* current
value */ 4.2
void main()
{
out0 = VALUE;
})");
}
TEST_F(ProgramUtilTest, PreprocessRedefineObjectMultilineCommentCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
# define VAL1 1.0
#define VAL2 2.0
#define RES2 /* fdsjklfdsjkl
dsfjkhfdsjkh
fdsjklhfdsjkh */ (RES1 * VAL2)
#define RES1 (VAL2 / VAL1)
#define RES2 /* ewrlkjhsadf */ (RES1 * VAL2)
#define VALUE (RES2 + RES1)
void main()
{
out0 = VALUE;
})");
}
TEST_F(ProgramUtilTest, PreprocessFunctionMacroRedefinitionMultilineCommentCompiles) {
ExpectCompiles(ShaderStage::Fragment, GL_FRAGMENT_SHADER,
R"(#version 330
precision mediump float;
out float out0;
# define FUNC(a,b) (a +b)
# define FUNC(a,b)(a /* comment
*/ +b)
void main()
{
out0 = FUNC(1.0, 2.0);
})");
}
// Note: KHR-GL3x.shaders.preprocessor.conditional_inclusion.basic_2 (`#define AAA defined(BBB)` used
// in `#if !AAA`) is intentionally NOT handled here. Generating the `defined` operator via macro
// expansion is undefined per the C/GLSL preprocessor spec, and glslang deliberately rejects it
// ("'defined' : cannot use in preprocessor expression when expanded from macros"). Making it pass
// would require MobileGL to run its own macro expansion ahead of glslang, which is exactly the
// preprocessing we defer to glslang; the two cases stay failing by design.
TEST_F(ProgramUtilTest, PreprocessLegacyFragmentShaderModernizesGlmarkStyleSource) {
using namespace MG_Util::ShaderTranspiler;
@@ -426,6 +497,62 @@ void main() {
verifyVersion("#version 460 core");
}
// KHR-GL33.shaders.preprocessor.directive.version_* (also re-run verbatim under GL40-GL44): the
// compiler must REJECT a malformed #version line. MobileGL used to rewrite the whole line to
// "#version 330 core" whenever it could scrape a leading integer - or treat an unknown profile token
// as core - which silently legalized every form below. CTS compiles the shader's own #version
// verbatim, so the rejection has to survive preprocessing (and the 460 retry).
TEST_F(ProgramUtilTest, PreprocessRejectsMalformedVersionDirectives) {
using namespace MG_Util::ShaderTranspiler;
const char* body = "\nout vec4 fragColor;\nvoid main() { fragColor = vec4(1.0); }\n";
const auto rejects = [](const String& fullSource) {
String src = fullSource;
PreprocessShaderSource(ShaderStage::Fragment, src);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
return res ? false : true; // "rejects" == compile failed
};
// Silently legalized today - the five this fix must flip to rejection:
EXPECT_TRUE(rejects(String("#version 329") + body)) << "329 is not a real version";
EXPECT_TRUE(rejects(String("#version 331") + body)) << "331 is not a real version";
EXPECT_TRUE(rejects(String("#version 330 foo") + body)) << "unknown profile keyword";
EXPECT_TRUE(rejects(String("#version 330.0") + body)) << "float literal, not an int token";
EXPECT_TRUE(rejects(String("#version 330 foobar") + body)) << "trailing tokens after a valid decl";
// Already rejected (no leading integer, or #version is not the first token) - pinned so a future
// change to the normalizer cannot start legalizing them either:
EXPECT_TRUE(rejects(String("#version") + body)) << "missing version number";
EXPECT_TRUE(rejects(String("#version foobar") + body)) << "identifier where the int belongs";
EXPECT_TRUE(rejects(String("#version AAA") + body)) << "identifier where the int belongs";
EXPECT_TRUE(rejects(String("precision mediump float;\n#version 330") + body))
<< "#version must be the first statement";
EXPECT_TRUE(rejects(String("#define FOO BAR\n#version 330") + body))
<< "#version must precede a #define";
}
// The PASS half of the same CTS group: a valid decl, and #version preceded only by whitespace or a
// comment, must still compile. Guards the fix above from over-rejecting.
TEST_F(ProgramUtilTest, PreprocessKeepsValidVersionDirectivesCompiling) {
using namespace MG_Util::ShaderTranspiler;
const char* body = "\nout vec4 fragColor;\nvoid main() { fragColor = vec4(1.0); }\n";
const auto compiles = [](const String& fullSource) {
String src = fullSource;
PreprocessShaderSource(ShaderStage::Fragment, src);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
return res ? true : false;
};
EXPECT_TRUE(compiles(String("#version 330 core") + body));
EXPECT_TRUE(compiles(String("\n#version 330 core") + body))
<< "leading whitespace is legal before #version";
EXPECT_TRUE(compiles(String("// test\n#version 330 core") + body))
<< "a leading comment is legal before #version";
}
TEST_F(ProgramUtilTest, PreprocessUsesRealSpacedVersionDirectiveForInjectedOutput) {
using namespace MG_Util::ShaderTranspiler;
@@ -446,6 +573,9 @@ void main() {
EXPECT_NE(versionPos, String::npos);
EXPECT_EQ(outputPos, versionPos + std::strlen("#version 330 core\n"));
EXPECT_NE(source.find("// #version 460 core"), String::npos);
// This #line sits ahead of the version directive, where GLSL would never have honoured it, so
// it is still dropped. Directives that follow the version line are kept - see
// PreprocessKeepsPlainLineDirectivesAndSparesLookalikeIdentifiers.
EXPECT_EQ(source.find("#line"), String::npos);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
@@ -455,6 +585,105 @@ void main() {
}
}
// A banner line like "//*** NOTE ***" contains "/*" at offset 1 and no "*/" anywhere after it. The
// old hand-rolled comment stripper searched for "/*" with no lexical state, found that, failed to
// find a terminator, and erased everything from there to the end of the file - deleting the entire
// shader. Banner comments in that exact shape are common in Iris and OptiFine packs.
TEST_F(ProgramUtilTest, PreprocessKeepsShaderBodyAfterAStarredLineComment) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
//*** lighting pass ***
out vec4 fragColor;
void main() {
fragColor = vec4(1.0);
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("void main()"), String::npos) << "shader body was truncated:\n" << source;
EXPECT_NE(source.find("fragColor = vec4(1.0);"), String::npos);
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
// The builtin-shadowing rename only fires when the shader really defines its own round/tanh/etc.
// Deciding that from a commented-out definition renames every genuine call to the builtin to a
// mg_ name that nothing defines, which fails to link.
TEST_F(ProgramUtilTest, PreprocessIgnoresCommentedOutBuiltinShadowingDefinition) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
// float round(float x) { return floor(x + 0.5); }
out vec4 fragColor;
void main() {
fragColor = vec4(round(1.25));
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("round(1.25)"), String::npos) << "call was renamed from a comment:\n" << source;
EXPECT_EQ(source.find("mg_round"), String::npos);
}
// A block-commented extension directive must not be treated as a real one - the int64 filter turns
// unsupported directives into #error, so reading one out of a comment manufactures a compile
// failure for a shader that never asked for the extension.
TEST_F(ProgramUtilTest, PreprocessIgnoresBlockCommentedExtensionDirectives) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
/*
#extension GL_ARB_gpu_shader_int64 : require
*/
out vec4 fragColor;
void main() {
fragColor = vec4(1.0);
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_EQ(source.find("#error"), String::npos) << "#error synthesized from a comment:\n" << source;
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
// KHR-GL33.shaders.preprocessor.builtin.line_* checks that __LINE__ follows #line. That only works
// if the directive reaches glslang, so a plain integer form must pass through untouched - while
// "#linear" and friends must not be mistaken for it.
TEST_F(ProgramUtilTest, PreprocessKeepsPlainLineDirectivesAndSparesLookalikeIdentifiers) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
out vec4 fragColor;
#line 42
float linear(float x) { return x; }
void main() {
#line 100
fragColor = vec4(linear(float(__LINE__)));
}
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("#line 42"), String::npos) << source;
EXPECT_NE(source.find("#line 100"), String::npos) << source;
EXPECT_NE(source.find("float linear(float x)"), String::npos) << "identifier lookalike was eaten:\n" << source;
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = source};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) {
FAIL() << "errc: " << res.error().errc << "\nlog: " << res.error().log << "\nsource:\n" << source;
}
}
TEST_F(ProgramUtilTest, PreprocessModernSampleQualifierStaysAtVersion460) {
using namespace MG_Util::ShaderTranspiler;
@@ -745,6 +974,16 @@ TEST_F(ProgramUtilTest, RetargetLegacyVersionDirectiveOnlyTouchesNormalizedDeskt
String commented = "// #version 330 core\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(commented));
EXPECT_EQ(commented.find("#version 460"), String::npos);
// A malformed directive must NOT be rescued to 460 - that is what silently legalized the CTS
// directive.version_* rejection cases. The bad version stays put so glslang keeps rejecting it.
String badNumber = "#version 331\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(badNumber));
EXPECT_EQ(badNumber.find("#version 460"), String::npos);
String badProfile = "#version 330 foo\nvoid main() {}\n";
EXPECT_FALSE(RetargetLegacyVersionDirectiveTo460(badProfile));
EXPECT_EQ(badProfile.find("#version 460"), String::npos);
}
const char* fs = R"(#version 150
@@ -921,6 +1160,353 @@ TEST_F(ProgramUtilTest, CompileFragmentShaderWithDiscard) {
}
}
// noperspective is core desktop GLSL (1.30+) and maps to the SPIR-V NoPerspective decoration. It must
// reach glslang (not be stripped as text) so the SPIR-V carries the decoration; SPIRV-Cross then emits
// ESSL `noperspective` + the GL_NV_shader_noperspective_interpolation extension. Shader packs
// (Iris/Complementary) depend on it, and KHR-GL33.glsl_noperspective fails if the result matches
// smooth. This is the DirectGLES path with the NV extension available (SPIRV-Cross's default).
TEST_F(ProgramUtilTest, NoperspectiveInterpolationSurvivesToEssl) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 fragColor;
void main() { fragColor = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile errc: " << res.error().errc << "\nlog: " << res.error().log;
ProgramAttrib programAttrib{.shaders = {res.value()}};
auto program_res = ShaderCompiler::LinkProgram(programAttrib);
if (!program_res) FAIL() << "link errc: " << program_res.error().errc << "\nlog: " << program_res.error().log;
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *program_res.value()};
auto bin_res = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
if (!bin_res) FAIL() << "spirv errc: " << bin_res.error().errc << "\nlog: " << bin_res.error().log;
ASSERT_EQ(bin_res.value().size(), 1u);
SpvcSession session(bin_res.value()[0], SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc << "\nlog: " << essl.error().log;
EXPECT_NE(essl.value().find("noperspective"), String::npos)
<< "noperspective was lost before it reached SPIR-V:\n" << essl.value();
EXPECT_NE(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos)
<< "SPIRV-Cross must require the NV extension for ES noperspective:\n" << essl.value();
}
// The old handling was a naked substring erase of "noperspective", so any identifier that merely
// contained those characters (a uniform named noperspectiveBlend, say) got mangled. Removing the
// strip fixes it - glslang, which is identifier-aware, is the only thing that should see the keyword.
TEST_F(ProgramUtilTest, PreprocessDoesNotCorruptIdentifiersContainingNoperspective) {
using namespace MG_Util::ShaderTranspiler;
String source = R"(#version 330 core
uniform float noperspectiveBlend;
out vec4 fragColor;
void main() { fragColor = vec4(noperspectiveBlend); }
)";
PreprocessShaderSource(ShaderStage::Fragment, source);
EXPECT_NE(source.find("noperspectiveBlend"), String::npos)
<< "identifier was corrupted by substring stripping:\n" << source;
}
// The DirectGLES fallback for devices without GL_NV_shader_noperspective_interpolation: stripping the
// NoPerspective decoration makes SPIRV-Cross emit a plain smooth varying with no `#extension … :
// require`, so the shader still compiles (rendering as smooth) instead of being rejected by the driver.
TEST_F(ProgramUtilTest, StripNoPerspectiveFallbackProducesPlainEssl) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 fragColor;
void main() { fragColor = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile errc: " << res.error().errc << "\nlog: " << res.error().log;
ProgramAttrib programAttrib{.shaders = {res.value()}};
auto program_res = ShaderCompiler::LinkProgram(programAttrib);
if (!program_res) FAIL() << "link errc: " << program_res.error().errc << "\nlog: " << program_res.error().log;
ProgramBinaryAttrib binaryAttrib{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *program_res.value()};
auto bin_res = ShaderCompiler::GetSpirvBinaryFromProgram(binaryAttrib);
if (!bin_res) FAIL() << "spirv errc: " << bin_res.error().errc << "\nlog: " << bin_res.error().log;
ASSERT_EQ(bin_res.value().size(), 1u);
// Precondition: with the decoration present the default decompile requires the NV extension.
{
SpvcSession session(bin_res.value()[0], SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc;
ASSERT_NE(essl.value().find("noperspective"), String::npos) << essl.value();
}
// The fallback strips the decoration -> plain smooth ESSL, no extension require.
Vector<Uint32> stripped;
ASSERT_TRUE(ShaderCompiler::StripNoPerspectiveForEssl(bin_res.value()[0], stripped));
ASSERT_FALSE(stripped.empty());
SpvcSession session(stripped, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile errc: " << essl.error().errc << "\nlog: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos)
<< "the decoration should be gone:\n" << essl.value();
EXPECT_EQ(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos)
<< "no extension require without the decoration:\n" << essl.value();
}
// Directly exercises BOTH decoration forms StripNoPerspectivePass handles: a plain-variable
// OpDecorate NoPerspective (in-operand 1) and an interface-block-member OpMemberDecorate NoPerspective
// (in-operand 2). The ESSL round-trip tests above use only a scalar input, so they never reach the
// member-decorate branch, which a block varying like `in Block { noperspective vec4 c; }` (common in
// shader packs) produces. Unrelated decorations (Flat, Location) must survive untouched.
TEST_F(ProgramUtilTest, StripNoPerspectivePassRemovesBothDecorateForms) {
using namespace MG_Util::ShaderTranspiler;
const String spirvText = R"(
OpCapability Shader
OpMemoryModel Logical GLSL450
OpEntryPoint Fragment %main "main" %plainVar %blockVar %flatVar
OpExecutionMode %main OriginUpperLeft
OpName %main "main"
OpDecorate %plainVar Location 0
OpDecorate %plainVar NoPerspective
OpMemberDecorate %Block 0 NoPerspective
OpDecorate %blockVar Location 1
OpDecorate %flatVar Location 2
OpDecorate %flatVar Flat
%void = OpTypeVoid
%mainFn = OpTypeFunction %void
%float = OpTypeFloat 32
%v4float = OpTypeVector %float 4
%int = OpTypeInt 32 1
%inV4Ptr = OpTypePointer Input %v4float
%plainVar = OpVariable %inV4Ptr Input
%Block = OpTypeStruct %v4float
%inBlockPtr = OpTypePointer Input %Block
%blockVar = OpVariable %inBlockPtr Input
%inIntPtr = OpTypePointer Input %int
%flatVar = OpVariable %inIntPtr Input
%main = OpFunction %void None %mainFn
%mainBody = OpLabel
OpReturn
OpFunctionEnd
)";
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
Vector<uint32_t> inputBinary;
ASSERT_TRUE(tools.Assemble(spirvText, &inputBinary));
const auto countNoPerspective = [](const String& text) {
SizeT count = 0, offset = 0;
while ((offset = text.find("NoPerspective", offset)) != String::npos) {
++count;
offset += std::strlen("NoPerspective");
}
return count;
};
String inputText;
ASSERT_TRUE(tools.Disassemble(inputBinary, &inputText));
ASSERT_EQ(countNoPerspective(inputText), 2u)
<< "fixture must carry both a plain and a member NoPerspective:\n" << inputText;
Vector<uint32_t> outputBinary;
ASSERT_TRUE(ShaderCompiler::StripNoPerspectiveForEssl(inputBinary, outputBinary));
ASSERT_FALSE(outputBinary.empty());
String outputText;
ASSERT_TRUE(tools.Disassemble(outputBinary, &outputText));
EXPECT_EQ(countNoPerspective(outputText), 0u)
<< "both NoPerspective decorations (OpDecorate and OpMemberDecorate) must be stripped:\n" << outputText;
EXPECT_NE(outputText.find("Flat"), String::npos)
<< "the unrelated Flat decoration must survive:\n" << outputText;
EXPECT_NE(outputText.find("Location"), String::npos)
<< "Location decorations must survive:\n" << outputText;
}
// Phase 2 emulation - fragment side. On a device without the NV extension the NoPerspective input is
// recovered as `load * gl_FragCoord.w` and the decoration removed; gl_FragCoord is synthesized because
// the shader did not otherwise use it. The emulated SPIR-V must validate and decompile without the
// extension require.
TEST_F(ProgramUtilTest, EmulateNoperspectiveFragmentRecoversWithFragCoordW) {
using namespace MG_Util::ShaderTranspiler;
String fs = R"(#version 330 core
noperspective in vec4 vColor;
out vec4 f;
void main() { f = vColor; }
)";
ShaderAttrib attrib{.shaderType = GL_FRAGMENT_SHADER, .sourceStr = fs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile: " << res.error().log;
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
if (!pr) FAIL() << "link: " << pr.error().log;
ProgramBinaryAttrib ba{.shaderTypes = {GL_FRAGMENT_SHADER}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
if (!br) FAIL() << "spirv: " << br.error().log;
ASSERT_EQ(br.value().size(), 1u);
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(br.value()[0], emulated));
ASSERT_FALSE(emulated.empty());
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << "emulated SPIR-V must be valid:\n" << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << "decoration must be stripped:\n" << dis;
EXPECT_NE(dis.find("FragCoord"), String::npos) << "gl_FragCoord must be synthesized:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos) << "the recovery multiply must be present:\n" << dis;
SpvcSession session(emulated, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos) << essl.value();
EXPECT_EQ(essl.value().find("GL_NV_shader_noperspective_interpolation"), String::npos) << essl.value();
EXPECT_NE(essl.value().find("gl_FragCoord"), String::npos) << "recovery must reference gl_FragCoord:\n" << essl.value();
}
// Phase 2 emulation - vertex side. The NoPerspective output is pre-multiplied by gl_Position.w before
// return and the decoration removed. Emulated SPIR-V must validate and decompile without the extension.
TEST_F(ProgramUtilTest, EmulateNoperspectiveVertexPreMultipliesByPositionW) {
using namespace MG_Util::ShaderTranspiler;
String vs = R"(#version 330 core
in vec4 pos;
noperspective out vec4 vColor;
void main() { gl_Position = pos; vColor = pos; }
)";
ShaderAttrib attrib{.shaderType = GL_VERTEX_SHADER, .sourceStr = vs};
auto res = ShaderCompiler::CompileShader(attrib);
if (!res) FAIL() << "compile: " << res.error().log;
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
if (!pr) FAIL() << "link: " << pr.error().log;
ProgramBinaryAttrib ba{.shaderTypes = {GL_VERTEX_SHADER}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
if (!br) FAIL() << "spirv: " << br.error().log;
ASSERT_EQ(br.value().size(), 1u);
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(br.value()[0], emulated));
ASSERT_FALSE(emulated.empty());
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << "emulated SPIR-V must be valid:\n" << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << "decoration must be stripped:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos) << "the pre-multiply must be present:\n" << dis;
SpvcSession session(emulated, SessionUsageBit::Transpile);
auto essl = ShaderCompiler::DecompileShader(session);
if (!essl) FAIL() << "decompile: " << essl.error().log;
EXPECT_EQ(essl.value().find("noperspective"), String::npos) << essl.value();
EXPECT_NE(essl.value().find("gl_Position"), String::npos) << "pre-multiply must reference gl_Position:\n" << essl.value();
}
namespace {
// Compiles one shader stage through the full pipeline and returns its SPIR-V, or fails the test.
MobileGL::Vector<uint32_t> CompileStageSpirv(GLenum type, const char* src) {
using namespace MG_Util::ShaderTranspiler;
ShaderAttrib attrib{.shaderType = type, .sourceStr = src};
auto res = ShaderCompiler::CompileShader(attrib);
EXPECT_TRUE(static_cast<bool>(res)) << (res ? "" : res.error().log);
if (!res) return {};
ProgramAttrib pa{.shaders = {res.value()}};
auto pr = ShaderCompiler::LinkProgram(pa);
EXPECT_TRUE(static_cast<bool>(pr)) << (pr ? "" : pr.error().log);
if (!pr) return {};
ProgramBinaryAttrib ba{.shaderTypes = {type}, .program = *pr.value()};
auto br = ShaderCompiler::GetSpirvBinaryFromProgram(ba);
EXPECT_TRUE(static_cast<bool>(br)) << (br ? "" : br.error().log);
if (!br || br.value().empty()) return {};
return br.value()[0];
}
} // namespace
// Regression: the vertex pre-multiply must be applied exactly once (in main), not once per function.
// glslang does not inline, so a helper function survives as its own OpFunction; instrumenting its
// return too would scale the varying by gl_Position.w twice (w^2).
TEST_F(ProgramUtilTest, EmulateNoperspectiveVertexWithHelperScalesExactlyOnce) {
using namespace MG_Util::ShaderTranspiler;
// helper() returns via OpReturnValue and adds (no vector*scalar), so the ONLY OpVectorTimesScalar
// in the module is the emulation's pre-multiply. The old all-functions code injected it at both
// helper's and main's return -> count 2; restricted to the entry function it is 1.
auto spirv = CompileStageSpirv(GL_VERTEX_SHADER, R"(#version 330 core
in vec4 pos;
noperspective out vec4 vColor;
vec4 helper(vec4 x) { return x + vec4(1.0); }
void main() { gl_Position = pos; vColor = helper(pos); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
SizeT count = 0, off = 0;
while ((off = dis.find("OpVectorTimesScalar", off)) != String::npos) {
++count;
off += std::strlen("OpVectorTimesScalar");
}
EXPECT_EQ(count, 1u) << "the gl_Position.w pre-multiply must happen exactly once, not per function:\n" << dis;
}
// Regression: a single-component read (vColor.x), which glslang lowers via OpAccessChain, must still be
// recovered with gl_FragCoord.w - not silently left un-scaled.
TEST_F(ProgramUtilTest, EmulateNoperspectiveFragmentComponentReadIsRecovered) {
using namespace MG_Util::ShaderTranspiler;
auto spirv = CompileStageSpirv(GL_FRAGMENT_SHADER, R"(#version 330 core
noperspective in vec4 vColor;
out vec4 f;
void main() { f = vec4(vColor.x); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << dis;
EXPECT_NE(dis.find("FragCoord"), String::npos)
<< "the component read must still be recovered via gl_FragCoord.w:\n" << dis;
}
// Coverage: a scalar float varying exercises the OpFMul path; a vector varying the OpVectorTimesScalar
// path; multiple noperspective varyings in one stage are all handled.
TEST_F(ProgramUtilTest, EmulateNoperspectiveHandlesScalarAndMultipleVaryings) {
using namespace MG_Util::ShaderTranspiler;
auto spirv = CompileStageSpirv(GL_FRAGMENT_SHADER, R"(#version 330 core
noperspective in float a;
noperspective in vec2 b;
out vec4 f;
void main() { f = vec4(a, b, 1.0); }
)");
ASSERT_FALSE(spirv.empty());
Vector<uint32_t> emulated;
ASSERT_TRUE(ShaderCompiler::EmulateNoPerspectiveForEssl(spirv, emulated));
spvtools::SpirvTools tools(SPV_ENV_VULKAN_1_1);
String dis;
ASSERT_TRUE(tools.Disassemble(emulated, &dis));
ASSERT_TRUE(tools.Validate(emulated)) << dis;
EXPECT_EQ(dis.find("NoPerspective"), String::npos) << dis;
EXPECT_NE(dis.find("OpFMul"), String::npos) << "the scalar varying must scale with OpFMul:\n" << dis;
EXPECT_NE(dis.find("OpVectorTimesScalar"), String::npos)
<< "the vector varying must scale with OpVectorTimesScalar:\n" << dis;
}
const char* vs_location = R"(#version 460
in vec4 Position;
@@ -1623,4 +2209,26 @@ TEST_F(ProgramUtilTest, RewriteLinearSubgroupPrefixScanRejectsPartialOrUnsafeTem
ASSERT_NE(consumerEnd, String::npos);
nestedScan.insert(consumerEnd + 1, "\n }");
expectUnchanged(std::move(nestedScan));
// ARB/NV spellings of lane-width-sensitive builtins must block the rewrite exactly
// like their KHR counterparts.
String arbSubgroupBuiltin = MakeLinearSubgroupPrefixScanShader();
arbSubgroupBuiltin.insert(arbSubgroupBuiltin.find("float importance"),
"uint arbLane = gl_SubGroupInvocationARB;\n ");
expectUnchanged(std::move(arbSubgroupBuiltin));
String arbBallotCall = MakeLinearSubgroupPrefixScanShader();
arbBallotCall.insert(arbBallotCall.find("float importance"),
"uint64_t arbMask = ballotARB(true);\n ");
expectUnchanged(std::move(arbBallotCall));
String nvWarpBuiltin = MakeLinearSubgroupPrefixScanShader();
nvWarpBuiltin.insert(nvWarpBuiltin.find("float importance"),
"uint warpSize = gl_WarpSizeNV;\n ");
expectUnchanged(std::move(nvWarpBuiltin));
String nvShuffleCall = MakeLinearSubgroupPrefixScanShader();
nvShuffleCall.insert(nvShuffleCall.find("float importance"),
"float other = shuffleNV(1.0f, 0u, 32u);\n ");
expectUnchanged(std::move(nvShuffleCall));
}
+351 -1
View File
@@ -794,6 +794,33 @@ TEST(DirectVulkanSanity, ReadbackConvertsRgba8AndRgba16fPixels) {
EXPECT_FLOAT_EQ(rgba16fFloatResult[3], 1.0f);
}
TEST(DirectVulkanSanity, ReadbackDecodesSingleChannel32BitFormats) {
using MobileGL::MG_Backend::DirectVulkan::VulkanRenderer;
// The reinterpretation feature makes R32F/R32UI-class images common readback sources
// (iterationRP custom images). Missing channels take GL defaults: 0 for GB, 1 for alpha.
const MobileGL::Float r32f[] = {0.75f, -2.0f};
MobileGL::Float r32fResult[8]{};
ASSERT_TRUE(VulkanRenderer::ConvertReadbackPixels(
reinterpret_cast<const MobileGL::Uint8*>(r32f), VK_FORMAT_R32_SFLOAT,
2, 1, GL_RGBA, GL_FLOAT, sizeof(MobileGL::Float) * 8,
reinterpret_cast<MobileGL::Uint8*>(r32fResult)));
EXPECT_FLOAT_EQ(r32fResult[0], 0.75f);
EXPECT_FLOAT_EQ(r32fResult[1], 0.0f);
EXPECT_FLOAT_EQ(r32fResult[2], 0.0f);
EXPECT_FLOAT_EQ(r32fResult[3], 1.0f);
EXPECT_FLOAT_EQ(r32fResult[4], -2.0f);
const MobileGL::Uint32 r32ui[] = {12345u};
MobileGL::Float r32uiResult[4]{};
ASSERT_TRUE(VulkanRenderer::ConvertReadbackPixels(
reinterpret_cast<const MobileGL::Uint8*>(r32ui), VK_FORMAT_R32_UINT,
1, 1, GL_RGBA, GL_FLOAT, sizeof(MobileGL::Float) * 4,
reinterpret_cast<MobileGL::Uint8*>(r32uiResult)));
EXPECT_FLOAT_EQ(r32uiResult[0], 12345.0f);
EXPECT_FLOAT_EQ(r32uiResult[3], 1.0f);
}
TEST(DirectVulkanSanity, DrawIndexedIndirectCommandMatchesGlAndVulkanLayout) {
using namespace MobileGL::MG_Backend::DirectVulkan;
@@ -999,9 +1026,27 @@ TEST(DirectVulkanSanity, SampledViewFormatMatchesSamplerNumericDomainWithoutChan
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_B10G11R11_UFLOAT_PACK32, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_UNDEFINED);
// Depth/stencil formats never resolve through color-class reinterpretation; they pass
// through unchanged so the existing depth-aspect sampled view is used. Combined
// depth-stencil formats are multi-numeric (vkuFormatIsSampledFloat is false for them),
// so without the passthrough a plain sampler2D/sampler2DShadow on GL_DEPTH24_STENCIL8
// would resolve to UNDEFINED and the draw would be dropped.
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D24_UNORM_S8_UINT, SamplerNumericDomain::Float),
VK_FORMAT_D24_UNORM_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT_S8_UINT, SamplerNumericDomain::Float),
VK_FORMAT_D32_SFLOAT_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT, SamplerNumericDomain::Float),
VK_FORMAT_D32_SFLOAT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D24_UNORM_S8_UINT, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_D24_UNORM_S8_UINT);
EXPECT_EQ(VkTextureManager::ResolveSampledImageViewFormat(
VK_FORMAT_D32_SFLOAT, SamplerNumericDomain::UnsignedInteger),
VK_FORMAT_UNDEFINED);
VK_FORMAT_D32_SFLOAT);
EXPECT_TRUE(VkTextureManager::AreSampledImageViewFormatsCompatible(
VK_FORMAT_R32_SFLOAT, VK_FORMAT_R32_UINT));
@@ -1386,3 +1431,308 @@ TEST(RenderStateSanity, PrimitiveRestartIndexStoresAndReadsBack) {
MG_State::pGLContext.reset();
}
// ---- DirectGLES readback driver-state shadows ----------------------------------------------------
// Regression coverage for the readback-path state-leak overhaul: the pixel-PBO
// binding cache, the framebuffer-binding shadow, the PACK pixel-store shadow and
// the scratch-FBO attachment shadow must (a) leave the driver in the documented
// resting state, (b) skip redundant GL calls, and (c) scrub correctly on
// deletion. All drive the real Managers.cpp implementations against a recording
// mock GLES table.
namespace {
struct StateGuardCallLog {
MobileGL::Vector<MobileGL::String> calls;
MobileGL::SizeT Count(const MobileGL::String& prefix) const {
MobileGL::SizeT n = 0;
for (const auto& c : calls) {
if (c.compare(0, prefix.size(), prefix) == 0) ++n;
}
return n;
}
};
StateGuardCallLog* g_stateGuardLog = nullptr;
GLuint g_nextStateGuardFBOId = 201;
void SG_Log(MobileGL::String entry) {
if (g_stateGuardLog) g_stateGuardLog->calls.push_back(MobileGL::Move(entry));
}
void SG_BindBuffer(GLenum target, GLuint buffer) {
SG_Log("BindBuffer:" + std::to_string(target) + ":" + std::to_string(buffer));
}
void SG_BindFramebuffer(GLenum target, GLuint framebuffer) {
SG_Log("BindFramebuffer:" + std::to_string(target) + ":" + std::to_string(framebuffer));
}
void SG_GetIntegerv(GLenum pname, GLint* data) {
SG_Log("GetIntegerv:" + std::to_string(pname));
if (data) *data = 0;
}
void SG_PixelStorei(GLenum pname, GLint param) {
SG_Log("PixelStorei:" + std::to_string(pname) + ":" + std::to_string(param));
}
void SG_GenFramebuffers(GLsizei count, GLuint* framebuffers) {
for (GLsizei i = 0; i < count; ++i) framebuffers[i] = g_nextStateGuardFBOId++;
}
void SG_FramebufferTexture2D(GLenum target, GLenum attachment, GLenum textarget, GLuint texture, GLint level) {
SG_Log("FramebufferTexture2D:" + std::to_string(target) + ":" + std::to_string(attachment) + ":" +
std::to_string(textarget) + ":" + std::to_string(texture) + ":" + std::to_string(level));
}
void SG_FramebufferTextureLayer(GLenum target, GLenum attachment, GLuint texture, GLint level, GLint layer) {
SG_Log("FramebufferTextureLayer:" + std::to_string(target) + ":" + std::to_string(attachment) + ":" +
std::to_string(texture) + ":" + std::to_string(level) + ":" + std::to_string(layer));
}
void SG_ReadBuffer(GLenum src) {
SG_Log("ReadBuffer:" + std::to_string(src));
}
void SG_DrawBuffers(GLsizei n, const GLenum* bufs) {
SG_Log("DrawBuffers:" + std::to_string(n) + ":" + std::to_string(n > 0 && bufs ? bufs[0] : 0));
}
GLenum SG_NoError() {
return GL_NO_ERROR;
}
// Installs the recording table and resets every readback driver-state shadow on
// both ends, so these tests cannot bleed into (or inherit from) other tests.
struct ScopedStateGuardMocks {
ScopedStateGuardMocks(): previousFunctions(MobileGL::MG_Backend::DirectGLES::g_GLESFuncs) {
ResetShadows();
MobileGL::MG_External::GLESFunctionsTable functions{};
functions.glBindBuffer = SG_BindBuffer;
functions.glBindFramebuffer = SG_BindFramebuffer;
functions.glGetIntegerv = SG_GetIntegerv;
functions.glPixelStorei = SG_PixelStorei;
functions.glGenFramebuffers = SG_GenFramebuffers;
functions.glFramebufferTexture2D = SG_FramebufferTexture2D;
functions.glFramebufferTextureLayer = SG_FramebufferTextureLayer;
functions.glReadBuffer = SG_ReadBuffer;
functions.glDrawBuffers = SG_DrawBuffers;
functions.glGetError = SG_NoError;
MobileGL::MG_Backend::DirectGLES::SetGLESFuncsTable(functions);
g_stateGuardLog = &log;
}
~ScopedStateGuardMocks() {
g_stateGuardLog = nullptr;
MobileGL::MG_Backend::DirectGLES::SetGLESFuncsTable(previousFunctions);
ResetShadows();
}
ScopedStateGuardMocks(const ScopedStateGuardMocks&) = delete;
ScopedStateGuardMocks& operator=(const ScopedStateGuardMocks&) = delete;
static void ResetShadows() {
MobileGL::MG_Backend::DirectGLES::BufferImpl::InvalidatePixelBufferBindingCaches();
MobileGL::MG_Backend::DirectGLES::FramebufferImpl::InvalidateFramebufferBindingCache();
MobileGL::MG_Backend::DirectGLES::PixelStoreImpl::InvalidatePackStateCache();
MobileGL::MG_Backend::DirectGLES::ScratchFBOImpl::OnBackendContextDestroyed();
}
StateGuardCallLog log;
MobileGL::MG_External::GLESFunctionsTable previousFunctions;
};
} // namespace
TEST(DirectGLESStateGuards, PixelPackBindingCacheSkipsRedundantBindsAndRestsAtZero) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
BufferImpl::BindPixelPackBufferId(5);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 1u);
BufferImpl::BindPixelPackBufferId(5); // redundant: must not reach the driver
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 1u);
BufferImpl::BindPixelPackBufferId(0); // scope exit: resting state
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 2u);
BufferImpl::BindPixelPackBufferId(0);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 2u);
// After invalidation (MakeCurrent / context reset) the first bind must reach
// the driver again even for the same value.
BufferImpl::InvalidatePixelBufferBindingCaches();
BufferImpl::BindPixelPackBufferId(0);
EXPECT_EQ(mocks.log.Count("BindBuffer:"), 3u);
}
TEST(DirectGLESStateGuards, FramebufferBindingShadowPinsOnceThenSkips) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// Cold path: one driver query pins the shadow; further reads are free.
(void)FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Read);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u);
(void)FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Read);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 1u);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 1u);
// GL_FRAMEBUFFER touches both targets; DRAW is still unknown so it must bind.
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 2u);
// Both halves now match: no further calls for either single target.
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_READ_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_FRAMEBUFFER, 7);
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 2u);
EXPECT_EQ(FramebufferImpl::CurrentFramebufferBinding(MobileGL::FramebufferTarget::Draw), 7u);
EXPECT_EQ(mocks.log.Count("GetIntegerv:"), 1u); // shadow answered, no new query
}
TEST(DirectGLESStateGuards, PackStateShadowAppliesMinimalDeltas) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// First application pins all four parameters.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{4, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 4u);
// Identical state: zero driver calls.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{4, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 4u);
// One field changed: exactly one driver call.
PixelStoreImpl::ApplyPackState(PixelStoreImpl::PackState{1, 0, 0, 0});
EXPECT_EQ(mocks.log.Count("PixelStorei:"), 5u);
const auto current = PixelStoreImpl::CurrentPackState();
EXPECT_EQ(current.Alignment, 1);
EXPECT_EQ(current.RowLength, 0);
EXPECT_EQ(current.SkipRows, 0);
EXPECT_EQ(current.SkipPixels, 0);
}
TEST(DirectGLESStateGuards, ScratchFBODetachesCrossAspectResidue) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::TempFramebuffer();
EXPECT_NE(ScratchFBOImpl::EnsureId(fb), 0u);
// A depth copy leaves a DEPTH_STENCIL attachment (the pre-fix code never
// detached it, wedging every later color readback through this FBO).
ScratchFBOImpl::EnsureDepthAttachment2D(fb, GL_DRAW_FRAMEBUFFER, 11, GL_TEXTURE_2D, 0, /*withStencil=*/true);
const MobileGL::String dsAttach = "FramebufferTexture2D:" + std::to_string(GL_DRAW_FRAMEBUFFER) + ":" +
std::to_string(GL_DEPTH_STENCIL_ATTACHMENT);
EXPECT_EQ(mocks.log.Count(dsAttach), 1u);
// The next color use must detach the stale depth-stencil attachment exactly once.
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
const MobileGL::String dsDetach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_DEPTH_STENCIL_ATTACHMENT) + ":" +
std::to_string(GL_TEXTURE_2D) + ":0:0";
const MobileGL::String colorAttach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_COLOR_ATTACHMENT0) + ":" +
std::to_string(GL_TEXTURE_2D) + ":22:0";
EXPECT_EQ(mocks.log.Count(dsDetach), 1u);
EXPECT_EQ(mocks.log.Count(colorAttach), 1u);
// Back-to-back identical color use: no driver traffic at all.
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
EXPECT_EQ(mocks.log.Count("FramebufferTexture2D:"), 0u);
}
TEST(DirectGLESStateGuards, ScratchFBOTextureDeletionForcesFullScrub) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::TempFramebuffer();
ScratchFBOImpl::EnsureId(fb);
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
// The attached texture id dies: the shadow can no longer vouch for the FBO
// (ES does not auto-detach from unbound FBOs, and the name may be recycled),
// so the next use must scrub and re-attach instead of skipping.
ScratchFBOImpl::NoteTextureIdDeleted(22);
mocks.log.calls.clear();
ScratchFBOImpl::EnsureColorAttachment2D(fb, GL_READ_FRAMEBUFFER, 22, GL_TEXTURE_2D, 0);
EXPECT_GE(mocks.log.Count("FramebufferTexture2D:"), 2u); // scrub (color + depth) ...
const MobileGL::String colorAttach = "FramebufferTexture2D:" + std::to_string(GL_READ_FRAMEBUFFER) + ":" +
std::to_string(GL_COLOR_ATTACHMENT0) + ":" +
std::to_string(GL_TEXTURE_2D) + ":22:0";
EXPECT_EQ(mocks.log.Count(colorAttach), 1u); // ... then the real re-attach
}
TEST(DirectGLESStateGuards, ScratchFBOReadDrawBufferStateCached) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
auto& fb = ScratchFBOImpl::BlitReadFramebuffer();
ScratchFBOImpl::EnsureId(fb);
// Fresh FBOs default to COLOR_ATTACHMENT0 for both buffers: no call needed.
ScratchFBOImpl::EnsureReadBuffer(fb, GL_COLOR_ATTACHMENT0);
EXPECT_EQ(mocks.log.Count("ReadBuffer:"), 0u);
// Depth blits want GL_NONE; the transition costs one call, repeats are free.
ScratchFBOImpl::EnsureReadBuffer(fb, GL_NONE);
ScratchFBOImpl::EnsureReadBuffer(fb, GL_NONE);
EXPECT_EQ(mocks.log.Count("ReadBuffer:"), 1u);
ScratchFBOImpl::EnsureDrawBuffer(fb, GL_NONE);
ScratchFBOImpl::EnsureDrawBuffer(fb, GL_NONE);
EXPECT_EQ(mocks.log.Count("DrawBuffers:"), 1u);
}
namespace {
MobileGL::Vector<GLuint>* g_deletedTextureIds = nullptr;
void SG_DeleteTextures(GLsizei count, const GLuint* textures) {
if (!g_deletedTextureIds) return;
for (GLsizei i = 0; i < count; ++i) g_deletedTextureIds->push_back(textures[i]);
}
// Clears the recording hook even when a gtest assertion unwinds the test body
// (a dangling pointer to the dead stack vector would corrupt later tests).
struct ScopedDeletedTextureRecording {
explicit ScopedDeletedTextureRecording(MobileGL::Vector<GLuint>& sink) { g_deletedTextureIds = &sink; }
~ScopedDeletedTextureRecording() { g_deletedTextureIds = nullptr; }
ScopedDeletedTextureRecording(const ScopedDeletedTextureRecording&) = delete;
ScopedDeletedTextureRecording& operator=(const ScopedDeletedTextureRecording&) = delete;
};
} // namespace
TEST(DirectGLESBackendTexture, DestructorDeletesIdAndScrubsBindingCache) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedDirectGLESTextureBindings scoped; // installs glGenTextures/glBindTexture mocks + resets caches
MobileGL::Vector<GLuint> deleted;
ScopedDeletedTextureRecording recording(deleted);
auto functions = g_GLESFuncs;
functions.glDeleteTextures = SG_DeleteTextures;
SetGLESFuncsTable(functions);
const auto texture2DSlot = static_cast<MobileGL::SizeT>(MobileGL::TextureTarget::Texture2D);
GLuint id = 0;
{
auto backendTexture = MobileGL::MakeShared<TextureImpl::BackendTextureObject>();
id = backendTexture->GetBackendTextureId();
ASSERT_NE(id, 0u);
backendTexture->Bind(GL_TEXTURE_2D, 0);
ASSERT_EQ(TextureImpl::g_boundTexturesCache[0][texture2DSlot], backendTexture.get());
}
// Frontend glDeleteTextures used to leak the backend id forever and leave the
// cache pointer dangling (heap-address reuse then false-skips a later Bind).
ASSERT_EQ(deleted.size(), 1u);
EXPECT_EQ(deleted[0], id);
EXPECT_EQ(TextureImpl::g_boundTexturesCache[0][texture2DSlot], nullptr);
// A wrapper whose context died must NOT delete a foreign (recycled) name.
{
auto backendTexture = MobileGL::MakeShared<TextureImpl::BackendTextureObject>();
++TextureImpl::g_textureContextGeneration;
backendTexture.reset();
--TextureImpl::g_textureContextGeneration; // restore for later tests
EXPECT_EQ(deleted.size(), 1u);
}
}
TEST(DirectGLESStateGuards, DefaultFramebufferBindGoesThroughShadow) {
using namespace MobileGL::MG_Backend::DirectGLES;
ScopedStateGuardMocks mocks;
// The regression this guards against: binding framebuffer 0 raw while the
// shadow keeps a user-FBO id makes the next re-bind of that FBO false-skip.
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7);
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 0); // default-FBO path must use this API
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7); // must reach the driver again
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 3u);
}
@@ -0,0 +1,25 @@
cmake_minimum_required(VERSION 3.14)
add_executable(
SpirvPassTest
SpirvPassTest.cpp
)
target_include_directories(SpirvPassTest PRIVATE
${MGL_ROOT}/include
${MGL_ROOT}/MobileGL
${MGL_ROOT}/3rdparty/SPIRV-Reflect
)
target_link_libraries(
SpirvPassTest PRIVATE
GTest::gtest_main
${LINK_LIBRARIES}
)
if (MSVC)
target_compile_options(SpirvPassTest PRIVATE /Zc:preprocessor)
endif()
include(GoogleTest)
gtest_discover_tests(SpirvPassTest DISCOVERY_TIMEOUT 30 PROPERTIES LABELS unit)
@@ -0,0 +1,170 @@
// MobileGL - MobileGL/MG_Test/ShaderTranspiler/SpirvPassTest.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include <gtest/gtest.h>
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <spirv_reflect.h>
using namespace MobileGL;
using MobileGL::MG_Util::ShaderTranspiler::ShaderCompiler;
namespace {
// glslangValidator -V output. Both are vertex shaders writing gl_Position through
// the gl_PerVertex block, i.e. the Position builtin arrives as OpMemberDecorate rather
// than a plain OpDecorate - the shape real glslang output actually takes.
// #version 450
// layout(location = 0) in vec4 inPos;
// void main() { gl_Position = inPos; }
constexpr Uint32 kPlainVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x00000015u, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u, 0x00000000u,
0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu, 0x006e6f69u,
0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu, 0x657a6953u,
0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u, 0x4470696cu,
0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u, 0x435f6c67u,
0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du, 0x00000000u,
0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00030047u, 0x0000000bu,
0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu, 0x00000000u,
0x00050048u, 0x0000000bu, 0x00000001u, 0x0000000bu, 0x00000001u, 0x00050048u,
0x0000000bu, 0x00000002u, 0x0000000bu, 0x00000003u, 0x00050048u, 0x0000000bu,
0x00000003u, 0x0000000bu, 0x00000004u, 0x00040047u, 0x00000011u, 0x0000001eu,
0x00000000u, 0x00020013u, 0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u,
0x00030016u, 0x00000006u, 0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u,
0x00000004u, 0x00040015u, 0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu,
0x00000008u, 0x00000009u, 0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u,
0x00000009u, 0x0006001eu, 0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au,
0x0000000au, 0x00040020u, 0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu,
0x0000000cu, 0x0000000du, 0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u,
0x00000001u, 0x0004002bu, 0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u,
0x00000010u, 0x00000001u, 0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u,
0x00000001u, 0x00040020u, 0x00000013u, 0x00000003u, 0x00000007u, 0x00050036u,
0x00000002u, 0x00000004u, 0x00000000u, 0x00000003u, 0x000200f8u, 0x00000005u,
0x0004003du, 0x00000007u, 0x00000012u, 0x00000011u, 0x00050041u, 0x00000013u,
0x00000014u, 0x0000000du, 0x0000000fu, 0x0003003eu, 0x00000014u, 0x00000012u,
0x000100fdu, 0x00010038u,
};
// ... plus `invariant gl_Position;` - already carries OpMemberDecorate %gl_PerVertex 0
// Invariant, so the pass must not add a duplicate.
constexpr Uint32 kAlreadyInvariantVertexSpirv[] = {
0x07230203u, 0x00010000u, 0x0008000bu, 0x00000015u, 0x00000000u, 0x00020011u,
0x00000001u, 0x0006000bu, 0x00000001u, 0x4c534c47u, 0x6474732eu, 0x3035342eu,
0x00000000u, 0x0003000eu, 0x00000000u, 0x00000001u, 0x0007000fu, 0x00000000u,
0x00000004u, 0x6e69616du, 0x00000000u, 0x0000000du, 0x00000011u, 0x00030003u,
0x00000002u, 0x000001c2u, 0x00040005u, 0x00000004u, 0x6e69616du, 0x00000000u,
0x00060005u, 0x0000000bu, 0x505f6c67u, 0x65567265u, 0x78657472u, 0x00000000u,
0x00060006u, 0x0000000bu, 0x00000000u, 0x505f6c67u, 0x7469736fu, 0x006e6f69u,
0x00070006u, 0x0000000bu, 0x00000001u, 0x505f6c67u, 0x746e696fu, 0x657a6953u,
0x00000000u, 0x00070006u, 0x0000000bu, 0x00000002u, 0x435f6c67u, 0x4470696cu,
0x61747369u, 0x0065636eu, 0x00070006u, 0x0000000bu, 0x00000003u, 0x435f6c67u,
0x446c6c75u, 0x61747369u, 0x0065636eu, 0x00030005u, 0x0000000du, 0x00000000u,
0x00040005u, 0x00000011u, 0x6f506e69u, 0x00000073u, 0x00030047u, 0x0000000bu,
0x00000002u, 0x00050048u, 0x0000000bu, 0x00000000u, 0x0000000bu, 0x00000000u,
0x00040048u, 0x0000000bu, 0x00000000u, 0x00000012u, 0x00050048u, 0x0000000bu,
0x00000001u, 0x0000000bu, 0x00000001u, 0x00050048u, 0x0000000bu, 0x00000002u,
0x0000000bu, 0x00000003u, 0x00050048u, 0x0000000bu, 0x00000003u, 0x0000000bu,
0x00000004u, 0x00040047u, 0x00000011u, 0x0000001eu, 0x00000000u, 0x00020013u,
0x00000002u, 0x00030021u, 0x00000003u, 0x00000002u, 0x00030016u, 0x00000006u,
0x00000020u, 0x00040017u, 0x00000007u, 0x00000006u, 0x00000004u, 0x00040015u,
0x00000008u, 0x00000020u, 0x00000000u, 0x0004002bu, 0x00000008u, 0x00000009u,
0x00000001u, 0x0004001cu, 0x0000000au, 0x00000006u, 0x00000009u, 0x0006001eu,
0x0000000bu, 0x00000007u, 0x00000006u, 0x0000000au, 0x0000000au, 0x00040020u,
0x0000000cu, 0x00000003u, 0x0000000bu, 0x0004003bu, 0x0000000cu, 0x0000000du,
0x00000003u, 0x00040015u, 0x0000000eu, 0x00000020u, 0x00000001u, 0x0004002bu,
0x0000000eu, 0x0000000fu, 0x00000000u, 0x00040020u, 0x00000010u, 0x00000001u,
0x00000007u, 0x0004003bu, 0x00000010u, 0x00000011u, 0x00000001u, 0x00040020u,
0x00000013u, 0x00000003u, 0x00000007u, 0x00050036u, 0x00000002u, 0x00000004u,
0x00000000u, 0x00000003u, 0x000200f8u, 0x00000005u, 0x0004003du, 0x00000007u,
0x00000012u, 0x00000011u, 0x00050041u, 0x00000013u, 0x00000014u, 0x0000000du,
0x0000000fu, 0x0003003eu, 0x00000014u, 0x00000012u, 0x000100fdu, 0x00010038u,
};
// OpMemberDecorate <struct-id> <member> <decoration>
constexpr Uint32 kOpMemberDecorate = 72;
constexpr Uint32 kDecorationInvariant = 18;
constexpr Uint32 kSpirvHeaderWordCount = 5;
// Test-side reference walker. Deliberately independent of the production code so a bug in
// the pass cannot hide behind the same helper; only used to count what the pass emitted.
Uint32 CountInvariantMemberDecorations(const Vector<Uint32>& spirv) {
Uint32 count = 0;
for (SizeT i = kSpirvHeaderWordCount; i < spirv.size();) {
const Uint32 wordCount = spirv[i] >> 16;
const Uint32 opcode = spirv[i] & 0xFFFFu;
if (wordCount == 0 || i + wordCount > spirv.size()) {
break;
}
if (opcode == kOpMemberDecorate && wordCount >= 4 && spirv[i + 3] == kDecorationInvariant) {
++count;
}
i += wordCount;
}
return count;
}
template <SizeT WordCount>
Vector<Uint32> ToVector(const Uint32 (&words)[WordCount]) {
return Vector<Uint32>(words, words + WordCount);
}
} // namespace
// --- DecoratePositionInvariantPass ---
TEST(DecoratePositionInvariant, AddsInvariantToThePositionMember) {
const Vector<Uint32> input = ToVector(kPlainVertexSpirv);
ASSERT_EQ(CountInvariantMemberDecorations(input), 0u);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(input, output));
EXPECT_EQ(CountInvariantMemberDecorations(output), 1u);
}
TEST(DecoratePositionInvariant, DoesNotDuplicateAnExistingInvariant) {
const Vector<Uint32> input = ToVector(kAlreadyInvariantVertexSpirv);
ASSERT_EQ(CountInvariantMemberDecorations(input), 1u);
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(input, output));
EXPECT_EQ(CountInvariantMemberDecorations(output), 1u);
// The pass reports SuccessWithoutChange here, and SPIRV-Tools asserts (in assert-enabled
// builds) that such a run round-trips byte-identically. Pin that from the outside so an
// assert-enabled CI build cannot be the first thing to discover a violation.
EXPECT_EQ(output, input);
}
TEST(DecoratePositionInvariant, IsIdempotent) {
Vector<Uint32> once;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(ToVector(kPlainVertexSpirv), once));
Vector<Uint32> twice;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(once, twice));
EXPECT_EQ(CountInvariantMemberDecorations(twice), 1u);
}
TEST(DecoratePositionInvariant, OutputStaysAReflectableModule) {
Vector<Uint32> output;
ASSERT_TRUE(ShaderCompiler::DecoratePositionInvariantForVulkan(ToVector(kPlainVertexSpirv), output));
SpvReflectShaderModule module{};
ASSERT_EQ(spvReflectCreateShaderModule(output.size() * sizeof(Uint32), output.data(), &module),
SPV_REFLECT_RESULT_SUCCESS);
EXPECT_EQ(module.entry_point_count, 1u);
spvReflectDestroyShaderModule(&module);
}
TEST(DecoratePositionInvariant, RejectsGarbageInput) {
const Vector<Uint32> notSpirv{0xdeadbeefu, 0u, 0u, 0u, 0u};
Vector<Uint32> output;
EXPECT_FALSE(ShaderCompiler::DecoratePositionInvariantForVulkan(notSpirv, output));
}
+195
View File
@@ -318,6 +318,49 @@ TEST_F(TextureTest, CopyTextureSubImage2DUsesNamedObjectAndRestoresBinding) {
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
TEST_F(TextureTest, CopyTextureSubImage2DRejectsCubeMapTargets) {
const ScopedTextureBackendFunctionsOverride backendGuard;
MG_Backend::gBackendFunctionsTable.GL.CopyTexSubImage2D = RecordCopyTexSubImage2D;
g_copyTexSubImage2DCall = {};
GLuint cubeTexture = 0;
MG_Impl::GLImpl::CreateTextures(GL_TEXTURE_CUBE_MAP, 1, &cubeTexture);
MG_Impl::GLImpl::CopyTextureSubImage2D(cubeTexture, 0, 0, 0, 0, 0, 1, 1);
// GL 4.6 sec. 8.8: the 2D form only accepts 2D/1D-array/rectangle effective targets.
EXPECT_FALSE(g_copyTexSubImage2DCall.Called);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
}
TEST_F(TextureTest, ClearTexImageErrorContracts) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 2, 2, 0,
GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
// Zero texture name is INVALID_OPERATION (ARB_clear_texture).
MG_Impl::GLImpl::ClearTexImage(0, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
// A negative level is INVALID_VALUE...
MG_Impl::GLImpl::ClearTexImage(texture, -1, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_VALUE));
// ...but clearing a level that was never defined is INVALID_OPERATION.
MG_Impl::GLImpl::ClearTexImage(texture, 5, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_OPERATION));
// A clear region outside the level is INVALID_VALUE.
MG_Impl::GLImpl::ClearTexSubImage(texture, 0, 1, 1, 0, 4, 4, 1,
GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_VALUE));
// An invalid pixel-transfer format is INVALID_ENUM from the shared validators.
MG_Impl::GLImpl::ClearTexImage(texture, 0, GL_NONE, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), static_cast<GLenum>(GL_INVALID_ENUM));
}
// GL_MAX_TEXTURE_MAX_ANISOTROPY_EXT is float state that must answer every numeric query: GetFloatv
// is authoritative and GetIntegerv would otherwise fall through to its INVALID_ENUM default.
TEST_F(TextureTest, MaxTextureMaxAnisotropyIsAnsweredFromTheBackendLimit) {
@@ -1384,6 +1427,158 @@ TEST_F(TextureTest, TextureStorage1DAndSubImageModifyNamedObjectOnly) {
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// Building a mip chain top-down - upload level N, then level 0 - must not destroy the levels
// already uploaded. AllocateLevel used to resize() the storage down to level+1 on every call, so
// the level-0 upload truncated the chain to a single level; the higher level then read back as
// {0,0,0}, IsComplete() rejected the zero-then-nonzero pattern, and DirectGLES answered that by
// skipping the texture's sync entirely. This is the shape KHR-GL33.texture_repeat_mode uses, and
// it accounted for 108 CTS failures in every GL version.
TEST_F(TextureTest, TexImage2DOnLevelZeroKeepsAnAlreadyUploadedHigherLevel) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 49, 23, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 98, 46, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 2u);
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 0), IntVec3(98, 46, 1));
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 1), IntVec3(49, 23, 1));
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// The other half of the contract: respecifying a level 0 that already held an image still drops
// the chain, exactly as before. Minecraft rebinds the block-atlas name and calls glTexImage2D on
// level 0 before uploading the new levels; leaving the previous chain in place would strand a tail
// at the wrong sizes and - because Mojang terminates its chains with a 0x0 level - reproduce the
// same incomplete-texture black atlas the fix above exists to prevent.
TEST_F(TextureTest, TexImage2DRespecifyingAnExistingLevelZeroDropsTheStaleChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 2, 2, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
ASSERT_EQ(mipmapObject->GetMipmapLevelCount(), 3u);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 16, 16, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 1u);
EXPECT_EQ(mipmapObject->GetMipmapTexelSize(TextureUploadTarget::Texture2D, 0), IntVec3(16, 16, 1));
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// Same-size respecification has to drop the chain too. The Mipmap Levels video setting rebuilds
// the atlas at identical dimensions with a different level count, so a size-change-only test would
// let the old tail survive.
TEST_F(TextureTest, TexImage2DRespecifyingLevelZeroAtTheSameSizeStillDropsTheChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 1u);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// glTexStorage2D defines exactly `levels` levels. AllocateStorage only grows now, so the immutable
// path has to drop a longer pre-existing chain explicitly.
TEST_F(TextureTest, TexStorage2DTrimsALongerPreExistingMipChain) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, GL_RGBA8, 8, 8, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 1, GL_RGBA8, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 2, 2, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 3, GL_RGBA8, 1, 1, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
MG_Impl::GLImpl::TexStorage2D(GL_TEXTURE_2D, 2, GL_RGBA8, 8, 8);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
auto* mipmapObject = static_cast<MG_State::GLState::TextureObjectMipmap*>(textureObject.get());
ASSERT_NE(mipmapObject, nullptr);
EXPECT_EQ(mipmapObject->GetMipmapLevelCount(), 2u);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
// glTexImage2D used to reject every GL_COMPRESSED_* internal format with GL_INVALID_ENUM, because
// none of them mapped to a TextureInternalFormat and the "unknown format" gate fired. They now
// resolve to the uncompressed storage that backs them - what GL prescribes for the generic formats,
// and a deliberate deviation for RGTC, which ES cannot compress. The (format, type) pairs below are
// the ones KHR-GL33.packed_pixels uploads with, so this table doubles as a pin for those 480 cases.
TEST_F(TextureTest, CompressedInternalFormatsResolveToTheirUncompressedStorage) {
struct Case {
GLenum internalFormat;
GLenum format;
GLenum type;
TextureInternalFormat expected;
};
const Case cases[] = {
{GL_COMPRESSED_RED, GL_RED, GL_UNSIGNED_BYTE, TextureInternalFormat::R8},
{GL_COMPRESSED_RG, GL_RG, GL_UNSIGNED_BYTE, TextureInternalFormat::RG8},
{GL_COMPRESSED_RGB, GL_RGB, GL_UNSIGNED_BYTE, TextureInternalFormat::RGB8},
{GL_COMPRESSED_RGBA, GL_RGBA, GL_UNSIGNED_BYTE, TextureInternalFormat::RGBA8},
{GL_COMPRESSED_SRGB, GL_RGB, GL_UNSIGNED_BYTE, TextureInternalFormat::SRGB8},
{GL_COMPRESSED_SRGB_ALPHA, GL_RGBA, GL_UNSIGNED_BYTE, TextureInternalFormat::SRGB8Alpha8},
{GL_COMPRESSED_RED_RGTC1, GL_RED, GL_UNSIGNED_BYTE, TextureInternalFormat::R8},
{GL_COMPRESSED_RG_RGTC2, GL_RG, GL_UNSIGNED_BYTE, TextureInternalFormat::RG8},
// The signed RGTC pair is uploaded as GL_BYTE and must land on SNORM storage - resolving
// them to plain R8/RG8 would silently reinterpret negative texels.
{GL_COMPRESSED_SIGNED_RED_RGTC1, GL_RED, GL_BYTE, TextureInternalFormat::R8Snorm},
{GL_COMPRESSED_SIGNED_RG_RGTC2, GL_RG, GL_BYTE, TextureInternalFormat::RG8Snorm},
};
MG_Impl::GLImpl::PixelStorei(GL_UNPACK_ALIGNMENT, 1);
for (const auto& c : cases) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_2D, texture);
MG_Impl::GLImpl::TexImage2D(GL_TEXTURE_2D, 0, c.internalFormat, 4, 4, 0, c.format, c.type, nullptr);
const auto textureObject = MG_State::pGLContext->GetTextureObject(texture);
ASSERT_NE(textureObject, nullptr) << "internalFormat 0x" << std::hex << c.internalFormat;
EXPECT_EQ(textureObject->GetFormat(), c.expected) << "internalFormat 0x" << std::hex << c.internalFormat;
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR) << "internalFormat 0x" << std::hex << c.internalFormat;
}
}
// RGTC compresses 4x4 blocks of a 2D image and has no 3D form, so glTexImage3D must reject it even
// though the same enum is accepted on a 2D target. The generic compressed formats carry no such
// restriction and stay legal in 3D.
TEST_F(TextureTest, RgtcInternalFormatsAreRejectedOnThreeDimensionalTargets) {
const GLenum rgtc[] = {GL_COMPRESSED_RED_RGTC1, GL_COMPRESSED_SIGNED_RED_RGTC1, GL_COMPRESSED_RG_RGTC2,
GL_COMPRESSED_SIGNED_RG_RGTC2};
for (const GLenum internalFormat : rgtc) {
GLuint texture = 0;
MG_Impl::GLImpl::GenTextures(1, &texture);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_3D, texture);
MG_Impl::GLImpl::TexImage3D(GL_TEXTURE_3D, 0, internalFormat, 4, 4, 4, 0, GL_RED, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_INVALID_OPERATION)
<< "internalFormat 0x" << std::hex << internalFormat;
}
GLuint generic = 0;
MG_Impl::GLImpl::GenTextures(1, &generic);
MG_Impl::GLImpl::BindTexture(GL_TEXTURE_3D, generic);
MG_Impl::GLImpl::TexImage3D(GL_TEXTURE_3D, 0, GL_COMPRESSED_RGBA, 4, 4, 4, 0, GL_RGBA, GL_UNSIGNED_BYTE, nullptr);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
}
TEST_F(TextureTest, TextureStorage3DAndSubImageModifyNamedObjectOnly) {
GLuint texture = 0;
MG_Impl::GLImpl::CreateTextures(GL_TEXTURE_3D, 1, &texture);
@@ -811,6 +811,9 @@ namespace MobileGL::MG_Util::BackendLoader {
if (std::strcmp(extension, "GL_EXT_blend_func_extended") == 0) {
caps.SupportsDualSourceBlend = true;
}
if (std::strcmp(extension, "GL_NV_shader_noperspective_interpolation") == 0) {
caps.SupportsNoperspectiveInterpolation = true;
}
}
}
@@ -1050,6 +1050,11 @@ namespace MobileGL {
// factors and layout(index = 1) fragment outputs. GLES core has no dual-source blending,
// so without this a draw using a SRC1 factor cannot proceed.
Bool SupportsDualSourceBlend = false;
// GL_NV_shader_noperspective_interpolation is present: the driver accepts the
// `noperspective` interpolation qualifier in ESSL. GLES core has none, so without this
// SPIRV-Cross's `#extension ... : require` would fail to compile and MobileGL falls back
// to stripping the NoPerspective decoration (smooth interpolation) via StripNoPerspectivePass.
Bool SupportsNoperspectiveInterpolation = false;
// GL_RENDERER contains "ANGLE".
Bool IsAngleRenderer = false;
// GL_RENDERER contains both "ANGLE" and "llvmpipe".
@@ -121,6 +121,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.VulkanAPIVersion = DecodeApiVersion(p.apiVersion);
caps.DeviceName = p.deviceName;
caps.DriverVersionString = DecodeDriverVersion(p.driverVersion);
caps.VendorId = p.vendorID;
caps.UniformBufferOffsetAlignment = static_cast<int>(p.limits.minUniformBufferOffsetAlignment);
caps.AliasedLineWidthRangeMin = p.limits.lineWidthRange[0];
caps.AliasedLineWidthRangeMax = p.limits.lineWidthRange[1];
@@ -210,6 +211,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.VulkanAPIVersion = DecodeApiVersion(properties.apiVersion);
caps.DeviceName = properties.deviceName;
caps.DriverVersionString = DecodeDriverVersion(properties.driverVersion);
caps.VendorId = properties.vendorID;
caps.UniformBufferOffsetAlignment = static_cast<int>(properties.limits.minUniformBufferOffsetAlignment);
caps.AliasedLineWidthRangeMin = properties.limits.lineWidthRange[0];
caps.AliasedLineWidthRangeMax = properties.limits.lineWidthRange[1];
@@ -15,6 +15,8 @@ namespace MobileGL {
Version VulkanAPIVersion{1, 0, 0};
String DeviceName;
String DriverVersionString;
// VkPhysicalDeviceProperties::vendorID, for device-quirk vendor gating.
Uint32 VendorId = 0;
Int UniformBufferOffsetAlignment = 256;
Float AliasedLineWidthRangeMin = 1.0f;
Float AliasedLineWidthRangeMax = 1.0f;
@@ -255,6 +255,37 @@ namespace MobileGL {
return TextureInternalFormat::DepthComponent;
case GL_DEPTH_STENCIL:
return TextureInternalFormat::DepthStencil;
// Compressed internal formats resolve to the uncompressed storage that backs them.
//
// For the six generic formats this is exactly what GL prescribes: the implementation
// picks a specific compressed format, and when none is available it falls back to the
// corresponding base format. Nothing downstream ever sees a compressed enum, so the
// metrics, pixel-store and backend tables keep their "one format, N bytes per texel"
// invariant instead of each needing a compressed-aware arm.
//
// The four RGTC formats are a deliberate deviation: they are specific formats that GL
// 3.3 requires, but ES exposes no RGTC compressor to hand the data to. Storing the
// texels uncompressed keeps them renderable at the cost of the memory saving, which is
// strictly better than the INVALID_ENUM the application used to get. Note the signed
// variants must land on SNORM storage - CTS uploads them as GL_BYTE.
case GL_COMPRESSED_RED:
case GL_COMPRESSED_RED_RGTC1:
return TextureInternalFormat::R8;
case GL_COMPRESSED_SIGNED_RED_RGTC1:
return TextureInternalFormat::R8Snorm;
case GL_COMPRESSED_RG:
case GL_COMPRESSED_RG_RGTC2:
return TextureInternalFormat::RG8;
case GL_COMPRESSED_SIGNED_RG_RGTC2:
return TextureInternalFormat::RG8Snorm;
case GL_COMPRESSED_RGB:
return TextureInternalFormat::RGB8;
case GL_COMPRESSED_RGBA:
return TextureInternalFormat::RGBA8;
case GL_COMPRESSED_SRGB:
return TextureInternalFormat::SRGB8;
case GL_COMPRESSED_SRGB_ALPHA:
return TextureInternalFormat::SRGB8Alpha8;
case GL_ALPHA:
case GL_RED:
return TextureInternalFormat::Red;
+225
View File
@@ -401,6 +401,230 @@ namespace MobileGL::MG_Util::SelfTest {
disabledNote);
}
// Compiles + links a two-stage program on the probe context. Returns 0 on failure and writes a
// human-readable reason into |detail|.
GLuint CompileLinkProgram(const MG_External::GLESFunctionsTable& g, const char* vs, const char* fs,
String& detail) {
const auto compile = [&](GLenum stage, const char* src, GLuint& out) -> bool {
out = g.glCreateShader(stage);
if (out == 0) {
detail = "glCreateShader returned 0";
return false;
}
g.glShaderSource(out, 1, &src, nullptr);
g.glCompileShader(out);
GLint ok = GL_FALSE;
g.glGetShaderiv(out, GL_COMPILE_STATUS, &ok);
if (ok != GL_TRUE) {
GLchar log[512] = {};
GLsizei len = 0;
g.glGetShaderInfoLog(out, static_cast<GLsizei>(sizeof(log) - 1), &len, log);
detail = format("{} shader compile failed: {}",
stage == GL_VERTEX_SHADER ? "vertex" : "fragment",
len > 0 ? log : "(no info log)");
return false;
}
return true;
};
GLuint v = 0, f = 0;
const ScopeGuard delV([&]() { if (v) g.glDeleteShader(v); });
const ScopeGuard delF([&]() { if (f) g.glDeleteShader(f); });
if (!compile(GL_VERTEX_SHADER, vs, v) || !compile(GL_FRAGMENT_SHADER, fs, f)) {
return 0;
}
const GLuint prog = g.glCreateProgram();
if (prog == 0) {
detail = "glCreateProgram returned 0";
return 0;
}
g.glAttachShader(prog, v);
g.glAttachShader(prog, f);
g.glLinkProgram(prog);
GLint linked = GL_FALSE;
g.glGetProgramiv(prog, GL_LINK_STATUS, &linked);
if (linked != GL_TRUE) {
detail = "program link failed";
g.glDeleteProgram(prog);
return 0;
}
return prog;
}
// "noperspective interpolation" row - a real correctness render, not just a compile. A viewport-
// filling quad is drawn with strong perspective (left clip-w 1, right clip-w 8) and a varying that
// runs 0..1 across it. At the screen centre screen-linear interpolation gives 0.5 while perspective-
// correct gives 1/(w+1) ~= 0.11, so reading the centre texel tells the two apart. The varying is
// carried either through the native `noperspective` qualifier (extension present) or through the
// exact gl_Position.w / gl_FragCoord.w rewrite MobileGL applies when it is absent. Verdict:
// PASS - extension present and the native noperspective result is screen-linear;
// WARN - extension absent but the gl_Position.w/gl_FragCoord.w emulation renders screen-linear
// (correct, just the fallback path shipping shader packs hit on such devices);
// FAIL - either path renders perspective-correct / wrong (noperspective does not actually work),
// or the program will not compile/link, or the render errors.
// Requires the probe context to still be current.
void ProbeGlesNoperspective(ReportBuilder& builder, const MG_External::GLESCapabilities& caps,
const MG_External::GLESFunctionsTable& g) {
const Bool native = caps.SupportsNoperspectiveInterpolation;
const String pathNote = native ? "GL_NV_shader_noperspective_interpolation present (native path)"
: "GL_NV_shader_noperspective_interpolation absent (gl_Position.w / "
"gl_FragCoord.w emulation path)";
const auto fail = [&](const String& detail) {
builder.Fail("noperspective interpolation", pathNote + "; " + detail);
};
if (!g.glCreateShader || !g.glShaderSource || !g.glCompileShader || !g.glGetShaderiv ||
!g.glGetShaderInfoLog || !g.glDeleteShader || !g.glCreateProgram || !g.glAttachShader ||
!g.glLinkProgram || !g.glGetProgramiv || !g.glUseProgram || !g.glDeleteProgram ||
!g.glGenFramebuffers || !g.glBindFramebuffer || !g.glDeleteFramebuffers ||
!g.glGenRenderbuffers || !g.glBindRenderbuffer || !g.glRenderbufferStorage ||
!g.glFramebufferRenderbuffer || !g.glDeleteRenderbuffers || !g.glCheckFramebufferStatus ||
!g.glGenBuffers || !g.glBindBuffer || !g.glBufferData || !g.glDeleteBuffers ||
!g.glGetAttribLocation || !g.glVertexAttribPointer || !g.glEnableVertexAttribArray ||
!g.glViewport || !g.glClearColor || !g.glClear || !g.glDrawArrays || !g.glReadPixels ||
!g.glFinish || !g.glGetError) {
fail("the render entry points did not resolve through eglGetProcAddress");
return;
}
// Match MobileGL's own ESSL target (the device's version). At #version 300 es some drivers
// (Adreno) still treat `noperspective` as reserved even with the extension enabled; the ES 3.2
// form the backend actually emits compiles. Emulated shaders are version-agnostic but use the
// same header for consistency.
const Int esslVer = caps.GLESVersion.Major * 100 + caps.GLESVersion.Minor * 10;
const String header = format("#version {} es\n", esslVer >= 300 ? esslVer : 300);
static const char* const kVsNativeBody =
"#extension GL_NV_shader_noperspective_interpolation : require\n"
"in vec4 a_pos;\n"
"in float a_v;\n"
"noperspective out highp float v_out;\n"
"void main() { gl_Position = a_pos; v_out = a_v; }\n";
static const char* const kFsNativeBody =
"#extension GL_NV_shader_noperspective_interpolation : require\n"
"precision highp float;\n"
"noperspective in highp float v_out;\n"
"out vec4 fragColor;\n"
"void main() { fragColor = vec4(v_out, 0.0, 0.0, 1.0); }\n";
// Exactly MobileGL's emulation (verified against EmulateNoPerspectivePass output): pre-multiply
// the varying by clip-w in the vertex stage, recover with gl_FragCoord.w in the fragment stage,
// no noperspective qualifier (so the driver interpolates it perspective-correct).
static const char* const kVsEmuBody =
"in vec4 a_pos;\n"
"in float a_v;\n"
"out highp float v_out;\n"
"void main() { gl_Position = a_pos; v_out = a_v * gl_Position.w; }\n";
static const char* const kFsEmuBody =
"precision highp float;\n"
"in highp float v_out;\n"
"out vec4 fragColor;\n"
"void main() { fragColor = vec4(v_out * gl_FragCoord.w, 0.0, 0.0, 1.0); }\n";
while (g.glGetError() != GL_NO_ERROR) {
}
const String vsSrc = header + (native ? kVsNativeBody : kVsEmuBody);
const String fsSrc = header + (native ? kFsNativeBody : kFsEmuBody);
String linkDetail;
const GLuint prog = CompileLinkProgram(g, vsSrc.c_str(), fsSrc.c_str(), linkDetail);
if (prog == 0) {
fail(native ? "a noperspective program failed to build though the extension is advertised: " +
linkDetail
: "the emulation program failed to build: " + linkDetail);
return;
}
const ScopeGuard delProg([&]() { g.glDeleteProgram(prog); });
// 9x9 so the centre texel (4,4) sits exactly at NDC (0,0).
constexpr GLsizei kDim = 9;
GLuint rbo = 0, fbo = 0, vbo = 0;
g.glGenRenderbuffers(1, &rbo);
const ScopeGuard delRbo([&]() { if (rbo) g.glDeleteRenderbuffers(1, &rbo); });
g.glBindRenderbuffer(GL_RENDERBUFFER, rbo);
g.glRenderbufferStorage(GL_RENDERBUFFER, GL_RGBA8, kDim, kDim);
g.glGenFramebuffers(1, &fbo);
const ScopeGuard delFbo([&]() {
if (fbo) {
g.glBindFramebuffer(GL_FRAMEBUFFER, 0);
g.glDeleteFramebuffers(1, &fbo);
}
});
g.glBindFramebuffer(GL_FRAMEBUFFER, fbo);
g.glFramebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, rbo);
if (g.glCheckFramebufferStatus(GL_FRAMEBUFFER) != GL_FRAMEBUFFER_COMPLETE) {
fail("the probe framebuffer is incomplete");
return;
}
// Interleaved [vec4 clip-pos, float v]. Left w=1, right w=8; x/y pre-multiplied by w so the quad
// still fills NDC after the perspective divide.
const GLfloat verts[] = {
-1.f, -1.f, 0.f, 1.f, 0.f, //
8.f, -8.f, 0.f, 8.f, 1.f, //
-1.f, 1.f, 0.f, 1.f, 0.f, //
8.f, 8.f, 0.f, 8.f, 1.f, //
};
g.glGenBuffers(1, &vbo);
const ScopeGuard delVbo([&]() { if (vbo) g.glDeleteBuffers(1, &vbo); });
g.glBindBuffer(GL_ARRAY_BUFFER, vbo);
g.glBufferData(GL_ARRAY_BUFFER, sizeof(verts), verts, GL_STATIC_DRAW);
g.glUseProgram(prog);
const GLint posLoc = g.glGetAttribLocation(prog, "a_pos");
const GLint vLoc = g.glGetAttribLocation(prog, "a_v");
if (posLoc < 0 || vLoc < 0) {
fail("the probe vertex attributes did not resolve");
return;
}
g.glEnableVertexAttribArray(static_cast<GLuint>(posLoc));
g.glVertexAttribPointer(static_cast<GLuint>(posLoc), 4, GL_FLOAT, GL_FALSE, 5 * sizeof(GLfloat),
reinterpret_cast<const void*>(0));
g.glEnableVertexAttribArray(static_cast<GLuint>(vLoc));
g.glVertexAttribPointer(static_cast<GLuint>(vLoc), 1, GL_FLOAT, GL_FALSE, 5 * sizeof(GLfloat),
reinterpret_cast<const void*>(4 * sizeof(GLfloat)));
g.glViewport(0, 0, kDim, kDim);
g.glClearColor(0.f, 0.f, 0.f, 1.f);
g.glClear(GL_COLOR_BUFFER_BIT);
g.glDrawArrays(GL_TRIANGLE_STRIP, 0, 4);
g.glFinish();
const GLenum drawError = g.glGetError();
if (drawError != GL_NO_ERROR) {
fail(format("GL error 0x{:x} while rendering the probe quad", drawError));
return;
}
GLubyte center[4] = {};
g.glReadPixels(kDim / 2, kDim / 2, 1, 1, GL_RGBA, GL_UNSIGNED_BYTE, center);
const GLenum readError = g.glGetError();
if (readError != GL_NO_ERROR) {
fail(format("GL error 0x{:x} while reading the probe pixel back", readError));
return;
}
// At the centre: screen-linear -> 0.5 (~128); perspective-correct -> 1/(8+1) ~= 0.111 (~28).
const float observed = static_cast<float>(center[0]) / 255.0f;
const int observedByte = center[0];
constexpr float kScreenLinear = 0.5f;
const bool screenLinear = observed > 0.5f * (kScreenLinear + 1.0f / 9.0f); // midpoint ~= 0.306
if (!screenLinear) {
fail(format("the centre texel read {} (~{:.3f}); expected the screen-linear ~0.5 - "
"interpolation came out perspective-correct, so noperspective does not work here",
observedByte, observed));
return;
}
if (native) {
builder.Pass("noperspective interpolation",
pathNote + format("; native noperspective renders screen-linear (centre {} ~= 0.5)",
observedByte));
} else {
builder.Warn("noperspective interpolation",
pathNote +
format("; the emulation renders screen-linear correctly (centre {} ~= 0.5), "
"but this is the fallback path with less driver coverage",
observedByte));
}
}
// Everything the "MobileGL reported ..." rows need from the GLES device probe.
struct GlesProbeSummary {
Bool capsValid = false;
@@ -527,6 +751,7 @@ namespace MobileGL::MG_Util::SelfTest {
builder.report.rendererInfo = format("{} ({})", caps.GLESRendererString, caps.GLESVersionString);
EvaluateGlesChecklist(builder, caps, glesFuncs);
ProbeGlesTimerQuery(builder, caps, glesFuncs);
ProbeGlesNoperspective(builder, caps, glesFuncs);
builder.report.formatCapabilities.emplace();
MG_Backend::DirectGLES::PopulateFormatCapabilities(
glesFuncs, caps, builder.report.formatCapabilities.value());
@@ -16,9 +16,12 @@
#include "SpirvPasses/FlattenInterfaceStructPass.h"
#include "SpirvPasses/RenameSamplerFunctionParameterPass.h"
#include "SpirvPasses/DecomposeWorkgroupVec3Pass.h"
#include "SpirvPasses/DecoratePositionInvariantPass.h"
#include "SpirvPasses/LowerDrawParametersPass.h"
#include "SpirvPasses/RebaseInstanceIndexPass.h"
#include "SpirvPasses/StripUboMemberRelaxedPrecisionPass.h"
#include "SpirvPasses/StripNoPerspectivePass.h"
#include "SpirvPasses/EmulateNoPerspectivePass.h"
#include "spirv-tools/libspirv.h"
#include "spirv-tools/optimizer.hpp"
@@ -26,6 +29,7 @@
#include <MG_Backend/BackendObjects.h>
#include <MG_Util/Converters/GLToStr/GLEnumConverter.h>
#include <MG_Util/Converters/GLToGlslang/ProgramEnumConverter.h>
#include <cstdlib>
namespace MobileGL {
namespace MG_Util {
@@ -332,6 +336,30 @@ namespace MobileGL {
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::StripNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(StripNoPerspectivePass::CreateStripNoPerspectivePass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::EmulateNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(EmulateNoPerspectivePass::CreateEmulateNoPerspectivePass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
@@ -344,6 +372,18 @@ namespace MobileGL {
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary) {
using namespace spvtools;
OptimizerOptions options;
options.set_run_validator(false);
Optimizer optimizer(SPV_ENV_VULKAN_1_1);
optimizer.RegisterPass(DecoratePositionInvariantPass::CreateDecoratePositionInvariantPass());
return optimizer.Run(inputBinary.data(), inputBinary.size(), &outputBinary, options);
}
bool ShaderCompiler::UseUnformattedFloatStorageImagesForVulkan(
const Vector<Uint32>& inputBinary, Vector<uint32_t>& outputBinary) {
constexpr SizeT kSpirvHeaderWordCount = 5;
@@ -33,12 +33,29 @@ namespace MobileGL {
// Only for the DirectGLES transpile path.
static bool StripUboMemberRelaxedPrecisionForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Removes NoPerspective decorations so SPIRV-Cross emits plain (smooth) ESSL varyings.
// DirectGLES fallback only, for devices lacking GL_NV_shader_noperspective_interpolation
// (SPIRV-Cross would otherwise require that extension and the driver would reject it).
static bool StripNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Emulates noperspective (screen-linear) interpolation via gl_Position.w / gl_FragCoord.w
// so no NV extension is needed; strips what it cannot emulate. DirectGLES fallback for
// devices lacking GL_NV_shader_noperspective_interpolation. See EmulateNoPerspectivePass.
static bool EmulateNoPerspectiveForEssl(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Rebases loads of the InstanceIndex builtin to (InstanceIndex - BaseInstance) so
// shaders see GL's zero-based gl_InstanceID. Vertex shaders only; DirectVulkan
// backend only (glslang's relaxed mode aliases gl_InstanceID to gl_InstanceIndex,
// which wrongly includes baseInstance).
static bool RebaseInstanceIndexForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Adds the Invariant decoration to every Position builtin output. GL apps
// routinely rely on cross-program position invariance for multi-pass
// equality depth tests (e.g. GEQUAL re-draws of the same geometry), and
// mobile drivers that optimize per-pipeline break that without the
// decoration. DirectVulkan only.
static bool DecoratePositionInvariantForVulkan(const Vector<Uint32>& inputBinary,
Vector<uint32_t>& outputBinary);
// Replaces the declared format of float storage images with Unknown and adds the
// matching SPIR-V capabilities. DirectVulkan uses this only when both Vulkan
// shaderStorageImage*WithoutFormat features are enabled, allowing the
@@ -12,6 +12,7 @@
#include <cctype>
#include <initializer_list>
#include <utility>
#include <Config.h>
#include <MG_Backend/BackendObjects.h>
namespace {
@@ -95,6 +96,77 @@ namespace {
return masked;
}
// Blank out block comments in place, leaving line comments and every other byte where it is.
//
// The passes that follow scan the source as raw text, so block comments have to stop being
// visible to them - but they must not be *deleted*: replacing the bytes with spaces keeps every
// later offset valid and keeps newlines, so glslang's diagnostics still point at the line the
// application wrote. It also has to be lexically aware. A banner line such as
//
// //*** lighting pass ***
//
// contains "/*" one byte in, and a naive search for that opener treats the rest of the file as
// an unterminated comment.
void BlankBlockComments(MobileGL::String& source) {
enum class Region { Code, SingleLineComment, MultiLineComment, QuotedText };
Region region = Region::Code;
char quote = '\0';
bool escaped = false;
for (SizeT pos = 0; pos < source.size(); pos++) {
const char ch = source[pos];
const char next = pos + 1 < source.size() ? source[pos + 1] : '\0';
if (region == Region::Code) {
if (ch == '/' && next == '/') {
pos++;
region = Region::SingleLineComment;
} else if (ch == '/' && next == '*') {
source[pos] = ' ';
source[pos + 1] = ' ';
pos++;
region = Region::MultiLineComment;
} else if (ch == '"' || ch == '\'') {
quote = ch;
escaped = false;
region = Region::QuotedText;
}
continue;
}
if (region == Region::SingleLineComment) {
if (ch == '\n' || ch == '\r') region = Region::Code;
continue;
}
if (region == Region::MultiLineComment) {
if (ch == '*' && next == '/') {
source[pos] = ' ';
source[pos + 1] = ' ';
pos++;
region = Region::Code;
} else if (ch != '\n' && ch != '\r') {
source[pos] = ' ';
}
continue;
}
// GLSL has no multi-line string literals, so a quote that reaches end of line was never
// a literal to begin with - most likely an apostrophe in a #error or #pragma message.
// Ending the region here keeps one stray apostrophe from swallowing the rest of the file.
if (ch == '\n' || ch == '\r') {
region = Region::Code;
} else if (escaped) {
escaped = false;
} else if (ch == '\\') {
escaped = true;
} else if (ch == quote) {
region = Region::Code;
}
}
}
struct CodeToken {
String text;
SizeT begin = 0;
@@ -365,7 +437,21 @@ namespace {
HasIdentifierWithPrefixOutsideAllowed(tokens, "subgroup", {"subgroupInclusiveAdd"}) ||
HasIdentifierWithPrefixOutsideAllowed(
tokens, "gl_Subgroup",
{"gl_SubgroupInvocationID", "gl_SubgroupSize", "gl_SubgroupID", "gl_NumSubgroups"})) {
{"gl_SubgroupInvocationID", "gl_SubgroupSize", "gl_SubgroupID", "gl_NumSubgroups"}) ||
// ARB/NV spellings of lane-width-sensitive builtins and functions
// (gl_SubGroupSizeARB, ballotARB, gl_WarpSizeNV, shuffleNV, ...) must block the
// rewrite just like their KHR counterparts: they would silently keep native-width
// semantics in a module rewritten to the virtual 32-lane model.
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_SubGroup", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_Warp", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_Thread", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "gl_SMID", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "ballot", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "shuffle", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "readInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "readFirstInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "anyInvocation", {}) ||
HasIdentifierWithPrefixOutsideAllowed(tokens, "allInvocations", {})) {
return false;
}
@@ -487,6 +573,23 @@ namespace {
static_cast<unsigned char>(source[1]) == 0xbb && static_cast<unsigned char>(source[2]) == 0xbf;
}
// The GLSL versions MobileGL is willing to normalize. Anything else in a #version line - a number
// that is not a real language version (329, 331), a bad profile keyword, a float/identifier where
// the integer belongs, or trailing tokens - is left untouched so glslang rejects it, matching
// KHR-GL33.shaders.preprocessor.directive.version_*. The set is deliberately generous (every real
// desktop and ES version) so the normalizer never starts rejecting a form it used to accept.
bool IsRecognizedGlslVersion(unsigned version) {
switch (version) {
case 100: case 110: case 120: case 130: case 140: case 150:
case 300: case 310: case 320:
case 330: case 400: case 410: case 420: case 430:
case 440: case 450: case 460:
return true;
default:
return false;
}
}
struct ShaderLanguageInfo {
unsigned version = 110;
MobileGL::ShaderProfile profile = MobileGL::ShaderProfile::Core;
@@ -494,6 +597,9 @@ namespace {
SizeT versionDirectiveEnd = MobileGL::String::npos;
bool hasUtf8Bom = false;
bool enablesGpuShader5 = false;
// Whether the parsed #version directive is a well-formed one MobileGL should rewrite. A
// malformed directive (see IsRecognizedGlslVersion) is left alone for glslang to reject.
bool hasValidVersionDirective = false;
bool HasVersionDirective() const { return versionDirectiveStart != MobileGL::String::npos; }
};
@@ -537,13 +643,25 @@ namespace {
info.versionDirectiveEnd = lineEnd + (hasLineBreak ? 1 : 0);
SkipDirectiveWhitespace(code, probe, lineEnd);
const MobileGL::String profile = ReadDirectiveIdentifier(code, probe, lineEnd);
if (profile == "es" || profile == "ES") {
bool profileTokenValid = true;
if (profile.empty() || profile == "core") {
info.profile = MobileGL::ShaderProfile::Core;
} else if (profile == "es" || profile == "ES") {
info.profile = MobileGL::ShaderProfile::ES;
} else if (profile == "compatibility") {
info.profile = MobileGL::ShaderProfile::Compatibility;
} else {
// "#version 330 foo": an unrecognized profile keyword. Keep Core for any
// downstream routing, but mark the directive malformed.
info.profile = MobileGL::ShaderProfile::Core;
profileTokenValid = false;
}
// Comments are already masked to spaces, so anything non-blank left on the
// line is real trailing garbage: "#version 330 foobar" / "#version 330.0".
SkipDirectiveWhitespace(code, probe, lineEnd);
const bool hasTrailingTokens = probe < lineEnd;
info.hasValidVersionDirective =
IsRecognizedGlslVersion(info.version) && profileTokenValid && !hasTrailingTokens;
}
} else if (directive == "extension") {
SkipDirectiveWhitespace(code, probe, lineEnd);
@@ -590,6 +708,17 @@ namespace {
}
void NormalizeVersionDirective(MobileGL::String& source, const ShaderLanguageInfo& info) {
// A malformed #version (329, 331, bad profile, float/trailing tokens) is left exactly as the
// application wrote it so glslang rejects it - rewriting it to "#version 330 core" would
// silently legalize the CTS directive.version_* rejection cases. Still drop a leading BOM so
// the reported error is the bad version rather than a stray byte-order mark.
if (info.HasVersionDirective() && !info.hasValidVersionDirective) {
if (info.hasUtf8Bom) {
source.erase(0, 3);
}
return;
}
const MobileGL::String replacement = GetNormalizedVersionDirective(info);
if (info.HasVersionDirective()) {
source.replace(info.versionDirectiveStart, info.versionDirectiveEnd - info.versionDirectiveStart,
@@ -671,7 +800,10 @@ namespace {
void RenameBuiltinShadowingFunction(MobileGL::String& source, const char* from, const char* to) {
const MobileGL::String fromName = from;
if (!HasSingleLineFunctionDefinition(source, fromName)) {
// Decide from a comment-free view. A commented-out definition is not a definition, and
// acting on one renames every genuine call to the builtin to a name nothing defines - which
// then fails to resolve. Line comments survive BlankBlockComments, so this matters.
if (!HasSingleLineFunctionDefinition(MaskCommentsAndQuotedText(source), fromName)) {
return;
}
@@ -744,6 +876,58 @@ namespace {
return info.HasVersionDirective() ? info.versionDirectiveEnd : 0;
}
// GLSL's #line takes integer expressions only, but plenty of shader-pack preprocessors emit the
// C form with a quoted filename. Deleting every #line outright made those harmless - at the cost
// of __LINE__ reporting the position in MobileGL's rewritten text rather than the one the pack
// author wrote, and of every later diagnostic pointing at the wrong line. Dropping just the
// quoted operand keeps the directive doing its job and still hands glslang something it accepts.
void NormalizeLineDirectives(MobileGL::String& source) {
const MobileGL::String masked = MaskCommentsAndQuotedText(source);
const SizeT versionEnd = FindAfterVersionDirective(source);
MobileGL::String result;
result.reserve(source.size());
SizeT lineStart = 0;
while (lineStart <= source.size()) {
SizeT lineEnd = source.find('\n', lineStart);
const bool lastLine = lineEnd == MobileGL::String::npos;
if (lastLine) lineEnd = source.size();
SizeT probe = lineStart;
while (probe < lineEnd && (source[probe] == ' ' || source[probe] == '\t')) probe++;
const bool isLineDirective = masked.compare(probe, 5, "#line") == 0 &&
(probe + 5 >= lineEnd || !IsIdentifierChar(source[probe + 5]));
if (isLineDirective && lineStart < versionEnd) {
// #version has to be the first token in the shader, so a #line ahead of it could
// never have taken effect. Drop it rather than hand glslang a source it must reject
// - some pack preprocessors emit their directives before the version line.
} else if (isLineDirective) {
// Keep everything up to the first quote that the masker identified as string text.
SizeT quotePos = MobileGL::String::npos;
for (SizeT i = probe + 5; i < lineEnd; i++) {
if (source[i] == '"' || source[i] == '\'') {
quotePos = i;
break;
}
}
if (quotePos != MobileGL::String::npos) {
result.append(source, lineStart, quotePos - lineStart);
} else {
result.append(source, lineStart, lineEnd - lineStart);
}
} else {
result.append(source, lineStart, lineEnd - lineStart);
}
if (lastLine) break;
result.push_back('\n');
lineStart = lineEnd + 1;
}
source = std::move(result);
}
bool IsExtensionAdvertised(MobileGL::GLExtension extension) {
const auto& activeBackendObject = MobileGL::MG_Backend::pActiveBackendObject;
if (!activeBackendObject) {
@@ -772,15 +956,28 @@ namespace {
return;
}
// Detect the directive on a comment/string-masked copy so a commented-out
// "#extension GL_ARB_gpu_shader_int64" is never turned into a synthesized #error. Comments are
// no longer blanked in the delivered source (glslang handles them), so this pass must mask
// locally like its siblings. Masking preserves offsets, so edits collected against the scan
// apply verbatim to `source`; they are applied back-to-front to keep earlier offsets valid.
const MobileGL::String scan = MaskCommentsAndQuotedText(source);
struct DirectiveEdit {
SizeT pos;
SizeT len;
MobileGL::String replacement;
};
Vector<DirectiveEdit> edits;
SizeT lineStart = 0;
while (lineStart < source.size()) {
SizeT lineEnd = source.find('\n', lineStart);
while (lineStart < scan.size()) {
SizeT lineEnd = scan.find('\n', lineStart);
const bool hasLineBreak = lineEnd != MobileGL::String::npos;
if (!hasLineBreak) {
lineEnd = source.size();
lineEnd = scan.size();
}
const MobileGL::String line = source.substr(lineStart, lineEnd - lineStart);
const MobileGL::String line = scan.substr(lineStart, lineEnd - lineStart);
SizeT probe = 0;
while (probe < line.size() && std::isspace(static_cast<unsigned char>(line[probe]))) {
probe++;
@@ -822,16 +1019,12 @@ namespace {
const MobileGL::String behavior = TrimDirectiveToken(line.substr(probe));
const SizeT replaceLen = lineEnd - lineStart + (hasLineBreak ? 1 : 0);
if (behavior == "require") {
const MobileGL::String replacement =
"#error GL_ARB_gpu_shader_int64 is not advertised by MobileGL\n";
source.replace(lineStart, replaceLen, replacement);
lineStart += replacement.size();
edits.push_back({lineStart, replaceLen,
"#error GL_ARB_gpu_shader_int64 is not advertised by MobileGL\n"});
} else if (behavior == "enable" || behavior == "warn") {
source.replace(lineStart, replaceLen, "\n");
lineStart++;
} else {
lineStart = lineEnd + (hasLineBreak ? 1 : 0);
edits.push_back({lineStart, replaceLen, "\n"});
}
lineStart = lineEnd + (hasLineBreak ? 1 : 0);
continue;
}
}
@@ -841,6 +1034,10 @@ namespace {
lineStart = lineEnd + (hasLineBreak ? 1 : 0);
}
for (auto it = edits.rbegin(); it != edits.rend(); ++it) {
source.replace(it->pos, it->len, it->replacement);
}
ReplaceIdentifier(source, "GL_ARB_gpu_shader_int64", "MG_DISABLED_GL_ARB_gpu_shader_int64");
}
@@ -967,6 +1164,14 @@ namespace MobileGL {
const Vector<CodeToken> tokens = TokenizeCode(source);
LinearPrefixScanMatch match;
if (!ParseLinearPrefixScanTemplate(tokens, match)) {
// Diagnosability: when the trigger op is present but the template no longer
// matches (e.g. the pack shipped a new shader revision), the affected device
// silently falls back to the driver's miscompiled path. Make that visible.
if (CountToken(tokens, "subgroupInclusiveAdd") > 0) {
MGLOG_W("%s: subgroupInclusiveAdd present but the linear prefix-scan template "
"did not match; the wide-subgroup rewrite was NOT applied",
__func__);
}
return false;
}
@@ -979,48 +1184,98 @@ namespace MobileGL {
return true;
}
namespace {
struct ShaderSourceQuirkContext {
ShaderStage stage = ShaderStage::Unknown;
BackendType backend = BackendType::Unknown;
MG_Backend::GpuVendorKind vendor = MG_Backend::GpuVendorKind::Unknown;
Uint32 subgroupSize = 0;
};
// Device-quirk registry. Every entry is a narrowly scoped source rewrite that
// works around a specific driver defect. A quirk runs when its env override
// forces it on, or when the override is Auto and DeviceApplies matches the
// detected device. ForceOn bypasses only the device gate - each Apply keeps
// its own structural safety checks. Add new per-device workarounds here
// instead of open-coding them in PreprocessShaderSource.
struct ShaderSourceQuirk {
const char* name;
MG_Config::QuirkOverride (*GetOverride)();
Bool (*DeviceApplies)(const ShaderSourceQuirkContext&);
Bool (*Apply)(const ShaderSourceQuirkContext&, String&);
};
constexpr ShaderSourceQuirk kShaderSourceQuirks[] = {
{
// MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN
"subgroup-prefix-scan-rewrite",
[] { return MG_Config::Features.SubgroupPrefixScanQuirk; },
[](const ShaderSourceQuirkContext& ctx) {
// Qualcomm's Vulkan driver miscompiles the recognized float
// InclusiveScan pattern for native subgroups wider than the
// captured 32 lanes; other vendors compile it correctly and
// should keep their native scan.
return ctx.backend == BackendType::DirectVulkan &&
ctx.vendor == MG_Backend::GpuVendorKind::Qualcomm;
},
[](const ShaderSourceQuirkContext& ctx, String& source) {
return RewriteLinearSubgroupPrefixScanForVulkan(ctx.stage, ctx.subgroupSize,
source);
},
},
};
void ApplyShaderSourceQuirks(ShaderStage stage, String& source) {
const auto& activeBackend = MG_Backend::pActiveBackendObject;
if (!activeBackend) {
return;
}
const auto& dynamicParameters = activeBackend->GetDynamicParameters();
const ShaderSourceQuirkContext quirkContext{
stage,
activeBackend->GetBackendType(),
dynamicParameters.GpuVendor,
dynamicParameters.SubgroupSize,
};
for (const ShaderSourceQuirk& quirk : kShaderSourceQuirks) {
const MG_Config::QuirkOverride quirkOverride = quirk.GetOverride();
if (quirkOverride == MG_Config::QuirkOverride::ForceOff) {
continue;
}
if (quirkOverride == MG_Config::QuirkOverride::Auto &&
!quirk.DeviceApplies(quirkContext)) {
continue;
}
if (quirk.Apply(quirkContext, source)) {
MGLOG_I("ApplyShaderSourceQuirks: applied '%s'%s", quirk.name,
quirkOverride == MG_Config::QuirkOverride::ForceOn ? " (forced on)" : "");
}
}
}
} // namespace
void PreprocessShaderSource(ShaderStage stage, String& source) {
// Normalize while the inspector's source span still refers to the untouched input. Later passes
// remove comments and directives, so any subsequent insertion re-inspects the current source.
const ShaderLanguageInfo originalLanguage = InspectShaderLanguage(source);
NormalizeVersionDirective(source, originalLanguage);
// remove multi-line comment
size_t commentStartPos = source.find("/*");
while (commentStartPos != String::npos) {
size_t commentEndPos = source.find("*/", commentStartPos);
if (commentEndPos == String::npos) {
source.erase(commentStartPos);
break;
}
// + length of "*/"
source = source.replace(commentStartPos, commentEndPos - commentStartPos + 2, "");
commentStartPos = source.find("/*", commentStartPos);
}
// Comments are left intact for glslang's own preprocessor: a block comment is a single
// preprocessing token that collapses to one space even across newlines and inside a
// directive, so blanking it here (which preserved the interior newlines) truncated
// multi-line #define bodies and broke otherwise-valid shaders (KHR-GL3x.shaders.
// preprocessor multiline_comment_define / redefine_object / function_redefinition).
// Every MobileGL pass that must ignore comment/string text already masks them locally
// via MaskCommentsAndQuotedText/TokenizeCode, so the source we hand glslang keeps them.
NormalizeLineDirectives(source);
// remove #line directives
SizeT linedirPos = source.find("#line");
while (linedirPos != String::npos) {
SizeT newlinePos = source.find('\n', linedirPos);
if (newlinePos == String::npos) {
source.erase(linedirPos);
break;
}
// Preserve a line break so adjacent preprocessor directives do not merge.
source = source.replace(linedirPos, newlinePos - linedirPos + 1, "\n");
linedirPos = source.find("#line", linedirPos);
}
// remove "noperspective"
const char* str_np = "noperspective";
const SizeT len_np = strlen(str_np);
SizeT noperspectivePos = source.find(str_np);
while (noperspectivePos != String::npos) {
// + length of "\n"
source = source.replace(noperspectivePos, len_np, "");
noperspectivePos = source.find(str_np);
}
// noperspective is intentionally NOT touched here. It is core in desktop GLSL (1.30+)
// and maps to the core SPIR-V NoPerspective decoration, which DirectVulkan renders
// natively and SPIRV-Cross turns into ESSL `noperspective` + the
// GL_NV_shader_noperspective_interpolation extension. The old naked substring erase
// both discarded that interpolation (shader packs need it) and corrupted any
// identifier that merely contained the word. The GLES fallback for devices without
// the extension lives in the backend, where device capabilities are known.
FilterUnsupportedGpuShaderInt64(source);
CoerceUniformBlockPackingToStd140(source);
@@ -1035,12 +1290,7 @@ namespace MobileGL {
ModernizeLegacyGLSL(stage, source);
InjectDepthRangeBuiltinShim(stage, source);
const auto& activeBackend = MG_Backend::pActiveBackendObject;
if (stage == ShaderStage::Compute && activeBackend &&
activeBackend->GetBackendType() == BackendType::DirectVulkan) {
RewriteLinearSubgroupPrefixScanForVulkan(stage, activeBackend->GetDynamicParameters().SubgroupSize,
source);
}
ApplyShaderSourceQuirks(stage, source);
}
Bool RetargetLegacyVersionDirectiveTo460(String& source) {
@@ -1049,6 +1299,10 @@ namespace MobileGL {
// must not be mistaken for the real one.
const ShaderLanguageInfo info = InspectShaderLanguage(source);
if (!info.HasVersionDirective()) return false;
// Never rescue a malformed directive to 460: that is precisely what re-legalized the
// CTS directive.version_* rejection cases after the first compile failed. The shader-
// pack retry this exists for only ever sees a valid low version (a real "#version 330").
if (!info.hasValidVersionDirective) return false;
// Only the set NormalizeVersionDirective downgraded: desktop core below 400. ES and
// compatibility shaders keep whatever they declared.
if (info.profile != ShaderProfile::Core || info.version >= 400) return false;
@@ -27,8 +27,10 @@ namespace MobileGL {
// wider than the capture's 32 lanes. For the narrowly recognized, uniform-control-
// flow template, replace the subgroup-local scan with a shared-memory, strict
// left-fold over virtual 32-lane segments. Returns true only when the complete safe
// template was recognized and rewritten. DirectVulkan calls this through
// PreprocessShaderSource; the explicit entry point exists for deterministic tests.
// template was recognized and rewritten. PreprocessShaderSource reaches this through
// its device-quirk registry: by default only on detected Qualcomm Vulkan devices,
// overridable either way with MOBILEGL_QUIRK_SUBGROUP_PREFIX_SCAN=1/0. The explicit
// entry point exists for deterministic tests.
Bool RewriteLinearSubgroupPrefixScanForVulkan(ShaderStage stage, Uint32 nativeSubgroupSize, String& source);
// Rewrites a "#version 330 core" directive that PreprocessShaderSource normalized down
@@ -0,0 +1,139 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "DecoratePositionInvariantPass.h"
#include "spirv.hpp"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/util/make_unique.h"
#include <algorithm>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
using spvtools::opt::Operand;
// Identifies one member of a decorated struct (gl_PerVertex's Position slot).
struct MemberKey {
uint32_t id = 0;
uint32_t member = 0;
bool operator==(const MemberKey& other) const {
return id == other.id && member == other.member;
}
};
// OpDecorate <target-id> <decoration> [literals...]
// OpMemberDecorate <struct-id> <member> <decoration> [literals...]
constexpr uint32_t kDecorateTargetOperand = 0;
constexpr uint32_t kDecorateDecorationOperand = 1;
constexpr uint32_t kDecorateBuiltInOperand = 2;
constexpr uint32_t kMemberDecorateStructOperand = 0;
constexpr uint32_t kMemberDecorateMemberOperand = 1;
constexpr uint32_t kMemberDecorateDecorationOperand = 2;
constexpr uint32_t kMemberDecorateBuiltInOperand = 3;
} // namespace
spvtools::opt::Pass::Status DecoratePositionInvariantPass::Process() {
auto* irContext = context();
// Collect first: AddAnnotationInst mutates the list being walked.
std::vector<uint32_t> invariantIds;
std::vector<MemberKey> invariantMembers;
std::vector<uint32_t> positionIds;
std::vector<MemberKey> positionMembers;
for (const Instruction& annotation : irContext->annotations()) {
if (annotation.opcode() == spv::Op::OpDecorate) {
if (annotation.NumInOperands() <= kDecorateDecorationOperand) {
continue;
}
const auto decoration = static_cast<spv::Decoration>(
annotation.GetSingleWordInOperand(kDecorateDecorationOperand));
const uint32_t target = annotation.GetSingleWordInOperand(kDecorateTargetOperand);
if (decoration == spv::Decoration::Invariant) {
invariantIds.push_back(target);
} else if (decoration == spv::Decoration::BuiltIn &&
annotation.NumInOperands() > kDecorateBuiltInOperand &&
static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(
kDecorateBuiltInOperand)) == spv::BuiltIn::Position) {
positionIds.push_back(target);
}
} else if (annotation.opcode() == spv::Op::OpMemberDecorate) {
if (annotation.NumInOperands() <= kMemberDecorateDecorationOperand) {
continue;
}
const auto decoration = static_cast<spv::Decoration>(
annotation.GetSingleWordInOperand(kMemberDecorateDecorationOperand));
const MemberKey key{
annotation.GetSingleWordInOperand(kMemberDecorateStructOperand),
annotation.GetSingleWordInOperand(kMemberDecorateMemberOperand)};
if (decoration == spv::Decoration::Invariant) {
invariantMembers.push_back(key);
} else if (decoration == spv::Decoration::BuiltIn &&
annotation.NumInOperands() > kMemberDecorateBuiltInOperand &&
static_cast<spv::BuiltIn>(annotation.GetSingleWordInOperand(
kMemberDecorateBuiltInOperand)) == spv::BuiltIn::Position) {
positionMembers.push_back(key);
}
}
}
Bool changed = false;
for (const uint32_t target : positionIds) {
if (std::find(invariantIds.begin(), invariantIds.end(), target) != invariantIds.end()) {
continue;
}
irContext->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {target}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::Invariant)}}}));
// Guard against a second Position decoration on the same target.
invariantIds.push_back(target);
changed = true;
}
for (const MemberKey& key : positionMembers) {
if (std::find(invariantMembers.begin(), invariantMembers.end(), key) !=
invariantMembers.end()) {
continue;
}
irContext->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
irContext, spv::Op::OpMemberDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {key.id}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {key.member}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::Invariant)}}}));
invariantMembers.push_back(key);
changed = true;
}
if (!changed) {
return Status::SuccessWithoutChange;
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken
DecoratePositionInvariantPass::CreateDecoratePositionInvariantPass() {
return spvtools::Optimizer::PassToken(MakeUnique<DecoratePositionInvariantPass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,35 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/DecoratePositionInvariantPass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Adds the Invariant decoration to every Position builtin output. GL apps
// routinely rely on cross-program position invariance for multi-pass equality
// depth tests - MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the
// depth its own first pass wrote - and a driver that optimizes each pipeline
// separately may otherwise vary the position math between passes, dropping whole
// primitives from the later ones. Both the plain (OpDecorate on a Position
// variable) and the block-member (OpMemberDecorate on gl_PerVertex) spellings are
// handled; targets that already carry Invariant are left alone. DirectVulkan only.
class DecoratePositionInvariantPass : public spvtools::opt::Pass {
public:
const char* name() const override { return "decorate-position-invariant"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateDecoratePositionInvariantPass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,407 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "EmulateNoPerspectivePass.h"
#include "spirv.hpp"
#include "source/opt/constants.h"
#include "source/opt/def_use_manager.h"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/opt/module.h"
#include "source/opt/type_manager.h"
#include "source/opt/types.h"
#include "source/util/make_unique.h"
#include <algorithm>
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
using spvtools::opt::Operand;
namespace analysis = spvtools::opt::analysis;
spv::ExecutionModel EntryExecutionModel(IRContext* ctx) {
for (Instruction& ep : ctx->module()->entry_points()) {
return static_cast<spv::ExecutionModel>(ep.GetSingleWordInOperand(0));
}
return spv::ExecutionModel::Max;
}
uint32_t VariablePointeeType(IRContext* ctx, Instruction* var) {
Instruction* ptrType = ctx->get_def_use_mgr()->GetDef(var->type_id());
// OpTypePointer <storage-class> <pointee>
return ptrType->GetSingleWordInOperand(1);
}
// If |typeId| is float or a vector of float, returns true and reports the scalar float
// type and whether it is a vector. Matrices, structs, ints etc. are not emulatable.
bool IsFloatScalarOrVector(IRContext* ctx, uint32_t typeId, uint32_t& floatTypeId, bool& isVector) {
Instruction* t = ctx->get_def_use_mgr()->GetDef(typeId);
if (t == nullptr) return false;
if (t->opcode() == spv::Op::OpTypeFloat) {
floatTypeId = typeId;
isVector = false;
return true;
}
if (t->opcode() == spv::Op::OpTypeVector) {
const uint32_t comp = t->GetSingleWordInOperand(0);
Instruction* ct = ctx->get_def_use_mgr()->GetDef(comp);
if (ct != nullptr && ct->opcode() == spv::Op::OpTypeFloat) {
floatTypeId = comp;
isVector = true;
return true;
}
}
return false;
}
uint32_t PointerTypeTo(IRContext* ctx, uint32_t pointeeId, spv::StorageClass sc) {
analysis::Type* pointee = ctx->get_type_mgr()->GetType(pointeeId);
analysis::Pointer ptr(pointee, sc);
return ctx->get_type_mgr()->GetTypeInstruction(&ptr);
}
uint32_t V4FloatType(IRContext* ctx) {
analysis::Float f(32);
analysis::Type* freg = ctx->get_type_mgr()->GetRegisteredType(&f);
analysis::Vector v(freg, 4);
return ctx->get_type_mgr()->GetTypeInstruction(&v);
}
uint32_t FloatType(IRContext* ctx) {
analysis::Float f(32);
return ctx->get_type_mgr()->GetTypeInstruction(&f);
}
uint32_t SignedIntConstant(IRContext* ctx, int32_t value) {
analysis::Integer i(32, true);
analysis::Type* reg = ctx->get_type_mgr()->GetRegisteredType(&i);
const analysis::Constant* c =
ctx->get_constant_mgr()->GetConstant(reg, {static_cast<uint32_t>(value)});
return ctx->get_constant_mgr()->GetDefiningInstruction(c)->result_id();
}
// Multiply |valueId| (of type |valueTypeId|) by the scalar |scalarId|, inserting the op
// before |before|. Returns the product's id.
uint32_t InsertScale(IRContext* ctx, Instruction* before, uint32_t valueTypeId,
uint32_t valueId, uint32_t scalarId, bool isVector) {
const uint32_t productId = ctx->TakeNextId();
const spv::Op op = isVector ? spv::Op::OpVectorTimesScalar : spv::Op::OpFMul;
before->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, op, valueTypeId, productId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {valueId}},
{SPV_OPERAND_TYPE_ID, {scalarId}}}));
return productId;
}
// --- Vertex stage: gl_Position discovery ------------------------------------------
// Finds gl_Position as member |memberIndex| of a gl_PerVertex-style block whose Output
// variable is |blockVarId|; |v4floatTypeId| is that member's (vec4) type. Returns false
// if gl_Position is not a block member (older plain-variable form is left to the strip).
bool FindPositionBlock(IRContext* ctx, uint32_t& blockVarId, uint32_t& memberIndex,
uint32_t& v4floatTypeId) {
uint32_t structId = 0;
uint32_t member = 0;
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpMemberDecorate && ann.NumInOperands() >= 4 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(2)) ==
spv::Decoration::BuiltIn &&
static_cast<spv::BuiltIn>(ann.GetSingleWordInOperand(3)) ==
spv::BuiltIn::Position) {
structId = ann.GetSingleWordInOperand(0);
member = ann.GetSingleWordInOperand(1);
break;
}
}
if (structId == 0) return false;
Instruction* structType = ctx->get_def_use_mgr()->GetDef(structId);
if (structType == nullptr || member >= structType->NumInOperands()) return false;
v4floatTypeId = structType->GetSingleWordInOperand(member);
for (Instruction& inst : ctx->module()->types_values()) {
if (inst.opcode() == spv::Op::OpVariable &&
static_cast<spv::StorageClass>(inst.GetSingleWordInOperand(0)) ==
spv::StorageClass::Output &&
VariablePointeeType(ctx, &inst) == structId) {
blockVarId = inst.result_id();
memberIndex = member;
return true;
}
}
return false;
}
// --- Fragment stage: gl_FragCoord discovery/synthesis -----------------------------
Instruction* FindBuiltinInput(IRContext* ctx, spv::BuiltIn builtin) {
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() != spv::Op::OpDecorate || ann.NumInOperands() < 3) continue;
if (static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) !=
spv::Decoration::BuiltIn)
continue;
if (static_cast<spv::BuiltIn>(ann.GetSingleWordInOperand(2)) != builtin) continue;
Instruction* var = ctx->get_def_use_mgr()->GetDef(ann.GetSingleWordInOperand(0));
if (var != nullptr && var->opcode() == spv::Op::OpVariable &&
static_cast<spv::StorageClass>(var->GetSingleWordInOperand(0)) ==
spv::StorageClass::Input) {
return var;
}
}
return nullptr;
}
uint32_t SynthesizeFragCoord(IRContext* ctx, uint32_t v4floatTypeId) {
const uint32_t ptrType = PointerTypeTo(ctx, v4floatTypeId, spv::StorageClass::Input);
const uint32_t varId = ctx->TakeNextId();
ctx->AddGlobalValue(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpVariable, ptrType, varId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_STORAGE_CLASS,
{static_cast<uint32_t>(spv::StorageClass::Input)}}}));
ctx->AddAnnotationInst(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpDecorate, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {varId}},
{SPV_OPERAND_TYPE_DECORATION,
{static_cast<uint32_t>(spv::Decoration::BuiltIn)}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER,
{static_cast<uint32_t>(spv::BuiltIn::FragCoord)}}}));
for (Instruction& ep : ctx->module()->entry_points()) {
ep.AddOperand({SPV_OPERAND_TYPE_ID, {varId}});
}
return varId;
}
} // namespace
spvtools::opt::Pass::Status EmulateNoPerspectivePass::Process() {
auto* ctx = context();
const spv::ExecutionModel model = EntryExecutionModel(ctx);
const bool isVertex = model == spv::ExecutionModel::Vertex;
const bool isFragment = model == spv::ExecutionModel::Fragment;
// Collect NoPerspective-decorated plain variables and every NoPerspective annotation.
std::vector<uint32_t> plainVarIds;
std::vector<Instruction*> decorationsToKill;
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpDecorate && ann.NumInOperands() >= 2 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) ==
spv::Decoration::NoPerspective) {
plainVarIds.push_back(ann.GetSingleWordInOperand(0));
decorationsToKill.push_back(&ann);
} else if (ann.opcode() == spv::Op::OpMemberDecorate && ann.NumInOperands() >= 3 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(2)) ==
spv::Decoration::NoPerspective) {
// Block-member noperspective is not emulated here; the decoration is stripped
// (smooth fallback) so SPIRV-Cross does not require the NV extension.
decorationsToKill.push_back(&ann);
}
}
if (decorationsToKill.empty()) {
return Status::SuccessWithoutChange;
}
const spv::StorageClass wantStorage =
isVertex ? spv::StorageClass::Output : spv::StorageClass::Input;
// Emulatable = plain variable of the stage's interface direction, float or floatN.
struct Target {
Instruction* var;
uint32_t typeId;
uint32_t floatTypeId;
bool isVector;
};
std::vector<Target> targets;
if (isVertex || isFragment) {
for (const uint32_t id : plainVarIds) {
Instruction* var = ctx->get_def_use_mgr()->GetDef(id);
if (var == nullptr || var->opcode() != spv::Op::OpVariable) continue;
if (static_cast<spv::StorageClass>(var->GetSingleWordInOperand(0)) != wantStorage)
continue;
const uint32_t pointee = VariablePointeeType(ctx, var);
uint32_t floatTypeId = 0;
bool isVector = false;
if (IsFloatScalarOrVector(ctx, pointee, floatTypeId, isVector)) {
targets.push_back({var, pointee, floatTypeId, isVector});
}
}
}
// Force highp on the varyings we emulate: the a*w round-trip overflows a mediump (fp16)
// varying at large clip-space w. Dropping RelaxedPrecision makes SPIRV-Cross emit them
// highp on both stages, keeping the emulation exact. Only touches emulated variables.
if (!targets.empty()) {
std::vector<uint32_t> targetIds;
targetIds.reserve(targets.size());
for (const Target& t : targets) targetIds.push_back(t.var->result_id());
for (Instruction& ann : ctx->annotations()) {
if (ann.opcode() == spv::Op::OpDecorate && ann.NumInOperands() >= 2 &&
static_cast<spv::Decoration>(ann.GetSingleWordInOperand(1)) ==
spv::Decoration::RelaxedPrecision &&
std::find(targetIds.begin(), targetIds.end(),
ann.GetSingleWordInOperand(0)) != targetIds.end()) {
decorationsToKill.push_back(&ann);
}
}
}
if (isVertex && !targets.empty()) {
uint32_t blockVarId = 0;
uint32_t memberIndex = 0;
uint32_t v4floatTypeId = 0;
if (FindPositionBlock(ctx, blockVarId, memberIndex, v4floatTypeId)) {
const uint32_t ptrOutV4 =
PointerTypeTo(ctx, v4floatTypeId, spv::StorageClass::Output);
const uint32_t memberConst = SignedIntConstant(ctx, static_cast<int32_t>(memberIndex));
const uint32_t floatTy = FloatType(ctx);
uint32_t entryFuncId = 0;
for (Instruction& ep : ctx->module()->entry_points()) {
// OpEntryPoint <model> <function> "name" <interface...>
entryFuncId = ep.GetSingleWordInOperand(1);
break;
}
// Pre-multiply every target output by gl_Position.w before each return of the
// ENTRY function only. glslang does not inline, so a called helper survives as
// its own OpFunction; instrumenting its returns too would scale the varying
// more than once (w^2), breaking the identity.
for (auto funcIt = ctx->module()->begin(); funcIt != ctx->module()->end(); ++funcIt) {
if (funcIt->result_id() != entryFuncId) continue;
funcIt->ForEachInst([&](Instruction* inst) {
if (inst->opcode() != spv::Op::OpReturn &&
inst->opcode() != spv::Op::OpReturnValue) {
return;
}
const uint32_t posPtrId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpAccessChain, ptrOutV4, posPtrId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {blockVarId}},
{SPV_OPERAND_TYPE_ID, {memberConst}}}));
const uint32_t posId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, v4floatTypeId, posId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {posPtrId}}}));
const uint32_t wId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpCompositeExtract, floatTy, wId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {posId}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {3u}}}));
for (const Target& t : targets) {
const uint32_t valId = ctx->TakeNextId();
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, t.typeId, valId,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {t.var->result_id()}}}));
const uint32_t scaledId =
InsertScale(ctx, inst, t.typeId, valId, wId, t.isVector);
inst->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpStore, 0, 0,
std::initializer_list<Operand>{
{SPV_OPERAND_TYPE_ID, {t.var->result_id()}},
{SPV_OPERAND_TYPE_ID, {scaledId}}}));
}
});
}
}
}
if (isFragment && !targets.empty()) {
Instruction* fragCoord = FindBuiltinInput(ctx, spv::BuiltIn::FragCoord);
uint32_t fragCoordId = 0;
uint32_t v4floatTypeId = 0;
if (fragCoord != nullptr) {
fragCoordId = fragCoord->result_id();
v4floatTypeId = VariablePointeeType(ctx, fragCoord);
} else {
v4floatTypeId = V4FloatType(ctx);
fragCoordId = SynthesizeFragCoord(ctx, v4floatTypeId);
}
const uint32_t floatTy = FloatType(ctx);
auto* defUse = ctx->get_def_use_mgr();
for (const Target& t : targets) {
// Collect every load that reads the varying. glslang lowers a whole-variable
// read to OpLoad(var), but a single-component read (v.x) to
// OpAccessChain(var) + OpLoad(chain). Both must be scaled; the identity is
// per-component, so scaling one loaded component by gl_FragCoord.w is valid.
std::vector<Instruction*> loads;
defUse->ForEachUser(t.var, [&](Instruction* user) {
if (user->opcode() == spv::Op::OpLoad &&
user->GetSingleWordInOperand(0) == t.var->result_id()) {
loads.push_back(user);
} else if (user->opcode() == spv::Op::OpAccessChain &&
user->GetSingleWordInOperand(0) == t.var->result_id()) {
const uint32_t chainId = user->result_id();
defUse->ForEachUser(user, [&](Instruction* chainUser) {
if (chainUser->opcode() == spv::Op::OpLoad &&
chainUser->GetSingleWordInOperand(0) == chainId) {
loads.push_back(chainUser);
}
});
}
});
// Rewrite `%r = OpLoad %ty %ptr` into
// %orig = OpLoad %ty %ptr
// %fc = OpLoad %v4float %fragCoord
// %w = OpCompositeExtract %float %fc 3
// %r = OpVectorTimesScalar/OpFMul %ty %orig %w (reuse %r: uses stay intact)
// The op is chosen from the LOAD's own result type: a whole-vector load scales
// with OpVectorTimesScalar, a scalar component load with OpFMul.
for (Instruction* load : loads) {
const uint32_t loadType = load->type_id();
uint32_t componentFloat = 0;
bool loadIsVector = false;
if (!IsFloatScalarOrVector(ctx, loadType, componentFloat, loadIsVector)) {
continue;
}
const uint32_t ptrId = load->GetSingleWordInOperand(0);
const uint32_t origId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, loadType, origId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {ptrId}}}));
const uint32_t fcId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpLoad, v4floatTypeId, fcId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {fragCoordId}}}));
const uint32_t wId = ctx->TakeNextId();
load->InsertBefore(spvtools::MakeUnique<Instruction>(
ctx, spv::Op::OpCompositeExtract, floatTy, wId,
std::initializer_list<Operand>{{SPV_OPERAND_TYPE_ID, {fcId}},
{SPV_OPERAND_TYPE_LITERAL_INTEGER, {3u}}}));
load->SetOpcode(loadIsVector ? spv::Op::OpVectorTimesScalar : spv::Op::OpFMul);
load->SetInOperands(Instruction::OperandList{
{SPV_OPERAND_TYPE_ID, {origId}}, {SPV_OPERAND_TYPE_ID, {wId}}});
}
}
}
// Strip every NoPerspective decoration: emulated varyings now transport smooth, and
// non-emulatable ones fall back to smooth.
for (Instruction* dec : decorationsToKill) {
ctx->KillInst(dec);
}
ctx->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken EmulateNoPerspectivePass::CreateEmulateNoPerspectivePass() {
return spvtools::Optimizer::PassToken(MakeUnique<EmulateNoPerspectivePass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,41 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/EmulateNoPerspectivePass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Emulates 'noperspective' (screen-linear) interpolation on GLES devices that lack
// GL_NV_shader_noperspective_interpolation, so no NV extension is required. The hardware
// interpolates perspective-correct; screen-linear L(a) is recovered from the identity
// L(a) = P(a * w) * gl_FragCoord.w
// where P is perspective-correct interpolation and w is the vertex clip-space w. So each
// NoPerspective-decorated output is pre-multiplied by gl_Position.w in the vertex stage
// and each NoPerspective-decorated input is multiplied by gl_FragCoord.w in the fragment
// stage; the decoration is then removed so the varying transports smooth. This is exact
// (modulo float precision - the emulated varyings want highp).
//
// Scope: plain interface variables of float or floatN type. Anything it cannot emulate
// (interface-block members, matrices, or a stage lacking the needed builtin) has its
// NoPerspective decoration stripped instead, degrading to smooth - the same result the
// extension-less fallback produced before, and never invalid SPIR-V. DirectGLES only.
class EmulateNoPerspectivePass : public spvtools::opt::Pass {
public:
const char* name() const override { return "emulate-noperspective"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateEmulateNoPerspectivePass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,72 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.cpp
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#include "StripNoPerspectivePass.h"
#include "spirv.hpp"
#include "source/opt/instruction.h"
#include "source/opt/ir_context.h"
#include "source/util/make_unique.h"
#include <vector>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
namespace {
using spvtools::opt::Instruction;
using spvtools::opt::IRContext;
// OpDecorate <target-id> <decoration> [literals...]
// OpMemberDecorate <struct-id> <member> <decoration> [literals...]
constexpr uint32_t kDecorateDecorationOperand = 1;
constexpr uint32_t kMemberDecorateDecorationOperand = 2;
} // namespace
spvtools::opt::Pass::Status StripNoPerspectivePass::Process() {
auto* irContext = context();
// Collect first: KillInst mutates the annotation list being walked.
std::vector<Instruction*> toKill;
for (Instruction& annotation : irContext->annotations()) {
uint32_t decorationOperand = 0;
if (annotation.opcode() == spv::Op::OpDecorate) {
decorationOperand = kDecorateDecorationOperand;
} else if (annotation.opcode() == spv::Op::OpMemberDecorate) {
decorationOperand = kMemberDecorateDecorationOperand;
} else {
continue;
}
if (annotation.NumInOperands() <= decorationOperand) {
continue;
}
if (static_cast<spv::Decoration>(annotation.GetSingleWordInOperand(decorationOperand)) ==
spv::Decoration::NoPerspective) {
toKill.push_back(&annotation);
}
}
if (toKill.empty()) {
return Status::SuccessWithoutChange;
}
for (Instruction* inst : toKill) {
irContext->KillInst(inst);
}
irContext->InvalidateAnalysesExceptFor(IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
spvtools::Optimizer::PassToken StripNoPerspectivePass::CreateStripNoPerspectivePass() {
return spvtools::Optimizer::PassToken(MakeUnique<StripNoPerspectivePass>());
}
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
@@ -0,0 +1,35 @@
// MobileGL - MobileGL/MG_Util/ShaderTranspiler/SpirvPasses/StripNoPerspectivePass.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include "source/opt/pass.h"
#include "spirv-tools/optimizer.hpp"
#include <Includes.h>
namespace MobileGL {
namespace MG_Util {
namespace ShaderTranspiler {
// Removes the NoPerspective decoration from every interface variable and block member.
// DirectGLES fallback only, for devices that lack GL_NV_shader_noperspective_interpolation:
// SPIRV-Cross renders a NoPerspective-decorated varying as ESSL `noperspective` plus
// `#extension GL_NV_shader_noperspective_interpolation : require`, which such a driver
// rejects. Dropping the decoration falls the varying back to smooth (perspective-correct)
// interpolation - the same visible result the old text-level strip produced, but without
// corrupting identifiers and without touching DirectVulkan, where NoPerspective is native.
// (The exact screen-linear emulation via gl_Position.w / gl_FragCoord.w is a later step.)
class StripNoPerspectivePass : public spvtools::opt::Pass {
public:
const char* name() const override { return "strip-noperspective"; }
Status Process() override;
static spvtools::Optimizer::PassToken CreateStripNoPerspectivePass();
};
} // namespace ShaderTranspiler
} // namespace MG_Util
} // namespace MobileGL
+9
View File
@@ -93,6 +93,15 @@ The bundled fixtures cover:
- minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
iterationT after entering a singleplayer world, with Iris' DSA path disabled.
![Minecraft 1.21.4 Fabric Iris iterationT no-DSA in-world golden](fixtures/minecraft-1.21.4-fabric-iris-iterationt-nodsa-in-world.0000115019.png)
- minecraft-1.21.4-fabric-iris-iterationrp-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
iterationRP after entering a singleplayer world, framing the iterationRP name overlay over a lake with far-shore
tree reflections. iterationRP's temporal auto-exposure makes a single-frame trim overexpose and drop the overlay,
so the fixture is a prefix trace (all calls up to the target frame) that replays the temporal state. The pack also
gates an NVIDIA-only shadow path (`subgroupPartitionNV`, `GL_NV_shader_subgroup_partitioned`) on the GL vendor
string, so the capture reports a masked vendor and the trace carries the portable `subgroupShuffleXor` path that
non-NVIDIA GPUs take.
The trace archive and golden are not committed yet (the repository's Git LFS quota rejects new objects with
`GH009`); the case stays registered and its fixture files are hydrated from the trace fixture mirror.
- minecraft-1.21.4-fabric-iris-photon-v1.1-in-world: captured from Minecraft 1.21.4 Fabric with Sodium, Iris, and
Photon v1.1 after entering a singleplayer world.
![Minecraft 1.21.4 Fabric Iris Photon v1.1 in-world golden](fixtures/minecraft-1.21.4-fabric-iris-photon-v1.1-in-world.0000159866.png)
@@ -1,4 +0,0 @@
interface:
display_name: "Capture RenderDoc Trace Frame"
short_description: "Capture exact Android Vulkan/GLES trace frames"
default_prompt: "Use $renderdoc-capture-trace-frame to capture and validate an exact Android retrace frame."
@@ -19,7 +19,12 @@ RESULT_ROOT = ROOT / ".trace-work" / "macos-window-retrace-result"
WORK_ROOT = ROOT / ".trace-work" / "macos-window-retrace-work"
SUMMARY_DIR = ROOT / ".trace-work" / "macos-window-retrace-summary"
SUMMARY_HTML = "mobilegl-macos-window-vulkan-retrace-overview.html"
DEFAULT_MIRROR_BASE = "https://repo.miawa.cn/mgl/tools/trace_replay/fixtures"
# Fixture mirrors, tried in order before Git LFS.
DEFAULT_MIRROR_BASES = [
"https://git.hit.moe/swung0x48/MobileGL/media/branch/dev/tools/trace_replay/fixtures",
"https://repo.miawa.cn/mgl/tools/trace_replay/fixtures",
]
DEFAULT_MIRROR_BASE = DEFAULT_MIRROR_BASES[0]
BACKENDS = ("DirectVulkan",)
CASES = load_trace_cases()
@@ -100,34 +105,35 @@ def default_vulkan_icd(mobilegl_library):
return ""
def download_fixture(path, mirror_base):
def download_fixture(path, mirror_bases):
"""Try each mirror in order; return on the first that serves a good file."""
if isinstance(mirror_bases, str):
mirror_bases = [mirror_bases]
path.parent.mkdir(parents=True, exist_ok=True)
url = f"{mirror_base.rstrip('/')}/{path.name}"
tmp = path.with_suffix(path.suffix + ".tmp")
command = [
"curl",
"-L",
"--fail",
"--retry",
"3",
"--retry-delay",
"2",
"--continue-at",
"-",
"-o",
str(tmp),
url,
]
result = subprocess.run(command, text=True, capture_output=True)
if result.returncode != 0:
return path, False, result.stderr.strip() or result.stdout.strip()
tmp.replace(path)
if is_bad_fixture(path):
return path, False, "downloaded fixture is empty or still an LFS pointer"
return path, True, ""
token = os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_TOKEN", "")
errors = []
for mirror_base in mirror_bases:
url = f"{mirror_base.rstrip('/')}/{path.name}"
command = ["curl", "-L", "--fail", "--retry", "3", "--retry-delay", "2",
"--continue-at", "-"]
if token:
command += ["--header", f"Authorization: token {token}"]
command += ["-o", str(tmp), url]
result = subprocess.run(command, text=True, capture_output=True)
if result.returncode != 0:
errors.append(f"{mirror_base}: {result.stderr.strip() or result.stdout.strip()}")
tmp.unlink(missing_ok=True)
continue
tmp.replace(path)
if is_bad_fixture(path):
errors.append(f"{mirror_base}: downloaded fixture is empty or still an LFS pointer")
continue
return path, True, ""
return path, False, "; ".join(errors)
def hydrate_fixtures(cases, fetch, mirror_base, download_jobs):
def hydrate_fixtures(cases, fetch, mirror_bases, download_jobs):
required = []
seen = set()
for case in cases:
@@ -149,7 +155,7 @@ def hydrate_fixtures(cases, fetch, mirror_base, download_jobs):
print(f"fixtures: downloading {len(required)} file(s) with {download_jobs} worker(s)", flush=True)
failures = []
with concurrent.futures.ThreadPoolExecutor(max_workers=download_jobs) as executor:
futures = [executor.submit(download_fixture, path, mirror_base) for path in required]
futures = [executor.submit(download_fixture, path, mirror_bases) for path in required]
for future in concurrent.futures.as_completed(futures):
path, ok, message = future.result()
if ok:
@@ -434,7 +440,9 @@ def parse_args():
parser.add_argument("--keep-results", action="store_true", help="Do not clear previous result/work roots.")
parser.add_argument("--no-render", action="store_true", help="Do not render the HTML summary.")
parser.add_argument("--fetch-fixtures", action=argparse.BooleanOptionalAction, default=True, help="Hydrate missing fixtures before running.")
parser.add_argument("--fixture-mirror-base", default=os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASE", DEFAULT_MIRROR_BASE))
parser.add_argument("--fixture-mirror-base", action="append", dest="fixture_mirror_bases",
help="Fixture mirror base URL; repeatable, tried in order. "
"Defaults to the hit.moe mirror then the miawa mirror.")
parser.add_argument("--download-jobs", type=int, default=4, help="Parallel fixture download count.")
parser.add_argument("--continue-after-fatal", action="store_true", help="Continue launching later cases after a fatal native-window replay.")
return parser.parse_args()
@@ -463,7 +471,10 @@ def main():
print(f"missing MobileGL library: {mobilegl_library}", file=sys.stderr)
return 2
if not hydrate_fixtures(selected_cases, args.fetch_fixtures, args.fixture_mirror_base, max(1, args.download_jobs)):
mirror_bases = args.fixture_mirror_bases or [
b for b in os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASES", "").split() or []
] or ([os.environ["MOBILEGL_TRACE_FIXTURE_MIRROR_BASE"]] if os.environ.get("MOBILEGL_TRACE_FIXTURE_MIRROR_BASE") else DEFAULT_MIRROR_BASES)
if not hydrate_fixtures(selected_cases, args.fetch_fixtures, mirror_bases, max(1, args.download_jobs)):
return 2
if not args.keep_results:
+20
View File
@@ -0,0 +1,20 @@
# MobileGL trace-replay skills
Task-focused skills for capturing, replaying, debugging, and authoring MobileGL
apitrace fixtures. Each skill is a self-contained package:
- `SKILL.md` — the skill (frontmatter `name` + `description`, then the body). The
directory name equals the frontmatter `name`.
- `agents/openai.yaml` — OpenAI agent descriptor (`display_name`,
`short_description`, `default_prompt`).
- `scripts/` and/or `references/` — bundled tooling and supporting docs, when the
skill has them.
## Skills
| Skill | What it does |
| --- | --- |
| [trace-fixture-authoring-on-android-fcl](trace-fixture-authoring-on-android-fcl/SKILL.md) | Capture an on-device Android apitrace from FCL's MobileGL renderers (DirectGLES / Magma / SimpleFPEWrapper), mark the defect frame, and pull `full.trace`. |
| [renderdoc-debug-on-trace-replay](renderdoc-debug-on-trace-replay/SKILL.md) | Capture and validate an exact frame from a MobileGL retrace on a connected Android device with RenderDoc / rdc-cli. |
| [mismatch-retrace-debugging](mismatch-retrace-debugging/SKILL.md) | Localize the first divergent render pass and draw call when a fixture replays correctly in a golden environment but renders differently on a target backend. |
| [trace-fixture-authoring](trace-fixture-authoring/SKILL.md) | Author a deterministic trace-replay fixture — trim, golden, package under the size budget, register in `trace_cases.json`, and validate on Linux and Android. |
@@ -1,3 +1,8 @@
---
name: mismatch-retrace-debugging
description: Localize the first divergent render pass and draw call when a MobileGL apitrace fixture replays correctly in a golden environment but renders differently under mobilegl_trace_replay, Android trace replay, or another backend. Use to binary-search pass/draw endpoints, diff GL state around the first bad call, and classify the fault as a vertex/VS, fragment/FS, or framebuffer/composition mismatch.
---
# Mismatch retrace debugging
Use this when an apitrace fixture replays correctly on one environment but
@@ -0,0 +1,4 @@
interface:
display_name: "MobileGL Mismatch Retrace Debugging"
short_description: "Localize the first divergent draw in a mismatching MobileGL retrace"
default_prompt: "Use $mismatch-retrace-debugging to find the first divergent render pass and draw call in a MobileGL retrace mismatch."
@@ -1,5 +1,5 @@
---
name: renderdoc-capture-trace-frame
name: renderdoc-debug-on-trace-replay
description: Capture and validate an exact frame from a MobileGL apitrace retrace on a connected Android device with RenderDoc/rdc-cli. Use for DirectVulkan or DirectGLES trace replay, mapping a target API call to an eglSwapBuffers frame, producing an .rdc plus a complete command manifest, checking capture stability, or troubleshooting Android TargetControl timing and replay failures.
---
@@ -19,20 +19,20 @@ adb -s SERIAL shell pm path top.mobilegl.plugin.trace
rdc doctor
```
4. Pass the unpacked `trace.trace`, its golden PNG, the fixture target call, backend, and output path to `tools/trace_replay/capture_android_retrace.py`.
4. Pass the unpacked `trace.trace`, its golden PNG, the fixture target call, backend, and output path to `tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py`.
## Capture
Let the tool infer the zero-based target swap from `eglSwapBuffers` calls:
```powershell
python tools/trace_replay/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectVulkan --output captures/case-vulkan.rdc --serial SERIAL --json
python tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectVulkan --output captures/case-vulkan.rdc --serial SERIAL --json
```
Change only the backend and output for GLES:
```powershell
python tools/trace_replay/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectGLES --output captures/case-gles.rdc --serial SERIAL --json
python tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/capture_android_retrace.py --trace .trace-work/case/trace.trace --golden tools/trace_replay/fixtures/case.0002667619.png --target-call 2667619 --backend DirectGLES --output captures/case-gles.rdc --serial SERIAL --json
```
Use `--target-swap N` when the mapping is already known. Use `--capture-frame N` only to override the backend rule deliberately.
@@ -0,0 +1,4 @@
interface:
display_name: "RenderDoc Debug on Trace Replay"
short_description: "Capture and debug an exact MobileGL trace-replay frame in RenderDoc"
default_prompt: "Use $renderdoc-debug-on-trace-replay to capture and validate an exact Android trace-replay frame in RenderDoc."
@@ -2,7 +2,7 @@
## TargetControl timing
- Start `tools/trace_replay/queue_android_frame.py` before launching `TraceReplayActivity`. A fast retrace can pass the requested frame before a late client connects.
- Start `tools/trace_replay/skills/renderdoc-debug-on-trace-replay/scripts/queue_android_frame.py` before launching `TraceReplayActivity`. A fast retrace can pass the requested frame before a late client connects.
- Keep TargetControl connected until `NewCapture` arrives. A queued request alone is not sufficient evidence that the RDC finished.
- Drain the asynchronous `RegisterAPI` and `CapturableWindowCount` messages before calling `QueueCapture`; otherwise `NewCapture` can be lost.
- Do not use the daemon-backed `rdc script` path for a capture that may exceed 30 seconds. Its outer RPC times out even when the device later writes a valid RDC. The repository helper imports the RenderDoc module discovered by `rdc` directly and has an independent capture timeout.
@@ -0,0 +1,137 @@
---
name: trace-fixture-authoring-on-android-fcl
description: Capture an on-device Android apitrace from FCL's MobileGL renderers. Use when preparing a reproducible MobileGL DirectGLES, Magma (DirectVulkan), or SimpleFPEWrapper rendering trace, marking the frame of a visual defect, pulling the resulting full.trace, or turning a device capture into a replay fixture.
---
# MobileGL Android trace capture
## Overview
Use FCL's Android `egltrace.so` wrapper, not Perfetto. When enabled before
launch, it records the complete EGL/GL call stream to `full.trace`. The game's
**MobileGL Trace → Capture** control marks the next swap frame in
`capture-result.json`; it does not start or stop recording and does not produce
a one-frame trace by itself.
Run commands from the FoldCraftLauncher repository root:
```sh
export REPO="$PWD"
export CAPTURE="$REPO/MobileGL/tools/trace_replay/skills/trace-fixture-authoring-on-android-fcl/scripts"
export SERIAL=<adb-device-serial> # omit --serial only if exactly one device is attached
```
The capture scripts are bundled inside this skill under `scripts/`; they
auto-detect the FoldCraftLauncher repository root from their own location, so
`--repo` only needs to be passed for a non-standard checkout layout.
## Prerequisites
- Use an FCL build containing `MobileGLTraceCapture` and the in-game Capture
menu entry.
- Select one of these renderers: MobileGL (DirectGLES), MobileGL Magma
(DirectVulkan), or SimpleFPEWrapper. MobileGlues is not supported by this
capture wrapper.
- Install `adb` and make it available on `PATH`; authorize USB debugging.
- Build the wrapper with Android NDK, CMake, Ninja, Python 3, and the checked
out in-tree `MobileGL/3rdparty/apitrace` submodule.
Confirm the attached device and ABI before building. The wrapper ABI must match
the device process ABI.
```sh
adb devices -l
adb -s "$SERIAL" shell getprop ro.product.cpu.abi
```
Use `arm64-v8a` for the usual `arm64-v8a` result; use the matching NDK ABI for
other devices.
## Build and install the wrapper
Build once per ABI or after changing apitrace/wrapper sources:
```sh
python3 "$CAPTURE/build_android_egltrace.py" --abi arm64-v8a
```
This generates `egltrace.so` under the skill's `scripts/out/` directory. Push
it and write FCL's enable sentinel before launching the game:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" install-wrapper
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" enable
```
The device-side control directory is `/sdcard/FCL/mobilegl-trace`. FCL copies
the shared `egltrace.so` into its private files directory at launch, replaces
the renderer's EGL library with it, and forwards to the real MobileGL library.
## Capture a reproduction
1. Start FCL after the wrapper and enable sentinel are in place. Select the
intended supported MobileGL renderer and launch the game.
2. Trace mode forces the game to `854x480`; account for that when reproducing
and comparing output.
3. Reproduce the issue. Start close to the target scene because tracing starts
when the game launches and trace files can grow rapidly.
4. At the desired visual state, open FCL's right-side game menu and press
**MobileGL Trace → Capture**. Let at least one frame present afterward.
5. Exit the game cleanly, then pull the latest session:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" pull-latest
```
The default local result is:
```text
.trace-work/pulled-mobilegl-captures/capture-YYYYMMDD-HHMMSS-<renderer>/
full.trace
capture-status.json
capture-result.json
```
`capture-result.json` must show `"status": "captured"`. Its `targetFrame`
is the one-based swap count used by `gltrim`; `zeroBasedFrame` is included for
tools that use zero-based indexing.
## Diagnose setup failures
Inspect the active device session directly:
```sh
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/latest-session.txt
adb -s "$SERIAL" shell ls -lh /sdcard/FCL/mobilegl-trace
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/capture-*/capture-status.json
adb -s "$SERIAL" shell cat /sdcard/FCL/mobilegl-trace/capture-*/capture-result.json
```
If `capture-status.json` reports a missing `egltrace.so`, rebuild/push the
correct ABI and relaunch. If `capture-result.json` is absent, Capture was
pressed without an active trace session, or no subsequent `eglSwapBuffers`
occurred. The menu button itself only writes `capture-once.request`.
Disable tracing when finished; otherwise the next supported MobileGL launch
will trace again:
```sh
python3 "$CAPTURE/adb_capture.py" --serial "$SERIAL" disable
```
## Create a replay fixture (optional)
Keep the raw `full.trace` until replay validation succeeds. To frame-trim and
package the marked frame for MobileGL trace replay, use the existing helper:
```sh
python3 "$CAPTURE/package_capture_fixture.py" \
--serial "$SERIAL" \
--case <case-name> \
--apitrace <path-to-in-tree-apitrace>
```
It pulls the latest capture if necessary, uses `capture-result.json` to select
the frame, runs `apitrace gltrim`, creates a golden image, and enforces the
fixture archive-size limit. Follow `../trace-fixture-authoring/SKILL.md` for
deterministic scene setup, verification, and registry changes.
@@ -0,0 +1,4 @@
interface:
display_name: "Trace Fixture Authoring on Android (FCL)"
short_description: "Capture and package a MobileGL trace fixture on Android FCL"
default_prompt: "Use $trace-fixture-authoring-on-android-fcl to capture a MobileGL trace on my Android device and package it into a replay fixture."
@@ -0,0 +1,3 @@
# Build output produced by build_android_egltrace.py (ABI-specific, regenerated).
out/
__pycache__/
@@ -0,0 +1,65 @@
#!/usr/bin/env python3
import argparse
import subprocess
from pathlib import Path
REMOTE_ROOT = "/sdcard/FCL/mobilegl-trace"
SCRIPT_DIR = Path(__file__).resolve().parent
DEFAULT_WRAPPER = SCRIPT_DIR / "out" / "egltrace.so"
def adb(serial, args):
cmd = ["adb"]
if serial:
cmd += ["-s", serial]
cmd += args
print("+", " ".join(cmd), flush=True)
subprocess.run(cmd, check=True)
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--serial")
sub = parser.add_subparsers(dest="command", required=True)
install = sub.add_parser("install-wrapper")
install.add_argument("--wrapper", default=str(DEFAULT_WRAPPER))
sub.add_parser("enable")
sub.add_parser("disable")
sub.add_parser("capture-once")
pull = sub.add_parser("pull-latest")
pull.add_argument("--output", default=".trace-work/pulled-mobilegl-captures")
args = parser.parse_args()
serial = args.serial
if args.command == "install-wrapper":
wrapper = Path(args.wrapper)
if not wrapper.exists():
raise SystemExit(f"missing wrapper: {wrapper}")
adb(serial, ["shell", "mkdir", "-p", REMOTE_ROOT])
adb(serial, ["push", str(wrapper), f"{REMOTE_ROOT}/egltrace.so"])
elif args.command == "enable":
adb(serial, ["shell", "mkdir", "-p", REMOTE_ROOT])
adb(serial, ["shell", f"printf enabled > {REMOTE_ROOT}/enable"])
elif args.command == "disable":
adb(serial, ["shell", "rm", "-f", f"{REMOTE_ROOT}/enable"])
elif args.command == "capture-once":
adb(serial, ["shell", "mkdir", "-p", REMOTE_ROOT])
adb(serial, ["shell", f"date +%s%3N > {REMOTE_ROOT}/capture-once.request"])
elif args.command == "pull-latest":
tmp = subprocess.check_output((["adb"] + (["-s", serial] if serial else []) +
["shell", "cat", f"{REMOTE_ROOT}/latest-session.txt"]),
text=True, encoding="utf-8", errors="replace").strip()
if not tmp:
raise SystemExit("no latest-session.txt on device")
output = Path(args.output)
output.mkdir(parents=True, exist_ok=True)
adb(serial, ["pull", tmp, str(output / Path(tmp).name)])
if __name__ == "__main__":
main()
@@ -0,0 +1,191 @@
cmake_minimum_required(VERSION 3.22.1)
project(mobilegl_android_egltrace)
set(APITRACE_ROOT "" CACHE PATH "Path to apitrace source tree")
set(PATCHED_EGLTRACE_CPP "" CACHE FILEPATH "Generated and patched egltrace.cpp")
set(PATCHED_GLPROC_EGL_CPP "" CACHE FILEPATH "Patched glproc_egl.cpp")
if(NOT EXISTS "${APITRACE_ROOT}/wrappers")
message(FATAL_ERROR "APITRACE_ROOT must point to apitrace")
endif()
if(NOT EXISTS "${PATCHED_EGLTRACE_CPP}")
message(FATAL_ERROR "PATCHED_EGLTRACE_CPP is required")
endif()
if(NOT EXISTS "${PATCHED_GLPROC_EGL_CPP}")
message(FATAL_ERROR "PATCHED_GLPROC_EGL_CPP is required")
endif()
set(APITRACE_BINARY_DIR "${CMAKE_CURRENT_BINARY_DIR}/apitrace")
set(APITRACE_VERSION "mobilegl-capture")
find_package(Python3 REQUIRED)
find_package(Threads REQUIRED)
include("${APITRACE_ROOT}/cmake/ConvenienceLibrary.cmake")
set(BUILD_TESTING OFF CACHE BOOL "" FORCE)
set(ENABLE_STATIC_SNAPPY ON CACHE BOOL "" FORCE)
set(DOC_INSTALL_DIR "doc" CACHE PATH "" FORCE)
set(HAVE_X86 OFF CACHE BOOL "" FORCE)
set(ZLIB_FOUND OFF CACHE BOOL "" FORCE)
set(PNG_FOUND OFF CACHE BOOL "" FORCE)
set(Snappy_FOUND OFF CACHE BOOL "" FORCE)
set(BROTLIDEC_FOUND OFF CACHE BOOL "" FORCE)
set(BROTLIENC_FOUND OFF CACHE BOOL "" FORCE)
set(ZSTD_FOUND OFF CACHE BOOL "" FORCE)
set(CMAKE_EXECUTABLE_FORMAT "MobileGLAndroid" CACHE INTERNAL "" FORCE)
add_custom_target(check)
add_subdirectory("${APITRACE_ROOT}/thirdparty" "${APITRACE_BINARY_DIR}/thirdparty")
set(APITRACE_GENERATED_DIR "${APITRACE_BINARY_DIR}/generated")
file(MAKE_DIRECTORY "${APITRACE_GENERATED_DIR}")
configure_file("${APITRACE_ROOT}/version.h.in" "${APITRACE_GENERATED_DIR}/version.h" @ONLY)
add_custom_command(
OUTPUT
"${APITRACE_GENERATED_DIR}/glproc.hpp"
"${APITRACE_GENERATED_DIR}/glproc.cpp"
COMMAND ${Python3_EXECUTABLE}
"${APITRACE_ROOT}/dispatch/glproc.py"
"${APITRACE_GENERATED_DIR}/glproc.hpp"
"${APITRACE_GENERATED_DIR}/glproc.cpp"
DEPENDS
"${APITRACE_ROOT}/dispatch/glproc.py"
"${APITRACE_ROOT}/dispatch/dispatch.py"
"${APITRACE_ROOT}/specs/wglapi.py"
"${APITRACE_ROOT}/specs/glxapi.py"
"${APITRACE_ROOT}/specs/cglapi.py"
"${APITRACE_ROOT}/specs/eglapi.py"
"${APITRACE_ROOT}/specs/glapi.py"
"${APITRACE_ROOT}/specs/gltypes.py"
"${APITRACE_ROOT}/specs/stdapi.py")
add_library(apitrace_os STATIC
"${APITRACE_ROOT}/lib/os/os_backtrace.cpp"
"${APITRACE_ROOT}/lib/os/os_crtdbg.cpp"
"${APITRACE_ROOT}/lib/os/os_posix.cpp")
target_include_directories(apitrace_os PUBLIC
"${APITRACE_ROOT}/compat"
"${APITRACE_ROOT}/thirdparty"
"${APITRACE_ROOT}/lib/os"
"${APITRACE_ROOT}/lib/trace")
target_link_libraries(apitrace_os PUBLIC Threads::Threads)
add_library(glproc STATIC
"${APITRACE_GENERATED_DIR}/glproc.cpp"
"${PATCHED_GLPROC_EGL_CPP}")
target_include_directories(glproc PUBLIC
"${APITRACE_GENERATED_DIR}"
"${APITRACE_ROOT}/wrappers"
"${APITRACE_ROOT}/dispatch"
"${APITRACE_ROOT}/lib/os"
"${APITRACE_ROOT}/thirdparty/khronos")
target_link_libraries(glproc PUBLIC apitrace_os dl)
add_library(highlight STATIC "${APITRACE_ROOT}/lib/highlight/highlight.cpp")
target_include_directories(highlight PUBLIC "${APITRACE_ROOT}/lib/highlight")
add_library(guids STATIC "${APITRACE_ROOT}/lib/guids/guids.cpp")
target_include_directories(guids PUBLIC
"${APITRACE_ROOT}/lib/guids"
"${APITRACE_ROOT}/lib/os")
add_library(common STATIC
"${APITRACE_ROOT}/lib/trace/trace_callset.cpp"
"${APITRACE_ROOT}/lib/trace/trace_dump.cpp"
"${APITRACE_ROOT}/lib/trace/trace_fast_callset.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_read.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_zlib.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_brotli.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_snappy.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_zstd.cpp"
"${APITRACE_ROOT}/lib/trace/trace_file_zstd_seekable.cpp"
"${APITRACE_ROOT}/lib/trace/trace_model.cpp"
"${APITRACE_ROOT}/lib/trace/trace_option.cpp"
"${APITRACE_ROOT}/lib/trace/trace_ostream_snappy.cpp"
"${APITRACE_ROOT}/lib/trace/trace_ostream_zlib.cpp"
"${APITRACE_ROOT}/lib/trace/trace_ostream_zstd.cpp"
"${APITRACE_ROOT}/lib/trace/trace_parser.cpp"
"${APITRACE_ROOT}/lib/trace/trace_parser_flags.cpp"
"${APITRACE_ROOT}/lib/trace/trace_parser_loop.cpp"
"${APITRACE_ROOT}/lib/trace/trace_profiler.cpp"
"${APITRACE_ROOT}/lib/trace/trace_writer.cpp"
"${APITRACE_ROOT}/lib/trace/trace_writer_local.cpp"
"${APITRACE_ROOT}/lib/trace/trace_writer_model.cpp")
target_include_directories(common PUBLIC
"${APITRACE_ROOT}/compat"
"${APITRACE_ROOT}/thirdparty"
"${APITRACE_ROOT}/lib/guids"
"${APITRACE_ROOT}/lib/highlight"
"${APITRACE_ROOT}/lib/os"
"${APITRACE_ROOT}/lib/trace"
"${APITRACE_ROOT}/lib/ubjson")
target_link_libraries(common PUBLIC
guids
highlight
apitrace_os
Snappy::snappy
ZLIB::ZLIB
PkgConfig::BROTLIDEC
PkgConfig::ZSTD
zstd_seekable)
add_convenience_library(trace
"${APITRACE_ROOT}/wrappers/memtrace.hpp"
"${APITRACE_ROOT}/wrappers/memtrace.cpp")
target_include_directories(trace PUBLIC
"${APITRACE_ROOT}/thirdparty/crc32c")
target_link_libraries(trace
common
guids
crc32c)
add_library(glhelpers STATIC
"${APITRACE_ROOT}/helpers/glfeatures.cpp"
"${APITRACE_ROOT}/helpers/eglsize.cpp")
target_include_directories(glhelpers PUBLIC
"${APITRACE_GENERATED_DIR}"
"${APITRACE_ROOT}/dispatch"
"${APITRACE_ROOT}/helpers"
"${APITRACE_ROOT}/lib/os"
"${APITRACE_ROOT}/thirdparty/khronos")
target_link_libraries(glhelpers PUBLIC glproc apitrace_os)
add_convenience_library(gltrace_common
"${APITRACE_ROOT}/wrappers/glcaps.cpp"
"${APITRACE_ROOT}/wrappers/config.cpp"
"${APITRACE_ROOT}/wrappers/gltrace_arrays.cpp"
"${APITRACE_ROOT}/wrappers/gltrace_state.cpp"
"${APITRACE_ROOT}/wrappers/glmemshadow.hpp"
"${APITRACE_ROOT}/wrappers/glmemshadow.cpp"
"${APITRACE_ROOT}/wrappers/gltrace_unpack_compressed.hpp"
"${APITRACE_ROOT}/wrappers/gltrace_unpack_compressed.cpp")
add_dependencies(gltrace_common glproc)
target_include_directories(gltrace_common PUBLIC
"${APITRACE_ROOT}/wrappers")
target_link_libraries(gltrace_common
glhelpers
trace)
add_library(egltrace SHARED
"${PATCHED_EGLTRACE_CPP}"
"${APITRACE_ROOT}/wrappers/dlsym.cpp"
"${PATCHED_GLPROC_EGL_CPP}")
add_dependencies(egltrace glproc)
set_target_properties(egltrace PROPERTIES PREFIX "")
target_compile_definitions(egltrace PRIVATE -DEGLTRACE=1)
target_include_directories(egltrace PRIVATE
"${APITRACE_ROOT}/wrappers"
"${APITRACE_GENERATED_DIR}"
"${APITRACE_ROOT}/helpers"
"${APITRACE_ROOT}/dispatch"
"${APITRACE_ROOT}/lib/os"
"${APITRACE_ROOT}/lib/trace"
"${APITRACE_ROOT}/thirdparty/khronos")
target_link_libraries(egltrace
gltrace_common
glproc
Threads::Threads
dl)
@@ -0,0 +1,234 @@
#!/usr/bin/env python3
import argparse
import os
import shutil
import subprocess
import sys
from pathlib import Path
# This script is bundled inside the trace-fixture-authoring-on-android-fcl skill at
# <FCL>/MobileGL/tools/trace_replay/skills/trace-fixture-authoring-on-android-fcl/scripts/.
SCRIPT_DIR = Path(__file__).resolve().parent
DEFAULT_REPO = SCRIPT_DIR.parents[5] # -> FoldCraftLauncher repo root
DEFAULT_OUTPUT = SCRIPT_DIR / "out" / "egltrace.so"
def run(cmd, cwd=None):
print("+", " ".join(str(part) for part in cmd), flush=True)
subprocess.run(cmd, cwd=cwd, check=True)
def find_ndk(repo):
for key in ("ANDROID_NDK_HOME", "ANDROID_NDK_ROOT"):
value = os.environ.get(key)
if value:
return Path(value)
for key in ("ANDROID_HOME", "ANDROID_SDK_ROOT"):
value = os.environ.get(key)
if value:
ndk_root = Path(value) / "ndk"
if ndk_root.exists():
versions = sorted([p for p in ndk_root.iterdir() if p.is_dir()])
if versions:
return versions[-1]
local = repo / "local.properties"
if local.exists():
sdk = None
ndk = None
for line in local.read_text(encoding="utf-8", errors="ignore").splitlines():
if line.startswith("sdk.dir="):
sdk = Path(line.split("=", 1)[1].replace("\\:", ":"))
if line.startswith("ndk.dir="):
ndk = Path(line.split("=", 1)[1].replace("\\:", ":"))
if ndk:
return ndk
if sdk:
ndk_root = sdk / "ndk"
if ndk_root.exists():
versions = sorted([p for p in ndk_root.iterdir() if p.is_dir()])
if versions:
return versions[-1]
raise SystemExit("Android NDK not found; set ANDROID_NDK_HOME or local.properties sdk.dir/ndk.dir")
def generate_and_patch(repo, build_dir):
wrapper_dir = repo / "MobileGL" / "3rdparty" / "apitrace" / "wrappers"
generated = build_dir / "patched" / "egltrace.cpp"
generated.parent.mkdir(parents=True, exist_ok=True)
with generated.open("w", encoding="utf-8", newline="\n") as out:
subprocess.run([sys.executable, str(wrapper_dir / "egltrace.py")], cwd=wrapper_dir, stdout=out, check=True)
text = generated.read_text(encoding="utf-8")
helper = r'''
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <time.h>
#include <unistd.h>
static unsigned long long mobilegl_capture_swap_count = 0;
static long long mobilegl_capture_time_ms(void) {
struct timespec ts;
clock_gettime(CLOCK_REALTIME, &ts);
return (long long) ts.tv_sec * 1000LL + ts.tv_nsec / 1000000LL;
}
static int mobilegl_capture_exists(const char *path) {
return path != NULL && path[0] != '\0' && access(path, F_OK) == 0;
}
static void mobilegl_capture_record_request(void) {
++mobilegl_capture_swap_count;
const char *request = getenv("MOBILEGL_TRACE_CAPTURE_REQUEST_FILE");
const char *output = getenv("MOBILEGL_TRACE_CAPTURE_FRAME_FILE");
if (!mobilegl_capture_exists(request) || output == NULL || output[0] == '\0') {
return;
}
unlink(request);
FILE *file = fopen(output, "w");
if (file == NULL) {
return;
}
const char *trace_file = getenv("TRACE_FILE");
fprintf(file,
"{\n"
" \"status\": \"captured\",\n"
" \"swapCount\": %llu,\n"
" \"targetFrame\": %llu,\n"
" \"zeroBasedFrame\": %llu,\n"
" \"capturedAtMs\": %lld,\n"
" \"traceFile\": \"%s\"\n"
"}\n",
mobilegl_capture_swap_count,
mobilegl_capture_swap_count,
mobilegl_capture_swap_count == 0 ? 0 : mobilegl_capture_swap_count - 1,
mobilegl_capture_time_ms(),
trace_file == NULL ? "" : trace_file);
fclose(file);
}
'''
insert_at = text.find("#include")
if insert_at < 0:
raise SystemExit("generated egltrace.cpp has no include block")
next_block = text.find("\n\n", insert_at)
text = text[:next_block] + "\n" + helper + text[next_block:]
needle = "EGLBoolean EGLAPIENTRY eglSwapBuffers(EGLDisplay dpy, EGLSurface surface)"
start = text.find(needle)
if start < 0:
raise SystemExit("generated egltrace.cpp has no eglSwapBuffers wrapper to patch")
brace = text.find("{", start)
if brace < 0:
raise SystemExit("eglSwapBuffers wrapper has no function body")
text = text[:brace + 1] + "\n mobilegl_capture_record_request();" + text[brace + 1:]
generated.write_text(text, encoding="utf-8", newline="\n")
return generated
def patch_glproc_egl(repo, build_dir):
source = repo / "MobileGL" / "3rdparty" / "apitrace" / "wrappers" / "glproc_egl.cpp"
patched = build_dir / "patched" / "glproc_egl.cpp"
patched.parent.mkdir(parents=True, exist_ok=True)
text = source.read_text(encoding="utf-8")
text = text.replace('#include "dlopen.hpp"\n', '#include "dlopen.hpp"\n#include <stdlib.h>\n')
needle = """void *
_getPublicProcAddress(const char *procName)
{
void *proc;
"""
replacement = """void *
_getPublicProcAddress(const char *procName)
{
void *proc;
static void *traceLibGL = NULL;
static bool triedTraceLibGL = false;
if (!triedTraceLibGL) {
triedTraceLibGL = true;
const char *traceLibGLName = getenv("TRACE_LIBGL");
if (traceLibGLName && traceLibGLName[0]) {
traceLibGL = _dlopen(traceLibGLName, RTLD_GLOBAL | RTLD_LAZY | RTLD_DEEPBIND);
}
}
if (traceLibGL) {
proc = dlsym(traceLibGL, procName);
if (proc) {
return proc;
}
}
"""
if needle not in text:
raise SystemExit("glproc_egl.cpp patch point not found")
text = text.replace(needle, replacement, 1)
patched.write_text(text, encoding="utf-8", newline="\n")
return patched
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--repo", default=str(DEFAULT_REPO), help="FoldCraftLauncher repo root")
parser.add_argument("--abi", default="arm64-v8a")
parser.add_argument("--android-platform", default="android-23")
parser.add_argument("--build-dir", default=".trace-work/build-android-egltrace")
parser.add_argument("--output", default=str(DEFAULT_OUTPUT))
args = parser.parse_args()
repo = Path(args.repo).resolve()
ndk = find_ndk(repo)
build_dir = (repo / args.build_dir / args.abi).resolve()
apitrace = repo / "MobileGL" / "3rdparty" / "apitrace"
source_dir = SCRIPT_DIR / "android_egltrace"
toolchain = ndk / "build" / "cmake" / "android.toolchain.cmake"
if not apitrace.exists():
raise SystemExit(f"missing apitrace checkout: {apitrace}")
if not source_dir.exists():
raise SystemExit(f"missing wrapper CMake project: {source_dir}")
if not toolchain.exists():
raise SystemExit(f"missing Android toolchain: {toolchain}")
cache = build_dir / "CMakeCache.txt"
if cache.exists() and "CMAKE_GENERATOR:INTERNAL=Ninja" not in cache.read_text(encoding="utf-8", errors="ignore"):
shutil.rmtree(build_dir)
elif cache.exists() and str(source_dir).replace("\\", "/") not in cache.read_text(encoding="utf-8", errors="ignore").replace("\\", "/"):
shutil.rmtree(build_dir)
ninja = shutil.which("ninja")
if ninja is None:
cmake_ninjas = sorted((Path(os.environ.get("ANDROID_HOME", "")) / "cmake").glob("*/bin/ninja.exe"))
ninja = str(cmake_ninjas[-1]) if cmake_ninjas else None
if ninja is None:
raise SystemExit("ninja not found; install Ninja or Android SDK CMake")
patched_egltrace = generate_and_patch(repo, build_dir)
patched_glproc_egl = patch_glproc_egl(repo, build_dir)
run([
"cmake", "-G", "Ninja", "-S", str(source_dir), "-B", str(build_dir),
"-DCMAKE_BUILD_TYPE=Release",
f"-DCMAKE_TOOLCHAIN_FILE={toolchain}",
f"-DCMAKE_MAKE_PROGRAM={ninja}",
f"-DANDROID_ABI={args.abi}",
f"-DANDROID_PLATFORM={args.android_platform}",
f"-DAPITRACE_ROOT={apitrace}",
f"-DPATCHED_EGLTRACE_CPP={patched_egltrace}",
f"-DPATCHED_GLPROC_EGL_CPP={patched_glproc_egl}",
])
run(["cmake", "--build", str(build_dir), "--target", "egltrace", "--parallel"])
output = Path(args.output)
if not output.is_absolute():
output = repo / output
output = output.resolve()
output.parent.mkdir(parents=True, exist_ok=True)
candidates = list(build_dir.rglob("egltrace.so"))
if not candidates:
raise SystemExit("egltrace.so was not produced")
shutil.copy2(candidates[0], output)
print(output)
if __name__ == "__main__":
main()
@@ -0,0 +1,168 @@
#!/usr/bin/env python3
import argparse
import json
import re
import shutil
import subprocess
import tarfile
from pathlib import Path
DEFAULT_MAX_ARCHIVE_BYTES = 20 * 1024 * 1024
# Bundled under the skill at .../skills/trace-fixture-authoring-on-android-fcl/scripts/.
SCRIPT_DIR = Path(__file__).resolve().parent
DEFAULT_REPO = SCRIPT_DIR.parents[5] # -> FoldCraftLauncher repo root
def run(cmd, cwd=None, capture=False):
print("+", " ".join(str(part) for part in cmd), flush=True)
if capture:
return subprocess.check_output(cmd, cwd=cwd, text=True, encoding="utf-8", errors="replace")
subprocess.run(cmd, cwd=cwd, check=True)
return ""
def adb(args, serial=None):
cmd = ["adb"]
if serial:
cmd += ["-s", serial]
cmd += args
return run(cmd, capture=True)
def pull_latest(serial, dest):
latest = adb(["shell", "cat", "/sdcard/FCL/mobilegl-trace/latest-session.txt"], serial).strip()
if not latest:
raise SystemExit("device has no /sdcard/FCL/mobilegl-trace/latest-session.txt")
dest.mkdir(parents=True, exist_ok=True)
local = dest / Path(latest).name
if local.exists():
shutil.rmtree(local)
adb(["pull", latest, str(local)], serial)
return local
def choose_target_frame(capture_dir, explicit_frame):
if explicit_frame is not None:
return explicit_frame
result = capture_dir / "capture-result.json"
if not result.exists():
raise SystemExit(f"missing {result}; press the FCL capture button or create capture-once.request first")
data = json.loads(result.read_text(encoding="utf-8"))
if "targetFrame" not in data:
raise SystemExit(f"{result} has no targetFrame")
return int(data["targetFrame"])
def choose_snapshot(golden_dir):
pngs = sorted(golden_dir.glob("*.png"))
if not pngs:
raise SystemExit(f"no snapshots produced in {golden_dir}")
def call_no(path):
match = re.search(r"\.(\d+)\.png$", path.name)
return int(match.group(1)) if match else -1
return max(pngs, key=call_no)
def choose_target_call(explicit_call, golden):
if explicit_call is not None:
return explicit_call
if golden is None:
return None
match = re.search(r"\.(\d+)\.png$", golden.name)
return int(match.group(1)) if match else None
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--repo", default=str(DEFAULT_REPO), help="FoldCraftLauncher repo root")
parser.add_argument("--serial", help="adb serial; when set, pull latest capture from device")
parser.add_argument("--capture-dir", help="local capture directory; defaults to pulled latest")
parser.add_argument("--pull-root", default=".trace-work/pulled-mobilegl-captures")
parser.add_argument("--case", required=True)
parser.add_argument("--target-frame", type=int)
parser.add_argument("--target-call", type=int)
parser.add_argument("--golden", help="existing golden PNG, normally produced by Android replay")
parser.add_argument("--skip-desktop-golden", action="store_true",
help="skip apitrace replay --headless; requires --golden and --target-call")
parser.add_argument("--apitrace", default="apitrace")
parser.add_argument("--fixtures-dir", default="MobileGL/tools/trace_replay/fixtures")
parser.add_argument("--width", type=int, default=854)
parser.add_argument("--height", type=int, default=480)
parser.add_argument("--ssim-threshold", default="0.99")
parser.add_argument("--max-archive-bytes", type=int, default=DEFAULT_MAX_ARCHIVE_BYTES)
args = parser.parse_args()
repo = Path(args.repo).resolve()
if args.capture_dir:
capture_dir = Path(args.capture_dir).resolve()
elif args.serial:
capture_dir = pull_latest(args.serial, repo / args.pull_root)
else:
raise SystemExit("pass --capture-dir or --serial")
full_trace = capture_dir / "full.trace"
if not full_trace.exists():
raise SystemExit(f"missing trace: {full_trace}")
target_frame = choose_target_frame(capture_dir, args.target_frame)
work = capture_dir / "fixture-work"
if work.exists():
shutil.rmtree(work)
work.mkdir(parents=True)
frames_txt = work / "frames.txt"
frames_txt.write_text(run([args.apitrace, "dump", "--calls=frame", str(full_trace)], capture=True), encoding="utf-8")
trimmed = work / "trace.trace"
run([args.apitrace, "gltrim", "-f", str(target_frame), "--output", str(trimmed), str(full_trace)])
supplied_golden = Path(args.golden).resolve() if args.golden else None
target_call = choose_target_call(args.target_call, supplied_golden)
if args.skip_desktop_golden:
if supplied_golden is None or target_call is None:
raise SystemExit("--skip-desktop-golden requires --golden and --target-call")
golden = supplied_golden
else:
golden_dir = work / "golden"
golden_dir.mkdir()
prefix = golden_dir / f"{args.case}."
run([args.apitrace, "replay", "--headless", "--snapshot-prefix", str(prefix), "--call-nos", str(trimmed)])
golden = choose_snapshot(golden_dir)
target_call = choose_target_call(args.target_call, golden)
if target_call is None:
raise SystemExit(f"cannot infer target call from {golden}")
fixtures = (repo / args.fixtures_dir).resolve()
fixtures.mkdir(parents=True, exist_ok=True)
archive_root = work / "archive"
archive_root.mkdir()
shutil.copy2(trimmed, archive_root / "trace.trace")
tgz = fixtures / f"{args.case}.tgz"
with tarfile.open(tgz, "w:gz") as tar:
tar.add(archive_root / "trace.trace", arcname="trace.trace")
archive_size = tgz.stat().st_size
if archive_size > args.max_archive_bytes:
raise SystemExit(
f"{tgz} is {archive_size} bytes, over the {args.max_archive_bytes} byte fixture limit; "
"choose an earlier/smaller frame and re-run gltrim"
)
golden_out = fixtures / f"{args.case}.{target_call:010d}.png"
shutil.copy2(golden, golden_out)
manifest = {
"name": args.case,
"trace_archive": tgz.name,
"trace_file": "trace.trace",
"golden": golden_out.name,
"target_call": target_call,
"width": args.width,
"height": args.height,
"ssim_threshold": float(args.ssim_threshold),
"archive_size": archive_size,
}
manifest_path = capture_dir / f"{args.case}.fixture.json"
manifest_path.write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
print(json.dumps(manifest, indent=2))
if __name__ == "__main__":
main()
@@ -1,3 +1,8 @@
---
name: trace-fixture-authoring
description: Author a deterministic MobileGL trace-replay fixture from a captured apitrace - build the in-tree apitrace fork, capture a reproducible scene, frame-trim with gltrim, generate and verify a golden image, package under the archive-size budget, register the case in trace_cases.json, and validate on Linux and Android. Use when adding or re-trimming a trace_replay regression fixture.
---
# Trace fixture authoring
## Variables
@@ -75,10 +80,19 @@ Minecraft specifics that keep the capture deterministic and small:
`doMobSpawning`, `randomTickSpeed 0`, a fixed `DayTime`, and the player
`Rotation` that frames the intended subject. The camera snaps to the saved
rotation on world join, so composition is edited in the save, not in-game.
- `options.txt`: `pauseOnLostFocus:false`, a low `maxFps` (10 works), and a
small `renderDistance` (3). Frame rate and render distance are the two main
levers on trace size; a ~35 s in-world session at 10 fps lands well under
the archive budget after repack.
- `options.txt`: `pauseOnLostFocus:false`, a low `maxFps` (10 works), a small
`renderDistance` (3), and the capture resolution pinned to `WIDTH` x `HEIGHT`
(854x480) via `overrideWidth`/`overrideHeight` (or `--width`/`--height`).
These are the main levers on fixture size: a low frame rate keeps the full
trace short, a small render distance keeps per-frame geometry down, and the
854x480 resolution keeps every render target the frame references small (a
trimmed frame's framebuffer/attachment textures scale with resolution
squared). A ~35 s in-world session at 10 fps and 854x480 lands well under the
archive budget after repack.
- `maxFps` has a practical floor: Minecraft ignores values below ~10 and falls
back to unlimited/vsync (a `maxFps:1` capture rendered ~60 fps and ballooned
the trace). 10 is as low as this lever goes, so do not count on a lower frame
rate to shrink the frame count further.
- Enter the world non-interactively with `--quickPlaySingleplayer <world>` so
every capture takes the same path from boot to gameplay.
- Keep the game window UNFOCUSED for the whole capture (focus the desktop
@@ -111,6 +125,61 @@ For Java:
An `@argfile` with the full JVM+game command line keeps the invocation
reproducible across recaptures.
NEVER put a real credential on the traced command line. apitrace records the
traced process's argv into the trace as a `process.commandLine` property, so
anything passed there - `--accessToken`, session tokens, API keys - is embedded
in the trace and ships inside the committed fixture. Minecraft never validates
`--accessToken` for singleplayer, so pass a placeholder (`--accessToken 0`);
`--username`/`--uuid` are public and may stay real. Before packaging, grep the
UNCOMPRESSED trace for the secret to confirm it is absent:
```sh
"$APITRACE" repack "$WORK/$CASE/trace.trace" /tmp/plain.trace # decompress
grep -ac "<secret-prefix>" /tmp/plain.trace # must be 0
```
If a secret has already been captured, it can be scrubbed in place instead of
recapturing: apitrace's snappy container is `[length][raw snappy]` chunks with
no checksum, and a high-entropy secret is stored as literal bytes, so replacing
those bytes with an EQUAL-LENGTH filler keeps the container valid and leaves the
GL call stream byte-identical. Blank every maximal run of the secret (it splits
across chunks), then verify: frame count unchanged, the decompressed trace no
longer contains the secret, and the replayed target frame still matches the
golden. Treat any already-pushed trace as leaked regardless - rotate the
credential, since a force-push does not purge the LFS object from the remote.
Watch for vendor-gated shader paths. Shader packs branch on the GL vendor that
Iris injects (`MC_GL_VENDOR_NVIDIA` / `_AMD` / ...) and compile a
vendor-exclusive path, so capturing on an NVIDIA card can bake NVIDIA-only GLSL
into the fixture (iterationRP selects `subgroupPartitionNV` /
`GL_NV_shader_subgroup_partitioned` instead of the portable
`subgroupShuffleXor`). Iris resolves the `#ifdef` before `glShaderSource`, so
only the taken branch is in the trace and the fixture cannot replay on the
mobile GPUs MobileGL targets. Rather than hunting for a second GPU (the Windows
per-app GPU preference does NOT change which OpenGL ICD is loaded), mask the
vendor at capture time with apitrace's own config - point `GLTRACE_CONF` at a
file containing:
```
GL_VENDOR = "NoVIDIA (MobileGL spoof)"
GL_RENDERER = "NoVIDIA (MobileGL spoof)"
```
The wrapper then returns that from `glGetString`, so the pack compiles the
portable path while still running on the fast driver. Pick a string that does
NOT contain the real vendor name as a substring (Iris matches by substring, so
"Not NVIDIA ..." would still match) and that is self-describing, so nobody later
mistakes the trace for a capture on different hardware. Afterwards, grep the
decoded trace to confirm the vendor-exclusive symbols are gone:
```sh
"$APITRACE" dump "$WORK/$CASE/full.trace" | grep -c subgroupPartitionNV # must be 0
```
Software rasterisers are not a substitute here: llvmpipe exposes no
`GL_KHR_shader_subgroup` at all, and packs that use subgroup ops unguarded
cannot run on it in any vendor configuration.
Keep `full.trace` until both backends are validated.
Persistent-mapped buffers: apps may legally write a `GL_MAP_PERSISTENT_BIT`
@@ -172,12 +241,37 @@ name"-style retrace warnings. If content is missing from the trimmed trace
but present in the full trace, the fix belongs in `3rdparty/apitrace`'s
frametrim, not in the fixture.
Temporal shaders (auto-exposure / eye adaptation, TAA, temporal reflections -
e.g. the iterationRP shader pack) break a single-frame `gltrim -f`: the target
frame reads its predecessors' feedback buffers, which the isolated frame no
longer contains, so a mid-sequence frame replays overexposed to white (and any
timed name/version overlay the pack draws in its first seconds silently drops).
The symptom is a trimmed frame that looks blown-out or washed while the same
frame of `full.trace` renders correctly, and it gets worse the later the frame.
When a single-frame trim of such a pack cannot be made to render correctly, keep
the temporal history instead of the dependency slice: select an early in-world
target frame and trim a PREFIX with `apitrace trim --calls=0-<target-swap-call>`
(it preserves call numbers, so `target_call` is just that swap call). The prefix
replays every frame up to the target, so its temporal buffers are correct.
Prefer the earliest frame that already shows the intended subject - fewer lead-in
frames means a smaller archive and a faster CI replay. This deviates from the
single-frame rule deliberately; note it in the README entry.
## Generate golden
Generate frame snapshots from the trimmed trace, then choose the snapshot that
matches the selected frame. The target call used by replay registration must
come from the trimmed trace, not from a call-filtered full-trace selection.
Generate the golden with the same GL stack the scene was captured on. A headless
software renderer (llvmpipe) is fine for vanilla and light packs, but heavy
ray-traced shader packs (compute-driven atmosphere LUTs, screen-space tracing -
e.g. iterationRP) render as solid black or blown-out white under llvmpipe. Drive
the golden from a real GPU instead: on Windows a stock `glretrace.exe` (an
upstream apitrace release works for replay even on an in-tree-fork trace) replays
the trace on the discrete GPU and snapshots the target call. Read the resulting
PNG back and confirm the subject actually rendered before trusting it as golden.
```sh
mkdir -p "$WORK/$CASE/golden"
"$APITRACE" replay --headless \
@@ -255,6 +349,15 @@ Check the final archive size. The committed fixture archive should be less than
with a shorter run or a lower frame rate / render distance instead of adding
call-based filtering.
Some packs have an irreducible size floor: a large static lookup table baked
into the pack (iterationRP ships a ~17 MiB half-float atmosphere LUT that the
target frame samples) lands in the trace once and does not compress, so every
variant - single frame, prefix, or full - sits near the same size regardless of
frame count. When the floor alone exceeds the budget, neither a lower frame rate
nor fewer frames helps; confirm the fixture is worth the exception and record the
measured size in the case's README entry rather than chasing an unreachable
target.
```sh
du -h "$REPO/tools/trace_replay/fixtures/$CASE.tgz"
tar -tzf "$REPO/tools/trace_replay/fixtures/$CASE.tgz"
@@ -0,0 +1,4 @@
interface:
display_name: "MobileGL Trace Fixture Authoring"
short_description: "Author and register a MobileGL trace-replay fixture"
default_prompt: "Use $trace-fixture-authoring to author and register a MobileGL trace replay fixture."
+1 -2
View File
@@ -276,12 +276,11 @@
},
{
"name": "improved-transparency-minecraft-26.3",
"ci": false,
"trace_archive": "improved-transparency-minecraft-26.3.tgz",
"golden": "improved-transparency-minecraft-26.3.0002667619.png",
"target_call": 2667619,
"timeout_seconds": 1800,
"coherent_as_flush": true
"ssim_threshold": 0.995
},
{
"name": "minecraft-1.21.4-fabric-iris-iterationrp-in-world",