Compare commits

...
38 Commits
Author SHA1 Message Date
BZLZHH e86a9bbec5 [Chore] (MG_Backend): restore target GL version to 3.3 2026-07-31 16:45:19 +08:00
swung0x48 c5569e71b3 [Feat] (tools/cts): automate Windows WGL conformance runs 2026-07-30 23:00:13 -04:00
swung0x48 7e8c32a063 [Feat] (MG_Backend, MG_Impl): expose experimental GL 4.6 CTS limits 2026-07-30 23:00:13 -04:00
swung0x48 77bd03d962 [Chore]: remove unnecessary doc 2026-07-30 21:56:17 -04:00
swung0x48 37111ae992 [Perf] (DirectVulkan): snapshot-gated consecutive-draw fast path skips SetupDraw re-resolution 2026-07-30 09:40:54 -04:00
swung0x48 6b0c2a15ab [Perf] (DirectVulkan): reuse unchanged global-UBO slices and skip identical descriptor binds 2026-07-30 08:18:21 -04:00
swung0x48 e9ffd99313 [Perf] (DirectVulkan): bake the attribute location mask and memoize the explicit-LOD eligibility probe 2026-07-30 08:18:20 -04:00
swung0x48 7c01ddea0c [Perf] (DirectVulkan): drop per-draw weak-ptr locks, re-resolves and rebuilt masks from the sampled-texture and vertex paths 2026-07-30 08:01:07 -04:00
swung0x48 2d4d6e9cfb [Perf] (DirectVulkan): skip pending-clear probes through a lock-free empty check 2026-07-30 07:02:03 -04:00
swung0x48 76b8957b99 [Perf] (DirectVulkan): memoize sampled-texture resources across draws 2026-07-30 07:02:02 -04:00
swung0x48 ec685b9fa7 [Perf] (DirectVulkan): memoize resolved vertex-input state on the VAO and dedupe vertex/index binds 2026-07-30 07:02:01 -04:00
swung0x48 a12068df52 [Perf] (DirectVulkan): reuse pipelines across per-chunk buffers and skip redundant pipeline binds 2026-07-30 05:01:46 -04:00
swung0x48 0b344792cc [Fix] (DirectVulkan): stop fence-waiting out-of-band texture uploads; reclaim transients asynchronously 2026-07-30 05:01:46 -04:00
swung0x48 9fa32bdad0 [Fix] (DirectVulkan): declare only the used colour attachment span per subpass
- every render pass declared colorAttachmentCount=8 (the full GL draw-buffer
  slot span) with trailing VK_ATTACHMENT_UNUSED references, and Adreno
  configures its per-pixel render-backend/export path from the DECLARED
  count - so every fragment of every pass paid an 8-render-target export
  cost; this was the bulk of the 1.5x per-pixel gap against
  MobileGlues+ANGLE on the same Qualcomm driver (their subpasses declare
  exactly the used span)
- measured on Adreno 650 / MC 26.2 / 1440x3044: total GPU frame time
  11.9 -> 7.5 ms (-37%, now below ANGLE's 7.87 ms), the single-quad
  swapchain blit pass alone 1.26 -> 0.40 ms, steady in-world FPS 82.8 -> 123
  under the standard cooled-start protocol, matching the
  MobileGlues+ANGLE+system-Vulkan benchmark of 123.8
- trailing UNUSED references are popped before the subpass is built (the
  entry's colorAttachmentCount and every pipeline's colour-blend span follow
  it); interior GL_NONE holes keep their slots so fragment-output locations
  still line up
- the pipeline-side fragmentOutputMask check downgrades from assert to a
  debug log: an output at a location past the trimmed span is discarded,
  which is GL's defined behaviour for a draw buffer set to GL_NONE
2026-07-30 02:45:39 -04:00
swung0x48 a4980f2b56 [Fix] (DirectVulkan): skip redundant per-draw dynamic-state commands
- viewport, scissor, blend constants, depth bias, line width and the six
  stencil parameters were re-emitted unconditionally for EVERY draw (~1500
  vkCmdSet* per frame in MC 26.2, where ANGLE emits a handful), costing CPU
  record time and GPU command-processor work for values that almost never
  change between draws
- a recording-scoped shadow now drops any vkCmdSet* whose values match what
  the command buffer already holds; valid because every PipelineFactory
  pipeline declares the same eight dynamic states, so set values persist
  across those binds
- the shadow resets at every command-buffer (re)begin (dynamic state does
  not survive the boundary) and after binding the blit or depth-mipmap
  pipelines, whose narrower dynamic sets make the untouched states undefined
  and whose raw viewport/scissor writes bypass the shadow
2026-07-30 02:45:07 -04:00
swung0x48 8ca20e28ca [Fix] (DirectVulkan): size texture backings by their defined mip level count
- every non-MSAA texture was allocated with a full mip chain regardless of
  how many levels the GL texture actually defines, so MC's 3044x1440 main
  colour and depth render targets each carried 12 levels where ANGLE
  allocates one; a level-0-only texture now gets a single-level backing and
  upgrades to the full chain exactly once when a second level is first
  defined, through the existing preserve-copy recreation path
- saves a third of the memory of every mip-less texture and keeps
  single-level render targets off the multi-mip image layout entirely, which
  also removes the surface the Adreno 650 implicit-LOD overread workaround
  (ForceExplicitLod0SamplePass) exists to defend
- measured perf-neutral on Adreno 650 / MC 26.2 (the driver keeps full UBWC
  on multi-mip render targets), so this is a memory/robustness fix, not a
  speed one
2026-07-30 02:43:32 -04:00
swung0x48 c353a2055f [Feat] (DirectVulkan): pre-pass command stream for reorderable out-of-pass work
- a draw whose sampled texture needs out-of-pass work (deferred clear
  materialization or a sampled-layout transition) used to end the active
  render pass - a full-target store+reload on a tiler - even when the only
  ordering the work needs is 'before this draw'; MC 26.2 clears an overlay
  texture every frame and samples it mid-pass, splitting the main scene pass
  once per frame for nothing
- every frame slot now carries a second primary command buffer, submitted
  strictly AHEAD of the frame command buffer in the same vkQueueSubmit; when
  the open recording has not referenced the image yet (tracked via a
  recording-generation stamp on the texture resource, advanced on every
  frame-command-buffer begin and stamped at every recorded reference:
  attachments at BeginRenderPass/attachment-write, sampled reads per draw,
  layout transitions), the clear/transition is recorded there and the active
  pass stays open - ANGLE's outside-render-pass command stream, restricted
  to the provably reorderable case
- mid-frame flushes and readback submits close and carry the pre stream with
  the frame buffer (it must never be submitted later than the recording it
  was paired with), retiring both under the same submit index; dropped
  recordings (present suspension, swapchain recreation) abandon it
- MaterializePendingClearForTexture's no-active-render-pass assert now
  applies only to the frame command buffer, since the pre stream records
  while a pass is open on the frame buffer by design
2026-07-30 02:43:13 -04:00
swung0x48 421c20984e [Fix] (DirectVulkan): stop loading and carrying dead default-framebuffer content
- EGL swap semantics make the presented colour buffer's content undefined at
  its next acquire (EGL_BUFFER_DESTROYED, the implementation default) and
  every ancillary depth/stencil buffer's content undefined after ANY swap,
  yet the default-FBO render pass reloaded both with LOAD_OP_LOAD every
  frame; SwapchainObject now tracks per-image content validity (defined when
  a pass stores into the attachment, invalidated at present) and the
  render-pass manager turns an undefined attachment's tile load into
  LOAD_OP_DONT_CARE with initialLayout=UNDEFINED, keyed into both hashes so
  the cached LOAD variants cannot be hit by mistake
- the default framebuffer's depth attachment is now attached ON DEMAND: a
  draw with depth test and stencil test both disabled (GL: a disabled test
  neither reads nor writes its buffer), and no pending depth/stencil clear,
  resolves to a depth-less pass flavour, dropping the D24S8 tile load AND
  store outright - MC 26.2 renders its GUI into its own FBO and only ever
  blits colour to the default framebuffer, so its swapchain pass carried a
  full-screen depth round-trip for nothing
- the flavour only escalates: an active depth-full pass absorbs depth-less
  draws unchanged, while a depth-using draw against a depth-less pass
  resolves to an incompatible entry and splits, its depth loading DONT_CARE
  (the content was undefined all along); the depth-less flavour is folded
  into ComputeHash and the per-draw fast-path memo so the two flavours can
  never alias
2026-07-30 02:40:47 -04:00
swung0x48 fc4cd980f2 [Fix] (DirectVulkan): bound image mutability so Adreno keeps UBWC compression
- Every storage-capable colour texture was created MUTABLE_FORMAT, and Adreno
  gives up bandwidth compression on an image that may be viewed as any format in
  its compatibility class. MC's main render target therefore ran uncompressed;
  in a fill-bound scene that is the whole frame budget. Measured on Adreno 650,
  MC 26.2, same scene and camera, device cooled to 38-40C before each run:
  65.3 -> 80.9 fps (+23.9%), GPU busy ~93% in both.
- VK_KHR_image_format_list (enabled when present) fixes it without giving up
  mutability: VkImageFormatListCreateInfo names the exact formats a view may
  use, so the driver can keep the image compressed. The set must be exhaustive
  or the result is undefined - for sampled views it is exactly what
  ResolveSampledImageViewFormat can return over the three numeric domains.
- glBindImageTexture may name any compatible format, which cannot be enumerated
  ahead of time, so a texture bound to an image unit gets no format list. That
  is what VK_IMAGE_USAGE_STORAGE_BIT becoming on-demand is for: it makes
  "unmarked" mean "will never receive an arbitrary-format storage view", which
  is what makes the list sound. Removing STORAGE is worth nothing on its own
  (65.4 fps, measured) - only the mutability bound pays.
- MarkStorageImageTexture runs over every collected image-unit texture before
  the probe loop in PrepareStorageImageTextures, because that loop stops at the
  first texture needing work and would leave the rest unmarked. The mark makes
  NeedsStorageImagePreparation report true, which is what ends the render pass,
  so the recreate lands outside it.
- storageUsageResolved separates "not upgraded yet" from "this format can never
  carry STORAGE", so a format whose optimalTilingFeatures lack STORAGE_IMAGE
  cannot ask for a recreate that will never happen. SyncTexture's cross-draw
  early-out also has to break on a pending upgrade or the recreate never runs.
- An upgrade recreates the image and carries its contents forward through
  PreserveTextureContentsOnRecreate, which submits its own command buffer and
  waits. Whatever the frame already recorded into the old image is still
  unsubmitted, so that copy would read pre-frame content and this frame's
  rendering into the texture would be lost - exactly the render-target-then-
  image-unit case. PrepareStorageImageTextures now flushes first; it takes the
  FrameData rather than a command buffer because the flush retires the current
  one, and drops the sampled-descriptor-set memo that described it.
2026-07-29 07:07:50 -04:00
swung0x48 992d16267c [Fix] (DirectVulkan): rewrite implicit-LOD fragment samples to explicit LOD 0 when every bound sampler is pinned to a single mip level - Adreno 650 (driver 512.502) reads outside a full-screen colour render target's allocation on its implicit-LOD sampling path and faults the GPU, which killed MC 26.2 on its own blit shader (texture(InSampler, texCoord)) between frames 344-421 on every run; this is the same driver defect the default-framebuffer blit shader already works around with textureLod, but an application's shader cannot be edited, so ForceExplicitLod0SamplePass converts OpImageSample*ImplicitLod to the explicit form at the SPIR-V level under a new CompileOptionBit that is only requested when the rewrite provably cannot move a texel (every sampler binding on a single-level view, no anisotropy, and either a LOD clamp that already pins lambda at 0 or min and mag filters that agree - an explicit LOD 0 always takes the magnification side of the min/mag decision); a single-level view now also clamps its sampler to mipmapMode NEAREST with maxLod min(maxLod, 0.25) rather than 0, since collapsing the clamp would make every fragment magnify and quietly retire the min filter; and the program's backend hash memo grows from one slot to four so a program resolved under two compile-flag sets in the same frame stops re-hashing every stage's SPIR-V once per draw 2026-07-29 03:26:57 -04:00
swung0x48 0ea9e6de5f [Fix] (DirectVulkan): follow surface resizes instead of rebuilding the swapchain on VK_SUBOPTIMAL_KHR - a per-frame surface-capabilities comparison (ANGLE's model) is now the only thing that schedules a rebuild, so a launcher-side resolution change reaches the swapchain and the compositor scales the smaller image up to the view, while a driver that merely reports the surface as suboptimal can no longer rebuild every frame (each rebuild destroys every pipeline, resets the render-pass manager and reallocates the default framebuffer, which showed as flicker, then corruption, then a crash); the comparison runs in SURFACE space against the extent the live swapchain was created from, since comparing against the swapchain's own quarter-turn-swapped extent reports a difference on every rotated frame 2026-07-28 21:06:57 -04:00
swung0x48 241ed377b4 [Fix] (macOS): harden Cocoa context setup and isolate embedded glslang 2026-07-28 11:54:50 -04:00
swung0x48 bf312a4b67 [Fix] (DirectVulkan): explicit-LOD blit sampling and present-path hardening - the default-framebuffer blit shader now samples with textureLod 0 (a blit reads exactly the selected level; Adreno 650's implicit-LOD path reads past a single-mip UBWC render target's allocation despite maxLod=0, page-faulting the GPU on MC 26.2's second startup frame once the neighbouring startup staging memory is returned - the invalidated context then failed the next Present submit with EDEADLK/DEVICE_LOST), TransitionToPresent appends the present barrier into the frame's open recording instead of silently dropping it whenever anything was recorded (frames without a default-FBO render pass presented images stuck in their acquired layout), VK_SUBOPTIMAL_KHR acquires are treated as the success they are (image acquired, semaphore signal armed - the early return skipped the fence reset and consumed-flag clear, and callers re-acquired on the same binary semaphore; rebuilds now defer to after the signal is consumed), and validation builds report through VK_EXT_debug_report when VK_EXT_debug_utils is absent instead of aborting instance creation 2026-07-28 06:00:20 -04:00
swung0x48 56b31a9587 [Fix] (FastSTL): bump submodule for the erase(iterator) double-advance fix and add erase-while-iterating regression tests - the old semantics skipped one live element per erase and ran past end() when erasing the highest occupied bucket, sending the new mass pipeline-cache eviction sweeps off the bucket array (device crash on first eviction during world load: garbage handles fed to vkDestroyPipeline) 2026-07-27 23:50:25 -04:00
swung0x48 8a0a8a0274 [Fix] (DirectVulkan): harden the leak-fix round after adversarial review - pipeline memo now drops at every command-buffer boundary (a flush-loop-memoized pipeline could age out and be destroyed while its submission was in flight), mid-frame drains no longer rewind the arena or advance the cache-aging clocks in presenting apps (gated to every 8th drain since the last Present, so readback/fence-heavy frames neither churn conversions nor shrink the 1024-boundary retire window), render-pass eviction notifies the pipeline cache once per sweep batch instead of once per dying pass, descriptor pools use FREE_DESCRIPTOR_SET_BIT so a destroyed layout's cached sets are freed back and credited instead of abandoning pool slots (the live-layout age sweep that could orphan slots is removed - layout destruction is the sole purge path), and renderbuffer respecify parks the old backing for aged destruction instead of destroying it while possibly in flight 2026-07-27 22:51:26 -04:00
swung0x48 d076c29146 [Fix] (DirectVulkan): bound the vertex-input and sampler caches and sweep undeleted GL syncs - both caches age out entries idle >1024 frame boundaries (animated LOD bias no longer mints a VkSampler per float value, buffer/VAO churn no longer grows the vertex-input map for the whole session), and library teardown drains the live-sync registry exactly as glDeleteSync would since GL requires syncs to die with their context 2026-07-27 22:16:16 -04:00
swung0x48 930a607bdf [Fix] (DirectVulkan): make texture/renderbuffer GC reach every dead resource - name-deleted textures register via weak_from_this so first-sync-after-delete can no longer orphan a TextureResource, an orphan sweep makes GC authoritative over the resource map, dead-texture pruning moves to a frame-boundary gate (64 frames) so churn through clears/readbacks reclaims without draws, and dead renderbuffers age past frames-in-flight before their VkImage/view is destroyed instead of leaking until shutdown (or being freed while in flight) 2026-07-27 22:16:15 -04:00
swung0x48 34685b4bb0 [Fix] (DirectVulkan): age-based eviction for the content-addressed cache family - ProgramFactory entries (shader modules/layouts), PipelineFactory graphics pipelines, compute pipelines and per-layout descriptor-set tracking now retire after ~1024 idle frame boundaries (render-pass-manager sweep precedent), render-pass eviction purges pipelines hashed on the dying handle (closes a handle-recycling stale-pipeline hazard), and the program reflection cache is lifetime-id-keyed and cleared at EGL teardown - shader/program churn no longer grows Vulkan objects without bound 2026-07-27 22:05:25 -04:00
swung0x48 c540fb88ee [Fix] (DirectVulkan): drain frame transients on present-less paths - readback waits, suspended presentation, blocking sync waits and flush completion polls now run Present's per-frame drains (deferred buffer/texture releases, transient arena rewind, descriptor cursors, retired command buffers, conversion caches) whenever every submission is provably complete, so offscreen/minimized workloads stay bounded; never blocks, frames-in-flight overlap untouched 2026-07-27 21:45:21 -04:00
swung0x48 6ae3245a0d [Test] (CTS): raise the no-output abort threshold - consecutive instant-crash cases are real progress once device liveness is confirmed 2026-07-26 19:23:05 -04:00
swung0x48 7e048fc2bf [Fix] (DirectVulkan): map RGB10_A2(UI) to A2B10G10R10 - GL 2_10_10_10_REV puts R in bits 0-9 so the A2R10G10B10 mapping silently swapped R/B on upload; also decode both 1010102 variants in readback 2026-07-26 19:13:46 -04:00
swung0x48 83cdfd6bdd [Fix] (DirectVulkan): GetTexImage reads all 3D slices/array layers with PACK_IMAGE_HEIGHT/SKIP_IMAGES semantics, and sRGB readback returns raw sRGB-encoded bytes instead of linearizing 2026-07-26 18:30:43 -04:00
swung0x48 1c76f886cf [Fix] (DirectVulkan): back legacy low-bit formats (RGB565/RGB5A1/RGBA4/R3G3B2/RGB4/RGBA2/RGB10/12) with their UNorm8/16 canonical shadow layouts and add capability fallbacks - they mapped to VK_FORMAT_UNDEFINED and crashed or wedged the GPU on upload; also admit 2DMSArray/CubeMap/3D color attachment targets in the render pass 2026-07-26 18:30:42 -04:00
swung0x48 a2e109beff [Fix] (DirectVulkan): general (format,type) readback conversion - hoist the CTS-verified StoreWideRowsToClient into shared ReadbackImpl and decode any color VkFormat to wide RGBA rows; readback previously supported only RGB/BGR/RGBA/BGRA x UNSIGNED_BYTE/FLOAT and silently returned zeros for everything else 2026-07-26 18:30:41 -04:00
swung0x48 63f0756644 [Fix] (DirectVulkan): support UBO instance arrays as arrayed descriptors - uniform Block{...}b[N] reflected as one binding with descriptorCount=N, per-element GL block mapping, per-element buffer infos and dynamic offsets; non-UBO descriptor arrays now fail program creation cleanly instead of continuing corrupt 2026-07-26 18:30:41 -04:00
swung0x48 450215d12c [Fix] (DirectVulkan): implement color renderbuffer attachments - render pass/pipeline/blit/copy/readback/clear paths treated color renderbuffers as absent (writes masked to VK_ATTACHMENT_UNUSED, glClear dropped, readback zeros) 2026-07-26 18:30:40 -04:00
swung0x48 3a9e520170 [Test] (CTS): isolate the DirectVulkan renderbuffer-FBO readback defect so the rest of KHR-GL33 can be measured 2026-07-26 18:30:39 -04:00
swung0x48 d2996ba1cf [Test] (CTS): run VK-GL-CTS KHR-GL33 against MobileGL on Android via a standalone glcts binary 2026-07-26 18:30:39 -04:00
82 changed files with 10824 additions and 543 deletions
+2
View File
@@ -25,3 +25,5 @@ MobileGL/MG*/cmake-build*
/android-plugin/app/src/trace/jniLibs /android-plugin/app/src/trace/jniLibs
/android-plugin/local.properties /android-plugin/local.properties
tools/trace_replay/work/ tools/trace_replay/work/
__pycache__/
*.py[cod]
+14
View File
@@ -455,8 +455,21 @@ if (ANDROID)
endif() endif()
if (APPLE AND NOT MOBILEGL_IOS) if (APPLE AND NOT MOBILEGL_IOS)
# MobileGL statically embeds glslang, SPIRV-Tools, and SPIRV-Cross. When
# this dylib is injected with DYLD_INSERT_LIBRARIES, exporting those C++
# symbols interposes incompatible copies embedded by host libraries such
# as shaderc. Keep only the public GL/EGL/CGL loader surface globally
# visible; GetProcAddress can still return pointers to hidden internals.
set(MOBILEGL_MACOS_EXPORTED_SYMBOLS
"${CMAKE_CURRENT_SOURCE_DIR}/MobileGL/MG_Impl/DyldInterpose/ExportedSymbols.txt")
target_link_options(${CMAKE_PROJECT_NAME} PRIVATE
"LINKER:-exported_symbols_list,${MOBILEGL_MACOS_EXPORTED_SYMBOLS}")
set_property(TARGET ${CMAKE_PROJECT_NAME} APPEND PROPERTY
LINK_DEPENDS "${MOBILEGL_MACOS_EXPORTED_SYMBOLS}")
target_link_libraries(${CMAKE_PROJECT_NAME} PUBLIC target_link_libraries(${CMAKE_PROJECT_NAME} PUBLIC
"-framework Cocoa" "-framework Cocoa"
"-framework CoreVideo"
"-framework QuartzCore" "-framework QuartzCore"
"-framework Foundation" "-framework Foundation"
"-framework OpenGL" "-framework OpenGL"
@@ -464,6 +477,7 @@ if (APPLE AND NOT MOBILEGL_IOS)
if(TARGET ${CMAKE_PROJECT_NAME}_s) if(TARGET ${CMAKE_PROJECT_NAME}_s)
target_link_libraries(${CMAKE_PROJECT_NAME}_s PUBLIC target_link_libraries(${CMAKE_PROJECT_NAME}_s PUBLIC
"-framework Cocoa" "-framework Cocoa"
"-framework CoreVideo"
"-framework QuartzCore" "-framework QuartzCore"
"-framework Foundation" "-framework Foundation"
"-framework OpenGL" "-framework OpenGL"
+13 -4
View File
@@ -14,6 +14,7 @@
#include <MG_State/EGLState/Core.h> #include <MG_State/EGLState/Core.h>
#include <MG_Impl/GLImpl/Texture/ProxyTexture.h> #include <MG_Impl/GLImpl/Texture/ProxyTexture.h>
#include <MG_Impl/GLImpl/Framebuffer/GL_Framebuffer.h> #include <MG_Impl/GLImpl/Framebuffer/GL_Framebuffer.h>
#include <MG_Impl/GLImpl/Sync/GL_Sync.h>
#include <atomic> #include <atomic>
#include <mutex> #include <mutex>
@@ -37,6 +38,12 @@ namespace MobileGL {
MGLOG_I("MobileGL closing..."); MGLOG_I("MobileGL closing...");
} }
glslang::FinalizeProcess(); glslang::FinalizeProcess();
// GL syncs die with their contexts, and every context is gone by the
// time full teardown runs: drain the live-sync registry while the
// backend function table can still release the backend handles (and
// before a re-initialized library could pair them with the wrong
// backend's DeleteSync).
MG_Impl::GLImpl::DestroyAllSyncObjects();
MG_Backend::pActiveBackendObject.reset(); MG_Backend::pActiveBackendObject.reset();
MG_State::pGLContext.reset(); MG_State::pGLContext.reset();
MG_State::pEGLContext.reset(); MG_State::pEGLContext.reset();
@@ -100,9 +107,11 @@ namespace MobileGL {
// (EGL/WGL/CGL): initialization happens lazily on the first entry point // (EGL/WGL/CGL): initialization happens lazily on the first entry point
// via EnsureInitialized(), and full teardown happens deterministically // via EnsureInitialized(), and full teardown happens deterministically
// when the last EGL display is terminated with nothing current (EGLImpl // when the last EGL display is terminated with nothing current (EGLImpl
// calls Destroy()). There is intentionally no static constructor, no // calls Destroy()). There is intentionally no backend-initializing static
// static destructor, and no DllMain: the global singletons use // constructor, no static destructor, and no DllMain: the global singletons
// leak-at-exit storage (see GlobalObjects.cpp), so a process that exits // use leak-at-exit storage (see GlobalObjects.cpp), so a process that exits
// without eglTerminate simply leaks them to the OS instead of running // without eglTerminate simply leaks them to the OS instead of running
// backend destructors during static teardown. // backend destructors during static teardown. macOS has a lightweight
// dyld constructor that installs NSOpenGL dispatch hooks only; full backend
// initialization still enters here from the first hooked CGL context.
} // namespace MobileGL } // namespace MobileGL
+4 -3
View File
@@ -13,9 +13,10 @@ namespace MobileGL {
void Initialize(); void Initialize();
// Thread-safe, idempotent, and re-entrant wrapper around Initialize(). // Thread-safe, idempotent, and re-entrant wrapper around Initialize().
// Host layers (EGL/WGL/CGL entry points) call this lazily on first use so // Host layers (EGL/WGL/CGL entry points) call this lazily on first use so
// MobileGL's lifecycle never depends on ELF/DLL static constructors, and // full backend initialization never depends on ELF/DLL static constructors,
// so a fresh init can follow a full Destroy() (e.g. after the last // and so a fresh init can follow a full Destroy() (e.g. after the last
// eglTerminate). // eglTerminate). The macOS dyld bootstrap installs only lightweight
// NSOpenGL method hooks.
void EnsureInitialized(); void EnsureInitialized();
void Destroy(); void Destroy();
+7
View File
@@ -300,6 +300,13 @@ namespace MobileGL {
Float ViewportBoundsRangeMin = 0.0f; Float ViewportBoundsRangeMin = 0.0f;
Float ViewportBoundsRangeMax = 0.0f; Float ViewportBoundsRangeMax = 0.0f;
Int ViewportSubpixelBits = 0; Int ViewportSubpixelBits = 0;
// GL 4.x fragment-interpolation offset limits. These defaults are the
// core minimums and are replaced by live GLES/Vulkan device limits.
Float MinFragmentInterpolationOffset = -0.5f;
// For four fractional bits the greatest required legal offset is
// 0.5 - 2^-4 = 0.4375 (GL 4.6 table 23.70).
Float MaxFragmentInterpolationOffset = 0.4375f;
Int FragmentInterpolationOffsetBits = 4;
Bool SupportsWideLines = false; Bool SupportsWideLines = false;
SizeT MaxShaderStorageBlockSize = 128 * 1024 * 1024; SizeT MaxShaderStorageBlockSize = 128 * 1024 * 1024;
Uint32 SubgroupSize = 0; Uint32 SubgroupSize = 0;
@@ -20,6 +20,7 @@
#include <MG_Util/Texture/TextureFormatProcessor.h> #include <MG_Util/Texture/TextureFormatProcessor.h>
#include <Config.h> #include <Config.h>
#include <algorithm> #include <algorithm>
#include <cmath>
#include <format> #include <format>
namespace MobileGL::MG_Backend::DirectGLES { namespace MobileGL::MG_Backend::DirectGLES {
@@ -603,7 +604,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
.ExtraVendor = Nullopt, // Extra vendor .ExtraVendor = Nullopt, // Extra vendor
.RendererGLInfo = .RendererGLInfo =
{ {
.TargetGLVersion = {3, 3, 0}, // Target OpenGL Version .TargetGLVersion = {3, 3, 0}, // GL target version
.TargetGLSLVersion = {4, 6, 0}, // Target Shading Language Version .TargetGLSLVersion = {4, 6, 0}, // Target Shading Language Version
// Baseline advertisement (no timer queries / anisotropy yet); reconciled // Baseline advertisement (no timer queries / anisotropy yet); reconciled
// once the ES capabilities exist, see UpdateAdvertisedCapabilityExtensions. // once the ES capabilities exist, see UpdateAdvertisedCapabilityExtensions.
@@ -1034,6 +1035,24 @@ namespace MobileGL::MG_Backend::DirectGLES {
m_dynamicParameters.ViewportBoundsRangeMin = m_GLESCapabilities.ViewportBoundsRangeMin; m_dynamicParameters.ViewportBoundsRangeMin = m_GLESCapabilities.ViewportBoundsRangeMin;
m_dynamicParameters.ViewportBoundsRangeMax = m_GLESCapabilities.ViewportBoundsRangeMax; m_dynamicParameters.ViewportBoundsRangeMax = m_GLESCapabilities.ViewportBoundsRangeMax;
m_dynamicParameters.ViewportSubpixelBits = m_GLESCapabilities.ViewportSubpixelBits; m_dynamicParameters.ViewportSubpixelBits = m_GLESCapabilities.ViewportSubpixelBits;
m_dynamicParameters.MinFragmentInterpolationOffset =
std::isfinite(m_GLESCapabilities.MinFragmentInterpolationOffset) &&
m_GLESCapabilities.MinFragmentInterpolationOffset <= -0.5f
? m_GLESCapabilities.MinFragmentInterpolationOffset
: -0.5f;
m_dynamicParameters.MaxFragmentInterpolationOffset = 0.4375f;
m_dynamicParameters.FragmentInterpolationOffsetBits = 4;
if (m_GLESCapabilities.FragmentInterpolationOffsetBits >= 4 &&
std::isfinite(m_GLESCapabilities.MaxFragmentInterpolationOffset)) {
const Float requiredMaxOffset =
0.5f - std::ldexp(1.0f, -m_GLESCapabilities.FragmentInterpolationOffsetBits);
if (m_GLESCapabilities.MaxFragmentInterpolationOffset >= requiredMaxOffset) {
m_dynamicParameters.MaxFragmentInterpolationOffset =
m_GLESCapabilities.MaxFragmentInterpolationOffset;
m_dynamicParameters.FragmentInterpolationOffsetBits =
m_GLESCapabilities.FragmentInterpolationOffsetBits;
}
}
m_dynamicParameters.SupportsWideLines = m_dynamicParameters.SupportsWideLines =
m_GLESCapabilities.AliasedLineWidthRangeMax > 1.0f || m_GLESCapabilities.SmoothLineWidthRangeMax > 1.0f; m_GLESCapabilities.AliasedLineWidthRangeMax > 1.0f || m_GLESCapabilities.SmoothLineWidthRangeMax > 1.0f;
+2 -86
View File
@@ -3435,90 +3435,6 @@ namespace MobileGL::MG_Backend::DirectGLES {
return componentType != 0 ? static_cast<GLenum>(componentType) : GL_UNSIGNED_NORMALIZED; return componentType != 0 ? static_cast<GLenum>(componentType) : GL_UNSIGNED_NORMALIZED;
} }
// Repacks wide RGBA(_INTEGER) rows into the client's (format, type) layout, honoring the
// client-side PACK parameters and the bound pixel-pack buffer. `wide` holds
// `sliceHeight * sliceCount` rows of `width` texels (slice-major, tightly stacked),
// 4 components x GetReadbackComponentSize(wideType) bytes each.
// applyPackImageParams: GL_PACK_IMAGE_HEIGHT / GL_PACK_SKIP_IMAGES apply only to GetTexImage
// of 3D/array images; ReadPixels and 2D GetTexImage ignore them (GL 3.3 sections 4.3.1, 6.1.4).
// Per the GL addressing rules, slice k row j lands at
// SKIP_IMAGES*imageStride + SKIP_ROWS*rowStride + SKIP_PIXELS*pixelBytes
// + k*imageStride + j*rowStride, with imageStride = max(IMAGE_HEIGHT, sliceHeight)*rowStride.
static Bool StoreWideRowsToClient(const Uint8* wide, GLenum wideType, GLsizei width, GLsizei sliceHeight,
GLsizei sliceCount, const ReadbackChannelMapping& mapping, GLenum type,
void* pixels, Bool applyPackImageParams) {
const SizeT dstPixelBytes = GetReadbackDstPixelSize(mapping, type);
if (dstPixelBytes == 0) {
return false;
}
ReadbackImpl::PackedReadbackLayout packedLayout{};
const Bool isPackedType = ReadbackImpl::GetPackedReadbackLayout(type, packedLayout);
const SizeT dstComponentSize = GetReadbackComponentSize(type);
const auto& pixelPackBufferObject =
MG_State::pGLContext->GetBufferBindingSlot(BufferTarget::PixelPack).GetBoundObject();
// Destination layout is computed from the client-side PACK parameters; only the actual pixel
// rows are written so skip regions of the destination stay untouched.
const auto packParams = MG_State::pGLContext->GetPixelStoreParameters(false);
const SizeT rowPixels = static_cast<SizeT>(packParams.RowLength > 0 ? packParams.RowLength : width);
const SizeT dstRowStride = AlignPixelRow(rowPixels * dstPixelBytes, packParams.Alignment);
const SizeT imageRows =
applyPackImageParams && packParams.ImageHeight > 0
? static_cast<SizeT>(packParams.ImageHeight)
: static_cast<SizeT>(sliceHeight);
const SizeT dstImageStride = imageRows * dstRowStride;
const SizeT skipImages =
applyPackImageParams ? static_cast<SizeT>(std::max(packParams.SkipImages, 0)) : SizeT{0};
const SizeT dstSkipOffset = skipImages * dstImageStride +
static_cast<SizeT>(std::max(packParams.SkipRows, 0)) * dstRowStride +
static_cast<SizeT>(std::max(packParams.SkipPixels, 0)) * dstPixelBytes;
const SizeT dstRowBytes = static_cast<SizeT>(width) * dstPixelBytes;
const SizeT pboBaseOffset = reinterpret_cast<SizeT>(pixels); // with a PBO, `pixels` is an offset
if (pixelPackBufferObject) {
const SizeT requiredSize = pboBaseOffset + dstSkipOffset +
static_cast<SizeT>(sliceCount - 1) * dstImageStride +
static_cast<SizeT>(sliceHeight - 1) * dstRowStride + dstRowBytes;
if (requiredSize > pixelPackBufferObject->GetSize()) {
MGLOG_E("Readback conversion: pixel pack buffer is too small");
return true;
}
}
const SizeT srcComponentSize = GetReadbackComponentSize(wideType);
const SizeT srcPixelBytes = 4 * srcComponentSize;
Vector<Uint8> convertedRow(dstRowBytes);
for (GLsizei slice = 0; slice < sliceCount; ++slice) {
for (GLsizei row = 0; row < sliceHeight; ++row) {
const SizeT flatRow = static_cast<SizeT>(slice) * static_cast<SizeT>(sliceHeight) +
static_cast<SizeT>(row);
const Uint8* srcRow = wide + flatRow * static_cast<SizeT>(width) * srcPixelBytes;
ReadbackImpl::ConvertWideReadbackRow(srcRow, convertedRow.data(), static_cast<SizeT>(width), wideType,
mapping, type);
if (packParams.SwapBytes) {
const SizeT groupSize = isPackedType ? packedLayout.byteSize : dstComponentSize;
if (groupSize > 1) {
for (SizeT offset = 0; offset + groupSize <= dstRowBytes; offset += groupSize) {
std::reverse(convertedRow.data() + offset, convertedRow.data() + offset + groupSize);
}
}
}
const SizeT dstOffset = dstSkipOffset + static_cast<SizeT>(slice) * dstImageStride +
static_cast<SizeT>(row) * dstRowStride;
if (pixelPackBufferObject) {
pixelPackBufferObject->WritebackFromBackend({convertedRow.data(), dstRowBytes},
pboBaseOffset + dstOffset);
} else {
Memcpy(static_cast<Uint8*>(pixels) + dstOffset, convertedRow.data(), dstRowBytes);
}
}
}
return true;
}
// Reads the current READ framebuffer as wide RGBA(_INTEGER) and repacks the pixels into the client's // Reads the current READ framebuffer as wide RGBA(_INTEGER) and repacks the pixels into the client's
// (format, type) layout. Returns false when the combination is not convertible (the caller keeps its // (format, type) layout. Returns false when the combination is not convertible (the caller keeps its
@@ -3651,7 +3567,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
ExpandNarrowWideRead(wide, static_cast<SizeT>(width) * static_cast<SizeT>(height), readChannels, wideType); ExpandNarrowWideRead(wide, static_cast<SizeT>(width) * static_cast<SizeT>(height), readChannels, wideType);
} }
if (!StoreWideRowsToClient(wide.data(), wideType, width, height, /*sliceCount=*/1, mapping, type, pixels, if (!ReadbackImpl::StoreWideRowsToClient(wide.data(), wideType, width, height, /*sliceCount=*/1, mapping, type, pixels,
honorPackImageParams)) { honorPackImageParams)) {
return false; return false;
} }
@@ -3705,7 +3621,7 @@ namespace MobileGL::MG_Backend::DirectGLES {
return false; return false;
} }
const GLenum wideType = isInteger ? (isSigned ? GL_INT : GL_UNSIGNED_INT) : GL_FLOAT; const GLenum wideType = isInteger ? (isSigned ? GL_INT : GL_UNSIGNED_INT) : GL_FLOAT;
if (!StoreWideRowsToClient(wide.data(), wideType, width, sliceHeight, sliceCount, mapping, type, pixels, if (!ReadbackImpl::StoreWideRowsToClient(wide.data(), wideType, width, sliceHeight, sliceCount, mapping, type, pixels,
applyPackImageParams)) { applyPackImageParams)) {
return false; return false;
} }
+90
View File
@@ -764,5 +764,95 @@ namespace MobileGL::MG_Backend::DirectGLES {
} }
} }
} }
static SizeT AlignReadbackRow(SizeT rowBytes, Int alignment) {
const SizeT align = alignment > 0 ? static_cast<SizeT>(alignment) : 1;
return (rowBytes + align - 1) / align * align;
}
// Repacks wide RGBA(_INTEGER) rows into the client's (format, type) layout, honoring the
// client-side PACK parameters and the bound pixel-pack buffer. `wide` holds
// `sliceHeight * sliceCount` rows of `width` texels (slice-major, tightly stacked),
// 4 components x GetReadbackComponentSize(wideType) bytes each.
// applyPackImageParams: GL_PACK_IMAGE_HEIGHT / GL_PACK_SKIP_IMAGES apply only to GetTexImage
// of 3D/array images; ReadPixels and 2D GetTexImage ignore them (GL 3.3 sections 4.3.1, 6.1.4).
// Per the GL addressing rules, slice k row j lands at
// SKIP_IMAGES*imageStride + SKIP_ROWS*rowStride + SKIP_PIXELS*pixelBytes
// + k*imageStride + j*rowStride, with imageStride = max(IMAGE_HEIGHT, sliceHeight)*rowStride.
Bool StoreWideRowsToClient(const Uint8* wide, GLenum wideType, GLsizei width, GLsizei sliceHeight,
GLsizei sliceCount, const ReadbackChannelMapping& mapping, GLenum type,
void* pixels, Bool applyPackImageParams) {
const SizeT dstPixelBytes = GetReadbackDstPixelSize(mapping, type);
if (dstPixelBytes == 0) {
return false;
}
PackedReadbackLayout packedLayout{};
const Bool isPackedType = GetPackedReadbackLayout(type, packedLayout);
const SizeT dstComponentSize = GetReadbackComponentSize(type);
const auto& pixelPackBufferObject =
MG_State::pGLContext->GetBufferBindingSlot(BufferTarget::PixelPack).GetBoundObject();
// Destination layout is computed from the client-side PACK parameters; only the actual pixel
// rows are written so skip regions of the destination stay untouched.
const auto packParams = MG_State::pGLContext->GetPixelStoreParameters(false);
const SizeT rowPixels = static_cast<SizeT>(packParams.RowLength > 0 ? packParams.RowLength : width);
const SizeT dstRowStride = AlignReadbackRow(rowPixels * dstPixelBytes, packParams.Alignment);
const SizeT imageRows =
applyPackImageParams && packParams.ImageHeight > 0
? static_cast<SizeT>(packParams.ImageHeight)
: static_cast<SizeT>(sliceHeight);
const SizeT dstImageStride = imageRows * dstRowStride;
const SizeT skipImages =
applyPackImageParams ? static_cast<SizeT>(std::max(packParams.SkipImages, 0)) : SizeT{0};
const SizeT dstSkipOffset = skipImages * dstImageStride +
static_cast<SizeT>(std::max(packParams.SkipRows, 0)) * dstRowStride +
static_cast<SizeT>(std::max(packParams.SkipPixels, 0)) * dstPixelBytes;
const SizeT dstRowBytes = static_cast<SizeT>(width) * dstPixelBytes;
const SizeT pboBaseOffset = reinterpret_cast<SizeT>(pixels); // with a PBO, `pixels` is an offset
if (pixelPackBufferObject) {
const SizeT requiredSize = pboBaseOffset + dstSkipOffset +
static_cast<SizeT>(sliceCount - 1) * dstImageStride +
static_cast<SizeT>(sliceHeight - 1) * dstRowStride + dstRowBytes;
if (requiredSize > pixelPackBufferObject->GetSize()) {
MGLOG_E("Readback conversion: pixel pack buffer is too small");
return true;
}
}
const SizeT srcComponentSize = GetReadbackComponentSize(wideType);
const SizeT srcPixelBytes = 4 * srcComponentSize;
Vector<Uint8> convertedRow(dstRowBytes);
for (GLsizei slice = 0; slice < sliceCount; ++slice) {
for (GLsizei row = 0; row < sliceHeight; ++row) {
const SizeT flatRow = static_cast<SizeT>(slice) * static_cast<SizeT>(sliceHeight) +
static_cast<SizeT>(row);
const Uint8* srcRow = wide + flatRow * static_cast<SizeT>(width) * srcPixelBytes;
ConvertWideReadbackRow(srcRow, convertedRow.data(), static_cast<SizeT>(width), wideType,
mapping, type);
if (packParams.SwapBytes) {
const SizeT groupSize = isPackedType ? packedLayout.byteSize : dstComponentSize;
if (groupSize > 1) {
for (SizeT offset = 0; offset + groupSize <= dstRowBytes; offset += groupSize) {
std::reverse(convertedRow.data() + offset, convertedRow.data() + offset + groupSize);
}
}
}
const SizeT dstOffset = dstSkipOffset + static_cast<SizeT>(slice) * dstImageStride +
static_cast<SizeT>(row) * dstRowStride;
if (pixelPackBufferObject) {
pixelPackBufferObject->WritebackFromBackend({convertedRow.data(), dstRowBytes},
pboBaseOffset + dstOffset);
} else {
Memcpy(static_cast<Uint8*>(pixels) + dstOffset, convertedRow.data(), dstRowBytes);
}
}
}
return true;
}
} // namespace ReadbackImpl } // namespace ReadbackImpl
} // namespace MobileGL::MG_Backend::DirectGLES } // namespace MobileGL::MG_Backend::DirectGLES
+8
View File
@@ -88,6 +88,14 @@ namespace MobileGL::MG_Backend::DirectGLES {
// bytes, dst receives width * GetReadbackDstPixelSize(mapping, type) bytes. // bytes, dst receives width * GetReadbackDstPixelSize(mapping, type) bytes.
void ConvertWideReadbackRow(const Uint8* src, Uint8* dst, SizeT width, GLenum wideType, void ConvertWideReadbackRow(const Uint8* src, Uint8* dst, SizeT width, GLenum wideType,
const ReadbackChannelMapping& mapping, GLenum type); const ReadbackChannelMapping& mapping, GLenum type);
// Stores wide RGBA(_INTEGER) rows into the client pointer or the bound PACK pixel buffer,
// honoring the client-side PACK pixel-store parameters (row length, alignment, skips,
// swap-bytes, and - when applyPackImageParams - image height/skip images). Shared by the
// DirectGLES and DirectVulkan readback conversion paths.
Bool StoreWideRowsToClient(const Uint8* wide, GLenum wideType, GLsizei width, GLsizei sliceHeight,
GLsizei sliceCount, const ReadbackChannelMapping& mapping, GLenum type,
void* pixels, Bool applyPackImageParams);
} // namespace ReadbackImpl } // namespace ReadbackImpl
namespace PrgramImpl { namespace PrgramImpl {
@@ -18,6 +18,7 @@
#include "MG_Util/Texture/TextureFormatProcessor.h" #include "MG_Util/Texture/TextureFormatProcessor.h"
#include <Config.h> #include <Config.h>
#include <cmath>
#include <cstdlib> #include <cstdlib>
#include <cstring> #include <cstring>
@@ -140,6 +141,20 @@ namespace MobileGL::MG_Backend::DirectVulkan {
case TextureInternalFormat::RGB: case TextureInternalFormat::RGB:
case TextureInternalFormat::RGB8: case TextureInternalFormat::RGB8:
return TextureInternalFormat::RGBA8; return TextureInternalFormat::RGBA8;
// Legacy low-bit-depth formats with no (or rarely supported) native Vulkan
// encoding; a wider normalized fallback keeps at least the required precision.
case TextureInternalFormat::R3G3B2:
case TextureInternalFormat::RGB4:
case TextureInternalFormat::RGB5:
case TextureInternalFormat::RGBA2:
case TextureInternalFormat::RGBA4:
case TextureInternalFormat::RGB5A1:
return TextureInternalFormat::RGBA8;
case TextureInternalFormat::RGB10:
return TextureInternalFormat::RGB10A2;
case TextureInternalFormat::RGB12:
case TextureInternalFormat::RGBA12:
return TextureInternalFormat::RGBA16;
case TextureInternalFormat::SRGB8: case TextureInternalFormat::SRGB8:
return TextureInternalFormat::SRGB8Alpha8; return TextureInternalFormat::SRGB8Alpha8;
case TextureInternalFormat::RGB8Snorm: case TextureInternalFormat::RGB8Snorm:
@@ -455,6 +470,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// treat them as signaled/available with zero results from here on. // treat them as signaled/available with zero results from here on.
BumpRendererGeneration(); BumpRendererGeneration();
pVulkanRenderer.reset(); pVulkanRenderer.reset();
// The reflection cache is file-scope, not renderer-owned; without this the
// deleted programs' reflection strings survive full context teardown.
ClearProgramResourceCaches();
BackendObject::ReleaseEGLResources(); BackendObject::ReleaseEGLResources();
} }
@@ -464,6 +482,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// treat them as signaled/available with zero results from here on. // treat them as signaled/available with zero results from here on.
BumpRendererGeneration(); BumpRendererGeneration();
pVulkanRenderer.reset(); pVulkanRenderer.reset();
// The reflection cache is file-scope, not renderer-owned; without this the
// deleted programs' reflection strings survive full context teardown.
ClearProgramResourceCaches();
} }
const RendererInfo& BackendObject_DirectVulkan::GetRendererInfo() const { const RendererInfo& BackendObject_DirectVulkan::GetRendererInfo() const {
@@ -776,6 +797,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_dynamicParameters.ViewportBoundsRangeMin = m_vulkanCaps.ViewportBoundsRangeMin; m_dynamicParameters.ViewportBoundsRangeMin = m_vulkanCaps.ViewportBoundsRangeMin;
m_dynamicParameters.ViewportBoundsRangeMax = m_vulkanCaps.ViewportBoundsRangeMax; m_dynamicParameters.ViewportBoundsRangeMax = m_vulkanCaps.ViewportBoundsRangeMax;
m_dynamicParameters.ViewportSubpixelBits = m_vulkanCaps.ViewportSubpixelBits; m_dynamicParameters.ViewportSubpixelBits = m_vulkanCaps.ViewportSubpixelBits;
m_dynamicParameters.MinFragmentInterpolationOffset =
std::isfinite(m_vulkanCaps.MinFragmentInterpolationOffset) &&
m_vulkanCaps.MinFragmentInterpolationOffset <= -0.5f
? m_vulkanCaps.MinFragmentInterpolationOffset
: -0.5f;
m_dynamicParameters.MaxFragmentInterpolationOffset = 0.4375f;
m_dynamicParameters.FragmentInterpolationOffsetBits = 4;
if (m_vulkanCaps.FragmentInterpolationOffsetBits >= 4 &&
std::isfinite(m_vulkanCaps.MaxFragmentInterpolationOffset)) {
const Float requiredMaxOffset =
0.5f - std::ldexp(1.0f, -m_vulkanCaps.FragmentInterpolationOffsetBits);
if (m_vulkanCaps.MaxFragmentInterpolationOffset >= requiredMaxOffset) {
m_dynamicParameters.MaxFragmentInterpolationOffset = m_vulkanCaps.MaxFragmentInterpolationOffset;
m_dynamicParameters.FragmentInterpolationOffsetBits =
m_vulkanCaps.FragmentInterpolationOffsetBits;
}
}
m_dynamicParameters.SupportsWideLines = m_vulkanCaps.SupportsWideLines; m_dynamicParameters.SupportsWideLines = m_vulkanCaps.SupportsWideLines;
m_dynamicParameters.MaxShaderStorageBlockSize = m_dynamicParameters.MaxShaderStorageBlockSize =
std::min(m_vulkanCaps.MaxShaderStorageBlockSize, kMaxAdvertisedShaderStorageBlockSize); std::min(m_vulkanCaps.MaxShaderStorageBlockSize, kMaxAdvertisedShaderStorageBlockSize);
@@ -61,6 +61,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}; };
struct ProgramResourceCache { struct ProgramResourceCache {
// Lifetime id of the program the cached reflection belongs to. GL names are
// recycled (IndexGenerator hands freed indices straight back), and a
// recreated program's backendStateVersion restarts at the same small values,
// so the version alone can collide; the never-reused lifetime id makes the
// slot's ownership unambiguous.
Uint64 programLifetimeId = 0;
Uint32 backendStateVersion = 0; Uint32 backendStateVersion = 0;
Vector<StorageBlockResource> storageBlocks; Vector<StorageBlockResource> storageBlocks;
Vector<BufferVariableResource> bufferVariables; Vector<BufferVariableResource> bufferVariables;
@@ -82,6 +88,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 baseInstance = 0; Uint32 baseInstance = 0;
}; };
// Keyed by GL program name so the freed-name reuse in IndexGenerator bounds the
// map at the peak-simultaneous-program high-water mark; each slot's ownership is
// checked against the program's lifetime id before it is served (see
// GetProgramResourceCache). Cleared wholesale at EGL teardown via
// ClearProgramResourceCaches.
UnorderedMap<GLuint, ProgramResourceCache> g_programResourceCaches; UnorderedMap<GLuint, ProgramResourceCache> g_programResourceCaches;
void ClearReadPixelsOutput(GLsizei width, GLsizei height, GLenum format, GLenum type, void* pixels) { void ClearReadPixelsOutput(GLsizei width, GLsizei height, GLenum format, GLenum type, void* pixels) {
@@ -142,13 +153,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ProgramResourceCache& GetProgramResourceCache(const MG_State::GLState::ProgramObject& program) { ProgramResourceCache& GetProgramResourceCache(const MG_State::GLState::ProgramObject& program) {
auto& cache = g_programResourceCaches[program.GetExternalIndex()]; auto& cache = g_programResourceCaches[program.GetExternalIndex()];
const Uint64 programLifetimeId = program.GetLifetimeId();
const Uint32 backendStateVersion = program.GetBackendStateVersion(); const Uint32 backendStateVersion = program.GetBackendStateVersion();
if (cache.backendStateVersion == backendStateVersion && // The lifetime id must match too: a new program that reuses a deleted
// program's name and happens to land on the same backendStateVersion (both
// count from zero) would otherwise be served the dead program's reflection.
if (cache.programLifetimeId == programLifetimeId &&
cache.backendStateVersion == backendStateVersion &&
(!cache.storageBlocks.empty() || !cache.bufferVariables.empty())) { (!cache.storageBlocks.empty() || !cache.bufferVariables.empty())) {
return cache; return cache;
} }
cache = {}; cache = {};
cache.programLifetimeId = programLifetimeId;
cache.backendStateVersion = backendStateVersion; cache.backendStateVersion = backendStateVersion;
Vector<SpvReflectShaderModule> modules; Vector<SpvReflectShaderModule> modules;
@@ -366,6 +383,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} // namespace } // namespace
void ClearProgramResourceCaches() {
// Called from EGL teardown while the backend's m_eglStateMutex is held; GL
// calls are serialized in this codebase (contexts migrate threads but never
// run concurrently), so no other thread can be inside the unsynchronized map.
// Live programs in another context self-heal: their entry rebuilds from the
// retained generated SPIR-V on the next resource query.
g_programResourceCaches.clear();
}
GLuint GetShaderStorageBlockIndex(const MG_State::GLState::ProgramObject& program, const String& name) { GLuint GetShaderStorageBlockIndex(const MG_State::GLState::ProgramObject& program, const String& name) {
auto& cache = GetProgramResourceCache(program); auto& cache = GetProgramResourceCache(program);
const auto it = std::find_if(cache.storageBlocks.begin(), cache.storageBlocks.end(), const auto it = std::find_if(cache.storageBlocks.begin(), cache.storageBlocks.end(),
@@ -23,6 +23,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 GetRendererGeneration(); Uint64 GetRendererGeneration();
void BumpRendererGeneration(); void BumpRendererGeneration();
// Drops every cached program-resource reflection entry (CPU-side strings/vectors
// only, no Vulkan handles). Called at EGL teardown next to the renderer reset;
// safe because GL calls are serialized in this codebase, and any still-live
// program rebuilds its entry from the retained generated SPIR-V on demand.
void ClearProgramResourceCaches();
void ClearBufferfi(GLenum buffer, GLint drawbuffer, GLfloat depth, GLint stencil); void ClearBufferfi(GLenum buffer, GLint drawbuffer, GLfloat depth, GLint stencil);
void ClearBufferfv(GLenum buffer, GLint drawbuffer, const GLfloat* value); void ClearBufferfv(GLenum buffer, GLint drawbuffer, const GLfloat* value);
void ClearBufferuiv(GLenum buffer, GLint drawbuffer, const GLuint* value); void ClearBufferuiv(GLenum buffer, GLint drawbuffer, const GLuint* value);
@@ -16,18 +16,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_device = device; m_device = device;
m_commandPool = commandPool; m_commandPool = commandPool;
Vector<VkCommandBuffer> commandBuffers(frameCount, VK_NULL_HANDLE); Vector<VkCommandBuffer> commandBuffers(frameCount * 2, VK_NULL_HANDLE);
VkCommandBufferAllocateInfo allocInfo{}; VkCommandBufferAllocateInfo allocInfo{};
allocInfo.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_ALLOCATE_INFO; allocInfo.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_ALLOCATE_INFO;
allocInfo.commandPool = commandPool; allocInfo.commandPool = commandPool;
allocInfo.level = VK_COMMAND_BUFFER_LEVEL_PRIMARY; allocInfo.level = VK_COMMAND_BUFFER_LEVEL_PRIMARY;
allocInfo.commandBufferCount = frameCount; allocInfo.commandBufferCount = frameCount * 2;
VkResult result = vkAllocateCommandBuffers(device, &allocInfo, commandBuffers.data()); VkResult result = vkAllocateCommandBuffers(device, &allocInfo, commandBuffers.data());
if (result != VK_SUCCESS) { if (result != VK_SUCCESS) {
return result; return result;
} }
for (Uint32 i = 0; i < frameCount; ++i) { for (Uint32 i = 0; i < frameCount; ++i) {
m_frames[i].commandBuffer = commandBuffers[i]; m_frames[i].commandBuffer = commandBuffers[i];
m_frames[i].preCommandBuffer = commandBuffers[frameCount + i];
} }
VkSemaphoreCreateInfo semaphoreInfo{VK_STRUCTURE_TYPE_SEMAPHORE_CREATE_INFO}; VkSemaphoreCreateInfo semaphoreInfo{VK_STRUCTURE_TYPE_SEMAPHORE_CREATE_INFO};
@@ -47,9 +48,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void FrameContext::Destroy(VkDevice device, VkCommandPool commandPool) { void FrameContext::Destroy(VkDevice device, VkCommandPool commandPool) {
const Uint32 frameCount = static_cast<Uint32>(m_frames.size()); const Uint32 frameCount = static_cast<Uint32>(m_frames.size());
Vector<VkCommandBuffer> commandBuffers(frameCount, VK_NULL_HANDLE); Vector<VkCommandBuffer> commandBuffers(frameCount * 2, VK_NULL_HANDLE);
for (Uint32 i = 0; i < frameCount; ++i) { for (Uint32 i = 0; i < frameCount; ++i) {
commandBuffers[i] = m_frames[i].commandBuffer; commandBuffers[i] = m_frames[i].commandBuffer;
commandBuffers[frameCount + i] = m_frames[i].preCommandBuffer;
} }
for (Uint32 i = 0; i < frameCount; ++i) { for (Uint32 i = 0; i < frameCount; ++i) {
@@ -60,7 +62,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
for (auto& frame : m_frames) { for (auto& frame : m_frames) {
FreeRetiredCommandBuffers(frame); FreeRetiredCommandBuffers(frame);
} }
vkFreeCommandBuffers(device, commandPool, frameCount, commandBuffers.data()); vkFreeCommandBuffers(device, commandPool, frameCount * 2, commandBuffers.data());
} }
m_frames.clear(); m_frames.clear();
currentFrameIndex = 0; currentFrameIndex = 0;
@@ -87,6 +89,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
currentFrameIndex = (currentFrameIndex + 1) % static_cast<Uint32>(m_frames.size()); currentFrameIndex = (currentFrameIndex + 1) % static_cast<Uint32>(m_frames.size());
GetCurrent().isCommandRecording = false; GetCurrent().isCommandRecording = false;
GetCurrent().hasCommandBufferRecorded = false; GetCurrent().hasCommandBufferRecorded = false;
GetCurrent().isPreCommandRecording = false;
GetCurrent().hasPreCommandBufferRecorded = false;
} }
VkCommandBuffer& FrameContext::BeginCommandRecording(VkCommandBufferUsageFlags flags, VkCommandBuffer& FrameContext::BeginCommandRecording(VkCommandBufferUsageFlags flags,
@@ -118,6 +122,41 @@ namespace MobileGL::MG_Backend::DirectVulkan {
frame.hasCommandBufferRecorded = true; frame.hasCommandBufferRecorded = true;
} }
VkCommandBuffer FrameContext::BeginPreCommandRecording() {
auto& frame = GetCurrent();
if (frame.isPreCommandRecording) {
return frame.preCommandBuffer;
}
MOBILEGL_ASSERT(!frame.hasPreCommandBufferRecorded,
"BeginPreCommandRecording: a recorded pre stream is still awaiting submission");
VK_VERIFY(vkResetCommandBuffer(frame.preCommandBuffer, 0), "BeginPreCommandRecording, vkResetCommandBuffer");
VkCommandBufferBeginInfo beginInfo{};
beginInfo.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_BEGIN_INFO;
VK_VERIFY(vkBeginCommandBuffer(frame.preCommandBuffer, &beginInfo),
"BeginPreCommandRecording, vkBeginCommandBuffer");
frame.isPreCommandRecording = true;
return frame.preCommandBuffer;
}
void FrameContext::EndPreCommandRecordingIfOpen() {
auto& frame = GetCurrent();
if (!frame.isPreCommandRecording) {
return;
}
VK_VERIFY(vkEndCommandBuffer(frame.preCommandBuffer), "EndPreCommandRecordingIfOpen, vkEndCommandBuffer");
frame.isPreCommandRecording = false;
frame.hasPreCommandBufferRecorded = true;
}
void FrameContext::AbandonPreCommandRecording() {
auto& frame = GetCurrent();
if (frame.isPreCommandRecording) {
VK_VERIFY(vkEndCommandBuffer(frame.preCommandBuffer), "AbandonPreCommandRecording, vkEndCommandBuffer");
}
frame.isPreCommandRecording = false;
frame.hasPreCommandBufferRecorded = false;
}
VkResult FrameContext::InitializeSwapchainSemaphores(VkDevice device, Uint32 swapchainImageCount) { VkResult FrameContext::InitializeSwapchainSemaphores(VkDevice device, Uint32 swapchainImageCount) {
DestroySwapchainSemaphores(device); DestroySwapchainSemaphores(device);
if (swapchainImageCount == 0) { if (swapchainImageCount == 0) {
@@ -150,12 +189,30 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool FrameContext::TransitionToPresent(VkImage image, VkImageLayout oldLayout, VkImageLayout presentLayout) { Bool FrameContext::TransitionToPresent(VkImage image, VkImageLayout oldLayout, VkImageLayout presentLayout) {
auto& frame = GetCurrent(); auto& frame = GetCurrent();
if (frame.hasCommandBufferRecorded || frame.isCommandRecording || oldLayout == presentLayout || if (oldLayout == presentLayout || oldLayout == VK_IMAGE_LAYOUT_SHARED_PRESENT_KHR) {
oldLayout == VK_IMAGE_LAYOUT_SHARED_PRESENT_KHR) {
return false; return false;
} }
auto& commandBuffer = BeginCommandRecording(); // The barrier belongs in the frame's own recording. Bailing out because
// something was already recorded (the previous behaviour) dropped the
// transition entirely for every frame that never ran a default-framebuffer
// render pass - the only other thing that carries the image to
// PRESENT_SRC_KHR, via that pass's finalLayout - so the swapchain image was
// handed to the WSI still in the layout it was acquired in.
// A closed-but-unsubmitted buffer can only come from a submit that already
// failed (SubmitPendingCommandBuffer leaves the flag set on error), and
// appending to it is illegal while reopening would reset the frame's own
// commands away. The device is gone on that path anyway - stay silent-safe
// rather than trade a lost device for a barrier into a closed buffer.
if (frame.hasCommandBufferRecorded) {
MGLOG_E("TransitionToPresent: command buffer already closed; skipping the present barrier");
return false;
}
// Reopening a recording here would vkResetCommandBuffer this frame's own
// commands away, so append to the open one and let the caller close it.
const Bool openedRecording = !frame.isCommandRecording;
VkCommandBuffer commandBuffer = openedRecording ? BeginCommandRecording() : frame.commandBuffer;
VkImageMemoryBarrier presentBarrier{}; VkImageMemoryBarrier presentBarrier{};
presentBarrier.sType = VK_STRUCTURE_TYPE_IMAGE_MEMORY_BARRIER; presentBarrier.sType = VK_STRUCTURE_TYPE_IMAGE_MEMORY_BARRIER;
@@ -174,7 +231,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
vkCmdPipelineBarrier(commandBuffer, VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT, VK_PIPELINE_STAGE_BOTTOM_OF_PIPE_BIT, 0, 0, vkCmdPipelineBarrier(commandBuffer, VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT, VK_PIPELINE_STAGE_BOTTOM_OF_PIPE_BIT, 0, 0,
nullptr, 0, nullptr, 1, &presentBarrier); nullptr, 0, nullptr, 1, &presentBarrier);
if (openedRecording) {
EndCommandRecording(); EndCommandRecording();
}
return true; return true;
} }
@@ -182,17 +241,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 swapchainImageIndex) const { Uint32 swapchainImageIndex) const {
const auto& frame = GetCurrent(); const auto& frame = GetCurrent();
MOBILEGL_ASSERT(!frame.isCommandRecording, "GetSubmitInfo called while command buffer recording is still active"); MOBILEGL_ASSERT(!frame.isCommandRecording, "GetSubmitInfo called while command buffer recording is still active");
MOBILEGL_ASSERT(!frame.isPreCommandRecording,
"GetSubmitInfo called while the pre-pass stream is still recording");
AssertValidSwapchainImageIndex(swapchainImageIndex); AssertValidSwapchainImageIndex(swapchainImageIndex);
SubmitInfoPacket packet{}; SubmitInfoPacket packet{};
packet.waitSemaphore = frame.imageAvailableSemaphore; packet.waitSemaphore = frame.imageAvailableSemaphore;
packet.signalSemaphore = m_swapchainImageRenderFinishedSemaphores[swapchainImageIndex]; packet.signalSemaphore = m_swapchainImageRenderFinishedSemaphores[swapchainImageIndex];
packet.commandBuffer = frame.commandBuffer;
Uint32 commandBufferCount = 0;
// The pre-pass stream executes strictly before the frame's commands.
if (frame.hasPreCommandBufferRecorded) {
packet.commandBuffers[commandBufferCount++] = frame.preCommandBuffer;
}
if (shouldSubmitCommandBuffer) {
packet.commandBuffers[commandBufferCount++] = frame.commandBuffer;
}
packet.submitInfo.waitSemaphoreCount = frame.imageAvailableSemaphoreConsumed ? 0U : 1U; packet.submitInfo.waitSemaphoreCount = frame.imageAvailableSemaphoreConsumed ? 0U : 1U;
packet.submitInfo.pWaitSemaphores = frame.imageAvailableSemaphoreConsumed ? nullptr : &packet.waitSemaphore; packet.submitInfo.pWaitSemaphores = frame.imageAvailableSemaphoreConsumed ? nullptr : &packet.waitSemaphore;
packet.submitInfo.pWaitDstStageMask = frame.imageAvailableSemaphoreConsumed ? nullptr : &packet.waitDstStageMask; packet.submitInfo.pWaitDstStageMask = frame.imageAvailableSemaphoreConsumed ? nullptr : &packet.waitDstStageMask;
packet.submitInfo.commandBufferCount = shouldSubmitCommandBuffer ? 1U : 0U; packet.submitInfo.commandBufferCount = commandBufferCount;
packet.submitInfo.pCommandBuffers = shouldSubmitCommandBuffer ? &packet.commandBuffer : nullptr; packet.submitInfo.pCommandBuffers = commandBufferCount > 0 ? packet.commandBuffers : nullptr;
packet.submitInfo.signalSemaphoreCount = 1; packet.submitInfo.signalSemaphoreCount = 1;
packet.submitInfo.pSignalSemaphores = &packet.signalSemaphore; packet.submitInfo.pSignalSemaphores = &packet.signalSemaphore;
return packet; return packet;
@@ -227,12 +296,21 @@ namespace MobileGL::MG_Backend::DirectVulkan {
result = vkAcquireNextImageKHR(device, swapchain, timeout, frame.imageAvailableSemaphore, acquireFence, result = vkAcquireNextImageKHR(device, swapchain, timeout, frame.imageAvailableSemaphore, acquireFence,
&outImageIndex); &outImageIndex);
if (result != VK_SUCCESS) { // VK_SUBOPTIMAL_KHR is a success code: an image *was* acquired and
// imageAvailableSemaphore *will* be signaled. Bailing out on it skipped both
// the consumed-flag reset (leaving a stale "already consumed", so the next
// submit never waited on the pending signal) and the fence reset (leaving
// the slot's fence signaled for the next submit to reuse). Only a genuine
// failure - VK_ERROR_OUT_OF_DATE_KHR and friends, where nothing is acquired
// and nothing is signaled - skips the bookkeeping.
if (result != VK_SUCCESS && result != VK_SUBOPTIMAL_KHR) {
return result; return result;
} }
frame.imageAvailableSemaphoreConsumed = false; frame.imageAvailableSemaphoreConsumed = false;
return vkResetFences(device, 1, &frame.imageInFlightFence); const VkResult resetResult = vkResetFences(device, 1, &frame.imageInFlightFence);
// Hand the acquire's own code back so the caller can schedule a rebuild.
return resetResult == VK_SUCCESS ? result : resetResult;
} }
Uint32 FrameContext::GetCurrentFrameIndex() const { Uint32 FrameContext::GetCurrentFrameIndex() const {
@@ -247,12 +325,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_recordingObserver = observer; m_recordingObserver = observer;
} }
VkResult FrameContext::RetireCurrentCommandBuffer() { VkResult FrameContext::RetireCurrentCommandBuffer(Bool retirePreCommandBuffer) {
MOBILEGL_ASSERT(m_device != VK_NULL_HANDLE && m_commandPool != VK_NULL_HANDLE, MOBILEGL_ASSERT(m_device != VK_NULL_HANDLE && m_commandPool != VK_NULL_HANDLE,
"RetireCurrentCommandBuffer requires an initialized FrameContext"); "RetireCurrentCommandBuffer requires an initialized FrameContext");
auto& frame = GetCurrent(); auto& frame = GetCurrent();
MOBILEGL_ASSERT(!frame.isCommandRecording, MOBILEGL_ASSERT(!frame.isCommandRecording,
"RetireCurrentCommandBuffer called while the command buffer is still recording"); "RetireCurrentCommandBuffer called while the command buffer is still recording");
MOBILEGL_ASSERT(!frame.isPreCommandRecording,
"RetireCurrentCommandBuffer called while the pre-pass stream is still recording");
VkCommandBufferAllocateInfo allocInfo{}; VkCommandBufferAllocateInfo allocInfo{};
allocInfo.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_ALLOCATE_INFO; allocInfo.sType = VK_STRUCTURE_TYPE_COMMAND_BUFFER_ALLOCATE_INFO;
@@ -260,11 +340,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
allocInfo.level = VK_COMMAND_BUFFER_LEVEL_PRIMARY; allocInfo.level = VK_COMMAND_BUFFER_LEVEL_PRIMARY;
allocInfo.commandBufferCount = 1; allocInfo.commandBufferCount = 1;
VkCommandBuffer replacement = VK_NULL_HANDLE; VkCommandBuffer replacement = VK_NULL_HANDLE;
const VkResult result = vkAllocateCommandBuffers(m_device, &allocInfo, &replacement); VkResult result = vkAllocateCommandBuffers(m_device, &allocInfo, &replacement);
if (result != VK_SUCCESS) { if (result != VK_SUCCESS) {
return result; return result;
} }
frame.retiredCommandBuffers.push_back(frame.commandBuffer); if (retirePreCommandBuffer) {
VkCommandBuffer preReplacement = VK_NULL_HANDLE;
result = vkAllocateCommandBuffers(m_device, &allocInfo, &preReplacement);
if (result != VK_SUCCESS) {
vkFreeCommandBuffers(m_device, m_commandPool, 1, &replacement);
return result;
}
frame.retiredCommandBuffers.push_back({frame.preCommandBuffer, frame.lastSubmitIndex});
frame.preCommandBuffer = preReplacement;
}
// lastSubmitIndex was just written by the renderer for the submission
// that carried this command buffer.
frame.retiredCommandBuffers.push_back({frame.commandBuffer, frame.lastSubmitIndex});
frame.commandBuffer = replacement; frame.commandBuffer = replacement;
return VK_SUCCESS; return VK_SUCCESS;
} }
@@ -274,12 +366,40 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return; return;
} }
if (m_device != VK_NULL_HANDLE && m_commandPool != VK_NULL_HANDLE) { if (m_device != VK_NULL_HANDLE && m_commandPool != VK_NULL_HANDLE) {
vkFreeCommandBuffers(m_device, m_commandPool, static_cast<Uint32>(frame.retiredCommandBuffers.size()), for (const auto& retired : frame.retiredCommandBuffers) {
frame.retiredCommandBuffers.data()); vkFreeCommandBuffers(m_device, m_commandPool, 1, &retired.commandBuffer);
}
} }
frame.retiredCommandBuffers.clear(); frame.retiredCommandBuffers.clear();
} }
void FrameContext::FreeRetiredCommandBuffersCompletedUpTo(Uint64 completedSubmitIndex) {
if (m_device == VK_NULL_HANDLE || m_commandPool == VK_NULL_HANDLE) {
return;
}
for (auto& frame : m_frames) {
// Retired buffers are appended in submit order, so the completed
// ones form a prefix.
SizeT completedCount = 0;
while (completedCount < frame.retiredCommandBuffers.size() &&
frame.retiredCommandBuffers[completedCount].submitIndex <= completedSubmitIndex) {
vkFreeCommandBuffers(m_device, m_commandPool, 1,
&frame.retiredCommandBuffers[completedCount].commandBuffer);
++completedCount;
}
if (completedCount > 0) {
frame.retiredCommandBuffers.erase(frame.retiredCommandBuffers.begin(),
frame.retiredCommandBuffers.begin() + completedCount);
}
}
}
void FrameContext::FreeAllRetiredCommandBuffers() {
for (auto& frame : m_frames) {
FreeRetiredCommandBuffers(frame);
}
}
void FrameContext::AssertValidFrameIndex(Uint32 frameIndex) const { void FrameContext::AssertValidFrameIndex(Uint32 frameIndex) const {
MOBILEGL_ASSERT(frameIndex < m_frames.size(), "FrameContext index out of range"); MOBILEGL_ASSERT(frameIndex < m_frames.size(), "FrameContext index out of range");
} }
@@ -29,7 +29,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipelineStageFlags waitDstStageMask = VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT; VkPipelineStageFlags waitDstStageMask = VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT;
VkSemaphore waitSemaphore = VK_NULL_HANDLE; VkSemaphore waitSemaphore = VK_NULL_HANDLE;
VkSemaphore signalSemaphore = VK_NULL_HANDLE; VkSemaphore signalSemaphore = VK_NULL_HANDLE;
VkCommandBuffer commandBuffer = VK_NULL_HANDLE; // [0] = pre-pass command buffer (when recorded), then the frame
// command buffer; submitInfo.pCommandBuffers points here.
VkCommandBuffer commandBuffers[2] = {VK_NULL_HANDLE, VK_NULL_HANDLE};
VkSubmitInfo submitInfo{VK_STRUCTURE_TYPE_SUBMIT_INFO}; VkSubmitInfo submitInfo{VK_STRUCTURE_TYPE_SUBMIT_INFO};
}; };
@@ -40,17 +42,35 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPresentInfoKHR presentInfo{VK_STRUCTURE_TYPE_PRESENT_INFO_KHR}; VkPresentInfoKHR presentInfo{VK_STRUCTURE_TYPE_PRESENT_INFO_KHR};
}; };
// A command buffer submitted mid-frame (FlushPendingCommands), tagged
// with the submit-tracker index it was submitted under so it can be
// freed as soon as that submission is observed complete - without
// waiting for the slot's fence to be waited again (present-less flush
// loops never wait it).
struct RetiredCommandBuffer {
VkCommandBuffer commandBuffer = VK_NULL_HANDLE;
Uint64 submitIndex = 0;
};
struct FrameData { struct FrameData {
VkCommandBuffer commandBuffer = VK_NULL_HANDLE; VkCommandBuffer commandBuffer = VK_NULL_HANDLE;
// Pre-pass work stream: out-of-pass commands (deferred clear
// materialization, sampled-layout transitions) for resources the
// frame's recording has not touched yet. Submitted immediately
// BEFORE commandBuffer in the same vkQueueSubmit, so recording
// into it never has to split the frame's active render pass.
VkCommandBuffer preCommandBuffer = VK_NULL_HANDLE;
VkSemaphore imageAvailableSemaphore = VK_NULL_HANDLE; VkSemaphore imageAvailableSemaphore = VK_NULL_HANDLE;
VkFence imageInFlightFence = VK_NULL_HANDLE; VkFence imageInFlightFence = VK_NULL_HANDLE;
Bool isCommandRecording = false; Bool isCommandRecording = false;
Bool hasCommandBufferRecorded = false; Bool hasCommandBufferRecorded = false;
Bool isPreCommandRecording = false;
Bool hasPreCommandBufferRecorded = false;
Bool imageAvailableSemaphoreConsumed = false; Bool imageAvailableSemaphoreConsumed = false;
// Command buffers submitted mid-frame (FlushPendingCommands) whose // Command buffers submitted mid-frame (FlushPendingCommands),
// execution is only known complete once this slot's fence has been // appended in submit order; freed once their submission is known
// waited again; freed at that point. // complete (fence wait or completion poll).
Vector<VkCommandBuffer> retiredCommandBuffers; Vector<RetiredCommandBuffer> retiredCommandBuffers;
// Submit-tracker index of this slot's most recent queue submission // Submit-tracker index of this slot's most recent queue submission
// (written by the renderer at submit time). // (written by the renderer at submit time).
Uint64 lastSubmitIndex = 0; Uint64 lastSubmitIndex = 0;
@@ -67,6 +87,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkCommandBuffer& BeginCommandRecording(VkCommandBufferUsageFlags flags = 0, VkCommandBuffer& BeginCommandRecording(VkCommandBufferUsageFlags flags = 0,
const VkCommandBufferInheritanceInfo* pInheritanceInfo = nullptr); const VkCommandBufferInheritanceInfo* pInheritanceInfo = nullptr);
void EndCommandRecording(); void EndCommandRecording();
// Lazily opens the pre-pass work stream (see FrameData::preCommandBuffer).
VkCommandBuffer BeginPreCommandRecording();
// Closes the pre stream if open, marking it for submission ahead of the
// frame command buffer. Safe to call when it never opened.
void EndPreCommandRecordingIfOpen();
// Drops an in-progress or recorded-but-unsubmitted pre stream (dropped
// frame recordings, swapchain recreation).
void AbandonPreCommandRecording();
VkResult InitializeSwapchainSemaphores(VkDevice device, Uint32 swapchainImageCount); VkResult InitializeSwapchainSemaphores(VkDevice device, Uint32 swapchainImageCount);
void DestroySwapchainSemaphores(VkDevice device); void DestroySwapchainSemaphores(VkDevice device);
Bool TransitionToPresent(VkImage image, VkImageLayout oldLayout, Bool TransitionToPresent(VkImage image, VkImageLayout oldLayout,
@@ -79,8 +107,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// Parks the current (already ended and submitted) command buffer on the // Parks the current (already ended and submitted) command buffer on the
// slot's retired list and installs a freshly allocated one, so recording // slot's retired list and installs a freshly allocated one, so recording
// can restart while the submitted buffer is still executing. Retired // can restart while the submitted buffer is still executing. Retired
// buffers are freed after the slot's fence is next waited. // buffers are freed after the slot's fence is next waited, or as soon
VkResult RetireCurrentCommandBuffer(); // as their submission is observed complete.
VkResult RetireCurrentCommandBuffer(Bool retirePreCommandBuffer = false);
// Frees every retired command buffer whose tagged submission index is
// known complete. Driven by the renderer's submit tracker on completion
// events (fence waits and non-blocking polls), so present-less flush
// loops reclaim their buffers without any extra wait.
void FreeRetiredCommandBuffersCompletedUpTo(Uint64 completedSubmitIndex);
// Frees every slot's retired command buffers. Only valid when the
// caller has proven every queue submission complete.
void FreeAllRetiredCommandBuffers();
Uint32 GetCurrentFrameIndex() const; Uint32 GetCurrentFrameIndex() const;
Uint32 GetFrameCount() const; Uint32 GetFrameCount() const;
@@ -8,6 +8,7 @@
#include "PipelineFactory.h" #include "PipelineFactory.h"
#include <algorithm>
namespace MobileGL::MG_Backend::DirectVulkan { namespace MobileGL::MG_Backend::DirectVulkan {
static const char* PrimitiveTopologyToString(VkPrimitiveTopology topology) { static const char* PrimitiveTopologyToString(VkPrimitiveTopology topology) {
@@ -243,23 +244,108 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const HashType hash = ComputeHash(payload); const HashType hash = ComputeHash(payload);
auto it = m_cache.find(hash); auto it = m_cache.find(hash);
if (it != m_cache.end()) { if (it != m_cache.end()) {
return it->second; it->second.lastUsedFrame = m_frameCounter;
return it->second.pipeline;
} }
VkPipeline pipeline = CreatePipeline(payload); VkPipeline pipeline = CreatePipeline(payload);
m_cache.emplace(hash, pipeline); m_cache.emplace(hash, PipelineCacheEntry{pipeline, payload.programHash, payload.renderPass,
m_frameCounter});
return pipeline; return pipeline;
} }
void PipelineFactory::DestroyAll() { void PipelineFactory::DestroyAll() {
for (auto& pair : m_cache) { for (auto& pair : m_cache) {
if (pair.second != VK_NULL_HANDLE) { if (pair.second.pipeline != VK_NULL_HANDLE) {
vkDestroyPipeline(m_device, pair.second, nullptr); vkDestroyPipeline(m_device, pair.second.pipeline, nullptr);
} }
} }
m_cache.clear(); m_cache.clear();
} }
Uint32 PipelineFactory::OnFrameBoundary() {
++m_frameCounter;
// Sweep cadence and retire age mirror VkRenderPassManager::OnPresent: an entry
// idle for more than kRetireAgeFrames frame boundaries cannot be referenced by
// any in-flight command buffer (frames-in-flight <= MOBILEGL_MAGMA_FRAMESINFLIGHT),
// so immediate vkDestroyPipeline is safe. The caller must drop its "last
// pipeline" memo when this returns non-zero: the memo can return a cached
// handle without touching this cache, so an evicted pipeline may still be
// memoized (present-less flush loops never reset the memo per frame).
constexpr Uint64 kSweepInterval = 256;
constexpr Uint64 kRetireAgeFrames = 1024;
if ((m_frameCounter % kSweepInterval) != 0) {
return 0;
}
Uint32 evicted = 0;
for (auto it = m_cache.begin(); it != m_cache.end();) {
if (m_frameCounter - it->second.lastUsedFrame > kRetireAgeFrames) {
if (it->second.pipeline != VK_NULL_HANDLE) {
vkDestroyPipeline(m_device, it->second.pipeline, nullptr);
}
it = m_cache.erase(it);
++evicted;
} else {
++it;
}
}
if (evicted > 0) {
MGLOG_D("PipelineFactory::OnFrameBoundary: evicted %u idle pipelines (%zu remain)", evicted,
m_cache.size());
}
return evicted;
}
Uint32 PipelineFactory::EvictByRenderPasses(const Vector<VkRenderPass>& renderPasses) {
if (renderPasses.empty() || m_cache.empty()) {
return 0;
}
// Sorted-batch membership test keeps a mass eviction (shader-pack switch,
// dimension exit) at one O(cache * log batch) scan instead of one full scan
// per dying pass.
Vector<VkRenderPass> sortedPasses = renderPasses;
std::sort(sortedPasses.begin(), sortedPasses.end());
Uint32 evicted = 0;
for (auto it = m_cache.begin(); it != m_cache.end();) {
if (std::binary_search(sortedPasses.begin(), sortedPasses.end(), it->second.renderPass)) {
if (it->second.pipeline != VK_NULL_HANDLE) {
vkDestroyPipeline(m_device, it->second.pipeline, nullptr);
}
it = m_cache.erase(it);
++evicted;
} else {
++it;
}
}
if (evicted > 0) {
MGLOG_D("PipelineFactory::EvictByRenderPasses: evicted %u pipelines for %zu destroyed render passes",
evicted, sortedPasses.size());
}
return evicted;
}
Uint32 PipelineFactory::EvictByProgramHash(HashType programHash) {
Uint32 evicted = 0;
for (auto it = m_cache.begin(); it != m_cache.end();) {
if (it->second.programHash == programHash) {
if (it->second.pipeline != VK_NULL_HANDLE) {
vkDestroyPipeline(m_device, it->second.pipeline, nullptr);
}
it = m_cache.erase(it);
++evicted;
} else {
++it;
}
}
if (evicted > 0) {
MGLOG_D("PipelineFactory::EvictByProgramHash: evicted %u pipelines for program hash 0x%llx",
evicted, static_cast<unsigned long long>(programHash));
}
return evicted;
}
VkPipeline PipelineFactory::CreatePipeline(const PipelineCreatePayload& payload) const { VkPipeline PipelineFactory::CreatePipeline(const PipelineCreatePayload& payload) const {
MOBILEGL_ASSERT(payload.stages != nullptr && !payload.stages->empty(), "PipelineFactory: stages are empty"); MOBILEGL_ASSERT(payload.stages != nullptr && !payload.stages->empty(), "PipelineFactory: stages are empty");
MOBILEGL_ASSERT(payload.vertexInputState != nullptr, "PipelineFactory: vertexInputState is null"); MOBILEGL_ASSERT(payload.vertexInputState != nullptr, "PipelineFactory: vertexInputState is null");
@@ -65,6 +65,26 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkPipeline GetOrCreatePipeline(const PipelineCreatePayload& payload); VkPipeline GetOrCreatePipeline(const PipelineCreatePayload& payload);
void DestroyAll(); void DestroyAll();
// Frame boundary hook: ages the pipeline cache and destroys long-unused entries
// (their command buffers retired many frames ago), mirroring
// VkRenderPassManager::OnPresent's sweep. Returns the number of pipelines
// destroyed so the caller can drop any memoized VkPipeline handle.
Uint32 OnFrameBoundary();
// Destroys every cached pipeline hashed on one of `renderPasses`. Only safe
// when the caller guarantees GPU idleness for them - the render-pass manager
// calls this (via the renderer) for passes its own >1024-boundary-idle sweep
// just evicted, and a pipeline hashed on those handles is only ever bound by
// draws that also hit the render-pass entries. Also closes the handle-recycling
// hazard: a recycled VkRenderPass value must never serve a stale pipeline.
// Batched: one cache scan regardless of how many passes died in the sweep.
// Returns the number destroyed (callers invalidate memos when non-zero).
Uint32 EvictByRenderPasses(const Vector<VkRenderPass>& renderPasses);
// Destroys every cached pipeline built from the program with content hash
// `programHash`. Called from the ProgramFactory eviction path, which proves the
// same >1024-boundary idleness (the program's pipelines are only bound by draws
// that stamp its factory entry). Returns the number destroyed.
Uint32 EvictByProgramHash(HashType programHash);
// Driver quirk: suppress depth writes on accumulation-blended pipelines. Multi-pass // Driver quirk: suppress depth writes on accumulation-blended pipelines. Multi-pass
// depth-equality rendering (a blended prepass writes depth that later passes re-test // depth-equality rendering (a blended prepass writes depth that later passes re-test
// with an equality-inclusive compare on the re-rasterized geometry) requires // with an equality-inclusive compare on the re-rasterized geometry) requires
@@ -87,12 +107,26 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static Bool ShouldSuppressDepthWrite(const PipelineCreatePayload& payload); static Bool ShouldSuppressDepthWrite(const PipelineCreatePayload& payload);
private: private:
struct PipelineCacheEntry {
VkPipeline pipeline = VK_NULL_HANDLE;
// The hashed inputs the eviction paths key on: programHash ties the entry to
// its ProgramFactory entry, renderPass records the exact handle the hash
// folded in (the hash is one-way, so targeted eviction needs them verbatim).
HashType programHash = 0;
VkRenderPass renderPass = VK_NULL_HANDLE;
// Frame-boundary counter value of the last GetOrCreatePipeline hit; drives
// cache eviction (see OnFrameBoundary).
Uint64 lastUsedFrame = 0;
};
VkPipeline CreatePipeline(const PipelineCreatePayload& payload) const; VkPipeline CreatePipeline(const PipelineCreatePayload& payload) const;
VkDevice m_device = VK_NULL_HANDLE; VkDevice m_device = VK_NULL_HANDLE;
const VulkanRendererConfig& m_config; const VulkanRendererConfig& m_config;
VkPipelineCache m_pipelineCache = VK_NULL_HANDLE; VkPipelineCache m_pipelineCache = VK_NULL_HANDLE;
UnorderedMap<HashType, VkPipeline> m_cache; UnorderedMap<HashType, PipelineCacheEntry> m_cache;
// Monotonic frame-boundary counter (bumped in OnFrameBoundary) for cache aging.
Uint64 m_frameCounter = 0;
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
static inline Bool s_suppressBlendedDepthWrite = false; static inline Bool s_suppressBlendedDepthWrite = false;
}; };
@@ -923,6 +923,206 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ProgramFactory::CompileOptionFlags m_transformFlags; ProgramFactory::CompileOptionFlags m_transformFlags;
}; };
// Adreno 650 (driver 512.502) faults the GPU on an implicit-LOD sample of a full-screen
// colour render target: the texture unit's derivative path reads outside the image's
// allocation even though the sampler clamps LOD to 0 and the mapping is 1:1. MobileGL's
// own default-framebuffer blit shader works around it with textureLod, but an
// application's shader (Minecraft's blit.fsh is `texture(InSampler, texCoord)`) cannot be
// edited - so rewrite the sample at the SPIR-V level instead.
//
// The rewrite is only requested for draws whose every sampler binding is clamped to one
// mip level, where explicit LOD 0 is exactly what the implicit form must already produce:
// lambda' = clamp(lambda + bias, minLod, maxLod) with minLod = maxLod = 0. Bias and MinLod
// operands are therefore dropped rather than translated.
class ForceExplicitLod0SamplePass final : public spvtools::opt::Pass {
public:
const char* name() const override { return "force-explicit-lod0-sample"; }
Status Process() override {
Bool isFragment = false;
for (auto& entryPoint : get_module()->entry_points()) {
if (entryPoint.opcode() != spv::Op::OpEntryPoint) continue;
if (static_cast<spv::ExecutionModel>(entryPoint.GetSingleWordInOperand(0)) ==
spv::ExecutionModel::Fragment) {
isFragment = true;
break;
}
}
if (!isFragment) return Status::SuccessWithoutChange;
// Plan first, mutate second. Materializing the LOD constant is itself a module
// change, so it must not happen unless at least one rewrite is going to follow -
// otherwise the pass would grow the binary while reporting SuccessWithoutChange.
Vector<RewritePlan> plans;
for (auto& function : *get_module()) {
for (auto& block : function) {
for (auto& inst : block) {
RewritePlan plan{};
if (PlanRewrite(&inst, plan)) plans.push_back(Move(plan));
}
}
}
if (plans.empty()) return Status::SuccessWithoutChange;
const Uint32 zeroId = GetFloatZeroId();
if (zeroId == 0) return Status::SuccessWithoutChange;
for (auto& plan : plans) {
plan.operands.push_back({SPV_OPERAND_TYPE_ID, {zeroId}});
for (auto& operand : plan.trailingOperands) {
plan.operands.push_back(operand);
}
plan.instruction->SetOpcode(plan.opcode);
plan.instruction->SetInOperands(Move(plan.operands));
}
// Opcodes and operand lists changed underneath every cached analysis.
context()->InvalidateAnalysesExceptFor(spvtools::opt::IRContext::kAnalysisNone);
return Status::SuccessWithChange;
}
private:
struct RewritePlan {
spvtools::opt::Instruction* instruction = nullptr;
spv::Op opcode = spv::Op::OpNop;
// Everything up to and including the Image Operands mask; the Lod id and the
// trailing operand values are appended once the constant exists.
Vector<spvtools::opt::Operand> operands;
Vector<spvtools::opt::Operand> trailingOperands;
};
// Image Operands bits that may accompany an implicit-LOD sample, in the canonical
// ascending order SPIR-V requires the operand values to appear in.
static constexpr Uint32 kBias = 0x1;
static constexpr Uint32 kLod = 0x2;
static constexpr Uint32 kGrad = 0x4;
static constexpr Uint32 kConstOffset = 0x8;
static constexpr Uint32 kOffset = 0x10;
static constexpr Uint32 kConstOffsets = 0x20;
static constexpr Uint32 kSample = 0x40;
static constexpr Uint32 kMinLod = 0x80;
static constexpr Uint32 kKnownMask = 0xFF;
Uint32 GetFloatZeroId() {
// Reuse a 32-bit float type already in the module; a shader that samples always has
// one, and looking it up avoids depending on type-creation API details.
Uint32 floatTypeId = 0;
for (auto& inst : get_module()->types_values()) {
if (inst.opcode() == spv::Op::OpTypeFloat && inst.NumInOperands() >= 1 &&
inst.GetSingleWordInOperand(0) == 32) {
floatTypeId = inst.result_id();
break;
}
}
if (floatTypeId == 0) return 0;
const auto* floatType = context()->get_type_mgr()->GetType(floatTypeId);
if (floatType == nullptr) return 0;
const auto zeroBits = std::bit_cast<Uint32>(0.0f);
const auto* zeroConst = context()->get_constant_mgr()->GetConstant(floatType, {zeroBits});
if (zeroConst == nullptr) return 0;
auto* zeroInst = context()->get_constant_mgr()->GetDefiningInstruction(zeroConst);
return zeroInst != nullptr ? zeroInst->result_id() : 0;
}
static Bool MapOpcode(spv::Op op, spv::Op& outOpcode, Uint32& outFixedOperandCount) {
switch (op) {
case spv::Op::OpImageSampleImplicitLod:
outOpcode = spv::Op::OpImageSampleExplicitLod;
outFixedOperandCount = 2; // sampled image, coordinate
return true;
case spv::Op::OpImageSampleProjImplicitLod:
outOpcode = spv::Op::OpImageSampleProjExplicitLod;
outFixedOperandCount = 2;
return true;
case spv::Op::OpImageSampleDrefImplicitLod:
outOpcode = spv::Op::OpImageSampleDrefExplicitLod;
outFixedOperandCount = 3; // sampled image, coordinate, Dref
return true;
case spv::Op::OpImageSampleProjDrefImplicitLod:
outOpcode = spv::Op::OpImageSampleProjDrefExplicitLod;
outFixedOperandCount = 3;
return true;
default:
return false;
}
}
static Bool PlanRewrite(spvtools::opt::Instruction* inst, RewritePlan& outPlan) {
spv::Op newOpcode = spv::Op::OpNop;
Uint32 fixedCount = 0;
if (!MapOpcode(inst->opcode(), newOpcode, fixedCount)) return false;
if (inst->NumInOperands() < fixedCount) return false;
Uint32 mask = 0;
Uint32 next = fixedCount;
if (inst->NumInOperands() > fixedCount) {
mask = inst->GetSingleWordInOperand(fixedCount);
next = fixedCount + 1;
}
// An operand this pass does not model would be silently reordered or dropped, and
// Grad cannot legally accompany an implicit-LOD sample: leave such an instruction be.
if ((mask & ~kKnownMask) != 0 || (mask & kGrad) != 0) return false;
Vector<spvtools::opt::Operand> fixedOperands;
fixedOperands.reserve(fixedCount + 1);
for (Uint32 i = 0; i < fixedCount; ++i) {
fixedOperands.push_back(inst->GetInOperand(i));
}
// Collect the surviving operand values in the same ascending-bit order they were
// encoded in, so the rebuilt list stays canonical.
Uint32 keptMask = kLod;
Vector<spvtools::opt::Operand> keptOperands;
static constexpr Uint32 kOrderedBits[] = {kBias, kLod, kGrad, kConstOffset,
kOffset, kConstOffsets, kSample, kMinLod};
for (const Uint32 bit : kOrderedBits) {
if ((mask & bit) == 0) continue;
if (next >= inst->NumInOperands()) return false;
const spvtools::opt::Operand value = inst->GetInOperand(next++);
// Bias and MinLod only shift a lambda that is already clamped to 0, and any
// original Lod is replaced by the constant the caller appends.
if (bit == kBias || bit == kMinLod || bit == kLod) continue;
keptMask |= bit;
keptOperands.push_back(value);
}
fixedOperands.push_back({SPV_OPERAND_TYPE_IMAGE, {keptMask}});
outPlan.instruction = inst;
outPlan.opcode = newOpcode;
outPlan.operands = Move(fixedOperands);
outPlan.trailingOperands = Move(keptOperands);
return true;
}
};
spvtools::Optimizer::PassToken CreateForceExplicitLod0SamplePass() {
return spvtools::Optimizer::PassToken(MakeUnique<ForceExplicitLod0SamplePass>());
}
Bool TransformSpirvForExplicitLod0Sampling(const Vector<Uint>& input, Vector<Uint>& output) {
if (input.empty()) {
output.clear();
return true;
}
spvtools::Optimizer optimizer(SPV_ENV_VULKAN_1_3);
spvtools::OptimizerOptions options;
// Matches the position-fix pass: this build of spirv-tools asserts rather than
// reporting, so validation stays off in the shipping path.
options.set_run_validator(false);
optimizer.SetMessageConsumer([](spv_message_level_t, const char*, const spv_position_t&,
const char* message) {
MGLOG_E("Vulkan: explicit-LOD0 pass: %s", message != nullptr ? message : "");
});
optimizer.RegisterPass(CreateForceExplicitLod0SamplePass());
const Bool success = optimizer.Run(input.data(), input.size(), &output, options);
if (!success) {
MGLOG_E("Vulkan: explicit-LOD0 sampling pass failed; keeping the original module");
output = input;
}
return success;
}
spvtools::Optimizer::PassToken CreateGlToVulkanPositionFixPass( spvtools::Optimizer::PassToken CreateGlToVulkanPositionFixPass(
ProgramFactory::CompileOptionFlags transformFlags) { ProgramFactory::CompileOptionFlags transformFlags) {
return spvtools::Optimizer::PassToken(MakeUnique<GlToVulkanPositionFixPass>(transformFlags)); return spvtools::Optimizer::PassToken(MakeUnique<GlToVulkanPositionFixPass>(transformFlags));
@@ -1094,9 +1294,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
for (auto* binding : bindings) { for (auto* binding : bindings) {
MOBILEGL_ASSERT(binding != nullptr, "ProgramFactory: null descriptor binding reflection record"); MOBILEGL_ASSERT(binding != nullptr, "ProgramFactory: null descriptor binding reflection record");
const auto kind = ReflectDescriptorTypeToBindingKind(binding->descriptor_type); const auto kind = ReflectDescriptorTypeToBindingKind(binding->descriptor_type);
MOBILEGL_ASSERT(binding->count == 1, // UBO instance arrays (uniform Block {...} b[N];) occupy one binding with
"ProgramFactory: descriptor arrays are unsupported (name='%s' count=%u)", // descriptorCount = N; other descriptor arrays stay unsupported and must
binding->name ? binding->name : "<null>", binding->count); // fail program creation cleanly rather than continue with corrupt state.
if (binding->count != 1 && kind != ProgramFactory::DescriptorBindingKind::UniformBufferDynamic) {
MGLOG_E("ProgramFactory: descriptor arrays are unsupported for this descriptor "
"kind (name='%s' count=%u type=%d)",
binding->name ? binding->name : "<null>", binding->count,
static_cast<Int>(binding->descriptor_type));
destroyReflectModules();
return false;
}
DescriptorKey key{}; DescriptorKey key{};
key.kind = kind; key.kind = kind;
@@ -1613,6 +1821,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
entry.storageBlockIndexByBinding.assign(m_maxBindings, -1); entry.storageBlockIndexByBinding.assign(m_maxBindings, -1);
entry.globalUboBinding = -1; entry.globalUboBinding = -1;
entry.dynamicBindings.clear(); entry.dynamicBindings.clear();
entry.bindingDescriptorCounts.assign(m_maxBindings, 1);
entry.arrayedUniformBlockIndicesByBinding.clear();
// Use SpvcSession (Reflection mode) to reflect all SPIR-V modules in a single pass per module // Use SpvcSession (Reflection mode) to reflect all SPIR-V modules in a single pass per module
for (const auto& module : spirv) { for (const auto& module : spirv) {
@@ -1628,6 +1838,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"ProgramFactory::ReflectLayout: failed to create reflection module (result=%d)", "ProgramFactory::ReflectLayout: failed to create reflection module (result=%d)",
static_cast<Int>(createReflectResult)); static_cast<Int>(createReflectResult));
// Descriptor counts per binding (UBO instance arrays reflect count > 1).
UnorderedMap<Uint32, Uint32> descriptorCountByBinding;
{
uint32_t countProbe = 0;
if (spvReflectEnumerateDescriptorBindings(&reflectModule, &countProbe, nullptr) ==
SPV_REFLECT_RESULT_SUCCESS &&
countProbe > 0) {
Vector<SpvReflectDescriptorBinding*> probeBindings(countProbe);
if (spvReflectEnumerateDescriptorBindings(&reflectModule, &countProbe,
probeBindings.data()) ==
SPV_REFLECT_RESULT_SUCCESS) {
for (const auto* probeBinding : probeBindings) {
if (probeBinding != nullptr) {
descriptorCountByBinding[probeBinding->binding] =
std::max<Uint32>(1, probeBinding->count);
}
}
}
}
}
// Reflect uniform buffers // Reflect uniform buffers
auto ubos = session.GetShaderInterface(SPVC_RESOURCE_TYPE_UNIFORM_BUFFER); auto ubos = session.GetShaderInterface(SPVC_RESOURCE_TYPE_UNIFORM_BUFFER);
for (const auto& ubo : ubos) { for (const auto& ubo : ubos) {
@@ -1653,6 +1884,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
continue; continue;
} }
const auto countIt = descriptorCountByBinding.find(binding);
const Uint32 descriptorCount =
countIt != descriptorCountByBinding.end() ? countIt->second : 1u;
if (descriptorCount <= 1) {
const Uint blockIndex = program.GetUniformBlockIndex(ubo.name.c_str()); const Uint blockIndex = program.GetUniformBlockIndex(ubo.name.c_str());
if (blockIndex == 0xFFFFFFFFu) { if (blockIndex == 0xFFFFFFFFu) {
MGLOG_D("ProgramFactory::ReflectLayout: skipping inactive UBO '%s' at binding %u", MGLOG_D("ProgramFactory::ReflectLayout: skipping inactive UBO '%s' at binding %u",
@@ -1673,6 +1909,56 @@ namespace MobileGL::MG_Backend::DirectVulkan {
"ProgramFactory::ReflectLayout: descriptor binding %u maps to conflicting UBO blocks (%d vs %u)", "ProgramFactory::ReflectLayout: descriptor binding %u maps to conflicting UBO blocks (%d vs %u)",
binding, entry.uniformBlockIndexByBinding[binding], blockIndex); binding, entry.uniformBlockIndexByBinding[binding], blockIndex);
entry.uniformBlockIndexByBinding[binding] = static_cast<Int>(blockIndex); entry.uniformBlockIndexByBinding[binding] = static_cast<Int>(blockIndex);
continue;
}
// UBO instance array: one binding, descriptorCount elements. GL exposes each
// element as its own active block named "Name[i]"; map every element to its
// GL block index so the descriptor write can gather per-element buffer ranges.
if (descriptorCount > m_maxBindings) {
MGLOG_E("ProgramFactory::ReflectLayout: UBO array '%s' count %u exceeds maxBindings=%u; "
"leaving binding %u unmapped",
ubo.name.c_str(), descriptorCount, m_maxBindings, binding);
continue;
}
Vector<Int> elementBlockIndices;
elementBlockIndices.reserve(descriptorCount);
for (Uint32 element = 0; element < descriptorCount; ++element) {
String elementName = ubo.name + "[" + std::to_string(element) + "]";
Uint elementBlockIndex = program.GetUniformBlockIndex(elementName.c_str());
if (elementBlockIndex == 0xFFFFFFFFu && element == 0) {
// Some frontends report the first element under the bare block name.
elementBlockIndex = program.GetUniformBlockIndex(ubo.name.c_str());
}
if (elementBlockIndex == 0xFFFFFFFFu) {
// Degrade rather than corrupt: reuse element 0's block if we have one,
// otherwise give up on the binding (same observable behavior as an
// inactive block: wrong values, but no crash).
MGLOG_E("ProgramFactory::ReflectLayout: UBO array '%s' element %u has no active "
"GL uniform block",
ubo.name.c_str(), element);
if (!elementBlockIndices.empty()) {
elementBlockIndex = static_cast<Uint>(elementBlockIndices.front());
} else {
break;
}
}
elementBlockIndices.push_back(static_cast<Int>(elementBlockIndex));
}
if (elementBlockIndices.size() != descriptorCount) {
MGLOG_E("ProgramFactory::ReflectLayout: skipping unresolved UBO array '%s' at binding %u",
ubo.name.c_str(), binding);
continue;
}
MOBILEGL_ASSERT(entry.bindingKinds[binding] == DescriptorBindingKind::None ||
entry.bindingKinds[binding] == DescriptorBindingKind::UniformBufferDynamic,
"ProgramFactory::ReflectLayout: descriptor binding %u has conflicting kinds for UBO '%s'",
binding, ubo.name.c_str());
entry.bindingKinds[binding] = DescriptorBindingKind::UniformBufferDynamic;
entry.bindingDescriptorCounts[binding] = static_cast<Uint16>(descriptorCount);
entry.uniformBlockIndexByBinding[binding] = elementBlockIndices[0];
entry.arrayedUniformBlockIndicesByBinding[binding] = Move(elementBlockIndices);
} }
// Reflect sampled images, storage images, samplerBuffer uniforms, and SSBOs. // Reflect sampled images, storage images, samplerBuffer uniforms, and SSBOs.
@@ -1819,7 +2105,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkDescriptorSetLayoutBinding layoutBinding{}; VkDescriptorSetLayoutBinding layoutBinding{};
layoutBinding.binding = binding; layoutBinding.binding = binding;
layoutBinding.descriptorCount = 1; layoutBinding.descriptorCount = entry.bindingDescriptorCounts[binding];
layoutBinding.stageFlags = VK_SHADER_STAGE_ALL; layoutBinding.stageFlags = VK_SHADER_STAGE_ALL;
layoutBinding.pImmutableSamplers = nullptr; layoutBinding.pImmutableSamplers = nullptr;
if (kind == DescriptorBindingKind::UniformBufferDynamic) { if (kind == DescriptorBindingKind::UniformBufferDynamic) {
@@ -1864,11 +2150,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
auto it = m_cache.find(hash); auto it = m_cache.find(hash);
if (it != m_cache.end()) { if (it != m_cache.end()) {
// Every draw/dispatch funnels through this lookup (the renderer memos only
// skip re-hashing, never the factory lookup), so an actively-used entry is
// stamped at least once per frame boundary and can never be aged out while
// any in-flight command buffer still references it.
it->second.lastUsedFrame = m_frameCounter;
return it->second; return it->second;
} }
auto& entry = m_cache[hash]; auto& entry = m_cache[hash];
entry.hash = hash; entry.hash = hash;
entry.lastUsedFrame = m_frameCounter;
auto& shaders = program.GetAttachedShaders(); auto& shaders = program.GetAttachedShaders();
auto& spirv = program.GetGeneratedSpirv(); auto& spirv = program.GetGeneratedSpirv();
Vector<Vector<Uint>> moduleSpirvs(spirv.size()); Vector<Vector<Uint>> moduleSpirvs(spirv.size());
@@ -1886,6 +2178,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
moduleSpirvs[i] = spv; moduleSpirvs[i] = spv;
} }
if ((flags & ProgramFactory::CompileOptionBit::ExplicitLod0Sampling) && shaders[i] &&
shaders[i]->GetShaderStage() == ShaderStage::Fragment) {
Vector<Uint> explicitLodSpirv;
if (TransformSpirvForExplicitLod0Sampling(moduleSpirvs[i], explicitLodSpirv)) {
moduleSpirvs[i] = Move(explicitLodSpirv);
}
}
// GL apps depend on cross-program position invariance for multi-pass equality // GL apps depend on cross-program position invariance for multi-pass equality
// depth tests (MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the // depth tests (MC 26.3's OIT re-draws the cloud geometry with GEQUAL against the
// depth its own first pass wrote); decorate Position outputs Invariant so // depth its own first pass wrote); decorate Position outputs Invariant so
@@ -1982,4 +2282,42 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return entry; return entry;
} }
void ProgramFactory::OnFrameBoundary() {
++m_frameCounter;
// Sweep cadence and retire age mirror VkRenderPassManager::OnPresent: an entry
// idle for more than kRetireAgeFrames frame boundaries cannot be referenced by
// any in-flight command buffer (frames-in-flight <= MOBILEGL_MAGMA_FRAMESINFLIGHT),
// so its shader modules and layouts are destroyed immediately - no deferred-
// destroy machinery needed. Eviction is content-based, never tied to
// glDeleteProgram: the cache is content-hash-shared across GL programs, so a
// delete-driven erase could free an entry another live program still resolves.
// An evicted entry self-heals - the frontend program keeps its generated
// SPIR-V, so the next GetOrCreateProgram rebuilds it (this also covers the
// renderer's internal blit/depth-mipmap programs).
constexpr Uint64 kSweepInterval = 256;
constexpr Uint64 kRetireAgeFrames = 1024;
if ((m_frameCounter % kSweepInterval) != 0) {
return;
}
for (auto it = m_cache.begin(); it != m_cache.end();) {
if (m_frameCounter - it->second.lastUsedFrame > kRetireAgeFrames) {
const HashType hash = it->first;
const VkDescriptorSetLayout descriptorSetLayout = it->second.descriptorSetLayout;
MGLOG_D("ProgramFactory::OnFrameBoundary: evicting idle program entry hash=0x%llx",
static_cast<unsigned long long>(hash));
// erase runs ~VkProgramObject (modules/layouts destroyed); notify after
// so an observer never observes a half-destroyed entry through a lookup.
// Observers only need the handle values to purge their keyed caches.
it = m_cache.erase(it);
if (m_evictionObserver != nullptr) {
m_evictionObserver->OnProgramEvicted(hash, descriptorSetLayout);
}
} else {
++it;
}
}
}
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -42,6 +42,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
SurfaceRotate90 = 1 << 2, SurfaceRotate90 = 1 << 2,
SurfaceRotate180 = 1 << 3, SurfaceRotate180 = 1 << 3,
SurfaceRotate270 = 1 << 4, SurfaceRotate270 = 1 << 4,
// Rewrites the fragment stage's implicit-LOD image samples to explicit LOD 0.
// Only ever set for a draw whose every sampler binding is clamped to a single mip
// level, which makes the two forms produce identical texels (the implicit lambda is
// clamped into [minLod, maxLod] = [0, 0] regardless of derivatives or bias).
ExplicitLod0Sampling = 1 << 5,
}; };
using CompileOptionFlags = Flags<CompileOptionBit>; using CompileOptionFlags = Flags<CompileOptionBit>;
using HashType = Uint64; using HashType = Uint64;
@@ -59,6 +64,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Vector<DescriptorBindingKind> bindingKinds; Vector<DescriptorBindingKind> bindingKinds;
Vector<Uint32> dynamicBindings; Vector<Uint32> dynamicBindings;
Vector<Int> uniformBlockIndexByBinding; Vector<Int> uniformBlockIndexByBinding;
// Descriptor count per binding (1 except for UBO instance arrays, which occupy one
// binding with descriptorCount = N).
Vector<Uint16> bindingDescriptorCounts;
// Per-element GL uniform block indices for arrayed UBO bindings (count > 1);
// element 0 of a non-arrayed binding stays in uniformBlockIndexByBinding.
UnorderedMap<Uint32, Vector<Int>> arrayedUniformBlockIndicesByBinding;
Vector<String> samplerNameByBinding; Vector<String> samplerNameByBinding;
Vector<Int> samplerUniformLocationByBinding; Vector<Int> samplerUniformLocationByBinding;
Vector<TextureTarget> samplerTextureTargetByBinding; Vector<TextureTarget> samplerTextureTargetByBinding;
@@ -82,6 +93,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// gl_FragDepth); shader-computed depth is immune to the cross-pipeline // gl_FragDepth); shader-computed depth is immune to the cross-pipeline
// position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite). // position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false; Bool fragmentReplacesDepth = false;
// Frame-boundary counter value of the last GetOrCreateProgram hit; drives
// cache eviction (see OnFrameBoundary).
Uint64 lastUsedFrame = 0;
static inline VkDevice s_device = VK_NULL_HANDLE; static inline VkDevice s_device = VK_NULL_HANDLE;
@@ -97,6 +111,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
bindingKinds = std::move(other.bindingKinds); bindingKinds = std::move(other.bindingKinds);
dynamicBindings = std::move(other.dynamicBindings); dynamicBindings = std::move(other.dynamicBindings);
uniformBlockIndexByBinding = std::move(other.uniformBlockIndexByBinding); uniformBlockIndexByBinding = std::move(other.uniformBlockIndexByBinding);
bindingDescriptorCounts = std::move(other.bindingDescriptorCounts);
arrayedUniformBlockIndicesByBinding = std::move(other.arrayedUniformBlockIndicesByBinding);
samplerNameByBinding = std::move(other.samplerNameByBinding); samplerNameByBinding = std::move(other.samplerNameByBinding);
samplerUniformLocationByBinding = std::move(other.samplerUniformLocationByBinding); samplerUniformLocationByBinding = std::move(other.samplerUniformLocationByBinding);
samplerTextureTargetByBinding = std::move(other.samplerTextureTargetByBinding); samplerTextureTargetByBinding = std::move(other.samplerTextureTargetByBinding);
@@ -116,6 +132,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth; fragmentReplacesDepth = other.fragmentReplacesDepth;
lastUsedFrame = other.lastUsedFrame;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE; other.pipelineLayout = VK_NULL_HANDLE;
@@ -127,6 +144,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false; other.fragmentReplacesDepth = false;
other.lastUsedFrame = 0;
} }
VkProgramObject& operator=(VkProgramObject&& other) noexcept { VkProgramObject& operator=(VkProgramObject&& other) noexcept {
if (this == &other) { if (this == &other) {
@@ -141,6 +159,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
bindingKinds = std::move(other.bindingKinds); bindingKinds = std::move(other.bindingKinds);
dynamicBindings = std::move(other.dynamicBindings); dynamicBindings = std::move(other.dynamicBindings);
uniformBlockIndexByBinding = std::move(other.uniformBlockIndexByBinding); uniformBlockIndexByBinding = std::move(other.uniformBlockIndexByBinding);
bindingDescriptorCounts = std::move(other.bindingDescriptorCounts);
arrayedUniformBlockIndicesByBinding = std::move(other.arrayedUniformBlockIndicesByBinding);
samplerNameByBinding = std::move(other.samplerNameByBinding); samplerNameByBinding = std::move(other.samplerNameByBinding);
samplerUniformLocationByBinding = std::move(other.samplerUniformLocationByBinding); samplerUniformLocationByBinding = std::move(other.samplerUniformLocationByBinding);
samplerTextureTargetByBinding = std::move(other.samplerTextureTargetByBinding); samplerTextureTargetByBinding = std::move(other.samplerTextureTargetByBinding);
@@ -160,6 +180,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
producerOutputComponentCount = other.producerOutputComponentCount; producerOutputComponentCount = other.producerOutputComponentCount;
fragmentInputComponentCount = other.fragmentInputComponentCount; fragmentInputComponentCount = other.fragmentInputComponentCount;
fragmentReplacesDepth = other.fragmentReplacesDepth; fragmentReplacesDepth = other.fragmentReplacesDepth;
lastUsedFrame = other.lastUsedFrame;
other.hash = 0; other.hash = 0;
other.descriptorSetLayout = VK_NULL_HANDLE; other.descriptorSetLayout = VK_NULL_HANDLE;
other.pipelineLayout = VK_NULL_HANDLE; other.pipelineLayout = VK_NULL_HANDLE;
@@ -171,6 +192,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
other.producerOutputComponentCount = 0; other.producerOutputComponentCount = 0;
other.fragmentInputComponentCount = 0; other.fragmentInputComponentCount = 0;
other.fragmentReplacesDepth = false; other.fragmentReplacesDepth = false;
other.lastUsedFrame = 0;
return *this; return *this;
} }
@@ -200,6 +222,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
}; };
// Notified when the OnFrameBoundary sweep destroys an aged-out cache entry,
// carrying the entry's content hash and the VkDescriptorSetLayout it owned.
// Dependent caches (compute pipelines, PipelineFactory entries, UniformManager's
// per-layout descriptor sets) must purge in the same step: after vkDestroy the
// layout handle value may be recycled for an unrelated layout, and the program
// hash may be re-inserted by a later rebuild of the same content.
class IEvictionObserver {
public:
virtual ~IEvictionObserver() = default;
virtual void OnProgramEvicted(HashType programHash, VkDescriptorSetLayout descriptorSetLayout) = 0;
};
explicit ProgramFactory(VkDevice device, const VulkanRendererConfig& config, Uint32 maxBindings = 16, explicit ProgramFactory(VkDevice device, const VulkanRendererConfig& config, Uint32 maxBindings = 16,
Bool shaderDrawParametersEnabled = false, Bool shaderDrawParametersEnabled = false,
Bool unformattedFloatStorageImagesEnabled = false) Bool unformattedFloatStorageImagesEnabled = false)
@@ -215,6 +249,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const VkProgramObject& GetOrCreateProgram( const VkProgramObject& GetOrCreateProgram(
const MG_State::GLState::ProgramObject& program, CompileOptionFlags flags); const MG_State::GLState::ProgramObject& program, CompileOptionFlags flags);
// Observer may be null (no notifications). Not owned.
void SetEvictionObserver(IEvictionObserver* observer) { m_evictionObserver = observer; }
// Frame boundary hook: ages the program cache and evicts long-unused entries
// (their command buffers retired many frames ago), mirroring
// VkRenderPassManager::OnPresent's sweep.
void OnFrameBoundary();
static VkShaderStageFlagBits ToVkStage(ShaderStage stage); static VkShaderStageFlagBits ToVkStage(ShaderStage stage);
static VkFormat ConvertSpirvImageFormatToVkFormat(SpvImageFormat format); static VkFormat ConvertSpirvImageFormatToVkFormat(SpvImageFormat format);
static SamplerNumericDomain UniformTypeToSamplerNumericDomain(GLenum glType); static SamplerNumericDomain UniformTypeToSamplerNumericDomain(GLenum glType);
@@ -256,6 +297,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// shaderStorageImageReadWithoutFormat and shaderStorageImageWriteWithoutFormat. // shaderStorageImageReadWithoutFormat and shaderStorageImageWriteWithoutFormat.
Bool m_unformattedFloatStorageImagesEnabled = false; Bool m_unformattedFloatStorageImagesEnabled = false;
mutable ProgramLookupCache m_lastLookup; mutable ProgramLookupCache m_lastLookup;
// Monotonic frame-boundary counter (bumped in OnFrameBoundary) for cache aging.
Uint64 m_frameCounter = 0;
IEvictionObserver* m_evictionObserver = nullptr;
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -247,6 +247,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_surfaceFormat = {createInfo.imageFormat, createInfo.imageColorSpace}; m_surfaceFormat = {createInfo.imageFormat, createInfo.imageColorSpace};
m_extent = createInfo.imageExtent; m_extent = createInfo.imageExtent;
// The surface-space extent this swapchain was built from, i.e. before the
// quarter-turn swap above. Out-of-date checks must compare in THIS space: comparing a
// freshly queried currentExtent against the swapped m_extent flips axes every rotation
// and makes the comparison alternate forever.
m_surfaceExtent = defaultFramebufferExtent;
m_preTransform = createInfo.preTransform; m_preTransform = createInfo.preTransform;
VK_VERIFY(vkCreateSwapchainKHR(device, &createInfo, nullptr, &m_swapchain)); VK_VERIFY(vkCreateSwapchainKHR(device, &createInfo, nullptr, &m_swapchain));
@@ -257,6 +262,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_images.resize(imageCount, VK_NULL_HANDLE); m_images.resize(imageCount, VK_NULL_HANDLE);
VK_VERIFY(vkGetSwapchainImagesKHR(device, m_swapchain, &imageCount, m_images.data())); VK_VERIFY(vkGetSwapchainImagesKHR(device, m_swapchain, &imageCount, m_images.data()));
m_imageLayouts.assign(imageCount, VK_IMAGE_LAYOUT_UNDEFINED); m_imageLayouts.assign(imageCount, VK_IMAGE_LAYOUT_UNDEFINED);
// Fresh swapchain images hold garbage until a render pass stores into them.
m_imageContentDefined.assign(imageCount, false);
m_depthStencilContentDefined.assign(imageCount, false);
CreateImageViews(device); CreateImageViews(device);
CreateDepthStencilResources(device, physicalDevice); CreateDepthStencilResources(device, physicalDevice);
@@ -428,9 +436,39 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_images.clear(); m_images.clear();
m_imageLayouts.clear(); m_imageLayouts.clear();
m_imageContentDefined.clear();
m_depthStencilContentDefined.clear();
m_preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR; m_preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR;
} }
Bool SwapchainObject::IsImageContentDefined(Uint32 index) const {
MOBILEGL_ASSERT(index < m_imageContentDefined.size(), "Swapchain image content index out of range");
return m_imageContentDefined[index];
}
void SwapchainObject::SetImageContentDefined(Uint32 index, Bool defined) {
MOBILEGL_ASSERT(index < m_imageContentDefined.size(), "Swapchain image content index out of range");
m_imageContentDefined[index] = defined;
}
Bool SwapchainObject::IsDepthStencilContentDefined(Uint32 index) const {
MOBILEGL_ASSERT(index < m_depthStencilContentDefined.size(),
"Swapchain depth/stencil content index out of range");
return m_depthStencilContentDefined[index];
}
void SwapchainObject::SetDepthStencilContentDefined(Uint32 index, Bool defined) {
MOBILEGL_ASSERT(index < m_depthStencilContentDefined.size(),
"Swapchain depth/stencil content index out of range");
m_depthStencilContentDefined[index] = defined;
}
void SwapchainObject::SetAllDepthStencilContentUndefined() {
for (SizeT i = 0; i < m_depthStencilContentDefined.size(); ++i) {
m_depthStencilContentDefined[i] = false;
}
}
VkImage SwapchainObject::GetImage(Uint32 index) const { VkImage SwapchainObject::GetImage(Uint32 index) const {
MOBILEGL_ASSERT(index < m_images.size(), "Swapchain image index out of range"); MOBILEGL_ASSERT(index < m_images.size(), "Swapchain image index out of range");
return m_images[index]; return m_images[index];
@@ -35,6 +35,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkSwapchainKHR GetHandle() const { return m_swapchain; } VkSwapchainKHR GetHandle() const { return m_swapchain; }
const VkSurfaceFormatKHR& GetSurfaceFormat() const { return m_surfaceFormat; } const VkSurfaceFormatKHR& GetSurfaceFormat() const { return m_surfaceFormat; }
VkExtent2D GetExtent() const { return m_extent; } VkExtent2D GetExtent() const { return m_extent; }
// Surface-space extent (before the pre-rotation quarter-turn swap) this swapchain was
// created from - the value to compare a freshly queried currentExtent against.
VkExtent2D GetSurfaceExtent() const { return m_surfaceExtent; }
VkSurfaceTransformFlagBitsKHR GetPreTransform() const { return m_preTransform; } VkSurfaceTransformFlagBitsKHR GetPreTransform() const { return m_preTransform; }
const Vector<VkImage>& GetImages() const { return m_images; } const Vector<VkImage>& GetImages() const { return m_images; }
const Vector<VkImageView>& GetImageViews() const { return m_imageViews; } const Vector<VkImageView>& GetImageViews() const { return m_imageViews; }
@@ -49,6 +52,21 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void SetImageLayout(Uint32 index, VkImageLayout layout); void SetImageLayout(Uint32 index, VkImageLayout layout);
SizeT GetImageCount() const { return m_images.size(); } SizeT GetImageCount() const { return m_images.size(); }
// EGL content-validity tracking for the default framebuffer. A color
// buffer's content is undefined once its image has been presented
// (EGL_BUFFER_DESTROYED swap behaviour, the implementation default),
// and every ancillary (depth/stencil) buffer's content is undefined
// after ANY swap regardless of swap behaviour (EGL 1.5 §3.10.1). The
// render-pass manager turns an undefined attachment's tile load into
// LOAD_OP_DONT_CARE. Flags start false (a fresh swapchain image holds
// garbage) and a render pass storing into an attachment sets it back
// to defined.
Bool IsImageContentDefined(Uint32 index) const;
void SetImageContentDefined(Uint32 index, Bool defined);
Bool IsDepthStencilContentDefined(Uint32 index) const;
void SetDepthStencilContentDefined(Uint32 index, Bool defined);
void SetAllDepthStencilContentUndefined();
private: private:
void CreateImageViews(VkDevice device); void CreateImageViews(VkDevice device);
void CreateDepthStencilResources(VkDevice device, VkPhysicalDevice physicalDevice); void CreateDepthStencilResources(VkDevice device, VkPhysicalDevice physicalDevice);
@@ -63,6 +81,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkSwapchainKHR m_swapchain = VK_NULL_HANDLE; VkSwapchainKHR m_swapchain = VK_NULL_HANDLE;
VkSurfaceFormatKHR m_surfaceFormat{}; VkSurfaceFormatKHR m_surfaceFormat{};
VkExtent2D m_extent{}; VkExtent2D m_extent{};
VkExtent2D m_surfaceExtent{};
VkSurfaceTransformFlagBitsKHR m_preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR; VkSurfaceTransformFlagBitsKHR m_preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR;
Vector<VkImage> m_images; Vector<VkImage> m_images;
Vector<VkImageView> m_imageViews; Vector<VkImageView> m_imageViews;
@@ -73,5 +92,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Vector<VkDeviceMemory> m_depthStencilImageMemories; Vector<VkDeviceMemory> m_depthStencilImageMemories;
Vector<VkImageView> m_depthStencilImageViews; Vector<VkImageView> m_depthStencilImageViews;
Vector<VkImageLayout> m_depthStencilImageLayouts; Vector<VkImageLayout> m_depthStencilImageLayouts;
Vector<Bool> m_imageContentDefined;
Vector<Bool> m_depthStencilContentDefined;
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -18,6 +18,7 @@
#include "MG_Util/Converters/MGToVk/TextureEnumConverter.h" #include "MG_Util/Converters/MGToVk/TextureEnumConverter.h"
#include "MG_Util/Metrics/TextureMetrics.h" #include "MG_Util/Metrics/TextureMetrics.h"
#include <Config.h> #include <Config.h>
#include <algorithm>
#include <cstdio> #include <cstdio>
#include <cstdlib> #include <cstdlib>
#include <cstring> #include <cstring>
@@ -204,6 +205,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// The frame's descriptor sets are recycled above, so last frame's reuse target // The frame's descriptor sets are recycled above, so last frame's reuse target
// is gone: start the per-draw descriptor-reuse cache fresh this frame. // is gone: start the per-draw descriptor-reuse cache fresh this frame.
m_hasLastDescriptor = false; m_hasLastDescriptor = false;
m_lastBindValid = false;
// Re-fingerprint the bound sampler set fresh this frame so any GL object address // Re-fingerprint the bound sampler set fresh this frame so any GL object address
// reuse cannot outlive a single frame (see SamplerResolveMemo). // reuse cannot outlive a single frame (see SamplerResolveMemo).
for (auto& memo : m_samplerResolveMemo) { for (auto& memo : m_samplerResolveMemo) {
@@ -211,6 +213,40 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} }
void UniformManager::OnDescriptorSetLayoutDestroyed(VkDescriptorSetLayout descriptorSetLayout) {
SizeT purgedSets = 0;
for (auto& frame : m_frames) {
const auto it = frame.descriptorSetCacheByLayout.find(descriptorSetLayout);
if (it == frame.descriptorSetCacheByLayout.end()) {
continue;
}
// Free the sets back to their pools and credit the bucket accounting, so
// program churn recycles pool capacity instead of abandoning the slots.
// GPU-safe: the layout only dies after >1024 idle frame boundaries, so no
// in-flight command buffer references these sets.
for (const auto& cached : it->second.sets) {
if (cached.set == VK_NULL_HANDLE) {
continue;
}
vkFreeDescriptorSets(m_device, cached.pool, 1, &cached.set);
const auto bucket = std::find_if(
frame.descriptorPools.begin(), frame.descriptorPools.end(),
[&cached](const DescriptorPoolBucket& candidate) { return candidate.handle == cached.pool; });
if (bucket != frame.descriptorPools.end() && bucket->allocatedSets > 0) {
--bucket->allocatedSets;
}
}
purgedSets += it->second.sets.size();
frame.descriptorSetCacheByLayout.erase(it);
}
if (purgedSets > 0) {
// The per-draw reuse memo folds the layout handle into its signature; drop
// it so a recycled handle value cannot revive a purged set mid-frame.
m_hasLastDescriptor = false;
MGLOG_D("UniformDescriptorBinder: freed %zu descriptor sets for destroyed layout", purgedSets);
}
}
Bool UniformManager::ResolveSamplerDescriptor(VkCommandBuffer commandBuffer, Bool UniformManager::ResolveSamplerDescriptor(VkCommandBuffer commandBuffer,
const MG_State::GLState::ProgramObject& program, const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj, const ProgramFactory::VkProgramObject& programObj,
@@ -350,24 +386,28 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const Uint16 samplerVersion = samplerToUse->GetVersion(); const Uint16 samplerVersion = samplerToUse->GetVersion();
const Uint64 textureLifetimeId = texture->GetLifetimeId(); const Uint64 textureLifetimeId = texture->GetLifetimeId();
const Uint16 textureParamsVersion = texture->GetTextureParamsVersion(); const Uint16 textureParamsVersion = texture->GetTextureParamsVersion();
// The sampler's LOD clamp depends on how many levels the sampled view exposes, and that
// follows uploads as well as GL parameters - so it belongs in the memo key too.
const Uint32 viewLevelCount = resource->sampledLevelCount;
if (memo.valid && memo.samplerLifetimeId == samplerLifetimeId && memo.samplerVersion == samplerVersion && if (memo.valid && memo.samplerLifetimeId == samplerLifetimeId && memo.samplerVersion == samplerVersion &&
memo.textureLifetimeId == textureLifetimeId && memo.textureParamsVersion == textureParamsVersion && memo.textureLifetimeId == textureLifetimeId && memo.textureParamsVersion == textureParamsVersion &&
memo.forceNearestFiltering == forceNearestFiltering) { memo.forceNearestFiltering == forceNearestFiltering && memo.viewLevelCount == viewLevelCount) {
resolvedSampler = memo.sampler; resolvedSampler = memo.sampler;
} else { } else {
resolvedSampler = resolvedSampler = m_samplerManager->GetOrCreateSampler(*samplerToUse, *texture,
m_samplerManager->GetOrCreateSampler(*samplerToUse, *texture, forceNearestFiltering); forceNearestFiltering, viewLevelCount);
memo.samplerLifetimeId = samplerLifetimeId; memo.samplerLifetimeId = samplerLifetimeId;
memo.samplerVersion = samplerVersion; memo.samplerVersion = samplerVersion;
memo.textureLifetimeId = textureLifetimeId; memo.textureLifetimeId = textureLifetimeId;
memo.textureParamsVersion = textureParamsVersion; memo.textureParamsVersion = textureParamsVersion;
memo.forceNearestFiltering = forceNearestFiltering; memo.forceNearestFiltering = forceNearestFiltering;
memo.viewLevelCount = viewLevelCount;
memo.sampler = resolvedSampler; memo.sampler = resolvedSampler;
memo.valid = true; memo.valid = true;
} }
} else { } else {
resolvedSampler = resolvedSampler = m_samplerManager->GetOrCreateSampler(*samplerToUse, *texture, forceNearestFiltering,
m_samplerManager->GetOrCreateSampler(*samplerToUse, *texture, forceNearestFiltering); resource->sampledLevelCount);
} }
outImageInfo = { outImageInfo = {
.sampler = resolvedSampler, .sampler = resolvedSampler,
@@ -408,6 +448,48 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return outImageInfo.sampler != VK_NULL_HANDLE; return outImageInfo.sampler != VK_NULL_HANDLE;
} }
Bool UniformManager::ProgramSamplesOnlySingleLevelTextures(
const MG_State::GLState::ProgramObject& program, const ProgramFactory::VkProgramObject& programObj) {
Bool sawSampler = false;
for (Uint32 binding = 0; binding < programObj.bindingKinds.size(); ++binding) {
if (programObj.bindingKinds[binding] != ProgramFactory::DescriptorBindingKind::CombinedImageSampler) {
continue;
}
const auto* texture = ResolveSamplerTextureRaw(program, programObj, binding);
if (texture == nullptr) return false;
const auto& levelRange = texture->GetLevelRange();
if (levelRange.x() != levelRange.y()) return false;
// An explicit-LOD sample is a single filtered tap, so it also gives up anisotropic
// filtering - which a single-level view can still have. Resolve the sampler exactly
// the way ResolveSamplerDescriptor does and bail if anisotropy would apply.
const Int location = programObj.samplerUniformLocationByBinding[binding];
const Int unit = ResolveSamplerUnitIndex(program, location, binding);
const auto& samplerOverride = MG_State::pGLContext->GetTextureUnitObject(unit).GetSamplerObject();
const auto* effectiveSampler =
samplerOverride ? samplerOverride.get() : texture->GetSamplerObject().get();
if (effectiveSampler == nullptr) return false;
if (effectiveSampler->GetMaxAnisotropy() > 1.0f &&
effectiveSampler->GetMinFilter() == SamplerFilterMode::Linear &&
effectiveSampler->GetMagFilter() == SamplerFilterMode::Linear) {
return false;
}
// An explicit LOD 0 makes lambda exactly 0, which is the magnification side of the
// min/mag decision. That only matches the implicit form when lambda could not have been
// positive anyway (the LOD clamp already pins it at or below 0), or when the two
// filters are the same and the choice cannot be observed.
const Float effectiveMaxLod = effectiveSampler->GetMipmapMode() == SamplerMipmapMode::None
? 0.0f
: effectiveSampler->GetMaxLod();
if (effectiveMaxLod > 0.0f && effectiveSampler->GetMinFilter() != effectiveSampler->GetMagFilter()) {
return false;
}
sawSampler = true;
}
return sawSampler;
}
Bool UniformManager::ResolveSamplerTexture(const MG_State::GLState::ProgramObject& program, Bool UniformManager::ResolveSamplerTexture(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj, Uint32 binding, const ProgramFactory::VkProgramObject& programObj, Uint32 binding,
SharedPtr<MG_State::GLState::ITextureObject>& outTexture) { SharedPtr<MG_State::GLState::ITextureObject>& outTexture) {
@@ -765,7 +847,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool UniformManager::ResolveUniformBufferPayload(const MG_State::GLState::ProgramObject& program, Bool UniformManager::ResolveUniformBufferPayload(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj, Uint32 binding, const ProgramFactory::VkProgramObject& programObj, Uint32 binding,
UboBindResult& out) const { Uint32 arrayElement, UboBindResult& out) const {
const void* outData = nullptr; const void* outData = nullptr;
VkDeviceSize outSize = 0; VkDeviceSize outSize = 0;
@@ -791,7 +873,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
MOBILEGL_ASSERT(binding < programObj.uniformBlockIndexByBinding.size(), MOBILEGL_ASSERT(binding < programObj.uniformBlockIndexByBinding.size(),
"ResolveUniformBufferPayload: UBO mapping binding %u out of range", binding); "ResolveUniformBufferPayload: UBO mapping binding %u out of range", binding);
const Int blockIndex = programObj.uniformBlockIndexByBinding[binding]; Int blockIndex = programObj.uniformBlockIndexByBinding[binding];
if (arrayElement > 0) {
const auto arrayIt = programObj.arrayedUniformBlockIndicesByBinding.find(binding);
const Bool elementValid = arrayIt != programObj.arrayedUniformBlockIndicesByBinding.end() &&
arrayElement < arrayIt->second.size();
MOBILEGL_ASSERT(elementValid,
"ResolveUniformBufferPayload: UBO binding %u has no array element %u", binding,
arrayElement);
if (!elementValid) {
return false;
}
blockIndex = arrayIt->second[arrayElement];
}
MOBILEGL_ASSERT(blockIndex >= 0, MOBILEGL_ASSERT(blockIndex >= 0,
"ResolveUniformBufferPayload: no uniform block mapped to descriptor binding %u", binding); "ResolveUniformBufferPayload: no uniform block mapped to descriptor binding %u", binding);
@@ -899,6 +993,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkDescriptorPoolCreateInfo poolInfo{}; VkDescriptorPoolCreateInfo poolInfo{};
poolInfo.sType = VK_STRUCTURE_TYPE_DESCRIPTOR_POOL_CREATE_INFO; poolInfo.sType = VK_STRUCTURE_TYPE_DESCRIPTOR_POOL_CREATE_INFO;
// FREE_DESCRIPTOR_SET_BIT lets a destroyed layout's cached sets be freed back
// (OnDescriptorSetLayoutDestroyed) so program churn recycles pool capacity.
// The cost is on set allocation only, which happens when a layout's per-frame
// cache grows - never on the per-draw reuse path.
poolInfo.flags = VK_DESCRIPTOR_POOL_CREATE_FREE_DESCRIPTOR_SET_BIT;
poolInfo.maxSets = maxSets; poolInfo.maxSets = maxSets;
poolInfo.poolSizeCount = static_cast<Uint32>(std::size(poolSizes)); poolInfo.poolSizeCount = static_cast<Uint32>(std::size(poolSizes));
poolInfo.pPoolSizes = poolSizes; poolInfo.pPoolSizes = poolSizes;
@@ -978,7 +1077,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
auto& frame = m_frames[frameIndex]; auto& frame = m_frames[frameIndex];
auto& cache = frame.descriptorSetCacheByLayout[programObj.descriptorSetLayout]; auto& cache = frame.descriptorSetCacheByLayout[programObj.descriptorSetLayout];
if (cache.cursor < cache.sets.size()) { if (cache.cursor < cache.sets.size()) {
outDescriptorSet = cache.sets[cache.cursor++]; outDescriptorSet = cache.sets[cache.cursor++].set;
} else { } else {
VkResult allocResult = AllocateDescriptorSetsFromActivePool(frameIndex, programObj, outDescriptorSet); VkResult allocResult = AllocateDescriptorSetsFromActivePool(frameIndex, programObj, outDescriptorSet);
if (allocResult == VK_ERROR_OUT_OF_POOL_MEMORY || allocResult == VK_ERROR_FRAGMENTED_POOL) { if (allocResult == VK_ERROR_OUT_OF_POOL_MEMORY || allocResult == VK_ERROR_FRAGMENTED_POOL) {
@@ -992,7 +1091,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return allocResult; return allocResult;
} }
cache.sets.push_back(outDescriptorSet); // The successful allocation came from the bucket the alloc helper left
// active; record it so a layout-destroyed purge can free the set back.
cache.sets.push_back({outDescriptorSet, frame.descriptorPools[frame.activeDescriptorPoolIndex].handle});
++cache.cursor; ++cache.cursor;
MGLOG_D("UniformDescriptorBinder: cached descriptor set count for frame=%u grew to %zu", frameIndex, MGLOG_D("UniformDescriptorBinder: cached descriptor set count for frame=%u grew to %zu", frameIndex,
cache.sets.size()); cache.sets.size());
@@ -1038,11 +1139,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
imageInfos.clear(); imageInfos.clear();
texelBufferViews.clear(); texelBufferViews.clear();
dynamicOffsets.clear(); dynamicOffsets.clear();
// Arrayed UBO bindings contribute extra buffer infos and dynamic offsets; reserve for
// the worst case so the pBufferInfo pointers taken below never dangle on reallocation.
Uint32 uboArrayExtra = 0;
for (const auto& arrayEntry : programObj.arrayedUniformBlockIndicesByBinding) {
uboArrayExtra += static_cast<Uint32>(arrayEntry.second.size()) - 1u;
}
writes.reserve(m_maxBindings); writes.reserve(m_maxBindings);
bufferInfos.reserve(m_maxBindings); bufferInfos.reserve(m_maxBindings + uboArrayExtra);
imageInfos.reserve(m_maxBindings); imageInfos.reserve(m_maxBindings);
texelBufferViews.reserve(m_maxBindings); texelBufferViews.reserve(m_maxBindings);
dynamicOffsets.reserve(programObj.dynamicBindings.size()); dynamicOffsets.reserve(programObj.dynamicBindings.size() + uboArrayExtra);
const Uint32 bindingCount = const Uint32 bindingCount =
std::min<Uint32>(m_maxBindings, static_cast<Uint32>(programObj.bindingKinds.size())); std::min<Uint32>(m_maxBindings, static_cast<Uint32>(programObj.bindingKinds.size()));
@@ -1060,11 +1167,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
write.descriptorCount = 1; write.descriptorCount = 1;
if (kind == ProgramFactory::DescriptorBindingKind::UniformBufferDynamic) { if (kind == ProgramFactory::DescriptorBindingKind::UniformBufferDynamic) {
const Uint32 descriptorCount =
binding < programObj.bindingDescriptorCounts.size()
? std::max<Uint32>(1, programObj.bindingDescriptorCounts[binding])
: 1u;
const SizeT firstBufferInfoIndex = bufferInfos.size();
for (Uint32 element = 0; element < descriptorCount; ++element) {
UboBindResult ubo{}; UboBindResult ubo{};
const Bool hasPayload = ResolveUniformBufferPayload(program, programObj, binding, ubo); const Bool hasPayload =
ResolveUniformBufferPayload(program, programObj, binding, element, ubo);
MOBILEGL_ASSERT(hasPayload && ubo.payload != nullptr && ubo.payloadSize > 0, MOBILEGL_ASSERT(hasPayload && ubo.payload != nullptr && ubo.payloadSize > 0,
"UniformDescriptorBinder::BindProgramUniformBuffers failed: missing UBO payload on binding %u", "UniformDescriptorBinder::BindProgramUniformBuffers failed: missing UBO payload on binding %u element %u",
binding); binding, element);
VkDescriptorBufferInfo bufferInfo{}; VkDescriptorBufferInfo bufferInfo{};
// Keep offset 0 (sub-range selected via the dynamic offset) so the hashed bufferInfo // Keep offset 0 (sub-range selected via the dynamic offset) so the hashed bufferInfo
@@ -1077,23 +1191,58 @@ namespace MobileGL::MG_Backend::DirectVulkan {
bufferInfo.range = ubo.range; bufferInfo.range = ubo.range;
dynOffset = static_cast<Uint32>(ubo.dynamicOffset); dynOffset = static_cast<Uint32>(ubo.dynamicOffset);
} else { } else {
// Global-UBO slice reuse (see GlobalUboSliceMemo): unchanged
// uniform bytes re-use the slice already uploaded this frame.
const Bool isGlobalUbo =
programObj.globalUboBinding == static_cast<Int>(binding) && element == 0;
const Uint64 uboFrameSerial = m_bufferManager->GetFrameSerial();
const Uint64 uboProgramLifetimeId = program.GetLifetimeId();
const Uint32 uboContentVersion = program.GetUBOContentVersion();
Bool reusedSlice = false;
if (isGlobalUbo) {
for (const auto& memo : m_globalUboMemo) {
if (memo.buffer != VK_NULL_HANDLE &&
memo.programLifetimeId == uboProgramLifetimeId &&
memo.frameSerial == uboFrameSerial &&
memo.uboContentVersion == uboContentVersion &&
memo.range == static_cast<VkDeviceSize>(ubo.payloadSize)) {
bufferInfo.buffer = memo.buffer;
bufferInfo.range = memo.range;
dynOffset = static_cast<Uint32>(memo.offset);
reusedSlice = true;
break;
}
}
}
if (!reusedSlice) {
BufferSlice slice{}; BufferSlice slice{};
if (!m_bufferManager->UploadTransient(BufferKind::Uniform, frameIndex, ubo.payload, if (!m_bufferManager->UploadTransient(BufferKind::Uniform, frameIndex, ubo.payload,
ubo.payloadSize, m_minDynamicOffsetAlignment, slice)) { ubo.payloadSize, m_minDynamicOffsetAlignment, slice)) {
MOBILEGL_ASSERT(false, "UniformDescriptorBinder::BindProgramUniformBuffers failed: UBO upload failed on binding %u", MOBILEGL_ASSERT(false, "UniformDescriptorBinder::BindProgramUniformBuffers failed: UBO upload failed on binding %u element %u",
binding); binding, element);
return false; return false;
} }
bufferInfo.buffer = slice.buffer; bufferInfo.buffer = slice.buffer;
bufferInfo.range = ubo.payloadSize; bufferInfo.range = ubo.payloadSize;
dynOffset = static_cast<Uint32>(slice.offset); dynOffset = static_cast<Uint32>(slice.offset);
if (isGlobalUbo) {
m_globalUboMemo[m_globalUboMemoNext] = GlobalUboSliceMemo{
uboProgramLifetimeId, uboFrameSerial, uboContentVersion,
slice.buffer, slice.offset, static_cast<VkDeviceSize>(ubo.payloadSize)};
m_globalUboMemoNext = (m_globalUboMemoNext + 1) % kGlobalUboMemoSize;
}
}
} }
bufferInfos.push_back(bufferInfo); bufferInfos.push_back(bufferInfo);
// Dynamic offsets are consumed in binding order, then array element order,
// matching Vulkan's dynamic-offset consumption rules.
dynamicOffsets.push_back(dynOffset);
}
write.descriptorType = VK_DESCRIPTOR_TYPE_UNIFORM_BUFFER_DYNAMIC; write.descriptorType = VK_DESCRIPTOR_TYPE_UNIFORM_BUFFER_DYNAMIC;
write.pBufferInfo = &bufferInfos.back(); write.descriptorCount = descriptorCount;
write.pBufferInfo = &bufferInfos[firstBufferInfoIndex];
writes.push_back(write); writes.push_back(write);
dynamicOffsets.push_back(dynOffset);
} else if (kind == ProgramFactory::DescriptorBindingKind::UniformTexelBuffer) { } else if (kind == ProgramFactory::DescriptorBindingKind::UniformTexelBuffer) {
VkBufferView bufferView = VK_NULL_HANDLE; VkBufferView bufferView = VK_NULL_HANDLE;
if (!ResolveTexelBufferDescriptor(program, programObj, binding, frameIndex, bufferView) || if (!ResolveTexelBufferDescriptor(program, programObj, binding, frameIndex, bufferView) ||
@@ -1221,8 +1370,34 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_hasLastDescriptor = cacheable; m_hasLastDescriptor = cacheable;
} }
// Skip the driver call when this exact binding is already live on the
// command buffer (see the bind-dedup shadow in the header).
const Uint32 offsetCount = static_cast<Uint32>(dynamicOffsets.size());
Bool identicalBind = m_lastBindValid && m_lastBindSet == descriptorSet &&
m_lastBindLayout == programObj.pipelineLayout && m_lastBindPoint == bindPoint &&
m_lastBindOffsetCount == offsetCount && offsetCount <= kMaxShadowedDynamicOffsets;
if (identicalBind) {
for (Uint32 i = 0; i < offsetCount; ++i) {
if (m_lastBindOffsets[i] != dynamicOffsets[i]) {
identicalBind = false;
break;
}
}
}
if (!identicalBind) {
vkCmdBindDescriptorSets(commandBuffer, bindPoint, programObj.pipelineLayout, 0, 1, vkCmdBindDescriptorSets(commandBuffer, bindPoint, programObj.pipelineLayout, 0, 1,
&descriptorSet, static_cast<Uint32>(dynamicOffsets.size()), dynamicOffsets.data()); &descriptorSet, offsetCount, dynamicOffsets.data());
if (offsetCount <= kMaxShadowedDynamicOffsets) {
m_lastBindValid = true;
m_lastBindSet = descriptorSet;
m_lastBindLayout = programObj.pipelineLayout;
m_lastBindPoint = bindPoint;
m_lastBindOffsetCount = offsetCount;
std::copy_n(dynamicOffsets.data(), offsetCount, m_lastBindOffsets);
} else {
m_lastBindValid = false;
}
}
return true; return true;
} }
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -39,6 +39,20 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void Shutdown(); void Shutdown();
void BeginFrame(Uint32 frameIndex); void BeginFrame(Uint32 frameIndex);
// A command buffer (re)began recording: descriptor bindings recorded into
// the previous buffer do not carry over, so drop the bind-dedup shadow.
void OnCommandBufferBoundary() { m_lastBindValid = false; }
// A ProgramFactory eviction just destroyed this layout: purge every frame
// slot's cached descriptor sets for it, so a recycled handle value can never
// stale-hit sets written for the dead layout's bindings. The sets are
// vkFreeDescriptorSets'd back to their pools (created with
// FREE_DESCRIPTOR_SET_BIT) and the pool accounting is credited, so program
// churn recycles pool capacity instead of abandoning it. GPU-safe: the layout
// only dies after >1024 idle frame boundaries, so no in-flight command buffer
// references its sets. This is the only eviction path for the per-layout
// caches - a live layout's entry must never be purged (its sets would be
// unreachable pool slots), so there is deliberately no age-based sweep here.
void OnDescriptorSetLayoutDestroyed(VkDescriptorSetLayout descriptorSetLayout);
Bool CollectSampledTextures(const MG_State::GLState::ProgramObject& program, Bool CollectSampledTextures(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj, const ProgramFactory::VkProgramObject& programObj,
Vector<MG_State::GLState::ITextureObject*>& outTextures); Vector<MG_State::GLState::ITextureObject*>& outTextures);
@@ -58,6 +72,16 @@ namespace MobileGL::MG_Backend::DirectVulkan {
static VkFormat ResolveStorageImageViewFormat(VkFormat reflectedFormat, GLenum bindingFormat, static VkFormat ResolveStorageImageViewFormat(VkFormat reflectedFormat, GLenum bindingFormat,
VkFormat resourceFormat, Bool useBindingFormat); VkFormat resourceFormat, Bool useBindingFormat);
// True when the program reads at least one sampler and every one of them is bound to a
// texture whose GL level range is a single level. Such a sampler resolves to
// minLod = maxLod = 0 (see VkSamplerManager::GetOrCreateSampler), so an implicit-LOD sample
// and an explicit LOD 0 sample must read the same texel - which is what makes the
// ExplicitLod0Sampling SPIR-V rewrite safe to request. Deliberately conservative: it reads
// only GL state, so a texture that ends up single-level for another reason (one uploaded
// level under a wide level range) merely misses the rewrite.
static Bool ProgramSamplesOnlySingleLevelTextures(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj);
private: private:
struct DescriptorPoolBucket { struct DescriptorPoolBucket {
VkDescriptorPool handle = VK_NULL_HANDLE; VkDescriptorPool handle = VK_NULL_HANDLE;
@@ -65,8 +89,16 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 allocatedSets = 0; Uint32 allocatedSets = 0;
}; };
// A cached descriptor set together with the pool it was allocated from, so a
// layout-destroyed purge can vkFreeDescriptorSets it back and credit the
// owning bucket's accounting.
struct CachedDescriptorSet {
VkDescriptorSet set = VK_NULL_HANDLE;
VkDescriptorPool pool = VK_NULL_HANDLE;
};
struct DescriptorSetCacheEntry { struct DescriptorSetCacheEntry {
Vector<VkDescriptorSet> sets; Vector<CachedDescriptorSet> sets;
Uint32 cursor = 0; Uint32 cursor = 0;
}; };
@@ -116,7 +148,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}; };
Bool ResolveUniformBufferPayload(const MG_State::GLState::ProgramObject& program, Bool ResolveUniformBufferPayload(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj, Uint32 binding, const ProgramFactory::VkProgramObject& programObj, Uint32 binding,
UboBindResult& out) const; Uint32 arrayElement, UboBindResult& out) const;
Bool CreateDescriptorPool(Uint32 maxSets, VkDescriptorPool& outPool) const; Bool CreateDescriptorPool(Uint32 maxSets, VkDescriptorPool& outPool) const;
Bool GrowFrameDescriptorPool(FrameResources& frame, Uint32 frameIndex); Bool GrowFrameDescriptorPool(FrameResources& frame, Uint32 frameIndex);
VkResult AllocateDescriptorSetsFromActivePool( VkResult AllocateDescriptorSetsFromActivePool(
@@ -156,6 +188,35 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 m_lastDescriptorSignature = 0; Uint64 m_lastDescriptorSignature = 0;
Bool m_hasLastDescriptor = false; Bool m_hasLastDescriptor = false;
// vkCmdBindDescriptorSets dedup: consecutive draws with a static uniform
// block resolve to the same set AND the same dynamic offsets, so the
// driver call can be skipped outright. Command-buffer-scope state; reset
// via OnCommandBufferBoundary whenever a recording (re)begins. Keyed on
// layout+bind point, so a pipeline-layout switch always rebinds.
static constexpr Uint32 kMaxShadowedDynamicOffsets = 8;
Bool m_lastBindValid = false;
VkDescriptorSet m_lastBindSet = VK_NULL_HANDLE;
VkPipelineLayout m_lastBindLayout = VK_NULL_HANDLE;
VkPipelineBindPoint m_lastBindPoint = VK_PIPELINE_BIND_POINT_GRAPHICS;
Uint32 m_lastBindOffsetCount = 0;
Uint32 m_lastBindOffsets[kMaxShadowedDynamicOffsets] = {};
// Global-UBO transient-slice reuse: MC leaves the default uniform block
// untouched across long GUI/terrain runs, so the per-draw re-upload of
// the same bytes can reuse the slice uploaded earlier THIS frame (frame
// serial guards arena recycling; the content version guards writes).
struct GlobalUboSliceMemo {
Uint64 programLifetimeId = 0;
Uint64 frameSerial = 0;
Uint32 uboContentVersion = 0;
VkBuffer buffer = VK_NULL_HANDLE;
VkDeviceSize offset = 0;
VkDeviceSize range = 0;
};
static constexpr Uint32 kGlobalUboMemoSize = 4;
GlobalUboSliceMemo m_globalUboMemo[kGlobalUboMemoSize];
Uint32 m_globalUboMemoNext = 0;
// Per-binding fast path over VkSamplerManager's content-hashed sampler cache, which // Per-binding fast path over VkSamplerManager's content-hashed sampler cache, which
// stays the source of truth: its key hashes all sampler+texture state, so two distinct // stays the source of truth: its key hashes all sampler+texture state, so two distinct
// sampler objects with identical state still resolve to one VkSampler. This memo only // sampler objects with identical state still resolve to one VkSampler. This memo only
@@ -172,6 +233,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 samplerLifetimeId = 0; Uint64 samplerLifetimeId = 0;
Uint64 textureLifetimeId = 0; Uint64 textureLifetimeId = 0;
VkSampler sampler = VK_NULL_HANDLE; VkSampler sampler = VK_NULL_HANDLE;
Uint32 viewLevelCount = 0;
Uint16 samplerVersion = 0; Uint16 samplerVersion = 0;
Uint16 textureParamsVersion = 0; Uint16 textureParamsVersion = 0;
Bool forceNearestFiltering = false; Bool forceNearestFiltering = false;
@@ -32,6 +32,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
XXHASH_VERIFY(XXH64_update(m_hashState, &attr.IsBgra, sizeof(attr.IsBgra))); XXHASH_VERIFY(XXH64_update(m_hashState, &attr.IsBgra, sizeof(attr.IsBgra)));
XXHASH_VERIFY(XXH64_update(m_hashState, &attr.Divisor, sizeof(attr.Divisor))); XXHASH_VERIFY(XXH64_update(m_hashState, &attr.Divisor, sizeof(attr.Divisor)));
// The buffer's heap address is an identity component of the key: a freed
// buffer's reused address can alias an old cache entry, but only under a
// byte-identical attribute layout - and the entry payload is a pure function
// of the hashed inputs, with the draw path re-resolving bindingBufferKeys
// against the live VAO attribute pointers, so an aliased hit returns exactly
// what a rebuild would. Address drift only grows the map; the OnFrameBoundary
// aging sweep bounds that.
const SizeT bufferKey = reinterpret_cast<SizeT>(attr.Buffer.get()); const SizeT bufferKey = reinterpret_cast<SizeT>(attr.Buffer.get());
XXHASH_VERIFY(XXH64_update(m_hashState, &bufferKey, sizeof(bufferKey))); XXHASH_VERIFY(XXH64_update(m_hashState, &bufferKey, sizeof(bufferKey)));
} }
@@ -51,14 +58,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const VertexInputStateFactory::BackendVertexInputState& VertexInputStateFactory::GetOrCreateVertexInputState( const VertexInputStateFactory::BackendVertexInputState& VertexInputStateFactory::GetOrCreateVertexInputState(
const MG_State::GLState::VertexArrayObject& vao) { const MG_State::GLState::VertexArrayObject& vao) {
return GetOrCreateVertexInputState(vao, GetOrComputeHash(vao)); // Per-draw fast path: the VAO carries a pointer to its resolved entry,
// valid while its config version and the cache's eviction epoch both
// match - no re-hash, no map lookup.
const void* memoState = nullptr;
Uint64 memoEpoch = 0;
if (vao.GetBackendStateMemo(memoState, memoEpoch) && memoEpoch == m_evictionEpoch) {
const auto* entry = static_cast<const BackendVertexInputState*>(memoState);
entry->lastUsedFrameBoundary = m_frameBoundaryCounter;
return *entry;
}
const BackendVertexInputState& entry = GetOrCreateVertexInputState(vao, GetOrComputeHash(vao));
vao.SetBackendStateMemo(&entry, m_evictionEpoch);
return entry;
} }
const VertexInputStateFactory::BackendVertexInputState& VertexInputStateFactory::GetOrCreateVertexInputState( const VertexInputStateFactory::BackendVertexInputState& VertexInputStateFactory::GetOrCreateVertexInputState(
const MG_State::GLState::VertexArrayObject& vao, HashType hash) { const MG_State::GLState::VertexArrayObject& vao, HashType hash) {
auto it = m_cache.find(hash); auto it = m_cache.find(hash);
if (it != m_cache.end()) { if (it != m_cache.end()) {
return it->second; it->second->lastUsedFrameBoundary = m_frameBoundaryCounter;
return *it->second;
} }
VertexInputStateBuilder builder; VertexInputStateBuilder builder;
@@ -164,10 +184,37 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const auto& state = builder.Build(); const auto& state = builder.Build();
auto& entry = m_cache[hash]; auto& slot = m_cache[hash];
if (!slot) {
slot = MakeUnique<BackendVertexInputState>();
}
BackendVertexInputState& entry = *slot;
entry.hash = hash; entry.hash = hash;
entry.lastUsedFrameBoundary = m_frameBoundaryCounter;
entry.bindings = builder.GetBindings(); entry.bindings = builder.GetBindings();
entry.attributes = builder.GetAttributes(); entry.attributes = builder.GetAttributes();
// See the layoutHash declaration: hash only the resolved layout, never
// buffer identities, so identical layouts across VAOs/buffers agree.
XXHASH_VERIFY(XXH64_reset(m_hashState, 0));
for (const auto& binding : entry.bindings) {
XXHASH_VERIFY(XXH64_update(m_hashState, &binding.binding, sizeof(binding.binding)));
XXHASH_VERIFY(XXH64_update(m_hashState, &binding.stride, sizeof(binding.stride)));
XXHASH_VERIFY(XXH64_update(m_hashState, &binding.inputRate, sizeof(binding.inputRate)));
}
for (const auto& attribute : entry.attributes) {
XXHASH_VERIFY(XXH64_update(m_hashState, &attribute.location, sizeof(attribute.location)));
XXHASH_VERIFY(XXH64_update(m_hashState, &attribute.binding, sizeof(attribute.binding)));
XXHASH_VERIFY(XXH64_update(m_hashState, &attribute.format, sizeof(attribute.format)));
XXHASH_VERIFY(XXH64_update(m_hashState, &attribute.offset, sizeof(attribute.offset)));
}
XXHASH_VERIFY(XXH64_update(m_hashState, &unsupportedAttribMask, sizeof(unsupportedAttribMask)));
entry.layoutHash = XXH64_digest(m_hashState);
entry.attributeLocationMask = 0;
for (const auto& attribute : entry.attributes) {
if (attribute.location < 32u) {
entry.attributeLocationMask |= (1u << attribute.location);
}
}
entry.bindingBufferKeys = std::move(bindingBufferKeys); entry.bindingBufferKeys = std::move(bindingBufferKeys);
entry.bindingBaseOffsets = std::move(bindingBaseOffsets); entry.bindingBaseOffsets = std::move(bindingBaseOffsets);
entry.bindingAttributeLocations = std::move(bindingAttributeLocations); entry.bindingAttributeLocations = std::move(bindingAttributeLocations);
@@ -180,6 +227,33 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return entry; return entry;
} }
void VertexInputStateFactory::OnFrameBoundary() {
++m_frameBoundaryCounter;
// Sweep occasionally; evict entries whose last hit is far in the past.
// Erasure happens only here, never mid-frame: the draw path holds a
// reference into the current entry across its setup, and unordered_map
// erase would invalidate it. Entries are CPU-side only, so no GPU-idle
// proof is needed; an evicted entry that is used again is simply rebuilt
// from the VAO state (same hash, same content).
constexpr Uint64 kSweepInterval = 256;
constexpr Uint64 kRetireAgeBoundaries = 1024;
if ((m_frameBoundaryCounter % kSweepInterval) != 0) {
return;
}
for (auto it = m_cache.begin(); it != m_cache.end();) {
if (m_frameBoundaryCounter - it->second->lastUsedFrameBoundary > kRetireAgeBoundaries) {
it = m_cache.erase(it);
// Invalidate every VAO's state-pointer memo: the erased node's
// address may be reused by a future insert.
++m_evictionEpoch;
} else {
++it;
}
}
}
VkFormat VertexInputStateFactory::ToVkVertexFormat(DataType type, Int size, Bool normalized, Bool isInteger, VkFormat VertexInputStateFactory::ToVkVertexFormat(DataType type, Int size, Bool normalized, Bool isInteger,
Bool isBgra) { Bool isBgra) {
if (isBgra) { if (isBgra) {
@@ -27,6 +27,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
struct BackendVertexInputState { struct BackendVertexInputState {
HashType hash = 0; HashType hash = 0;
// Hash of the resolved Vulkan vertex layout only (bindings, attributes,
// unsupported mask) - NO buffer identities. `hash` mixes buffer heap
// addresses so per-chunk VBOs mint a fresh identity per buffer; keying
// pipelines on that minted one VkPipeline per chunk section for an
// identical layout, defeating pipeline reuse and the per-draw memo.
// Pipelines depend only on the layout, so they key on this instead.
HashType layoutHash = 0;
// Frame boundary of the last cache hit; entries idle past the
// OnFrameBoundary retirement age are evicted (CPU heap only).
// Mutable: the VAO's state-pointer memo fast path stamps it through
// a const entry reference.
mutable Uint64 lastUsedFrameBoundary = 0;
Vector<VkVertexInputBindingDescription> bindings; Vector<VkVertexInputBindingDescription> bindings;
Vector<VkVertexInputAttributeDescription> attributes; Vector<VkVertexInputAttributeDescription> attributes;
Vector<SizeT> bindingBufferKeys; Vector<SizeT> bindingBufferKeys;
@@ -38,6 +50,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// absent from `attributes`, so without this mask the draw path cannot tell them apart from // absent from `attributes`, so without this mask the draw path cannot tell them apart from
// a genuinely disabled array and would silently feed the shader the current attribute value. // a genuinely disabled array and would silently feed the shader the current attribute value.
Uint32 unsupportedAttribMask = 0; Uint32 unsupportedAttribMask = 0;
// Bitmask of `attributes[i].location` - the draw path needs it up to
// three times per draw, so it is baked once at build time.
Uint32 attributeLocationMask = 0;
VkPipelineVertexInputStateCreateInfo state{ VkPipelineVertexInputStateCreateInfo state{
VK_STRUCTURE_TYPE_PIPELINE_VERTEX_INPUT_STATE_CREATE_INFO VK_STRUCTURE_TYPE_PIPELINE_VERTEX_INPUT_STATE_CREATE_INFO
}; };
@@ -55,6 +70,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const BackendVertexInputState& GetOrCreateVertexInputState( const BackendVertexInputState& GetOrCreateVertexInputState(
const MG_State::GLState::VertexArrayObject& vao, HashType hash); const MG_State::GLState::VertexArrayObject& vao, HashType hash);
const BackendVertexInputState& GetOrCreateVertexInputState(const MG_State::GLState::VertexArrayObject& vao); const BackendVertexInputState& GetOrCreateVertexInputState(const MG_State::GLState::VertexArrayObject& vao);
// Frame boundary hook: ages the cache and evicts entries not hit for many
// frames. The key mixes buffer heap addresses, so buffer/VAO churn keeps
// minting fresh keys; without eviction the map grows for the whole session.
// Entries hold no Vulkan handles (pipeline creation copies the descriptions)
// and the draw path's entry reference never spans a frame boundary, so
// eviction here needs no GPU-idle proof. Self-gated: one counter bump and
// compare except on sweep boundaries.
void OnFrameBoundary();
static SizeT GetComponentSize(DataType type); static SizeT GetComponentSize(DataType type);
// Tightly-packed byte size of one vertex element for this attribute: componentSize * size for // Tightly-packed byte size of one vertex element for this attribute: componentSize * size for
// normal types, and 4 (one packed word) for the 2_10_10_10 types and GL_BGRA. Returns 0 for // normal types, and 4 (one packed word) for the 2_10_10_10 types and GL_BGRA. Returns 0 for
@@ -69,7 +92,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const VulkanRendererConfig& m_config; const VulkanRendererConfig& m_config;
VkPhysicalDevice m_physicalDevice = VK_NULL_HANDLE; VkPhysicalDevice m_physicalDevice = VK_NULL_HANDLE;
UnorderedMap<HashType, BackendVertexInputState> m_cache; // Values are heap-allocated: FastSTL::unordered_map is open-addressing,
// so INSERT invalidates references to stored values. The draw path (and
// the VAOs' state-pointer memos) hold entry pointers across inserts;
// only the unique_ptr cell moves, never the pointee.
UnorderedMap<HashType, UniquePtr<BackendVertexInputState>> m_cache;
// Monotonic frame-boundary counter (bumped in OnFrameBoundary) for cache aging.
Uint64 m_frameBoundaryCounter = 0;
// Bumped whenever any cache entry is erased. VAOs memo a raw pointer to
// their heap-allocated entry (stable across map insert/rehash by
// construction); a memo is honored only while its recorded epoch
// matches, so an evicted entry can never be dereferenced through a
// stale memo.
Uint64 m_evictionEpoch = 1;
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -141,6 +141,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_transientUploadArena.BeginFrame(frameIndex); m_transientUploadArena.BeginFrame(frameIndex);
} }
void VkBufferManager::CollectAllDeferredReleases() {
for (Uint32 frameIndex = 0; frameIndex < m_deferredBufferReleases.size(); ++frameIndex) {
CollectDeferredReleases(frameIndex);
}
for (Uint32 frameIndex = 0; frameIndex < m_transientUploadArena.GetFrameCount(); ++frameIndex) {
m_transientUploadArena.CollectDeferredReleases(frameIndex);
}
}
void VkBufferManager::NotifyDeviceIdle() { void VkBufferManager::NotifyDeviceIdle() {
// Everything submitted so far has completed. Work recorded for the // Everything submitted so far has completed. Work recorded for the
// current frame has not been submitted yet, so the current serial // current frame has not been submitted yet, so the current serial
@@ -77,6 +77,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// Recreate all per-frame transient arenas // Recreate all per-frame transient arenas
Bool RecreateTransientArenas(Uint32 frameCount); Bool RecreateTransientArenas(Uint32 frameCount);
void BeginFrame(Uint32 frameIndex); void BeginFrame(Uint32 frameIndex);
// Drains every frame slot's deferred buffer/resource releases (and the
// transient arena's parked superseded blocks). Only valid when the
// caller has proven every queue submission complete; used by the
// present-less frame-boundary drain.
void CollectAllDeferredReleases();
// All previously submitted GPU work has completed (vkDeviceWaitIdle). // All previously submitted GPU work has completed (vkDeviceWaitIdle).
void NotifyDeviceIdle(); void NotifyDeviceIdle();
// A frame slot's submission fence has been waited: every serial up to // A frame slot's submission fence has been waited: every serial up to
@@ -93,6 +93,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
m_pendingClears.clear(); m_pendingClears.clear();
m_aliveObjects.clear(); m_aliveObjects.clear();
m_pendingCount.store(static_cast<Uint32>(m_pendingClears.size()), std::memory_order_relaxed);
} }
TextureIdentity VkClearManager::MakeTextureIdentity(MG_State::GLState::ITextureObject* texture) { TextureIdentity VkClearManager::MakeTextureIdentity(MG_State::GLState::ITextureObject* texture) {
@@ -127,6 +128,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_pendingClears.erase(key); m_pendingClears.erase(key);
} }
m_aliveObjects.erase(identity); m_aliveObjects.erase(identity);
m_pendingCount.store(static_cast<Uint32>(m_pendingClears.size()), std::memory_order_relaxed);
} }
Bool VkClearManager::LockTextureIdentityLocked(const TextureIdentity& identity, Bool VkClearManager::LockTextureIdentityLocked(const TextureIdentity& identity,
@@ -221,6 +223,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_aliveObjects[MakeTextureIdentity(texture.get())] = texture; m_aliveObjects[MakeTextureIdentity(texture.get())] = texture;
auto& pending = m_pendingClears[key]; auto& pending = m_pendingClears[key];
MergeClearPayload(pending, clearPayload); MergeClearPayload(pending, clearPayload);
m_pendingCount.store(static_cast<Uint32>(m_pendingClears.size()), std::memory_order_relaxed);
} }
void VkClearManager::QueueClear(const ClearAttachmentPayload& clearPayload, void VkClearManager::QueueClear(const ClearAttachmentPayload& clearPayload,
@@ -238,6 +241,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_aliveObjects[MakeTextureIdentity(texture.get())] = texture; m_aliveObjects[MakeTextureIdentity(texture.get())] = texture;
auto& pending = m_pendingClears[key]; auto& pending = m_pendingClears[key];
MergeClearPayload(pending, clearPayload); MergeClearPayload(pending, clearPayload);
m_pendingCount.store(static_cast<Uint32>(m_pendingClears.size()), std::memory_order_relaxed);
} }
Bool VkClearManager::HasPendingClear(MG_State::GLState::ITextureObject* texture) { Bool VkClearManager::HasPendingClear(MG_State::GLState::ITextureObject* texture) {
@@ -245,6 +249,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return false; return false;
} }
if (m_pendingCount.load(std::memory_order_relaxed) == 0) {
return false; // per-draw hot path: nothing pending anywhere
}
const Uint64 lifetimeId = texture->GetLifetimeId(); const Uint64 lifetimeId = texture->GetLifetimeId();
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
for (auto it = m_pendingClears.begin(); it != m_pendingClears.end(); ++it) { for (auto it = m_pendingClears.begin(); it != m_pendingClears.end(); ++it) {
@@ -260,6 +268,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (key.texture == nullptr) { if (key.texture == nullptr) {
return false; return false;
} }
if (m_pendingCount.load(std::memory_order_relaxed) == 0) {
return false; // per-draw hot path: nothing pending anywhere
}
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
if (m_pendingClears.find(key) == m_pendingClears.end()) { if (m_pendingClears.find(key) == m_pendingClears.end()) {
@@ -287,6 +298,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (key.texture == nullptr) { if (key.texture == nullptr) {
return false; return false;
} }
if (m_pendingCount.load(std::memory_order_relaxed) == 0) {
return false; // per-draw hot path: nothing pending anywhere
}
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
if (!LockTextureLocked(key, outTexture)) { if (!LockTextureLocked(key, outTexture)) {
@@ -325,6 +339,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (texture == nullptr) { if (texture == nullptr) {
return false; return false;
} }
if (m_pendingCount.load(std::memory_order_relaxed) == 0) {
return false; // per-draw hot path: nothing pending anywhere
}
const Uint64 lifetimeId = texture->GetLifetimeId(); const Uint64 lifetimeId = texture->GetLifetimeId();
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
@@ -345,6 +362,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return; return;
} }
if (m_pendingCount.load(std::memory_order_relaxed) == 0) {
return; // per-draw hot path: nothing pending anywhere
}
const TextureIdentity identity = MakeTextureIdentity(texture); const TextureIdentity identity = MakeTextureIdentity(texture);
MGLOG_D("%s: Pop all pending clears for texture %d", __func__, texture->GetExternalIndex()); MGLOG_D("%s: Pop all pending clears for texture %d", __func__, texture->GetExternalIndex());
const std::lock_guard<std::mutex> lock(m_mutex); const std::lock_guard<std::mutex> lock(m_mutex);
@@ -361,6 +381,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
auto it = m_pendingClears.find(key); auto it = m_pendingClears.find(key);
if (it != m_pendingClears.end()) { if (it != m_pendingClears.end()) {
m_pendingClears.erase(it); m_pendingClears.erase(it);
m_pendingCount.store(static_cast<Uint32>(m_pendingClears.size()), std::memory_order_relaxed);
} }
} }
@@ -14,6 +14,7 @@
#include "MG_Util/Math/VectorTypes.h" #include "MG_Util/Math/VectorTypes.h"
#include <Includes.h> #include <Includes.h>
#include <atomic>
#include <unordered_map> #include <unordered_map>
namespace MobileGL::MG_Backend::DirectVulkan { namespace MobileGL::MG_Backend::DirectVulkan {
@@ -120,7 +121,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
SharedPtr<MG_State::GLState::ITextureObject>& outTexture); SharedPtr<MG_State::GLState::ITextureObject>& outTexture);
Uint8 m_gcCounter = 0; Uint8 m_gcCounter = 0;
public:
// Lock-free probe for the consecutive-draw fast path: any pending clear
// forces the full SetupDraw path (which materializes/consumes it).
Bool HasAnyPendingClears() const { return m_pendingCount.load(std::memory_order_relaxed) != 0; }
private:
mutable std::mutex m_mutex; mutable std::mutex m_mutex;
// Lock-free mirror of m_pendingClears.size(), maintained under m_mutex
// by every mutation. The per-draw probes (HasPendingClear/GetPending*)
// read it before taking the lock: during draw batches the pending set
// is almost always empty, so this turns several locked map probes per
// draw into one relaxed load.
std::atomic<Uint32> m_pendingCount{0};
std::unordered_map<PendingClearKey, ClearAttachmentPayload, PendingClearKeyHash> m_pendingClears; std::unordered_map<PendingClearKey, ClearAttachmentPayload, PendingClearKeyHash> m_pendingClears;
std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects; std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects;
}; };
@@ -180,6 +180,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
sampleCount = VK_SAMPLE_COUNT_1_BIT; sampleCount = VK_SAMPLE_COUNT_1_BIT;
internalFormat = TextureInternalFormat::Unknown; internalFormat = TextureInternalFormat::Unknown;
samples = 0; samples = 0;
deadSinceFrame = kNeverObservedDead;
} }
VkRenderPassManager::VkRenderPassManager(VkDevice device, VkRenderPassManager::VkRenderPassManager(VkDevice device,
@@ -206,6 +207,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource.Destroy(m_device, m_allocator); resource.Destroy(m_device, m_allocator);
} }
m_renderbufferResources.clear(); m_renderbufferResources.clear();
CollectDeferredRenderbufferReleases(/*destroyAll=*/true); // caller guarantees device idle
m_pendingRenderbufferClears.clear(); m_pendingRenderbufferClears.clear();
RenderPassEntry::s_textureResourcesScratch.clear(); RenderPassEntry::s_textureResourcesScratch.clear();
s_activeRenderPass = {}; s_activeRenderPass = {};
@@ -213,22 +215,75 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_rpFastValid = false; m_rpFastValid = false;
} }
Uint64 VkRenderPassManager::RetireAgeFrames() const {
// MaxFramesInFlight + 2 covers the frame ring plus one boundary for the
// recording-to-submit gap and one because OnPresent runs ahead of Present's
// fence wait; the floor of 8 keeps a margin over the default ring of 3 while
// still releasing multi-MB attachment memory promptly (the render-pass cache's
// 1024-frame retirement would pin it for no additional safety).
return std::max<Uint64>(8, static_cast<Uint64>(m_config.MaxFramesInFlight) + 2);
}
void VkRenderPassManager::DeferRenderbufferBackingRelease(RenderbufferResource& resource) {
// The superseded backing may still be referenced by in-flight command buffers
// (glRenderbufferStorage can respecify a renderbuffer drawn this very frame),
// so it is parked and destroyed only after RetireAgeFrames() boundaries.
if (resource.image == VK_NULL_HANDLE && resource.view == VK_NULL_HANDLE) {
return;
}
m_deferredRenderbufferReleases.push_back({resource.image, resource.allocation, resource.view, m_frameCounter});
resource.image = VK_NULL_HANDLE;
resource.allocation = nullptr;
resource.view = VK_NULL_HANDLE;
}
void VkRenderPassManager::CollectDeferredRenderbufferReleases(Bool destroyAll) {
if (m_deferredRenderbufferReleases.empty()) {
return;
}
const Uint64 retireAgeFrames = RetireAgeFrames();
std::erase_if(m_deferredRenderbufferReleases, [&](DeferredRenderbufferRelease& release) {
if (!destroyAll && m_frameCounter - release.deferredAtFrame < retireAgeFrames) {
return false;
}
if (release.view != VK_NULL_HANDLE) {
vkDestroyImageView(m_device, release.view, nullptr);
}
if (release.image != VK_NULL_HANDLE) {
vmaDestroyImage(m_allocator, release.image, release.allocation);
}
return true;
});
}
void VkRenderPassManager::CollectRenderbufferGarbage() { void VkRenderPassManager::CollectRenderbufferGarbage() {
Vector<MG_State::GLState::RenderbufferObject*> deadRenderbuffers; // Two-phase reclamation: a dead renderbuffer's VkImage may still be referenced by
deadRenderbuffers.reserve(m_renderbufferResources.size()); // command buffers submitted up to frames-in-flight frames ago (it was legally
for (auto& [renderbuffer, resource] : m_renderbufferResources) { // attached and drawn right up to its deletion), so the first observation of an
// expired weak reference only stamps the current frame counter; Destroy runs once
// enough frame boundaries have passed that the stamping frame's submission fence
// has provably been waited (see RetireAgeFrames).
const Uint64 retireAgeFrames = RetireAgeFrames();
for (auto it = m_renderbufferResources.begin(); it != m_renderbufferResources.end();) {
auto& resource = it->second;
const auto liveRenderbuffer = resource.renderbuffer.lock(); const auto liveRenderbuffer = resource.renderbuffer.lock();
if (!liveRenderbuffer || liveRenderbuffer.get() != renderbuffer) { if (liveRenderbuffer && liveRenderbuffer.get() == it->first) {
deadRenderbuffers.emplace_back(renderbuffer); resource.deadSinceFrame = RenderbufferResource::kNeverObservedDead;
++it;
continue;
} }
if (resource.deadSinceFrame == RenderbufferResource::kNeverObservedDead) {
resource.deadSinceFrame = m_frameCounter;
++it;
continue;
} }
for (auto* renderbuffer : deadRenderbuffers) { if (m_frameCounter - resource.deadSinceFrame < retireAgeFrames) {
auto resourceIt = m_renderbufferResources.find(renderbuffer); ++it;
if (resourceIt != m_renderbufferResources.end()) { continue;
resourceIt->second.Destroy(m_device, m_allocator);
m_renderbufferResources.erase(resourceIt);
} }
m_pendingRenderbufferClears.erase(renderbuffer); m_pendingRenderbufferClears.erase(it->first);
resource.Destroy(m_device, m_allocator);
it = m_renderbufferResources.erase(it);
} }
} }
@@ -251,11 +306,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const auto internalFormat = renderbuffer->GetInternalFormat(); const auto internalFormat = renderbuffer->GetInternalFormat();
const VkFormat format = MG_Util::ConvertTextureInternalFormatToVkEnum(internalFormat); const VkFormat format = MG_Util::ConvertTextureInternalFormatToVkEnum(internalFormat);
const VkImageAspectFlags aspect = ResolveImageAspectMaskForFormat(format); const VkImageAspectFlags aspect = ResolveImageAspectMaskForFormat(format);
if ((aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0) { // Renderbuffers are never sampled (GL has no way to bind one to a sampler), so the
MGLOG_E("GetOrCreateRenderbufferResource: color renderbuffer %u is not supported by DirectVulkan render passes yet", // usage set is attachment + transfer: transfer covers readback (vkCmdCopyImageToBuffer),
renderbuffer->GetExternalIndex()); // BlitFramebuffer, CopyTexImage sources, and out-of-render-pass clear materialization.
return nullptr; const VkImageUsageFlags imageUsage =
} ((aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 ? VK_IMAGE_USAGE_COLOR_ATTACHMENT_BIT
: VK_IMAGE_USAGE_DEPTH_STENCIL_ATTACHMENT_BIT) |
VK_IMAGE_USAGE_TRANSFER_SRC_BIT | VK_IMAGE_USAGE_TRANSFER_DST_BIT;
auto& resource = m_renderbufferResources[renderbuffer.get()]; auto& resource = m_renderbufferResources[renderbuffer.get()];
const Bool needsCreate = const Bool needsCreate =
@@ -268,9 +325,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource.samples != renderbuffer->GetSamples(); resource.samples != renderbuffer->GetSamples();
if (!needsCreate) { if (!needsCreate) {
resource.renderbuffer = renderbuffer; resource.renderbuffer = renderbuffer;
// A new renderbuffer at a recycled address may adopt a compatible entry that
// was already stamped dead; it is alive again, so cancel the aging.
resource.deadSinceFrame = RenderbufferResource::kNeverObservedDead;
return &resource; return &resource;
} }
// Respecify: park the old backing for aged destruction instead of destroying
// inline - it may still be referenced by in-flight command buffers.
DeferRenderbufferBackingRelease(resource);
resource.Destroy(m_device, m_allocator); resource.Destroy(m_device, m_allocator);
resource.renderbuffer = renderbuffer; resource.renderbuffer = renderbuffer;
@@ -285,7 +348,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
imageInfo.format = format; imageInfo.format = format;
imageInfo.tiling = VK_IMAGE_TILING_OPTIMAL; imageInfo.tiling = VK_IMAGE_TILING_OPTIMAL;
imageInfo.initialLayout = VK_IMAGE_LAYOUT_UNDEFINED; imageInfo.initialLayout = VK_IMAGE_LAYOUT_UNDEFINED;
imageInfo.usage = VK_IMAGE_USAGE_DEPTH_STENCIL_ATTACHMENT_BIT; imageInfo.usage = imageUsage;
imageInfo.samples = sampleCount; imageInfo.samples = sampleCount;
imageInfo.sharingMode = VK_SHARING_MODE_EXCLUSIVE; imageInfo.sharingMode = VK_SHARING_MODE_EXCLUSIVE;
@@ -386,6 +449,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void VkRenderPassManager::QueueRenderbufferClear( void VkRenderPassManager::QueueRenderbufferClear(
GLbitfield mask, const ClearFramebufferPayload& clearPayload, GLbitfield mask, const ClearFramebufferPayload& clearPayload,
const MG_State::GLState::FramebufferObject& drawFbo) { const MG_State::GLState::FramebufferObject& drawFbo) {
if ((mask & GL_COLOR_BUFFER_BIT) != 0) {
// Color renderbuffer draw buffers take the framebuffer-level clear too; texture
// attachments are skipped by the per-attachment overload's IsRenderbuffer guard.
for (const auto attachmentType : drawFbo.GetDrawBuffers()) {
if (attachmentType == FramebufferAttachmentType::None) {
continue;
}
QueueRenderbufferClear(
ClearAttachmentPayload{.mask = GL_COLOR_BUFFER_BIT, .color = clearPayload.color},
drawFbo.GetAttachment(attachmentType));
}
}
if ((mask & GL_DEPTH_BUFFER_BIT) != 0) { if ((mask & GL_DEPTH_BUFFER_BIT) != 0) {
QueueRenderbufferClear( QueueRenderbufferClear(
ClearAttachmentPayload{.mask = GL_DEPTH_BUFFER_BIT, .depth = clearPayload.depth}, ClearAttachmentPayload{.mask = GL_DEPTH_BUFFER_BIT, .depth = clearPayload.depth},
@@ -406,7 +481,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
VkRenderPassManager::HashType VkRenderPassManager::ComputeHash( VkRenderPassManager::HashType VkRenderPassManager::ComputeHash(
const MG_State::GLState::FramebufferObject& fbo, Uint32 swapchainImageIndex, Bool includePendingClear) { const MG_State::GLState::FramebufferObject& fbo, Uint32 swapchainImageIndex, Bool includePendingClear,
Bool includeDefaultFboDepthStencil) {
XXHASH_VERIFY(XXH64_reset(m_hashState, m_config.CacheVersion)); XXHASH_VERIFY(XXH64_reset(m_hashState, m_config.CacheVersion));
const Bool isDefaultFbo = fbo.IsDefaultFramebuffer(); const Bool isDefaultFbo = fbo.IsDefaultFramebuffer();
if (isDefaultFbo) { if (isDefaultFbo) {
@@ -485,9 +561,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
attachment <= FramebufferAttachmentType::BackRight); attachment <= FramebufferAttachmentType::BackRight);
if (isDefaultColorAttachment) { if (isDefaultColorAttachment) {
currentLayout = m_swapchainObject.GetImageLayout(swapchainImageIndex); currentLayout = m_swapchainObject.GetImageLayout(swapchainImageIndex);
// Content validity feeds the attachment's loadOp (see the
// creation path), so it must key the cache as well.
if (!m_swapchainObject.IsImageContentDefined(swapchainImageIndex)) {
currentLayout = VK_IMAGE_LAYOUT_UNDEFINED;
}
} else if (attachment == FramebufferAttachmentType::Depth || } else if (attachment == FramebufferAttachmentType::Depth ||
attachment == FramebufferAttachmentType::Stencil) { attachment == FramebufferAttachmentType::Stencil) {
currentLayout = m_swapchainObject.GetDepthStencilImageLayout(swapchainImageIndex); currentLayout = m_swapchainObject.GetDepthStencilImageLayout(swapchainImageIndex);
if (!m_swapchainObject.IsDepthStencilContentDefined(swapchainImageIndex)) {
currentLayout = VK_IMAGE_LAYOUT_UNDEFINED;
}
} }
} else { } else {
auto* textureResource = m_textureManager.SyncTextureAndGetDescriptor(*texture); auto* textureResource = m_textureManager.SyncTextureAndGetDescriptor(*texture);
@@ -542,14 +626,49 @@ namespace MobileGL::MG_Backend::DirectVulkan {
combineFramebufferAttachmentObjHash(drawbuf); combineFramebufferAttachmentObjHash(drawbuf);
} }
// The depth-less default-FBO flavor omits the depth/stencil attachment
// entirely, so it must hash differently from the depth-full flavor.
const Bool depthStencilIncluded = !isDefaultFbo || includeDefaultFboDepthStencil;
XXHASH_VERIFY(XXH64_update(m_hashState, &depthStencilIncluded, sizeof(depthStencilIncluded)));
if (depthStencilIncluded) {
combineFramebufferAttachmentObjHash(FramebufferAttachmentType::Depth); combineFramebufferAttachmentObjHash(FramebufferAttachmentType::Depth);
combineFramebufferAttachmentObjHash(FramebufferAttachmentType::Stencil); combineFramebufferAttachmentObjHash(FramebufferAttachmentType::Stencil);
}
return XXH64_digest(m_hashState); return XXH64_digest(m_hashState);
} }
RenderPassEntry& VkRenderPassManager::GetOrCreateRenderPass(const MG_State::GLState::FramebufferObject& fbo, RenderPassEntry& VkRenderPassManager::GetOrCreateRenderPass(const MG_State::GLState::FramebufferObject& fbo,
Uint32 swapchainImageIndex) { Uint32 swapchainImageIndex,
Bool drawUsesDepthStencil) {
// Resolve the default-FBO depth flavor (see the header comment): keep the
// depth attachment when the caller needs it, when a depth/stencil clear is
// pending, or when the active pass already carries it (escalate-only, so
// alternating depth-less draws never split an established depth pass).
Bool includeDefaultFboDepthStencil = true;
if (fbo.IsDefaultFramebuffer()) {
Bool activeDefaultHasDepthStencil = false;
if (const auto* active = GetActiveRenderPass()) {
Bool activeIsSwapchainPass = false;
Bool activeHasSwapchainDepthStencil = false;
for (const auto& tracked : active->trackedAttachmentLayouts) {
activeIsSwapchainPass |= tracked.target == TrackedAttachmentTarget::SwapchainColor;
activeHasSwapchainDepthStencil |=
tracked.target == TrackedAttachmentTarget::SwapchainDepthStencil;
}
activeDefaultHasDepthStencil = activeIsSwapchainPass && activeHasSwapchainDepthStencil;
}
const auto& defaultDepthAtt = fbo.GetAttachment(FramebufferAttachmentType::Depth);
const auto& defaultStencilAtt = fbo.GetAttachment(FramebufferAttachmentType::Stencil);
const Bool pendingDepthStencilClear =
(defaultDepthAtt.IsTexture() && m_clearManager.HasPendingClear(defaultDepthAtt)) ||
HasPendingRenderbufferClear(defaultDepthAtt) ||
(defaultStencilAtt.IsTexture() && m_clearManager.HasPendingClear(defaultStencilAtt)) ||
HasPendingRenderbufferClear(defaultStencilAtt);
includeDefaultFboDepthStencil =
drawUsesDepthStencil || activeDefaultHasDepthStencil || pendingDepthStencilClear;
}
auto hasPendingClearOnFramebuffer = [&]() -> Bool { auto hasPendingClearOnFramebuffer = [&]() -> Bool {
const auto& drawBuffers = fbo.GetDrawBuffers(); const auto& drawBuffers = fbo.GetDrawBuffers();
for (auto attachment : drawBuffers) { for (auto attachment : drawBuffers) {
@@ -599,6 +718,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_rpFastFboVersion == fbo.GetObjectVersion() && m_rpFastSwapchainIndex == swapchainImageIndex && m_rpFastFboVersion == fbo.GetObjectVersion() && m_rpFastSwapchainIndex == swapchainImageIndex &&
m_rpFastTexEpoch == m_textureManager.GetTextureImageEpoch() && m_rpFastTexEpoch == m_textureManager.GetTextureImageEpoch() &&
m_rpFastRbEpoch == m_renderbufferImageEpoch && m_rpFastRbEpoch == m_renderbufferImageEpoch &&
(!fbo.IsDefaultFramebuffer() || m_rpFastHadDepthStencil == includeDefaultFboDepthStencil) &&
m_rpFastRenderPassHash == activeRenderPass->hash && !hasPendingClearOnFramebuffer()) { m_rpFastRenderPassHash == activeRenderPass->hash && !hasPendingClearOnFramebuffer()) {
auto activeIt = m_renderPasses.find(activeRenderPass->hash); auto activeIt = m_renderPasses.find(activeRenderPass->hash);
if (activeIt != m_renderPasses.end()) { if (activeIt != m_renderPasses.end()) {
@@ -607,7 +727,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} }
auto compatibilityHash = ComputeHash(fbo, swapchainImageIndex, false); auto compatibilityHash = ComputeHash(fbo, swapchainImageIndex, false, includeDefaultFboDepthStencil);
if (activeRenderPass != nullptr && if (activeRenderPass != nullptr &&
activeRenderPass->CompatibleWith(compatibilityHash) && activeRenderPass->CompatibleWith(compatibilityHash) &&
!hasPendingClearOnFramebuffer()) { !hasPendingClearOnFramebuffer()) {
@@ -624,10 +744,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_rpFastTexEpoch = m_textureManager.GetTextureImageEpoch(); m_rpFastTexEpoch = m_textureManager.GetTextureImageEpoch();
m_rpFastRbEpoch = m_renderbufferImageEpoch; m_rpFastRbEpoch = m_renderbufferImageEpoch;
m_rpFastRenderPassHash = activeRenderPass->hash; m_rpFastRenderPassHash = activeRenderPass->hash;
m_rpFastHadDepthStencil = activeIt->second.hasDepthStencilAttachment;
activeIt->second.lastUsedFrame = m_frameCounter; activeIt->second.lastUsedFrame = m_frameCounter;
return activeIt->second; return activeIt->second;
} }
auto hash = ComputeHash(fbo, swapchainImageIndex, true); auto hash = ComputeHash(fbo, swapchainImageIndex, true, includeDefaultFboDepthStencil);
auto it = m_renderPasses.find(hash); auto it = m_renderPasses.find(hash);
if (it != m_renderPasses.end()) { if (it != m_renderPasses.end()) {
it->second.lastUsedFrame = m_frameCounter; it->second.lastUsedFrame = m_frameCounter;
@@ -682,6 +803,83 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// assuming default FBO has the right param // assuming default FBO has the right param
for (Uint32 i = 0; i < colorAttachmentSlotCount; ++i) { for (Uint32 i = 0; i < colorAttachmentSlotCount; ++i) {
auto drawbuf = drawbufs[i]; auto drawbuf = drawbufs[i];
// Renderbuffer color attachments mirror the texture path below, with the
// resource (image/view/format/layout) coming from the render-pass manager's
// renderbuffer store instead of the texture manager.
if (drawbuf != FramebufferAttachmentType::None && !isDefaultFbo) {
const auto& rbAtt = fbo.GetAttachment(drawbuf);
if (rbAtt.IsRenderbuffer() && rbAtt.IsComplete()) {
const auto& renderbuffer = rbAtt.GetRenderbuffer();
auto* rbResource = GetOrCreateRenderbufferResource(renderbuffer);
if (rbResource == nullptr || (rbResource->aspect & VK_IMAGE_ASPECT_COLOR_BIT) == 0) {
MGLOG_E("GetOrCreateRenderPass: draw buffer slot %u on FBO %u has an unsupported color "
"renderbuffer %u; using VK_ATTACHMENT_UNUSED",
i, fbo.GetExternalIndex(), renderbuffer->GetExternalIndex());
continue;
}
const Uint32 rbAttachmentIndex = static_cast<Uint32>(attachmentDescriptions.size());
attachmentDescriptions.emplace_back();
VkAttachmentDescription& rbDesc = attachmentDescriptions.back();
ClearAttachmentPayload rbClearPayload{};
Bool rbHasClear = GetPendingRenderbufferClear(renderbuffer.get(), rbClearPayload) &&
(rbClearPayload.mask & GL_COLOR_BUFFER_BIT) != 0;
if (rbHasClear &&
MG_Util::GetBaseInternalFormatComponentCount(renderbuffer->GetInternalFormat()) == 3) {
// RGB renderbuffers are backed by an RGBA image; the missing alpha reads as 1.
rbClearPayload.color =
FloatVec4(rbClearPayload.color.x(), rbClearPayload.color.y(),
rbClearPayload.color.z(), 1.0f);
}
const VkImageLayout trackedRbLayout = rbResource->layout;
rbDesc.flags = 0;
rbDesc.format = rbResource->format;
rbDesc.samples = rbResource->sampleCount;
rbDesc.loadOp = rbHasClear ? VK_ATTACHMENT_LOAD_OP_CLEAR :
(trackedRbLayout == VK_IMAGE_LAYOUT_UNDEFINED ? VK_ATTACHMENT_LOAD_OP_DONT_CARE
: VK_ATTACHMENT_LOAD_OP_LOAD);
rbDesc.storeOp = VK_ATTACHMENT_STORE_OP_STORE;
rbDesc.stencilLoadOp = VK_ATTACHMENT_LOAD_OP_DONT_CARE;
rbDesc.stencilStoreOp = VK_ATTACHMENT_STORE_OP_DONT_CARE;
rbDesc.initialLayout = (rbHasClear || trackedRbLayout == VK_IMAGE_LAYOUT_UNDEFINED) ?
VK_IMAGE_LAYOUT_UNDEFINED : trackedRbLayout;
rbDesc.finalLayout = VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL;
adoptRenderPassSampleCount(rbResource->sampleCount, "color",
static_cast<Int>(renderbuffer->GetExternalIndex()));
if (rbHasClear) {
pendingClearAttachments.emplace_back(PendingClearAttachmentInfo {
.attachmentIndex = rbAttachmentIndex,
.colorAttachmentSlot = i,
.renderbuffer = renderbuffer.get(),
.hasInlinePayload = true,
.inlinePayload = rbClearPayload,
});
}
if (width == 0)
width = static_cast<Int>(rbResource->extent.width);
if (height == 0)
height = static_cast<Int>(rbResource->extent.height);
trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo {
.target = TrackedAttachmentTarget::Renderbuffer,
.renderbuffer = renderbuffer,
.finalLayout = rbDesc.finalLayout,
});
textureResources.emplace_back(nullptr);
attachmentViews.emplace_back(rbResource->view);
MOBILEGL_ASSERT(attachmentViews.back() != VK_NULL_HANDLE,
"GetOrCreateRenderPass: renderbuffer view missing at color attachment %d", i);
colorAttachmentRefs[i].attachment = rbAttachmentIndex;
continue;
}
}
auto* texture = ResolveCompleteColorAttachmentTexture(fbo, drawbuf, i); auto* texture = ResolveCompleteColorAttachmentTexture(fbo, drawbuf, i);
if (texture == nullptr) if (texture == nullptr)
continue; continue;
@@ -700,6 +898,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
case TextureTarget::Texture2D: case TextureTarget::Texture2D:
case TextureTarget::Texture2DArray: case TextureTarget::Texture2DArray:
case TextureTarget::Texture2DMultisample: case TextureTarget::Texture2DMultisample:
case TextureTarget::Texture2DMultisampleArray:
case TextureTarget::Texture3D:
case TextureTarget::TextureCubeMap:
case TextureTarget::TextureCubeMapArray:
case TextureTarget::TextureRectangle: { case TextureTarget::TextureRectangle: {
desc.flags = 0; desc.flags = 0;
desc.format = isDefaultFbo ? desc.format = isDefaultFbo ?
@@ -738,6 +940,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
MOBILEGL_ASSERT(swapchainImageIndex < swapchainViews.size(), MOBILEGL_ASSERT(swapchainImageIndex < swapchainViews.size(),
"GetOrCreateRenderPass: swapchain image index out of range"); "GetOrCreateRenderPass: swapchain image index out of range");
trackedColorLayout = m_swapchainObject.GetImageLayout(swapchainImageIndex); trackedColorLayout = m_swapchainObject.GetImageLayout(swapchainImageIndex);
// EGL: a presented color buffer's content is undefined when its
// image comes back around (EGL_BUFFER_DESTROYED, the default
// swap behaviour) - skip the tile load instead of reloading
// stale pixels nobody may rely on.
if (!hasClear && !m_swapchainObject.IsImageContentDefined(swapchainImageIndex)) {
trackedColorLayout = VK_IMAGE_LAYOUT_UNDEFINED;
}
trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo { trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo {
.target = TrackedAttachmentTarget::SwapchainColor, .target = TrackedAttachmentTarget::SwapchainColor,
.swapchainImageIndex = swapchainImageIndex, .swapchainImageIndex = swapchainImageIndex,
@@ -757,6 +966,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo { trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo {
.target = TrackedAttachmentTarget::Texture, .target = TrackedAttachmentTarget::Texture,
.texture = att.GetTexture(), .texture = att.GetTexture(),
.textureRaw = att.GetTexture().get(),
.textureMipLevel = attachmentMipLevel, .textureMipLevel = attachmentMipLevel,
.finalLayout = desc.finalLayout, .finalLayout = desc.finalLayout,
}); });
@@ -820,6 +1030,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}; };
const auto* selectedDepthStencilAttachment = isUsableDepthStencilAttachment(depthAtt) ? &depthAtt : const auto* selectedDepthStencilAttachment = isUsableDepthStencilAttachment(depthAtt) ? &depthAtt :
(isUsableDepthStencilAttachment(stencilAtt) ? &stencilAtt : nullptr); (isUsableDepthStencilAttachment(stencilAtt) ? &stencilAtt : nullptr);
// Depth-less default-FBO flavor: nothing in this pass touches depth/stencil
// and their content is undefined anyway (EGL swap), so drop the attachment
// and its whole tile load + store.
if (isDefaultFbo && !includeDefaultFboDepthStencil) {
selectedDepthStencilAttachment = nullptr;
}
const Bool hasDistinctDepthAndStencilAttachments = const Bool hasDistinctDepthAndStencilAttachments =
isUsableDepthStencilAttachment(depthAtt) && isUsableDepthStencilAttachment(stencilAtt) && isUsableDepthStencilAttachment(depthAtt) && isUsableDepthStencilAttachment(stencilAtt) &&
!sameDepthStencilAttachmentObject(depthAtt, stencilAtt); !sameDepthStencilAttachmentObject(depthAtt, stencilAtt);
@@ -838,6 +1054,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkImageLayout trackedDepthLayout = isDefaultFbo ? VkImageLayout trackedDepthLayout = isDefaultFbo ?
m_swapchainObject.GetDepthStencilImageLayout(swapchainImageIndex) : m_swapchainObject.GetDepthStencilImageLayout(swapchainImageIndex) :
VK_IMAGE_LAYOUT_DEPTH_STENCIL_ATTACHMENT_OPTIMAL; VK_IMAGE_LAYOUT_DEPTH_STENCIL_ATTACHMENT_OPTIMAL;
// EGL 1.5 §3.10.1: every ancillary (depth/stencil) buffer's content is
// undefined after a swap, so the first default-FBO pass of a frame can
// skip the depth/stencil tile load outright.
if (isDefaultFbo && !m_swapchainObject.IsDepthStencilContentDefined(swapchainImageIndex)) {
trackedDepthLayout = VK_IMAGE_LAYOUT_UNDEFINED;
}
depthAttachmentDescription.flags = 0; depthAttachmentDescription.flags = 0;
VkSampleCountFlagBits depthAttachmentSampleCount = VK_SAMPLE_COUNT_1_BIT; VkSampleCountFlagBits depthAttachmentSampleCount = VK_SAMPLE_COUNT_1_BIT;
Int depthAttachmentId = 0; Int depthAttachmentId = 0;
@@ -921,6 +1143,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo { trackedAttachmentLayouts.emplace_back(TrackedAttachmentLayoutInfo {
.target = TrackedAttachmentTarget::Texture, .target = TrackedAttachmentTarget::Texture,
.texture = selectedDepthStencilAttachment->GetTexture(), .texture = selectedDepthStencilAttachment->GetTexture(),
.textureRaw = selectedDepthStencilAttachment->GetTexture().get(),
.textureMipLevel = attachmentMipLevel, .textureMipLevel = attachmentMipLevel,
.finalLayout = depthAttachmentDescription.finalLayout, .finalLayout = depthAttachmentDescription.finalLayout,
}); });
@@ -966,6 +1189,22 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
const Bool hasDepthStencilAttachment = depthAttachmentRef.attachment != VK_ATTACHMENT_UNUSED; const Bool hasDepthStencilAttachment = depthAttachmentRef.attachment != VK_ATTACHMENT_UNUSED;
// Declare only the used colour-reference span. The GL draw-buffer array
// always spans 8 slots, so passes used to declare colorAttachmentCount=8
// with trailing VK_ATTACHMENT_UNUSED holes - and Adreno configures its
// per-pixel render-backend/export path from the DECLARED count, so every
// fragment of every pass paid the 8-target export cost (measured on
// Adreno 650 / MC 26.2: 11.9 -> 7.5 ms of GPU time per frame, with the
// single-quad swapchain blit pass alone dropping 1.26 -> 0.40 ms).
// Interior GL_NONE holes keep their slots so fragment-output locations
// still line up; a fragment output at a location past the trimmed count
// is discarded, which is exactly GL's semantic for writing to a draw
// buffer set to GL_NONE.
while (!colorAttachmentRefs.empty() &&
colorAttachmentRefs.back().attachment == VK_ATTACHMENT_UNUSED) {
colorAttachmentRefs.pop_back();
}
// Subpass // Subpass
VkSubpassDescription subpassDesc; VkSubpassDescription subpassDesc;
subpassDesc.flags = 0; subpassDesc.flags = 0;
@@ -1083,6 +1322,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void VkRenderPassManager::OnPresent() { void VkRenderPassManager::OnPresent() {
++m_frameCounter; ++m_frameCounter;
// Runs every frame boundary, ahead of the render-pass sweep gate below: the walk
// is O(#renderbuffer resources) — single digits in practice — and per-frame
// invocation keeps dead-resource reclaim latency at the aging bound instead of
// coupling it to renderbuffer *use* (the GetOrCreateRenderbufferResource call
// site never runs again once an app stops using renderbuffers).
CollectRenderbufferGarbage();
CollectDeferredRenderbufferReleases(/*destroyAll=*/false);
// Sweep occasionally; evict entries whose last use is far past every // Sweep occasionally; evict entries whose last use is far past every
// in-flight frame so their VkRenderPass/VkFramebuffer can be destroyed // in-flight frame so their VkRenderPass/VkFramebuffer can be destroyed
// safely (RenderPassEntry's destructor releases the handles). // safely (RenderPassEntry's destructor releases the handles).
@@ -1092,6 +1339,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return; return;
} }
// Collect the dying handles and notify once after the loop: pipelines hashed
// on them share the entries' >kRetireAgeFrames idleness (they are only bound
// by draws that hit those entries), so the observer may destroy them
// immediately - and a single batched notification costs one pipeline-cache
// scan instead of one per evicted pass.
Vector<VkRenderPass> destroyedRenderPasses;
const Uint64 activeHash = s_hasActiveRenderPass ? s_activeRenderPass.hash : 0; const Uint64 activeHash = s_hasActiveRenderPass ? s_activeRenderPass.hash : 0;
for (auto it = m_renderPasses.begin(); it != m_renderPasses.end();) { for (auto it = m_renderPasses.begin(); it != m_renderPasses.end();) {
const Bool isActive = s_hasActiveRenderPass && it->first == activeHash; const Bool isActive = s_hasActiveRenderPass && it->first == activeHash;
@@ -1099,11 +1352,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (m_rpFastValid && m_rpFastRenderPassHash == it->first) { if (m_rpFastValid && m_rpFastRenderPassHash == it->first) {
m_rpFastValid = false; m_rpFastValid = false;
} }
destroyedRenderPasses.push_back(it->second.renderPass);
it = m_renderPasses.erase(it); it = m_renderPasses.erase(it);
} else { } else {
++it; ++it;
} }
} }
if (!destroyedRenderPasses.empty() && m_evictionObserver != nullptr) {
m_evictionObserver->OnRenderPassesDestroyed(destroyedRenderPasses);
}
} }
Bool VkRenderPassManager::BeginRenderPass(VkCommandBuffer commandBuffer, RenderPassEntry& renderPassEntry) { Bool VkRenderPassManager::BeginRenderPass(VkCommandBuffer commandBuffer, RenderPassEntry& renderPassEntry) {
@@ -1156,6 +1413,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
renderPassBeginInfo.pClearValues = clearValues.data(); renderPassBeginInfo.pClearValues = clearValues.data();
vkCmdBeginRenderPass(commandBuffer, &renderPassBeginInfo, VK_SUBPASS_CONTENTS_INLINE); vkCmdBeginRenderPass(commandBuffer, &renderPassBeginInfo, VK_SUBPASS_CONTENTS_INLINE);
// Pre-pass stream bookkeeping: this pass's attachment images are now
// referenced by the open frame recording.
if (s_textureManager != nullptr) {
for (const auto& tracked : renderPassEntry.trackedAttachmentLayouts) {
if (tracked.target == TrackedAttachmentTarget::Texture) {
if (const auto texture = tracked.texture.lock()) {
s_textureManager->StampTextureRecordingUse(texture.get());
}
}
}
}
for (const auto& pending: renderPassEntry.pendingClearAttachments) { for (const auto& pending: renderPassEntry.pendingClearAttachments) {
if (pending.hasInlinePayload) { if (pending.hasInlinePayload) {
if (s_renderPassManager != nullptr) { if (s_renderPassManager != nullptr) {
@@ -1208,11 +1476,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
case TrackedAttachmentTarget::SwapchainColor: case TrackedAttachmentTarget::SwapchainColor:
MOBILEGL_ASSERT(s_swapchainObject != nullptr, "EndRenderPass: swapchain object is null"); MOBILEGL_ASSERT(s_swapchainObject != nullptr, "EndRenderPass: swapchain object is null");
s_swapchainObject->SetImageLayout(trackedAttachment.swapchainImageIndex, trackedAttachment.finalLayout); s_swapchainObject->SetImageLayout(trackedAttachment.swapchainImageIndex, trackedAttachment.finalLayout);
// The pass stored into the attachment: its content is defined
// until the image is next presented.
s_swapchainObject->SetImageContentDefined(trackedAttachment.swapchainImageIndex, true);
break; break;
case TrackedAttachmentTarget::SwapchainDepthStencil: case TrackedAttachmentTarget::SwapchainDepthStencil:
MOBILEGL_ASSERT(s_swapchainObject != nullptr, "EndRenderPass: swapchain object is null"); MOBILEGL_ASSERT(s_swapchainObject != nullptr, "EndRenderPass: swapchain object is null");
s_swapchainObject->SetDepthStencilImageLayout(trackedAttachment.swapchainImageIndex, s_swapchainObject->SetDepthStencilImageLayout(trackedAttachment.swapchainImageIndex,
trackedAttachment.finalLayout); trackedAttachment.finalLayout);
s_swapchainObject->SetDepthStencilContentDefined(trackedAttachment.swapchainImageIndex, true);
break; break;
default: default:
MOBILEGL_ASSERT(false, "EndRenderPass: unsupported tracked attachment target=%d", MOBILEGL_ASSERT(false, "EndRenderPass: unsupported tracked attachment target=%d",
@@ -42,6 +42,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
struct TrackedAttachmentLayoutInfo { struct TrackedAttachmentLayoutInfo {
TrackedAttachmentTarget target = TrackedAttachmentTarget::Texture; TrackedAttachmentTarget target = TrackedAttachmentTarget::Texture;
WeakPtr<MG_State::GLState::ITextureObject> texture; WeakPtr<MG_State::GLState::ITextureObject> texture;
// Identity-compare shortcut for the per-draw "does the active pass use
// this sampled texture" probe: comparing this against a LIVE texture's
// address needs no weak_ptr::lock (two refcount atomics per probe).
// May dangle once the texture dies - compare only, never dereference.
MG_State::GLState::ITextureObject* textureRaw = nullptr;
WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer; WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer;
Uint32 textureMipLevel = 0; Uint32 textureMipLevel = 0;
Uint32 swapchainImageIndex = 0; Uint32 swapchainImageIndex = 0;
@@ -157,19 +162,53 @@ namespace MobileGL::MG_Backend::DirectVulkan {
class VkRenderPassManager { class VkRenderPassManager {
public: public:
using HashType = Uint64; using HashType = Uint64;
// Notified once per OnPresent sweep with every aged-out entry's VkRenderPass
// value: pipelines are hashed on the raw handle, and once destroyed the value
// may be recycled for an incompatible pass, so dependent caches must purge
// everything keyed on them before any new pass can be created (the sweep and
// the notification run back-to-back with no creation in between; observers
// compare the values, never dereference them). Batched so a mass-idle cohort
// (shader-pack switch, dimension exit) costs the observer one pipeline-cache
// scan, not one per dying pass. The wholesale paths
// (Shutdown/RecreateSwapchain) do not notify - their callers already drop
// every pipeline outright.
class IEvictionObserver {
public:
virtual ~IEvictionObserver() = default;
virtual void OnRenderPassesDestroyed(const Vector<VkRenderPass>& renderPasses) = 0;
};
VkRenderPassManager(VkDevice device, VkRenderPassManager(VkDevice device,
VkPhysicalDevice physicalDevice, VmaAllocator allocator, const VulkanRendererConfig& config, VkPhysicalDevice physicalDevice, VmaAllocator allocator, const VulkanRendererConfig& config,
VkClearManager& clearManager, VkTextureManager& textureManager, SwapchainObject& swapchainObject); VkClearManager& clearManager, VkTextureManager& textureManager, SwapchainObject& swapchainObject);
~VkRenderPassManager(); ~VkRenderPassManager();
// Observer may be null (no notifications). Not owned.
void SetEvictionObserver(IEvictionObserver* observer) { m_evictionObserver = observer; }
Bool Initialize(); Bool Initialize();
void Shutdown(); void Shutdown();
HashType ComputeHash( HashType ComputeHash(
const MG_State::GLState::FramebufferObject& fbo, const MG_State::GLState::FramebufferObject& fbo,
Uint32 swapchainImageIndex, Uint32 swapchainImageIndex,
Bool includePendingClear = true); Bool includePendingClear = true,
RenderPassEntry& GetOrCreateRenderPass(const MG_State::GLState::FramebufferObject& fbo, Uint32 swapchainImageIndex); Bool includeDefaultFboDepthStencil = true);
// drawUsesDepthStencil: whether the operation about to run inside the pass
// reads or writes the depth/stencil buffer (depth test or stencil test
// enabled, or a depth/stencil clear). Only consulted for the DEFAULT
// framebuffer: EGL undefines its ancillary buffers at every swap, so a
// default-FBO pass whose draws provably never touch depth/stencil is
// created WITHOUT the depth attachment - on a tiler that skips the whole
// depth tile load AND store. The flavor only escalates: once a pass with
// depth is active, later depth-less draws keep using it, and a depth-using
// draw against a depth-less active pass resolves to a new (incompatible)
// entry, which the caller's compatibility check turns into a pass split;
// the new pass's depth loads DONT_CARE (content was undefined all along).
RenderPassEntry& GetOrCreateRenderPass(const MG_State::GLState::FramebufferObject& fbo,
Uint32 swapchainImageIndex,
Bool drawUsesDepthStencil = true);
void QueueRenderbufferClear(GLbitfield mask, const ClearFramebufferPayload& clearPayload, void QueueRenderbufferClear(GLbitfield mask, const ClearFramebufferPayload& clearPayload,
const MG_State::GLState::FramebufferObject& drawFbo); const MG_State::GLState::FramebufferObject& drawFbo);
void QueueRenderbufferClear(const ClearAttachmentPayload& clearPayload, void QueueRenderbufferClear(const ClearAttachmentPayload& clearPayload,
@@ -192,12 +231,20 @@ namespace MobileGL::MG_Backend::DirectVulkan {
UnorderedMap<Uint64, RenderPassEntry> m_renderPasses; UnorderedMap<Uint64, RenderPassEntry> m_renderPasses;
// Monotonic frame counter (bumped in OnPresent) for render-pass cache aging. // Monotonic frame counter (bumped in OnPresent) for render-pass cache aging.
Uint64 m_frameCounter = 0; Uint64 m_frameCounter = 0;
IEvictionObserver* m_evictionObserver = nullptr;
// Bumped whenever a renderbuffer VkImage is (re)created; together with the texture // Bumped whenever a renderbuffer VkImage is (re)created; together with the texture
// manager's image epoch this invalidates the render-pass fast path on any attachment // manager's image epoch this invalidates the render-pass fast path on any attachment
// image recreation. // image recreation.
Uint64 m_renderbufferImageEpoch = 1; Uint64 m_renderbufferImageEpoch = 1;
public:
// Bumped whenever a renderbuffer backing is (re)created; consecutive-draw
// snapshots include it so an attachment respecify forces a re-resolve.
Uint64 GetRenderbufferImageEpoch() const { return m_renderbufferImageEpoch; }
private:
// Per-draw fast-path memo for GetOrCreateRenderPass (dirty-flag state tracking): when the // Per-draw fast-path memo for GetOrCreateRenderPass (dirty-flag state tracking): when the
// framebuffer state is provably unchanged since the last resolution, the active render pass // framebuffer state is provably unchanged since the last resolution, the active render pass
// is reused WITHOUT recomputing the expensive per-draw hash. Invalidated by FBO switch / // is reused WITHOUT recomputing the expensive per-draw hash. Invalidated by FBO switch /
@@ -210,8 +257,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 m_rpFastTexEpoch = 0; Uint64 m_rpFastTexEpoch = 0;
Uint64 m_rpFastRbEpoch = 0; Uint64 m_rpFastRbEpoch = 0;
Uint64 m_rpFastRenderPassHash = 0; Uint64 m_rpFastRenderPassHash = 0;
// Whether the memoized entry carries a depth/stencil attachment; a
// default-FBO resolution whose effective depth request differs must
// miss the memo (the depth-less/depth-full flavors hash differently).
Bool m_rpFastHadDepthStencil = false;
public:
struct RenderbufferResource { struct RenderbufferResource {
// deadSinceFrame sentinel: the owning weak reference has not been observed
// expired. Dead resources age past every in-flight frame before Destroy
// (see CollectRenderbufferGarbage); the GPU may still reference the image
// for frames-in-flight frames after the GL object dies.
static constexpr Uint64 kNeverObservedDead = UINT64_MAX;
WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer; WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer;
VkImage image = VK_NULL_HANDLE; VkImage image = VK_NULL_HANDLE;
VmaAllocation allocation = nullptr; VmaAllocation allocation = nullptr;
@@ -223,25 +281,47 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkSampleCountFlagBits sampleCount = VK_SAMPLE_COUNT_1_BIT; VkSampleCountFlagBits sampleCount = VK_SAMPLE_COUNT_1_BIT;
TextureInternalFormat internalFormat = TextureInternalFormat::Unknown; TextureInternalFormat internalFormat = TextureInternalFormat::Unknown;
Int samples = 0; Int samples = 0;
// m_frameCounter value at which the weak reference was first seen expired.
Uint64 deadSinceFrame = kNeverObservedDead;
void Destroy(VkDevice device, VmaAllocator allocator); void Destroy(VkDevice device, VmaAllocator allocator);
}; };
// Public so the renderer's blit/copy/readback bindings can source renderbuffer
// attachments the same way texture attachments go through the texture manager.
RenderbufferResource* GetOrCreateRenderbufferResource(
const SharedPtr<MG_State::GLState::RenderbufferObject>& renderbuffer);
Bool GetPendingRenderbufferClear(MG_State::GLState::RenderbufferObject* renderbuffer,
ClearAttachmentPayload& outPayload) const;
private:
struct PendingRenderbufferClear { struct PendingRenderbufferClear {
WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer; WeakPtr<MG_State::GLState::RenderbufferObject> renderbuffer;
ClearAttachmentPayload payload{}; ClearAttachmentPayload payload{};
}; };
// A superseded renderbuffer backing (glRenderbufferStorage respecify) parked
// until enough frame boundaries have passed that no in-flight command buffer
// can still reference it; destroyed in OnPresent (see RetireAgeFrames).
struct DeferredRenderbufferRelease {
VkImage image = VK_NULL_HANDLE;
VmaAllocation allocation = nullptr;
VkImageView view = VK_NULL_HANDLE;
Uint64 deferredAtFrame = 0;
};
UnorderedMap<MG_State::GLState::RenderbufferObject*, RenderbufferResource> m_renderbufferResources; UnorderedMap<MG_State::GLState::RenderbufferObject*, RenderbufferResource> m_renderbufferResources;
UnorderedMap<MG_State::GLState::RenderbufferObject*, PendingRenderbufferClear> m_pendingRenderbufferClears; UnorderedMap<MG_State::GLState::RenderbufferObject*, PendingRenderbufferClear> m_pendingRenderbufferClears;
Vector<DeferredRenderbufferRelease> m_deferredRenderbufferReleases;
RenderbufferResource* GetOrCreateRenderbufferResource(
const SharedPtr<MG_State::GLState::RenderbufferObject>& renderbuffer);
Bool GetPendingRenderbufferClear(MG_State::GLState::RenderbufferObject* renderbuffer,
ClearAttachmentPayload& outPayload) const;
Bool HasPendingRenderbufferClear( Bool HasPendingRenderbufferClear(
const MG_State::GLState::FramebufferAttachmentObject& attachment) const; const MG_State::GLState::FramebufferAttachmentObject& attachment) const;
void CollectRenderbufferGarbage(); void CollectRenderbufferGarbage();
// Frame-boundary margin after which a resource last referenced by a retired
// GL object (or superseded backing) is provably past every in-flight frame.
Uint64 RetireAgeFrames() const;
void DeferRenderbufferBackingRelease(RenderbufferResource& resource);
void CollectDeferredRenderbufferReleases(Bool destroyAll);
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
static inline ActiveRenderPassInfo s_activeRenderPass{}; static inline ActiveRenderPassInfo s_activeRenderPass{};
@@ -51,6 +51,18 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Float ResolveEffectiveMinLod(const MG_State::GLState::SamplerObject& sampler, Float effectiveMaxLod) { Float ResolveEffectiveMinLod(const MG_State::GLState::SamplerObject& sampler, Float effectiveMaxLod) {
return std::min(sampler.GetMinLod(), effectiveMaxLod); return std::min(sampler.GetMinLod(), effectiveMaxLod);
} }
// A single-level view can only ever deliver the base level, but the LOD clamp must not be
// collapsed to exactly 0: both GL and Vulkan pick magFilter over minFilter from the
// *clamped* lambda, so maxLod = 0 would make every fragment magnify and quietly retire the
// min filter. 0.25 is the value VkSamplerCreateInfo's own note prescribes for emulating
// GL's non-mipmapped minification - large enough for lambda to stay positive, small enough
// that a NEAREST mip mode still rounds down to level 0. Clamped rather than assigned, so a
// texture whose GL_TEXTURE_MAX_LOD really is 0 keeps magnifying as GL says it must.
Float ResolveSingleLevelMaxLod(const MG_State::GLState::SamplerObject& sampler, Bool singleLevelView) {
const Float maxLod = ResolveEffectiveMaxLod(sampler);
return singleLevelView ? std::min(maxLod, 0.25f) : maxLod;
}
} // namespace } // namespace
Bool VkSamplerManager::Initialize(const InitInfo& initInfo) { Bool VkSamplerManager::Initialize(const InitInfo& initInfo) {
@@ -89,15 +101,43 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_device = VK_NULL_HANDLE; m_device = VK_NULL_HANDLE;
m_config = nullptr; m_config = nullptr;
m_frameBoundaryCounter = 0;
}
void VkSamplerManager::OnFrameBoundary() {
++m_frameBoundaryCounter;
// Sweep occasionally; destroy samplers whose last use is far past every
// in-flight frame. Destroy and erase must stay atomic, or Shutdown would
// double-free the handle; an evicted key that recurs simply re-creates
// its sampler on the next miss.
constexpr Uint64 kSweepInterval = 256;
constexpr Uint64 kRetireAgeBoundaries = 1024;
if ((m_frameBoundaryCounter % kSweepInterval) != 0) {
return;
}
for (auto it = m_samplers.begin(); it != m_samplers.end();) {
auto& entry = it->second;
if (m_frameBoundaryCounter - entry.lastUsedFrameBoundary > kRetireAgeBoundaries) {
if (m_device != VK_NULL_HANDLE && entry.handle != VK_NULL_HANDLE) {
vkDestroySampler(m_device, entry.handle, nullptr);
}
it = m_samplers.erase(it);
} else {
++it;
}
}
} }
Uint64 VkSamplerManager::BuildSamplerKey(const MG_State::GLState::SamplerObject& sampler, Uint64 VkSamplerManager::BuildSamplerKey(const MG_State::GLState::SamplerObject& sampler,
const MG_State::GLState::ITextureObject& texture, const MG_State::GLState::ITextureObject& texture,
Bool forceNearestFiltering) const { Bool forceNearestFiltering, Bool singleLevelView) const {
MOBILEGL_ASSERT(m_config != nullptr, "VkSamplerManager::BuildSamplerKey: m_config is null"); MOBILEGL_ASSERT(m_config != nullptr, "VkSamplerManager::BuildSamplerKey: m_config is null");
XXHASH_VERIFY(XXH64_reset(m_hashState, m_config->CacheVersion)); XXHASH_VERIFY(XXH64_reset(m_hashState, m_config->CacheVersion));
XXHASH_VERIFY(XXH64_update(m_hashState, &forceNearestFiltering, sizeof(forceNearestFiltering))); XXHASH_VERIFY(XXH64_update(m_hashState, &forceNearestFiltering, sizeof(forceNearestFiltering)));
XXHASH_VERIFY(XXH64_update(m_hashState, &singleLevelView, sizeof(singleLevelView)));
const auto minFilter = sampler.GetMinFilter(); const auto minFilter = sampler.GetMinFilter();
XXHASH_VERIFY(XXH64_update(m_hashState, &minFilter, sizeof(minFilter))); XXHASH_VERIFY(XXH64_update(m_hashState, &minFilter, sizeof(minFilter)));
@@ -111,7 +151,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
XXHASH_VERIFY(XXH64_update(m_hashState, &wrapT, sizeof(wrapT))); XXHASH_VERIFY(XXH64_update(m_hashState, &wrapT, sizeof(wrapT)));
const auto wrapR = sampler.GetWrapR(); const auto wrapR = sampler.GetWrapR();
XXHASH_VERIFY(XXH64_update(m_hashState, &wrapR, sizeof(wrapR))); XXHASH_VERIFY(XXH64_update(m_hashState, &wrapR, sizeof(wrapR)));
const auto maxLod = ResolveEffectiveMaxLod(sampler); const auto maxLod = ResolveSingleLevelMaxLod(sampler, singleLevelView);
const auto minLod = ResolveEffectiveMinLod(sampler, maxLod); const auto minLod = ResolveEffectiveMinLod(sampler, maxLod);
XXHASH_VERIFY(XXH64_update(m_hashState, &minLod, sizeof(minLod))); XXHASH_VERIFY(XXH64_update(m_hashState, &minLod, sizeof(minLod)));
XXHASH_VERIFY(XXH64_update(m_hashState, &maxLod, sizeof(maxLod))); XXHASH_VERIFY(XXH64_update(m_hashState, &maxLod, sizeof(maxLod)));
@@ -133,10 +173,20 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkSampler VkSamplerManager::GetOrCreateSampler(const MG_State::GLState::SamplerObject& sampler, VkSampler VkSamplerManager::GetOrCreateSampler(const MG_State::GLState::SamplerObject& sampler,
const MG_State::GLState::ITextureObject& texture, const MG_State::GLState::ITextureObject& texture,
Bool forceNearestFiltering) { Bool forceNearestFiltering, Uint32 viewLevelCount) {
const Uint64 key = BuildSamplerKey(sampler, texture, forceNearestFiltering); // A view that exposes a single mip level has no second level to blend with, so GL's
// *_MIPMAP_* minification filters degenerate to plain filtering on the base level -
// sampling is unchanged by pinning the Vulkan sampler to NEAREST mip mode at LOD 0.
// It is not cosmetic: MobileGL backs such a view with a fully allocated mip chain whose
// tail is never written, and a LINEAR mip mode lets the texture unit issue the level+1
// fetch anyway. On Adreno that fetch lands in uninitialized UBWC pages (or past the
// allocation for a genuinely single-level image) and faults the GPU - the same failure
// the default-framebuffer blit shader had to work around with an explicit-LOD sample.
const Bool singleLevelView = viewLevelCount == 1;
const Uint64 key = BuildSamplerKey(sampler, texture, forceNearestFiltering, singleLevelView);
auto it = m_samplers.find(key); auto it = m_samplers.find(key);
if (it != m_samplers.end()) { if (it != m_samplers.end()) {
it->second.lastUsedFrameBoundary = m_frameBoundaryCounter;
return it->second.handle; return it->second.handle;
} }
@@ -144,7 +194,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
samplerInfo.sType = VK_STRUCTURE_TYPE_SAMPLER_CREATE_INFO; samplerInfo.sType = VK_STRUCTURE_TYPE_SAMPLER_CREATE_INFO;
samplerInfo.magFilter = forceNearestFiltering ? VK_FILTER_NEAREST : ToVkFilter(sampler.GetMagFilter()); samplerInfo.magFilter = forceNearestFiltering ? VK_FILTER_NEAREST : ToVkFilter(sampler.GetMagFilter());
samplerInfo.minFilter = forceNearestFiltering ? VK_FILTER_NEAREST : ToVkFilter(sampler.GetMinFilter()); samplerInfo.minFilter = forceNearestFiltering ? VK_FILTER_NEAREST : ToVkFilter(sampler.GetMinFilter());
samplerInfo.mipmapMode = forceNearestFiltering ? VK_SAMPLER_MIPMAP_MODE_NEAREST samplerInfo.mipmapMode = (forceNearestFiltering || singleLevelView)
? VK_SAMPLER_MIPMAP_MODE_NEAREST
: ToVkMipmapMode(sampler.GetMipmapMode()); : ToVkMipmapMode(sampler.GetMipmapMode());
samplerInfo.addressModeU = ToVkAddressMode(sampler.GetWrapS()); samplerInfo.addressModeU = ToVkAddressMode(sampler.GetWrapS());
samplerInfo.addressModeV = ToVkAddressMode(sampler.GetWrapT()); samplerInfo.addressModeV = ToVkAddressMode(sampler.GetWrapT());
@@ -157,7 +208,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
samplerInfo.maxAnisotropy = maxAnisotropy; samplerInfo.maxAnisotropy = maxAnisotropy;
samplerInfo.compareEnable = sampler.GetCompareMode() == SamplerCompareMode::CompareToTexture ? VK_TRUE : VK_FALSE; samplerInfo.compareEnable = sampler.GetCompareMode() == SamplerCompareMode::CompareToTexture ? VK_TRUE : VK_FALSE;
samplerInfo.compareOp = ToVkCompareOp(ResolveCompareFunc(sampler, texture)); samplerInfo.compareOp = ToVkCompareOp(ResolveCompareFunc(sampler, texture));
samplerInfo.maxLod = ResolveEffectiveMaxLod(sampler); // Must match BuildSamplerKey's resolution exactly.
samplerInfo.maxLod = ResolveSingleLevelMaxLod(sampler, singleLevelView);
samplerInfo.minLod = ResolveEffectiveMinLod(sampler, samplerInfo.maxLod); samplerInfo.minLod = ResolveEffectiveMinLod(sampler, samplerInfo.maxLod);
samplerInfo.borderColor = ResolveVkBorderColor(sampler, texture); samplerInfo.borderColor = ResolveVkBorderColor(sampler, texture);
samplerInfo.unnormalizedCoordinates = VK_FALSE; samplerInfo.unnormalizedCoordinates = VK_FALSE;
@@ -169,6 +221,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
entry.handle = vkSampler; entry.handle = vkSampler;
entry.externalIndex = sampler.GetExternalIndex(); entry.externalIndex = sampler.GetExternalIndex();
entry.version = sampler.GetVersion(); entry.version = sampler.GetVersion();
entry.lastUsedFrameBoundary = m_frameBoundaryCounter;
m_samplers[key] = entry; m_samplers[key] = entry;
return vkSampler; return vkSampler;
} }
@@ -33,20 +33,38 @@ public:
Bool Initialize(const InitInfo& initInfo); Bool Initialize(const InitInfo& initInfo);
void Shutdown(); void Shutdown();
// viewLevelCount is the mip-level count of the image view this sampler will be paired
// with; 0 means "unknown, do not narrow". See GetOrCreateSampler for why it matters.
VkSampler GetOrCreateSampler(const MG_State::GLState::SamplerObject& sampler, VkSampler GetOrCreateSampler(const MG_State::GLState::SamplerObject& sampler,
const MG_State::GLState::ITextureObject& texture, const MG_State::GLState::ITextureObject& texture,
Bool forceNearestFiltering = false); Bool forceNearestFiltering = false,
Uint32 viewLevelCount = 0);
// Frame boundary hook: ages the sampler cache and destroys samplers not used
// for many frames. The key hashes continuous float state (lodBias, LOD clamps,
// anisotropy), so an app animating those would otherwise mint an unbounded
// stream of never-destroyed VkSamplers and eventually exhaust the device's
// maxSamplerAllocationCount. A sampler idle for over a thousand frame
// boundaries cannot be referenced by any in-flight command buffer (frames in
// flight are single digits), and every descriptor set the GPU consumes is
// written that same frame with live handles (the per-binding resolve memo and
// descriptor-set reuse are both frame-reset), so destruction here needs no
// fence wait. Self-gated: one counter bump and compare except on sweep
// boundaries.
void OnFrameBoundary();
private: private:
struct SamplerCacheEntry { struct SamplerCacheEntry {
VkSampler handle = VK_NULL_HANDLE; VkSampler handle = VK_NULL_HANDLE;
Uint externalIndex = 0; Uint externalIndex = 0;
Uint16 version = 0; Uint16 version = 0;
// Frame boundary of the last cache hit; entries idle past the
// OnFrameBoundary retirement age have their VkSampler destroyed.
Uint64 lastUsedFrameBoundary = 0;
}; };
Uint64 BuildSamplerKey(const MG_State::GLState::SamplerObject& sampler, Uint64 BuildSamplerKey(const MG_State::GLState::SamplerObject& sampler,
const MG_State::GLState::ITextureObject& texture, const MG_State::GLState::ITextureObject& texture,
Bool forceNearestFiltering) const; Bool forceNearestFiltering, Bool singleLevelView) const;
static VkFilter ToVkFilter(SamplerFilterMode mode); static VkFilter ToVkFilter(SamplerFilterMode mode);
static VkSamplerMipmapMode ToVkMipmapMode(SamplerMipmapMode mode); static VkSamplerMipmapMode ToVkMipmapMode(SamplerMipmapMode mode);
static VkSamplerAddressMode ToVkAddressMode(SamplerWrapMode mode); static VkSamplerAddressMode ToVkAddressMode(SamplerWrapMode mode);
@@ -67,6 +85,8 @@ private:
Bool m_samplerAnisotropySupported = false; Bool m_samplerAnisotropySupported = false;
Float m_maxSamplerAnisotropy = 1.0f; Float m_maxSamplerAnisotropy = 1.0f;
UnorderedMap<Uint64, SamplerCacheEntry> m_samplers; UnorderedMap<Uint64, SamplerCacheEntry> m_samplers;
// Monotonic frame-boundary counter (bumped in OnFrameBoundary) for cache aging.
Uint64 m_frameBoundaryCounter = 0;
static inline XXH64_state_t* m_hashState = XXH64_createState(); static inline XXH64_state_t* m_hashState = XXH64_createState();
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
@@ -375,7 +375,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
switch (format) { switch (format) {
case TextureInternalFormat::RGB: case TextureInternalFormat::RGB:
case TextureInternalFormat::RGB8: case TextureInternalFormat::RGB8:
// Legacy low-bit RGB formats share the UNorm8 canonical shadow layout (see
// TextureFormatProcessor), so they upload exactly like RGB8 with an alpha expand.
case TextureInternalFormat::R3G3B2:
case TextureInternalFormat::RGB4:
case TextureInternalFormat::RGB5:
return {VK_FORMAT_R8G8B8A8_UNORM, true, 1, {0xFF, 0x00, 0x00, 0x00}}; return {VK_FORMAT_R8G8B8A8_UNORM, true, 1, {0xFF, 0x00, 0x00, 0x00}};
// Low-bit RGBA formats: UNorm8x4 canonical shadow, no expansion needed.
case TextureInternalFormat::RGBA2:
case TextureInternalFormat::RGBA4:
case TextureInternalFormat::RGB5A1:
return {VK_FORMAT_R8G8B8A8_UNORM, false, 0, {0, 0, 0, 0}};
// 10/12-bit RGB(A): UNorm16 canonical shadow.
case TextureInternalFormat::RGB10:
case TextureInternalFormat::RGB12:
return {VK_FORMAT_R16G16B16A16_UNORM, true, 2, {0xFF, 0xFF, 0x00, 0x00}};
case TextureInternalFormat::RGBA12:
return {VK_FORMAT_R16G16B16A16_UNORM, false, 0, {0, 0, 0, 0}};
case TextureInternalFormat::SRGB8: case TextureInternalFormat::SRGB8:
return {VK_FORMAT_R8G8B8A8_SRGB, true, 1, {0xFF, 0x00, 0x00, 0x00}}; return {VK_FORMAT_R8G8B8A8_SRGB, true, 1, {0xFF, 0x00, 0x00, 0x00}};
case TextureInternalFormat::RGB8Snorm: case TextureInternalFormat::RGB8Snorm:
@@ -571,6 +587,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_allocator = initInfo.allocator; m_allocator = initInfo.allocator;
m_commandPool = initInfo.commandPool; m_commandPool = initInfo.commandPool;
m_graphicsQueue = initInfo.graphicsQueue; m_graphicsQueue = initInfo.graphicsQueue;
m_imageFormatListSupported = initInfo.imageFormatListSupported;
m_currentFrameIndex = 0; m_currentFrameIndex = 0;
m_deferredReleases.clear(); m_deferredReleases.clear();
m_deferredReleases.resize(initInfo.frameCount); m_deferredReleases.resize(initInfo.frameCount);
@@ -590,9 +607,14 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
void VkTextureManager::Shutdown() { void VkTextureManager::Shutdown() {
if (m_device != VK_NULL_HANDLE) {
ReclaimCompletedUploads(/*waitAll=*/true);
}
DestroyDeferredReleases(); DestroyDeferredReleases();
++m_resourceEraseEpoch; // every memoized resource pointer dies with the map
m_textureResources.clear(); m_textureResources.clear();
m_aliveObjects.clear(); m_aliveObjects.clear();
m_storageImageTextures.clear();
m_device = VK_NULL_HANDLE; m_device = VK_NULL_HANDLE;
m_physicalDevice = VK_NULL_HANDLE; m_physicalDevice = VK_NULL_HANDLE;
@@ -611,6 +633,27 @@ namespace MobileGL::MG_Backend::DirectVulkan {
frameIndex, m_deferredViewReleases.size()); frameIndex, m_deferredViewReleases.size());
m_currentFrameIndex = frameIndex; m_currentFrameIndex = frameIndex;
CollectDeferredReleases(frameIndex); CollectDeferredReleases(frameIndex);
ReclaimCompletedUploads();
// Frame-boundary GC: every 64 frame boundaries (~1 s at 60 fps) bounds the reclaim
// latency for dead textures regardless of draw traffic — workloads that churn
// textures through clears/readbacks alone never reach the draw-gated
// CollectGarbage. Must run after CollectDeferredReleases above: the prune defers
// its releases into this frame's slot, which was just drained, so they are
// destroyed only after the slot's fence has been waited again one full frame-ring
// cycle from now (never while an in-flight frame may still reference them).
constexpr Uint32 kGcFrameInterval = 64;
++m_gcFrameCounter;
if (m_gcFrameCounter % kGcFrameInterval == 0) {
PruneDeadTextures();
}
}
void VkTextureManager::CollectAllDeferredReleases() {
const SizeT frameCount = std::min(m_deferredReleases.size(), m_deferredViewReleases.size());
for (SizeT frameIndex = 0; frameIndex < frameCount; ++frameIndex) {
CollectDeferredReleases(static_cast<Uint32>(frameIndex));
}
} }
void VkTextureManager::EraseTrackedTexture(const TextureIdentity& identity) { void VkTextureManager::EraseTrackedTexture(const TextureIdentity& identity) {
@@ -620,6 +663,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_textureResources.erase(resourceIt); m_textureResources.erase(resourceIt);
} }
m_aliveObjects.erase(identity); m_aliveObjects.erase(identity);
m_storageImageTextures.erase(identity);
// Invalidate every cross-draw sampled-texture memo: the erased
// resource's address may be reused by a future emplace.
++m_resourceEraseEpoch;
} }
void VkTextureManager::PruneStaleTextureAliases(MG_State::GLState::ITextureObject* texture) { void VkTextureManager::PruneStaleTextureAliases(MG_State::GLState::ITextureObject* texture) {
@@ -675,6 +722,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} }
// Cross-draw memo probe (see SyncedTextureMemoEntry): skips both map
// lookups and the (re)registration path for repeat-bound textures.
TextureResource* resourcePtr = nullptr;
for (Uint32 i = 0; i < kSyncedTextureMemoSize; ++i) {
const SyncedTextureMemoEntry& memo = m_syncedTextureMemo[i];
if (memo.texture == &texture && memo.lifetimeId == identity.lifetimeId &&
memo.eraseEpoch == m_resourceEraseEpoch) {
resourcePtr = memo.resource;
break;
}
}
if (resourcePtr == nullptr) {
auto aliveIt = m_aliveObjects.find(identity); auto aliveIt = m_aliveObjects.find(identity);
if (aliveIt != m_aliveObjects.end() && aliveIt->second.expired()) { if (aliveIt != m_aliveObjects.end() && aliveIt->second.expired()) {
EraseTrackedTexture(aliveIt->first); EraseTrackedTexture(aliveIt->first);
@@ -686,9 +746,22 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// construction introduces a new identity. Doing this unconditionally made every // construction introduces a new identity. Doing this unconditionally made every
// sampled-texture sync scan the entire alive-texture map per draw. // sampled-texture sync scan the entire alive-texture map per draw.
if (aliveIt == m_aliveObjects.end()) { if (aliveIt == m_aliveObjects.end()) {
WeakPtr<MG_State::GLState::ITextureObject> aliveTexture;
const auto& liveTexture = MG_State::pGLContext->GetTextureObject(texture.GetExternalIndex()); const auto& liveTexture = MG_State::pGLContext->GetTextureObject(texture.GetExternalIndex());
if (liveTexture && liveTexture.get() == &texture) { if (liveTexture && liveTexture.get() == &texture) {
m_aliveObjects[identity] = WeakPtr<MG_State::GLState::ITextureObject>(liveTexture); aliveTexture = liveTexture;
} else {
// The name lookup legally fails while the object is alive: the name was
// deleted with the texture still attached to an FBO (the attachment's
// SharedPtr keeps it alive), or the name was reused by a new texture, or
// this is a default texture object (name 0 lives outside the name map).
// Register through the object's own control block so the resource created
// below still participates in weak-expiry GC instead of becoming an
// orphan no reclamation path can reach until Shutdown.
aliveTexture = texture.weak_from_this();
}
if (!aliveTexture.expired()) {
m_aliveObjects[identity] = Move(aliveTexture);
PruneStaleTextureAliases(&texture); PruneStaleTextureAliases(&texture);
} }
} }
@@ -699,8 +772,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
auto [insertIt, _] = m_textureResources.emplace(identity, Move(initial)); auto [insertIt, _] = m_textureResources.emplace(identity, Move(initial));
it = insertIt; it = insertIt;
} }
resourcePtr = &(it->second);
m_syncedTextureMemo[m_syncedTextureMemoNext] =
SyncedTextureMemoEntry{&texture, identity.lifetimeId, m_resourceEraseEpoch, resourcePtr};
m_syncedTextureMemoNext = (m_syncedTextureMemoNext + 1) % kSyncedTextureMemoSize;
}
if (!SyncTexture(texture, it->second)) { if (!SyncTexture(texture, *resourcePtr)) {
MGLOG_D("%s: Syncing texture %d failed", __func__, texture.GetExternalIndex()); MGLOG_D("%s: Syncing texture %d failed", __func__, texture.GetExternalIndex());
return nullptr; return nullptr;
} }
@@ -714,11 +792,11 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
} }
if (!recorded) { if (!recorded) {
m_drawSyncedThisDraw.push_back({identity, &(it->second)}); m_drawSyncedThisDraw.push_back({identity, resourcePtr});
} }
} }
return &(it->second); return resourcePtr;
} }
VkImageView VkTextureManager::GetOrCreateViewAtMipLevel(MG_State::GLState::ITextureObject& texture, Uint32 mipLevel) { VkImageView VkTextureManager::GetOrCreateViewAtMipLevel(MG_State::GLState::ITextureObject& texture, Uint32 mipLevel) {
@@ -997,6 +1075,16 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return view; return view;
} }
void VkTextureManager::StampTextureRecordingUse(MG_State::GLState::ITextureObject* texture) {
if (texture == nullptr) {
return;
}
auto it = m_textureResources.find(MakeTextureIdentity(texture));
if (it != m_textureResources.end()) {
it->second.lastRecordingGeneration = m_recordingGeneration;
}
}
void VkTextureManager::UpdateTrackedImageLayout(MG_State::GLState::ITextureObject* texture, VkImageLayout newLayout) { void VkTextureManager::UpdateTrackedImageLayout(MG_State::GLState::ITextureObject* texture, VkImageLayout newLayout) {
MOBILEGL_ASSERT(texture != nullptr, "UpdateTrackedImageLayout: texture is null"); MOBILEGL_ASSERT(texture != nullptr, "UpdateTrackedImageLayout: texture is null");
auto it = m_textureResources.find(MakeTextureIdentity(texture)); auto it = m_textureResources.find(MakeTextureIdentity(texture));
@@ -1024,6 +1112,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
MOBILEGL_ASSERT(writtenMipLevel < resource.mipLevels, MOBILEGL_ASSERT(writtenMipLevel < resource.mipLevels,
"UpdateTrackedImageLayoutAfterAttachmentWrite: textureId=%d mipLevel=%u out of range %u", "UpdateTrackedImageLayoutAfterAttachmentWrite: textureId=%d mipLevel=%u out of range %u",
texture->GetExternalIndex(), writtenMipLevel, resource.mipLevels); texture->GetExternalIndex(), writtenMipLevel, resource.mipLevels);
// Pre-pass stream bookkeeping: the render pass that just ended wrote this image.
StampResourceRecordingUse(resource);
if (resource.layout != newLayout && resource.mipLevels > 1) { if (resource.layout != newLayout && resource.mipLevels > 1) {
VkPipelineStageFlags srcStageMask = VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT; VkPipelineStageFlags srcStageMask = VK_PIPELINE_STAGE_TOP_OF_PIPE_BIT;
@@ -1108,6 +1198,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VK_ACCESS_SHADER_READ_BIT, resource->aspect, 0, resource->mipLevels, VK_ACCESS_SHADER_READ_BIT, resource->aspect, 0, resource->mipLevels,
resource->arrayLayers); resource->arrayLayers);
MOBILEGL_ASSERT(ok, "TransitionTextureForSampling: transition failed for textureId=%d", texture.GetExternalIndex()); MOBILEGL_ASSERT(ok, "TransitionTextureForSampling: transition failed for textureId=%d", texture.GetExternalIndex());
// Pre-pass stream bookkeeping: a command referencing the image was recorded.
StampResourceRecordingUse(*resource);
return ok; return ok;
} }
@@ -1137,11 +1229,30 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource->aspect, 0, resource->mipLevels, resource->arrayLayers); resource->aspect, 0, resource->mipLevels, resource->arrayLayers);
MOBILEGL_ASSERT(ok, "TransitionTextureForStorageImage: transition failed for textureId=%d", MOBILEGL_ASSERT(ok, "TransitionTextureForStorageImage: transition failed for textureId=%d",
texture.GetExternalIndex()); texture.GetExternalIndex());
// Pre-pass stream bookkeeping: a command referencing the image was recorded.
StampResourceRecordingUse(*resource);
return ok; return ok;
} }
void VkTextureManager::MarkStorageImageTexture(MG_State::GLState::ITextureObject& texture) {
m_storageImageTextures.insert(MakeTextureIdentity(&texture));
}
Bool VkTextureManager::NeedsStorageUsageUpgrade(MG_State::GLState::ITextureObject& texture) const {
const TextureIdentity identity = MakeTextureIdentity(&texture);
if (m_storageImageTextures.find(identity) == m_storageImageTextures.end()) {
return false;
}
const auto it = m_textureResources.find(identity);
// No image yet: the first sync creates it with STORAGE straight away, so there is nothing
// to preserve and nothing to order against.
return it != m_textureResources.end() && it->second.image != VK_NULL_HANDLE &&
!it->second.storageUsageResolved;
}
Bool VkTextureManager::NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const { Bool VkTextureManager::NeedsStorageImagePreparation(MG_State::GLState::ITextureObject& texture) const {
const auto it = m_textureResources.find(MakeTextureIdentity(&texture)); const TextureIdentity identity = MakeTextureIdentity(&texture);
const auto it = m_textureResources.find(identity);
if (it == m_textureResources.end()) { if (it == m_textureResources.end()) {
return true; return true;
} }
@@ -1149,6 +1260,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (resource.image == VK_NULL_HANDLE || resource.layout != VK_IMAGE_LAYOUT_GENERAL) { if (resource.image == VK_NULL_HANDLE || resource.layout != VK_IMAGE_LAYOUT_GENERAL) {
return true; return true;
} }
// The image predates this texture's first image-unit binding, so it was created without
// STORAGE usage and has to be recreated - which is illegal inside a render pass.
if (!resource.storageUsageResolved &&
m_storageImageTextures.find(identity) != m_storageImageTextures.end()) {
return true;
}
// Mirror SyncTexture's cross-draw skip condition: any version drift means the sync // Mirror SyncTexture's cross-draw skip condition: any version drift means the sync
// path may upload or rebuild, both of which need the render pass ended first. // path may upload or rebuild, both of which need the render pass ended first.
const auto* mipTexture = MG_State::GLState::AsMipmapTexture(&texture); const auto* mipTexture = MG_State::GLState::AsMipmapTexture(&texture);
@@ -1196,10 +1313,22 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
SizeT VkTextureManager::CollectGarbage() { SizeT VkTextureManager::CollectGarbage() {
// Draw-gated stagger (1 in 256 calls): keeps the per-draw cost at one counter
// bump. The guaranteed reclaim path is the frame-boundary prune in BeginFrame;
// this remains as a cheap assist so draw-heavy workloads reclaim sooner.
m_gcCounter++; m_gcCounter++;
if (m_gcCounter != 0) { if (m_gcCounter != 0) {
return 0; return 0;
} }
return PruneDeadTextures();
}
SizeT VkTextureManager::PruneDeadTextures() {
// Erasing entries would dangle the raw TextureResource pointers memoized for the
// current draw; every call path (BeginFrame, and CollectGarbage at the top of a
// freshly opened draw-sync scope) runs before any memo entry is recorded.
MOBILEGL_ASSERT(m_drawSyncedThisDraw.empty(),
"PruneDeadTextures: draw-sync memo holds raw resource pointers an erase would dangle");
Vector<MG_State::GLState::ITextureObject*> expiredTextures; Vector<MG_State::GLState::ITextureObject*> expiredTextures;
expiredTextures.reserve(m_aliveObjects.size()); expiredTextures.reserve(m_aliveObjects.size());
@@ -1211,7 +1340,25 @@ namespace MobileGL::MG_Backend::DirectVulkan {
for (auto* texture : expiredTextures) { for (auto* texture : expiredTextures) {
PruneStaleTextureAliases(texture); PruneStaleTextureAliases(texture);
} }
return expiredTextures.size(); SizeT prunedCount = expiredTextures.size();
// Orphan sweep: after the pass above, m_aliveObjects holds only live entries.
// Registration in SyncTextureAndGetDescriptor cannot fail for a SharedPtr-owned
// texture (weak_from_this fallback), so a resource whose identity has no alive
// entry has no trackable owner: its GL-side object is gone, or was never
// shared-owned, in which case recreation on a later sync is the safe fallback.
// Destruction goes through the per-frame deferred queues, never immediate.
Vector<TextureIdentity> orphanIdentities;
for (auto it = m_textureResources.begin(); it != m_textureResources.end(); ++it) {
if (m_aliveObjects.find(it->first) == m_aliveObjects.end()) {
orphanIdentities.emplace_back(it->first);
}
}
for (const auto& identity : orphanIdentities) {
EraseTrackedTexture(identity);
}
prunedCount += orphanIdentities.size();
return prunedCount;
} }
Bool VkTextureManager::SyncTexture(MG_State::GLState::ITextureObject &texture, Bool VkTextureManager::SyncTexture(MG_State::GLState::ITextureObject &texture,
@@ -1225,7 +1372,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const auto* syncingMipTexture = MG_State::GLState::AsMipmapTexture(&texture); const auto* syncingMipTexture = MG_State::GLState::AsMipmapTexture(&texture);
const Uint32 syncingMipLevelCount = const Uint32 syncingMipLevelCount =
syncingMipTexture != nullptr ? syncingMipTexture->GetMipmapLevelCount() : 0u; syncingMipTexture != nullptr ? syncingMipTexture->GetMipmapLevelCount() : 0u;
if (outResource.image != VK_NULL_HANDLE && // A pending storage-usage upgrade also has to bust the skip: nothing about the texture's
// content or params changed, but the image itself must be recreated with STORAGE usage
// before it can back an image-unit descriptor.
const Bool storageUpgradePending =
!outResource.storageUsageResolved &&
m_storageImageTextures.find(MakeTextureIdentity(&texture)) != m_storageImageTextures.end();
if (outResource.image != VK_NULL_HANDLE && !storageUpgradePending &&
outResource.syncedContentVersion == syncingContentVersion && outResource.syncedContentVersion == syncingContentVersion &&
outResource.syncedTextureParamsVersion == texture.GetTextureParamsVersion() && outResource.syncedTextureParamsVersion == texture.GetTextureParamsVersion() &&
outResource.syncedMipLevelCount == syncingMipLevelCount) { outResource.syncedMipLevelCount == syncingMipLevelCount) {
@@ -1309,8 +1462,17 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return false; return false;
} }
const Bool isMultisampleTexture = IsMultisampleTextureUploadTarget(uploadTarget); const Bool isMultisampleTexture = IsMultisampleTextureUploadTarget(uploadTarget);
// A texture that has only ever defined level 0 gets a single-level backing
// (ANGLE's model). Preallocating the full chain put every render target
// onto Adreno's multi-mip image layout and grew each texture by a third
// for levels most textures never define. Once a second level is defined
// the backing is recreated ONE time with the full chain (the
// preserve-copy path below carries the pixels over), so sequentially-
// defined atlas mips do not recreate per level, and glGenerateMipmap -
// which defines every level before syncing - works unchanged.
const Uint32 backingMipLevels = const Uint32 backingMipLevels =
isMultisampleTexture ? 1u : std::max(mipLevels, ComputeFullMipLevelCount(texelSize)); isMultisampleTexture ? 1u
: (mipLevels > 1 ? std::max(mipLevels, ComputeFullMipLevelCount(texelSize)) : 1u);
TextureShapeInfo shapeInfo{}; TextureShapeInfo shapeInfo{};
const Bool supportedShape = TryResolveTextureShapeInfo(texture, uploadTarget, texelSize, shapeInfo); const Bool supportedShape = TryResolveTextureShapeInfo(texture, uploadTarget, texelSize, shapeInfo);
MOBILEGL_ASSERT(supportedShape, MOBILEGL_ASSERT(supportedShape,
@@ -1336,16 +1498,44 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const VkImageAspectFlags aspect = GetAspectMaskForFormat(format); const VkImageAspectFlags aspect = GetAspectMaskForFormat(format);
VkFormatProperties formatProperties{}; VkFormatProperties formatProperties{};
vkGetPhysicalDeviceFormatProperties(m_physicalDevice, format, &formatProperties); vkGetPhysicalDeviceFormatProperties(m_physicalDevice, format, &formatProperties);
const Bool supportsStorageImage = // Only textures that have actually been bound to a GL image unit get STORAGE usage (and
// the MUTABLE_FORMAT it drags in for format-reinterpreting image views). Requesting it
// for every storage-capable colour texture costs real bandwidth: Adreno cannot keep UBWC
// compression on an image that may be written through a storage descriptor, so the whole
// render target - MC's included - runs uncompressed. MarkStorageImageTexture upgrades a
// texture before its first image-unit draw, and the usage below feeds the compatibility
// check so the upgrade recreates the image.
const Bool markedAsStorageImage =
m_storageImageTextures.find(MakeTextureIdentity(
const_cast<MG_State::GLState::ITextureObject*>(&texture))) != m_storageImageTextures.end();
// Storage-image CAPABILITY (does the format allow it at all) is deliberately separate from
// whether this texture actually needs the usage. MUTABLE_FORMAT keys off capability, as
// before: format-reinterpreting views are not a storage-only concern - the SAMPLED path
// needs them too (GetOrCreateSampledImageView bails out without it, see ~line 892), so
// tying MUTABLE_FORMAT to the image-unit mark would break sampled format reinterpretation
// for every texture that never becomes a storage image.
const Bool storageImageCapable =
!isMultisampleTexture && !isMultisampleTexture &&
(aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 && (aspect & VK_IMAGE_ASPECT_COLOR_BIT) != 0 &&
(formatProperties.optimalTilingFeatures & VK_FORMAT_FEATURE_STORAGE_IMAGE_BIT) != 0; (formatProperties.optimalTilingFeatures & VK_FORMAT_FEATURE_STORAGE_IMAGE_BIT) != 0;
const Bool supportsStorageImage = storageImageCapable && markedAsStorageImage;
VkImageCreateFlags imageCreateFlags = shapeInfo.imageFlags; VkImageCreateFlags imageCreateFlags = shapeInfo.imageFlags;
if (supportsStorageImage && IsMutableStorageImageFormat(format) && if (storageImageCapable && IsMutableStorageImageFormat(format) &&
m_mutableFormatUnsupported.find(format) == m_mutableFormatUnsupported.end()) { m_mutableFormatUnsupported.find(format) == m_mutableFormatUnsupported.end()) {
imageCreateFlags |= VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT; imageCreateFlags |= VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT;
} }
VkImageUsageFlags desiredUsage =
VK_IMAGE_USAGE_SAMPLED_BIT |
(supportsStorageImage ? VK_IMAGE_USAGE_STORAGE_BIT : 0) |
((aspect & VK_IMAGE_ASPECT_COLOR_BIT) ? VK_IMAGE_USAGE_COLOR_ATTACHMENT_BIT : 0) |
(((aspect & VK_IMAGE_ASPECT_DEPTH_BIT) || (aspect & VK_IMAGE_ASPECT_STENCIL_BIT)) ?
VK_IMAGE_USAGE_DEPTH_STENCIL_ATTACHMENT_BIT :
0);
if (!isMultisampleTexture) {
desiredUsage |= VK_IMAGE_USAGE_TRANSFER_DST_BIT | VK_IMAGE_USAGE_TRANSFER_SRC_BIT;
}
const Bool compatible = resource.image != VK_NULL_HANDLE && resource.format == format && const Bool compatible = resource.image != VK_NULL_HANDLE && resource.format == format &&
resource.extent.width == static_cast<Uint32>(texelSize.x()) && resource.extent.width == static_cast<Uint32>(texelSize.x()) &&
resource.extent.height == static_cast<Uint32>(texelSize.y()) && resource.extent.height == static_cast<Uint32>(texelSize.y()) &&
@@ -1354,6 +1544,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource.viewType == shapeInfo.viewType && resource.viewType == shapeInfo.viewType &&
resource.sampleCount == resolvedSampleCount && resource.sampleCount == resolvedSampleCount &&
resource.imageCreateFlags == imageCreateFlags && resource.imageCreateFlags == imageCreateFlags &&
resource.usageFlags == desiredUsage &&
resource.mipLevels == backingMipLevels; resource.mipLevels == backingMipLevels;
if (compatible) { if (compatible) {
if (resource.perMipViews.size() != backingMipLevels) { if (resource.perMipViews.size() != backingMipLevels) {
@@ -1362,6 +1553,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
if (resource.perMipSampledViews.size() != backingMipLevels) { if (resource.perMipSampledViews.size() != backingMipLevels) {
resource.perMipSampledViews.resize(backingMipLevels, VK_NULL_HANDLE); resource.perMipSampledViews.resize(backingMipLevels, VK_NULL_HANDLE);
} }
// Keeping the image is itself the answer to the mark: either it already carries
// STORAGE, or this format can never carry it. Either way there is nothing left to
// recreate, so stop reporting the texture as needing preparation.
resource.storageUsageResolved = markedAsStorageImage;
return true; return true;
} }
@@ -1376,7 +1571,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource.sampleCount == resolvedSampleCount && resource.sampleCount == resolvedSampleCount &&
resource.imageCreateFlags == imageCreateFlags && resource.imageCreateFlags == imageCreateFlags &&
resolvedSampleCount == VK_SAMPLE_COUNT_1_BIT && resolvedSampleCount == VK_SAMPLE_COUNT_1_BIT &&
resource.mipLevels < backingMipLevels && // '<=' rather than '<': a storage-usage upgrade recreates the image with an
// unchanged mip count, and its contents (a render target's pixels live only on the
// GPU) still have to survive. The vkCmdCopyImage below copies min(mipLevels).
resource.mipLevels <= backingMipLevels &&
resource.layout != VK_IMAGE_LAYOUT_UNDEFINED; resource.layout != VK_IMAGE_LAYOUT_UNDEFINED;
std::unique_ptr<TextureResource> preservedResource; std::unique_ptr<TextureResource> preservedResource;
@@ -1398,16 +1596,37 @@ namespace MobileGL::MG_Backend::DirectVulkan {
imageInfo.format = format; imageInfo.format = format;
imageInfo.tiling = VK_IMAGE_TILING_OPTIMAL; imageInfo.tiling = VK_IMAGE_TILING_OPTIMAL;
imageInfo.initialLayout = VK_IMAGE_LAYOUT_UNDEFINED; imageInfo.initialLayout = VK_IMAGE_LAYOUT_UNDEFINED;
imageInfo.usage = VK_IMAGE_USAGE_SAMPLED_BIT | imageInfo.usage = desiredUsage;
(supportsStorageImage ? VK_IMAGE_USAGE_STORAGE_BIT : 0) |
((aspect & VK_IMAGE_ASPECT_COLOR_BIT) ? VK_IMAGE_USAGE_COLOR_ATTACHMENT_BIT : 0) |
(((aspect & VK_IMAGE_ASPECT_DEPTH_BIT) || (aspect & VK_IMAGE_ASPECT_STENCIL_BIT)) ?
VK_IMAGE_USAGE_DEPTH_STENCIL_ATTACHMENT_BIT :
0);
if (!isMultisampleTexture) {
imageInfo.usage |= VK_IMAGE_USAGE_TRANSFER_DST_BIT | VK_IMAGE_USAGE_TRANSFER_SRC_BIT;
}
imageInfo.samples = resolvedSampleCount; imageInfo.samples = resolvedSampleCount;
// Bound the mutability. A blindly-mutable image has to be laid out so that ANY format in
// its compatibility class can be viewed, which costs bandwidth compression on tilers;
// naming the exact set instead lets the driver keep it. Only safe when that set really is
// exhaustive, so it is restricted to textures that are not image-unit bound: sampled views
// can only ever ask for ResolveSampledImageViewFormat's output, whereas glBindImageTexture
// may name any compatible format, which nothing here can enumerate ahead of time.
Vector<VkFormat> viewFormats;
VkImageFormatListCreateInfo formatListInfo{};
if (m_imageFormatListSupported && !supportsStorageImage &&
(imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
viewFormats.push_back(format);
for (const SamplerNumericDomain domain : {SamplerNumericDomain::Float,
SamplerNumericDomain::SignedInteger,
SamplerNumericDomain::UnsignedInteger}) {
const VkFormat viewFormat = ResolveSampledImageViewFormat(format, domain);
if (viewFormat == VK_FORMAT_UNDEFINED) {
continue;
}
if (std::find(viewFormats.begin(), viewFormats.end(), viewFormat) == viewFormats.end()) {
viewFormats.push_back(viewFormat);
}
}
formatListInfo.sType = VK_STRUCTURE_TYPE_IMAGE_FORMAT_LIST_CREATE_INFO;
formatListInfo.viewFormatCount = static_cast<Uint32>(viewFormats.size());
formatListInfo.pViewFormats = viewFormats.data();
imageInfo.pNext = &formatListInfo;
}
if (isMultisampleTexture || (imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) { if (isMultisampleTexture || (imageInfo.flags & VK_IMAGE_CREATE_MUTABLE_FORMAT_BIT) != 0) {
VkImageFormatProperties imageFormatProperties{}; VkImageFormatProperties imageFormatProperties{};
VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties( VkResult imageFormatResult = vkGetPhysicalDeviceImageFormatProperties(
@@ -1464,6 +1683,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
resource.viewType = shapeInfo.viewType; resource.viewType = shapeInfo.viewType;
resource.sampleCount = resolvedSampleCount; resource.sampleCount = resolvedSampleCount;
resource.imageCreateFlags = imageCreateFlags; resource.imageCreateFlags = imageCreateFlags;
resource.usageFlags = imageInfo.usage;
resource.storageUsageResolved = markedAsStorageImage;
resource.syncedTextureParamsVersion = 0; resource.syncedTextureParamsVersion = 0;
if (preservedResource) { if (preservedResource) {
@@ -1513,6 +1734,28 @@ namespace MobileGL::MG_Backend::DirectVulkan {
m_deferredViewReleases[frameIndex].clear(); m_deferredViewReleases[frameIndex].clear();
} }
void VkTextureManager::ReclaimCompletedUploads(Bool waitAll) {
if (m_pendingUploadReclaims.empty()) {
return;
}
SizeT completed = 0;
for (; completed < m_pendingUploadReclaims.size(); ++completed) {
PendingUploadReclaim& entry = m_pendingUploadReclaims[completed];
if (waitAll) {
VK_VERIFY(vkWaitForFences(m_device, 1, &entry.fence, VK_TRUE, UINT64_MAX),
"vkWaitForFences(texture upload reclaim)");
} else if (vkGetFenceStatus(m_device, entry.fence) != VK_SUCCESS) {
break;
}
vkDestroyFence(m_device, entry.fence, nullptr);
vkFreeCommandBuffers(m_device, m_commandPool, 1, &entry.commandBuffer);
vmaDestroyBuffer(m_allocator, entry.stagingBuffer, entry.stagingAllocation);
}
m_pendingUploadReclaims.erase(m_pendingUploadReclaims.begin(),
m_pendingUploadReclaims.begin() + static_cast<std::ptrdiff_t>(completed));
}
void VkTextureManager::DestroyDeferredReleases() { void VkTextureManager::DestroyDeferredReleases() {
for (auto& deferredReleases : m_deferredReleases) { for (auto& deferredReleases : m_deferredReleases) {
deferredReleases.clear(); deferredReleases.clear();
@@ -1819,11 +2062,23 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VK_VERIFY(vkCreateFence(m_device, &fenceInfo, nullptr, &uploadFence), "vkCreateFence(texture upload)"); VK_VERIFY(vkCreateFence(m_device, &fenceInfo, nullptr, &uploadFence), "vkCreateFence(texture upload)");
VK_VERIFY(vkQueueSubmit(m_graphicsQueue, 1, &submitInfo, uploadFence), "vkQueueSubmit(texture)"); VK_VERIFY(vkQueueSubmit(m_graphicsQueue, 1, &submitInfo, uploadFence), "vkQueueSubmit(texture)");
VK_VERIFY(vkWaitForFences(m_device, 1, &uploadFence, VK_TRUE, UINT64_MAX), "vkWaitForFences(texture upload)"); // Do NOT wait the fence here: this submit sits behind the previous
vkDestroyFence(m_device, uploadFence, nullptr); // frame's rendering on the queue, so a synchronous wait stalls the CPU
vkFreeCommandBuffers(m_device, m_commandPool, 1, &commandBuffer); // until the GPU drains - a per-frame vkQueueWaitIdle for any workload
// with animated textures. Ordering against the current frame's draws is
vmaDestroyBuffer(m_allocator, stagingBuffer, stagingAllocation); // already guaranteed (its command buffer is submitted later, at
// present), so only the transient objects need to survive execution;
// park them until the fence signals.
m_pendingUploadReclaims.push_back({uploadFence, commandBuffer, stagingBuffer, stagingAllocation});
ReclaimCompletedUploads();
// Backstop for pathological upload storms: bound in-flight staging
// memory by blocking on the oldest upload only once the list is deep.
constexpr SizeT kMaxPendingTextureUploads = 16;
if (m_pendingUploadReclaims.size() > kMaxPendingTextureUploads) {
VK_VERIFY(vkWaitForFences(m_device, 1, &m_pendingUploadReclaims.front().fence, VK_TRUE, UINT64_MAX),
"vkWaitForFences(texture upload backstop)");
ReclaimCompletedUploads();
}
if (!ok) { if (!ok) {
MGLOG_D("%s: texture upload cmd failed", __func__); MGLOG_D("%s: texture upload cmd failed", __func__);
@@ -28,6 +28,9 @@ public:
// manager keys its per-draw fast path on this so an attachment's image recreation // manager keys its per-draw fast path on this so an attachment's image recreation
// invalidates the cached render pass (dirty-flag tracking; portable to Vulkan 1.1). // invalidates the cached render pass (dirty-flag tracking; portable to Vulkan 1.1).
Uint64 GetTextureImageEpoch() const { return m_textureImageEpoch; } Uint64 GetTextureImageEpoch() const { return m_textureImageEpoch; }
// Bumped whenever any tracked texture resource is erased; cached
// TextureResource pointers are valid only while this is unchanged.
Uint64 GetResourceEraseEpoch() const { return m_resourceEraseEpoch; }
struct TextureIdentity { struct TextureIdentity {
MG_State::GLState::ITextureObject* texture = nullptr; MG_State::GLState::ITextureObject* texture = nullptr;
@@ -53,6 +56,9 @@ public:
VkCommandPool commandPool = VK_NULL_HANDLE; VkCommandPool commandPool = VK_NULL_HANDLE;
VkQueue graphicsQueue = VK_NULL_HANDLE; VkQueue graphicsQueue = VK_NULL_HANDLE;
Uint32 frameCount = 0; Uint32 frameCount = 0;
// VK_KHR_image_format_list is enabled: MUTABLE_FORMAT images can name the exact set of
// formats they will be viewed as, which is what lets a tiler keep them compressed.
Bool imageFormatListSupported = false;
}; };
struct TextureResource { struct TextureResource {
@@ -157,7 +163,25 @@ public:
VkImageViewType viewType = VK_IMAGE_VIEW_TYPE_2D; VkImageViewType viewType = VK_IMAGE_VIEW_TYPE_2D;
VkSampleCountFlagBits sampleCount = VK_SAMPLE_COUNT_1_BIT; VkSampleCountFlagBits sampleCount = VK_SAMPLE_COUNT_1_BIT;
VkImageCreateFlags imageCreateFlags = 0; VkImageCreateFlags imageCreateFlags = 0;
// Usage the live image was created with. STORAGE is only requested for textures that
// have actually been bound to a GL image unit, because on Adreno a storage-capable
// image loses UBWC bandwidth compression; a later image binding upgrades the usage
// and recreates the image, so the resolved usage has to be part of the compatibility
// check that decides whether the existing image can be kept.
VkImageUsageFlags usageFlags = 0;
// True once this image was (re)resolved while the texture was already marked as an
// image-unit texture. Distinguishes "not upgraded yet" from "cannot be upgraded"
// (a format whose optimalTilingFeatures lack STORAGE_IMAGE never gains the bit), so
// NeedsStorageImagePreparation cannot ask for a recreate that will never happen.
Bool storageUsageResolved = false;
Uint16 syncedTextureParamsVersion = 0; Uint16 syncedTextureParamsVersion = 0;
// Recording generation (VkTextureManager::GetRecordingGeneration) of the last
// command referencing this image that was recorded into the CURRENT frame
// command buffer. An image untouched by the open recording may have its
// out-of-pass work (deferred clears, sampled-layout transitions) recorded
// into the frame's PRE command buffer - which executes strictly before the
// frame's commands - instead of splitting the active render pass.
Uint64 lastRecordingGeneration = 0;
// Snapshot of ITextureObject::GetContentVersion() at the last successful sync; // Snapshot of ITextureObject::GetContentVersion() at the last successful sync;
// lets SyncTexture skip the whole re-check/re-upload when content is unchanged. // lets SyncTexture skip the whole re-check/re-upload when content is unchanged.
Uint64 syncedContentVersion = 0; Uint64 syncedContentVersion = 0;
@@ -190,7 +214,10 @@ public:
std::swap(this->viewType, that.viewType); std::swap(this->viewType, that.viewType);
std::swap(this->sampleCount, that.sampleCount); std::swap(this->sampleCount, that.sampleCount);
std::swap(this->imageCreateFlags, that.imageCreateFlags); std::swap(this->imageCreateFlags, that.imageCreateFlags);
std::swap(this->usageFlags, that.usageFlags);
std::swap(this->storageUsageResolved, that.storageUsageResolved);
std::swap(this->syncedTextureParamsVersion, that.syncedTextureParamsVersion); std::swap(this->syncedTextureParamsVersion, that.syncedTextureParamsVersion);
std::swap(this->lastRecordingGeneration, that.lastRecordingGeneration);
std::swap(this->syncedContentVersion, that.syncedContentVersion); std::swap(this->syncedContentVersion, that.syncedContentVersion);
std::swap(this->syncedMipLevelCount, that.syncedMipLevelCount); std::swap(this->syncedMipLevelCount, that.syncedMipLevelCount);
} }
@@ -251,6 +278,8 @@ public:
viewType = VK_IMAGE_VIEW_TYPE_2D; viewType = VK_IMAGE_VIEW_TYPE_2D;
sampleCount = VK_SAMPLE_COUNT_1_BIT; sampleCount = VK_SAMPLE_COUNT_1_BIT;
imageCreateFlags = 0; imageCreateFlags = 0;
usageFlags = 0;
storageUsageResolved = false;
syncedTextureParamsVersion = 0; syncedTextureParamsVersion = 0;
syncedContentVersion = 0; syncedContentVersion = 0;
syncedMipLevelCount = 0; syncedMipLevelCount = 0;
@@ -267,6 +296,10 @@ public:
Bool Initialize(const InitInfo& initInfo); Bool Initialize(const InitInfo& initInfo);
void Shutdown(); void Shutdown();
void BeginFrame(Uint32 frameIndex); void BeginFrame(Uint32 frameIndex);
// Drains every frame slot's deferred image/view releases. Only valid when
// the caller has proven every queue submission complete; used by the
// present-less frame-boundary drain.
void CollectAllDeferredReleases();
TextureResource* SyncTextureAndGetDescriptor( TextureResource* SyncTextureAndGetDescriptor(
MG_State::GLState::ITextureObject& texture); MG_State::GLState::ITextureObject& texture);
@@ -285,6 +318,32 @@ public:
VkImageLayout newLayout); VkImageLayout newLayout);
Bool TransitionTextureForSampling(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture); Bool TransitionTextureForSampling(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
Bool TransitionTextureForStorageImage(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture); Bool TransitionTextureForStorageImage(VkCommandBuffer commandBuffer, MG_State::GLState::ITextureObject& texture);
// Recording-generation bookkeeping for the pre-pass command stream. The
// generation advances every time the frame command buffer (re)begins
// recording; a resource whose stamp does not match was not referenced by
// any command in the open recording, so its out-of-pass work may safely
// execute ahead of the whole recording (in the pre command buffer).
void AdvanceRecordingGeneration() { ++m_recordingGeneration; }
void StampResourceRecordingUse(TextureResource& resource) const {
resource.lastRecordingGeneration = m_recordingGeneration;
}
// Map-lookup variant for callers that only hold the GL texture object.
void StampTextureRecordingUse(MG_State::GLState::ITextureObject* texture);
Bool WasTouchedThisRecording(const TextureResource& resource) const {
return resource.lastRecordingGeneration == m_recordingGeneration;
}
// Records that this texture is bound to a GL image unit, so its image must carry
// VK_IMAGE_USAGE_STORAGE_BIT. Must be called before NeedsStorageImagePreparation, and
// therefore before the render pass is committed: an image that has to be upgraded is
// recreated, which is illegal inside a render pass. Sticky for the texture's lifetime -
// GL lets an image binding come and go, and re-creating the image every time it does
// would cost far more than the compression it wins back.
void MarkStorageImageTexture(MG_State::GLState::ITextureObject& texture);
// True when this texture is marked but its live image predates the mark, i.e. the next sync
// will recreate it with STORAGE usage and copy the old contents forward. Callers use this to
// submit their pending recording first, so that copy cannot read pre-flush content.
Bool NeedsStorageUsageUpgrade(MG_State::GLState::ITextureObject& texture) const;
// Non-mutating probe for the per-draw storage-image fast path: true when preparing this // Non-mutating probe for the per-draw storage-image fast path: true when preparing this
// texture as a storage image may need work that is illegal inside a render pass (resource // texture as a storage image may need work that is illegal inside a render pass (resource
// creation, dirty-content upload, or a layout transition to GENERAL). Unknown state reports // creation, dirty-content upload, or a layout transition to GENERAL). Unknown state reports
@@ -331,6 +390,9 @@ public:
private: private:
// Bumped in SyncTextureResource right after vmaCreateImage(texture). See GetTextureImageEpoch(). // Bumped in SyncTextureResource right after vmaCreateImage(texture). See GetTextureImageEpoch().
Uint64 m_textureImageEpoch = 1; Uint64 m_textureImageEpoch = 1;
// See AdvanceRecordingGeneration. Starts above every resource's default
// stamp of 0 so a fresh resource counts as untouched.
Uint64 m_recordingGeneration = 1;
Bool SyncTexture(MG_State::GLState::ITextureObject &texture, Bool SyncTexture(MG_State::GLState::ITextureObject &texture,
TextureResource &outResource); TextureResource &outResource);
@@ -361,18 +423,28 @@ private:
void DeferViewRelease(VkImageView view); void DeferViewRelease(VkImageView view);
void CollectDeferredReleases(Uint32 frameIndex); void CollectDeferredReleases(Uint32 frameIndex);
void DestroyDeferredReleases(); void DestroyDeferredReleases();
// Frees the fence/command buffer/staging buffer of every in-flight texture
// upload whose fence has signaled (submission order = completion order on
// the single queue, so the scan stops at the first still-pending entry).
// waitAll blocks on every entry - Shutdown's drain.
void ReclaimCompletedUploads(Bool waitAll = false);
static TextureIdentity MakeTextureIdentity(MG_State::GLState::ITextureObject* texture); static TextureIdentity MakeTextureIdentity(MG_State::GLState::ITextureObject* texture);
void EraseTrackedTexture(const TextureIdentity& identity); void EraseTrackedTexture(const TextureIdentity& identity);
void PruneStaleTextureAliases(MG_State::GLState::ITextureObject* texture); void PruneStaleTextureAliases(MG_State::GLState::ITextureObject* texture);
SizeT PruneDeadTextures();
VkDevice m_device = VK_NULL_HANDLE; VkDevice m_device = VK_NULL_HANDLE;
VkPhysicalDevice m_physicalDevice = VK_NULL_HANDLE; VkPhysicalDevice m_physicalDevice = VK_NULL_HANDLE;
VmaAllocator m_allocator = nullptr; VmaAllocator m_allocator = nullptr;
VkCommandPool m_commandPool = VK_NULL_HANDLE; VkCommandPool m_commandPool = VK_NULL_HANDLE;
VkQueue m_graphicsQueue = VK_NULL_HANDLE; VkQueue m_graphicsQueue = VK_NULL_HANDLE;
Bool m_imageFormatListSupported = false;
Uint32 m_currentFrameIndex = 0; Uint32 m_currentFrameIndex = 0;
Uint8 m_gcCounter = 0; Uint8 m_gcCounter = 0;
// Frame-boundary GC gate: counts BeginFrame calls, not draws, so texture churn
// through non-draw paths (FBO clears, readbacks) still reaches the prune.
Uint32 m_gcFrameCounter = 0;
// Active only between BeginDrawSyncScope/EndDrawSyncScope; identities of // Active only between BeginDrawSyncScope/EndDrawSyncScope; identities of
// textures already fully synced in the current draw (small N -> flat scan). // textures already fully synced in the current draw (small N -> flat scan).
Bool m_drawSyncScopeActive = false; Bool m_drawSyncScopeActive = false;
@@ -385,12 +457,42 @@ private:
TextureResource* resource = nullptr; TextureResource* resource = nullptr;
}; };
Vector<DrawSyncedTexture> m_drawSyncedThisDraw; Vector<DrawSyncedTexture> m_drawSyncedThisDraw;
// Cross-draw sampled-texture memo: the same few textures (atlas, lightmap)
// are resolved on every draw, so cache their resource pointers and skip the
// alive/resource map lookups. Node-based std::unordered_map keeps the
// pointees stable across inserts; erases bump m_resourceEraseEpoch, which
// every memo entry must match. SyncTexture still runs on memo hits, so
// content/param freshness is unaffected. A dead-then-reused texture address
// cannot false-hit: the new object carries a new lifetime id.
struct SyncedTextureMemoEntry {
const MG_State::GLState::ITextureObject* texture = nullptr;
Uint64 lifetimeId = 0;
Uint64 eraseEpoch = 0;
TextureResource* resource = nullptr;
};
static constexpr Uint32 kSyncedTextureMemoSize = 8;
SyncedTextureMemoEntry m_syncedTextureMemo[kSyncedTextureMemoSize];
Uint32 m_syncedTextureMemoNext = 0;
Uint64 m_resourceEraseEpoch = 1;
// Formats whose mutable-image probe failed on this device; their images are created // Formats whose mutable-image probe failed on this device; their images are created
// without MUTABLE_FORMAT_BIT so repeat syncs neither re-probe nor flag-mismatch. // without MUTABLE_FORMAT_BIT so repeat syncs neither re-probe nor flag-mismatch.
std::unordered_set<VkFormat> m_mutableFormatUnsupported; std::unordered_set<VkFormat> m_mutableFormatUnsupported;
std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects; std::unordered_map<TextureIdentity, WeakPtr<MG_State::GLState::ITextureObject>, TextureIdentityHash> m_aliveObjects;
std::unordered_map<TextureIdentity, TextureResource, TextureIdentityHash> m_textureResources; std::unordered_map<TextureIdentity, TextureResource, TextureIdentityHash> m_textureResources;
// Textures that have been bound to a GL image unit (see MarkStorageImageTexture).
std::unordered_set<TextureIdentity, TextureIdentityHash> m_storageImageTextures;
Vector<Vector<TextureResource>> m_deferredReleases; Vector<Vector<TextureResource>> m_deferredReleases;
Vector<Vector<VkImageView>> m_deferredViewReleases; Vector<Vector<VkImageView>> m_deferredViewReleases;
// Texture uploads are submitted out-of-band but NOT waited on (waiting
// behind the queue serialized the CPU against the previous frame's GPU
// work every time an animated atlas re-uploaded). Their transient objects
// are parked here and reclaimed once the upload fence signals.
struct PendingUploadReclaim {
VkFence fence = VK_NULL_HANDLE;
VkCommandBuffer commandBuffer = VK_NULL_HANDLE;
VkBuffer stagingBuffer = VK_NULL_HANDLE;
VmaAllocation stagingAllocation = nullptr;
};
Vector<PendingUploadReclaim> m_pendingUploadReclaims;
}; };
} // namespace MobileGL::MG_Backend::DirectVulkan } // namespace MobileGL::MG_Backend::DirectVulkan
File diff suppressed because it is too large Load Diff
@@ -114,7 +114,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
} }
}; };
class VulkanRenderer : public IBufferCopyCommandProvider, public FrameContext::IRecordingObserver { class VulkanRenderer : public IBufferCopyCommandProvider,
public FrameContext::IRecordingObserver,
public VkRenderPassManager::IEvictionObserver,
public ProgramFactory::IEvictionObserver {
public: public:
VulkanRenderer(NativeWindowType window, const VulkanRendererConfig& cfg = {}); VulkanRenderer(NativeWindowType window, const VulkanRendererConfig& cfg = {});
~VulkanRenderer(); ~VulkanRenderer();
@@ -131,9 +134,31 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// recording, before any render pass. // recording, before any render pass.
void OnFrameCommandRecordingBegan(VkCommandBuffer commandBuffer) override; void OnFrameCommandRecordingBegan(VkCommandBuffer commandBuffer) override;
// VkRenderPassManager::IEvictionObserver: the render-pass aging sweep just
// destroyed these VkRenderPasses; evict every graphics pipeline hashed on a
// dying handle (they share its >1024-boundary idleness, so immediate
// destruction is safe) and drop the last-pipeline memo if any went.
void OnRenderPassesDestroyed(const Vector<VkRenderPass>& renderPasses) override;
// ProgramFactory::IEvictionObserver: an aged-out program entry was
// destroyed; evict its compute pipeline and graphics pipelines (same
// idleness guarantee - they are only bound through draws/dispatches that
// stamp the program entry) and purge the descriptor-set cache entries
// keyed by its now-recyclable VkDescriptorSetLayout handle.
void OnProgramEvicted(ProgramFactory::HashType programHash,
VkDescriptorSetLayout descriptorSetLayout) override;
Bool SetupDraw(FrameContext::FrameData& frame, GLenum mode, Flags<DrawSetupAspect> aspects, Bool SetupDraw(FrameContext::FrameData& frame, GLenum mode, Flags<DrawSetupAspect> aspects,
const DrawCmdParam& drawParams, const DrawCmdParam& drawParams,
const IndexBufferView* pIndexBufferView = nullptr); const IndexBufferView* pIndexBufferView = nullptr);
// ANGLE-style consecutive-draw fast path: SetupDraw snapshots the fully
// resolved draw configuration; the next draw whose cheap version/identity
// checks all match skips the resolution half (LOD probe, sampled-set
// walk, render-pass and pipeline resolution) and jumps straight to the
// per-draw tail. Returns false (leaving no side effects that the full
// path cannot redo idempotently) whenever anything might have changed.
Bool TrySetupDrawFastPath(FrameContext::FrameData& frame, GLenum mode, Flags<DrawSetupAspect> aspects,
const DrawCmdParam& drawParams, const IndexBufferView* pIndexBufferView);
void ClearAttachmentsOnActiveRenderPass(VkCommandBuffer commandBuffer, void ClearAttachmentsOnActiveRenderPass(VkCommandBuffer commandBuffer,
const RenderPassEntry& compatibleRenderPassEntry); const RenderPassEntry& compatibleRenderPassEntry);
@@ -257,6 +282,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 GetTimerQueryTimestampNs(const VkTimerQueryManager::TimestampRecord& record) const; Uint64 GetTimerQueryTimestampNs(const VkTimerQueryManager::TimestampRecord& record) const;
void RequestSwapchainResize(Uint32 width, Uint32 height); void RequestSwapchainResize(Uint32 width, Uint32 height);
// Re-query the surface and report whether the live swapchain no longer matches it
// (size or orientation). This - not a VK_SUBOPTIMAL_KHR result - is what decides a
// rebuild, so a surface the driver merely considers suboptimal cannot thrash.
Bool SwapchainIsOutOfDate();
// Returns false when the surface is zero-area (minimized/hidden window): // Returns false when the surface is zero-area (minimized/hidden window):
// no new swapchain is installed and presentation must stay suspended. // no new swapchain is installed and presentation must stay suspended.
Bool RecreateSwapchain(); Bool RecreateSwapchain();
@@ -345,11 +374,26 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkFence AcquirePooledSubmitFence(); VkFence AcquirePooledSubmitFence();
void DestroySubmitFencePool(); void DestroySubmitFencePool();
Bool HasPendingRecordedWork() const; Bool HasPendingRecordedWork() const;
// Frame-boundary housekeeping for paths that never reach Present's
// tail (present-less readback loops, suspended presentation, blocking
// sync waits): runs the same per-frame drains Present performs, but
// only when every queue submission has been observed complete AND no
// recorded-but-unsubmitted commands exist - i.e. when CPU-GPU overlap
// is provably already zero. Never blocks (non-blocking fence poll
// only), so the presenting path's frames-in-flight pipelining is
// untouched. Returns true when the drain ran.
Bool TryDrainFrameTransients();
Vector<SubmitRecord> m_inFlightSubmits; Vector<SubmitRecord> m_inFlightSubmits;
Vector<VkFence> m_freeSubmitFences; Vector<VkFence> m_freeSubmitFences;
Uint64 m_submitCounter = 0; Uint64 m_submitCounter = 0;
Uint64 m_completedSubmitCounter = 0; Uint64 m_completedSubmitCounter = 0;
// Drains since the last Present, gating the drain's frame-boundary-equivalent
// work (arena rewind + cache aging): a presenting app's mid-frame
// readbacks/waits must neither churn the transient caches nor accelerate the
// aging clocks, while present-less loops still cross a boundary every few
// iterations. Reset in Present.
Uint32 m_drainsSinceLastPresent = 0;
NativeWindowType m_window = 0; NativeWindowType m_window = 0;
void* m_platformDisplay = nullptr; void* m_platformDisplay = nullptr;
@@ -367,6 +411,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Vector<VkExtensionProperties> m_extensions; Vector<VkExtensionProperties> m_extensions;
VkInstance m_instance = VK_NULL_HANDLE; VkInstance m_instance = VK_NULL_HANDLE;
VkDebugUtilsMessengerEXT m_debugMessenger = VK_NULL_HANDLE; VkDebugUtilsMessengerEXT m_debugMessenger = VK_NULL_HANDLE;
// Fallback reporting channel for drivers that ship the validation layers but
// only expose the older VK_EXT_debug_report (Adreno 650 / Vulkan 1.1.128).
VkDebugReportCallbackEXT m_debugReportCallback = VK_NULL_HANDLE;
PhysicalDevice m_physicalDevice; PhysicalDevice m_physicalDevice;
VkDevice m_device = VK_NULL_HANDLE; VkDevice m_device = VK_NULL_HANDLE;
VmaAllocator m_allocator = nullptr; VmaAllocator m_allocator = nullptr;
@@ -416,14 +463,30 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// gather + synthetic vertex-input rebuild + payload hash + lookup) when the full pipeline // gather + synthetic vertex-input rebuild + payload hash + lookup) when the full pipeline
// state is unchanged from the previous draw. The key provably covers every pipeline field. // state is unchanged from the previous draw. The key provably covers every pipeline field.
// Reset per-frame and on pipeline destruction so the cached handle can never dangle. // Reset per-frame and on pipeline destruction so the cached handle can never dangle.
Bool m_lastPipelineValid = false; // Small N-way pipeline-resolution memo (round-robin replacement). A
GLenum m_lastPipelineMode = 0; // single-entry memo thrashed on draw sequences that alternate a few
Uint64 m_lastPipelineProgramHash = 0; // pipelines (GUI text/quad program ping-pong), paying the full
Uint64 m_lastPipelineVertexInputHash = 0; // payload-hash lookup per draw; eight entries cover such working sets
Uint64 m_lastPipelineRenderPassHash = 0; // while keeping the hit path a trivial linear scan.
Uint m_lastPipelineRenderStateVersion = 0; struct PipelineMemoEntry {
ProgramFactory::CompileOptionFlags m_lastPipelineTransformFlags = {}; GLenum mode = 0;
VkPipeline m_lastPipelineResult = VK_NULL_HANDLE; Uint64 programHash = 0;
Uint64 vertexInputHash = 0;
Uint64 renderPassHash = 0;
Uint renderStateVersion = 0;
ProgramFactory::CompileOptionFlags transformFlags = {};
VkPipeline pipeline = VK_NULL_HANDLE;
};
static constexpr Uint32 kPipelineMemoSize = 8;
PipelineMemoEntry m_pipelineMemo[kPipelineMemoSize];
Uint32 m_pipelineMemoCount = 0;
Uint32 m_pipelineMemoNext = 0;
// Drops every memoized pipeline handle. Required at command-buffer
// boundaries and whenever any pipeline may have been destroyed.
void InvalidatePipelineMemo() {
m_pipelineMemoCount = 0;
m_pipelineMemoNext = 0;
}
UnorderedMap<ProgramFactory::HashType, VkPipeline> m_computePipelines; UnorderedMap<ProgramFactory::HashType, VkPipeline> m_computePipelines;
UniquePtr<ProgramFactory> m_programFactory; UniquePtr<ProgramFactory> m_programFactory;
UniquePtr<UniformManager> m_uniformManager; UniquePtr<UniformManager> m_uniformManager;
@@ -452,9 +515,61 @@ namespace MobileGL::MG_Backend::DirectVulkan {
ProgramFactory::CompileOptionFlags m_lastSampledSetTransformFlags = {}; ProgramFactory::CompileOptionFlags m_lastSampledSetTransformFlags = {};
Uint64 m_lastSampledSetBindGeneration = 0; Uint64 m_lastSampledSetBindGeneration = 0;
// Memo for the per-draw explicit-LOD-0 eligibility probe
// (ProgramSamplesOnlySingleLevelTextures): same key family as the
// sampled-set memo, plus the sampled textures' params-version sum so a
// level-range or filter change re-probes. On a hit the resolved
// transform flags are reused, which also collapses the two
// GetOrCreateProgram lookups into one.
Bool m_lastLodDecisionValid = false;
Uint64 m_lastLodProgramLifetimeId = 0;
Uint32 m_lastLodProgramVersion = 0;
Uint64 m_lastLodBindGeneration = 0;
Uint64 m_lastLodParamsSum = 0;
ProgramFactory::CompileOptionFlags m_lastLodBaseFlags = {};
ProgramFactory::CompileOptionFlags m_lastLodResultFlags = {};
// Snapshot behind TrySetupDrawFastPath. Values only: the program and
// render-pass caches are open-addressing maps whose entries move on
// insert, so no pointers into them are cached; the pipeline handle is
// protected by the command-buffer-boundary reset plus the mid-frame
// pipeline-destruction resets, and monotonic epochs guard everything
// that can be destroyed or recreated between draws.
struct SetupDrawSnapshot {
Bool valid = false;
Uint8 aspects = 0;
GLenum mode = 0;
Uint64 programLifetimeId = 0;
Uint32 programVersion = 0;
const void* vao = nullptr;
Uint32 vaoConfigVersion = 0;
const void* drawFbo = nullptr;
Uint16 fboVersion = 0;
Bool drawFboIsDefault = false;
Uint renderStateVersion = 0;
Uint64 bindGeneration = 0;
Uint32 baseTransformFlags = 0;
Uint32 resolvedTransformFlags = 0;
Uint64 renderPassHash = 0;
Uint32 imageIndex = 0;
Uint64 textureEraseEpoch = 0;
Uint64 textureImageEpoch = 0;
Uint64 renderbufferImageEpoch = 0;
Uint64 sampledContentSum = 0;
Uint64 sampledParamsSum = 0;
IntVec2 renderPassExtent = {0, 0};
VkPipeline pipeline = VK_NULL_HANDLE;
};
SetupDrawSnapshot m_setupDrawSnapshot;
// Per-draw scratch buffers (clear keeps capacity) — these paths run for every // Per-draw scratch buffers (clear keeps capacity) — these paths run for every
// draw call and must not allocate. // draw call and must not allocate.
Vector<MG_State::GLState::ITextureObject*> m_sampledTexturesScratch; Vector<MG_State::GLState::ITextureObject*> m_sampledTexturesScratch;
// Parallel to m_sampledTexturesScratch, refilled by every SetupDraw's
// first sampled-texture loop: the resolved backend resources, so the
// post-transition loop can skip re-resolving textures whose layout is
// already sampleable.
Vector<VkTextureManager::TextureResource*> m_sampledResourcesScratch;
Vector<MG_State::GLState::ITextureObject*> m_storageImageTexturesScratch; Vector<MG_State::GLState::ITextureObject*> m_storageImageTexturesScratch;
Vector<VkBuffer> m_vertexBuffersScratch; Vector<VkBuffer> m_vertexBuffersScratch;
Vector<VkDeviceSize> m_vertexOffsetsScratch; Vector<VkDeviceSize> m_vertexOffsetsScratch;
@@ -517,6 +632,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
void CreateInstance(); void CreateInstance();
VkResult SetupDebugMessenger(); VkResult SetupDebugMessenger();
VkResult DestroyDebugMessenger(); VkResult DestroyDebugMessenger();
VkResult SetupDebugReportCallback();
void DestroyDebugReportCallback();
VkDebugUtilsMessengerCreateInfoEXT PopulateDebugMessengerCreateInfo(); VkDebugUtilsMessengerCreateInfoEXT PopulateDebugMessengerCreateInfo();
void CreateSurface(); void CreateSurface();
void PickPhysicalDevice(); void PickPhysicalDevice();
@@ -535,8 +652,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const RenderPassEntry& renderPassEntry); const RenderPassEntry& renderPassEntry);
VkPipeline GetOrCreateComputePipeline(const ProgramFactory::VkProgramObject& programObj); VkPipeline GetOrCreateComputePipeline(const ProgramFactory::VkProgramObject& programObj);
void DestroyComputePipelines(); void DestroyComputePipelines();
// Takes the frame rather than a command buffer: a first-time storage-usage upgrade has to
// flush the pending recording (see the body), which retires the current command buffer.
Bool PrepareStorageImageTextures( Bool PrepareStorageImageTextures(
VkCommandBuffer commandBuffer, FrameContext::FrameData& frame,
const MG_State::GLState::ProgramObject& program, const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj); const ProgramFactory::VkProgramObject& programObj);
@@ -561,6 +680,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
GLenum filter); GLenum filter);
Bool MaterializePendingClearForTexture(VkCommandBuffer commandBuffer, Bool MaterializePendingClearForTexture(VkCommandBuffer commandBuffer,
MG_State::GLState::ITextureObject& texture); MG_State::GLState::ITextureObject& texture);
Bool MaterializePendingClearForRenderbuffer(
VkCommandBuffer commandBuffer,
const SharedPtr<MG_State::GLState::RenderbufferObject>& renderbuffer);
VkPipeline GetOrCreateBlitPipeline(const RenderPassEntry& renderPassEntry); VkPipeline GetOrCreateBlitPipeline(const RenderPassEntry& renderPassEntry);
Bool GenerateDepthMipmapWithShader(FrameContext::FrameData& frame, Bool GenerateDepthMipmapWithShader(FrameContext::FrameData& frame,
MG_State::GLState::ITextureObject& texture, MG_State::GLState::ITextureObject& texture,
@@ -595,6 +717,10 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const PhysicalDevice& compareWithDevice, const PhysicalDevice& compareWithDevice,
PhysicalDevice& outBetterDevice); PhysicalDevice& outBetterDevice);
static constexpr const char* s_validationLayerNames[] = {"VK_LAYER_KHRONOS_validation"}; static constexpr const char* s_validationLayerNames[] = {"VK_LAYER_KHRONOS_validation"};
// VK_KHR_image_format_list: lets MUTABLE_FORMAT images declare their exact view-format
// set so the driver can keep bandwidth compression (see CreateLogicalDeviceAndQueues).
Bool m_imageFormatListExtensionEnabled = false;
static constexpr const char* s_deviceExtensionNames[] = {VK_KHR_SWAPCHAIN_EXTENSION_NAME}; static constexpr const char* s_deviceExtensionNames[] = {VK_KHR_SWAPCHAIN_EXTENSION_NAME};
static Bool CheckValidationLayerSupport(); static Bool CheckValidationLayerSupport();
+33
View File
@@ -24,6 +24,7 @@ namespace MobileGL::MG_Impl::CGLImpl {
GLint Samples = 0; GLint Samples = 0;
GLint Profile = kCGLOGLPVersion_3_2_Core; GLint Profile = kCGLOGLPVersion_3_2_Core;
GLint RendererId = 0x4d474c; GLint RendererId = 0x4d474c;
GLint DisplayMask = 0;
}; };
struct ContextObject { struct ContextObject {
@@ -134,6 +135,9 @@ namespace MobileGL::MG_Impl::CGLImpl {
case kCGLPFARendererID: case kCGLPFARendererID:
pixelFormat.RendererId = value; pixelFormat.RendererId = value;
break; break;
case kCGLPFADisplayMask:
pixelFormat.DisplayMask = value;
break;
default: default:
break; break;
} }
@@ -343,6 +347,9 @@ namespace MobileGL::MG_Impl::CGLImpl {
case kCGLPFARendererID: case kCGLPFARendererID:
*value = pixelFormat->RendererId; *value = pixelFormat->RendererId;
return kCGLNoError; return kCGLNoError;
case kCGLPFADisplayMask:
*value = pixelFormat->DisplayMask;
return kCGLNoError;
case kCGLPFAOpenGLProfile: case kCGLPFAOpenGLProfile:
*value = pixelFormat->Profile; *value = pixelFormat->Profile;
return kCGLNoError; return kCGLNoError;
@@ -481,6 +488,32 @@ namespace MobileGL::MG_Impl::CGLImpl {
return it == currentContexts.end() ? nullptr : it->second; return it == currentContexts.end() ? nullptr : it->second;
} }
CGLError SetVirtualScreen(CGLContextObj ctx, GLint screen) {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto* object = TryGetContext(ctx);
if (!object) {
return kCGLBadContext;
}
if (screen != 0) {
return kCGLBadValue;
}
object->VirtualScreen = screen;
return kCGLNoError;
}
CGLError GetVirtualScreen(CGLContextObj ctx, GLint* screen) {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto* object = TryGetContext(ctx);
if (!object) {
return kCGLBadContext;
}
if (!screen) {
return kCGLBadAddress;
}
*screen = object->VirtualScreen;
return kCGLNoError;
}
CGLError SetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params) { CGLError SetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params) {
const std::lock_guard<std::recursive_mutex> lock(RegistryMutex()); const std::lock_guard<std::recursive_mutex> lock(RegistryMutex());
auto* object = TryGetContext(ctx); auto* object = TryGetContext(ctx);
+2
View File
@@ -32,6 +32,8 @@ namespace MobileGL::MG_Impl::CGLImpl {
CGLError SetCurrentContext(CGLContextObj ctx); CGLError SetCurrentContext(CGLContextObj ctx);
CGLContextObj GetCurrentContext(); CGLContextObj GetCurrentContext();
CGLError SetVirtualScreen(CGLContextObj ctx, GLint screen);
CGLError GetVirtualScreen(CGLContextObj ctx, GLint* screen);
CGLError SetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params); CGLError SetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params);
CGLError GetParameter(CGLContextObj ctx, CGLContextParameter pname, GLint* params); CGLError GetParameter(CGLContextObj ctx, CGLContextParameter pname, GLint* params);
CGLError UpdateContext(CGLContextObj ctx); CGLError UpdateContext(CGLContextObj ctx);
@@ -71,6 +71,14 @@ MOBILEGL_CGL_API CGLContextObj CGLGetCurrentContext(void) {
return MobileGL::MG_Impl::CGLImpl::GetCurrentContext(); return MobileGL::MG_Impl::CGLImpl::GetCurrentContext();
} }
MOBILEGL_CGL_API CGLError CGLSetVirtualScreen(CGLContextObj ctx, GLint screen) {
return MobileGL::MG_Impl::CGLImpl::SetVirtualScreen(ctx, screen);
}
MOBILEGL_CGL_API CGLError CGLGetVirtualScreen(CGLContextObj ctx, GLint* screen) {
return MobileGL::MG_Impl::CGLImpl::GetVirtualScreen(ctx, screen);
}
MOBILEGL_CGL_API CGLError CGLSetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params) { MOBILEGL_CGL_API CGLError CGLSetParameter(CGLContextObj ctx, CGLContextParameter pname, const GLint* params) {
return MobileGL::MG_Impl::CGLImpl::SetParameter(ctx, pname, params); return MobileGL::MG_Impl::CGLImpl::SetParameter(ctx, pname, params);
} }
@@ -10,8 +10,12 @@
#if defined(__APPLE__) #if defined(__APPLE__)
#include "MG_Impl/CGLImpl/CGLImpl.h"
#include "MG_Impl/GetProcAddress.h" #include "MG_Impl/GetProcAddress.h"
#include <CoreGraphics/CoreGraphics.h>
#include <CoreVideo/CVDisplayLink.h>
#include <cstdint>
#include <dlfcn.h> #include <dlfcn.h>
namespace { namespace {
@@ -47,10 +51,52 @@ namespace {
return dlsym(handle, symbol); return dlsym(handle, symbol);
} }
CGDirectDisplayID DisplayForMask(GLint displayMask) {
constexpr std::uint32_t MaxDisplays = sizeof(CGOpenGLDisplayMask) * 8;
CGDirectDisplayID displays[MaxDisplays] = {};
std::uint32_t displayCount = 0;
if (displayMask != 0 &&
CGGetActiveDisplayList(MaxDisplays, displays, &displayCount) == kCGErrorSuccess) {
const auto mask = static_cast<CGOpenGLDisplayMask>(displayMask);
for (std::uint32_t i = 0; i < displayCount; ++i) {
if ((CGDisplayIDToOpenGLDisplayMask(displays[i]) & mask) != 0) {
return displays[i];
}
}
}
return CGMainDisplayID();
}
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wdeprecated-declarations"
CVReturn MobileGLCVDisplayLinkSetCurrentCGDisplayFromOpenGLContext(
CVDisplayLinkRef displayLink,
CGLContextObj context,
CGLPixelFormatObj pixelFormat) {
GLint virtualScreen = 0;
if (MobileGL::MG_Impl::CGLImpl::GetVirtualScreen(context, &virtualScreen) == kCGLNoError) {
GLint displayMask = 0;
if (!displayLink ||
MobileGL::MG_Impl::CGLImpl::DescribePixelFormat(
pixelFormat, virtualScreen, kCGLPFADisplayMask, &displayMask) != kCGLNoError) {
return kCVReturnInvalidArgument;
}
return CVDisplayLinkSetCurrentCGDisplay(displayLink, DisplayForMask(displayMask));
}
using OriginalFunction = CVReturn (*)(CVDisplayLinkRef, CGLContextObj, CGLPixelFormatObj);
static const auto original = reinterpret_cast<OriginalFunction>(
dlsym(RTLD_NEXT, "CVDisplayLinkSetCurrentCGDisplayFromOpenGLContext"));
return original ? original(displayLink, context, pixelFormat) : kCVReturnError;
}
__attribute__((used)) static const DyldInterposeEntry kMobileGLDyldInterpose[] __attribute__((used)) static const DyldInterposeEntry kMobileGLDyldInterpose[]
__attribute__((section("__DATA,__interpose"))) = { __attribute__((section("__DATA,__interpose"))) = {
{reinterpret_cast<const void*>(MobileGLDlsym), reinterpret_cast<const void*>(dlsym)}, {reinterpret_cast<const void*>(MobileGLDlsym), reinterpret_cast<const void*>(dlsym)},
{reinterpret_cast<const void*>(MobileGLCVDisplayLinkSetCurrentCGDisplayFromOpenGLContext),
reinterpret_cast<const void*>(CVDisplayLinkSetCurrentCGDisplayFromOpenGLContext)},
}; };
#pragma clang diagnostic pop
} // namespace } // namespace
#endif #endif
@@ -0,0 +1,10 @@
# Public CGL entry points.
_CGL*
# Public EGL entry points.
_egl*
# Public OpenGL and GLX entry points. OpenGL function names always use an
# uppercase letter or digit after the "gl" prefix; excluding lowercase here
# deliberately prevents glslang_* from matching this pattern.
_gl[A-Z0-9]*
@@ -465,6 +465,14 @@ namespace MobileGL::MG_Impl::GLImpl {
case GL_STENCIL_TEST: case GL_STENCIL_TEST:
*params = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::StencilTest) ? GL_TRUE : GL_FALSE; *params = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::StencilTest) ? GL_TRUE : GL_FALSE;
return; return;
case GL_MIN_FRAGMENT_INTERPOLATION_OFFSET:
case GL_MAX_FRAGMENT_INTERPOLATION_OFFSET:
case GL_FRAGMENT_INTERPOLATION_OFFSET_BITS: {
GLfloat value = 0.0f;
GetFloatv(pname, &value);
*params = value != 0.0f ? GL_TRUE : GL_FALSE;
return;
}
default: default:
break; break;
} }
@@ -520,6 +528,19 @@ namespace MobileGL::MG_Impl::GLImpl {
params[1] = dynamicParameters.ViewportBoundsRangeMax; params[1] = dynamicParameters.ViewportBoundsRangeMax;
return; return;
} }
case GL_MIN_FRAGMENT_INTERPOLATION_OFFSET:
case GL_MAX_FRAGMENT_INTERPOLATION_OFFSET:
case GL_FRAGMENT_INTERPOLATION_OFFSET_BITS: {
const auto& dynamicParameters = MG_Backend::pActiveBackendObject->GetDynamicParameters();
if (pname == GL_MIN_FRAGMENT_INTERPOLATION_OFFSET) {
params[0] = dynamicParameters.MinFragmentInterpolationOffset;
} else if (pname == GL_MAX_FRAGMENT_INTERPOLATION_OFFSET) {
params[0] = dynamicParameters.MaxFragmentInterpolationOffset;
} else {
params[0] = static_cast<GLfloat>(dynamicParameters.FragmentInterpolationOffsetBits);
}
return;
}
case GL_DEPTH_CLEAR_VALUE: case GL_DEPTH_CLEAR_VALUE:
params[0] = MG_State::pGLContext->GetClearDepth(); params[0] = MG_State::pGLContext->GetClearDepth();
return; return;
@@ -1957,6 +1978,15 @@ namespace MobileGL::MG_Impl::GLImpl {
case GL_SUBPIXEL_BITS: case GL_SUBPIXEL_BITS:
*params = std::max(dynamicParameters.ViewportSubpixelBits, kFrontendSubpixelBits); *params = std::max(dynamicParameters.ViewportSubpixelBits, kFrontendSubpixelBits);
break; break;
case GL_MIN_FRAGMENT_INTERPOLATION_OFFSET:
*params = static_cast<GLint>(std::lround(dynamicParameters.MinFragmentInterpolationOffset));
break;
case GL_MAX_FRAGMENT_INTERPOLATION_OFFSET:
*params = static_cast<GLint>(std::lround(dynamicParameters.MaxFragmentInterpolationOffset));
break;
case GL_FRAGMENT_INTERPOLATION_OFFSET_BITS:
*params = dynamicParameters.FragmentInterpolationOffsetBits;
break;
case GL_UNIFORM_BUFFER_OFFSET_ALIGNMENT: case GL_UNIFORM_BUFFER_OFFSET_ALIGNMENT:
*params = static_cast<Int>(dynamicParameters.UniformBufferOffsetAlignment); *params = static_cast<Int>(dynamicParameters.UniformBufferOffsetAlignment);
break; break;
+29
View File
@@ -133,4 +133,33 @@ namespace MobileGL::MG_Impl::GLImpl {
values[0] = value; values[0] = value;
} }
} }
void DestroyAllSyncObjects() {
// Detach the registry under the lock, release outside it. Entries the app
// already deleted were erased by DeleteSync, so nothing here double-frees;
// a DeleteSync racing this sweep finds an empty registry and returns. A
// thread still blocked inside ClientWaitSync/GetSynciv during teardown
// holds a raw SyncObject* these deletes invalidate - the same undefined
// race an app-driven DeleteSync already has.
UnorderedMap<GLsync, SyncObject*> orphans;
{
const std::lock_guard<std::mutex> lock(g_syncObjectsMutex);
orphans.swap(g_liveSyncObjects);
}
if (orphans.empty()) {
return;
}
// Both backends' DeleteSync only free the heap wrapper once their GL
// context/renderer is gone (generation/current-thread guards), so this is
// safe after the backend has released its EGL resources - but not after
// the function table itself is cleared.
const auto backendDeleteSync = MG_Backend::gBackendFunctionsTable.GL.DeleteSync;
for (const auto& [_, syncObject] : orphans) {
if (backendDeleteSync && syncObject->backendHandle) {
backendDeleteSync(syncObject->backendHandle);
}
delete syncObject;
}
MGLOG_D("DestroyAllSyncObjects: reclaimed %zu sync object(s) the app left undeleted", orphans.size());
}
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
+8
View File
@@ -16,4 +16,12 @@ namespace MobileGL::MG_Impl::GLImpl {
void WaitSync(GLsync sync, GLbitfield flags, GLuint64 timeout); void WaitSync(GLsync sync, GLbitfield flags, GLuint64 timeout);
void DeleteSync(GLsync sync); void DeleteSync(GLsync sync);
void GetSynciv(GLsync sync, GLenum pname, GLsizei bufSize, GLsizei* length, GLint* values); void GetSynciv(GLsync sync, GLenum pname, GLsizei bufSize, GLsizei* length, GLint* values);
// Destroys every still-registered sync object exactly as DeleteSync would.
// GL requires syncs to die with their context; called only from full library
// teardown (DestroyImpl), where no context survives on any thread, so the
// process-global registry can be drained wholesale. Must run while the
// backend function table is still populated: each backend handle has to be
// released by the backend that created it, never by a later re-initialized
// one.
void DestroyAllSyncObjects();
} // namespace MobileGL::MG_Impl::GLImpl } // namespace MobileGL::MG_Impl::GLImpl
+2
View File
@@ -85,6 +85,8 @@ namespace MobileGL::MG_Impl {
GETPROC(CGLGetPixelFormat, name); GETPROC(CGLGetPixelFormat, name);
GETPROC(CGLSetCurrentContext, name); GETPROC(CGLSetCurrentContext, name);
GETPROC(CGLGetCurrentContext, name); GETPROC(CGLGetCurrentContext, name);
GETPROC(CGLSetVirtualScreen, name);
GETPROC(CGLGetVirtualScreen, name);
GETPROC(CGLSetParameter, name); GETPROC(CGLSetParameter, name);
GETPROC(CGLGetParameter, name); GETPROC(CGLGetParameter, name);
GETPROC(CGLUpdateContext, name); GETPROC(CGLUpdateContext, name);
+36 -4
View File
@@ -29,10 +29,19 @@ namespace MobileGL::MG_Impl::NSOpenGLImpl {
char kContextViewKey; char kContextViewKey;
char kContextLayerKey; char kContextLayerKey;
std::once_flag g_installOnce;
IMP g_pixelFormatDealloc = nullptr; IMP g_pixelFormatDealloc = nullptr;
IMP g_contextDealloc = nullptr; IMP g_contextDealloc = nullptr;
std::mutex& HookInstallMutex() {
static auto* mutex = new std::mutex();
return *mutex;
}
Bool& HooksInstalled() {
static auto* installed = new Bool(false);
return *installed;
}
template <typename Fn> template <typename Fn>
Fn ObjcMsgSend() { Fn ObjcMsgSend() {
return reinterpret_cast<Fn>(objc_msgSend); return reinterpret_cast<Fn>(objc_msgSend);
@@ -431,12 +440,12 @@ namespace MobileGL::MG_Impl::NSOpenGLImpl {
method_setImplementation(method, replacement); method_setImplementation(method, replacement);
} }
void InstallHooksOnce() { Bool InstallHooksOnce() {
Class pixelFormatClass = objc_getClass("NSOpenGLPixelFormat"); Class pixelFormatClass = objc_getClass("NSOpenGLPixelFormat");
Class contextClass = objc_getClass("NSOpenGLContext"); Class contextClass = objc_getClass("NSOpenGLContext");
if (!pixelFormatClass || !contextClass) { if (!pixelFormatClass || !contextClass) {
MGLOG_W("NSOpenGLImpl: NSOpenGL classes are not loaded; hooks not installed"); MGLOG_W("NSOpenGLImpl: NSOpenGL classes are not loaded; hooks not installed");
return; return false;
} }
ReplaceInstanceMethod(pixelFormatClass, "initWithAttributes:", ReplaceInstanceMethod(pixelFormatClass, "initWithAttributes:",
@@ -471,11 +480,34 @@ namespace MobileGL::MG_Impl::NSOpenGLImpl {
ReplaceInstanceMethod(contextClass, "dealloc", reinterpret_cast<IMP>(ContextDealloc), &g_contextDealloc); ReplaceInstanceMethod(contextClass, "dealloc", reinterpret_cast<IMP>(ContextDealloc), &g_contextDealloc);
MGLOG_I("NSOpenGLImpl hooks installed"); MGLOG_I("NSOpenGLImpl hooks installed");
return true;
} }
} // namespace } // namespace
void InstallHooks() { void InstallHooks() {
std::call_once(g_installOnce, InstallHooksOnce); const std::lock_guard<std::mutex> lock(HookInstallMutex());
if (!HooksInstalled()) {
// Do not permanently consume the install attempt when the OpenGL
// framework has not registered its Objective-C classes yet. The
// dyld bootstrap normally runs after framework dependencies, but
// an explicitly loaded/static-linked MobileGL can arrive earlier.
HooksInstalled() = InstallHooksOnce();
}
} }
} // namespace MobileGL::MG_Impl::NSOpenGLImpl } // namespace MobileGL::MG_Impl::NSOpenGLImpl
namespace {
// SDL's Cocoa backend creates NSOpenGLPixelFormat/NSOpenGLContext before
// its first dlsym("glGetString") or other MobileGL host-API call. Install
// only the lightweight Objective-C dispatch hooks while the injected dylib
// is loading so those first Cocoa objects are routed through CGLImpl. The
// hooked context constructor reaches EGLImpl::GetDisplay(), which performs
// the full, thread-safe MobileGL initialization outside this bootstrap.
//
// There is intentionally no matching destructor: backend teardown remains
// owned by the EGL lifecycle and process-exit globals remain leak-at-exit.
__attribute__((constructor)) void BootstrapNSOpenGLHooks() {
MobileGL::MG_Impl::NSOpenGLImpl::InstallHooks();
}
} // namespace
#endif #endif
@@ -334,16 +334,32 @@ namespace MobileGL::MG_State::GLState {
// draw. The memo is keyed by (backendStateVersion, flags); ResetLinkArtifacts and // draw. The memo is keyed by (backendStateVersion, flags); ResetLinkArtifacts and
// the binding setters below invalidate it by bumping m_backendStateVersion. // the binding setters below invalidate it by bumping m_backendStateVersion.
Bool GetBackendHashMemo(Uint flags, Uint64& outHash) const { Bool GetBackendHashMemo(Uint flags, Uint64& outHash) const {
if (m_backendHashMemoVersion != m_backendStateVersion || m_backendHashMemoFlags != flags) { if (m_backendHashMemoVersion != m_backendStateVersion) return false;
return false; for (const auto& slot : m_backendHashMemoSlots) {
} if (slot.valid && slot.flags == flags) {
outHash = m_backendHashMemo; outHash = slot.hash;
return true; return true;
} }
}
return false;
}
void SetBackendHashMemo(Uint flags, Uint64 hash) const { void SetBackendHashMemo(Uint flags, Uint64 hash) const {
m_backendHashMemo = hash; if (m_backendHashMemoVersion != m_backendStateVersion) {
for (auto& slot : m_backendHashMemoSlots) slot.valid = false;
m_backendHashMemoVersion = m_backendStateVersion; m_backendHashMemoVersion = m_backendStateVersion;
m_backendHashMemoFlags = flags; m_backendHashMemoNextSlot = 0;
}
for (auto& slot : m_backendHashMemoSlots) {
if (slot.valid && slot.flags == flags) {
slot.hash = hash;
return;
}
}
auto& slot = m_backendHashMemoSlots[m_backendHashMemoNextSlot];
slot.flags = flags;
slot.hash = hash;
slot.valid = true;
m_backendHashMemoNextSlot = (m_backendHashMemoNextSlot + 1) % kBackendHashMemoSlotCount;
} }
void SetUniformSamplerOrImageUnitIndex(Uint location, Int unit) { void SetUniformSamplerOrImageUnitIndex(Uint location, Int unit) {
@@ -527,10 +543,19 @@ namespace MobileGL::MG_State::GLState {
Uint32 m_backendStateVersion = 0; Uint32 m_backendStateVersion = 0;
// Backend-owned content-hash memo (see GetBackendHashMemo): valid only while // Backend-owned content-hash memo (see GetBackendHashMemo): valid only while
// m_backendStateVersion and the compile flags match the recorded values. // m_backendStateVersion matches. Several slots, not one: a backend may resolve the same
mutable Uint64 m_backendHashMemo = 0; // program under more than one compile-flag set within a frame (surface rotation, and the
// explicit-LOD sampling variant), and a single slot would then miss on every lookup and
// re-hash the program's whole SPIR-V once per draw.
static constexpr SizeT kBackendHashMemoSlotCount = 4;
struct BackendHashMemoSlot {
Uint64 hash = 0;
Uint flags = 0;
Bool valid = false;
};
mutable Array<BackendHashMemoSlot, kBackendHashMemoSlotCount> m_backendHashMemoSlots{};
mutable SizeT m_backendHashMemoNextSlot = 0;
mutable Uint32 m_backendHashMemoVersion = ~0u; mutable Uint32 m_backendHashMemoVersion = ~0u;
mutable Uint m_backendHashMemoFlags = 0;
Uint32 m_uboContentVersion = 0; Uint32 m_uboContentVersion = 0;
Uint32 m_linkVersion = 0; Uint32 m_linkVersion = 0;
}; };
@@ -15,7 +15,11 @@
#include <MG_Util/Math/VectorTypes.h> #include <MG_Util/Math/VectorTypes.h>
namespace MobileGL::MG_State::GLState { namespace MobileGL::MG_State::GLState {
class ITextureObject { // Texture objects are always SharedPtr-owned (TextureState creates every instance via
// MakeShared, including the per-target default objects). enable_shared_from_this lets
// backends that only receive a reference (e.g. syncing a name-deleted texture kept
// alive by an FBO attachment) still register a weak liveness reference for GC.
class ITextureObject : public std::enable_shared_from_this<ITextureObject> {
public: public:
using TargetEnum = TextureTarget; using TargetEnum = TextureTarget;
virtual ~ITextureObject() = default; virtual ~ITextureObject() = default;
@@ -104,6 +104,23 @@ namespace MobileGL {
m_backendHashMemoVersion = m_configVersion; m_backendHashMemoVersion = m_configVersion;
} }
// Backend-owned resolved-state memo: an opaque pointer into the
// backend's vertex-input-state cache plus the cache's eviction
// epoch, valid while the config version matches. Lets the
// per-draw path skip the content hash AND the cache lookup; the
// epoch guards against the cache evicting the pointee.
Bool GetBackendStateMemo(const void*& outState, Uint64& outEpoch) const {
if (m_backendStateMemoVersion != m_configVersion) return false;
outState = m_backendStateMemo;
outEpoch = m_backendStateMemoEpoch;
return true;
}
void SetBackendStateMemo(const void* state, Uint64 epoch) const {
m_backendStateMemo = state;
m_backendStateMemoEpoch = epoch;
m_backendStateMemoVersion = m_configVersion;
}
private: private:
void BumpAttributeFormatVersion(Uint index); void BumpAttributeFormatVersion(Uint index);
void BumpAttributeBufferVersion(Uint index); void BumpAttributeBufferVersion(Uint index);
@@ -137,6 +154,9 @@ namespace MobileGL {
Uint32 m_configVersion = 0; Uint32 m_configVersion = 0;
mutable Uint64 m_backendHashMemo = 0; mutable Uint64 m_backendHashMemo = 0;
mutable Uint32 m_backendHashMemoVersion = ~0u; mutable Uint32 m_backendHashMemoVersion = ~0u;
mutable const void* m_backendStateMemo = nullptr;
mutable Uint64 m_backendStateMemoEpoch = 0;
mutable Uint32 m_backendStateMemoVersion = ~0u;
}; };
} // namespace GLState } // namespace GLState
} // namespace MG_State } // namespace MG_State
@@ -34,6 +34,11 @@ namespace {
GLint maxFragmentImageUniforms = 4; GLint maxFragmentImageUniforms = 4;
GLint maxComputeImageUniforms = 5; GLint maxComputeImageUniforms = 5;
bool maxGeometryImageUniformsQueried = false; bool maxGeometryImageUniformsQueried = false;
GLfloat minFragmentInterpolationOffset = -0.75f;
GLfloat maxFragmentInterpolationOffset = 0.625f;
GLint fragmentInterpolationOffsetBits = 6;
bool fragmentInterpolationLimitsQueried = false;
bool fragmentInterpolationQueryRaisesError = false;
// Emulates ANGLE-on-Vulkan: the draw reads the indirect command's // Emulates ANGLE-on-Vulkan: the draw reads the indirect command's
// baseInstance word and exposes it through gl_InstanceID. // baseInstance word and exposes it through gl_InstanceID.
bool drawLeaksBaseInstanceWord = false; bool drawLeaksBaseInstanceWord = false;
@@ -108,6 +113,14 @@ namespace {
case GL_MAX_COMPUTE_IMAGE_UNIFORMS: case GL_MAX_COMPUTE_IMAGE_UNIFORMS:
*data = g_fake.maxComputeImageUniforms; *data = g_fake.maxComputeImageUniforms;
break; break;
case GL_FRAGMENT_INTERPOLATION_OFFSET_BITS:
g_fake.fragmentInterpolationLimitsQueried = true;
if (g_fake.fragmentInterpolationQueryRaisesError) {
g_fake.pendingError = GL_INVALID_ENUM;
} else {
*data = g_fake.fragmentInterpolationOffsetBits;
}
break;
// FillInGLESCapabilities reads the context version before running the // FillInGLESCapabilities reads the context version before running the
// baseInstance probe, which requires ES >= 3.1. // baseInstance probe, which requires ES >= 3.1.
case GL_MAJOR_VERSION: case GL_MAJOR_VERSION:
@@ -155,6 +168,22 @@ namespace {
g_fake.maxTextureMaxAnisotropyQueried = true; g_fake.maxTextureMaxAnisotropyQueried = true;
data[0] = g_fake.maxTextureMaxAnisotropy; data[0] = g_fake.maxTextureMaxAnisotropy;
break; break;
case GL_MIN_FRAGMENT_INTERPOLATION_OFFSET:
g_fake.fragmentInterpolationLimitsQueried = true;
if (g_fake.fragmentInterpolationQueryRaisesError) {
g_fake.pendingError = GL_INVALID_ENUM;
} else {
data[0] = g_fake.minFragmentInterpolationOffset;
}
break;
case GL_MAX_FRAGMENT_INTERPOLATION_OFFSET:
g_fake.fragmentInterpolationLimitsQueried = true;
if (g_fake.fragmentInterpolationQueryRaisesError) {
g_fake.pendingError = GL_INVALID_ENUM;
} else {
data[0] = g_fake.maxFragmentInterpolationOffset;
}
break;
// Two-component range queries. // Two-component range queries.
case GL_ALIASED_LINE_WIDTH_RANGE: case GL_ALIASED_LINE_WIDTH_RANGE:
case GL_SMOOTH_LINE_WIDTH_RANGE: case GL_SMOOTH_LINE_WIDTH_RANGE:
@@ -462,6 +491,52 @@ TEST(ImageUniformCapabilities, QueriesRealPerStageLimitsAndConservativelyGatesGe
EXPECT_TRUE(g_fake.maxGeometryImageUniformsQueried); EXPECT_TRUE(g_fake.maxGeometryImageUniformsQueried);
} }
TEST(FragmentInterpolationCapabilities, QueriesOnlyWhenSupportedAndPreservesDriverLimits) {
const auto funcs = MakeFakeGLESFunctions();
ResetFakeDriver();
g_fake.maxVertexSsboBlocks = 0;
MobileGL::MG_External::GLESCapabilities unsupportedCaps;
ASSERT_TRUE(MobileGL::MG_Util::BackendLoader::FillInGLESCapabilities(unsupportedCaps, funcs));
EXPECT_FALSE(unsupportedCaps.SupportsShaderMultisampleInterpolation);
EXPECT_FALSE(g_fake.fragmentInterpolationLimitsQueried);
EXPECT_FLOAT_EQ(unsupportedCaps.MinFragmentInterpolationOffset, -0.5f);
EXPECT_FLOAT_EQ(unsupportedCaps.MaxFragmentInterpolationOffset, 0.4375f);
EXPECT_EQ(unsupportedCaps.FragmentInterpolationOffsetBits, 4);
ResetFakeDriver();
g_fake.maxVertexSsboBlocks = 0;
g_fake.extensions.emplace_back("GL_OES_shader_multisample_interpolation");
// A stale error from an earlier capability probe must not make the optional
// interpolation query look like it failed.
g_fake.pendingError = GL_INVALID_OPERATION;
MobileGL::MG_External::GLESCapabilities supportedCaps;
ASSERT_TRUE(MobileGL::MG_Util::BackendLoader::FillInGLESCapabilities(supportedCaps, funcs));
EXPECT_TRUE(supportedCaps.SupportsShaderMultisampleInterpolation);
EXPECT_TRUE(g_fake.fragmentInterpolationLimitsQueried);
EXPECT_FLOAT_EQ(supportedCaps.MinFragmentInterpolationOffset, g_fake.minFragmentInterpolationOffset);
EXPECT_FLOAT_EQ(supportedCaps.MaxFragmentInterpolationOffset, g_fake.maxFragmentInterpolationOffset);
EXPECT_EQ(supportedCaps.FragmentInterpolationOffsetBits, g_fake.fragmentInterpolationOffsetBits);
EXPECT_EQ(funcs.glGetError(), GL_NO_ERROR);
}
TEST(FragmentInterpolationCapabilities, QueryErrorIsDrainedAndFallsBackToCoreMinimums) {
ResetFakeDriver();
g_fake.maxVertexSsboBlocks = 0;
g_fake.extensions.emplace_back("GL_OES_shader_multisample_interpolation");
g_fake.fragmentInterpolationQueryRaisesError = true;
const auto funcs = MakeFakeGLESFunctions();
MobileGL::MG_External::GLESCapabilities caps;
ASSERT_TRUE(MobileGL::MG_Util::BackendLoader::FillInGLESCapabilities(caps, funcs));
EXPECT_TRUE(g_fake.fragmentInterpolationLimitsQueried);
EXPECT_FLOAT_EQ(caps.MinFragmentInterpolationOffset, -0.5f);
EXPECT_FLOAT_EQ(caps.MaxFragmentInterpolationOffset, 0.4375f);
EXPECT_EQ(caps.FragmentInterpolationOffsetBits, 4);
EXPECT_EQ(funcs.glGetError(), GL_NO_ERROR);
}
// The extension string is what apps gate on (LWJGL builds GLCapabilities from it), so advertising // The extension string is what apps gate on (LWJGL builds GLCapabilities from it), so advertising
// it on a driver that cannot filter anisotropically would leave them silently on trilinear. // it on a driver that cannot filter anisotropically would leave them silently on trilinear.
TEST(TextureAnisotropyCapabilities, ExtensionIsAdvertisedOnlyWhenTheHostDriverSupportsIt) { TEST(TextureAnisotropyCapabilities, ExtensionIsAdvertisedOnlyWhenTheHostDriverSupportsIt) {
+151 -6
View File
@@ -33,6 +33,7 @@
#include <MG_Util/ShaderTranspiler/ShaderCompiler.h> #include <MG_Util/ShaderTranspiler/ShaderCompiler.h>
#include <MG_Util/ShaderTranspiler/ShaderSourceProcessor.h> #include <MG_Util/ShaderTranspiler/ShaderSourceProcessor.h>
#include <MG_Util/Debug/Log.h> #include <MG_Util/Debug/Log.h>
#include <FastSTL/UnorderedMap.h>
namespace { namespace {
class DynamicParameterBackend final : public MobileGL::MG_Backend::BackendObject { class DynamicParameterBackend final : public MobileGL::MG_Backend::BackendObject {
@@ -196,13 +197,13 @@ TEST(DirectGLESSanity, AdvertisesDepthTextureForGlmarkShadowScenes) {
EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_depth_texture), extensions.end()); EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_depth_texture), extensions.end());
} }
TEST(DirectGLESSanity, AdvertisesVoxyRequiredRenderingExtensionsWithoutRaisingGLVersion) { TEST(DirectGLESSanity, AdvertisesVoxyRequiredRenderingExtensionsAtExperimentalCTSVersion) {
MobileGL::MG_Backend::DirectGLES::BackendObject_DirectGLES backend; MobileGL::MG_Backend::DirectGLES::BackendObject_DirectGLES backend;
const auto& rendererInfo = backend.GetRendererInfo().RendererGLInfo; const auto& rendererInfo = backend.GetRendererInfo().RendererGLInfo;
const auto& extensions = rendererInfo.Extensions; const auto& extensions = rendererInfo.Extensions;
EXPECT_EQ(rendererInfo.TargetGLVersion.Major, 3); EXPECT_EQ(rendererInfo.TargetGLVersion.Major, 4);
EXPECT_EQ(rendererInfo.TargetGLVersion.Minor, 3); EXPECT_EQ(rendererInfo.TargetGLVersion.Minor, 6);
EXPECT_EQ(rendererInfo.TargetGLVersion.Patch, 0); EXPECT_EQ(rendererInfo.TargetGLVersion.Patch, 0);
EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_compute_shader), EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_compute_shader),
@@ -396,13 +397,13 @@ TEST(DirectVulkanSanity, RenderPassExtentUsesSwapchainSizeOnlyForDefaultFramebuf
MobileGL::IntVec2(512, 512)); MobileGL::IntVec2(512, 512));
} }
TEST(DirectVulkanSanity, AdvertisesVoxyRequiredRenderingExtensionsWithoutRaisingGLVersion) { TEST(DirectVulkanSanity, AdvertisesVoxyRequiredRenderingExtensionsAtExperimentalCTSVersion) {
MobileGL::MG_Backend::DirectVulkan::BackendObject_DirectVulkan backend; MobileGL::MG_Backend::DirectVulkan::BackendObject_DirectVulkan backend;
const auto& rendererInfo = backend.GetRendererInfo().RendererGLInfo; const auto& rendererInfo = backend.GetRendererInfo().RendererGLInfo;
const auto& extensions = rendererInfo.Extensions; const auto& extensions = rendererInfo.Extensions;
EXPECT_EQ(rendererInfo.TargetGLVersion.Major, 3); EXPECT_EQ(rendererInfo.TargetGLVersion.Major, 4);
EXPECT_EQ(rendererInfo.TargetGLVersion.Minor, 3); EXPECT_EQ(rendererInfo.TargetGLVersion.Minor, 6);
EXPECT_EQ(rendererInfo.TargetGLVersion.Patch, 0); EXPECT_EQ(rendererInfo.TargetGLVersion.Patch, 0);
EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_compute_shader), EXPECT_NE(std::find(extensions.begin(), extensions.end(), MobileGL::E_GL_ARB_compute_shader),
@@ -515,6 +516,50 @@ TEST(DirectGLESSanity, PreservesHostPerStageImageUniformLimits) {
EXPECT_EQ(params.MaxComputeImageUniforms, 5); EXPECT_EQ(params.MaxComputeImageUniforms, 5);
} }
TEST(FragmentInterpolationCapabilities, PlumbsGLESAndBothVulkanPropertyPaths) {
using namespace MobileGL;
MG_External::GLESCapabilities glesCaps;
glesCaps.MinFragmentInterpolationOffset = -0.75f;
glesCaps.MaxFragmentInterpolationOffset = 0.625f;
glesCaps.FragmentInterpolationOffsetBits = 6;
MG_Backend::DirectGLES::BackendObject_DirectGLES glesBackend;
glesBackend.ApplyGLESCapabilitiesForTesting(glesCaps);
EXPECT_FLOAT_EQ(glesBackend.GetDynamicParameters().MinFragmentInterpolationOffset, -0.75f);
EXPECT_FLOAT_EQ(glesBackend.GetDynamicParameters().MaxFragmentInterpolationOffset, 0.625f);
EXPECT_EQ(glesBackend.GetDynamicParameters().FragmentInterpolationOffsetBits, 6);
VkPhysicalDeviceProperties properties{};
// A common Vulkan limit pair: max is one representable 4-bit step below 0.5.
properties.limits.minInterpolationOffset = -0.5f;
properties.limits.maxInterpolationOffset = 0.4375f;
properties.limits.subPixelInterpolationOffsetBits = 4;
MG_External::VulkanCapabilities vkCaps;
MG_Util::BackendLoader::FillInVulkanCapabilities(vkCaps, properties);
EXPECT_FLOAT_EQ(vkCaps.MinFragmentInterpolationOffset, -0.5f);
EXPECT_FLOAT_EQ(vkCaps.MaxFragmentInterpolationOffset, 0.4375f);
EXPECT_EQ(vkCaps.FragmentInterpolationOffsetBits, 4);
vkCaps.MinFragmentInterpolationOffset = -0.875f;
vkCaps.MaxFragmentInterpolationOffset = 0.75f;
vkCaps.FragmentInterpolationOffsetBits = 7;
MG_Backend::DirectVulkan::BackendObject_DirectVulkan vkBackend;
vkBackend.ApplyVulkanCapabilitiesForTesting(vkCaps);
EXPECT_FLOAT_EQ(vkBackend.GetDynamicParameters().MinFragmentInterpolationOffset, -0.875f);
EXPECT_FLOAT_EQ(vkBackend.GetDynamicParameters().MaxFragmentInterpolationOffset, 0.75f);
EXPECT_EQ(vkBackend.GetDynamicParameters().FragmentInterpolationOffsetBits, 7);
// Invalid/zero host data cannot under-advertise the OpenGL 4 minimums.
MG_External::VulkanCapabilities invalidCaps;
invalidCaps.MinFragmentInterpolationOffset = 0.0f;
invalidCaps.MaxFragmentInterpolationOffset = 0.0f;
invalidCaps.FragmentInterpolationOffsetBits = 0;
vkBackend.ApplyVulkanCapabilitiesForTesting(invalidCaps);
EXPECT_LE(vkBackend.GetDynamicParameters().MinFragmentInterpolationOffset, -0.5f);
EXPECT_FLOAT_EQ(vkBackend.GetDynamicParameters().MaxFragmentInterpolationOffset, 0.4375f);
EXPECT_EQ(vkBackend.GetDynamicParameters().FragmentInterpolationOffsetBits, 4);
}
TEST(DirectVulkanSanity, AdvertisesSubgroupOnlyWhenVulkanReportsUsableSupport) { TEST(DirectVulkanSanity, AdvertisesSubgroupOnlyWhenVulkanReportsUsableSupport) {
using namespace MobileGL; using namespace MobileGL;
@@ -637,6 +682,48 @@ TEST(GetterSanity, ClampsMaxVertexAttribsToCurrentValueStorageCapacity) {
MG_State::pGLContext.reset(); MG_State::pGLContext.reset();
} }
TEST(GetterSanity, ReportsFragmentInterpolationLimitsForFloatAndIntegerQueries) {
using namespace MobileGL;
auto previousContext = Move(MG_State::pGLContext);
auto previousBackend = Move(MG_Backend::pActiveBackendObject);
MG_State::pGLContext = MakeUnique<MG_State::GLState::GLContext>();
MG_Backend::DynamicBackendParameters params;
params.MinFragmentInterpolationOffset = -0.75f;
params.MaxFragmentInterpolationOffset = 0.4375f;
params.FragmentInterpolationOffsetBits = 6;
MG_Backend::pActiveBackendObject = MakeUnique<DynamicParameterBackend>(params);
GLfloat floatValue = 0.0f;
MG_Impl::GLImpl::GetFloatv(GL_MIN_FRAGMENT_INTERPOLATION_OFFSET, &floatValue);
EXPECT_FLOAT_EQ(floatValue, -0.75f);
MG_Impl::GLImpl::GetFloatv(GL_MAX_FRAGMENT_INTERPOLATION_OFFSET, &floatValue);
EXPECT_FLOAT_EQ(floatValue, 0.4375f);
MG_Impl::GLImpl::GetFloatv(GL_FRAGMENT_INTERPOLATION_OFFSET_BITS, &floatValue);
EXPECT_FLOAT_EQ(floatValue, 6.0f);
GLint intValue = 0;
MG_Impl::GLImpl::GetIntegerv(GL_MIN_FRAGMENT_INTERPOLATION_OFFSET, &intValue);
EXPECT_EQ(intValue, -1);
MG_Impl::GLImpl::GetIntegerv(GL_MAX_FRAGMENT_INTERPOLATION_OFFSET, &intValue);
EXPECT_EQ(intValue, 0);
MG_Impl::GLImpl::GetIntegerv(GL_FRAGMENT_INTERPOLATION_OFFSET_BITS, &intValue);
EXPECT_EQ(intValue, 6);
GLboolean boolValue = GL_FALSE;
MG_Impl::GLImpl::GetBooleanv(GL_MIN_FRAGMENT_INTERPOLATION_OFFSET, &boolValue);
EXPECT_EQ(boolValue, GL_TRUE);
MG_Impl::GLImpl::GetBooleanv(GL_MAX_FRAGMENT_INTERPOLATION_OFFSET, &boolValue);
EXPECT_EQ(boolValue, GL_TRUE);
MG_Impl::GLImpl::GetBooleanv(GL_FRAGMENT_INTERPOLATION_OFFSET_BITS, &boolValue);
EXPECT_EQ(boolValue, GL_TRUE);
EXPECT_EQ(MG_Impl::GLImpl::GetError(), GL_NO_ERROR);
MG_Backend::pActiveBackendObject = Move(previousBackend);
MG_State::pGLContext = Move(previousContext);
}
TEST(GetterSanity, PerStageImageUniformQueriesMatchShaderCompilerLimits) { TEST(GetterSanity, PerStageImageUniformQueriesMatchShaderCompilerLimits) {
using namespace MobileGL; using namespace MobileGL;
@@ -1736,3 +1823,61 @@ TEST(DirectGLESStateGuards, DefaultFramebufferBindGoesThroughShadow) {
FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7); // must reach the driver again FramebufferImpl::BindFramebufferId(GL_DRAW_FRAMEBUFFER, 7); // must reach the driver again
EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 3u); EXPECT_EQ(mocks.log.Count("BindFramebuffer:"), 3u);
} }
// FastSTL::unordered_map::erase(iterator) regression coverage. The open-addressing
// iterator constructor snaps forward from a tombstoned slot to the successor, so
// erase must NOT advance the rebuilt iterator again: the old double-advance skipped
// one live element per erase, and erasing the element in the highest occupied
// bucket pushed the returned index past bucket_count where it never compared equal
// to end() again - erase-while-iterating sweeps (pipeline/program cache eviction)
// then ran off the bucket array and fed garbage handles to vkDestroyPipeline
// (device crash on first mass eviction during world load).
TEST(FastSTLSanity, EraseWhileIteratingVisitsEveryElementExactlyOnce) {
FastSTL::unordered_map<MobileGL::Uint64, MobileGL::Uint64> map;
constexpr MobileGL::Uint64 kCount = 1000;
for (MobileGL::Uint64 key = 0; key < kCount; ++key) {
map.emplace(key * 0x9e3779b97f4a7c15ull, key);
}
ASSERT_EQ(map.size(), kCount);
MobileGL::SizeT visited = 0;
for (auto it = map.begin(); it != map.end();) {
it = map.erase(it);
++visited;
ASSERT_LE(visited, kCount); // old code: runaway past end / skipped entries
}
EXPECT_EQ(visited, kCount);
EXPECT_EQ(map.size(), 0u);
}
TEST(FastSTLSanity, EraseReturnsTheSuccessorElement) {
FastSTL::unordered_map<MobileGL::Uint32, MobileGL::Uint32> map;
for (MobileGL::Uint32 key = 1; key <= 64; ++key) {
map.emplace(key, key);
}
// Erasing every other visited element must still visit all 64 exactly once:
// the iterator returned by erase names the very next element, not one past it.
MobileGL::SizeT visited = 0;
MobileGL::SizeT erased = 0;
for (auto it = map.begin(); it != map.end();) {
++visited;
if ((visited & 1) != 0) {
it = map.erase(it);
++erased;
} else {
++it;
}
ASSERT_LE(visited, 64u);
}
EXPECT_EQ(visited, 64u);
EXPECT_EQ(map.size(), 64u - erased);
}
TEST(FastSTLSanity, ErasingTheOnlyElementReturnsEnd) {
FastSTL::unordered_map<MobileGL::Uint32, MobileGL::Uint32> map;
map.emplace(42u, 1u);
auto next = map.erase(map.begin());
EXPECT_EQ(next, map.end());
EXPECT_TRUE(map.empty());
}
@@ -9,6 +9,7 @@
#include "Loader.h" #include "Loader.h"
#include "MG_Util/Types.h" #include "MG_Util/Types.h"
#include <Config.h> #include <Config.h>
#include <cmath>
#if defined(_WIN32) #if defined(_WIN32)
#ifndef WIN32_LEAN_AND_MEAN #ifndef WIN32_LEAN_AND_MEAN
#define WIN32_LEAN_AND_MEAN 1 #define WIN32_LEAN_AND_MEAN 1
@@ -838,8 +839,14 @@ namespace MobileGL::MG_Util::BackendLoader {
if (std::strcmp(extension, "GL_NV_shader_noperspective_interpolation") == 0) { if (std::strcmp(extension, "GL_NV_shader_noperspective_interpolation") == 0) {
caps.SupportsNoperspectiveInterpolation = true; caps.SupportsNoperspectiveInterpolation = true;
} }
if (std::strcmp(extension, "GL_OES_shader_multisample_interpolation") == 0) {
caps.SupportsShaderMultisampleInterpolation = true;
} }
} }
}
caps.SupportsShaderMultisampleInterpolation =
caps.SupportsShaderMultisampleInterpolation || caps.GLESVersion.Major > 3 ||
(caps.GLESVersion.Major == 3 && caps.GLESVersion.Minor >= 2);
// Detect optional raster/color-mask entry points by whether they loaded. glColorMaski is GLES // Detect optional raster/color-mask entry points by whether they loaded. glColorMaski is GLES
// 3.2 core (no extension string), so pointer presence is the reliable signal for all of these. // 3.2 core (no extension string), so pointer presence is the reliable signal for all of these.
@@ -900,6 +907,9 @@ namespace MobileGL::MG_Util::BackendLoader {
GLint maxColorAttachments = 8; GLint maxColorAttachments = 8;
GLint maxClipDistances = 8; GLint maxClipDistances = 8;
GLint maxViewports = 16; GLint maxViewports = 16;
GLfloat minFragmentInterpolationOffset = -0.5f;
GLfloat maxFragmentInterpolationOffset = 0.4375f;
GLint fragmentInterpolationOffsetBits = 4;
glesFuncs.glGetFloatv(GL_ALIASED_LINE_WIDTH_RANGE, aliasedLineWidthRange); glesFuncs.glGetFloatv(GL_ALIASED_LINE_WIDTH_RANGE, aliasedLineWidthRange);
glesFuncs.glGetFloatv(GL_SMOOTH_LINE_WIDTH_RANGE, smoothLineWidthRange); glesFuncs.glGetFloatv(GL_SMOOTH_LINE_WIDTH_RANGE, smoothLineWidthRange);
glesFuncs.glGetFloatv(GL_SMOOTH_LINE_WIDTH_GRANULARITY, &smoothLineWidthGranularity); glesFuncs.glGetFloatv(GL_SMOOTH_LINE_WIDTH_GRANULARITY, &smoothLineWidthGranularity);
@@ -950,6 +960,29 @@ namespace MobileGL::MG_Util::BackendLoader {
glesFuncs.glGetIntegerv(GL_MAX_VIEWPORTS, &maxViewports); glesFuncs.glGetIntegerv(GL_MAX_VIEWPORTS, &maxViewports);
glesFuncs.glGetIntegerv(GL_MAX_VIEWPORT_DIMS, maxViewportDims); glesFuncs.glGetIntegerv(GL_MAX_VIEWPORT_DIMS, maxViewportDims);
glesFuncs.glGetIntegerv(GL_VIEWPORT_SUBPIXEL_BITS, &viewportSubpixelBits); glesFuncs.glGetIntegerv(GL_VIEWPORT_SUBPIXEL_BITS, &viewportSubpixelBits);
if (caps.SupportsShaderMultisampleInterpolation && glesFuncs.glGetFloatv) {
const auto drainErrors = [&glesFuncs]() {
Bool hadError = false;
if (glesFuncs.glGetError) {
while (glesFuncs.glGetError() != GL_NO_ERROR) hadError = true;
}
return hadError;
};
// Isolate these optional queries from errors raised by preceding capability
// probes, then consume any query error so initialization never leaks it into
// the application's first glGetError call.
drainErrors();
glesFuncs.glGetFloatv(GL_MIN_FRAGMENT_INTERPOLATION_OFFSET, &minFragmentInterpolationOffset);
glesFuncs.glGetFloatv(GL_MAX_FRAGMENT_INTERPOLATION_OFFSET, &maxFragmentInterpolationOffset);
glesFuncs.glGetIntegerv(GL_FRAGMENT_INTERPOLATION_OFFSET_BITS, &fragmentInterpolationOffsetBits);
if (drainErrors()) {
MGLOG_W("Fragment interpolation limit query failed; using OpenGL minimums");
minFragmentInterpolationOffset = -0.5f;
maxFragmentInterpolationOffset = 0.4375f;
fragmentInterpolationOffsetBits = 4;
}
}
// Only legal to query once the extension has been seen in the loop above, hence not batched // Only legal to query once the extension has been seen in the loop above, hence not batched
// with the unconditional probes: on a driver without it this raises GL_INVALID_ENUM. // with the unconditional probes: on a driver without it this raises GL_INVALID_ENUM.
if (caps.SupportsTextureFilterAnisotropy) { if (caps.SupportsTextureFilterAnisotropy) {
@@ -1007,6 +1040,19 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.ViewportBoundsRangeMin = viewportBoundsRange[0]; caps.ViewportBoundsRangeMin = viewportBoundsRange[0];
caps.ViewportBoundsRangeMax = viewportBoundsRange[1]; caps.ViewportBoundsRangeMax = viewportBoundsRange[1];
caps.ViewportSubpixelBits = viewportSubpixelBits; caps.ViewportSubpixelBits = viewportSubpixelBits;
caps.MinFragmentInterpolationOffset =
std::isfinite(minFragmentInterpolationOffset) && minFragmentInterpolationOffset <= -0.5f
? minFragmentInterpolationOffset
: -0.5f;
caps.MaxFragmentInterpolationOffset = 0.4375f;
caps.FragmentInterpolationOffsetBits = 4;
if (fragmentInterpolationOffsetBits >= 4 && std::isfinite(maxFragmentInterpolationOffset)) {
const Float requiredMaxOffset = 0.5f - std::ldexp(1.0f, -fragmentInterpolationOffsetBits);
if (maxFragmentInterpolationOffset >= requiredMaxOffset) {
caps.MaxFragmentInterpolationOffset = maxFragmentInterpolationOffset;
caps.FragmentInterpolationOffsetBits = fragmentInterpolationOffsetBits;
}
}
MGLOG_I(" GL_ALIASED_LINE_WIDTH_RANGE: [%.3f, %.3f]", caps.AliasedLineWidthRangeMin, MGLOG_I(" GL_ALIASED_LINE_WIDTH_RANGE: [%.3f, %.3f]", caps.AliasedLineWidthRangeMin,
caps.AliasedLineWidthRangeMax); caps.AliasedLineWidthRangeMax);
MGLOG_I(" GL_SMOOTH_LINE_WIDTH_RANGE: [%.3f, %.3f]", caps.SmoothLineWidthRangeMin, MGLOG_I(" GL_SMOOTH_LINE_WIDTH_RANGE: [%.3f, %.3f]", caps.SmoothLineWidthRangeMin,
@@ -1055,6 +1055,9 @@ namespace MobileGL {
// SPIRV-Cross's `#extension ... : require` would fail to compile and MobileGL falls back // SPIRV-Cross's `#extension ... : require` would fail to compile and MobileGL falls back
// to stripping the NoPerspective decoration (smooth interpolation) via StripNoPerspectivePass. // to stripping the NoPerspective decoration (smooth interpolation) via StripNoPerspectivePass.
Bool SupportsNoperspectiveInterpolation = false; Bool SupportsNoperspectiveInterpolation = false;
// GLES 3.2 core or GL_OES_shader_multisample_interpolation exposes
// interpolateAtOffset and the three fragment-offset limit queries.
Bool SupportsShaderMultisampleInterpolation = false;
// GL_RENDERER contains "ANGLE". // GL_RENDERER contains "ANGLE".
Bool IsAngleRenderer = false; Bool IsAngleRenderer = false;
// GL_RENDERER contains both "ANGLE" and "llvmpipe". // GL_RENDERER contains both "ANGLE" and "llvmpipe".
@@ -1120,6 +1123,9 @@ namespace MobileGL {
Float ViewportBoundsRangeMin = 0.0f; Float ViewportBoundsRangeMin = 0.0f;
Float ViewportBoundsRangeMax = 0.0f; Float ViewportBoundsRangeMax = 0.0f;
Int ViewportSubpixelBits = 0; Int ViewportSubpixelBits = 0;
Float MinFragmentInterpolationOffset = -0.5f;
Float MaxFragmentInterpolationOffset = 0.4375f;
Int FragmentInterpolationOffsetBits = 4;
}; };
} // namespace MG_External } // namespace MG_External
@@ -9,6 +9,7 @@
#include "Loader.h" #include "Loader.h"
#include <Config.h> #include <Config.h>
#include <cmath>
namespace MobileGL::MG_Util::BackendLoader { namespace MobileGL::MG_Util::BackendLoader {
namespace { namespace {
@@ -40,6 +41,25 @@ namespace MobileGL::MG_Util::BackendLoader {
return MaxSampleCountFromFlags(commonFlags); return MaxSampleCountFromFlags(commonFlags);
} }
void FillFragmentInterpolationLimits(MG_External::VulkanCapabilities& caps,
const VkPhysicalDeviceLimits& limits) {
caps.MinFragmentInterpolationOffset =
std::isfinite(limits.minInterpolationOffset) && limits.minInterpolationOffset <= -0.5f
? limits.minInterpolationOffset
: -0.5f;
caps.MaxFragmentInterpolationOffset = 0.4375f;
caps.FragmentInterpolationOffsetBits = 4;
const Int bits = static_cast<Int>(limits.subPixelInterpolationOffsetBits);
if (bits >= 4 && std::isfinite(limits.maxInterpolationOffset)) {
const Float requiredMaxOffset = 0.5f - std::ldexp(1.0f, -bits);
if (limits.maxInterpolationOffset >= requiredMaxOffset) {
caps.MaxFragmentInterpolationOffset = limits.maxInterpolationOffset;
caps.FragmentInterpolationOffsetBits = bits;
}
}
}
VulkanDynamicFunctions LoadVulkanDynamicFunctions(VkInstance instance) { VulkanDynamicFunctions LoadVulkanDynamicFunctions(VkInstance instance) {
VulkanDynamicFunctions loaded{}; VulkanDynamicFunctions loaded{};
if (instance == VK_NULL_HANDLE) { if (instance == VK_NULL_HANDLE) {
@@ -171,6 +191,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.ViewportBoundsRangeMin = p.limits.viewportBoundsRange[0]; caps.ViewportBoundsRangeMin = p.limits.viewportBoundsRange[0];
caps.ViewportBoundsRangeMax = p.limits.viewportBoundsRange[1]; caps.ViewportBoundsRangeMax = p.limits.viewportBoundsRange[1];
caps.ViewportSubpixelBits = static_cast<Int>(p.limits.viewportSubPixelBits); caps.ViewportSubpixelBits = static_cast<Int>(p.limits.viewportSubPixelBits);
FillFragmentInterpolationLimits(caps, p.limits);
VkPhysicalDeviceFeatures supportedFeatures{}; VkPhysicalDeviceFeatures supportedFeatures{};
vkGetPhysicalDeviceFeatures(physicalDevice, &supportedFeatures); vkGetPhysicalDeviceFeatures(physicalDevice, &supportedFeatures);
@@ -261,6 +282,7 @@ namespace MobileGL::MG_Util::BackendLoader {
caps.ViewportBoundsRangeMin = properties.limits.viewportBoundsRange[0]; caps.ViewportBoundsRangeMin = properties.limits.viewportBoundsRange[0];
caps.ViewportBoundsRangeMax = properties.limits.viewportBoundsRange[1]; caps.ViewportBoundsRangeMax = properties.limits.viewportBoundsRange[1];
caps.ViewportSubpixelBits = static_cast<Int>(properties.limits.viewportSubPixelBits); caps.ViewportSubpixelBits = static_cast<Int>(properties.limits.viewportSubPixelBits);
FillFragmentInterpolationLimits(caps, properties.limits);
caps.SupportsWideLines = false; caps.SupportsWideLines = false;
// This helper only receives properties, not VkPhysicalDeviceFeatures. Leave optional // This helper only receives properties, not VkPhysicalDeviceFeatures. Leave optional
// stage writes disabled rather than inferring them from descriptor limits alone. // stage writes disabled rather than inferring them from descriptor limits alone.
@@ -68,6 +68,9 @@ namespace MobileGL {
Float ViewportBoundsRangeMin = 0.0f; Float ViewportBoundsRangeMin = 0.0f;
Float ViewportBoundsRangeMax = 0.0f; Float ViewportBoundsRangeMax = 0.0f;
Int ViewportSubpixelBits = 0; Int ViewportSubpixelBits = 0;
Float MinFragmentInterpolationOffset = -0.5f;
Float MaxFragmentInterpolationOffset = 0.4375f;
Int FragmentInterpolationOffsetBits = 4;
Bool SupportsWideLines = false; Bool SupportsWideLines = false;
// Storage-image descriptors are limited per stage by // Storage-image descriptors are limited per stage by
// maxPerStageDescriptorStorageImages, but writes/atomics outside compute additionally // maxPerStageDescriptorStorageImages, but writes/atomics outside compute additionally
@@ -133,9 +133,11 @@ namespace MobileGL {
case TextureInternalFormat::RGBA8Snorm: case TextureInternalFormat::RGBA8Snorm:
return VK_FORMAT_R8G8B8A8_SNORM; return VK_FORMAT_R8G8B8A8_SNORM;
case TextureInternalFormat::RGB10A2: case TextureInternalFormat::RGB10A2:
return VK_FORMAT_A2R10G10B10_UNORM_PACK32; // GL_UNSIGNED_INT_2_10_10_10_REV puts R in bits 0-9, which is Vulkan's
// A2B10G10R10 layout - A2R10G10B10 silently swaps R and B on upload.
return VK_FORMAT_A2B10G10R10_UNORM_PACK32;
case TextureInternalFormat::RGB10A2UI: case TextureInternalFormat::RGB10A2UI:
return VK_FORMAT_A2R10G10B10_UINT_PACK32; return VK_FORMAT_A2B10G10R10_UINT_PACK32;
case TextureInternalFormat::RGBA16: case TextureInternalFormat::RGBA16:
return VK_FORMAT_R16G16B16A16_UNORM; return VK_FORMAT_R16G16B16A16_UNORM;
case TextureInternalFormat::RGBA16Snorm: case TextureInternalFormat::RGBA16Snorm:
@@ -1,78 +0,0 @@
# MobileGL POST Format Capability Tables
## Goal
Expose the format-capability results used during MobileGL backend startup in the Android plugin's driver POST screen. The screen must show the exact `Full`, `Caveat`, or `None` result for every backend, target, internal format, and capability without duplicating the backend's detection rules.
## Existing Architecture
- `DriverPost.cpp` probes the device GLES and Vulkan drivers before `MobileGL::Initialize()` and returns a `BackendPostReport` for each backend.
- `DriverPostJni.cpp` serializes those reports to JSON for `PostActivity`.
- `PostActivity` uses platform Android views and already supports collapsible check details and a collapsible raw report.
- Backend startup fills a `FormatCapabilityCache` in the DirectGLES and DirectVulkan `InitCapabilities()` paths. `FullCaps` takes precedence over `CaveatCaps`; an absent bit means `None`.
- The capability matrix contains 12 targets, 75 internal formats, and 14 capability columns per backend.
## Selected Approach
Extract callable format-probe entry points from the existing DirectGLES and DirectVulkan implementations. Backend startup and the POST will call these same functions, so their results cannot drift.
The POST will run each probe while its temporary driver resources are still valid:
- DirectGLES: after the GLES function table and capabilities have been populated, while the 1x1 pbuffer context is current.
- DirectVulkan: after selecting the physical device, while the Vulkan instance and physical device handles remain valid.
The resulting optional `FormatCapabilityCache` will be stored in each `BackendPostReport`. Failure to obtain a format table will not discard the existing POST checks or change their verdict; the UI will instead report that the format table is unavailable.
## JSON Contract
The JNI report will add an optional `formatCapabilities` object to each backend. To avoid repeating tens of thousands of status strings, the object will contain:
- one ordered capability-name array;
- one entry per target;
- one compact row per internal format containing the format name, a Full bitmask, and a Caveat bitmask.
Java resolves each cell in this order:
1. Full bit present: `Full`.
2. Otherwise Caveat bit present: `Caveat`.
3. Otherwise: `None`.
This preserves the backend's current precedence and keeps the raw JSON reasonably small.
## Android UI
Each backend section keeps its existing verdict, renderer string, and check table. A new `Format capabilities` subsection follows it.
- Each of the 11 texture targets and `Renderbuffer` is a separate, initially collapsed table.
- Target headers can be expanded independently.
- Table content is created on expansion and removed when collapsed, preventing the activity from retaining roughly 27,000 status views.
- Each expanded table is placed in a horizontal scroll container.
- The first column contains internal-format names. The remaining columns use the ordered capability names from the JSON report.
- Every status cell displays its status text and uses the conventional color mapping:
- `Full`: green background with white text.
- `Caveat`: yellow background with black text.
- `None`: red background with white text.
- Header and format-name cells use neutral dark backgrounds consistent with the existing POST theme.
- The existing raw-report toggle remains available at the end of the screen.
## Performance and Lifecycle
- The existing single-flight native POST and cached JSON behavior remains unchanged.
- Format tables are lazily materialized and discarded on collapse.
- The JSON carries bitmasks rather than repeated `Full`, `Caveat`, and `None` strings.
- Existing configuration-change handling remains unchanged.
## Validation
1. Run focused source checks and `git diff --check`.
2. Build the Android plugin APK with the repository's current Gradle workflow.
3. If an Android target is connected, install the APK and open `PostActivity`.
4. Verify both backend sections, all target toggles, horizontal scrolling, visible cell text, and the green/yellow/red mapping.
5. Confirm collapsing a table removes its generated content and expanding it recreates the same values.
## Non-Goals
- Changing the meanings of `Full`, `Caveat`, or `None`.
- Changing POST verdict rules.
- Displaying sample-count vectors in this iteration.
- Replacing the existing platform-view UI with Compose, AppCompat, or WebView.
+112
View File
@@ -0,0 +1,112 @@
# Running the OpenGL CTS against MobileGL
This directory contains two supported paths:
- Android arm64 / MobileGL EGL: the KHR-GL33 workflow documented below and in
`skills/gl-cts-on-mobilegl/SKILL.md`.
- Windows x64 / MobileGL WGL: the GL30-GL46 pipeline in
`scripts/wgl_glcts_pipeline.py`, documented by
`skills/wgl-gl-cts-on-mobilegl/SKILL.md`.
Windows prerequisites are Git, Python 3.9+, CMake, Visual Studio 2022's Desktop
C++ workload, and a Vulkan SDK visible to CMake. DirectVulkan also needs a
working Vulkan loader plus a GPU-vendor ICD and driver; the SDK alone is not a
GPU driver.
For Windows, start with:
```powershell
python tools\cts\scripts\wgl_glcts_pipeline.py --help
```
The pipeline builds MobileGL as a drop-in `opengl32.dll`, builds or reuses
`glcts.exe`, checks that WGL loaded MobileGL rather than the system driver,
resumes individual suites after crashes/timeouts, and writes Markdown plus JSON
reports below the printed `runs/<first-16-of-run-fingerprint>` directory. Its
manifest records provenance and the runner settings used to validate a resume.
## Android KHR-GL33 workflow
Goal: measure how much of the OpenGL 3.3 core-profile conformance suite MobileGL
passes, separately for each backend (`DirectGLES`, `DirectVulkan`).
## How MobileGL is reached from a test binary
MobileGL ships its own EGL implementation alongside its desktop-GL implementation
in a single `libMobileGL.so`. A plain arm64 ELF in `/data/local/tmp` can therefore
drive it with no APK and no Activity:
1. `setenv("MOBILEGL_BACKEND_TYPE", "DirectGLES"|"DirectVulkan")` **before** the
library is mapped — MobileGL parses its configuration from an ELF constructor.
2. `dlopen("libMobileGL.so")`, then `dlsym` the `egl*` and `gl*` entry points.
MobileGL exports 45 EGL symbols and the desktop GL functions directly;
`eglGetProcAddress` resolves the same set.
3. `eglBindAPI(EGL_OPENGL_API)`, choose a config with `EGL_RENDERABLE_TYPE =
EGL_OPENGL_BIT`, then `eglCreateContext` with
`EGL_CONTEXT_OPENGL_PROFILE_MASK = EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT` and
major/minor `3`/`3`.
This yields a genuine GL 3.3 core context (`GL_CONTEXT_PROFILE_MASK == 0x1`).
## Surface type, per backend
| backend | pbuffer (headless) | window |
|---|---|---|
| `DirectGLES` | works | works |
| `DirectVulkan` | **unusable** | works |
`DirectVulkan`'s pbuffer path builds a headless `VkSurfaceKHR` and so requires the
`VK_EXT_headless_surface` instance extension, which Adreno's Android driver does
not expose. It fails inside `eglMakeCurrent`, not at surface creation.
The workaround that keeps everything in a shell process: obtain a real
`ANativeWindow` from **`AImageReader`** (`AImageReader_newWithUsage` +
`AImageReader_getWindow`). It is an ordinary BufferQueue producer, so
`vkCreateAndroidSurfaceKHR` accepts it, and no Activity is involved. Register an
`onImageAvailable` listener that acquires and deletes each image — otherwise the
producer blocks once `maxImages` buffers are in flight and the next swap hangs.
## Why the suite must render into an FBO
On a window surface, `DirectVulkan`'s `glReadPixels` from the **default
framebuffer** returns all zeros, with no GL error, both before and after
`eglSwapBuffers`. `DirectGLES` on the identical window is correct, and readback
from a **user FBO is correct on both backends**.
Verified on two SoCs and two drivers, so this is MobileGL's behaviour rather than
a driver quirk:
| device | GPU | driver | default-FB | user FBO |
|---|---|---|---|---|
| Xiaomi 24129PN74C | Adreno 830 | Vulkan 1.3.284 / 512.800.46 | zeros | ok |
| Lenovo TB321FU | Adreno 750 | Vulkan 1.3.128 / 512.762.28 | zeros | ok |
dEQP verifies nearly every case through `glReadPixels`, so running it against the
default framebuffer would score `DirectVulkan` near zero for a reason unrelated to
conformance. The runs therefore use `--deqp-surface-type=fbo`, uniformly for both
backends so the two numbers stay comparable.
## Other constraints the harness must respect
- `eglMakeCurrent` requires **draw == read** and rejects `EGL_NO_SURFACE` with
`EGL_BAD_MATCH`. dEQP's `surfaceless` platform is therefore unusable, which is
why this port supplies its own `tcu::Platform`.
- MobileGL aborts during static teardown (`FORTIFY: pthread_mutex_lock called on a
destroyed mutex`) *after* all work completes. Flush and `_exit()` so the exit
code and the `.qpa` log survive.
## Contents
probe/mgprobe.c preflight gate: one backend x one surface type, checks
context version/profile and both readback paths
scripts/qpa_report.py .qpa -> pass rate, status histogram, worst groups
### Preflight
aarch64-linux-android26-clang -O1 -o mgprobe mgprobe.c -ldl -llog -landroid -lmediandk
adb push mgprobe libMobileGL.so /data/local/tmp/mgcts/
adb shell 'cd /data/local/tmp/mgcts && LD_LIBRARY_PATH=. ./mgprobe \
--backend DirectVulkan --surface imagereader --lib ./libMobileGL.so'
Exit status is 0 when a 3.3 core context came up and FBO readback is correct.
Default-framebuffer readback is reported but deliberately does not gate.
@@ -0,0 +1,109 @@
diff --git a/framework/opengl/gluFboRenderContext.cpp b/framework/opengl/gluFboRenderContext.cpp
index 588cf7d2a..0721ffee7 100644
--- a/framework/opengl/gluFboRenderContext.cpp
+++ b/framework/opengl/gluFboRenderContext.cpp
@@ -132,6 +132,7 @@ FboRenderContext::FboRenderContext(RenderContext *context, const RenderConfig &c
: m_context(context)
, m_framebuffer(0)
, m_colorBuffer(0)
+ , m_colorIsTexture(false)
, m_depthStencilBuffer(0)
, m_renderTarget()
{
@@ -151,6 +152,7 @@ FboRenderContext::FboRenderContext(const ContextFactory &factory, const RenderCo
: m_context(nullptr)
, m_framebuffer(0)
, m_colorBuffer(0)
+ , m_colorIsTexture(false)
, m_depthStencilBuffer(0)
, m_renderTarget()
{
@@ -215,19 +217,41 @@ void FboRenderContext::createFramebuffer(const RenderConfig &config)
height = (height == glu::RenderConfig::DONT_CARE) ? maxSize : height;
}
+ // MOBILEGL: allow the colour attachment to be a texture instead of a
+ // renderbuffer. MobileGL's DirectVulkan backend returns zeros when reading
+ // back a renderbuffer-attached FBO, which makes every image comparison fail
+ // for one reason and hides everything else. Setting
+ // MOBILEGL_CTS_FBO_COLOR_TEXTURE=1 isolates that single defect so the rest
+ // of the suite can be measured. Off by default: stock behaviour.
{
- pixelFormat = getPixelFormat(colorFormat);
+ const char *useTexEnv = getenv("MOBILEGL_CTS_FBO_COLOR_TEXTURE");
+ m_colorIsTexture = (useTexEnv && useTexEnv[0] == '1' && config.numSamples <= 0);
- gl.genRenderbuffers(1, &m_colorBuffer);
- gl.bindRenderbuffer(GL_RENDERBUFFER, m_colorBuffer);
+ pixelFormat = getPixelFormat(colorFormat);
- if (config.numSamples > 0)
- gl.renderbufferStorageMultisample(GL_RENDERBUFFER, config.numSamples, colorFormat, width, height);
+ if (m_colorIsTexture)
+ {
+ gl.genTextures(1, &m_colorBuffer);
+ gl.bindTexture(GL_TEXTURE_2D, m_colorBuffer);
+ gl.texStorage2D(GL_TEXTURE_2D, 1, colorFormat, width, height);
+ gl.texParameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_NEAREST);
+ gl.texParameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_NEAREST);
+ gl.bindTexture(GL_TEXTURE_2D, 0);
+ GLU_EXPECT_NO_ERROR(gl.getError(), "Creating color texture");
+ }
else
- gl.renderbufferStorage(GL_RENDERBUFFER, colorFormat, width, height);
-
- gl.bindRenderbuffer(GL_RENDERBUFFER, 0);
- GLU_EXPECT_NO_ERROR(gl.getError(), "Creating color renderbuffer");
+ {
+ gl.genRenderbuffers(1, &m_colorBuffer);
+ gl.bindRenderbuffer(GL_RENDERBUFFER, m_colorBuffer);
+
+ if (config.numSamples > 0)
+ gl.renderbufferStorageMultisample(GL_RENDERBUFFER, config.numSamples, colorFormat, width, height);
+ else
+ gl.renderbufferStorage(GL_RENDERBUFFER, colorFormat, width, height);
+
+ gl.bindRenderbuffer(GL_RENDERBUFFER, 0);
+ GLU_EXPECT_NO_ERROR(gl.getError(), "Creating color renderbuffer");
+ }
}
if (depthStencilFormat != GL_NONE)
@@ -250,7 +274,12 @@ void FboRenderContext::createFramebuffer(const RenderConfig &config)
gl.bindFramebuffer(GL_FRAMEBUFFER, m_framebuffer);
if (m_colorBuffer)
- gl.framebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, m_colorBuffer);
+ {
+ if (m_colorIsTexture)
+ gl.framebufferTexture2D(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, m_colorBuffer, 0);
+ else
+ gl.framebufferRenderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, m_colorBuffer);
+ }
if (m_depthStencilBuffer)
{
@@ -290,7 +319,10 @@ void FboRenderContext::destroyFramebuffer(void)
if (m_colorBuffer)
{
- gl.deleteRenderbuffers(1, &m_colorBuffer);
+ if (m_colorIsTexture)
+ gl.deleteTextures(1, &m_colorBuffer);
+ else
+ gl.deleteRenderbuffers(1, &m_colorBuffer);
m_colorBuffer = 0;
}
}
diff --git a/framework/opengl/gluFboRenderContext.hpp b/framework/opengl/gluFboRenderContext.hpp
index 75a0ff6b7..09ff1e7a9 100644
--- a/framework/opengl/gluFboRenderContext.hpp
+++ b/framework/opengl/gluFboRenderContext.hpp
@@ -80,6 +80,7 @@ private:
RenderContext *m_context;
uint32_t m_framebuffer;
uint32_t m_colorBuffer;
+ bool m_colorIsTexture;
uint32_t m_depthStencilBuffer;
tcu::RenderTarget m_renderTarget;
};
+473
View File
@@ -0,0 +1,473 @@
/*-------------------------------------------------------------------------
* dEQP platform port for MobileGL on Android
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*
*//*!
* \file
* \brief MobileGL platform.
*
* Modelled on the surfaceless platform, but adapted to MobileGL, which ships
* its own EGL implementation inside libMobileGL.so:
*
* - Every EGL call goes through the dynamically loaded library. The
* surfaceless port mixes wrapper calls with globally linked egl* symbols;
* doing that here would silently reach Android's system EGL instead.
* - Desktop-GL configs are selected with EGL_OPENGL_BIT. The surfaceless port
* always asks for an ES bit, which cannot satisfy a GL 3.3 core context.
* - A real surface is always created. MobileGL rejects EGL_NO_SURFACE with
* EGL_BAD_MATCH, and --deqp-surface-type=fbo asks the platform for
* SURFACETYPE_DONT_CARE, so "no surface" is not an option.
* - Window surfaces are backed by an AImageReader rather than an Activity,
* which is what lets the suite run as a plain adb-shell binary. DirectVulkan
* needs this: its pbuffer path requires VK_EXT_headless_surface, which
* Adreno's Android driver does not expose.
*
* Environment:
* MOBILEGL_CTS_LIB path/soname of the MobileGL library (default libMobileGL.so)
* MOBILEGL_CTS_SURFACE "window" (default) or "pbuffer"
* MOBILEGL_BACKEND_TYPE read by MobileGL itself; set it before launching
*//*--------------------------------------------------------------------*/
#include "tcuMobileGLPlatform.hpp"
#include <cstdlib>
#include <string>
#include <vector>
#include "deDynamicLibrary.hpp"
#include "egluUtil.hpp"
#include "eglwEnums.hpp"
#include "eglwLibrary.hpp"
#include "gluPlatform.hpp"
#include "gluRenderConfig.hpp"
#include "gluRenderContext.hpp"
#include "glwInitFunctions.hpp"
#include "tcuCommandLine.hpp"
#include "tcuPixelFormat.hpp"
#include "tcuPlatform.hpp"
#include "tcuRenderTarget.hpp"
#include <android/hardware_buffer.h>
#include <android/native_window.h>
#include <media/NdkImageReader.h>
using std::string;
using std::vector;
#if !defined(EGL_CONTEXT_OPENGL_PROFILE_MASK_KHR)
#define EGL_CONTEXT_FLAGS_KHR 0x30FC
#define EGL_CONTEXT_MAJOR_VERSION_KHR 0x3098
#define EGL_CONTEXT_MINOR_VERSION_KHR 0x30FB
#define EGL_CONTEXT_OPENGL_COMPATIBILITY_PROFILE_BIT_KHR 0x00000002
#define EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT_KHR 0x00000001
#define EGL_CONTEXT_OPENGL_DEBUG_BIT_KHR 0x00000001
#define EGL_CONTEXT_OPENGL_FORWARD_COMPATIBLE_BIT_KHR 0x00000002
#define EGL_CONTEXT_OPENGL_PROFILE_MASK_KHR 0x30FD
#define EGL_CONTEXT_OPENGL_ROBUST_ACCESS_BIT_KHR 0x00000004
#endif
namespace tcu
{
namespace mobilegl
{
static string getLibraryName(void)
{
const char *env = std::getenv("MOBILEGL_CTS_LIB");
return (env && env[0]) ? string(env) : string("libMobileGL.so");
}
//! Window surfaces default on: they are the only kind DirectVulkan can use.
static bool useWindowSurface(void)
{
const char *env = std::getenv("MOBILEGL_CTS_SURFACE");
return !(env && string(env) == "pbuffer");
}
/*--------------------------------------------------------------------*//*!
* \brief A real ANativeWindow with no Activity behind it.
*
* AImageReader's window is an ordinary BufferQueue producer, so both
* eglCreateWindowSurface and vkCreateAndroidSurfaceKHR accept it. The image
* listener must drain the queue: without it the producer blocks once maxImages
* buffers are in flight and the next swap deadlocks.
*//*--------------------------------------------------------------------*/
class ImageReaderWindow
{
public:
ImageReaderWindow(int width, int height) : m_reader(nullptr), m_window(nullptr)
{
const media_status_t status =
AImageReader_newWithUsage(width, height, AIMAGE_FORMAT_RGBA_8888,
AHARDWAREBUFFER_USAGE_GPU_SAMPLED_IMAGE |
AHARDWAREBUFFER_USAGE_GPU_COLOR_OUTPUT,
kMaxImages, &m_reader);
if (status != AMEDIA_OK || m_reader == nullptr)
throw tcu::ResourceError("AImageReader_newWithUsage() failed");
AImageReader_ImageListener listener = {this, onImageAvailable};
AImageReader_setImageListener(m_reader, &listener);
if (AImageReader_getWindow(m_reader, &m_window) != AMEDIA_OK || m_window == nullptr)
{
AImageReader_delete(m_reader);
m_reader = nullptr;
throw tcu::ResourceError("AImageReader_getWindow() failed");
}
ANativeWindow_acquire(m_window);
}
~ImageReaderWindow(void)
{
if (m_window != nullptr)
ANativeWindow_release(m_window);
if (m_reader != nullptr)
{
AImageReader_setImageListener(m_reader, nullptr);
AImageReader_delete(m_reader);
}
}
ANativeWindow *getWindow(void) const
{
return m_window;
}
private:
static const int kMaxImages = 4;
static void onImageAvailable(void *, AImageReader *reader)
{
AImage *image = nullptr;
if (AImageReader_acquireNextImage(reader, &image) == AMEDIA_OK && image != nullptr)
AImage_delete(image);
}
ImageReaderWindow(const ImageReaderWindow &);
ImageReaderWindow &operator=(const ImageReaderWindow &);
AImageReader *m_reader;
ANativeWindow *m_window;
};
class GetProcFuncLoader : public glw::FunctionLoader
{
public:
GetProcFuncLoader(const eglw::Library &egl) : m_egl(egl)
{
}
glw::GenericFuncType get(const char *name) const
{
return (glw::GenericFuncType)m_egl.getProcAddress(name);
}
protected:
const eglw::Library &m_egl;
};
class EglRenderContext : public glu::RenderContext
{
public:
EglRenderContext(const glu::RenderConfig &config, const tcu::CommandLine &cmdLine,
const glu::RenderContext *sharedContext);
~EglRenderContext(void);
glu::ContextType getType(void) const
{
return m_contextType;
}
eglw::EGLContext getEglContext(void) const
{
return m_eglContext;
}
const glw::Functions &getFunctions(void) const
{
return m_glFunctions;
}
const tcu::RenderTarget &getRenderTarget(void) const
{
return m_renderTarget;
}
void postIterate(void);
void makeCurrent(void);
glw::GenericFuncType getProcAddress(const char *name) const
{
return (glw::GenericFuncType)m_egl.getProcAddress(name);
}
private:
const eglw::DefaultLibrary m_egl;
const glu::ContextType m_contextType;
eglw::EGLDisplay m_eglDisplay;
eglw::EGLContext m_eglContext;
eglw::EGLSurface m_eglSurface;
ImageReaderWindow *m_window;
glw::Functions m_glFunctions;
tcu::RenderTarget m_renderTarget;
eglw::EGLContext m_sharedEglContext;
};
class ContextFactory : public glu::ContextFactory
{
public:
ContextFactory(void) : glu::ContextFactory("default", "MobileGL EGL context")
{
}
glu::RenderContext *createContext(const glu::RenderConfig &config, const tcu::CommandLine &cmdLine,
const glu::RenderContext *sharedContext) const
{
return new EglRenderContext(config, cmdLine, sharedContext);
}
};
class Platform : public tcu::Platform, public glu::Platform
{
public:
Platform(void)
{
m_contextFactoryRegistry.registerFactory(new ContextFactory());
}
const glu::Platform &getGLPlatform(void) const
{
return *this;
}
};
EglRenderContext::EglRenderContext(const glu::RenderConfig &config, const tcu::CommandLine &cmdLine,
const glu::RenderContext *sharedContext)
: m_egl(getLibraryName().c_str())
, m_contextType(config.type)
, m_eglDisplay(EGL_NO_DISPLAY)
, m_eglContext(EGL_NO_CONTEXT)
, m_eglSurface(EGL_NO_SURFACE)
, m_window(nullptr)
, m_renderTarget(config.width, config.height,
tcu::PixelFormat(config.redBits, config.greenBits, config.blueBits, config.alphaBits),
config.depthBits, config.stencilBits, config.numSamples)
, m_sharedEglContext(EGL_NO_CONTEXT)
{
DE_UNREF(cmdLine);
const glu::ContextType &contextType = config.type;
const bool isES = glu::isContextTypeES(contextType);
eglw::EGLint eglMajorVersion = 0;
eglw::EGLint eglMinorVersion = 0;
m_eglDisplay = m_egl.getDisplay(EGL_DEFAULT_DISPLAY);
EGLU_CHECK_MSG(m_egl, "eglGetDisplay()");
if (m_eglDisplay == EGL_NO_DISPLAY)
throw tcu::ResourceError("eglGetDisplay() failed");
EGLU_CHECK_CALL(m_egl, initialize(m_eglDisplay, &eglMajorVersion, &eglMinorVersion));
// MobileGL cannot make a context current without a surface, so
// SURFACETYPE_DONT_CARE (which is what --deqp-surface-type=fbo requests)
// still gets a real one.
bool wantWindow = false;
switch (config.surfaceType)
{
case glu::RenderConfig::SURFACETYPE_WINDOW:
wantWindow = true;
break;
case glu::RenderConfig::SURFACETYPE_OFFSCREEN_NATIVE:
case glu::RenderConfig::SURFACETYPE_OFFSCREEN_GENERIC:
wantWindow = false;
break;
case glu::RenderConfig::SURFACETYPE_DONT_CARE:
wantWindow = useWindowSurface();
break;
default:
TCU_CHECK_INTERNAL(false);
}
const int width = (config.width == glu::RenderConfig::DONT_CARE) ? 256 : config.width;
const int height = (config.height == glu::RenderConfig::DONT_CARE) ? 256 : config.height;
vector<eglw::EGLint> cfgAttribs;
cfgAttribs.push_back(EGL_RENDERABLE_TYPE);
if (isES)
{
switch (contextType.getMajorVersion())
{
case 3:
cfgAttribs.push_back(EGL_OPENGL_ES3_BIT);
break;
case 2:
cfgAttribs.push_back(EGL_OPENGL_ES2_BIT);
break;
default:
cfgAttribs.push_back(EGL_OPENGL_ES_BIT);
}
}
else
{
// Desktop GL, which is the whole point of this port.
cfgAttribs.push_back(EGL_OPENGL_BIT);
}
cfgAttribs.push_back(EGL_SURFACE_TYPE);
cfgAttribs.push_back(wantWindow ? EGL_WINDOW_BIT : EGL_PBUFFER_BIT);
static const struct
{
eglw::EGLint attrib;
int glu::RenderConfig::*field;
} s_sizeAttribs[] = {
{EGL_RED_SIZE, &glu::RenderConfig::redBits}, {EGL_GREEN_SIZE, &glu::RenderConfig::greenBits},
{EGL_BLUE_SIZE, &glu::RenderConfig::blueBits}, {EGL_ALPHA_SIZE, &glu::RenderConfig::alphaBits},
{EGL_DEPTH_SIZE, &glu::RenderConfig::depthBits}, {EGL_STENCIL_SIZE, &glu::RenderConfig::stencilBits},
{EGL_SAMPLES, &glu::RenderConfig::numSamples},
};
for (size_t ndx = 0; ndx < DE_LENGTH_OF_ARRAY(s_sizeAttribs); ndx++)
{
const int value = config.*(s_sizeAttribs[ndx].field);
if (value != glu::RenderConfig::DONT_CARE)
{
cfgAttribs.push_back(s_sizeAttribs[ndx].attrib);
cfgAttribs.push_back(value);
}
}
cfgAttribs.push_back(EGL_NONE);
eglw::EGLConfig eglConfig = nullptr;
eglw::EGLint numConfigs = 0;
EGLU_CHECK_CALL(m_egl, chooseConfig(m_eglDisplay, &cfgAttribs[0], &eglConfig, 1, &numConfigs));
if (numConfigs < 1)
throw tcu::NotSupportedError("No matching EGL config for the requested context");
if (wantWindow)
{
m_window = new ImageReaderWindow(width, height);
eglw::EGLint visualId = 0;
if (m_egl.getConfigAttrib(m_eglDisplay, eglConfig, EGL_NATIVE_VISUAL_ID, &visualId) && visualId != 0)
ANativeWindow_setBuffersGeometry(m_window->getWindow(), width, height, visualId);
m_eglSurface = m_egl.createWindowSurface(m_eglDisplay, eglConfig,
(eglw::EGLNativeWindowType)m_window->getWindow(), nullptr);
EGLU_CHECK_MSG(m_egl, "eglCreateWindowSurface()");
}
else
{
const eglw::EGLint surfaceAttribs[] = {EGL_WIDTH, width, EGL_HEIGHT, height, EGL_NONE};
m_eglSurface = m_egl.createPbufferSurface(m_eglDisplay, eglConfig, surfaceAttribs);
EGLU_CHECK_MSG(m_egl, "eglCreatePbufferSurface()");
}
if (m_eglSurface == EGL_NO_SURFACE)
throw tcu::ResourceError("Failed to create EGL surface");
vector<eglw::EGLint> ctxAttribs;
ctxAttribs.push_back(EGL_CONTEXT_MAJOR_VERSION_KHR);
ctxAttribs.push_back(contextType.getMajorVersion());
ctxAttribs.push_back(EGL_CONTEXT_MINOR_VERSION_KHR);
ctxAttribs.push_back(contextType.getMinorVersion());
switch (contextType.getProfile())
{
case glu::PROFILE_ES:
EGLU_CHECK_CALL(m_egl, bindAPI(EGL_OPENGL_ES_API));
break;
case glu::PROFILE_CORE:
EGLU_CHECK_CALL(m_egl, bindAPI(EGL_OPENGL_API));
ctxAttribs.push_back(EGL_CONTEXT_OPENGL_PROFILE_MASK_KHR);
ctxAttribs.push_back(EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT_KHR);
break;
case glu::PROFILE_COMPATIBILITY:
EGLU_CHECK_CALL(m_egl, bindAPI(EGL_OPENGL_API));
ctxAttribs.push_back(EGL_CONTEXT_OPENGL_PROFILE_MASK_KHR);
ctxAttribs.push_back(EGL_CONTEXT_OPENGL_COMPATIBILITY_PROFILE_BIT_KHR);
break;
default:
TCU_CHECK_INTERNAL(false);
}
eglw::EGLint flags = 0;
if ((contextType.getFlags() & glu::CONTEXT_DEBUG) != 0)
flags |= EGL_CONTEXT_OPENGL_DEBUG_BIT_KHR;
if ((contextType.getFlags() & glu::CONTEXT_ROBUST) != 0)
flags |= EGL_CONTEXT_OPENGL_ROBUST_ACCESS_BIT_KHR;
if ((contextType.getFlags() & glu::CONTEXT_FORWARD_COMPATIBLE) != 0)
flags |= EGL_CONTEXT_OPENGL_FORWARD_COMPATIBLE_BIT_KHR;
if (flags != 0)
{
ctxAttribs.push_back(EGL_CONTEXT_FLAGS_KHR);
ctxAttribs.push_back(flags);
}
ctxAttribs.push_back(EGL_NONE);
const EglRenderContext *sharedEglRenderContext = dynamic_cast<const EglRenderContext *>(sharedContext);
m_sharedEglContext = sharedEglRenderContext ? sharedEglRenderContext->getEglContext() : EGL_NO_CONTEXT;
m_eglContext = m_egl.createContext(m_eglDisplay, eglConfig, m_sharedEglContext, &ctxAttribs[0]);
EGLU_CHECK_MSG(m_egl, "eglCreateContext()");
if (!m_eglContext)
throw tcu::ResourceError("eglCreateContext() failed");
// MobileGL requires draw == read.
EGLU_CHECK_CALL(m_egl, makeCurrent(m_eglDisplay, m_eglSurface, m_eglSurface, m_eglContext));
// MobileGL advertises EGL 1.5, so eglGetProcAddress resolves core entry
// points too; there is no separate GL library to dlopen.
GetProcFuncLoader funcLoader(m_egl);
glu::initCoreFunctions(&m_glFunctions, &funcLoader, contextType.getAPI());
glu::initExtensionFunctions(&m_glFunctions, &funcLoader, contextType.getAPI());
}
EglRenderContext::~EglRenderContext(void)
{
try
{
if (m_eglDisplay != EGL_NO_DISPLAY)
{
m_egl.makeCurrent(m_eglDisplay, EGL_NO_SURFACE, EGL_NO_SURFACE, EGL_NO_CONTEXT);
if (m_eglContext != EGL_NO_CONTEXT)
m_egl.destroyContext(m_eglDisplay, m_eglContext);
if (m_eglSurface != EGL_NO_SURFACE)
m_egl.destroySurface(m_eglDisplay, m_eglSurface);
if (m_sharedEglContext == EGL_NO_CONTEXT)
m_egl.terminate(m_eglDisplay);
}
}
catch (...)
{
}
delete m_window;
}
void EglRenderContext::makeCurrent(void)
{
EGLU_CHECK_CALL(m_egl, makeCurrent(m_eglDisplay, m_eglSurface, m_eglSurface, m_eglContext));
}
void EglRenderContext::postIterate(void)
{
m_glFunctions.finish();
}
} // namespace mobilegl
} // namespace tcu
tcu::Platform *createPlatform(void)
{
return new tcu::mobilegl::Platform();
}
@@ -0,0 +1,33 @@
#ifndef _TCUMOBILEGLPLATFORM_HPP
#define _TCUMOBILEGLPLATFORM_HPP
/*-------------------------------------------------------------------------
* dEQP platform port for MobileGL on Android
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*
*//*!
* \file
* \brief MobileGL platform - drives libMobileGL.so's own EGL from a bare
* Android process, with no Activity and no system EGL involved.
*//*--------------------------------------------------------------------*/
#include "tcuDefs.hpp"
namespace tcu
{
class Platform;
}
tcu::Platform *createPlatform(void);
#endif // _TCUMOBILEGLPLATFORM_HPP
+2
View File
@@ -0,0 +1,2 @@
mgprobe
*.o
+359
View File
@@ -0,0 +1,359 @@
/* mgprobe - preflight gate for running a GL conformance suite against MobileGL
* from a bare adb-shell process (no APK, no Activity).
*
* Verifies, for one backend and one surface type, that MobileGL can hand out a
* GL 3.3 core context and that pixels read back correctly - both from the
* default framebuffer and from a user FBO. Run this before burning hours on a
* CTS run; it catches a broken device/library pairing in about a second.
*
* mgprobe --backend DirectGLES|DirectVulkan --surface pbuffer|imagereader
* [--lib /path/to/libMobileGL.so]
*
* Exit status: 0 if a context came up and FBO readback is correct, non-zero
* otherwise. Default-framebuffer readback is reported but does NOT gate, because
* DirectVulkan is known to return zeros there while FBO readback is sound.
*
* Build (NDK, arm64):
* $NDK/toolchains/llvm/prebuilt/<host>/bin/aarch64-linux-android26-clang \
* -O1 -o mgprobe mgprobe.c -ldl -llog -landroid -lmediandk
*/
#include <android/native_window.h>
#include <dlfcn.h>
#include <media/NdkImageReader.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
typedef void *EGLDisplay;
typedef void *EGLConfig;
typedef void *EGLSurface;
typedef void *EGLContext;
typedef int EGLint;
typedef unsigned int EGLBoolean;
typedef unsigned int EGLenum;
typedef void *EGLNativeDisplayType;
typedef void *EGLNativeWindowType;
#define EGL_DEFAULT_DISPLAY ((EGLNativeDisplayType)0)
#define EGL_NO_CONTEXT ((EGLContext)0)
#define EGL_NO_SURFACE ((EGLSurface)0)
#define EGL_NONE 0x3038
#define EGL_WIDTH 0x3057
#define EGL_HEIGHT 0x3056
#define EGL_RENDERABLE_TYPE 0x3040
#define EGL_SURFACE_TYPE 0x3033
#define EGL_WINDOW_BIT 0x0004
#define EGL_PBUFFER_BIT 0x0001
#define EGL_OPENGL_BIT 0x0008
#define EGL_OPENGL_API 0x30A2
#define EGL_RED_SIZE 0x3024
#define EGL_GREEN_SIZE 0x3023
#define EGL_BLUE_SIZE 0x3022
#define EGL_ALPHA_SIZE 0x3021
#define EGL_DEPTH_SIZE 0x3025
#define EGL_STENCIL_SIZE 0x3026
#define EGL_NATIVE_VISUAL_ID 0x302E
#define EGL_CONTEXT_MAJOR_VERSION 0x3098
#define EGL_CONTEXT_MINOR_VERSION 0x30FB
#define EGL_CONTEXT_OPENGL_PROFILE_MASK 0x30FD
#define EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT 0x00000001
#define GL_VENDOR 0x1F00
#define GL_RENDERER 0x1F01
#define GL_VERSION 0x1F02
#define GL_SHADING_LANGUAGE_VERSION 0x8B8C
#define GL_CONTEXT_PROFILE_MASK 0x9126
#define GL_MAJOR_VERSION 0x821B
#define GL_MINOR_VERSION 0x821C
#define GL_COLOR_BUFFER_BIT 0x00004000
#define GL_RGBA 0x1908
#define GL_RGBA8 0x8058
#define GL_UNSIGNED_BYTE 0x1401
#define GL_TEXTURE_2D 0x0DE1
#define GL_FRAMEBUFFER 0x8D40
#define GL_COLOR_ATTACHMENT0 0x8CE0
#define GL_FRAMEBUFFER_COMPLETE 0x8CD5
#define GL_TEXTURE_MIN_FILTER 0x2801
#define GL_TEXTURE_MAG_FILTER 0x2800
#define GL_NEAREST 0x2600
#define GL_RENDERBUFFER 0x8D41
typedef EGLDisplay (*P_getDisplay)(EGLNativeDisplayType);
typedef EGLBoolean (*P_initialize)(EGLDisplay, EGLint *, EGLint *);
typedef EGLBoolean (*P_bindAPI)(EGLenum);
typedef EGLBoolean (*P_chooseConfig)(EGLDisplay, const EGLint *, EGLConfig *, EGLint, EGLint *);
typedef EGLBoolean (*P_getConfigAttrib)(EGLDisplay, EGLConfig, EGLint, EGLint *);
typedef EGLSurface (*P_createWindowSurface)(EGLDisplay, EGLConfig, EGLNativeWindowType, const EGLint *);
typedef EGLSurface (*P_createPbufferSurface)(EGLDisplay, EGLConfig, const EGLint *);
typedef EGLContext (*P_createContext)(EGLDisplay, EGLConfig, EGLContext, const EGLint *);
typedef EGLBoolean (*P_makeCurrent)(EGLDisplay, EGLSurface, EGLSurface, EGLContext);
typedef EGLint (*P_getError)(void);
typedef const unsigned char *(*P_glGetString)(unsigned int);
typedef void (*P_glGetIntegerv)(unsigned int, int *);
typedef void (*P_glClearColor)(float, float, float, float);
typedef void (*P_glClear)(unsigned int);
typedef void (*P_glFinish)(void);
typedef void (*P_glReadPixels)(int, int, int, int, unsigned int, unsigned int, void *);
typedef unsigned int (*P_glGetError)(void);
typedef void (*P_glGenTextures)(int, unsigned int *);
typedef void (*P_glBindTexture)(unsigned int, unsigned int);
typedef void (*P_glTexImage2D)(unsigned int, int, int, int, int, int, unsigned int, unsigned int, const void *);
typedef void (*P_glTexParameteri)(unsigned int, unsigned int, int);
typedef void (*P_glGenFramebuffers)(int, unsigned int *);
typedef void (*P_glBindFramebuffer)(unsigned int, unsigned int);
typedef void (*P_glFramebufferTexture2D)(unsigned int, unsigned int, unsigned int, unsigned int, int);
typedef unsigned int (*P_glCheckFramebufferStatus)(unsigned int);
typedef void (*P_glViewport)(int, int, int, int);
typedef void (*P_glGenRenderbuffers)(int, unsigned int *);
typedef void (*P_glBindRenderbuffer)(unsigned int, unsigned int);
typedef void (*P_glRenderbufferStorage)(unsigned int, unsigned int, int, int);
typedef void (*P_glFramebufferRenderbuffer)(unsigned int, unsigned int, unsigned int, unsigned int);
static void *g_lib;
static void *S(const char *n) { return dlsym(g_lib, n); }
static void on_image(void *ctx, AImageReader *r) {
(void)ctx;
AImage *img = NULL;
/* Drain the queue, or the producer blocks once maxImages are in flight. */
if (AImageReader_acquireNextImage(r, &img) == AMEDIA_OK && img) AImage_delete(img);
}
#define DIM 256
static int near8(unsigned got, int want, int tol) {
int d = (int)got - want;
return d <= tol && d >= -tol;
}
int main(int argc, char **argv) {
const char *backend = "DirectGLES";
const char *surface = "pbuffer";
const char *libpath = "libMobileGL.so";
for (int i = 1; i < argc; ++i) {
if (!strcmp(argv[i], "--backend") && i + 1 < argc) backend = argv[++i];
else if (!strcmp(argv[i], "--surface") && i + 1 < argc) surface = argv[++i];
else if (!strcmp(argv[i], "--lib") && i + 1 < argc) libpath = argv[++i];
else {
fprintf(stderr, "usage: %s [--backend DirectGLES|DirectVulkan]"
" [--surface pbuffer|imagereader] [--lib path]\n", argv[0]);
return 2;
}
}
setvbuf(stdout, NULL, _IONBF, 0);
/* MobileGL parses its config from an ELF constructor, so the backend must be
* selected before the library is mapped. */
setenv("MOBILEGL_BACKEND_TYPE", backend, 1);
printf("mgprobe backend=%s surface=%s lib=%s\n", backend, surface, libpath);
int useWindow = !strcmp(surface, "imagereader");
ANativeWindow *win = NULL;
AImageReader *reader = NULL;
if (useWindow) {
if (AImageReader_newWithUsage(DIM, DIM, AIMAGE_FORMAT_RGBA_8888,
AHARDWAREBUFFER_USAGE_GPU_SAMPLED_IMAGE |
AHARDWAREBUFFER_USAGE_GPU_COLOR_OUTPUT,
4, &reader) != AMEDIA_OK || !reader) {
printf("FAIL AImageReader_newWithUsage\n");
return 3;
}
AImageReader_ImageListener l = {NULL, on_image};
AImageReader_setImageListener(reader, &l);
if (AImageReader_getWindow(reader, &win) != AMEDIA_OK || !win) {
printf("FAIL AImageReader_getWindow\n");
return 3;
}
}
g_lib = dlopen(libpath, RTLD_NOW | RTLD_LOCAL);
if (!g_lib) {
printf("FAIL dlopen: %s\n", dlerror());
return 4;
}
P_getDisplay eglGetDisplay_ = (P_getDisplay)S("eglGetDisplay");
P_initialize eglInitialize_ = (P_initialize)S("eglInitialize");
P_bindAPI eglBindAPI_ = (P_bindAPI)S("eglBindAPI");
P_chooseConfig eglChooseConfig_ = (P_chooseConfig)S("eglChooseConfig");
P_getConfigAttrib eglGetConfigAttrib_ = (P_getConfigAttrib)S("eglGetConfigAttrib");
P_createWindowSurface eglCreateWindowSurface_ = (P_createWindowSurface)S("eglCreateWindowSurface");
P_createPbufferSurface eglCreatePbufferSurface_ = (P_createPbufferSurface)S("eglCreatePbufferSurface");
P_createContext eglCreateContext_ = (P_createContext)S("eglCreateContext");
P_makeCurrent eglMakeCurrent_ = (P_makeCurrent)S("eglMakeCurrent");
P_getError eglGetError_ = (P_getError)S("eglGetError");
if (!eglGetDisplay_ || !eglInitialize_ || !eglChooseConfig_ || !eglCreateContext_ || !eglMakeCurrent_) {
printf("FAIL missing core EGL exports\n");
return 5;
}
EGLDisplay dpy = eglGetDisplay_(EGL_DEFAULT_DISPLAY);
EGLint vmaj = 0, vmin = 0;
if (!eglInitialize_(dpy, &vmaj, &vmin)) {
printf("FAIL eglInitialize err=0x%x\n", eglGetError_ ? eglGetError_() : 0);
return 6;
}
if (eglBindAPI_ && !eglBindAPI_(EGL_OPENGL_API)) {
printf("FAIL eglBindAPI(EGL_OPENGL_API) err=0x%x\n", eglGetError_ ? eglGetError_() : 0);
return 7;
}
const EGLint cfgAttribs[] = {
EGL_SURFACE_TYPE, useWindow ? EGL_WINDOW_BIT : EGL_PBUFFER_BIT,
EGL_RENDERABLE_TYPE, EGL_OPENGL_BIT,
EGL_RED_SIZE, 8, EGL_GREEN_SIZE, 8, EGL_BLUE_SIZE, 8, EGL_ALPHA_SIZE, 8,
EGL_DEPTH_SIZE, 24, EGL_STENCIL_SIZE, 8,
EGL_NONE};
EGLConfig cfg = 0;
EGLint ncfg = 0;
if (!eglChooseConfig_(dpy, cfgAttribs, &cfg, 1, &ncfg) || ncfg < 1) {
printf("FAIL eglChooseConfig n=%d err=0x%x\n", ncfg, eglGetError_ ? eglGetError_() : 0);
return 8;
}
EGLSurface surf;
if (useWindow) {
EGLint vis = 0;
if (eglGetConfigAttrib_ && eglGetConfigAttrib_(dpy, cfg, EGL_NATIVE_VISUAL_ID, &vis) && vis)
ANativeWindow_setBuffersGeometry(win, DIM, DIM, vis);
surf = eglCreateWindowSurface_(dpy, cfg, (EGLNativeWindowType)win, NULL);
} else {
const EGLint sa[] = {EGL_WIDTH, DIM, EGL_HEIGHT, DIM, EGL_NONE};
surf = eglCreatePbufferSurface_(dpy, cfg, sa);
}
if (surf == EGL_NO_SURFACE) {
printf("FAIL create%sSurface err=0x%x\n", useWindow ? "Window" : "Pbuffer",
eglGetError_ ? eglGetError_() : 0);
return 9;
}
const EGLint ctxAttribs[] = {
EGL_CONTEXT_MAJOR_VERSION, 3, EGL_CONTEXT_MINOR_VERSION, 3,
EGL_CONTEXT_OPENGL_PROFILE_MASK, EGL_CONTEXT_OPENGL_CORE_PROFILE_BIT, EGL_NONE};
EGLContext ctx = eglCreateContext_(dpy, cfg, EGL_NO_CONTEXT, ctxAttribs);
if (ctx == EGL_NO_CONTEXT) {
printf("FAIL eglCreateContext(3.3 core) err=0x%x\n", eglGetError_ ? eglGetError_() : 0);
return 10;
}
/* MobileGL requires draw == read and rejects EGL_NO_SURFACE. */
if (!eglMakeCurrent_(dpy, surf, surf, ctx)) {
printf("FAIL eglMakeCurrent err=0x%x\n", eglGetError_ ? eglGetError_() : 0);
return 11;
}
P_glGetString glGetString_ = (P_glGetString)S("glGetString");
P_glGetIntegerv glGetIntegerv_ = (P_glGetIntegerv)S("glGetIntegerv");
P_glClearColor glClearColor_ = (P_glClearColor)S("glClearColor");
P_glClear glClear_ = (P_glClear)S("glClear");
P_glFinish glFinish_ = (P_glFinish)S("glFinish");
P_glReadPixels glReadPixels_ = (P_glReadPixels)S("glReadPixels");
P_glGetError glGetError_ = (P_glGetError)S("glGetError");
P_glGenTextures glGenTextures_ = (P_glGenTextures)S("glGenTextures");
P_glBindTexture glBindTexture_ = (P_glBindTexture)S("glBindTexture");
P_glTexImage2D glTexImage2D_ = (P_glTexImage2D)S("glTexImage2D");
P_glTexParameteri glTexParameteri_ = (P_glTexParameteri)S("glTexParameteri");
P_glGenFramebuffers glGenFramebuffers_ = (P_glGenFramebuffers)S("glGenFramebuffers");
P_glBindFramebuffer glBindFramebuffer_ = (P_glBindFramebuffer)S("glBindFramebuffer");
P_glFramebufferTexture2D glFramebufferTexture2D_ = (P_glFramebufferTexture2D)S("glFramebufferTexture2D");
P_glCheckFramebufferStatus glCheckFramebufferStatus_ = (P_glCheckFramebufferStatus)S("glCheckFramebufferStatus");
P_glViewport glViewport_ = (P_glViewport)S("glViewport");
P_glGenRenderbuffers glGenRenderbuffers_ = (P_glGenRenderbuffers)S("glGenRenderbuffers");
P_glBindRenderbuffer glBindRenderbuffer_ = (P_glBindRenderbuffer)S("glBindRenderbuffer");
P_glRenderbufferStorage glRenderbufferStorage_ = (P_glRenderbufferStorage)S("glRenderbufferStorage");
P_glFramebufferRenderbuffer glFramebufferRenderbuffer_ = (P_glFramebufferRenderbuffer)S("glFramebufferRenderbuffer");
int major = -1, minor = -1, profile = -1;
glGetIntegerv_(GL_MAJOR_VERSION, &major);
glGetIntegerv_(GL_MINOR_VERSION, &minor);
glGetIntegerv_(GL_CONTEXT_PROFILE_MASK, &profile);
printf(" GL_VENDOR %s\n", (const char *)glGetString_(GL_VENDOR));
printf(" GL_RENDERER %s\n", (const char *)glGetString_(GL_RENDERER));
printf(" GL_VERSION %s\n", (const char *)glGetString_(GL_VERSION));
printf(" GLSL %s\n", (const char *)glGetString_(GL_SHADING_LANGUAGE_VERSION));
printf(" version %d.%d profile_mask 0x%x %s\n", major, minor, profile,
(profile & 1) ? "(core)" : "(NOT CORE)");
unsigned char px[4];
/* Default framebuffer. */
glClearColor_(0.25f, 0.5f, 0.75f, 1.0f);
glClear_(GL_COLOR_BUFFER_BIT);
if (glFinish_) glFinish_();
memset(px, 0, sizeof px);
glReadPixels_(DIM / 2, DIM / 2, 1, 1, GL_RGBA, GL_UNSIGNED_BYTE, px);
int defOk = near8(px[0], 64, 10) && near8(px[1], 128, 10) && near8(px[2], 191, 10);
printf(" default-FB readback (%u,%u,%u,%u) %s\n", px[0], px[1], px[2], px[3],
defOk ? "ok" : "BROKEN");
/* User FBO - this is what dEQP uses with --deqp-surface-type=fbo. */
unsigned int tex = 0, fbo = 0;
glGenTextures_(1, &tex);
glBindTexture_(GL_TEXTURE_2D, tex);
glTexImage2D_(GL_TEXTURE_2D, 0, GL_RGBA8, DIM, DIM, 0, GL_RGBA, GL_UNSIGNED_BYTE, NULL);
glTexParameteri_(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_NEAREST);
glTexParameteri_(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_NEAREST);
glGenFramebuffers_(1, &fbo);
glBindFramebuffer_(GL_FRAMEBUFFER, fbo);
glFramebufferTexture2D_(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, tex, 0);
unsigned int fbst = glCheckFramebufferStatus_(GL_FRAMEBUFFER);
int fboOk = 0;
if (fbst == GL_FRAMEBUFFER_COMPLETE) {
glViewport_(0, 0, DIM, DIM);
glClearColor_(0.9f, 0.2f, 0.4f, 1.0f);
glClear_(GL_COLOR_BUFFER_BIT);
if (glFinish_) glFinish_();
memset(px, 0, sizeof px);
glReadPixels_(DIM / 2, DIM / 2, 1, 1, GL_RGBA, GL_UNSIGNED_BYTE, px);
fboOk = near8(px[0], 230, 10) && near8(px[1], 51, 10) && near8(px[2], 102, 10);
printf(" user-FBO readback (%u,%u,%u,%u) %s\n", px[0], px[1], px[2], px[3],
fboOk ? "ok" : "BROKEN");
} else {
printf(" user-FBO incomplete status=0x%x\n", fbst);
}
/* FBO with a RENDERBUFFER colour attachment. This is what dEQP's
* FboRenderContext allocates for --deqp-surface-type=fbo, so it is the path
* that actually decides a conformance run - a texture-attached FBO working
* says nothing about it. */
unsigned int rbo = 0, rfbo = 0;
int rboOk = 0;
if (glGenRenderbuffers_ && glBindRenderbuffer_ && glRenderbufferStorage_ && glFramebufferRenderbuffer_) {
glGenRenderbuffers_(1, &rbo);
glBindRenderbuffer_(GL_RENDERBUFFER, rbo);
glRenderbufferStorage_(GL_RENDERBUFFER, GL_RGBA8, DIM, DIM);
glBindRenderbuffer_(GL_RENDERBUFFER, 0);
glGenFramebuffers_(1, &rfbo);
glBindFramebuffer_(GL_FRAMEBUFFER, rfbo);
glFramebufferRenderbuffer_(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, rbo);
unsigned int rst = glCheckFramebufferStatus_(GL_FRAMEBUFFER);
if (rst == GL_FRAMEBUFFER_COMPLETE) {
glViewport_(0, 0, DIM, DIM);
glClearColor_(0.1f, 0.7f, 0.3f, 1.0f);
glClear_(GL_COLOR_BUFFER_BIT);
if (glFinish_) glFinish_();
memset(px, 0, sizeof px);
glReadPixels_(DIM / 2, DIM / 2, 1, 1, GL_RGBA, GL_UNSIGNED_BYTE, px);
rboOk = near8(px[0], 26, 10) && near8(px[1], 179, 10) && near8(px[2], 77, 10);
printf(" rbo-FBO readback (%u,%u,%u,%u) %s\n", px[0], px[1], px[2], px[3],
rboOk ? "ok" : "BROKEN");
} else {
printf(" rbo-FBO incomplete status=0x%x\n", rst);
}
} else {
printf(" rbo-FBO skipped (renderbuffer entry points unavailable)\n");
}
unsigned glerr = glGetError_ ? glGetError_() : 0;
int ok = fboOk && rboOk && (major > 3 || (major == 3 && minor >= 3)) && (profile & 1) && glerr == 0;
printf("%s backend=%s surface=%s default_fb=%s user_fbo=%s rbo_fbo=%s glerr=0x%x\n",
ok ? "PASS" : "FAIL", backend, surface, defOk ? "ok" : "broken",
fboOk ? "ok" : "broken", rboOk ? "ok" : "broken", glerr);
fflush(stdout);
/* MobileGL aborts in static teardown; leave before that runs. */
_exit(ok ? 0 : 1);
}
+545
View File
@@ -0,0 +1,545 @@
#!/usr/bin/env python
"""Build a GL 3.0--3.3 CTS conformance matrix from dEQP QPA logs.
The report deliberately scores against the unique cases in each supplied
caselist. A case that has not produced a result therefore cannot disappear
from the denominator and make a partial run look conformant.
QPA parsing and crash/hang sidecar handling follow :mod:`qpa_report`:
* a later QPA observation of a case wins;
* ``crashed.txt`` upgrades a missing/incomplete result to ``Crash``;
* ``hung.txt`` upgrades a missing/incomplete/crash result to ``DeviceHang``.
Example::
python cts_matrix_report.py \
--gl30-caselist gl30-main.txt --gl30-results runs/gl30 \
--gl31-caselist gl31-main.txt --gl31-results runs/gl31 \
--gl32-caselist gl32-main.txt --gl32-results runs/gl32 \
--gl33-caselist gl33-main.txt --gl33-results runs/gl33 \
--json runs/cts-matrix.json
"""
from __future__ import annotations
import argparse
import json
import os
import re
import sys
from collections import Counter, defaultdict
from datetime import datetime, timezone
from typing import Iterable, Optional, Sequence
try: # Works both as a directly executed script and as a package import.
from . import qpa_report
except ImportError: # pragma: no cover - exercised by the command-line tests
import qpa_report
VERSIONS = ("gl30", "gl31", "gl32", "gl33")
ACCEPTED_STATUSES = (
"Pass",
"NotSupported",
"QualityWarning",
"CompatibilityWarning",
"Waiver",
)
ACCEPTED = frozenset(ACCEPTED_STATUSES)
CHUNK_QPA = re.compile(r"^chunk(\d+)\.qpa$", re.IGNORECASE)
class ReportInputError(ValueError):
"""An input path cannot be used to construct a meaningful report."""
def _read_non_comment_lines(path: str) -> list[str]:
try:
# Match run_cts_windows.py: Khronos lists are UTF-8 and may carry a BOM.
with open(path, "r", encoding="utf-8-sig", errors="strict") as fh:
return [
line.strip()
for line in fh
if line.strip() and not line.lstrip().startswith("#")
]
except (OSError, UnicodeError) as exc:
raise ReportInputError(f"cannot read {path}: {exc}") from exc
def read_caselist(path: str) -> tuple[list[str], dict[str, int]]:
"""Return unique cases in file order and repeated caselist entries.
The mustpass files consumed by glcts and ``run_cts.py`` are one case per
non-empty, non-comment line, so this intentionally uses the same syntax.
"""
entries = _read_non_comment_lines(path)
counts = Counter(entries)
unique = list(dict.fromkeys(entries))
duplicates = {case: count for case, count in counts.items() if count > 1}
return unique, duplicates
def _collect_qpa_files(paths: Sequence[str]) -> list[str]:
missing = [path for path in paths if not os.path.exists(path)]
if missing:
raise ReportInputError(
"result path(s) do not exist: " + ", ".join(sorted(missing))
)
# qpa_report.collect provides the established directory-recursion rules.
# De-duplicate aliases so specifying the same directory twice does not
# manufacture duplicate observations.
files = qpa_report.collect(paths)
by_identity: dict[str, str] = {}
for path in files:
if not os.path.isfile(path):
raise ReportInputError(f"QPA input is not a file: {path}")
absolute = os.path.abspath(path)
by_identity.setdefault(os.path.normcase(absolute), absolute)
def order_key(value: str) -> tuple[str, str, int, str]:
absolute = os.path.abspath(value)
directory = os.path.normcase(os.path.dirname(absolute))
filename = os.path.normcase(os.path.basename(absolute))
match = CHUNK_QPA.fullmatch(filename)
if match:
# run_cts_windows.py uses a minimum width of four digits, not a
# fixed width. Numeric ordering is therefore required once a run
# reaches chunk10000; lexical ordering would put it before
# chunk9999 and break the later-observation-wins rule.
return directory, "chunk", int(match.group(1)), filename
# Preserve a deterministic, name-based position for foreign/legacy
# QPA files while grouping numeric runner chunks at the lexical
# position occupied by the "chunk" basename.
return directory, filename, -1, filename
return sorted(by_identity.values(), key=order_key)
def _ratio(numerator: int, denominator: int) -> float:
return numerator / denominator if denominator else 0.0
def _sidecar(paths: Sequence[str], name: str) -> set[str]:
"""Load a run_cts.py sidecar with qpa_report-compatible lookup rules."""
return qpa_report.load_sidecar(paths, name)
def build_version_report(
version: str, caselist: str, result_paths: Sequence[str]
) -> dict:
"""Build the serialisable report for one GL mustpass version."""
expected_cases, expected_duplicates = read_caselist(caselist)
expected = set(expected_cases)
qpa_files = _collect_qpa_files(result_paths)
results: dict[str, str] = {}
observation_history: dict[str, list[dict[str, str]]] = defaultdict(list)
for qpa_file in qpa_files:
for case, status in qpa_report.parse_qpa(qpa_file):
observation_history[case].append(
{"file": qpa_file, "status": status}
)
results[case] = status
crashed = _sidecar(result_paths, "crashed.txt")
hung = _sidecar(result_paths, "hung.txt")
explicit_unrun = _sidecar(result_paths, "unrun.txt")
skipped = _sidecar(result_paths, "skipped.txt")
# Keep this order and these guards in lock-step with qpa_report.py.
for case in crashed:
if results.get(case, "Incomplete") == "Incomplete":
results[case] = "Crash"
for case in hung:
if results.get(case, "Incomplete") in ("Incomplete", "Crash"):
results[case] = "DeviceHang"
# A begin/end pair without <Result>, or a QPA truncated mid-case, is not a
# completed observation. Sidecars above may upgrade it to Crash/Hang;
# anything still Incomplete must stay in the expected denominator as unrun.
incomplete_results = {
case for case, status in results.items() if status == "Incomplete"
}
for case in incomplete_results:
del results[case]
expected_results = {
case: status for case, status in results.items() if case in expected
}
unexpected_results = {
case: status for case, status in results.items() if case not in expected
}
# Missing cases are inferred from the caselist even if unrun.txt itself is
# missing or stale. This is the invariant that prevents partial-run rate
# inflation.
unrun_cases = expected - set(expected_results)
declared_not_measured = explicit_unrun | skipped
undeclared_unrun = unrun_cases - declared_not_measured
stale_unrun = (explicit_unrun | skipped) & set(expected_results)
counts = Counter(expected_results.values())
strict_pass = counts["Pass"]
accepted = sum(counts[status] for status in ACCEPTED)
result_count = len(expected_results)
expected_count = len(expected)
crash_count = counts["Crash"]
hang_count = counts["DeviceHang"]
duplicate_cases = {
case: {
"observations": len(history),
"extra_observations": len(history) - 1,
"final_status": results.get(case, "Incomplete"),
"history": history,
}
for case, history in sorted(observation_history.items())
if len(history) > 1
}
duplicate_observations = sum(
item["extra_observations"] for item in duplicate_cases.values()
)
sidecar_unknown = {
name: sorted(cases - expected)
for name, cases in (
("crashed.txt", crashed),
("hung.txt", hung),
("unrun.txt", explicit_unrun),
("skipped.txt", skipped),
)
if cases - expected
}
errors: list[str] = []
warnings: list[str] = []
if not expected_count:
errors.append("caselist has no cases")
if expected_duplicates:
errors.append(
f"caselist has {sum(n - 1 for n in expected_duplicates.values())} "
"duplicate entry/entries"
)
if not qpa_files:
errors.append("no .qpa files found")
if unexpected_results:
errors.append(
f"{len(unexpected_results)} result case(s) are absent from the caselist"
)
if sidecar_unknown:
errors.append("one or more sidecars name cases absent from the caselist")
if undeclared_unrun:
errors.append(
f"{len(undeclared_unrun)} missing result case(s) are not declared by "
"unrun.txt/skipped.txt"
)
if stale_unrun:
warnings.append(
f"{len(stale_unrun)} case(s) declared unrun/skipped also have a result"
)
if duplicate_observations:
warnings.append(
f"{duplicate_observations} duplicate QPA observation(s); last result wins"
)
if incomplete_results:
warnings.append(
f"{len(incomplete_results)} QPA case(s) ended without a final result and were treated as unrun"
)
if errors:
state = "ERROR"
elif unrun_cases:
state = "INCOMPLETE"
else:
state = "OK"
return {
"version": version,
"inputs": {
"caselist": os.path.abspath(caselist),
"result_paths": [os.path.abspath(path) for path in result_paths],
"qpa_files": qpa_files,
},
"expected": expected_count,
"result": result_count,
"pass": strict_pass,
"accepted": accepted,
"crash": crash_count,
"hang": hang_count,
"unrun": len(unrun_cases),
"duplicate": duplicate_observations,
"counts": dict(sorted(counts.items())),
"coverage": {
"numerator": result_count,
"denominator": expected_count,
"rate": _ratio(result_count, expected_count),
},
"rates": {
# These are the report's conformance rates. Expected, not merely
# measured results, is the denominator.
"denominator": "expected",
"strict_pass_only": _ratio(strict_pass, expected_count),
"conformance_accepted": _ratio(accepted, expected_count),
# Useful for comparison with qpa_report.py, whose denominator is
# cases with a result. Never presented as the conformance rate.
"measured_only_strict_pass": _ratio(strict_pass, result_count),
"measured_only_conformance_accepted": _ratio(
accepted, result_count
),
},
"strict_pass_rate": _ratio(strict_pass, expected_count),
"conformance_accepted_rate": _ratio(accepted, expected_count),
"validation": {
"state": state,
"ok": state == "OK",
"errors": errors,
"warnings": warnings,
"invariant_expected_equals_result_plus_unrun": (
expected_count == result_count + len(unrun_cases)
),
"undeclared_unrun": sorted(undeclared_unrun),
"stale_unrun_or_skipped": sorted(stale_unrun),
"sidecar_cases_absent_from_caselist": sidecar_unknown,
},
"cases": {
"results": dict(sorted(expected_results.items())),
"unrun": sorted(unrun_cases),
"unexpected_results": dict(sorted(unexpected_results.items())),
"incomplete_results": sorted(incomplete_results),
"duplicate_results": duplicate_cases,
"duplicate_caselist_entries": dict(sorted(expected_duplicates.items())),
},
}
def build_matrix(suites: dict[str, tuple[str, Sequence[str]]]) -> dict:
"""Build all four version reports and their case-weighted aggregate."""
version_reports = {
version: build_version_report(version, *suites[version])
for version in VERSIONS
}
totals = {
key: sum(report[key] for report in version_reports.values())
for key in (
"expected",
"result",
"pass",
"accepted",
"crash",
"hang",
"unrun",
"duplicate",
)
}
status_counts: Counter[str] = Counter()
for report in version_reports.values():
status_counts.update(report["counts"])
states = {report["validation"]["state"] for report in version_reports.values()}
if "ERROR" in states:
overall_state = "ERROR"
elif "INCOMPLETE" in states:
overall_state = "INCOMPLETE"
else:
overall_state = "OK"
overall = {
**totals,
"counts": dict(sorted(status_counts.items())),
"aggregation": "weighted_by_expected_cases",
"coverage": {
"numerator": totals["result"],
"denominator": totals["expected"],
"rate": _ratio(totals["result"], totals["expected"]),
},
"rates": {
"denominator": "expected",
"strict_pass_only": _ratio(totals["pass"], totals["expected"]),
"conformance_accepted": _ratio(
totals["accepted"], totals["expected"]
),
"measured_only_strict_pass": _ratio(
totals["pass"], totals["result"]
),
"measured_only_conformance_accepted": _ratio(
totals["accepted"], totals["result"]
),
},
"strict_pass_rate": _ratio(totals["pass"], totals["expected"]),
"conformance_accepted_rate": _ratio(
totals["accepted"], totals["expected"]
),
"validation": {
"state": overall_state,
"ok": overall_state == "OK",
"invariant_expected_equals_result_plus_unrun": (
totals["expected"] == totals["result"] + totals["unrun"]
),
},
}
return {
"schema_version": 1,
"generated_at": datetime.now(timezone.utc).isoformat(),
"accepted_statuses": list(ACCEPTED_STATUSES),
"rate_policy": {
"denominator": "unique expected cases from each caselist",
"unrun_cases": "included in the denominator and never accepted",
"duplicate_results": "last QPA result wins, matching qpa_report.py",
},
"versions": version_reports,
"overall": overall,
}
def _percent(numerator: int, denominator: int) -> str:
if not denominator:
return "n/a"
return f"{100.0 * numerator / denominator:.2f}% ({numerator}/{denominator})"
def render_markdown(report: dict) -> str:
"""Render the compact terminal-facing conformance table."""
header = (
"| Suite | Expected | Result | Pass | Accepted | Crash | Hang | Unrun | "
"Duplicate | Coverage | Strict Pass-only | Conformance-accepted | Validation |"
)
separator = (
"|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|:---:|"
)
rows = [header, separator]
for version in VERSIONS:
item = report["versions"][version]
rows.append(
"| {version} | {expected} | {result} | {pass_count} | {accepted} | "
"{crash} | {hang} | {unrun} | {duplicate} | {coverage} | {strict} | "
"{accepted_rate} | {state} |".format(
version=version.upper(),
expected=item["expected"],
result=item["result"],
pass_count=item["pass"],
accepted=item["accepted"],
crash=item["crash"],
hang=item["hang"],
unrun=item["unrun"],
duplicate=item["duplicate"],
coverage=_percent(item["result"], item["expected"]),
strict=_percent(item["pass"], item["expected"]),
accepted_rate=_percent(item["accepted"], item["expected"]),
state=item["validation"]["state"],
)
)
overall = report["overall"]
rows.append(
"| **Overall (weighted)** | **{expected}** | **{result}** | **{pass_count}** | "
"**{accepted}** | **{crash}** | **{hang}** | **{unrun}** | **{duplicate}** | "
"**{coverage}** | **{strict}** | **{accepted_rate}** | **{state}** |".format(
expected=overall["expected"],
result=overall["result"],
pass_count=overall["pass"],
accepted=overall["accepted"],
crash=overall["crash"],
hang=overall["hang"],
unrun=overall["unrun"],
duplicate=overall["duplicate"],
coverage=_percent(overall["result"], overall["expected"]),
strict=_percent(overall["pass"], overall["expected"]),
accepted_rate=_percent(overall["accepted"], overall["expected"]),
state=overall["validation"]["state"],
)
)
rows.extend(
(
"",
"Rates use unique **Expected** caselist cases as the denominator; unrun cases "
"remain in that denominator and are not accepted.",
"Accepted statuses: " + ", ".join(f"`{s}`" for s in ACCEPTED_STATUSES) + ".",
"Duplicate is the number of extra QPA observations; the last observation wins.",
)
)
details: list[str] = []
for version in VERSIONS:
validation = report["versions"][version]["validation"]
messages = validation["errors"] + validation["warnings"]
if messages:
details.append(
f"- **{version.upper()} {validation['state']}**: " + "; ".join(messages)
)
if details:
rows.extend(("", "Validation details:", "", *details))
return "\n".join(rows)
def _write_json(path: str, report: dict) -> None:
parent = os.path.dirname(os.path.abspath(path))
os.makedirs(parent, exist_ok=True)
with open(path, "w", encoding="utf-8", newline="\n") as fh:
json.dump(report, fh, indent=2, sort_keys=True)
fh.write("\n")
def _parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description=__doc__)
for version in VERSIONS:
parser.add_argument(
f"--{version}-caselist",
f"--{version}-case-list",
required=True,
help=f"{version.upper()} mustpass caselist",
)
parser.add_argument(
f"--{version}-results",
f"--{version}-result-dir",
f"--{version}-results-dir",
action="append",
required=True,
help=f"{version.upper()} result directory or QPA file (repeatable)",
)
parser.add_argument(
"--json",
dest="json_out",
default="cts_matrix_report.json",
help="JSON output path (default: ./cts_matrix_report.json)",
)
parser.add_argument(
"--allow-incomplete",
action="store_true",
help="return success even when validation is ERROR/INCOMPLETE",
)
return parser
def main(argv: Optional[Iterable[str]] = None) -> int:
args = _parser().parse_args(argv)
suites = {
version: (
getattr(args, f"{version}_caselist"),
getattr(args, f"{version}_results"),
)
for version in VERSIONS
}
try:
report = build_matrix(suites)
_write_json(args.json_out, report)
except (OSError, ReportInputError) as exc:
print(f"cts_matrix_report: {exc}", file=sys.stderr)
return 2
print(render_markdown(report))
print(f"\nJSON: {os.path.abspath(args.json_out)}")
if not args.allow_incomplete and not report["overall"]["validation"]["ok"]:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
+540
View File
@@ -0,0 +1,540 @@
#!/usr/bin/env python
"""Summarise any number of GL CTS suites and MobileGL backends.
Each repeatable suite specification consists of four values: backend, label,
caselist, and result directory. For example::
python cts_multi_report.py \
--suite DirectGLES gl30 gl30-main.txt runs/gles/gl30 \
--suite DirectVulkan gl30 gl30-main.txt runs/vulkan/gl30 \
--markdown cts-summary.md --json cts-summary.json
A compact comma form is accepted as well::
--suite=DirectGLES,gl31,gl31-main.txt,runs/gles/gl31
Per-suite parsing and validation deliberately delegate to
``cts_matrix_report`` so QPA ordering, sidecar upgrades, accepted statuses,
unrun handling, and duplicate-result semantics cannot drift between reports.
All conformance rates use unique expected caselist cases as their denominator.
Backend subtotals and the overall total are therefore case-weighted, not an
unweighted average of suite percentages.
"""
from __future__ import annotations
import argparse
from collections import Counter
from dataclasses import dataclass
from datetime import datetime, timezone
import hashlib
import json
import os
import sys
from typing import Iterable, Optional, Sequence
try: # Direct script execution and package imports are both supported.
from . import cts_matrix_report
except ImportError: # pragma: no cover - covered through CLI-style tests
import cts_matrix_report
SUPPORTED_BACKENDS = ("DirectGLES", "DirectVulkan")
SUM_FIELDS = (
"expected",
"result",
"pass",
"accepted",
"crash",
"hang",
"unrun",
"duplicate",
)
class MultiReportInputError(ValueError):
"""Suite specifications cannot produce an unambiguous report."""
@dataclass(frozen=True)
class SuiteSpec:
backend: str
label: str
caselist: str
result_dir: str
@property
def suite_id(self) -> str:
return f"{self.backend}/{self.label}"
def _ratio(numerator: int, denominator: int) -> float:
return numerator / denominator if denominator else 0.0
def _validation_state(items: Sequence[dict]) -> str:
states = {item["validation"]["state"] for item in items}
if "ERROR" in states:
return "ERROR"
if "INCOMPLETE" in states:
return "INCOMPLETE"
return "OK"
def aggregate_reports(items: Sequence[dict], suite_state_keys: Sequence[str]) -> dict:
"""Return an expected-case-weighted aggregate for suite reports."""
if len(items) != len(suite_state_keys):
raise MultiReportInputError("internal suite/state key count mismatch")
totals = {
field: sum(int(item[field]) for item in items)
for field in SUM_FIELDS
}
status_counts: Counter[str] = Counter()
for item in items:
status_counts.update(item["counts"])
state = _validation_state(items)
expected = totals["expected"]
result = totals["result"]
suite_states = {
key: item["validation"]["state"]
for key, item in zip(suite_state_keys, items)
}
return {
**totals,
"suite_count": len(items),
"counts": dict(sorted(status_counts.items())),
"aggregation": "weighted_by_expected_cases",
"coverage": {
"numerator": result,
"denominator": expected,
"rate": _ratio(result, expected),
},
"rates": {
"denominator": "expected",
"strict_pass_only": _ratio(totals["pass"], expected),
"conformance_accepted": _ratio(totals["accepted"], expected),
"measured_only_strict_pass": _ratio(totals["pass"], result),
"measured_only_conformance_accepted": _ratio(
totals["accepted"], result
),
},
# Keep the convenient aliases used by cts_matrix_report consumers.
"strict_pass_rate": _ratio(totals["pass"], expected),
"conformance_accepted_rate": _ratio(totals["accepted"], expected),
"validation": {
"state": state,
"ok": state == "OK",
"suite_states": suite_states,
"invariant_expected_equals_result_plus_unrun": (
expected == result + totals["unrun"]
),
},
}
def _caselist_fingerprint(path: str) -> tuple[str, int]:
cases, _duplicates = cts_matrix_report.read_caselist(path)
payload = "\n".join(cases).encode("utf-8") + b"\n"
return hashlib.sha256(payload).hexdigest(), len(cases)
def _read_provenance(
spec: SuiteSpec,
require_run_state: bool,
expected_run_identity: Optional[str],
) -> dict:
path = os.path.join(spec.result_dir, "run_state.json")
if not os.path.isfile(path):
if require_run_state or expected_run_identity is not None:
raise MultiReportInputError(
"suite result directory has no run_state.json; invocation provenance "
f"cannot be verified: {spec.result_dir}"
)
return {"state": "UNVERIFIED", "run_state": None}
try:
with open(path, "r", encoding="utf-8") as handle:
state = json.load(handle)
except (OSError, UnicodeError, json.JSONDecodeError) as exc:
raise MultiReportInputError(f"cannot read suite run identity {path}: {exc}") from exc
if not isinstance(state, dict):
raise MultiReportInputError(f"suite run identity must be a JSON object: {path}")
if state.get("backend") != spec.backend:
raise MultiReportInputError(
f"suite {spec.suite_id} is labelled {spec.backend}, but run_state.json "
f"records {state.get('backend')!r}"
)
fingerprint, case_count = _caselist_fingerprint(spec.caselist)
if state.get("caselist_sha256") != fingerprint or state.get("case_count") != case_count:
raise MultiReportInputError(
f"suite {spec.suite_id} run_state.json belongs to a different caselist"
)
invocation_identity = state.get("invocation_identity")
if (
expected_run_identity is not None
and invocation_identity != expected_run_identity
):
raise MultiReportInputError(
f"suite {spec.suite_id} run_state.json belongs to a different CTS invocation"
)
return {
"state": "VERIFIED",
"run_state": os.path.abspath(path),
"invocation_identity": invocation_identity,
}
def _validate_specs(
specs: Sequence[SuiteSpec],
require_run_state: bool,
expected_run_identity: Optional[str],
) -> dict[str, dict]:
if not specs:
raise MultiReportInputError("at least one --suite specification is required")
seen: set[tuple[str, str]] = set()
seen_result_dirs: list[tuple[str, str]] = []
provenance: dict[str, dict] = {}
for spec in specs:
if spec.backend not in SUPPORTED_BACKENDS:
raise MultiReportInputError(
f"unsupported backend {spec.backend!r}; expected one of "
+ ", ".join(SUPPORTED_BACKENDS)
)
if not spec.label.strip():
raise MultiReportInputError("suite label cannot be empty")
identity = (spec.backend, spec.label)
if identity in seen:
raise MultiReportInputError(
f"duplicate suite specification for {spec.backend}/{spec.label}"
)
seen.add(identity)
if not os.path.isdir(spec.result_dir):
raise MultiReportInputError(
f"suite result directory does not exist: {spec.result_dir}"
)
physical_result_dir = os.path.normcase(
os.path.realpath(os.path.abspath(spec.result_dir))
)
for previous_dir, previous_suite in seen_result_dirs:
try:
common_dir = os.path.commonpath(
[previous_dir, physical_result_dir]
)
except ValueError:
continue
if common_dir in (previous_dir, physical_result_dir):
raise MultiReportInputError(
f"suite {spec.suite_id} uses a result directory which overlaps "
f"{previous_suite}: {spec.result_dir}"
)
seen_result_dirs.append((physical_result_dir, spec.suite_id))
provenance[spec.suite_id] = _read_provenance(
spec, require_run_state, expected_run_identity
)
return provenance
def build_report(
specs: Sequence[SuiteSpec],
require_run_state: bool = True,
expected_run_identity: Optional[str] = None,
) -> dict:
"""Build suite, per-backend, and overall serialisable reports."""
provenance = _validate_specs(
specs, require_run_state, expected_run_identity
)
suite_reports: list[dict] = []
backend_order: list[str] = []
for spec in specs:
if spec.backend not in backend_order:
backend_order.append(spec.backend)
item = cts_matrix_report.build_version_report(
spec.label, spec.caselist, [spec.result_dir]
)
# ``version`` is the generic label argument in build_version_report;
# expose explicit multi-report terminology while retaining all of its
# validation and case-level evidence.
item.pop("version", None)
item["backend"] = spec.backend
item["label"] = spec.label
item["suite_id"] = spec.suite_id
item["provenance"] = provenance[spec.suite_id]
if item["provenance"]["state"] == "UNVERIFIED":
item["validation"]["warnings"].append(
"result directory has no run_state.json; backend provenance is unverified"
)
suite_reports.append(item)
backends: dict[str, dict] = {}
for backend in backend_order:
backend_items = [
item for item in suite_reports if item["backend"] == backend
]
labels = [item["label"] for item in backend_items]
aggregate = aggregate_reports(backend_items, labels)
aggregate["backend"] = backend
aggregate["suite_labels"] = labels
backends[backend] = aggregate
overall = aggregate_reports(
suite_reports, [item["suite_id"] for item in suite_reports]
)
overall["backend_count"] = len(backends)
overall["backends"] = backend_order
return {
"schema_version": 1,
"generated_at": datetime.now(timezone.utc).isoformat(),
"accepted_statuses": list(cts_matrix_report.ACCEPTED_STATUSES),
"rate_policy": {
"denominator": "unique expected cases from each suite caselist",
"unrun_cases": "included in the denominator and never accepted",
"backend_aggregation": "weighted by expected cases",
"overall_aggregation": "weighted by expected cases across backend-suite pairs",
"duplicate_results": "last QPA result wins, matching qpa_report.py",
},
"suites": suite_reports,
"backends": backends,
"overall": overall,
}
def _percent(numerator: int, denominator: int) -> str:
if not denominator:
return "n/a"
return f"{100.0 * numerator / denominator:.2f}% ({numerator}/{denominator})"
def _markdown_cell(value: object) -> str:
return str(value).replace("|", r"\|").replace("\r", " ").replace("\n", " ")
def _table_row(backend: str, label: str, item: dict, bold: bool = False) -> str:
values = [
backend,
label,
str(item["expected"]),
str(item["result"]),
str(item["pass"]),
str(item["accepted"]),
str(item["crash"]),
str(item["hang"]),
str(item["unrun"]),
str(item["duplicate"]),
_percent(item["result"], item["expected"]),
_percent(item["pass"], item["expected"]),
_percent(item["accepted"], item["expected"]),
item["validation"]["state"],
]
values = [_markdown_cell(value) for value in values]
if bold:
values = [f"**{value}**" for value in values]
return "| " + " | ".join(values) + " |"
def render_markdown(report: dict) -> str:
lines = [
"# GL CTS multi-suite conformance report",
"",
(
"| Backend | Suite | Expected | Result | Pass | Accepted | Crash | Hang | "
"Unrun | Duplicate | Coverage | Strict Pass-only | Conformance-accepted | Validation |"
),
"|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|:---:|",
]
for backend in report["backends"]:
for item in report["suites"]:
if item["backend"] == backend:
lines.append(_table_row(backend, item["label"], item))
subtotal = report["backends"][backend]
lines.append(
_table_row(backend, f"{backend} weighted subtotal", subtotal, bold=True)
)
lines.append(
_table_row(
"All backends",
"Overall weighted",
report["overall"],
bold=True,
)
)
lines.extend(
[
"",
(
"Rates use unique **Expected** caselist cases as the denominator. "
"Unrun cases remain in the denominator and are not accepted."
),
"Accepted statuses: "
+ ", ".join(
f"`{status}`" for status in report["accepted_statuses"]
)
+ ".",
(
"Duplicate is the number of extra QPA observations; the final "
"observation wins."
),
]
)
details: list[str] = []
for item in report["suites"]:
validation = item["validation"]
messages = validation["errors"] + validation["warnings"]
if messages:
details.append(
f"- **{_markdown_cell(item['suite_id'])} {validation['state']}**: "
+ "; ".join(_markdown_cell(message) for message in messages)
)
if details:
lines.extend(["", "## Validation details", "", *details])
return "\n".join(lines) + "\n"
def _write_text(path: str, contents: str) -> None:
absolute = os.path.abspath(path)
os.makedirs(os.path.dirname(absolute), exist_ok=True)
with open(absolute, "w", encoding="utf-8", newline="\n") as handle:
handle.write(contents)
def _write_json(path: str, report: dict) -> None:
_write_text(path, json.dumps(report, indent=2, sort_keys=True) + "\n")
def _normalise_compact_suite_args(argv: Sequence[str]) -> list[str]:
"""Expand ``--suite=b,l,c,r`` into the four-value argparse form."""
result: list[str] = []
index = 0
while index < len(argv):
token = argv[index]
if token.startswith("--suite="):
compact = token.split("=", 1)[1]
parts = compact.split(",", 3)
if len(parts) != 4:
raise MultiReportInputError(
"compact --suite expects backend,label,caselist,result-dir"
)
result.extend(["--suite", *parts])
index += 1
continue
if token == "--suite" and index + 1 < len(argv) and argv[index + 1].count(",") >= 3:
parts = argv[index + 1].split(",", 3)
result.extend(["--suite", *parts])
index += 2
continue
result.append(token)
index += 1
return result
def _parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--suite",
action="append",
nargs=4,
required=True,
metavar=("BACKEND", "LABEL", "CASELIST", "RESULT_DIR"),
help=(
"suite specification; repeat for every backend/suite pair "
f"(backends: {', '.join(SUPPORTED_BACKENDS)})"
),
)
parser.add_argument(
"--markdown",
default="cts_multi_report.md",
help="Markdown output path (default: ./cts_multi_report.md)",
)
parser.add_argument(
"--json",
dest="json_out",
default="cts_multi_report.json",
help="JSON output path (default: ./cts_multi_report.json)",
)
parser.add_argument(
"--allow-incomplete",
action="store_true",
help="return success even when one or more suites are ERROR/INCOMPLETE",
)
parser.add_argument(
"--adopt-legacy",
dest="allow_unverified_provenance",
action="store_true",
help="accept legacy result directories without run_state.json (provenance remains unverified)",
)
parser.add_argument(
"--allow-unverified-provenance",
dest="allow_unverified_provenance",
action="store_true",
help=argparse.SUPPRESS,
)
parser.add_argument(
"--expected-run-identity",
help="require every suite run_state.json to contain this controller fingerprint",
)
return parser
def _specs_from_args(values: Sequence[Sequence[str]]) -> list[SuiteSpec]:
return [
SuiteSpec(
backend=backend.strip(),
label=label.strip(),
caselist=caselist,
result_dir=result_dir,
)
for backend, label, caselist, result_dir in values
]
def main(argv: Optional[Iterable[str]] = None) -> int:
raw_argv = list(argv) if argv is not None else sys.argv[1:]
try:
normalised = _normalise_compact_suite_args(raw_argv)
except MultiReportInputError as exc:
print(f"cts_multi_report: {exc}", file=sys.stderr)
return 2
args = _parser().parse_args(normalised)
if os.path.normcase(os.path.abspath(args.markdown)) == os.path.normcase(
os.path.abspath(args.json_out)
):
print("cts_multi_report: Markdown and JSON paths must differ", file=sys.stderr)
return 2
try:
report = build_report(
_specs_from_args(args.suite),
require_run_state=not args.allow_unverified_provenance,
expected_run_identity=args.expected_run_identity,
)
markdown = render_markdown(report)
_write_text(args.markdown, markdown)
_write_json(args.json_out, report)
except (
OSError,
MultiReportInputError,
cts_matrix_report.ReportInputError,
) as exc:
print(f"cts_multi_report: {exc}", file=sys.stderr)
return 2
print(markdown, end="")
print(f"\nMarkdown: {os.path.abspath(args.markdown)}")
print(f"JSON: {os.path.abspath(args.json_out)}")
if not args.allow_incomplete and not report["overall"]["validation"]["ok"]:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
+204
View File
@@ -0,0 +1,204 @@
#!/usr/bin/env python
"""Summarise dEQP/glcts .qpa logs into a conformance pass rate.
Handles the two ways a case can end in a .qpa: a normal
``#beginTestCaseResult``/``#endTestCaseResult`` pair carrying a
``<Result StatusCode="...">`` element, and ``#terminateTestCaseResult <reason>``,
which is what the log contains when the process died partway through a case.
Cases that were started but never terminated (the run was killed) are reported
separately so a truncated chunk is never silently scored as a pass.
Usage:
python qpa_report.py <file-or-dir> [<file-or-dir> ...] [--json out.json] [--top N]
"""
import argparse
import json
import os
import re
import sys
from collections import Counter, defaultdict
# Khronos conformance treats these as non-failures: the test either passed or
# the implementation legitimately does not expose the feature under test.
NON_FAILURE = {
"Pass",
"NotSupported",
"QualityWarning",
"CompatibilityWarning",
"Waiver",
}
# Statuses that indicate the case did not merely fail but destabilised the run.
HARD = {"Crash", "Timeout", "InternalError", "ResourceError", "DeviceHang"}
CASE_START = re.compile(r"^#beginTestCaseResult\s+(\S+)")
CASE_END = re.compile(r"^#endTestCaseResult")
CASE_TERM = re.compile(r"^#terminateTestCaseResult\s+(.*)")
RESULT = re.compile(r'<Result\s+StatusCode="([^"]+)"')
def parse_qpa(path):
"""Yield (case_name, status) for every case recorded in one .qpa file."""
current = None
status = None
with open(path, "r", encoding="utf-8", errors="replace") as fh:
for line in fh:
m = CASE_START.match(line)
if m:
if current is not None:
# A new case started before the previous one closed.
yield current, status or "Incomplete"
current, status = m.group(1), None
continue
if current is None:
continue
m = RESULT.search(line)
if m:
status = m.group(1)
continue
m = CASE_TERM.match(line)
if m:
reason = m.group(1).strip() or "Terminated"
# dEQP writes e.g. "Crash" / "Timeout" here.
yield current, reason if reason in HARD else "Crash"
current, status = None, None
continue
if CASE_END.match(line):
yield current, status or "Incomplete"
current, status = None, None
if current is not None:
# File ended mid-case: the runner was killed.
yield current, "Incomplete"
def collect(paths):
files = []
for p in paths:
if os.path.isdir(p):
for root, _dirs, names in os.walk(p):
files.extend(os.path.join(root, n) for n in sorted(names) if n.endswith(".qpa"))
else:
files.append(p)
return files
def group_of(case):
"""The case's parent group, e.g. KHR-GL33.shaders.arrays for ...arrays.foo."""
parts = case.split(".")
return ".".join(parts[:-1]) if len(parts) > 1 else case
def load_sidecar(paths, name):
"""Case names run_cts.py recorded in one of its sidecar lists."""
out = set()
for p in paths:
d = p if os.path.isdir(p) else os.path.dirname(p)
f = os.path.join(d, name)
if os.path.isfile(f):
with open(f, "r", encoding="utf-8") as fh:
out.update(l.strip() for l in fh if l.strip() and not l.strip().startswith("#"))
return out
def main():
ap = argparse.ArgumentParser()
ap.add_argument("paths", nargs="+")
ap.add_argument("--json", dest="json_out")
ap.add_argument("--top", type=int, default=25)
ap.add_argument("--label", default="")
args = ap.parse_args()
files = collect(args.paths)
if not files:
print("no .qpa files found", file=sys.stderr)
return 2
# Later chunks may re-run a case; last result wins.
results = {}
for f in files:
for case, status in parse_qpa(f):
results[case] = status
# A case the runner saw take the process down is a Crash, not merely an
# unterminated log entry - but a real result from a later retry wins.
for case in load_sidecar(args.paths, "crashed.txt"):
if results.get(case, "Incomplete") == "Incomplete":
results[case] = "Crash"
# Worse than a crash: these rebooted the device.
for case in load_sidecar(args.paths, "hung.txt"):
if results.get(case, "Incomplete") in ("Incomplete", "Crash"):
results[case] = "DeviceHang"
# Cases excluded up front, and cases the run never reached, are not results.
# Report them separately so a partial run is never read as a complete one.
skipped = load_sidecar(args.paths, "skipped.txt")
unrun = load_sidecar(args.paths, "unrun.txt") - set(results)
counts = Counter(results.values())
total = len(results)
non_fail = sum(counts[s] for s in NON_FAILURE)
strict_pass = counts["Pass"]
failures = total - non_fail
by_group_fail = defaultdict(int)
by_group_total = defaultdict(int)
for case, status in results.items():
g = group_of(case)
by_group_total[g] += 1
if status not in NON_FAILURE:
by_group_fail[g] += 1
label = f" [{args.label}]" if args.label else ""
print(f"=== glcts conformance summary{label} ===")
print(f"files parsed : {len(files)}")
print(f"cases with result : {total}")
print()
for status, n in counts.most_common():
mark = " " if status in NON_FAILURE else " ! "
print(f"{mark}{status:<22} {n:>7} {100.0 * n / total:6.2f}%")
print()
if total:
print(f"conformance pass rate (Pass+NotSupported+warnings) : {100.0 * non_fail / total:6.2f}% ({non_fail}/{total})")
print(f"strict pass rate (Pass only) : {100.0 * strict_pass / total:6.2f}% ({strict_pass}/{total})")
print(f"failures : {failures}")
if skipped or unrun:
print("\n--- NOT MEASURED (excluded from the rates above) ---")
if skipped:
print(f" quarantined up front : {len(skipped)}")
if unrun:
print(f" never reached : {len(unrun)}")
print(" The rates above cover only cases that produced a result.")
if failures:
print(f"\n--- worst groups (of {len(by_group_total)}) ---")
worst = sorted(by_group_fail.items(), key=lambda kv: -kv[1])[: args.top]
for g, nf in worst:
nt = by_group_total[g]
print(f" {g:<52} {nf:>6}/{nt:<6} fail ({100.0 * nf / nt:5.1f}%)")
if args.json_out:
with open(args.json_out, "w", encoding="utf-8") as fh:
json.dump(
{
"label": args.label,
"files": len(files),
"total": total,
"counts": dict(counts),
"non_failure": non_fail,
"strict_pass": strict_pass,
"failures": failures,
"pass_rate": (non_fail / total) if total else 0.0,
"strict_pass_rate": (strict_pass / total) if total else 0.0,
"results": results,
},
fh,
indent=1,
)
print(f"\nwrote {args.json_out}")
return 0
if __name__ == "__main__":
sys.exit(main())
+269
View File
@@ -0,0 +1,269 @@
#!/usr/bin/env python
"""Drive a glcts run on a device, resuming across crashes.
MobileGL crashes on some cases, and glcts takes the whole process down with it.
A single invocation would therefore stop at the first crash and leave most of
the suite unmeasured. This runner re-invokes glcts with only the cases that have
not produced a result yet, records each crashed case as "Crash", and repeats
until the list is exhausted, so one bad case costs one case rather than the run.
Usage:
python run_cts.py --serial <adb-serial> --backend DirectGLES|DirectVulkan \\
--caselist <host-path-to-mustpass.txt> --outdir <host-dir> [--device-dir /data/local/tmp/mgcts]
"""
import argparse
import os
import re
import subprocess
import sys
import time
CASE_START = re.compile(r"^#beginTestCaseResult\s+(\S+)")
CASE_END = re.compile(r"^#endTestCaseResult")
CASE_TERM = re.compile(r"^#terminateTestCaseResult\s+(.*)")
def adb(serial, *args, timeout=None):
try:
return subprocess.run(["adb", "-s", serial, *args], capture_output=True, text=True, timeout=timeout)
except subprocess.TimeoutExpired:
return subprocess.CompletedProcess(args, returncode=124, stdout="", stderr="adb timeout")
def device_alive(serial, timeout=30):
"""True only if the device answers a trivial shell command.
Distinguishes "glcts crashed" from "the device fell over". Without this a
dead device looks like every remaining case crashing, which silently turns a
broken run into a plausible-looking conformance number.
"""
r = adb(serial, "shell", "echo alive", timeout=timeout)
return r.returncode == 0 and "alive" in (r.stdout or "")
def wait_for_device(serial, attempts=20, delay=15):
for i in range(attempts):
if device_alive(serial):
return True
print(f"[run_cts] device {serial} unresponsive, waiting ({i + 1}/{attempts})")
time.sleep(delay)
return False
def mem_available_kb(serial):
r = adb(serial, "shell", "grep MemAvailable /proc/meminfo", timeout=30)
m = re.search(r"(\d+)", r.stdout or "")
return int(m.group(1)) if m else None
def completed_cases(qpa_path):
"""Return (finished_case_names, last_started_case_or_None).
A case that was started but never closed is the one the process died in.
"""
finished = []
current = None
if not os.path.exists(qpa_path):
return finished, None
with open(qpa_path, "r", encoding="utf-8", errors="replace") as fh:
for line in fh:
m = CASE_START.match(line)
if m:
current = m.group(1)
continue
if current is not None and (CASE_END.match(line) or CASE_TERM.match(line)):
finished.append(current)
current = None
return finished, current
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--serial", required=True)
ap.add_argument("--backend", required=True, choices=["DirectGLES", "DirectVulkan"])
ap.add_argument("--caselist", required=True)
ap.add_argument("--outdir", required=True)
ap.add_argument("--device-dir", default="/data/local/tmp/mgcts")
ap.add_argument("--surface", default="fbo", help="--deqp-surface-type value")
ap.add_argument("--max-rounds", type=int, default=4000)
ap.add_argument("--max-empty-streak", type=int, default=64,
help="abort after this many consecutive chunks that produce no log at all")
ap.add_argument("--min-mem-kb", type=int, default=400000,
help="pause when the device drops below this much available memory")
ap.add_argument("--chunk-timeout", type=int, default=900,
help="seconds before giving up on one glcts invocation (a GPU hang never returns)")
ap.add_argument("--skip-file", default=None,
help="file of case names to exclude, e.g. cases known to hang the device")
ap.add_argument("--env", action="append", default=[], metavar="K=V",
help="extra environment variable for glcts (repeatable)")
args = ap.parse_args()
os.makedirs(args.outdir, exist_ok=True)
with open(args.caselist, "r", encoding="utf-8") as fh:
remaining = [l.strip() for l in fh if l.strip() and not l.strip().startswith("#")]
skipped = []
if args.skip_file and os.path.isfile(args.skip_file):
with open(args.skip_file, "r", encoding="utf-8") as fh:
skip = {l.strip() for l in fh if l.strip() and not l.strip().startswith("#")}
skipped = [c for c in remaining if c in skip]
remaining = [c for c in remaining if c not in skip]
print(f"[run_cts] skipping {len(skipped)} case(s) from {args.skip_file}")
total = len(remaining)
print(f"[run_cts] {args.backend} on {args.serial}: {total} cases")
crashed = []
hung = []
done = set()
chunk = 0
started = time.time()
empty_streak = 0
if not wait_for_device(args.serial):
print("[run_cts] device not responding before start; aborting", file=sys.stderr)
return 3
while remaining and chunk < args.max_rounds:
listfile = os.path.join(args.outdir, "remaining.txt")
with open(listfile, "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(remaining) + "\n")
# Repeated process launches plus crash tombstones can drive the device
# into memory pressure; give it room rather than pushing it over.
mem = mem_available_kb(args.serial)
if mem is not None and mem < args.min_mem_kb:
print(f"[run_cts] low memory ({mem} kB available); pausing 30 s")
time.sleep(30)
dev_list = f"{args.device_dir}/remaining.txt"
dev_qpa = f"{args.device_dir}/chunk.qpa"
push = adb(args.serial, "push", listfile, dev_list, timeout=120)
if push.returncode != 0:
print(f"[run_cts] push failed ({push.stderr.strip()}); treating as device trouble",
file=sys.stderr)
if not wait_for_device(args.serial):
print("[run_cts] ABORTING: device unreachable.", file=sys.stderr)
break
continue
adb(args.serial, "shell", f"rm -f {dev_qpa}", timeout=60)
extra_env = "".join(f"{kv} " for kv in args.env)
cmd = (
f"cd {args.device_dir} && "
f"MOBILEGL_BACKEND_TYPE={args.backend} LD_LIBRARY_PATH=. {extra_env}"
f"./glcts --deqp-caselist-file={dev_list} "
f"--deqp-surface-type={args.surface} "
f"--deqp-terminate-on-device-lost=disable "
f"--deqp-log-images=disable --deqp-log-shader-sources=disable "
f"--deqp-log-filename={dev_qpa} > /dev/null 2>&1; echo RC=$?"
)
run = adb(args.serial, "shell", cmd, timeout=args.chunk_timeout)
if run.returncode == 124:
print(f"[run_cts] chunk {chunk:04d} timed out after {args.chunk_timeout}s "
f"(likely a GPU hang)", file=sys.stderr)
# Some cases hang the GPU hard enough to reboot the device. The log on
# /data/local/tmp survives that, so wait for the device to come back and
# pull it anyway rather than losing the whole chunk.
rebooted = False
if not device_alive(args.serial, timeout=30):
print(f"[run_cts] device went away during chunk {chunk:04d}; waiting for it",
file=sys.stderr)
if not wait_for_device(args.serial, attempts=40, delay=15):
print("[run_cts] ABORTING: device never came back. Results are incomplete; "
"do NOT treat the remaining cases as failures.", file=sys.stderr)
break
rebooted = True
print("[run_cts] device is back")
local_qpa = os.path.join(args.outdir, f"chunk{chunk:04d}.qpa")
pull = adb(args.serial, "pull", dev_qpa, local_qpa, timeout=300)
if pull.returncode != 0 and rebooted:
time.sleep(10)
adb(args.serial, "pull", dev_qpa, local_qpa, timeout=300)
finished, in_flight = completed_cases(local_qpa)
for c in finished:
done.add(c)
progressed = len(finished)
if progressed > 0:
empty_streak = 0
if in_flight is not None:
# The case that was open when the process (or the device) died.
if rebooted:
# It took the whole device down: quarantine it, or the next
# invocation walks straight back into it.
print(f"[run_cts] DEVICE HANG in {in_flight} - quarantining it")
hung.append(in_flight)
else:
crashed.append(in_flight)
done.add(in_flight)
progressed += 1
elif progressed == 0:
# Nothing at all came back. Either the first remaining case takes
# the process down before the log is flushed, or the device died.
# Those look identical from here, so confirm the device is alive
# before blaming the test.
if not device_alive(args.serial):
print(f"[run_cts] device went away during chunk {chunk:04d}", file=sys.stderr)
if not wait_for_device(args.serial):
print("[run_cts] ABORTING: device never came back. Results are "
"incomplete; do NOT treat the remaining cases as crashes.", file=sys.stderr)
break
print("[run_cts] device recovered; retrying the same chunk")
continue
empty_streak += 1
if empty_streak >= args.max_empty_streak:
print(f"[run_cts] ABORTING: {empty_streak} consecutive chunks produced no output "
f"while the device stayed reachable. Something systemic is wrong; refusing "
f"to label the rest of the suite as crashes.", file=sys.stderr)
break
victim = remaining[0]
print(f"[run_cts] no output at all; recording {victim} as Crash")
crashed.append(victim)
done.add(victim)
progressed = 1
remaining = [c for c in remaining if c not in done]
elapsed = time.time() - started
print(
f"[run_cts] chunk {chunk:04d}: +{progressed} (done {len(done)}/{total}, "
f"crashes {len(crashed)}, {elapsed / 60:.1f} min)"
)
chunk += 1
with open(os.path.join(args.outdir, "crashed.txt"), "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(crashed) + ("\n" if crashed else ""))
# Cases that rebooted the device. Feed this back in via --skip-file to avoid
# paying for the same reboot on the next run.
with open(os.path.join(args.outdir, "hung.txt"), "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(hung) + ("\n" if hung else ""))
if hung:
print(f"[run_cts] {len(hung)} case(s) hung the device (see hung.txt):")
for c in hung:
print(f" {c}")
# Anything still in `remaining` was never measured. Record it so the report
# cannot quietly present a partial run as a complete one.
with open(os.path.join(args.outdir, "unrun.txt"), "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(remaining) + ("\n" if remaining else ""))
if skipped:
with open(os.path.join(args.outdir, "skipped.txt"), "w", encoding="utf-8", newline="\n") as fh:
fh.write("\n".join(skipped) + "\n")
if remaining:
print(f"[run_cts] WARNING: {len(remaining)} cases were never run (see unrun.txt)", file=sys.stderr)
print(f"[run_cts] finished: {len(done)}/{total} cases, {len(crashed)} crashes, {chunk} invocations")
print(f"[run_cts] qpa chunks in {args.outdir}")
return 0
if __name__ == "__main__":
sys.exit(main())
+878
View File
@@ -0,0 +1,878 @@
#!/usr/bin/env python
"""Run a local Windows glcts executable and resume across process failures.
The desktop CTS normally runs a complete caselist in one process. That is a
poor fit for testing a developing OpenGL implementation: one access violation
or GPU hang prevents every later case from running. This driver gives each
invocation the cases which have not produced a result yet, preserves one QPA
and stdout/stderr pair per invocation, and starts another process after a
crash.
Existing ``chunkNNNN.qpa`` files and ``crashed.txt``/``hung.txt`` sidecars are
read on startup, so invoking the same command and output directory resumes an
interrupted run. A timeout is based on *idle QPA time*, not total process wall
time: a healthy invocation may legitimately run thousands of cases for hours.
Example (values beginning with ``--`` use argparse's ``=`` spelling)::
py run_cts_windows.py \
--exe D:\\glcts\\glcts.exe --workdir D:\\glcts \
--caselist D:\\glcts\\mustpass\\gl30.txt --outdir D:\\results\\gl30 \
--backend DirectVulkan \
--deqp-arg=--deqp-surface-type=window
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
from pathlib import Path
import re
import signal
import subprocess
import sys
import tempfile
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Iterable, Optional, Sequence
CASE_START = re.compile(r"^#beginTestCaseResult\s+(\S+)")
CASE_END = re.compile(r"^#endTestCaseResult(?:\s|$)")
CASE_TERM = re.compile(r"^#terminateTestCaseResult(?:\s|$)")
CASE_RESULT = re.compile(r'<Result\s+StatusCode="[^"]+"')
CHUNK_ARTIFACT = re.compile(r"^chunk(\d+)(?:\.|$)", re.IGNORECASE)
CHUNK_META = re.compile(r"^chunk(\d+)\.meta\.json$", re.IGNORECASE)
RECOVERY_SIDECAR_NAMES = frozenset(
{"crashed.txt", "hung.txt", "unrun.txt", "skipped.txt", "remaining.txt"}
)
CONTROLLED_DEQP_OPTIONS = {
"--deqp-caselist-file",
"--deqp-log-filename",
}
ATOMIC_REPLACE_ATTEMPTS = 8
ATOMIC_REPLACE_INITIAL_BACKOFF_SECONDS = 0.025
ATOMIC_REPLACE_MAX_BACKOFF_SECONDS = 0.2
class RunnerError(Exception):
"""A user/configuration error which should not be attributed to a case."""
@dataclass
class QpaProgress:
"""Cases recorded by a QPA and its unterminated tail, if any."""
recorded: list[str]
in_flight: Optional[str]
begin_count: int
@dataclass
class ProcessOutcome:
returncode: Optional[int]
duration_seconds: float
timed_out: bool = False
timeout_reason: Optional[str] = None
interrupted: bool = False
launch_error: Optional[str] = None
def utc_now() -> str:
return datetime.now(timezone.utc).isoformat(timespec="seconds")
def read_caselist(path: Path) -> list[str]:
"""Read a dEQP text caselist, preserving order and removing duplicates."""
try:
lines = path.read_text(encoding="utf-8-sig", errors="strict").splitlines()
except (OSError, UnicodeError) as exc:
raise RunnerError(f"cannot read caselist {path}: {exc}") from exc
cases: list[str] = []
seen: set[str] = set()
for raw in lines:
case = raw.strip()
if not case or case.startswith("#") or case in seen:
continue
cases.append(case)
seen.add(case)
if not cases:
raise RunnerError(f"caselist contains no test cases: {path}")
return cases
def read_name_set(path: Path) -> set[str]:
if not path.is_file():
return set()
try:
return {
line.strip()
for line in path.read_text(encoding="utf-8-sig", errors="replace").splitlines()
if line.strip() and not line.lstrip().startswith("#")
}
except OSError as exc:
raise RunnerError(f"cannot read recovery file {path}: {exc}") from exc
def _replace_with_retry(source: Path, destination: Path) -> None:
"""Replace a state file, tolerating brief Windows access-denied races.
Antivirus/indexing tools can momentarily open ``remaining.txt`` without
delete sharing. Windows then reports either ``PermissionError`` or a
generic ``OSError`` carrying ``winerror == 5``. Retry only those cases;
disk, path, and programming errors remain immediately visible.
"""
for attempt in range(ATOMIC_REPLACE_ATTEMPTS):
try:
os.replace(source, destination)
return
except OSError as exc:
retryable = isinstance(exc, PermissionError) or getattr(exc, "winerror", None) == 5
if not retryable or attempt + 1 >= ATOMIC_REPLACE_ATTEMPTS:
raise
delay = min(
ATOMIC_REPLACE_INITIAL_BACKOFF_SECONDS * (2**attempt),
ATOMIC_REPLACE_MAX_BACKOFF_SECONDS,
)
time.sleep(delay)
def atomic_write_text(path: Path, text: str) -> None:
"""Replace a small state file without exposing a partially-written copy."""
path.parent.mkdir(parents=True, exist_ok=True)
fd, temporary = tempfile.mkstemp(prefix=f".{path.name}.", suffix=".tmp", dir=str(path.parent))
temporary_path = Path(temporary)
try:
with os.fdopen(fd, "w", encoding="utf-8", newline="\n") as handle:
handle.write(text)
handle.flush()
os.fsync(handle.fileno())
_replace_with_retry(temporary_path, path)
finally:
try:
temporary_path.unlink()
except FileNotFoundError:
pass
def atomic_write_json(path: Path, value: object) -> None:
atomic_write_text(path, json.dumps(value, indent=2, sort_keys=True) + "\n")
def write_case_file(path: Path, cases: Iterable[str]) -> None:
values = list(cases)
atomic_write_text(path, "\n".join(values) + ("\n" if values else ""))
def scan_qpa(path: Path) -> QpaProgress:
"""Return cases with a final result and the unfinished tail, if any.
``#terminateTestCaseResult`` is a completed result (usually Crash or
Timeout). ``#endTestCaseResult`` only completes a case when its XML carried
a ``<Result StatusCode=...>``. A truncated case that already wrote Result is
also recoverable; a case with no Result remains eligible for a retry.
"""
if not path.is_file():
return QpaProgress([], None, 0)
recorded: list[str] = []
current: Optional[str] = None
has_result = False
begin_count = 0
try:
with path.open("r", encoding="utf-8", errors="replace") as handle:
for raw_line in handle:
line = raw_line.lstrip("\ufeff")
match = CASE_START.match(line)
if match:
if current is not None and has_result:
recorded.append(current)
current = match.group(1)
has_result = False
begin_count += 1
continue
if current is not None and CASE_RESULT.search(line):
has_result = True
continue
if current is not None and CASE_TERM.match(line):
recorded.append(current)
current = None
has_result = False
continue
if current is not None and CASE_END.match(line):
if has_result:
recorded.append(current)
current = None
has_result = False
except OSError as exc:
raise RunnerError(f"cannot read QPA {path}: {exc}") from exc
if current is not None and has_result:
recorded.append(current)
current = None
return QpaProgress(recorded, current, begin_count)
def numbered_files(outdir: Path, pattern: re.Pattern[str]) -> list[tuple[int, Path]]:
found: list[tuple[int, Path]] = []
try:
children = list(outdir.iterdir())
except OSError as exc:
raise RunnerError(f"cannot list output directory {outdir}: {exc}") from exc
for path in children:
match = pattern.match(path.name)
if match:
found.append((int(match.group(1)), path))
found.sort(key=lambda item: item[0])
return found
def next_chunk_number(outdir: Path) -> int:
numbers = [number for number, _path in numbered_files(outdir, CHUNK_ARTIFACT)]
return max(numbers, default=-1) + 1
def load_meta_classifications(outdir: Path, expected: set[str]) -> tuple[set[str], set[str]]:
"""Recover an atomic classification written just before sidecar updates."""
crashed: set[str] = set()
hung: set[str] = set()
for _number, path in numbered_files(outdir, CHUNK_META):
try:
value = json.loads(path.read_text(encoding="utf-8"))
except (OSError, UnicodeError, json.JSONDecodeError):
# A damaged metadata file is diagnostic only. QPA and sidecars are
# authoritative and must still allow recovery.
continue
if not isinstance(value, dict):
continue
case = value.get("classified_case")
classification = value.get("classification")
if not isinstance(case, str) or case not in expected:
continue
if classification == "DeviceHang":
hung.add(case)
elif classification == "Crash":
crashed.add(case)
crashed.difference_update(hung)
return crashed, hung
def result_qpa_files(outdir: Path) -> list[Path]:
"""Return every QPA a directory-based report would consume."""
found: list[Path] = []
try:
for root, directories, names in os.walk(outdir):
directories.sort(key=str.casefold)
for name in sorted(names, key=str.casefold):
if name.casefold().endswith(".qpa"):
found.append(Path(root) / name)
except OSError as exc:
raise RunnerError(f"cannot scan output directory {outdir}: {exc}") from exc
return found
def recover_results(outdir: Path, expected: set[str]) -> tuple[set[str], set[str], set[str]]:
recorded: set[str] = set()
for path in result_qpa_files(outdir):
progress = scan_qpa(path)
recorded.update(case for case in progress.recorded if case in expected)
crashed = read_name_set(outdir / "crashed.txt") & expected
hung = read_name_set(outdir / "hung.txt") & expected
meta_crashed, meta_hung = load_meta_classifications(outdir, expected)
crashed.update(meta_crashed)
hung.update(meta_hung)
crashed.difference_update(hung)
return recorded, crashed, hung
def caselist_fingerprint(cases: Sequence[str]) -> str:
payload = "\n".join(cases).encode("utf-8") + b"\n"
return hashlib.sha256(payload).hexdigest()
def recovery_artifacts(outdir: Path) -> list[Path]:
"""Return prior-run evidence which must not be adopted implicitly."""
found = set(result_qpa_files(outdir))
try:
children = list(outdir.iterdir())
except OSError as exc:
raise RunnerError(f"cannot list output directory {outdir}: {exc}") from exc
found.update(
path
for path in children
if CHUNK_ARTIFACT.match(path.name)
or path.name.casefold() in RECOVERY_SIDECAR_NAMES
)
return sorted(
found,
key=lambda path: str(path.relative_to(outdir)).casefold(),
)
def check_run_identity(
outdir: Path,
backend: str,
cases: Sequence[str],
invocation_identity: Optional[str] = None,
adopt_legacy: bool = False,
) -> None:
"""Refuse to silently mix different suites/backends in one result dir."""
path = outdir / "run_state.json"
fingerprint = caselist_fingerprint(cases)
if path.is_file():
try:
state = json.loads(path.read_text(encoding="utf-8"))
except (OSError, UnicodeError, json.JSONDecodeError) as exc:
raise RunnerError(f"cannot read run identity {path}: {exc}") from exc
if not isinstance(state, dict):
raise RunnerError(f"run identity must be a JSON object: {path}")
if state.get("backend") != backend:
raise RunnerError(
f"output directory belongs to backend {state.get('backend')!r}, not {backend!r}: {outdir}"
)
if state.get("caselist_sha256") != fingerprint:
raise RunnerError(f"output directory belongs to a different caselist: {outdir}")
stored_invocation_identity = state.get("invocation_identity")
if (
stored_invocation_identity is not None
or invocation_identity is not None
) and stored_invocation_identity != invocation_identity:
raise RunnerError(f"output directory belongs to a different CTS invocation: {outdir}")
return
legacy_artifacts = recovery_artifacts(outdir)
if legacy_artifacts and not adopt_legacy:
examples = ", ".join(path.name for path in legacy_artifacts[:3])
raise RunnerError(
"output directory contains CTS recovery artifacts but no run_state.json; "
f"refusing to adopt unverified legacy results ({examples}). Re-run with "
"--adopt-legacy only after verifying the backend, caselist, and invocation."
)
atomic_write_json(
path,
{
"version": 1,
"backend": backend,
"case_count": len(cases),
"caselist_sha256": fingerprint,
"invocation_identity": invocation_identity,
"adopted_legacy": bool(legacy_artifacts),
"created_utc": utc_now(),
},
)
def persist_sidecars(
outdir: Path,
ordered_cases: Sequence[str],
crashed: set[str],
hung: set[str],
remaining: Sequence[str],
) -> None:
write_case_file(outdir / "crashed.txt", (case for case in ordered_cases if case in crashed))
write_case_file(outdir / "hung.txt", (case for case in ordered_cases if case in hung))
write_case_file(outdir / "unrun.txt", remaining)
write_case_file(outdir / "remaining.txt", remaining)
def parse_environment(values: Sequence[str]) -> dict[str, str]:
result: dict[str, str] = {}
for value in values:
if "=" not in value:
raise RunnerError(f"--env expects NAME=VALUE, got {value!r}")
name, contents = value.split("=", 1)
if not name or "\x00" in name or "=" in name:
raise RunnerError(f"invalid environment variable name in {value!r}")
result[name] = contents
return result
def validate_deqp_args(values: Sequence[str]) -> None:
for value in values:
option = value.split("=", 1)[0].lower()
if option in CONTROLLED_DEQP_OPTIONS:
raise RunnerError(f"{option} is controlled by this runner and cannot be supplied via --deqp-arg")
def resolve_paths(
exe_value: str,
workdir_value: Optional[str],
caselist_value: str,
outdir_value: str,
) -> tuple[Path, Path, Path, Path]:
launch_dir = Path.cwd()
requested_exe = Path(exe_value).expanduser()
if workdir_value:
workdir = Path(workdir_value).expanduser().resolve()
elif requested_exe.is_absolute():
workdir = requested_exe.resolve().parent
else:
workdir = launch_dir
if requested_exe.is_absolute():
exe = requested_exe.resolve()
else:
in_workdir = (workdir / requested_exe).resolve()
in_launch_dir = (launch_dir / requested_exe).resolve()
exe = in_workdir if in_workdir.is_file() else in_launch_dir
caselist = Path(caselist_value).expanduser().resolve()
outdir = Path(outdir_value).expanduser().resolve()
if not exe.is_file():
raise RunnerError(f"glcts executable does not exist: {exe}")
if not workdir.is_dir():
raise RunnerError(f"working directory does not exist: {workdir}")
if not caselist.is_file():
raise RunnerError(f"caselist does not exist: {caselist}")
return exe, workdir, caselist, outdir
def qpa_signature(path: Path) -> Optional[tuple[int, int]]:
try:
stat = path.stat()
except FileNotFoundError:
return None
except OSError:
# A transient sharing violation must not kill a healthy process. The
# next poll will retry and the idle clock retains its previous value.
return None
return stat.st_size, stat.st_mtime_ns
def kill_process_tree(process: subprocess.Popen[bytes]) -> None:
"""Force-stop the process and descendants, with a parent-only fallback."""
if process.poll() is not None:
return
if os.name == "nt":
# /T is essential: CTS/platform helpers can outlive the top-level
# process, retain the QPA/DLL, and poison the next continuation round.
taskkill = Path(os.environ.get("SystemRoot", r"C:\Windows")) / "System32" / "taskkill.exe"
command = [str(taskkill), "/PID", str(process.pid), "/T", "/F"]
try:
subprocess.run(
command,
stdin=subprocess.DEVNULL,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
timeout=20,
check=False,
creationflags=getattr(subprocess, "CREATE_NO_WINDOW", 0),
)
except (OSError, subprocess.TimeoutExpired):
pass
else:
try:
os.killpg(process.pid, signal.SIGKILL)
except (ProcessLookupError, PermissionError, OSError):
pass
try:
process.wait(timeout=10)
return
except subprocess.TimeoutExpired:
pass
try:
process.kill()
except OSError:
pass
try:
process.wait(timeout=10)
except subprocess.TimeoutExpired:
pass
def run_process(
command: Sequence[str],
workdir: Path,
environment: dict[str, str],
qpa_path: Path,
stdout_path: Path,
stderr_path: Path,
idle_timeout: float,
max_round_seconds: float,
poll_seconds: float,
) -> ProcessOutcome:
"""Run one CTS chunk, killing its tree only after QPA progress stalls."""
started = time.monotonic()
with stdout_path.open("wb") as stdout_handle, stderr_path.open("wb") as stderr_handle:
popen_options: dict[str, object] = {
"cwd": str(workdir),
"env": environment,
"stdin": subprocess.DEVNULL,
"stdout": stdout_handle,
"stderr": stderr_handle,
}
if os.name == "nt":
popen_options["creationflags"] = getattr(subprocess, "CREATE_NEW_PROCESS_GROUP", 0)
else:
popen_options["start_new_session"] = True
try:
process = subprocess.Popen(list(command), **popen_options) # type: ignore[arg-type]
except OSError as exc:
message = f"failed to launch {command[0]}: {exc}\n"
stderr_handle.write(message.encode("utf-8", errors="replace"))
stderr_handle.flush()
return ProcessOutcome(None, time.monotonic() - started, launch_error=str(exc))
last_signature = qpa_signature(qpa_path)
last_progress = time.monotonic()
timed_out = False
timeout_reason: Optional[str] = None
interrupted = False
try:
while True:
try:
returncode = process.wait(timeout=poll_seconds)
break
except subprocess.TimeoutExpired:
pass
now = time.monotonic()
signature = qpa_signature(qpa_path)
if signature is not None and signature != last_signature:
last_signature = signature
last_progress = now
if idle_timeout > 0 and now - last_progress >= idle_timeout:
timed_out = True
timeout_reason = "qpa-idle"
kill_process_tree(process)
returncode = process.poll()
break
if max_round_seconds > 0 and now - started >= max_round_seconds:
timed_out = True
timeout_reason = "max-round"
kill_process_tree(process)
returncode = process.poll()
break
except KeyboardInterrupt:
interrupted = True
kill_process_tree(process)
returncode = process.poll()
return ProcessOutcome(
returncode=returncode,
duration_seconds=time.monotonic() - started,
timed_out=timed_out,
timeout_reason=timeout_reason,
interrupted=interrupted,
)
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Run Windows glcts against MobileGL, resuming across crashes and GPU hangs."
)
parser.add_argument("--exe", required=True, help="path to glcts.exe")
parser.add_argument(
"--workdir",
help="glcts working directory (default: executable directory, or current directory for a relative exe)",
)
parser.add_argument("--caselist", required=True, help="mustpass/caselist text file")
parser.add_argument("--outdir", required=True, help="persistent result directory")
parser.add_argument("--backend", required=True, choices=("DirectGLES", "DirectVulkan"))
parser.add_argument(
"--run-identity",
help="controller fingerprint for executable, data, arguments, and environment",
)
parser.add_argument(
"--adopt-legacy",
action="store_true",
help=(
"adopt existing chunk/sidecar results which predate run_state.json; "
"disabled by default because their provenance cannot be verified"
),
)
parser.add_argument(
"--env",
action="append",
default=[],
metavar="NAME=VALUE",
help="extra child environment variable (repeatable)",
)
parser.add_argument(
"--deqp-arg",
action="append",
default=[],
metavar="ARG",
help="extra glcts argument; repeat and use --deqp-arg=--option=value for leading dashes",
)
parser.add_argument(
"--idle-timeout",
type=float,
default=300.0,
metavar="SECONDS",
help="kill a chunk after this many seconds with no QPA size/mtime change (0 disables; default: 300)",
)
parser.add_argument(
"--max-round-seconds",
type=float,
default=0.0,
metavar="SECONDS",
help="optional total wall limit for one invocation (0 disables; default: 0)",
)
parser.add_argument("--poll-seconds", type=float, default=1.0, help=argparse.SUPPRESS)
parser.add_argument(
"--max-rounds",
type=int,
default=10000,
help="maximum glcts invocations in this runner process (default: 10000)",
)
parser.add_argument(
"--max-empty-streak",
type=int,
default=3,
help="abort after this many invocations record no case at all; no case is blamed (default: 3)",
)
return parser
def execute(args: argparse.Namespace) -> int:
if args.idle_timeout < 0 or args.max_round_seconds < 0:
raise RunnerError("timeout values must be non-negative")
if args.poll_seconds <= 0:
raise RunnerError("--poll-seconds must be greater than zero")
if args.max_rounds <= 0 or args.max_empty_streak <= 0:
raise RunnerError("--max-rounds and --max-empty-streak must be greater than zero")
validate_deqp_args(args.deqp_arg)
extra_environment = parse_environment(args.env)
exe, workdir, caselist_path, outdir = resolve_paths(
args.exe, args.workdir, args.caselist, args.outdir
)
outdir.mkdir(parents=True, exist_ok=True)
cases = read_caselist(caselist_path)
expected = set(cases)
check_run_identity(
outdir,
args.backend,
cases,
args.run_identity,
adopt_legacy=args.adopt_legacy,
)
recorded, crashed, hung = recover_results(outdir, expected)
accounted = recorded | crashed | hung
remaining = [case for case in cases if case not in accounted]
persist_sidecars(outdir, cases, crashed, hung, remaining)
existing_qpas = len(result_qpa_files(outdir))
print(
f"[run_cts_windows] {args.backend}: expected {len(cases)}, recovered {len(accounted)} "
f"({existing_qpas} QPA chunk(s), {len(crashed)} crash, {len(hung)} hang)"
)
if not remaining:
print(f"[run_cts_windows] complete: all {len(cases)} expected cases are accounted")
return 0
environment = os.environ.copy()
environment.update(extra_environment)
# --backend is authoritative even if the inherited or extra environment
# already contains a different value.
environment["MOBILEGL_BACKEND_TYPE"] = args.backend
common_arguments = [
"--deqp-terminate-on-device-lost=disable",
"--deqp-log-images=disable",
"--deqp-log-shader-sources=disable",
]
chunk_number = next_chunk_number(outdir)
rounds = 0
empty_streak = 0
interrupted = False
fatal_launch_error = False
started_all = time.monotonic()
while remaining and rounds < args.max_rounds:
prefix = f"chunk{chunk_number:04d}"
remaining_path = outdir / "remaining.txt"
qpa_path = outdir / f"{prefix}.qpa"
stdout_path = outdir / f"{prefix}.stdout.log"
stderr_path = outdir / f"{prefix}.stderr.log"
meta_path = outdir / f"{prefix}.meta.json"
# The number allocator considers every chunk artifact, so these should
# be new. Refuse to truncate evidence if a foreign file races us.
for artifact in (qpa_path, stdout_path, stderr_path, meta_path):
if artifact.exists():
raise RunnerError(f"refusing to overwrite existing chunk artifact: {artifact}")
write_case_file(remaining_path, remaining)
command = [
str(exe),
f"--deqp-caselist-file={remaining_path}",
f"--deqp-log-filename={qpa_path}",
*common_arguments,
*args.deqp_arg,
]
print(
f"[run_cts_windows] {prefix}: launching {len(remaining)} remaining case(s); "
f"idle timeout {args.idle_timeout:g}s"
)
chunk_started_utc = utc_now()
outcome = run_process(
command,
workdir,
environment,
qpa_path,
stdout_path,
stderr_path,
args.idle_timeout,
args.max_round_seconds,
args.poll_seconds,
)
progress = scan_qpa(qpa_path)
before = set(accounted)
for case in progress.recorded:
if case in expected:
recorded.add(case)
accounted.add(case)
classification: Optional[str] = None
classified_case: Optional[str] = None
in_flight = progress.in_flight if progress.in_flight in expected else None
if not outcome.interrupted and in_flight is not None and in_flight not in accounted:
classified_case = in_flight
if outcome.timed_out:
classification = "DeviceHang"
hung.add(in_flight)
crashed.discard(in_flight)
else:
classification = "Crash"
crashed.add(in_flight)
accounted.add(in_flight)
new_accounted = len(accounted - before)
if new_accounted:
empty_streak = 0
elif progress.begin_count == 0:
# No #begin marker means there is no evidence that the first
# remaining case was reached. Retry the identical caselist, then
# abort rather than manufacturing a string of false Crash results.
empty_streak += 1
else:
# A log containing only already-accounted cases is also no forward
# progress, but it is a different failure mode. Bound it with the
# same guard while retaining the QPA evidence.
empty_streak += 1
remaining = [case for case in cases if case not in accounted]
metadata = {
"version": 1,
"chunk": chunk_number,
"started_utc": chunk_started_utc,
"finished_utc": utc_now(),
"duration_seconds": round(outcome.duration_seconds, 3),
"returncode": outcome.returncode,
"timed_out": outcome.timed_out,
"timeout_reason": outcome.timeout_reason,
"interrupted": outcome.interrupted,
"launch_error": outcome.launch_error,
"qpa_begin_count": progress.begin_count,
"qpa_recorded_count": len(progress.recorded),
"in_flight": progress.in_flight,
"classification": classification,
"classified_case": classified_case,
"new_accounted": new_accounted,
"remaining": len(remaining),
}
# Metadata is committed first. If the runner itself dies between this
# write and the sidecars, recovery can reconstruct the classification.
atomic_write_json(meta_path, metadata)
persist_sidecars(outdir, cases, crashed, hung, remaining)
rounds += 1
elapsed_minutes = (time.monotonic() - started_all) / 60.0
detail = ""
if classification:
detail = f", {classification}={classified_case}"
if outcome.timed_out:
detail += f", timeout={outcome.timeout_reason}"
print(
f"[run_cts_windows] {prefix}: +{new_accounted}, accounted "
f"{len(accounted)}/{len(cases)}, remaining {len(remaining)}{detail} "
f"({elapsed_minutes:.1f} min)"
)
chunk_number += 1
if outcome.interrupted:
interrupted = True
print("[run_cts_windows] interrupted; process tree stopped and state preserved", file=sys.stderr)
break
if outcome.launch_error:
fatal_launch_error = True
print(
f"[run_cts_windows] launch failed; see {stderr_path.name}: {outcome.launch_error}",
file=sys.stderr,
)
break
if empty_streak >= args.max_empty_streak:
print(
f"[run_cts_windows] aborting after {empty_streak} consecutive chunks made no "
"case progress; no unobserved case was labelled Crash/Hang",
file=sys.stderr,
)
break
# Recompute from the persisted evidence so the final completeness claim is
# subject to the exact same recovery path as a later invocation.
final_recorded, final_crashed, final_hung = recover_results(outdir, expected)
final_accounted = final_recorded | final_crashed | final_hung
final_remaining = [case for case in cases if case not in final_accounted]
persist_sidecars(outdir, cases, final_crashed, final_hung, final_remaining)
if not final_remaining and final_accounted == expected:
print(
f"[run_cts_windows] complete: all {len(cases)} expected cases are accounted "
f"({len(final_crashed)} crash, {len(final_hung)} hang, {rounds} new invocation(s))"
)
return 0
print(
f"[run_cts_windows] INCOMPLETE: {len(final_accounted)}/{len(cases)} accounted; "
f"{len(final_remaining)} listed in {outdir / 'unrun.txt'}",
file=sys.stderr,
)
if interrupted:
return 130
if fatal_launch_error:
return 3
return 4
def main(argv: Optional[Sequence[str]] = None) -> int:
parser = build_parser()
args = parser.parse_args(argv)
try:
return execute(args)
except RunnerError as exc:
print(f"[run_cts_windows] ERROR: {exc}", file=sys.stderr)
return 2
except OSError as exc:
print(f"[run_cts_windows] ERROR: filesystem/process operation failed: {exc}", file=sys.stderr)
return 2
if __name__ == "__main__":
sys.exit(main())
+53
View File
@@ -0,0 +1,53 @@
#!/usr/bin/env python
"""Copy the MobileGL dEQP platform port into a VK-GL-CTS checkout.
The port is version-controlled here, in the MobileGL repo, so it survives a
throwaway CTS clone. This drops it into the places VK-GL-CTS expects:
framework/platform/mobilegl/ <- platform sources
targets/mobilegl/mobilegl.cmake <- target definition (-DDEQP_TARGET=mobilegl)
Usage:
python sync_to_cts.py <path-to-VK-GL-CTS>
"""
import os
import shutil
import sys
HERE = os.path.dirname(os.path.abspath(__file__))
CTS_TOOLS = os.path.dirname(HERE)
COPIES = [
(os.path.join(CTS_TOOLS, "platform"), "framework/platform/mobilegl", None),
(os.path.join(CTS_TOOLS, "targets"), "targets/mobilegl", ["mobilegl.cmake", "ndk-modern.cmake"]),
]
def main():
if len(sys.argv) != 2:
print(__doc__)
return 2
cts = sys.argv[1]
if not os.path.isfile(os.path.join(cts, "CMakeLists.txt")):
print(f"error: {cts} does not look like a VK-GL-CTS checkout", file=sys.stderr)
return 1
for src, reldst, only in COPIES:
dst = os.path.join(cts, reldst)
os.makedirs(dst, exist_ok=True)
for name in sorted(os.listdir(src)):
if only is not None and name not in only:
continue
s = os.path.join(src, name)
if not os.path.isfile(s):
continue
shutil.copy2(s, os.path.join(dst, name))
print(f" {reldst}/{name}")
print("\nsynced. configure with -DDEQP_TARGET=mobilegl")
return 0
if __name__ == "__main__":
sys.exit(main())
+251
View File
@@ -0,0 +1,251 @@
import contextlib
import io
import json
import tempfile
import unittest
from pathlib import Path
try:
from . import cts_matrix_report as report
except ImportError: # Allows `python test_cts_matrix_report.py`.
import cts_matrix_report as report
def qpa_case(case, status):
return (
f"#beginTestCaseResult {case}\n"
f'<Result StatusCode="{status}"/>\n'
"#endTestCaseResult\n"
)
class MatrixReportTests(unittest.TestCase):
def test_utf8_bom_caselist_matches_runner_semantics(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_bytes(b"\xef\xbb\xbfcase.a\n")
results = root / "results"
results.mkdir()
(results / "run.qpa").write_text(
qpa_case("case.a", "Pass"), encoding="utf-8"
)
item = report.build_version_report("gl30", str(caselist), [str(results)])
self.assertEqual(1, item["expected"])
self.assertEqual({"case.a": "Pass"}, item["cases"]["results"])
self.assertEqual("OK", item["validation"]["state"])
def test_incomplete_qpa_is_unrun_not_a_completed_result(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_text("case.a\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "run.qpa").write_text(
"#beginTestCaseResult case.a\n#endTestCaseResult\n",
encoding="utf-8",
)
(results / "unrun.txt").write_text("case.a\n", encoding="utf-8")
item = report.build_version_report("gl46", str(caselist), [str(results)])
self.assertEqual(0, item["result"])
self.assertEqual(1, item["unrun"])
self.assertEqual(["case.a"], item["cases"]["incomplete_results"])
self.assertEqual("INCOMPLETE", item["validation"]["state"])
def test_incomplete_qpa_is_upgraded_by_crash_sidecar(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_text("case.a\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "run.qpa").write_text(
"#beginTestCaseResult case.a\n", encoding="utf-8"
)
(results / "crashed.txt").write_text("case.a\n", encoding="utf-8")
item = report.build_version_report("gl46", str(caselist), [str(results)])
self.assertEqual("Crash", item["cases"]["results"]["case.a"])
self.assertEqual([], item["cases"]["incomplete_results"])
self.assertEqual("OK", item["validation"]["state"])
def test_chunk_numbers_above_four_digits_use_numeric_order(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_text("case.a\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "alpha.qpa").write_text("# no results\n", encoding="utf-8")
(results / "chunk9999.qpa").write_text(
qpa_case("case.a", "Fail"), encoding="utf-8"
)
(results / "chunk10000.qpa").write_text(
qpa_case("case.a", "Pass"), encoding="utf-8"
)
(results / "zeta.qpa").write_text("# no results\n", encoding="utf-8")
item = report.build_version_report(
"gl46", str(caselist), [str(results)]
)
self.assertEqual(
["alpha.qpa", "chunk9999.qpa", "chunk10000.qpa", "zeta.qpa"],
[Path(path).name for path in item["inputs"]["qpa_files"]],
)
self.assertEqual("Pass", item["cases"]["results"]["case.a"])
self.assertEqual(1, item["duplicate"])
def test_qpa_sidecars_duplicates_and_expected_denominator(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "gl30.txt"
caselist.write_text("\n".join("abcdefg") + "\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "chunk0000.qpa").write_text(
qpa_case("a", "Fail") + "#beginTestCaseResult e\n",
encoding="utf-8",
)
(results / "chunk0001.qpa").write_text(
qpa_case("a", "Pass")
+ qpa_case("b", "NotSupported")
+ qpa_case("c", "QualityWarning")
+ qpa_case("d", "Fail"),
encoding="utf-8",
)
(results / "crashed.txt").write_text("e\n", encoding="utf-8")
(results / "hung.txt").write_text("f\n", encoding="utf-8")
(results / "unrun.txt").write_text("g\n", encoding="utf-8")
item = report.build_version_report(
"gl30", str(caselist), [str(results)]
)
self.assertEqual(item["expected"], 7)
self.assertEqual(item["result"], 6)
self.assertEqual(item["pass"], 1)
self.assertEqual(item["accepted"], 3)
self.assertEqual(item["crash"], 1)
self.assertEqual(item["hang"], 1)
self.assertEqual(item["unrun"], 1)
self.assertEqual(item["duplicate"], 1)
self.assertEqual(item["cases"]["results"]["a"], "Pass")
self.assertEqual(item["cases"]["results"]["e"], "Crash")
self.assertEqual(item["cases"]["results"]["f"], "DeviceHang")
self.assertAlmostEqual(item["strict_pass_rate"], 1 / 7)
self.assertAlmostEqual(item["conformance_accepted_rate"], 3 / 7)
self.assertAlmostEqual(
item["rates"]["measured_only_conformance_accepted"], 3 / 6
)
self.assertEqual(item["validation"]["state"], "INCOMPLETE")
self.assertEqual(item["validation"]["errors"], [])
self.assertTrue(
item["validation"]["invariant_expected_equals_result_plus_unrun"]
)
def test_missing_result_is_inferred_and_rejected_when_not_declared(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_text("a\nb\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "run.qpa").write_text(qpa_case("a", "Pass"), encoding="utf-8")
item = report.build_version_report(
"gl31", str(caselist), [str(results)]
)
self.assertEqual(item["unrun"], 1)
self.assertEqual(item["cases"]["unrun"], ["b"])
self.assertEqual(item["validation"]["state"], "ERROR")
self.assertEqual(item["validation"]["undeclared_unrun"], ["b"])
self.assertIn("not declared", item["validation"]["errors"][0])
self.assertEqual(item["strict_pass_rate"], 0.5)
def test_cli_emits_markdown_json_and_weighted_overall(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
statuses = {
"gl30": "Pass",
"gl31": "Fail",
"gl32": "NotSupported",
"gl33": None,
}
argv = []
for version, status in statuses.items():
caselist = root / f"{version}.txt"
caselist.write_text(f"{version}.case\n", encoding="utf-8")
result_dir = root / f"{version}-results"
result_dir.mkdir()
qpa = result_dir / "run.qpa"
qpa.write_text(
qpa_case(f"{version}.case", status) if status else "# empty run\n",
encoding="utf-8",
)
if status is None:
(result_dir / "unrun.txt").write_text(
f"{version}.case\n", encoding="utf-8"
)
argv.extend(
[
f"--{version}-caselist",
str(caselist),
f"--{version}-results",
str(result_dir),
]
)
json_path = root / "matrix.json"
argv.extend(["--json", str(json_path)])
stdout = io.StringIO()
with contextlib.redirect_stdout(stdout):
rc = report.main(argv)
self.assertEqual(rc, 1) # GL33 is explicitly incomplete.
markdown = stdout.getvalue()
self.assertIn("| Suite | Expected | Result", markdown)
self.assertIn("| **Overall (weighted)**", markdown)
payload = json.loads(json_path.read_text(encoding="utf-8"))
overall = payload["overall"]
self.assertEqual(overall["expected"], 4)
self.assertEqual(overall["result"], 3)
self.assertEqual(overall["pass"], 1)
self.assertEqual(overall["accepted"], 2)
self.assertEqual(overall["unrun"], 1)
self.assertEqual(overall["strict_pass_rate"], 0.25)
self.assertEqual(overall["conformance_accepted_rate"], 0.5)
self.assertEqual(overall["aggregation"], "weighted_by_expected_cases")
self.assertEqual(overall["validation"]["state"], "INCOMPLETE")
def test_duplicate_caselist_and_unexpected_result_are_validation_errors(self):
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
caselist = root / "cases.txt"
caselist.write_text("a\na\n", encoding="utf-8")
results = root / "results"
results.mkdir()
(results / "run.qpa").write_text(
qpa_case("a", "Pass") + qpa_case("outside", "Pass"),
encoding="utf-8",
)
item = report.build_version_report(
"gl32", str(caselist), [str(results)]
)
self.assertEqual(item["cases"]["duplicate_caselist_entries"], {"a": 2})
self.assertEqual(item["cases"]["unexpected_results"], {"outside": "Pass"})
self.assertEqual(item["validation"]["state"], "ERROR")
self.assertEqual(len(item["validation"]["errors"]), 2)
if __name__ == "__main__":
unittest.main()
+342
View File
@@ -0,0 +1,342 @@
import contextlib
import io
import json
import tempfile
import unittest
from pathlib import Path
try:
from . import cts_multi_report as report
except ImportError: # Allows `python test_cts_multi_report.py`.
import cts_multi_report as report
def qpa_case(case: str, status: str) -> str:
return (
f"#beginTestCaseResult {case}\n"
f'<Result StatusCode="{status}"/>\n'
"#endTestCaseResult\n"
)
def write_run_state(
caselist: Path,
result_dir: Path,
backend: str,
invocation_identity=None,
) -> None:
fingerprint, case_count = report._caselist_fingerprint(str(caselist))
(result_dir / "run_state.json").write_text(
json.dumps(
{
"version": 1,
"backend": backend,
"case_count": case_count,
"caselist_sha256": fingerprint,
"invocation_identity": invocation_identity,
}
),
encoding="utf-8",
)
def make_inputs(
root: Path,
name: str,
cases: list[str],
qpa: str,
backend: str = "DirectGLES",
):
caselist = root / f"{name}.txt"
caselist.write_text("\n".join(cases) + "\n", encoding="utf-8")
result_dir = root / f"{name}-results"
result_dir.mkdir()
(result_dir / "chunk0000.qpa").write_text(qpa, encoding="utf-8")
write_run_state(caselist, result_dir, backend)
return caselist, result_dir
class MultiReportTests(unittest.TestCase):
def test_backend_aggregate_is_weighted_by_expected_cases(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
small_cases, small_results = make_inputs(
root, "small", ["small.pass"], qpa_case("small.pass", "Pass")
)
large_cases, large_results = make_inputs(
root,
"large",
["large.pass", "large.crash", "large.hang"],
qpa_case("large.pass", "Pass"),
)
(large_results / "crashed.txt").write_text(
"large.crash\n", encoding="utf-8"
)
(large_results / "hung.txt").write_text(
"large.hang\n", encoding="utf-8"
)
payload = report.build_report(
[
report.SuiteSpec(
"DirectGLES", "small", str(small_cases), str(small_results)
),
report.SuiteSpec(
"DirectGLES", "large", str(large_cases), str(large_results)
),
]
)
aggregate = payload["backends"]["DirectGLES"]
self.assertEqual(4, aggregate["expected"])
self.assertEqual(4, aggregate["result"])
self.assertEqual(2, aggregate["pass"])
self.assertEqual(2, aggregate["accepted"])
self.assertEqual(1, aggregate["crash"])
self.assertEqual(1, aggregate["hang"])
self.assertEqual(0, aggregate["unrun"])
# (1 accepted + 1 accepted) / (1 expected + 3 expected), not
# the unweighted mean of 100% and 33.3%.
self.assertEqual(0.5, aggregate["conformance_accepted_rate"])
self.assertEqual("weighted_by_expected_cases", aggregate["aggregation"])
self.assertEqual("OK", aggregate["validation"]["state"])
def test_declared_unrun_is_incomplete_and_cli_returns_nonzero(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root, "missing", ["case.a", "case.b"], qpa_case("case.a", "Pass")
)
(result_dir / "unrun.txt").write_text("case.b\n", encoding="utf-8")
markdown_path = root / "report.md"
json_path = root / "report.json"
argv = [
"--suite",
"DirectGLES",
"gl30",
str(caselist),
str(result_dir),
"--markdown",
str(markdown_path),
"--json",
str(json_path),
]
with contextlib.redirect_stdout(io.StringIO()):
returncode = report.main(argv)
self.assertEqual(1, returncode)
self.assertTrue(markdown_path.is_file())
payload = json.loads(json_path.read_text(encoding="utf-8"))
suite = payload["suites"][0]
self.assertEqual(2, suite["expected"])
self.assertEqual(1, suite["result"])
self.assertEqual(1, suite["unrun"])
self.assertEqual("INCOMPLETE", suite["validation"]["state"])
self.assertEqual("INCOMPLETE", payload["overall"]["validation"]["state"])
self.assertEqual(0.5, payload["overall"]["conformance_accepted_rate"])
def test_dual_backend_cli_outputs_markdown_and_json(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist = root / "gl30.txt"
caselist.write_text("gl30.case\n", encoding="utf-8")
gles = root / "gles"
vulkan = root / "vulkan"
gles.mkdir()
vulkan.mkdir()
(gles / "run.qpa").write_text(
qpa_case("gl30.case", "Pass"), encoding="utf-8"
)
(vulkan / "run.qpa").write_text(
qpa_case("gl30.case", "Fail"), encoding="utf-8"
)
write_run_state(caselist, gles, "DirectGLES")
write_run_state(caselist, vulkan, "DirectVulkan")
markdown_path = root / "dual.md"
json_path = root / "dual.json"
argv = [
f"--suite=DirectGLES,gl30,{caselist},{gles}",
"--suite",
"DirectVulkan",
"gl30",
str(caselist),
str(vulkan),
"--markdown",
str(markdown_path),
"--json",
str(json_path),
]
stdout = io.StringIO()
with contextlib.redirect_stdout(stdout):
returncode = report.main(argv)
self.assertEqual(0, returncode)
payload = json.loads(json_path.read_text(encoding="utf-8"))
self.assertEqual({"DirectGLES", "DirectVulkan"}, set(payload["backends"]))
self.assertEqual(1, payload["backends"]["DirectGLES"]["accepted"])
self.assertEqual(0, payload["backends"]["DirectVulkan"]["accepted"])
self.assertEqual(2, payload["overall"]["expected"])
self.assertEqual(1, payload["overall"]["accepted"])
self.assertEqual(0.5, payload["overall"]["conformance_accepted_rate"])
markdown = markdown_path.read_text(encoding="utf-8")
self.assertIn("DirectGLES weighted subtotal", markdown)
self.assertIn("DirectVulkan weighted subtotal", markdown)
self.assertIn("Overall weighted", markdown)
self.assertIn("Markdown:", stdout.getvalue())
def test_duplicate_qpa_result_uses_last_observation(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root,
"duplicate",
["case.a"],
qpa_case("case.a", "Fail"),
backend="DirectVulkan",
)
(result_dir / "chunk0001.qpa").write_text(
qpa_case("case.a", "Pass"), encoding="utf-8"
)
payload = report.build_report(
[
report.SuiteSpec(
"DirectVulkan", "gl33", str(caselist), str(result_dir)
)
]
)
suite = payload["suites"][0]
self.assertEqual("Pass", suite["cases"]["results"]["case.a"])
self.assertEqual(1, suite["duplicate"])
self.assertEqual(1, payload["overall"]["duplicate"])
self.assertEqual(1.0, payload["overall"]["strict_pass_rate"])
self.assertEqual("OK", suite["validation"]["state"])
self.assertIn("last result wins", suite["validation"]["warnings"][0])
def test_backend_provenance_mismatch_is_rejected(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root,
"provenance",
["case.a"],
qpa_case("case.a", "Pass"),
backend="DirectVulkan",
)
with self.assertRaises(report.MultiReportInputError):
report.build_report(
[report.SuiteSpec("DirectGLES", "gl30", str(caselist), str(result_dir))]
)
def test_missing_provenance_requires_explicit_legacy_opt_in(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root, "legacy", ["case.a"], qpa_case("case.a", "Pass")
)
(result_dir / "run_state.json").unlink()
spec = report.SuiteSpec("DirectGLES", "gl30", str(caselist), str(result_dir))
with self.assertRaises(report.MultiReportInputError):
report.build_report([spec])
payload = report.build_report([spec], require_run_state=False)
self.assertEqual("UNVERIFIED", payload["suites"][0]["provenance"]["state"])
def test_expected_run_identity_accepts_match_and_rejects_mismatch(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root, "identity", ["case.a"], qpa_case("case.a", "Pass")
)
write_run_state(
caselist, result_dir, "DirectGLES", invocation_identity="identity-a"
)
spec = report.SuiteSpec(
"DirectGLES", "gl30", str(caselist), str(result_dir)
)
payload = report.build_report(
[spec], expected_run_identity="identity-a"
)
self.assertEqual(
"identity-a",
payload["suites"][0]["provenance"]["invocation_identity"],
)
with self.assertRaises(report.MultiReportInputError):
report.build_report([spec], expected_run_identity="identity-b")
def test_expected_identity_rejects_legacy_state_and_missing_state(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root, "legacy-identity", ["case.a"], qpa_case("case.a", "Pass")
)
spec = report.SuiteSpec(
"DirectGLES", "gl30", str(caselist), str(result_dir)
)
with self.assertRaises(report.MultiReportInputError):
report.build_report([spec], expected_run_identity="identity-a")
(result_dir / "run_state.json").unlink()
with self.assertRaises(report.MultiReportInputError):
report.build_report(
[spec],
require_run_state=False,
expected_run_identity="identity-a",
)
def test_duplicate_physical_result_directory_is_rejected(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist, result_dir = make_inputs(
root, "duplicate-dir", ["case.a"], qpa_case("case.a", "Pass")
)
with self.assertRaises(report.MultiReportInputError):
report.build_report(
[
report.SuiteSpec(
"DirectGLES", "gl30", str(caselist), str(result_dir)
),
report.SuiteSpec(
"DirectGLES",
"gl31",
str(caselist),
str(result_dir / "."),
),
]
)
def test_ancestor_and_descendant_result_directories_are_rejected(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist = root / "cases.txt"
caselist.write_text("case.a\n", encoding="utf-8")
parent = root / "results"
child = parent / "nested"
child.mkdir(parents=True)
(parent / "chunk0000.qpa").write_text(
qpa_case("case.a", "Pass"), encoding="utf-8"
)
(child / "chunk0000.qpa").write_text(
qpa_case("case.a", "Fail"), encoding="utf-8"
)
write_run_state(caselist, parent, "DirectGLES")
write_run_state(caselist, child, "DirectVulkan")
with self.assertRaises(report.MultiReportInputError):
report.build_report(
[
report.SuiteSpec(
"DirectGLES", "gl30", str(caselist), str(parent)
),
report.SuiteSpec(
"DirectVulkan", "gl30", str(caselist), str(child)
),
]
)
if __name__ == "__main__":
unittest.main()
+401
View File
@@ -0,0 +1,401 @@
import sys
import tempfile
import time
import unittest
from pathlib import Path
from unittest import mock
import run_cts_windows as runner
def qpa_closed(case: str, status: str = "Pass") -> str:
return (
f"#beginTestCaseResult {case}\n"
f'<Result StatusCode="{status}">ok</Result>\n'
"#endTestCaseResult\n"
)
def command_path(command, option):
prefix = option + "="
return Path(next(value[len(prefix) :] for value in command if value.startswith(prefix)))
class AtomicWriteTests(unittest.TestCase):
def test_access_denied_retries_then_replace_succeeds(self):
with tempfile.TemporaryDirectory() as temporary:
target = Path(temporary) / "remaining.txt"
real_replace = runner.os.replace
attempts = 0
def flaky_replace(source, destination):
nonlocal attempts
attempts += 1
if attempts == 1:
raise PermissionError(13, "temporarily denied", str(destination))
if attempts == 2:
error = OSError("temporary WinError 5")
error.winerror = 5
raise error
real_replace(source, destination)
with mock.patch.object(runner.os, "replace", side_effect=flaky_replace), mock.patch.object(
runner.time, "sleep"
) as sleep:
runner.atomic_write_text(target, "case.a\n")
self.assertEqual(3, attempts)
self.assertEqual("case.a\n", target.read_text(encoding="utf-8"))
self.assertEqual(2, sleep.call_count)
self.assertEqual(
[
mock.call(runner.ATOMIC_REPLACE_INITIAL_BACKOFF_SECONDS),
mock.call(runner.ATOMIC_REPLACE_INITIAL_BACKOFF_SECONDS * 2),
],
sleep.call_args_list,
)
self.assertEqual([], list(target.parent.glob(".remaining.txt.*.tmp")))
def test_permanent_access_denied_stops_after_bounded_attempts(self):
with tempfile.TemporaryDirectory() as temporary:
target = Path(temporary) / "remaining.txt"
def always_denied(_source, _destination):
error = OSError("persistent WinError 5")
error.winerror = 5
raise error
with mock.patch.object(
runner.os, "replace", side_effect=always_denied
) as replace, mock.patch.object(runner.time, "sleep") as sleep:
with self.assertRaises(OSError) as raised:
runner.atomic_write_text(target, "case.a\n")
self.assertEqual(5, raised.exception.winerror)
self.assertEqual(runner.ATOMIC_REPLACE_ATTEMPTS, replace.call_count)
self.assertEqual(runner.ATOMIC_REPLACE_ATTEMPTS - 1, sleep.call_count)
self.assertFalse(target.exists())
self.assertEqual([], list(target.parent.glob(".remaining.txt.*.tmp")))
def test_non_access_error_is_not_retried(self):
with tempfile.TemporaryDirectory() as temporary:
target = Path(temporary) / "remaining.txt"
error = OSError(28, "disk full")
with mock.patch.object(
runner.os, "replace", side_effect=error
) as replace, mock.patch.object(runner.time, "sleep") as sleep:
with self.assertRaises(OSError):
runner.atomic_write_text(target, "case.a\n")
self.assertEqual(1, replace.call_count)
sleep.assert_not_called()
class QpaParsingTests(unittest.TestCase):
def test_terminate_is_a_completed_result(self):
with tempfile.TemporaryDirectory() as temporary:
path = Path(temporary) / "chunk0000.qpa"
path.write_text(
"#beginTestCaseResult KHR-GL30.a\n"
"#terminateTestCaseResult Crash\n"
"#beginTestCaseResult KHR-GL30.b\n",
encoding="utf-8",
)
progress = runner.scan_qpa(path)
self.assertEqual(["KHR-GL30.a"], progress.recorded)
self.assertEqual("KHR-GL30.b", progress.in_flight)
self.assertEqual(2, progress.begin_count)
def test_end_without_result_is_not_accounted(self):
with tempfile.TemporaryDirectory() as temporary:
path = Path(temporary) / "chunk0000.qpa"
path.write_text(
"#beginTestCaseResult KHR-GL46.incomplete\n"
"#endTestCaseResult\n"
+ qpa_closed("KHR-GL46.complete"),
encoding="utf-8",
)
progress = runner.scan_qpa(path)
self.assertEqual(["KHR-GL46.complete"], progress.recorded)
self.assertIsNone(progress.in_flight)
def test_result_written_before_truncated_eof_is_recovered(self):
with tempfile.TemporaryDirectory() as temporary:
path = Path(temporary) / "chunk0000.qpa"
path.write_text(
"#beginTestCaseResult KHR-GL46.complete\n"
'<Result StatusCode="Pass">ok</Result>\n',
encoding="utf-8",
)
progress = runner.scan_qpa(path)
self.assertEqual(["KHR-GL46.complete"], progress.recorded)
self.assertIsNone(progress.in_flight)
class RunIdentityTests(unittest.TestCase):
def test_non_object_run_state_is_a_controlled_error(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
(outdir / "run_state.json").write_text("null\n", encoding="utf-8")
with self.assertRaises(runner.RunnerError):
runner.check_run_identity(outdir, "DirectVulkan", ["case.a"])
def test_controller_identity_prevents_mixed_invocations(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a"], invocation_identity="identity-a"
)
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a"], invocation_identity="identity-a"
)
with self.assertRaises(runner.RunnerError):
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a"], invocation_identity="identity-b"
)
def test_controller_identity_cannot_be_downgraded_by_omission(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a"], invocation_identity="identity-a"
)
with self.assertRaises(runner.RunnerError):
runner.check_run_identity(outdir, "DirectVulkan", ["case.a"])
def test_legacy_artifacts_require_explicit_adoption(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
(outdir / "chunk0000.qpa").write_text(
qpa_closed("case.a"), encoding="utf-8"
)
with self.assertRaises(runner.RunnerError):
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a"], invocation_identity="identity-a"
)
self.assertFalse((outdir / "run_state.json").exists())
runner.check_run_identity(
outdir,
"DirectVulkan",
["case.a"],
invocation_identity="identity-a",
adopt_legacy=True,
)
state = runner.json.loads(
(outdir / "run_state.json").read_text(encoding="utf-8")
)
self.assertTrue(state["adopted_legacy"])
def test_foreign_nested_qpa_and_skipped_sidecar_are_legacy_evidence(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
nested = outdir / "old"
nested.mkdir()
qpa = nested / "legacy.qpa"
qpa.write_text(qpa_closed("case.a"), encoding="utf-8")
skipped = outdir / "skipped.txt"
skipped.write_text("case.b\n", encoding="utf-8")
self.assertEqual(
{qpa, skipped}, set(runner.recovery_artifacts(outdir))
)
with self.assertRaises(runner.RunnerError):
runner.check_run_identity(
outdir, "DirectVulkan", ["case.a", "case.b"]
)
recorded, crashed, hung = runner.recover_results(
outdir, {"case.a", "case.b"}
)
self.assertEqual({"case.a"}, recorded)
self.assertEqual(set(), crashed)
self.assertEqual(set(), hung)
def test_non_object_or_non_string_meta_classification_is_ignored(self):
with tempfile.TemporaryDirectory() as temporary:
outdir = Path(temporary)
(outdir / "chunk0000.meta.json").write_text("null\n", encoding="utf-8")
(outdir / "chunk0001.meta.json").write_text("[]\n", encoding="utf-8")
(outdir / "chunk0002.meta.json").write_text(
'{"classified_case": [], "classification": "Crash"}\n', encoding="utf-8"
)
self.assertEqual(
(set(), set()), runner.load_meta_classifications(outdir, {"case.a"})
)
class RunnerRecoveryTests(unittest.TestCase):
def run_args(self, root: Path, caselist: Path, outdir: Path, *extra: str):
return [
"--exe",
sys.executable,
"--workdir",
str(root),
"--caselist",
str(caselist),
"--outdir",
str(outdir),
"--backend",
"DirectVulkan",
*extra,
]
def test_crash_tail_is_quarantined_and_next_chunk_resumes(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist = root / "cases.txt"
outdir = root / "results"
caselist.write_text("KHR-GL30.a\nKHR-GL30.b\nKHR-GL30.c\n", encoding="utf-8")
seen_remaining = []
def fake_run(command, workdir, environment, qpa_path, stdout_path, stderr_path, *timeouts):
del workdir, timeouts
seen_remaining.append(
command_path(command, "--deqp-caselist-file")
.read_text(encoding="utf-8")
.splitlines()
)
stdout_path.write_text("fake stdout\n", encoding="utf-8")
stderr_path.write_text("fake stderr\n", encoding="utf-8")
self.assertEqual("DirectVulkan", environment["MOBILEGL_BACKEND_TYPE"])
if len(seen_remaining) == 1:
qpa_path.write_text(
qpa_closed("KHR-GL30.a") + "#beginTestCaseResult KHR-GL30.b\n",
encoding="utf-8",
)
return runner.ProcessOutcome(0xC0000005, 0.1)
qpa_path.write_text(qpa_closed("KHR-GL30.c"), encoding="utf-8")
return runner.ProcessOutcome(0, 0.1)
with mock.patch.object(runner, "run_process", side_effect=fake_run):
result = runner.main(self.run_args(root, caselist, outdir))
self.assertEqual(0, result)
self.assertEqual(
[
["KHR-GL30.a", "KHR-GL30.b", "KHR-GL30.c"],
["KHR-GL30.c"],
],
seen_remaining,
)
self.assertEqual("KHR-GL30.b\n", (outdir / "crashed.txt").read_text(encoding="utf-8"))
self.assertEqual("", (outdir / "hung.txt").read_text(encoding="utf-8"))
self.assertEqual("", (outdir / "unrun.txt").read_text(encoding="utf-8"))
def test_existing_qpa_and_sidecar_are_recovered(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist = root / "cases.txt"
outdir = root / "results"
outdir.mkdir()
caselist.write_text("KHR-GL31.a\nKHR-GL31.b\nKHR-GL31.c\n", encoding="utf-8")
(outdir / "chunk0000.qpa").write_text(qpa_closed("KHR-GL31.a"), encoding="utf-8")
(outdir / "crashed.txt").write_text("KHR-GL31.b\n", encoding="utf-8")
seen_remaining = []
def fake_run(command, workdir, environment, qpa_path, stdout_path, stderr_path, *timeouts):
del workdir, environment, stdout_path, stderr_path, timeouts
seen_remaining.extend(
command_path(command, "--deqp-caselist-file")
.read_text(encoding="utf-8")
.splitlines()
)
qpa_path.write_text(qpa_closed("KHR-GL31.c"), encoding="utf-8")
return runner.ProcessOutcome(0, 0.1)
with mock.patch.object(runner, "run_process", side_effect=fake_run):
result = runner.main(
self.run_args(root, caselist, outdir, "--adopt-legacy")
)
self.assertEqual(0, result)
self.assertEqual(["KHR-GL31.c"], seen_remaining)
self.assertTrue((outdir / "chunk0001.qpa").is_file())
def test_repeated_no_output_aborts_without_false_case_blame(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
caselist = root / "cases.txt"
outdir = root / "results"
caselist.write_text("KHR-GL32.a\nKHR-GL32.b\n", encoding="utf-8")
seen_remaining = []
def fake_run(command, workdir, environment, qpa_path, stdout_path, stderr_path, *timeouts):
del workdir, environment, stdout_path, stderr_path, timeouts
seen_remaining.append(
command_path(command, "--deqp-caselist-file")
.read_text(encoding="utf-8")
.splitlines()
)
qpa_path.write_text("#sessionInfo releaseName fake\n", encoding="utf-8")
return runner.ProcessOutcome(
1, 0.1, timed_out=True, timeout_reason="qpa-idle"
)
with mock.patch.object(runner, "run_process", side_effect=fake_run):
result = runner.main(
self.run_args(root, caselist, outdir, "--max-empty-streak", "2")
)
self.assertEqual(4, result)
self.assertEqual(
[["KHR-GL32.a", "KHR-GL32.b"], ["KHR-GL32.a", "KHR-GL32.b"]],
seen_remaining,
)
self.assertEqual("", (outdir / "crashed.txt").read_text(encoding="utf-8"))
self.assertEqual("", (outdir / "hung.txt").read_text(encoding="utf-8"))
self.assertEqual(
"KHR-GL32.a\nKHR-GL32.b\n",
(outdir / "unrun.txt").read_text(encoding="utf-8"),
)
class ProcessTimeoutTests(unittest.TestCase):
def test_qpa_activity_prevents_idle_timeout(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
qpa = root / "active.qpa"
helper = (
"import pathlib,sys,time\n"
"path=pathlib.Path(sys.argv[1])\n"
"for size in range(1, 9):\n"
" path.write_text('x' * size, encoding='utf-8')\n"
" time.sleep(0.08)\n"
)
outcome = runner.run_process(
[sys.executable, "-c", helper, str(qpa)],
root,
dict(runner.os.environ),
qpa,
root / "stdout.log",
root / "stderr.log",
idle_timeout=0.2,
max_round_seconds=0,
poll_seconds=0.03,
)
self.assertFalse(outcome.timed_out)
self.assertEqual(0, outcome.returncode)
def test_idle_timeout_really_stops_process(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
started = time.monotonic()
outcome = runner.run_process(
[sys.executable, "-c", "import time; time.sleep(30)"],
root,
dict(runner.os.environ),
root / "never-created.qpa",
root / "stdout.log",
root / "stderr.log",
idle_timeout=0.2,
max_round_seconds=0,
poll_seconds=0.05,
)
elapsed = time.monotonic() - started
self.assertTrue(outcome.timed_out)
self.assertEqual("qpa-idle", outcome.timeout_reason)
self.assertLess(elapsed, 10)
if __name__ == "__main__":
unittest.main()
@@ -0,0 +1,294 @@
import argparse
import json
from pathlib import Path
import struct
import tempfile
import unittest
import wgl_glcts_pipeline as pipeline
def write_fake_pe(path: Path, machine: int = pipeline.PE_MACHINE_AMD64, payload: bytes = b"") -> None:
data = bytearray(0x88)
data[0:2] = b"MZ"
struct.pack_into("<I", data, 0x3C, 0x80)
data[0x80:0x84] = b"PE\0\0"
struct.pack_into("<H", data, 0x84, machine)
path.write_bytes(bytes(data) + payload)
class ArgumentTests(unittest.TestCase):
def test_versions_accept_gl_and_dotted_spellings(self):
self.assertEqual("30", pipeline.normalize_version("GL30"))
self.assertEqual("46", pipeline.normalize_version("4.6"))
with self.assertRaises(argparse.ArgumentTypeError):
pipeline.normalize_version("4.7")
def test_environment_assignment_validation(self):
self.assertEqual(("MOBILEGL_TEST", "a=b"), pipeline.parse_assignment("MOBILEGL_TEST=a=b"))
with self.assertRaises(argparse.ArgumentTypeError):
pipeline.parse_assignment("9BAD=value")
def test_windows_environment_keys_are_canonical_and_last_wins(self):
self.assertEqual(
{"FOO": "x", "PATH": "second"},
pipeline.canonicalize_windows_environment(
[("Path", "first"), ("FOO", "x"), ("pAtH", "second")]
),
)
class CommandTests(unittest.TestCase):
def test_visual_studio_configure_commands_are_x64_and_wgl_default(self):
mobilegl = pipeline.mobilegl_configure_command(
Path("C:/src/MobileGL"), Path("D:/work/mg"), "Visual Studio 17 2022", "x64", []
)
self.assertIn("-A", mobilegl)
self.assertIn("x64", mobilegl)
self.assertIn("-DMOBILEGL_BUILD_TEST=OFF", mobilegl)
cts = pipeline.cts_configure_command(
Path("D:/src/VK-GL-CTS"), Path("D:/work/cts"), "Visual Studio 17 2022", "x64", []
)
self.assertIn("-DDEQP_TARGET=default", cts)
self.assertNotIn("-DDEQP_TARGET=mobilegl", cts)
def test_runner_command_contains_identity_preserving_wgl_flags(self):
command = pipeline.runner_command(
Path("runner.py"),
Path("runtime/glcts.exe"),
Path("cts/modules"),
Path("gl46-main.txt"),
Path("results/gl46"),
"DirectVulkan",
300,
0,
10000,
{"MOBILEGL_LOG_FILE_PATH": "result/mobilegl.log"},
pipeline.DEFAULT_DEQP_ARGS,
"run-fingerprint",
)
self.assertIn("--backend", command)
self.assertIn("DirectVulkan", command)
self.assertIn("--deqp-arg=--deqp-gl-context-type=wgl", command)
self.assertIn("--deqp-arg=--deqp-surface-type=fbo", command)
self.assertIn("--env", command)
self.assertIn("MOBILEGL_LOG_FILE_PATH=result/mobilegl.log", command)
self.assertIn("--run-identity", command)
self.assertIn("run-fingerprint", command)
def test_report_command_requires_the_pipeline_run_identity(self):
command = pipeline.report_command(
Path("report.py"),
[("DirectVulkan", "gl46", Path("gl46.txt"), Path("results/gl46"))],
Path("summary.md"),
Path("summary.json"),
False,
"run-fingerprint",
)
self.assertIn("--expected-run-identity", command)
self.assertIn("run-fingerprint", command)
class RuntimeTests(unittest.TestCase):
def test_pe_machine_rejects_non_x64(self):
with tempfile.TemporaryDirectory() as temporary:
path = Path(temporary) / "x86.dll"
write_fake_pe(path, machine=0x14C)
self.assertEqual(0x14C, pipeline.pe_machine(path))
with self.assertRaises(pipeline.PipelineError):
pipeline.require_x64_pe(path, "test DLL")
def test_runtime_is_hash_keyed_and_copies_only_declared_files(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
glcts = root / "source-glcts.exe"
mobilegl = root / "source-MobileGL.dll"
write_fake_pe(glcts, payload=b"glcts")
write_fake_pe(mobilegl, payload=b"mobilegl")
sources = {"glcts.exe": glcts, "opengl32.dll": mobilegl}
fingerprint, hashes = pipeline.runtime_fingerprint(sources)
runtime = pipeline.assemble_runtime(root / "work", sources, fingerprint, hashes)
self.assertEqual(fingerprint[:16], runtime.name)
self.assertEqual(hashes["glcts.exe"], pipeline.sha256_file(runtime / "glcts.exe"))
self.assertEqual(hashes["opengl32.dll"], pipeline.sha256_file(runtime / "opengl32.dll"))
manifest = json.loads((runtime / "manifest.json").read_text(encoding="utf-8"))
self.assertEqual(fingerprint, manifest["fingerprint"])
self.assertFalse((runtime / "libEGL.dll").exists())
def test_run_fingerprint_changes_with_execution_semantics(self):
base = pipeline.run_fingerprint(
"runtime", "data", {"runner": "tool"}, {"30": "caselist"}, ["--deqp-surface-type=fbo"], {"FLAG": "1"}
)
self.assertEqual(
base,
pipeline.run_fingerprint(
"runtime", "data", {"runner": "tool"}, {"30": "caselist"}, ["--deqp-surface-type=fbo"], {"FLAG": "1"}
),
)
self.assertNotEqual(
base,
pipeline.run_fingerprint(
"runtime", "data", {"runner": "tool"}, {"30": "caselist"}, ["--deqp-surface-type=window"], {"FLAG": "1"}
),
)
self.assertNotEqual(
base,
pipeline.run_fingerprint(
"runtime", "data", {"runner": "tool"}, {"30": "different"}, ["--deqp-surface-type=fbo"], {"FLAG": "1"}
),
)
self.assertNotEqual(
base,
pipeline.run_fingerprint(
"runtime", "different-data", {"runner": "tool"}, {"30": "caselist"}, ["--deqp-surface-type=fbo"], {"FLAG": "1"}
),
)
self.assertNotEqual(
base,
pipeline.run_fingerprint(
"runtime", "data", {"runner": "different-tool"}, {"30": "caselist"}, ["--deqp-surface-type=fbo"], {"FLAG": "1"}
),
)
timeout_baseline = pipeline.run_fingerprint(
"runtime",
"data",
{"runner": "tool"},
{"30": "caselist"},
["--deqp-surface-type=fbo"],
{"FLAG": "1"},
{"idle_timeout_seconds": 300.0, "max_round_seconds": 0.0},
)
idle_changed = pipeline.run_fingerprint(
"runtime",
"data",
{"runner": "tool"},
{"30": "caselist"},
["--deqp-surface-type=fbo"],
{"FLAG": "1"},
{"idle_timeout_seconds": 1.0, "max_round_seconds": 0.0},
)
max_round_changed = pipeline.run_fingerprint(
"runtime",
"data",
{"runner": "tool"},
{"30": "caselist"},
["--deqp-surface-type=fbo"],
{"FLAG": "1"},
{"idle_timeout_seconds": 300.0, "max_round_seconds": 60.0},
)
self.assertNotEqual(timeout_baseline, idle_changed)
self.assertNotEqual(timeout_baseline, max_round_changed)
def test_tracked_environment_is_case_insensitive_and_narrow(self):
ambient, effective = pipeline.tracked_run_environment(
{
"Path": "ambient-path",
"mobilegl_debug": "0",
"LibGL_Driver": "ambient-libgl",
"vK_iCd_fIlEnAmEs": "ambient-icd",
"ANGLE_DEFAULT_PLATFORM": "vulkan",
"Egl_Test": "1",
"D3D_Feature": "1",
"DxVk_Config": "ambient-dxvk",
"HOME": "ignored",
"PATH_EXTRA": "ignored",
"MOBILEGL": "ignored",
},
{"pAtH": "explicit-path", "vk_icd_filenames": "explicit-icd", "CUSTOM": "kept"},
)
self.assertEqual("ambient-path", ambient["PATH"])
self.assertNotIn("HOME", ambient)
self.assertNotIn("PATH_EXTRA", ambient)
self.assertNotIn("MOBILEGL", ambient)
self.assertEqual("explicit-path", effective["PATH"])
self.assertEqual("explicit-icd", effective["VK_ICD_FILENAMES"])
self.assertEqual("kept", effective["CUSTOM"])
def test_reserved_environment_names_cannot_hide_behind_case(self):
overrides = pipeline.canonicalize_windows_environment(
[("mobilegl_backend_type", "DirectGLES")]
)
self.assertEqual(
{"MOBILEGL_BACKEND_TYPE"},
pipeline.CONTROLLED_ENVIRONMENT_NAMES & set(overrides),
)
def test_directgles_requires_complete_angle_runtime(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
glcts = root / "glcts.exe"
mobilegl = root / "MobileGL.dll"
write_fake_pe(glcts)
write_fake_pe(mobilegl)
with self.assertRaises(pipeline.PipelineError):
pipeline.runtime_source_files(glcts, mobilegl, ["DirectGLES"], root / "angle")
angle = root / "angle"
angle.mkdir()
for name in pipeline.ANGLE_REQUIRED_DLLS:
write_fake_pe(angle / name, payload=name.encode("ascii"))
files = pipeline.runtime_source_files(glcts, mobilegl, ["DirectGLES"], angle)
self.assertEqual(
{"glcts.exe", "opengl32.dll", *pipeline.ANGLE_REQUIRED_DLLS}, set(files)
)
class CaselistTests(unittest.TestCase):
def test_preflight_prefers_small_buffer_case(self):
with tempfile.TemporaryDirectory() as temporary:
caselist = Path(temporary) / "gl30-main.txt"
caselist.write_text(
"KHR-GL30.api.coverage\nKHR-GL30.buffer_objects.gen_buffers\n",
encoding="utf-8",
)
self.assertEqual("KHR-GL30.buffer_objects.gen_buffers", pipeline.choose_preflight_case(caselist))
class PreflightTests(unittest.TestCase):
def write_identity(self, root: Path, renderer: str, version: str = "4.6"):
qpa = root / "chunk0000.qpa"
qpa.write_text(
'#sessionInfo vendor "MobileGL-Dev"\n'
f'#sessionInfo renderer "{renderer}"\n'
'#sessionInfo commandLineParameters "--deqp-gl-context-type=wgl --deqp-surface-type=fbo"\n',
encoding="utf-8",
)
log = root / "mobilegl.log"
log.write_text(f"Target OpenGL Version: {version}\n", encoding="utf-8")
return qpa, log
def test_identity_accepts_both_mobilegl_renderers(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
qpa, log = self.write_identity(root, "Magma (MobileGL Core)")
identity = pipeline.parse_preflight_identity([qpa], log, "DirectVulkan", (4, 6))
self.assertEqual("4.6", identity["target_gl_version"])
qpa, log = self.write_identity(root, "Espryt (MobileGL Core)")
identity = pipeline.parse_preflight_identity([qpa], log, "DirectGLES", (3, 3))
self.assertIn("Espryt", identity["renderer"])
def test_identity_rejects_system_driver_or_low_version(self):
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
qpa, log = self.write_identity(root, "NVIDIA GeForce RTX", version="4.6")
qpa.write_text(
'#sessionInfo vendor "NVIDIA Corporation"\n'
'#sessionInfo renderer "NVIDIA GeForce RTX"\n'
'#sessionInfo commandLineParameters "--deqp-gl-context-type=wgl"\n',
encoding="utf-8",
)
with self.assertRaises(pipeline.PipelineError):
pipeline.parse_preflight_identity([qpa], log, "DirectVulkan", (4, 6))
qpa, log = self.write_identity(root, "Magma (MobileGL Core)", version="4.5")
with self.assertRaises(pipeline.PipelineError):
pipeline.parse_preflight_identity([qpa], log, "DirectVulkan", (4, 6))
if __name__ == "__main__":
unittest.main()
+914
View File
@@ -0,0 +1,914 @@
#!/usr/bin/env python
"""Build MobileGL's Windows WGL shim and run Khronos OpenGL CTS suites.
The pipeline intentionally keeps the build, runtime, results, and reports in an
explicit work root. Each run is keyed by the hashes of glcts.exe, opengl32.dll,
and (for DirectGLES) the ANGLE runtime, so resuming can never silently combine
results from different binaries.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
from pathlib import Path
import re
import shutil
import struct
import subprocess
import sys
from datetime import datetime, timezone
from typing import Iterable, Mapping, Optional, Sequence
SUPPORTED_VERSIONS = ("30", "31", "32", "33", "40", "41", "42", "43", "44", "45", "46")
SUPPORTED_BACKENDS = ("DirectGLES", "DirectVulkan")
ANGLE_REQUIRED_DLLS = ("libEGL.dll", "libGLESv2.dll", "d3dcompiler_47.dll")
ANGLE_OPTIONAL_DLLS = ("dxcompiler.dll", "dxil.dll")
PE_MACHINE_AMD64 = 0x8664
TRACKED_ENVIRONMENT_NAMES = frozenset({"PATH"})
TRACKED_ENVIRONMENT_PREFIXES = (
"MOBILEGL_",
"LIBGL_",
"VK_",
"ANGLE_",
"EGL_",
"D3D_",
"DXVK_",
)
CONTROLLED_ENVIRONMENT_NAMES = frozenset(
{"MOBILEGL_BACKEND_TYPE", "MOBILEGL_LOG_FILE_PATH"}
)
DEFAULT_DEQP_ARGS = (
"--deqp-gl-context-type=wgl",
"--deqp-surface-type=fbo",
"--deqp-gl-config-name=rgba8888d24s8",
"--deqp-surface-width=64",
"--deqp-surface-height=-1",
"--deqp-base-seed=3",
"--deqp-visibility=hidden",
"--deqp-watchdog=enable",
"--deqp-crashhandler=enable",
)
SESSION_VENDOR = re.compile(r'^#sessionInfo vendor "([^"]*)"', re.MULTILINE)
SESSION_RENDERER = re.compile(r'^#sessionInfo renderer "([^"]*)"', re.MULTILINE)
SESSION_COMMAND_LINE = re.compile(r'^#sessionInfo commandLineParameters "([^"]*)"', re.MULTILINE)
TARGET_GL_VERSION = re.compile(r"Target OpenGL Version:\s*(\d+)\.(\d+)")
class PipelineError(RuntimeError):
"""A configuration, build, or identity error."""
def repository_root() -> Path:
return Path(__file__).resolve().parents[3]
def utc_now() -> str:
return datetime.now(timezone.utc).isoformat(timespec="seconds")
def normalize_version(value: str) -> str:
normalized = value.strip().lower().removeprefix("gl").replace(".", "")
if normalized not in SUPPORTED_VERSIONS:
supported = ", ".join(f"gl{version}" for version in SUPPORTED_VERSIONS)
raise argparse.ArgumentTypeError(f"unsupported GL suite {value!r}; choose one of: {supported}")
return normalized
def gl_version_tuple(version: str) -> tuple[int, int]:
return int(version[0]), int(version[1])
def parse_assignment(value: str) -> tuple[str, str]:
name, separator, setting = value.partition("=")
if not separator or not name or "\x00" in value:
raise argparse.ArgumentTypeError(f"expected NAME=VALUE, got {value!r}")
if not re.fullmatch(r"[A-Za-z_][A-Za-z0-9_]*", name):
raise argparse.ArgumentTypeError(f"invalid environment variable name {name!r}")
return name, setting
def canonicalize_windows_environment(
items: Iterable[tuple[str, str]],
) -> dict[str, str]:
"""Canonicalize environment keys using Windows' case-insensitive rules."""
result: dict[str, str] = {}
for name, value in items:
result[name.upper()] = value
return result
def tracked_run_environment(
inherited: Mapping[str, str], overrides: Mapping[str, str]
) -> tuple[dict[str, str], dict[str, str]]:
"""Return tracked ambient values and the effective values used for identity."""
canonical_inherited = canonicalize_windows_environment(inherited.items())
ambient = {
name: value
for name, value in canonical_inherited.items()
if name in TRACKED_ENVIRONMENT_NAMES
or name.startswith(TRACKED_ENVIRONMENT_PREFIXES)
}
effective = dict(ambient)
effective.update(canonicalize_windows_environment(overrides.items()))
return dict(sorted(ambient.items())), dict(sorted(effective.items()))
def command_text(command: Sequence[object]) -> str:
return subprocess.list2cmdline([str(part) for part in command])
def run_command(
command: Sequence[object],
*,
cwd: Optional[Path] = None,
env: Optional[Mapping[str, str]] = None,
check: bool = True,
capture: bool = False,
) -> subprocess.CompletedProcess[str]:
rendered = command_text(command)
location = f" (cwd={cwd})" if cwd else ""
print(f"[wgl_glcts_pipeline] $ {rendered}{location}", flush=True)
completed = subprocess.run(
[str(part) for part in command],
cwd=str(cwd) if cwd else None,
env=dict(env) if env else None,
text=True,
capture_output=capture,
check=False,
)
if check and completed.returncode != 0:
detail = ""
if capture:
detail = f"\nstdout:\n{completed.stdout}\nstderr:\n{completed.stderr}"
raise PipelineError(f"command failed with exit code {completed.returncode}: {rendered}{detail}")
return completed
def require_file(path: Path, label: str) -> Path:
resolved = path.expanduser().resolve()
if not resolved.is_file():
raise PipelineError(f"{label} does not exist or is not a file: {resolved}")
return resolved
def require_directory(path: Path, label: str) -> Path:
resolved = path.expanduser().resolve()
if not resolved.is_dir():
raise PipelineError(f"{label} does not exist or is not a directory: {resolved}")
return resolved
def sha256_file(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as stream:
for block in iter(lambda: stream.read(1024 * 1024), b""):
digest.update(block)
return digest.hexdigest()
def sha256_directory(path: Path) -> str:
digest = hashlib.sha256()
files = sorted(
(candidate for candidate in path.rglob("*") if candidate.is_file()),
key=lambda candidate: candidate.relative_to(path).as_posix(),
)
if not files:
raise PipelineError(f"directory contains no files to fingerprint: {path}")
for candidate in files:
relative = candidate.relative_to(path).as_posix()
digest.update(relative.encode("utf-8"))
digest.update(b"\0")
digest.update(sha256_file(candidate).encode("ascii"))
digest.update(b"\n")
return digest.hexdigest()
def pe_machine(path: Path) -> int:
with path.open("rb") as stream:
if stream.read(2) != b"MZ":
raise PipelineError(f"not a PE executable: {path}")
stream.seek(0x3C)
offset_bytes = stream.read(4)
if len(offset_bytes) != 4:
raise PipelineError(f"truncated PE header: {path}")
pe_offset = struct.unpack("<I", offset_bytes)[0]
stream.seek(pe_offset)
if stream.read(4) != b"PE\0\0":
raise PipelineError(f"invalid PE signature: {path}")
machine_bytes = stream.read(2)
if len(machine_bytes) != 2:
raise PipelineError(f"truncated PE COFF header: {path}")
return struct.unpack("<H", machine_bytes)[0]
def require_x64_pe(path: Path, label: str) -> None:
machine = pe_machine(path)
if machine != PE_MACHINE_AMD64:
raise PipelineError(f"{label} must be an x64 PE (machine 0x8664), got 0x{machine:04x}: {path}")
def git_snapshot(path: Path) -> dict[str, object]:
snapshot: dict[str, object] = {"path": str(path)}
try:
head = run_command(
["git", "-C", path, "rev-parse", "HEAD"], check=True, capture=True
).stdout.strip()
status = run_command(
["git", "-C", path, "status", "--porcelain"], check=True, capture=True
).stdout
snapshot.update({"head": head, "dirty": bool(status.strip())})
except (OSError, PipelineError):
snapshot.update({"head": None, "dirty": None})
return snapshot
def generator_arguments(generator: str, architecture: str) -> list[str]:
arguments = ["-G", generator]
if generator.lower().startswith("visual studio"):
arguments.extend(["-A", architecture])
return arguments
def mobilegl_configure_command(
repo_root: Path,
build_dir: Path,
generator: str,
architecture: str,
extra: Iterable[str],
) -> list[str]:
return [
"cmake",
"-S",
str(repo_root),
"-B",
str(build_dir),
*generator_arguments(generator, architecture),
"-DMOBILEGL_BUILD_TEST=OFF",
"-DMOBILEGL_BUILD_BENCHMARK=OFF",
"-DMOBILEGL_BUILD_TRACE_REPLAY=OFF",
"-DMOBILEGL_ENABLE_TRACY=OFF",
"-DMOBILEGL_FORCE_RELEASE_OPT=ON",
*extra,
]
def cts_configure_command(
cts_source: Path,
build_dir: Path,
generator: str,
architecture: str,
extra: Iterable[str],
) -> list[str]:
return [
"cmake",
"-S",
str(cts_source),
"-B",
str(build_dir),
*generator_arguments(generator, architecture),
"-DDEQP_TARGET=default",
"-DDEQP_SUPPORT_DRM=OFF",
*extra,
]
def build_command(build_dir: Path, configuration: str, target: str, jobs: int) -> list[str]:
command = ["cmake", "--build", str(build_dir), "--config", configuration, "--target", target]
if jobs > 0:
command.extend(["--parallel", str(jobs)])
return command
def verify_mobilegl_sources(repo_root: Path) -> None:
required = (
repo_root / "CMakeLists.txt",
repo_root / "MobileGL" / "MG_Impl" / "WGLImpl" / "WGLImpl.cpp",
repo_root / "3rdparty" / "glslang" / "CMakeLists.txt",
repo_root / "3rdparty" / "SPIRV-Cross" / "CMakeLists.txt",
repo_root / "3rdparty" / "Vulkan-Headers" / "CMakeLists.txt",
)
missing = [str(path) for path in required if not path.is_file()]
if missing:
raise PipelineError(
"MobileGL source/submodules are incomplete:\n "
+ "\n ".join(missing)
+ f"\nRun: git -C {repo_root} submodule update --init --recursive"
)
def verify_cts_sources(cts_source: Path) -> None:
required = (
cts_source / "CMakeLists.txt",
cts_source / "external" / "openglcts" / "CMakeLists.txt",
)
missing = [str(path) for path in required if not path.is_file()]
if missing:
raise PipelineError(
"VK-GL-CTS source/external packages are incomplete:\n "
+ "\n ".join(missing)
+ f"\nRun: {sys.executable} {cts_source / 'external' / 'fetch_sources.py'}"
)
def discover_mobilegl_dll(build_dir: Path, configuration: str) -> Path:
preferred = (
build_dir / configuration / "opengl32.dll",
build_dir / "MobileGL" / configuration / "opengl32.dll",
build_dir / "opengl32.dll",
)
for candidate in preferred:
if candidate.is_file():
return candidate.resolve()
candidates = sorted({path.resolve() for path in build_dir.rglob("opengl32.dll") if path.is_file()})
if len(candidates) == 1:
return candidates[0]
if not candidates:
raise PipelineError(f"MobileGL build produced no opengl32.dll under {build_dir}")
raise PipelineError("multiple opengl32.dll candidates; pass --mobilegl-dll explicitly:\n " + "\n ".join(map(str, candidates)))
def discover_glcts_exe(build_dir: Path, configuration: str) -> Path:
preferred = (
build_dir / "external" / "openglcts" / "modules" / configuration / "glcts.exe",
build_dir / "external" / "openglcts" / "modules" / "glcts.exe",
)
for candidate in preferred:
if candidate.is_file():
return candidate.resolve()
candidates = sorted({path.resolve() for path in build_dir.rglob("glcts.exe") if path.is_file()})
if len(candidates) == 1:
return candidates[0]
if not candidates:
raise PipelineError(f"CTS build produced no glcts.exe under {build_dir}")
raise PipelineError("multiple glcts.exe candidates; pass --glcts-exe explicitly:\n " + "\n ".join(map(str, candidates)))
def default_cts_modules_dir(cts_build_dir: Path) -> Path:
return cts_build_dir / "external" / "openglcts" / "modules"
def find_caselist_root(cts_modules_dir: Path, cts_source: Path) -> Path:
relative = Path("gl_cts/data/mustpass/gl/khronos_mustpass/main")
candidates = (cts_modules_dir / relative, cts_source / "external" / "openglcts" / "modules" / relative)
for candidate in candidates:
if candidate.is_dir():
return candidate.resolve()
raise PipelineError("Khronos GL mustpass directory was not found; checked:\n " + "\n ".join(map(str, candidates)))
def caselist_for(caselist_root: Path, version: str) -> Path:
return require_file(caselist_root / f"gl{version}-main.txt", f"GL{version} mustpass caselist")
def runtime_source_files(
glcts_exe: Path,
mobilegl_dll: Path,
backends: Sequence[str],
angle_dir: Optional[Path],
) -> dict[str, Path]:
files = {"glcts.exe": glcts_exe, "opengl32.dll": mobilegl_dll}
if "DirectGLES" in backends:
if angle_dir is None:
raise PipelineError("--angle-dir is required when DirectGLES is selected")
angle_dir = require_directory(angle_dir, "ANGLE runtime directory")
for name in ANGLE_REQUIRED_DLLS:
files[name] = require_file(angle_dir / name, f"ANGLE {name}")
for name in ANGLE_OPTIONAL_DLLS:
candidate = angle_dir / name
if candidate.is_file():
files[name] = candidate.resolve()
return files
def runtime_fingerprint(files: Mapping[str, Path]) -> tuple[str, dict[str, str]]:
hashes = {name: sha256_file(path) for name, path in sorted(files.items())}
digest = hashlib.sha256()
for name, file_hash in hashes.items():
digest.update(f"{name}\0{file_hash}\n".encode("utf-8"))
return digest.hexdigest(), hashes
def run_fingerprint(
runtime_hash: str,
cts_data_hash: str,
tool_hashes: Mapping[str, str],
caselist_hashes: Mapping[str, str],
deqp_args: Sequence[str],
environment: Mapping[str, str],
result_semantics: Optional[Mapping[str, object]] = None,
) -> str:
identity = {
"version": 2,
"runtime_fingerprint": runtime_hash,
"cts_data_sha256": cts_data_hash,
"tool_hashes": dict(sorted(tool_hashes.items())),
"caselist_hashes": dict(sorted(caselist_hashes.items())),
"deqp_args": list(deqp_args),
"environment": dict(sorted(environment.items())),
"result_semantics": dict(sorted((result_semantics or {}).items())),
}
return hashlib.sha256(
json.dumps(identity, sort_keys=True, separators=(",", ":")).encode("utf-8")
).hexdigest()
def assemble_runtime(
work_root: Path,
sources: Mapping[str, Path],
fingerprint: str,
hashes: Mapping[str, str],
) -> Path:
runtime_dir = work_root / "runtime" / fingerprint[:16]
runtime_dir.mkdir(parents=True, exist_ok=True)
for name, source in sources.items():
require_x64_pe(source, name)
destination = runtime_dir / name
if source.resolve() != destination.resolve():
shutil.copy2(source, destination)
if sha256_file(destination) != hashes[name]:
raise PipelineError(f"runtime copy hash mismatch: {destination}")
manifest = {
"version": 1,
"created_utc": utc_now(),
"fingerprint": fingerprint,
"files": {
name: {"source": str(source), "sha256": hashes[name]}
for name, source in sorted(sources.items())
},
}
write_json(runtime_dir / "manifest.json", manifest)
return runtime_dir
def write_json(path: Path, value: object) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
temporary = path.with_name(path.name + ".tmp")
temporary.write_text(json.dumps(value, indent=2, sort_keys=True) + "\n", encoding="utf-8")
os.replace(temporary, path)
def read_cases(caselist: Path) -> list[str]:
cases: list[str] = []
for raw_line in caselist.read_text(encoding="utf-8-sig").splitlines():
line = raw_line.strip()
if line and not line.startswith("#"):
cases.append(line)
if not cases:
raise PipelineError(f"caselist contains no cases: {caselist}")
return cases
def choose_preflight_case(caselist: Path) -> str:
cases = read_cases(caselist)
preferred_suffixes = (
".buffer_objects.gen_buffers",
".CommonBugs.CommonBug_GetProgramivActiveUniformBlockMaxNameLength",
)
for suffix in preferred_suffixes:
for case in cases:
if case.endswith(suffix):
return case
return cases[0]
def runner_command(
runner: Path,
runtime_exe: Path,
workdir: Path,
caselist: Path,
outdir: Path,
backend: str,
idle_timeout: float,
max_round_seconds: float,
max_rounds: int,
environment: Mapping[str, str],
deqp_args: Sequence[str],
run_identity: Optional[str] = None,
) -> list[str]:
command = [
sys.executable,
str(runner),
"--exe",
str(runtime_exe),
"--workdir",
str(workdir),
"--caselist",
str(caselist),
"--outdir",
str(outdir),
"--backend",
backend,
"--idle-timeout",
str(idle_timeout),
"--max-round-seconds",
str(max_round_seconds),
"--max-rounds",
str(max_rounds),
]
if run_identity is not None:
command.extend(["--run-identity", run_identity])
for name, value in sorted(environment.items()):
command.extend(["--env", f"{name}={value}"])
command.extend(f"--deqp-arg={argument}" for argument in deqp_args)
return command
def parse_preflight_identity(
qpa_files: Sequence[Path],
mobilegl_log: Path,
backend: str,
minimum_version: tuple[int, int],
) -> dict[str, object]:
if not qpa_files:
raise PipelineError(f"{backend} preflight produced no QPA file")
qpa_text = "\n".join(path.read_text(encoding="utf-8", errors="replace") for path in qpa_files)
vendors = SESSION_VENDOR.findall(qpa_text)
renderers = SESSION_RENDERER.findall(qpa_text)
command_lines = SESSION_COMMAND_LINE.findall(qpa_text)
if not vendors or "MobileGL" not in vendors[-1]:
raise PipelineError(f"{backend} preflight did not load MobileGL (vendor={vendors[-1] if vendors else None!r})")
expected_renderer = "Espryt" if backend == "DirectGLES" else "Magma"
if not renderers or expected_renderer not in renderers[-1]:
raise PipelineError(
f"{backend} preflight renderer mismatch: expected {expected_renderer!r}, "
f"got {renderers[-1] if renderers else None!r}"
)
if not command_lines or "--deqp-gl-context-type=wgl" not in command_lines[-1]:
raise PipelineError(f"{backend} preflight did not record a WGL context")
if not mobilegl_log.is_file():
raise PipelineError(f"{backend} preflight did not create MobileGL log: {mobilegl_log}")
log_text = mobilegl_log.read_text(encoding="utf-8", errors="replace")
versions = [(int(major), int(minor)) for major, minor in TARGET_GL_VERSION.findall(log_text)]
if not versions:
raise PipelineError(f"{backend} preflight log contains no target OpenGL version")
actual_version = versions[-1]
if actual_version < minimum_version:
raise PipelineError(
f"{backend} reports GL {actual_version[0]}.{actual_version[1]}, "
f"but selected suites require at least {minimum_version[0]}.{minimum_version[1]}"
)
return {
"backend": backend,
"vendor": vendors[-1],
"renderer": renderers[-1],
"target_gl_version": f"{actual_version[0]}.{actual_version[1]}",
"qpa_files": [str(path) for path in qpa_files],
"mobilegl_log": str(mobilegl_log),
}
def run_preflight(
*,
runner: Path,
runtime_exe: Path,
cts_modules_dir: Path,
caselist: Path,
preflight_root: Path,
backend: str,
idle_timeout: float,
environment: Mapping[str, str],
deqp_args: Sequence[str],
minimum_version: tuple[int, int],
run_identity: str,
) -> dict[str, object]:
outdir = preflight_root / backend.lower()
outdir.mkdir(parents=True, exist_ok=True)
case_file = outdir / "case.txt"
case_file.write_text(choose_preflight_case(caselist) + "\n", encoding="utf-8")
log_path = outdir / "mobilegl.log"
child_environment = dict(environment)
child_environment["MOBILEGL_LOG_FILE_PATH"] = str(log_path)
command = runner_command(
runner,
runtime_exe,
cts_modules_dir,
case_file,
outdir,
backend,
min(idle_timeout, 120.0) if idle_timeout > 0 else 120.0,
180.0,
1,
child_environment,
deqp_args,
run_identity,
)
# A developing driver may fail the chosen case or terminate during deinit.
# Identity is the gate: the QPA and MobileGL log must prove which WGL driver ran.
run_command(command, cwd=repository_root(), check=False)
identity = parse_preflight_identity(sorted(outdir.glob("chunk*.qpa")), log_path, backend, minimum_version)
print(
f"[wgl_glcts_pipeline] preflight {backend}: {identity['renderer']} | "
f"GL {identity['target_gl_version']}"
)
return identity
def report_command(
reporter: Path,
suites: Sequence[tuple[str, str, Path, Path]],
markdown: Path,
json_out: Path,
allow_incomplete: bool,
expected_run_identity: Optional[str] = None,
) -> list[str]:
command = [sys.executable, str(reporter)]
for backend, label, caselist, results in suites:
command.extend(["--suite", backend, label, str(caselist), str(results)])
command.extend(["--markdown", str(markdown), "--json", str(json_out)])
if expected_run_identity is not None:
command.extend(["--expected-run-identity", expected_run_identity])
if allow_incomplete:
command.append("--allow-incomplete")
return command
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Build MobileGL WGL, assemble a local glcts runtime, run GL30-GL46, and report results."
)
parser.add_argument("--repo-root", type=Path, default=repository_root(), help="MobileGL worktree root")
parser.add_argument("--cts-source", type=Path, required=True, help="VK-GL-CTS source checkout")
parser.add_argument("--work-root", type=Path, required=True, help="build/result root (kept outside source)")
parser.add_argument("--angle-dir", type=Path, help="x64 ANGLE directory for DirectGLES")
parser.add_argument("--backends", nargs="+", choices=SUPPORTED_BACKENDS, default=list(SUPPORTED_BACKENDS))
parser.add_argument("--versions", nargs="+", type=normalize_version, default=list(SUPPORTED_VERSIONS))
parser.add_argument("--configuration", default="Release")
parser.add_argument("--generator", default="Visual Studio 17 2022")
parser.add_argument("--architecture", default="x64")
parser.add_argument("--jobs", type=int, default=max(1, os.cpu_count() or 1))
parser.add_argument("--mobilegl-build-dir", type=Path)
parser.add_argument("--cts-build-dir", type=Path)
parser.add_argument("--mobilegl-dll", type=Path, help="reuse an existing MobileGL/opengl32 DLL")
parser.add_argument("--glcts-exe", type=Path, help="reuse an existing glcts.exe")
parser.add_argument("--cts-modules-dir", type=Path, help="glcts working directory containing gl_cts data")
parser.add_argument("--skip-mobilegl-build", action="store_true")
parser.add_argument("--skip-cts-build", action="store_true")
parser.add_argument("--skip-preflight", action="store_true")
parser.add_argument("--skip-run", action="store_true")
parser.add_argument("--skip-report", action="store_true")
parser.add_argument("--allow-incomplete-report", action="store_true")
parser.add_argument("--continue-on-suite-error", action="store_true")
parser.add_argument("--idle-timeout", type=float, default=300.0)
parser.add_argument("--max-round-seconds", type=float, default=0.0)
parser.add_argument("--max-rounds", type=int, default=10000)
parser.add_argument("--env", action="append", type=parse_assignment, default=[], metavar="NAME=VALUE")
parser.add_argument(
"--deqp-arg", action="append", default=[], metavar="ARG", help="extra glcts option; use --deqp-arg=--x=y"
)
parser.add_argument(
"--mobilegl-cmake-arg", action="append", default=[], metavar="ARG", help="extra MobileGL configure option"
)
parser.add_argument("--cts-cmake-arg", action="append", default=[], metavar="ARG", help="extra CTS configure option")
return parser
def execute(args: argparse.Namespace) -> int:
if os.name != "nt":
raise PipelineError("this pipeline builds and exercises the Windows WGL target and must run on Windows")
if shutil.which("cmake") is None and (not args.skip_mobilegl_build or not args.skip_cts_build):
raise PipelineError("cmake was not found on PATH")
if args.jobs < 0 or args.max_rounds <= 0:
raise PipelineError("--jobs must be >= 0 and --max-rounds must be > 0")
repo_root = require_directory(args.repo_root, "MobileGL worktree")
cts_source = require_directory(args.cts_source, "VK-GL-CTS checkout")
work_root = args.work_root.expanduser().resolve()
work_root.mkdir(parents=True, exist_ok=True)
versions = list(dict.fromkeys(args.versions))
backends = list(dict.fromkeys(args.backends))
minimum_version = max(gl_version_tuple(version) for version in versions)
extra_environment = canonicalize_windows_environment(args.env)
reserved_environment = CONTROLLED_ENVIRONMENT_NAMES & set(extra_environment)
if reserved_environment:
raise PipelineError("the pipeline controls these environment variables: " + ", ".join(sorted(reserved_environment)))
configuration = args.configuration
mobilegl_build_dir = (args.mobilegl_build_dir or work_root / f"mobilegl-build-{configuration.lower()}").resolve()
cts_build_dir = (args.cts_build_dir or work_root / f"cts-build-wgl-{configuration.lower()}").resolve()
if not args.skip_mobilegl_build:
verify_mobilegl_sources(repo_root)
mobilegl_build_dir.mkdir(parents=True, exist_ok=True)
run_command(
mobilegl_configure_command(
repo_root, mobilegl_build_dir, args.generator, args.architecture, args.mobilegl_cmake_arg
),
cwd=repo_root,
)
run_command(build_command(mobilegl_build_dir, configuration, "MobileGL", args.jobs), cwd=repo_root)
mobilegl_dll = (
require_file(args.mobilegl_dll, "MobileGL DLL")
if args.mobilegl_dll
else discover_mobilegl_dll(mobilegl_build_dir, configuration)
)
if not args.skip_cts_build:
verify_cts_sources(cts_source)
cts_build_dir.mkdir(parents=True, exist_ok=True)
run_command(
cts_configure_command(cts_source, cts_build_dir, args.generator, args.architecture, args.cts_cmake_arg),
cwd=cts_source,
)
run_command(build_command(cts_build_dir, configuration, "glcts", args.jobs), cwd=cts_source)
glcts_exe = (
require_file(args.glcts_exe, "glcts executable")
if args.glcts_exe
else discover_glcts_exe(cts_build_dir, configuration)
)
cts_modules_dir = require_directory(
args.cts_modules_dir or default_cts_modules_dir(cts_build_dir), "CTS modules/working directory"
)
cts_data_dir = require_directory(cts_modules_dir / "gl_cts" / "data", "CTS gl_cts data directory")
cts_data_hash = sha256_directory(cts_data_dir)
caselist_root = find_caselist_root(cts_modules_dir, cts_source)
caselists = {version: caselist_for(caselist_root, version) for version in versions}
runtime_sources = runtime_source_files(glcts_exe, mobilegl_dll, backends, args.angle_dir)
fingerprint, hashes = runtime_fingerprint(runtime_sources)
runtime_dir = assemble_runtime(work_root, runtime_sources, fingerprint, hashes)
runtime_exe = runtime_dir / "glcts.exe"
runner = require_file(repo_root / "tools" / "cts" / "scripts" / "run_cts_windows.py", "Windows CTS runner")
reporter = require_file(repo_root / "tools" / "cts" / "scripts" / "cts_multi_report.py", "CTS reporter")
matrix_reporter = require_file(
repo_root / "tools" / "cts" / "scripts" / "cts_matrix_report.py", "CTS matrix reporter"
)
qpa_reporter = require_file(repo_root / "tools" / "cts" / "scripts" / "qpa_report.py", "QPA parser")
tool_hashes = {
"wgl_glcts_pipeline.py": sha256_file(Path(__file__).resolve()),
"run_cts_windows.py": sha256_file(runner),
"cts_multi_report.py": sha256_file(reporter),
"cts_matrix_report.py": sha256_file(matrix_reporter),
"qpa_report.py": sha256_file(qpa_reporter),
}
deqp_args = [*DEFAULT_DEQP_ARGS, *args.deqp_arg]
caselist_hashes = {version: sha256_file(path) for version, path in caselists.items()}
ambient_environment, identity_environment = tracked_run_environment(
os.environ, extra_environment
)
result_semantics = {
"idle_timeout_seconds": args.idle_timeout,
"max_round_seconds": args.max_round_seconds,
}
execution_settings = {
**result_semantics,
"max_rounds": args.max_rounds,
"continue_on_suite_error": args.continue_on_suite_error,
}
execution_fingerprint = run_fingerprint(
fingerprint,
cts_data_hash,
tool_hashes,
caselist_hashes,
deqp_args,
identity_environment,
result_semantics,
)
run_root = work_root / "runs" / execution_fingerprint[:16]
report_root = run_root / "reports"
manifest = {
"version": 1,
"created_utc": utc_now(),
"run_fingerprint": execution_fingerprint,
"runtime_fingerprint": fingerprint,
"runtime_hashes": hashes,
"tool_hashes": tool_hashes,
"cts_data": {"path": str(cts_data_dir), "sha256": cts_data_hash},
"runtime_dir": str(runtime_dir),
"mobilegl": git_snapshot(repo_root),
"vk_gl_cts": git_snapshot(cts_source),
"configuration": configuration,
"generator": args.generator,
"architecture": args.architecture,
"backends": backends,
"versions": versions,
"caselists": {version: {"path": str(path), "sha256": caselist_hashes[version]} for version, path in caselists.items()},
"deqp_args": deqp_args,
"environment_overrides": extra_environment,
"ambient_environment": ambient_environment,
"identity_environment": identity_environment,
"result_semantics": result_semantics,
"execution_settings": execution_settings,
}
write_json(run_root / "manifest.json", manifest)
print(f"[wgl_glcts_pipeline] runtime fingerprint: {fingerprint}")
print(f"[wgl_glcts_pipeline] run fingerprint: {execution_fingerprint}")
print(f"[wgl_glcts_pipeline] run root: {run_root}")
identities: list[dict[str, object]] = []
if not args.skip_preflight:
first_caselist = caselists[versions[0]]
preflight_root = run_root / "preflight"
for backend in backends:
identities.append(
run_preflight(
runner=runner,
runtime_exe=runtime_exe,
cts_modules_dir=cts_modules_dir,
caselist=first_caselist,
preflight_root=preflight_root,
backend=backend,
idle_timeout=args.idle_timeout,
environment=extra_environment,
deqp_args=deqp_args,
minimum_version=minimum_version,
run_identity=execution_fingerprint,
)
)
manifest["preflight"] = identities
write_json(run_root / "manifest.json", manifest)
suites: list[tuple[str, str, Path, Path]] = []
suite_errors: list[dict[str, object]] = []
for backend in backends:
for version in versions:
label = f"gl{version}"
result_dir = run_root / "results" / backend.lower() / label
suites.append((backend, label, caselists[version], result_dir))
if args.skip_run:
continue
result_dir.mkdir(parents=True, exist_ok=True)
child_environment = dict(extra_environment)
child_environment["MOBILEGL_LOG_FILE_PATH"] = str(result_dir / "mobilegl.log")
command = runner_command(
runner,
runtime_exe,
cts_modules_dir,
caselists[version],
result_dir,
backend,
args.idle_timeout,
args.max_round_seconds,
args.max_rounds,
child_environment,
deqp_args,
execution_fingerprint,
)
completed = run_command(command, cwd=repo_root, check=False)
if completed.returncode != 0:
suite_errors.append({"backend": backend, "suite": label, "returncode": completed.returncode})
if not args.continue_on_suite_error:
break
if suite_errors and not args.continue_on_suite_error:
break
report_returncode: Optional[int] = None
if not args.skip_report:
report_root.mkdir(parents=True, exist_ok=True)
completed = run_command(
report_command(
reporter,
suites,
report_root / "gl-cts-summary.md",
report_root / "gl-cts-summary.json",
args.allow_incomplete_report or bool(suite_errors),
execution_fingerprint,
),
cwd=repo_root,
check=False,
)
report_returncode = completed.returncode
manifest["suite_errors"] = suite_errors
manifest["report_returncode"] = report_returncode
manifest["finished_utc"] = utc_now()
write_json(run_root / "manifest.json", manifest)
if suite_errors:
print(f"[wgl_glcts_pipeline] {len(suite_errors)} suite runner(s) incomplete; see manifest/report", file=sys.stderr)
returncodes = {int(item["returncode"]) for item in suite_errors}
if 130 in returncodes:
return 130
if 2 in returncodes:
return 2
if 3 in returncodes:
return 3
return 4
if report_returncode is not None and report_returncode != 0:
print(f"[wgl_glcts_pipeline] report validation failed with exit code {report_returncode}", file=sys.stderr)
return report_returncode
return 0
def main(argv: Optional[Sequence[str]] = None) -> int:
parser = build_parser()
args = parser.parse_args(argv)
try:
return execute(args)
except PipelineError as exc:
print(f"[wgl_glcts_pipeline] ERROR: {exc}", file=sys.stderr)
return 2
except OSError as exc:
print(f"[wgl_glcts_pipeline] ERROR: filesystem/process operation failed: {exc}", file=sys.stderr)
return 2
if __name__ == "__main__":
sys.exit(main())
+19
View File
@@ -0,0 +1,19 @@
# MobileGL conformance-suite skills
Task-focused skills for running Khronos conformance suites against MobileGL.
Each skill is a self-contained package, matching the layout used by
`tools/trace_replay/skills/`:
- `SKILL.md` — the skill (frontmatter `name` + `description`, then the body). The
directory name equals the frontmatter `name`.
- `agents/openai.yaml` — OpenAI agent descriptor (`display_name`,
`short_description`, `default_prompt`).
- `scripts/` and/or `references/` — bundled tooling and supporting docs, when the
skill has them.
## Skills
| Skill | What it does |
| --- | --- |
| [gl-cts-on-mobilegl](gl-cts-on-mobilegl/SKILL.md) | Build VK-GL-CTS `glcts` as a standalone Android arm64 binary against MobileGL's own EGL, run KHR-GL33, and report a per-backend OpenGL 3.3 core conformance rate. |
| [wgl-gl-cts-on-mobilegl](wgl-gl-cts-on-mobilegl/SKILL.md) | Build MobileGL's Windows x64 WGL drop-in, run GL30-GL46 core CTS against DirectGLES and DirectVulkan, resume safely, and emit validated reports. |
@@ -0,0 +1,214 @@
---
name: gl-cts-on-mobilegl
description: Run the Khronos OpenGL CTS (VK-GL-CTS glcts, KHR-GL33) against MobileGL on an Android device and compute a per-backend conformance rate. Use when measuring OpenGL 3.3 core conformance for DirectGLES or DirectVulkan, building glcts for Android arm64, porting a dEQP tcu::Platform onto MobileGL, or triaging CTS failures, crashes, and cases that hang the device.
---
# OpenGL CTS on MobileGL (Android)
## Overview
`glcts` from VK-GL-CTS is built as a **standalone arm64 executable** and run from
`adb shell`. It reaches OpenGL only through `libMobileGL.so`, which supplies both
EGL and desktop GL, so a result is unambiguously MobileGL's and never the system
GL stack's. No APK and no Activity are involved.
The port lives in this repository under `MobileGL/tools/cts/` and is copied into
a VK-GL-CTS checkout by `scripts/sync_to_cts.py`, so it survives a throwaway CTS
clone.
Set up paths first:
```sh
export MG=<path-to-MobileGL-worktree> # do builds in a worktree, not the shared tree
export CTS=<path-to-VK-GL-CTS-checkout>
export NDK="$ANDROID_HOME/ndk/27.3.13750724"
export SERIAL=<adb-device-serial>
```
## Prerequisites
- Android NDK r27 (the repo builds MobileGL with 27.3.13750724), CMake, Ninja, Python 3.
- A rooted-or-not Android device with `adb`; ~600 MB free under `/data/local/tmp`.
- **A device you can physically power-cycle.** Some cases hang the GPU hard
enough to reboot it — see "Cases that take the device down".
- On Windows, invoke `python`, not `python3`: the latter resolves to the
Microsoft Store alias stub and exits 49.
## Step 1 — build libMobileGL.so
Build in a git worktree (other agents share the main tree). A fresh worktree is
missing glslang's bundled SPIR-V Tools, which is a hard configure blocker
because `ENABLE_OPT` is forced on:
```sh
cp -r <main-tree>/3rdparty/glslang/External/* "$MG/3rdparty/glslang/External/"
./gradlew -p "$MG/android-plugin" :app:assembleTraceRelease
```
The stripped library lands in
`android-plugin/app/build/intermediates/stripped_native_libs/traceRelease/.../arm64-v8a/libMobileGL.so`.
## Step 2 — get VK-GL-CTS and its externals
Use a **release tag**, not `main`, so the mustpass list — and therefore the
reported rate — is citable:
```sh
git -C "$CTS" checkout opengl-cts-4.6.8.1
cd "$CTS" && python external/fetch_sources.py
```
## Step 3 — build glcts for Android arm64
```sh
python "$MG/tools/cts/scripts/sync_to_cts.py" "$CTS"
cmake -S "$CTS" -B build-cts-a64 -G Ninja \
-DDEQP_TARGET=mobilegl -DDEQP_TARGET_TOOLCHAIN=ndk-modern \
-DANDROID_NDK_PATH="$NDK" -DDE_ANDROID_API=26 -DANDROID_ABI=arm64-v8a \
-DCMAKE_BUILD_TYPE=Release
ninja -C build-cts-a64 glcts
"$NDK"/toolchains/llvm/prebuilt/*/bin/llvm-strip build-cts-a64/external/openglcts/modules/glcts
```
Confirm the configure output says `DE_OS = DE_OS_ANDROID`, `DE_CPU =
DE_CPU_ARM_64` and `DEQP_ANDROID_BUILD = EXE`. Two things make that work and
both are easy to get wrong:
- `DEQP_TARGET_TOOLCHAIN=ndk-modern` is required. dEQP includes `Defs.cmake`
*before* the target file, so a target cannot set `DE_OS` itself. Without the
toolchain hook the build mis-detects as `DE_OS_UNIX`/`x86_64` and dies on
`__assert_fail` (bionic has `__assert2`).
- The target sets `DEQP_ANDROID_EXE ON`. Otherwise dEQP builds the modules into
the `libdeqp.so` an APK would load and no `glcts` executable exists.
`KHR-GL33` needs no ungating — the package registry registers it unconditionally;
only the `dEQP-*` packages are `#if DE_OS != DE_OS_ANDROID`.
## Step 4 — deploy
```sh
adb -s $SERIAL shell mkdir -p /data/local/tmp/mgcts
adb -s $SERIAL push build-cts-a64/external/openglcts/modules/glcts /data/local/tmp/mgcts/
adb -s $SERIAL push build-cts-a64/external/openglcts/modules/gl_cts /data/local/tmp/mgcts/
adb -s $SERIAL push <libMobileGL.so> /data/local/tmp/mgcts/
adb -s $SERIAL shell chmod 755 /data/local/tmp/mgcts/glcts
```
## Step 5 — preflight
Never start a multi-hour run without this. It proves the device/library pair
yields a 3.3 core context and that FBO readback is correct, in about a second:
```sh
adb -s $SERIAL shell 'cd /data/local/tmp/mgcts && LD_LIBRARY_PATH=. ./mgprobe \
--backend DirectVulkan --surface imagereader --lib ./libMobileGL.so'
```
Expect `PASS ... user_fbo=ok`. `default_fb=broken` on DirectVulkan is expected
and does not gate — see below.
## Step 6 — run
```sh
python "$MG/tools/cts/scripts/run_cts.py" \
--serial $SERIAL --backend DirectGLES \
--caselist .../mustpass/gl/khronos_mustpass/main/gl33-main.txt \
--outdir runs/gles --skip-file runs/skip.txt
```
The runner re-invokes `glcts` with only the cases that have no result yet, so a
crash costs one case rather than the run. It distinguishes a crashed *case* from
a dead *device* by checking the device still answers a shell command — without
that check a dead device looks like every remaining case crashing, which yields
a completely bogus but plausible-looking conformance number. On a device reboot
it waits, re-pulls the partial `.qpa` (which survives on `/data/local/tmp`),
records the case that was open as `DeviceHang`, and quarantines it.
## Step 7 — report
```sh
python "$MG/tools/cts/scripts/qpa_report.py" runs/gles --label DirectGLES
```
Pass rate counts `Pass`, `NotSupported`, `QualityWarning`, `CompatibilityWarning`
and `Waiver` as non-failures, matching how Khronos scores a submission; the
strict rate counts only `Pass`. Quarantined and never-reached cases are reported
separately and excluded from the rates, so a partial run cannot read as a
complete one.
## Required flags, and why
| Flag | Why it is not optional |
| --- | --- |
| `--deqp-surface-type=fbo` | On DirectVulkan, `glReadPixels` from the **default framebuffer returns all zeros** with no GL error. dEQP verifies nearly everything through `glReadPixels`, so rendering to the surface scores DirectVulkan near zero for a reason unrelated to conformance. Use it for **both** backends so the two numbers stay comparable. |
| `MOBILEGL_CTS_FBO_COLOR_TEXTURE=1` | **`--deqp-surface-type=fbo` alone is not enough.** dEQP's `FboRenderContext` allocates a *renderbuffer* colour attachment, and DirectVulkan returns zeros from a renderbuffer-attached FBO too — only a *texture*-attached FBO reads back correctly. This env var (a patch to `framework/opengl/gluFboRenderContext.cpp`, off by default) switches the attachment to a texture and isolates that single defect. Measured effect: `KHR-GL33.shaders.loops.for_constant_iterations.*` goes 0/62 → 62/62, and the whole-suite DirectVulkan conformance rate goes 46.15% → 72.74%. DirectGLES is bit-identical either way (93.05%), which is the control proving the switch is neutral where readback works. |
| `--deqp-terminate-on-device-lost=disable` | Defaults to *enable*, which calls `glGetGraphicsResetStatus()` after every case. That is GL 4.5 / `KHR_robustness`, absent from GL 3.3 core, so the pointer is null and the process segfaults on the first case. Desktop drivers expose the extension, which is why upstream never trips on it. |
## Cases that take the device down
Some cases hang the GPU hard enough that the device reboots or stops answering
adb entirely. Keep them in a `--skip-file`, and expect to find more:
- `KHR-GL33.clip_distance.functional` — wedged an Adreno 750 tablet; it rebooted
and then stopped responding to adb altogether.
- `KHR-GL33.framebuffer_blit.multisampled_to_singlesampled_blit_color_config_test`
— rebooted an Adreno 830 phone after 862 cases, on DirectGLES.
- `KHR-GL33.framebuffer_blit.multisampled_to_singlesampled_blit_depth_config_test`
— same, on both backends (found and quarantined automatically by the runner).
- `KHR-GL33.texture_repeat_mode.rgb565_11x131_0_clamp_to_edge` — on DirectVulkan.
The whole `framebuffer_blit.multisampled_to_singlesampled_*` family is suspect;
treat a new variant as a device-hang candidate rather than a normal failure.
When a run dies, pull `/data/local/tmp/mgcts/chunk.qpa` — it survives the reboot,
and the last `#beginTestCaseResult` with no matching `#endTestCaseResult` names
the case that did it.
## Reference results
`opengl-cts-4.6.8.1`, KHR-GL33 mustpass (`gl33-main.txt`, 9886 cases), Adreno 830
/ Android 15, MobileGL `dev`@199164c2, 9884 measured / 0 unrun / 2 quarantined.
Conformance rate = Pass + NotSupported, as Khronos scores a submission.
| backend | conformance | strict Pass | Fail | Crash | InternalError | DeviceHang |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| DirectGLES | **93.05%** | 85.94% | 679 | 1 | 6 | 1 |
| DirectVulkan (texture FBO) | **72.74%** | 65.71% | 2394 | 294 | 5 | 1 |
| DirectVulkan (stock renderbuffer FBO) | 46.15% | 39.11% | 5024 | 292 | 5 | 2 |
The third row is what stock dEQP reports; the gap to the second row is entirely
the renderbuffer-FBO readback defect.
## MobileGL constraints the port works around
- **DirectVulkan cannot use an EGL pbuffer.** That path needs
`VK_EXT_headless_surface`, which Adreno's Android driver does not expose; it
fails inside `eglMakeCurrent`. The platform therefore gets a real
`ANativeWindow` from **`AImageReader`** — an ordinary BufferQueue producer that
`vkCreateAndroidSurfaceKHR` accepts, with no Activity. An `onImageAvailable`
listener must drain the queue or the producer blocks once `maxImages` buffers
are in flight and the next swap deadlocks.
- **`eglMakeCurrent` requires draw == read** and rejects `EGL_NO_SURFACE` with
`EGL_BAD_MATCH`, so dEQP's `surfaceless` platform cannot be used at all, and
`--deqp-surface-type=fbo` (which asks the platform for `SURFACETYPE_DONT_CARE`)
must still be given a real surface.
- **Every EGL call must go through the dynamically loaded library.** dEQP's
`surfaceless` platform mixes wrapper calls with globally linked `egl*` symbols;
copying that on Android silently reaches the system EGL and invalidates the
measurement. The `mobilegl` target links no `libEGL`/`libGLESv*` at all.
- **Desktop-GL configs need `EGL_OPENGL_BIT`.** The surfaceless port always asks
for an ES bit, which can never satisfy a GL 3.3 core context.
- MobileGL aborts during static teardown (`FORTIFY: pthread_mutex_lock called on
a destroyed mutex`) *after* the work is done; flush and `_exit()` in any small
tool, or its exit code and output are lost.
## Contents
platform/tcuMobileGLPlatform.{cpp,hpp} dEQP tcu::Platform for MobileGL
targets/mobilegl.cmake VK-GL-CTS target (-DDEQP_TARGET=mobilegl)
targets/ndk-modern.cmake NDK toolchain hook (sets DE_OS/DE_CPU)
probe/mgprobe.c preflight gate
scripts/sync_to_cts.py inject the port into a CTS checkout
scripts/run_cts.py crash- and reboot-resuming runner
scripts/qpa_report.py .qpa -> conformance rate
@@ -0,0 +1,4 @@
interface:
display_name: "OpenGL CTS on MobileGL (Android)"
short_description: "Build and run VK-GL-CTS KHR-GL33 against MobileGL and report per-backend conformance"
default_prompt: "Use $gl-cts-on-mobilegl to run the OpenGL 3.3 core CTS against MobileGL on my Android device and report the conformance rate for DirectGLES and DirectVulkan."
@@ -0,0 +1,143 @@
---
name: wgl-gl-cts-on-mobilegl
description: Build MobileGL as a Windows x64 WGL drop-in opengl32.dll, build or reuse VK-GL-CTS glcts, run Khronos GL30 through GL46 core suites against DirectGLES and DirectVulkan, resume after crashes or idle timeouts, and produce validated Markdown/JSON conformance reports. Use when Codex needs to compile MobileGL's WGL target, connect it to desktop OpenGL CTS on Windows, rerun selected GL core mustpass lists, or verify that CTS loaded MobileGL instead of the system OpenGL driver.
---
# WGL OpenGL CTS on MobileGL
Use `tools/cts/scripts/wgl_glcts_pipeline.py` as the single entry point. It
configures both Visual Studio builds, assembles a private runtime, verifies the
loaded WGL implementation, calls the crash-resuming runner, and generates the
multi-suite report.
## Prerequisites
- Run on Windows x64 with Git, Python 3.9 or newer, CMake, Visual Studio 2022's
Desktop C++ workload, and a Vulkan SDK visible to MobileGL's CMake configure.
- Use the already selected MobileGL worktree. Inspect `git status` first and do
not create another worktree unless the user explicitly asks.
- Initialize MobileGL submodules:
```powershell
git submodule update --init --recursive
```
- Prepare a VK-GL-CTS checkout at a citable release tag, then fetch its pinned
externals:
```powershell
git -C D:\VK-GL-CTS checkout opengl-cts-4.6.8.1
python D:\VK-GL-CTS\external\fetch_sources.py
```
- For DirectGLES, provide one x64 ANGLE directory containing matching
`libEGL.dll`, `libGLESv2.dll`, and `d3dcompiler_47.dll`. DirectVulkan does not
need ANGLE.
- For DirectVulkan, provide a working Vulkan loader plus a GPU-vendor ICD and
driver. The Vulkan SDK alone does not provide a usable GPU device.
## Run the full matrix
From the MobileGL worktree root:
```powershell
python tools\cts\scripts\wgl_glcts_pipeline.py `
--cts-source D:\VK-GL-CTS `
--work-root D:\MobileGL-WGL-CTS `
--angle-dir C:\path\to\angle-x64 `
--backends DirectGLES DirectVulkan `
--versions gl30 gl31 gl32 gl33 gl40 gl41 gl42 gl43 gl44 gl45 gl46
```
The defaults deliberately reproduce the proven Windows setup:
- Visual Studio 17 2022, x64, Release;
- VK-GL-CTS `DEQP_TARGET=default`, which selects desktop WGL on Windows;
- MobileGL copied beside `glcts.exe` as `opengl32.dll`;
- WGL context, FBO surface, `rgba8888d24s8`, hidden window, watchdog and crash
handler enabled;
- `--deqp-terminate-on-device-lost=disable` supplied by the underlying runner.
Do not replace the FBO surface with the default framebuffer when comparing
backends; it changes the readback path and invalidates comparison with the
established runs.
## Mandatory preflight
Leave preflight enabled for a new binary. It must prove all of the following
before the full suite starts:
- QPA vendor contains `MobileGL`;
- DirectGLES renderer contains `Espryt`, or DirectVulkan contains `Magma`;
- QPA command line records `--deqp-gl-context-type=wgl`;
- MobileGL's log reports a GL version at least as high as the highest selected
suite.
Treat any preflight failure as a hard stop. It commonly means `glcts.exe` loaded
the system `opengl32.dll`, ANGLE DLLs have the wrong architecture, or the driver
still reports too low a GL version.
## Common variants
Run DirectVulkan only, without ANGLE:
```powershell
python tools\cts\scripts\wgl_glcts_pipeline.py `
--cts-source D:\VK-GL-CTS --work-root D:\MobileGL-WGL-CTS `
--backends DirectVulkan --versions gl43 gl44 gl45 gl46
```
Build and preflight without starting a multi-hour CTS run:
```powershell
python tools\cts\scripts\wgl_glcts_pipeline.py `
--cts-source D:\VK-GL-CTS --work-root D:\MobileGL-WGL-CTS `
--angle-dir C:\path\to\angle-x64 `
--skip-run --skip-report
```
Reuse previously built binaries with `--skip-mobilegl-build`,
`--skip-cts-build`, `--mobilegl-dll`, `--glcts-exe`, and
`--cts-modules-dir`. Continue to use a modules directory containing the
`gl_cts` data tree; the executable directory alone is insufficient.
Pass extra dEQP options with the equals form so argparse does not consume the
leading dashes:
```powershell
--deqp-arg=--deqp-log-images=enable
```
## Resume and artifacts
Repeat the exact command and `--work-root` to resume. The runner recovers
completed QPA cases plus `crashed.txt` and `hung.txt`, then schedules only
unaccounted cases.
The pipeline prints full SHA-256 fingerprints and uses their first 16
hexadecimal characters as directory names: `runtime/<runtime-prefix>` and
`runs/<run-prefix>`. The run fingerprint covers the CTS data tree,
runner/report tools, caselists, dEQP arguments, explicit `--env` values, tracked
ambient GL/Vulkan environment, and the timeout settings that determine
Crash/Hang classification. The manifest records those inputs plus
`max_rounds`, suite-error continuation policy, source commits and dirty state,
preflight identity, suite errors, and report status.
The runner refuses to attach a new `run_state.json` to old QPA or sidecar files
by default. Invoke `run_cts_windows.py --adopt-legacy` directly only after
verifying that those artifacts match the backend, caselist, binaries, and dEQP
arguments. Pipeline reports require every suite state to match the current run
fingerprint.
Read the final outputs at:
```text
<work-root>/runs/<run-prefix>/reports/gl-cts-summary.md
<work-root>/runs/<run-prefix>/reports/gl-cts-summary.json
```
Use the exact run root printed by the pipeline.
Do not claim completeness when the report validation is incomplete or when the
manifest records suite errors. Keep `NotSupported` separate from hard failures
when prioritizing implementation work.
@@ -0,0 +1,4 @@
interface:
display_name: "WGL GL CTS on MobileGL"
short_description: "Build MobileGL WGL and run GL30-GL46 CTS on Windows"
default_prompt: "Use $wgl-gl-cts-on-mobilegl to build MobileGL's WGL DLL and run the selected OpenGL CTS suites on Windows."
+36
View File
@@ -0,0 +1,36 @@
#-------------------------------------------------------------------------
# VK-GL-CTS target: MobileGL on Android
#
# Builds a standalone arm64 ELF that reaches OpenGL exclusively through
# libMobileGL.so, loaded at runtime. Nothing here links libEGL or libGLESv*:
# the whole point is that the system GL stack must not be reachable, so that a
# conformance result is unambiguously MobileGL's.
#-------------------------------------------------------------------------
message("*** Using MobileGL target")
set(DEQP_TARGET_NAME "MobileGL")
# Build the modules as standalone executables instead of the libdeqp.so an APK
# would load. The suite runs from adb shell, with no Activity.
set(DEQP_ANDROID_EXE ON)
# EGL comes from libMobileGL.so via the eglw dynamic wrapper, so the support
# flag is on but no import library is supplied.
set(DEQP_SUPPORT_EGL ON)
set(DEQP_EGL_LIBRARIES)
set(DEQP_GLES2_LIBRARIES)
set(DEQP_GLES3_LIBRARIES)
set(TCUTIL_PLATFORM_SRCS
mobilegl/tcuMobileGLPlatform.cpp
mobilegl/tcuMobileGLPlatform.hpp
)
find_library(LOG_LIBRARY NAMES log)
find_library(ANDROID_LIBRARY NAMES android)
find_library(MEDIANDK_LIBRARY NAMES mediandk)
# libmediandk supplies AImageReader, which is how a process with no Activity
# gets a real ANativeWindow.
list(APPEND TCUTIL_PLATFORM_LIBS ${ANDROID_LIBRARY} ${MEDIANDK_LIBRARY} ${LOG_LIBRARY})
+61
View File
@@ -0,0 +1,61 @@
#-------------------------------------------------------------------------
# drawElements CMake utilities
# ----------------------------
#
# Copyright 2016 The Android Open Source Project
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#
#-------------------------------------------------------------------------
# Delegate most things to the NDK's cmake toolchain script
if (NOT DEFINED ANDROID_NDK_PATH)
message(FATAL_ERROR "Please provide ANDROID_NDK_PATH")
endif ()
set(ANDROID_PLATFORM "android-${DE_ANDROID_API}")
set(ANDROID_STL c++_static)
set(ANDROID_CPP_FEATURES "rtti exceptions")
include(${ANDROID_NDK_PATH}/build/cmake/android.toolchain.cmake)
# The try_compile() used to verify the C/C++ compilers are sane tries to
# generate an executable, but doesn't seem to use the right compiler/linker
# options when cross-compiling, so it fails even when building an actual
# shared library or executable succeeds.
#
# I don't know why this doesn't affect simpler projects that use the NDK
# toolchain.
set(CMAKE_TRY_COMPILE_TARGET_TYPE STATIC_LIBRARY)
# Set variables used by other parts of dEQP's build scripts
set(DE_OS "DE_OS_ANDROID")
if (NOT DEFINED DE_COMPILER)
set(DE_COMPILER "DE_COMPILER_CLANG")
endif ()
if (ANDROID_ABI STREQUAL "x86")
set(DE_CPU "DE_CPU_X86")
elseif (ANDROID_ABI STREQUAL "armeabi" OR
ANDROID_ABI STREQUAL "armeabi-v7a")
set(DE_CPU "DE_CPU_ARM")
elseif (ANDROID_ABI STREQUAL "arm64-v8a")
set(DE_CPU "DE_CPU_ARM_64")
elseif (ANDROID_ABI STREQUAL "x86_64")
set(DE_CPU "DE_CPU_X86_64")
else ()
message(FATAL_ERROR "Unknown ABI \"${ANDROID_ABI}\"")
endif ()