mirror of
https://github.com/MobileGL-Dev/MobileGL
synced 2026-09-08 04:08:32 +09:00
25a8f51db55d2ea7c25cd2f428ef8a6ea3564d3b
2008
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
25a8f51db5 |
[Perf] (MG_Backend): make DirectGLES program switches remember their own bindings
mc_use_program cycles programs whose texture bindings never change, yet every switch re-walked the units. Six fixes, one theme: a switch back to a known program should find its own state waiting. Per-program 4-entry resolved-texture-binding memo (round-robin, shadow memcmp on hit) skips the unit walk when a program returns with its bindings intact. The whole sampler-uniform pass in BindCurrentProgramWithResources is memoized per program twin behind (context, unitBindingsEpoch, samplingGeneration, backendStateVersion, textureContextGeneration) plus a per-sampled-unit sampler-shadow row compare, invalidated on relink/backend rebuild; the BindCurrentUnitSamplers walk sits behind the same keys. Every unit assignment, sampler-parameter change and bind path was verified to bump one of those inputs. UboRingAllocate's common path is now a generation check, a power-of-two mask, an overrun check and a head bump - the duplicate availability probe, frame-mark retirement and divisions moved to the wrap slow path. The per-context framebuffer binding slots (the frontend getter linear-scans per call) are cached as direct pointers - slots are by-value members of GLContext, so the pointers are stable by construction - feeding SyncCurrentFBO, SyncNeccessaryTextures and the broadcast memo; BindCurrentFBO's per-draw registry hash Find became a TwinLookupMemo probe. The VAO config-version cold-line load is hoisted to the top of PrepareForDraw to overlap its miss. Quiet-box load-gated 6-round order-alternating A/B, all nine cases, both backends: use_program -27.7%, vanilla_draw -16.2%, ubo_range -13.4%, pass_switch -11.1%, sampler_churn -10.7%, state_toggle -9.2%, sodium_multidraw -5.3%, rest flat. No regression on either backend (magma's one matrix flag disproved by isolated re-runs against byte-identical DirectVulkan sources). Unit tests 421/421. |
||
|
|
8f2b766b56 |
[Perf] (MG_Backend): let DirectVulkan trust across frames what it proved once
The draw fast path still paid for its own proofs: the hottest single load (20% of TrySetupDrawFastPath) was chasing the cold VertexInputStateFactory heap entry just to answer "same vertex-input layout?". That answer now comes from the frontend VAO's config-guarded aux memo (layout hash + attribute masks), and a VAO-cycling stream with a stable layout skips the pre-flight AND pipeline re-resolution entirely. The VkProgramObject* is memoized on the snapshot behind a new ProgramFactory cache-structure epoch (bumped on every insert/erase; use is re-stamped so the idle sweep can never evict a live entry). A render-state version move no longer forces the full path: the pipeline value hash is refreshed in place and the 8-entry memo probed directly (the GL_BLEND-toggle case). The resolved-vertex-bindings memo now revalidates all-resident unmapped entries ACROSS frames via per-binding slice epochs - minted from a process-lifetime counter so a recycled address can never revalidate, with every mutation path funnelled through BumpSliceEpoch - while stamping each resource's GPU-use serial exactly as the skipped acquire would, preserving the busy-tracking that glBufferSubData's host-write-vs-staged-copy choice depends on. Resident index buffers get the same treatment through an EBO slice memo. The six-part dynamic-state tail (viewport/scissor/blend constants/depth bias/line width/stencil) is gated behind one render-state-parameters version + pass-geometry compare per command buffer. GetSlice is inlined; SampledBindingsUnchanged walks only the program's declared bindings. Quiet-box 6-round order-alternating A/B (on top of the frontend VAO-bind commit): vanilla_draw -20.8% (790 -> 626 ns/op, 3.4x native to 2.5x), sampler_churn -28.2%, ubo_range -8.4%, state_toggle -3.8%; tex_param's matrix flag (+10%) was adjudicated by an isolated alternating re-run at +1.0% - position bias, not regression. Unit tests 421/421. |
||
|
|
b9d8ad0421 |
[Perf] (MG_Backend): give DirectGLES one epoch that says no buffer moved
Four draw-path costs, one theme: re-proving what nothing invalidated. A manager-wide buffer-mutation epoch (atomic; bumped with release AFTER every mutation lands: all six BufferBackendOps via tracking wrappers, every backend-initiated writeback - XFB readback/scatter, the five pack-PBO readbacks - registry registration changes, and backend context destruction; the full site inventory lives in a comment at the accessor) lets the per-VAO resolved-buffers memo stamp the epoch after one all-clean probe pass and skip every IsBufferDrawClean probe while it holds. The IBO keeps its bound-object identity compare - only the probe is elided. Non-bumping paths are enumerated with why they are safe: GPU-authoritative writes are ignored by the probe, persistent-mapped resources are clean by construction, and draws on non-persistent maps are frontend-rejected GL errors. GetProgramForDraw is hoisted to one call per PrepareForDraw and handed to the four consumers that each re-derived it. The enabled-draw-buffers walk feeding the fragColor broadcast count is memoized on the (FBO, slot version, object version) trio. The UBO-binding loop probes IsBufferDrawClean before falling back to EnsureBufferResource. The texture chain captures (context, maxTouchedUnit, samplingGeneration, unitBindingsEpoch) once per draw - shared by SyncNeccessaryTextures and BindCurrentTextures, halving the epoch computations - and an aggregate gate that is the exact conjunction of the three Sync*ToBackend early-outs skips the per-texture cross-TU calls. The t_egl* thread_local verification pair became owner-thread-guarded atomics reset by MakeCurrent/ReleaseCurrent, removing __tls_get_addr from the draw loop. Quiet-box 6-round order-alternating A/B (with the frontend VAO-bind commit): all NINE Espryt cases improved - sampler_churn -11.0%, state_toggle -9.3%, ubo_range -9.1%, vanilla_draw -5.4%, pass_switch -3.5%, the rest -1% to -2.5%. Unit tests 421/421. |
||
|
|
f8069c0624 |
[Perf] (MG_State): stop paying two atomic refcounts for every glBindVertexArray
perf annotate put 94% of VertexArrayState::Bind's 10.5% self time on the two lock-prefixed shared_ptr refcount RMWs each bind performs. The bound VAO is now stored as a slot index into m_vertexArrays - no SharedPtr copy, no atomics on the bind path. The lifetime invariant (the bound object is kept alive by its slot; any cold path that clobbers a bound slot - delete-while-bound including slot 0, create-over-bound-slot - detaches the old object into m_boundDetached so GetBoundVertexArray keeps answering with it) is enforced in MarkVertexArrayForDeletion / CreateVertexArrayObject rather than assumed, and documented at the change. Out-of-range binds and null slots keep their exact old semantics. VertexArrayObject also gains two opaque config-version-guarded backend aux memo words, letting a backend answer "same vertex-input layout?" from the frontend object instead of chasing its own cold cache entry. After this change the frontend Bind drops out of the DirectVulkan draw profile entirely (11.5% -> 0.5%). Measured jointly with the two backend rounds that land on top: quiet-box 6-round order-alternating A/B, all nine cases, no case worse than noise on either backend. Unit tests 421/421. |
||
|
|
d0aae85da2 |
[Perf] (MG_Backend): let DirectVulkan's draw fast path survive a VAO swap
TrySetupDrawFastPath declined on its VAO pointer check for every draw of a 512-VAO cycle - the Blaze3D chunk-render shape - so the fast path was dead exactly where it mattered: full SetupDraw, per-draw ResolveSamplerDescriptor, SyncTextureAndGetDescriptor and render-pass re-fetch, for draws whose only change was the VAO. Three fixes. A moved VAO now re-runs only the vertex-input pre-flight and re-resolves the pipeline instead of declining to the full path. That resolution probes the value-keyed pipeline memo directly off a cached pipeline-state hash and snapshot render-pass hash, skipping GetOrCreateRenderPass and its GetPendingRenderbufferClear probes per draw; a stale cached hash can only miss, never false-hit. And when the sampler-descriptor hint holds and the program's single dynamic UBO re-resolves to the same VkBuffer and range - only the dynamic offset moved, the per-draw glUniform case - the descriptor walk collapses to one offset recompute and a vkCmdBindDescriptorSets of the same recorded set with new pDynamicOffsets. The rebind memo is invalidated at BeginFrame, layout destruction and override walks; the program lifetime id never repeats, and per-frame descriptor sets are never rewritten within their frame. mc_vanilla_draw -36.9% (1260 -> 795 ns/op, 4.6x native to 3.4x), sodium_multidraw -18.0%, state_toggle -14.9%, sampler_churn -7.8%, use_program -7.7%, ubo_range -7.3%, tex_param -7.3%. All nine cases on both backends, interleaved A/B; no attributable regression. Unit tests 421/421. |
||
|
|
b904658b10 |
[Perf] (MG_Backend): stop DirectGLES re-resolving the same VAO's buffers and twins every draw
Four per-draw costs, all lookups that re-answer the same question. SyncNeccessaryBuffers walked all 32 attribute slots cold and ran EnsureBufferResource per buffer on every draw. The backend VAO twin now hosts a resolved-draw-buffers memo: the deduped enabled-attribute buffers and the index buffer resolve once per VAO config version, and each hit re-validates every entry with IsBufferDrawClean - a shadow probe mirroring every no-op branch of EnsureBufferResource (resource identity, context generation, pending ops, change serial) - falling back to the full path for just the dirty entries. The IBO entry is checked against the live bound object each draw, so slot-version wrap cannot false-hit. The registry hash Finds that resolve state objects to their backend twins ran several times per draw. TwinLookupMemo - a direct-mapped, Fibonacci-hashed table (4096 VAO / 256 program slots) with weak-ptr owner equality against address reuse - answers them in one probe; collisions fall back to the registry. A live entry's twin is never replaced once set, so owner equality proves the raw pointer. SyncCurrentVertexAttributeValues' pending-mask memo was a function-static single entry that missed every draw once the app cycled VAOs; it now lives on the twin. CurrentXfb()'s per-draw FastSTL map lookup became a cached pointer invalidated at every map mutation (open addressing moves values on any insert/erase/clear). mc_vanilla_draw -12.6% (3.4x native to 3.0x), ubo_range -9.8%, sampler_churn -7.5%, pass_switch -6.5%, sodium_multidraw -6.0%, state_toggle -4.6%. All nine cases measured on both backends, interleaved A/B; no case regressed. Unit tests 421/421. |
||
|
|
4b3fd11462 |
[Perf] (MG_Backend): key DirectVulkan's pipeline memo on state values, not a version that never repeats
Two per-draw churn costs, one cause each. A blend toggle switched pipelines through a memo keyed on a monotonic pipeline-state version - which never repeats, so flipping GL_BLEND off and back on produced a "new" key both times, forced the full SetupDraw and rebuilt the whole pipeline payload for a pipeline the cache already held. The memo now keys on a value hash of the pipeline-relevant fixed-function state, recomputed only when the state version moved, and the consecutive-draw fast path re-resolves just the pipeline through it when nothing but render state changed. Blaze3D brackets every batch with exactly this toggle; mc_state_toggle drops 36% (6629 -> 4230 ns/op, 4.8x native to 3.7x). The sampler-churn cost had the same shape as the Espryt side fixed separately: glBindSampler bumps the frontend texture-bind generation even when it re-binds the sampler the unit already holds, so the per-draw fast path died every draw. The fast path now proves each binding's descriptor inputs unchanged - texture and sampler lifetime ids, parameter and content sums, the sampling-resolution generation, image epochs and exact layouts - and reuses the binding's cached VkDescriptorImageInfo instead of re-running the resolve chain. mc_sampler_churn drops 30% (1597 -> 1125), and the proof machinery pays for itself on the uniform-range case too (-17%). mc_tex_param stays where it is on this backend deliberately: profiling shows its remaining cost is frontend validation with zero backend work, unreachable from Renderer/. All nine cases measured on both backends, interleaved A/B, no case worse than noise. Unit tests 421/421. |
||
|
|
9be5d95440 |
[Perf] (MG_Backend): give DirectGLES unit bindings an epoch the sampler churn cannot fake
The texture-binding memos added earlier keyed on the frontend texture-bind generation, and 26.2-style unit switching defeats them: glBindSampler bumps the generation even when it re-binds the sampler the unit already carries, so a frame that cycles active units re-ran the full two-pass, eleven-slot alias resolution and the unbind walks on every draw. mc_sampler_churn sat at 1674 ns/op against the native driver's 239 - the worst multiplier left on this backend - with about half the time in two virtual calls per binding slot. The units now carry an epoch: a snapshot of each touched unit's slot objects and sampler object, compared by weak_ptr OWNERSHIP rather than raw pointer - a held weak_ptr pins its control block, so a freed-and-recycled object can never owner-equal its predecessor, which is the ABA hole a pointer key would have and the reason version keying was rejected (WithTemporarilyBoundNamedTexture bumps slot versions without touching the bind generation). The (context id, bind generation, high-water mark) triple gates the snapshot walk to at most once per draw; the epoch moves only when a binding really changed. Both per-draw memos key on the epoch plus the sampling-resolution generation, which carries what the epoch cannot see: a default texture's image appearing, and every completeness input. Two smaller memos ride along: the per-unit sampler-registry lookup (owner-keyed, misses never cached - the backend object may be created later in the same draw), and the pending-vertex-attribute mask, whose first version scanned all 32 slots and put +10% on the VAO-cycling case before being restricted to the program's active locations. ns per op, DriverBench on a GTX 1660 SUPER, isolated A/B, all nine cases on both backends: mc_sampler_churn 1673 -> 732, mc_use_program 4513 -> 4279, mc_state_toggle 2365 -> 2247, everything else within noise and nothing worse. 7.0x native to 3.1x on the churn case. Unit tests 421/421. |
||
|
|
d49d79a64b |
[Perf] (MG_Backend): pool DirectVulkan's upload staging and batch its submits
Every dirty texture bought itself a fresh staging buffer (vmaCreateBuffer + vmaMapMemory), a fresh command buffer, a fresh fence, and its own vkQueueSubmit. A perf profile of the sprite-animation case put 41% of the whole run in the kernel on the resulting ioctl traffic; the reclaim list already avoided waiting on the fences, so the cost was the allocation and submission machinery itself, paid per texture per frame. Staging now comes from a pool of persistently-mapped blocks (1 MiB minimum, exact-size beyond that, bump-allocated, 32 MiB idle cap), and uploads record into one shared batch command buffer from a dedicated command pool, going out as one submit with one pooled fence per flush. Fences, command buffers and blocks all recycle through the existing fence-list reclaim instead of being destroyed. Flush points: before every frame command buffer submission (which is what preserves the old ordering argument - the batch reaches the queue strictly before anything that could sample its images), on the glFlush finite-time path, when a batch would outgrow its staging bound, and eagerly at 128 KiB, which measured faster because the GPU overlaps the copy with the rest of the frame's CPU recording. The mid-frame upload-draw-upload-again sequence detects itself through the batch image list and flushes first, reproducing the old two-submit granularity exactly; a deferred image release flushes any open batch that still references the image, because drain proofs only cover submitted work. ns per op, DriverBench on a GTX 1660 SUPER: mc_tex_stream 9405 -> 5373 (2.3x the native driver, from 3.9x), atlas_sprite -57%, lightmap -89%, chunk_upload -10%; draw-path cases unchanged. The suite's sampler-churn number reads a few percent worse right after the now-much-faster upload case, which was chased to schedutil downclocking during the newly-blocking-free frames - isolated and frequency-pinned runs measure parity; noted here so the next person does not re-chase it. Unit tests 421/421; Vulkan validation layer clean across draw and upload cases. |
||
|
|
f5761ea1f3 |
[Perf] (MG_Backend): diff only the render-state span that moved, and gate the per-draw walks
Four per-draw costs in DirectGLES, all of the same species: work re-done for an answer that had not changed. SyncRenderState was guarded by a single version compare, so one blend toggle - the way Blaze3D brackets every batch - re-diffed the whole ~40-field render state block and copied the full struct back into the shadow, every draw. The parameter struct is now split into three contiguous byte spans, each gated by a memcmp against the backend shadow; a per-draw blend flip touches only the blend span. The shadow is byte-cloned after each sync so the span compares stay exact, padding included. Blocks whose inputs live outside the parameter struct (the surface-size viewport fallback, the sRGB context capability) stay ungated, and the dual-source-blend hard-fail still fires every draw because a throwing sync never stamps the shadow. SyncMipmapsToBackend gained a first-level clean gate on (context id, sampling-resolution generation, content version, params version) that skips the IsComplete walk and the eight-field shape probe outright; every shape mutation funnels through BumpShapeVersion, which is what makes the gate sound. SyncToBackend for vertex arrays compares one aggregate config version instead of three stamps per attribute slot. And SyncNeccessaryTextures memoises the draw-framebuffer attachment list, keyed the same way the framebuffer sync memo already is, instead of re-walking attachments per draw. ns per draw, DriverBench on a GTX 1660 SUPER, isolated A/B: mc_state_toggle 3151 -> 2397, mc_ubo_range 792 -> 579, mc_vanilla_draw 1111 -> 881, mc_sampler_churn 2019 -> 1676, mc_use_program 5132 -> 4356; every one of the nine cases improved. Against the native driver Espryt now stands at 3.6x on the plain draw path, 2.8x on the per-draw uniform-range path and 2.1x on the blend toggle, from 8.7x / 9.1x / 7.2x when this effort began. Unit tests 421/421. |
||
|
|
b3f774d2c0 |
[Fix] (CI): name the EGL vendor library the benchmark job runs on
The benchmark job is the only one that brings a real GL context up - DriverBench dlopens libEGL.so.1 and renders through it - but its apt list only asks for libegl1, which is glvnd's dispatch layer and nothing more. The vendor library behind it, libegl-mesa0, has been arriving as a Recommends of libegl1 rather than because anything asked for it. That is too quiet a dependency for the one job whose whole purpose is running a driver: a base image change, or --no-install-recommends turning up anywhere upstream, would leave eglInitialize with no vendor to dispatch to and fail the job for a reason nothing in the workflow explains. Name it, next to libgl1-mesa-dri, which is listed for exactly the same reason. Verified with a full headless ctest -C Release -L benchmark - no $DISPLAY, no $EGL_PLATFORM, mesa as the only EGL vendor: SanityBench, ProgramBench, BufferBench and DriverBench all pass. |
||
|
|
d524330032 |
[Test] (MG_Benchmark, MG_Util): model four more Minecraft frame patterns in the driver bench
The captured traces contain per-frame patterns the bench did not exercise, and first measurements show two of them are now the worst remaining multipliers - which is exactly what the missing cases were hiding. mc_pass_switch: the 26.2 snapshot switches render targets 132 times a frame and re-declares draw buffers 198 times. Render-target churn is where a Vulkan backend pays for render-pass breaks and where a tiler pays most on device, and no case measured it. mc_state_toggle: Blaze3D brackets batches with blend toggles - 46 enable/disable pairs and 28 blend-func changes per vanilla frame. mc_tex_param: 26.2 re-sets texture parameters 612 times a frame, almost always to the value already in place, so this measures redundant-parameter filtering. mc_use_program: Sodium switches programs 62 times a frame with a mat4 upload on each, roughly one switch per multi-draw. All four live in the shared case file at the measured per-frame rates, so the desktop harness, the on-device harness and the POST screen's Run Bench report comparable numbers. First desktop measurements (ns/op, native / Espryt / Magma): pass_switch 8877 / 18502 / 13896, state_toggle 1182 / 8526 / 8305, tex_param 42 / 102 / 197, use_program 2182 / 10648 / 5096. The state-toggle multiplier - 7x on both backends - is the largest newly exposed gap and the next optimization target. Unit tests 421/421; the Android JNI translation unit compiles against the extended case set. |
||
|
|
f2d210b12d |
[Perf] (MG_Backend): memoise DirectVulkan's per-draw vertex binding resolution
Every draw re-resolved its whole vertex binding array: for each enabled binding, look up the buffer, acquire a slice from the buffer manager, apply the binding's base offset, fill the VkBuffer and offset arrays, bind. In the Minecraft-shaped benchmark the same few hundred vertex array objects cycle for the whole run and each one's answer is stable, so UploadAndBindVertexBuffers was the single largest cost in the backend at 7.9% of the render thread, with AcquireResidentSlice another 3.8% underneath it. The resolved array is now kept per vertex array object and revalidated instead of rebuilt. Validation is two-tier. The vertex array's own configuration version already invalidates its backend vertex-input state, so a changed attribute, format, buffer or base offset yields a different state object - the memo compares both that object's address and its hash, which mixes the bound buffers and the whole layout. What that does not cover is the slice moving underneath an unchanged configuration, so the buffer manager now carries a monotonic epoch that every writer of slice-deciding state bumps: resident storage creation, respecify, sub-data, flush of a mapped range, the promotion and demotion between streamed and resident storage, each fresh arena allocation, and bulk release. The counter is manager-wide and never reset, so a resource created at a recycled address cannot reproduce a value some memo still holds. The miss path was the thing to get right, because the previous attempt in this area regressed the texture-upload and sampler-churn cases by 60-85%: it added a verification pass that re-ran the resolution work it was trying to skip, so every miss paid for it twice. Here a miss is one pointer-keyed lookup and a few stores, and nothing else runs that the full path would not have run anyway. ns per draw, DriverBench on a GTX 1660 SUPER: mc_ubo_range 924 -> 767, mc_vanilla_draw 1346 -> 1227, mc_sampler_churn 1397 -> 1279, mc_sodium_multidraw 3365 -> 3266. Magma is now 4.1x the native driver on the per-draw uniform-range case, from 5.4x when this round started. No case regressed on either backend. Unit tests 421/421. |
||
|
|
fd40960f70 |
[Perf] (MG_Backend): revive DirectGLES's dead framebuffer-sync guard, and stop probing twice
SyncCurrentFBO has an early-out that compares three memos, and it could never fire. One of the three, g_fboBindVersions, was only ever stamped by ForceBindCurrentFBO - which runs from glBlitFramebuffer and the DSA glClearNamedFramebuffer* paths and nowhere else. An application that touches neither leaves that memo at 0 while the binding slot's version is at least 1 from its first glBindFramebuffer, so the first term mismatched forever and the guard was dead code rather than merely too coarse. Every draw therefore re-walked all 40-odd attachment slots and rebuilt the 8-slot snorm/unorm clamp mask for a framebuffer that had not changed since the previous draw. SyncCurrentFBO now stamps all three memos itself, through one helper, on every path that leaves the target synced - including the default-framebuffer "nothing to do" path, which previously returned without stamping anything. The memo is renamed to say what it now records (a sync, not a bind). Instrumenting a throwaway build put it at 539998 hits against 2 misses, the misses being the first bind of each target; it was 0 hits before. Skipping the sync also skips the Bind() inside it, so all eleven call sites were checked: every one issues its own bind afterwards (PrepareForDraw and the glClearBuffer* paths bind Draw, ReadPixels and the CopyTexSubImage paths bind Read, BlitFramebuffer binds both, GetTexImage uses its own scoped binder). The global snorm/unorm clamp masks written inside the sync stay correct because they can only be stale if a different framebuffer was synced as Draw in between, which moves the pointer or slot version and forces the re-sync that rewrites them. InvalidateFramebufferBindingCache now also clears these memos: both its callers mean the ES context may have been reset, and a live early-out must not survive that. Two smaller items in the same pass. StateBackendObjectRegistry kept the backend twin and its liveness weak_ptr in two maps, so every lookup cost two hash probes and the draw path does ten to twenty of them; they are one map with one entry type now, one probe. The weak_ptr check itself is load-bearing and stays - glDeleteVertexArrays followed by glGenVertexArrays recycles heap addresses readily. And SyncNeccessaryBuffers ran the full EnsureBufferResource check once per enabled vertex attribute, which on an interleaved Minecraft-shaped VAO means four to eight times over the same VBO; it is deduplicated per distinct buffer now. ns per draw, DriverBench on a GTX 1660 SUPER, A/B against a build differing only by this diff: mc_vanilla_draw 1403 -> 1113, mc_ubo_range 983 -> 797, mc_sampler_churn 2309 -> 2003, mc_sodium_multidraw 3232 -> 3023. Against the native driver Espryt is now 4.3x on both the plain draw and the per-draw uniform-range case, from 8.7x and 9.1x at the start of this work. Unit tests 421/421. Also replayed all 38 locally-available DirectGLES trace fixtures against a baseline library: every one produced bit-identical ssim and mismatched-pixel counts, including the improved-transparency OIT trace whose scratch clear framebuffer is exactly the draw-buffer hazard the code comments warn about. |
||
|
|
49aab57f03 |
[Perf] (MG_Backend, MG_State): stop re-resolving texture unit bindings on every draw
DirectGLES re-derived the whole texture binding state for every draw: for each touched unit, two alias-resolution passes over all binding slots, then a third walk to unbind native targets nothing claimed, then the sampler. With the Minecraft-shaped bench that was 13.2% of the render thread in BindCurrentTextures alone, plus 4.6% in SyncNeccessaryTextures deciding which textures to consider. The answer is identical across a whole terrain batch. The resolution is now memoised, and what makes replaying it as a no-op legitimate is that the memo does not merely trust a key: it compares the backend's own bound texture shadow against the one resolution left behind. Every path that binds a texture behind this function's back already maintains that shadow - the scratch bind an upload does on the temp unit, CopyTexSubImage2D and GenerateMipmap binding on the active unit, the glBindTextures fast path, the scrub a backend texture performs when it is destroyed or respecified - so a memcmp catches all of them without having to enumerate them. On top of that the key covers the texture bind generation, the program that arbitrates aliased targets (pointer, lifetime id, backend state version, link status), and the ES context generation. Two invalidation sources had no signal at all and needed one. Mipmap completeness decides whether a texture is bound in the first place, and it moves with texture shape and with the effective sampler's filter - so a sampling-resolution generation now moves with both, routed through single choke points (TextureObjectBase::BumpShapeVersion, SamplerObject::BumpVersion) so a future bump site cannot forget it. A texture context id was needed because both generations restart at zero in a new GLContext, which can land on the old heap address. This also closes a pre-existing hole rather than working around it: glDeleteSamplers unbinds the sampler from every unit straight through TextureUnit::SetSamplerObject, bypassing the touch bookkeeping, so that setter now bumps the bind generation on a real change. The sampler bind step itself stays outside the memo and runs every draw - the program's raw-depth-fetch substitution rewrites unit samplers immediately afterwards, so a memo there could never hit. ns per draw, DriverBench on a GTX 1660 SUPER (native / Espryt): mc_vanilla_draw 253 / 2037->1315, mc_ubo_range 202 / 1684->955, mc_sodium_multidraw 739 / 3939->3150. Espryt goes from 8.3x to 4.7x the native driver on the per-draw uniform-range case. Magma is unaffected (the MG_State additions are counter bumps), and no case regressed. Unit tests 421/421. |
||
|
|
62dea3bea4 |
[Perf] (MG_State): answer texture sampling completeness from a memo
Every draw asks, for every bound texture, whether it is mipmap-complete for the filter in use, and the answer was recomputed from scratch each time: walk the level chain, read each level's texel size, verify each is half the previous. With the Minecraft-shaped bench that walk plus the GetTexelSize calls under it measured about 8% of the render thread on both backends. The answer depends only on the texture's shape - internal format, stored level set, level sizes, level range - and never on its texel content, which is the thing that actually changes between draws. A shape version now moves on exactly those four mutations (SetInternalFormat, SetBaseLevel/SetMaxLevel, and the AllocateStorage/TruncateMipmapLevels pair on both mipmap storage classes), and the completeness answer is memoised against it, one slot for the mipmapped question and one for the plain one. An upload leaves the memo standing, which is the whole point; anything that could change the answer invalidates it. ns per draw, DriverBench on a GTX 1660 SUPER (native / Espryt / Magma): mc_vanilla_draw 257 / 2201->2037 / 1550->1346, mc_ubo_range 203 / 1832->1684 / 1089->934, mc_sampler_churn 272 / 2349->2325 / 1533->1396. Texture-upload cases are unchanged, as expected - they were never asking this question in a loop. Unit tests 421/421. |
||
|
|
57aeeec053 |
[Perf] (MG_State, MG_Impl, MG_Backend): stop paying per draw and per upload for work already known
A per-draw CPU profile of a real Minecraft frame (perf on the render thread, which sits at 100% of one core on both backends) said the deficit is translation overhead, not the GPU, and named where it goes. This removes the largest items it found, on both backends and in the shared frontend they both feed. The single biggest one was not translation at all: IsBackendContextCurrentOnThisThread called eglGetCurrentContext on every invocation, and glvnd answers that with a getpid() fork check - a real syscall. The predicate sits two and three deep in every draw (the deferred-release drain, the global-UBO ring availability check, and the ring allocation), so it accounted for 16.3% of the render thread. EGL is still the ground truth, but re-verifying it once per thread per frame catches an external migration at the next frame boundary rather than the next call, which recovers the same bookkeeping. Texture uploads now carry a dirty region instead of a per-level flag. Minecraft animates atlas sprites with 16x16 glTexSubImage2D calls into a 1024x512 atlas and respecifies the lightmap every frame; a per-level flag turned each of those into a full-level re-upload - about 3.6 MB a frame of texels nobody changed. MipmapStorage accumulates the written box, Espryt uploads it with UNPACK_ROW_LENGTH striding into the level shadow, and Magma stages just that box. The box is a union, not a range list: repeated writes to one level widen it and it degrades to exactly the old whole-level upload, which is the honest worst case. glBufferData(NULL) is the orphaning idiom, and the backend was answering it by uploading the stale CPU shadow - turning a rename the driver does for free into a full synchronized upload. BufferObject now records that a NULL respecify leaves the store undefined, and the upload is skipped until content is actually written. The rest are smaller and of a kind: the deferred-release queue is probed without taking its mutex, the UBO ring waits on the frame fence that frees the space it needs instead of draining the whole pipeline with glFinish at the size cap, VAO binds go through a shadow so a draw's second bind of the same object does not reach the driver, the per-draw clean-texture probe short-circuits on the content version before rebuilding shape info, glUniform drops byte-identical writes (which otherwise dirty the whole UBO for the next draw), re-binding the texture or VAO a slot already holds no longer bumps the generation counters a backend fast path is keyed on, and the texture validators stopped taking shared_ptr by value. On Magma: descriptor-set reuse keeps four entries instead of one, because draws alternating between two programs - the chunk/entity ping-pong - thrashed a single slot into a full re-allocate and re-write every draw; a DynamicDraw buffer whose contents survive two frame boundaries is promoted to resident storage instead of being re-copied into the per-frame arena forever; and sampled-read barriers name only the shader stages whose device feature is enabled, which also removes a latent VUID violation (ALL_GRAPHICS names geometry and tessellation stages a device need not have). Measured with the Minecraft rig (render distance 32, p50 fps, same machine, single sample each): vanilla 1.21.1 Espryt 10.8 -> 36.3 and Magma 31.3 -> 44.6; 26.2 snapshot Magma 114.5 -> 210.5. Fabric+Sodium moved inside noise on Magma (854 -> 766) with the native baseline itself moving 838 -> 1031 between the two sessions, so treat that cell as unresolved rather than a regression measured. Unit tests 421/421. The CTS A/B was not run: these numbers and the test suite are the whole of the evidence, and a conformance regression would not have been caught here. |
||
|
|
9c0144d24a |
[Test] (MG_Benchmark, MG_Util, MG_Backend, android-plugin): run the driver benchmark on a phone
The Minecraft-shaped driver benchmark could only be run from a desktop shell against a desktop driver, which is the wrong machine: MobileGL exists to run on mobile GPUs, and nothing said what its translation costs there. This puts the same cases on an Android device, both in the plugin's POST screen and from a shell, and adds the native-driver baseline they have to be read against. The cases move into DriverBenchCases.inc so both harnesses run byte-identical bodies - the desktop program resolving entry points from one EGL provider, and DriverBenchJni.cpp calling MobileGL's frontend in-process. The JNI file binds every gl*/egl* name to MG_Impl by macro rather than by linkage: this library legitimately has the platform libEGL and libGLESv3 in its own lookup scope, and a benchmark that quietly measured the device driver instead of the translation layer would have looked like very good news. Frames are now closed with a fence wait instead of glFinish. MobileGL implements glFinish and glFlush as no-ops, so the old loop timed submit-plus-GPU on a native driver and submit-only on a MobileGL backend, and the two numbers did not describe the same work. To measure a device's own driver the cases needed to be expressible in GLES: ESSL 3.20 twins of the four shaders (chosen at runtime from GL_VERSION, since MobileGL is deliberately still fed desktop GLSL - translating it is the thing under test), a multi-draw hook that loops DrawElementsBaseVertex where the multi-draw entry point does not exist, and an EGL bootstrap that falls back from desktop GL to GLES 3. The binary cross-compiles for arm64 unchanged. BenchService hosts each run in its own process and exits afterwards. That is not caution: the backend is latched from MOBILEGL_BACKEND_TYPE at initialization, so Espryt and Magma can never share a process, and Espryt's teardown terminates the process-default EGL display, which would take the POST activity's own EGL objects with it. Running it found that Magma could not create a windowless context on Mali at all - CreateInstance required VK_EXT_headless_surface, which no mobile driver here exposes, and aborted the process. The Xlib path already probes and falls back to a hidden window for the same reason on NVIDIA; Android now probes too and hands the WSI an AImageReader's ANativeWindow, a real producer surface attached to no display whose images are never acquired. DriverPost reports the extension's absence as a WARN so the fallback is visible rather than silent. Measured on a Mali-G77 MC9 (native / Espryt / Magma, ns per operation): 5495 chunk draws 14397 / 36934 / 33763, the 26.2 per-draw uniform-range pattern 13710 / 31205 / 21252, sodium-style multi-draw 256956 / 238389 / 209527. The translation costs about 2.4x per draw here against 5-9x on the desktop, because the mobile driver's own per-call cost dwarfs it - and both backends beat the native driver on multi-draw, which it has to emulate. Desktop unit tests 421/421; the POST screen and both Run Bench buttons verified on the device. |
||
|
|
1e45958e01 |
[Test] (MG_Benchmark): measure the driver work a real Minecraft frame asks for
The benchmark tree had nothing that exercised a driver: SanityBench times std::vector, and the Buffer/Program benches call into MobileGL_s directly, so neither can say what a backend costs against the native driver. This adds a headless EGL client that can, and shapes its cases from measured traces rather than guesses. DriverBench dlopens exactly one EGL provider - the system libEGL.so.1, or a libMobileGL.so with MOBILEGL_BACKEND_TYPE selecting Espryt or Magma - so the same binary measures all three stacks with no LD_LIBRARY_PATH shadowing, which matters because MobileGL's own loader has to keep finding the real driver underneath. It renders into its own renderbuffer FBO on a 64x64 pbuffer and paces frames with glFinish, so it needs no window and no compositor. The six mc_* cases replay the per-frame call mix of 30-second render-distance-32 captures of three Minecraft versions, at the rates those captures measured: vanilla 1.21.1 issues 5495 glDrawElements per frame, each preceded by its own glBindVertexArray and glUniform3fv; Fabric+Sodium collapses the same scene into 132 glMultiDrawElementsBaseVertex; the 26.2 snapshot issues 3401 glDrawElementsBaseVertex, each preceded by glBindBufferRange + glBindBuffer. The texture case wraps every 16x16 atlas upload in the four glPixelStorei and two glTexParameteri calls Blaze3D re-sets around it, because that wrapper is a large part of what an upload costs a translation layer. One bench frame therefore costs what one real frame of that version costs, and ns_per_op is directly comparable across renderers. run_driver_bench.sh pins __EGL_VENDOR_LIBRARY_FILENAMES and VK_ICD_FILENAMES. Without that, eglGetDisplay(EGL_DEFAULT_DISPLAY) on this glvnd system resolves to Mesa llvmpipe and the "native" numbers silently describe a software rasteriser - the first run of this bench reported 11 us per draw before the pin, versus 250 ns on the real GPU. Verified against the NVIDIA 610.43.03 driver, Espryt and Magma on a GTX 1660 SUPER; the CMake target builds and runs from a clean configure. |
||
|
|
6e6f5268fb |
[Fix] (MG_Backend): let a default-visual X11 window match an alpha-free config
ChooseConfigForSurface prefilters candidate configs with eglChooseConfig requiring EGL_ALPHA_SIZE 8, then tries to match the window's X visual. On NVIDIA's X11 EGL every alpha-8 config lives on the 32-bit ARGB visual, and the default depth-24 TrueColor visual only appears on alpha-0 configs - so for any window created with the default visual the match loop scanned a list that could not contain its visual, fell through to a 32-bit-visual config, and eglCreateWindowSurface failed with EGL_BAD_CONFIG. Keep the alpha-8 list as the first tier and add an alpha-relaxed second tier used only for the visual match; the sizeless fallbacks below still run on the alpha-8 list. Mesa is unaffected (its default-visual configs carry alpha), and a destination-alpha-free default framebuffer is exactly what native GLX hands out on these visuals anyway. Found by running Minecraft through the new GLXImpl on Espryt: NVIDIA EGL also needs EGL_PLATFORM=x11 under a Wayland session or eglGetDisplay itself returns no display, which is a launcher-environment concern, not a library one. |
||
|
|
08f98ad9ce |
[Feat] (MG_Impl): implement GLX 1.4 on the EGL layer so GLFW apps run on Linux
Desktop Linux GL apps (GLFW/LWJGL, glxgears, anything X11) create contexts through GLX, and MobileGL only spoke EGL - the two exported glX symbols were proc-address stubs that could resolve GL entry points but never produce a context. GLXImpl is the missing sibling of WGLImpl/CGLImpl: the same window-system-binding pattern, calling the internal MG_Impl::EGLImpl namespace directly. The surface covers exactly what GLFW 3.4 resolves via dlsym plus the legacy visual API: FBConfig enumeration mirrors the two EGLState configs (stencil-8 first so stencil-wanting choosers land on it), glXGetVisualFromFBConfig answers with the screen's default visual (falling back to any 24-bit TrueColor one), and glXCreateContextAttribsARB maps the ARB attribs onto EGL context attribs the way WGL's Ext_CreateContextAttribsARB does - profile mask only emitted for 3.2+ or an explicit profile request, since that bit is what keys MobileGL's relaxed-semantics compatibility mode. Legacy glXCreateContext/CreateNewContext hand out 3.3 compatibility contexts, matching wglCreateContext. Drawables follow the WGL HWND model: the GLXWindow is the X window itself, the EGL window surface is created lazily on first MakeCurrent and cached per XID, and the GLX layer owns size discovery per the platform-layer contract - it pushes changes through EGLImpl::ResizePlatformWindowSurface, polling XGetGeometry on MakeCurrent and on swaps throttled to 250ms so a fast-swapping app is not paying a server round trip per frame. libX11 is dlopen'd at runtime like everywhere else in the tree; Xlib.h is already in every TU via the vulkan include, so XVisualInfo gets an ABI mirror struct (Xutil.h needs the Bool and Status macros that Includes.h deliberately pops) and the caller's XFree pairs with our malloc. glXGetProcAddress now resolves glX names from the export table before falling through to the shared GL resolver, which previously returned nullptr for every glX extension entry point - GLFW requires glXCreateContextAttribsARB and glXSwapIntervalEXT to arrive that way. Verified with a smoke test replaying GLFW's exact call sequence (dlsym-only resolution, manual FBConfig filtering, 3.2 core forward-compatible context, glXCreateWindow, 60 swapped frames, clean glGetError) on both backends against the real NVIDIA driver, then with Minecraft 1.21.1, 1.21.4+Fabric+Sodium and 26.2-snapshot-6 reaching in-world rendering on both Espryt and Magma. |
||
|
|
d39a706d57 |
[Perf] (MG_Backend): stop paying for descriptor slots and mip barriers nobody asked for
Five independent bits of per-draw and per-operation waste in the DirectVulkan backend, all removing work whose answer was already known. The per-draw descriptor walk iterated all 256 slots of bindingKinds to find the one to eight bindings a real GL program declares, because that vector is sized to the binding cap rather than to the program. Reflection now records the bindings it actually assigned, and the draw path iterates that. It is built at the end of ReflectLayout, not where bindingKinds is sized - at that point the vector is only zero-initialised and the kinds are assigned further down, so a list built there would be empty. It has to stay ascending: Vulkan consumes pDynamicOffsets in binding order and the writer pushes them in iteration order, so an unordered list would silently mis-pair dynamic offsets with their uniform blocks. Descriptor pools were sized maxSets * the 256-binding cap, declaring 81,920 descriptors per pool and 245,760 across the frames in flight, for sets that hold what shader reflection found. Sized from eight now; an outlier program is absorbed by the VK_ERROR_OUT_OF_POOL_MEMORY path that already exists, which works because pool sizes are aggregate budgets rather than per-set limits. TrackLiveResource swept the whole live-buffer vector on every insert once it passed 256 entries, and when the buffers are all live the sweep removes nothing and the vector grows by one - so creating N live buffers cost about N^2/2 expired() checks. It sweeps on a doubling watermark now, with the same reclamation semantics. GenerateMipmap transitioned each destination level individually inside its loop, but every generated level starts in the same layout and the loop only moves a level out of TRANSFER_DST after writing it, so the whole range can be prepared in one barrier - 3(N-1)+1 barrier commands become 2(N-1)+2. Each level is still transitioned to TRANSFER_SRC before it is read, so the dependency between consecutive levels is unchanged. WaitForFrameSerial drained the entire graphics queue, as its own comment admitted. Every submission records the frame serial it was made under, so it now waits on the first fence at or past the requested serial. The narrow path deliberately does not call NotifyDeviceIdle(): that claims every submission has retired, which is only true after a real drain, so it stays on the fallback. Verified with an 8213-case A/B (textures, buffers, queries, mipmaps, uniforms and the whole direct_state_access suite): the Espryt failure list is identical, the Magma failure list differs by one case, and both crash sets are unchanged on Magma. That one case, buffer_storage.map_persistent_draw, does not reproduce in isolation - running the buffer_storage group alone gives byte-identical results on both builds (the same three failures, not including it), and it reports NotSupported when run on its own. It is the same ordering-dependent behaviour this suite shows elsewhere, and the three Espryt crash-set differences are the known copy_image cluster moving chunk position. Flagging rather than hiding it. direct_state_access stays at Espryt 370/371 and Magma 371/371; unit tests 421/421. |
||
|
|
f3d52faad4 |
[Perf] (MG_State, MG_Backend): stop glViewport from evicting a cached VkPipeline
RenderState kept one version counter for all render state, and DirectVulkan read it in three places: the pipeline memo key, the SetupDrawSnapshot fast-path guard, and that guard's store. So glViewport, glScissor, glBlendColor, glStencilMask, glClearColor, glPolygonOffset, glLineWidth and the point-size family - none of which can alter a VkPipeline, all of which an application changes between draws - knocked the next draw off both fast paths and made it rebuild a pipeline lookup that was already correct. The counter is now split. m_version still moves on every state change, because the draw snapshot really does depend on all of it. m_pipelineStateVersion moves only for the state a backend bakes into a pipeline object, and it is what the three DirectVulkan sites read. The exclusion list is the eight VkDynamicState entries PipelineFactory declares plus the state that is not pipeline state at all (the clear values, hints, the point-size family, clamp read colour, the primitive restart index). glStencilFunc is the one setter that had to be split rather than classified: Func is in the pipeline payload but Ref and ValueMask are dynamic state, so it bumps the pipeline version only when Func actually changes. Capabilities are deliberately NOT in the exclusion list even though several look like dynamic state: GL_FRAMEBUFFER_SRGB feeds the render-pass hash, depth and stencil test feed drawUsesDepthStencil, and scissor test, blend, cull face, polygon offset fill, primitive restart, colour logic op and rasterizer discard all feed the pipeline payload. Two smaller draw-path wins ride along, both removing work whose answer was already in hand. UploadAndBindVertexStreams searched all 32 VAO attribute slots for the SharedPtr matching a binding's buffer key, once per binding per draw - but VertexInputStateFactory writes bindingBufferKeys[b] and bindingAttributeLocations[b] from the same loop iteration, one binding per attribute with no merging, so the attribute at that location IS the buffer, by construction. UploadAndBindIndexBuffer round-tripped the element-array buffer's raw pointer back through the GL name table on every indexed draw, costing a map lookup and an atomic refcount pair, when the binding slot's SharedPtr was already in scope forty lines above - where a comment says exactly that about the vertex path. Behaviour-neutral by construction and verified as such: a 13355-case subset of GL30-GL45 covering viewport, scissor, blend, stencil, depth, polygon offset, clear, multisample, cull, logic op, line width and point state, plus the whole direct_state_access suite, is identical before and after on both backends - in the failure list and in the crashed-case set. direct_state_access stays at Espryt 370/371 and Magma 371/371. |
||
|
|
ba81ee114e |
[Feat] (MG_Backend, MG_Impl, MG_Util): attach one layer of any layered texture on DirectVulkan
Whether a backend can attach a single layer of a texture to a framebuffer was one Bool, so it could only give the most conservative answer any target needed. DirectVulkan therefore declined every layer of every target and direct_state_access.framebuffers_texture_layer_attachment failed with 542 messages across four targets. The three ways a GL layer maps onto Vulkan are independent capabilities, so the flag becomes a per-TextureTarget mask. A 2D or 2D multisample array layer IS a VkImage array layer and needed nothing but the gate opened. A cube map array is one 2D image with arrayLayers = 6 * cubeCount and CUBE_COMPATIBLE, which is a shape VkTextureManager simply did not have - it is declined softly when the depth is not a whole number of cubes or the level is not square, because that function's Bool return exists for unrepresentable shapes and asserting there would abort on ordinary input, GL_PROXY_TEXTURE_CUBE_MAP_ARRAY above all. A 3D texture's layer is a z slice, which needs a 2D-array-compatible image and a per-slice clear, because vkCmdClearColorImage cannot address a subset of a 3D image's slices - a render pass whose only content is its LOAD_OP_CLEAR can, since its attachment is a 2D view over that one slice. VK_IMAGE_CREATE_2D_ARRAY_COMPATIBLE_BIT is asked for per format and withdrawn per format, mirroring the MUTABLE_FORMAT pattern already in this file: the capability is per format+usage, so a single global probe answers a different question than the one the frontend goes on to ask. Losing it costs per-slice attachment for that format; failing creation would lose the texture. Three things found on the way that are not the headline: glFramebufferTextureLayer, the non-DSA twin, had no gate at all and additionally refused cube map arrays that GL 4.5 requires it to accept. GL 4.6 core 9.2.8 makes the two entry points equivalent, so they now decline in the same places - leaving one ungated is what let an unrepresentable attachment reach the renderer. ComputeFullMipLevelCount takes max(x, y, z), and for every array shape z is the layer count rather than a mip-able axis, so a 4x4 array with 192 layers asked for six mip levels on an image whose legal maximum is three (VUID-VkImageCreateInfo-mipLevels-00958). Only the image's own extent can bound it. lavapipe had been letting that through. A layered GL clear queues layerCount = depth, which is illegal for a VK_IMAGE_TYPE_3D image (VUID-vkCmdClearColorImage-baseArrayLayer-01472 pins it to 0/1, read as the whole mip level) and the old code passed it straight through. Takes framebuffers_texture_layer_attachment green on DirectVulkan, so the whole direct_state_access suite is 371/371 there; Espryt stays 370/371, the remaining case being the fp64 one it declines by design. Known and deliberately not fixed here, with a FIXME at the site: KHR-GL44/45/46.geometry_shader.layered_framebuffer.clear_call_support now fails on DirectVulkan - a layered clear of a 3D texture reads back zeros. Those cases exist only in the GL44+ lists, above the 4.0 this backend reports. An A/B of a 6935-case subset (cube map array, texture storage, framebuffer, 3D, the full DSA suite and the GL33 texture group) is otherwise clean on both backends: 16 cases fixed and none broken on Espryt, 15 fixed and those 2 broken on Magma, and zero difference anywhere at GL 4.0 or below. The FIXME records which causes were already ruled out by bisection so the next reader does not repeat them. |
||
|
|
c8c7b19579 |
[Feat] (MG_Backend, MG_Util): give DirectVulkan GL's provoking vertex
Vulkan's built-in convention is "provoking vertex first"; GL's default is LAST_VERTEX_CONVENTION, and GL derives both flat shading and the transform feedback vertex order from it. DirectVulkan had no way to say so, which is why direct_state_access.queries_functional failed on a value with nothing in its log - the primitives came back counted against a strip recorded in the wrong vertex order. VK_EXT_provoking_vertex is now enabled when present, and the mode is a hashed field of the pipeline payload rather than dynamic state, because it is baked into VkPipelineRasterizationStateCreateInfo: two draws differing only in it must not collide on one cached VkPipeline, or whichever mode built first would stick for the rest of the frame. The pNext is chained only when the mode is not Vulkan's default, so a device without the extension produces a byte-identical VkGraphicsPipelineCreateInfo to before. Two carve-outs, both measured rather than reasoned: A geometry shader already emits its triangles in GL's vertex order, so asking for LAST rotates them a second time and transform_feedback.geometry reads back the wrong vertices. The mode is one pipeline bit and the input-assembler path wants the opposite, so the two cannot both be satisfied: a program that runs a geometry shader and captures transform feedback keeps Vulkan's own convention. That test is read off the program's own shader list, not programObj.rasterizationProducerStage - the latter is filled by the clip-fixup analysis, which does not run for every program and reads Unknown for exactly the programs this guard exists to catch. Both halves are link-time facts folded into programObj.hash, so no pipeline memo can hand back one built for the other mode; keying on IsTransformFeedbackActive() instead would be a live bug, since neither memo key moves on glBeginTransformFeedback. transformFeedbackPreservesProvokingVertex is deliberately not requested. It buys nothing here - the capture order queries_functional needs comes from provokingVertexLast alone - and leaving it off keeps VUID-VkGraphicsPipelineCreateInfo-topology-04884 disarmed, so a TRIANGLE_FAN pipeline may take LAST on any device. The blit pipeline routes through the same selector: it has no flat varying and no capture, but on a device without provokingVertexModePerPipeline a blit left on FIRST inside a render pass whose draws are LAST is an illegal mix. Per the POST rule the new extension gets rows for provokingVertexLast and for the two properties that change what MobileGL can promise. Fixes queries_functional on Magma (370/371). An A/B over a 976-case transform feedback / geometry shader / layered rendering subset of GL30-GL45 is otherwise identical on both backends and additionally takes 14 geometry_shader rendering and layered_rendering cases from failing to passing on Magma. |
||
|
|
0e7692251d |
[Feat] (MG_State, MG_Impl, MG_Util): store a compressed texture image and hand it back
glCompressedTexImage2D rejected every internalformat with GL_INVALID_ENUM, so direct_state_access.textures_get_image threw at its first compressed call and reported InternalError with nothing in the log at all - the uncompressed half of the case had already passed. The compressed bytes are now kept verbatim, in a side-channel beside the texel shadow rather than in place of it. That placement is the load-bearing decision: both backends pair MapMipmapData with GetMipmapByteSize while sizing their copy regions from GetMipmapTexelSize, and DirectGLES additionally divides the byte size by the texel count to recover bytes-per-texel, so putting 16 bytes where a 4x4 RGBA8 extent says 64 would be an out-of-bounds read on both. The texel storage therefore stays uncompressed and correctly sized - the image samples as zeros, which is the same deviation the RGTC/BPTC/ETC2 arms of ConvertGLEnumToTextureInternalFormat already document - while glGetCompressedTexImage returns the image *as stored*, which GL 4.6 core 8.11 requires and which no re-encode could satisfy byte for byte. Nothing ever hands the compressed bytes to GLES or Vulkan, so the shadow is authoritative rather than potentially stale, which is why the readback never asks a backend. The accepted set is exactly the RGTC/BPTC/ETC2-EAC formats core GL requires, and it is deliberately the same set ConvertGLEnumToTextureInternalFormat can back with uncompressed storage, so the upload can never accept a format whose texel shadow it cannot allocate. imageSize is checked against the block arithmetic, which is also what keeps the copy in bounds. Three things the shape depends on. AllocateStorage clears the compressed tag, so a glTexImage2D or glTexStorage2D over the level un-compresses it - without that, textures_compressed_subimage would flip branches and start asking for data MobileGL cannot produce. GL_TEXTURE_COMPRESSED and GL_TEXTURE_COMPRESSED_IMAGE_SIZE are answered per level rather than per texture, because a compressed internalformat handed to glTexImage2D resolves to uncompressed storage and must keep reading as uncompressed. And GL_TEXTURE_INTERNAL_FORMAT now reports the compressed token for such a level, or it would claim GL_RGBA8 while GL_TEXTURE_COMPRESSED said true. Still rejected on purpose: glCompressedTexImage1D/3D and every glCompressedTexSubImage*, which caps the blast radius. Fixes textures_get_image on both backends (Espryt 370/371, Magma 369/371). A/B over a 1210-case compressed/texture-storage/texture-view/buffer-storage subset of KHR-GL45 is identical before and after on both backends but for get_texture_sub_image.errors_test, which stops throwing and fails on a value instead. |
||
|
|
34f09291da |
[Feat] (MG_State, MG_Backend, MG_Util): feed a 64-bit vertex attribute on DirectVulkan
glVertexAttribLFormat validated its arguments and then refused unconditionally with "64-bit vertex attributes are not supported", so direct_state_access.vertex_arrays_attribute_format failed every GL_DOUBLE subcase on both backends - the format never landed, the draw fetched whatever the attribute held before, and the captured values came back as reinterpreted garbage. The attribute is now real state. IsLong is its own bit rather than being inferred from Float64, because glVertexAttribFormat(GL_DOUBLE) also reads doubles - it just asks for them converted to float - so the type alone cannot tell the two apart. It participates in the format comparison, so an L-format call over a plain one still bumps the version, and glVertexAttribPointer clears it inside the mutation block so the clear and the bump stay atomic. GL_VERTEX_ATTRIB_ARRAY_LONG stops being hardcoded false, and the pname is now accepted by the attribute queries at all. Support is detected, never assumed. SupportsFloat64VertexAttributes comes from VkPhysicalDeviceFeatures::shaderFloat64 on DirectVulkan and is false on DirectGLES - not a driver question there and never will be, since ES has no GL_DOUBLE vertex format and ESSL has no fp64 type to consume one with. A backend without it declines in the entry point, with the GL error and a log line naming the reason, rather than accepting state no draw could honour. Both cases get a DriverPost row so the loss is named at startup instead of at draw setup. On DirectVulkan the attribute deliberately does not use VK_FORMAT_R64*_SFLOAT: those are optional and lavapipe advertises zero features for all four of them. It is fetched as its 32-bit word pair (R32G32_UINT / R32G32B32A32_UINT) and bitcast back to double in the shader by a new SPIR-V pass, which is bit-exact and needs no format capability at all. The pass re-declares the input as uvec2 / uvec4, demotes the original variable to a Private global and seeds it once at the top of the entry point, so every existing load keeps its id and its double type and no other instruction is rewritten. Both halves branch on nothing but "is this attribute long", so they cannot disagree - and if the pass ever fails, the assertion fires rather than letting a UINT format sit under a double input. The pointer types are all created before any variable that names them and the demoted variable is moved after them, since the types-and-variables section may not forward-reference a type. dvec3/dvec4 are declined rather than fetched wrong: six or eight uint32 components have no single VkFormat, and GL spreads such an input over two attribute locations, which the location-per-index model here does not express. Fixes vertex_arrays_attribute_format on Magma (369/371). On Espryt it stays failing, now as a detected and explained decline rather than a blanket refusal. |
||
|
|
3b65e646e1 |
[Fix] (MG_Backend): give every colour attachment its own backend slot on DirectGLES
ES only accepts glDrawBuffers bufs[s] == GL_COLOR_ATTACHMENTs, so a desktop glDrawBuffer(GL_COLOR_ATTACHMENT3) cannot be expressed directly and DirectGLES compacts: it physically relocates the draw buffer's image onto backend point 0 so ES's output-0-to-attachment-0 rule lands on the right image. The clears were therefore always correct. The read side was not. GetBackendAttachmentType derived the attachment-to-point map by searching the draw-buffer array and falling back to the identity point for anything it did not find. That derivation is not injective against the compaction: after clearing attachments 0..7 one at a time, every one of them has been relocated onto point 0 in turn, so a later glReadBuffer(GL_COLOR_ATTACHMENT0) - not a draw buffer any more - takes the identity fallback to point 0 and reads attachment 7's image. Hence the single mismatch, 0.875 where 0 was expected: 7/8 is attachment 7's clear colour. The map is now stored state rather than a re-derivation, and kept a permutation: a draw buffer takes the point ES forces on it, everything else keeps its identity point when that point survived, and an attachment evicted from its identity point is parked on the lowest free one so it stays addressable for glReadBuffer and blits. With identity draw buffers nothing moves and not one extra GL call is issued, which is what keeps ordinary rendering untouched. Two things the permutation depends on. The attachment loop now detaches a colour point whose frontend owner is empty - SyncAttachmentObject only ever attaches, so without this a point handed to an empty attachment would still hold the previous owner's image and hand it back. And QueryReadColorAttachmentInternalFormat asked GL_COLOR_ATTACHMENT0 for the format it sizes the multisample-resolve scratch renderbuffer from; it now asks the point the read buffer actually names, since that is only CA0 when the map happens to be identity. Fixes framebuffers_read_draw_buffer on Espryt. A 5677-case readback and framebuffer subset of GL30-33 stays at zero failures on both backends. |
||
|
|
25b9370815 |
[Fix] (MG_Backend): stop a renderbuffer blit reading a freed image layout
VkRenderPassManager kept m_renderbufferResources on FastSTL's open-addressing UnorderedMap while BlitFramebuffer caches a raw pointer into one of its elements - ResolveColorBlitBinding stores &rbResource->layout - and then calls MaterializePendingClearForRenderbuffer, which looks that same resource up again. FastSTL's operator[] runs its load-factor check before find_key and reallocates the whole bucket array when occupancy crosses it, so even a plain lookup relocates every element; erase only tombstones and never lowers the occupancy, so the doubling keeps firing. After a relocation the cached pointer names freed storage still holding the pre-clear VK_IMAGE_LAYOUT_UNDEFINED, BlitFramebuffer takes its "source image layout is undefined" early return, and the blit is silently dropped - glReadPixels then returns the zero-filled fresh allocation. That is why the failures looked arbitrary: which iteration breaks is pure arithmetic on the table's occupancy, and the observed set (GL_R8 at k=0,1,3,7, GL_R16 at k=6, GL_RG16 at k=4) is exactly the doubling ladder. Padding the map with unrelated live renderbuffers moves the failures to the positions the model predicts and every previously failing format then passes, so nothing else hides behind it. Reordering the materialize ahead of the resolves - the fix ReadPixels got, see the note at its call site - does not cover this, because BlitFramebuffer resolves two bindings and the second resolve still runs after the first pointer is taken. The depth blit, GetOrCreateRenderPass's depthRenderbufferResource and ReadDepthStencilPixels cache the same kind of pointer, so the invariant belongs in the container rather than in a per-call-site ordering rule. m_textureResources was already node-based for exactly this reason; this is the map that was left behind. Fixes renderbuffers_storage_multisample on DirectVulkan. |
||
|
|
4ce808b9f2 |
[Feat] (MG_State, MG_Impl, MG_Backend): let a bound program pipeline actually draw
The pipeline object bookkeeping landed already - names, stage slots, queries - but nothing consumed it. Every draw asked the context for the current program, got null because a pipeline is used with program zero, and drew nothing; glCreateShaderProgramv was still a stub returning zero, so direct_state_access.program_pipelines_functional could not even build its stage programs and reported InternalError on both backends. glCreateShaderProgramv is written as the exact call sequence the spec defines it to be, with one deviation that matters: the link goes straight to ProgramObject::Link(false) rather than through LinkProgram, because LinkProgram injects a default fragment shader into a program that has none - correct for a whole program, wrong for a separable vertex-stage one whose fragment stage comes from the pipeline. glDetachShader defers removal to the next link, so the program keeps the shader object it was built from while correctly no longer reporting it attached. GL_PROGRAM_SEPARABLE joins glProgramParameteri and glGetProgramiv. Everything downstream of a draw - both backends, the uniform plumbing, the draw validation - is written against one linked program, so rather than teach all of it about stages, the pipeline is flattened: GetProgramForDraw() composites the stage programs' shaders into a single hidden program object and caches it against a signature of each stage program's lifetime id and link generation, so it is rebuilt exactly when a stage or a stage's link changes. The composite carries no GL name - it must not answer glIsProgram, and it must not consume a name the application could be handed. Uniform entry points get their own resolver rather than sharing that one: glUniform* addresses the pipeline's active program, not the composited draw program. GL_CURRENT_PROGRAM still reads the program in use, which is zero here. Fixes program_pipelines_functional on both backends. |
||
|
|
5545d31c37 |
[Feat] (MG_Backend, MG_Util): give a cube map array real storage on DirectGLES
TextureCubeMapArray was missing from every storage and upload switch in the DirectGLES texture sync, so a cube map array reached the driver with no storage at all - and from the glFramebufferTextureLayer branch, so attaching one of its layers fell through to glFramebufferTexture2D and raised INVALID_ENUM. Every GL_TEXTURE_CUBE_MAP_ARRAY colour check in direct_state_access.framebuffers_texture_layer_attachment read nothing. ES 3.2 has GL_TEXTURE_CUBE_MAP_ARRAY natively and it stores exactly like a 2D array whose depth is six times the cube count, so each switch gains the case beside Texture2DArray and nothing else changes. 1D arrays join the layer branch for the same reason - their backend image is a 2D array. Per the POST rule the new GLES dependency gets a capability (SupportsTextureCubeMapArray, ES 3.2 core or EXT/OES_texture_cube_map_array) and a DriverPost row saying what a user loses without it. Takes framebuffers_texture_layer_attachment from failing to passing on Espryt. It still fails on DirectVulkan, which declines a layered attachment outright. |
||
|
|
588ddba722 |
[Fix] (MG_Backend): scale a depth blit, keep going after one declines, and mip a 1D texture
Three DirectVulkan gaps found together.
glBlitFramebuffer's depth/stencil path refused any blit whose source and
destination extents differ, because vkCmdCopyImage cannot resize. vkCmdBlitImage
can, and VK_FILTER_NEAREST is the only filter Vulkan allows for depth/stencil
anyway - which is what the GL front end already requires. A same-size pair keeps
the cheaper copy.
Worse, that refusal and four others were `return`, not `continue`, so a
depth/stencil aspect this backend could not handle abandoned the whole function -
including the colour blit that only starts after the aspect loop. The CTS's
scaling blits therefore lost their colour as well, which is why
direct_state_access.framebuffers_blit failed all three of its checks rather than
one.
VulkanRenderer::GenerateMipmap declined GL_TEXTURE_1D. It needed nothing else:
the blit loop derives every offset from the storage extent, and a 1D texture's is
{width, 1, 1}, which is exactly the y and z offsets a 1D image requires.
Also: IsTimerQueryResultReady now asks the query pool before the frame serial.
The pool polls with VK_QUERY_RESULT_WITH_AVAILABILITY_BIT and is the authority;
the frame serial only advances at Present and neither completion notifier will
mark the current serial done, so a timestamp written and fence-waited inside one
GL frame could never be read back within it.
Takes framebuffers_blit and textures_generate_mipmaps from failing to passing on
DirectVulkan. queries_functional still fails there on a value.
|
||
|
|
62301b1061 |
[Fix] (MG_State): let a double-typed varying be captured by transform feedback
ResolveXfbSymbolType accepted only float, int and uint, and its caller reports anything it rejects as "Transform feedback varying 'x' is not an output of the vertex stage" - which is a misleading thing to say about a varying that is right there in the shader, just declared `double`. Program linkage failed outright. Doubles are now resolved to the GL_DOUBLE* types, in vector and matrix form, and the per-element size is computed from an 8-byte component rather than a hardcoded 4 (GL 4.6 core 11.1.2.1), so the byte-based limit checks charge a double what GL says it costs. direct_state_access.vertex_arrays_attribute_format stops throwing on both backends and fails on the captured values instead: the capture layout still owes the 8-byte alignment doubles require, and neither backend feeds a 64-bit vertex attribute yet - DirectGLES cannot at all, ESSL having no double. |
||
|
|
f3a846d336 |
[Docs] (README): carry the 4.2 short-term target into the status note
The compatibility section already said 4.2; the status note at the top of the README still said 3.3, so the two disagreed depending on how far a reader got. |
||
|
|
9cdc82fbdd |
[Fix] (MG_Backend): actually bind the sampler object DirectGLES just synced
BindCurrentTextures' program-driven path synced a bound sampler object's parameters to its backend object and then never put it on the texture unit, so every sampler object was inert and the driver kept sampling with the texture's own parameters - direct_state_access.samplers_functional read black where the sampler's NEAREST filtering should have given red. The bind alone is a regression, and the CTS says so loudly: a sampler left on a unit by an earlier draw keeps being applied, and a multisample texture takes no sampler object at all, so the next draw against one is rejected and all 27 textures_storage_multisample_3d_* cases fail. The sibling path in the same function had an empty else branch where the unbind belonged; it now unbinds, making the two symmetric. Takes samplers_functional from failing to passing on Espryt, with no other case moving in either direction. |
||
|
|
4a9d20c49f |
[Fix] (MG_Backend): resolve a framebuffer attachment's layer in the Vulkan blit bindings
ResolveAttachmentBaseArrayLayer answered zero for everything but a cube map face, so every blit, copy and glReadPixels against a layered attachment read layer zero whatever was attached. It reads the attachment's layer now. A 3D texture needs the other half of the distinction: its image has arrayLayers == 1 and the GL layer is a z slice, which VkBufferImageCopy will not take as a base array layer. BlitImageBinding carries it separately as depthOffset, and the readback copy region uses it as the image offset's z. Takes textures_copy from failing to passing on DirectVulkan, which is what glCopyTextureSubImage3D needs to see the slice the CTS attached rather than slice zero. |
||
|
|
394d1ce748 |
[Feat] (MG_Impl, MG_Util): copy into 1D and 3D textures, and accept the BPTC and ETC2 enums
Two unrelated texture gaps. glCopyTextureSubImage1D and 3D validated their arguments and then did nothing: CopyTexSubImage1D_State and CopyTexSubImage3D_State were empty TODOs and no backend exposes anything but a 2D blit. But a texture's contents live in its CPU storage - the backends sync from it - so the copy does not need a blit at all. CopyReadFramebufferIntoMipmapRegion reads the region out of the read framebuffer through the existing ReadPixels path, in the destination's own canonical client layout so the bytes need no second conversion, and writes them straight into the level. GL 4.6 core 8.6 says the copy ignores pixel-store state and any bound pack buffer, which the borrowed readback does not, so both are neutralised for the duration and restored after. A cube map destination addresses its faces as separate upload targets, so its zoffset picks the target rather than a slice. ConvertGLEnumToTextureInternalFormat had arms for the six generic compressed formats and the four RGTC ones, all resolving to uncompressed storage, but none for BPTC or ETC2/EAC - so glTexImage2D with one of those fourteen enums answered INVALID_ENUM, which was never a legal reply for formats core GL has required since 4.2 and 4.3. They follow the same deviation for the same reason: nothing in this stack can compress them, and uncompressed storage is the trade the RGTC formats already take. Takes textures_compressed_subimage from failing to passing on both backends and textures_copy on Espryt. textures_copy still fails on Magma, where the readback of a layered attachment does not yet resolve the attached layer. |
||
|
|
300b458132 |
[Feat] (MG_State, MG_Impl): give program pipelines their object and their state
Every program pipeline entry point was an export stub, and the stub macro's `return (type)1` made glIsProgramPipeline answer GL_TRUE for anything - including the names glGenProgramPipelines had never written. All four direct_state_access.program_pipelines cases failed. ProgramPipelineObject holds what GL 4.6 core 7.4 says a pipeline is: a program reference per shader stage, the active program glProgramUniform* addresses, a validate status and an info log. Its validate status starts false, unlike ProgramObject's, because a pipeline that has never been validated must report GL_VALIDATE_STATUS as 0. The name rules follow the shape queries and transform feedbacks already use, and which the CTS checks first: glGenProgramPipelines only RESERVES a name and glIsProgramPipeline answers GL_FALSE for it; the object appears on first bind, or immediately from glCreateProgramPipelines. Map membership is object existence - a pipeline, unlike a transform feedback, has no stateful default object zero, so no everBound flag is needed. glGet(GL_PROGRAM_PIPELINE_BINDING) reports the real binding now instead of a hardcoded zero whose comment said the entry points were stubbed. This is the state half only. program_pipelines_functional needs mixed-stage rendering - a vertex-only and a fragment-only program drawn together - and stays failing; glCreateShaderProgramv is deliberately left stubbed until that lands, so nothing can half-work in between. Takes program_pipelines_creation, _defaults and _errors from failing to passing on both backends. |
||
|
|
1f1a331a44 |
[Feat] (MG_Impl, MG_State): implement the framebuffer parameter getters and setters
glFramebufferParameteri, glGetFramebufferParameteriv and their two by-name siblings were all export stubs - the GL_ARB_framebuffer_no_attachments entry points. The stub raises no error and writes nothing, so direct_state_access.framebuffers_get_parameter_errors saw GL_NO_ERROR for all three conditions it checks. FramebufferObject gains the five DEFAULT_* parameters as real state, initialised to GL 4.6 core table 23.24 and bumping the object version on a write like the read buffer does. The getter answers those plus the six derived names - GL_SAMPLES and GL_SAMPLE_BUFFERS from the attachments' sample counts, GL_IMPLEMENTATION_COLOR_READ_FORMAT/_TYPE from the read buffer's internal format, GL_DOUBLEBUFFER true only for the window-system framebuffer, GL_STEREO false because stereo surfaces are not exposed - which is what glGetIntegerv already reports for the bound framebuffer. The pname rules live in ValidateFramebufferParameterPname, and their ORDER is load-bearing: a name outside the table is INVALID_ENUM, and only a name that IS in the table but that the default framebuffer cannot answer is INVALID_OPERATION. Testing the framebuffer kind first would answer INVALID_ENUM for GL_FRAMEBUFFER_DEFAULT_WIDTH on framebuffer zero, which is exactly the third thing the case checks. The by-name forms take zero as the default framebuffer, like the other DSA framebuffer entry points. Rendering to a framebuffer with no attachments is deliberately NOT enabled by this: CheckCompleteness still reports INCOMPLETE_MISSING_ATTACHMENT, because no backend can rasterize one. The state is real and the queries are honest; the draw path is a separate piece of work. Takes framebuffers_get_parameter_errors from failing to passing on both backends, with framebuffers_get_parameters - which passed only because both getters were stubs leaving the CTS's zero-initialised comparands untouched - still passing. |
||
|
|
817091641c |
[Fix] (MG_Impl): give a cube map the storage and the layered attachment it asks for
direct_state_access.framebuffers_texture_attachment threw on both backends, and three separate things were wrong on the way to a cube map framebuffer. glTexStorage1D/2D/3D validated their target by converting it to a single TextureUploadTarget. GL_TEXTURE_CUBE_MAP has no single upload target - it allocates all six faces - so the conversion produced Unknown and a legal glTexStorage2D(GL_TEXTURE_CUBE_MAP, ...) was rejected with INVALID_ENUM, which is where the case threw. The accepted set for these entry points is the dimension's storage targets, which IsTextureStorageTargetForDimension already spells out, so that is what they check now. TextureStorage2D then allocated only the primary upload target, leaving a cube map with one face out of six - cube-incomplete, so every framebuffer it was attached to answered GL_FRAMEBUFFER_INCOMPLETE_ATTACHMENT. It allocates every upload target the object has; for every other 2D target that is the same single target as before. ResolveRepresentableFramebufferTextureUploadTarget declined every layered target but 2D array, so glNamedFramebufferTexture on a cube map reported "not represented by the current framebuffer attachment model". Cube maps, cube map arrays, 1D arrays, 2D multisample arrays and 3D textures are all the same shape as the 2D array that already worked - glFramebufferTexture binds the whole texture and the attachment records a representative upload target - so they are all handled now. DirectGLES routes a layered attachment to glFramebufferTexture, which is exactly this. Takes framebuffers_texture_attachment from failing to passing on both backends. |
||
|
|
e64c7c7e65 |
[Fix] (MG_Backend): never back a multisample texture with a one-sample Vulkan image
Every one of the sixty direct_state_access.textures_storage_multisample_2d_* and _3d_* cases failed on DirectVulkan, for every internal format, with no GL error anywhere - a pure data mismatch. The CTS asks for glTextureStorage2DMultisample(tex, samples = 1, ...), which is legal GL, and MobileGL carried the 1 faithfully through to VkImageCreateInfo::samples = VK_SAMPLE_COUNT_1_BIT. It then binds that image to the auxiliary program's sampler2DMS, whose SPIR-V is OpTypeImage with MS = 1. VUID-RuntimeSpirv-samples-08726 forbids exactly that pairing: an MS access must come from an image created with more than one sample. The texelFetch therefore read undefined data - which is why it looked format-independent and raised nothing. GL only promises "at least the requested number of samples", so a multisample texture is now floored at two. GL_TEXTURE_SAMPLES still reports what the application asked for; that is read off the texture object, not off the image. The device-capability round below it is bounded at two for the same reason - letting it land back on one sample would recreate the violation silently for any format whose only supported count is one. Takes all 60 textures_storage_multisample_* cases from failing to passing on DirectVulkan, which goes from 296/371 to 356/371. DirectGLES is untouched. |
||
|
|
dd60ff39ce |
[Feat] (MG_State, MG_Impl, MG_Backend, MG_Util): make the border colour real sampler state
glGetSamplerParameterfv(sampler, GL_TEXTURE_BORDER_COLOR) raised INVALID_ENUM,
because MobileGL kept the border colour on the texture object and
GetSamplerParam_State had no case for it at all. That is the first thing
direct_state_access.samplers_defaults asks, so the case threw before reaching
any of the defaults it was written to check.
GL 4.6 core table 23.18 lists TEXTURE_BORDER_COLOR as sampler state, so it moves
to SamplerParameters and TextureObjectBase reaches it through the SamplerObject
it already owns - one source of truth, and a sampler object bound over a texture
now supplies its own border colour, which is what GL says should happen. The
texture params version still moves on a write, because the DirectGLES texture
sync memoises on it. glSamplerParameter{fv,Iiv,Iuiv} and their getters read and
write all four components in whichever representation the caller used, and the
three representations are kept in step so any getter has an answer. The bogus
[0,1] and [0,255] range checks are gone: GL clamps a border colour when a
fixed-point format is sampled, it does not reject it.
DirectVulkan's ResolveVkBorderColor now reads the sampler rather than the
texture. DirectGLES gained a glSamplerParameterfv in its sampler sync, and both
that and the pre-existing glTexParameterfv are gated on a new
SupportsTextureBorderClamp capability - ES 3.2 core, or EXT/OES_texture_border_clamp
before it - since without the extension every such call is INVALID_ENUM on the
driver. DriverPost gains the matching row per the POST rule, saying what a user
actually loses when it is missing.
Takes direct_state_access.samplers_defaults from failing to passing on both
backends.
|
||
|
|
96ad7ca0cc |
[Fix] (MG_Impl): asking a renderbuffer for more samples than it has is INVALID_OPERATION
ValidateRenderbufferStorageSamples_State answered INVALID_VALUE for a sample count above GL_MAX_SAMPLES. GL 4.6 core 9.2.4 reserves INVALID_VALUE for a negative count: a count that is well formed but larger than the format can deliver is INVALID_OPERATION, because the argument is fine and the format is what cannot honour it. Takes direct_state_access.renderbuffers_storage_multisample_errors from failing to passing on both backends. |
||
|
|
e80a23eae6 |
[Fix] (MG_Backend): read a multi-slice glGetTexImage off the GPU instead of the CPU shadow
DirectGLES served every multi-slice glGetTexImage from the CPU shadow copy, on the grounds that its scratch FBO can only expose one layer at a time. But the shadow only holds what was uploaded, so any slice that was rendered to rather than written by glTexSubImage came back stale - and a layered framebuffer produces exactly that. The scratch FBO can expose one layer at a time repeatedly. The read now attaches each layer in turn and takes the slice off the GPU, walking the destination over GL_PACK_SKIP_IMAGES / GL_PACK_IMAGE_HEIGHT itself so each per-slice call packs a plain 2D image with the same layout StoreWideRowsToClient computes for the whole stack. The shadow stays as the fallback for the formats a colour attachment cannot represent at all, and for any slice whose attachment comes back incomplete. Takes all 27 remaining direct_state_access.textures_storage_multisample_3d_* cases from failing to passing on Espryt - they render into a TEXTURE_2D_MULTISAMPLE_ARRAY one layer per colour attachment and then read the whole array back. DirectVulkan is untouched. |
||
|
|
088f263495 |
[Feat] (MG_Impl): answer the two query parameters the getters were missing
GetQueryObjectValue implemented GL_QUERY_RESULT_AVAILABLE and GL_QUERY_RESULT and rejected everything else, so direct_state_access.queries_functional threw on its very first probe - GL_QUERY_TARGET - and never reached any of the checks it was written for. GL_QUERY_TARGET is state the object has carried all along; it just had no case. GL_QUERY_RESULT_NO_WAIT is GL_QUERY_RESULT with the backend asked not to block, and it brings a wrinkle the shared getter could not express: when the result has not landed, GL_ARB_query_buffer_object leaves the destination untouched rather than writing a placeholder. GetQueryObjectValue now reports "succeeded but produced no value" through an optional out-parameter, and all five callers - the four buffer forms and the four client-memory forms - skip the write on it. The switch is deliberately widened by exactly these two names: its default INVALID_ENUM is what the GL33 and GL40 query error cases rely on. queries_functional passes on Espryt. On Magma it stops throwing and fails on a value instead, which is a separate problem in the query results themselves. |
||
|
|
3b3b6e5b8b |
[Fix] (MG_Backend): read back the stencil half, and clear an sRGB target to the value asked for
Two reasons a framebuffer's contents came back wrong, both on the read/clear side rather than the write side. Stencil, on both backends. The CTS reads stencil with glReadPixels(GL_STENCIL_INDEX, GL_INT), which is as legal as the unsigned widths, and neither backend accepted it: DirectGLES's ReadPixelsStencilViaNative rejected every signed type, after which the call fell through to a native ES read the driver refuses and nothing was written at all, so the caller kept its zeros; DirectVulkan's pack switch had no GL_INT case, and of the cases it did have only GL_UNSIGNED_INT sourced the stencil plane - GL_FLOAT and GL_UNSIGNED_SHORT emitted a depth value, which is meaningless for a stencil-only image. Both now take the signed and float widths, and DirectVulkan decides "this is a stencil read" once rather than per type. DirectGLES also gains the GL_FLOAT_32_UNSIGNED_INT_24_8_REV fallback a DEPTH32F_STENCIL8 attachment needs, which rejects the 24_8 packed type. sRGB, on DirectVulkan. Every other write path goes through the UNORM twin view while GL_FRAMEBUFFER_SRGB is off, storing the raw value GL asked for, but a deferred clear is materialised with vkCmdClearColorImage - which names the image, so the driver applied the sRGB transfer function and a clear to 0.25 landed at 0.537. PreCompensateSrgbClearColor hands it the linear colour whose encoding is the requested value instead. It is a no-op for non-sRGB destinations, for integer clear encodings, and when GL_FRAMEBUFFER_SRGB is on and GL really does want the encode. Takes renderbuffers_storage from failing to passing on both backends, plus renderbuffers_storage_multisample and framebuffers_blit on Espryt. |
||
|
|
9eda2147b1 |
[Fix] (MG_Impl, MG_Backend): let the backend that can honour a layered attachment have it
NamedFramebufferTextureLayer declined every attachment but layer zero, on both backends. That was right for DirectVulkan, which maps a GL layer onto a Vulkan array layer with no notion of a 3D depth slice, but wrong for DirectGLES: SyncAttachmentObject already routes a layered upload target to glFramebufferTextureLayer with the attachment's layer passed straight through, and array storage already carries the real layer count into glTexStorage3D. The one backend that could render to the layer was being told it could not. The decision now lives in a DynamicBackendParameters flag, so it is the backend that answers rather than the entry point guessing. DirectGLES sets it when the driver resolved glFramebufferTextureLayer; DirectVulkan leaves it false until VkRenderPassManager tells a depth slice from an array layer. framebuffers_texture_layer_attachment's colour checks now pass on Espryt for 3D, 2D array and 2D multisample array textures - the case still fails there on cube map arrays, which DirectGLES gives no storage at all, and on the depth and stencil halves. No case changes on DirectVulkan, which keeps the old behaviour. |
||
|
|
a63699cde6 |
[Fix] (MG_Impl, MG_Backend): reject incomplete cube maps in mipmap generation instead of crashing on them
Both direct_state_access.textures_generate_mipmap* cases crashed DirectVulkan. Two causes, neither of them a broken invariant: glGenerateMipmap and glGenerateTextureMipmap never checked cube completeness, so an incomplete cube map went straight to the backend, which asserts that the texture it is handed is complete. GL 4.6 core 8.14.4 makes that call INVALID_OPERATION - there is no consistent set of faces to filter down - and both entry points now say so through a shared check. VulkanRenderer::GenerateMipmap asserted that the target was one of the four it implements. 1D, 1D array and cube map array are legal GL and the front end passes them through, so meeting one is a gap in this backend's coverage; it now logs and declines, leaving the generated levels unwritten rather than aborting. textures_generate_mipmap_errors passes on both backends now. textures_generate_mipmaps stops crashing but still fails: DirectVulkan does not generate the 1D mip chain the case checks - the frontend's storage allocation gives the levels the right sizes, which is why the case passes when run on its own, but not the descending content the full-run state leaves it looking for. |
||
|
|
765aaec6dc |
[Fix] (MG_Impl, MG_Backend): stop the new layer attachment from reaching backends that cannot back it
Implementing NamedFramebufferTextureLayer made layered attachments reachable for the first time, and direct_state_access.framebuffers_texture_layer_attachment went from Fail to Crash on DirectVulkan. Two separate gaps sat behind it, both of them asserted on rather than reported: - The renderer resolves an attachment's GL layer straight onto a Vulkan array layer. A 3D texture's z-slice therefore lands outside its image, which has one array layer by construction, and the array texture objects are still the one-image stubs in TextureObjectStubs.h, so their image has a single layer whatever GL believes. MaterializePendingClearForTexture tripped over a clear whose layer span was outside the image it was given. - A cube map array has no image shape in VkTextureManager at all, so SyncTextureAndGetDescriptor returns null for it. NamedFramebufferTextureLayer now answers the full error set for every target and layer - which is what took the two error cases green - and then declines to attach anything but layer zero of a non-cube-array texture, through the same RecordUnsupportedFramebufferTextureAttachmentError the by-target entry point already uses. Layer zero of the other targets is the plain first-slice attachment glFramebufferTextureLayer already backs, so it still goes through. SyncTextureResource's assertion on an unsupported texture shape is also gone: it is a gap in this backend's coverage, not a broken invariant, and the code below it already handles the failure by declining the sync. It logs a warning instead. framebuffers_texture_layer_attachment goes back to Fail on DirectVulkan rather than Crash; no case changes in either direction beyond that. |
||
|
|
bcd669bd25 |
[Feat] (MG_Impl): complete the by-name framebuffer attachment and buffer-selection entry points
Four direct_state_access framebuffer cases failed on one shared cause and three local ones. The shared cause: every DSA framebuffer entry point resolved its name through GetNamedFramebufferObject_State, which rejects zero outright. But zero names the default framebuffer to these functions, so glGetNamedFramebufferAttachmentParameteriv, glNamedFramebufferDrawBuffer(s) and glNamedFramebufferReadBuffer answered INVALID_VALUE for every default-framebuffer query the CTS makes. They now resolve zero to the default framebuffer object and tell the two kinds apart explicitly, which is what the accepted-name rules key off anyway. Attachment queries: the accepted attachment names differ between the default framebuffer (FRONT/BACK variants, DEPTH, STENCIL) and a framebuffer object (COLOR_ATTACHMENTi, DEPTH/STENCIL/DEPTH_STENCIL_ATTACHMENT), and a name outside the relevant list is INVALID_ENUM. Both getters share ResolveAttachmentQueryName for that, so the by-target form no longer aliases GL_FRONT onto a framebuffer object's colour attachment 0. The TEXTURE_* parameters are also rejected with INVALID_ENUM when the attached object is a renderbuffer. Buffer selection: naming a buffer that belongs to the other kind of framebuffer is INVALID_OPERATION, not INVALID_ENUM - the enum is accepted, the framebuffer just has no such buffer. glDrawBuffers additionally rejects the multi-buffer names (FRONT, LEFT, RIGHT, FRONT_AND_BACK) with INVALID_ENUM on both kinds, takes BACK only when n is one, and glReadBuffer treats the multi-buffer names as accepted-but-unselectable. Both colour-attachment range checks now go through ValidateColorAttachmentInRange instead of comparing against MAX_DRAW_BUFFERS with an off-by-one. NamedFramebufferTextureLayer was a stub that reported "not represented by the current framebuffer attachment model" for every call, even though the attachment model stores a layer and the by-target glFramebufferTextureLayer already uses it. It is implemented against the same model, with the per-target layer limits and the INVALID_OPERATION-for-a-bad-name rule that separates it from NamedFramebufferTexture. NamedFramebufferTexture itself gained the two checks it lacked: colour attachment range, and a negative level. Takes framebuffers_get_attachment_parameters, framebuffers_get_attachment_parameter_errors, framebuffers_texture_attachment_errors and framebuffers_draw_read_buffers_errors from failing to passing on both backends. |