[Fix] (MG_State/MG_Impl/MG_Backend): render Flywheel instanced+indirect on both backends

Create 6 / Flywheel 1.0.6 now renders correctly with both flywheel:instancing
and flywheel:indirect on DirectGLES and DirectVulkan (verified in-game on
Adreno 830: waterwheels and cogwheels solid, animated, correct pairing, no
crashes across all four combinations).

- MG_State/MG_Impl: sync explicitly-ranged SSBO bindings of FLUSH_EXPLICIT
  persistent maps to the backend before compute dispatches. Flywheel writes
  its scatter-copy descriptors into the staging ring's persistent map and
  never flushes that span (UB per spec, works on drivers whose maps alias
  GPU-visible memory); our maps alias the CPU shadow, so the descriptors
  never reached the GPU: the scatter compute copied nothing (GLES: empty
  draw commands) or stale garbage (Vulkan: wild indirect commands ending in
  VK_ERROR_DEVICE_LOST).
- MG_Impl/MG_Backend: real glFenceSync objects backed by backend fences
  (GLES: native ES syncs guarded by context generation and owner thread;
  Vulkan: buffer-manager frame serials), replacing always-signaled stubs
  that let Flywheel reclaim staging memory the GPU still reads.
- MG_Backend/DirectGLES: compute dispatches now run the same per-program
  resource sync as draws (uniform-block bindings and sampler units must be
  re-established through the API because layout(binding) is stripped from
  transpiled ESSL) and rebind texture units afterwards; the cull shader
  used to read a stale _FlwFrameUniforms binding and the depth-pyramid
  downsample sampled a stale unit-0 texture, zeroing the Hi-Z pyramid and
  occlusion-culling all Flywheel geometry. Image uniforms are excluded from
  glUniform1i (ES bakes their unit via layout(binding)); image-unit sync is
  clamped to the device limit; eliminated/SSBO-classified uniform blocks
  are skipped.
- MG_Backend/DirectGLES: gl_BaseInstance in native indirect draws reads the
  GPU-written command buffer through an injected mg_IndirectParams SSBO
  view addressed per draw instead of the zero CPU shadow; layout(binding)
  is preserved for SSBO/image declarations (ES has no API rebinding for
  them); the ES context ownership claim moved to a global atomic owner
  thread with an EGL ground-truth check, and deferred buffer op state is
  mutex-guarded, so ops cannot silently no-op after context migration.
- MG_Backend/DirectVulkan: new RebaseInstanceIndexPass rewrites vertex
  InstanceIndex loads to (InstanceIndex - BaseInstance). glslang's relaxed
  Vulkan mode aliases gl_InstanceID to InstanceIndex, which includes
  firstInstance, but GL's gl_InstanceID is zero-based - draws with nonzero
  baseInstance paired meshes with wrong instance data (cogwheel drawn as a
  waterwheel, another wheel collapsed invisible). Gated on the
  shaderDrawParameters device feature. Sampled-read barriers additionally
  cover the compute stage (the Hi-Z downsample samples the depth
  attachment from compute), and short uniform-buffer ranges keep the
  existing zero-padding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-09 06:30:10 +00:00
co-authored by Claude Fable 5
parent 2395a6ded2
commit 139de76347
27 changed files with 973 additions and 130 deletions
@@ -1342,6 +1342,63 @@ namespace MobileGL::MG_Backend::DirectVulkan {
dstY0, dstX1, dstY1, mask, filter);
}
namespace {
// Backend fence handle: the VkBufferManager frame serial captured at
// fence creation. The fence is signaled once every command recorded
// under that serial has completed on the GPU (the same busy-tracking
// horizon used to recycle buffer resources).
struct VulkanSyncObject {
Uint64 frameSerial = 0;
};
} // namespace
BackendSyncHandle FenceSync() {
if (!pVulkanRenderer) {
return nullptr;
}
return new VulkanSyncObject{pVulkanRenderer->GetCurrentFrameSerial()};
}
GLenum ClientWaitSync(BackendSyncHandle handle, GLbitfield flags, GLuint64 timeout) {
// Commands are only submitted at Present, so GL_SYNC_FLUSH_COMMANDS_BIT
// cannot force progress mid-frame; WaitForFrameSerial reports whether
// waiting can succeed at all.
(void)flags;
const auto* sync = static_cast<VulkanSyncObject*>(handle);
if (sync == nullptr || !pVulkanRenderer) {
return GL_ALREADY_SIGNALED;
}
if (pVulkanRenderer->IsFrameSerialComplete(sync->frameSerial)) {
return GL_ALREADY_SIGNALED;
}
if (timeout == 0) {
return GL_TIMEOUT_EXPIRED;
}
return pVulkanRenderer->WaitForFrameSerial(sync->frameSerial, timeout) ? GL_CONDITION_SATISFIED
: GL_TIMEOUT_EXPIRED;
}
void WaitSync(BackendSyncHandle handle, GLbitfield flags, GLuint64 timeout) {
// Server-side waits are implicit: the single graphics queue executes
// submissions in order, so later GPU work already observes everything
// recorded before the fence.
(void)handle;
(void)flags;
(void)timeout;
}
void DeleteSync(BackendSyncHandle handle) {
delete static_cast<VulkanSyncObject*>(handle);
}
Bool GetSyncStatus(BackendSyncHandle handle) {
const auto* sync = static_cast<VulkanSyncObject*>(handle);
if (sync == nullptr || !pVulkanRenderer) {
return true;
}
return pVulkanRenderer->IsFrameSerialComplete(sync->frameSerial);
}
void Present() {
MOBILEGL_ASSERT(pVulkanRenderer, "DirectVulkan::Present called with null VulkanRenderer");
pVulkanRenderer->Present();