[Perf] (MG_Backend): let DirectVulkan trust across frames what it proved once

The draw fast path still paid for its own proofs: the hottest single
load (20% of TrySetupDrawFastPath) was chasing the cold
VertexInputStateFactory heap entry just to answer "same vertex-input
layout?". That answer now comes from the frontend VAO's config-guarded
aux memo (layout hash + attribute masks), and a VAO-cycling stream with
a stable layout skips the pre-flight AND pipeline re-resolution
entirely. The VkProgramObject* is memoized on the snapshot behind a new
ProgramFactory cache-structure epoch (bumped on every insert/erase; use
is re-stamped so the idle sweep can never evict a live entry). A
render-state version move no longer forces the full path: the pipeline
value hash is refreshed in place and the 8-entry memo probed directly
(the GL_BLEND-toggle case).

The resolved-vertex-bindings memo now revalidates all-resident unmapped
entries ACROSS frames via per-binding slice epochs - minted from a
process-lifetime counter so a recycled address can never revalidate,
with every mutation path funnelled through BumpSliceEpoch - while
stamping each resource's GPU-use serial exactly as the skipped acquire
would, preserving the busy-tracking that glBufferSubData's
host-write-vs-staged-copy choice depends on. Resident index buffers get
the same treatment through an EBO slice memo.

The six-part dynamic-state tail (viewport/scissor/blend constants/depth
bias/line width/stencil) is gated behind one render-state-parameters
version + pass-geometry compare per command buffer. GetSlice is inlined;
SampledBindingsUnchanged walks only the program's declared bindings.

Quiet-box 6-round order-alternating A/B (on top of the frontend
VAO-bind commit): vanilla_draw -20.8% (790 -> 626 ns/op, 3.4x native to
2.5x), sampler_churn -28.2%, ubo_range -8.4%, state_toggle -3.8%;
tex_param's matrix flag (+10%) was adjudicated by an isolated
alternating re-run at +1.0% - position bias, not regression. Unit tests
421/421.
This commit is contained in:
BZLZHH
2026-08-06 13:18:38 -04:00
parent b9d8ad0421
commit 8f2b766b56
9 changed files with 431 additions and 161 deletions
@@ -2379,6 +2379,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return it->second;
}
// Structural change: the insert below can move every entry of this
// open-addressing map, so all memoised entry pointers die here.
++m_cacheStructureEpoch;
auto& entry = m_cache[hash];
entry.hash = hash;
entry.lastUsedFrame = m_frameCounter;
@@ -2580,6 +2583,7 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// erase runs ~VkProgramObject (modules/layouts destroyed); notify after
// so an observer never observes a half-destroyed entry through a lookup.
// Observers only need the handle values to purge their keyed caches.
++m_cacheStructureEpoch; // erase moves/kills entries: memoised pointers die
it = m_cache.erase(it);
if (m_evictionObserver != nullptr) {
m_evictionObserver->OnProgramEvicted(hash, descriptorSetLayout);
@@ -105,8 +105,9 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// position-invariance quirk (see PipelineFactory::ShouldSuppressDepthWrite).
Bool fragmentReplacesDepth = false;
// Frame-boundary counter value of the last GetOrCreateProgram hit; drives
// cache eviction (see OnFrameBoundary).
Uint64 lastUsedFrame = 0;
// cache eviction (see OnFrameBoundary). Mutable: the draw snapshot's memoised
// entry pointer re-stamps use through a const reference (StampProgramUse).
mutable Uint64 lastUsedFrame = 0;
static inline VkDevice s_device = VK_NULL_HANDLE;
@@ -262,6 +263,16 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const VkProgramObject& GetOrCreateProgram(
const MG_State::GLState::ProgramObject& program, CompileOptionFlags flags);
// Bumped whenever m_cache's STRUCTURE changes (any insert or erase): the cache is
// an open-addressing map holding entries by value, so both moves existing entries.
// A caller that memoised a VkProgramObject* may keep dereferencing it only while
// this is unchanged; on a bump it must re-run GetOrCreateProgram.
Uint64 GetCacheStructureEpoch() const { return m_cacheStructureEpoch; }
// A memoised entry pointer bypasses GetOrCreateProgram, whose per-lookup stamp is
// what keeps an in-use entry out of OnFrameBoundary's idle sweep - so such a
// caller must re-stamp the entry itself, at least once per frame boundary.
void StampProgramUse(const VkProgramObject& entry) const { entry.lastUsedFrame = m_frameCounter; }
// Observer may be null (no notifications). Not owned.
void SetEvictionObserver(IEvictionObserver* observer) { m_evictionObserver = observer; }
// Frame boundary hook: ages the program cache and evicts long-unused entries
@@ -312,6 +323,8 @@ namespace MobileGL::MG_Backend::DirectVulkan {
mutable ProgramLookupCache m_lastLookup;
// Monotonic frame-boundary counter (bumped in OnFrameBoundary) for cache aging.
Uint64 m_frameCounter = 0;
// See GetCacheStructureEpoch(). Starts at 1 so a zero-initialized memo can never match.
Uint64 m_cacheStructureEpoch = 1;
IEvictionObserver* m_evictionObserver = nullptr;
static inline XXH64_state_t* m_hashState = XXH64_createState();
};
@@ -907,10 +907,15 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Bool UniformManager::SampledBindingsUnchanged(const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj,
const Vector<SampledBindingRecord>& previousRecords) const {
const Uint32 bindingCount =
std::min<Uint32>(m_maxBindings, static_cast<Uint32>(programObj.bindingKinds.size()));
SizeT recordIndex = 0;
for (Uint32 binding = 0; binding < bindingCount; ++binding) {
// Iterate only the bindings this program declares (ascending), exactly like
// BindProgramUniformBuffers: this runs per draw whenever the texture bind
// generation moved, and walking all m_maxBindings slots to find the 1-8 real
// ones dominated it.
for (const Uint32 binding : programObj.activeBindings) {
if (binding >= m_maxBindings) {
break; // ascending, so nothing past the cap can follow
}
if (programObj.bindingKinds[binding] != ProgramFactory::DescriptorBindingKind::CombinedImageSampler) {
continue;
}
@@ -71,6 +71,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
}
const BackendVertexInputState& entry = GetOrCreateVertexInputState(vao, GetOrComputeHash(vao));
vao.SetBackendStateMemo(&entry, m_evictionEpoch);
// Also mirror the layout identity and the two per-draw masks into the VAO's aux
// memo (pure VALUES derived from the VAO configuration, so config-version
// guarding alone is sound). The draw fast path reads them from the VAO object it
// already touched instead of chasing into this entry - see PackVertexInputAuxMemo.
vao.SetBackendAuxMemo(entry.layoutHash,
PackVertexInputAuxMasks(entry.unsupportedAttribMask, entry.attributeLocationMask));
return entry;
}
@@ -71,6 +71,13 @@ namespace MobileGL::MG_Backend::DirectVulkan {
~VertexInputStateFactory() = default;
VertexInputStateFactory(const VertexInputStateFactory&) = delete;
// The VAO aux-memo payload GetOrCreateVertexInputState(vao) stamps: aux0 is the
// entry's layoutHash, aux1 packs (unsupportedAttribMask << 32) | attributeLocationMask.
// Readers that find the aux memo valid can use these without resolving the entry.
static Uint64 PackVertexInputAuxMasks(Uint32 unsupportedAttribMask, Uint32 attributeLocationMask) {
return (static_cast<Uint64>(unsupportedAttribMask) << 32) | attributeLocationMask;
}
HashType ComputeHash(const MG_State::GLState::VertexArrayObject& vao) const;
// Memoized ComputeHash: reuses the VAO's cached hash while its config version
// is unchanged. Use this on per-draw paths.
@@ -176,16 +176,4 @@ namespace MobileGL::MG_Backend::DirectVulkan {
return true;
}
BufferSlice VkBufferObject::GetSlice(VkDeviceSize offset, VkDeviceSize size) const {
MOBILEGL_ASSERT(offset <= m_size, "VkBufferObject::GetSlice offset out of range");
const VkDeviceSize resolvedSize = (size == VK_WHOLE_SIZE) ? (m_size - offset) : size;
MOBILEGL_ASSERT(offset + resolvedSize <= m_size, "VkBufferObject::GetSlice range out of bounds");
BufferSlice slice{};
slice.buffer = m_buffer;
slice.offset = offset;
slice.size = resolvedSize;
slice.mapped = (m_mappedData != nullptr) ? static_cast<Uint8*>(m_mappedData) + offset : nullptr;
return slice;
}
} // namespace MobileGL::MG_Backend::DirectVulkan
@@ -48,7 +48,20 @@ namespace MobileGL::MG_Backend::DirectVulkan {
VkBuffer GetHandle() const { return m_buffer; }
VkDeviceSize GetSize() const { return m_size; }
BufferSlice GetSlice(VkDeviceSize offset = 0, VkDeviceSize size = VK_WHOLE_SIZE) const;
// Inline: runs on the per-draw acquire path (a resident buffer bind is a
// GetSlice per binding), where an out-of-line call was measurable.
BufferSlice GetSlice(VkDeviceSize offset = 0, VkDeviceSize size = VK_WHOLE_SIZE) const {
MOBILEGL_ASSERT(offset <= m_size, "VkBufferObject::GetSlice offset out of range");
const VkDeviceSize resolvedSize = (size == VK_WHOLE_SIZE) ? (m_size - offset) : size;
MOBILEGL_ASSERT(offset + resolvedSize <= m_size, "VkBufferObject::GetSlice range out of bounds");
BufferSlice slice{};
slice.buffer = m_buffer;
slice.offset = offset;
slice.size = resolvedSize;
slice.mapped = (m_mappedData != nullptr) ? static_cast<Uint8*>(m_mappedData) + offset : nullptr;
return slice;
}
void* GetMappedData() const { return m_mappedData; }
Bool IsMapped() const { return m_mappedData != nullptr; }
Bool IsValid() const { return m_allocator != nullptr && m_buffer != VK_NULL_HANDLE && m_allocation != nullptr; }
@@ -264,6 +264,21 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint32 stencilBackWriteMask = 0;
Uint32 stencilFrontReference = 0;
Uint32 stencilBackReference = 0;
// Gate over the whole per-draw dynamic-state tail (viewport, scissor, blend
// constants, depth bias, line width, stencil) - see ApplyDynamicDrawStateTail.
// Every GL input of that tail lives in RenderState's value-shadowed parameters:
// each setter early-outs on an equal value and bumps the parameters version
// otherwise, and capability toggles (scissor test) bump it too. So an unchanged
// version + unchanged pass geometry means re-running the tail could only
// re-derive the exact values already applied on this command buffer. The
// remaining input, the swapchain pre-transform, cannot change mid-recording
// (a swapchain recreate retires the command buffer, and recording begin resets
// this whole shadow).
Bool dynamicTailValid = false;
Uint dynamicTailParamsVersion = 0;
Int dynamicTailExtentX = 0;
Int dynamicTailExtentY = 0;
Bool dynamicTailIsDefaultFbo = false;
};
static DynamicStateShadow g_dynamicStateShadow;
@@ -3047,55 +3062,98 @@ void main() {
Bool VulkanRenderer::TryBindResolvedVertexBindings(
VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ResolvedVertexBindings& entry,
const VertexInputStateFactory::BackendVertexInputState& vertexInputState, Uint32 activeAttribMask,
Uint32 bindingCount, Uint64 frameSerial) {
// Frame-scoped: a streamed binding's slice moves to a new arena block every
// frame by design, and the frame's first resolve is also what stamps every
// bound buffer's GPU-use serial, which the buffer manager's busy tracking (and
// therefore glBufferSubData's choice between a host write and a staged copy)
// depends on.
if (entry.frameSerial != frameSerial || entry.vertexInputState != &vertexInputState ||
entry.vertexInputHash != vertexInputState.hash || entry.activeAttribMask != activeAttribMask ||
entry.bindingCount != bindingCount) {
ResolvedVertexBindings& entry, Uint64 vaoContentHash, Uint32 activeAttribMask,
Uint64 frameSerial) {
// Layout + buffer identity in two loads from data the caller already has: the
// VAO's content hash mixes every enabled attribute's format AND its bound
// buffer's address (any change bumps the config version, invalidating the hash
// memo the caller read), and the program's active-location mask fixes the
// synthetic-binding set. Together they pin bindings.size(), every base offset
// and which buffer each binding reads, so the hit path never has to resolve the
// vertex-input factory entry at all - that chase was the dominant cost of a
// VAO-cycling frame's memo hit.
if (entry.frameSerial == 0 || entry.vertexInputHash != vaoContentHash ||
entry.activeAttribMask != activeAttribMask) {
return false;
}
// The layout identity above already fixes which buffer each binding reads: the
// factory entry is reached through the VAO's config-version-gated memo and its
// hash mixes the bound buffers' addresses, so rebinding a buffer resolves to a
// different entry. What is left to establish is that those buffers still hand
// back the slices recorded here, and that none of them is a host map whose
// shadow needs pushing down. An unmoved manager-wide epoch counter says both.
if (!entry.anyBufferMapped && entry.sliceEpochCounter == m_bufferManager.GetSliceEpochCounter()) {
ShadowedBindVertexBuffers(commandBuffer, entry.vkBuffers, entry.vkOffsets, bindingCount);
if (entry.frameSerial == frameSerial) {
// What is left to establish is that the buffers still hand back the slices
// recorded here, and that none of them is a host map whose shadow needs
// pushing down. An unmoved manager-wide epoch counter says both.
if (!entry.anyBufferMapped && entry.sliceEpochCounter == m_bufferManager.GetSliceEpochCounter()) {
ShadowedBindVertexBuffers(commandBuffer, entry.vkBuffers, entry.vkOffsets, entry.bindingCount);
return true;
}
// Something moved somewhere; ask the buffers themselves.
const auto& attributes = vao.GetAllAttributes();
const MG_State::GLState::BufferObject* synced = nullptr;
for (Uint32 binding = 0; binding < entry.bindingCount; ++binding) {
auto* bufferObject = attributes[entry.attributeLocations[binding]].Buffer.get();
if (bufferObject != entry.buffers[binding]) {
return false;
}
// What the resolving path does before every acquire: a persistent map the
// backend could not adopt into coherent GPU storage mutates its shadow with
// no API call, so the write range has to be pushed down here too. It is a
// no-op for every buffer that is not such a map; when it is not, it dispatches
// a SubData that retires the epoch below, and this draw resolves in full.
if (bufferObject != synced) {
bufferObject->SyncPersistentMappedRange();
synced = bufferObject;
}
const auto* resource =
static_cast<const VkBufferResource*>(bufferObject->GetBackendResource().get());
if (resource == nullptr || resource->sliceEpoch != entry.sliceEpochs[binding]) {
return false;
}
}
ShadowedBindVertexBuffers(commandBuffer, entry.vkBuffers, entry.vkOffsets, entry.bindingCount);
return true;
}
// Something moved somewhere; ask the buffers themselves.
// Cross-frame revalidation. Only all-resident, unmapped layouts qualify:
// - a streamed binding's slice moves to a new arena block every frame BY DESIGN,
// but every such move funnels through BumpSliceEpoch, so the per-binding epoch
// compares below catch it (as they do a respecify, a sub-data update, a
// resident<->streamed promotion and a buffer becoming persistently mapped);
// - a mapped buffer mutates its shadow with no API call and must re-run the
// Sync/acquire path every frame, so it is excluded outright;
// - epochs are minted from a process-lifetime counter, so a deleted buffer (or
// resource) recycled at the same address can never reproduce a recorded epoch.
if (entry.anyBufferMapped) {
return false;
}
const auto& attributes = vao.GetAllAttributes();
const MG_State::GLState::BufferObject* synced = nullptr;
for (Uint32 binding = 0; binding < bindingCount; ++binding) {
VkBufferResource* resources[ResolvedVertexBindings::kMaxBindings] = {};
for (Uint32 binding = 0; binding < entry.bindingCount; ++binding) {
// Compare against the LIVE pointer before any dereference: entry.buffers may
// dangle if the app deleted a buffer since (the VAO unbind that deletion
// performs bumps the config version, so the hash compare above already
// declined - this is defense in depth for the aliased-address case).
auto* bufferObject = attributes[entry.attributeLocations[binding]].Buffer.get();
if (bufferObject != entry.buffers[binding]) {
return false;
}
// What the resolving path does before every acquire: a persistent map the
// backend could not adopt into coherent GPU storage mutates its shadow with
// no API call, so the write range has to be pushed down here too. It is a
// no-op for every buffer that is not such a map; when it is not, it dispatches
// a SubData that retires the epoch below, and this draw resolves in full.
if (bufferObject != synced) {
bufferObject->SyncPersistentMappedRange();
synced = bufferObject;
}
const auto* resource = static_cast<const VkBufferResource*>(bufferObject->GetBackendResource().get());
auto* resource = static_cast<VkBufferResource*>(bufferObject->GetBackendResource().get());
if (resource == nullptr || resource->sliceEpoch != entry.sliceEpochs[binding]) {
return false;
}
resources[binding] = resource;
}
ShadowedBindVertexBuffers(commandBuffer, entry.vkBuffers, entry.vkOffsets, bindingCount);
// Every binding still resolves to the recorded slice. The per-frame resolve this
// replaces had one side effect the busy tracking depends on (glBufferSubData's
// host-write-vs-staged-copy choice): stamping each resource's GPU-use serial.
// Do exactly that, then re-arm the entry so the rest of the frame's draws take
// the one-compare path above.
for (Uint32 binding = 0; binding < entry.bindingCount; ++binding) {
resources[binding]->lastUseSerial = frameSerial;
}
entry.frameSerial = frameSerial;
entry.sliceEpochCounter = m_bufferManager.GetSliceEpochCounter();
ShadowedBindVertexBuffers(commandBuffer, entry.vkBuffers, entry.vkOffsets, entry.bindingCount);
return true;
}
@@ -3127,14 +3185,7 @@ void main() {
}
return drawElementBound;
};
// programObj is resolved once in SetupDraw and passed in; re-resolving it here would repeat
// the GetCurrentProgram + GetOrCreateProgram hash lookup every draw.
auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
const Uint32 activeAttribMask = programObj.activeVertexInputLocationMask;
const Uint32 vertexInputAttribMask = vertexInputState.attributeLocationMask;
const Uint32 missingAttribMask = activeAttribMask & ~vertexInputAttribMask;
const auto bindingCount = vertexInputState.bindings.size() + static_cast<SizeT>(std::popcount(missingAttribMask));
const Uint64 frameSerial = m_bufferManager.GetFrameSerial();
if (frameSerial != m_resolvedVertexBindingsFrameSerial) {
@@ -3143,17 +3194,32 @@ void main() {
m_resolvedVertexBindings.clear();
}
}
// Probe the memo BEFORE resolving the vertex-input entry: a hit needs nothing
// from it (the VAO's own hash memo pins layout and buffers - see
// TryBindResolvedVertexBindings), and skipping the resolve also skips its
// per-draw cold chase into the factory's heap entry.
m_currentDrawResolvedEntry = nullptr;
ResolvedVertexBindings* memo = nullptr;
if (auto found = m_resolvedVertexBindings.find(&vao); found != m_resolvedVertexBindings.end()) {
memo = &found->second;
if (TryBindResolvedVertexBindings(commandBuffer, vao, *memo, vertexInputState, activeAttribMask,
static_cast<Uint32>(bindingCount), frameSerial)) {
return true;
Uint64 vaoContentHash = 0;
const Bool vaoHashKnown = vao.GetBackendHashMemo(vaoContentHash);
if (vaoHashKnown) {
if (auto found = m_resolvedVertexBindings.find(&vao); found != m_resolvedVertexBindings.end()) {
memo = &found->second;
if (TryBindResolvedVertexBindings(commandBuffer, vao, *memo, vaoContentHash,
activeAttribMask, frameSerial)) {
m_currentDrawResolvedEntry = memo;
return true;
}
// Whatever it described is stale; a resolve that bails out below must not
// leave the old contents matchable either.
memo->frameSerial = 0;
}
// Whatever it described is stale; a resolve that bails out below must not
// leave the old contents matchable either.
memo->frameSerial = 0;
}
auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
const Uint32 vertexInputAttribMask = vertexInputState.attributeLocationMask;
const Uint32 missingAttribMask = activeAttribMask & ~vertexInputAttribMask;
const auto bindingCount = vertexInputState.bindings.size() + static_cast<SizeT>(std::popcount(missingAttribMask));
// Anything the memo cannot key on (see ResolvedVertexBindings) clears this as
// the resolve below discovers it.
Bool memoisable = missingAttribMask == 0 && bindingCount > 0 &&
@@ -3433,7 +3499,8 @@ void main() {
if (memoisable && memo != nullptr) {
std::copy_n(vkBuffers.data(), count, memo->vkBuffers);
std::copy_n(vkOffsets.data(), count, memo->vkOffsets);
memo->vertexInputState = &vertexInputState;
// The factory keys entries on the VAO content hash, so this is the same
// value the hit path reads back from the VAO's own hash memo.
memo->vertexInputHash = vertexInputState.hash;
memo->activeAttribMask = activeAttribMask;
memo->bindingCount = count;
@@ -3443,6 +3510,9 @@ void main() {
// Published last: the entry is only matchable once every field above is
// the one this completed resolve produced.
memo->frameSerial = frameSerial;
// The same draw's UploadAndBindIndexBuffer may extend this entry with the
// EBO slice memo; the pointer dies at the map's next insert (next draw).
m_currentDrawResolvedEntry = memo;
}
ShadowedBindVertexBuffers(commandBuffer, vkBuffers.data(), vkOffsets.data(), count);
}
@@ -3574,6 +3644,39 @@ void main() {
MOBILEGL_ASSERT(pIndexBufferView->indexByteOffset + indexDataSizeBytes <= indexBuffer->GetSize(),
"DrawElements index range out of bounds");
// EBO slice memo (see ResolvedVertexBindings): skips the per-draw
// AcquireResidentSlice when the live bound EBO and its resource epoch still
// match what the recording draw resolved. Restart substitution re-uploads per
// draw and never stores a memo, so a hit requires it off. The index TYPE is not
// memo state: it flows from the draw's view into the shadowed bind below.
ResolvedVertexBindings* indexMemo = m_currentDrawResolvedEntry;
if (indexMemo != nullptr && !substituteRestart && indexMemo->indexFrameSerial != 0 &&
indexMemo->indexBuffer == indexBuffer) {
auto* resource = static_cast<VkBufferResource*>(
indexBufferShared->GetBackendResource().get());
if (resource != nullptr && resource->sliceEpoch == indexMemo->indexSliceEpoch) {
const Uint64 frameSerial = m_bufferManager.GetFrameSerial();
if (indexMemo->indexFrameSerial != frameSerial) {
// Same busy-tracking stamp the skipped acquire would have made.
resource->lastUseSerial = frameSerial;
indexMemo->indexFrameSerial = frameSerial;
}
const VkDeviceSize memoBindOffset = indexMemo->indexSliceOffset +
static_cast<VkDeviceSize>(pIndexBufferView->indexByteOffset);
auto& shadow = g_dynamicStateShadow;
if (!shadow.indexBindValid || shadow.indexBuffer != indexMemo->indexVkBuffer ||
shadow.indexOffset != memoBindOffset || shadow.indexType != vkIndexType) {
vkCmdBindIndexBuffer(frame.commandBuffer, indexMemo->indexVkBuffer, memoBindOffset,
vkIndexType);
shadow.indexBindValid = true;
shadow.indexBuffer = indexMemo->indexVkBuffer;
shadow.indexOffset = memoBindOffset;
shadow.indexType = vkIndexType;
}
return true;
}
}
BufferSlice slice{};
MOBILEGL_ASSERT(indexBufferShared != nullptr, "UploadAndBindIndexBuffer failed to resolve shared EBO");
if (substituteRestart) {
@@ -3598,6 +3701,19 @@ void main() {
} else if (!m_bufferManager.AcquireResidentSlice(BufferKind::Index, indexBufferShared, slice)) {
MGLOG_E("DrawElements skipped: failed to sync resident index buffer");
return false;
} else if (indexMemo != nullptr && !substituteRestart) {
// Resident acquire succeeded: record the slice for the next draw of this VAO.
// Read the epoch AFTER the acquire - it is the acquire that mints the epoch
// this slice belongs to.
const auto* resource = static_cast<const VkBufferResource*>(
indexBufferShared->GetBackendResource().get());
if (resource != nullptr) {
indexMemo->indexBuffer = indexBuffer;
indexMemo->indexSliceEpoch = resource->sliceEpoch;
indexMemo->indexVkBuffer = slice.buffer;
indexMemo->indexSliceOffset = slice.offset;
indexMemo->indexFrameSerial = m_bufferManager.GetFrameSerial();
}
}
const VkDeviceSize indexBindOffset =
slice.offset + static_cast<VkDeviceSize>(pIndexBufferView->indexByteOffset);
@@ -4893,6 +5009,42 @@ void main() {
}
void VulkanRenderer::ApplyDynamicDrawStateTail(FrameContext::FrameData& frame, const IntVec2& extent,
Bool isDefaultFbo) {
auto& shadow = g_dynamicStateShadow;
// One compare for the whole tail: see the gate's declaration in
// DynamicStateShadow for why (version, extent, default-FBO flag) pins every
// input the six Apply* below read.
const Uint paramsVersion = MG_State::pGLContext->GetRenderStateParametersVersion();
if (shadow.dynamicTailValid && shadow.dynamicTailParamsVersion == paramsVersion &&
shadow.dynamicTailExtentX == extent.x() && shadow.dynamicTailExtentY == extent.y() &&
shadow.dynamicTailIsDefaultFbo == isDefaultFbo) {
return;
}
ApplyGLViewportState(frame.commandBuffer, extent, m_swapchainObject.GetPreTransform(), isDefaultFbo);
ApplyBlendConstants(frame.commandBuffer);
ApplyPolygonOffsetState(frame.commandBuffer);
ApplyLineWidthState(frame.commandBuffer);
ApplyStencilState(frame.commandBuffer);
const Bool scissorEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::ScissorTest);
VkRect2D scissor{};
if (scissorEnabled) {
const auto& scissorBox = MG_State::pGLContext->GetScissorBox();
scissor = isDefaultFbo
? MakeDefaultFramebufferScissorRect(scissorBox, extent, m_swapchainObject.GetPreTransform())
: MakeClampedScissorRect(scissorBox, extent);
} else {
scissor.offset = {0, 0};
scissor.extent = { (Uint)extent.x(), (Uint)extent.y() };
}
ShadowedSetScissor(frame.commandBuffer, scissor);
shadow.dynamicTailValid = true;
shadow.dynamicTailParamsVersion = paramsVersion;
shadow.dynamicTailExtentX = extent.x();
shadow.dynamicTailExtentY = extent.y();
shadow.dynamicTailIsDefaultFbo = isDefaultFbo;
}
Bool VulkanRenderer::TrySetupDrawFastPath(FrameContext::FrameData& frame, GLenum mode,
Flags<DrawSetupAspect> aspects, const DrawCmdParam& drawParams,
const IndexBufferView* pIndexBufferView) {
@@ -4959,28 +5111,67 @@ void main() {
m_renderPassManager->GetRenderbufferImageEpoch() != snap.renderbufferImageEpoch) {
return false;
}
const auto& programObj = m_programFactory->GetOrCreateProgram(
program, ProgramFactory::CompileOptionFlags(snap.resolvedTransformFlags));
// Program entry: (lifetimeId, backend-state version, resolved flags) were proven
// equal above, and those pin the factory hash - so the snapshot's memoised entry
// pointer IS this draw's entry while the factory's open-addressing cache has not
// moved entries (structure epoch). Bypassing GetOrCreateProgram skips its use
// stamp, so re-stamp here or the idle sweep could evict a live entry.
const ProgramFactory::VkProgramObject* programObjPtr = snap.programObj;
if (programObjPtr != nullptr &&
snap.programFactoryEpoch == m_programFactory->GetCacheStructureEpoch()) {
m_programFactory->StampProgramUse(*programObjPtr);
} else {
programObjPtr = &m_programFactory->GetOrCreateProgram(
program, ProgramFactory::CompileOptionFlags(snap.resolvedTransformFlags));
snap.programObj = programObjPtr;
snap.programFactoryEpoch = m_programFactory->GetCacheStructureEpoch();
}
const auto& programObj = *programObjPtr;
// The pipeline and the vertex-input pre-flight depend on the VAO only through
// its resolved LAYOUT (layoutHash folds the attribute formats, bindings and the
// unsupported mask; the masks below are functions of the same configuration),
// never its identity. A VAO-cycling stream (Minecraft chunk rendering) swaps
// hundreds of VAOs sharing one layout per frame: answer "same layout?" from the
// VAO's aux memo - it sits next to the config-version word this compare chain
// already loaded - instead of chasing the vertex-input factory's cold heap entry.
Uint64 vaoLayoutHash = snap.vaoLayoutHash;
Bool vaoLayoutMoved = false;
if (vaoMoved) {
// Vertex-input pre-flight for the changed VAO, mirroring the full path:
// a bad attribute must never be baked into a cached VkPipeline, and the
// current-value synthesis in UploadAndBindVertexBuffers must never see
// an unsupported generic-attribute type. Declining routes the draw
// through the full path's loud failure reporting.
const auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
const Uint32 activeAttribMask = programObj.activeVertexInputLocationMask;
if ((vertexInputState.unsupportedAttribMask & activeAttribMask) != 0) {
return false;
Uint64 auxMasks = 0;
if (!vao.GetBackendAuxMemo(vaoLayoutHash, auxMasks)) {
// First sight of this VAO configuration: resolve (which stamps the aux
// memo for every later draw) and read the same facts from the entry.
const auto& vertexInputState = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
vaoLayoutHash = vertexInputState.layoutHash;
auxMasks = VertexInputStateFactory::PackVertexInputAuxMasks(
vertexInputState.unsupportedAttribMask, vertexInputState.attributeLocationMask);
}
const Uint32 missingAttribMask = activeAttribMask & ~vertexInputState.attributeLocationMask;
if (missingAttribMask != 0) {
for (Uint32 location = 0; location < kMaxVertexAttribs; ++location) {
if ((missingAttribMask & (1u << location)) == 0) {
continue;
}
if (MG_State::GLState::ClassifyVertexAttribType(programObj.vertexInputTypes[location])
.baseType == MG_State::GLState::VertexAttribBaseType::Unsupported) {
return false;
vaoLayoutMoved = vaoLayoutHash != snap.vaoLayoutHash;
if (vaoLayoutMoved) {
// Vertex-input pre-flight for the changed layout, mirroring the full
// path: a bad attribute must never be baked into a cached VkPipeline,
// and the current-value synthesis in UploadAndBindVertexBuffers must
// never see an unsupported generic-attribute type. Declining routes the
// draw through the full path's loud failure reporting. An UNMOVED layout
// needs no pre-flight: the snapshotting draw passed it with identical
// inputs (same program; masks pinned by the layout hash).
const Uint32 unsupportedAttribMask = static_cast<Uint32>(auxMasks >> 32);
const Uint32 attributeLocationMask = static_cast<Uint32>(auxMasks);
const Uint32 activeAttribMask = programObj.activeVertexInputLocationMask;
if ((unsupportedAttribMask & activeAttribMask) != 0) {
return false;
}
const Uint32 missingAttribMask = activeAttribMask & ~attributeLocationMask;
if (missingAttribMask != 0) {
for (Uint32 location = 0; location < kMaxVertexAttribs; ++location) {
if ((missingAttribMask & (1u << location)) == 0) {
continue;
}
if (MG_State::GLState::ClassifyVertexAttribType(programObj.vertexInputTypes[location])
.baseType == MG_State::GLState::VertexAttribBaseType::Unsupported) {
return false;
}
}
}
}
@@ -5043,35 +5234,41 @@ void main() {
}
// Everything the full path would re-resolve is provably unchanged - or, for
// a moved pipeline-state version or a changed VAO, reduces to re-resolving
// just the pipeline through the value-keyed memo against the still-active
// render pass. Run only the per-draw tail.
// a moved pipeline-state version or a changed vertex-input LAYOUT, reduces to
// re-resolving just the pipeline through the value-keyed memo against the
// still-active render pass. A changed VAO with the SAME layout keeps the
// snapshot's pipeline outright (the layout is the pipeline's only VAO input).
// Run only the per-draw tail.
VkPipeline pipeline = snap.pipeline;
if (renderStateMoved || vaoMoved) {
if (renderStateMoved || vaoLayoutMoved) {
pipeline = VK_NULL_HANDLE;
if (!renderStateMoved && m_pipelineStateHashValid &&
m_pipelineStateHashVersion == renderStateVersion) {
// VAO-only movement: the render pass is provably the snapshot's (hash
// match above) and the pipeline-state hash is cached for this
// untouched state version, so probe the value-keyed pipeline memo
// directly - no render-pass-entry re-fetch (whose pending-clear
// probes cost more than the whole probe below). A miss, or a cached
// hash computed against another pass's attachment count (the hash
// folds it in, so such a mismatch can only produce a miss, never a
// false hit), falls through to the full lookup.
const auto& vis = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao);
const auto memoTransformFlags =
ProgramFactory::CompileOptionFlags(snap.resolvedTransformFlags);
for (Uint32 i = 0; i < m_pipelineMemoCount; ++i) {
const PipelineMemoEntry& entry = m_pipelineMemo[i];
if (entry.pipeline != VK_NULL_HANDLE && entry.mode == mode &&
entry.programHash == programObj.hash && entry.vertexInputHash == vis.layoutHash &&
entry.renderPassHash == snap.renderPassHash &&
entry.pipelineStateHash == m_pipelineStateHash &&
entry.transformFlags == memoTransformFlags) {
pipeline = entry.pipeline;
break;
}
// The render pass is provably the snapshot's (hash match above), so probe
// the value-keyed pipeline memo directly - no render-pass-entry re-fetch
// (whose pending-clear probes cost more than the whole probe below). For a
// moved state version, first refresh the pipeline-state VALUE hash exactly
// as GetOrCreatePipeline would (same inputs: the snapshot pins the pass, so
// its color attachment count is the right hash input); the value hash is
// what lets a per-draw GL_BLEND toggle alternate between two memo entries
// instead of missing forever on a monotonic version. A miss falls through
// to the full lookup.
if (!m_pipelineStateHashValid || m_pipelineStateHashVersion != renderStateVersion ||
m_pipelineStateHashColorCount != snap.renderPassColorCount) {
m_pipelineStateHash = ComputePipelineStateHash(snap.renderPassColorCount);
m_pipelineStateHashVersion = renderStateVersion;
m_pipelineStateHashColorCount = snap.renderPassColorCount;
m_pipelineStateHashValid = true;
}
const auto memoTransformFlags =
ProgramFactory::CompileOptionFlags(snap.resolvedTransformFlags);
for (Uint32 i = 0; i < m_pipelineMemoCount; ++i) {
const PipelineMemoEntry& entry = m_pipelineMemo[i];
if (entry.pipeline != VK_NULL_HANDLE && entry.mode == mode &&
entry.programHash == programObj.hash && entry.vertexInputHash == vaoLayoutHash &&
entry.renderPassHash == snap.renderPassHash &&
entry.pipelineStateHash == m_pipelineStateHash &&
entry.transformFlags == memoTransformFlags) {
pipeline = entry.pipeline;
break;
}
}
if (pipeline == VK_NULL_HANDLE) {
@@ -5098,6 +5295,7 @@ void main() {
snap.bindGeneration = bindGeneration;
snap.vao = static_cast<const void*>(&vao);
snap.vaoConfigVersion = vao.GetConfigVersion();
snap.vaoLayoutHash = vaoLayoutHash;
snap.pipeline = pipeline;
if (!g_dynamicStateShadow.graphicsPipelineValid ||
g_dynamicStateShadow.graphicsPipeline != pipeline) {
@@ -5118,25 +5316,7 @@ void main() {
const Bool idxUploadOk = UploadAndBindIndexBuffer(frame, vao, pIndexBufferView);
MOBILEGL_ASSERT(idxUploadOk, "SetupDraw fast path: failed to upload index buffer");
}
ApplyGLViewportState(frame.commandBuffer, snap.renderPassExtent, m_swapchainObject.GetPreTransform(),
snap.drawFboIsDefault);
ApplyBlendConstants(frame.commandBuffer);
ApplyPolygonOffsetState(frame.commandBuffer);
ApplyLineWidthState(frame.commandBuffer);
ApplyStencilState(frame.commandBuffer);
const Bool scissorEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::ScissorTest);
VkRect2D scissor{};
if (scissorEnabled) {
const auto& scissorBox = MG_State::pGLContext->GetScissorBox();
scissor = snap.drawFboIsDefault
? MakeDefaultFramebufferScissorRect(scissorBox, snap.renderPassExtent,
m_swapchainObject.GetPreTransform())
: MakeClampedScissorRect(scissorBox, snap.renderPassExtent);
} else {
scissor.offset = {0, 0};
scissor.extent = { (Uint)snap.renderPassExtent.x(), (Uint)snap.renderPassExtent.y() };
}
ShadowedSetScissor(frame.commandBuffer, scissor);
ApplyDynamicDrawStateTail(frame, snap.renderPassExtent, snap.drawFboIsDefault);
return true;
}
@@ -5215,6 +5395,10 @@ void main() {
}
}
const auto& programObj = m_programFactory->GetOrCreateProgram(program, transformFlags);
// For the snapshot's memoised entry pointer: if anything below inserts into the
// program cache (blit/aux program compiles), the epoch moves and the snapshot
// stores no pointer for this draw - the fast path then re-looks-up once.
const Uint64 programFactoryEpochAtResolve = m_programFactory->GetCacheStructureEpoch();
// Begin command recording if not yet
if (!frame.isCommandRecording) {
@@ -5478,26 +5662,7 @@ void main() {
MOBILEGL_ASSERT(idxUploadOk, "SetupDraw skipped: failed to upload index buffer");
}
ApplyGLViewportState(frame.commandBuffer, renderPassEntry->extent,
m_swapchainObject.GetPreTransform(), drawFbo->IsDefaultFramebuffer());
ApplyBlendConstants(frame.commandBuffer);
ApplyPolygonOffsetState(frame.commandBuffer);
ApplyLineWidthState(frame.commandBuffer);
ApplyStencilState(frame.commandBuffer);
Bool scissorEnabled = MG_State::pGLContext->IsCapabilityEnabled(CapabilityInput::ScissorTest);
VkRect2D scissor{};
if (scissorEnabled) {
const auto& scissorBox = MG_State::pGLContext->GetScissorBox();
scissor = drawFbo->IsDefaultFramebuffer()
? MakeDefaultFramebufferScissorRect(scissorBox, renderPassEntry->extent,
m_swapchainObject.GetPreTransform())
: MakeClampedScissorRect(scissorBox, renderPassEntry->extent);
} else {
scissor.offset = {0, 0};
scissor.extent = { (Uint)renderPassEntry->extent.x(), (Uint)renderPassEntry->extent.y() };
}
ShadowedSetScissor(frame.commandBuffer, scissor);
ApplyDynamicDrawStateTail(frame, renderPassEntry->extent, drawFbo->IsDefaultFramebuffer());
// Snapshot the fully resolved configuration for the consecutive-draw
// fast path (see TrySetupDrawFastPath).
@@ -5526,7 +5691,21 @@ void main() {
snap.renderbufferImageEpoch = m_renderPassManager->GetRenderbufferImageEpoch();
snap.drawUsesDepthStencil = drawUsesDepthStencil;
snap.renderPassExtent = renderPassEntry->extent;
snap.renderPassColorCount = renderPassEntry->colorAttachmentCount;
snap.pipeline = pipeline;
// The layout identity the fast path's aux-memo compare answers against.
// A memo hit here, not a rebuild: the pre-flight above resolved this
// VAO's entry already, so this re-reads the VAO's stamped state memo.
snap.vaoLayoutHash = m_vertexInputStateFactory->GetOrCreateVertexInputState(vao).layoutHash;
// Entry pointer memo: only when nothing since the resolve restructured
// the factory cache (see programFactoryEpochAtResolve above).
if (m_programFactory->GetCacheStructureEpoch() == programFactoryEpochAtResolve) {
snap.programObj = &programObj;
snap.programFactoryEpoch = programFactoryEpochAtResolve;
} else {
snap.programObj = nullptr;
snap.programFactoryEpoch = 0;
}
snap.samplingResolutionGeneration = MG_State::pGLContext->GetSamplingResolutionGeneration();
Uint64 snapContentSum = 0;
Uint64 snapParamsSum = 0;
@@ -738,7 +738,26 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// flips it must fall back to the full path's pass selection.
Bool drawUsesDepthStencil = false;
IntVec2 renderPassExtent = {0, 0};
// colorAttachmentCount of the snapshotting draw's render pass: the
// pipeline-state hash input, so the fast path can refresh that hash and
// probe the pipeline memo after a state change without re-fetching the
// render-pass entry (the pass itself is pinned by renderPassHash above).
Uint32 renderPassColorCount = 0;
VkPipeline pipeline = VK_NULL_HANDLE;
// layoutHash of the snapshotting draw's vertex-input state. The pipeline and
// the vertex-input pre-flight depend on the VAO only through this (plus the
// program, pinned separately), so a changed VAO whose aux memo carries the
// same layoutHash re-uses the snapshot's pipeline and pre-flight verdict
// outright - the VAO-cycling case Minecraft chunk rendering hits every draw.
Uint64 vaoLayoutHash = 0;
// Memoised ProgramFactory entry of the snapshotting draw, valid while
// (programLifetimeId, programVersion, resolvedTransformFlags) match - all
// checked above - AND the factory's cache structure epoch is unchanged (the
// cache is open-addressing and holds entries by value, so any insert/erase
// moves them). The fast path must re-stamp use through StampProgramUse when
// it bypasses GetOrCreateProgram, or the idle sweep could evict a live entry.
const ProgramFactory::VkProgramObject* programObj = nullptr;
Uint64 programFactoryEpoch = 0;
};
SetupDrawSnapshot m_setupDrawSnapshot;
@@ -838,13 +857,19 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// either, so a wider layout resolves per draw. Minecraft-shaped layouts use four.
static constexpr Uint32 kMaxBindings = 8;
// Frame serial of the last completed resolve OR cross-frame revalidation.
// Zero until a resolve completes, and reset to zero before one starts, so a
// resolve that bails out midway cannot leave a half-filled entry matchable.
// Unlike the original frame-scoped memo, an entry whose buffers are all
// resident and unmapped is revalidated across frames (per-binding slice
// epoch compares) instead of re-resolved - see TryBindResolvedVertexBindings.
Uint64 frameSerial = 0;
// Identity of the resolved Vulkan layout: fixes bindings.size(), each
// binding's base offset, which bindings are client/converted, and (through
// the hash, which mixes the bound buffers' addresses) the VAO configuration.
const VertexInputStateFactory::BackendVertexInputState* vertexInputState = nullptr;
// Identity of the resolved Vulkan layout: the VAO's content hash
// (VertexInputStateFactory::GetOrComputeHash - the same value the factory
// keys its entries on) fixes bindings.size(), each binding's base offset,
// which bindings are client/converted, and (through the mixed-in buffer
// addresses) which buffer each binding reads. Compared against the VAO's
// own hash memo on the hit path, so a hit never touches the factory entry.
VertexInputStateFactory::HashType vertexInputHash = 0;
// The program's vertex input layout: decides the synthetic-binding set and
// hence the total binding count.
@@ -865,6 +890,22 @@ namespace MobileGL::MG_Backend::DirectVulkan {
Uint64 sliceEpochs[kMaxBindings] = {};
VkBuffer vkBuffers[kMaxBindings] = {};
VkDeviceSize vkOffsets[kMaxBindings] = {};
// Resident element-buffer slice memo (skips the per-draw AcquireResidentSlice
// for the VAO's EBO, which cold-chases 500+ distinct resources in a
// chunk-cycling frame). Self-validating exactly like the bindings above: a hit
// requires the LIVE bound EBO pointer to equal indexBuffer AND that buffer's
// resource to still carry indexSliceEpoch (epochs are minted from a
// process-lifetime counter, so a recycled address can never revalidate).
// Restart-substituted and streamed EBOs are never stored. indexFrameSerial
// tracks the last frame the resource's GPU-use serial was stamped through this
// memo; 0 means no index memo. Independent of the vertex half: both are
// (pointer, epoch)-validated, so neither can serve stale state for the other.
const MG_State::GLState::BufferObject* indexBuffer = nullptr;
Uint64 indexSliceEpoch = 0;
VkBuffer indexVkBuffer = VK_NULL_HANDLE;
VkDeviceSize indexSliceOffset = 0;
Uint64 indexFrameSerial = 0;
};
// Keyed on the VAO address purely as a lookup hint - an entry is only ever
// compared against, never dereferenced through, so a recycled address cannot
@@ -878,6 +919,12 @@ namespace MobileGL::MG_Backend::DirectVulkan {
// Entries of deleted VAOs are never hit again but still occupy the map; drop the
// lot at a frame boundary once they could outweigh a large frame's working set.
static constexpr SizeT kMaxResolvedVertexBindings = 4096;
// The current draw's memo entry, set by UploadAndBindVertexBuffers and consumed
// by the same draw's UploadAndBindIndexBuffer (the EBO memo lives in the same
// entry). Valid ONLY within that window: the map is open-addressing, so the next
// insert (i.e. the next draw's resolve of a new VAO) can move it. Null when the
// draw's layout is not memoisable.
ResolvedVertexBindings* m_currentDrawResolvedEntry = nullptr;
void CreateInstance();
VkResult SetupDebugMessenger();
@@ -909,17 +956,25 @@ namespace MobileGL::MG_Backend::DirectVulkan {
const MG_State::GLState::ProgramObject& program,
const ProgramFactory::VkProgramObject& programObj);
// The per-draw dynamic-state tail (viewport, scissor, blend constants, depth
// bias, line width, stencil), gated behind one render-state-parameters-version
// compare per command buffer - see the gate fields in DynamicStateShadow.
void ApplyDynamicDrawStateTail(FrameContext::FrameData& frame, const IntVec2& extent, Bool isDefaultFbo);
Bool UploadAndBindVertexBuffers(VkCommandBuffer commandBuffer, const MG_State::GLState::VertexArrayObject& vao,
const ProgramFactory::VkProgramObject& programObj,
const DrawCmdParam& drawParams,
const IndexBufferView* pIndexBufferView);
// Binds `entry`'s memoised buffers when every input it was resolved from is
// still live and unchanged, else returns false and leaves nothing bound.
// vaoContentHash is the VAO's memoised content hash (GetBackendHashMemo), which
// pins the layout AND the bound buffers without resolving the factory entry.
// Non-const entry: a cross-frame revalidation refreshes its serial/epoch stamps.
Bool TryBindResolvedVertexBindings(VkCommandBuffer commandBuffer,
const MG_State::GLState::VertexArrayObject& vao,
const ResolvedVertexBindings& entry,
const VertexInputStateFactory::BackendVertexInputState& vertexInputState,
Uint32 activeAttribMask, Uint32 bindingCount, Uint64 frameSerial);
ResolvedVertexBindings& entry,
Uint64 vaoContentHash,
Uint32 activeAttribMask, Uint64 frameSerial);
Bool UploadAndBindIndexBuffer(FrameContext::FrameData& frame,
const MG_State::GLState::VertexArrayObject& vao,
const IndexBufferView* pIndexBufferView = nullptr);