Files
MobileGL/MobileGL/MG_State/GLState/ProgramState/ShaderPreprocessCache.h
T
BZLZHH c93e5fa409 [Refactor] (MG_State, MG_Util): join-by-construction link/compile artifacts (P1 stage 2)
Still fully synchronous - EnsureLinkJoined()/EnsureCompileJoined() are empty
inline no-ops (verified to fold away at every one of the ~1200 call sites;
this project builds without LTO) - but every read of link- or compile-produced
state now goes through a private accessor the compiler enforces, so when
stage 4 moves the bodies onto pool workers, 'which reads must join' is a
type-system fact instead of a 400-line audit.

- ProgramObject: the 31 fields ResetLinkArtifacts clears plus the 5 link
  outputs it forgot (infoLog, linkedFragData{Location,Index}, the geometry
  strip-capture pair) move into a nested LinkArtifacts behind Artifacts().
  ResetLinkArtifacts is now a worker-safe pure clear; the link-observable
  version bumps (backendState/link/uboContent) move to a GL-thread-only
  BumpLinkObservableVersions() called once from Link()'s prologue and from
  glProgramBinary's mandated failure - the link body never writes them, so
  a stage-4 worker cannot lose an invalidation against the draw path.
- ShaderObject: compile artifacts (TShader, preprocessed source, side-channel
  maps, status/log, consume-once flag) behind Compiled(); the P0b layer-1
  memo trio deliberately stays outside as the future non-joining
  COMPLETION_STATUS_KHR fast path.
- CompileEnv (new): a GL-thread snapshot of everything the compile pipeline
  used to read live from the backend mid-parse - compute limits (the
  GetIntegeri_v reach-back is gone from the worker path), advertised
  extensions, device quirks, TBuiltInResource inputs. Captured lazily per
  backend activation; the consume-once re-parse now runs against the same
  env as the original parse.
- The GL-thread prologue / worker-body boundary is marked in Link() where
  the stage sort ends; everything below is a pure function of the snapshot.

Public getter signatures unchanged - MG_Impl and both backends compile
untouched. Unit 476/476, Program suites 117/117, DirectGLES retrace 38/39 on
llvmpipe (the one failure is the known pre-existing non-CI iterationrp case;
the NVIDIA userspace driver was updated out from under the running kernel
module mid-session, so GLX there is down until a reboot).
2026-08-08 07:12:37 -04:00

166 lines
8.2 KiB
C++

// MobileGL - MobileGL/MG_State/GLState/ProgramState/ShaderPreprocessCache.h
// Copyright (c) 2025-2026 MobileGL-Dev
// Licensed under the GNU Lesser General Public License v3.0:
// https://www.gnu.org/licenses/gpl-3.0.txt
// https://www.gnu.org/licenses/lgpl-3.0.txt
// SPDX-License-Identifier: LGPL-3.0-only
// End of Source File Header
#pragma once
#include <Includes.h>
#include <list>
#include <mutex>
#include <MG_State/GLState/ProgramState/ShaderObject.h>
namespace MobileGL::MG_State::GLState {
// Where the shared, source-only half of ShaderObject::Compile() stopped. The two
// rejection verdicts are kept apart (rather than collapsed into "failed") so a hit
// reproduces the original diagnosis, not just the original info log.
enum class ShaderPreprocessOutcome : Uint8 {
// The source-only half ran clean; preprocessedSource and both maps are valid.
Preprocessed,
// ValidateComputeLocalSizeLimits rejected it (compute only).
ComputeLocalSizeRejected,
// FindReservedIdentifierViolation rejected it.
ReservedIdentifierRejected,
// The source-only half was clean but glslang rejected the preprocessed source.
// Memoizing this saves the parse itself on every later object with that source.
ParseFailed,
};
// Everything ShaderObject::Compile() derives from the source text alone, i.e.
// everything that is identical for two shader objects holding byte-identical source.
struct ShaderPreprocessResult {
ShaderPreprocessOutcome outcome = ShaderPreprocessOutcome::Preprocessed;
// Valid unless the preprocessor itself never ran; kept even for the rejection
// outcomes because that is the text the diagnostics refer to.
String preprocessedSource;
UnorderedMap<String, Int> explicitUniformLocations;
UnorderedMap<String, Uint> explicitOpaqueBindings;
// The compile info log to publish; empty when outcome == Preprocessed.
String infoLog;
Bool Preprocessed() const { return outcome == ShaderPreprocessOutcome::Preprocessed; }
};
// Cache hits hand out shared ownership, not a raw pointer into the entry list. That is
// what makes the cache safe once compiles run concurrently: a reader keeps its payload
// alive across any eviction, and a 107 KB preprocessedSource is never copied on a hit.
using ShaderPreprocessResultPtr = SharedPtr<const ShaderPreprocessResult>;
// P0b layer 2: a per-context, bounded memo of the source-only half of shader
// compilation, keyed by (stage, xxhash64(source), source length).
//
// Motivation: in the Iris shader-pack corpus ~21% of every glCompileShader in a trace
// is a *different* shader object holding byte-identical source (packs glue the same
// common/composite GLSL into many program stages), so the preprocess + reserved-
// identifier scan + explicit-location/binding extraction runs over the same megabytes
// again and again. Layer 1 (in ShaderObject) covers the same object recompiled with
// unchanged source; this covers the cross-object case.
//
// What is NOT cached: the glslang parse. glslang's TShader is consume-once (mapIO
// mutates the aliased intermediate at link), so every shader object still needs its
// own parse; only the text-processing half is shared.
//
// Correctness: the 64-bit hash is a lookup accelerator only. Every hit re-compares the
// full stored original source with memcmp before it is honored, so a hash collision
// degrades to a miss, never to a wrong answer. That is why the full original text is
// stored rather than a prefix/suffix digest - the cache is bounded, so the cost is.
//
// Eviction: FIFO (insertion order), bounded by BOTH an entry count and a stored-source
// byte budget, whichever binds first. FIFO rather than LRU because shader-pack loading
// is a burst of mostly-distinct sources whose reuse clusters around insertion time;
// LRU's extra list splice on every hit buys nothing measurable here, and FIFO keeps
// Find() a genuinely const, read-only operation.
class ShaderPreprocessCache {
public:
static constexpr SizeT kMaxEntries = 128;
static constexpr SizeT kMaxStoredSourceBytes = 8u * 1024u * 1024u;
// Returns the memoized result for this exact source under this exact compile
// environment, or null on a miss. The returned SharedPtr owns its payload, so it
// stays valid for as long as the caller holds it - across Insert(), Clear(), and
// across the destruction of the cache itself.
//
// envFingerprint joins the key because the source-only pipeline's compute
// local-size verdict is computed against CompileEnv's device limits: a memo must
// never outlive the environment it was computed against (memo-hazard rule).
ShaderPreprocessResultPtr Find(ShaderStage stage, Uint64 sourceHash, const String& source,
Uint64 envFingerprint) const;
// Memoizes `result` for this source. A source whose own storage cost already
// exceeds the byte budget is simply not cached (caching it would evict everything
// else and then itself).
void Insert(ShaderStage stage, Uint64 sourceHash, const String& source, Uint64 envFingerprint,
ShaderPreprocessResultPtr result);
void Clear();
static Uint64 HashSource(const String& source) {
return static_cast<Uint64>(XXH64(source.data(), source.length(), 0));
}
SizeT GetEntryCount() const {
const std::lock_guard<std::mutex> lock(m_mutex);
return m_entries.size();
}
SizeT GetStoredSourceBytes() const {
const std::lock_guard<std::mutex> lock(m_mutex);
return m_storedSourceBytes;
}
private:
struct Key {
ShaderStage stage = ShaderStage::Unknown;
Uint64 sourceHash = 0;
SizeT sourceLength = 0;
Uint64 envFingerprint = 0;
Bool operator==(const Key& other) const {
return stage == other.stage && sourceHash == other.sourceHash &&
sourceLength == other.sourceLength && envFingerprint == other.envFingerprint;
}
};
struct KeyHasher {
SizeT operator()(const Key& key) const {
// The source hash already spreads well; fold the two discriminators in so
// that same-hash-different-stage/length keys land in different buckets.
Uint64 mixed = key.sourceHash;
mixed ^= static_cast<Uint64>(key.sourceLength) + 0x9e3779b97f4a7c15ull + (mixed << 6) + (mixed >> 2);
mixed ^= static_cast<Uint64>(static_cast<Int>(key.stage)) * 0xff51afd7ed558ccdull;
mixed ^= key.envFingerprint + 0x9e3779b97f4a7c15ull + (mixed << 6) + (mixed >> 2);
return static_cast<SizeT>(mixed);
}
};
struct Entry {
Key key;
// The full original (pre-preprocess) source, kept so a hit can be confirmed by
// comparison instead of trusting the hash.
String originalSource;
ShaderPreprocessResultPtr result;
};
using EntryList = std::list<Entry>;
static SizeT EntryBytes(const String& source, const ShaderPreprocessResult& result) {
return source.length() + result.preprocessedSource.length();
}
void EvictUntilWithinBudgetLocked();
void EraseEntryLocked(EntryList::iterator it);
// P1: every public entry point takes this. The lock alone would NOT have been
// enough - the old Find() handed back a raw pointer into an entry that a
// concurrent Insert()'s FIFO eviction could erase while the caller was still
// reading it. Shared ownership of the payload is what closes that hole; the mutex
// only protects the containers below.
mutable std::mutex m_mutex;
EntryList m_entries; // front = oldest (FIFO victim)
UnorderedMap<Key, EntryList::iterator, KeyHasher> m_index;
SizeT m_storedSourceBytes = 0;
};
} // namespace MobileGL::MG_State::GLState