mirror of
https://github.com/MobileGL-Dev/MobileGL
synced 2026-09-09 20:58:31 +09:00
ES has neither glMultiDrawElements nor glMultiDrawElementsBaseVertex, so both are emulated. DirectGLES had two ways of doing it - one glMultiDrawElementsBaseVertexEXT where the driver has the extension interaction, otherwise a per-draw loop. This adds the five MobileGlues uses (gl/multidraw.cpp), so the ladder is now: one glMultiDrawElementsBaseVertexEXT; one glMultiDrawElementsIndirectEXT over a synthesized command buffer; one glDrawElementsIndirect per command over that same buffer; the base-vertex replay; plain glDrawElements over a CPU-rewritten index stream, for drivers with no base-vertex draw at all; and a compute shader that flattens the whole batch into one rebased index buffer drawn by a single glDrawElements. They live in their own translation unit that owns the entry point outright, preparation included - the compute tier has to dispatch BEFORE PrepareForDraw, or it would have to unpick the program, storage-block and index bindings the preparation just made, and a dispatch inside an open transform-feedback span is not legal at all. The auto ladder is ext -> basevertex -> multiindirect -> indirect -> drawelements, which is NOT MobileGlues' order (it puts the indirect tiers first). Measured on mc_sodium_multidraw, ns/op, median of three: NVIDIA ES 3.2 basevertex 2500 vs multiindirect 5700 and indirect 5800; Mesa llvmpipe ext 19300, basevertex 25200, multiindirect 27600, drawelements 28700, indirect 31000. Ring-allocating the command staging instead of respecifying per batch was tried first and moved the indirect tiers by less than noise, so the cost is the indirect draw path itself, not the upload; only a real multi-draw entry point beats replaying the sub-draws. auto therefore resolves to basevertex on this box - byte for byte the behaviour that shipped - and the new tiers are what a driver with the ext interaction, or without base vertex at all, now gets. compute is never chosen by auto (nor by MobileGlues'): it rewrites the primitive stream rather than replaying it, and it measured slowest here. Four places this deliberately does not follow MobileGlues, each a correctness bug there. A rewritten stream is emitted as GL_UNSIGNED_INT whatever came in, because GL adds baseVertex at full precision and folding it into ushort indices wraps. The restart sentinel is carried across a rebase unrebased, or an enabled primitive restart is lost. The flattening tier declines strip/loop/fan modes, any sub-draw whose count is not a whole number of primitives, and any batch at all while primitive restart is enabled - a restart ends a primitive, so leftover vertices would find a third vertex in the next sub-draw and become a triangle GL never draws. And the indirect tiers decline client-memory index arrays, which have no buffer to address. gl_DrawID gets better rather than worse: the unrolled tiers now feed each sub-draw its index (the spec's value, where the old loop left the uniform untouched), and a program that actually reads it demotes the batched tiers, which can only hold one value for the whole batch. The per-batch cost is nil for the programs that do not read it. Verified: the five DirectGLES retraces are byte-identical (md5) across all six tiers on NVIDIA and on Mesa, each tier proven to have really executed rather than silently demoted, via a per-tier announcement in the log. Unit suite 421/421. The full retrace suite's five failures all reproduce unchanged on a stashed tree, so none are new.
65 lines
3.8 KiB
C++
65 lines
3.8 KiB
C++
// MobileGL - MobileGL/MG_Backend/DirectGLES/MultiDraw.h
|
|
// Copyright (c) 2025-2026 MobileGL-Dev
|
|
// Licensed under the GNU Lesser General Public License v3.0:
|
|
// https://www.gnu.org/licenses/gpl-3.0.txt
|
|
// https://www.gnu.org/licenses/lgpl-3.0.txt
|
|
// SPDX-License-Identifier: LGPL-3.0-only
|
|
// End of Source File Header
|
|
|
|
#pragma once
|
|
#include <Includes.h>
|
|
#include <Config.h>
|
|
#include "DirectGLES.h"
|
|
|
|
// Emulation of the desktop glMultiDrawElements / glMultiDrawElementsBaseVertex entry
|
|
// points on OpenGL ES, which has neither in core.
|
|
//
|
|
// Every strategy below is an emulation; they differ only in which driver capability
|
|
// they lean on and in how many driver entries a batch of N sub-draws costs. The design
|
|
// follows MobileGlues (MobileGL-Dev/MobileGlues, gl/multidraw.cpp) tier for tier, plus
|
|
// the native GL_EXT_multi_draw_arrays interaction that MobileGL already had:
|
|
//
|
|
// Ext one glMultiDrawElementsBaseVertexEXT 1 driver entry
|
|
// MultiIndirect one glMultiDrawElementsIndirectEXT 1 driver entry + 1 upload
|
|
// Indirect N x glDrawElementsIndirect N + 1 upload
|
|
// BaseVertex N x glDrawElementsBaseVertex N
|
|
// DrawElements N x glDrawElements over CPU-rebased indices N + 1 upload
|
|
// Compute 1 x glDrawElements over a GPU-flattened, 1 dispatch + 1 entry
|
|
// rebased index stream
|
|
//
|
|
// Which one runs is resolved once per ES context from the driver's capabilities,
|
|
// capped by MOBILEGL_ESPRYT_MULTIDRAW_MODE, and can additionally be demoted per batch
|
|
// when the batch's own shape rules a tier out (see ResolveTierForBatch in the .cpp).
|
|
namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl {
|
|
// The tier this ES context resolved to, computed on first use and stable after.
|
|
MG_Config::GLESMultiDrawMode ResolvedTier();
|
|
// "multiindirect", "compute", ... - stable identifiers, also used by the POST row.
|
|
const char* TierName(MG_Config::GLESMultiDrawMode tier);
|
|
// One line naming the resolved tier, the tiers the driver can support, and the env
|
|
// clamp if one applied. For DriverPost and the startup log.
|
|
String DescribeTierResolution();
|
|
|
|
// The resolution itself, as a pure function of a capability set: the backend feeds
|
|
// it the live ES context's capabilities, DriverPost feeds it the ones it probed
|
|
// standalone, and both therefore report the same tier. `explanation`, when non-null,
|
|
// receives the "requested -> resolved (driver supports: ...)" line.
|
|
MG_Config::GLESMultiDrawMode ResolveTier(const MG_External::GLESCapabilities& caps,
|
|
const MG_External::GLESFunctionsTable& funcs,
|
|
MG_Config::GLESMultiDrawMode requested, String* explanation);
|
|
// Whether one tier is runnable on the given capability set, for per-row POST output.
|
|
Bool IsTierSupported(const MG_External::GLESCapabilities& caps, const MG_External::GLESFunctionsTable& funcs,
|
|
MG_Config::GLESMultiDrawMode tier);
|
|
|
|
// Runs `drawcount` indexed sub-draws as one glMultiDrawElements(BaseVertex) call
|
|
// would. `basevertex` is null for the plain glMultiDrawElements entry point (every
|
|
// base vertex is 0). Owns the whole draw, preparation included: callers must not
|
|
// have run PrepareForDraw, because the compute tier has to dispatch before the
|
|
// draw state is established.
|
|
void DrawElementsBatch(GLenum mode, const GLsizei* count, GLenum type, const GLvoid* const* indices,
|
|
GLsizei drawcount, const GLint* basevertex);
|
|
|
|
// The ES context is gone: every scratch buffer and the compute program belonged to
|
|
// it, so drop the names without deleting them (the dead context reclaims them).
|
|
void OnBackendContextDestroyed();
|
|
} // namespace MobileGL::MG_Backend::DirectGLES::MultiDrawImpl
|