[Fix, Test] (Espryt): drive a unit's sampler through the CSO's own twin on the handle arm - the record arm looked a content-addressed handle up in the identity-keyed registry and could never hit, so only the pre-handle program pass ever put a glBindSampler'd object on the driver (P4a seam F-4) - #17
Closed
Swung0x48 wants to merge 582 commits into
Closed
Swung0x48 wants to merge 582 commits into
Swung0x48 wants to merge 582 commits into
Conversation
…ailing fields and give CallMask, the server's backend type and the ABI fingerprint the carriers they never had - tableSlotMask could not address the 69-slot table its comment named and the two compute limits already ride inside dynamicParameters
…ndshake and both sessions' construction and teardown - the ABI assertion runs in Accept before a record is decoded and is Fatal{AbiMismatch} rather than a downgrade, and the two doorbell accessors live on the session rather than on ITransport
…s across two real threads, one case per watermark rule, the kRecPad rule against the session's own counter, the reply slot's seq stamp and a shutdown whose join is bounded at five seconds so a lost wakeup fails red
…7 internal formats is ~16 KiB of dense table, not half a megabyte
…edSeq has exactly one writer and it is the session, so the decoder keeps its own tally and the two are compared instead
…he live-host-writer fact is content-shaped, and MGPResourceDesc's only carrier costs a buffer a spurious reallocation on every map
…rsistent-map block tracker, on the client, where the decision SyncGpuWrites makes synchronously actually lives
… blocks, publish the live-host-writes edge, and let the writeback rather than the request clear a GPU write
…persistent-map acquisition under split, and retire the verify pin that was waiting for exactly this
… and refuse tier 1 of the flush ladder when the map and the queued range disagree in either direction
…he bytes the client ships, rewrite the inventory line that called it unwired, and correct the persistent-map census from 20 to 21
…re targets at glEndTransformFeedback, in place of a stall and an unbounded fence wait
…ites pin from both sides, the widened INVALIDATE_RANGE edge and a non-zero persistent-map-push
…ing the bytes spelled Pad - a codec that normalised one would silently delete the fields P5 and b1 have just put there
…ario landed in through a peek at IsBackendPersistentMapped, not by inferring it from the push byte counter - "the push was never wired" and "the store was adopted" are different defects with different owners and the counter reads the same for both
…t its OWN message - one of the eleven was a silent duplicate of another because the harness asked only whether something exited, and the harness now has a control of its own
… typo - every verb-shaped call now has a row or a named exemption, because a record with no stamp point applies under the previous verb's serial, mask and name
…rather than inherited from the client's fill, and name the argument in an argument-narrowed abort so the two halves of the pixel store do not print the same line
…wned server takes, and four named fields the stamp case must not re-derive from the table it is testing
…est in pipe-gates - R-7.1 rests on a generated table that nothing in CI diffed, and the build assertion only catches an unclassified row
… make an unset mask Fatal{UnsetCallMask} - the consumer default read MGPipeGetResourceOps, a process-wide global, so under inproc the server answered R-8's client-side gate with the client's own registration and under spawn it collapsed to P2's mask, dropping five P4a families with the dirty flags cleared anyway and the lane green
…deprecated instead of deleting them - a deletion FREES the vtable slot, so callMask took slot 14 from an [int] vector as an 8-byte ulong with no ABI-major bump and a fingerprint that mixes struct sizes rather than the schema
… adds a control page, a backwards watermark is Fatal rather than a permanent hang, RetireThrough stops at a borrowed slot, the inproc peer gets a real second mapping through dup and Adopt, and the producer remembers its own published seq
…shake guards, close the server session on every Start failure that ran after Accept succeeded, and drain teardown against the producer's own published seq rather than a watermark the ring lets lag
… frozen CapsSnapshot field ids, the verifiable frame with a null union, the borrowed slot that is not handed back, the peer's own mapping, the backwards-watermark abort, the lagging watermark and the refused reply geometry
…oducer writing SEG_STAGEs consumer cursors, bound the stage-mark queue, and answer a reply slot only for rows the catalogue flags kReplySlot
…cess - MG_Config::Transport, ClientSession::Active() and ImplementedVerbCount() - instead of from a grep over c0's stub messages, which a one-word rename disarmed into eight green monolith runs, and require the encoder's record ordinal to move before any split case may pass
…gative control it was cloned without, match ConfigLoader's whole INFO sentence rather than a KEY=VALUE an env dump also prints, read the three-arm names file the identity step was writing and discarding, and refuse an unknown manifest key that silently empties a CI matrix
… linear allocator and not a ring - drop the stage RingProducer and its power-of-two rounding, pin the stage cursor triple dead, require a wrap filler to carry kind kRingPadRecordKind as well as the flag it shares with kVarTail, and name every kRecBorrowSlot sighting
…applier entry points have always answered through, so the flags table is the single thing that sizes the reply pool and the decoder posts against
…ll seven pointer-backed residual rows - EmittedCallSuppliesTheWholeField answers false for eight rows; seven of them are pointer-backed and this residual fill is Magma's ONLY source for them, dereferenced on its first draw (Magma is lockstep for the whole phase, ID-90). NoP4aFamilyFieldIsWhollySupplied is keyed on the EMITTER'S SUBSYSTEM, so it covered five of the seven: GetBoundVertexArray (P3a's, BindVertexElements) and GetBufferBindingPoint (sb's, SetShaderBuffers) sit outside kMGPipeP4aFamilySubsystems and had no compile-time protection at all. - A package that tidied one of those two away would not break the build - it would SIGSEGV Magma on a device, which is the shape that already cost this phase 38 scenarios (ID-107). The seven are now named in one list with their own assertion, so the most likely mistake of every remaining package is a build break instead. - The comment says where the fired assertion sends you: a row retires in the phase that stops the backend reading a frontend object (P7/P8), never in a consumer-side cleanup. - Red-once: deleting the GetBoundVertexArray case fails the build - "PipeFill.cpp:2720:23: error: static assertion failed due to requirement 'NoPointerBackedRowIsWhollySupplied()': a pointer-backed residual row stopped being pulled (ID-112)". The pre-existing P4a assertion stayed silent on the same edit, which is the gap ID-112 names.
…ch program rows leave the frontend - PrepareForDraw's single GetProgramForDraw() and PrepareForCompute's GetProgramForDispatch() are the 68 red entries this package owns (61 + 7 in the base census; 62 + 7 once the PersistentMapArm entry is re-run under its own lane environment). Merely CALLING the accessor is what trips MGP_INPUT_CHECK, so the row retires by not making the call. - DrawProgramFromRecords() states the arm ONCE and as a CONJUNCTION: ProgramHandleArm() AND TextureImpl::UnitTexturesByHandle(). The hoisted value serves four callees and the arm has to name every one of them - the second conjunct is BindCurrentTextures', whose legacy memo key and whose sampledTargetForUnit lambda both read a real object. The family bits are independent A/Bs, so the
…lowlist is generated, not hand-kept - ra landed a hand-kept 13-row copy of CONTRACT-P5E section 7 in .github/workflows/test.yml, and it was wrong in BOTH directions at once: it omitted GetFramebufferBindingSlot@ReadPixels (21 entries, the lane's largest survivor) and carried GetTextureObject@CopyImageSubData, which resource_copy_region's kWaitNone makes a genuine unbarriered read of client memory. Neither error is visible by reading the list. Both are obvious once the list is derived. - gen_pipe_field_ownership.py now joins the three tables ID-84 already states the rule in - kMGPipeFieldOwnership, kMGPipeWaitClasses (generated/PipeWire.inc) and kMGPipeClassFieldMask / kMGPipeVerbClass (generated/PipeFillPoints.inc) - and emits kMGPipeAdmittedPullMask plus MGPipeBarrierPullAdmitted(field, verb): 101 admitted pairs across 12 barriered verbs. - The three inputs are READ, never restated, so the allowlist moves with the code: retiring a verb's wait class narrows it with no edit anywhere, and a new BARRIER_PULLED row widens it only on verbs that wait. A verb with no stamp row admits nothing - only MGP_VERB_OP_LIST's ops can ever be the record a pull is reported against. - --print-admitted gives the CI shell the SAME table the C++ answers from, which is what makes the hand copy unnecessary rather than merely corrected. - Three new negative controls (each source renamed away must stop the script - a missing source yields a silently EMPTY allowlist, which reads as rigour and is blindness) and five derivation controls that move an input and read the allowlist back. - Red-once: dropping the wait-class conjunct from build_admitted fails the self-test by name - "3 admission-derivation control(s) did not hold: unbarriering ReadPixels' op removes every @ReadPixels pair; unbarriering ResourceCopyRegion removes every @CopyImageSubData pair; no pair is admitted on an unbarriered verb (DrawArrays)" - and the new unit case fails with "GetBoundVertexArray@DrawArrays is admitted on a verb no barriered op stamps".
… of its own - MOBILEGL_IPC_STRICT_ERRORS=1 had no registered ctest entry anywhere: the phase's gate driver sets the knob by hand and pg's declared red-once for GetProgramForDraw never landed, so the two rows pa retires could come back and every gate would stay green. ID-109's finding, one scale down: the gap is a missing gated entry, not a missing case. - DirectGLES.Split.StrictProgramArm. registers the only two scenarios that are strict-CLEAN after pa - one per half. MultiDrawScenario.BaseVertexDrawsReject- MalformedArguments prepares a draw and asserts on GL errors rather than pixels; AtomicCounterScenario's three cases are the dispatch half. Everything else in the lane still ends in a readback, whose GetFramebufferBindingSlot@ReadPixels is an allowlisted row only P7 retires. - integration-split alone, and the prefix says which arm: the monolith arm of these scenarios is registered under DirectGLES. already and the knob is a no-op there, so labelling them integration-gpu would add entries that cannot fail for the reason they exist (R-16). - Negative control run: with the arm reverted all four abort, two with GetProgramForDraw@DrawArrays and two with GetProgramForDispatch@DispatchCompute. - Each entry takes a private log path through SplitLogPaths, for vi's reason: an entry that exists to carry a marker cannot share the file the marker is in.
…ains an admitted state
- CountBarrierPull aborted under MOBILEGL_IPC_STRICT_ERRORS on EVERY barrier-pulled read,
admitted or not. The run's rc is therefore non-zero before CI ever reaches its allowlist
comparison, so that comparison was unreachable code and "hard green" was impossible with the
knob as written - the lane could not tell a debt this phase leaves standing from a defect.
- Third state, with the other two untouched: unbarriered stays an unconditional Fatal; barriered
and NOT admitted stays a Fatal under strict; barriered and admitted logs one MGLOG_W and the
entry completes. Admission is MGPipeBarrierPullAdmitted, generated from the three tables
(ID-116), so there is no list here to drift.
- The marker keeps the grammar and changes only the tag: Admitted{UnmigratedPipeInput,
"<F>@<V>"} [BARRIER-PULLED, ADMITTED, retires in <phase>]. Every filter written since P5c
matches Fatal{UnmigratedPipeInput and therefore still means exactly "red".
- Deduped per (field, verb) per process. An 852-draw frame reaches an admitted readback row once
per draw; undeduped, one frame writes 852 identical lines and the marker census is unreadable.
- CHECKED, NOT ASSUMED (the brief's instruction): StrictErrorsTurnsABarrierPulledReadIntoANamedAbort
and StrictErrorsAlsoPromotesTheStickyForwards both still go red. Both stamp DrawArrays, whose
op is kWaitNone, so neither pair is admitted.
- ID-119: the old "rsp = 0 on unbarriered records" comment at this function is withdrawn as
vacuous - the unbarriered arm is [[noreturn]] and runs BEFORE ++g_residualPulls, so the
sentence was true of every possible implementation.
- Red-once: AnAdmittedBarrierPullIsLoudOnceAndNotFatal is the same sticky forward as the case
above it under a barriered verb; making the strict arm unconditional again kills it -
"FieldOwnershipTest.cpp:867: Failure / Actual: false / signal 6".
…tries of their own - The resolved multi-draw tier is ONE process-wide read of MOBILEGL_ESPRYT_MULTIDRAW_MODE (MultiDraw.cpp's ResolveTierOnce) and the d1 block leaves it at "auto", which on this driver lands on "ext" - the one tier that asks the bound index buffer nothing. Three of the five sites the previous commit migrated therefore had no entry in the set the gate runs: ID-107's shape one level down, with the env knob in place of MG_Config::Transport. - Five blocks, one per tier "auto" cannot reach (basevertex, multiindirect, indirect, drawelements, compute), prefix DirectGLES.Split.MultiDrawTier<Tier>., both labels, and MOBILEGL_ESPRYT_MULTIDRAW_MODE pinned as a ctest ENVIRONMENT property rather than a job-level export - review finding N-6's reason, and the PushMonolithArm block's precedent. - "ext" is deliberately NOT registered: the d1 block already IS the ext lane. Every other tier was confirmed to resolve to ITSELF here rather than falling back to ext, so no entry in this set is a green that silently means "ext again". - The filter is the tier-sensitive workload: MultiDrawScenario (minus the client-indices case d1 already excludes by name) plus IndexedDrawFamilyScenario's three MultiDraw cases. 12 entries per tier, 60 in total. - SplitLogPaths.cmake.in gives each of them a private log path. That is the rule every DirectGLES.Split. entry follows (SplitLogPaths.PrivateAndDistinct), and for these it also repairs the census blind spot: a third name segment means strict_census2.sh would re-run them without the very pin under test, so the marker is read where the lane wrote it. - integration-split 116 -> 176, integration-gpu 1294 -> 1354, both fully green. Under the strict lane all 60 stop on GetProgramForDraw@DrawArrays, which is pa's row: PrepareForDraw runs before the tier dispatch, so no new marker class appears and mv's row stays at 0.
… becomes barriered - CopyImageSubData's apply reads GetTextureObject, a BARRIER_PULLED row that retires in P7, not in P5e. resource_copy_region was kWaitNone, so after the flip that is a genuinely UNBARRIERED record reading client memory: an unconditional Fatal, no knob, standing for two phases. - The wait class moves rather than the field being special-cased into the allowlist. The allowlist is derived (ID-116), so an exception written into it would be a real defect kept quiet instead of a debt stated - and with the class moved the derivation admits the pair with no exception at all: 101 -> 111 pairs, CopyImageSubData joins the 12 barriered verbs. - kWaitApplied AND NOT kWaitReply: the row posts no reply slot, so the client parks until the record applies. That is exactly what copy_framebuffer_to_texture - the two framebuffer-sourced copies, same kBlitOrCopy verb class - already does. THE COST IS MEASURED: glCopyImageSubData is zero per frame in the Minecraft 26.3-rc-3 scene, so the wait is on no per-frame path. - PipeCatalogue's wait-class census moves with it, both halves in this commit: the named kWaitApplied set gains the row, the count goes 9 -> 10, and the kWaitNone offset 24 -> 25. - Red-once: putting the row back to kWaitNone drops CopyImageSubData out of the derived allowlist (111 -> 101 pairs, 0 @CopyImageSubData) and fails "FieldOwnershipTest.cpp:422: Failure / resource_copy_region is unbarriered again: after the flip, a copy record reading the client's texture object is an unconditional Fatal and P7 is two phases away", with PipeCatalogueTest red on three counts beside it.
… compile again Three sites, all pre-existing at 7c6f688, all the same disease: something guarded by MOBILEGL_BUILD_DISAGGREGATED used from code that is not. Nothing in the gate built either flavour, so nobody knew. - Managers.cpp: StagedTextureTargetForPipeTarget is a pure MGPipeResourceTarget -> TextureTarget switch with no wire, no record and no role in it, but it sat inside a nested #if MOBILEGL_BUILD_DISAGGREGATED while its only caller MGB_TEXPARAM_TARGET is guarded by #if MOBILEGL_PIPE_PUSH alone. Let out of the inner block without moving any text. The narrower fix - guarding the macro block too - is WRONG and was tried: that block also carries TextureDiagName and MGB_TEXTURE_RECORD_ARM_SELECTED, which the push flavour does need, and removing them broke it differently. - DirectGLES.cpp:5575: the storage-block-signature clause called ComputeShaderStorageBlockBindingSignatureOf, declared only under MOBILEGL_PIPE_PUSH, from an UNGUARDED condition list. The clause is push-only in substance too - the signature it compares is over the override map the wire carries. - DirectGLES.cpp BindCurrentTextures: keysMatch was declared inside #if MOBILEGL_BUILD_DISAGGREGATED while the ternary's else-arm and every use sat outside it. The legacy arm becomes a LAZY lambda, not an eager Bool: it reads currentProgram, which is null on the handle arm after pa, so the laziness the ternary gave for free is load-bearing (ID-81 / ID-110 - the arm decides, and the other arm's reads must not happen at all). Gate: flavours rc=0 - pull 915 TUs (push define absent), push 929 TUs (push=1), both arms asserted against compile_commands.json rather than against the flags, because the first version of this step passed -DMOBILEGL_PIPE_PUSH=OFF beside an option that implies it and reported a clean 982-TU "pull" build that was a second copy of the split build. The cache entry still read OFF. Split build unaffected.
…the retiring-phase disjunct - ID-116's static rule failed on a row no package in this phase may touch. pa retired the per-draw program pull, which unmasked GetTransformFeedbackProgram@DrawArrays (11 entries): BARRIER_PULLED, on kWaitNone draw_vbo, and ruled out of P5e entirely by CONTRACT-P5E section 5.7, with a retiring phase of "P3b/P4b (Espryt), P7 (Magma)". The rule was wrong, not the code. - A pull is now admitted if the field is BARRIER_PULLED, the verb has a stamp row, the field is in the verb class's may-read mask, and EITHER the verb's op is statically barriered OR the field's retiring phase does not name P5e. This table is only ever consulted on a record the server already stamped BARRIERED - CountBarrierPull's unbarriered arm is [[noreturn]] and fires first - so the pull is legal by construction and the only open question is whose debt it is. The retiring-phase column is exactly that answer and was already in the table. - A verb with NO stamp row still admits nothing under either disjunct: the server stamps only MGP_VERB_OP_LIST's ops, so any other bit would be a row nothing can exercise. 111 -> 174 pairs. - The phase column is matched as a TOKEN, not a substring, so a "P5" or a future "P5e2" cannot answer for "P5e". - All six of ID-125's survivors checked: GetFramebufferBindingSlot@ReadPixels and GetTextureUnitObject@CopyTexImage2D by disjunct 1; GetTransformFeedbackProgram@DrawArrays (P3b/P4b), GetTextureObject@CopyImageSubData (P7) and ValidateProgramName@ShaderStorageBlockBinding (P9) by disjunct 2. GetTextureObject@CopyImageSubData holds by BOTH, asserted separately. GetProgramForDraw@DrawArrays, GetBoundVertexArray@DrawArrays and GetProgramForDispatch@DispatchCompute stay rejected. - Magma's GetFramebufferBindingSlot@Clear is NOT admitted and cannot be: the retiring-phase column is one string for both backends and it names P5e for Espryt. The DirectVulkan entries get their own lane label instead (item 7), which keeps the Espryt allowlist about Espryt. - StrictErrorsAlsoPromotesTheStickyForwards is RENAMED to StrictErrorsReachTheStickyForwardsAndNameThemByField and restated. Not one of the seven sticky forwards names P5e (P7, P7/P9, P7/P13, P9), so on a stamped verb they are all admitted now. The claim that still has teeth is that the detector REACHES them - asserted by rsp moving and by the marker naming the field and its phase, so "did not abort" cannot pass for "never detected". The Fatal half keeps its own case next door on GetBoundVertexArray, a P5e row. - Six new derivation controls, including one that makes GetTransformFeedbackProgram this phase's debt and requires it to leave the unbarriered verb's allowlist - without which disjunct 2 could be a constant `true` and every other control would still hold.
… - admission gains the escalation disjunct
- The integrator's census over the merged tree left two markers nothing admitted, both
DirectGLES.Split.ClientVertexArrayScenario: GetBoundVertexArray@DrawArrays. Disjunct 1 fails
(draw_vbo is kWaitNone) and disjunct 2 fails (that row's retiring phase names P5e, correctly,
for the ORDINARY draw path vi and mv retired). But these two entries pull the pair on a record
the client IS parked behind: MGPipeBarriered escalates a draw carrying client vertex arrays
(ID-83), ID-82 puts client arrays outside P5e entirely, and the read site
(SyncClientSideVertexArraysForDrawArrays) aborts by its own name if it is ever applied
unbarriered. The pull is legal; no static table could see it.
- Third disjunct: the record was barriered BY ESCALATION, i.e. wireSaysBarriered(op, payload,
applierState) && MGPipeWaitClassFor(op) == kWaitNone. Both halves were already computed on
every record - ApplyOne computes the predicate unconditionally for ID-103's reason - so this is
one more bit stamped beside the barriered stamp, not new plumbing.
- It is a SEPARATE flag rather than a refinement of the barriered stamp, because the two answer
different questions: the stamp answers "is the client parked" (the allocator and registry
guards ask that), this answers "and for a reason the static table cannot see" (only the strict
knob asks). Its default is false - the answer that admits nothing - which is the opposite of
the stamp's default and deliberately so.
- The marker says WHICH disjunct: ADMITTED when the generated table answered, ADMITTED-ESCALATED
when only the record did. That keeps the lane's comparison against --print-admitted exact
instead of widening the static list with every pair that could ever escalate - and the tag sits
in the slot the grammar already has for the reason, so Fatal{UnmigratedPipeInput filters are
untouched.
- Escalation also happens to be the truer justification for GetTransformFeedbackProgram@DrawArrays,
which ID-125 admits by phase. Both now hold; ID-125's disjunct is kept, as ruled.
- Red-once: AnEscalatedRecordsPullIsAdmittedAndSaysWhy. Dropping `&& !escalated` from the strict
arm kills it - "FieldOwnershipTest.cpp:938: Failure / Actual: false / signal 6" with
'Fatal{UnmigratedPipeInput, "GetBoundVertexArray@DrawArrays"} [BARRIER-PULLED,
MOBILEGL_IPC_STRICT_ERRORS=1, retires in P5e (Espryt unbarriered), P7 (Magma)]'. Its control is
StrictErrorsTurnsABarrierPulledReadIntoANamedAbort, which drives the identical pull with the
flag at its default and must keep aborting.
…P5e ID-119/ID-115 - the lane reads every log, counts, and has a positive control
Three measured defects in the strict lane, one commit each in spirit and one in fact because
they are the same step.
1. THE LOG SET CAME FROM A DIRECTORY. The CI step grepped MG_IntegrationTest/split-logs/, which
holds 90 of the lane's entries. The ones it could not see were the F1. readback block, both
NamedBlit pairs, the four Ct. entries and PersistentMapArm - precisely the readback population
the allowlist is ABOUT, plus the only DirectVulkan entries in the lane. split_log_paths.py
gains a `markers` mode that walks ctest's own {entry: MOBILEGL_LOG_FILE_PATH} map and
classifies every Fatal{ and Admitted{ line. Measured on this head: 113 of 113 entries read.
The lane is now a TWO-SIDED RATCHET - an unadmitted marker fails, an admitted pair not in
Harness/strict-expected-markers.txt fails, AND an expected pair that no longer appears fails,
so the lane cannot rot green by outliving its own debts.
2. THE `rsp` CHECK GREPPED FOR A COUNTER NO SPLIT ENTRY EMITS, and the pin behind it was vacuous:
CountBarrierPull's unbarriered arm is [[noreturn]] and runs BEFORE ++g_residualPulls, so
"rsp = 0 on unbarriered records" was true of every possible implementation. Replaced by an
assertion where the counter actually is - one entry, DirectGLES.Split.StrictArming., with
MOBILEGL_PIPE_STATS=1, MOBILEGL_PIPE_STATS_PERIOD=1 and a private log path of its own (a
whole-lane env flip races under -j: the library opens that path "w"). It asserts
rsp == 0 || rsp >= draws, which is the device evidence's shape - rsp ~= the draw count under
inproc, 0 under monolith - so the draw path's retirement is visible as rsp ceasing to scale.
3. THE LANE HAD NO POSITIVE CONTROL (ID-115). Of the seven entries that passed strict, three were
monolith transport - where nothing stamps a verb boundary, so the whole mechanism is
structurally unreachable - two self-skipped, one was a death test and one was a Python check.
Not one was a record-carrying split GL scenario, so "green" and "strict was never armed" were
the same observation. New push-only counter `vbs` (PipeStats::ServerVerbBoundaries), counted
at MGPipeServerStampVerbBoundary behind Enabled(), and the same entry asserts vbs > 0 on a
frame that drew. It is deliberately NOT rsp: rsp reaches zero when the phase SUCCEEDS, so a
control built on it would start failing on the day the debt is paid.
- Measured on this head (base 7c6f688, before pa and mv): isplit 119/119; under strict the
census is Fatal{ GetProgramForDraw@DrawArrays 52, GetBoundVertexArray@DrawArrays 9,
GetProgramForDispatch@DispatchCompute 7 - exactly what pa and mv owe - and Admitted{ on ten
pairs, three of them ADMITTED-ESCALATED. The expected-marker file carries all ten and says in
its header that the wave-3 merge retires three of them.
- CONTRACT-P5E §7 rewritten around what is now checkable: the derived allowlist and its three
disjuncts, the two-sided ratchet, the restated rsp pin, and the positive control.
…lit entries get a lane of their own
- Both DirectVulkan.Split.NamedBlit entries carried LABELS "integration-split" and abort under
strict on GetFramebufferBindingSlot@Clear (VulkanRenderer.cpp:7648). P5e does not unbarrier that
field on Magma - it retires there in P7 - so the Espryt lane could not go hard green with them
in it.
- ARGUED RATHER THAN ASSUMED, because the brief allowed either reading: admitting the row instead
is the WORSE option. FieldOwnership.def carries ONE retiring-phase string for both backends
("P5e (Espryt unbarriered), P7 (Magma)"), so a rule that admitted the pair for Magma would admit
it for Espryt too - and an Espryt Clear reading the frontend's binding slot again is exactly the
regression fb's handle arm exists to prevent. The lane split costs two entries; the admission
would cost the gate.
- THE LABEL IS `integration-magma-split`, NOT `integration-split-magma`, and the spelling is
load-bearing: `ctest -L` takes a REGEX, so `-L integration-split` matches a label spelled
`integration-split-magma` as a substring. Measured with the obvious name first: both DirectVulkan
entries were still listed by `ctest -N -L integration-split`, i.e. the split looked done in every
listing and had not happened.
- Its CI step asserts the run is RED **and** that a private log carries the named marker, so
"Magma is expected red" cannot decay into "Magma did not run". Measured: 0/2 under strict with
two logs carrying Fatal{UnmigratedPipeInput, "GetFramebufferBindingSlot@Clear"}, and 2/2 green
without the knob. integration-split is 117 and integration-magma-split is 2.
- CONTRACT-P5E section 7 states the split and the reason it is not an allowlist entry.
…rward clause is marked UNLANDED - CONTRACT-P5E section 8 item 8 says GetBufferBindingPointCount's sticky forward becomes FATAL under split. It is not in the code and must not be put there this phase: FieldOwnership.def:147 (field row) and :179 (forward row) both say BARRIER_PULLED, and the generator refuses a field/forward pair that disagrees - pinned by FieldOwnershipTest.TheSevenStickyForwardsAgreeWithTheirFieldRows. - Landing it as written would either trip the generator or drag the FIELD row to FATAL with it, and that second reading moves 23 pairs of the derived admitted set: the row's retiring phase is "P7/P13", which is exactly what admits it on every stamped verb under ID-125, and a FATAL row is admitted nowhere - so 23 reads this phase never promised to migrate would start aborting. - MEASURED, NOT INHERITED: ID-120 said 10 pairs; that count predates ID-125's retiring-phase disjunct. `--print-admitted | grep -c '^GetBufferBindingPointCount@'` says 23 on this head, and the note records the command as well as the number. - Annotated in place, struck through with the phase that lands it (P7/P13), for ID-105's reason: a contract clause that is arithmetically impossible against the code is a clause that changes, and the divergence is better stated than discovered from a red gate two phases from now.
…rker set, re-measured on the merge The file carried gl's base measurement (7c6f688, before pa and mv landed) and said so rather than hiding it. The merge discharged that note exactly as written, and the two-sided ratchet named both deltas by itself rather than being told: - GetProgramForDraw@DrawArrays no longer appears - pa retired it, 62 entries to 0, and it was the read that fired once per draw and kept every draw barriered. Removed here, in the wave that retired it, which is the whole point of the other side of the ratchet: an expected set that keeps rows nothing writes any more is how a lane rots green. - GetBufferBindingSlot@DrawArrays is new, 18 entries, all of them the two indirect multi-draw tiers. mv NAMED it as a latent row before it could fire - BoundDrawIndirectBufferId's GetBufferBindingSlot(DrawIndirect), which "would surface as GetBufferBindingSlot@DrawArrays if an indirect tier were ever forced under strict" - and then gated the five tiers, which forced it. It retires in P8 (FieldOwnership.def:78), so it is a debt owed by a later phase and not by this one. Added with that reason. Measured on the integration tree, two independent instruments agreeing (Harness/split_log_paths.py markers, and a separate per-entry census that drives each entry's own command and ENVIRONMENT from ctest --show-only=json-v1): integration-split 181 / 181 green Fatal{UnmigratedPipeInput 0 distinct pairs Admitted{ 10 distinct pairs - 8 by the generated table, 2 by escalation So "the strict lane is hard green" is now a statement a script evaluates, not one a reader argues: every entry completes, no Fatal marker over the log set ctest itself declares, every Admitted marker checked against the same table the C++ answers from, and the ratchet fails in BOTH directions. Gate on this head: build rc=0, unit 2254/2254, integration-split 181/181, integration-gpu 1359/1359, gens all rc=0, flavours rc=0 (pull 915 TUs push-absent, push 929 TUs push=1).
…stants (NOT FOR MERGE YET) kMGPipeP5eRunAheadReady (MG_Backend/Init.cpp) and kMGPipeP5eClientWaitRuleLanded (MG_Pipe/PipeApply.h) both to true. They are ONE switch with two spellings and the integration commit throws both. THIS BRANCH IS NOT MERGEABLE AS IT STANDS and is parked so the defect it exposes can be worked on against a reproducible head. Measured on this commit: build rc=0 · unit 2254/2254 · gens rc=0 · flavours rc=0 (pull 915 TUs, push 929 TUs) integration-split 112 / 181 (69 failed) integration-gpu 1290 / 1359 (69 failed) and the failures are FLAKY: 71 failed at -j 4, 68 at -j 1, and a single entry run alone passes and then aborts on a repeat. A census over all 181 entries under each entry's own command and ENVIRONMENT gives 70 red on 17 distinct <field>@<verb> pairs, over just FOUR fields: 27 GetRenderStateParametersVersion 25 GetTextureContextId 16 GetBufferBindingSlot 2 IsCapabilityEnabled spread across eleven verbs - Clear, DrawArrays, DrawElements, DrawElementsBaseVertex, DrawElementsIndirect, DrawElementsInstancedBaseVertex, EndTransformFeedback, GetIntegeri_v, GetSyncStatus, MultiDrawElementsBaseVertex, ReadPixels. Several of those verbs are BARRIERED, which is what makes this interesting: 16 of the 70 abort on the UNBARRIERED arm (why=UNBARRIERED) and the other 54 are the plain unfresh-read abort with no [BARRIER-PULLED] bracket at all, i.e. a RECORD-SUPPLIED field that the freshness stamp says was never filled. One field at one verb, consistently, would be a missing migration. Four fields scattered over eleven verbs, flaky run to run, including barriered ones, is a RACE - and the obvious candidate is that gPipeInputs is shared mutable state whose mutual exclusion WAS the lockstep. The client still fills it for barriered verbs while the apply thread is draining earlier unbarriered records, so the client's fill and the server's verb-boundary stamp can now overlap in time. That is consistent with P5d's own research conclusion, recorded in the roadmap, that retiring this barrier needs gPipeInputs VERSIONING or double-buffering, deferred to P11 after P3b/P4b's handle-keyed twin tables. The strict lane being hard green (181/181, zero Fatal) is necessary and NOT sufficient, and this commit is the proof: strict runs under lockstep, where ApplyOne stamps every record barriered, so it cannot see anything that only happens when the client stops waiting. The lane measures remaining DEBT; it does not measure whether run-ahead is correct.
…fill must be quiescent, not merely about to park
- ID-132's finding, diagnosed: with the flip thrown the 70 red lane entries are ONE race
and two debts. The race is the client's residual fill writing gPipeInputs while the apply
thread is inside an earlier UNBARRIERED record. Measured interleaving, verbatim:
`[MobileGLIntegra/ERROR] RA2-PROBE fill verb=DrawElements wireop=81 waitclass=-1
applierInside=1` then `[mgl-srv-apply/ERROR] RA2-PROBE unfresh field=IsCapabilityEnabled
verb=DrawElements serverStamped=0 ownership=RECORD-SUPPLIED` then
`[mgl-srv-apply/FATAL] MGPipe: Fatal{UnmigratedPipeInput, "IsCapabilityEnabled@DrawElements"}`.
The GL thread's MGPipeServerClearVerbBoundary withdrew the applier's own stamp mid-record.
- The marker's verb is the CLIENT's verb, which is what identifies the writer without a
probe at all: the applier can only ever put a verb of MGP_VERB_OP_LIST into the block, and
40 of the 54 unbracketed aborts named a verb that is not in it.
- A/B control on the same binary and the same records: MOBILEGL_IPC_RUN_AHEAD=0 is 8/8 green
where the default arm is 2/6 red, so the defect is in the wait rule and not in a record.
- RefusePipeInputsTouchWhileApplierOwnsIt could not fire on any of it: its run-ahead arm read
`if (isBarrieredFill) return;`, and that sentence is a claim about the FUTURE ("this thread
is about to park behind the record"). The order at the validate point is fill, then emit,
then park. So the arm now tests the fact - `isBarrieredFill && !ApplyThreadIsInsideApplier()`
- and the fill sites MAKE it true first with §2.5's forced wait (QuiesceApplierBeforeFill).
- TWICE PER VERB, because the block is written in two phases that straddle publication: the
serial bump / stamp withdrawal / verb rename run before the fourteen emitters and the
63-field walk of step 4 runs after them, and the records those emitters published are
records whose apply reads the block.
- ClientVerbIsBarriered: MGP_VERB_OP_LIST joins the whole draw family through one row
(DrawVbo -> DrawArrays), so inverting it verb-first answered kOpCount for DrawElements and
eighteen other draw verbs, and the "unknown verb answers barriered" default filled for each
while the record it published was a kWaitNone draw_vbo it never parked behind. The family's
row is applied to the family; every other unknown verb keeps the conservative answer, which
the wait above now makes honest. Measured alone it is worth 112 -> 133 isplit; with the wait
in place the lane is 161 either way, so this half is the PERFORMANCE half - without it every
indexed draw takes a forced drain and run-ahead is lockstep by another name on the draw path.
- Red-once A: put the guard's arm back to `if (isBarrieredFill) return;` and
RemoteGuards.ClientBarrieredFillWhileTheApplierIsInsideUnderRunAheadIsFatalByName exits 0
instead of aborting (`Value of: DiedOfAbort(child) Actual: false`), while its green control
keeps passing. Red-once B: delete the kDraw fallback and
RemoteRunAhead.AnIndexedDrawVerbIsUnbarrieredAndTouchesPipeInputsNotAtAll dies of signal 6
with `Fatal{RoleViolation, "gPipeInputs"} - the GL thread touched gPipeInputs
(MGPipeValidateForVerb) on a RUN-AHEAD session in a barriered fill that did not first wait
for the applier to catch up`.
- Both new cases run in the ordinary unit lane with run-ahead armed by the fixture, which is
what ID-132(1) asks for: the strict lane runs under lockstep and cannot see this class.
- Gate on the flipped head: unit 2256/2256, isplit 161/181 (was 112), gpu 1339/1359 (was 1290),
gens rc=0, flavours rc=0. The remaining 20 are two named debts, not this race: 18 are
`GetBufferBindingSlot@DrawArrays [UNBARRIERED]` on the two indirect multi-draw tiers, whose
retiring phase is P8, and 2 are ID-82's own refusal of client vertex arrays.
- The arm is stated and not inferred (ID-81): WaitForApplyToCatchUp's first test is
`m_runAheadArmed && m_started`, which reduces to Transport != Monolith. Under lockstep,
under monolith and under MOBILEGL_IPC_RUN_AHEAD=0 every line added here is a call that
returns, so those arms are unchanged.
… ra2 - ID-133's escalation, ID-134's lane, and the contract sentence that could not be checked
- ID-133: escalation (iii) in MGPipeBarriered - a draw_vbo with NumDraws > 1 that is NOT
kDrawIsIndirect is barriered. Espryt's indirect multi-draw tiers reach
MultiDrawImpl::RunIndirect, whose apply reads the client's GL_DRAW_INDIRECT_BUFFER binding
(BoundDrawIndirectBufferId, MultiDraw.cpp:87) - GetBufferBindingSlot, retiring in P8 - and
unbarriered that is Fatal{UnmigratedPipeInput, "GetBufferBindingSlot@DrawArrays"} on 18 lane
entries. Backtrace of the aborting thread, verbatim frames: CountBarrierPull <-
MultiDrawImpl::RunIndirect <- MultiDrawImpl::DrawElementsBatch <- ServerVerbSink::OnDrawVbo
<- PipeWireDecoder::DecodeAndApply <- PipeApplier::ApplyOne.
- TWO PREMISE CORRECTIONS, both measured, both in the report and in the contract. (a) There is
no indirect multi-draw OP: all twenty draw entry points collapse onto draw_vbo, and the tier
is resolved on the server per batch (ResolveTierForBatch), so the escalation cannot be keyed
on the tier and its cost is NOT confined to the opt-in tier lanes - a plain glMultiDraw* now
waits on the auto -> ext default too. NumDraws > 1 is the narrowest wire fact that contains
the reaching set; kDrawIsIndirect is excluded so Sodium's indirect draw, which resolves by
handle and pulls nothing, keeps running ahead. (b) The read is a SAVE/RESTORE of a GL binding
name around the tier's own scratch buffer, not a data dependency - so it can be retired
outright by giving BoundDrawIndirectBufferId the handle arm its neighbour
ResolveBoundIndexBuffer already has, and this escalation withdrawn. That is three lines in
MultiDraw.cpp, which ID-113 gives to mv, so it is reported rather than edited here.
- The client's half is verb-keyed because the fill decision is asked before the record exists:
{MultiDrawArrays, MultiDrawElements, MultiDrawElementsBaseVertex} is exactly the set whose
record can carry NumDraws > 1 without kDrawIsIndirect. The inclusion goes that way round on
purpose - a record the server calls barriered whose fields the client did not fill would be
an ADMITTED pull of a stale value, which is worse than the abort.
- ID-134: the two ClientVertexArrayScenario entries get the label integration-clientarrays-split
and an expected-red CI step that asserts the MARKER, shaped like gl's Magma lane. ID-131 heeded
by COUNTING, not by assuming: -L integration-split 181 -> 179, -L integration-gpu 1359 -> 1357,
the new label selects exactly 2. The step branches on whether run-ahead is actually ARMED,
because kMGPipeP5eRunAheadReady is a BUILD constant: on a head where it is false the refusal is
unreachable and the entries are legitimately green, so an unconditional demand for red would
fail for the one reason that is not a defect. The client says which world it is in in its own
log line; delete the green branch when the flip lands for good.
- CONTRACT-P5E §3.5 amended to the code: "outside a barriered fill" rested on the caller's claim
about the FUTURE, and a guard whose exemption cannot be false on the class it exists for is not
a guard - it never fired on any of the 70 red entries. The landed rule is "not while the applier
is inside, barriered fill or not", with the fill sites taking §2.5's forced wait first, twice
per filling verb because the block is written either side of the emitters.
- AND §3.2 IS MARKED UNLANDED, which nobody had noticed: "the server's stamp is server-private
… never into the shared block" describes storage the applier owns, and
MGPipeServerStampVerbBoundary still writes m_filled / m_currentVerb / m_serverStampedVerb into
gPipeInputs itself. That gap is what made the race possible. Splitting the stamp is the P11
item - but per-role stamps WITHOUT versioning the BARRIER_PULLED values would be worse than
today, because it turns a loud Fatal into a silent stale read.
- §2.1 carries escalation (iii) as a row, per its own rule that a new escalation is a ruling and
a row in the contract rather than a condition at a call site.
- Gate on the flipped head: unit 2256/2256, isplit 179/179 HARD GREEN, gens rc=0, doc citations
rc=0 (the three remaining are CONTRACT-P5.md's and pre-date this branch).
…- retire the indirect-binding pull instead of escalating for it, and withdraw escalation (iii)
- TOUCHING mv's FILE UNDER ID-136, deliberately and by ruling: MG_Backend/DirectGLES/
MultiDraw.cpp is mv's per ID-113, and the integrator overrode that boundary for this one
change because the two halves must not exist apart on any head. The boundary has not drifted.
- BoundDrawIndirectBufferId (MultiDraw.cpp) gets the handle arm its neighbour
ResolveBoundIndexBuffer already had, in mv's own idiom rather than a second one beside it:
the arm is `MG_Config::Transport != Monolith`, STATED (ID-81) and never inferred from a null
handle (ID-110); a null handle is "this verb bound no indirect buffer", the same 0 the
monolith arm returns for an empty slot; a missing backend resource is the named refusal
Fatal{RoleViolation, "multidraw-indirect-buffer-arm"}, a sibling of mv's
RefuseMissingIndexBufferRecord, never a quiet fall-back to the frontend. The monolith body
below it is unchanged token for token. The function moved down beside BoundIndexBufferId,
which is the same question about the other target.
- WHY A HANDLE ARM AND NOT A MIGRATION: the tier never reads a byte of the client's indirect
buffer. It binds its OWN scratch command buffer and wants to put back the name that was
there - a save/restore. With a transport the only other writer of this process's
GL_DRAW_INDIRECT_BUFFER is DirectGLES.cpp's DrawSyncBit::IndirectBuffer arm, which binds from
MGPipeApplier().VerbIndirectBuffer and nothing else, so "what was bound" IS "what the verb's
record named".
- Escalation (iii) withdrawn in the same commit - MGPipeBarriered's clause, the client's
verb-keyed half in ClientVerbIsBarriered, and the contract row. It was never wrong, it was
the wrong instrument: there is one draw opcode and the tier is chosen server-side per batch,
so the narrowest key both roles could compute charged EVERY plain glMultiDraw* on the
auto -> ext default - a per-batch rendezvous on the arm the phone ships, added to make a lane
green. THE CHECK ON AN ESCALATION IS "WHICH ARM PAYS FOR IT", NOT "WHICH LANE GOES GREEN";
the rule is recorded in PipeApply.h beside the three that remain and in CONTRACT-P5E §2.1.
- Red-once D: `if (false && ...)` on the new arm and the entry dies 3/3, deterministically, on
`Fatal{UnmigratedPipeInput, "GetBufferBindingSlot@DrawArrays"} [BARRIER-PULLED, UNBARRIERED,
the client did not fill it, retires in P8 (indirect), P9 (readback), P13 (transfer)]`.
Its counterpart is this head itself: the escalation is GONE and the lane is green, so nothing
needed it - and the strict census's admitted set drops from nine rows back to ID-128's eight,
because the pull no longer happens at all rather than being legalised.
- Gate on the flipped head: unit 2256/2256, isplit 179/179, gens rc=0, lane counts unchanged
(179 / 1357 / 2 / 2).
- CURRENT_STAGE_PROGRESS.md 新增 2.6 节:四个包(pa / mv / gl / ra2)、采纳规则的三个析取项 及其各自是被哪条真实条目逼出来的、覆盖面从 116 涨到 179 条 + 1357 条、翻开开关后暴露的竞态 与其根因,以及"strict 硬绿是必要而不充分"这条(ID-132)。代码头更新为 25fba0d。 - MEASUREMENTS.md 新增 11 节:设备矩阵。该节先讲定频,因为第一轮矩阵没定频就作废了——monolith 自己那一臂逐帧 CPU 在两次之间 4.43 -> 7.07 ms,而慢的那次温度更低,是 governor 把 policy0 从 2745 拉到 748 MHz。附 pin_redmi.sh 的写序与回读验证,以及为什么 pin_device.sh 拒绝猜本机 serial 是对的。结论:两个定频档下 monolith 与 run-ahead 都撞 120 Hz 面板上限,只有 lockstep 撞不到。 - ROADMAP.md 的 P5e 行与状态行改为"十二包全落地、run-ahead 已武装"。 阶段仍未收官:面板上限之上谁更快没有回答(本机定不住更高频率),BRIEF-P5E 3 节的三条阴性对照与 5 节的逐包 red-once 重跑按 ID-121 已重排到车道变绿之后,尚未执行。
11 节的矩阵在 VD12 下跑,而 monolith 与 run-ahead 都撞到了 120 Hz 面板上限,所以那一节回答的是 "能不能撞到上限"而不是"谁更快"。把渲染距离拉到 32 后负载涨到 ~3550 draws/帧(4.2 倍),两条臂都 远离任何上限(游戏自身 enableVsync:false / maxFps:260,实测 max 仅 37-71),这才是可比的一轮。 - 每臂五次、交错、CPU 定频(跑前跑后各 check 一次,均 PINNED)、GPU 钉 pwrlevel 0、风扇开, 进世界后静置 75-150 s 再取 40 s 窗口——VD32 的区块流式加载远比 VD12 长,沿用 20 s 会把加载期 算进稳态。 - monolith p50 中位 58.5(范围 56.3-67.9),inproc run-ahead p50 中位 59.9(范围 56.8-61.1), client CPU 16.19 vs 15.72 ms/帧。齐平,run-ahead 略前,且离散度只有 monolith 的一半。 - 相对它取代的 lockstep:p50 +69%、client CPU -40%、apply CPU -43%。负载越重越值钱,符合预期: 会合次数随 draw 数走,VD12 每帧约 900 次,VD32 约 3600 次。 - 两条臂都是 client 线程打满(约 94%),所以 client ms/帧 这一列是各自真正的瓶颈,可直接比。 新增的 11.1 节另记一条机会而非缺陷:apply 线程 11-12 ms 对 client 15.7 ms,两侧不平衡,把工作 从 client 挪到 apply 会直接降瓶颈。这在 lockstep 下没有意义(两侧串行相加),是 run-ahead 解锁 出来的方向,留给后续阶段。 options.txt 已从 .p5e-backup 还原为 renderDistance:12;设备已解频、风扇关、GPU 恢复 governed。
- 新增 P5E-RUNAHEAD.md:P5e 的阶段报告,与 P5D-INPROC-PERFORMANCE.md 同类。含十二包清单、三类 缺陷及其教训(未被门禁的运行时臂 / 未被门禁的构建 flavour 且第一次补门禁是假绿 / 一个永远打不响 的守卫)、侧写与两档设备数字、未完成项、以及 run-ahead 解锁出来的下一个机会。 - CURRENT_STAGE_PROGRESS.md:阶段状态表的 P5e 行改为十二包 + run-ahead 已武装 + ID-80..136;新增 2.7 节 P5e 落地内容表(与 P5c/P5b 同形制);2.5 节改标题为历史;修掉两个都编号为 4 的小节; 开放项补 P5e 五条;下一步 item 1 从 c0e 改写为收官所缺的证据;裁定索引从 ID-79 延到 ID-136 的 承重条目。 - ROADMAP.md:P11 行加 P5e 的更正——退役 draw lockstep 没有需要 gPipeInputs 版本化,当时那次竞态 在元数据标量上;该行对一般情形仍成立,契约 3.2 的 per-role stamp 仍未落地。 - README.md:状态行加 P5e 与 run-ahead,并把 下一个是 P6 改为 P5e 之后才是 P6。 doc 引用检查:3 条问题全部是 CONTRACT-P5.md 的既有项,与本次改动无关。
上一次只把状态格改成"十二包全落地、run-ahead 已武装",而范围格与出口门格仍写着八包、19 条裁定、 "集成 commit 翻 kMGPipeP5eRunAheadReady",以及一份在开工前拟的出口门——其中数条后来证明按原样 不可执行。三格都重写: - 范围格:十二个包分三波(含 fix1 / pa / mv / gl / ra2),裁定 ID-80..136,两个常量均已翻;并点出 最大的那次迁移(逐 draw 的 GetProgramForDraw)没有改动任何线上结构——记录早就带着逐 link 的 ProgramArchive,缺的是仍去问前端对象的读者。链到 P5E-RUNAHEAD.md。 - 出口门格:改成实测已达成的部分(unit 2256、integration-split 179/179、integration-gpu 1357/1357、 flavours 三个且每臂断言生效的编译宏、strict 硬绿零 Fatal、设备 VD32 与 monolith 齐平、相对 lockstep p50 +69%),加上尚未执行的部分(三条阴性对照、逐包 red-once 重跑、E1 对照自检失败、 G1 仍由 CI 断言),再加上已作废的判据(rsp = 0 按构造为空,ID-119 改写;VERB_BARRIER=0 不得做 配对对照,ID-114,它会顺手关掉 run-ahead)。 - 债务表的 rsp 行:口径已被 P5e 重写,改成硬绿车道与八行 Admitted,并说明原判据为何是空的—— CountBarrierPull 的未设障臂在 ++g_residualPulls 之前且是 [[noreturn]]。 - 债务表的未迁移行:client vertex arrays 在 run-ahead 下是具名拒绝(ID-82),并有自己的预期红 车道 integration-clientarrays-split(ID-134);staging 仍归 P8。 - 开放问题新增 18:split 两侧负载不平衡(VD32 下 apply 11-12 ms vs client 15.7 ms,client 打满 约 94%)。把工作从 client 挪到 apply 会直接降瓶颈——lockstep 下毫无意义(两侧串行相加), run-ahead 才使它成为一个杠杆。需先重新侧写:等待消失后尾部各项的分母都变了。 - 头部状态行补上 P5E-RUNAHEAD.md 的链接。 doc 引用检查:docs/Disaggregated 七个文件 58 条引用,0 problem。
BRIEF-P5E 3 节 item 4 要求跑三条阴性对照。跑起来才发现 E1(MOBILEGL_IPC_VERB_BARRIER=0)身上叠了
三层问题,前两层把第三层遮住了。这次修掉前两层,并把第三层的诊断收窄成可执行的描述。
一、选中了一个在本车道永远跑不起来的用例
E1 的选择正则是 (Triangle|ClearThenReadPixels),而 gl 在 ID-115 下新增的阳性对照
TriangleScenario.TheServerStampedAVerbBoundaryOnThisDrawingFrame 正好被它匹配到。那个用例要从车道
自己的私有日志里读 PipeStats 窗口,所以只在 strict-arming 车道里跑,在别处**按设计**跳过,与这个
旋钮无关。而 ID-62 规定"被选中的条目被跳过"是硬失败——这是对的,一个会悄悄丢掉受试对象的对照什么
也证明不了。于是 E1 报"the knob killed the pre-flight, not the entry",而旋钮其实在起作用:同一次
运行里另有四条被选条目带着 abort 变红了。
修法是排除一个在这里永远跑不起来的用例,而不是教 E1 容忍跳过(那等于把 ID-62 扔掉)。
ctest -R 是 POSIX ERE,没有负向先行断言,所以 run_control 增加第五个位置参数:一个真正的 ctest -E。
二、ctest 与日志助手对"谁被选中"的看法不一致
run_control 把 -E 传给了 ctest,但 split_log_paths.py 仍只按 -R 过滤,于是它把被排除的条目算作
选中、再报告它们"did not run"。助手现在读同一个排除式(SPLIT_LOG_EXCLUDE)。用环境变量而不是再加
一个位置参数,是因为该助手各 mode 的位置参数布局不同;而一个对照的两半对自己的受试集合意见不一,
正是这个脚本存在要防的那类缺陷。
三、真正的病灶,ID-122 当初说对了,现在有了更锐利的说法
修掉前两层之后 E1 仍然红,且红在"10 selected entries did not fail":14 条里只有 4 条变红,
**而且整次运行里 Fatal{BarrierViolation} 出现 0 次**——那 4 条是因为别的原因红的。
也就是说 E1 索要的证据从来没有出现过。ClientSession.cpp:878 的那条 Fatal 断言的是"与 apply 线程
观察到的重叠",而 MOBILEGL_IPC_BATCH_WAITS 默认值 1 下 barrier 是在下一个读残余的 verb 上取的,
那个重叠根本不发生。
更进一步:run-ahead 武装之后,VERB_BARRIER=0 还会顺手关掉 run-ahead(它是 RunAheadArmed() 的第一个
合取项,ID-114),所以这个旋钮产生的那条臂既不是 lockstep 也不是 run-ahead——用它做任何一边的阴性
对照都不成立。
E1 需要的是**重新定义它要证伪什么**,不是再调阈值。留作具名未决项,写进 ID-122 与开放项。
本次不动 E1 的计数规则("全部选中条目都必须变红"):那是当初为了挡住"任何非零退出都算数"的伪绿而
刻意设的,在知道 E1 该证伪什么之前放宽它,只会换一种伪绿。
之前标"进行中"是因为 BRIEF-P5E 3 节的出口门还没跑,不是因为工作没做完。现在跑了,逐条记在
ROADMAP 的 P5e 出口门格里:
- item 1:unit 在 strict 两臂均 2256/2256。
- item 2/3:integration-split 179/179;逐条普查零 Fatal{,Admitted{ 恰为八行。
- item 4(a):RUN_AHEAD=0 作为配对对照 179/179 绿——同一份服务端代码、同一批记录,唯一差别是客户端
等不等。
- item 4(b):VERB_BARRIER=0 变红(车道 62/179)。
- item 4(c):Magma 逐字记下 "run-ahead requested, server does not publish kCapRunAheadApply -
running lockstep",Espryt 对照记下 "run-ahead ARMED"。两条都从条目自己的私有日志里取的。
- item 5:在集成树上重跑 red-once,用 ID-121 重排后的形式——倒掉 pa 的臂得 GetProgramForDraw@
DrawArrays,倒掉 mv 的臂得 GetBoundVertexArray@DrawArrays,**恰好各一个具名对**,还原后各自转绿。
- item 7:MOBILEGL_IPC_AUDIT=1 179/179。
- item 6 的 G1 仍由 CI 断言(ID-123)。
唯一未过的是 E1 对照脚本本身(ID-122),且诊断比原来锐利:修掉它两个遮蔽性缺陷之后仍然红,而它索要
的 Fatal{BarrierViolation} 在整次运行里出现 0 次——BATCH_WAITS 默认值 1 下那个重叠根本不发生;且
run-ahead 武装后 VERB_BARRIER=0 会顺手关掉 run-ahead(ID-114),它产生的那条臂既不是 lockstep 也不是
run-ahead。E1 要的是重新定义它证伪什么,不是调阈值。
所以状态改为"已收官,附一条具名未决",而不是无条件的已收官:把一条没跑通的对照藏在"收官"后面,
正是这个阶段反复记的那类错误。
两件事一起落:P6 开工前的计划与契约草稿,以及把此前只存在于一台机器上的工作笔记收进仓库。
一、P6 的计划与契约草稿(均未开工,一行代码没写)
docs/Disaggregated/P6-SPAWN-PLAN.md 包计划
docs/Disaggregated/P6-CONTRACT-DRAFT.md 契约草稿(c6 落地时移为 MG_Remote/CONTRACT-P6.md)
调查基线头之后最要紧的三个结论,都写进了文档:
1. P6 比路线图那一格小。传输原语已经在树上并且有测试:SocketDoorbell(含挂断检测与两道
防丢唤醒的 fence,Doorbell.cpp:211-360,FdPassingTest 四例)、SCM_RIGHTS fd 传递、
ShmSegment::Adopt、protocol.fbs 里已定好的 SurfaceOp / WindowKind。inproc 走真实的
第二份 shm 映射而不是别名,SessionRings.h:22-28 说明当初就是为了今天不是第一次跑。
P6 不是"写一个 socket 传输",是装配 + 造控制面 + 造进程。
2. "P6 只是传输替换"这句话的证据已经过期。它是 P5c 花一个阶段换来的,而那次审计在一棵
看起来已经很干净的树上找出 59 处直接内存读写。P5e 之后树又变了。所以第一个包 a6 是
只读审计,不写代码;契约里凡依赖它的行都显式标成 PENDING a6,而不是先替未来下结论。
3. P6 之后手机上跑不了游戏。ANativeWindow* 在子进程里没有意义,protocol.fbs:186-196 的
注释自己写着 Android 的窗口 transfers out of band——那就是 P12。P6 落 pbuffer /
surfaceless / 离屏(够 HeadlessGL、等价 split 车道与 trace replay,也正是 P6 出口门
量的东西),真窗口到达 server 时具名拒绝。单列一节,因为"P6 完了就能入世界"是这份
计划最容易被读成的意思。
契约加一条规则 G:角色不得指名进程局部句柄。它是 P5c 规则 E(角色不得指名对方的内存)在
身份上的延伸——EGLDisplay / EGLSurface / EGLContext / ANativeWindow* / fd 号 / pid / 映射
地址,要么作为本契约定义的令牌过线,要么不过线;fd 只经 SCM_RIGHTS,由内核重新编号。
P5e 给 P6 留下的新问题单独成节(§5):闩锁语义 ARCHITECTURE.md:504 已经设计好,照抄;新的是
run-ahead 之后客户端不再等待而对端进程会死——它手里会有一批已发布、永远不会被 apply 的记录,
而 Android 的 freezer 会冻住子进程且不产生挂断事件。所以"慢"与"死"不能用超时阈值区分。
这是 P6 里唯一的新设计,其余都是搬运,文档里这么写了。它也正好和 E1 对照未决项(ID-122)
共用同一个谓词,两者建议一起想。
出口门里加了两条 P5e 学来的:G2 要求 integration-spawn 的用例名集合与 integration-split
逐名相同(名集合不同即意味着偷偷少跑);臂证明门要求每条用例在自己的私有日志里留下子进程
pid 与 transport=spawn——ID-124 已经证明这类假绿真会发生,而一条实际跑了 monolith 的 spawn
车道能通过其余所有门。S2 阴性对照的形状要先想清楚再写,理由写在文档里:E1 就是前车之鉴。
二、工作笔记移进仓库(docs/Disaggregated/notes/,415 份 .md)
ROADMAP / CURRENT_STAGE_PROGRESS / MEASUREMENTS 大量引用 ~/w7/notes/ 下的裁定编号、审计行号
与逐门数字出处,而那个目录只存在于一台机器上、且未被任何版本控制跟踪。一份被引用几百次却
只在本地的记录等于没有记录,所以整体收进来,目录结构原样保留。
只移 .md。脚本、运行日志、普查与门的结果 JSON、补丁、APK 与 nm 转储留在原处:它们是证据数据
而不是文档,体积大、每次运行重新产生,且多数带绝对路径与设备序列号,只在产生它的机器上有意义。
映射规则与这条取舍写在 docs/Disaggregated/notes/README.md。
两个例外也记在那里:P6 的 BRIEF 已提升为正式文档,本目录不留副本以免两份同源文档各自漂移;
recovered/ 是从失败 workflow 里捞回的材料,与同名阶段目录可能重复,保留是因为其中几份是当时
唯一的完整版本。
CI 的文档引用 lint 只吃 docs/Disaggregated/*.md(非递归),notes/ 不在其中,这是有意的:
这些笔记的 file:line 写在各自当时的代码头上,拿今天的 HEAD 解析必然大面积失败,而那不是错误,
是它们作为历史记录的正确状态。核对某份笔记用它自己声明的 revision(--rev)。
三、顺带修掉的不一致
- README.md 的状态行仍写着"P5e 进行中",而阶段表已是已收官。
- CURRENT_STAGE_PROGRESS 的"下一步"第 1 条仍把出口门列为待办,其中 (a)(b) 上一个 commit
已经跑过了;改写为已收官 + 实际剩下的那一条(E1)。
- ROADMAP 的 P6 行此前三格仍是开工前的计划,现在填上包序、落地边界与新增的两条门。
逐条引用过 check_doc_citations.py:docs/Disaggregated/*.md 共 125 处引用,0 问题。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.