diff --git a/perforce/agents/perforce.md b/perforce/agents/perforce.md index 602abba..156ef3b 100644 --- a/perforce/agents/perforce.md +++ b/perforce/agents/perforce.md @@ -95,6 +95,24 @@ shell environment: `p4 login` themselves in their terminal, then re-check `p4 login -s`. Never run `p4 login -p` (prints a reusable ticket ≈ credential), never read `P4TICKETS`/ `p4 tickets` into context, and don't suggest `p4 login -a` (all-hosts ticket). +- **`-Mj` writes errors to STDOUT, not stderr.** In tagged-JSON mode (`p4 -Mj`, + and `-Mj -ztag`) the client emits **every** record - including error records - + as JSON on stdout. stderr is empty. Read stderr alone and a failed command + gives you nothing but the shell's `exit status 1`, which names no cause at + all. Parse stdout instead: one JSON object per line, and the human message is + the `data` field of the first record whose `severity` is >= 3 (p4's scale is + 0 empty, 1 info, 2 warning, 3 failed, 4 fatal). Two traps in that stream - a + failing command often emits **successful** records before the error, so the + first record is not the error record; and `severity` appears as a number in + some versions and a quoted string in others, so parse both. Keep reading + stderr as a fallback: a connection-level failure never gets far enough to + produce a record. Without `-Mj` the behaviour is the ordinary one and errors + do go to stderr, so this applies specifically to the mode you reach for when + you want machine-readable output. This cost four separate misdiagnoses in a + real webhook-ingestion connector (a missing trust file and an expired ticket, + then two changelist sync failures) before it was found: every one of them + reported as `perforce: exit status 1`, and the table below was unusable + because the message it keys on never arrived. - **Distinguish failure classes** - retry/backoff only *network/server* failures, never *user* failures: @@ -113,6 +131,7 @@ shell environment: | `Client ... unknown` / `must create client` | workspace missing/misconfigured | user → stop | | `You don't have permission` | protections | user → stop (check `p4 protects -m`) | | (unfamiliar wording) | a **broker** may be rewriting messages | show the user the raw output | + | `exit status N` and nothing else | you are running `-Mj` and reading stderr; the real record is on stdout | **not** a p4 failure class at all - fix the caller (see above) before classifying | ## The read-only checkout model - game-critical diff --git a/unity/LEARNINGS.md b/unity/LEARNINGS.md index 737cf89..fadb170 100644 --- a/unity/LEARNINGS.md +++ b/unity/LEARNINGS.md @@ -126,11 +126,40 @@ No secrets - reference credential *locations*, never paste them. calls against it work fine, because Unity's live-control server is a loopback TCP port that doesn't care which session issued the connection. ⤳skill: on a build node reached only over SSH, don't propose the live-control path without - a persistence mechanism for the Editor itself. + a persistence mechanism for the Editor itself. **Validated 2026-09-23:** the + mechanism that actually works is a one-shot scheduled task - a `.cmd` wrapper + that runs the Editor and records its exit code to a marker file, launched with + `schtasks /create ... /sc once /st 00:00 /f && schtasks /run`, which runs + outside the SSH session's process tree; poll the marker and the log from a + fresh connection, then `schtasks /delete`. Verified end to end on Unity + `6000.3.22f1`: a cold import and a WebGL `-executeMethod` build both survived + the SSH session that started them. Two mechanisms that look like a fix are + not: `cmd /c start "" /min ` hangs the SSH session itself and never + launches anything, and `Start-Process` without `-Wait` returns a PID that + dies at session exit exactly like the inline case. Don't point the real run's + log at stdout (`-logFile -`) either - it's fine for a quick startup probe, but + if the reading pipe closes the Editor dies with it. ⤳skill: `unity-build` §8. - **The Hub's "add project" folder picker wants the parent directory, not the project directory** - it silently rejects the project folder itself and scans for children. Registering a project directly by path sidesteps the picker - entirely. + entirely. The same Hub's **"Add project from repository" flow only recognizes + three providers (Unity Version Control, GitHub, GitLab)** and clones the + repository root, not a subfolder - a poor fit for a project that lives at + `/unity/` rather than the repo root. Use "Add project from disk" (or + `unity projects add `) against an existing clone instead. +- **A synchronous headless Hub install undersells its own uncertainty, and its + progress text oversells completeness.** `Unity Hub.exe -- --headless install + --version --changeset -m ` blocks until done - about 9 + minutes for an Editor plus two modules on a wired gigabit test machine - but + Unity's CDN caps the download at roughly 12 MB/s regardless of link speed, so + don't infer link health from it. The "installed successfully" lines are + printed per component, not proof of anything: verify by listing + `Editor\Data\PlaybackEngines\` for the expected module folders rather than + trusting the summary (the progress output also carries raw ANSI cursor codes + that need stripping before grepping it). Once installed, real timings beat + the "up to an hour" folklore: a project with 328 scripts and 644 other assets + cold-imported in about a minute and a WebGL build finished in under 3 - + budget timeouts around that and treat 20+ minutes as a stall, not normal. - **Disk, not CPU, was the binding constraint on the primary test rig** - a single volume with limited free space before the install, consumed significantly by an Editor plus a probe project. A second machine on hand had much more disk but diff --git a/unity/agents/unity.md b/unity/agents/unity.md index a72bb44..b34fbeb 100644 --- a/unity/agents/unity.md +++ b/unity/agents/unity.md @@ -163,6 +163,7 @@ project-affecting decisions, not tooling conveniences - gate them. | Build/test fails to start; lock errors | another Editor/batchmode process holds the project (`Temp/UnityLockfile`) | environment -> gate 1: ask the user to close it, wait; never kill it | | CLI install fetches nothing / `latest.json` 404 | there is no stable channel - beta only | environment -> pin `UNITY_CLI_CHANNEL=beta` (`unity-cli` §2) | | Build works locally, build node fails with a missing-platform error | the target's module isn't installed for that Editor there | environment -> gated `unity install-modules` / `unity editors module` for that version | +| Editor started over SSH exits instantly with an empty or missing log | Windows OpenSSH killed the process tree at session exit | environment -> launch via a one-shot scheduled task, `unity-build` §8 | | "This project was created with a different version of the editor" (or silent reserialization) | Editor/project version mismatch | version -> confirm intent; a forward open+save is one-way - a named decision, not a default | ## How to reason about a request diff --git a/unity/skills/unity-build/SKILL.md b/unity/skills/unity-build/SKILL.md index f247f2a..bd0cb58 100644 --- a/unity/skills/unity-build/SKILL.md +++ b/unity/skills/unity-build/SKILL.md @@ -249,6 +249,52 @@ if any, test pass/fail counts, warnings - not "exit code 0". State the cost in the confirmation prompt; if a previous log exists, its timestamps beat these estimates. +## 8. Build node reached over Windows OpenSSH + +A build/test/run driven over Windows OpenSSH has one more failure mode on top +of everything above: **OpenSSH kills the whole process tree at session exit.** +A long Editor run started inline - `ssh node 'unity build ...'` or a raw +`Unity.exe -batchmode ...` - dies the instant the SSH session ends, with a +zero-byte or missing log. This reads exactly like "Unity failed to start"; it +is not - the Editor never got the chance to fail, it was killed. + +Two things that look like a fix and are not - don't propose either: + +- `cmd /c start "" /min .cmd` - hangs the SSH session itself and + never launches anything. +- `Start-Process` without `-Wait` - returns a PID immediately, but that + process dies at session exit exactly like the inline case (no log written). + +The pattern that works is a one-shot scheduled task: a `.cmd` wrapper that +runs the Editor and records its exit code to a marker file, launched via +`schtasks` so it runs outside the SSH session's process tree, polled from a +fresh connection, and cleaned up after: + +```bat +:: C:\Users\\.cmd +"" -batchmode -nographics -quit ^ + -projectPath -executeMethod Builder.PerformBuild ^ + -logFile C:\Users\\.log +echo EXITCODE=%ERRORLEVEL% > C:\Users\\.exit +``` + +```sh +schtasks /create /tn /tr C:\Users\\.cmd /sc once /st 00:00 /f && schtasks /run /tn +``` + +The `WARNING: Task may not run because /ST is earlier than current time` line +is harmless - `/run` fires the task immediately regardless. Poll the `.exit` +marker and the `-logFile` from a fresh SSH connection (or from WSL via +`/mnt/c/...`), and `schtasks /delete /tn /f` once done. + +This changes *how* the job runs, not what "done" means: still verify by +parsing the log and checking the artifact, never by trusting the marker +file's exit code alone (§6) - the exit-0 trap (§2) applies exactly as it does +to any other batchmode invocation. Background and evidence: `LEARNINGS.md` +§B (the process-tree kill, the two mechanisms that don't work, why CLI calls +against an already-running Editor don't need this - only starting the Editor +does - and the verified end-to-end sequence this section distills). + ## Validation status Command surface, flags, examples, and the licensing/exit-code behavior in this