From 4991ee7f6db5b31fca8506d7c8e2d3bc6b2ecd57 Mon Sep 17 00:00:00 2001 From: "claude[bot]" <3247365+butterstack[bot]@users.noreply.github.com> Date: Mon, 21 Sep 2026 19:32:02 +0000 Subject: [PATCH 1/2] perforce: mirror the -Mj stdout-not-stderr agent instruction The private repo's agent instructions for perforce learned that tagged-JSON mode (p4 -Mj) writes every record, including errors, to stdout and leaves stderr empty - so an agent that classifies failures by reading stderr gets nothing but a bare exit status. Mirrors that guidance into the public perforce agent instructions verbatim, including the parse recipe, the two traps in the JSON stream, and a new row in the failure-class table for the "exit status N and nothing else" symptom. This is an instruction-file change (agents/perforce.md), not a LEARNINGS.md entry, so it is mirrored per DISTILLING.md section (f) rather than compressed: no step or table row was dropped. The private source names an internal connector by name; this generalizes it to "a real webhook-ingestion connector," consistent with how the same validation story is already described elsewhere in this file's LEARNINGS.md. Co-Authored-By: Claude Fable 5 --- perforce/agents/perforce.md | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/perforce/agents/perforce.md b/perforce/agents/perforce.md index 602abba..156ef3b 100644 --- a/perforce/agents/perforce.md +++ b/perforce/agents/perforce.md @@ -95,6 +95,24 @@ shell environment: `p4 login` themselves in their terminal, then re-check `p4 login -s`. Never run `p4 login -p` (prints a reusable ticket ≈ credential), never read `P4TICKETS`/ `p4 tickets` into context, and don't suggest `p4 login -a` (all-hosts ticket). +- **`-Mj` writes errors to STDOUT, not stderr.** In tagged-JSON mode (`p4 -Mj`, + and `-Mj -ztag`) the client emits **every** record - including error records - + as JSON on stdout. stderr is empty. Read stderr alone and a failed command + gives you nothing but the shell's `exit status 1`, which names no cause at + all. Parse stdout instead: one JSON object per line, and the human message is + the `data` field of the first record whose `severity` is >= 3 (p4's scale is + 0 empty, 1 info, 2 warning, 3 failed, 4 fatal). Two traps in that stream - a + failing command often emits **successful** records before the error, so the + first record is not the error record; and `severity` appears as a number in + some versions and a quoted string in others, so parse both. Keep reading + stderr as a fallback: a connection-level failure never gets far enough to + produce a record. Without `-Mj` the behaviour is the ordinary one and errors + do go to stderr, so this applies specifically to the mode you reach for when + you want machine-readable output. This cost four separate misdiagnoses in a + real webhook-ingestion connector (a missing trust file and an expired ticket, + then two changelist sync failures) before it was found: every one of them + reported as `perforce: exit status 1`, and the table below was unusable + because the message it keys on never arrived. - **Distinguish failure classes** - retry/backoff only *network/server* failures, never *user* failures: @@ -113,6 +131,7 @@ shell environment: | `Client ... unknown` / `must create client` | workspace missing/misconfigured | user → stop | | `You don't have permission` | protections | user → stop (check `p4 protects -m`) | | (unfamiliar wording) | a **broker** may be rewriting messages | show the user the raw output | + | `exit status N` and nothing else | you are running `-Mj` and reading stderr; the real record is on stdout | **not** a p4 failure class at all - fix the caller (see above) before classifying | ## The read-only checkout model - game-critical From 483c223ef04babba1d3c9733568e4162171153c7 Mon Sep 17 00:00:00 2001 From: "claude[bot]" <3247365+butterstack[bot]@users.noreply.github.com> Date: Mon, 28 Sep 2026 20:36:46 +0000 Subject: [PATCH 2/2] unity: distill the schtasks OpenSSH fix and Hub 3.x headless-install learnings The private repo validated the fix for the OpenSSH process-tree-kill trap already noted in this file: a one-shot scheduled task launched via schtasks runs outside the SSH session's process tree and survives it, verified end to end with a cold import and a WebGL build on Unity 6000.3.22f1. Also records the two mechanisms that look like fixes and are not (cmd start hangs the session, Start-Process without -Wait dies at session exit anyway), Hub 3.53's headless install behavior (CDN rate cap, per-component progress text that isn't proof of a finished install, real cold-import/build timings), and the narrow provider list on Hub's "Add project from repository" flow. Mirrored (not compressed, per DISTILLING.md section f): unity-build SKILL.md gains section 8 with the wrapper/schtasks recipe verbatim, and agents/unity.md gains the matching symptom-table row. The private source's numbered LEARNINGS.md cross-reference (B8/B10/B12) is rewritten to the public file's unnumbered convention (LEARNINGS.md paragraph B), consistent with how unity-pipeline/SKILL.md already references this file. Sanitization: the private source names the rig `beast` in the Hub-install example; genericized to "a wired gigabit test machine," consistent with how the existing "primary test rig" bullet in this file already anonymizes the same machine. Final grep sweep on the changed lines: zero hits. No secrets or credential values were present in the source material. Excluded from this pass: the private repo also added two entirely new plugins, teamcity and buildkite, since the last distillation. Per the rationale already recorded on this PR, bringing a new plugin into this repo needs more than the skills/agents/commands/LEARNINGS files this job is scoped to (plugin.json, hooks.json, the guard script, README, and marketplace registration all need to land together, plus a security review of the guard scripts) - recommend a separate, deliberate distillation pass for those once reviewed. Line counts: unity/LEARNINGS.md 170 -> 199, unity/agents/unity.md 231 -> 232, unity/skills/unity-build/SKILL.md 260 -> 306. Co-Authored-By: Claude Fable 5 --- unity/LEARNINGS.md | 33 ++++++++++++++++++++-- unity/agents/unity.md | 1 + unity/skills/unity-build/SKILL.md | 46 +++++++++++++++++++++++++++++++ 3 files changed, 78 insertions(+), 2 deletions(-) diff --git a/unity/LEARNINGS.md b/unity/LEARNINGS.md index 737cf89..fadb170 100644 --- a/unity/LEARNINGS.md +++ b/unity/LEARNINGS.md @@ -126,11 +126,40 @@ No secrets - reference credential *locations*, never paste them. calls against it work fine, because Unity's live-control server is a loopback TCP port that doesn't care which session issued the connection. ⤳skill: on a build node reached only over SSH, don't propose the live-control path without - a persistence mechanism for the Editor itself. + a persistence mechanism for the Editor itself. **Validated 2026-09-23:** the + mechanism that actually works is a one-shot scheduled task - a `.cmd` wrapper + that runs the Editor and records its exit code to a marker file, launched with + `schtasks /create ... /sc once /st 00:00 /f && schtasks /run`, which runs + outside the SSH session's process tree; poll the marker and the log from a + fresh connection, then `schtasks /delete`. Verified end to end on Unity + `6000.3.22f1`: a cold import and a WebGL `-executeMethod` build both survived + the SSH session that started them. Two mechanisms that look like a fix are + not: `cmd /c start "" /min ` hangs the SSH session itself and never + launches anything, and `Start-Process` without `-Wait` returns a PID that + dies at session exit exactly like the inline case. Don't point the real run's + log at stdout (`-logFile -`) either - it's fine for a quick startup probe, but + if the reading pipe closes the Editor dies with it. ⤳skill: `unity-build` §8. - **The Hub's "add project" folder picker wants the parent directory, not the project directory** - it silently rejects the project folder itself and scans for children. Registering a project directly by path sidesteps the picker - entirely. + entirely. The same Hub's **"Add project from repository" flow only recognizes + three providers (Unity Version Control, GitHub, GitLab)** and clones the + repository root, not a subfolder - a poor fit for a project that lives at + `/unity/` rather than the repo root. Use "Add project from disk" (or + `unity projects add `) against an existing clone instead. +- **A synchronous headless Hub install undersells its own uncertainty, and its + progress text oversells completeness.** `Unity Hub.exe -- --headless install + --version --changeset -m ` blocks until done - about 9 + minutes for an Editor plus two modules on a wired gigabit test machine - but + Unity's CDN caps the download at roughly 12 MB/s regardless of link speed, so + don't infer link health from it. The "installed successfully" lines are + printed per component, not proof of anything: verify by listing + `Editor\Data\PlaybackEngines\` for the expected module folders rather than + trusting the summary (the progress output also carries raw ANSI cursor codes + that need stripping before grepping it). Once installed, real timings beat + the "up to an hour" folklore: a project with 328 scripts and 644 other assets + cold-imported in about a minute and a WebGL build finished in under 3 - + budget timeouts around that and treat 20+ minutes as a stall, not normal. - **Disk, not CPU, was the binding constraint on the primary test rig** - a single volume with limited free space before the install, consumed significantly by an Editor plus a probe project. A second machine on hand had much more disk but diff --git a/unity/agents/unity.md b/unity/agents/unity.md index a72bb44..b34fbeb 100644 --- a/unity/agents/unity.md +++ b/unity/agents/unity.md @@ -163,6 +163,7 @@ project-affecting decisions, not tooling conveniences - gate them. | Build/test fails to start; lock errors | another Editor/batchmode process holds the project (`Temp/UnityLockfile`) | environment -> gate 1: ask the user to close it, wait; never kill it | | CLI install fetches nothing / `latest.json` 404 | there is no stable channel - beta only | environment -> pin `UNITY_CLI_CHANNEL=beta` (`unity-cli` §2) | | Build works locally, build node fails with a missing-platform error | the target's module isn't installed for that Editor there | environment -> gated `unity install-modules` / `unity editors module` for that version | +| Editor started over SSH exits instantly with an empty or missing log | Windows OpenSSH killed the process tree at session exit | environment -> launch via a one-shot scheduled task, `unity-build` §8 | | "This project was created with a different version of the editor" (or silent reserialization) | Editor/project version mismatch | version -> confirm intent; a forward open+save is one-way - a named decision, not a default | ## How to reason about a request diff --git a/unity/skills/unity-build/SKILL.md b/unity/skills/unity-build/SKILL.md index f247f2a..bd0cb58 100644 --- a/unity/skills/unity-build/SKILL.md +++ b/unity/skills/unity-build/SKILL.md @@ -249,6 +249,52 @@ if any, test pass/fail counts, warnings - not "exit code 0". State the cost in the confirmation prompt; if a previous log exists, its timestamps beat these estimates. +## 8. Build node reached over Windows OpenSSH + +A build/test/run driven over Windows OpenSSH has one more failure mode on top +of everything above: **OpenSSH kills the whole process tree at session exit.** +A long Editor run started inline - `ssh node 'unity build ...'` or a raw +`Unity.exe -batchmode ...` - dies the instant the SSH session ends, with a +zero-byte or missing log. This reads exactly like "Unity failed to start"; it +is not - the Editor never got the chance to fail, it was killed. + +Two things that look like a fix and are not - don't propose either: + +- `cmd /c start "" /min .cmd` - hangs the SSH session itself and + never launches anything. +- `Start-Process` without `-Wait` - returns a PID immediately, but that + process dies at session exit exactly like the inline case (no log written). + +The pattern that works is a one-shot scheduled task: a `.cmd` wrapper that +runs the Editor and records its exit code to a marker file, launched via +`schtasks` so it runs outside the SSH session's process tree, polled from a +fresh connection, and cleaned up after: + +```bat +:: C:\Users\\.cmd +"" -batchmode -nographics -quit ^ + -projectPath -executeMethod Builder.PerformBuild ^ + -logFile C:\Users\\.log +echo EXITCODE=%ERRORLEVEL% > C:\Users\\.exit +``` + +```sh +schtasks /create /tn /tr C:\Users\\.cmd /sc once /st 00:00 /f && schtasks /run /tn +``` + +The `WARNING: Task may not run because /ST is earlier than current time` line +is harmless - `/run` fires the task immediately regardless. Poll the `.exit` +marker and the `-logFile` from a fresh SSH connection (or from WSL via +`/mnt/c/...`), and `schtasks /delete /tn /f` once done. + +This changes *how* the job runs, not what "done" means: still verify by +parsing the log and checking the artifact, never by trusting the marker +file's exit code alone (§6) - the exit-0 trap (§2) applies exactly as it does +to any other batchmode invocation. Background and evidence: `LEARNINGS.md` +§B (the process-tree kill, the two mechanisms that don't work, why CLI calls +against an already-running Editor don't need this - only starting the Editor +does - and the verified end-to-end sequence this section distills). + ## Validation status Command surface, flags, examples, and the licensing/exit-code behavior in this