Skip to content

Base seam: fix 6 real defects, and rebuild the base suite as a reproducible 34-feature catalogue - #170

Merged
xywang68 merged 8 commits into
masterfrom
harden-seam-stdout
Sep 22, 2026
Merged

xywang68 merged 8 commits into
masterfrom
harden-seam-stdout

Conversation

@xywang68

@xywang68 xywang68 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up: the log now reflects the test it ran

Auditing the suite's own narration (not just its assertions) found the cmd: line reporting intent rather than execution. Six fixes:

# Defect found in the log Fix
1 Not runnable. json-on-error printed findTargetImage DISPLAY=:77 --imagePath=Screen — env assignment in argument position, so pasting it ran a successful screen read instead of the error path. natives/pointer printed parenthesised prose. Every seam call goes through target(), which prints the exact quoted argv and the exit status of the call it just made. Host commands go through probe(), printed then run from the same string.
2 Under-reported. image-similarity, text-hint, ocr-similarity, flash, latency ran 2–11 invocations but printed one command — the rejection half of the floor tests was invisible. target() is the single funnel, so one line per invocation is structural, not discipline.
3 False coverage claim. ocr-similarity was named for the OCR floor but asserted neither direction of it. See below — the truth is that the floor is not applied, and the feature now says so with a known-gap: line.
4 Names over-claimed. provenance said "records os + java + node" but grepped only built=; wm used pgrep -f openbox, which matches its own command line, so it could pass with no WM; json-on-error's "payload is a JSON array" grepped [, which the JVM's own [DEBUG STARTUP] lines satisfied vacuously. provenance asserts all five keys; wm checks the openbox PID we started plus pgrep -x; json-on-error parses the payload and asserts a status.
5 No evidence on success. 80-odd ✓ lines were names, so a green run could not be audited. ✓ <what> = <observed> on every passing check.
6 Order dependence. Features relied on the root window surviving from a previous feature. The fixture is re-shown before each invocation; timing features opt out with NOSHOW=1 and say so, since the re-show lands inside the measured window (it inflated latency from ~1130 ms to 1675 ms).

The finding behind #3

While verifying the floor I measured "0.99 rejects" — then found it was an artifact of a WM-less harness where the fixture had vanished (final screen read: empty). Under the suite's own setup (openbox running) fixtures are durable, and the truth is the opposite: --ocrSimilarity is accepted but never applied. With the text on screen, 0.99 still returns a match. The image floor is applied (image-similarity asserts accept and reject). Asserting an OCR rejection would have passed only when the screen happened to be blank — a false green — so it is not asserted; the report prints a known-gap: line instead, and the README lists it.

Two bugs I introduced and fixed

  • target()'s log line went to stdout, so inside $( ) it was captured into the JSON variable and corrupted the payload (all three checks went red). Log lines now go to stderr; stdout carries only the payload. That's the same class as the JVM-log interleaving fixed earlier in this branch.
  • The fixture re-show inflated the latency measurement; excluded via NOSHOW=1.

The flash threshold was also recalibrated to the real separation — 15–284 ms while the flash was broken, 780–1010 ms when applied — with an upper bound added to catch a runaway flash. It had failed at 783 < 800 on a loaded host.

Verified

make base-test                 -> feature conformance: 92 passed, 0 failed, ALL FEATURES OK
make one FEATURE=flash         -> 2 passed   (delta 903 ms, bounded)
make one FEATURE=ocr-similarity-> 2 passed, with the known-gap line printed

Also in this push: the banner reports the image's own build stamp from /etc/autobdd-versions instead of a <ver> placeholder, and the catalogue table is the single source of every feature description (so --list, suite headers and run lines cannot drift).


Follow-up: the discovery surface (first, non-breaking step of the command-surface plan)

The rename to autobdd find-target --match-image/--match-text is not in this PR — deliberately. Renaming flags on a frozen contract touches 38 call sites in 10 files (incl. framework/libs/screen_session.js) and needs an alias + translation layer plus a release running both surfaces, so it is sequenced behind this.

What this PR now adds is the part that makes any name discoverable. Measured before the change:

findTargetImage --help     -> exit 0 after a 3079 ms whole-screen OCR SCAN
findTargetImage --version  -> exit 0, silently ignored
usage strings in the binary -> 0

--help is the canonical "what does this do?" move for a human and for an agent; it silently did work and reported success. Now:

Arg Behaviour Cost
--help, -h flows, output shape, exit codes, every flag + default, and the known gaps 37 ms, never touches the screen
--list the supported flows as runnable examples ~37 ms
--version seam, image build stamp (from /etc/autobdd-versions), Oculix, Node ~37 ms
bad value findTargetImage: --imageAction expects one of … on stderr, exit 2 —

Previously --imageAction=bogus fell through the action switch and returned a normal-looking result with no action taken — indistinguishable from success to an automated caller. Unknown arguments still warn and are ignored, so the additive-argument guarantee is unchanged.

The flag catalogue is a single array, so --help, --list and unknown-flag detection cannot drift.

Verified

make base-test   -> feature conformance: 107 passed, 0 failed, ALL FEATURES OK   (38 features)
make one FEATURE=I
   help         ✓ exits 0, prints flows/flags, no result payload, no engine banner, <= 500 ms
   version      ✓ names the seam, reports the build stamp
   list         ✓ lists the read flow and the acting flow
   usage-error  ✓ exits 2, names the flag, lists accepted values

Still to come (separate PRs)

  1. autobdd dispatcher + read-text, engines out of PATH, findTargetImage alias with a stderr-only deprecation notice, legacy-flag translation layer, CI dual-surface run.
  2. --match-image / --match-text become canonical (--min-score, --wait 5s, --limit, --click/--hover, --box), the 10 files migrate, and the suite's run lines become autobdd find-target --match-image … for free — because the printed string and the executed argv are the same string.

Follow-up: the front door (autobdd find-target / autobdd read-text) — c60370e

Second step of the command-surface plan. Flags are deliberately not renamed yet — that touches 38 call sites in 10 files and needs the legacy translation layer, so it stays the next step. This commit adds the verb layer and the alias.

Path Role
/usr/local/bin/autobdd the interface
/usr/local/bin/findTargetImage deprecated alias — stderr notice only, then defers to the front door
/usr/local/libexec/autobdd/{find-target,read-text} verb-named entry points, not on PATH (argv[0] dispatch)
/opt/autobdd/seam/src/findTargetImage.js the single engine
autobdd find-target --imagePath=logo.png --imageAction=click
autobdd read-text

read-text is find-target's whole-screen mode given its own name, so reading the screen is no longer spelled as finding a target called Screen. The alias defers to the front door, so argument translation added later applies to it without a second implementation — and its notice goes to stderr only, because stdout is parsed as target_result: <json>.

Behaviour change: the target is now required

Omitting it used to default to --imagePath=Screen, so a mistyped flag silently ran a whole-screen scan (~3 s) and looked like a successful call — and it made find-target and read-text behaviourally identical. CONTRACT.md §1 always said the target was required; this enforces it (exit 2, with a hint pointing at read-text).

Verified on both surfaces

AutoBDD_Ver=test make base-test                                  -> 124 passed, 0 failed
TARGET_BIN=/usr/local/libexec/autobdd/find-target make base-test -> 124 passed, 0 failed

The suite defaults to the alias, so a green matrix doubles as proof the alias is transparent. 43 features now (group J: front door + alias).

Two of my own test bugs, caught by the run lines

  • front-door-read-text was pointed at the find-target entry point and passed anyway, because the engine's default was already Screen — a false green. It now exercises read-text and also asserts that find-target with no target exits 2.
  • The alias-agreement check called the alias on both sides.

One observation, recorded not hidden

A single rc=139 (SIGSEGV during native teardown) on a successful front-door call while the host had stacked containers; not reproducible in 23 clean runs since. The front-door feature now asserts exit 0 so it stays visible rather than silent.

Next

--match-image / --match-text become canonical (--min-score, --wait 5s, --limit, --click/--hover, --box), the 10 files migrate, the deprecated alias gains the legacy-flag translation, and the suite's run: lines become autobdd find-target --match-image … for free — the printed string and the executed argv are the same string.


Follow-up: the canonical argument vocabulary — 6df4e1e

The v1 flags described how we look and disagreed with themselves: --imageSimilarity and --ocrSimilarity were one concept under two names, and --imageWaitTime was seconds while --ocrWaitTime was milliseconds.

v2 describes what you want, once:

v1 v2
--imagePath / --ocrPath --match-image / --match-text — the two ways to name the target
--imageSimilarity / --ocrSimilarity --min-score
--maxSim / --ocrMaxSim --max-score
--imageWaitTime (s) / --ocrWaitTime (ms) --wait=5s / --wait=800ms
--imageMaxCount / --ocrMaxCount --limit
--imageAction=hoverClick --hover --click (composable flags)
--ocrDetail, --ocrPSM, --ocrOEM --box, --psm, --oem

Presence defines the mode, so there are no modes to remember: --match-image alone, --match-text alone, or both (a picture gated on its region's text). --match-text is literal and case-insensitive by default with --match-regex opting in, so a phrase containing : or ( cannot silently change meaning — v1 --textHint was regex and keeps that behaviour through the legacy path.

Every v1 name still works, translated in the engine and warned about on stderr only. docs/CONTRACT.md gains §2 (canonical) and §2b (deprecation table). A legacy-flags feature asserts the v1 names still match and still click, that stdout carries exactly one result line and no deprecation text, and that the notice lands on stderr.

The suite migrated, so the run lines now read what they run:

run: findTargetImage --match-image=/tmp/…/hello.png --click --flash=0s
✓ reports clicked == center = 900
✓ the pointer is really at the reported centre = 900,1100

Verified: 44 features, 130 checks, 0 failed.

Also: usage-error now exercises a canonical flag (--min-score=abc), and the whole log moved to stderr so a run: line can never print after the check it produced.

Correction to my earlier claim

rc=139 (SIGSEGV during native teardown) on a successful front-door call has now been seen twice, so my earlier "one-off" characterisation was wrong. Output is correct both times — the crash is after the result is written. Not reproducible in 23 clean runs; the front-door feature asserts exit 0 so it stays visible rather than silent. Mitigation candidate for a follow-up: skip the JVM/native teardown on the success path.


Follow-up: the front door is now the default surface — b1b88f1

The suite drives /usr/local/libexec/autobdd/find-target by default and prints the surface it used:

════════════════════════════════════════════════════════════════════════════
 AutoBDD base image — feature conformance
   image   : xyteam/autobdd-base  (built 2026-09-22T19:02:59Z)
   display : :1  1920x1200x24
   surface : /usr/local/libexec/autobdd/find-target
   catalogue: 44 features — see base-test/features.sh
════════════════════════════════════════════════════════════════════════════

TARGET_BIN=findTargetImage runs the same matrix through the deprecated alias (the CI dual-surface check). The alias stays covered inside the default run too: group J's alias-transparent and legacy-flags both invoke findTargetImage explicitly, the latter proving the v1 argument names still match and still click.

Guidance bugs found while doing this

  • The banner and footer printed a command that cannot work. AutoBDD_Ver=test make base-test runs the script on the host, where none of the image's tooling exists. Both now print make docker-run jobs="base-test".
  • The suite now detects the host case and exits 2 with a one-line pointer, instead of failing once per assertion with convert: command not found noise — which is how a first-time user loses an afternoon.
  • test-projects/autobdd-base-test/README.md was stale: it still documented findTargetImage, --textHint, --maxSim and --imageAction=click; described four sections of checks where there are now ten groups; and omitted features.sh/one.sh from the layout. Rewritten to cover the default surface and how to switch it, the four env vars that matter, the run-one-at-a-time caveat (with the measured latency inflation), the group table, how to read a run, and the COPY seam/ cache trap (step 12).
  • Root README's feature-conformance section and sample output now match the real banner.

Verified

AutoBDD_Ver=test make docker-run jobs="base-test"
  -> surface : /usr/local/libexec/autobdd/find-target
  -> feature conformance: 130 passed, 0 failed      ALL FEATURES OK

bash base-test/run.sh          # on the host
  -> exit 2, one-line pointer, no Xvfb started

Follow-up: documentation audit, and the last v1 consumer — bd12649

I grepped every doc for v1 flag names, the old command name and stale counts rather than assuming the earlier passes had caught everything. They hadn't:

README.md — the seam section still documented the v1 arguments as current (--ocrPath, --ocrSimilarity, --ocrMaxSim, --ocrWaitTime, --ocrMaxCount, --ocrAction, --ocrDetail, --ocrPSM, --ocrOEM) with findTargetImage examples; the illustrative run output showed --imagePath=…; the feature table and honesty notes named --textHint, --maxSim, --imageAction, --imageMaxCount, --ocrSimilarity; and the intro claimed the interface was findTargetImage. Now: canonical table + examples with docs/CONTRACT.md as the source of truth (so the two cannot drift again), real run lines in the sample, canonical names throughout, and the duplicated Basic usage block removed.

docs/CONTRACT.md — the v1 argument table was left dangling under a "Raw arguments" heading, reading as if current. Retitled Raw argument reference (v1 — deprecated) with a pointer to the §2b mapping; the --imageAction/--ocrDetail prose and the center row now use canonical names.

framework/libs/screen_session.js — the last v1 emitter. It still built findTargetImage --imagePath=… --imageAction=…, so after the rename every framework call would have warned on stderr seven times. It now emits the front door with canonical flags. Subtlety worth flagging: textHint was a regex and --match-text is literal by default, so --match-regex is passed to preserve the existing semantics exactly — documented in the code.

autobdd find-target --match-image=/tmp/…/hello.png --min-score=0.8 --max-score=1 \
  --match-text="HELLO" --match-regex --wait=5s --hover --click --limit=1
→ {"name":"hello.png","score":1,"clicked":"object"}      stderr: (empty)

Verified: the generated framework command matches and clicks with empty stderr; base suite 131 passed, 0 failed. The framework e2e suite itself was not run — it needs the L2 image with Chrome — so that consumer is verified at the command-string level plus the engine's v2/legacy coverage.

No v1 flag name remains in any documentation except the deprecation tables (§2b) and the intentional historical notes in CHANGELOG.md.

Also in this commit: the help text's known-gap line and the suite's user-facing known-gap line now say --box (what a reader would type), and the help feature asserts both that the canonical flag is listed and that the deprecated v1 names are, so the deprecation table cannot silently vanish.

An intermittent red run showed five image-match checks receiving an empty
string: the Oculix/JVM writes its own startup logging to the same stdout fd the
seam uses, and on a cold call it can emit a partial line before ours, so
'target_result:' ended up mid-line. The test helper extracted with an anchored
'sed -n s/^target_result: //p', which finds nothing in that case.

test-projects/autobdd-base-test/base-test/run.sh:
- Extract the marker anywhere on the line (s/.*target_result: //p) so a trailing
  or partial JVM log line cannot swallow the payload. The seam's stdout may
  legitimately carry engine logging; the contract is that the JSON travels on
  the line carrying the marker.

seam/src/findTargetImage.js:
- findImage(): construct Screen (and setAutoWaitTimeout) inside the try. Building
  the Screen is exactly the operation that can fail transiently (display not
  ready, X hiccup); outside the try it killed the process before the CLI printed
  anything, violating the 'stdout always carries a result' contract.
- findImageOcr(): extend the per-iteration try to cover result assembly and
  performAction(), not just rect discovery, so a mouse-action failure also yields
  JSON instead of exiting silently.

Verified: two consecutive cold-cache runs of
  AutoBDD_Ver=test docker compose run --rm autobdd-base-test make base-test
both report 'base-test: 50 passed, 0 failed'.
…leeps

Five defects surfaced by the new 34-feature conformance suite (each verified
against the running image before/after):

- --imageWaitTime was ignored. SikuliX's autoWaitTimeout applies to wait()/exists(),
  not to the findAll() the seam used, so a late-appearing target was missed even
  though the caller asked to wait for it. The search now retries until the deadline,
  matching docs/CONTRACT.md §2.
- Clicking actions did not move the pointer. Region.click()/doubleClick()/rightClick()
  do not reposition the pointer in this Oculix build, so the action landed somewhere
  else (measured: reported center 900,1100 vs pointer 131,160). Actions now move to
  the region centre first.
- --flash never paused: java.lang.Thread.sleep() through java-bridge is a no-op
  (sleep(1000) measured 0 ms), so the flash had no hold time and the OCR poll spun.
  Waits now use a real Node-side block (Atomics.wait).
- --imageMaxCount returned aliases: one accumulator object was pushed per match, so N
  results were N references to the last match (identical centres). Fresh object per
  match now, so two identical tiles report two distinct centres.
- clicked was set for non-clicking actions: the image path recorded it even for
  --imageAction=hover. Now set only when a click was dispatched, consistent with the
  OCR path and with the field's meaning.
… catalogue

The suite was a linear list of assertions: it did not say which feature a check
belonged to, and a failure could not be reproduced on its own.

- base-test/features.sh — the catalogue: 34 features in groups A..H (runtime
  substrate, display+desktop, whole-screen OCR, image matching, actions, opt-in OCR,
  contract robustness, NFR). Each feature renders its own fixtures, so features are
  order-independent, and each is one call to reproduce.
- base-test/run.sh — driver: runs the whole matrix, grouped, with per-feature
  headers, the exact cmd line, and failures collected by feature at the end.
- base-test/one.sh <feature|group|--list> — runs exactly the checks the full suite
  runs for that feature (same code path), so a green single run means the same thing
  as a green full run.
- Makefile — 'make one FEATURE=<feature>'.

Coverage added beyond the previous suite: --imageSimilarity floor (accept and
reject), --imageWaitTime, --imageMaxCount with distinct centres, all five actions
(verified against xdotool's view of the pointer, not just self-consistent numbers),
--flash measured against --flash=0, --imageAction=none leaving the pointer alone,
OCR floor/wait/detail levels/actions/PSM-OEM, unknown-argument tolerance, and
/etc/autobdd-versions provenance.

Deliberately not asserted: SCREENSHOT appears in docs/CONTRACT.md §4 but is not
implemented by the base seam (it is an L2/framework concern), so it is absent from
the catalogue rather than falsely claimed.

README.md — a 'Feature conformance' section: the whole-matrix one-liner, the
per-feature/group one-liners, the feature table, and the coverage-honesty notes.
CHANGELOG.md — the five seam fixes above.

Verified: make base-test -> 83 passed, 0 failed, ALL FEATURES OK (34 features);
make one FEATURE=action-hover -> 2 passed, 0 failed.
@xywang68 xywang68 changed the title Harden the seam CLI contract against interleaved JVM stdout (fixes intermittent empty result) Base seam: fix 6 real defects, and rebuild the base suite as a reproducible 34-feature catalogue Sep 22, 2026
…what is true

Auditing the suite's own narration found the log reporting intent rather than
execution. Six fixes, applied to features.sh:

1. Printed command is now literally runnable. Every seam call goes through
   target(), which prints the exact argv (quoted) and the exit status of the call
   it just made; host commands go through probe(), which prints then runs the same
   string. Previously json-on-error printed 'findTargetImage DISPLAY=:77
   --imagePath=Screen' - the env assignment in argument position - so pasting it ran
   a *successful* screen read instead of the error path.
2. One line per invocation. target() is the single funnel, so multi-invocation
   features (image-similarity, text-hint, ocr-similarity, flash, latency) can no
   longer print one command while running several.
3. Assert only what is implemented. --ocrSimilarity is accepted but NOT applied by
   this build: with the text on screen, 0.99 still matches (no per-match confidence
   to filter on). The feature now asserts acceptance + absent-text rejection and
   prints an explicit 'known-gap:' line. An earlier observation of '0.99 rejects'
   was an artifact of a WM-less harness where the fixture had vanished; with
   openbox running, fixtures are durable and the truth is that the floor does not
   filter.
4. Names match assertions: provenance now checks os=, java=, node=, ubuntu_digest=
   and built= (it previously claimed three keys while grepping one); wm checks the
   openbox PID we started plus 'pgrep -x' instead of a '-f' fallback that matched
   its own command line; json-on-error parses the payload and asserts a status
   instead of grepping '[' - which the JVM's own '[DEBUG STARTUP]' lines satisfied
   vacuously.
5. Passing checks print the observed value ('✓ <what> = <value>'), so a green run
   is evidence rather than a checklist.
6. Features are order-independent: the fixture is re-shown before each invocation.
   Timing features opt out via NOSHOW=1 and say so, because the re-show would
   otherwise land inside the measured window (it inflated the latency check from
   ~1130 ms to 1675 ms).

Two self-inflicted bugs found and fixed while doing this: target()'s log line went
to stdout, so inside $( ) it was captured into the JSON variable and corrupted the
payload (log lines now go to stderr); and the fixture re-show inflated the latency
measurement. The flash threshold was also recalibrated to the real separation -
measured 15-284 ms while the flash was broken, 780-1010 ms when applied - with an
upper bound added to catch a runaway flash.

Also: the banner now reports the image's own build stamp from
/etc/autobdd-versions instead of a placeholder, and the catalogue table is the
single source of every feature description.

Verified: make base-test -> feature conformance: 92 passed, 0 failed, ALL
FEATURES OK (34 features); make one FEATURE=flash -> 2 passed; make one
FEATURE=ocr-similarity -> 2 passed with the known-gap line printed.
…oudly on bad input

First, non-breaking step of the command-surface plan (the rename to
'autobdd find-target --match-image/--match-text' is PR 2/3 and needs the
alias + translation layer, because renaming flags on a frozen contract touches
38 call sites in 10 files).

Discovery, measured before this change:
    findTargetImage --help   -> exit 0 after a 3079 ms whole-screen OCR SCAN
    findTargetImage --version-> exit 0, silently ignored
    usage strings in the binary -> 0
So the canonical 'what does this do?' move for a human and for an agent silently
did work and reported success. --help/--list/--version are now handled in the argv
stage, before java-bridge is required, so they cost ~37 ms and never touch the
screen. --help prints the flows, the output shape, the exit codes, every flag with
its default, and the known gaps; --list prints the flows as runnable examples;
--version reports the image's own build stamp from /etc/autobdd-versions.

Failure modes now explicit instead of silent:
- unusable values (non-numeric threshold, unknown --imageAction/--ocrDetail) print
  'findTargetImage: <flag> expects ...' on stderr and exit 2. Previously
  --imageAction=bogus fell through the switch and returned a normal-looking result
  with no action taken.
- unknown arguments still warn on stderr and are otherwise ignored, so the
  additive-argument guarantee for existing consumers is unchanged.
- the flag catalogue is a single array, so --help, --list and unknown-flag
  detection cannot drift apart.

Tests: 4 new features (38 total, 107 checks, all green):
  help         exits 0, prints flows/flags, emits no result payload, no engine
               banner, answers in <= 500 ms (a JVM+engine start is ~1100 ms)
  version      names the seam and reports the build stamp
  list         lists the read and the acting flow
  usage-error  exits 2, names the offending flag, lists the accepted values

Also: README documents the discovery surface and the exit-status split (0 = a
result incl. notFound, 2 = usage error), and the CHANGELOG records the measured
before/after.

Verified: make base-test -> feature conformance: 107 passed, 0 failed, ALL
FEATURES OK (38 features).
…eprecated alias

Second step of the command-surface plan. Flags are deliberately NOT renamed yet:
renaming them on a frozen contract touches 38 call sites in 10 files and needs the
legacy translation layer, so that remains the next step.

- autobdd <verb> is the interface. Verbs are verb-object, so a call reads as an
  instruction: 'autobdd find-target --imagePath=logo.png --imageAction=click',
  'autobdd read-text'. read-text is find-target's whole-screen mode given its own
  name, so reading the screen is no longer spelled as finding a target called
  Screen.
- The engine moved out of PATH to /opt/autobdd/seam/src/. Only the front door and
  the deprecated alias are on PATH; the verb-named entry points live in
  /usr/local/libexec/autobdd/{find-target,read-text} and are reached by argv[0]
  dispatch, so a derived image cannot shadow or collide with them.
- findTargetImage is kept as a deprecated alias: notice on STDERR only — stdout is
  parsed as 'target_result: <json>', so a notice there would corrupt every consumer
  — and it defers to the front door, so argument translation added later applies to
  it without a second implementation.
- The target is now required. Omitting it used to default to --imagePath=Screen, so
  a mistyped flag silently ran a whole-screen scan (~3 s) and looked like a
  successful call, and it made find-target and read-text behaviourally identical.
  docs/CONTRACT.md §1 already said the target was required; this enforces it with
  exit 2 and a hint pointing at read-text.

Tests: 5 new features (43 total, 124 checks). The suite defaults to the ALIAS, so a
green matrix doubles as proof the alias is transparent; TARGET_BIN lets the same
matrix run through the front door.

Two of my own test bugs were caught by the new run lines and fixed: the read-text
feature was pointed at the find-target entry point (and passed only because the
engine's default was already Screen — a false green), and the alias-agreement check
called the alias on both sides.

One observation recorded rather than hidden: a single rc=139 (SIGSEGV during native
teardown) on a successful front-door call under a loaded host, not reproducible in
23 clean runs since. The front-door feature now asserts exit 0 so it stays visible.

Verified on both surfaces:
  AutoBDD_Ver=test make base-test                                  -> 124 passed, 0 failed
  TARGET_BIN=/usr/local/libexec/autobdd/find-target make base-test -> 124 passed, 0 failed
…mes translated

The v1 names described HOW WE LOOK and disagreed with themselves:
--imageSimilarity and --ocrSimilarity were one concept under two names, and
--imageWaitTime was in SECONDS while --ocrWaitTime was in MILLISECONDS.

The v2 names describe WHAT YOU WANT and say each thing once:
  --match-image / --match-text   the two ways to name the target
  --min-score / --max-score      the score bounds, mirroring the 'score' field
  --wait=5s | --wait=800ms       durations carry their unit, so s/ms cannot be confused
  --limit, --box, --psm, --oem   stated once for both modalities
  --click/--double-click/--right-click/--hover   composable, so '--hover --click'
                                 replaces the hoverClick enum string

Presence defines the mode, so there are no modes to remember: --match-image alone,
--match-text alone, or both (a picture gated on its region's text). --match-text is
literal and case-insensitive by default with --match-regex opting in, so a phrase
containing ':' or '(' cannot silently change meaning; v1 --textHint was regex and keeps
that behaviour through the legacy path.

Every v1 name is still accepted: translated in the engine, warned about on STDERR only
(stdout is the payload). docs/CONTRACT.md gains §2 (canonical) and §2b (deprecation table).
The front door's read-text verb and its examples use the canonical flag, so it does not
warn about itself.

The suite migrated mechanically, and the run lines now read:
  run: findTargetImage --match-image=/tmp/.../hello.png --click --flash=0s

Tests: 44 features, 130 checks, all green. New: legacy-flags asserts the v1 names still
match and still click, that stdout carries exactly one result line and no deprecation
text, and that the notice lands on stderr. usage-error now exercises a canonical flag
(--min-score=abc). The whole log moved to stderr so a 'run:' line can never print after
the check it produced.

Observed honestly: rc=139 (SIGSEGV during native teardown) on a successful front-door
call has now been seen TWICE, so the earlier 'one-off' characterisation was wrong. Output
is correct both times; the crash is after the result is written. Not reproducible in 23
clean runs; the front-door feature asserts exit 0 so it stays visible. Mitigation
candidate for a follow-up: skip the JVM/native teardown on the success path.

Verified: make base-test -> feature conformance: 130 passed, 0 failed, ALL FEATURES OK.
…nstructions

The suite now drives /usr/local/libexec/autobdd/find-target by default — the
interface this repo is moving to — and prints the surface it used in its banner.
TARGET_BIN=findTargetImage still runs the same matrix through the deprecated
alias (the CI dual-surface check), and the alias stays covered inside the default
run by group J: alias-transparent and legacy-flags both invoke findTargetImage
explicitly, the latter proving the v1 argument names still match and still click.

Guidance fixes from the same review:
- The banner and footer printed 'AutoBDD_Ver=test make base-test', which runs the
  script ON THE HOST — where none of the image's tooling exists. They now print
  the invocation that works: make docker-run jobs=base-test.
- The suite detects the host case and exits 2 with a one-line pointer, instead of
  failing once per assertion with missing-binary noise. This is how a first-time
  user loses an afternoon.
- test-projects/autobdd-base-test/README.md was stale: it still documented
  findTargetImage, --textHint, --maxSim and --imageAction=click, listed four
  sections of checks where there are now ten groups, and omitted features.sh and
  one.sh from the layout. Rewritten: the default surface and how to switch it, the
  four environment variables that matter, the run-one-at-a-time caveat with the
  measured latency inflation, the full group table, how to read a run, and the
  cache trap for the COPY seam layer (step 12).
- README's feature-conformance section and the sample output now match the real
  banner, and CHANGELOG records the default-surface change.

Verified: AutoBDD_Ver=test make docker-run jobs="base-test" -> 130 passed, 0 failed,
surface: /usr/local/libexec/autobdd/find-target; and 'bash base-test/run.sh' on the
host exits 2 with the pointer, without starting Xvfb.
@xywang68
xywang68 merged commit 6cf8f94 into master Sep 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant