From c660ea739259e1659001b10694a3f6c49cbbbf14 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 10:13:44 +0000 Subject: [PATCH 1/6] perf(ui-loop): skip the redundant baseline mxbuild, wait on page state, legible screenshots - exec.sh skips the pre-flight baseline mxbuild when the model is unchanged since its last clean pass (stamp gains errors:, check --clean). - New tests/e2e/settle.js replaces fixed sleeps in journey-runner and page-audit; SETTLE_MODE=fixed restores the old waits for an A/B run. - page-audit saves a viewport shot beside the full-page one; ui-loop.md says to look at viewport shots. - shrink-image-read.sh works off macOS (ImageMagick/Pillow, no-jq path). Field run 2026-10-07, requirements-driven build. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- CHANGELOG.md | 1 + bin/lib/install-manifest.sh | 2 +- claude-hooks/hooks/shrink-image-read.sh | 76 +++++++++++++++------ project-bin/exec.sh | 23 +++++-- project-bin/model-stamp.sh | 14 +++- project-bin/verify-model.sh | 2 +- project-tests/e2e/journey-runner.js | 9 ++- project-tests/e2e/page-audit.js | 13 +++- project-tests/e2e/settle.js | 90 +++++++++++++++++++++++++ skills/journey-examples.md | 4 +- skills/ui-loop.md | 5 +- 11 files changed, 204 insertions(+), 35 deletions(-) create mode 100644 project-tests/e2e/settle.js diff --git a/CHANGELOG.md b/CHANGELOG.md index 1a79d3c9..f99f456b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -17,6 +17,7 @@ Sections dated before 2026-09-19 predate the cycle and stay as they are. ## Unreleased +- perf(ui-loop): **fewer mxbuilds, no fixed sleeps, legible screenshots.** Field run 2026-10-07 (a requirements-driven build): `exec.sh` was 22% of tool time at two mxbuilds per exec, journeys and page waits slept a fixed time, and screenshots were 11.5% of re-read tokens. `exec.sh` now skips the pre-flight baseline mxbuild when the model is byte-identical to its last *clean* pass: the verification stamp gains an `errors:` field (`model-stamp.sh check --clean` accepts only `errors: 0`, so a delta-gate pass over pre-existing errors never counts; `SKIP_BASELINE=0` is unchanged). New `tests/e2e/settle.js` waits for the page's own state (no visible Mendix progress/loading indicator, DOM size stable) instead of a fixed time; `journey-runner.js` uses it after each action (an explicit `settleMs` still wins) and `page-audit.js` per page. Probed in a real browser: on a page that loads in 1.8 s the old 1200 ms sleep and networkidle+600 both captured it half-loaded, settle waited for it; on a fast page settle took 775 ms vs 1200. `SETTLE_MODE=fixed` restores the old waits for an A/B run. `page-audit.js` also saves a viewport shot `page-audit-.top.png` beside the full-page one, and `ui-loop.md` says to look at viewport shots: a tall full-page shot reaches the model at ~1024 px high, text unreadable. `shrink-image-read.sh` now works off macOS (ImageMagick or Pillow, python when jq is absent; it was a silent no-op without `sips`). — MendixMau - feat(permissions): **intake now picks how much Claude Code may run without a prompt, and `doctor.sh` shows it.** Until now the allow-list covered only the safe wrappers (`bin/exec.sh`, `bin/*.sh`, `./mxcli`, `mx`), so every edit, test run, `npm install` or `git commit` in a project still prompted. New `bin/permission-profile.sh` resolves a profile (env, `.mxtk/permission-profile`, PROJECT.md `Permission profile:`, default `project`; a typo falls back to `wrappers`): `wrappers` is the old list; `project` adds edits anywhere inside the project folder (`Edit(/**)`) and the everyday test/npm/docker/git-add/commit commands, with force-push, `git reset --hard`, `git clean` and `sudo` behind `ask` rules that prompt in every mode; `full` adds bare `Bash`/`Edit`/`Read`/`WebFetch`/`WebSearch` to the per-machine `settings.local.json` only, never the shared file. `install-claude-permissions.sh` / `install-harness-permissions.sh` take `--profile`, and switching to a smaller profile removes only what the script itself added. Intake gains Q12, `init-project.sh` writes `Permission profile: project`, and `doctor.sh ` reports the profile, whether the settings files carry it, and warns on a `defaultMode` of auto/bypass in a project file, which Claude Code ignores there. No profile can switch on bypass or auto mode from the project: doctor says how to do that from the user's own settings. `test-install-claude-permissions.sh` gains T5b. — MendixMau - fix(small): **five small field defects (#146, #149, #152, #153, #156) and the missing v0.24 ledger entries (#168).** `gate-check.sh --closeout ` now writes `index.html` when that stage PASSes or is WAIVED, so the board moves at every gate the runbook closes; plain stage queries stay read-only (#146). `journey-runner.js` refuses a journey whose `persona` differs from the signed-in user — INVALID, not walked, naming the `TEST_USER` to use — where before a non-admin journey walked green under admin; `journey-proof.md` says one file, one run per role, and `test-journey-persona.sh` is new (#149). `design-audit.js` rung 7 exempts the three `spacing-outer-bottom*` classes `design-spacing.md` prescribes (still FAIL when the widget also sets the Spacing property), and the skill names the list (#152). `verify-module.sh` runs `REFRESH CATALOG FULL` before the model-side instruments when the catalog is older than the `.mpr`, so `test --attach` then verify no longer reads INCOMPLETE (probed on a scratch model: stale, refreshed, fresh, then a no-op) (#153). `wiring-sweep.md` adds N/A (by design) for a control its own state disables and UNREACHED for a conditional widget the sweep never saw, and enumerates `Visible:`/`Editable:` widgets from `describe page` so the denominator is not the DOM's alone (#156). `mxcli-bugs.md` gains four BUG-DRAFT entries from a field comparison build (2026-09-29): false page-setting warnings, the unwritable pluggable `file` property, `index on createdDate`, and the `describe` replay CE1613; the inbox note is promoted and removed. Upstream asks in #168 stay open. — MendixMau - fix(gates): **a red gate under an approval, a `done-` deadlock and unfilled agent slots are now visible.** A field project had Stage 4 CONFIRMED in `PROJECT.md` while its gate FAILED, 47 scripts at gate-pass with none renamed `done-`, and 35 unfilled `{{...}}` slots in its agent stubs, and no instrument said so (a requirements-driven field build, 2026-10-06; #207, #210, #211). `gate-check.sh` now prefixes a FAIL on a CONFIRMED stage with `APPROVED OVER A FAILING GATE` and adds a summary line under Needs attention (verdict and exit code unchanged). `status.sh` gains a `WATCH` line for all three plus a stale `build-plan.html`, and its NEXT names the first one; its not-yet-`done-` list now skips BUILD-LOG rows for scripts no longer on disk. New `bin/lib/placeholders.sh` counts slots (excluding the `{{DOUBLE_BRACE}}` prose); `init-project.sh` and `sync-project.sh` list them per file. Slots stay unfilled until each agent's stage starts, as `agent-roles.md` says: these are reports, not failures. — MendixMau diff --git a/bin/lib/install-manifest.sh b/bin/lib/install-manifest.sh index e385ea68..4bb29c67 100644 --- a/bin/lib/install-manifest.sh +++ b/bin/lib/install-manifest.sh @@ -176,7 +176,7 @@ MXTK_LINT_RULES_NOINSTALL="README.md STOCK-HASHES.txt" # # project.config.template.js is listed separately because it is the one file a project is # EXPECTED to edit — installing it over an edited copy would silently revert the port. -MXTK_PROJECT_TESTS="config.js helpers.js otel.js journey-runner.js journey-runner.selftest.js journey-rung4-scope.test.js monkey.js monkey.selftest.js page-audit.js page-audit-rules.js design-audit.js full-app-walkthrough.js report-normalize.js report-render.js review-report.js example.journey.json" +MXTK_PROJECT_TESTS="config.js helpers.js settle.js otel.js journey-runner.js journey-runner.selftest.js journey-rung4-scope.test.js monkey.js monkey.selftest.js page-audit.js page-audit-rules.js design-audit.js full-app-walkthrough.js report-normalize.js report-render.js review-report.js example.journey.json" MXTK_PROJECT_TESTS_TEMPLATE="project.config.template.js" # Deliberately not installed. Same contract as the lists above. diff --git a/claude-hooks/hooks/shrink-image-read.sh b/claude-hooks/hooks/shrink-image-read.sh index 939cf777..6a1fc391 100755 --- a/claude-hooks/hooks/shrink-image-read.sh +++ b/claude-hooks/hooks/shrink-image-read.sh @@ -8,6 +8,11 @@ # This rewrites the Read's file_path to point at a downscaled copy. Fail-open: # any error passes the original Read through untouched. # +# Resizer, first one found: sips (macOS), ImageMagick (magick / convert), python3 with +# Pillow. JSON via jq, else python3. WHY (field run 2026-10-07): this hook needed sips, +# jq and md5, all macOS-only or often absent, so on Linux and Git Bash it was a silent +# no-op — while screenshot reads were 11.5% of one run's re-read tokens. +# # Tunables: CLAUDE_IMG_MAX_BYTES (default 150000), CLAUDE_IMG_MAX_DIM (default 1024) # Bypass entirely: CLAUDE_IMG_SHRINK=0 @@ -19,15 +24,32 @@ MAX_BYTES=${CLAUDE_IMG_MAX_BYTES:-150000} MAX_DIM=${CLAUDE_IMG_MAX_DIM:-1024} CACHE="${TMPDIR:-/tmp}/claude-shrunk" -command -v jq >/dev/null 2>&1 || exit 0 -command -v sips >/dev/null 2>&1 || exit 0 +PY="" +for c in python3 python py; do # portability-ok: installed to ~/.claude/hooks, cannot source portable.sh; same probe list + command -v "$c" >/dev/null 2>&1 && "$c" -c 'import sys; sys.exit(sys.version_info[0] < 3)' 2>/dev/null \ + && { PY="$c"; break; } +done +HAVE_JQ=0; command -v jq >/dev/null 2>&1 && HAVE_JQ=1 +[ "$HAVE_JQ" = 1 ] || [ -n "$PY" ] || exit 0 + +RESIZER="" +if command -v sips >/dev/null 2>&1; then RESIZER=sips +elif command -v magick >/dev/null 2>&1; then RESIZER=magick +elif command -v convert >/dev/null 2>&1 && convert -version 2>/dev/null | grep -q ImageMagick; then RESIZER=convert +elif [ -n "$PY" ] && "$PY" -c 'import PIL' 2>/dev/null; then RESIZER=pil +fi +[ -n "$RESIZER" ] || exit 0 input=$(cat) -path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // empty' 2>/dev/null) +if [ "$HAVE_JQ" = 1 ]; then + path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // empty' 2>/dev/null) +else + path=$(printf '%s' "$input" | "$PY" -c 'import json,sys; print((json.load(sys.stdin).get("tool_input") or {}).get("file_path") or "")' 2>/dev/null) +fi [ -n "$path" ] || exit 0 [ -f "$path" ] || exit 0 -# Image extensions sips can resample. +# Image extensions the resizers can read. ext=$(printf '%s' "${path##*.}" | tr '[:upper:]' '[:lower:]') case "$ext" in png|jpg|jpeg|tif|tiff|bmp|gif|heic|webp) ;; @@ -42,16 +64,24 @@ bytes=$(stat -f%z "$path" 2>/dev/null || stat -c%s "$path" 2>/dev/null \ [ "$bytes" -gt "$MAX_BYTES" ] 2>/dev/null || exit 0 mtime=$(stat -f%m "$path" 2>/dev/null || stat -c%Y "$path" 2>/dev/null) -key=$(printf '%s:%s:%s:%s' "$path" "$mtime" "$bytes" "$MAX_DIM" | md5 -q 2>/dev/null) +key=$(printf '%s:%s:%s:%s' "$path" "$mtime" "$bytes" "$MAX_DIM" \ + | { md5 -q 2>/dev/null || md5sum 2>/dev/null || cksum; } | awk '{print $1}') [ -n "$key" ] || exit 0 mkdir -p "$CACHE" 2>/dev/null || exit 0 out="$CACHE/$key.png" if [ ! -f "$out" ]; then - sips --setProperty format png \ - --resampleHeightWidthMax "$MAX_DIM" \ - "$path" --out "$out" >/dev/null 2>&1 || exit 0 + case "$RESIZER" in + sips) sips --setProperty format png --resampleHeightWidthMax "$MAX_DIM" \ + "$path" --out "$out" >/dev/null 2>&1 ;; + magick|convert) "$RESIZER" "$path[0]" -resize "${MAX_DIM}x${MAX_DIM}>" "png:$out" >/dev/null 2>&1 ;; + pil) "$PY" -c 'import sys +from PIL import Image +im = Image.open(sys.argv[1]); im.thumbnail((int(sys.argv[3]),) * 2) +(im if im.mode in ("RGB", "RGBA", "L", "LA", "P") else im.convert("RGBA")).save(sys.argv[2], "PNG", optimize=True)' \ + "$path" "$out" "$MAX_DIM" >/dev/null 2>&1 ;; + esac || { rm -f "$out"; exit 0; } fi [ -s "$out" ] || exit 0 @@ -64,15 +94,23 @@ newbytes=$(stat -f%z "$out" 2>/dev/null || stat -c%s "$out" 2>/dev/null \ # original beside it: project-bin/look-ledger.sh reads this to record which screenshot was seen. printf '%s\n' "$path" > "$CACHE/$key.src" 2>/dev/null || true -printf '%s' "$input" | jq -c \ - --arg p "$out" \ - --arg orig "$path" \ - --arg dim "$MAX_DIM" \ - '{ - hookSpecificOutput: { - hookEventName: "PreToolUse", - updatedInput: (.tool_input + {file_path: $p}), - additionalContext: ("Downscaled copy of \($orig) (max \($dim)px). Refer to the original path in your output.") - } - }' +if [ "$HAVE_JQ" = 1 ]; then + printf '%s' "$input" | jq -c \ + --arg p "$out" \ + --arg orig "$path" \ + --arg dim "$MAX_DIM" \ + '{ + hookSpecificOutput: { + hookEventName: "PreToolUse", + updatedInput: (.tool_input + {file_path: $p}), + additionalContext: ("Downscaled copy of \($orig) (max \($dim)px). Refer to the original path in your output.") + } + }' +else + printf '%s' "$input" | "$PY" -c 'import json,sys +d = json.load(sys.stdin); ti = dict(d.get("tool_input") or {}); ti["file_path"] = sys.argv[1] +print(json.dumps({"hookSpecificOutput": {"hookEventName": "PreToolUse", "updatedInput": ti, + "additionalContext": "Downscaled copy of %s (max %spx). Refer to the original path in your output." % (sys.argv[2], sys.argv[3])}}))' \ + "$out" "$path" "$MAX_DIM" || exit 0 +fi exit 0 diff --git a/project-bin/exec.sh b/project-bin/exec.sh index 761cef31..5e84a87f 100755 --- a/project-bin/exec.sh +++ b/project-bin/exec.sh @@ -91,6 +91,7 @@ BUILD_LOG="$PROJECT_ROOT/docs/BUILD-LOG.md" # Declared HERE, above log_build, not next to the gate: a row must be able to # carry a gate verdict even when the gate block below is never reached. GATE_STATE="not-run" # not-run | skipped | unverified | pass | fail +GATE_ERRORS="?" # mxbuild error count behind a pass (0 = clean); recorded in the stamp MXB_WHY="" # mxbuild's errors[] reason when it exited without checking the model # ── Exec approval (auto records itself) ────────────────────────────────────── @@ -682,8 +683,22 @@ fi # Capture the CURRENT error set before touching anything, so the gate can tell # "this script broke it" from "it was already broken". Costs one extra mxbuild; # skip with SKIP_BASELINE=1 when you know the tree is clean. +# +# The extra mxbuild is also skipped when the verification stamp (bin/model-stamp.sh) +# says THIS exact model state already passed mxbuild with 0 errors: the baseline is +# then known to be empty without measuring it. Any change since (Studio Pro, a test +# run, a bare exec) changes the fingerprint and the baseline runs as before. +# WHY (field run 2026-10-07, requirements-driven build): exec.sh was 22% of all tool +# time, with two mxbuilds per exec; the baseline one re-measured a model the previous +# exec had just verified, in most of 100 execs. BASELINE_SET="" BASE_KNOWN=0 +if [ "${SKIP_BASELINE:-0}" != "1" ] && [ -x ./bin/model-stamp.sh ] \ + && ./bin/model-stamp.sh check -q --clean >/dev/null 2>&1; then + echo "→ Pre-flight: model unchanged since its last clean mxbuild (verification stamp) — baseline 0, not re-measured" + _BC=0; _BCODES=""; BASE_KNOWN=1 + SKIP_BASELINE=1 +fi if [ "${SKIP_BASELINE:-0}" != "1" ] && [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; then echo "→ Pre-flight: checking whether the model already has errors..." _BF=$(mktemp /tmp/mxbuild-baseline.XXXXXX) @@ -823,7 +838,7 @@ if [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; then # A guard that tested only "file is non-empty" therefore took the restore # branch on a clean build and rolled back good work, while printing # "0 error(s) found". Test the parsed count, never the file's existence. - GATE_STATE="pass" + GATE_STATE="pass"; GATE_ERRORS=0 echo " ✓ mxbuild: 0 errors — model is clean." [ "$EXEC_STATUS" -eq 0 ] && log_applied "mxbuild clean" else @@ -848,7 +863,7 @@ if [ -x "$MXBUILD" ] && [ -x "$JAVA_EXE" ]; then DELTA_KIND="subset" fi if [ -n "$DELTA_KIND" ]; then - GATE_STATE="pass" + GATE_STATE="pass"; GATE_ERRORS="$CE_COUNT" KEEP_CODES=$(err_codes "$ERRORS_FILE") if [ "$DELTA_KIND" = "subset" ]; then echo " ⚠ mxbuild: $CE_COUNT error(s) — a STRICT SUBSET of the $_BC pre-flight baseline error(s), no new ones." @@ -959,7 +974,7 @@ EOF # --patch logs its own row below, after rolling back; one row per run. [ "$PATCH_MODE" = "1" ] || log_build "❌ gate could not run" "mxbuild exit $MXBUILD_EXIT" else - GATE_STATE="pass" + GATE_STATE="pass"; GATE_ERRORS=0 echo " ✓ mxbuild: 0 errors — model is clean." # No errors file, or an empty one. The comment at the CE_COUNT=0 branch says # mxbuild ALWAYS writes one; on Mendix 11.13 it does not, so THIS is where @@ -1071,7 +1086,7 @@ fi # is no longer the one anything verified. if [ -x ./bin/model-stamp.sh ]; then if [ "$GATE_STATE" = "pass" ] && [ "$EXEC_STATUS" -eq 0 ]; then - ./bin/model-stamp.sh write pass "exec.sh $([ "$PATCH_MODE" = 1 ] && echo '--patch ')$(basename "$SCRIPT")" || true + MXTK_STAMP_ERRORS="$GATE_ERRORS" ./bin/model-stamp.sh write pass "exec.sh $([ "$PATCH_MODE" = 1 ] && echo '--patch ')$(basename "$SCRIPT")" || true else ./bin/model-stamp.sh clear >/dev/null 2>&1 || true fi diff --git a/project-bin/model-stamp.sh b/project-bin/model-stamp.sh index e0027882..aaf6d693 100755 --- a/project-bin/model-stamp.sh +++ b/project-bin/model-stamp.sh @@ -3,7 +3,8 @@ # # ./bin/model-stamp.sh fingerprint [--staged] # print the model fingerprint # ./bin/model-stamp.sh write # record that the current model was verified -# ./bin/model-stamp.sh check [--staged] [-q] # exit 0 iff the (staged) model matches a PASS stamp +# ./bin/model-stamp.sh check [--staged] [-q] [--clean] # exit 0 iff the (staged) model matches a PASS stamp +# (--clean: and that pass saw 0 mxbuild errors) # ./bin/model-stamp.sh clear # forget the stamp (the model changed unverified) # ./bin/model-stamp.sh paths # the repo-relative model paths this covers # @@ -219,6 +220,10 @@ case "$cmd" in { printf 'fingerprint: %s\n' "$fp" printf 'state: %s\n' "$state" printf 'source: %s\n' "$*" + # mxbuild error count the pass saw. exec.sh's delta gate also passes a model that + # still carries pre-existing errors, so "pass" alone does not mean clean; exec.sh + # skips its pre-flight mxbuild only on a stamp that says 0 here (check --clean). + printf 'errors: %s\n' "${MXTK_STAMP_ERRORS:-?}" printf 'at: %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" } > "$STAMP" # Machine-local, like the doctor receipt: never let it travel with the repo. @@ -230,8 +235,8 @@ case "$cmd" in ;; clear) rm -f "$STAMP"; echo " verification stamp cleared" ;; check) - staged=0; quiet=0 - for a in "$@"; do case "$a" in --staged) staged=1 ;; -q|--quiet) quiet=1 ;; esac; done + staged=0; quiet=0; clean=0 + for a in "$@"; do case "$a" in --staged) staged=1 ;; -q|--quiet) quiet=1 ;; --clean) clean=1 ;; esac; done say() { [ "$quiet" = 1 ] || echo "$@"; } if [ "$staged" = 1 ]; then in_git || exit 0 @@ -263,6 +268,9 @@ case "$cmd" in if [ "$st" != "pass" ]; then say " ✗ $what is UNVERIFIED: last verification was '$st' ($src, $at)"; exit 1 fi + if [ "$clean" = 1 ] && [ "$(stamp_field errors)" != "0" ]; then + say " ✗ $what passed with pre-existing errors ($(stamp_field errors)) — not a clean stamp"; exit 1 + fi say " ✓ $what verified ($src, $at)" ;; ""|-h|--help) sed -n '2,7p' "$0"; exit 0 ;; diff --git a/project-bin/verify-model.sh b/project-bin/verify-model.sh index 865d6e96..816df6f6 100755 --- a/project-bin/verify-model.sh +++ b/project-bin/verify-model.sh @@ -84,5 +84,5 @@ for x in [p for p in d.get('problems',[]) if p.get('severity')=='Error'][:15]: fi echo "✓ mxbuild: 0 errors — model is clean." rm -f "$ERR" "$OUT" -[ "$STAMP" = 1 ] && ./bin/model-stamp.sh write pass "verify-model.sh" +[ "$STAMP" = 1 ] && MXTK_STAMP_ERRORS=0 ./bin/model-stamp.sh write pass "verify-model.sh" exit 0 diff --git a/project-tests/e2e/journey-runner.js b/project-tests/e2e/journey-runner.js index 0fffb151..14831616 100644 --- a/project-tests/e2e/journey-runner.js +++ b/project-tests/e2e/journey-runner.js @@ -33,6 +33,7 @@ const fs = require('fs'); const path = require('path'); const H = require(__dirname + '/helpers.js'); +const { settle } = require(__dirname + '/settle.js'); const O = require(__dirname + '/otel.js'); const { cfg } = require(__dirname + '/config.js'); @@ -288,7 +289,11 @@ async function act(page, a, vars, note) { case 'wait': await page.waitForTimeout(a.ms || 1000); break; default: throw new Error(`unknown action "${a.do}"`); } - await page.waitForTimeout(a.settleMs ?? 1200); + // An explicit settleMs is a fixed wait, as before. Otherwise wait for the page to + // finish (settle.js) rather than 1200 ms after every action; SETTLE_MODE=fixed + // restores the old 1200 ms for an A/B run. + if (a.settleMs != null) await page.waitForTimeout(a.settleMs); + else await settle(page, { timeout: 10000, stableMs: 300, networkIdleMs: 0, fallbackMs: 1200 }); } // ── Rung 3 helpers ────────────────────────────────────────────────────────── @@ -575,7 +580,7 @@ function controlMutants(j) { // carry-over inside a walk. Here, carry-over between walks is precisely the bug. async function resetToStart(page) { await page.goto(cfg.baseUrl, { waitUntil: 'domcontentloaded' }); - await page.waitForTimeout(2500); + await settle(page, { fallbackMs: 2500 }); } async function runJourney(page, j) { diff --git a/project-tests/e2e/page-audit.js b/project-tests/e2e/page-audit.js index d241f655..b8d664f7 100644 --- a/project-tests/e2e/page-audit.js +++ b/project-tests/e2e/page-audit.js @@ -11,6 +11,7 @@ // tests/e2e/artifacts/page-audit-control.json (--positive-control; own file — // a control run must never overwrite a real run) // tests/e2e/artifacts/page-audit-.png (one screenshot per page) +// tests/e2e/artifacts/page-audit-.top.png (first screen — read this one) // // WHAT THIS IS FOR, AND HOW IT DIFFERS FROM design-audit.js // `design-audit.js` sweeps the corpus for class correctness, accessibility and @@ -53,6 +54,7 @@ const path = require('path'); const { execFileSync } = require('child_process'); const R = require('./page-audit-rules.js'); +const { settle } = require('./settle.js'); // project.config.js is the only project-aware file in tests/e2e/. It is safe to // require here where ./config is not: it has NO side effects at require time @@ -547,8 +549,9 @@ async function liveSweep(pages, perPage) { } const item = page.locator(`text="${target.item}"`).first(); await item.click({ timeout: 10000 }); - await page.waitForLoadState('networkidle', { timeout: 20000 }).catch(() => {}); - await page.waitForTimeout(600); + // Page state, not 'networkidle' (a polling client only ends that by timing out, + // up to 20 s per page) plus a fixed 600 ms. settle.js; SETTLE_MODE=fixed = old waits. + await settle(page, { timeout: 20000, fallbackMs: 600 }); navigated = true; } catch (e) { why = `clicking nav item "${target.item}" failed: ${String(e.message).slice(0, 160)}`; @@ -570,8 +573,14 @@ async function liveSweep(pages, perPage) { } await page.screenshot({ path: shot, fullPage: true }).catch(() => {}); + // Plus the first screen (1440x900) — the shot an agent should LOOK at. A tall + // full-page shot is downscaled to ~1024 px high before a model sees it, which + // leaves text unreadable; the viewport shot stays legible at ~1.7k tokens. + const top = shot.replace(/\.png$/, '.top.png'); + await page.screenshot({ path: top, fullPage: false }).catch(() => {}); const captured = fs.existsSync(shot); if (pp && captured) pp.screenshot = path.relative(ROOT, shot); + if (pp && fs.existsSync(top)) pp.screenshotTop = path.relative(ROOT, top); rows.push(row({ id: `${INSTRUMENT}/live/screenshot/${qn}`, module: mod, page: qn, category: 'live', severity: 'P2', title: 'page screenshot captured', verdict: captured ? 'pass' : 'fault', diff --git a/project-tests/e2e/settle.js b/project-tests/e2e/settle.js new file mode 100644 index 00000000..2f77889a --- /dev/null +++ b/project-tests/e2e/settle.js @@ -0,0 +1,90 @@ +'use strict'; +// ============================================================================ +// settle.js — wait until a Mendix page has finished, instead of sleeping a fixed time. +// ---------------------------------------------------------------------------- +// A page counts as SETTLED when, for `stableMs` in a row: +// - no visible loading indicator (Mendix progress bar, data view / list view / grid +// loading state, spinner, aria-busy) — see BUSY; +// - the DOM size (document.body.innerHTML length) did not change; +// after at most one bounded try at Playwright's 'networkidle' (a Mendix client that +// polls never idles, so that try is capped and never fatal). +// +// settle(page) never throws. It returns { ok, ms, reason, stuck }: ok false means the +// page was still loading when the time ran out. STUCK: the time ran out with a loading +// indicator still shown but the DOM unchanged for `stuckMs` (a Data grid 2 with 0 rows +// can keep its progress bar) — the caller decides what that means. +// +// WHY (field run 2026-10-07, requirements-driven build). journey-runner slept 1200 ms +// after every action and page-audit waited for 'networkidle' up to 20 s per page, which +// a polling client only ends by timing out. Waiting on the page's own state is faster +// on a quick page and still correct on a slow one, where a fixed sleep screenshots or +// asserts on a half-loaded screen. Ported from a field-proven upgrade-test runner, +// where the fixed 1.5 s sleep had captured a progress bar instead of the page. +// +// SETTLE_MODE=fixed restores the old fixed waits (callers pass `fallbackMs`) — keep it +// for an A/B run: the same journeys must reach the same verdicts either way. +// ============================================================================ + +const BUSY = [ + '.mx-progress', '.mx-progress-indicator', '.mx-progress-line', + '.mx-dataview-loading', '.mx-listview-loading', '.mx-templategrid-loading', '.mx-datagrid-loading', + '.widget-datagrid-loading', '.widget-gallery-loading', '.mx-loading', + '.spinner', '.spinner-border', '.loading-spinner', '[aria-busy="true"]', +]; + +// Runs in the browser (page.evaluate). Self-contained: no closure over node values. +function pageSettled(busySel) { + try { + const sel = busySel.join(', '); + const shown = (e) => { + try { + if (e.getClientRects && !e.getClientRects().length) return false; + const s = getComputedStyle(e); + return s.display !== 'none' && s.visibility !== 'hidden' && s.opacity !== '0'; + } catch (x) { return false; } + }; + const busy = [...document.querySelectorAll(sel)].filter(shown) + .map((e) => (e.className && String(e.className).split(/\s+/)[0]) || e.tagName || '?'); + return { busy, size: document.body ? document.body.innerHTML.length : 0 }; + } catch (e) { + return { busy: ['error: ' + String(e && e.message || e).slice(0, 80)], size: -1 }; + } +} + +async function settle(page, opts = {}) { + if (process.env.SETTLE_MODE === 'fixed' && opts.fallbackMs != null) { + await page.waitForTimeout(opts.fallbackMs).catch(() => {}); + return { ok: true, ms: opts.fallbackMs, reason: 'SETTLE_MODE=fixed' }; + } + const timeout = opts.timeout ?? Number(process.env.SETTLE_MS || 15000); + const stableMs = opts.stableMs ?? 500; + const minMs = opts.minMs ?? 300; // let an action's own request start before judging + const stuckMs = opts.stuckMs ?? 5000; + const idleMs = opts.networkIdleMs ?? 1500; + const poll = 150; + const t0 = Date.now(); + let last = null, stableSince = 0, sizeSince = 0, lastBusy = []; + if (idleMs > 0) await page.waitForLoadState('networkidle', { timeout: Math.min(idleMs, timeout) }).catch(() => {}); + while (Date.now() - t0 < timeout) { + const s = await page.evaluate(pageSettled, BUSY) + .catch((e) => ({ busy: ['evaluate: ' + String(e.message).split('\n')[0]], size: -1 })); + lastBusy = s.busy; + if (s.size !== last || s.size < 0) sizeSince = Date.now(); + if (!s.busy.length && s.size === last) { + if (!stableSince) stableSince = Date.now(); + if (Date.now() - stableSince >= stableMs && Date.now() - t0 >= minMs) return { ok: true, ms: Date.now() - t0 }; + } else { + stableSince = 0; + } + last = s.size; + await page.waitForTimeout(poll).catch(() => {}); + } + const ms = Date.now() - t0, still = Date.now() - sizeSince; + const busy = [...new Set(lastBusy)].slice(0, 4).join(', '); + const real = lastBusy.length && !lastBusy.some((b) => /^(error|evaluate):/.test(b)); + if (real && sizeSince && still >= stuckMs) + return { ok: false, stuck: true, ms, reason: `stuck: ${busy} shown, DOM unchanged for ${still}ms` }; + return { ok: false, ms, reason: lastBusy.length ? 'still loading: ' + busy : 'DOM still changing' }; +} + +module.exports = { BUSY, pageSettled, settle }; diff --git a/skills/journey-examples.md b/skills/journey-examples.md index 0ebf02ca..72061965 100644 --- a/skills/journey-examples.md +++ b/skills/journey-examples.md @@ -123,7 +123,7 @@ row means "the walk did not get there", never "not applicable". | `wait` | `ms` | **`ms`, not `settleMs`.** | | anything else | — | throws `unknown action` → step `FAIL`. | -Every action also reads `settleMs` (default 1200, applied *after* the action) and `label` +Every action also reads `settleMs` (a fixed wait *after* the action; when absent the runner waits until the page is idle — no progress bar, DOM still — via `tests/e2e/settle.js`, and `SETTLE_MODE=fixed` restores the old 1200) and `label` (**display only** — the walk narrates `Click "Submit for approval"` while keeping `.mx-name-btnSubmit` for whoever has to fix it). `widget` is a bare Mendix widget name; the runner prepends `.mx-name-`. @@ -254,7 +254,7 @@ Each of these is either an error path in the runner or a correction recorded in | `oql` / `outcome` with no `expect` or `atLeast` | `INVALID` on every run, forever | declare a bar; prefer `atLeast` on a fixture with history | | Two independent `LIMIT 1` seeds | the journey selects one row and asserts about another | key the second seed off the first | | Unaliased seed SQL | seed resolves to `null` → whole journey `INVALID` | `SELECT Name AS n …` | -| `wait` with `settleMs` | the wait is 1200ms, not what you wrote | `wait` reads `ms` | +| `wait` with `settleMs` | the wait is `settleMs` after a 1000ms `wait`, not what you wrote | `wait` reads `ms` | | A `_note` object inside `textPresent` / `ordered` / `checks` | the comment becomes an assertion | comment on the parent object, never inside an iterated array | | Treating "seeds returned nothing" as a feature failure | someone debugs the app for a fixture gap | `INVALID`; see `fixture-seeding.md`'s four-cause table | | `persona` drifting from `TEST_USER` | the report names an identity that never walked | keep them in step; check `usedFallback` in the findings file | diff --git a/skills/ui-loop.md b/skills/ui-loop.md index 2151a828..2dd7d075 100644 --- a/skills/ui-loop.md +++ b/skills/ui-loop.md @@ -47,7 +47,10 @@ The gate was not missing. Its cadence was too coarse to catch anything early. After a script that creates or changes a page: 1. **Run it and open the page** — through real navigation, as a real user role, not a direct URL. -2. **Screenshot it, then open the PNG** (the Read tool). Name the file after the page — +2. **Screenshot it, then open the PNG** (the Read tool). Take the viewport (`fullPage: false`, + e.g. 1440×900), not the whole page: a tall full-page shot reaches you shrunk to ~1024 px high + with unreadable text. Below the fold, scroll and take a second viewport shot. + Name the file after the page — `order-overview.png` or `Order_Overview_1280.png` for `Orders.Order_Overview`. `bin/exec.sh` recorded the page as owed a look when the script landed; opening a screenshot whose name contains the page name, taken after that build, is what clears it. Until every built page is From 85afba1aa880d683b78206faf290523b35de6392 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 12:47:15 +0000 Subject: [PATCH 2/6] perf(ui-loop): reload a stale app, settle journeys, worst-first look order - test-stack-up.sh: record the model the app was built from; on change, `mxcli docker reload` (+ --skip-check when exec.sh already passed it), restart only on failure. Header carries the measured loop table. - journey-runner.js: combobox steps wait on page state; the toggle is scoped to the named widget (it opened the page's first combobox). - page-audit.js: click the first nav copy a user can click (a collapsed sidebar covered .first()); write page-audit-look-order.txt, all pages worst first. - design-audit.js: drop the fixed 1.2 s wait. - ui-loop.md / module-review.md: warm-loop table; order, not selection. Field run 2026-10-07, existing app: same 7-step journey 56 s fixed vs 24 s settle, 11/11 pass both, 3 runs each; reload --skip-check 93 s vs restart 179 s; run --local --watch ~18 s per page change. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- CHANGELOG.md | 1 + project-bin/test-stack-up.sh | 117 +++++++++++++++++++++++++++- project-tests/e2e/design-audit.js | 4 +- project-tests/e2e/journey-runner.js | 22 ++++-- project-tests/e2e/page-audit.js | 48 +++++++++++- skills/module-review.md | 7 ++ skills/ui-loop.md | 14 ++++ 7 files changed, 200 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f99f456b..4d65e66e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,7 @@ Sections dated before 2026-09-19 predate the cycle and stay as they are. ## Unreleased - perf(ui-loop): **fewer mxbuilds, no fixed sleeps, legible screenshots.** Field run 2026-10-07 (a requirements-driven build): `exec.sh` was 22% of tool time at two mxbuilds per exec, journeys and page waits slept a fixed time, and screenshots were 11.5% of re-read tokens. `exec.sh` now skips the pre-flight baseline mxbuild when the model is byte-identical to its last *clean* pass: the verification stamp gains an `errors:` field (`model-stamp.sh check --clean` accepts only `errors: 0`, so a delta-gate pass over pre-existing errors never counts; `SKIP_BASELINE=0` is unchanged). New `tests/e2e/settle.js` waits for the page's own state (no visible Mendix progress/loading indicator, DOM size stable) instead of a fixed time; `journey-runner.js` uses it after each action (an explicit `settleMs` still wins) and `page-audit.js` per page. Probed in a real browser: on a page that loads in 1.8 s the old 1200 ms sleep and networkidle+600 both captured it half-loaded, settle waited for it; on a fast page settle took 775 ms vs 1200. `SETTLE_MODE=fixed` restores the old waits for an A/B run. `page-audit.js` also saves a viewport shot `page-audit-.top.png` beside the full-page one, and `ui-loop.md` says to look at viewport shots: a tall full-page shot reaches the model at ~1024 px high, text unreadable. `shrink-image-read.sh` now works off macOS (ImageMagick or Pillow, python when jq is absent; it was a silent no-op without `sips`). — MendixMau +- perf(ui-loop): **a stale app reloads instead of restarting, journeys stop sleeping, worst pages are looked at first.** Field run 2026-10-07, existing app. `test-stack-up.sh` records which model the running app was built from (`.claude/loop/served-model`); when the model has changed since, it hot-reloads with `mxcli docker reload` (adding `--skip-check` when the exec.sh gate already passed this model) and restarts only if that fails: 93 s instead of 179 s. `ui-loop.md` and the script header carry the measured table and point at `mxcli run --local --watch` as the fast loop (~18 s per page change). `journey-runner.js` waits on page state after a combobox action instead of fixed pauses, and a combobox step now opens the widget it names: it used to open the page's first combobox, so on a form with three every step but the first failed. `page-audit.js` clicks the first copy of a nav item a user could actually click, not just the first match: with the sidebar collapsed, every in-nav page faulted while the same item sat in the top menu bar. It also writes `page-audit-look-order.txt`, all pages worst first (`module-review.md` §4a: order, not selection, every page is still looked at). `design-audit.js` drops its fixed 1.2 s wait. Same 7-step journey on the same app: 56 s with `SETTLE_MODE=fixed`, 24 s with settle, 11/11 pass both ways, 3 runs each. — MendixMau - feat(permissions): **intake now picks how much Claude Code may run without a prompt, and `doctor.sh` shows it.** Until now the allow-list covered only the safe wrappers (`bin/exec.sh`, `bin/*.sh`, `./mxcli`, `mx`), so every edit, test run, `npm install` or `git commit` in a project still prompted. New `bin/permission-profile.sh` resolves a profile (env, `.mxtk/permission-profile`, PROJECT.md `Permission profile:`, default `project`; a typo falls back to `wrappers`): `wrappers` is the old list; `project` adds edits anywhere inside the project folder (`Edit(/**)`) and the everyday test/npm/docker/git-add/commit commands, with force-push, `git reset --hard`, `git clean` and `sudo` behind `ask` rules that prompt in every mode; `full` adds bare `Bash`/`Edit`/`Read`/`WebFetch`/`WebSearch` to the per-machine `settings.local.json` only, never the shared file. `install-claude-permissions.sh` / `install-harness-permissions.sh` take `--profile`, and switching to a smaller profile removes only what the script itself added. Intake gains Q12, `init-project.sh` writes `Permission profile: project`, and `doctor.sh ` reports the profile, whether the settings files carry it, and warns on a `defaultMode` of auto/bypass in a project file, which Claude Code ignores there. No profile can switch on bypass or auto mode from the project: doctor says how to do that from the user's own settings. `test-install-claude-permissions.sh` gains T5b. — MendixMau - fix(small): **five small field defects (#146, #149, #152, #153, #156) and the missing v0.24 ledger entries (#168).** `gate-check.sh --closeout ` now writes `index.html` when that stage PASSes or is WAIVED, so the board moves at every gate the runbook closes; plain stage queries stay read-only (#146). `journey-runner.js` refuses a journey whose `persona` differs from the signed-in user — INVALID, not walked, naming the `TEST_USER` to use — where before a non-admin journey walked green under admin; `journey-proof.md` says one file, one run per role, and `test-journey-persona.sh` is new (#149). `design-audit.js` rung 7 exempts the three `spacing-outer-bottom*` classes `design-spacing.md` prescribes (still FAIL when the widget also sets the Spacing property), and the skill names the list (#152). `verify-module.sh` runs `REFRESH CATALOG FULL` before the model-side instruments when the catalog is older than the `.mpr`, so `test --attach` then verify no longer reads INCOMPLETE (probed on a scratch model: stale, refreshed, fresh, then a no-op) (#153). `wiring-sweep.md` adds N/A (by design) for a control its own state disables and UNREACHED for a conditional widget the sweep never saw, and enumerates `Visible:`/`Editable:` widgets from `describe page` so the denominator is not the DOM's alone (#156). `mxcli-bugs.md` gains four BUG-DRAFT entries from a field comparison build (2026-09-29): false page-setting warnings, the unwritable pluggable `file` property, `index on createdDate`, and the `describe` replay CE1613; the inbox note is promoted and removed. Upstream asks in #168 stay open. — MendixMau - fix(gates): **a red gate under an approval, a `done-` deadlock and unfilled agent slots are now visible.** A field project had Stage 4 CONFIRMED in `PROJECT.md` while its gate FAILED, 47 scripts at gate-pass with none renamed `done-`, and 35 unfilled `{{...}}` slots in its agent stubs, and no instrument said so (a requirements-driven field build, 2026-10-06; #207, #210, #211). `gate-check.sh` now prefixes a FAIL on a CONFIRMED stage with `APPROVED OVER A FAILING GATE` and adds a summary line under Needs attention (verdict and exit code unchanged). `status.sh` gains a `WATCH` line for all three plus a stale `build-plan.html`, and its NEXT names the first one; its not-yet-`done-` list now skips BUILD-LOG rows for scripts no longer on disk. New `bin/lib/placeholders.sh` counts slots (excluding the `{{DOUBLE_BRACE}}` prose); `init-project.sh` and `sync-project.sh` list them per file. Slots stay unfilled until each agent's stage starts, as `agent-roles.md` says: these are reports, not failures. — MendixMau diff --git a/project-bin/test-stack-up.sh b/project-bin/test-stack-up.sh index 8d1a0b08..09abf59f 100755 --- a/project-bin/test-stack-up.sh +++ b/project-bin/test-stack-up.sh @@ -10,15 +10,43 @@ # 1. NEVER touches Studio Pro. No force-quit, no kill. It detects SP and reports; that is all. # restart-sp.sh owns crash recovery and warns that unsaved work is lost. This must be safe to # run while someone else is working in SP. -# 2. NEVER writes to the .mpr. `mxcli docker build` reads the model and writes only build output. +# 2. Never writes model CONTENT. Measured (field run 2026-10-07, existing app): a `mxcli docker +# run` / `docker build` runs `mx update-widgets` first, which rewrites widget definitions in the +# .mpr — git showed it modified after a build that failed. That is a widget-definition refresh, +# not a change anyone authored, but it IS a write: commit or snapshot before a build, and never +# run one while Studio Pro has the model open. # 3. Proves liveness with an actual HTTP response. "No error" is not "up". -# 4. Idempotent. If the app already answers, it does nothing and exits 0. +# 4. Idempotent, but never STALE. If the app already answers and was built from the current model, +# it does nothing and exits 0. If the model changed since this script last built it, an app +# that answers is serving the OLD model — every test then measures yesterday's build and a +# fixed bug "still fails". So it hot-reloads (`mxcli docker reload`: build + reload_model, +# no container restart) and falls back to a restart only if the reload fails. See WARM LOOP. # 5. Starts an auxiliary service if it is down; NEVER stops one (another session may be using it). # # Usage: # bin/test-stack-up.sh --check # report only, change nothing. Exit 0 iff required deps are up. # bin/test-stack-up.sh # report, then bring up what is missing (may run a Docker build) # bin/test-stack-up.sh --no-docker # bring up auxiliaries only; never build. "SP is already running it". +# bin/test-stack-up.sh --restart # skip the hot reload; rebuild and restart the container (old path) +# +# WARM LOOP. After a model change the old loop was: exec.sh gate (mxbuild) -> `docker run` (mx check + +# mxbuild again) -> container restart -> boot wait. Now: the app's model is recorded in +# .claude/loop/served-model (the model-stamp fingerprint at the moment this script built it); when +# the current fingerprint differs, `mxcli docker reload` rebuilds and swaps the model into the +# running runtime instead of restarting it, and passes --skip-check when model-stamp says the +# current model already passed an mxbuild (the exec.sh gate) — that check was the duplicate. +# Entity/association changes need a DB sync that reload_model may not do; if the reload fails or the +# app stops answering, it falls back to the restart. MXTK_WARM_RELOAD=off keeps the old idempotent +# no-op. Measured (field run 2026-10-07, existing app, cloud container, one page change): +# docker run --wait, cold 190 s +# restart (docker run --wait over a running app) 179 s +# docker reload, with the duplicate check 140 s +# docker reload --skip-check 93 s (reload_model itself 0.6 s; the build is the rest) +# mxcli run --local --watch, page change ~18 s from exec to applied (security/nav change: ~60 s restart) +# So the fastest loop is not this script: keep `mxcli run --local --watch` running beside the +# session and let it apply each exec (skills/ui-loop.md). This script is the Docker path and the +# proof-of-ownership step. `docker reload --css` is NOT a theme loop: it copies the theme without +# compiling SCSS, so a main.scss edit does not show. # # Exit codes: 0 = required stack is up · 1 = not up and could not fix · 2 = usage/env error # @@ -70,8 +98,9 @@ MODE="up" case "${1:-}" in --check) MODE="check" ;; --no-docker) MODE="nodocker" ;; + --restart) MODE="restart" ;; "") ;; - *) echo "usage: $0 [--check|--no-docker]" >&2; exit 2 ;; + *) echo "usage: $0 [--check|--no-docker|--restart]" >&2; exit 2 ;; esac ok() { printf ' \033[32m✓\033[0m %s\n' "$1"; } @@ -201,6 +230,29 @@ publish_stack_env() { } > "$STACK_ENV" } +# --- Served model: which model state is the running app built from? ---------- +# Written ONLY by this script, right after it built the app, so it names the model the build read +# (after mx update-widgets, which can rewrite the .mpr — fingerprint before the build would be +# stale the moment the build finished). An app started any other way (Studio Pro, a manual +# `docker run`) has no record: freshness UNKNOWN, reported, never guessed. +SERVED_FILE="${SERVED_FILE:-$ROOT/.claude/loop/served-model}" +STAMP_SH="$(dirname "${BASH_SOURCE[0]}")/model-stamp.sh" +model_fp() { [ -f "$STAMP_SH" ] && bash "$STAMP_SH" fingerprint 2>/dev/null; } +record_served() { + local fp; fp="$(model_fp)" || fp="" + [ -n "$fp" ] || return 0 + mkdir -p "$(dirname "$SERVED_FILE")" 2>/dev/null && printf '%s\n' "$fp" > "$SERVED_FILE" +} +# Echoes fresh | stale | unknown. +served_state() { + local want have + have="$(head -1 "$SERVED_FILE" 2>/dev/null)" + want="$(model_fp)" || want="" + if [ -z "$have" ] || [ -z "$want" ]; then echo unknown + elif [ "$have" = "$want" ]; then echo fresh + else echo stale; fi +} + echo "── Test stack: $PROJ ──────────────────────────────────" # --- Studio Pro: report only, never act ------------------------------------- @@ -236,6 +288,15 @@ if [ $APP_FOUND -eq 0 ]; then else bad "App not serving on any of: $APP_PORTS" fi +SERVED="none" +if [ $APP_FOUND -eq 0 ]; then + SERVED="$(served_state)" + case "$SERVED" in + fresh) ok "App serves the current model" ;; + stale) warn "App serves an OLDER model — the model changed since it was built; tests would measure the old build" ;; + unknown) warn "App freshness unknown — not built by this script (Studio Pro / manual run); it may serve an old model" ;; + esac +fi # --- Jaeger (needed only for Trace assertions) ------------------------------ JAEGER_OK=1 @@ -293,6 +354,10 @@ if [ "$MODE" = "check" ]; then echo "APP UP, BUT REST-FED SPECS WILL BE INVALID — see the container warning above" exit 1 fi + if [ "$SERVED" = "stale" ] && [ "$APP_OWNERSHIP" = "verified" ]; then + echo "APP UP BUT STALE — re-run without --check to hot-reload the current model" + exit 1 + fi if [ $APP_FOUND -eq 0 ] && [ $MOCK_OK -eq 0 ]; then echo "READY (app :$APP_PORT · jaeger $([ $JAEGER_OK -eq 0 ] && echo up || echo DOWN))" exit 0 @@ -317,6 +382,51 @@ if [ $MOCK_REQUIRED -eq 1 ] && [ $MOCK_OK -ne 0 ]; then fi fi +# --- Warm loop: refresh a running app instead of restarting it -------------- +# Only for OUR container (ownership verified): a reload swaps the model in a runtime, and doing that +# to an unverified port could be another project's app or a Studio Pro run. +if [ $APP_FOUND -eq 0 ] && [ "$APP_OWNERSHIP" = "verified" ] && [ -n "$MXCLI" ] \ + && [ "${MXTK_WARM_RELOAD:-on}" != "off" ] && [ "$MODE" != "nodocker" ] \ + && { [ "$SERVED" != "fresh" ] || [ "$MODE" = "restart" ]; }; then + T0=$(date +%s) + RELOADED=1 + if [ "$MODE" != "restart" ]; then + # --skip-check only when this exact model already passed an mxbuild (exec.sh's gate or + # verify-model.sh). Otherwise the check is the only build-error report before the build. + SKIP="" + if [ -f "$STAMP_SH" ] && bash "$STAMP_SH" check -q >/dev/null 2>&1; then SKIP="--skip-check"; fi + echo "→ Model changed since the app was built — hot-reloading (build + reload_model${SKIP:+, check skipped: model already gate-verified})..." + "$MXCLI" docker reload -p "$MPR" $SKIP >"$DOCKER_LOG" 2>&1 + RELOADED=$? + if [ $RELOADED -eq 0 ]; then + # reload_model returned; prove the app still answers before trusting it. + for _ in $(seq 1 20); do + APP_FIND="$(find_app_port)" && break + sleep 1 + done + if mendix_at "${APP_FIND%%|*}"; then + record_served + ok "Hot reload done in $(( $(date +%s) - T0 ))s — app serves the current model" + SERVED=fresh + else + warn "reload_model returned but the app stopped answering — falling back to a restart" + RELOADED=1 + fi + else + warn "docker reload failed (rc=$RELOADED) — falling back to a restart. Log: $DOCKER_LOG" + tail -5 "$DOCKER_LOG" | sed 's/^/ /' + fi + fi + if [ $RELOADED -ne 0 ]; then + echo "→ Rebuilding and restarting the container..." + "$MXCLI" docker run -p "$MPR" --wait >"$DOCKER_LOG" 2>&1 + if [ $? -ne 0 ]; then + bad "mxcli docker run failed — see $DOCKER_LOG"; tail -20 "$DOCKER_LOG"; exit 1 + fi + APP_FOUND=1 # re-proved below by the boot-wait loop + fi +fi + # --- Bring up the app ------------------------------------------------------- if [ $APP_FOUND -ne 0 ]; then if [ "$MODE" = "nodocker" ]; then @@ -374,6 +484,7 @@ if [ $APP_FOUND -ne 0 ]; then APP_OWNERSHIP="${APP_FIND##*|}" if [ $APP_FOUND -eq 0 ]; then + record_served ok "App serving on :$APP_PORT after ${WAITED}s (ownership $APP_OWNERSHIP)" else bad "App still not answering after ${BOOT_TIMEOUT}s — see $DOCKER_LOG" diff --git a/project-tests/e2e/design-audit.js b/project-tests/e2e/design-audit.js index 4c92d8e4..d66c0598 100644 --- a/project-tests/e2e/design-audit.js +++ b/project-tests/e2e/design-audit.js @@ -47,6 +47,7 @@ const os = require('os'); // project.config.js is the only project-aware file in tests/e2e/, and it has no // side effects at require time (unlike ./config, which can process.exit(2)). const PROJ = require('./project.config'); +const { settle } = require(__dirname + '/settle.js'); const ROOT = PROJ.root; const MPR = PROJ.mprName; @@ -771,7 +772,8 @@ async function appSweep(pages, navMap, controlNotes) { try { await page.setViewportSize({ width: 1440, height: 900 }); await helpers.navTo(page, nav.group, nav.item); - await page.waitForTimeout(1200); + // wait for the page, not 1200 ms (SETTLE_MODE=fixed keeps the 1200 ms for an A/B run) + await settle(page, { timeout: 10000, networkIdleMs: 0, fallbackMs: 1200 }); const st = await checkStructure(page); const shot = path.join(ARTIFACTS, `design-audit-${qn}.png`); diff --git a/project-tests/e2e/journey-runner.js b/project-tests/e2e/journey-runner.js index 14831616..d0140248 100644 --- a/project-tests/e2e/journey-runner.js +++ b/project-tests/e2e/journey-runner.js @@ -115,6 +115,11 @@ const subst = (v, vars) => // meant to stay declarative enough to generate from a module brief's golden-path // table. const PAUSE = ms => new Promise(r => setTimeout(r, ms)); +// The combobox waits below used to be fixed sleeps (400/1200/120/500 ms per field). Each now +// waits for the thing it was sleeping for; SETTLE_MODE=fixed keeps the old sleeps for an A/B run. +const FIXED = process.env.SETTLE_MODE === 'fixed'; +const frames = page => page.evaluate(() => new Promise(r => + requestAnimationFrame(() => requestAnimationFrame(r)))).catch(() => {}); // A human-readable account of what the action WILL do, in the vocabulary a reviewer // uses ("click Receive"), not the harness's (".mx-name-bReceive"). Both are kept: the @@ -200,12 +205,18 @@ async function act(page, a, vars, note) { if (!(await box.isVisible({ timeout: 6000 }).catch(() => false))) throw new Error(`.mx-name-${a.widget} not visible`); await box.scrollIntoViewIfNeeded().catch(() => {}); - await PAUSE(400); - await page.locator('[id^="downshift-"][id$="-toggle-button"]').first() + if (FIXED) await PAUSE(400); else await frames(page); + // The toggle is INSIDE the widget (only the menu renders outside it). It used to be + // looked up page-wide with .first(), which opened the page's FIRST combobox whatever + // `widget` said: on a form with three comboboxes every step read the first one's menu + // and failed "no option matches" (field run 2026-10-07, existing app). + await box.locator('[id^="downshift-"][id$="-toggle-button"]').first() .click({ force: true }).catch(() => {}); - await PAUSE(1200); + if (FIXED) await PAUSE(1200); const opts = page.locator('[role="option"]'); await opts.first().waitFor({ state: 'visible', timeout: 5000 }).catch(() => {}); + // options can arrive in batches (async datasource): wait until the list stops growing + if (!FIXED) await settle(page, { timeout: 3000, stableMs: 200, minMs: 0, networkIdleMs: 0 }); const texts = await opts.evaluateAll(e => e.map(x => x.textContent.trim())).catch(() => []); const want = String(subst(a.match, vars)); const idx = texts.findIndex(t => t.includes(want)); @@ -261,7 +272,7 @@ async function act(page, a, vars, note) { const cur = (await activeText() || '').trim(); if (cur && cur.includes(want)) { landed = true; break; } await page.keyboard.press('ArrowDown'); - await PAUSE(120); + if (FIXED) await PAUSE(120); else await frames(page); } if (!landed) { throw new Error( @@ -269,7 +280,8 @@ async function act(page, a, vars, note) { `"${want}". The option is listed but not reachable by keyboard.`); } await page.keyboard.press('Enter'); - await PAUSE(500); + if (FIXED) await PAUSE(500); + else await settle(page, { timeout: 3000, stableMs: 150, minMs: 0, networkIdleMs: 0 }); // Selecting is not committing. Downshift can close the menu without writing the // value back, so the committed input is read and compared to the seed — not diff --git a/project-tests/e2e/page-audit.js b/project-tests/e2e/page-audit.js index b8d664f7..7fbf847c 100644 --- a/project-tests/e2e/page-audit.js +++ b/project-tests/e2e/page-audit.js @@ -547,8 +547,22 @@ async function liveSweep(pages, perPage) { const g = page.locator(`.mx-navigationtree >> text="${target.group}"`).first(); if (await g.count()) { await g.click({ timeout: 5000 }).catch(() => {}); await page.waitForTimeout(400); } } - const item = page.locator(`text="${target.item}"`).first(); - await item.click({ timeout: 10000 }); + // Try every element carrying the label, not just the first. A layout can render + // the same item twice (sidebar tree + top menu bar); with the sidebar collapsed + // to its 32 px rail the content placeholder covers it, so .first() timed out on + // every page while the menu-bar copy was one click away (field run 2026-10-07, + // existing app). A trial click finds the one a user could actually click. + const items = page.locator(`text="${target.item}"`); + await items.first().waitFor({ state: 'attached', timeout: 10000 }); + let clicked = false; let lastErr = null; + for (let i = 0, n = await items.count(); i < n && !clicked; i++) { + try { + await items.nth(i).click({ trial: true, timeout: 2000 }); + await items.nth(i).click({ timeout: 5000 }); + clicked = true; + } catch (err) { lastErr = err; } + } + if (!clicked) throw lastErr; // Page state, not 'networkidle' (a polling client only ends that by timing out, // up to 20 s per page) plus a fixed 600 ms. settle.js; SETTLE_MODE=fixed = old waits. await settle(page, { timeout: 20000, fallbackMs: 600 }); @@ -911,7 +925,25 @@ function main() { }); } +// Look order: every page, worst first — fault, then fail, then by its P1/P2/P3 finding counts. +// Triage only changes which screenshot a reviewer opens FIRST; the list always holds all N pages, +// because a page with zero findings is not a page nobody needs to look at (the rules cannot see +// layout, and module-review.md stage 4 owes a look at every one). +const SEV_RANK = ['P1', 'P2', 'P3']; +function lookOrder(perPage) { + const key = (p) => [PRECEDENCE[p.verdict] || 0, + ...SEV_RANK.map((s) => p.findings.filter((f) => f.severity === s).length), p.findings.length]; + return perPage.map((p, i) => ({ p, i, k: key(p) })) + .sort((a, b) => { for (let j = 0; j < a.k.length; j++) if (b.k[j] !== a.k[j]) return b.k[j] - a.k[j]; return a.i - b.i; }) + .map(({ p }, n) => ({ + rank: n + 1, page: p.page, verdict: p.verdict, + findings: SEV_RANK.map((s) => `${s}:${p.findings.filter((f) => f.severity === s).length}`).join(' '), + screenshotTop: p.screenshotTop || null, screenshot: p.screenshot || null, + })); +} + function finish({ startedAt, checks, instruments, perPage, coverage, app, controlRows }) { + const order = lookOrder(perPage); const tally = checks.reduce((a, c) => { a[c.verdict] = (a[c.verdict] || 0) + 1; return a; }, {}); const artifact = { schemaVersion: SCHEMA_VERSION, @@ -938,6 +970,7 @@ function finish({ startedAt, checks, instruments, perPage, coverage, app, contro }, humanJudgement: R.HUMAN_JUDGEMENT, instruments, + lookOrder: order, perPage, checks, controls: controlRows, @@ -948,6 +981,11 @@ function finish({ startedAt, checks, instruments, perPage, coverage, app, contro const out = OPT.out || path.join(ARTIFACTS, OPT.page ? `page-audit-${OPT.page}.json` : 'page-audit.json'); fs.writeFileSync(out, JSON.stringify(artifact, null, 2)); + if (!OPT.page && !OPT.out && order.length) { + fs.writeFileSync(path.join(ARTIFACTS, 'page-audit-look-order.txt'), + `# Look at ALL ${order.length} pages; this only sets which first (worst first).\n` + + order.map((o) => `${o.rank}\t${o.verdict}\t${o.findings}\t${o.page}\t${o.screenshotTop || o.screenshot || '(no screenshot)'}`).join('\n') + '\n'); + } // ── console summary ──────────────────────────────────────────────────────── console.log(`\npage-audit — ${artifact.run.mode} — ${perPage.length} pages, ${checks.length} checks`); @@ -955,12 +993,14 @@ function finish({ startedAt, checks, instruments, perPage, coverage, app, contro console.log(`wireframe coverage: ${coverage.withWireframe}/${coverage.inScopePages} in-scope pages have one ` + `(${coverage.withoutWireframe} do not: ${coverage.pagesWithNoWireframe.join(', ') || 'none'})`); } - for (const p of perPage) { + const byPage = new Map(perPage.map((p) => [p.page, p])); + for (const p of order.map((o) => byPage.get(o.page))) { const f = p.findings.length; console.log(` ${p.verdict.toUpperCase().padEnd(5)} ${p.page.padEnd(42)} ${p.checkCount || 0} checks, ${f} finding${f === 1 ? '' : 's'}` + (p.wireframe && !p.wireframe.exists ? ' [NO WIREFRAME]' : '')); } console.log(`\nverdicts: ${JSON.stringify(tally)} → ${out}`); + if (order.length) console.log(`look order (worst first, all ${order.length} pages): ${path.join(ARTIFACTS, 'page-audit-look-order.txt')}`); const anyFault = checks.some((c) => c.verdict === 'fault'); const anyFail = checks.some((c) => c.verdict === 'fail'); @@ -975,4 +1015,4 @@ if (require.main === module) { } catch (e) { console.error(e.stack); process.exitCode = 2; } } -module.exports = { auditPageStatic, readCompleteness, worst, DOM_PROBE, PAGE_WIREFRAME }; +module.exports = { auditPageStatic, readCompleteness, worst, lookOrder, DOM_PROBE, PAGE_WIREFRAME }; diff --git a/skills/module-review.md b/skills/module-review.md index fe653082..e2141a69 100644 --- a/skills/module-review.md +++ b/skills/module-review.md @@ -80,6 +80,13 @@ Both were invisible to a report that enumerated only what ran. Record the page set and its size in the report. `12 of 12 pages reviewed` is a claim. `reviewed the module` is not. +**Order, not selection.** When `tests/e2e/page-audit.js` has run, open the screenshots in the +order of `tests/e2e/artifacts/page-audit-look-order.txt`: worst first (fault, then fail, then by +P1/P2/P3 finding count), so attention is freshest where the rules already found trouble. The file +lists **all N pages**, its header says so, and the denominator does not change. A page at the +bottom with zero findings still gets its look: the rules cannot see layout, and "the audit was +clean" is exactly the "looked fine" this stage exists to refuse. + ### 4b. Run the mechanical sweep first, then look at what it cannot see ```bash diff --git a/skills/ui-loop.md b/skills/ui-loop.md index 2dd7d075..78b5e64d 100644 --- a/skills/ui-loop.md +++ b/skills/ui-loop.md @@ -169,6 +169,20 @@ the "fix" (`Class: 'form-vertical'`) landed in the model, passed the gate, and c nothing on screen, because the theme, not the page, decides the form orientation — the first row of the table below, met on the first attempt. +**Keep the app current without rebuilding it.** Same measure on an existing app (field run +2026-10-07, cloud container, one page change), from exec to the change on screen: + +| How the app picks up the change | Time | +|---|---| +| `mxcli run --local --watch` kept running (page change) | ~18 s | +| same, security or navigation change (it restarts itself) | ~60 s | +| `bin/test-stack-up.sh` → `docker reload --skip-check` | 93 s | +| restart (`docker run --wait`) | 179 s | + +So keep `run --local --watch` running beside the session; `test-stack-up.sh` notices a stale +Docker app and reloads it, but that path is five times slower. `docker reload --css` does not +compile SCSS: a `main.scss` edit needs the build. + ## When it finds something — which side is actually wrong The symptom is easy to see. **Where the fix goes is not, and the instinct is usually wrong.** From a908b4e565678ad3fadb11ae97ebe39b3d1c1325 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 13:19:40 +0000 Subject: [PATCH 3/6] test: journey fixtures copy settle.js with the engine journey-runner.js now requires settle.js; the two fixtures that copy the engine into a scratch project listed the files by hand and missed it, so the require failed (MODULE_NOT_FOUND) before any assertion ran. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- tests/wave2/test-journey-control-exit.sh | 2 +- tests/wave2/test-journey-persona.sh | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/wave2/test-journey-control-exit.sh b/tests/wave2/test-journey-control-exit.sh index 6d5a469f..8cbed98a 100755 --- a/tests/wave2/test-journey-control-exit.sh +++ b/tests/wave2/test-journey-control-exit.sh @@ -45,7 +45,7 @@ bad() { FAIL=$((FAIL+1)); echo " FAIL — $1"; [ -n "${2:-}" ] && echo " WORK="$(mktemp -d "${TMPDIR:-/tmp}/journey-exit.XXXXXX")" trap 'rm -rf "$WORK"' EXIT mkdir -p "$WORK/tests/e2e"; : > "$WORK/Fixture.mpr" -cp "$SUT" "$E2E/helpers.js" "$E2E/otel.js" "$E2E/config.js" "$WORK/tests/e2e/" +cp "$SUT" "$E2E/helpers.js" "$E2E/settle.js" "$E2E/otel.js" "$E2E/config.js" "$WORK/tests/e2e/" cp "$CFG" "$WORK/tests/e2e/project.config.js" IFS= read -r -d '' JS <<'EOF' || true diff --git a/tests/wave2/test-journey-persona.sh b/tests/wave2/test-journey-persona.sh index 33d3454d..c44e779c 100755 --- a/tests/wave2/test-journey-persona.sh +++ b/tests/wave2/test-journey-persona.sh @@ -35,7 +35,7 @@ bad() { FAIL=$((FAIL+1)); echo " FAIL — $1"; [ -n "${2:-}" ] && echo " WORK="$(mktemp -d "${TMPDIR:-/tmp}/journey-persona.XXXXXX")" trap 'rm -rf "$WORK"' EXIT mkdir -p "$WORK/tests/e2e"; : > "$WORK/Fixture.mpr" -cp "$SUT" "$E2E/helpers.js" "$E2E/otel.js" "$E2E/config.js" "$WORK/tests/e2e/" +cp "$SUT" "$E2E/helpers.js" "$E2E/settle.js" "$E2E/otel.js" "$E2E/config.js" "$WORK/tests/e2e/" cp "$CFG" "$WORK/tests/e2e/project.config.js" IFS= read -r -d '' JS <<'JSEOF' || true From 6444992116dbe4dc9bc506701069ce49db782185 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 13:56:25 +0000 Subject: [PATCH 4/6] perf(ui-loop): test-stack-up recognises mxcli run --local; restart builds once - A live `mxcli run --local` (its .mxcli/run-local.json beside this project's .mpr, pid alive) is now published as APP_OWNERSHIP=verified, APP_SOURCE=run-local instead of an unverified port-scan hit the e2e harness refuses. - Staleness knows the loop: with --watch a model edit reads "watch" (ready), without it "APP UP BUT STALE - restart with --watch". - Two-tree checkouts (app/X.mpr, with or without a root symlink) are resolved. - docker path uses `docker run --wait` instead of the 240s poll; --restart no longer runs the build twice. - Fixture tests/wave2/test-stack-up-local.sh over a golden captured from a real run and value-scrubbed (field run 2026-10-07, existing app). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- CHANGELOG.md | 1 + project-bin/test-stack-up.sh | 169 ++++++++++++++---- project-tests/e2e/config.js | 3 +- skills/harness-architecture.md | 6 +- skills/report-schema.md | 2 +- skills/ui-loop.md | 6 + tests/wave2/fixtures/run-local/run-local.json | 22 +++ tests/wave2/test-stack-up-local.sh | 147 +++++++++++++++ 8 files changed, 316 insertions(+), 40 deletions(-) create mode 100644 tests/wave2/fixtures/run-local/run-local.json create mode 100755 tests/wave2/test-stack-up-local.sh diff --git a/CHANGELOG.md b/CHANGELOG.md index 4d65e66e..5275228f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -17,6 +17,7 @@ Sections dated before 2026-09-19 predate the cycle and stay as they are. ## Unreleased +- perf(ui-loop): **`test-stack-up.sh` recognises a live `mxcli run --local` as this project's app, and `--restart` no longer builds twice.** The script read only Docker labels, so an app started with `./mxcli run --local --watch` (mxcli's own ~20s page loop) was found by port scan, published as `APP_OWNERSHIP=unverified`, and refused by the e2e harness unless `ALLOW_UNVERIFIED_APP=1`. It now reads mxcli's `.mxcli/run-local.json` handshake (v0.24+; beside the `.mpr` or the `-p` path on a two-tree checkout), checks its pid is alive, and publishes `verified` with a new `APP_SOURCE=run-local` line in `stack.env` (`docker` / `scan` otherwise). A model edited after boot reads as applied under `--watch` and as STALE without it — the script never restarts a loop it did not start. Docker bring-up and restart now wait on `mxcli docker run --wait` instead of a 240s poll loop, and still prove HTTP; `--restart` used to fall through into the bring-up block and run `docker run` a second time. New fixture `tests/wave2/test-stack-up-local.sh` over a redacted capture of the handshake. Field run 2026-10-07, existing app: same app, same port, `unverified` → `verified`; a touched model read `watch`; stopping the loop removed the handshake and the script fell back. — MendixMau - perf(ui-loop): **fewer mxbuilds, no fixed sleeps, legible screenshots.** Field run 2026-10-07 (a requirements-driven build): `exec.sh` was 22% of tool time at two mxbuilds per exec, journeys and page waits slept a fixed time, and screenshots were 11.5% of re-read tokens. `exec.sh` now skips the pre-flight baseline mxbuild when the model is byte-identical to its last *clean* pass: the verification stamp gains an `errors:` field (`model-stamp.sh check --clean` accepts only `errors: 0`, so a delta-gate pass over pre-existing errors never counts; `SKIP_BASELINE=0` is unchanged). New `tests/e2e/settle.js` waits for the page's own state (no visible Mendix progress/loading indicator, DOM size stable) instead of a fixed time; `journey-runner.js` uses it after each action (an explicit `settleMs` still wins) and `page-audit.js` per page. Probed in a real browser: on a page that loads in 1.8 s the old 1200 ms sleep and networkidle+600 both captured it half-loaded, settle waited for it; on a fast page settle took 775 ms vs 1200. `SETTLE_MODE=fixed` restores the old waits for an A/B run. `page-audit.js` also saves a viewport shot `page-audit-.top.png` beside the full-page one, and `ui-loop.md` says to look at viewport shots: a tall full-page shot reaches the model at ~1024 px high, text unreadable. `shrink-image-read.sh` now works off macOS (ImageMagick or Pillow, python when jq is absent; it was a silent no-op without `sips`). — MendixMau - perf(ui-loop): **a stale app reloads instead of restarting, journeys stop sleeping, worst pages are looked at first.** Field run 2026-10-07, existing app. `test-stack-up.sh` records which model the running app was built from (`.claude/loop/served-model`); when the model has changed since, it hot-reloads with `mxcli docker reload` (adding `--skip-check` when the exec.sh gate already passed this model) and restarts only if that fails: 93 s instead of 179 s. `ui-loop.md` and the script header carry the measured table and point at `mxcli run --local --watch` as the fast loop (~18 s per page change). `journey-runner.js` waits on page state after a combobox action instead of fixed pauses, and a combobox step now opens the widget it names: it used to open the page's first combobox, so on a form with three every step but the first failed. `page-audit.js` clicks the first copy of a nav item a user could actually click, not just the first match: with the sidebar collapsed, every in-nav page faulted while the same item sat in the top menu bar. It also writes `page-audit-look-order.txt`, all pages worst first (`module-review.md` §4a: order, not selection, every page is still looked at). `design-audit.js` drops its fixed 1.2 s wait. Same 7-step journey on the same app: 56 s with `SETTLE_MODE=fixed`, 24 s with settle, 11/11 pass both ways, 3 runs each. — MendixMau - feat(permissions): **intake now picks how much Claude Code may run without a prompt, and `doctor.sh` shows it.** Until now the allow-list covered only the safe wrappers (`bin/exec.sh`, `bin/*.sh`, `./mxcli`, `mx`), so every edit, test run, `npm install` or `git commit` in a project still prompted. New `bin/permission-profile.sh` resolves a profile (env, `.mxtk/permission-profile`, PROJECT.md `Permission profile:`, default `project`; a typo falls back to `wrappers`): `wrappers` is the old list; `project` adds edits anywhere inside the project folder (`Edit(/**)`) and the everyday test/npm/docker/git-add/commit commands, with force-push, `git reset --hard`, `git clean` and `sudo` behind `ask` rules that prompt in every mode; `full` adds bare `Bash`/`Edit`/`Read`/`WebFetch`/`WebSearch` to the per-machine `settings.local.json` only, never the shared file. `install-claude-permissions.sh` / `install-harness-permissions.sh` take `--profile`, and switching to a smaller profile removes only what the script itself added. Intake gains Q12, `init-project.sh` writes `Permission profile: project`, and `doctor.sh ` reports the profile, whether the settings files carry it, and warns on a `defaultMode` of auto/bypass in a project file, which Claude Code ignores there. No profile can switch on bypass or auto mode from the project: doctor says how to do that from the user's own settings. `test-install-claude-permissions.sh` gains T5b. — MendixMau diff --git a/project-bin/test-stack-up.sh b/project-bin/test-stack-up.sh index 09abf59f..01f467ac 100755 --- a/project-bin/test-stack-up.sh +++ b/project-bin/test-stack-up.sh @@ -48,6 +48,16 @@ # proof-of-ownership step. `docker reload --css` is NOT a theme loop: it copies the theme without # compiling SCSS, so a main.scss edit does not show. # +# RUN --LOCAL IS RECOGNISED, NOT REPLACED. `mxcli run --local` (v0.24+) publishes +# /.mxcli/run-local.json while it serves — pid and app port, written when the app is ready, +# removed on exit. A live pid in OUR project's directory plus a Mendix answer on its port is the same +# kind of claim as the compose label: it cannot be true of two projects at once. So that app is +# ownership `verified` (APP_SOURCE=run-local), where it used to scan as `unverified` and the e2e +# config refused it — the recommended fast loop was the one the harness would not test. This script +# never restarts or reloads it (it did not start it): `--watch` applies each change itself, and +# without --watch a model newer than the handshake is reported stale with the restart to run. Only +# pid and appPort are read; the file also holds the admin password and boot config, never copied. +# # Exit codes: 0 = required stack is up · 1 = not up and could not fix · 2 = usage/env error # # PROJECT CONFIGURATION — no project name, port or service path is hardcoded here. Put anything @@ -176,7 +186,7 @@ EOF return 1 } -# find_app_port emits "|" on stdout, NOT a bare port. +# find_app_port emits "||" on stdout, NOT a bare port. # # It is called in a command substitution, so it runs in a subshell and cannot set a variable in the # parent — the ownership verdict has to travel on stdout with the port or it is lost. @@ -189,11 +199,13 @@ EOF # someone else's software. find_app_port() { local p owned + # 0. A live `mxcli run --local` of this project (its handshake names the port). + if owned="$(local_loop_port)" && mendix_at "$owned"; then echo "$owned|verified|run-local"; return 0; fi # 1. Ownership: which port does THIS project's container publish? if owned="$(owned_app_port)" && [ -n "$owned" ]; then - if mendix_at "$owned"; then echo "$owned|verified"; return 0; fi + if mendix_at "$owned"; then echo "$owned|verified|docker"; return 0; fi warn "our container publishes :$owned but it is not serving Mendix yet" >&2 - echo "$owned|verified"; return 1 + echo "$owned|verified|docker"; return 1 fi # 2. No container of ours is running (SP-hosted run, or docker unavailable). Fall back to the # scan — but say so, because a hit here is UNVERIFIED ownership: it proves a Mendix answered, @@ -203,10 +215,62 @@ find_app_port() { warn "no container owned by $ROOT/.docker is running; :$p matched by scan — ownership UNVERIFIED" >&2 warn " → recorded as APP_OWNERSHIP=unverified in stack.env. A test harness should REFUSE an" >&2 warn " unverified port unless ALLOW_UNVERIFIED_APP=1, because :$p may be another project." >&2 - echo "$p|unverified"; return 0 + echo "$p|unverified|scan"; return 0 fi done - echo "|none"; return 1 + echo "|none|none"; return 1 +} + +# A live `mxcli run --local` serving THIS project: echoes its app port. See RUN --LOCAL above. +# Fields are read at MarshalIndent's two-space top level, so a key inside bootConfig cannot match. +MPR_DIR="$(cd "$(dirname "$MPR")" && pwd)" +# A two-tree checkout may carry a root symlink to app/X.mpr. mprcontents/ sits beside the real +# file, and mxcli writes the handshake beside whichever path -p was given — so look in both. +MODEL_DIR="$MPR_DIR" +if [ -L "$MPR" ]; then + _t="$(readlink "$MPR")" + case "$_t" in /*) ;; *) _t="$MPR_DIR/$_t" ;; esac + MODEL_DIR="$(cd "$(dirname "$_t")" 2>/dev/null && pwd -P)" || MODEL_DIR="$MPR_DIR" +fi +LOCAL_HS="$MPR_DIR/.mxcli/run-local.json" +pick_local_hs() { + local d + for d in "$MPR_DIR" "$MODEL_DIR"; do + [ -f "$d/.mxcli/run-local.json" ] && { LOCAL_HS="$d/.mxcli/run-local.json"; return 0; } + done + return 1 +} +pick_local_hs || true +hs_field() { sed -n "s/^ \"$1\": *\([0-9][0-9]*\).*/\1/p" "$LOCAL_HS" 2>/dev/null | head -1; } +pid_alive() { + [ -n "$1" ] || return 1 + case "$(mxtk_platform)" in + # mxcli.exe records a Windows pid; Git Bash's kill only knows MSYS pids. + windows) tasklist //FI "PID eq $1" //NH 2>/dev/null | grep -q " $1 " ;; + *) kill -0 "$1" 2>/dev/null ;; + esac +} +local_loop_port() { + local pid port + pick_local_hs || return 1 + pid="$(hs_field pid)"; port="$(hs_field appPort)" + [ -n "$port" ] && pid_alive "$pid" || return 1 + echo "$port" +} +# yes | no | unknown — was the loop started with --watch? Read from its command line where ps can +# show one; Windows processes do not expose it to Git Bash. +local_loop_watch() { + local pid args + pid="$(hs_field pid)" + args="$(ps -o args= -p "$pid" 2>/dev/null)" || { echo unknown; return; } + [ -n "$args" ] || { echo unknown; return; } + case "$args" in *--watch*) echo yes ;; *) echo no ;; esac +} +# Did the model change since the loop booted? The handshake is written when the app is ready, so +# its mtime is the boot's model; any .mpr or mprcontents file newer than it is a later edit. +local_model_changed() { + [ -n "$(find "$MODEL_DIR" -maxdepth 1 -name '*.mpr' -newer "$LOCAL_HS" 2>/dev/null | head -1)" ] && return 0 + [ -d "$MODEL_DIR/mprcontents" ] && [ -n "$(find "$MODEL_DIR/mprcontents" -type f -newer "$LOCAL_HS" 2>/dev/null | head -1)" ] } # Single source of truth for the discovered ports. This script is the only thing that PROVES which @@ -215,7 +279,7 @@ find_app_port() { # 8081 and a skill said 8080 — three "truths", and a spec pointed at a dead port reports the app as # broken rather than as unreachable. STACK_ENV="${STACK_ENV:-$ROOT/.claude/loop/stack.env}" -# publish_stack_env +# publish_stack_env # APP_OWNERSHIP is not decoration: it is the difference between "this port is ours" and "a Mendix # answered here". Consumers must be able to tell those apart from the file alone, because the stderr # warning that distinguishes them is gone by the time anyone reads it. @@ -225,6 +289,7 @@ publish_stack_env() { echo "# written by bin/test-stack-up.sh — $(date -u +%Y-%m-%dT%H:%M:%SZ). Do not edit." echo "APP_PORT=$1" echo "APP_OWNERSHIP=${3:-unknown}" + echo "APP_SOURCE=${4:-unknown}" [ -n "$MOCK_PORT" ] && echo "MOCK_PORT=$MOCK_PORT" [ "$2" -eq 0 ] && echo "JAEGER_PORT=$JAEGER_PORT" || echo "JAEGER_PORT=" } > "$STACK_ENV" @@ -246,6 +311,13 @@ record_served() { # Echoes fresh | stale | unknown. served_state() { local want have + # run --local: mxcli rebuilt nothing for us to record; the handshake's mtime is the boot. + if [ "${APP_SOURCE:-}" = "run-local" ]; then + pick_local_hs || true + if ! local_model_changed; then echo fresh + else case "$(local_loop_watch)" in yes) echo watch ;; no) echo stale ;; *) echo changed ;; esac; fi + return + fi have="$(head -1 "$SERVED_FILE" 2>/dev/null)" want="$(model_fp)" || want="" if [ -z "$have" ] || [ -z "$want" ]; then echo unknown @@ -253,6 +325,19 @@ served_state() { else echo stale; fi } +# After `docker run --wait` returned: prove over HTTP that OUR app answers (design rule 3). mxcli's +# wait reads the runtime log, so this normally passes on the first probe. Sets APP_* in this shell. +prove_app() { + local waited=0 + while :; do + APP_FIND="$(find_app_port)" && break + [ $waited -ge 30 ] && break + sleep 3; waited=$((waited + 3)) + done + IFS='|' read -r APP_PORT APP_OWNERSHIP APP_SOURCE <<<"$APP_FIND" + mendix_at "$APP_PORT" +} + echo "── Test stack: $PROJ ──────────────────────────────────" # --- Studio Pro: report only, never act ------------------------------------- @@ -277,10 +362,11 @@ fi # --- App -------------------------------------------------------------------- APP_FIND="$(find_app_port)" APP_FOUND=$? -APP_PORT="${APP_FIND%%|*}" -APP_OWNERSHIP="${APP_FIND##*|}" +IFS='|' read -r APP_PORT APP_OWNERSHIP APP_SOURCE <<<"$APP_FIND" if [ $APP_FOUND -eq 0 ]; then - if [ "$APP_OWNERSHIP" = "verified" ]; then + if [ "$APP_SOURCE" = "run-local" ]; then + ok "App serving on :$APP_PORT (ownership verified — this project's mxcli run --local)" + elif [ "$APP_OWNERSHIP" = "verified" ]; then ok "App serving on :$APP_PORT (ownership verified — our container)" else warn "App serving on :$APP_PORT — ownership $APP_OWNERSHIP, this may be another project" @@ -293,7 +379,10 @@ if [ $APP_FOUND -eq 0 ]; then SERVED="$(served_state)" case "$SERVED" in fresh) ok "App serves the current model" ;; - stale) warn "App serves an OLDER model — the model changed since it was built; tests would measure the old build" ;; + stale) warn "App serves an OLDER model — the model changed since it was built; tests would measure the old build" + [ "$APP_SOURCE" = "run-local" ] && warn " → run --local was started without --watch: stop it and start ./mxcli run --local --watch -p $(basename "$MPR")" ;; + watch) ok "Model changed since boot — run --local --watch applies it (~20s a page change, ~60s security/navigation)" ;; + changed) warn "Model changed since run --local booted; cannot see whether it runs with --watch. Without --watch it serves the old model" ;; unknown) warn "App freshness unknown — not built by this script (Studio Pro / manual run); it may serve an old model" ;; esac fi @@ -346,7 +435,7 @@ fi # Publish as soon as the port is MEASURED, not only when the whole stack is READY. A down mock makes # REST-fed specs invalid; it does not change which port the app answers on, and the non-REST specs # still need to know it. -[ $APP_FOUND -eq 0 ] && publish_stack_env "$APP_PORT" "$JAEGER_OK" "$APP_OWNERSHIP" +[ $APP_FOUND -eq 0 ] && publish_stack_env "$APP_PORT" "$JAEGER_OK" "$APP_OWNERSHIP" "$APP_SOURCE" if [ "$MODE" = "check" ]; then echo "────────────────────────────────────────────────────────" @@ -354,6 +443,10 @@ if [ "$MODE" = "check" ]; then echo "APP UP, BUT REST-FED SPECS WILL BE INVALID — see the container warning above" exit 1 fi + if [ "$SERVED" = "stale" ] && [ "$APP_SOURCE" = "run-local" ]; then + echo "APP UP BUT STALE — restart mxcli run --local with --watch (this script does not restart a loop it did not start)" + exit 1 + fi if [ "$SERVED" = "stale" ] && [ "$APP_OWNERSHIP" = "verified" ]; then echo "APP UP BUT STALE — re-run without --check to hot-reload the current model" exit 1 @@ -383,9 +476,10 @@ if [ $MOCK_REQUIRED -eq 1 ] && [ $MOCK_OK -ne 0 ]; then fi # --- Warm loop: refresh a running app instead of restarting it -------------- -# Only for OUR container (ownership verified): a reload swaps the model in a runtime, and doing that -# to an unverified port could be another project's app or a Studio Pro run. -if [ $APP_FOUND -eq 0 ] && [ "$APP_OWNERSHIP" = "verified" ] && [ -n "$MXCLI" ] \ +# Only for OUR container (ownership verified, source docker): a reload swaps the model in a runtime, +# and doing that to an unverified port could be another project's app or a Studio Pro run. A +# run --local loop is never reloaded from here — --watch does that, inside the process that owns it. +if [ $APP_FOUND -eq 0 ] && [ "$APP_SOURCE" = "docker" ] && [ -n "$MXCLI" ] \ && [ "${MXTK_WARM_RELOAD:-on}" != "off" ] && [ "$MODE" != "nodocker" ] \ && { [ "$SERVED" != "fresh" ] || [ "$MODE" = "restart" ]; }; then T0=$(date +%s) @@ -419,18 +513,25 @@ if [ $APP_FOUND -eq 0 ] && [ "$APP_OWNERSHIP" = "verified" ] && [ -n "$MXCLI" ] fi if [ $RELOADED -ne 0 ]; then echo "→ Rebuilding and restarting the container..." - "$MXCLI" docker run -p "$MPR" --wait >"$DOCKER_LOG" 2>&1 + "$MXCLI" docker run -p "$MPR" --wait --wait-timeout "$BOOT_TIMEOUT" >"$DOCKER_LOG" 2>&1 if [ $? -ne 0 ]; then bad "mxcli docker run failed — see $DOCKER_LOG"; tail -20 "$DOCKER_LOG"; exit 1 fi - APP_FOUND=1 # re-proved below by the boot-wait loop + # Proved here, not by falling into "Bring up the app" below: that block runs `docker run` + # again, and did — every restart built and booted the app twice. + if prove_app; then + record_served; SERVED=fresh + ok "Restarted in $(( $(date +%s) - T0 ))s — app serves the current model" + else + bad "mxcli reported the runtime started, but no Mendix answers over HTTP — see $DOCKER_LOG"; exit 1 + fi fi fi # --- Bring up the app ------------------------------------------------------- if [ $APP_FOUND -ne 0 ]; then if [ "$MODE" = "nodocker" ]; then - bad "App down and --no-docker given. Click Run Locally in Studio Pro (or, with no Studio Pro, ./mxcli run --local), then re-run --check." + bad "App down and --no-docker given. Click Run Locally in Studio Pro (or, with no Studio Pro, ./mxcli run --local --watch), then re-run --check." exit 1 fi # No reachable Docker (or Podman) daemon: say so and name the Docker-free route, instead of letting @@ -450,7 +551,8 @@ if [ $APP_FOUND -ne 0 ]; then bad "App down and no Docker/Podman reachable, so this script cannot build the app container." echo " Normal in a cloud container, and fine on a desktop without one. Run the app without it:" echo " Studio Pro: Run Locally. No Studio Pro:" - echo " ./mxcli run --local -p $(basename "$MPR") # flags: skills/cloud-dev-environment.md" + echo " ./mxcli run --local --watch -p $(basename "$MPR") # flags: skills/cloud-dev-environment.md" + echo " This script then sees it as yours (verified) and --watch keeps it current." echo " Snapshot first — mxcli run --local consolidates a split-model .mpr. Then re-run with --check." exit 1 fi @@ -463,9 +565,9 @@ if [ $APP_FOUND -ne 0 ]; then echo " This can take several minutes on a cold build." # Do NOT pipe mxcli: reading $? through a pipe measures the pipe, not the command. - # (this script runs without `set -e` on purpose — the boot-wait loop below relies on - # find_app_port returning non-zero repeatedly without killing the script) - "$MXCLI" docker run -p "$MPR" >"$DOCKER_LOG" 2>&1 + # --wait: mxcli follows the runtime log until it reports started, so the boot wait is mxcli's. + # The loop below is the HTTP proof (design rule 3) and normally passes on its first probe. + "$MXCLI" docker run -p "$MPR" --wait --wait-timeout "$BOOT_TIMEOUT" >"$DOCKER_LOG" 2>&1 DOCKER_RC=$? if [ $DOCKER_RC -ne 0 ]; then bad "mxcli docker run failed (rc=$DOCKER_RC) — see $DOCKER_LOG" @@ -473,33 +575,28 @@ if [ $APP_FOUND -ne 0 ]; then exit 1 fi - echo "→ Waiting for the app to answer (timeout ${BOOT_TIMEOUT}s)..." - WAITED=0 - while [ $WAITED -lt "$BOOT_TIMEOUT" ]; do - APP_FIND="$(find_app_port)" && { APP_FOUND=0; break; } - sleep 3; WAITED=$((WAITED + 3)) - [ $((WAITED % 30)) -eq 0 ] && echo " ...${WAITED}s" - done - APP_PORT="${APP_FIND%%|*}" - APP_OWNERSHIP="${APP_FIND##*|}" - - if [ $APP_FOUND -eq 0 ]; then + if prove_app; then + APP_FOUND=0 record_served - ok "App serving on :$APP_PORT after ${WAITED}s (ownership $APP_OWNERSHIP)" + ok "App serving on :$APP_PORT (ownership $APP_OWNERSHIP)" else - bad "App still not answering after ${BOOT_TIMEOUT}s — see $DOCKER_LOG" + bad "mxcli reported the runtime started, but no Mendix answers over HTTP — see $DOCKER_LOG" "$MXCLI" docker status -p "$MPR" 2>&1 | tail -5 exit 1 fi fi echo "────────────────────────────────────────────────────────" +if [ "$APP_SOURCE" = "run-local" ] && [ "$SERVED" = "stale" ]; then + echo "NOT READY — the run --local app serves an older model; restart it with --watch" + exit 1 +fi if [ $APP_FOUND -eq 0 ] && [ $MOCK_OK -eq 0 ]; then # re-publish: both the port AND the ownership can change across a Docker boot — before the boot no # container of ours was running (so any hit was an unverified scan); after it, one is. - publish_stack_env "$APP_PORT" "$JAEGER_OK" "$APP_OWNERSHIP" + publish_stack_env "$APP_PORT" "$JAEGER_OK" "$APP_OWNERSHIP" "$APP_SOURCE" echo "READY — ports published to ${STACK_ENV#$ROOT/}; the e2e config reads them automatically" - echo " app :$APP_PORT (ownership $APP_OWNERSHIP)" + echo " app :$APP_PORT (ownership $APP_OWNERSHIP, $APP_SOURCE)" [ -n "$MOCK_PORT" ] && echo " mock :$MOCK_PORT" echo " jaeger $([ $JAEGER_OK -eq 0 ] && echo ":$JAEGER_PORT" || echo 'DOWN — Trace assertions will not run')" exit 0 diff --git a/project-tests/e2e/config.js b/project-tests/e2e/config.js index 8b626a57..23c54cb1 100644 --- a/project-tests/e2e/config.js +++ b/project-tests/e2e/config.js @@ -34,7 +34,8 @@ const ROOT = PROJ.root; // // test-stack-up.sh now records APP_OWNERSHIP: `verified` means the port is published // by a container whose compose working_dir is THIS project (the one identity claim -// that cannot be true of two projects at once); `unverified` means a port scan found +// that cannot be true of two projects at once), or by a live `mxcli run --local` whose +// .mxcli/run-local.json sits beside THIS project's .mpr; `unverified` means a port scan found // a Mendix and nothing more. We refuse `unverified` rather than warn about it — // this whole harness exists because a warning on stderr is not a guard. function resolveStack() { diff --git a/skills/harness-architecture.md b/skills/harness-architecture.md index 08969e06..59925ca0 100644 --- a/skills/harness-architecture.md +++ b/skills/harness-architecture.md @@ -34,11 +34,13 @@ APP_PORT= APP_OWNERSHIP=verified | unverified | asserted | unknown ``` -`verified` means a container owned by this project's `.docker` is serving that port. `unverified` +`verified` means a container owned by this project's `.docker` is serving that port, or a live +`mxcli run --local` of this project is (its `.mxcli/run-local.json` names the port; `APP_SOURCE` +says which: `docker` | `run-local` | `scan`). `unverified` means a Mendix answered a port from the `APP_PORTS` fallback scan list — *a* Mendix, not necessarily yours. **A harness must refuse an unverified port** unless the operator sets `ALLOW_UNVERIFIED_APP=1`, and must record the ownership it acted on. **VERIFIED** -(`project-bin/test-stack-up.sh:150-198`). +(`project-bin/test-stack-up.sh`, `owned_app_port` / `local_loop_port` / `find_app_port`). This is not defensive decoration. On one measured occasion a published `stack.env` named a port that belonged to a *different* project's Mendix, which answered 200 with a real login page. Every diff --git a/skills/report-schema.md b/skills/report-schema.md index f8054752..b564068e 100644 --- a/skills/report-schema.md +++ b/skills/report-schema.md @@ -492,7 +492,7 @@ reported normally. | Value | Means | Basis | |---|---|---| -| `verified` | a container published this port **and** its compose `working_dir` label is this project — the one identity claim that cannot be true of two projects at once | `APP_OWNERSHIP=verified` in `.claude/loop/stack.env`, written by `bin/test-stack-up.sh` | +| `verified` | a container published this port **and** its compose `working_dir` label is this project — the one identity claim that cannot be true of two projects at once; or a live `mxcli run --local` whose `.mxcli/run-local.json`, beside this project's `.mpr`, names the port (`APP_SOURCE=run-local`) | `APP_OWNERSHIP=verified` in `.claude/loop/stack.env`, written by `bin/test-stack-up.sh` | | `asserted` | an operator said so and nothing checked | an `APP_PORT` env var, or the project's deployment config | | `unknown` | a Mendix answered and nothing more | a guess, or a `stack.env` predating the marker | diff --git a/skills/ui-loop.md b/skills/ui-loop.md index 78b5e64d..2efc2b43 100644 --- a/skills/ui-loop.md +++ b/skills/ui-loop.md @@ -183,6 +183,12 @@ So keep `run --local --watch` running beside the session; `test-stack-up.sh` not Docker app and reloads it, but that path is five times slower. `docker reload --css` does not compile SCSS: a `main.scss` edit needs the build. +`test-stack-up.sh` recognises the watch loop as this project's app: `mxcli run --local` (v0.24+) +writes `.mxcli/run-local.json` beside the `.mpr` while it serves, and a live pid there plus a +Mendix answer on its port publishes `APP_OWNERSHIP=verified`, `APP_SOURCE=run-local`, so the e2e +config tests it without `ALLOW_UNVERIFIED_APP=1`. It never reloads or restarts that loop. Started +without `--watch`, a model edited since boot is reported stale with the restart to run. + ## When it finds something — which side is actually wrong The symptom is easy to see. **Where the fix goes is not, and the instinct is usually wrong.** diff --git a/tests/wave2/fixtures/run-local/run-local.json b/tests/wave2/fixtures/run-local/run-local.json new file mode 100644 index 00000000..bd343462 --- /dev/null +++ b/tests/wave2/fixtures/run-local/run-local.json @@ -0,0 +1,22 @@ +{ + "project": "/work/Fixture/Fixture.mpr", + "pid": 4242, + "appPort": 8180, + "adminPort": 8190, + "adminPass": "REDACTED", + "bootConfig": { + "BasePath": "/work/Fixture/deployment", + "DTAPMode": "D", + "DatabaseHost": "", + "DatabaseName": "default", + "DatabasePassword": "REDACTED", + "DatabaseType": "HSQLDB", + "DatabaseUserName": "sa", + "MicroflowConstants": { + "MyFirstModule.LogNode": "REDACTED", + "MyFirstModule.ServiceUrl": "REDACTED" + }, + "RuntimePath": "/opt/mendix/runtime" + }, + "started": "2026-10-07T12:27:57Z" +} \ No newline at end of file diff --git a/tests/wave2/test-stack-up-local.sh b/tests/wave2/test-stack-up-local.sh new file mode 100755 index 00000000..8813e3d9 --- /dev/null +++ b/tests/wave2/test-stack-up-local.sh @@ -0,0 +1,147 @@ +#!/usr/bin/env bash +# Fixtures for test-stack-up.sh recognising a live `mxcli run --local` as THIS project's app. +# +# mxcli (v0.24+) writes /.mxcli/run-local.json when the app is ready and removes it on +# exit. Before this change the script never read it: a run --local app was found by port scan, +# published as APP_OWNERSHIP=unverified, and the e2e harness refused it unless the operator set +# ALLOW_UNVERIFIED_APP=1. Field run 2026-10-07, existing app: same app, same port, now `verified`. +# +# The handshake is fixtures/run-local/run-local.json — a real capture from that field run, every +# value replaced (project path, pid, ports, passwords, constants), byte layout kept. Each case +# rewrites only pid/appPort/adminPass. The "app" is a python http.server serving a login page with +# a Mendix marker; the "mxcli" is `sleep` with its argv[0] set, so `ps` shows the command line the +# script reads --watch from. docker/podman are stubbed out so no real container can answer. +# Not fixtured: the Windows tasklist branch of pid_alive (CI is Linux). +# +# usage: test-stack-up-local.sh /path/to/test-stack-up.sh +set -uo pipefail + +SRC="${1:?usage: test-stack-up-local.sh /path/to/test-stack-up.sh}" +case "$SRC" in /*) ;; *) SRC="$PWD/$SRC" ;; esac +GOLDEN="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fixtures/run-local/run-local.json" +[ -f "$SRC" ] && [ -f "$(dirname "$SRC")/_common.sh" ] && [ -f "$GOLDEN" ] \ + || { echo "FIXTURE ERROR: need $SRC, its _common.sh, and $GOLDEN"; exit 2; } +command -v python3 >/dev/null 2>&1 && command -v curl >/dev/null 2>&1 \ + || { echo "FIXTURE ERROR: needs python3 and curl"; exit 2; } + +WORK="$(mktemp -d "${TMPDIR:-/tmp}/stackup-local.XXXXXX")" +PIDS="" +cleanup() { for p in $PIDS; do kill "$p" 2>/dev/null; done; rm -rf "$WORK"; } +trap cleanup EXIT +PASS=0; FAIL=0 +ok() { PASS=$((PASS+1)); printf ' ok %s\n' "$1"; } +bad() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; [ -n "${2:-}" ] && printf '%s\n' "$2" | sed 's/^/ FAIL-detail: /'; } + +# --- no container runtime: only our fake app can answer ---------------------------------- +mkdir -p "$WORK/stubs" +for t in docker podman; do printf '#!/bin/sh\nexit 1\n' > "$WORK/stubs/$t"; chmod +x "$WORK/stubs/$t"; done +export PATH="$WORK/stubs:$PATH" +unset PROJECT_ROOT MPR_FILE STACK_ENV STACK_CONF SERVED_FILE + +# --- the fake app ------------------------------------------------------------------------- +mkdir -p "$WORK/www" +printf '\n' > "$WORK/www/login.html" +python3 -I - "$WORK/www" "$WORK/port" >/dev/null 2>&1 <<'EOF' & +import functools, http.server, sys +H = functools.partial(http.server.SimpleHTTPRequestHandler, directory=sys.argv[1]) +s = http.server.ThreadingHTTPServer(("127.0.0.1", 0), H) +open(sys.argv[2], "w").write(str(s.server_address[1])) +s.serve_forever() +EOF +PIDS="$PIDS $!" +for _ in $(seq 1 50); do [ -s "$WORK/port" ] && break; sleep 0.1; done +APP="$(cat "$WORK/port" 2>/dev/null)" +[ -n "$APP" ] && curl -s --max-time 4 "http://localhost:$APP/login.html" | grep -q mxui \ + || { echo "FIXTURE ERROR: fake app did not come up"; exit 2; } + +# A process whose command line reads like mxcli's: fake_mxcli . +fake_mxcli() { + bash -c "exec -a 'mxcli $2' sleep 600" >/dev/null 2>&1 & + PIDS="$PIDS $!"; printf -v "$1" '%s' "$!" +} +fake_mxcli WATCH_PID 'run --local --watch -p Fixture.mpr' +fake_mxcli PLAIN_PID 'run --local -p Fixture.mpr' +sleep 0.3 # let exec -a replace bash, so ps shows the mxcli command line +sleep 0 & DEAD_PID=$!; wait "$DEAD_PID" 2>/dev/null + +ADMIN_MARK="fixture-admin-marker-7c1" # stands in for the adminPass value +OLD=202001010000 + +# mkproj -> echoes the project root +mkproj() { + local p="$WORK/$1" m + mkdir -p "$p/bin" "$p/.claude/loop" + cp "$SRC" "$p/bin/test-stack-up.sh"; cp "$(dirname "$SRC")/_common.sh" "$p/bin/" + # Nothing scanned or traced may answer: the run --local handshake is the only way in. + printf 'APP_PORTS="1"\nJAEGER_PORT="1"\n' > "$p/.claude/loop/stack.conf" + case "$2" in + root) m="$p" ;; + app) m="$p/app" ;; + symlink) m="$p/app"; ln -s app/Fixture.mpr "$p/Fixture.mpr" ;; + esac + mkdir -p "$m/mprcontents"; printf 'mpr' > "$m/Fixture.mpr"; printf 'u' > "$m/mprcontents/a.mxunit" + touch -t "$OLD" "$m/Fixture.mpr" "$m/mprcontents/a.mxunit" + echo "$p" +} +# handshake -> the golden, pointed at our fake app, written into /.mxcli/ +handshake() { + mkdir -p "$1/.mxcli" + sed -e "s/^ \"pid\": [0-9]*/ \"pid\": $2/" \ + -e "s/^ \"appPort\": [0-9]*/ \"appPort\": $APP/" \ + -e "s/^ \"adminPass\": \"[^\"]*\"/ \"adminPass\": \"$ADMIN_MARK\"/" "$GOLDEN" > "$1/.mxcli/run-local.json" +} +run() { (cd "$1" && bash bin/test-stack-up.sh --check 2>&1); } +envv() { sed -n "s/^$2=//p" "$1/.claude/loop/stack.env" 2>/dev/null; } + +echo "== T0: the golden is what the script's reader expects ==" +grep -q '^ "pid": 4242,$' "$GOLDEN" && grep -q '^ "appPort": 8180,$' "$GOLDEN" \ + && ok "golden carries pid/appPort at MarshalIndent's two-space top level" || bad "golden shape changed" +grep -q "$ADMIN_MARK" "$GOLDEN" && bad "golden carries the test secret" || ok "golden carries no live secret" + +echo "== T1: live --watch loop, model unchanged -> verified, run-local, fresh ==" +P="$(mkproj t1 root)"; handshake "$P" "$WATCH_PID" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 0 ] && ok "--check exits 0" || bad "--check rc=$RC" "$OUT" +printf '%s' "$OUT" | grep -q "ownership verified — this project's mxcli run --local" && ok "reported as this project's run --local" || bad "not reported as run --local" "$OUT" +printf '%s' "$OUT" | grep -q "App serves the current model" && ok "fresh" || bad "not fresh" "$OUT" +[ "$(envv "$P" APP_PORT)" = "$APP" ] && [ "$(envv "$P" APP_OWNERSHIP)" = verified ] && [ "$(envv "$P" APP_SOURCE)" = run-local ] \ + && ok "stack.env: APP_PORT=$APP APP_OWNERSHIP=verified APP_SOURCE=run-local" || bad "stack.env wrong" "$(cat "$P/.claude/loop/stack.env" 2>&1)" +grep -rq "$ADMIN_MARK" "$P/.claude" && bad "adminPass leaked into .claude/" || ok "adminPass not copied anywhere under .claude/" +printf '%s' "$OUT" | grep -q "$ADMIN_MARK" && bad "adminPass printed" || ok "adminPass not printed" + +echo "== T2: dead pid -> the handshake is ignored ==" +P="$(mkproj t2 root)"; handshake "$P" "$DEAD_PID" +OUT="$(run "$P")"; RC=$? +[ "$RC" -ne 0 ] && printf '%s' "$OUT" | grep -q "App not serving" && ok "not trusted: app reported not serving" || bad "dead-pid handshake trusted (rc=$RC)" "$OUT" +[ "$(envv "$P" APP_SOURCE)" = run-local ] && bad "stack.env says run-local for a dead loop" || ok "stack.env not claimed for a dead loop" + +echo "== T3: model edited, loop without --watch -> stale, exit 1 ==" +P="$(mkproj t3 root)"; handshake "$P" "$PLAIN_PID"; touch "$P/Fixture.mpr" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 1 ] && ok "--check exits 1" || bad "--check rc=$RC" "$OUT" +printf '%s' "$OUT" | grep -q "APP UP BUT STALE — restart mxcli run --local with --watch" && ok "says restart with --watch" || bad "no stale verdict" "$OUT" + +echo "== T4: model edited, loop WITH --watch -> watch applies it, ready ==" +P="$(mkproj t4 root)"; handshake "$P" "$WATCH_PID"; touch "$P/mprcontents/a.mxunit" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 0 ] && printf '%s' "$OUT" | grep -q "run --local --watch applies it" && ok "watch state, exit 0" || bad "watch not recognised (rc=$RC)" "$OUT" + +echo "== T5: two-tree (app/), handshake beside app/Fixture.mpr ==" +P="$(mkproj t5 app)"; handshake "$P/app" "$PLAIN_PID" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 0 ] && [ "$(envv "$P" APP_SOURCE)" = run-local ] && ok "found under app/" || bad "two-tree handshake missed (rc=$RC)" "$OUT" +touch "$P/app/mprcontents/a.mxunit" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 1 ] && printf '%s' "$OUT" | grep -q "APP UP BUT STALE" && ok "an app/mprcontents edit reads stale" || bad "app/ edit not seen (rc=$RC)" "$OUT" + +echo "== T6: two-tree with a root symlink, run as -p Fixture.mpr from the root ==" +P="$(mkproj t6 symlink)"; handshake "$P" "$PLAIN_PID" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 0 ] && [ "$(envv "$P" APP_SOURCE)" = run-local ] && ok "found beside the symlink" || bad "symlinked handshake missed (rc=$RC)" "$OUT" +touch "$P/app/mprcontents/a.mxunit" +OUT="$(run "$P")"; RC=$? +[ "$RC" -eq 1 ] && printf '%s' "$OUT" | grep -q "APP UP BUT STALE" && ok "edit behind the symlink reads stale" || bad "edit behind symlink not seen (rc=$RC)" "$OUT" + +echo +echo "stack-up-local: $PASS passed, $FAIL failed" +[ "$FAIL" -eq 0 ] From 94532e4dddd22fae12ade8eb5e28f3bd3033faf4 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 13:59:54 +0000 Subject: [PATCH 5/6] test(stack-up-local): resolve Python via portable.sh require_py check-portability flagged the hard-coded python3; use "$PY" like the other fixtures. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- tests/wave2/test-stack-up-local.sh | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/tests/wave2/test-stack-up-local.sh b/tests/wave2/test-stack-up-local.sh index 8813e3d9..d9c5b761 100755 --- a/tests/wave2/test-stack-up-local.sh +++ b/tests/wave2/test-stack-up-local.sh @@ -21,8 +21,9 @@ case "$SRC" in /*) ;; *) SRC="$PWD/$SRC" ;; esac GOLDEN="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fixtures/run-local/run-local.json" [ -f "$SRC" ] && [ -f "$(dirname "$SRC")/_common.sh" ] && [ -f "$GOLDEN" ] \ || { echo "FIXTURE ERROR: need $SRC, its _common.sh, and $GOLDEN"; exit 2; } -command -v python3 >/dev/null 2>&1 && command -v curl >/dev/null 2>&1 \ - || { echo "FIXTURE ERROR: needs python3 and curl"; exit 2; } +. "$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)/bin/lib/portable.sh" +require_py +command -v curl >/dev/null 2>&1 || { echo "FIXTURE ERROR: needs curl"; exit 2; } WORK="$(mktemp -d "${TMPDIR:-/tmp}/stackup-local.XXXXXX")" PIDS="" @@ -41,7 +42,7 @@ unset PROJECT_ROOT MPR_FILE STACK_ENV STACK_CONF SERVED_FILE # --- the fake app ------------------------------------------------------------------------- mkdir -p "$WORK/www" printf '\n' > "$WORK/www/login.html" -python3 -I - "$WORK/www" "$WORK/port" >/dev/null 2>&1 <<'EOF' & +"$PY" -I - "$WORK/www" "$WORK/port" >/dev/null 2>&1 <<'EOF' & import functools, http.server, sys H = functools.partial(http.server.SimpleHTTPRequestHandler, directory=sys.argv[1]) s = http.server.ThreadingHTTPServer(("127.0.0.1", 0), H) From ff83fc8ffc01df758e13f75eba5ff332263f7b5a Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 7 Oct 2026 14:10:08 +0000 Subject: [PATCH 6/6] test(stack-up-local): backdate the handshake so a same-tick edit still reads newer CI T3 failed: the handshake write and the model touch landed in one clock tick, so find -newer saw no edit (reproduced 35/50 locally). The handshake is now dated between the model and the case's edit. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw --- tests/wave2/test-stack-up-local.sh | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/tests/wave2/test-stack-up-local.sh b/tests/wave2/test-stack-up-local.sh index d9c5b761..369bbd72 100755 --- a/tests/wave2/test-stack-up-local.sh +++ b/tests/wave2/test-stack-up-local.sh @@ -67,6 +67,7 @@ sleep 0 & DEAD_PID=$!; wait "$DEAD_PID" 2>/dev/null ADMIN_MARK="fixture-admin-marker-7c1" # stands in for the adminPass value OLD=202001010000 +BOOT=202101010000 # the handshake: newer than the model, older than any edit a case makes # mkproj -> echoes the project root mkproj() { @@ -90,6 +91,9 @@ handshake() { sed -e "s/^ \"pid\": [0-9]*/ \"pid\": $2/" \ -e "s/^ \"appPort\": [0-9]*/ \"appPort\": $APP/" \ -e "s/^ \"adminPass\": \"[^\"]*\"/ \"adminPass\": \"$ADMIN_MARK\"/" "$GOLDEN" > "$1/.mxcli/run-local.json" + # Between the model (OLD) and now: a write and an edit in the same clock tick share an mtime, + # so an edit made right after this line would not read as newer (35/50 on a probe here). + touch -t "$BOOT" "$1/.mxcli/run-local.json" } run() { (cd "$1" && bash bin/test-stack-up.sh --check 2>&1); } envv() { sed -n "s/^$2=//p" "$1/.claude/loop/stack.env" 2>/dev/null; }