Skip to content

integration 2026-10-08: thirteen green PRs combined for one merge - #258

Merged
MendixMau merged 34 commits into
masterfrom
integration/2026-10-08
Oct 8, 2026
Merged

MendixMau merged 34 commits into
masterfrom
integration/2026-10-08

Conversation

@MendixMau

Copy link
Copy Markdown
Owner

What changed and why

One branch that merges thirteen open PRs on top of master 2fcdf5e, so they can be merged in one go after review. Each merge is a plain merge commit; the only conflicts were CHANGELOG.md (all lines kept) and the routing block in bin/gate-check.sh, which was re-rendered from bin/lib/skill-routing.tsv with bin/render-routing.sh.

Included, in merge order:

PR Branch What
#225 perf/stack-up-local-loop test-stack-up recognises mxcli run --local
#242 fix/no-devcontainer init: local by default, Dev Container parked
#243 docs/lint-backlog lint rollout plan on disk
#247 feat/lint-ledger per-run lint ledger + lint-trend.sh
#231 fix/sync-project-dup-heading-exit sync-project no longer exits early
#249 feat/e2e-locators-that-lie skill promotion
#250 feat/unhappy-path-testing skill promotion
#251 docs/testing-shape-local-deploy #230 docs fix
#252 fix/obligation-valid-at-picker #155
#253 docs/lint-backlog-batch4 Batch 4 rows + seven skill notes
#254 fix/conv020-modules-feedback #229 #245
#255 fix/journey-rung2-mustnotfire #150
#256 fix/gate-surface-realpath-journeydir #233 #234 #235

Not included: #238, #248 (lint rules, draft until laptop calibration), #257 (exec guards, draft until a field run), #223, #217–#219, #145.

Two ways to land this: squash-merge this PR and close the thirteen as merged-via-integration, or merge the thirteen individually and close this one. Either is fine; the tree is identical.

Field evidence

Not a new instrument. Checks run locally on the merged tree: tests/run-tests.sh 33 passed / 0 failed, tests/test-lint-delivery.sh ALL GREEN, render-routing --check, check-scripts, check-portability, leak guard and private-citation guard all clean. The wave2 fixture suite result is posted below as a comment when it finishes.

Checklist

  • Size cap: exempt (integration branch; each member PR is within cap)
  • Test tier: T1 (both suites on the merged tree)
  • Instrument rules: per member PR
  • CHANGELOG lines: 26 under Unreleased, all member lines kept
  • No conflict markers in the tree

🤖 Generated with Claude Code

https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw


Generated by Claude Code

claude and others added 30 commits October 7, 2026 10:13
…e, legible screenshots

- exec.sh skips the pre-flight baseline mxbuild when the model is unchanged
  since its last clean pass (stamp gains errors:, check --clean).
- New tests/e2e/settle.js replaces fixed sleeps in journey-runner and
  page-audit; SETTLE_MODE=fixed restores the old waits for an A/B run.
- page-audit saves a viewport shot beside the full-page one; ui-loop.md
  says to look at viewport shots.
- shrink-image-read.sh works off macOS (ImageMagick/Pillow, no-jq path).

Field run 2026-10-07, requirements-driven build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…order

- test-stack-up.sh: record the model the app was built from; on change,
  `mxcli docker reload` (+ --skip-check when exec.sh already passed it),
  restart only on failure. Header carries the measured loop table.
- journey-runner.js: combobox steps wait on page state; the toggle is
  scoped to the named widget (it opened the page's first combobox).
- page-audit.js: click the first nav copy a user can click (a collapsed
  sidebar covered .first()); write page-audit-look-order.txt, all pages
  worst first.
- design-audit.js: drop the fixed 1.2 s wait.
- ui-loop.md / module-review.md: warm-loop table; order, not selection.

Field run 2026-10-07, existing app: same 7-step journey 56 s fixed vs
24 s settle, 11/11 pass both, 3 runs each; reload --skip-check 93 s vs
restart 179 s; run --local --watch ~18 s per page change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
journey-runner.js now requires settle.js; the two fixtures that copy the
engine into a scratch project listed the files by hand and missed it, so
the require failed (MODULE_NOT_FOUND) before any assertion ran.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…ilds once

- A live `mxcli run --local` (its .mxcli/run-local.json beside this project's
  .mpr, pid alive) is now published as APP_OWNERSHIP=verified, APP_SOURCE=run-local
  instead of an unverified port-scan hit the e2e harness refuses.
- Staleness knows the loop: with --watch a model edit reads "watch" (ready),
  without it "APP UP BUT STALE - restart with --watch".
- Two-tree checkouts (app/X.mpr, with or without a root symlink) are resolved.
- docker path uses `docker run --wait` instead of the 240s poll; --restart no
  longer runs the build twice.
- Fixture tests/wave2/test-stack-up-local.sh over a golden captured from a real
  run and value-scrubbed (field run 2026-10-07, existing app).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
check-portability flagged the hard-coded python3; use "$PY" like the other fixtures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…l reads newer

CI T3 failed: the handshake write and the model touch landed in one clock tick, so
find -newer saw no edit (reproduced 35/50 locally). The handshake is now dated
between the model and the case's edit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
… refresh

The CLAUDE.md duplicate-routing check looked up the "mxcli-project-toolkit
Integration" heading with an unguarded grep | head | cut under
set -euo pipefail. A CLAUDE.md that cites query-the-model.md or
learned-mdl-preflight.md without that heading (every mxcli init v0.25
CLAUDE.md) made grep exit 1, and errexit ended the whole sync there:
no bin/ crash-net refresh, no tests/e2e install, no lint-rule restore,
no message. Introduced by 9a0350b (2026-09-18).

Guard it with || true, and ledger_row_prefix too. New fixture
tests/wave2/test-sync-dup-heading-exit.sh: 6/6 on this script, 4 of 6
fail on the pre-fix one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MpdPDrT1dsCR54JrgWmNso
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
`mxcli new` / `mxcli init` always writes `.devcontainer/`, and VS Code then offers (and
remembers) "Reopen in Container". A first-time user followed the docs on 2026-10-08 and
landed inside that container with "claude: command not found" (the template installs
claude to ~/.local/bin and never adds it to PATH) and a fresh Claude login prompt, then
asked how to get rid of the container. Local use was always the intent.

- init-project.sh moves `.devcontainer/` to `.mxtk-backup/devcontainer-<stamp>` unless
  MXTK_KEEP_DEVCONTAINER=1. Nothing deleted, idempotent, probed on a scratch project.
- README: "say no" to the reopen prompt, how to get back out, the lane stays opt-in.
- CHANGELOG line under Unreleased.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
Twelve rule candidates mined from skills/, each with the evidence it rests on, the catalog
field it needs and a status; five rollout rules (two or three rules per PR, ship as warning,
one field run before merge, promote after two clean projects, blind rules say so); what is
not lintable and lives in a skill instead; four brain contradictions to settle before the
rules that touch them. Linked from lint-rules/README.md. Until today this plan existed only
in chat.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
project-bin/lint-gate.sh appends one row per rule per run to
.claude/loop/lint-ledger.tsv (ts, verdict, rule, count, severity, blind,
plus a _run row with the total). New project-bin/lint-trend.sh prints per
rule: runs fired in, first/last count with dates, best/worst, blind runs,
up/down/flat; --rule <id> for one series, --runs for the run rows.

Why: lint-baseline.json records the count that was accepted, not what each
run saw, so "did the rule we shipped catch anything, and did the count go
down" had no answer. The lint rollout plan's "promote after two clean
projects" step needs exactly that.

Probe: the gate's own parser on a crafted lint result (3 findings, one
blind rule) — ledger rows and trend output as expected, lint-last.json and
the ratchet verdict unchanged. check-portability and check-scripts clean.
Field run on a real model still owed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…s on Mendix pages)

Four measured defects (2026-09-17, user-group administration module) where
the Playwright harness reported a defect in the app: widget name unique per
page not per DOM, forced click on a tab swallowed without throwing, grid
re-render after the response not the click, fixture literal written by the
reset script. Routed test/review at Stages 5-6 (diagnose), pointer from
e2e-harness-base.md, CHANGELOG line in this commit. The shipped runner does
not yet apply rules 1-3; the skill states that and the runner is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e2e)

Every unhappy-path case cites a requirement before it runs; results are
PASS / GAP / UNSPECIFIED, and UNSPECIFIED is re-scored against seven house
rules as GAP-HOUSE so a silent requirement cannot hide a user-visible
defect. Measured on a warehouse-management demo project, 2026-09-24/25:
45 cases, 5 modules, 25 -> 27 of 45 PASS, 5 GAP, 13 UNSPECIFIED.

Wired: routing row (test,review; stages 5,6; ondemand; verify) and the
rendered surfaces, a pointer row in testing-shape.md's owner table, and
the CHANGELOG line. testing-shape.md:273 reworded (16-step -> 16 steps)
so the pre-commit numbering check no longer reads a measurement as a
list count.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…cal (#230)

The skill claimed a warm run --local app keeps serving during mxcli test --local.
Ports and DB are separate but mxbuild writes <app dir>/deployment for both, so a
building test run breaks the live app. Rewrote the passage: restart run --local
after test --local; --attach hot-applies into the live app and its database.
No other contradicting sentences found (--attach mentions at lines 317-333 are consistent).
Checks: check-docs-numbering (nothing staged to examine), leak guard clean,
check-pr-discipline clean. Docs only; no fixture covers this file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
# Conflicts:
#	CHANGELOG.md
#	bin/gate-check.sh
_ob_stale and the proof check took `head -c 8000 | ... | head -1`, the first stamp in the
first 8000 bytes, so a review file with dated per-build addenda was judged by its headline
stamp and went false-STALE. Now reads the whole file and takes `tail -1`.

New fixture tests/wave2/test-obligation-valid-at.sh (git scratch repo, synthetic reports):
control old-only stamp -> STALE; old then HEAD stamp -> not STALE; HEAD stamp past 8000 bytes
-> not STALE. Result 3/3 pass post-fix; pre-fix 1/3 (T2, T3 fail). Golden input is not a
tool-output parse here, so hand-built reports are the stamp shape from module-review.md.
Also ran bash -n, check-portability, check-scripts, leak guard, check-pr-discipline: clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
Adds 19 backlog rows (13-31) plus review notes to process/lint-backlog.md, the review record
process/quality-sources-review-2026-10-08.md, and seven skill-only additions (indexes, scheduled
events, Atlas 4 tokens, accessibility floor, workflow targeting, test isolation, rollback semantics).
Lintability read from mxcli source only; no rule probed on a model yet.
Checks: leak guard clean, render-routing --check exit 0, check-pr-discipline clean,
check-portability clean, check-docs-numbering (0 staged docs).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…ck (#229, #245)

PROJECT_MODULES defaults to "*" (all modules the lint run includes, vendor
list excluded by lint-gate) so a fresh install is no longer inert.
ShowPage/ClosePage/ShowHomePage count as feedback (presence only; action_type
strings verified against the mxcli catalog builder). lint-gate --update-baseline
drops _rule findings. Install-time module derivation not done: the existing
vendor-exclusion mechanism makes it unnecessary.

Checks: bash -n lint-gate.sh, tests/test-lint-delivery.sh (ALL GREEN, updated
for the new default), test-stock-hashes.sh (16/0), leak guard, portability,
check-scripts. Starlark behaviour fixture NOT added/run: no mxcli binary or
rule runner in this repo.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
Rung 2 required a non-empty spans.ordered, so a step proving only that a
microflow did NOT fire was skipped with no output. Gate is now ordered OR
mustNotFire non-empty; assertSequence only runs when ordered is non-empty.
Checks: node --check ok; gate-condition probe over 5 span shapes; no live
Jaeger run available, so no end-to-end execution. Selftest covers otel
assertions only and does not exercise the runner loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
claude added 4 commits October 8, 2026 15:14
…233 #234 #235)

#233: new test_surface_hit() in gate-check.sh is the one test behind check_stage_6 and the
Surface column, so the harness pair (docs/verification/report.html + docs/report.json)
no longer reads MISSING while the gate passes.
#234: source-sufficiency report title uses realpath basename, so "." / trailing slash
no longer yield an empty or "." title.
#235: artifact-manifest.tsv accepts a ${JOURNEY_DIR} token; artifact-check.sh resolves it
via project-bin/_common.sh (env > toolkit.env > ~/.mxcli-toolkit.env > journeys/), relative
or absolute. Ad-hoc checked: default, relative env, absolute env, toolkit.env file.
No golden-input capture applies (no parser added).

Checks: bash -n; check-portability, check-no-client-data (leak guard), check-scripts,
check-pr-discipline; fixtures test-source-sufficiency (51/0), test-source-sufficiency-gate
(7/0), test-bug03-gates (31/0), test-wrong-verdicts (20/0).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
… fixture

The fixture header named consuming projects by client name; the public
toolkit cites field runs without naming clients.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VJgWP5vEoAsNsJYqCDMGNw
…ation/2026-10-08

# Conflicts:
#	CHANGELOG.md
@MendixMau

Copy link
Copy Markdown
Owner Author

Test results on head 5f7d293

Check Result
tests/run-tests.sh passed=33 failed=0
tests/test-lint-delivery.sh ALL GREEN
tests/wave2/run-all.sh 75 fixtures, 73 ok, 2 failing, both environmental (below)
bin/render-routing.sh --check, check-scripts, check-portability, leak guard, check-no-private-citations, check-pr-discipline all rc=0
Scaffold smoke (init → gate-check 0/6 → sync → artifact-check, fresh dir, no .mpr) no crashes; CONV020 installs with PROJECT_MODULES = "*"; ${JOURNEY_DIR} row present in the artifact manifest

The two wave2 failures are the runner's environment, not the branch:

  • test-guide-reopen.sh — the suite ran with MXTK_NO_GUIDE=1 exported, so the cold-project run opened the guide 0 times. Re-run without that variable: PASS=10 FAIL=0.
  • test-harvest-learnings.sh — its own control step fails ("clone has no history for project-bin/snapshot-mpr.sh"): the checkout is a shallow clone. Not reproducible on a full clone; CI uses one.

One fix folded in since the first push: the smoke test's forbidden-word grep found a client name in the field-evidence note of tests/wave2/test-sync-dup-heading-exit.sh and its CHANGELOG line (both from #231). Replaced with neutral wording on the #231 branch (d3f1985) and merged here.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants