Skip to content
Merged
17 changes: 17 additions & 0 deletions .changeset/20289-os-test-names-tags.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
---
'@objectstack/cli': minor
'@objectstack/core': minor
'@objectstack/spec': patch
---

`os test` reports the suite and scenario names an author writes, and selects scenarios with `--tags` (#20289)

Clause-②: no

A Quality Protocol suite's `name`, each scenario's `name` and `description`, and scenario `tags` were parsed at load and then read by nothing: the report headed each suite with its file's basename, printed every scenario by its `id`, and `os test --tags critical` failed with `Nonexistent flag: --tags`.

- **Names in the report.** The suite heading is now the suite's `name` followed by its file — `📄 Running suite: Accounts smoke (accounts.test.json)` — and each scenario line is its `name` with the `id` in brackets — `✅ Scenario: An account can be created [acct-create] (12ms)` (the id alone when the two are equal). A failed scenario's `description` is printed under its line, before the error. A suite whose file fails to load is still headed by the file alone, since no name was parsed.
- **`--tags TAG[,TAG...]`** runs only the scenarios carrying AT LEAST ONE of the listed tags (any-of, exact, case-sensitive) — the comma-list reading of Odoo's `--test-tags` and the everyday use of Playwright's `--grep @a|@b`. With the flag, an untagged scenario is left out. Left-out scenarios are **deselected**: not run, counted on the summary (`--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).`), never counted as passed. A requested tag that no loaded scenario carries is named on the summary. An empty entry (`--tags smoke,`) is refused before anything runs. Without the flag nothing changes: every scenario runs.
- **Exit status.** A selection that matches no scenario takes the posture an empty pattern already has: exit `0` with `No scenario matched --tags …`, and exit `1` under `--fail-on-empty`, whose description now covers both cases. The `Found N test suites.` line and the `SUCCESS: All N scenarios passed.` / `FAILED: …` summary lines keep their spelling.
- **`@objectstack/core`:** `QA.TestResult` gains `scenarioName` and `description` on every result, and `suiteName` on every result `runSuite` produces (absent only from a lone `runScenario` call, which has no suite).
- **`@objectstack/spec`:** `TestScenario.requires` (`params`, `plugins`) is still checked by nothing — its describe() now says **NOT CHECKED** instead of reading as a guard, so a scenario that declares a plugin the target lacks still runs, and the unmet requirement surfaces only as whatever failure it causes, if any. The liveness ledger (`liveness/qa.json`) moves the four keys above to `live`, citing their readers.
29 changes: 29 additions & 0 deletions content/docs/deployment/cli.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -1651,6 +1651,7 @@ os test --url http://localhost:4000 # Custom server URL
os test --token my-api-key # With authentication
os test 'qa/**/*.test.json' # Recursive — quote it, or the shell expands it first
os test --fail-on-empty # Matching no suite is a failure, not a pass
os test --tags smoke,critical # Only scenarios tagged smoke OR critical
```

The pattern accepts `*` (one path segment) and `**` (any number of segments);
Expand All @@ -1677,6 +1678,30 @@ switch and reported ✅, so a `contains` against a missing path was a test that
silently deleted itself. Assert absence with `is_null`; compare a scalar with
`equals`.

The report prints the **names the suite author wrote**. Each suite is headed by
its `name` and the file it was loaded from —
`📄 Running suite: Accounts smoke (accounts.test.json)` — and each scenario line
leads with the scenario's `name` and carries its `id` in brackets:
`✅ Scenario: An account can be created [acct-create] (12ms)`, or the id alone when
the two are equal. A failed scenario also prints its `description` under its line,
before the error, so a red run says what the scenario was checking. A file that is
refused at load time is headed by its file name alone, because no suite name was
ever parsed from it.

**`--tags` selects scenarios by their `tags`.** Pass a comma-separated list: a
scenario runs when it carries **at least one** of the listed tags (any-of), so
`--tags smoke,critical` runs everything tagged `smoke` or `critical`. Matching is
exact and case-sensitive — a tag is a name, not a pattern — and while the flag is
given an untagged scenario is never selected. The scenarios the selection leaves
out are **deselected**: not run, counted on the summary
(`--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).`),
and never counted as passed. A listed tag that no loaded scenario carries is named
on the summary, since a typo narrows the run without failing it, and an empty entry
(`--tags smoke,`) is refused before anything runs. Without the flag every scenario
runs. A scenario's `requires` block is **not checked**: a scenario that names a
plugin the target does not load still runs, and the unmet requirement surfaces
only as whatever failure it causes, if any — the scenario can still pass.

**A pattern that matches no suite is not a failure by default.** The run prints
`Found 0 test suites.` — the same machine-readable line a full run prints, so a
caller can tell "every suite passed" from "there were no suites" — and exits
Expand All @@ -1685,6 +1710,10 @@ build. That is a posture, not an oversight, and it has the cost you would expect
a CI step whose glob stops matching (a renamed directory, a moved suite) reports
success forever. Pass **`--fail-on-empty`** to opt into the strict reading, where
an empty match exits 1 (#7848).
A `--tags` selection that matches no scenario takes the same posture: it prints
`No scenario matched --tags nightly.` and exits **0**, and **`--fail-on-empty`**
makes it exit 1 — a renamed tag in a CI step is the same trap as a renamed
directory.

The **record-shaped** action types — `create_record`, `read_record`,
`update_record`, `delete_record`, `query_records` — **ask the server where the
Expand Down
8 changes: 4 additions & 4 deletions content/docs/references/qa/testing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ A complete test scenario with setup, execution steps, and teardown
| **setup** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | optional | Steps to run before main test (preconditions) |
| **steps** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | ✅ | Main test sequence to execute |
| **teardown** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | optional | Steps to cleanup after test execution |
| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario |
| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario. NOT CHECKED by `os test` or the core TestRunner: the scenario runs whether or not they hold, and an unmet requirement surfaces only as the failure it causes |

### Nested Shape: `TestScenario.setup[number]`

Expand Down Expand Up @@ -157,8 +157,8 @@ A single step in a test scenario, consisting of an action and optional assertion

| Property | Type | Required | Description |
| :--- | :--- | :--- | :--- |
| **params** | `string[]` | optional | Required environment variables or parameters |
| **plugins** | `string[]` | optional | Required plugins that must be loaded |
| **params** | `string[]` | optional | Environment variables or parameters the scenario needs. Declared only: nothing checks them before the scenario runs |
| **plugins** | `string[]` | optional | Plugins the scenario needs loaded on the target. Declared only: nothing checks them before the scenario runs |


---
Expand Down Expand Up @@ -223,7 +223,7 @@ A complete test scenario with setup, execution steps, and teardown
| **setup** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | optional | Steps to run before main test (preconditions) |
| **steps** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | ✅ | Main test sequence to execute |
| **teardown** | `{ name: string; description?: string; action: object; assertions?: object[]; … }[]` | optional | Steps to cleanup after test execution |
| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario |
| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario. NOT CHECKED by `os test` or the core TestRunner: the scenario runs whether or not they hold, and an unmet requirement surfaces only as the failure it causes |


---
Expand Down
9 changes: 5 additions & 4 deletions docs/qa/platform-checklist/areas/cli.json
Original file line number Diff line number Diff line change
Expand Up @@ -438,7 +438,7 @@
"title": "os test: a Quality Protocol suite is validated at LOAD, executed against a booted app, and its verdict is the exit code — capture/interpolation thread state, an unevaluable assertion FAILS",
"since": "v17",
"status": "active",
"revision": 4,
"revision": 5,
"priority": "P1",
"surface": "cli",
"personas": ["operator (local shell)", "suite author (writes qa/*.test.json)"],
Expand All @@ -453,7 +453,7 @@
"knownGaps": [
"`run_script` has no adapter branch and fails by name (`Unsupported action type in HttpAdapter: run_script`) — the variant sweep records ONE refusal and that is the honest verdict, not a fixture gap to work around. The other seven members execute since #7848; the five record-shaped ones did not until then (see `negative`), so a sweep transcript predating that fix shows five 404s and is not comparable",
"the record action types resolve the Data Protocol mount from the server's own `/discovery` (`routes.data`) since #7983, so a deployment that sets `crud.dataPrefix` IS reached; what remains out of reach is `api.apiPath`, which moves the discovery document itself out from under the probe (measured: `{apiBase}/discovery` 404s, and `/.well-known/objectstack` reports the DISPATCHER's `/api/v1/data`, not the REST mount). Against such a host the adapter falls back to the convention and says so — it names the mount it addressed, the probe that failed and the remedy, on the warning AND on every 404 — and those steps must still be written as `api_call`. The fixture boot is stock, so neither case bites here",
"no scenario SELECTION exists: `os test` has exactly three flags (--url, --token, --fail-on-empty — commands/test.ts), none of which selects scenarios, `scenario.tags` filters nothing and `scenario.requires` is never checked (packages/spec/liveness/qa.json rows `tags`/`requires`), so the whole glob always runs and a suite cannot declare a precondition it will be skipped for"
"scenario SELECTION is by tag only, and preconditions are NOT honoured: `os test --tags a,b` runs the scenarios carrying at least one listed tag (any-of, exact, case-sensitive — commands/test.ts `selectScenariosByTags`), counts the rest as deselected — not run, never counted as passed — and without the flag the whole glob runs; `scenario.requires` (`params`/`plugins`) is declared but NOT CHECKED — its describe() says so and its row in packages/spec/liveness/qa.json stays `dead` pending the maintainer's decision on #20289 — so a suite still cannot declare a precondition it will be skipped for, and a scenario naming a plugin the target lacks still runs, the unmet requirement surfacing only as whatever failure it causes, if any"
]
},
"steps": [
Expand Down Expand Up @@ -548,7 +548,7 @@
"packages/core/src/qa/runner.ts#scenario (scenario sequencing, `capture` + `{{var}}` interpolation, the assertion operators, setup/teardown semantics, the #7256 unevaluable-`contains` fix)",
"packages/core/src/qa/http-adapter.ts#mount (the action-type switch — its case labels ARE the enum values; the record routes take their prefix from the one memoised `/discovery` probe per run, falling back to the RestApiConfigSchema + CrudEndpointsConfigSchema convention with a diagnostic that names the mount, #7848 / #7983)",
"packages/spec/src/qa/testing.zod.ts#TestSuiteSchema (TestSuiteSchema — the shape enforced at load; TestActionTypeSchema pinned above)",
"packages/spec/liveness/qa.json#tags (the ADR-0049 ledger whose existence this item is coverage.json's mapping for — its dead `tags`/`requires` rows are why no scenario selection exists)",
"packages/spec/liveness/qa.json#tags (the ADR-0049 ledger whose existence this item is coverage.json's mapping for — its `tags` row is live on `os test --tags`, and its `requires` row stays dead: the precondition is declared and not checked)",
"packages/cli/test/qa-suite-schema-load.test.ts, packages/cli/test/resolve-glob-lazy-walk.test.ts, packages/core/src/qa/runner.test.ts (the three unit pins — cited so a run knows what is already covered, NOT a substitute for driving a booted app)",
"content/docs/deployment/cli.mdx §os test (the documented command contract)",
"examples/app-showcase/qa/platform-smoke.test.json (the fixture suite this item drives)"
Expand All @@ -557,7 +557,8 @@
{ "revision": 1, "date": "2026-08-11", "change": "new item: the `qa` capability's coverage.json mapping, authored rather than waived (#7347 triage ruling). `os test` is a shipped, documented CLI command, so the honest mapping is a surface:cli item that authors a real qa/*.test.json suite and drives it against a booted app — the fixture suite examples/app-showcase/qa/platform-smoke.test.json lands with this item and is the repo's first Quality Protocol suite. Every clause was measured on showcase before it was written: the green path, capture/interpolation, the #6247 load refusal, the #7256 unevaluable-contains failure, teardown-after-failure, the 8-member action-type sweep and the #7363 glob. `since: v17` records the release in which the surface became GOVERNED (liveness ledger seeded + TestSuiteSchema enforced at the load site, #6247 / PR #7255); the command itself predates it. No `automated` entry: the three unit pins cover pieces, none of them proves a suite reaches a real server", "ref": "#7347" },
{"revision": 2, "date": "2026-08-12", "change": "the adapter was repaired, which this item's own `negative` clause declared to be a revision rather than a silent green (#7848). Item 1: the five record-shaped action types now round-trip against a stock server — the `${baseUrl}/api/data/:object` literal became a prefix DERIVED from the two schemas RestServer itself resolves from (RestApiConfigSchema `apiPath ?? {basePath}/{version}` + CrudEndpointsConfigSchema.dataPrefix), and `update_record` PATCHes where it used to PUT a route that has no PUT sibling. Re-measured on a booted showcase, one scenario per member with NO shared setup so no member's verdict is inferred from a sibling: 7 of 8 execute and assert, `run_script` still refuses by name. The `negative` clause is inverted accordingly — a 404 from a record action type is the regression now — and the knownGap it rested on is replaced by the narrower one that survives: the record types address the DEFAULT mount only, so a host that moves it with `api.apiPath`/`crud.dataPrefix` still needs `api_call`. Item 2: a zero-match glob still exits 0, deliberately (a repo that legitimately ships no suites must not start failing CI), but the posture is now DECLARED — stated in `--help`, opt out with the new `--fail-on-empty`, and `Found N test suites.` is emitted on EVERY run including `Found 0 test suites.`, which is the line this item's first acceptance clause already asks a run record to quote", "ref": "#7848"},
{"revision": 3, "date": "2026-08-17", "change": "the record action types stopped ASSUMING the mount (#7983). They now resolve it from the server: one memoised `GET {apiBase}/discovery` per run, addressing whatever `routes.data` advertises, with the RestApiConfigSchema + CrudEndpointsConfigSchema convention as the fallback — the `@objectstack/client` `getRoute` pattern, copied rather than re-invented. Measured on a booted stack (REST generator + dispatcher bridge) before and after, three configs: stock stays green; `crud.dataPrefix: '/objects'` went from `HTTP Error 404` to a created record, so that row of the gap is CLOSED; `api.apiPath: '/api/2026-01'` still 404s and is NOT closed — `apiPath` moves the discovery document itself, and the one fixed-path document (`/.well-known/objectstack`) advertises the dispatcher's `/api/v1/data` under all three configs, so trusting it would attach a false provenance to the same 404. That row is narrowed instead: the fallback is announced (a warning naming the mount, the failed probe and the remedy) and every 404 from a record action now carries the mount it addressed and where that mount came from, so the failure can no longer read as the suite author's own URL mistake. `api_call` is unchanged and probes nothing — it remains the escape hatch for a host the probe cannot reach", "ref": "#7983"},
{ "revision": 4, "date": "2026-08-18", "change": "corrected the flag count from two to three. packages/cli/src/commands/test.ts declares url, token AND fail-on-empty; the same revision's own acceptance already describes --fail-on-empty, so the knownGap contradicted its own item. The gap's point is unchanged and now stated directly: none of the three flags selects scenarios (#9417)", "ref": "#9386" }
{ "revision": 4, "date": "2026-08-18", "change": "corrected the flag count from two to three. packages/cli/src/commands/test.ts declares url, token AND fail-on-empty; the same revision's own acceptance already describes --fail-on-empty, so the knownGap contradicted its own item. The gap's point is unchanged and now stated directly: none of the three flags selects scenarios (#9417)", "ref": "#9386" },
{ "revision": 5, "date": "2026-09-28", "change": "the knownGap and the qa.json source note were falsified by #20289 (PR #20341) and are corrected: `os test` gained `--tags` (any-of, exact, case-sensitive; deselected scenarios are counted and never passed; a selection that matches no scenario exits 0, or 1 under --fail-on-empty), and the report now heads each suite with its `name` and prints each scenario's `name` beside its id. What stays true is narrower: `scenario.requires` is declared but NOT CHECKED, pending the maintainer's decision on #20289. Steps and acceptance are unchanged — the green-path clause's scenarioId is still on every scenario line, now in brackets after the name, or bare when the name equals the id", "ref": "#20289" }
]
},
{
Expand Down
Loading
Loading