diff --git a/.changeset/20289-os-test-names-tags.md b/.changeset/20289-os-test-names-tags.md new file mode 100644 index 00000000000..56d93b9abdf --- /dev/null +++ b/.changeset/20289-os-test-names-tags.md @@ -0,0 +1,17 @@ +--- +'@objectstack/cli': minor +'@objectstack/core': minor +'@objectstack/spec': patch +--- + +`os test` reports the suite and scenario names an author writes, and selects scenarios with `--tags` (#20289) + +Clause-โ‘ก: no + +A Quality Protocol suite's `name`, each scenario's `name` and `description`, and scenario `tags` were parsed at load and then read by nothing: the report headed each suite with its file's basename, printed every scenario by its `id`, and `os test --tags critical` failed with `Nonexistent flag: --tags`. + +- **Names in the report.** The suite heading is now the suite's `name` followed by its file โ€” `๐Ÿ“„ Running suite: Accounts smoke (accounts.test.json)` โ€” and each scenario line is its `name` with the `id` in brackets โ€” `โœ… Scenario: An account can be created [acct-create] (12ms)` (the id alone when the two are equal). A failed scenario's `description` is printed under its line, before the error. A suite whose file fails to load is still headed by the file alone, since no name was parsed. +- **`--tags TAG[,TAG...]`** runs only the scenarios carrying AT LEAST ONE of the listed tags (any-of, exact, case-sensitive) โ€” the comma-list reading of Odoo's `--test-tags` and the everyday use of Playwright's `--grep @a|@b`. With the flag, an untagged scenario is left out. Left-out scenarios are **deselected**: not run, counted on the summary (`--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).`), never counted as passed. A requested tag that no loaded scenario carries is named on the summary. An empty entry (`--tags smoke,`) is refused before anything runs. Without the flag nothing changes: every scenario runs. +- **Exit status.** A selection that matches no scenario takes the posture an empty pattern already has: exit `0` with `No scenario matched --tags โ€ฆ`, and exit `1` under `--fail-on-empty`, whose description now covers both cases. The `Found N test suites.` line and the `SUCCESS: All N scenarios passed.` / `FAILED: โ€ฆ` summary lines keep their spelling. +- **`@objectstack/core`:** `QA.TestResult` gains `scenarioName` and `description` on every result, and `suiteName` on every result `runSuite` produces (absent only from a lone `runScenario` call, which has no suite). +- **`@objectstack/spec`:** `TestScenario.requires` (`params`, `plugins`) is still checked by nothing โ€” its describe() now says **NOT CHECKED** instead of reading as a guard, so a scenario that declares a plugin the target lacks still runs, and the unmet requirement surfaces only as whatever failure it causes, if any. The liveness ledger (`liveness/qa.json`) moves the four keys above to `live`, citing their readers. diff --git a/content/docs/deployment/cli.mdx b/content/docs/deployment/cli.mdx index 247bfeadca4..30d11ba0750 100644 --- a/content/docs/deployment/cli.mdx +++ b/content/docs/deployment/cli.mdx @@ -1651,6 +1651,7 @@ os test --url http://localhost:4000 # Custom server URL os test --token my-api-key # With authentication os test 'qa/**/*.test.json' # Recursive โ€” quote it, or the shell expands it first os test --fail-on-empty # Matching no suite is a failure, not a pass +os test --tags smoke,critical # Only scenarios tagged smoke OR critical ``` The pattern accepts `*` (one path segment) and `**` (any number of segments); @@ -1677,6 +1678,30 @@ switch and reported โœ…, so a `contains` against a missing path was a test that silently deleted itself. Assert absence with `is_null`; compare a scalar with `equals`. +The report prints the **names the suite author wrote**. Each suite is headed by +its `name` and the file it was loaded from โ€” +`๐Ÿ“„ Running suite: Accounts smoke (accounts.test.json)` โ€” and each scenario line +leads with the scenario's `name` and carries its `id` in brackets: +`โœ… Scenario: An account can be created [acct-create] (12ms)`, or the id alone when +the two are equal. A failed scenario also prints its `description` under its line, +before the error, so a red run says what the scenario was checking. A file that is +refused at load time is headed by its file name alone, because no suite name was +ever parsed from it. + +**`--tags` selects scenarios by their `tags`.** Pass a comma-separated list: a +scenario runs when it carries **at least one** of the listed tags (any-of), so +`--tags smoke,critical` runs everything tagged `smoke` or `critical`. Matching is +exact and case-sensitive โ€” a tag is a name, not a pattern โ€” and while the flag is +given an untagged scenario is never selected. The scenarios the selection leaves +out are **deselected**: not run, counted on the summary +(`--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).`), +and never counted as passed. A listed tag that no loaded scenario carries is named +on the summary, since a typo narrows the run without failing it, and an empty entry +(`--tags smoke,`) is refused before anything runs. Without the flag every scenario +runs. A scenario's `requires` block is **not checked**: a scenario that names a +plugin the target does not load still runs, and the unmet requirement surfaces +only as whatever failure it causes, if any โ€” the scenario can still pass. + **A pattern that matches no suite is not a failure by default.** The run prints `Found 0 test suites.` โ€” the same machine-readable line a full run prints, so a caller can tell "every suite passed" from "there were no suites" โ€” and exits @@ -1685,6 +1710,10 @@ build. That is a posture, not an oversight, and it has the cost you would expect a CI step whose glob stops matching (a renamed directory, a moved suite) reports success forever. Pass **`--fail-on-empty`** to opt into the strict reading, where an empty match exits 1 (#7848). +A `--tags` selection that matches no scenario takes the same posture: it prints +`No scenario matched --tags nightly.` and exits **0**, and **`--fail-on-empty`** +makes it exit 1 โ€” a renamed tag in a CI step is the same trap as a renamed +directory. The **record-shaped** action types โ€” `create_record`, `read_record`, `update_record`, `delete_record`, `query_records` โ€” **ask the server where the diff --git a/content/docs/references/qa/testing.mdx b/content/docs/references/qa/testing.mdx index 36c0dbfcb6d..51ab7765103 100644 --- a/content/docs/references/qa/testing.mdx +++ b/content/docs/references/qa/testing.mdx @@ -115,7 +115,7 @@ A complete test scenario with setup, execution steps, and teardown | **setup** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | optional | Steps to run before main test (preconditions) | | **steps** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | โœ… | Main test sequence to execute | | **teardown** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | optional | Steps to cleanup after test execution | -| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario | +| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario. NOT CHECKED by `os test` or the core TestRunner: the scenario runs whether or not they hold, and an unmet requirement surfaces only as the failure it causes | ### Nested Shape: `TestScenario.setup[number]` @@ -157,8 +157,8 @@ A single step in a test scenario, consisting of an action and optional assertion | Property | Type | Required | Description | | :--- | :--- | :--- | :--- | -| **params** | `string[]` | optional | Required environment variables or parameters | -| **plugins** | `string[]` | optional | Required plugins that must be loaded | +| **params** | `string[]` | optional | Environment variables or parameters the scenario needs. Declared only: nothing checks them before the scenario runs | +| **plugins** | `string[]` | optional | Plugins the scenario needs loaded on the target. Declared only: nothing checks them before the scenario runs | --- @@ -223,7 +223,7 @@ A complete test scenario with setup, execution steps, and teardown | **setup** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | optional | Steps to run before main test (preconditions) | | **steps** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | โœ… | Main test sequence to execute | | **teardown** | `{ name: string; description?: string; action: object; assertions?: object[]; โ€ฆ }[]` | optional | Steps to cleanup after test execution | -| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario | +| **requires** | `{ params?: string[]; plugins?: string[] }` | optional | Environment requirements for this scenario. NOT CHECKED by `os test` or the core TestRunner: the scenario runs whether or not they hold, and an unmet requirement surfaces only as the failure it causes | --- diff --git a/docs/qa/platform-checklist/areas/cli.json b/docs/qa/platform-checklist/areas/cli.json index 1dd4608414e..d328057fed9 100644 --- a/docs/qa/platform-checklist/areas/cli.json +++ b/docs/qa/platform-checklist/areas/cli.json @@ -438,7 +438,7 @@ "title": "os test: a Quality Protocol suite is validated at LOAD, executed against a booted app, and its verdict is the exit code โ€” capture/interpolation thread state, an unevaluable assertion FAILS", "since": "v17", "status": "active", - "revision": 4, + "revision": 5, "priority": "P1", "surface": "cli", "personas": ["operator (local shell)", "suite author (writes qa/*.test.json)"], @@ -453,7 +453,7 @@ "knownGaps": [ "`run_script` has no adapter branch and fails by name (`Unsupported action type in HttpAdapter: run_script`) โ€” the variant sweep records ONE refusal and that is the honest verdict, not a fixture gap to work around. The other seven members execute since #7848; the five record-shaped ones did not until then (see `negative`), so a sweep transcript predating that fix shows five 404s and is not comparable", "the record action types resolve the Data Protocol mount from the server's own `/discovery` (`routes.data`) since #7983, so a deployment that sets `crud.dataPrefix` IS reached; what remains out of reach is `api.apiPath`, which moves the discovery document itself out from under the probe (measured: `{apiBase}/discovery` 404s, and `/.well-known/objectstack` reports the DISPATCHER's `/api/v1/data`, not the REST mount). Against such a host the adapter falls back to the convention and says so โ€” it names the mount it addressed, the probe that failed and the remedy, on the warning AND on every 404 โ€” and those steps must still be written as `api_call`. The fixture boot is stock, so neither case bites here", - "no scenario SELECTION exists: `os test` has exactly three flags (--url, --token, --fail-on-empty โ€” commands/test.ts), none of which selects scenarios, `scenario.tags` filters nothing and `scenario.requires` is never checked (packages/spec/liveness/qa.json rows `tags`/`requires`), so the whole glob always runs and a suite cannot declare a precondition it will be skipped for" + "scenario SELECTION is by tag only, and preconditions are NOT honoured: `os test --tags a,b` runs the scenarios carrying at least one listed tag (any-of, exact, case-sensitive โ€” commands/test.ts `selectScenariosByTags`), counts the rest as deselected โ€” not run, never counted as passed โ€” and without the flag the whole glob runs; `scenario.requires` (`params`/`plugins`) is declared but NOT CHECKED โ€” its describe() says so and its row in packages/spec/liveness/qa.json stays `dead` pending the maintainer's decision on #20289 โ€” so a suite still cannot declare a precondition it will be skipped for, and a scenario naming a plugin the target lacks still runs, the unmet requirement surfacing only as whatever failure it causes, if any" ] }, "steps": [ @@ -548,7 +548,7 @@ "packages/core/src/qa/runner.ts#scenario (scenario sequencing, `capture` + `{{var}}` interpolation, the assertion operators, setup/teardown semantics, the #7256 unevaluable-`contains` fix)", "packages/core/src/qa/http-adapter.ts#mount (the action-type switch โ€” its case labels ARE the enum values; the record routes take their prefix from the one memoised `/discovery` probe per run, falling back to the RestApiConfigSchema + CrudEndpointsConfigSchema convention with a diagnostic that names the mount, #7848 / #7983)", "packages/spec/src/qa/testing.zod.ts#TestSuiteSchema (TestSuiteSchema โ€” the shape enforced at load; TestActionTypeSchema pinned above)", - "packages/spec/liveness/qa.json#tags (the ADR-0049 ledger whose existence this item is coverage.json's mapping for โ€” its dead `tags`/`requires` rows are why no scenario selection exists)", + "packages/spec/liveness/qa.json#tags (the ADR-0049 ledger whose existence this item is coverage.json's mapping for โ€” its `tags` row is live on `os test --tags`, and its `requires` row stays dead: the precondition is declared and not checked)", "packages/cli/test/qa-suite-schema-load.test.ts, packages/cli/test/resolve-glob-lazy-walk.test.ts, packages/core/src/qa/runner.test.ts (the three unit pins โ€” cited so a run knows what is already covered, NOT a substitute for driving a booted app)", "content/docs/deployment/cli.mdx ยงos test (the documented command contract)", "examples/app-showcase/qa/platform-smoke.test.json (the fixture suite this item drives)" @@ -557,7 +557,8 @@ { "revision": 1, "date": "2026-08-11", "change": "new item: the `qa` capability's coverage.json mapping, authored rather than waived (#7347 triage ruling). `os test` is a shipped, documented CLI command, so the honest mapping is a surface:cli item that authors a real qa/*.test.json suite and drives it against a booted app โ€” the fixture suite examples/app-showcase/qa/platform-smoke.test.json lands with this item and is the repo's first Quality Protocol suite. Every clause was measured on showcase before it was written: the green path, capture/interpolation, the #6247 load refusal, the #7256 unevaluable-contains failure, teardown-after-failure, the 8-member action-type sweep and the #7363 glob. `since: v17` records the release in which the surface became GOVERNED (liveness ledger seeded + TestSuiteSchema enforced at the load site, #6247 / PR #7255); the command itself predates it. No `automated` entry: the three unit pins cover pieces, none of them proves a suite reaches a real server", "ref": "#7347" }, {"revision": 2, "date": "2026-08-12", "change": "the adapter was repaired, which this item's own `negative` clause declared to be a revision rather than a silent green (#7848). Item 1: the five record-shaped action types now round-trip against a stock server โ€” the `${baseUrl}/api/data/:object` literal became a prefix DERIVED from the two schemas RestServer itself resolves from (RestApiConfigSchema `apiPath ?? {basePath}/{version}` + CrudEndpointsConfigSchema.dataPrefix), and `update_record` PATCHes where it used to PUT a route that has no PUT sibling. Re-measured on a booted showcase, one scenario per member with NO shared setup so no member's verdict is inferred from a sibling: 7 of 8 execute and assert, `run_script` still refuses by name. The `negative` clause is inverted accordingly โ€” a 404 from a record action type is the regression now โ€” and the knownGap it rested on is replaced by the narrower one that survives: the record types address the DEFAULT mount only, so a host that moves it with `api.apiPath`/`crud.dataPrefix` still needs `api_call`. Item 2: a zero-match glob still exits 0, deliberately (a repo that legitimately ships no suites must not start failing CI), but the posture is now DECLARED โ€” stated in `--help`, opt out with the new `--fail-on-empty`, and `Found N test suites.` is emitted on EVERY run including `Found 0 test suites.`, which is the line this item's first acceptance clause already asks a run record to quote", "ref": "#7848"}, {"revision": 3, "date": "2026-08-17", "change": "the record action types stopped ASSUMING the mount (#7983). They now resolve it from the server: one memoised `GET {apiBase}/discovery` per run, addressing whatever `routes.data` advertises, with the RestApiConfigSchema + CrudEndpointsConfigSchema convention as the fallback โ€” the `@objectstack/client` `getRoute` pattern, copied rather than re-invented. Measured on a booted stack (REST generator + dispatcher bridge) before and after, three configs: stock stays green; `crud.dataPrefix: '/objects'` went from `HTTP Error 404` to a created record, so that row of the gap is CLOSED; `api.apiPath: '/api/2026-01'` still 404s and is NOT closed โ€” `apiPath` moves the discovery document itself, and the one fixed-path document (`/.well-known/objectstack`) advertises the dispatcher's `/api/v1/data` under all three configs, so trusting it would attach a false provenance to the same 404. That row is narrowed instead: the fallback is announced (a warning naming the mount, the failed probe and the remedy) and every 404 from a record action now carries the mount it addressed and where that mount came from, so the failure can no longer read as the suite author's own URL mistake. `api_call` is unchanged and probes nothing โ€” it remains the escape hatch for a host the probe cannot reach", "ref": "#7983"}, - { "revision": 4, "date": "2026-08-18", "change": "corrected the flag count from two to three. packages/cli/src/commands/test.ts declares url, token AND fail-on-empty; the same revision's own acceptance already describes --fail-on-empty, so the knownGap contradicted its own item. The gap's point is unchanged and now stated directly: none of the three flags selects scenarios (#9417)", "ref": "#9386" } + { "revision": 4, "date": "2026-08-18", "change": "corrected the flag count from two to three. packages/cli/src/commands/test.ts declares url, token AND fail-on-empty; the same revision's own acceptance already describes --fail-on-empty, so the knownGap contradicted its own item. The gap's point is unchanged and now stated directly: none of the three flags selects scenarios (#9417)", "ref": "#9386" }, + { "revision": 5, "date": "2026-09-28", "change": "the knownGap and the qa.json source note were falsified by #20289 (PR #20341) and are corrected: `os test` gained `--tags` (any-of, exact, case-sensitive; deselected scenarios are counted and never passed; a selection that matches no scenario exits 0, or 1 under --fail-on-empty), and the report now heads each suite with its `name` and prints each scenario's `name` beside its id. What stays true is narrower: `scenario.requires` is declared but NOT CHECKED, pending the maintainer's decision on #20289. Steps and acceptance are unchanged โ€” the green-path clause's scenarioId is still on every scenario line, now in brackets after the name, or bare when the name equals the id", "ref": "#20289" } ] }, { diff --git a/packages/cli/src/commands/test.ts b/packages/cli/src/commands/test.ts index 0365c7deec3..774d2436cf4 100644 --- a/packages/cli/src/commands/test.ts +++ b/packages/cli/src/commands/test.ts @@ -245,6 +245,90 @@ export function foundSuitesLine(count: number): string { return `Found ${count} test suites.`; } +/** + * Parse the `--tags` value: a comma-separated list of tag names, trimmed and + * de-duplicated in the order given. + * + * An empty entry (`--tags smoke,`, `--tags ""`) is refused rather than dropped. + * Dropping it would turn a typo into a narrower selection than the one asked + * for, and nothing in the run would say so. + */ +export function parseTagsFlag(raw: string): string[] { + const tags = raw.split(',').map((tag) => tag.trim()); + if (tags.some((tag) => tag.length === 0)) { + throw new Error( + `--tags "${raw}" contains an empty tag name. ` + + 'Pass one or more tag names separated by commas, e.g. --tags smoke,critical.', + ); + } + return [...new Set(tags)]; +} + +/** One suite after tag selection. */ +export interface TagSelection { + /** The suite with only the selected scenarios, in their authored order. */ + suite: QA.TestSuite; + /** How many of the suite's scenarios the selection left out. */ + deselected: number; + /** The requested tags at least one selected scenario carries. */ + matchedTags: string[]; +} + +/** + * Select a suite's scenarios by `TestScenario.tags`. + * + * ANY-OF: a scenario is selected when it carries at least one of `tags`, so + * `--tags smoke,critical` runs everything tagged smoke OR critical. That is + * the reading of a comma list in Odoo's `--test-tags` and in Cucumber's + * comma form, and the everyday use of Playwright's `--grep @smoke|@critical`: + * name the sets you want and get their union. Matching is exact and + * case-sensitive โ€” a tag is a name the author chose, not a pattern. An + * expression language (Cucumber tag expressions, pytest `-m`) is a grammar to + * specify and report errors against, which no suite has yet needed. + * + * `tags` undefined โ€” the flag was not passed โ€” selects every scenario: the + * unfiltered run is the default and does not consult `tags` at all. With the + * flag, an untagged scenario is never selected. + */ +export function selectScenariosByTags(suite: QA.TestSuite, tags: readonly string[] | undefined): TagSelection { + if (tags === undefined) return { suite, deselected: 0, matchedTags: [] }; + const matched = new Set(); + const scenarios = suite.scenarios.filter((scenario) => { + const hits = (scenario.tags ?? []).filter((tag) => tags.includes(tag)); + for (const tag of hits) matched.add(tag); + return hits.length > 0; + }); + return { + suite: { ...suite, scenarios }, + deselected: suite.scenarios.length - scenarios.length, + matchedTags: tags.filter((tag) => matched.has(tag)), + }; +} + +/** + * The suite heading: the `name` the author gave the suite, then the file it was + * loaded from โ€” the name is what a person recognises, the file is where to go + * and fix it. + */ +export function suiteHeading(suiteName: string, file: string): string { + return `๐Ÿ“„ Running suite: ${suiteName} (${path.basename(file)})`; +} + +/** + * A scenario's report label: its `name`, then its `id` in brackets. When an + * author made the two identical, the id is printed once. + */ +export function scenarioLabel(result: Pick): string { + return result.scenarioName === result.scenarioId + ? result.scenarioId + : `${result.scenarioName} [${result.scenarioId}]`; +} + +/** The selection line printed above the summary whenever `--tags` is given. */ +export function tagSelectionLine(selected: number, deselected: number, tags: readonly string[]): string { + return `--tags ${tags.join(',')} selected ${selected} of ${selected + deselected} scenarios; ${deselected} deselected (not run, not counted as passed).`; +} + export default class Test extends Command { /** * The empty-match posture is stated in the help text on purpose (#7848). @@ -263,7 +347,13 @@ export default class Test extends Command { '"Found 0 test suites." and exits 0, so a repository that legitimately ships no ' + 'suites does not fail CI. Pass --fail-on-empty for the strict reading, where a ' + 'pattern that has stopped matching (a renamed directory, a moved suite) fails the ' + - 'step instead of reporting success forever.'; + 'step instead of reporting success forever.\n' + + '--tags selects scenarios by their "tags": pass a comma-separated list and a ' + + 'scenario runs when it carries AT LEAST ONE of the listed tags (any-of; exact, ' + + 'case-sensitive). Untagged scenarios are left out whenever --tags is given. Left-out ' + + 'scenarios are counted as deselected โ€” never run and never counted as passed. A ' + + 'selection that matches no scenario exits 0 like an empty pattern, and 1 under ' + + '--fail-on-empty.'; static override args = { files: Args.string({ description: 'Glob pattern for test files (e.g. "qa/*.test.json")', required: false, default: 'qa/*.test.json' }), @@ -273,9 +363,18 @@ export default class Test extends Command { url: Flags.string({ description: 'Target base URL', default: 'http://localhost:3000' }), token: Flags.string({ description: 'Authentication token' }), 'fail-on-empty': Flags.boolean({ - description: 'Exit non-zero when the pattern matches no test suite (default: matching nothing exits 0)', + description: + 'Exit non-zero when the pattern matches no test suite, or --tags matches no scenario (default: both exit 0)', default: false, }), + // A custom flag so a malformed list is refused while the invocation is + // parsed, before a suite is loaded or the server is contacted. + tags: Flags.custom({ + description: + 'Run only scenarios carrying at least one of these comma-separated tags (e.g. "smoke,critical"); the rest are deselected', + helpValue: 'TAG[,TAG...]', + parse: async (input) => parseTagsFlag(input), + })(), }; async run(): Promise { @@ -309,31 +408,49 @@ export default class Test extends Command { console.log(foundSuitesLine(testFiles.length)); // 3. Run Tests + const tags = flags.tags; let totalPassed = 0; let totalFailed = 0; + let totalSelected = 0; + let totalDeselected = 0; + const matchedTags = new Set(); for (const file of testFiles) { - console.log(`\n๐Ÿ“„ Running suite: ${chalk.bold(path.basename(file))}`); - // Load and validate FIRST, and report a refusal on its own terms: a file // the schema rejects never had a chance to run, so folding it into the // run-failure branch below would report it as if the server had said no. - let suite: QA.TestSuite; + let loaded: QA.TestSuite; try { - suite = loadTestSuite(file); + loaded = loadTestSuite(file); } catch (e) { + // No suite name to print โ€” the file is all there is. + console.log(`\n๐Ÿ“„ Running suite: ${chalk.bold(path.basename(file))}`); console.error(chalk.red(e instanceof Error ? e.message : String(e))); totalFailed++; // Count suite failure continue; } + console.log(`\n${chalk.bold(suiteHeading(loaded.name, file))}`); + + const { suite, deselected, matchedTags: matchedHere } = selectScenariosByTags(loaded, tags); + totalSelected += suite.scenarios.length; + totalDeselected += deselected; + for (const tag of matchedHere) matchedTags.add(tag); + if (deselected > 0) { + console.log(chalk.dim(` ${deselected} of ${loaded.scenarios.length} scenarios deselected by --tags`)); + } + try { const results = await runner.runSuite(suite); for (const result of results) { const icon = result.passed ? 'โœ…' : 'โŒ'; - console.log(` ${icon} Scenario: ${result.scenarioId} (${result.duration}ms)`); + console.log(` ${icon} Scenario: ${scenarioLabel(result)} (${result.duration}ms)`); if (!result.passed) { + // What the scenario was checking, in the author's words โ€” the + // first thing a reader of a red run needs and the one thing + // the suite file alone used to hold. + if (result.description) console.log(chalk.dim(` ${result.description}`)); console.error(chalk.red(` Error: ${result.error}`)); result.steps.forEach(step => { if (!step.passed) { @@ -355,9 +472,30 @@ export default class Test extends Command { // 4. Summary console.log(chalk.dim(`\n-------------------------------------`)); + if (tags) { + console.log(tagSelectionLine(totalSelected, totalDeselected, tags)); + // A tag nothing carries is a typo or a renamed tag. It narrows the run + // without failing it, so it is named rather than left to be noticed. + const unmatched = tags.filter((tag) => !matchedTags.has(tag)); + if (unmatched.length > 0 && totalSelected > 0) { + console.warn(chalk.yellow(`--tags: no scenario in the loaded suites carries ${unmatched.map((t) => `"${t}"`).join(', ')}.`)); + } + } if (totalFailed > 0) { console.log(chalk.red(`FAILED: ${totalFailed} scenarios failed. ${totalPassed} passed.`)); process.exit(1); + } else if (tags && totalSelected === 0) { + // The --tags twin of the empty pattern above, and the same posture: a + // run that executed nothing is not a failure by default, and is one + // under --fail-on-empty โ€” a renamed tag in a CI step otherwise reports + // success forever. + console.warn(chalk.yellow(`No scenario matched --tags ${tags.join(',')}.`)); + if (flags['fail-on-empty']) { + console.error(chalk.red(`--fail-on-empty: a run whose --tags selected no scenario is a failed run.`)); + process.exit(1); + } + console.log(chalk.dim(`Exiting 0 โ€” a selection that matches nothing is not a failure. Pass --fail-on-empty to make it one.`)); + process.exit(0); } else { console.log(chalk.green(`SUCCESS: All ${totalPassed} scenarios passed.`)); process.exit(0); diff --git a/packages/cli/test/qa-names-and-tags-run.test.ts b/packages/cli/test/qa-names-and-tags-run.test.ts new file mode 100644 index 00000000000..3b5680d78f9 --- /dev/null +++ b/packages/cli/test/qa-names-and-tags-run.test.ts @@ -0,0 +1,181 @@ +// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license. + +/** + * PIN โ€” what `os test` PRINTS and how it EXITS, for suite/scenario names and + * `--tags`, measured on a real child process against a real HTTP target. + * + * Before: the suite heading was the file's basename and each scenario line + * was its `id`, so the `name`s an author wrote were parsed and shown nowhere; + * and `--tags` was an unknown-flag error. The selection's semantics are held + * by `qa-tags-selection.test.ts`; this file holds the wiring those unit pins + * cannot see โ€” the flag reaching the selection, the names reaching stdout, + * and the exit status a CI step reads. + * + * The fixture is built so each run's EXIT STATUS is itself evidence: the + * unfiltered run (the control) includes a scenario that fails on purpose and + * exits 1, and `--tags smoke` exits 0 only if that scenario was really left + * out โ€” not run, and not merely hidden from the report. + * + * A real child process, because `process.exit(1)` inside a vitest worker is + * not an exit status. Spawned through `bin/run-dev.js` + tsx so the suite does + * not depend on `packages/cli/dist`. The target is a stub `node:http` server + * in this process: the runner's HTTP adapter is what is exercised, and no + * ObjectStack server is booted. + */ + +import { describe, it, expect, beforeAll, afterAll } from 'vitest'; +import { execFile } from 'node:child_process'; +import { createServer, type Server } from 'node:http'; +import type { AddressInfo } from 'node:net'; +import { mkdtempSync, mkdirSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { CLI, TSX, childEnv } from './helpers/serve-process.js'; + +/** oclif + tsx cold start; a healthy run here is a few seconds. */ +const RUN_TIMEOUT_MS = 120_000; + +interface Run { + code: number; + stdout: string; + stderr: string; +} + +function runCli(args: string[], cwd: string): Promise { + return new Promise((resolvePromise) => { + execFile( + TSX, + [CLI, ...args], + { cwd, maxBuffer: 8 * 1024 * 1024, env: childEnv({ NO_COLOR: '1' }) }, + (err, stdout, stderr) => { + resolvePromise({ + // `err.code` is the real exit status; a signalled child is never read as 0. + code: err ? (typeof (err as { code?: unknown }).code === 'number' ? (err as unknown as { code: number }).code : 1) : 0, + stdout: String(stdout), + stderr: String(stderr), + }); + }, + ); + }); +} + +/** One `api_call` step against the stub's health route, asserting `data.status`. */ +function healthStep(expectedStatus: string) { + return { + name: `status is ${expectedStatus}`, + action: { type: 'api_call', target: '/api/v1/health', payload: { method: 'GET' } }, + assertions: [{ field: 'data.status', operator: 'equals', expectedValue: expectedStatus }], + }; +} + +const SUITE = { + name: 'Accounts smoke', + scenarios: [ + { id: 'acct-health', name: 'The server answers its health probe', tags: ['smoke'], steps: [healthStep('ok')] }, + { id: 'acct-regression', name: 'A regression-only check', tags: ['regression'], steps: [healthStep('ok')] }, + { id: 'acct-untagged', name: 'An untagged check', steps: [healthStep('ok')] }, + { + id: 'acct-failing', + name: 'A check that fails on purpose', + description: 'Asserts a status the stub never returns.', + tags: ['regression'], + steps: [healthStep('degraded')], + }, + ], +}; + +const PATTERN = 'qa/*.test.json'; + +let dir: string; +let server: Server; +let unfiltered: Run; +let smoke: Run; +let regression: Run; +let noMatch: Run; +let noMatchStrict: Run; +let malformed: Run; + +beforeAll(async () => { + server = createServer((req, res) => { + res.setHeader('content-type', 'application/json'); + res.end(JSON.stringify({ success: true, data: { status: 'ok' }, path: req.url })); + }); + await new Promise((resolveListen) => server.listen(0, '127.0.0.1', resolveListen)); + const url = `http://127.0.0.1:${(server.address() as AddressInfo).port}`; + + dir = mkdtempSync(join(tmpdir(), 'os-qa-names-tags-')); + mkdirSync(join(dir, 'qa')); + writeFileSync(join(dir, 'qa', 'accounts.test.json'), JSON.stringify(SUITE), 'utf-8'); + + const base = ['test', PATTERN, '--url', url]; + [unfiltered, smoke, regression, noMatch, noMatchStrict, malformed] = await Promise.all([ + runCli(base, dir), + runCli([...base, '--tags', 'smoke'], dir), + runCli([...base, '--tags', 'regression'], dir), + runCli([...base, '--tags', 'nomatch'], dir), + runCli([...base, '--tags', 'nomatch', '--fail-on-empty'], dir), + runCli([...base, '--tags', 'smoke,'], dir), + ]); +}, RUN_TIMEOUT_MS * 2); + +afterAll(async () => { + await new Promise((resolveClose) => server.close(() => resolveClose())); + rmSync(dir, { recursive: true, force: true }); +}); + +describe('`os test` prints the names the author wrote', () => { + it('heads the suite with its name and file', () => { + expect(unfiltered.stdout).toContain('๐Ÿ“„ Running suite: Accounts smoke (accounts.test.json)'); + }); + + it('prints each scenario by name, with its id beside it', () => { + expect(unfiltered.stdout).toContain('โœ… Scenario: The server answers its health probe [acct-health]'); + expect(unfiltered.stdout).toContain('โŒ Scenario: A check that fails on purpose [acct-failing]'); + }); + + it("prints a failed scenario's description", () => { + expect(unfiltered.stdout).toContain('Asserts a status the stub never returns.'); + }); +}); + +describe('`os test --tags`', () => { + it('CONTROL โ€” without --tags every scenario runs, the failing one included, and the run exits 1', () => { + for (const id of ['acct-health', 'acct-regression', 'acct-untagged', 'acct-failing']) { + expect(unfiltered.stdout).toContain(`[${id}]`); + } + expect(unfiltered.stdout).not.toContain('deselected'); + expect(unfiltered.stdout).toContain('FAILED: 1 scenarios failed. 3 passed.'); + expect(unfiltered.code).toBe(1); + }); + + it('--tags smoke runs only the smoke scenario and exits 0 โ€” the failing one never ran', () => { + expect(smoke.stdout).toContain('[acct-health]'); + for (const id of ['acct-regression', 'acct-untagged', 'acct-failing']) { + expect(smoke.stdout).not.toContain(`[${id}]`); + } + expect(smoke.stdout).toContain('--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).'); + expect(smoke.stdout).toContain('SUCCESS: All 1 scenarios passed.'); + expect(smoke.code).toBe(0); + }); + + it('--tags regression runs both regression scenarios and fails on the failing one', () => { + expect(regression.stdout).toContain('[acct-regression]'); + expect(regression.stdout).toContain('[acct-failing]'); + expect(regression.stdout).not.toContain('[acct-health]'); + expect(regression.stdout).toContain('FAILED: 1 scenarios failed. 1 passed.'); + expect(regression.code).toBe(1); + }); + + it('a selection that matches nothing exits 0 by default, and 1 under --fail-on-empty', () => { + expect(`${noMatch.stdout}${noMatch.stderr}`).toContain('No scenario matched --tags nomatch.'); + expect(noMatch.stdout).not.toContain('SUCCESS'); + expect(noMatch.code).toBe(0); + expect(noMatchStrict.code).toBe(1); + }); + + it('a malformed list is refused before anything runs', () => { + expect(malformed.code).not.toBe(0); + expect(malformed.stderr).toContain('empty tag name'); + expect(malformed.stdout).not.toContain('Running suite'); + }); +}); diff --git a/packages/cli/test/qa-tags-selection.test.ts b/packages/cli/test/qa-tags-selection.test.ts new file mode 100644 index 00000000000..87d7c896a52 --- /dev/null +++ b/packages/cli/test/qa-tags-selection.test.ts @@ -0,0 +1,124 @@ +// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license. + +/** + * PIN โ€” `os test --tags` selection semantics, and the labels a report prints. + * + * `TestScenario.tags` was declared "for filtering and categorization" and + * nothing filtered on it: `os test` had no selection flag at all, so + * `--tags critical` was an unknown-flag error and a suite tagged `regression` + * ran on every invocation. The selection now exists, and these pins hold its + * semantics โ€” the parts a CI step would be built on: + * + * - ANY-OF: a scenario carrying at least one listed tag is selected; + * - exact, case-sensitive matching โ€” a tag is a name, not a pattern; + * - with the flag, an untagged scenario is never selected; + * - WITHOUT the flag (the control), every scenario is selected and `tags` + * is not consulted; + * - a malformed list is refused, never narrowed silently. + * + * The run-level wiring โ€” the flag reaching the selection, the names reaching + * the printed report, the exit statuses โ€” is pinned by spawning the command + * in `qa-names-and-tags-run.test.ts`. + */ + +import { describe, it, expect } from 'vitest'; +import type * as QA from '@objectstack/spec/qa'; +import Test, { + parseTagsFlag, + scenarioLabel, + selectScenariosByTags, + suiteHeading, + tagSelectionLine, +} from '../src/commands/test'; + +const step: QA.TestStep = { name: 'probe', action: { type: 'api_call', target: '/api/v1/health' } }; + +const suite: QA.TestSuite = { + name: 'Tagged suite', + scenarios: [ + { id: 'smoke-only', name: 'Smoke only', tags: ['smoke'], steps: [step] }, + { id: 'regression-only', name: 'Regression only', tags: ['regression'], steps: [step] }, + { id: 'both', name: 'Smoke and critical', tags: ['critical', 'smoke'], steps: [step] }, + { id: 'untagged', name: 'Untagged', steps: [step] }, + { id: 'cased', name: 'Capitalised tag', tags: ['Smoke'], steps: [step] }, + ], +}; + +const ids = (s: QA.TestSuite) => s.scenarios.map((scenario) => scenario.id); + +describe('selectScenariosByTags โ€” the --tags filter', () => { + it('selects a scenario carrying ANY of the listed tags, in authored order', () => { + const { suite: selected, deselected } = selectScenariosByTags(suite, ['smoke', 'regression']); + expect(ids(selected)).toEqual(['smoke-only', 'regression-only', 'both']); + expect(deselected).toBe(2); + }); + + it('matches exactly and case-sensitively, and never selects an untagged scenario', () => { + const { suite: selected, deselected } = selectScenariosByTags(suite, ['smoke']); + expect(ids(selected)).toEqual(['smoke-only', 'both']); + expect(deselected).toBe(3); + }); + + it('reports which requested tags matched, so an unmatched one can be named', () => { + expect(selectScenariosByTags(suite, ['critical', 'smkoe']).matchedTags).toEqual(['critical']); + }); + + it('selects nothing, and says how much it left out, when no scenario carries the tag', () => { + const { suite: selected, deselected, matchedTags } = selectScenariosByTags(suite, ['nomatch']); + expect(ids(selected)).toEqual([]); + expect(deselected).toBe(suite.scenarios.length); + expect(matchedTags).toEqual([]); + }); + + it('CONTROL โ€” without the flag every scenario is selected, untagged ones included', () => { + const { suite: selected, deselected } = selectScenariosByTags(suite, undefined); + expect(ids(selected)).toEqual(ids(suite)); + expect(deselected).toBe(0); + }); + + it('keeps the suite name on the narrowed suite', () => { + expect(selectScenariosByTags(suite, ['smoke']).suite.name).toBe('Tagged suite'); + }); +}); + +describe('parseTagsFlag โ€” the --tags value', () => { + it('splits a comma list, trims each name and drops repeats', () => { + expect(parseTagsFlag('smoke, critical,smoke')).toEqual(['smoke', 'critical']); + }); + + it('refuses an empty entry instead of narrowing the selection silently', () => { + expect(() => parseTagsFlag('smoke,')).toThrow(/empty tag name/); + expect(() => parseTagsFlag('')).toThrow(/empty tag name/); + }); +}); + +describe('report labels โ€” the names the author wrote', () => { + it('prints the scenario name with its id', () => { + expect(scenarioLabel({ scenarioId: 'acct-create', scenarioName: 'An account can be created' })).toBe( + 'An account can be created [acct-create]', + ); + }); + + it('prints the id once when the author made the name the same', () => { + expect(scenarioLabel({ scenarioId: 'acct-create', scenarioName: 'acct-create' })).toBe('acct-create'); + }); + + it('heads a suite with its name and the file it came from', () => { + expect(suiteHeading('Accounts smoke', '/work/qa/accounts.test.json')).toBe( + '๐Ÿ“„ Running suite: Accounts smoke (accounts.test.json)', + ); + }); + + it('states the selection in counts, and that deselected is not passed', () => { + expect(tagSelectionLine(1, 3, ['smoke'])).toBe( + '--tags smoke selected 1 of 4 scenarios; 3 deselected (not run, not counted as passed).', + ); + }); +}); + +describe('os test declares the flag', () => { + it('has a --tags flag and states the any-of semantics in --help', () => { + expect(Test.flags.tags).toBeDefined(); + expect(Test.description).toContain('AT LEAST ONE of the listed tags'); + }); +}); diff --git a/packages/core/src/qa/runner.test.ts b/packages/core/src/qa/runner.test.ts index c9704149a47..77f7697f7be 100644 --- a/packages/core/src/qa/runner.test.ts +++ b/packages/core/src/qa/runner.test.ts @@ -219,3 +219,65 @@ describe('TestRunner โ€” the sibling operators are unchanged by #7256', () => { expect(error).toContain('Unknown assertion operator: not_contains'); }); }); + +// A result used to carry `scenarioId` and nothing else an author wrote: the +// suite's `name` and each scenario's `name` were parsed and read by nothing, so a +// careful human-readable title came back as the terse id in every report. These +// pin the names onto BOTH result envelopes โ€” the completed run and the +// setup-failure early return โ€” because a report prints the names of the +// scenarios that failed at least as often as of the ones that passed. +describe('TestRunner โ€” results carry the suite and scenario names the author wrote', () => { + const step: QA.TestStep = { + name: 'step-1', + action: { type: 'api_call', target: '/api/v1/health' }, + }; + + /** An adapter whose every action throws โ€” drives the setup-failure envelope. */ + class ThrowingAdapter implements TestExecutionAdapter { + async execute(): Promise { + throw new Error('target unreachable'); + } + } + + const suite: QA.TestSuite = { + name: 'Accounts smoke', + scenarios: [ + { + id: 'acct-create', + name: 'An account can be created', + description: 'Fails when the data API refuses a plain insert.', + steps: [step], + }, + { id: 'acct-read', name: 'An account reads back', steps: [step] }, + ], + }; + + it('runSuite stamps suiteName, scenarioName and description on every result', async () => { + const results = await new TestRunner(new StubAdapter({ ok: true })).runSuite(suite); + + expect(results.map((r) => [r.suiteName, r.scenarioId, r.scenarioName, r.description])).toEqual([ + ['Accounts smoke', 'acct-create', 'An account can be created', 'Fails when the data API refuses a plain insert.'], + ['Accounts smoke', 'acct-read', 'An account reads back', undefined], + ]); + expect(results.every((r) => r.passed)).toBe(true); + }); + + it('the setup-failure envelope carries the names too', async () => { + const [result] = await new TestRunner(new ThrowingAdapter()).runSuite({ + name: 'Setup suite', + scenarios: [{ id: 'with-setup', name: 'Setup that cannot run', setup: [step], steps: [step] }], + }); + + expect(result.passed).toBe(false); + expect(String(result.error)).toContain('Setup failed'); + expect(result.suiteName).toBe('Setup suite'); + expect(result.scenarioName).toBe('Setup that cannot run'); + }); + + it('runScenario on a lone scenario names the scenario and no suite', async () => { + const result = await new TestRunner(new StubAdapter({ ok: true })).runScenario(suite.scenarios[1]); + + expect(result.scenarioName).toBe('An account reads back'); + expect(result.suiteName).toBeUndefined(); + }); +}); diff --git a/packages/core/src/qa/runner.ts b/packages/core/src/qa/runner.ts index 2f6795f0b6c..faf69146dca 100644 --- a/packages/core/src/qa/runner.ts +++ b/packages/core/src/qa/runner.ts @@ -3,8 +3,27 @@ import * as QA from '@objectstack/spec/qa'; import { TestExecutionAdapter } from './adapter.js'; +/** + * One scenario's outcome, carrying the names a report leads with. + * + * `scenarioId` is the machine handle; `scenarioName` and `suiteName` are the + * human titles the author wrote (`TestScenario.name`, `TestSuite.name`) โ€” a + * report that printed only the id handed the author back the terse half of + * what they wrote. `description` rides along so a report can say what a + * FAILED scenario was checking without the reader opening the suite file. + */ export interface TestResult { + /** + * `TestSuite.name` of the suite the scenario ran in. Set by `runSuite`; + * absent only when `runScenario` is called on a lone scenario, which has no + * suite to name. + */ + suiteName?: string; scenarioId: string; + /** `TestScenario.name` โ€” the title a report prints for this scenario. */ + scenarioName: string; + /** `TestScenario.description`, when the author wrote one. */ + description?: string; passed: boolean; steps: StepResult[]; error?: unknown; @@ -59,7 +78,7 @@ export class TestRunner { async runSuite(suite: QA.TestSuite): Promise { const results: TestResult[] = []; for (const scenario of suite.scenarios) { - results.push(await this.runScenario(scenario)); + results.push({ suiteName: suite.name, ...(await this.runScenario(scenario)) }); } return results; } @@ -79,6 +98,8 @@ export class TestRunner { } catch (e) { return { scenarioId: scenario.id, + scenarioName: scenario.name, + description: scenario.description, passed: false, steps: [], error: `Setup failed: ${e instanceof Error ? e.message : String(e)}`, @@ -133,6 +154,8 @@ export class TestRunner { return { scenarioId: scenario.id, + scenarioName: scenario.name, + description: scenario.description, passed: scenarioPassed, steps: stepResults, error: scenarioError, diff --git a/packages/spec/liveness/README.md b/packages/spec/liveness/README.md index f846cf7798c..560b83313a8 100644 --- a/packages/spec/liveness/README.md +++ b/packages/spec/liveness/README.md @@ -927,7 +927,7 @@ marker where the Notes cell goes, never a guess at what belongs there. | mapping | seeded 2026-08-01 (#4488) at 8/11 live; **0 dead since #4509** retired the three that were not. The import half (#2611) is loudly enforced โ€” unsupported transforms/formats are 400s, `mode`/`upsertKey` default the request, the wizard picker renders `label`. RETIRED 17.0.0: `extractQuery` (authorWarn โ€” "for export only" promised an export path no exporter implements) + `errorPolicy`/`batchSize`, which were dead AND **unwarnable** (schema defaults materialize at parse, so presence โ‰  authored โ€” `_authorWarnSkipped`, the non-boolean instance of the default(true) rule). That unwarnability is why they went out in the 17.0.0 window rather than after a deprecation cycle: removal was the only channel that could ever reach the author. Rows DELETED, not tombstoned โ€” MappingSchema is strict, so the keys left the walked shape | | seed | seeded 2026-08-01 (#4488). Fully live via SeedLoaderService on both doors (boot/per-org replay + runtime-draft publish). `records` is the z.record walk boundary: the keys an author writes are the target object's fields, governed by that object's own definitions โ€” recorded in the entry, not silently skipped | | translation | seeded 2026-08-01 (#4488) โ€” after fixing the walker: the registered schema is a z.preprocess pipe (#3778 retired-dialect guard) whose transform side the unwrap always took, so the type was literally unwalkable. 11 of 12 groups live across spec resolvers, REST localization, objectui client resolvers and plugin-audit (whose composed-key `t()` calls make `messages` easy to mis-verify as dead) โ€” `flows` was the one that was not, and was `planned` at seeding. **#14253** added the twelfth, `datasets`, seeded LIVE and DRILLED (label / description / dimensions / measures) with its reader in the same change: `translateDataset` in the dispatch table, which is what `TRANSLATABLE_METADATA_TYPES` is derived from, so the REST boundary followed with nothing else to remember. The same change gave `objects.._views..bulkActions` and `objects.._validations..message` their first keys โ€” both beneath the walk boundary, so neither adds a row here. Dead 1 = `validationMessages` (authorWarn) at seeding: nothing resolved it, and #3778's own legacy-key migration table steered `errors:` authors into it โ€” a shipped false signpost, the capabilities.readOnly shape. **#4667**: `validationMessages` REMOVED (row deleted) โ€” removed from the shared translationDataShape(), so it retired at BOTH doors at once, closing the item-only asymmetry #3778's original guard had. #3778's own `errors` guidance was rewritten in the same change: it had been steering authors INTO this dead group. โš ๏ธ **What that left behind is this table's own worked example of the defect it warns about** (#7377): the same commit that deleted the `validationMessages` row wrote a count column of `dead 2` beside a sentence that named exactly one dead key โ€” and that one was the key it had just removed. The real two were `name` and `label`, which the cell never mentioned. Measured at that commit, not inferred: the ledger's dead set there is `{name, label}` and `validationMessages` is absent from `props`. The number was right and the prose was false, in the same cell, on the day it was written โ€” which is why the counts are now generated and this cell holds prose only. **#7131** (PR #7425) resolves it: `name` and `label` re-grade `dead` โ†’ `live` under the designer-previews-count-as-consumers ruling (objectui `TranslationPreview.tsx:67` reads `label` first and falls back to `name`, both rendering at `:100`), so the dead set is empty and there is no dead-set sentence left to keep true. As on `job`, the ADR-0033 docs-shaped exemption is untouched โ€” nothing about enforce-or-remove moved. **#19620** (ruling batch #210 item 2 letter B): `settings` row DELETED โ€” the strict-delete route, because `TranslationItemSchema` no longer declares the key and refuses it by name (the item door now takes the per-app face, as the file door has since #15178). โš ๏ธ The deleted row read `live`, and that verdict was TRUE and stays true of the platform: its evidence read the SERVED tree, which the platform bundle feeds, so the deletion retires the key from the application-authored item and nothing else โ€” the capability lives on `PlatformTranslationDataSchema`, outside this ledger. **#20296**: `flows` is now half-read. `flows.screens` re-graded `planned` โ†’ `live`: objectui's FlowRunner reads each screen's `title` and each field's `label` / `placeholder` at the `.objectui-sha` pin f8a9d0fb. `flows.label` stays `planned` because nothing reads it yet (#20318), and so does the container's `authorWarn`, whose `authorHint` now names the read half and the unread one. | -| qa | seeded 2026-08-10 (#6247) โ€” **not a metadata type**: `TestSuiteSchema` is the FILE surface of the shipped `os test` command (`qa/*.test.json`), governed through the same `SPEC_ONLY_SCHEMAS` override as `query`/`webhook`/`validation`. It is in the table as the clearest worked example of a **false `dead` measurement**: #6247 reported the whole domain declared-but-inert on a grep that scanned only `*Schema` identifiers, and every consumer here reads the **type** names (`QA.TestSuite`, `QA.TestStep`, `QA.TestAction`) โ€” so an entire execution chain (core's `TestRunner` + `HttpTestAdapter`, published via `export * as QA`, driven by a documented CLI command) read as zero consumers, and a retire ruling was issued on it before being withdrawn. The `evidenceScope` table one section up says no amount of specifier matching is sufficient for a negative claim; this is the same lesson for **identifier** matching. What was really wrong was narrower and real: the type was the contract and the schema had no `parse` site, so the CLI's `JSON.parse(content) as QA.TestSuite` cast admitted anything โ€” ENFORCED in the same change (`TestSuiteSchema.safeParse` at the load site, pinned). Dead 5 = `name` (the file name is the suite identity; the CLI prints `path.basename`), `scenarios.name` (describe() says "for test reports"; every report carries `scenarioId` instead), `scenarios.description` (docs-shaped, kept), and the two on the enforce-or-remove worklist โ€” `scenarios.tags` promises filtering that `os test`'s two flags cannot express, and `scenarios.requires` declares param/plugin preconditions nothing checks, so a suite naming a missing plugin runs anyway and fails as an unexplained HTTP error. Neither carries `authorWarn` and the omission is deliberate (`_authorWarnSkipped`): the lint walks stack **collections**, a QA suite is a loose file in no stack, so a warn flag here would emit nothing โ€” a silent no-op inside the mechanism built to catch silent no-ops | +| qa | seeded 2026-08-10 (#6247) โ€” **not a metadata type**: `TestSuiteSchema` is the FILE surface of the shipped `os test` command (`qa/*.test.json`), governed through the same `SPEC_ONLY_SCHEMAS` override as `query`/`webhook`/`validation`. It is in the table as the clearest worked example of a **false `dead` measurement**: #6247 reported the whole domain declared-but-inert on a grep that scanned only `*Schema` identifiers, and every consumer here reads the **type** names (`QA.TestSuite`, `QA.TestStep`, `QA.TestAction`) โ€” so an entire execution chain (core's `TestRunner` + `HttpTestAdapter`, published via `export * as QA`, driven by a documented CLI command) read as zero consumers, and a retire ruling was issued on it before being withdrawn. The `evidenceScope` table one section up says no amount of specifier matching is sufficient for a negative claim; this is the same lesson for **identifier** matching. What was really wrong was narrower and real: the type was the contract and the schema had no `parse` site, so the CLI's `JSON.parse(content) as QA.TestSuite` cast admitted anything โ€” ENFORCED in the same change (`TestSuiteSchema.safeParse` at the load site, pinned). The seeding recorded Dead 5 โ€” `name`, `scenarios.name`, `scenarios.description`, `scenarios.tags` and `scenarios.requires` โ€” and #20289 (family `qa-runner`, verdict ENFORCE) made FOUR of them live: the suite `name` heads the suite in `os test`'s report and is stamped as `suiteName` on every TestResult the suite produces; `scenarios.name` is printed beside the id and carried as `scenarioName`; `scenarios.description` is carried and printed under a failed scenario; and `scenarios.tags` is read by `os test --tags` (any-of, exact; deselected scenarios are counted and never passed). Dead 1 = `scenarios.requires`, which declares param/plugin preconditions nothing checks, so a suite naming a missing plugin runs anyway, and the unmet requirement surfaces only as whatever failure it causes, if any. It was put to the maintainer rather than guessed: over HTTP no server surface lists the loaded plugins, and whose environment `params` names is undefined โ€” its describe() now says NOT CHECKED in the meantime. It does not carry `authorWarn` and the omission is deliberate (`_authorWarnSkipped`): the lint walks stack **collections**, a QA suite is a loose file in no stack, so a warn flag here would emit nothing โ€” a silent no-op inside the mechanism built to catch silent no-ops | | validation | seeded 2026-08-01 (#4488). The ADR-0020 carrier: the evaluator honors active/events/priority/severity/type/condition/message (the zod header's "only reads type/condition/โ€ฆ" prose is STALE โ€” trust the ledger). Dead 3 = label/description/tags, declared governance metadata, kept unmarked. Union walk boundary recorded: only base + `script` keys walked; per-variant keys are governed by the evaluator's tests, not ledger rows. **No longer a registered metadata kind** โ€” #4509 retired it under ADR-0088 (a standalone rule had no object-binding key and every variant is `.strict()`, so it bound to nothing and gated no write; a state machine authored that way saved cleanly and did nothing). The rule VOCABULARY is untouched and fully live via `object.validations[]`, so the ledger keeps governing it through the gate's spec-only override, alongside `webhook` and `query`. The contrast with the two bridges in the same batch is the point: enforce-or-remove picked ENFORCE where the feature existed and only the wiring was missing, and REMOVE where the shape itself could not carry the feature | | api | seeded 2026-08-04 (#5271, part of #5206; PR #5312) โ€” **not a metadata type until that same change made it one**, which is the row's point: governance and registration landed together, the treatment `datasource` did not get (#4487) and paid for with six inert keys found by hand. What #5206 measured before the fix: `api` was in neither `DEFAULT_METADATA_TYPE_REGISTRY` nor `BUILTIN_METADATA_TYPE_SCHEMAS`, so `saveMetaItem`'s `resolveOverlaySchema('api', โ€ฆ)` โ†’ `getMetadataTypeSchema('api')` returned `undefined` and took its own documented branch โ€” an unregistered type is stored **unvalidated** โ€” while `getMetaTypes()` could not enumerate the type at all, so Studio rendered neither list nor form. That issue names the shape precisely and it is the inverse of this ledger's usual one: **enforced but undeclared** (the matcher was already indexing these entries, #5089), where `dead` is declared-but-unenforced. The seeding pass classified 27 keys โ€” live 25 / planned 2 / dead 0 โ€” each cited `file:line` at the consumer layer that reads it: the MATCHER (`packages/metadata/src/endpoint-matcher.ts`) indexes `name`/`path`/`method`; the EXECUTOR (`packages/runtime/src/endpoint-executor.ts`) dispatches on `type` and reads `target`/`objectParams`; the POLICY chain (`packages/runtime/src/endpoint-policy.ts` + `security/inbound-rate-limit.ts`) enforces `authRequired`/`rateLimit`/`cacheTtl`; the MAPPING layer (`packages/runtime/src/api-mapping.ts`) applies `inputMapping`/`outputMapping`; and OpenAPI enrichment (`packages/rest/src/openapi-endpoints.ts`) emits `summary`/`description`. Timing was the reason it was cheap: #5040's E-series had built every one of those consumers and all of it was on main, so each key had a real evidence path rather than a promise. **Planned 2 = `inputMapping.transform` + `outputMapping.transform`, and `planned` rather than `dead` is load-bearing**: `dead` here means parsed with no consumer โ€” a silent no-op โ€” and these are the opposite, parsed and then LOUDLY REFUSED at publish (`endpoint-publish-gate.ts` mappingGate) and again at runtime, because no transformation-function registry exists anywhere in the platform. An author who writes one is told so and told what to do instead, so there is nothing for enforce-or-remove to chase; they stay in the vocabulary because admitting them needs a function registry **and** a sandbox ruling (#5040 ยง3.4), which is a design decision, not a key to quietly delete. Zero dead | | capability | seeded 2026-08-08 (#5961; PR #6540) โ€” `CapabilityDeclarationSchema`, the DECLARATION side of ADR-0066 D1's three-way separation: packages DEFINE a capability, permission sets GRANT it via `systemPermissions`, resources REQUIRE it via `requiredPermissions`. **The gate's 12 and the seeding PR's 5 are the same measurement at two granularities** โ€” PR #6540 call-graph-closed **5 authorable properties**, every one to a real reader in `packages/plugins/plugin-security/src/bootstrap-declared-capabilities.ts` (the one consumer that turns a declaration into a `sys_capability` row), all `live`, with no `PENDING_GOVERNANCE` debt recorded; the other 7 are the ADR-0010 protection-envelope keys the gate auto-classifies `live` and which carry `null` verdicts in the file, exactly as on `permission`/`position`. The same worked example as `api` above and PR #6540 says so in those words โ€” **enforced but undeclared**, the mirror of the hole #5271 closed. What #5961 measured: absent from `DEFAULT_METADATA_TYPE_REGISTRY`, `BUILTIN_METADATA_TYPE_SCHEMAS` and `HAND_CRAFTED_SCHEMAS`, so `isRuntimeCreateAllowed()` took its no-static-entry fallback (permanently true) and `saveMetaItem` its no-schema branch โ€” `PUT /api/v1/meta/capability/:name` accepted **arbitrary JSON** onto an authorization surface whose names `systemPermissions`/`requiredPermissions` resolve by string, while `/meta/types` synthesised a false `allowRuntimeCreate: true` descriptor Studio drew a raw-JSON create form from. #5870 did not open that path (the write gate reads the registry, not the item store); it only made the type visible in `getMetaTypes()`, and both the issue and this row say so to stop the next reader filing it as a regression. Landed as ruling A on ADR-0066 D1's own authority: `allowRuntimeCreate: false` **and** `allowOrgOverride: false`, the second self-judged inside the ruling's rationale and flagged for veto โ€” a tenant overlay of a package declaration would lift `scope` from `org` to `platform`, which is the one field on this type that is an escalation rather than display. Its reverse verification is worth copying: deleting the registry entry gave 7 red / 3 green and measured something **sharper than predicted** โ€” a garbage payload turned 422 rather than resolving, i.e. the schema binding is a real second line of defence behind the registry row, not a restatement of it; deleting the schema binding alone gave exactly 3 red. `packageId` is the one key that reads oddly: deliberately a FALLBACK, not the primary, since #5870 added `capabilities` to the ObjectQL stamped-collection list so `_packageId` now reaches a declaration and wins โ€” it stays `live` because the fallback branch still decides materialization for any declaration arriving unstamped. Zero dead | diff --git a/packages/spec/liveness/qa.json b/packages/spec/liveness/qa.json index 239b2d6d0a5..cc9c16e037f 100644 --- a/packages/spec/liveness/qa.json +++ b/packages/spec/liveness/qa.json @@ -3,10 +3,11 @@ "_note": "TestSuiteSchema (packages/spec/src/qa/testing.zod.ts) โ€” the Quality Protocol file surface: an author writes `qa/*.test.json`, `os test` loads it, and core's TestRunner executes it. Seeded 2026-08-10 (#6247), the ENFORCE leg of an enforce-or-remove call that was ruled the other way first and then withdrawn, which is the lesson worth keeping. #6247 filed this domain as declared-but-inert on a grep that scanned only `*Schema` identifiers; every consumer here reads the TYPE names (`QA.TestSuite`, `QA.TestScenario`, `QA.TestStep`, `QA.TestAction`, `QA.TestAssertion`), so the search matched nothing and a complete execution chain read as zero consumers. The 2026-08-07 retire ruling rested on that reading and was WITHDRAWN on 2026-08-08 (issue comment 5225532429) once the sweep's pre-flight gate falsified it. A schema with no `parse` site is not the same finding as a schema with no consumer, and only the first one was true. The measured chain, by layer: the RUNNER (packages/core/src/qa/runner.ts) reads suite.scenarios, scenario.id/setup/steps/teardown and step.name/action/capture/assertions; the ADAPTER (packages/core/src/qa/http-adapter.ts) switches on action.type โ€” its case labels ARE the TestActionTypeSchema values โ€” and reads target/payload/user; both are published through packages/core/src/index.ts:25 (`export * as QA`); the driving entry point is the shipped oclif command `os test` (packages/cli/src/commands/test.ts), documented at content/docs/deployment/cli.mdx:987,1012-1020 and packages/cli/README.md:104. WALK BOUNDARY, recorded rather than silently skipped: the gate classifies one level and this file drills `scenarios` one more, so the verdicts here cover the suite and scenario levels only. Step / action / assertion keys sit BELOW the walk; they were measured in the same pass and their verdicts are recorded in the `setup`/`steps`/`teardown` notes instead of being fanned out into rows the gate would not check. AUTHOR-WARN CHANNEL: none exists for this type, and no entry is marked `authorWarn` for that reason (`_authorWarnSkipped`). The CLI lint (packages/lint/src/lint-liveness-properties.ts) walks stack COLLECTIONS โ€” `stack.flows`, `stack.views`, โ€ฆ โ€” and a QA suite is not part of a stack at all; it is a loose JSON file `os test` globs off disk. Marking an entry `authorWarn` here would produce a warning nothing can emit, which is the same silent no-op this ledger exists to catch, so the dead entries below carry their correction in `note` and the load-site parse (below) is what actually reaches the author. LOAD-SITE ENFORCEMENT: the same change that seeded this file replaced the CLI's `JSON.parse(content) as QA.TestSuite` cast โ€” the schema author's own `// Should validate with Zod` TODO โ€” with a real `TestSuiteSchema.safeParse`, so a malformed suite is named at load time instead of reaching the runner as a lie about its own shape. That is what makes the `live` rows below enforced rather than merely read. No `qa` property is a bound HIGH_RISK class in proof-registry.mts, so no entry carries a `proof`; none is invented to look thorough. 2026-08-28 (#13003): the four `path:NNN` citations in this file were re-anchored to `runScenario`, the method that reads all four keys. All four were wrong and all four IN RANGE โ€” they moved as one block when a diagnostic helper was added at the head of the file. The many bare `:NNN` suffixes in the sub-key notes below are prose, not citations: no check has ever resolved them, and they are left as the hand-measured record they are.", "props": { "name": { - "status": "dead", - "verifiedAt": "2026-08-10", + "status": "live", + "evidence": "packages/core/src/qa/runner.ts#runSuite (`suiteName: suite.name` stamped on every TestResult the suite produces) ยท packages/cli/src/commands/test.ts#suiteHeading (the suite heading `os test` prints: the name, then the file it was loaded from)", + "verifiedAt": "2026-09-27", "evidenceScope": "in-repo", - "note": "Parsed and never read. `TestRunner.runSuite` (packages/core/src/qa/runner.ts:25-31) touches only `suite.scenarios`, and the one caller prints `path.basename(file)` as the suite heading (packages/cli/src/commands/test.ts:95) โ€” so renaming a suite changes nothing an operator sees. Kept, not retired: it is display-shaped metadata whose describe() promises no capability ('Test suite name'), and it is the natural title the moment anything reports per-suite results. The honest reading is that the FILE NAME is the suite identity today. Recorded so the next reader does not have to re-derive that, and so that wiring it up counts as making a dead key live rather than as a no-op refactor." + "note": "The author is the producer โ€” both readers take the authored value directly, so no `producer` is owed. Was `dead` from the 2026-08-10 seeding until #20289: the suite heading was `path.basename(file)` and no result carried the name, so renaming a suite changed nothing an operator saw. Now the heading leads with the name and keeps the file beside it (the file is still where you go to fix the suite), and a suite whose file fails to load โ€” where no name was ever parsed โ€” is headed by the file alone. Measured before and after on a stub target: the heading read `probe.test.json`, and reads `Probe suite: names, tags and requires (probe.test.json)`." }, "scenarios": { "children": { @@ -15,25 +16,29 @@ "evidence": "packages/core/src/qa/runner.ts#runScenario (`scenarioId: scenario.id` on BOTH result envelopes โ€” the setup-failure path and the completed run)", "verifiedAt": "2026-08-28", "evidenceScope": "in-repo", - "note": "The scenario identity in every result the CLI prints (packages/cli/src/commands/test.ts:104) and the ONLY human-readable handle a failing run gives you โ€” `name` is not printed anywhere. Not deduplicated: two scenarios may declare the same id and both run, so an id collision shows up as two indistinguishable result lines rather than an error. 2026-08-28: RE-ANCHORED (#13003) and REPOINTED โ€” `:101` had rotted onto `stepName: step.name` in the per-step result, and the sibling `:47` was a bare line suffix with no path in front of it, which the scanner cannot resolve at all. Both id reads live in `runScenario`. Re-closed by hand against 8cb96ec41." + "note": "The scenario identity in every result the CLI prints, in brackets after the scenario's `name` (packages/cli/src/commands/test.ts#scenarioLabel). Until #20289 it was the ONLY handle a report gave โ€” `name` was printed nowhere. Not deduplicated: two scenarios may declare the same id and both run, so an id collision shows up as two indistinguishable result lines rather than an error. 2026-08-28: RE-ANCHORED (#13003) and REPOINTED โ€” `:101` had rotted onto `stepName: step.name` in the per-step result, and the sibling `:47` was a bare line suffix with no path in front of it, which the scanner cannot resolve at all. Both id reads live in `runScenario`. Re-closed by hand against 8cb96ec41." }, "name": { - "status": "dead", - "verifiedAt": "2026-08-10", + "status": "live", + "evidence": "packages/core/src/qa/runner.ts#runScenario (`scenarioName: scenario.name` on BOTH result envelopes โ€” the setup-failure path and the completed run) ยท packages/cli/src/commands/test.ts#scenarioLabel (every scenario line prints the name, then the id in brackets; the id once when the author made them equal)", + "verifiedAt": "2026-09-27", "evidenceScope": "in-repo", - "note": "Its describe() says 'Scenario name for test reports' and no report carries it: `TestResult` (packages/core/src/qa/runner.ts:6-12) has `scenarioId` and no name field, and the CLI prints `result.scenarioId` (packages/cli/src/commands/test.ts:104). So an author who writes a careful human-readable `name` and a terse `id` gets the terse one in every failure message. Mildly misleading rather than benign โ€” the describe() names an output that does not exist โ€” but it is a display key with an obvious enforcement route (print it beside the id), which is why the note carries the correction instead of a retirement." + "note": "The author is the producer. Was `dead` from the 2026-08-10 seeding until #20289: `TestResult` had `scenarioId` and no name field, so an author who wrote a careful human-readable `name` and a terse `id` got the terse one in every report. Its describe() ('Scenario name for test reports') is now true as written and was left unchanged. The id stays on the line: it is still the one unique handle, and the platform checklist's green-path clause reads it." }, "description": { - "status": "dead", - "verifiedAt": "2026-08-10", + "status": "live", + "evidence": "packages/core/src/qa/runner.ts#runScenario (`description: scenario.description` on both result envelopes) ยท packages/cli/src/commands/test.ts#run (printed under a FAILED scenario's line, before its error)", + "verifiedAt": "2026-09-27", "evidenceScope": "in-repo", - "note": "Docs-shaped annotation, no consumer โ€” the same disposition `flow.description` and `hook.label`/`description` carry and for the same reason: an author is not misled by a field that only claims to describe. Exempt from enforce-or-remove (ADR-0033)." + "note": "The author is the producer. Was `dead` (docs-shaped, exempt from enforce-or-remove under ADR-0033) until #20289 gave it a reader: a failed scenario's report now carries what the scenario was checking, in the author's words. Printed on failure only โ€” a passing scenario's description is context nobody needs mid-run. The one suite in this repo (examples/app-showcase/qa/platform-smoke.test.json) already writes its descriptions as failure-diagnosis text (`a failure here means the --url is wrong or the server is down`), which is the use this reader serves." }, "tags": { - "status": "dead", - "verifiedAt": "2026-08-10", + "status": "live", + "evidence": "packages/cli/src/commands/test.ts#selectScenariosByTags (ANY-OF, exact and case-sensitive: a scenario runs when it carries at least one tag named by `--tags`; the rest are deselected โ€” not run, counted in the summary, never counted as passed)", + "producer": "packages/cli/src/commands/test.ts#parseTagsFlag โ€” the `--tags` value (a comma list; an empty entry is refused at parse time). Without the flag the selection is not consulted and every scenario runs, so the read depends on the operator supplying it.", + "verifiedAt": "2026-09-27", "evidenceScope": "in-repo", - "note": "The sharpest row in this file. Its describe() promises 'Tags for filtering and categorization (e.g. \"critical\", \"regression\", \"crm\")' and NOTHING filters on it: `os test` has three flags โ€” `--url`, `--token` and `--fail-on-empty` (#7848) โ€” and not one of them selects scenarios, the runner never reads `scenario.tags`, and the only selection the command offers is the file glob. So `os test --tags critical` is not a narrower run, it is an unknown-flag error, and a suite tagged `regression` runs on every invocation. This is the entry that would carry `authorWarn` if the type had a channel for one (see `_authorWarnSkipped` in the file note): an author tagging scenarios is buying a filter that does not exist. Enforce-or-remove worklist โ€” the enforce route is a `--tags` filter in the command, the remove route drops the key; either is a decision, not a cleanup." + "note": "Was `dead` and 'the sharpest row in this file' from the 2026-08-10 seeding until #20289: its describe() promised filtering and `os test --tags critical` was an unknown-flag error. Any-of is the comma-list reading of Odoo's `--test-tags` and Cucumber's comma form, and Playwright's everyday `--grep @a|@b`; an expression language was not taken because no suite has needed one. A selection that matches no scenario takes the empty-pattern posture โ€” exit 0, and 1 under `--fail-on-empty` โ€” and a requested tag no loaded scenario carries is named on the summary. Measured before and after on a stub target: `--tags smoke` exited 2 with `Nonexistent flag: --tags`, and now runs the one smoke scenario and reports the others deselected." }, "setup": { "status": "live", @@ -58,9 +63,9 @@ }, "requires": { "status": "dead", - "verifiedAt": "2026-08-10", + "verifiedAt": "2026-09-27", "evidenceScope": "in-repo", - "note": "Declared environment preconditions โ€” `requires.params` (environment variables) and `requires.plugins` (plugins that must be loaded) โ€” that nothing checks. Neither the runner nor the CLI reads `scenario.requires`; there is no skip path and no precondition failure in the code at all, so a suite that declares `plugins: ['plugin-sharing']` runs unchanged against a server without it and fails later as an unexplained HTTP error. That is the misleading direction โ€” the author reads it as a guard and gets none โ€” so it belongs on the enforce-or-remove worklist beside `tags`, not in the docs-shaped bucket with `description`. Enforce route: check requirements before running and report the scenario as skipped-with-reason. Remove route: drop the block, since a precondition that is never checked is worse than an absent one." + "note": "Declared environment preconditions โ€” `requires.params` (environment variables or parameters) and `requires.plugins` (plugins that must be loaded) โ€” that nothing checks. Re-measured 2026-09-27 (#20289) on a stub target: a scenario declaring a plugin that does not exist and an environment variable that is unset reported PASSED, exactly like its no-requirements control. STILL DEAD, and deliberately: #20289 enforced the other four rows of this family and put this one to the maintainer, because neither sub-key has an honest judge yet. `plugins`: `os test` reaches the target over HTTP, and no server surface lists the loaded plugins โ€” the discovery document advertises `services` and `capabilities`, not plugins โ€” and the plugin spelling itself (package name or `plugin.name`) is undefined. `params`: whose environment is undefined โ€” the runner's is observable but no step can consume a runner variable (interpolation reads only captured context), and the server's is not observable at all. Guessing either would make a skip that means nothing. What changed now: the describe() of `requires` and both sub-keys says NOT CHECKED instead of reading as a guard, so an author (or an AI) is no longer told it is one. The enforce route still needs a SKIPPED result with its reason, counted separately and never passed; the remove route drops the block." } } } diff --git a/packages/spec/liveness/state-counts.md b/packages/spec/liveness/state-counts.md index 54ac2f5bb92..483ad0bfadf 100644 --- a/packages/spec/liveness/state-counts.md +++ b/packages/spec/liveness/state-counts.md @@ -56,7 +56,7 @@ for both corollaries. | `validation` | 18 | 0 | 0 | 0 | 0 | 18 | | `api` | 25 | 0 | 0 | 1 | 2 | 28 | | `capability` | 12 | 0 | 0 | 0 | 0 | 12 | -| `qa` | 4 | 0 | 0 | 5 | 0 | 9 | +| `qa` | 8 | 0 | 0 | 1 | 0 | 9 | | `manifest` | 23 | 0 | 1 | 15 | 0 | 39 | | `crud_endpoints` | 6 | 0 | 0 | 2 | 0 | 8 | | `metadata_endpoints` | 7 | 0 | 0 | 2 | 0 | 9 | @@ -67,4 +67,4 @@ for both corollaries. | `sharing_rule` | 16 | 0 | 0 | 0 | 1 | 17 | | `connector` | 29 | 0 | 0 | 44 | 1 | 74 | | `analytics_cube` | 17 | 0 | 0 | 10 | 0 | 27 | -| **total** | **936** | **5** | **1** | **166** | **9** | **1117** | +| **total** | **940** | **5** | **1** | **162** | **9** | **1117** | diff --git a/packages/spec/scripts/liveness/check-liveness.mts b/packages/spec/scripts/liveness/check-liveness.mts index 4ec7afb6781..6bd1d578df8 100644 --- a/packages/spec/scripts/liveness/check-liveness.mts +++ b/packages/spec/scripts/liveness/check-liveness.mts @@ -364,10 +364,12 @@ const PENDING_GOVERNANCE: Record = {}; // so the grep missed an entire execution chain and the 2026-08-07 retire ruling // was withdrawn on 2026-08-08 in favour of enforce. Governing it here is the // half of that ruling that keeps the surface honest going forward: the runner -// reads a real, measured subset of the declared keys, and the ones nothing reads -// (`suite.name`, `scenario.tags`, `scenario.requires`) are now recorded as such -// instead of being invisible. Like `query`, there is no registry to fold it back -// onto โ€” the override IS its governance. +// reads a real, measured subset of the declared keys, and a key nothing reads is +// recorded as such instead of being invisible. `suite.name` and `scenario.tags` +// were two of those and have since gained readers in `os test` (the suite +// heading; the `--tags` selection); `scenario.requires` is the one still unread โ€” +// declared, NOT CHECKED, and its row stays dead. Like `query`, there is no +// registry to fold it back onto โ€” the override IS its governance. // `manifest` is the THIRD category the override has had to reach, and the one // that showed the escape hatch was load-bearing rather than a webhook special // case. `ManifestSchema` (src/kernel/manifest.zod.ts) is what an author writes diff --git a/packages/spec/src/qa/testing.zod.ts b/packages/spec/src/qa/testing.zod.ts index 198a4236e9c..850fb87d2cc 100644 --- a/packages/spec/src/qa/testing.zod.ts +++ b/packages/spec/src/qa/testing.zod.ts @@ -68,11 +68,11 @@ export const TestScenarioSchema = lazySchema(() => z.object({ steps: z.array(TestStepSchema).describe('Main test sequence to execute'), teardown: z.array(TestStepSchema).optional().describe('Steps to cleanup after test execution'), - // Environment requirements + // Environment requirements โ€” declared, not yet checked by any reader. requires: z.object({ - params: z.array(z.string()).optional().describe('Required environment variables or parameters'), - plugins: z.array(z.string()).optional().describe('Required plugins that must be loaded') - }).optional().describe('Environment requirements for this scenario') + params: z.array(z.string()).optional().describe('Environment variables or parameters the scenario needs. Declared only: nothing checks them before the scenario runs'), + plugins: z.array(z.string()).optional().describe('Plugins the scenario needs loaded on the target. Declared only: nothing checks them before the scenario runs') + }).optional().describe('Environment requirements for this scenario. NOT CHECKED by `os test` or the core TestRunner: the scenario runs whether or not they hold, and an unmet requirement surfaces only as the failure it causes') }).describe('A complete test scenario with setup, execution steps, and teardown')); export const TestSuiteSchema = lazySchema(() => z.object({