Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions .changeset/20282-analytics-cube-format-granularities-enforced.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
---
'@objectstack/spec': minor
'@objectstack/service-analytics': minor
---

An authored analytics cube's measure `format` and time-dimension `granularities` now take effect on the analytics query doors, the way a compiled dataset's always have (#20282).

Clause-②: yes (narrowing)

<!-- adr-0087: registered analytics-cube-single-granularity-default-enforced -->

**BREAKING**: this narrows what `POST /api/v1/analytics/query` and `POST /api/v1/analytics/sql` answer for one class of request. When an authored cube's time dimension declares exactly one granularity, a query that groups by that dimension without stating a granularity is now bucketed at the declared one. The raw-SQL path declines every bucketed query, so such a query now runs on the engine aggregate path, which answers `400 INVALID_FIELD` for every member it cannot evaluate: a custom-SQL measure (a measure of type `number`, `string` or `boolean` whose `sql` is an expression); and, on a cube whose members resolve through its `joins`, a measure or a `where` field over a joined object, a `timeDimensions` entry over a joined object (bucketed or a `dateRange` window, so grouping by a one-granularity time dimension over a joined object is refused too), a dimension that traverses more than one relationship, and an `avg` or `count_distinct` measure beside any dimension over a joined object. The raw-SQL path answers every one of these, with one group per distinct timestamp; each is now refused, exactly as it already was when the caller stated that granularity by hand. On a host that overrides `queryCapabilities` to offer raw SQL with no engine aggregate bridge (the plugin's default wires both), no strategy remains for a bucketed query, so every newly bucketed query, a plain `count` included, now answers "No strategy can handle query" instead of grouping raw timestamps. The remedy: run such a query without grouping by that dimension, or, if the dimension is not meant to have one default bucket, declare the granularities it offers as a list of two or more (or omit the key); on a raw-SQL-only host, add the engine aggregate bridge. It ships as `minor` under the launch-window convention; the widening half is two authored keys taking effect.

Until this change both keys were read on the compiled-dataset path only. One cube shape has three producers — cubes authored with `defineCube()` / `defineStack({ analyticsCubes })`, cubes the dataset compiler mints, and cubes inferred for an ad-hoc query — and only a compiled dataset's cube reached the two readers:

- **`measures.<metric>.format`** reached a caller as `fields[].format` only because the dataset door copies it from the DATASET measure. An authored cube has no dataset, so `POST /api/v1/analytics/query` described its measure columns with `name` and `type` alone. Now every measure column a query names carries the `format` its cube measure declares, whichever strategy answered, and a column that declares none carries no `format` key at all. `GET /api/v1/analytics/meta` is unchanged: its projection stays `name`, `type` and `title`, and a client reads `format` off the query result's `fields[]`, as the Data API page already says. The value is relayed verbatim; the vocabulary `fields[].format` documents is a numeral pattern such as `"$0,0.00"` or `"0.0%"`.
- **`dimensions.<dimension>.granularities`** was the default bucket only for a compiled dataset, which the dataset executor filled in before querying. An authored cube's time dimension grouped raw timestamps whatever it declared. Now `query()` and the `generateSql()` dry run read it the same way, through the one rule both paths share: a single-entry list is the dimension's default bucket for a query that groups by it; a granularity the query states always wins, and one the list does not name is not refused (the dataset path compares against no list either); a list of two or more states no default; and a `timeDimensions` entry that carries only a `dateRange` for a dimension the query does not group stays a filter.

What to expect after upgrading:

- **A cube measure that declares `format`** now carries it on `POST /api/v1/analytics/query` results. A client that formats amounts from `fields[].format` starts formatting that column.
- **A cube time dimension that declares one granularity** (`granularities: ['month']`) is now bucketed by it when a query groups by it without stating one: one row per month where there was one row per timestamp. Name another granularity in the query's `timeDimensions` to bucket differently.
- **A cube time dimension that declares several, or none**, behaves exactly as before.
- **Compiled datasets** (`POST /api/v1/analytics/dataset/query`) answer exactly as before: the value read off their cube is the one the dataset door already used.

In `@objectstack/spec`, `MetricSchema.format` and `DimensionSchema.granularities` now carry descriptions that state what the analytics service does with them (the metric's example values move from the names "currency" / "percent" to numeral patterns, the vocabulary the `fields[].format` slot documents), and the liveness ledger rows `analytics_cube.measures.format` and `analytics_cube.dimensions.granularities` move from `dead` to `live`, citing the new readers.
8 changes: 4 additions & 4 deletions content/docs/references/data/analytics.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -148,7 +148,7 @@ Type: `[string, string]`
| **description** | `string` | optional | |
| **type** | `Enum<'count' \| 'sum' \| 'avg' \| 'min' \| 'max' \| 'count_distinct' \| 'number' \| 'string' \| 'boolean'>` | ✅ | |
| **sql** | `string` | ✅ | SQL expression or field reference |
| **format** | `string` | optional | |
| **format** | `string` | optional | Display format for this measure's result column: a numeral pattern such as "$0,0.00" or "0.0%". Relayed verbatim as fields[].format on POST /analytics/query results; not published by GET /analytics/meta. |

### Nested Shape: `Cube.dimensions[string]`

Expand All @@ -159,7 +159,7 @@ Type: `[string, string]`
| **description** | `string` | optional | |
| **type** | `Enum<'string' \| 'number' \| 'boolean' \| 'time' \| 'geo'>` | ✅ | |
| **sql** | `string` | ✅ | SQL expression or column reference |
| **granularities** | `Enum<'day' \| 'week' \| 'month' \| 'quarter' \| 'year'>[]` | optional | |
| **granularities** | `Enum<'day' \| 'week' \| 'month' \| 'quarter' \| 'year'>[]` | optional | For a time dimension. A single interval is its default bucket: a query that groups by this dimension without stating a granularity is bucketed at it. Two or more intervals state no default. A granularity the query states always wins, listed or not. |

### Nested Shape: `Cube.joins[string]`

Expand Down Expand Up @@ -199,7 +199,7 @@ Type: `[string, string]`
| **description** | `string` | optional | |
| **type** | `Enum<'string' \| 'number' \| 'boolean' \| 'time' \| 'geo'>` | ✅ | |
| **sql** | `string` | ✅ | SQL expression or column reference |
| **granularities** | `Enum<'day' \| 'week' \| 'month' \| 'quarter' \| 'year'>[]` | optional | |
| **granularities** | `Enum<'day' \| 'week' \| 'month' \| 'quarter' \| 'year'>[]` | optional | For a time dimension. A single interval is its default bucket: a query that groups by this dimension without stating a granularity is bucketed at it. Two or more intervals state no default. A granularity the query states always wins, listed or not. |


---
Expand Down Expand Up @@ -228,7 +228,7 @@ Type: `[string, string]`
| **description** | `string` | optional | |
| **type** | `Enum<'count' \| 'sum' \| 'avg' \| 'min' \| 'max' \| 'count_distinct' \| 'number' \| 'string' \| 'boolean'>` | ✅ | |
| **sql** | `string` | ✅ | SQL expression or field reference |
| **format** | `string` | optional | |
| **format** | `string` | optional | Display format for this measure's result column: a numeral pattern such as "$0,0.00" or "0.0%". Relayed verbatim as fields[].format on POST /analytics/query results; not published by GET /analytics/meta. |


---
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,154 @@
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.

/**
* `analytics_cube.measures.format` and `analytics_cube.dimensions.granularities`
* over the wire: an AUTHORED cube's two keys reach `POST /api/v1/analytics/query`
* and `POST /api/v1/analytics/sql` through the real dispatcher route.
*
* The service door is pinned beside the implementation
* (`service-analytics` `cube-authored-format-granularity.test.ts`, each case
* against a compiled-dataset control). This file asks the question that pin
* cannot: does what the service answers survive the route — `fields[].format`
* is a member `deps.success()` relays verbatim, and the dry-run door serves
* the bucketed statement `query()` runs.
*
* The cube is handed to the service as `AnalyticsServiceConfig.cubes`, the
* config key the CLI threads an app's `analyticsCubes` into — the authoring
* door, not a registered dataset.
*/

import { describe, it, expect } from 'vitest';
import { CubeSchema } from '@objectstack/spec/data';
import { AnalyticsService } from '@objectstack/service-analytics';

import { createDispatcherPlugin } from './dispatcher-plugin.js';

// ── harness (the shape `analytics-query-read-scope-withhold.test.ts` uses) ────

type Handler = (req: unknown, res: unknown) => unknown;

function makeFakeServer() {
const handlers: Record<string, Handler> = {};
const rec = (verb: string) => (path: string, handler: Handler) => {
handlers[`${verb} ${path}`] = handler;
};
return {
handlers,
server: { get: rec('GET'), post: rec('POST'), put: rec('PUT'), delete: rec('DELETE'), patch: rec('PATCH') },
};
}

function makeCtx(fakeServer: unknown, analytics: unknown) {
const kernel = {
getService: (name: string) => (name === 'analytics' ? analytics : undefined),
getServiceAsync: async (name: string) => (name === 'analytics' ? analytics : undefined),
};
return {
getKernel: () => kernel,
getService: (name: string) => (name === 'http.server' ? fakeServer : undefined),
environmentId: undefined,
logger: { info() {}, warn() {}, error() {}, debug() {} },
hook: () => {},
on: () => {},
} as any;
}

function makeRes() {
const res: any = {
statusCode: undefined as number | undefined,
body: undefined as any,
status(c: number) { res.statusCode = c; return res; },
header() { return res; },
json(b: unknown) { res.body = b; return res; },
};
return res;
}

/** Drive the REAL `POST /api/v1/analytics/<sub>` route against `analytics`. */
async function post(analytics: unknown, sub: 'query' | 'sql', body: unknown) {
const { server, handlers } = makeFakeServer();
const plugin = createDispatcherPlugin({ prefix: '/api/v1', securityHeaders: false });
await plugin.start?.(makeCtx(server, analytics));
const handler = handlers[`POST /api/v1/analytics/${sub}`];
expect(handler, `POST /api/v1/analytics/${sub} must be mounted`).toBeTypeOf('function');
const res = makeRes();
await handler({ body, query: {} }, res);
return res;
}

const silent = { debug() {}, info() {}, warn() {}, error() {} };

/** Parsed the way `defineCube()` and `defineStack({ analyticsCubes })` parse an authored cube. */
const orders = CubeSchema.parse({
name: 'orders',
sql: 'shop_order',
measures: {
count: { label: 'Orders', type: 'count', sql: '*' },
revenue: { label: 'Revenue', type: 'sum', sql: 'amount', format: '$0,0.00' },
},
dimensions: {
status: { label: 'Status', type: 'string', sql: 'status' },
placed_at: { label: 'Placed', type: 'time', sql: 'placed_at', granularities: ['month'] },
},
});

type GroupByItem = string | { field: string; dateGranularity?: string };

/** The composition `AnalyticsServicePlugin` wires by default: both strategies. */
function analytics() {
const groupBys: GroupByItem[][] = [];
const service = new AnalyticsService({
logger: silent,
cubes: [orders],
queryCapabilities: () => ({ nativeSql: true, objectqlAggregate: true, inMemory: false }),
executeRawSql: async () => [{ status: 'open', count: 2, revenue: 10 }],
executeAggregate: async (_object, options) => {
groupBys.push((options.groupBy ?? []) as GroupByItem[]);
return [{ placed_at: '2026-07', count: 2 }];
},
});
return { service, groupBys };
}

describe('POST /analytics/query — an authored cube measure\'s `format` reaches `fields[]`', () => {
it('the measure column carries the declared format; an undeclared one and a dimension carry none', async () => {
const res = await post(analytics().service, 'query', {
cube: 'orders',
measures: ['orders.revenue', 'orders.count'],
dimensions: ['orders.status'],
});

expect(res.statusCode).toBe(200);
const fields = res.body.data.fields as Array<{ name: string; format?: string }>;
expect(fields.find((f) => f.name === 'orders.revenue')?.format).toBe('$0,0.00');
expect(fields.find((f) => f.name === 'orders.count')).not.toHaveProperty('format');
expect(fields.find((f) => f.name === 'orders.status')).not.toHaveProperty('format');
});
});

describe('an authored time dimension\'s single declared granularity is its default bucket over the wire', () => {
it('POST /analytics/query groups the dimension at the declared granularity', async () => {
const { service, groupBys } = analytics();

const res = await post(service, 'query', { cube: 'orders', measures: ['count'], dimensions: ['placed_at'] });

expect(res.statusCode).toBe(200);
expect(groupBys).toEqual([[{ field: 'placed_at', dateGranularity: 'month' }]]);
expect(res.body.data.rows).toEqual([{ placed_at: '2026-07', count: 2 }]);
});

it('POST /analytics/sql dry-runs the bucketed statement, and a stated granularity still wins', async () => {
const declared = await post(analytics().service, 'sql', { cube: 'orders', measures: ['count'], dimensions: ['placed_at'] });
const stated = await post(analytics().service, 'sql', {
cube: 'orders',
measures: ['count'],
dimensions: ['placed_at'],
timeDimensions: [{ dimension: 'placed_at', granularity: 'year' }],
});

expect(declared.statusCode).toBe(200);
expect(declared.body.data.sql).toMatch(/date_trunc\('month'/i);
expect(stated.statusCode).toBe(200);
expect(stated.body.data.sql).toMatch(/date_trunc\('year'/i);
});
});
Loading
Loading