From c7adc9a2364a2d6ebc2ec41201cac3d91a38778e Mon Sep 17 00:00:00 2001 From: -z <11802769+Aloento@users.noreply.github.com> Date: Wed, 23 Sep 2026 22:48:22 +0200 Subject: [PATCH 1/2] docs: document the Zitadel service identity of the reporter The reporter no longer signs an HMAC JWT with a shared secret, it obtains a Zitadel access token for a machine user through the JWT Profile flow. - rewrite doc/api/authentication.md around the key file, the assertion and the token endpoint, and name the scopes a deployment has to request - document the required oidc_issuer, oidc_key_file and oidc_scopes keys, drop the status_dashboard.secret examples and regenerate the config schema - follow the rename through the reporter, config, deployment, troubleshooting and architecture pages --- doc/api/authentication.md | 280 +++++++++++++++-------- doc/architecture/data-flow.md | 6 +- doc/architecture/diagrams.md | 14 +- doc/architecture/overview.md | 13 +- doc/config.md | 34 ++- doc/configuration/examples.md | 7 +- doc/configuration/overview.md | 4 +- doc/configuration/schema.md | 13 +- doc/getting-started/development.md | 6 + doc/getting-started/project-structure.md | 2 +- doc/guides/deployment.md | 46 ++-- doc/guides/troubleshooting.md | 8 +- doc/reporter.md | 58 +++-- doc/schemas/config-schema.json | 21 +- 14 files changed, 342 insertions(+), 170 deletions(-) diff --git a/doc/api/authentication.md b/doc/api/authentication.md index a6c5fed..cce9099 100644 --- a/doc/api/authentication.md +++ b/doc/api/authentication.md @@ -1,153 +1,231 @@ -# Authentication - -This document describes the authentication mechanism used for integrating with the CloudMon Status Dashboard. +# Status Dashboard Authentication ## Overview -The metrics-processor uses JWT (JSON Web Token) authentication when reporting component status to the status-dashboard API. This is specifically used by the `cloudmon-metrics-reporter` component to securely communicate health status updates. +The `cloudmon-metrics-reporter` authenticates against the Status Dashboard with a Zitadel OIDC +service identity. MP does not implement the OAuth flow itself: the JWT Profile exchange is delegated +to the community [`zitadel`](https://crates.io/crates/zitadel) crate (`zitadel::credentials`). -## JWT Token Mechanism +The reporter loads the Zitadel machine user key file once at startup. For every report the crate -### Token Generation +1. discovers the token endpoint from `{issuer}/.well-known/openid-configuration`, +2. signs a short-lived JWT Profile assertion (RS256) with the private key of the key file, +3. exchanges it at the discovered token endpoint with the + `urn:ietf:params:oauth:grant-type:jwt-bearer` grant, -JWT tokens are generated using the HMAC-SHA256 algorithm with a shared secret key. +and MP only wraps the returned access token into an `Authorization: Bearer ` header. +No client secret is involved anywhere, and the Status Dashboard verifies the issued token against +the Zitadel issuer. -**Algorithm:** `HS256` (HMAC with SHA-256) +| Component | Responsibility | +|--------------------------|-------------------------------------------------------------------| +| MP (`src/oidc.rs`) | Load the key file once, call the crate per request, build headers | +| `zitadel` crate | OIDC discovery, JWKS fetch, assertion signing, token request | +| Status Dashboard backend | Token verification (`SD_OIDC_*` settings) | -**Token Structure:** +## Machine User Key File -The JWT token contains a simple claim structure: +The reporter authenticates as a Zitadel **machine user** (service user). The key file is downloaded +from the Zitadel Console (instance → *Users* → *Service Users* → select the machine user → *Keys* → +*New* → download the JSON key file): ```json { - "stackmon": "dummy" + "type": "serviceaccount", + "keyId": "392067695547252958", + "key": "-----BEGIN RSA PRIVATE KEY-----\n...\n-----END RSA PRIVATE KEY-----", + "userId": "392040635458125910" } ``` -**Signing Process:** - -1. The shared secret is loaded from configuration (`status_dashboard.secret`) -2. An HMAC-SHA256 key is created from the secret bytes -3. Claims are signed with the key to produce the JWT token -4. The token is included in the `Authorization` header as a Bearer token +| Field | Meaning | +|----------|--------------------------------------------------------------------------------------------------------| +| `type` | Must be `serviceaccount`; any other value aborts the reporter at startup with the configuration key named | +| `keyId` | Key id sent as the `kid` header of the assertion, used by Zitadel to select the public key | +| `key` | PEM encoded RSA private key; Zitadel emits PKCS#1 (`BEGIN RSA PRIVATE KEY`), PKCS#8 is accepted too | +| `userId` | Machine user id, used as `iss` and `sub` of the assertion | + +Notes: + +- The path is configured through `status_dashboard.oidc_key_file` + (`MP_STATUS_DASHBOARD__OIDC_KEY_FILE`). The file is read and validated exactly once at startup, so + it does not have to stay available after the reporter started. +- Zitadel *application* keys (`type: application`, `clientId`) are not accepted: the `zitadel` crate + implements the machine user JWT Profile flow, which is the key kind this deployment uses. Such a + key file aborts startup with an error naming `MP_STATUS_DASHBOARD__OIDC_KEY_FILE`. +- Keep the key file out of the repository and out of configuration files, for example by mounting it + as a secret volume or by writing it to a container file system path. +- The key material, the signed assertion and the access token are never logged, and the private key + never appears in an error message. `OidcIdentity` redacts the loaded credentials in its `Debug` + output. + +## Flow + +1. The reporter resolves its service identity from the configuration + (`status_dashboard.oidc_issuer`, `oidc_key_file`, `oidc_scopes`). A missing or unusable value + aborts the reporter at startup and the error names the offending configuration key. +2. Before every report the crate signs a new assertion and exchanges it for an access token. + Tokens are never cached in-process, so every report carries a fresh token. +3. The access token is sent as `Authorization: Bearer ` on every Status Dashboard + call. +4. The Status Dashboard validates the token with `go-oidc`: the issuer must match `SD_OIDC_ISSUER`, + the audience must contain `SD_OIDC_CLIENT_ID`, and the signing key is resolved from the JWKS + endpoint by `kid`. +5. The reporter role is read from the project roles claim + (`urn:zitadel:iam:org:project:roles`, configurable on the backend via `SD_OIDC_ROLES_CLAIM`) and + mapped to the reporter role configured by `SD_RBAC_ROLES_REPORTERS`. + +## Assertion + +The assertion signed by the crate for the machine user looks like this: -### Configuration - -Authentication is configured in the `status_dashboard` section of the configuration file: +```json +{ + "alg": "RS256", + "kid": "392067695547252958", + "typ": "JWT" +} +``` -```yaml -status_dashboard: - url: https://status-dashboard.example.com - secret: your-shared-secret-key +```json +{ + "iss": "392040635458125910", + "sub": "392040635458125910", + "aud": "https://zitadel.example.com", + "iat": 1705929045, + "exp": 1705932645 +} ``` -| Field | Type | Required | Description | -|-------|------|----------|-------------| -| `url` | string | Yes | Status dashboard base URL | -| `secret` | string | No | JWT signing secret. If not provided, requests are sent without authentication | +`iss` and `sub` are the `userId` of the machine user, the audience is the issuer URL and `exp` is one +hour after `iat`. -### Environment Variable Override +## Token Request -The secret can also be set via environment variable: +The reporter presents the assertion as the `assertion` parameter of the JWT bearer grant, without +any `Authorization` header and without a client id: ```bash -export MP_STATUS_DASHBOARD__SECRET="your-shared-secret-key" +curl -sS -X POST "https://zitadel.example.com/oauth/v2/token" \ + -H "Content-Type: application/x-www-form-urlencoded" \ + -d "grant_type=urn:ietf:params:oauth:grant-type:jwt-bearer" \ + -d "assertion=" \ + -d "scope=" ``` -Environment variables are merged with the configuration file, with environment variables taking precedence. +Notes: -## Token Usage +- The token endpoint is discovered per request from `{issuer}/.well-known/openid-configuration`, so + only the issuer is configured. +- The `zitadel` crate always adds the `openid` scope in front of the configured scopes and sends the + whole list space-joined as a single `scope` parameter, in the configured order. -### Authorization Header +### Required scopes -When making requests to the status-dashboard, the JWT token is included in the HTTP `Authorization` header using the Bearer scheme: +The configured scope list has to contain both of these scopes: -``` -Authorization: Bearer eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJzdGFja21vbiI6ImR1bW15In0. -``` +| Scope | Why it is needed | +|------------------------------------------------------------|-------------------------------------------------------------------------------------| +| `urn:zitadel:iam:org:project:role:sd_reporters` | Puts the reporter project role into the token, so the status page RBAC check passes | +| `urn:zitadel:iam:org:project:id::aud` | Makes the token audience the project id instead of the client id | -### Request Flow +`` is the Zitadel project that both the Status Dashboard and the machine user are part +of. Zitadel puts the client id into the token audience by default, and the Status Dashboard rejects +that with `expected audience ... got [...]`; requesting the audience scope switches the audience to +the project id, so the backend has to be configured with `SD_OIDC_CLIENT_ID=`. -1. **Configuration Load:** The reporter reads the `status_dashboard.secret` from configuration -2. **Token Generation:** If a secret is configured, a JWT token is generated at startup -3. **Request Authentication:** All POST requests to `/v1/component_status` include the Bearer token -4. **Server Validation:** The status-dashboard validates the token signature using the same shared secret - -## Token Validation - -On the server side (status-dashboard), tokens should be validated by: - -1. Extracting the token from the `Authorization` header -2. Verifying the HMAC-SHA256 signature using the shared secret -3. Optionally checking the claims (currently contains `{"stackmon": "dummy"}`) +```yaml +status_dashboard: + oidc_scopes: + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id:392066917738875090:aud" +``` -## Security Considerations +`oidc_scopes` has no default, because the audience scope contains the project id of the deployment. +The reporter fails at startup when the list misses either scope and names +`MP_STATUS_DASHBOARD__OIDC_SCOPES` in the error. -### Secret Management +Verified against the pre-production instance with those two scopes: the access token carried +`aud = ["390700708019568682"]` and `groups = ["sd_reporters"]`, with `iss` set to the issuer URL. -- **Never commit secrets** to version control -- Use environment variables (`MP_STATUS_DASHBOARD__SECRET`) in production -- Rotate secrets periodically -- Use strong, randomly-generated secrets (minimum 32 characters recommended) +## Report Authentication -### Transport Security +```http +POST /v2/events HTTP/1.1 +Host: status.example.com +Authorization: Bearer +Content-Type: application/json -- Always use HTTPS for the status-dashboard URL in production -- The JWT token is sent in clear text in the Authorization header -- Without TLS, tokens could be intercepted and reused +{ + "title": "System incident from monitoring system", + "description": "Object Storage Service is degraded", + "impact": 2, + "components": [218], + "start_date": "2024-01-22T10:30:44Z", + "system": true, + "type": "incident" +} +``` -### Token Characteristics +## Permission Boundary -- **Stateless:** Tokens are self-contained and don't require server-side session storage -- **No Expiration:** Current implementation does not include expiration claims -- **Single Use Case:** Tokens are specifically for machine-to-machine authentication between reporter and status-dashboard +The Zitadel service identity used by the reporter is restricted to a single operation: -### Best Practices +- Allowed: `POST /v2/events` with `"system": true` +- Denied: component read endpoints and any event without `system: true` -1. **Use Strong Secrets:** Generate cryptographically secure random strings - ```bash - openssl rand -base64 32 - ``` +The Status Dashboard backend enforces this boundary and rejects reporter-scoped tokens on every +other route, independently of the roles present in the token. -2. **Environment Separation:** Use different secrets for development, staging, and production +## Backend Verification Points -3. **Audit Logging:** Log authentication failures on the status-dashboard for monitoring +| Check | Backend setting | Expectation | +|---------------|--------------------------|------------------------------------------------------------| +| Issuer | `SD_OIDC_ISSUER` | Matches `status_dashboard.oidc_issuer` | +| Audience | `SD_OIDC_CLIENT_ID` | The project id, requested through the audience scope | +| Signing key | JWKS by `kid` | Resolved from the issuer | +| Roles claim | `SD_OIDC_ROLES_CLAIM` | `urn:zitadel:iam:org:project:roles` | +| Reporter role | `SD_RBAC_ROLES_REPORTERS` | Role key requested through `oidc_scopes` | -4. **Network Isolation:** Where possible, restrict network access between components +The reporter does not send an audience of its own: Zitadel decides the audience of the access token, +so the backend has to be configured with the audience the requested scopes produce: -## Example Implementation +- Without the audience scope the token audience is the client id of the machine user, which the + backend rejects with `expected audience ... got [...]`. +- With `urn:zitadel:iam:org:project:id::aud` the audience is the project id, which is + what `SD_OIDC_CLIENT_ID` has to be set to. -### Generating a Token (Rust) +## Failure Handling -```rust -use hmac::{Hmac, Mac}; -use jwt::SignWithKey; -use sha2::Sha256; -use std::collections::BTreeMap; +The reporter is fail-closed: -let secret = "your-shared-secret"; -let key: Hmac = Hmac::new_from_slice(secret.as_bytes()).unwrap(); +- A missing `oidc_issuer`, `oidc_key_file` or `oidc_scopes`, an unreadable or invalid key file, a key + type other than `serviceaccount`, an empty `keyId`/`key`/`userId` and a scope list without a role + scope or without the project audience scope abort the reporter at startup, and the error names the + configuration key that has to be fixed. +- A failing discovery or token request, and a token response without a usable `access_token`, abort + the report instead of sending it without authentication. +- The private key, the signed assertion and the access token are never part of an error message or + of the reporter output. -let mut claims = BTreeMap::new(); -claims.insert("stackmon", "dummy"); +## Configuration -let token_str = claims.sign_with_key(&key).unwrap(); -let bearer = format!("Bearer {}", token_str); +```yaml +status_dashboard: + url: "https://status.example.com" + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-identity.json" + oidc_scopes: + # Reports the reporter role and makes the token audience the project id + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id:392066917738875090:aud" ``` -### Validating a Token (Pseudocode) +Environment variable equivalents: -``` -function validate_token(authorization_header, secret): - # Extract token from "Bearer " - token = extract_bearer_token(authorization_header) - - # Verify signature - key = hmac_sha256_key(secret) - claims = verify_and_decode(token, key) - - if claims is valid: - return true - else: - return false -``` +| Environment Variable | Configuration path | +|--------------------------------------|----------------------------------| +| `MP_STATUS_DASHBOARD__URL` | `status_dashboard.url` | +| `MP_STATUS_DASHBOARD__OIDC_ISSUER` | `status_dashboard.oidc_issuer` | +| `MP_STATUS_DASHBOARD__OIDC_KEY_FILE` | `status_dashboard.oidc_key_file` | +| `MP_STATUS_DASHBOARD__OIDC_SCOPES` | `status_dashboard.oidc_scopes` | diff --git a/doc/architecture/data-flow.md b/doc/architecture/data-flow.md index e8723c1..8421498 100644 --- a/doc/architecture/data-flow.md +++ b/doc/architecture/data-flow.md @@ -313,7 +313,7 @@ flowchart TD Parse["Parse Response"] Check{"Last value > 0?"} Skip["Skip notification"] - Build["Build ComponentStatus"] + Build["Build IncidentData"] Post["POST to Dashboard"] Poll --> Parse @@ -340,7 +340,7 @@ if let Some(last) = data.metrics.pop() { ```http POST /v1/component_status -Authorization: Bearer +Authorization: Bearer Content-Type: application/json { @@ -402,7 +402,7 @@ sequenceDiagram API-->>Reporter: ServiceHealthResponse alt Health status > 0 - Reporter->>Reporter: Build ComponentStatus + Reporter->>Reporter: Build IncidentData Reporter->>Dashboard: POST /v1/component_status Dashboard-->>Reporter: 200 OK end diff --git a/doc/architecture/diagrams.md b/doc/architecture/diagrams.md index 6b30044..58ab2bf 100644 --- a/doc/architecture/diagrams.md +++ b/doc/architecture/diagrams.md @@ -25,7 +25,7 @@ graph TB subgraph Reporter["Reporter Binary"] Poller["Metric Poller
(60s interval)"] - Notifier["Dashboard Notifier
(JWT Auth)"] + Notifier["Dashboard Notifier
(OIDC Auth)"] end end @@ -168,8 +168,7 @@ graph TB end subgraph Auth["Authentication"] - JWT["jwt"] - HMAC["hmac + sha2"] + OidcToken["Zitadel OIDC token
(reqwest)"] end subgraph Utilities["Utilities"] @@ -190,8 +189,7 @@ graph TB Reporter --> Tokio Reporter --> Reqwest Reporter --> Serde - Reporter --> JWT - Reporter --> HMAC + Reporter --> OidcToken Reporter --> Tracing Axum --> Hyper @@ -256,7 +254,7 @@ graph TB end CM["ConfigMap
config.yaml"] - Secret["Secret
JWT credentials"] + Secret["Secret
OIDC client credentials"] end end @@ -429,7 +427,9 @@ graph TD HM --> HM_Exprs["expressions: Vec"] Status --> Status_URL["url: String"] - Status --> Status_Secret["secret: Option"] + Status --> Status_Issuer["oidc_issuer: Option"] + Status --> Status_KeyFile["oidc_key_file: Option"] + Status --> Status_Scopes["oidc_scopes: Option>"] end style Config fill:#e8eaf6 diff --git a/doc/architecture/overview.md b/doc/architecture/overview.md index 6b10383..3ef4db6 100644 --- a/doc/architecture/overview.md +++ b/doc/architecture/overview.md @@ -43,7 +43,7 @@ The Reporter is a background service that: - **Polls Convertor**: Queries the Convertor API at configurable intervals (default: 60s) - **Detects Issues**: Identifies when health status indicates degradation or outage - **Sends Notifications**: Posts status updates to external dashboards (e.g., Atlassian Statuspage) -- **Handles Authentication**: Manages JWT tokens for secure dashboard communication +- **Handles Authentication**: Obtains Zitadel OIDC service identity tokens for secure dashboard communication **Dependencies**: - Convertor API (localhost HTTP calls) @@ -63,7 +63,7 @@ The Reporter is a background service that: **Role**: External consumer of health status - **Supported**: CloudMon Status Dashboard (custom API) -- **Protocol**: REST API with JWT authentication +- **Protocol**: REST API authenticated with Zitadel OIDC service identity tokens - **Data Format**: Component status with name, impact level, and attributes ## Key Design Decisions @@ -256,15 +256,16 @@ src/ ```bash # Environment variable example -MP_STATUS_DASHBOARD__SECRET=my-jwt-secret -# Translates to: status_dashboard.secret = "my-jwt-secret" +MP_STATUS_DASHBOARD__OIDC_KEY_FILE=/etc/cloudmon/service-account.json +# Translates to: status_dashboard.oidc_key_file = "/etc/cloudmon/service-account.json" ``` ## Security Considerations ### Authentication -- **Status Dashboard**: JWT tokens signed with HMAC-SHA256 +- **Status Dashboard**: Zitadel OIDC service identity tokens (JWT Profile exchange delegated to the + `zitadel` crate, machine user key file) - **Internal APIs**: No authentication (expected behind firewall) ### Network Security @@ -274,7 +275,7 @@ MP_STATUS_DASHBOARD__SECRET=my-jwt-secret ### Configuration Secrets -- Sensitive values (JWT secrets) should be injected via environment variables +- Sensitive values (OIDC client secrets) should be injected via environment variables - Config files should not contain production secrets ## Performance Characteristics diff --git a/doc/config.md b/doc/config.md index 4fc5a18..876d12e 100644 --- a/doc/config.md +++ b/doc/config.md @@ -34,7 +34,8 @@ environments: status_dashboard: url: "https://status.cloudmon.com" - secret: "dev" + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-account.json" flag_metrics: ### Comp1 @@ -89,18 +90,37 @@ This section is providing capability to describe query templates to be later ref ## status_dashboard -Configures URL and JWT secret for communication with the status dashboard. +Configures the URL and the Zitadel OIDC service identity used for communication with the status dashboard. ```yaml status_dashboard: url: "https://status-dashboard.example.com" - secret: "your-jwt-secret" + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-account.json" + oidc_scopes: + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id::aud" ``` -| Property | Type | Required | Default | Description | -|----------|--------|----------|---------|---------------------------------------| -| `url` | string | Yes | - | Status Dashboard API URL | -| `secret` | string | No | - | JWT signing secret for authentication | +| Property | Type | Required | Default | Description | +|----------------|----------|----------|-----------------------------------------------------|------------------------------------------------| +| `url` | string | Yes | - | Status Dashboard API URL | +| `oidc_issuer` | string | Yes | - | Zitadel issuer URL | +| `oidc_key_file` | string | Yes | - | Path to the Zitadel machine user key file | +| `oidc_scopes` | string[] | Yes | - | Requested token scopes | + +The key file is the JSON key file downloaded from the Zitadel Console for a machine user +(`type: serviceaccount`); the reporter delegates the JWT Profile exchange (assertion signing and +token request) to the `zitadel` crate instead of authenticating with a client secret. The OIDC +token endpoint is discovered from `{oidc_issuer}/.well-known/openid-configuration`. + +`oidc_scopes` has no default and has to list both a project role scope +(`urn:zitadel:iam:org:project:role:`) and the audience scope of the Zitadel project shared +with the Status Dashboard (`urn:zitadel:iam:org:project:id::aud`, where `` is +the value the Status Dashboard uses for `SD_OIDC_CLIENT_ID`). The reporter rejects a list that +misses either scope at startup, because Zitadel only reports the project roles claim when the +audience scope is requested and the Status Dashboard rejects a token whose audience is the client +id. ## health_query diff --git a/doc/configuration/examples.md b/doc/configuration/examples.md index 367a2a9..e88a89a 100644 --- a/doc/configuration/examples.md +++ b/doc/configuration/examples.md @@ -596,7 +596,12 @@ server: status_dashboard: url: "https://status.example.com" - # Secret should be set via MP_STATUS_DASHBOARD__SECRET environment variable + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-account.json" + # Zitadel machine user key file mounted as a secret volume + oidc_scopes: + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id::aud" metric_templates: api_latency: diff --git a/doc/configuration/overview.md b/doc/configuration/overview.md index d2672a6..baaddb5 100644 --- a/doc/configuration/overview.md +++ b/doc/configuration/overview.md @@ -57,8 +57,8 @@ export MP_DATASOURCE__URL="http://graphite.example.com:8080" # Override server.port export MP_SERVER__PORT=3005 -# Set status_dashboard.secret (sensitive values) -export MP_STATUS_DASHBOARD__SECRET="your-jwt-secret" +# Point the reporter at the Zitadel machine user key file +export MP_STATUS_DASHBOARD__OIDC_KEY_FILE="/etc/cloudmon/service-account.json" ``` **Best Practice**: Use environment variables for sensitive values like secrets and for deployment-specific overrides in containerized environments. diff --git a/doc/configuration/schema.md b/doc/configuration/schema.md index c2a45de..142bb91 100644 --- a/doc/configuration/schema.md +++ b/doc/configuration/schema.md @@ -195,14 +195,23 @@ Optional status dashboard integration. | Property | Type | Required | Description | |----------|------|----------|-------------| | `url` | string | Yes | Status dashboard URL | -| `secret` | string | No | JWT signing secret | +| `oidc_issuer` | string | Yes | Zitadel OIDC issuer URL | +| `oidc_key_file` | string | Yes | Path to the Zitadel machine user key file (JWT profile) | +| `oidc_scopes` | string[] | Yes | Requested token scopes; has to list a `project:role:` scope and the `project:id::aud` scope of the Status Dashboard project | ```yaml status_dashboard: url: "https://status.example.com" - secret: "your-jwt-secret" # Use MP_STATUS_DASHBOARD__SECRET env var instead + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-account.json" + oidc_scopes: + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id::aud" ``` +A scope list that misses the role scope or the audience scope fails startup with +`MP_STATUS_DASHBOARD__OIDC_SCOPES` named in the error. + ## Comparison Operators The `op` field in templates accepts these values: diff --git a/doc/getting-started/development.md b/doc/getting-started/development.md index 710d7ad..5c85a73 100644 --- a/doc/getting-started/development.md +++ b/doc/getting-started/development.md @@ -21,6 +21,12 @@ pre-commit install cargo install mdbook ``` +Native build tools are required as well, because the `zitadel` crate builds `aws-lc-rs` for its JWT +support: + +- Linux/macOS: a C toolchain and CMake +- Windows: a C toolchain plus NASM on `PATH` (for example `scoop install nasm`) + ## Building the Project ### Development Build diff --git a/doc/getting-started/project-structure.md b/doc/getting-started/project-structure.md index b1cdf82..51c639a 100644 --- a/doc/getting-started/project-structure.md +++ b/doc/getting-started/project-structure.md @@ -76,7 +76,7 @@ Two independent executable binaries: - `GET /v1/health`: Query health metrics for services - `GET /v1/maintenances`: Query maintenance status - **Key types**: `HealthQuery`, `HealthResponse`, `MaintenancesResponse` -- **Authentication**: JWT token validation for status dashboard +- **Outbound authentication**: Zitadel OIDC service identity for reporter calls **When to edit**: Adding/modifying API endpoints, changing request/response formats diff --git a/doc/guides/deployment.md b/doc/guides/deployment.md index ac74ee4..054349b 100644 --- a/doc/guides/deployment.md +++ b/doc/guides/deployment.md @@ -93,8 +93,9 @@ docker run -d \ --name metrics-reporter \ --network host \ -v $(pwd)/config:/cloudmon/config:ro \ + -v $(pwd)/secrets/service-account.json:/cloudmon/service-account.json:ro \ -e RUST_LOG=info \ - -e MP_STATUS_DASHBOARD__SECRET=your-jwt-secret \ + -e MP_STATUS_DASHBOARD__OIDC_KEY_FILE=/cloudmon/service-account.json \ metrics-processor:latest \ /cloudmon/cloudmon-metrics-reporter ``` @@ -128,9 +129,10 @@ services: command: /cloudmon/cloudmon-metrics-reporter volumes: - ./config:/cloudmon/config:ro + - ./secrets/service-account.json:/cloudmon/service-account.json:ro environment: - RUST_LOG=info - - MP_STATUS_DASHBOARD__SECRET=${STATUS_DASHBOARD_SECRET} + - MP_STATUS_DASHBOARD__OIDC_KEY_FILE=/cloudmon/service-account.json depends_on: convertor: condition: service_healthy @@ -221,7 +223,14 @@ metadata: namespace: monitoring type: Opaque stringData: - status-dashboard-secret: "your-jwt-secret-here" + # JSON key file downloaded from the Zitadel Console for the machine user + service-account.json: | + { + "type": "serviceaccount", + "keyId": "81693565968962154", + "key": "-----BEGIN RSA PRIVATE KEY-----\n...\n-----END RSA PRIVATE KEY-----", + "userId": "392040635458125910" + } ``` ### Convertor Deployment @@ -343,16 +352,17 @@ spec: env: - name: RUST_LOG value: "info" - - name: MP_STATUS_DASHBOARD__SECRET - valueFrom: - secretKeyRef: - name: metrics-processor-secrets - key: status-dashboard-secret + - name: MP_STATUS_DASHBOARD__OIDC_KEY_FILE + value: /cloudmon/service-account.json volumeMounts: - name: config mountPath: /cloudmon/config.yaml subPath: config.yaml readOnly: true + - name: service-account-key + mountPath: /cloudmon/service-account.json + subPath: service-account.json + readOnly: true resources: requests: memory: "32Mi" @@ -364,6 +374,12 @@ spec: - name: config configMap: name: metrics-processor-config + - name: service-account-key + secret: + secretName: metrics-processor-secrets + items: + - key: service-account.json + path: service-account.json ``` ### Ingress Configuration @@ -461,8 +477,8 @@ health_metrics: Override configuration values using environment variables prefixed with `MP_`: ```bash -# Override status dashboard secret -export MP_STATUS_DASHBOARD__SECRET=production-secret +# Override the Zitadel key file path +export MP_STATUS_DASHBOARD__OIDC_KEY_FILE=/cloudmon/service-account.json # Override datasource URL export MP_DATASOURCE__URL=https://graphite-prod.example.com @@ -494,8 +510,8 @@ configMapGenerator: secretGenerator: - name: metrics-processor-secrets - literals: - - status-dashboard-secret=your-secret + files: + - service-account.json=secrets/service-account.json images: - name: metrics-processor @@ -682,7 +698,7 @@ The metrics-processor is **stateless**: 2. **Dependencies:** - [ ] Graphite TSDB URL and credentials - - [ ] Status Dashboard URL and JWT secret + - [ ] Status Dashboard URL and OIDC service identity credentials 3. **Recovery steps:** ```bash @@ -797,10 +813,10 @@ spec: target: name: metrics-processor-secrets data: - - secretKey: status-dashboard-secret + - secretKey: service-account.json remoteRef: key: secret/metrics-processor - property: jwt-secret + property: status-dashboard-service-account-key ``` ### Container Security diff --git a/doc/guides/troubleshooting.md b/doc/guides/troubleshooting.md index 5c39050..9246474 100644 --- a/doc/guides/troubleshooting.md +++ b/doc/guides/troubleshooting.md @@ -459,11 +459,15 @@ ERROR cloudmon_metrics: Error during posting component status: error sending req **Cause:** Status dashboard is unreachable or misconfigured. **Solution:** -1. Verify status dashboard URL: +1. Verify status dashboard URL and service identity: ```yaml status_dashboard: url: https://status.cloudmon.com - secret: your-jwt-secret + oidc_issuer: https://zitadel.example.com + oidc_key_file: /etc/cloudmon/service-account.json + oidc_scopes: + - urn:zitadel:iam:org:project:role:sd_reporters + - urn:zitadel:iam:org:project:id::aud ``` 2. Test connectivity: ```bash diff --git a/doc/reporter.md b/doc/reporter.md index 45e96d2..4987e51 100644 --- a/doc/reporter.md +++ b/doc/reporter.md @@ -9,7 +9,7 @@ The reporter acts as a bridge between the convertor's real-time health evaluatio 2. Polls convertor API at regular intervals (60 seconds) 3. Checks if service health has degraded (impact > 0) 4. Creates incidents via Status Dashboard API -5. Handles HMAC-JWT authentication +5. Authenticates with a Zitadel OIDC service identity **Key Characteristics**: - **Background service**: Runs as daemon or scheduled job @@ -138,18 +138,24 @@ Incidents are created with static, secure payloads: ### 5. Authentication -The reporter uses HMAC-JWT for authentication (unchanged from V1): +The reporter authenticates to the Status Dashboard with a Zitadel OIDC service identity whose JWT +Profile exchange is delegated to the community `zitadel` crate: ```rust -// Generate HMAC-JWT token -let headers = build_auth_headers(secret.as_deref()); -// Headers contain: Authorization: Bearer +// Request a service identity token and build the authorization headers +let headers = build_auth_headers(&service_identity).await?; ``` -**Token Format**: -- Algorithm: HMAC-SHA256 -- Claims: `{"stackmon": "dummy"}` -- Optional: No secret = no auth header (for environments without auth) +**Token acquisition**: +- Endpoint: discovered per request from `{status_dashboard.oidc_issuer}/.well-known/openid-configuration` +- Grant: `urn:ietf:params:oauth:grant-type:jwt-bearer` with an RS256 assertion the crate signs from + the machine user key file; never HTTP Basic auth and never a client secret +- Credentials: the Zitadel machine user key file configured through `oidc_key_file` +- Scope: `status_dashboard.oidc_scopes` (the crate prepends `openid`), space-joined in one field +- Fail-closed: a missing issuer, key file or scope list, an unreadable, invalid or mistyped key + file, and a scope list without a project role scope or without the project audience scope abort + the reporter at startup, naming the configuration key +- No token caching in-process, a fresh assertion and token are requested before every report ## Module Structure @@ -166,7 +172,7 @@ pub struct IncidentData { title, description, impact, components, start_date, sy pub type ComponentCache = HashMap<(String, Vec), u32>; // Authentication -pub fn build_auth_headers(secret: Option<&str>) -> HeaderMap +pub async fn build_auth_headers(identity: &OidcIdentity) -> anyhow::Result // V2 API Functions pub async fn fetch_components(...) -> Result> @@ -176,6 +182,9 @@ pub fn build_incident_data(...) -> IncidentData pub async fn create_incident(...) -> Result<()> ``` +The Zitadel side of the flow lives in `src/oidc.rs`: it loads the machine user key file once and +delegates discovery, assertion signing and the token request to the `zitadel` crate. + ## Configuration The reporter requires configuration for: @@ -193,13 +202,19 @@ convertor: ```yaml status_dashboard: url: "https://dashboard.example.com" - secret: "your-jwt-secret" + oidc_issuer: "https://zitadel.example.com" + oidc_key_file: "/etc/cloudmon/service-account.json" + oidc_scopes: + - "urn:zitadel:iam:org:project:role:sd_reporters" + - "urn:zitadel:iam:org:project:id::aud" ``` -| Property | Type | Required | Default | Description | -|----------|--------|----------|---------|---------------------------------------| -| `url` | string | Yes | - | Status Dashboard API URL | -| `secret` | string | No | - | JWT signing secret for authentication | +| Property | Type | Required | Default | Description | +|----------------|----------|----------|-----------------------------------------------------|---------------------------------------------| +| `url` | string | Yes | - | Status Dashboard API URL | +| `oidc_issuer` | string | Yes | - | Zitadel issuer URL | +| `oidc_key_file` | string | Yes | - | Path to the Zitadel machine user key file | +| `oidc_scopes` | string[] | Yes | - | Requested token scopes; needs a `project:role:` scope and the `project:id::aud` scope | ### Health Query Configuration @@ -282,7 +297,7 @@ spec: Override configuration: ```bash -MP_STATUS_DASHBOARD__SECRET=new-secret \ +MP_STATUS_DASHBOARD__OIDC_KEY_FILE=/etc/cloudmon/service-account.json \ MP_CONVERTOR__URL=http://convertor-svc:3005 \ cloudmon-metrics-reporter --config config.yaml ``` @@ -307,7 +322,7 @@ cloudmon-metrics-reporter --config config.yaml - Poll cycle duration - Notification success rate - API errors (convertor, dashboard) -- JWT token generation failures +- OIDC service token acquisition failures **Logging**: ```bash @@ -390,8 +405,9 @@ When the reporter decides to create an incident, it logs all the information nee ### Authentication Failures -**Cause**: Invalid JWT secret -**Solution**: Update `status_dashboard.secret` in configuration +**Cause**: Invalid Zitadel OIDC service identity credentials +**Solution**: Verify `status_dashboard.oidc_issuer` and that `oidc_key_file` points to the key file +of a `serviceaccount` machine user (or `application` client), whose key id, key and id are all present ## Use Cases @@ -429,8 +445,8 @@ curl http://localhost:3005/v1/health?service=api&environment=prod&from=2024-01-0 ### "Dashboard authentication failed" -**Cause**: Invalid JWT secret -**Solution**: Ensure `status_dashboard.secret` matches dashboard configuration +**Cause**: Invalid service identity credentials or a missing Reporter role +**Solution**: Verify the Zitadel machine user credentials and that the requested role scope maps to the dashboard reporter role ### "No services being polled" diff --git a/doc/schemas/config-schema.json b/doc/schemas/config-schema.json index e7fdaab..60d764f 100644 --- a/doc/schemas/config-schema.json +++ b/doc/schemas/config-schema.json @@ -309,13 +309,30 @@ "url" ], "properties": { - "secret": { - "description": "JWT token signature secret", + "oidc_issuer": { + "description": "Zitadel OIDC issuer URL", "type": [ "string", "null" ] }, + "oidc_key_file": { + "description": "Path to the Zitadel machine user key file.\n\nThe file is downloaded from the Zitadel Console for a machine user (service user) and has `type: serviceaccount`; it is read once at startup and any other key type fails startup.", + "type": [ + "string", + "null" + ] + }, + "oidc_scopes": { + "description": "OIDC scopes of the token request, sent as one space-joined `scope` parameter.\n\n`MP_STATUS_DASHBOARD__OIDC_SCOPES` has to contain both `urn:zitadel:iam:org:project:role:sd_reporters`, so that the roles are reported in the `groups` claim, and `urn:zitadel:iam:org:project:id::aud`, which makes `aud` the project id the Status Dashboard verifies against `SD_OIDC_CLIENT_ID`. `` is the Zitadel project shared by the Status Dashboard and the machine user. Without the audience scope Zitadel puts the client id into `aud`, which the backend rejects.\n\nThere is no default, because the audience scope carries the project id of the deployment.", + "type": [ + "array", + "null" + ], + "items": { + "type": "string" + } + }, "url": { "description": "Status dashboard URL", "type": "string" From 1f1c9f5cb4b5faf0c654beb1ca7b1e1a10783e9b Mon Sep 17 00:00:00 2001 From: -z <11802769+Aloento@users.noreply.github.com> Date: Wed, 23 Sep 2026 22:49:38 +0200 Subject: [PATCH 2/2] docs: drop the generated specifications and the duplicated module pages The repository carried 667 kB of markdown under `doc/` and `specs/` next to 158 kB of Rust source. Two parts of it are redundant: `specs/` holds the generated design artifacts of three past features (they restate what the code and `doc/` already say, and two of them still describe the V1 API and the HMAC flow), and `doc/modules/` restates module signatures that the source and `doc/reporter.md` cover, with an example that no longer compiles. - remove `specs/` (feature plans, tasks, checklists and API contracts) - remove `doc/modules/` and the module section of the summary - point the project structure, quickstart and README pages at the architecture page and `cargo doc` instead - fix the relative link to the architecture page in the quickstart --- README.md | 2 - doc/SUMMARY.md | 8 - doc/getting-started/project-structure.md | 5 +- doc/getting-started/quickstart.md | 6 +- doc/modules/api.md | 140 ---- doc/modules/common.md | 229 ----- doc/modules/config.md | 215 ----- doc/modules/graphite.md | 250 ------ doc/modules/overview.md | 101 --- doc/modules/sd.md | 231 ------ doc/modules/types.md | 308 ------- .../IMPLEMENTATION_REPORT.md | 273 ------ .../checklists/requirements.md | 65 -- .../contracts/README.md | 157 ---- .../contracts/config-schema.json | 217 ----- .../contracts/patterns.json | 190 ----- specs/001-project-documentation/data-model.md | 396 --------- specs/001-project-documentation/plan.md | 310 ------- specs/001-project-documentation/quickstart.md | 403 --------- specs/001-project-documentation/research.md | 387 --------- specs/001-project-documentation/spec.md | 173 ---- specs/001-project-documentation/tasks.md | 409 --------- .../IMPLEMENTATION_REPORT.md | 343 -------- .../IMPLEMENTATION_STATUS.txt | 206 ----- .../checklists/requirements.md | 127 --- specs/002-functional-test-suite/plan.md | 492 ----------- specs/002-functional-test-suite/spec.md | 245 ------ specs/002-functional-test-suite/tasks.md | 358 -------- .../test-implementation-summary.md | 249 ------ .../checklists/requirements.md | 54 -- .../contracts/README.md | 57 -- .../contracts/components-api.md | 296 ------- .../contracts/incidents-api.md | 468 ----------- .../create-incident-multi-component.json | 9 - .../create-incident-single-component.json | 9 - .../response-examples/components-list.json | 39 - .../response-examples/incident-created.json | 8 - specs/003-sd-api-v2-migration/data-model.md | 785 ------------------ specs/003-sd-api-v2-migration/plan.md | 340 -------- specs/003-sd-api-v2-migration/quickstart.md | 721 ---------------- specs/003-sd-api-v2-migration/research.md | 276 ------ specs/003-sd-api-v2-migration/spec.md | 202 ----- specs/003-sd-api-v2-migration/tasks.md | 370 --------- 43 files changed, 4 insertions(+), 10125 deletions(-) delete mode 100644 doc/modules/api.md delete mode 100644 doc/modules/common.md delete mode 100644 doc/modules/config.md delete mode 100644 doc/modules/graphite.md delete mode 100644 doc/modules/overview.md delete mode 100644 doc/modules/sd.md delete mode 100644 doc/modules/types.md delete mode 100644 specs/001-project-documentation/IMPLEMENTATION_REPORT.md delete mode 100644 specs/001-project-documentation/checklists/requirements.md delete mode 100644 specs/001-project-documentation/contracts/README.md delete mode 100644 specs/001-project-documentation/contracts/config-schema.json delete mode 100644 specs/001-project-documentation/contracts/patterns.json delete mode 100644 specs/001-project-documentation/data-model.md delete mode 100644 specs/001-project-documentation/plan.md delete mode 100644 specs/001-project-documentation/quickstart.md delete mode 100644 specs/001-project-documentation/research.md delete mode 100644 specs/001-project-documentation/spec.md delete mode 100644 specs/001-project-documentation/tasks.md delete mode 100644 specs/002-functional-test-suite/IMPLEMENTATION_REPORT.md delete mode 100644 specs/002-functional-test-suite/IMPLEMENTATION_STATUS.txt delete mode 100644 specs/002-functional-test-suite/checklists/requirements.md delete mode 100644 specs/002-functional-test-suite/plan.md delete mode 100644 specs/002-functional-test-suite/spec.md delete mode 100644 specs/002-functional-test-suite/tasks.md delete mode 100644 specs/002-functional-test-suite/test-implementation-summary.md delete mode 100644 specs/003-sd-api-v2-migration/checklists/requirements.md delete mode 100644 specs/003-sd-api-v2-migration/contracts/README.md delete mode 100644 specs/003-sd-api-v2-migration/contracts/components-api.md delete mode 100644 specs/003-sd-api-v2-migration/contracts/incidents-api.md delete mode 100644 specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-multi-component.json delete mode 100644 specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-single-component.json delete mode 100644 specs/003-sd-api-v2-migration/contracts/response-examples/components-list.json delete mode 100644 specs/003-sd-api-v2-migration/contracts/response-examples/incident-created.json delete mode 100644 specs/003-sd-api-v2-migration/data-model.md delete mode 100644 specs/003-sd-api-v2-migration/plan.md delete mode 100644 specs/003-sd-api-v2-migration/quickstart.md delete mode 100644 specs/003-sd-api-v2-migration/research.md delete mode 100644 specs/003-sd-api-v2-migration/spec.md delete mode 100644 specs/003-sd-api-v2-migration/tasks.md diff --git a/README.md b/README.md index 3ed1d75..04c1839 100644 --- a/README.md +++ b/README.md @@ -25,7 +25,6 @@ metrics-processor is there to address 2 primary needs: - `src/` - Rust source code - `doc/` - Documentation sources (mdbook) - `tests/` - Integration and validation tests -- `specs/` - Feature specifications and implementation plans - `playbooks/` - Operational playbooks ## Documentation @@ -63,7 +62,6 @@ mdbook serve doc/ | [API Reference](doc/api/) | REST endpoints, authentication, examples | | [Configuration](doc/configuration/) | Config schema, examples, validation | | [Integration](doc/integration/) | TSDB interface, adding new backends | -| [Modules](doc/modules/) | Rust module documentation | | [Guides](doc/guides/) | Troubleshooting, deployment | | [Testing](doc/testing.md) | Testing guide, fixtures, coverage | diff --git a/doc/SUMMARY.md b/doc/SUMMARY.md index f58cc1f..9e50c3b 100644 --- a/doc/SUMMARY.md +++ b/doc/SUMMARY.md @@ -36,14 +36,6 @@ - [Graphite Backend](integration/graphite.md) - [Adding New Backends](integration/adding-backends.md) -# Modules -- [Overview](modules/overview.md) -- [API Module](modules/api.md) -- [Config Module](modules/config.md) -- [Types Module](modules/types.md) -- [Graphite Module](modules/graphite.md) -- [Common Module](modules/common.md) - # Operational Guides - [Troubleshooting](guides/troubleshooting.md) - [Deployment](guides/deployment.md) diff --git a/doc/getting-started/project-structure.md b/doc/getting-started/project-structure.md index 51c639a..f97f7f8 100644 --- a/doc/getting-started/project-structure.md +++ b/doc/getting-started/project-structure.md @@ -27,12 +27,9 @@ metrics-processor/ │ ├── api/ # API reference │ ├── configuration/ # Configuration reference │ ├── integration/ # TSDB integration guides -│ ├── modules/ # Rust module documentation │ └── guides/ # Operational guides ├── tests/ # Integration and documentation tests │ └── documentation_validation.rs # Documentation example validation -├── specs/ # Feature specifications -│ └── 001-project-documentation/ # This feature's design docs ├── Cargo.toml # Rust dependencies and project metadata ├── build.rs # Build script (generates JSON schemas) ├── openapi-schema.yaml # OpenAPI 3.0 API specification @@ -245,4 +242,4 @@ All I/O operations use Tokio async runtime: - **New to the codebase?** Start with [Quickstart Guide](quickstart.md) - **Want to contribute?** Read [Development Workflow](development.md) - **Need to understand architecture?** See [Architecture Overview](../architecture/overview.md) -- **Adding a feature?** Check [Module Documentation](../modules/overview.md) for responsibilities +- **Adding a feature?** Check [Architecture Overview](../architecture/overview.md) for responsibilities diff --git a/doc/getting-started/quickstart.md b/doc/getting-started/quickstart.md index f7bbeae..3e89991 100644 --- a/doc/getting-started/quickstart.md +++ b/doc/getting-started/quickstart.md @@ -297,11 +297,11 @@ Now that you have a working environment, explore these documentation sections: | **Architecture** | System design, data flow | `doc/architecture/` | | **API Reference** | Endpoint details, authentication | `doc/api/` | | **Configuration** | All config fields, examples | `doc/configuration/` | -| **Module Docs** | Rust module responsibilities | `doc/modules/` | +| **Code Docs** | Generated module documentation | `cargo doc --open` | | **Troubleshooting** | Common issues, solutions | `doc/guides/troubleshooting.md` | **Next Steps**: -- Read [Architecture Overview](../../../doc/architecture/overview.md) to understand component interactions +- Read [Architecture Overview](../architecture/overview.md) to understand component interactions - Review [Configuration Schema](./contracts/config-schema.json) for full config reference - Check [patterns.json](./contracts/patterns.json) for coding conventions @@ -385,7 +385,7 @@ You've successfully completed onboarding if you can: ## Getting Help -- **Code Questions**: Check `doc/modules/` for module-specific documentation +- **Code Questions**: Run `cargo doc --open` for the generated module documentation - **Configuration Issues**: See `doc/configuration/schema.md` for field reference - **Architecture Questions**: Read `doc/architecture/overview.md` - **Bugs**: Check `doc/guides/troubleshooting.md` first, then file an issue diff --git a/doc/modules/api.md b/doc/modules/api.md deleted file mode 100644 index 0e66a8d..0000000 --- a/doc/modules/api.md +++ /dev/null @@ -1,140 +0,0 @@ -# API Module - -The API module (`src/api.rs` and `src/api/v1.rs`) provides HTTP endpoints for the metrics-processor service using the Axum web framework. - -## Module Structure - -``` -src/api.rs # Module declaration -src/api/v1.rs # V1 API implementation -``` - -## API v1 Routes - -The `get_v1_routes()` function constructs the router: - -```rust -pub fn get_v1_routes() -> Router { - return Router::new() - .route("/", get(root)) - .route("/info", get(info)) - .route("/health", get(handler_health)); -} -``` - -### Endpoints - -| Method | Path | Handler | Description | -|--------|------|---------|-------------| -| GET | `/v1/` | `root` | Returns API name | -| GET | `/v1/info` | `info` | Returns API info text | -| GET | `/v1/health` | `handler_health` | Returns service health metrics | - -## Request/Response Types - -### HealthQuery - -Query parameters for the `/health` endpoint: - -```rust -#[derive(Debug, Deserialize)] -pub struct HealthQuery { - /// Start point to query metrics (RFC3339 timestamp) - pub from: String, - /// End point to query metrics (RFC3339 timestamp) - pub to: String, - /// Maximum data points to return (default: 100) - #[serde(default = "default_max_data_points")] - pub max_data_points: u32, - /// Service name to query - pub service: String, - /// Environment name - pub environment: String, -} -``` - -**Example request:** -``` -GET /v1/health?from=2024-01-01T00:00:00Z&to=2024-01-02T00:00:00Z&service=srvA&environment=env1 -``` - -### ServiceHealthResponse - -Response structure for the `/health` endpoint: - -```rust -#[derive(Debug, Serialize, Deserialize)] -pub struct ServiceHealthResponse { - /// Service name - pub name: String, - /// Service category (e.g., "compute", "storage") - pub service_category: String, - /// Environment name - pub environment: String, - /// Health metric data points: Vec<(timestamp, health_value)> - pub metrics: ServiceHealthData, -} -``` - -**Example response:** -```json -{ - "name": "srvA", - "service_category": "compute", - "environment": "env1", - "metrics": [[1704067200, 1], [1704070800, 0]] -} -``` - -## Handler Implementation - -### handler_health - -The main health endpoint handler: - -```rust -pub async fn handler_health( - query: Query, - State(state): State -) -> Response { - // 1. Look up service in health_metrics config - // 2. Call get_service_health() from common module - // 3. Return ServiceHealthResponse or error -} -``` - -**Error Responses:** - -| Status | Condition | -|--------|-----------| -| 200 OK | Success | -| 409 Conflict | Service or environment not supported | -| 500 Internal Server Error | Expression evaluation or Graphite error | - -## State Management - -All handlers receive `AppState` via Axum's state extraction: - -```rust -State(state): State -``` - -The `AppState` contains: -- Processed configuration -- Pre-computed metric templates -- HTTP client for Graphite queries -- Flag and health metric definitions - -## Authentication - -Currently, the API does not implement authentication. The `status_dashboard.secret` configuration option suggests JWT-based authentication may be planned for integration with status dashboard services. - -## Integration with Other Modules - -``` -api::v1 - │ - ├──► types::AppState (state extraction) - ├──► common::get_service_health() (business logic) - └──► types::CloudMonError (error handling) -``` diff --git a/doc/modules/common.md b/doc/modules/common.md deleted file mode 100644 index 4c4f73d..0000000 --- a/doc/modules/common.md +++ /dev/null @@ -1,229 +0,0 @@ -# Common Module - -The common module (`src/common.rs`) provides shared utility functions for metric evaluation and service health calculation. - -## Functions - -### get_metric_flag_state() - -Converts a raw metric value to a boolean flag based on the configured threshold and comparison operator. - -```rust -pub fn get_metric_flag_state(value: &Option, metric: &FlagMetric) -> bool -``` - -**Parameters:** -- `value` - Raw metric value from Graphite (may be `None` for missing data) -- `metric` - Flag metric configuration with operator and threshold - -**Logic:** -```rust -return match *value { - Some(x) => match metric.op { - CmpType::Lt => x < metric.threshold, // value < threshold = healthy - CmpType::Gt => x > metric.threshold, // value > threshold = healthy - CmpType::Eq => x == metric.threshold, // value == threshold = healthy - }, - None => false, // Missing data = unhealthy -}; -``` - -**Usage Example:** -```rust -let metric = FlagMetric { - query: "service.latency".to_string(), - op: CmpType::Lt, - threshold: 1000.0, -}; - -// Latency of 500ms is healthy (500 < 1000) -assert!(get_metric_flag_state(&Some(500.0), &metric)); - -// Latency of 1500ms is unhealthy (1500 >= 1000) -assert!(!get_metric_flag_state(&Some(1500.0), &metric)); - -// Missing data is unhealthy -assert!(!get_metric_flag_state(&None, &metric)); -``` - -### get_service_health() - -Calculates aggregated health scores for a service based on multiple flag metrics and boolean expressions. - -```rust -pub async fn get_service_health( - state: &AppState, - service: &str, - environment: &str, - from: &str, - to: &str, - max_data_points: u16, -) -> Result -``` - -**Parameters:** -- `state` - Application state containing configurations and HTTP client -- `service` - Service name to evaluate -- `environment` - Environment name -- `from` / `to` - Time range (RFC3339 format or Graphite relative time) -- `max_data_points` - Maximum data points to return - -**Returns:** -- `Ok(ServiceHealthData)` - Vector of `(timestamp, health_score)` tuples -- `Err(CloudMonError)` - On service not found, environment not supported, or evaluation error - -## Health Calculation Algorithm - -### Step 1: Validate Service - -```rust -if !state.health_metrics.contains_key(service) { - return Err(CloudMonError::ServiceNotSupported); -} -``` - -### Step 2: Build Graphite Query Map - -```rust -let mut graphite_targets: HashMap = HashMap::new(); -for metric_name in metric_names.iter() { - if let Some(metric) = state.flag_metrics.get(metric_name) { - match metric.get(environment) { - Some(m) => { - graphite_targets.insert(metric_name.clone(), m.query.clone()); - } - _ => return Err(CloudMonError::EnvNotSupported), - }; - } -} -``` - -### Step 3: Fetch Data from Graphite - -```rust -let raw_data: Vec = graphite::get_graphite_data( - &state.req_client, - &state.config.datasource.url, - &graphite_targets, - // ... time parameters -).await?; -``` - -### Step 4: Organize Data by Timestamp - -```rust -let mut metrics_map: BTreeMap> = BTreeMap::new(); - -for data_element in raw_data.iter() { - let metric = metric_cfg.get(environment).unwrap(); - for (val, ts) in data_element.datapoints.iter() { - metrics_map.entry(*ts).or_insert(HashMap::new()).insert( - data_element.target.clone(), - get_metric_flag_state(val, metric), - ); - } -} -``` - -### Step 5: Evaluate Health Expressions - -Uses the `evalexpr` crate for boolean expression evaluation: - -```rust -for (ts, ts_val) in metrics_map.iter() { - let mut context = HashMapContext::new(); - - // Build context with all metric flags - for metric in hm_config.metrics.iter() { - let xval = ts_val.get(metric).unwrap_or(&false); - context.set_value( - metric.replace("-", "_").into(), // evalexpr doesn't support "-" - Value::from(*xval) - ).unwrap(); - } - - // Evaluate expressions in order of weight - let mut expression_res: u8 = 0; - for expr in hm_config.expressions.iter() { - if expr.weight as u8 <= expression_res { - continue; // Skip lower-weight expressions - } - if eval_boolean_with_context(expr.expression.as_str(), &context)? { - expression_res = expr.weight as u8; - } - } - - result.push((*ts, expression_res)); -} -``` - -## Expression Evaluation - -### Supported Operators - -The `evalexpr` crate supports standard boolean operators: -- `&&` - Logical AND -- `||` - Logical OR -- `!` - Logical NOT -- Parentheses for grouping - -### Example Expressions - -```yaml -expressions: - # Both metrics must be healthy - - expression: 'api_latency && availability' - weight: 1 - - # Either metric being unhealthy triggers warning - - expression: '!api_latency || !availability' - weight: 2 - - # Complex conditions - - expression: '(api_latency && availability) || backup_service' - weight: 1 -``` - -### Metric Name Handling - -Metric names with hyphens are automatically converted: -- Config: `service.metric-name` -- Expression context: `service.metric_name` - -```rust -context.set_value( - metric.replace("-", "_").into(), - Value::from(xval) -) -``` - -## Data Flow - -``` -get_service_health() - │ - ├──► Validate service exists in health_metrics - │ - ├──► Build target->query map from flag_metrics - │ - ├──► graphite::get_graphite_data() - │ │ - │ ▼ - │ Upstream Graphite TSDB - │ - ├──► Reorganize: [target, [(val, ts)]] → {ts: {target: bool}} - │ - ├──► For each timestamp: - │ ├──► Build evalexpr context - │ ├──► Evaluate expressions by weight - │ └──► Record highest matching weight - │ - └──► Return Vec<(timestamp, health_score)> -``` - -## Dependencies - -- `evalexpr` - Boolean expression evaluation -- `chrono` - Date/time parsing -- `crate::graphite` - Graphite data fetching -- `crate::types` - Core types (`AppState`, `FlagMetric`, etc.) diff --git a/doc/modules/config.md b/doc/modules/config.md deleted file mode 100644 index c8b9962..0000000 --- a/doc/modules/config.md +++ /dev/null @@ -1,215 +0,0 @@ -# Configuration Module - -The configuration module (`src/config.rs`) handles loading, parsing, and merging configuration from YAML files and environment variables. - -## Key Types - -### Config - -The main configuration structure: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct Config { - /// Datasource connection (Graphite TSDB) - pub datasource: Datasource, - /// Server binding configuration - pub server: ServerConf, - /// Metric templates for reuse across metrics - pub metric_templates: Option>, - /// Environment definitions - pub environments: Vec, - /// Flag metric definitions - pub flag_metrics: Vec, - /// Health metric definitions keyed by service name - pub health_metrics: HashMap, - /// Status Dashboard integration config - pub status_dashboard: Option, -} -``` - -### Datasource - -TSDB connection settings: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct Datasource { - /// Graphite URL (e.g., "https://graphite.example.com") - pub url: String, - /// Query timeout in seconds (default: 10) - #[serde(default = "default_timeout")] - pub timeout: u16, -} -``` - -### ServerConf - -HTTP server binding: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct ServerConf { - /// IP address to bind (default: "0.0.0.0") - #[serde(default = "default_address")] - pub address: String, - /// Port to bind (default: 3000) - #[serde(default = "default_port")] - pub port: u16, -} -``` - -### StatusDashboardConfig - -Optional status dashboard integration: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct StatusDashboardConfig { - /// Status dashboard URL - pub url: String, - /// JWT token signature secret - pub secret: Option, -} -``` - -## Configuration Loading - -### Primary Method: `Config::new()` - -```rust -pub fn new(config_file: &str) -> Result -``` - -Loading process: -1. **Load main config file** - YAML file at specified path -2. **Merge conf.d parts** - All `*.yaml` files in `{config_dir}/conf.d/` -3. **Merge environment variables** - Variables prefixed with `MP_` - -### Environment Variable Merging - -Environment variables use `MP_` prefix with `__` as sublevel separator: - -| Environment Variable | Config Path | -|---------------------|-------------| -| `MP_DATASOURCE__URL` | `datasource.url` | -| `MP_SERVER__PORT` | `server.port` | -| `MP_STATUS_DASHBOARD__SECRET` | `status_dashboard.secret` | - -```rust -Environment::with_prefix("MP") - .prefix_separator("_") - .separator("__") -``` - -### Alternative: `Config::from_config_str()` - -For testing, load config from a string: - -```rust -#[allow(dead_code)] -pub fn from_config_str(data: &str) -> Self -``` - -## Configuration File Format - -### Example Configuration - -```yaml ---- -datasource: - url: 'https://graphite.example.com' - timeout: 30 - -server: - address: '0.0.0.0' - port: 3005 - -metric_templates: - api_latency: - query: 'summarize($environment.$service.latency, "1h", "avg")' - op: lt - threshold: 1000 - -environments: - - name: production - attributes: - region: eu-de - - name: staging - -flag_metrics: - - name: api-latency - service: compute - template: - name: api_latency - environments: - - name: production - threshold: 500 - - name: staging - -health_metrics: - compute: - service: compute - category: compute - component_name: "Compute Service" - metrics: - - compute.api-latency - - compute.availability - expressions: - - expression: 'compute.api_latency && compute.availability' - weight: 1 - - expression: 'compute.api_latency || compute.availability' - weight: 2 - -status_dashboard: - url: 'https://status.example.com' - secret: ${MP_STATUS_DASHBOARD__SECRET} -``` - -### Modular Configuration (conf.d) - -Split configuration into multiple files: - -``` -config/ -├── config.yaml # Main config -└── conf.d/ - ├── compute.yaml # Compute service metrics - ├── storage.yaml # Storage service metrics - └── network.yaml # Network service metrics -``` - -Each conf.d file can contain partial configuration that gets merged. - -## Helper Methods - -### `get_socket_addr()` - -Returns a `SocketAddr` for server binding: - -```rust -pub fn get_socket_addr(&self) -> SocketAddr { - SocketAddr::from(( - self.server.address.as_str().parse::().unwrap(), - self.server.port, - )) -} -``` - -## Default Values - -| Field | Default | -|-------|---------| -| `server.address` | `"0.0.0.0"` | -| `server.port` | `3000` | -| `datasource.timeout` | `10` | - -## Validation - -Configuration validation happens during deserialization. Missing required fields or type mismatches will cause `ConfigError` to be returned. - -## Dependencies - -- `config` crate - Configuration loading and merging -- `serde` - Deserialization -- `glob` - Finding conf.d files diff --git a/doc/modules/graphite.md b/doc/modules/graphite.md deleted file mode 100644 index 69f6e3b..0000000 --- a/doc/modules/graphite.md +++ /dev/null @@ -1,250 +0,0 @@ -# Graphite Module - -The Graphite module (`src/graphite.rs`) provides a Graphite TSDB-compatible API interface, enabling integration with Grafana and other Graphite-compatible tools. - -## Overview - -This module implements: -- Graphite render API for time series data -- Metrics discovery API (`/metrics/find`) -- Grafana-compatible endpoints - -## Key Types - -### GraphiteData - -Response structure from Graphite queries: - -```rust -#[derive(Deserialize, Serialize, Debug)] -pub struct GraphiteData { - /// Metric target name - pub target: String, - /// Array of (value, timestamp) tuples - pub datapoints: Vec<(Option, u32)>, -} -``` - -### MetricsQuery - -Query parameters for metrics discovery: - -```rust -#[derive(Debug, Deserialize)] -pub struct MetricsQuery { - /// Query pattern (e.g., "flag.*", "health.env1.*") - pub query: String, - /// Optional start time - pub from: Option, - /// Optional end time - pub until: Option, -} -``` - -### Metric - -Metric metadata for discovery responses: - -```rust -#[derive(Debug, Eq, Ord, PartialEq, PartialOrd, Serialize)] -pub struct Metric { - #[serde(rename(serialize = "allowChildren"))] - pub allow_children: u8, - pub expandable: u8, - pub leaf: u8, - pub id: String, - pub text: String, -} -``` - -### RenderRequest - -Parameters for render API: - -```rust -#[derive(Default, Debug, Deserialize)] -pub struct RenderRequest { - /// Target metric path - pub target: Option, - /// Start time - pub from: Option, - /// End time - pub until: Option, - /// Maximum data points to return - #[serde(rename(deserialize = "maxDataPoints"))] - pub max_data_points: Option, -} -``` - -## Routes - -```rust -pub fn get_graphite_routes() -> Router { - Router::new() - .route("/functions", get(handler_functions)) - .route("/metrics/find", get(handler_metrics_find_get).post(handler_metrics_find_post)) - .route("/render", get(handler_render).post(handler_render)) - .route("/tags/autoComplete/tags", get(handler_tags)) -} -``` - -| Method | Path | Handler | Description | -|--------|------|---------|-------------| -| GET | `/functions` | `handler_functions` | Returns empty object (Grafana compatibility) | -| GET/POST | `/metrics/find` | `handler_metrics_find_*` | Discover available metrics | -| GET/POST | `/render` | `handler_render` | Render time series data | -| GET | `/tags/autoComplete/tags` | `handler_tags` | Returns empty array (Grafana compatibility) | - -## Metric Discovery - -### Virtual Metric Hierarchy - -The module exposes a virtual metric hierarchy: - -``` -├── flag -│ └── {environment} -│ └── {service} -│ └── {metric_name} -└── health - └── {environment} - └── {service_name} -``` - -### find_metrics() Function - -```rust -pub fn find_metrics(find_request: MetricsQuery, state: AppState) -> Vec -``` - -Query patterns: -- `*` - Returns top-level: `["flag", "health"]` -- `flag.*` or `health.*` - Returns environments -- `flag.{env}.*` - Returns services -- `flag.{env}.{service}.*` - Returns metric names -- `health.{env}.*` - Returns health metric names - -## Render API - -### handler_render - -Handles both GET and POST requests for time series data. - -**Flag Metrics** (`flag.{env}.{service}.{metric}`): -1. Looks up metric configuration -2. Queries upstream Graphite with resolved query -3. Converts raw values to binary flags (0/1) based on threshold - -**Health Metrics** (`health.{env}.{service}`): -1. Calls `get_service_health()` from common module -2. Returns aggregated health scores - -### Response Format - -```json -[ - { - "target": "service.metric-name", - "datapoints": [[1.0, 1704067200], [0.0, 1704070800]] - } -] -``` - -## Graphite Client - -### get_graphite_data() - -Core function for querying upstream Graphite: - -```rust -pub async fn get_graphite_data( - client: &reqwest::Client, - url: &str, - targets: &HashMap, // alias -> query - from: Option>, - from_raw: Option, - to: Option>, - to_raw: Option, - max_data_points: u16, -) -> Result, CloudMonError> -``` - -**Query Construction:** -```rust -let query_params: Vec<(_, String)> = [ - ("format", "json".to_string()), - ("maxDataPoints", max_data_points.to_string()), -].into(); -// Add from, until, and target parameters -``` - -**Query Aliasing:** -```rust -fn alias_graphite_query(query: &str, alias: &str) -> String { - format!("alias({},'{}')", query, alias) -} -``` - -This wraps each query with Graphite's `alias()` function to preserve the logical metric name in responses. - -## Request Extraction - -### JsonOrForm Extractor - -Custom Axum extractor that accepts both JSON and form-encoded bodies: - -```rust -#[derive(Default, Debug)] -pub struct JsonOrForm(T); - -#[async_trait] -impl FromRequest for JsonOrForm -where - // ... constraints -{ - async fn from_request(req: Request, _state: &S) -> Result { - let content_type = req.headers().get(CONTENT_TYPE); - - if content_type.starts_with("application/json") { - // Extract as JSON - } - if content_type.starts_with("application/x-www-form-urlencoded") { - // Extract as Form - } - - Err(StatusCode::UNSUPPORTED_MEDIA_TYPE) - } -} -``` - -This enables compatibility with various Graphite clients (Grafana uses form encoding). - -## Integration Flow - -``` -Grafana/Client - │ - ▼ -┌─────────────────┐ -│ /metrics/find │──► find_metrics() ──► AppState.flag_metrics -└─────────────────┘ AppState.health_metrics - AppState.environments - │ - ▼ -┌─────────────────┐ -│ /render │──► handler_render() -└─────────────────┘ - │ - ├──► Flag: get_graphite_data() ──► Upstream Graphite - │ │ - │ ▼ - │ Convert to 0/1 flags - │ - └──► Health: get_service_health() ──► Aggregate expressions -``` - -## Error Handling - -- Returns empty array `[]` for unrecognized metric paths -- Returns `CloudMonError::GraphiteError` for upstream failures -- Logs warnings for unknown targets in responses diff --git a/doc/modules/overview.md b/doc/modules/overview.md deleted file mode 100644 index ddd708b..0000000 --- a/doc/modules/overview.md +++ /dev/null @@ -1,101 +0,0 @@ -# Module Overview - -This document provides a high-level overview of the metrics-processor crate modules, their responsibilities, and relationships. - -## Module Responsibility Matrix - -| Module | Primary Responsibility | Key Types | Dependencies | Used By | -|--------|----------------------|-----------|--------------|---------| -| `lib` | Crate entry point | - | `api`, `common`, `config`, `graphite`, `sd`, `types` | External consumers | -| `api` | HTTP API routing | - | `api::v1` | `main` binary | -| `api::v1` | V1 REST endpoints | `HealthQuery`, `ServiceHealthResponse` | `common`, `types` | `api` | -| `config` | Configuration parsing | `Config`, `Datasource`, `ServerConf` | `types` | `types`, `main` | -| `types` | Core data structures | `AppState`, `FlagMetric`, `ServiceHealthDef` | `config` | All modules | -| `graphite` | Graphite TSDB interface | `GraphiteData`, `Metric`, `RenderRequest` | `common`, `types` | `common`, `api::v1` | -| `common` | Shared utilities | - | `types`, `graphite` | `api::v1`, `graphite` | -| `sd` | Status Dashboard API | `IncidentData`, `ComponentCache`, `StatusDashboardComponent` | `anyhow`, `hmac`, `jwt` | `reporter` binary | - -## Architecture Diagram - -``` -┌─────────────────────────────────────────────────────────────┐ -│ convertor binary │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ config │ -│ (Config, Datasource, ServerConf) │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ types │ -│ (AppState, FlagMetric, ServiceHealthDef, CloudMonError) │ -└─────────────────────────────────────────────────────────────┘ - │ - ┌───────────────────┼───────────────────┐ - ▼ ▼ ▼ -┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ -│ api::v1 │ │ graphite │ │ common │ -│ (REST API v1) │ │ (TSDB Client) │ │ (Utilities) │ -└─────────────────┘ └─────────────────┘ └─────────────────┘ - │ │ │ - └───────────────────┴───────────────────┘ - │ - ▼ - ┌─────────────────┐ - │ HTTP Server │ - │ (Axum) │ - └─────────────────┘ - -┌─────────────────────────────────────────────────────────────┐ -│ reporter binary │ -└─────────────────────────────────────────────────────────────┘ - │ - ┌───────────────────┼───────────────────┐ - ▼ ▼ ▼ -┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ -│ config │ │ sd │ │ api::v1 │ -│ (Config load) │ │ (Status Dash) │ │ (Query health) │ -└─────────────────┘ └─────────────────┘ └─────────────────┘ - │ - ▼ - ┌─────────────────┐ - │ Status Dashboard│ - │ V2 API │ - └─────────────────┘ -``` - -## Module Summaries - -### `lib.rs` -Entry point for the crate, re-exports all public modules: -- `api` - HTTP API handlers -- `common` - Shared utilities -- `config` - Configuration management -- `graphite` - Graphite TSDB communication -- `types` - Core type definitions - -### `api` / `api::v1` -Axum-based HTTP handlers for the REST API. Provides health metrics endpoints. - -### `config` -YAML configuration loading with environment variable merging. Supports `conf.d` style modular configuration. - -### `types` -Core domain types including metric definitions, application state, and error types. - -### `graphite` -Graphite TSDB client implementing the render and metrics/find APIs for Grafana compatibility. - -### `common` -Shared business logic for metric flag evaluation and service health calculation. - -## Data Flow - -1. **Startup**: `config::Config::new()` loads YAML + env vars -2. **Initialization**: `types::AppState::new()` builds runtime state with processed metrics -3. **Request Handling**: Axum routes to `api::v1` or `graphite` handlers -4. **Data Retrieval**: Handlers call `common` utilities which query `graphite` module -5. **Response**: Results transformed and returned as JSON diff --git a/doc/modules/sd.md b/doc/modules/sd.md deleted file mode 100644 index 5227292..0000000 --- a/doc/modules/sd.md +++ /dev/null @@ -1,231 +0,0 @@ -# Status Dashboard Module (`sd`) - -The `sd` module provides all functionality for integrating with the Status Dashboard API, including component management, incident creation, cache operations, and authentication. - -## Module Location - -- **Source**: `src/sd.rs` -- **Public export**: `cloudmon_metrics::sd` - -## Overview - -This module consolidates all Status Dashboard V2 API integration logic in one place, providing: - -- Component fetching and caching -- Component ID resolution with subset attribute matching -- Incident creation with static payloads -- HMAC-JWT authentication - -## Data Structures - -### ComponentAttribute - -Key-value pair for identifying components: - -```rust -#[derive(Clone, Deserialize, Serialize, Debug, PartialEq, Eq, Hash, Ord, PartialOrd)] -pub struct ComponentAttribute { - pub name: String, - pub value: String, -} -``` - -Derives `Ord` and `PartialOrd` for deterministic sorting in cache keys. - -### Component - -Component definition from configuration: - -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct Component { - pub name: String, - pub attributes: Vec, -} -``` - -### StatusDashboardComponent - -API response from `GET /v2/components`: - -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct StatusDashboardComponent { - pub id: u32, - pub name: String, - #[serde(default)] - pub attributes: Vec, -} -``` - -### IncidentData - -API request for `POST /v2/events`: - -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentData { - pub title: String, - pub description: String, - pub impact: u8, - pub components: Vec, - pub start_date: String, // RFC3339 format - pub system: bool, - #[serde(rename = "type")] - pub incident_type: String, -} -``` - -### ComponentCache - -Type alias for the component ID cache: - -```rust -pub type ComponentCache = HashMap<(String, Vec), u32>; -``` - -Key: `(component_name, sorted_attributes)` → Value: `component_id` - -## Functions - -### Authentication - -#### `build_auth_headers` - -```rust -pub fn build_auth_headers(secret: Option<&str>) -> HeaderMap -``` - -Generates HMAC-JWT authorization headers for Status Dashboard API. - -- Creates Bearer token using HMAC-SHA256 signing -- Returns empty HeaderMap if no secret provided (optional auth) - -**Example**: -```rust -let headers = build_auth_headers(Some("my-secret")); -// Headers contain: Authorization: Bearer eyJ... -``` - -### Component Management - -#### `fetch_components` - -```rust -pub async fn fetch_components( - client: &reqwest::Client, - base_url: &str, - headers: &HeaderMap, -) -> anyhow::Result> -``` - -Fetches all components from Status Dashboard API V2 (`GET /v2/components`). - -#### `build_component_id_cache` - -```rust -pub fn build_component_id_cache( - components: Vec -) -> ComponentCache -``` - -Builds component ID cache from fetched components. Sorts attributes for deterministic cache keys. - -#### `find_component_id` - -```rust -pub fn find_component_id( - cache: &ComponentCache, - target: &Component -) -> Option -``` - -Finds component ID in cache with **subset attribute matching**: -- Config attributes must be a subset of cache attributes -- Example: config `{region: "EU-DE"}` matches cache `{region: "EU-DE", category: "Storage"}` - -### Incident Management - -#### `build_incident_data` - -```rust -pub fn build_incident_data( - component_id: u32, - impact: u8, - timestamp: i64 -) -> IncidentData -``` - -Builds incident data structure for V2 API: -- **Static title**: "System incident from monitoring system" -- **Static description**: "System-wide incident affecting one or multiple components. Created automatically." -- **Timestamp**: RFC3339 format, minus 1 second from input -- **system**: true (indicates auto-generated) - -#### `create_incident` - -```rust -pub async fn create_incident( - client: &reqwest::Client, - base_url: &str, - headers: &HeaderMap, - incident_data: &IncidentData, -) -> anyhow::Result<()> -``` - -Creates incident via Status Dashboard API V2 (`POST /v2/events`). - -## Usage Example - -```rust -use cloudmon_metrics::sd::{ - build_auth_headers, build_component_id_cache, build_incident_data, - create_incident, fetch_components, find_component_id, - Component, ComponentAttribute, -}; - -// Build auth headers -let headers = build_auth_headers(config.secret.as_deref()); - -// Fetch and cache components -let components = fetch_components(&client, &url, &headers).await?; -let cache = build_component_id_cache(components); - -// Find component ID -let target = Component { - name: "Object Storage Service".to_string(), - attributes: vec![ComponentAttribute { - name: "region".to_string(), - value: "EU-DE".to_string(), - }], -}; -let component_id = find_component_id(&cache, &target)?; - -// Create incident -let incident = build_incident_data(component_id, 2, timestamp); -create_incident(&client, &url, &headers, &incident).await?; -``` - -## Testing - -Integration tests are in `tests/integration_sd.rs`: - -```bash -cargo test --test integration_sd -``` - -**Test coverage**: -- `test_fetch_components_success` - API fetching -- `test_build_component_id_cache` - Cache structure -- `test_find_component_id_subset_matching` - Subset matching logic -- `test_build_incident_data_structure` - Static payload generation -- `test_timestamp_rfc3339_minus_one_second` - Timestamp handling -- `test_create_incident_success` - API posting -- `test_build_auth_headers` - JWT generation -- Additional edge case tests - -## Related Documentation - -- [Reporter Overview](../reporter.md) - How reporter uses this module -- [API Contracts](../../specs/003-sd-api-v2-migration/contracts/) - V2 API specifications -- [Spec](../../specs/003-sd-api-v2-migration/spec.md) - Feature specification diff --git a/doc/modules/types.md b/doc/modules/types.md deleted file mode 100644 index 350a801..0000000 --- a/doc/modules/types.md +++ /dev/null @@ -1,308 +0,0 @@ -# Types Module - -The types module (`src/types.rs`) defines the core data structures used throughout the metrics-processor application. - -## Metric Comparison Types - -### CmpType - -Enum for metric comparison operations: - -```rust -#[derive(Clone, Debug, Deserialize, PartialEq)] -#[serde(rename_all = "lowercase")] -pub enum CmpType { - Lt, // Less than - Gt, // Greater than - Eq, // Equal to -} -``` - -Used to determine when a metric value should be flagged as "healthy" or "unhealthy". - -## Metric Definition Types - -### BinaryMetricRawDef - -Raw metric template definition (used in `metric_templates`): - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct BinaryMetricRawDef { - /// Graphite query template (supports $var substitution) - pub query: String, - /// Comparison operator - pub op: CmpType, - /// Threshold value for comparison - pub threshold: f32, -} -``` - -### BinaryMetricDef - -Metric definition with optional template reference: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct BinaryMetricDef { - pub query: Option, - pub op: Option, - pub threshold: Option, - pub template: Option, -} -``` - -### MetricTemplateRef - -Reference to a named template: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct MetricTemplateRef { - /// Template name (key in metric_templates) - pub name: String, - /// Optional variable substitutions - pub vars: Option>, -} -``` - -### FlagMetricDef - -Configuration definition for a flag metric: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct FlagMetricDef { - /// Metric name - pub name: String, - /// Service this metric belongs to - pub service: String, - /// Template reference - pub template: Option, - /// Per-environment overrides - pub environments: Vec, -} -``` - -### FlagMetric - -Processed/resolved flag metric (runtime): - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct FlagMetric { - /// Resolved Graphite query - pub query: String, - /// Comparison operator - pub op: CmpType, - /// Threshold value - pub threshold: f32, -} -``` - -### MetricEnvironmentDef - -Per-environment metric overrides: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct MetricEnvironmentDef { - /// Environment name - pub name: String, - /// Optional threshold override - pub threshold: Option, -} -``` - -## Environment Types - -### EnvironmentDef - -Environment definition: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct EnvironmentDef { - /// Environment name (e.g., "production", "staging") - pub name: String, - /// Optional attributes for template substitution - pub attributes: Option>, -} -``` - -## Health Metric Types - -### ServiceHealthDef - -Service health configuration: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct ServiceHealthDef { - /// Service identifier - pub service: String, - /// Optional display name - pub component_name: Option, - /// Category (e.g., "compute", "storage", "network") - pub category: String, - /// List of flag metric names to evaluate - pub metrics: Vec, - /// Expressions for health calculation - pub expressions: Vec, -} -``` - -### MetricExpressionDef - -Boolean expression for health evaluation: - -```rust -#[derive(Clone, Debug, Deserialize)] -pub struct MetricExpressionDef { - /// Boolean expression (e.g., "metric_a && metric_b") - pub expression: String, - /// Weight/severity (higher = more severe) - pub weight: i32, -} -``` - -## Data Types - -### MetricPoints / MetricData - -Time series data structures: - -```rust -/// Timestamp -> boolean flag mapping -pub type MetricPoints = BTreeMap; - -#[derive(Clone, Debug, Deserialize, Serialize)] -pub struct MetricData { - pub target: String, - #[serde(rename(serialize = "datapoints"))] - pub points: MetricPoints, -} - -/// Service health time series: Vec<(timestamp, health_value)> -pub type ServiceHealthData = Vec<(u32, u8)>; -``` - -## Error Types - -### CloudMonError - -Application-specific error enum: - -```rust -pub enum CloudMonError { - ServiceNotSupported, // Requested service not found - EnvNotSupported, // Environment not configured for service - ExpressionError, // Boolean expression evaluation failed - GraphiteError, // TSDB communication error -} -``` - -Implements `std::error::Error`, `Display`, and `Debug`. - -## Application State - -### AppState - -Central application state shared across handlers: - -```rust -#[derive(Clone)] -pub struct AppState { - /// Original configuration - pub config: Config, - /// Processed metric templates - pub metric_templates: HashMap, - /// HTTP client for Graphite queries - pub req_client: reqwest::Client, - /// Processed flag metrics: service.metric -> env -> FlagMetric - pub flag_metrics: HashMap>, - /// Health metric definitions by service name - pub health_metrics: HashMap, - /// Environment definitions - pub environments: Vec, - /// Set of known service names - pub services: HashSet, -} -``` - -### AppState Methods - -#### `new(config: Config) -> Self` - -Creates new state with configured HTTP client: - -```rust -impl AppState { - pub fn new(config: Config) -> Self { - let timeout = Duration::from_secs(config.datasource.timeout as u64); - Self { - config: config, - metric_templates: HashMap::new(), - flag_metrics: HashMap::new(), - req_client: ClientBuilder::new().timeout(timeout).build().unwrap(), - health_metrics: HashMap::new(), - environments: Vec::new(), - services: HashSet::new(), - } - } -} -``` - -#### `process_config(&mut self)` - -Processes configuration into runtime structures: - -1. **Copies metric templates** from config -2. **Resolves flag metrics** - substitutes `$service` and `$environment` variables in queries -3. **Processes health metrics** - replaces `-` with `_` in expression metric names (for evalexpr compatibility) -4. **Populates services set** for discovery endpoints - -```rust -pub fn process_config(&mut self) { - // Template variable substitution uses $var syntax - let custom_regex = Regex::new(r"(?mi)\$([^\.]+)").unwrap(); - - // Process flag_metrics with template resolution - for metric_def in self.config.flag_metrics.iter() { - // ... resolves template, substitutes variables - } - - // Process health_metrics, fixing expression syntax - for (metric_name, health_def) in self.config.health_metrics.iter() { - // ... replaces "-" with "_" for evalexpr - } -} -``` - -## Type Relationships - -``` -Config - │ - ├── metric_templates: HashMap - │ - ├── flag_metrics: Vec - │ │ - │ └── template: MetricTemplateRef - │ └──► BinaryMetricRawDef - │ - └── health_metrics: HashMap - │ - ├── metrics: Vec - └── expressions: Vec - - ▼ process_config() ▼ - -AppState - │ - ├── flag_metrics: HashMap> - │ │ │ - │ "service.metric" "environment" - │ - └── health_metrics: HashMap -``` diff --git a/specs/001-project-documentation/IMPLEMENTATION_REPORT.md b/specs/001-project-documentation/IMPLEMENTATION_REPORT.md deleted file mode 100644 index ec50e8f..0000000 --- a/specs/001-project-documentation/IMPLEMENTATION_REPORT.md +++ /dev/null @@ -1,273 +0,0 @@ -# Implementation Report: Project Documentation Feature - -**Feature ID**: 001-project-documentation -**Status**: ✅ COMPLETE -**Date**: January 20, 2025 -**Execution Time**: ~2 hours - -## Executive Summary - -Successfully implemented comprehensive project documentation for the metrics-processor project, delivering: -- **34 markdown files** (9,783 lines) of human-readable documentation -- **34 HTML pages** generated via mdbook -- **Auto-generated JSON schemas** for AI/IDE integration -- **Validation test framework** ensuring examples remain correct -- **Complete coverage** of architecture, API, configuration, and operations - -## Implementation Statistics - -| Metric | Value | -|--------|-------| -| Total Tasks | 71/71 (100%) | -| Documentation Files | 34 markdown files | -| Total Lines | 9,783 lines | -| HTML Pages | 34 pages | -| JSON Schemas | 2 files | -| Test Files | 1 validation suite | -| Mermaid Diagrams | 15+ diagrams | -| Git Changes | 19 files modified/added | - -## Phases Completed - -### ✅ Phase 1: Setup (4 tasks) -- Installed mdbook toolchain (mdbook, mdbook-mermaid, mdbook-linkcheck) -- Configured book.toml with preprocessors -- Added schemars dependency to Cargo.toml -- Updated SUMMARY.md structure - -### ✅ Phase 2: Foundational (7 tasks) -- Created build.rs for auto-schema generation -- Copied contracts (patterns.json, README.md) -- Created validation test framework (documentation_validation.rs) -- Enhanced doc/index.md with comprehensive overview - -### ✅ Phase 3: User Story 1 - Developer Onboarding (5 tasks) -- Migrated quickstart.md to getting-started/ -- Created project-structure.md (10,299 chars) -- Created development.md with workflow guide -- Added validation tests for examples - -### ✅ Phase 4: User Story 2 - AI Integration (5 tasks) -- Generated config-schema.json from Rust types -- Copied patterns.json for code generation -- Created schema validation tests -- Updated agent context - -### ✅ Phase 5: User Story 6 - Architecture (7 tasks) -- Created architecture/overview.md -- Created architecture/diagrams.md with Mermaid diagrams -- Created architecture/data-flow.md with sequences -- Enhanced convertor.md and reporter.md - -### ✅ Phase 6: User Story 3 - API Integration (5 tasks) -- Created api/endpoints.md from OpenAPI spec -- Created api/authentication.md for JWT -- Created api/examples.md with curl samples - -### ✅ Phase 7: User Story 4 - Configuration (9 tasks) -- Created 8 configuration documentation files -- Covered all config sections: datasource, templates, flags, health, environments -- Provided working YAML examples - -### ✅ Phase 8: User Story 5 - TSDB Integration (4 tasks) -- Created integration/interface.md -- Created integration/graphite.md -- Created integration/adding-backends.md with examples - -### ✅ Phase 9: Module Documentation (7 tasks) -- Created modules/overview.md with responsibility matrix -- Documented all Rust modules: api, config, types, graphite, common - -### ✅ Phase 10: Operational Guides (3 tasks) -- Created guides/troubleshooting.md -- Created guides/deployment.md with Kubernetes examples - -### ✅ Phase 11: Validation (6 tasks) -- Built documentation successfully (mdbook build) -- Generated 34 HTML pages -- Created validation test suite -- Verified schemas and examples - -### ✅ Phase 12: Polish (7 tasks) -- Updated SUMMARY.md with all sections -- Enhanced convertor.md and reporter.md -- Added cross-references -- Ensured consistent terminology - -## Success Criteria Achievement - -| Criterion | Target | Achievement | Status | -|-----------|--------|-------------|--------| -| SC-001: Onboarding time | <30 min | quickstart.md provides 30-min guide | ✅ | -| SC-002: AI accuracy | 90% | patterns.json + schema enable accurate gen | ✅ | -| SC-003: Support requests | Zero | Comprehensive docs for all use cases | ✅ | -| SC-004: Config success | 80% | Working examples + schema validation | ✅ | -| SC-006: API coverage | 100% | All endpoints, configs, modules documented | ✅ | -| SC-007: Examples work | 100% | Validation tests ensure correctness | ✅ | -| SC-008: Search speed | <3s | Built-in mdbook search <1s | ✅ | -| SC-009: OpenAPI sync | Zero discrepancies | API docs sourced from openapi-schema.yaml | ✅ | - -## Documentation Structure - -``` -doc/ -├── getting-started/ # Developer onboarding (3 files) -│ ├── quickstart.md -│ ├── project-structure.md -│ └── development.md -├── architecture/ # System design (3 files) -│ ├── overview.md -│ ├── diagrams.md -│ └── data-flow.md -├── api/ # API reference (3 files) -│ ├── endpoints.md -│ ├── authentication.md -│ └── examples.md -├── configuration/ # Config reference (8 files) -│ ├── overview.md -│ ├── schema.md -│ ├── datasource.md -│ ├── metric-templates.md -│ ├── flag-metrics.md -│ ├── health-metrics.md -│ ├── environments.md -│ └── examples.md -├── integration/ # TSDB backends (3 files) -│ ├── interface.md -│ ├── graphite.md -│ └── adding-backends.md -├── modules/ # Rust modules (6 files) -│ ├── overview.md -│ ├── api.md -│ ├── config.md -│ ├── types.md -│ ├── graphite.md -│ └── common.md -├── guides/ # Operations (2 files) -│ ├── troubleshooting.md -│ └── deployment.md -├── schemas/ # Machine-readable (3 files) -│ ├── config-schema.json -│ ├── patterns.json -│ └── README.md -├── convertor.md # Enhanced component doc -├── reporter.md # Enhanced component doc -└── index.md # Enhanced overview -``` - -## Key Features Delivered - -### For Human Developers -- 30-minute quickstart guide -- Complete architecture documentation with Mermaid diagrams -- API reference with curl examples -- Configuration reference with working YAML samples -- Troubleshooting guide with common issues -- Deployment guide with Kubernetes manifests - -### For AI/IDE Tools -- Auto-generated JSON schema (config-schema.json) -- Code patterns for AI generation (patterns.json) -- Machine-readable configuration structure -- IDE autocomplete support via JSON Schema - -### Quality Assurance -- Documentation validation tests -- YAML example parsing tests -- Schema validation tests -- Automated build via build.rs -- Link checking (mdbook-linkcheck) - -## Artifacts Delivered - -1. **Documentation Website**: 34 HTML pages in `docs/` -2. **JSON Schemas**: Auto-generated config-schema.json + patterns.json -3. **Validation Tests**: tests/documentation_validation.rs -4. **Build Automation**: build.rs (generates schemas on build) -5. **Enhanced Configuration**: doc/book.toml with mermaid + linkcheck - -## Technical Implementation - -### Build Script (build.rs) -- Auto-generates JSON Schema from Rust Config struct -- Runs on every `cargo build` -- Ensures schema stays in sync with code - -### Validation Tests (documentation_validation.rs) -- Validates YAML examples parse correctly -- Validates JSON schemas are well-formed -- Validates documentation structure -- Runs in CI/CD pipeline - -### Documentation Tooling -- **mdbook**: Static site generator -- **mdbook-mermaid**: Diagram rendering -- **mdbook-linkcheck**: Link validation -- **schemars**: JSON Schema generation - -## Validation Results - -✅ **mdbook build**: SUCCESS (34 HTML pages) -✅ **Schema generation**: SUCCESS (config-schema.json) -✅ **Documentation structure**: COMPLETE -✅ **Cross-references**: COMPLETE -✅ **Navigation**: COMPLETE -✅ **Examples**: VALID - -## Files Modified - -``` -Modified: -- Cargo.toml (added schemars) -- doc/SUMMARY.md (updated structure) -- doc/book.toml (added preprocessors) -- doc/convertor.md (enhanced) -- doc/reporter.md (enhanced) -- doc/index.md (enhanced) - -Created: -- build.rs -- tests/documentation_validation.rs -- doc/getting-started/ (3 files) -- doc/architecture/ (3 files) -- doc/api/ (3 files) -- doc/configuration/ (8 files) -- doc/integration/ (3 files) -- doc/modules/ (6 files) -- doc/guides/ (2 files) -- doc/schemas/ (3 files) -- docs/ (34 HTML files) -``` - -## Next Steps - -1. **Review**: Open docs/index.html in browser for full review -2. **Test**: Run `cargo test --test documentation_validation` -3. **Deploy**: Publish to GitHub Pages or internal docs site -4. **CI/CD**: Update pipeline to include `mdbook build doc/` -5. **Feedback**: Share with team and gather feedback - -## Access Points - -- **Local Documentation**: file:///Users/A107229207/dev/otc/stackmon/metrics-processor/docs/index.html -- **Source Files**: /Users/A107229207/dev/otc/stackmon/metrics-processor/doc/ -- **Schemas**: /Users/A107229207/dev/otc/stackmon/metrics-processor/doc/schemas/ - -## Constitution Compliance - -✅ **Principle I (Code Quality)**: Documentation does not introduce code quality issues -✅ **Principle II (Testing)**: Validation tests ensure documentation quality -✅ **Principle III (UX)**: Documentation enhances developer experience -✅ **Principle IV (Performance)**: Documentation generation <30s - -## Conclusion - -All 71 tasks completed successfully. The project documentation feature is fully implemented and ready for deployment. The documentation serves both human developers (onboarding, reference, troubleshooting) and AI-powered tools (schemas, patterns, structured data). - -**Status**: ✅ READY FOR DEPLOYMENT - ---- - -**Prepared by**: GitHub Copilot CLI -**Date**: January 20, 2025 -**Total Execution Time**: ~2 hours diff --git a/specs/001-project-documentation/checklists/requirements.md b/specs/001-project-documentation/checklists/requirements.md deleted file mode 100644 index 90a0c9f..0000000 --- a/specs/001-project-documentation/checklists/requirements.md +++ /dev/null @@ -1,65 +0,0 @@ -# Specification Quality Checklist: Comprehensive Project Documentation - -**Purpose**: Validate specification completeness and quality before proceeding to planning -**Created**: 2025-01-23 -**Feature**: [spec.md](../spec.md) - -## Content Quality - -- [x] No implementation details (languages, frameworks, APIs) -- [x] Focused on user value and business needs -- [x] Written for non-technical stakeholders -- [x] All mandatory sections completed - -## Requirement Completeness - -- [x] No [NEEDS CLARIFICATION] markers remain -- [x] Requirements are testable and unambiguous -- [x] Success criteria are measurable -- [x] Success criteria are technology-agnostic (no implementation details) -- [x] All acceptance scenarios are defined -- [x] Edge cases are identified -- [x] Scope is clearly bounded -- [x] Dependencies and assumptions identified - -## Feature Readiness - -- [x] All functional requirements have clear acceptance criteria -- [x] User scenarios cover primary flows -- [x] Feature meets measurable outcomes defined in Success Criteria -- [x] No implementation details leak into specification - -## Validation Results - -**Status**: ✅ PASSED - All validation items passed - -### Detailed Review: - -#### Content Quality -- ✅ Documentation is described in terms of what it should contain, not how to implement it -- ✅ Focus is on enabling users (developers, AI agents, operations teams) to achieve their goals -- ✅ Language is accessible to non-technical stakeholders - describes documentation needs without technical jargon -- ✅ All mandatory sections (User Scenarios, Requirements, Success Criteria) are complete - -#### Requirement Completeness -- ✅ No [NEEDS CLARIFICATION] markers present - all requirements are explicit and clear -- ✅ Each requirement is testable (e.g., FR-003: "provide complete API reference" is verifiable) -- ✅ All success criteria include measurable metrics (time, percentage, count) -- ✅ Success criteria focus on outcomes (e.g., "developers complete setup in 30 minutes") not implementation -- ✅ Each user story has detailed acceptance scenarios with Given-When-Then format -- ✅ Edge cases cover documentation maintenance, synchronization, and AI tool interaction -- ✅ Scope is bounded to documentation creation (doesn't include implementing the features being documented) -- ✅ Assumptions are implicit but reasonable (existing OpenAPI schema, current mdbook structure) - -#### Feature Readiness -- ✅ Each functional requirement maps to user stories and success criteria -- ✅ Six prioritized user stories cover the full spectrum from P1 (onboarding, AI) to P3 (extensibility) -- ✅ Measurable outcomes clearly define what "done" looks like -- ✅ Specification remains technology-agnostic (discusses "diagrams" not "Mermaid", "documentation" not "mdbook") - -## Notes - -- The specification successfully balances human and AI audience needs by including both traditional documentation and machine-readable structure requirements -- The prioritization of user stories is well-justified with P1 focusing on immediate team needs (onboarding, AI assistance) -- Edge cases appropriately address documentation maintenance concerns, which is often overlooked -- No implementation guidance is needed at this stage - the spec is ready for `/speckit.plan` diff --git a/specs/001-project-documentation/contracts/README.md b/specs/001-project-documentation/contracts/README.md deleted file mode 100644 index 99e7230..0000000 --- a/specs/001-project-documentation/contracts/README.md +++ /dev/null @@ -1,157 +0,0 @@ -# Contracts: Machine-Readable Documentation Schemas - -**Feature**: 001-project-documentation -**Date**: 2025-01-23 -**Purpose**: Define machine-readable schemas for AI tools and IDEs - -## Overview - -This directory contains JSON Schema definitions and pattern documentation designed for consumption by AI-powered development tools, IDEs, and code generators. - ---- - -## Schema Files - -### config-schema.json - -**Purpose**: Complete JSON Schema (Draft 7) for metrics-processor configuration files - -**Usage**: -- **IDE Integration**: Provides autocomplete and validation in VSCode/IntelliJ when editing YAML config files -- **Runtime Validation**: Can be used with `jsonschema` crate to validate configuration at startup -- **AI Code Generation**: Enables LLMs to generate valid configuration examples - -**Integration Example** (VSCode): -```json -// .vscode/settings.json -{ - "yaml.schemas": { - "specs/001-project-documentation/contracts/config-schema.json": "config*.yaml" - } -} -``` - -**Generated From**: Rust `Config` struct in `src/config.rs` (will be auto-generated via `schemars` crate in implementation) - -**Validation Rules**: -- All required fields must be present (datasource, server, flag_metrics, health_metrics) -- Service names must use lowercase_with_underscores pattern -- Environment names must use lowercase-with-dashes pattern -- Flag metric templates must reference existing metric_template entries -- Health metric expressions must reference existing flag metrics - ---- - -### patterns.json - -**Purpose**: Document code patterns, conventions, and domain-specific logic not captured in type schemas - -**Usage**: -- **AI Code Generation**: LLMs read this to understand project conventions and generate conformant code -- **Onboarding**: New developers reference to understand naming, error handling, and architectural patterns -- **IDE Plugins**: Can be parsed to provide context-aware suggestions - -**Sections**: - -1. **patterns[]**: Array of documented patterns with examples - - Variable substitution in metric templates - - Boolean expression evaluation in health metrics - - Error handling conventions - - Async/await usage patterns - - Configuration validation approach - - Logging conventions - -2. **conventions**: Naming and structural conventions - - File naming (Rust modules, docs, configs) - - Data structure patterns - - Testing organization - -3. **architectural_principles**: High-level design decisions - - Module separation of concerns - - Binary responsibilities (convertor vs reporter) - -**Example Usage** (AI prompt): -``` -Given the patterns.json file, generate a new flag metric configuration -that monitors database query latency for the storage service. -``` - ---- - -## Implementation Notes - -### Generation Strategy - -These schemas should be **auto-generated** during the build process to ensure they stay synchronized with code: - -```rust -// build.rs -use schemars::schema_for; -use cloudmon_metrics::config::Config; - -fn main() { - // Generate config-schema.json from Config struct - let schema = schema_for!(Config); - std::fs::write( - "doc/schemas/config-schema.json", - serde_json::to_string_pretty(&schema).unwrap() - ).expect("Failed to write schema"); - - println!("cargo:rerun-if-changed=src/config.rs"); -} -``` - -### Validation Testing - -Add tests to ensure schemas remain valid and examples conform: - -```rust -// tests/schema_validation.rs -#[test] -fn config_schema_validates_examples() { - let schema = include_str!("../doc/schemas/config-schema.json"); - let example = include_str!("../doc/configuration/examples.md"); - - // Extract YAML blocks and validate against schema - assert!(validate_yaml_against_schema(example, schema).is_ok()); -} -``` - ---- - -## Maintenance - -- **config-schema.json**: Auto-generated from `src/config.rs` - DO NOT EDIT MANUALLY -- **patterns.json**: Manually maintained - Update when conventions change -- **Version**: Both files should be versioned with project (currently 0.2.0) - ---- - -## Target Consumers - -| Consumer | File Used | Purpose | -|----------|-----------|---------| -| VSCode | config-schema.json | YAML autocomplete and validation | -| IntelliJ | config-schema.json | Configuration editing assistance | -| Claude/GPT | patterns.json | Understand code conventions for generation | -| GitHub Copilot | Both | Suggest conformant code and configuration | -| Custom tooling | config-schema.json | Runtime configuration validation | -| Documentation | Both | Generate reference documentation | - ---- - -## Future Extensions - -Potential additional schema files as project grows: - -- **api-types-schema.json**: Request/response types for `/v1/health` endpoint -- **tsdb-interface-schema.json**: Contract for implementing new TSDB backends -- **plugin-schema.json**: If plugin system added for custom metrics - ---- - -## References - -- JSON Schema specification: https://json-schema.org/draft-07/schema -- `schemars` crate: https://docs.rs/schemars/latest/schemars/ -- VSCode YAML extension: https://marketplace.visualstudio.com/items?itemName=redhat.vscode-yaml diff --git a/specs/001-project-documentation/contracts/config-schema.json b/specs/001-project-documentation/contracts/config-schema.json deleted file mode 100644 index 5aa55bf..0000000 --- a/specs/001-project-documentation/contracts/config-schema.json +++ /dev/null @@ -1,217 +0,0 @@ -{ - "$schema": "http://json-schema.org/draft-07/schema#", - "$id": "https://cloudmon.eco.tsi-dev.otc-service.com/schemas/config-schema.json", - "title": "CloudMon Metrics Processor Configuration", - "description": "Complete configuration schema for metrics-processor including datasource, server, metric templates, and health metrics", - "type": "object", - "required": ["datasource", "server", "flag_metrics", "health_metrics"], - "properties": { - "datasource": { - "type": "object", - "description": "Time-series database (TSDB) connection configuration", - "required": ["url", "type"], - "properties": { - "url": { - "type": "string", - "format": "uri", - "description": "TSDB base URL (e.g., https://graphite.example.com)", - "examples": ["https://graphite.example.com"] - }, - "type": { - "type": "string", - "enum": ["graphite"], - "description": "TSDB backend type (currently only Graphite supported)" - }, - "timeout": { - "type": "integer", - "minimum": 1, - "maximum": 300, - "default": 30, - "description": "Request timeout in seconds" - } - } - }, - "server": { - "type": "object", - "description": "HTTP API server configuration", - "required": ["address", "port"], - "properties": { - "address": { - "type": "string", - "format": "ipv4", - "description": "IP address to bind to", - "examples": ["0.0.0.0", "127.0.0.1"] - }, - "port": { - "type": "integer", - "minimum": 1024, - "maximum": 65535, - "description": "Port number for HTTP server", - "examples": [3005, 8080] - } - } - }, - "metric_templates": { - "type": "object", - "description": "Reusable query templates for TSDB metrics", - "additionalProperties": { - "type": "object", - "required": ["query", "op", "threshold"], - "properties": { - "query": { - "type": "string", - "description": "TSDB query with variable substitution ($service, $environment)", - "examples": ["stats.counters.api.$environment.$service.*.*.count"] - }, - "op": { - "type": "string", - "enum": ["lt", "gt", "eq"], - "description": "Comparison operator: lt (less than), gt (greater than), eq (equals)" - }, - "threshold": { - "type": "number", - "description": "Threshold value for comparison" - } - } - } - }, - "environments": { - "type": "array", - "description": "Monitoring environments (e.g., production, staging)", - "items": { - "type": "object", - "required": ["name"], - "properties": { - "name": { - "type": "string", - "pattern": "^[a-z0-9-]+$", - "description": "Environment name (lowercase with dashes)" - }, - "attributes": { - "type": "object", - "description": "Optional attributes for status dashboard integration", - "additionalProperties": { - "type": "string" - } - } - } - } - }, - "status_dashboard": { - "type": "object", - "description": "Status dashboard integration configuration", - "required": ["url", "secret"], - "properties": { - "url": { - "type": "string", - "format": "uri", - "description": "Status dashboard base URL" - }, - "secret": { - "type": "string", - "description": "JWT signing secret for authentication" - } - } - }, - "flag_metrics": { - "type": "array", - "description": "Individual flag metrics that evaluate to raised (true) or lowered (false)", - "minItems": 1, - "items": { - "type": "object", - "required": ["name", "service", "template", "environments"], - "properties": { - "name": { - "type": "string", - "pattern": "^[a-z_]+$", - "description": "Metric name (lowercase with underscores)" - }, - "service": { - "type": "string", - "pattern": "^[a-z_]+$", - "description": "Service identifier this metric belongs to" - }, - "template": { - "type": "object", - "required": ["name"], - "properties": { - "name": { - "type": "string", - "description": "Reference to metric_template name" - } - } - }, - "environments": { - "type": "array", - "description": "Environments this metric applies to", - "minItems": 1, - "items": { - "type": "object", - "required": ["name"], - "properties": { - "name": { - "type": "string", - "description": "Environment name matching environments array" - } - } - } - } - } - } - }, - "health_metrics": { - "type": "object", - "description": "Health metrics combining flag metrics into service health status", - "additionalProperties": { - "type": "object", - "required": ["service", "component_name", "category", "metrics", "expressions"], - "properties": { - "service": { - "type": "string", - "description": "Service identifier (must match flag_metrics service values)" - }, - "component_name": { - "type": "string", - "description": "Human-readable component name" - }, - "category": { - "type": "string", - "description": "Service category (e.g., compute, storage, network)" - }, - "metrics": { - "type": "array", - "description": "Flag metrics referenced in expressions (service.metric_name format)", - "minItems": 1, - "items": { - "type": "string", - "pattern": "^[a-z_]+\\.[a-z_]+$", - "examples": ["api.slow", "api.down"] - } - }, - "expressions": { - "type": "array", - "description": "Boolean expressions with weights determining health status", - "minItems": 1, - "items": { - "type": "object", - "required": ["expression", "weight"], - "properties": { - "expression": { - "type": "string", - "description": "Boolean expression using flag metric references (supports &&, ||, !)", - "examples": ["api.slow || api.success_rate_low", "api.down"] - }, - "weight": { - "type": "integer", - "minimum": 0, - "maximum": 2, - "description": "Health weight: 0=healthy, 1=degraded, 2=outage" - } - } - } - } - } - } - } - } -} diff --git a/specs/001-project-documentation/contracts/patterns.json b/specs/001-project-documentation/contracts/patterns.json deleted file mode 100644 index a72ab49..0000000 --- a/specs/001-project-documentation/contracts/patterns.json +++ /dev/null @@ -1,190 +0,0 @@ -{ - "$schema": "http://json-schema.org/draft-07/schema#", - "$id": "https://cloudmon.eco.tsi-dev.otc-service.com/schemas/patterns.json", - "title": "CloudMon Metrics Processor Code Patterns and Conventions", - "description": "Patterns and conventions for AI-powered code generation and analysis", - "version": "0.2.0", - "patterns": [ - { - "id": "metric_template_variable_substitution", - "name": "Metric Template Variable Substitution", - "context": "Templates use $variable syntax for dynamic query generation", - "syntax": "$service, $environment (case-sensitive)", - "available_variables": ["service", "environment"], - "example": "stats.counters.api.$environment.$service.*.*.count", - "implementation": "src/types.rs:172", - "notes": "Variable names must match exactly; no other variables supported" - }, - { - "id": "binary_metric_configuration", - "name": "Binary Metric Configuration Pattern", - "description": "Flag and health metrics must specify comparison operator and threshold", - "operators": [ - { - "name": "lt", - "description": "Less than - metric value < threshold triggers flag" - }, - { - "name": "gt", - "description": "Greater than - metric value > threshold triggers flag" - }, - { - "name": "eq", - "description": "Equals - metric value == threshold triggers flag" - } - ], - "reference": "src/types.rs:BinaryMetricRawDef", - "example": { - "query": "aggregate(stats.counters.api.*.mean)", - "op": "gt", - "threshold": 1000 - } - }, - { - "id": "health_expression_evaluation", - "name": "Health Expression Boolean Logic", - "description": "Health metrics use boolean expressions combining flag metrics", - "syntax": "service.metric_name with operators: &&, ||, !", - "expression_format": "flag_reference [operator flag_reference ...]", - "available_operators": { - "||": "Logical OR - any flag raised triggers expression", - "&&": "Logical AND - all flags must be raised", - "!": "Logical NOT - inverts flag state" - }, - "examples": [ - { - "expression": "api.slow || api.success_rate_low", - "description": "Triggers if either API is slow OR success rate is low" - }, - { - "expression": "api.down", - "description": "Triggers only if API is completely down" - }, - { - "expression": "!api.healthy && api.degraded", - "description": "Triggers if API is not healthy AND is degraded" - } - ], - "implementation": "Uses evalexpr crate for expression evaluation", - "reference": "src/types.rs:HealthMetric" - }, - { - "id": "service_environment_naming", - "name": "Service and Environment Naming Convention", - "services": { - "format": "lowercase_with_underscores", - "pattern": "^[a-z_]+$", - "examples": ["api", "compute", "block_storage"] - }, - "environments": { - "format": "lowercase-with-dashes", - "pattern": "^[a-z0-9-]+$", - "examples": ["production", "eu-de", "us-west-2"] - }, - "rationale": "Consistent naming enables predictable query construction" - }, - { - "id": "error_handling_pattern", - "name": "Error Handling Convention", - "approach": "Custom error types with context", - "common_errors": [ - { - "error": "ServiceNotSupported", - "cause": "Service requested not defined in health_metrics configuration", - "resolution": "Add service to health_metrics section or check spelling" - }, - { - "error": "EnvNotSupported", - "cause": "Environment requested not defined in environments configuration", - "resolution": "Add environment to environments array or check spelling" - }, - { - "error": "InvalidExpression", - "cause": "Health metric expression contains syntax error", - "resolution": "Check expression syntax: use &&, ||, ! operators and service.metric format" - }, - { - "error": "MetricNotFound", - "cause": "Flag metric referenced in health expression not defined", - "resolution": "Ensure flag metric exists in flag_metrics array" - } - ], - "implementation": "Custom error enums, avoid unwrap() in production code", - "reference": ".specify/memory/constitution.md:Principle I.4" - }, - { - "id": "async_patterns", - "name": "Async/Await Usage Pattern", - "description": "All I/O operations use async with tokio runtime", - "rules": [ - "HTTP requests to TSDB: use reqwest with async", - "API handlers: use axum async handlers", - "Never block tokio runtime with sync I/O", - "Use tokio::spawn for concurrent operations" - ], - "example_handler": "async fn get_health(Query(params): Query) -> Result, StatusCode>", - "reference": "src/api/v1.rs" - }, - { - "id": "configuration_validation", - "name": "Configuration Validation Pattern", - "approach": "Validate configuration at startup before server starts", - "validation_rules": [ - "All flag_metric templates must exist in metric_templates", - "All flag_metric environments must exist in environments array", - "All health_metric expressions must reference existing flag metrics", - "Service names in flag_metrics must match health_metrics keys" - ], - "implementation": "config::validate() function called during initialization", - "reference": "src/config.rs" - }, - { - "id": "logging_conventions", - "name": "Structured Logging Pattern", - "framework": "tracing crate with tower-http middleware", - "levels": { - "ERROR": "Actionable issues requiring operator intervention", - "WARN": "Degraded operation or unexpected but handled conditions", - "INFO": "Significant state changes (server startup, config reload)", - "DEBUG": "Detailed diagnostics for troubleshooting" - }, - "required_context": [ - "request_id: included via tower-http middleware", - "service: for metric-specific logs", - "environment: for environment-specific logs" - ], - "example": "tracing::info!(request_id = %request_id, service = %service, \"Processing health query\");", - "reference": ".specify/memory/constitution.md:Principle III.3" - } - ], - "conventions": { - "file_naming": { - "rust_modules": "lowercase_snake_case.rs", - "documentation": "kebab-case.md", - "configuration": "kebab-case.yaml" - }, - "data_structures": { - "config_structs": "Use serde derive for serialization/deserialization", - "api_types": "Derive JsonSchema for OpenAPI generation", - "internal_types": "Prefer owned types over lifetimes for simplicity" - }, - "testing": { - "unit_tests": "#[cfg(test)] modules within implementation files", - "integration_tests": "tests/ directory for HTTP endpoint testing", - "mocking": "Use mockito for TSDB backend mocking" - } - }, - "architectural_principles": { - "separation_of_concerns": [ - "lib.rs: Common types and utilities", - "api.rs: HTTP API handlers", - "config.rs: Configuration parsing and validation", - "graphite.rs: TSDB backend implementation", - "types.rs: Domain data structures" - ], - "binary_responsibilities": { - "convertor": "Evaluates metrics and provides HTTP API for queries", - "reporter": "Polls convertor and sends updates to status dashboard" - } - } -} diff --git a/specs/001-project-documentation/data-model.md b/specs/001-project-documentation/data-model.md deleted file mode 100644 index b829e96..0000000 --- a/specs/001-project-documentation/data-model.md +++ /dev/null @@ -1,396 +0,0 @@ -# Data Model: Project Documentation System - -**Feature**: 001-project-documentation -**Date**: 2025-01-23 -**Purpose**: Define the structure and relationships of documentation entities - -## Overview - -This document defines the logical data model for the comprehensive project documentation system. While documentation is not stored in a traditional database, understanding the entities, their attributes, and relationships helps structure the documentation consistently and enables tooling to process it effectively. - ---- - -## Core Entities - -### 1. Documentation Section - -**Description**: A top-level category of documentation content (e.g., Architecture, API, Configuration) - -**Attributes**: -- `id`: String - Unique identifier (e.g., "architecture", "api-reference") -- `title`: String - Human-readable title (e.g., "Architecture Overview") -- `order`: Integer - Display order in navigation -- `path`: Path - Filesystem location (e.g., "doc/architecture/") -- `summary_entry`: String - Line in SUMMARY.md linking to this section -- `audience`: Enum - Primary audience (Developer, Operator, AI-Tool) -- `status`: Enum - Completeness (Draft, Complete, Needs-Update) - -**Relationships**: -- Contains many **Documentation Pages** -- May contain **Subsections** (recursive) - -**Validation Rules**: -- `id` must be unique across all sections -- `path` must exist in filesystem -- `order` must be positive integer -- Each section must have at least one page - ---- - -### 2. Documentation Page - -**Description**: A single markdown file containing specific documentation content - -**Attributes**: -- `id`: String - Unique identifier (e.g., "architecture-overview") -- `title`: String - Page title (extracted from first # heading) -- `file_path`: Path - Filesystem location (e.g., "doc/architecture/overview.md") -- `parent_section`: Reference - Parent Documentation Section -- `frontmatter`: Object - Optional YAML frontmatter metadata - - `date`: DateTime - Last updated - - `author`: String - - `tags`: Array -- `content_blocks`: Array - Parsed markdown content -- `references`: Array - Links to other pages or external resources -- `code_examples`: Array - Embedded code snippets - -**Relationships**: -- Belongs to one **Documentation Section** -- Contains many **Content Blocks** -- May reference many **Diagrams** -- May reference many **Schema Definitions** - -**Validation Rules**: -- `file_path` must exist and be valid markdown -- First line must be level-1 heading (# Title) -- All internal links must resolve to existing pages -- All code examples must specify language - ---- - -### 3. Content Block - -**Description**: A logical unit of content within a page (paragraph, code, diagram, etc.) - -**Attributes**: -- `type`: Enum - BlockType (Paragraph, CodeBlock, Diagram, Table, List, Heading) -- `content`: String - Raw content -- `language`: String (optional) - For code blocks (rust, yaml, json, bash) -- `metadata`: Object - Type-specific metadata - - For code blocks: `testable: boolean`, `example_name: string` - - For diagrams: `diagram_type: string` (mermaid, graphviz) -- `order`: Integer - Position within page - -**Relationships**: -- Belongs to one **Documentation Page** - -**Validation Rules**: -- Code blocks must have valid syntax for specified language -- Mermaid diagrams must render without errors -- Tables must have consistent column counts - ---- - -### 4. Diagram - -**Description**: Visual representation of system architecture, data flow, or relationships - -**Attributes**: -- `id`: String - Unique identifier (e.g., "arch-system-overview") -- `title`: String - Diagram title -- `type`: Enum - DiagramType (Architecture, DataFlow, Sequence, ERD, ComponentDependency) -- `format`: Enum - Format (Mermaid, SVG, PNG) -- `source`: String - Diagram definition (for text-based formats) -- `rendered_path`: Path (optional) - For pre-rendered formats -- `entities`: Array - Key entities/components shown -- `description`: String - Text description for accessibility and AI parsing - -**Relationships**: -- Referenced by many **Documentation Pages** -- May be associated with **Architecture Components** - -**Validation Rules**: -- Mermaid diagrams must compile successfully -- All entities referenced must exist in codebase or configuration -- Must include text description for accessibility (FR-013) - ---- - -### 5. Schema Definition - -**Description**: Machine-readable schema for configuration, API types, or data structures - -**Attributes**: -- `id`: String - Unique identifier (e.g., "config-schema") -- `name`: String - Schema name (e.g., "Configuration") -- `format`: Enum - SchemaFormat (JSONSchema, TypeScript, OpenAPI) -- `file_path`: Path - Location of schema file (e.g., "doc/schemas/config-schema.json") -- `source_code_ref`: Path - Rust struct this schema derives from (e.g., "src/config.rs:Config") -- `version`: String - Schema version (e.g., "0.2.0") -- `generated`: Boolean - Whether auto-generated from code -- `validation_rules`: Array - Validation constraints - -**Relationships**: -- Derived from **Source Code Types** -- Referenced by **Configuration Pages** -- Used by **IDE Tools** (external) - -**Validation Rules**: -- Must be valid according to format specification -- If `generated: true`, must match source code structure -- Version must match project version - ---- - -### 6. Configuration Example - -**Description**: Working configuration sample demonstrating specific features - -**Attributes**: -- `id`: String - Unique identifier (e.g., "basic-setup") -- `title`: String - Example title (e.g., "Basic Multi-Service Setup") -- `description`: String - What this example demonstrates -- `content`: String - Full YAML configuration -- `use_cases`: Array - Scenarios this applies to -- `referenced_sections`: Array - Config sections demonstrated -- `validated`: Boolean - Whether example has passed validation tests - -**Relationships**: -- References **Schema Definition** for validation -- Documented in **Configuration Pages** - -**Validation Rules**: -- Must parse as valid YAML -- Must conform to configuration schema -- Must include all required fields from schema -- Should demonstrate at least one unique feature - ---- - -### 7. API Endpoint Documentation - -**Description**: Documentation for a specific HTTP API endpoint - -**Attributes**: -- `id`: String - Unique identifier (e.g., "get-v1-health") -- `path`: String - URL path (e.g., "/v1/health") -- `method`: Enum - HTTPMethod (GET, POST, PUT, DELETE) -- `summary`: String - Brief description -- `description`: String - Detailed explanation -- `parameters`: Array - Query/path/header parameters -- `request_body`: Schema Reference (optional) -- `responses`: Object - Possible responses -- `authentication`: String - Auth requirements (e.g., "JWT token") -- `examples`: Array - Request/response examples -- `openapi_ref`: String - Reference to OpenAPI schema definition - -**Relationships**: -- Defined in **OpenAPI Schema** (openapi-schema.yaml) -- Documented in **API Pages** -- May reference **Schema Definitions** for request/response types - -**Validation Rules**: -- Must exist in openapi-schema.yaml (FR-012) -- All parameters must have type and description -- Examples must match schema definitions -- Response schemas must match Rust types - ---- - -### 8. Module Documentation - -**Description**: Documentation for a Rust module/crate component - -**Attributes**: -- `id`: String - Unique identifier (e.g., "module-api") -- `module_name`: String - Rust module name (e.g., "api", "config") -- `path`: Path - Filesystem location (e.g., "src/api.rs" or "src/api/") -- `purpose`: String - Primary responsibility -- `public_items`: Array - Exported functions, structs, traits -- `dependencies`: Array - Other modules this depends on -- `used_by`: Array - Modules that depend on this -- `key_types`: Array - Important data structures -- `rustdoc_coverage`: Integer - Percentage of public items documented - -**Relationships**: -- Contains **Public API Items** -- Depends on other **Modules** -- Documented in **Module Pages** - -**Validation Rules**: -- Module path must exist in src/ -- All public items should have rustdoc comments (warn if <90%) -- Dependencies must form acyclic graph (except for lib.rs) - ---- - -### 9. Integration Guide - -**Description**: Documentation for integrating with external systems or extending the project - -**Attributes**: -- `id`: String - Unique identifier (e.g., "tsdb-integration") -- `title`: String - Guide title (e.g., "Adding a New TSDB Backend") -- `type`: Enum - GuideType (Integration, Extension, Migration) -- `target_audience`: String - Who this is for (e.g., "Backend developers") -- `prerequisites`: Array - Required knowledge or setup -- `steps`: Array - Sequential instructions -- `code_templates`: Array - Boilerplate code -- `testing_guidance`: String - How to test the integration -- `estimated_time`: Duration - Expected completion time - -**Relationships**: -- References **Module Documentation** -- May reference **Schema Definitions** -- Contains **Code Examples** - -**Validation Rules**: -- Steps must be numbered sequentially -- All code templates must be syntactically valid -- Prerequisites must reference existing documentation - ---- - -### 10. Troubleshooting Entry - -**Description**: A known issue, error, or problem with resolution steps - -**Attributes**: -- `id`: String - Unique identifier (e.g., "error-service-not-found") -- `symptom`: String - What the user observes -- `error_message`: String (optional) - Exact error text -- `cause`: String - Root cause explanation -- `solution`: String - Step-by-step resolution -- `related_config`: Array - Config sections involved -- `related_logs`: Array - Log patterns to look for -- `severity`: Enum - Severity (Critical, High, Medium, Low) - -**Relationships**: -- References **Configuration Pages** -- May reference **API Endpoint Documentation** -- May reference **Module Documentation** - -**Validation Rules**: -- Must include both cause and solution -- Related config/logs must reference real configuration fields -- Error messages should be searchable - ---- - -## Relationships Diagram - -```mermaid -erDiagram - DOCUMENTATION-SECTION ||--o{ DOCUMENTATION-PAGE : contains - DOCUMENTATION-SECTION ||--o{ DOCUMENTATION-SECTION : has-subsections - DOCUMENTATION-PAGE ||--o{ CONTENT-BLOCK : contains - DOCUMENTATION-PAGE }o--o{ DIAGRAM : references - DOCUMENTATION-PAGE }o--o{ SCHEMA-DEFINITION : references - - CONTENT-BLOCK }o--|| DOCUMENTATION-PAGE : belongs-to - - DIAGRAM }o--o{ DOCUMENTATION-PAGE : referenced-by - - SCHEMA-DEFINITION }o--|| SOURCE-CODE : generated-from - SCHEMA-DEFINITION }o--o{ CONFIGURATION-EXAMPLE : validates - SCHEMA-DEFINITION }o--o{ API-ENDPOINT : defines-types - - CONFIGURATION-EXAMPLE }o--o{ DOCUMENTATION-PAGE : documented-in - - API-ENDPOINT ||--|| OPENAPI-SCHEMA : defined-in - API-ENDPOINT }o--o{ DOCUMENTATION-PAGE : documented-in - API-ENDPOINT }o--o{ SCHEMA-DEFINITION : uses - - MODULE }o--o{ MODULE : depends-on - MODULE ||--o{ PUBLIC-ITEM : contains - MODULE }o--o{ DOCUMENTATION-PAGE : documented-in - - INTEGRATION-GUIDE }o--o{ MODULE : references - INTEGRATION-GUIDE }o--o{ SCHEMA-DEFINITION : references - - TROUBLESHOOTING }o--o{ CONFIGURATION-PAGE : references - TROUBLESHOOTING }o--o{ API-ENDPOINT : references -``` - ---- - -## State Transitions - -### Documentation Page States - -```mermaid -stateDiagram-v2 - [*] --> Planned - Planned --> Draft: Author starts writing - Draft --> Review: Author requests review - Review --> Draft: Reviewer requests changes - Review --> Complete: Reviewer approves - Complete --> NeedsUpdate: Code changes detected - NeedsUpdate --> Review: Author updates - Complete --> [*] -``` - ---- - -## Validation & Integrity Rules - -### Cross-Entity Validation - -1. **Link Integrity**: All internal references must resolve - - Documentation page links to other pages - - Schema references to source code - - Configuration examples to schema definitions - -2. **Schema Alignment**: Generated artifacts must match source - - JSON schemas match Rust struct definitions - - OpenAPI docs match API endpoint implementations - - Configuration examples conform to schema - -3. **Example Validation**: All code examples must be executable - - Rust code blocks compile successfully - - YAML configuration examples parse correctly - - API request examples match OpenAPI spec - -4. **Completeness**: Required documentation exists - - Every public module has documentation page - - Every API endpoint documented in API section - - Every configuration field has description - ---- - -## Implementation Notes - -### Tooling Support - -- **mdbook**: Renders Documentation Sections, Pages, and Content Blocks into web documentation -- **schemars**: Generates Schema Definitions from Rust types -- **cargo test**: Validates Code Examples and Configuration Examples -- **mdbook-linkcheck**: Validates link integrity across Documentation Pages -- **build.rs**: Automates Schema Definition generation from source code - -### File System Mapping - -``` -doc/ → Documentation Section (root) - architecture/ → Documentation Section - overview.md → Documentation Page - # Heading → Content Block (Heading) - Paragraph → Content Block (Paragraph) - ```mermaid ... ``` → Content Block (Diagram) - schemas/ → Schema Definitions storage - config-schema.json → Schema Definition - patterns.json → Schema Definition (conventions) -``` - ---- - -## Summary - -This data model defines 10 core entities organized into 4 logical layers: - -1. **Structure Layer**: Documentation Sections and Pages -2. **Content Layer**: Content Blocks, Diagrams, Examples -3. **Schema Layer**: Schema Definitions, API Endpoints -4. **Guidance Layer**: Module Docs, Integration Guides, Troubleshooting - -The model supports both human-readable documentation (rendered by mdbook) and machine-readable artifacts (JSON schemas, OpenAPI specs) while maintaining referential integrity through validation rules. diff --git a/specs/001-project-documentation/plan.md b/specs/001-project-documentation/plan.md deleted file mode 100644 index f4f9511..0000000 --- a/specs/001-project-documentation/plan.md +++ /dev/null @@ -1,310 +0,0 @@ -# Implementation Plan: Comprehensive Project Documentation - -**Branch**: `001-project-documentation` | **Date**: 2025-01-23 | **Spec**: [specs/001-project-documentation/spec.md](./spec.md) -**Input**: Feature specification from `/specs/001-project-documentation/spec.md` - -**Note**: This template is filled in by the `/speckit.plan` command. See `.specify/templates/commands/plan.md` for the execution workflow. - -## Summary - -Create comprehensive project documentation that serves both human developers and AI-powered tools (agents, LLMs, IDEs). Documentation will cover project architecture, API references, configuration schemas, developer onboarding, module organization, and TSDB integration patterns. The implementation leverages existing mdbook infrastructure (doc/) and OpenAPI schema (openapi-schema.yaml), extending them with architecture diagrams, data models, developer guides, and machine-readable schemas that enable AI assistants to provide accurate code suggestions and maintain project conventions. - -## Technical Context - -**Language/Version**: Rust 1.75+ (edition 2021, per Cargo.toml) -**Primary Dependencies**: axum (0.6), tokio (1.28), serde (1.0), tracing (0.1), reqwest (0.11) -**Storage**: N/A (documentation feature - targets file system) -**Testing**: cargo test with mockito (1.0) for HTTP mocking, tempfile (3.5) for test fixtures -**Target Platform**: Multi-platform documentation (mdbook for web, markdown for AI agents/IDEs) -**Project Type**: Single project (library + 2 binaries: convertor, reporter) -**Performance Goals**: Documentation generation <30s, search response <3s (per SC-008) -**Constraints**: Must align with existing OpenAPI schema, parseable by AI tools, maintain <2hr/month maintenance burden (per SC-012) -**Scale/Scope**: 9 Rust modules, 2 binaries, 2 HTTP endpoints, ~20 configuration fields, 5 user stories - -### Existing Documentation Infrastructure - -- **mdbook**: Already configured at `doc/` with book.toml, currently has basic component descriptions -- **OpenAPI Schema**: `openapi-schema.yaml` defines `/v1/health` and `/v1/maintenances` endpoints with full request/response schemas -- **Existing Docs**: `doc/convertor.md`, `doc/reporter.md`, `doc/config.md` - minimal coverage, needs expansion - -### Documentation Requirements Context - -- **Human Consumption**: Onboarding guides, architecture overviews, troubleshooting (User Stories 1, 3, 4, 6) -- **AI/Machine Consumption**: Structured schemas for IDE autocomplete, code generation patterns, type information (User Story 2) -- **Dual-purpose**: Examples must be executable/validated to serve both audiences (Edge Case: validate examples work) - -## Constitution Check - -*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.* - -### Principle I: Code Quality Standards -✅ **PASS** - Documentation feature does not introduce Rust code requiring clippy/rustdoc compliance. Focus is on markdown and tooling configuration. - -### Principle II: Testing Excellence -✅ **PASS (Resolved)** - Research identified validation strategy: -- Code examples: `cargo test --doc` validates Rust blocks -- YAML configs: Custom test harness with `serde_yaml` parsing -- Markdown links: `mdbook-linkcheck` plugin validates all references -- Schema alignment: Test assertions compare OpenAPI to generated schemas - -Implementation will include `tests/documentation_validation.rs` to enforce these checks. - -### Principle III: User Experience Consistency -✅ **PASS** - Documentation enhances UX by providing clear error message examples, configuration references, and troubleshooting guides. Aligns with existing logging/error handling patterns documented per constitution. - -### Principle IV: Performance Requirements -✅ **PASS** - Documentation generation is build-time activity, not runtime. Success criteria (SC-008) defines search response <3s which is achievable with mdbook's built-in search (<1s typical response time per research). - -### Overall Assessment - Post Phase 1 -✅ **ALL GATES PASSED** - Proceed to implementation (Phase 2 via /speckit.tasks command). All technical unknowns resolved, validation strategy confirmed, tooling selected. - -## Project Structure - -### Documentation (this feature) - -```text -specs/[###-feature]/ -├── plan.md # This file (/speckit.plan command output) -├── research.md # Phase 0 output (/speckit.plan command) -├── data-model.md # Phase 1 output (/speckit.plan command) -├── quickstart.md # Phase 1 output (/speckit.plan command) -├── contracts/ # Phase 1 output (/speckit.plan command) -└── tasks.md # Phase 2 output (/speckit.tasks command - NOT created by /speckit.plan) -``` - -### Source Code (repository root) - -```text -# Documentation structure (extends existing doc/ mdbook) -doc/ -├── book.toml # mdbook configuration (existing, may need updates) -├── SUMMARY.md # Table of contents (existing, extend with new sections) -├── index.md # Overview (existing, enhance per FR-001) -├── architecture/ # NEW: Architecture documentation (FR-002, FR-006) -│ ├── overview.md # System architecture, component relationships -│ ├── diagrams.md # Mermaid/SVG diagrams for architecture + data flow -│ └── data-flow.md # TSDB → flag metrics → health metrics flow (FR-005) -├── getting-started/ # NEW: Developer onboarding (FR-008, User Story 1) -│ ├── quickstart.md # Environment setup, first build -│ ├── project-structure.md # Module organization (FR-009) -│ └── development.md # Development workflow, testing, debugging -├── api/ # NEW: API reference (FR-003, FR-012) -│ ├── endpoints.md # /v1/health, /v1/maintenances details -│ ├── authentication.md # JWT mechanism (FR-017) -│ └── examples.md # Request/response examples (FR-011) -├── configuration/ # ENHANCE: Configuration reference (FR-004, FR-007) -│ ├── overview.md # Configuration structure overview -│ ├── schema.md # Complete field reference with types (FR-014) -│ ├── datasource.md # TSDB connection configuration -│ ├── metric-templates.md # Query templates, variables (FR-015) -│ ├── flag-metrics.md # Flag metric configuration, operators (FR-019) -│ ├── health-metrics.md # Health expressions (FR-016, FR-020) -│ ├── environments.md # Environment configuration -│ └── examples.md # Working configuration samples (FR-011, SC-007) -├── integration/ # NEW: TSDB backend integration (FR-010, User Story 5) -│ ├── interface.md # TSDB trait/interface requirements -│ ├── graphite.md # Existing Graphite implementation as reference -│ └── adding-backends.md # Guide for implementing new backends -├── modules/ # NEW: Rust module documentation (FR-009) -│ ├── overview.md # Module responsibilities, dependencies -│ ├── api.md # HTTP API module (axum handlers) -│ ├── config.md # Configuration parsing and validation -│ ├── types.md # Core data structures -│ ├── graphite.md # Graphite TSDB integration -│ └── common.md # Shared utilities -├── guides/ # NEW: Operational guides -│ ├── troubleshooting.md # Common issues, error resolution (FR-018) -│ └── deployment.md # Deployment patterns, configuration tips -├── convertor.md # ENHANCE: Expand existing convertor docs -├── reporter.md # ENHANCE: Expand existing reporter docs -└── config.md # DEPRECATE: Migrate content to configuration/ section - -# Generated artifacts (for AI consumption) -doc/schemas/ # NEW: Machine-readable schemas -├── config-schema.json # JSON Schema for configuration validation -├── types.json # Type definitions for AI autocomplete -└── patterns.json # Code patterns and conventions - -# Source code structure (unchanged - for reference) -src/ -├── lib.rs # Library entry point -├── api.rs # HTTP API module -├── api/ -│ └── v1.rs # v1 API handlers -├── config.rs # Configuration module -├── types.rs # Core data types -├── graphite.rs # Graphite TSDB integration -├── common.rs # Shared utilities -└── bin/ - ├── convertor.rs # Convertor binary - └── reporter.rs # Reporter binary - -tests/ # Integration tests (unchanged) -├── contract/ -├── integration/ -└── unit/ -``` - -**Structure Decision**: Extends existing mdbook documentation (`doc/`) with new sections for architecture, getting-started, api, integration, modules, and guides. Preserves existing mdbook.toml and SUMMARY.md structure while adding comprehensive coverage. New `doc/schemas/` directory provides machine-readable artifacts for AI tools (User Story 2). This single-project structure aligns with the Rust library + binaries pattern identified in Cargo.toml. - -## Complexity Tracking - -> **No violations - table not needed** - -All requirements align with Constitution principles. Documentation is additive and enhances existing project without introducing complexity violations. - ---- - -## Implementation Phases - -### Phase 0: Research ✅ COMPLETE - -**Objective**: Resolve all technical unknowns and select tooling - -**Completed**: -- ✅ Documentation validation approach (cargo test --doc, mdbook-linkcheck, custom YAML tests) -- ✅ Diagram tooling selection (Mermaid for version control + AI parsing) -- ✅ Schema generation strategy (schemars crate + build.rs automation) -- ✅ mdbook best practices (plugin configuration, structure, search) - -**Output**: `research.md` documenting all decisions and rationale - ---- - -### Phase 1: Design & Contracts ✅ COMPLETE - -**Objective**: Define data models, generate schemas, create quickstart guide - -**Completed**: -- ✅ `data-model.md`: 10 core entities with relationships and validation rules -- ✅ `contracts/config-schema.json`: JSON Schema for configuration validation -- ✅ `contracts/patterns.json`: AI-readable code patterns and conventions -- ✅ `contracts/README.md`: Schema usage guide for AI tools and IDEs -- ✅ `quickstart.md`: 30-minute developer onboarding guide -- ✅ Agent context updated: `.github/agents/copilot-instructions.md` - -**Architecture Decisions**: -1. **Documentation Structure**: Extends existing `doc/` mdbook with new sections -2. **Schema Generation**: Auto-generate from Rust types via build.rs -3. **Validation**: Multi-layered (cargo test, mdbook plugins, custom tests) -4. **AI Integration**: JSON schemas + patterns.json for code generation support - ---- - -### Phase 2: Implementation (Next - via /speckit.tasks) - -**Objective**: Create actual documentation content and tooling - -**Scope**: - -#### 2.1: Tooling Setup -- Install mdbook plugins (mermaid, linkcheck) -- Update `doc/book.toml` with plugin configuration -- Create `build.rs` for schema auto-generation -- Add `.vscode/settings.json` for IDE integration -- Configure pre-commit hooks for validation - -#### 2.2: Architecture Documentation -- Create `doc/architecture/overview.md` with system design -- Create `doc/architecture/diagrams.md` with Mermaid diagrams: - - System architecture (convertor, reporter, TSDB, dashboard) - - Module dependency graph -- Create `doc/architecture/data-flow.md` documenting TSDB → flag → health flow - -#### 2.3: Getting Started Section -- Migrate `quickstart.md` to `doc/getting-started/quickstart.md` -- Create `doc/getting-started/project-structure.md` -- Create `doc/getting-started/development.md` (testing, debugging, workflows) - -#### 2.4: API Documentation -- Create `doc/api/endpoints.md` from openapi-schema.yaml -- Create `doc/api/authentication.md` documenting JWT mechanism -- Create `doc/api/examples.md` with request/response samples - -#### 2.5: Configuration Documentation -- Expand `doc/config.md` → `doc/configuration/overview.md` -- Create `doc/configuration/schema.md` (reference from auto-generated JSON schema) -- Create individual pages: datasource.md, metric-templates.md, flag-metrics.md, health-metrics.md, environments.md -- Create `doc/configuration/examples.md` with working configurations -- Add validation test for all examples - -#### 2.6: Component Documentation -- Enhance `doc/convertor.md` with detailed flag metric evaluation process -- Enhance `doc/reporter.md` with polling logic and dashboard integration -- Add sequence diagrams for each component's workflow - -#### 2.7: Module Documentation -- Create `doc/modules/overview.md` with module responsibility matrix -- Create individual module pages: api.md, config.md, types.md, graphite.md, common.md -- Document public APIs, key types, and usage examples - -#### 2.8: Integration Guide -- Create `doc/integration/interface.md` defining TSDB trait requirements -- Create `doc/integration/graphite.md` as reference implementation -- Create `doc/integration/adding-backends.md` step-by-step guide - -#### 2.9: Operational Guides -- Create `doc/guides/troubleshooting.md` with common issues and solutions -- Create `doc/guides/deployment.md` with deployment patterns - -#### 2.10: Validation & Testing -- Create `tests/documentation_validation.rs`: - - Test all YAML examples parse correctly - - Test all code examples compile - - Test schema matches Rust structs -- Update CI pipeline to run documentation tests -- Add pre-commit hook for link checking - -#### 2.11: Schema Generation -- Generate `doc/schemas/config-schema.json` from Config struct -- Copy `contracts/patterns.json` → `doc/schemas/patterns.json` -- Create `doc/schemas/README.md` for consumers - -#### 2.12: Navigation & Polish -- Update `doc/SUMMARY.md` with all new sections -- Enhance `doc/index.md` with comprehensive overview -- Add cross-references between related documentation pages -- Ensure all diagrams render correctly -- Proofread all content for clarity and accuracy - -**Estimated Effort**: 16-24 hours across 12 subtasks - -**Success Criteria** (from spec.md): -- SC-001: New developers complete setup in <30 min ✅ (quickstart.md enables this) -- SC-002: AI agents generate correct code 90% of time ✅ (patterns.json + schemas) -- SC-003: Zero source-code-related support requests (comprehensive API docs) -- SC-004: 80% first-attempt config success (examples + schema validation) -- SC-006: 100% coverage of public APIs, configs, modules -- SC-007: All examples execute successfully (validation tests enforce) -- SC-008: Search response <3s (mdbook built-in achieves <1s) -- SC-009: Zero OpenAPI discrepancies (validation test enforces) - ---- - -## Next Steps - -1. **Review**: Stakeholders review plan.md, data-model.md, contracts/, quickstart.md -2. **Approve**: Obtain approval to proceed to implementation -3. **Task Generation**: Run `/speckit.tasks` command to generate dependency-ordered tasks.md -4. **Implementation**: Execute tasks via `/speckit.implement` or manual development -5. **Validation**: Run test suite and manual review against success criteria -6. **Merge**: Submit PR with all documentation and tooling changes - ---- - -## Artifacts Summary - -Generated by this planning phase: - -| Artifact | Location | Purpose | -|----------|----------|---------| -| Implementation Plan | plan.md | Overall strategy and phases | -| Research Findings | research.md | Technical decisions and rationale | -| Data Model | data-model.md | Entity definitions and relationships | -| Config Schema | contracts/config-schema.json | JSON Schema for validation | -| Patterns Documentation | contracts/patterns.json | AI-readable conventions | -| Contracts README | contracts/README.md | Schema usage guide | -| Quickstart Guide | quickstart.md | 30-min developer onboarding | -| Agent Context | .github/agents/copilot-instructions.md | Updated Copilot context | - -**Ready for**: Task generation and implementation diff --git a/specs/001-project-documentation/quickstart.md b/specs/001-project-documentation/quickstart.md deleted file mode 100644 index f7bbeae..0000000 --- a/specs/001-project-documentation/quickstart.md +++ /dev/null @@ -1,403 +0,0 @@ -# Quickstart: Developer Onboarding - -**Feature**: 001-project-documentation -**Target Audience**: New developers joining the metrics-processor project -**Time Estimate**: 30 minutes - -## What You'll Learn - -By the end of this guide, you will: -- ✅ Understand what metrics-processor does and why it exists -- ✅ Have a working local development environment -- ✅ Successfully build and run both binaries (convertor and reporter) -- ✅ Identify the major components and their responsibilities -- ✅ Know where to find key documentation sections - ---- - -## What is CloudMon Metrics Processor? - -**Problem**: Cloud monitoring produces many different metric types (latencies, status codes, rates). Visualizing overall service health from these disparate metrics is challenging. - -**Solution**: metrics-processor converts raw time-series metrics into simple semaphore-like health indicators: -- 🟢 **Green (0)**: Service up and running normally -- 🟡 **Yellow (1)**: Service degraded (slow, errors) -- 🔴 **Red (2)**: Service outage - -**Architecture** (high-level): - -``` -┌─────────────┐ ┌────────────────┐ ┌──────────────┐ ┌─────────────────┐ -│ Graphite │────▶│ Convertor │────▶│ Reporter │────▶│ Status Dashboard│ -│ (TSDB) │ │ (evaluates) │ │ (notifies) │ │ (displays) │ -└─────────────┘ └────────────────┘ └──────────────┘ └─────────────────┘ - Raw metrics Flag metrics Health status Semaphore UI -``` - -**Two Main Components**: -1. **Convertor**: Evaluates health from raw metrics, exposes HTTP API -2. **Reporter**: Polls convertor, sends updates to status dashboard - ---- - -## Prerequisites - -Before starting, ensure you have: - -- **Rust**: Version 1.75 or later (check with `rustc --version`) - - Install: `curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh` -- **Git**: For cloning the repository -- **Text Editor**: VSCode, IntelliJ, or vim with Rust support -- **Optional**: Docker (for containerized TSDB testing) - -**System Requirements**: -- Linux, macOS, or WSL2 on Windows -- 4GB RAM minimum -- 500MB disk space for dependencies - ---- - -## Step 1: Clone and Build (5 minutes) - -```bash -# Clone the repository -git clone https://github.com/your-org/metrics-processor.git -cd metrics-processor - -# Build all components -cargo build - -# Expected output: Compiling cloudmon-metrics v0.2.0 -# Should complete in 1-3 minutes depending on hardware -``` - -**Verify**: You should see two binaries created: -```bash -ls -lh target/debug/cloudmon-metrics-* -# cloudmon-metrics-convertor -# cloudmon-metrics-reporter -``` - ---- - -## Step 2: Understand the Project Structure (5 minutes) - -### Repository Layout - -``` -metrics-processor/ -├── src/ # Rust source code -│ ├── lib.rs # Library root -│ ├── api.rs # HTTP API module -│ ├── api/ -│ │ └── v1.rs # API v1 handlers -│ ├── config.rs # Configuration parsing -│ ├── types.rs # Domain data structures -│ ├── graphite.rs # Graphite TSDB integration -│ ├── common.rs # Shared utilities -│ └── bin/ -│ ├── convertor.rs # Convertor binary entry point -│ └── reporter.rs # Reporter binary entry point -├── doc/ # Documentation (mdbook) -├── tests/ # Integration tests -├── Cargo.toml # Rust dependencies -├── openapi-schema.yaml # API specification -└── README.md # Project overview -``` - -### Key Files to Know - -| File | Purpose | When to Edit | -|------|---------|-------------| -| `src/config.rs` | Configuration parsing & validation | Adding new config fields | -| `src/types.rs` | Core data structures (Config, FlagMetric, HealthMetric) | Changing data models | -| `src/api/v1.rs` | HTTP API handlers (`/v1/health`, `/v1/maintenances`) | Adding/modifying endpoints | -| `src/graphite.rs` | Graphite TSDB client | TSDB query changes | -| `openapi-schema.yaml` | API contract | Documenting API changes | - ---- - -## Step 3: Run Tests (3 minutes) - -```bash -# Run all tests (unit + integration) -cargo test - -# Expected output: test result: ok. X passed; 0 failed -``` - -**What's being tested**: -- Unit tests: Configuration parsing, metric evaluation logic -- Integration tests: HTTP endpoint behavior with mocked TSDB - -**If tests fail**: Check the output for specific failures. Common issues: -- Missing dependencies: Run `cargo build` first -- Port conflicts: Ensure ports 3005+ are available - ---- - -## Step 4: Run Convertor Locally (5 minutes) - -### Create a Minimal Configuration - -Create `config.yaml`: - -```yaml -datasource: - url: "http://localhost:8080" # Mock TSDB (won't actually connect yet) - type: graphite - -server: - address: "127.0.0.1" - port: 3005 - -metric_templates: - api_slow: - query: "stats.timers.api.$environment.$service.mean" - op: "gt" - threshold: 500 - -environments: - - name: "local-dev" - -flag_metrics: - - name: "api_slow" - service: "test_service" - template: - name: "api_slow" - environments: - - name: "local-dev" - -health_metrics: - test_service: - service: "test_service" - component_name: "Test Service" - category: "demo" - metrics: - - "test_service.api_slow" - expressions: - - expression: "test_service.api_slow" - weight: 1 -``` - -### Start Convertor - -```bash -cargo run --bin cloudmon-metrics-convertor -- --config config.yaml - -# Expected output: -# INFO cloudmon_metrics: Server starting on 127.0.0.1:3005 -``` - -### Test the API - -In another terminal: - -```bash -# Query health endpoint -curl "http://localhost:3005/v1/health?from=2024-01-01T00:00:00Z&to=2024-01-01T01:00:00Z&service=test_service&environment=local-dev" - -# Expected response (empty data since no real TSDB): -# {"name":"test_service","category":"demo","environment":"local-dev","metrics":[]} -``` - -**Success!** You've successfully run the convertor binary and made an API call. - ---- - -## Step 5: Explore the Codebase (10 minutes) - -### Understanding Flag Metrics - -Flag metrics are binary indicators (raised/lowered) based on raw TSDB queries: - -```rust -// src/types.rs -pub struct FlagMetric { - pub name: String, // "api_slow" - pub query: String, // TSDB query template - pub comparison: Comparison, // gt, lt, eq - pub threshold: f64, // 500.0 -} -``` - -**Flow**: -1. Query TSDB with variable substitution: `$service` → `test_service` -2. Compare result to threshold: `mean_latency > 500ms` -3. Set flag: `true` (raised) or `false` (lowered) - -### Understanding Health Metrics - -Health metrics combine flag metrics using boolean expressions: - -```rust -// src/types.rs -pub struct HealthMetric { - pub service: String, - pub expressions: Vec, // Boolean expressions -} - -pub struct Expression { - pub expr: String, // "api_slow || api_error_rate_high" - pub weight: u8, // 0=healthy, 1=degraded, 2=outage -} -``` - -**Flow**: -1. Evaluate each expression using flag states -2. Take maximum weight of matching expressions -3. Return semaphore value (0, 1, or 2) - ---- - -## Step 6: Common Development Tasks - -### Adding a New Metric Template - -1. Edit `config.yaml` → `metric_templates` section -2. Add new template with query, op, threshold -3. Reference in `flag_metrics` section -4. Restart convertor: `cargo run --bin cloudmon-metrics-convertor` - -### Running with Auto-Reload - -```bash -# Install cargo-watch -cargo install cargo-watch - -# Auto-rebuild on file changes -cargo watch -x 'run --bin cloudmon-metrics-convertor -- --config config.yaml' -``` - -### Debugging with Logs - -```bash -# Enable debug logging -RUST_LOG=debug cargo run --bin cloudmon-metrics-convertor -- --config config.yaml - -# Filter to specific module -RUST_LOG=cloudmon_metrics::api=debug cargo run ... -``` - -### Running Clippy (Linter) - -```bash -cargo clippy -# Fix all warnings before committing (Constitution requirement) -``` - ---- - -## Step 7: Key Documentation Sections - -Now that you have a working environment, explore these documentation sections: - -| Section | What to Learn | Location | -|---------|---------------|----------| -| **Architecture** | System design, data flow | `doc/architecture/` | -| **API Reference** | Endpoint details, authentication | `doc/api/` | -| **Configuration** | All config fields, examples | `doc/configuration/` | -| **Module Docs** | Rust module responsibilities | `doc/modules/` | -| **Troubleshooting** | Common issues, solutions | `doc/guides/troubleshooting.md` | - -**Next Steps**: -- Read [Architecture Overview](../../../doc/architecture/overview.md) to understand component interactions -- Review [Configuration Schema](./contracts/config-schema.json) for full config reference -- Check [patterns.json](./contracts/patterns.json) for coding conventions - ---- - -## Quick Reference Commands - -```bash -# Build -cargo build # Debug build -cargo build --release # Production build - -# Test -cargo test # All tests -cargo test --test integration # Integration tests only - -# Run -cargo run --bin cloudmon-metrics-convertor -- --config config.yaml -cargo run --bin cloudmon-metrics-reporter -- --config config.yaml - -# Lint -cargo clippy # Linter -cargo fmt # Auto-format - -# Documentation -cargo doc --open # Generate and open rustdoc -mdbook build doc/ && mdbook serve doc/ # Build user documentation -``` - ---- - -## Common Issues - -### "Could not compile `cloudmon-metrics`" - -**Cause**: Missing system dependencies or outdated Rust version - -**Solution**: -```bash -rustup update -cargo clean -cargo build -``` - -### "Address already in use" when starting convertor - -**Cause**: Port 3005 is occupied - -**Solution**: -```bash -# Change port in config.yaml -server: - port: 3006 - -# Or kill existing process -lsof -i :3005 -kill -``` - -### Tests failing with "connection refused" - -**Cause**: Integration tests expect mock TSDB, mockito setup issue - -**Solution**: Check that `mockito` dependency is present in `Cargo.toml` - ---- - -## Success Criteria - -You've successfully completed onboarding if you can: - -- ✅ Build the project without errors -- ✅ Run all tests successfully -- ✅ Start convertor binary and query the API -- ✅ Explain the difference between flag metrics and health metrics -- ✅ Identify which module to edit for different types of changes - -**Estimated Time**: If you completed this guide in ~30 minutes, you're ready to contribute! 🎉 - ---- - -## Getting Help - -- **Code Questions**: Check `doc/modules/` for module-specific documentation -- **Configuration Issues**: See `doc/configuration/schema.md` for field reference -- **Architecture Questions**: Read `doc/architecture/overview.md` -- **Bugs**: Check `doc/guides/troubleshooting.md` first, then file an issue - ---- - -## Next Steps - -1. Pick a starter issue from the issue tracker (look for "good first issue" label) -2. Read the relevant module documentation -3. Make your changes following the constitution guidelines (`.specify/memory/constitution.md`) -4. Run tests and clippy before committing -5. Submit a PR with clear description - -Welcome to the team! 🚀 diff --git a/specs/001-project-documentation/research.md b/specs/001-project-documentation/research.md deleted file mode 100644 index f0575c1..0000000 --- a/specs/001-project-documentation/research.md +++ /dev/null @@ -1,387 +0,0 @@ -# Research: Comprehensive Project Documentation - -**Feature**: 001-project-documentation -**Date**: 2025-01-23 -**Status**: Complete - -## Overview - -This document consolidates research findings for implementing comprehensive project documentation that serves both human developers and AI-powered tools. Research focused on four key areas: documentation validation, diagram tooling, AI-friendly schemas, and mdbook best practices. - ---- - -## 1. Documentation Validation & Testing - -### Decision: Multi-layered validation approach - -**Rationale**: Edge cases require ensuring documentation examples remain correct as code evolves. Rust ecosystem provides robust tooling for this. - -### Tools Selected - -| Validation Type | Tool | Purpose | -|----------------|------|---------| -| Code examples in docs | `cargo test --doc` | Validates Rust code blocks compile and run | -| YAML config examples | Custom test harness with `serde_yaml` | Ensures config examples parse correctly | -| OpenAPI sync | `utoipa` crate + test assertions | Auto-generate schema from code, validate against openapi-schema.yaml | -| Markdown links | `mdbook-linkcheck` plugin | Validates all internal/external links | - -### Implementation Pattern - -```rust -// tests/documentation_validation.rs -#[test] -fn validate_config_examples() { - let config_example = include_str!("../doc/configuration/examples.md"); - // Extract YAML blocks from markdown - let yaml_blocks = extract_yaml_from_markdown(config_example); - - for (i, yaml) in yaml_blocks.iter().enumerate() { - let parsed: Result = serde_yaml::from_str(yaml); - assert!(parsed.is_ok(), "Example {} failed to parse", i); - } -} -``` - -### Alternatives Considered - -- **Manual review**: Rejected due to high maintenance burden and human error risk -- **External validation services**: Rejected due to lack of Rust-specific tooling -- **CI-only validation**: Rejected - want pre-commit hooks to catch issues early - ---- - -## 2. Diagram Generation & Tooling - -### Decision: Mermaid diagrams for all architecture and data flow visualizations - -**Rationale**: Text-based format enables version control, AI parsing (FR-013), and browser rendering. Superior to binary formats (SVG/PNG) for maintainability. - -### Selected Format: Mermaid.js - -**Pros:** -- ✅ Text-based → Git diffs show semantic changes -- ✅ AI-parseable → LLMs can understand diagram structure -- ✅ Native browser rendering via `mdbook-mermaid` plugin -- ✅ No build step → Write in markdown, renders automatically -- ✅ Version control friendly → Merge conflicts are rare and readable - -**Cons:** -- Limited fine-grained styling (acceptable tradeoff) -- Complex diagrams become verbose (mitigated by splitting into multiple diagrams) - -### Plugin Configuration - -```toml -# doc/book.toml -[preprocessor.mermaid] -command = "mdbook-mermaid" - -[output.html] -additional-css = ["theme/mermaid.css"] -``` - -### Diagram Types Required - -| Diagram Type | Mermaid Type | Purpose | -|-------------|-------------|---------| -| Architecture Overview | `graph TB` | Show convertor, reporter, TSDB, dashboard relationships (FR-006) | -| Data Flow | `sequenceDiagram` | Show TSDB → flag metrics → health metrics flow (FR-005) | -| Configuration Structure | `graph LR` | Show configuration section relationships (FR-004) | -| Module Dependencies | `graph TD` | Show Rust module structure (FR-009) | - -### Alternatives Considered - -- **PlantUML**: Rejected - requires external server for rendering, less AI-friendly syntax -- **GraphViz**: Rejected - steeper learning curve, less intuitive for team -- **SVG/PNG**: Rejected - binary format breaks version control and AI parsing - ---- - -## 3. AI-Friendly Schema Generation - -### Decision: `schemars` crate + build.rs script for automatic JSON Schema generation - -**Rationale**: Enables IDE autocomplete and AI code generation (User Story 2) while maintaining single source of truth in Rust structs. - -### Implementation Architecture - -``` -Rust structs (src/config.rs) - ↓ [schemars derive] -JSON Schema (doc/schemas/config-schema.json) - ↓ [IDE reads] -Autocomplete in VSCode/IntelliJ - ↓ [AI tools read] -Code generation suggestions -``` - -### Tools Selected - -| Use Case | Tool | Format | -|---------|------|--------| -| Config struct → JSON Schema | `schemars` crate | JSON Schema Draft 7 | -| Runtime validation | `jsonschema` crate | Validates YAML against schema | -| Type exports for TypeScript | `ts-rs` crate (optional) | TypeScript interfaces | -| IDE integration | `.vscode/settings.json` | VSCode JSON schema mapping | - -### Generated Schema Structure - -```json -// doc/schemas/config-schema.json -{ - "$schema": "http://json-schema.org/draft-07/schema#", - "title": "CloudMon Metrics Configuration", - "type": "object", - "required": ["datasource", "server", "flag_metrics", "health_metrics"], - "properties": { - "datasource": { - "type": "object", - "properties": { - "url": {"type": "string", "format": "uri"}, - "type": {"type": "string", "enum": ["graphite"]} - } - } - // ... rest generated from Config struct - } -} -``` - -### Patterns Documentation for AI - -Created `doc/schemas/patterns.json` for conventions not captured in schemas: - -```json -{ - "patterns": [ - { - "name": "metric_template_variable_substitution", - "syntax": "$variable or ${variable}", - "available_variables": ["service", "environment"], - "example": "stats.timers.$service.$environment.mean" - }, - { - "name": "health_expression_syntax", - "operators": ["||", "&&", "!"], - "operands": "service.metric_name references", - "example": "api.slow || api.success_rate_low" - } - ], - "conventions": { - "naming": { - "services": "lowercase_with_underscores", - "environments": "lowercase-with-dashes", - "metrics": "service.metric_name format" - } - } -} -``` - -### Alternatives Considered - -- **Manual JSON Schema writing**: Rejected - high maintenance burden, prone to drift from code -- **OpenAPI only**: Rejected - doesn't cover configuration, only API -- **TypeScript-first approach**: Rejected - Rust is source of truth - ---- - -## 4. mdbook Configuration & Best Practices - -### Decision: Enhanced mdbook with search, mermaid, and collapsible sections - -**Rationale**: Existing `doc/` infrastructure works well. Enhance rather than replace to minimize migration effort. - -### Configuration Enhancements - -```toml -# doc/book.toml (enhancements) -[book] -title = "CloudMon Metrics Processor" -authors = ["CloudMon Team"] -language = "en" -multilingual = false -src = "." - -[preprocessor.mermaid] -command = "mdbook-mermaid" - -[preprocessor.linkcheck] -# Validates all links - -[output.html] -default-theme = "light" -preferred-dark-theme = "navy" -git-repository-url = "https://github.com/your-org/metrics-processor" - -[output.html.search] -enable = true -limit-results = 30 -use-boolean-and = true - -[output.html.fold] -enable = true # Collapsible sidebar sections -level = 1 # Fold chapters by default -``` - -### Documentation Structure - -``` -doc/ -├── SUMMARY.md # Navigation (enhanced) -├── index.md # Project overview (enhanced per FR-001) -├── architecture/ # NEW -│ ├── overview.md -│ ├── diagrams.md # Mermaid diagrams -│ └── data-flow.md -├── getting-started/ # NEW -│ ├── quickstart.md -│ ├── project-structure.md -│ └── development.md -├── api/ # NEW -│ ├── endpoints.md -│ ├── authentication.md -│ └── examples.md -├── configuration/ # ENHANCED (from config.md) -│ ├── overview.md -│ ├── schema.md -│ ├── datasource.md -│ ├── metric-templates.md -│ ├── flag-metrics.md -│ ├── health-metrics.md -│ ├── environments.md -│ └── examples.md -├── components/ # ENHANCED (from convertor.md, reporter.md) -│ ├── convertor.md -│ └── reporter.md -├── integration/ # NEW -│ ├── interface.md -│ ├── graphite.md -│ └── adding-backends.md -├── modules/ # NEW -│ ├── overview.md -│ ├── api.md -│ ├── config.md -│ ├── types.md -│ ├── graphite.md -│ └── common.md -└── guides/ # NEW - ├── troubleshooting.md - └── deployment.md -``` - -### Search Performance (SC-008) - -Built-in mdbook search achieves <1s response time for typical queries: -- Index size: ~200KB for 50 pages -- JavaScript-based client-side search -- No backend required - -### Versioning Strategy - -**Current**: Single version (0.2.0) -**Future**: Use Git tags + GitHub Pages branches when needed -**Not needed now**: Project is pre-1.0, breaking changes are acceptable - -### Alternatives Considered - -- **Docusaurus**: Rejected - requires Node.js ecosystem, heavier setup -- **Sphinx**: Rejected - Python-based, less natural for Rust projects -- **Custom static site**: Rejected - high maintenance burden -- **rustdoc only**: Rejected - not suitable for user guides and architecture docs - ---- - -## 5. Tooling Recommendations Summary - -### Development Tools - -```toml -# Cargo.toml additions -[dependencies] -schemars = "0.8" - -[build-dependencies] -schemars = "0.8" - -[dev-dependencies] -# (existing mockito, tempfile already suitable) -``` - -### Documentation Tools - -```bash -# Install once -cargo install mdbook -cargo install mdbook-mermaid -cargo install mdbook-linkcheck - -# Build documentation -mdbook build doc/ -mdbook serve doc/ # Local preview at http://localhost:3000 -``` - -### CI/CD Integration - -```yaml -# .github/workflows/docs.yml (or Zuul equivalent) -- name: Validate documentation - run: | - cargo test --doc - cargo test --test documentation_validation - mdbook build doc/ - # Check for broken links - mdbook test doc/ -``` - -### Pre-commit Hooks - -```yaml -# .pre-commit-config.yaml additions -- repo: local - hooks: - - id: mdbook-test - name: Validate documentation examples - entry: cargo test --test documentation_validation - language: system - pass_filenames: false -``` - ---- - -## 6. Implementation Phases - -### Phase 0: Research (COMPLETE) -✅ Validated tooling choices -✅ Identified implementation patterns -✅ Resolved technical unknowns - -### Phase 1: Design & Contracts (NEXT) -- Create data-model.md defining documentation structure entities -- Generate JSON schemas in contracts/ directory -- Write quickstart.md for developers -- Update agent context - -### Phase 2: Implementation (Future - via /speckit.tasks) -- Enhance mdbook.toml with plugins -- Create new documentation sections -- Generate schemas with build.rs -- Write mermaid diagrams -- Migrate existing docs -- Add validation tests - ---- - -## Decisions Log - -| Decision | Rationale | Risk Mitigation | -|---------|----------|-----------------| -| Mermaid over PlantUML | Version control + AI parsing | Document complex diagram patterns | -| schemars for schema gen | Single source of truth | Add validation tests | -| mdbook enhancement | Preserve existing work | Incremental migration | -| cargo test for validation | Native Rust tooling | Add pre-commit hooks | -| JSON Schema Draft 7 | Wide IDE support | Document VSCode setup | - ---- - -## Open Questions (None Remaining) - -All technical unknowns resolved. Ready to proceed to Phase 1: Design. diff --git a/specs/001-project-documentation/spec.md b/specs/001-project-documentation/spec.md deleted file mode 100644 index 588d5e9..0000000 --- a/specs/001-project-documentation/spec.md +++ /dev/null @@ -1,173 +0,0 @@ -# Feature Specification: Comprehensive Project Documentation - -**Feature Branch**: `001-project-documentation` -**Created**: 2025-01-23 -**Status**: Draft -**Input**: User description: "Create comprehensive project documentation including architecture diagrams, API documentation, data flow documentation, configuration reference, developer onboarding guide, code organization and module documentation, and integration documentation for TSDB backends" - -## User Scenarios & Testing *(mandatory)* - -### User Story 1 - New Developer Onboarding (Priority: P1) - -A new developer joins the team and needs to understand the metrics-processor project quickly to begin contributing. They need to understand what the project does, its architecture, how to set up their development environment, and where to find key components. - -**Why this priority**: This is the most critical use case because without proper onboarding documentation, new team members cannot effectively contribute, leading to productivity loss and increased onboarding time. This directly impacts team velocity and project maintenance. - -**Independent Test**: Can be fully tested by having a developer unfamiliar with the project follow only the onboarding documentation to set up their environment, understand the project purpose, and locate key components without external help. Success means they can identify the convertor and reporter components and explain their purpose within 30 minutes. - -**Acceptance Scenarios**: - -1. **Given** a new developer with Rust experience, **When** they read the project overview documentation, **Then** they understand the project converts raw TSDB metrics into semaphore-like health indicators -2. **Given** the developer needs to set up their environment, **When** they follow the setup instructions, **Then** they successfully build and run the project locally -3. **Given** the developer wants to understand component structure, **When** they review the architecture documentation, **Then** they can identify and explain the purpose of convertor and reporter binaries -4. **Given** the developer needs to modify configuration, **When** they consult the configuration reference, **Then** they can add a new flag metric or health metric correctly - ---- - -### User Story 2 - AI-Assisted Development (Priority: P1) - -AI agents, LLMs, and IDE assistants need to understand the project structure, conventions, and APIs to provide accurate code suggestions and automated refactoring without introducing bugs or violating project patterns. - -**Why this priority**: Modern development increasingly relies on AI-powered tools. Without machine-readable documentation (structured schemas, clear module boundaries, documented patterns), these tools provide inaccurate suggestions that reduce developer productivity and introduce technical debt. - -**Independent Test**: Can be tested by providing only the documentation to an AI agent and asking it to generate code for adding a new TSDB backend or creating a new metric template. Success means the generated code follows project conventions and correctly uses existing abstractions without guidance beyond the documentation. - -**Acceptance Scenarios**: - -1. **Given** an AI agent analyzing the codebase, **When** it reads the API schema documentation, **Then** it correctly understands request/response formats for all endpoints -2. **Given** an IDE assistant needs type information, **When** it accesses the data model documentation, **Then** it provides accurate autocomplete for all configuration structures -3. **Given** an LLM suggests refactoring, **When** it references the architecture documentation, **Then** it maintains proper separation between convertor and reporter components -4. **Given** an AI tool generates integration code, **When** it reads the TSDB integration documentation, **Then** it correctly implements the required interfaces for new backends - ---- - -### User Story 3 - API Integration (Priority: P2) - -External teams and services need to integrate with the metrics-processor API to query health metrics for their dashboards and monitoring systems. They need clear endpoint documentation, request/response examples, and error handling guidance. - -**Why this priority**: API consumers cannot integrate successfully without documentation. This is priority P2 because the API exists and works, but undocumented APIs block adoption and lead to support burden and integration errors. - -**Independent Test**: Can be tested by providing only the API documentation to a developer unfamiliar with the project and asking them to build a client that queries health metrics for multiple services. Success means they implement correct authentication, parameter handling, and response parsing without consulting the source code. - -**Acceptance Scenarios**: - -1. **Given** an external developer wants to query health metrics, **When** they read the API documentation, **Then** they understand all required and optional parameters for the `/v1/health` endpoint -2. **Given** they need to authenticate requests, **When** they consult the API reference, **Then** they correctly implement JWT token generation -3. **Given** they receive API responses, **When** they reference the response schema documentation, **Then** they correctly parse the ServiceData structure -4. **Given** an API request fails, **When** they check the error documentation, **Then** they understand the error and how to resolve it - ---- - -### User Story 4 - Configuration Management (Priority: P2) - -Operations teams and developers need to configure metrics-processor for new services, environments, or TSDB backends. They need comprehensive configuration reference with examples, validation rules, and troubleshooting guidance. - -**Why this priority**: Configuration errors are the primary source of runtime issues. Clear configuration documentation reduces deployment time, prevents misconfigurations, and enables self-service for operations teams. Priority P2 because configuration examples exist but lack comprehensive reference documentation. - -**Independent Test**: Can be tested by asking an operations engineer to configure monitoring for a new service with custom flag metrics and health expressions using only the configuration documentation. Success means they produce valid configuration without trial-and-error or source code inspection. - -**Acceptance Scenarios**: - -1. **Given** an ops engineer needs to add a new environment, **When** they read the configuration reference, **Then** they understand all fields in the `environments` section and their purposes -2. **Given** they want to create custom metric templates, **When** they consult the template documentation, **Then** they correctly use TSDB query syntax and comparison operators -3. **Given** they need to configure health expressions, **When** they reference the expression documentation, **Then** they write valid boolean expressions using available metrics -4. **Given** configuration validation fails, **When** they check the troubleshooting guide, **Then** they identify and fix the configuration error - ---- - -### User Story 5 - TSDB Backend Extension (Priority: P3) - -Developers need to add support for new TSDB backends beyond Graphite (e.g., Prometheus, InfluxDB). They need clear documentation on the integration interfaces, data transformation requirements, and testing approaches. - -**Why this priority**: Currently only Graphite is supported, limiting adoption. This is P3 because it's a future enhancement rather than current functionality, but documentation should guide future extensibility. - -**Independent Test**: Can be tested by asking a developer to implement a Prometheus backend using only the integration documentation. Success means they implement the correct trait/interface, handle query translation, and format responses correctly without extensive source code archaeology. - -**Acceptance Scenarios**: - -1. **Given** a developer wants to add Prometheus support, **When** they read the TSDB integration guide, **Then** they understand which traits or interfaces to implement -2. **Given** they need to translate queries, **When** they consult the query transformation documentation, **Then** they understand how to convert template queries to Prometheus PromQL -3. **Given** they implement the backend, **When** they reference the response format documentation, **Then** they correctly transform Prometheus responses to internal data structures -4. **Given** they complete implementation, **When** they follow the testing guide, **Then** they write appropriate integration tests matching existing patterns - ---- - -### User Story 6 - Architecture Understanding for Critical Changes (Priority: P2) - -Senior developers need to make architectural decisions or critical changes (refactoring, performance optimization, adding features) and must understand the system's design patterns, data flow, and key principles to avoid introducing issues. - -**Why this priority**: Critical changes without architectural understanding lead to technical debt, bugs, and maintenance issues. This is P2 because it enables safe evolution of the codebase and prevents costly mistakes during refactoring. - -**Independent Test**: Can be tested by asking a senior developer to design a performance optimization for the flag metrics evaluation using only the architecture and data flow documentation. Success means their design respects existing patterns, doesn't duplicate functionality, and correctly identifies performance bottlenecks. - -**Acceptance Scenarios**: - -1. **Given** a developer plans a refactoring, **When** they review the architecture documentation, **Then** they understand the separation between library code and binary components -2. **Given** they need to optimize query processing, **When** they study the data flow documentation, **Then** they identify where TSDB queries are executed and cached -3. **Given** they want to add a feature, **When** they read the design patterns documentation, **Then** they follow existing patterns for configuration, error handling, and logging -4. **Given** they need to understand dependencies, **When** they consult the module documentation, **Then** they identify which modules own which responsibilities and avoid tight coupling - ---- - -### Edge Cases - -- What happens when documentation describes features that no longer exist or have been significantly changed? -- How does the system ensure documentation stays synchronised with code changes over time? -- What happens when AI tools encounter ambiguous or conflicting documentation? -- How are examples in documentation validated to ensure they actually work? -- What happens when documentation needs to serve both human readers and machine parsers (AI agents)? -- How should documentation handle deprecated features or migration paths? - -## Requirements *(mandatory)* - -### Functional Requirements - -- **FR-001**: Documentation MUST include a project overview explaining the purpose of converting raw TSDB metrics into semaphore-like health indicators -- **FR-002**: Documentation MUST describe the architecture of the two main components (convertor and reporter) and their responsibilities -- **FR-003**: Documentation MUST provide a complete API reference for all HTTP endpoints including request parameters, response formats, and authentication -- **FR-004**: Documentation MUST include comprehensive configuration reference covering all sections (datasource, server, metric_templates, environments, flag_metrics, health_metrics, status_dashboard) -- **FR-005**: Documentation MUST describe data flow from TSDB query through flag metric evaluation to health metric calculation -- **FR-006**: Documentation MUST include diagrams illustrating system architecture and data flow -- **FR-007**: Documentation MUST document all configuration validation rules and constraints -- **FR-008**: Documentation MUST provide setup instructions for local development environment -- **FR-009**: Documentation MUST describe the module structure and purpose of each Rust module (api, config, types, graphite, common) -- **FR-010**: Documentation MUST document the integration interface for TSDB backends -- **FR-011**: Documentation MUST include working examples for common configuration scenarios (adding services, environments, custom metrics) -- **FR-012**: Documentation MUST align with the existing OpenAPI schema (openapi-schema.yaml) -- **FR-013**: Documentation MUST be structured to be parseable by AI agents and IDE assistants -- **FR-014**: Documentation MUST include type definitions for all configuration structures -- **FR-015**: Documentation MUST document the query template system and variable substitution -- **FR-016**: Documentation MUST explain the expression evaluation system for health metrics -- **FR-017**: Documentation MUST describe JWT authentication mechanism for status dashboard integration -- **FR-018**: Documentation MUST include troubleshooting guide for common configuration and runtime issues -- **FR-019**: Documentation MUST document all comparison operators (lt, gt, eq) and their usage -- **FR-020**: Documentation MUST describe the relationship between flag metrics and health metrics - -### Key Entities *(include if feature involves data)* - -- **Project Documentation Structure**: Overall organisation of documentation including sections for overview, architecture, API reference, configuration, guides, and integration -- **Architecture Diagram**: Visual representation showing convertor binary, reporter binary, TSDB backend, status dashboard, and their interactions -- **Data Flow Diagram**: Visual representation showing the flow from TSDB raw metrics → flag metrics → health metrics → status dashboard -- **API Endpoint Documentation**: Structured reference for each HTTP endpoint with parameters, responses, and examples -- **Configuration Schema**: Complete reference of all configuration sections with field descriptions, types, and validation rules -- **Module Documentation**: Description of each Rust module's purpose, public interfaces, and relationships -- **TSDB Integration Interface**: Abstract interface definition that TSDB backends must implement -- **Configuration Examples**: Working sample configurations demonstrating common use cases -- **Troubleshooting Guide**: Common issues, error messages, and resolution steps - -## Success Criteria *(mandatory)* - -### Measurable Outcomes - -- **SC-001**: New developers can complete environment setup and identify all major components within 30 minutes using only the documentation -- **SC-002**: AI agents correctly generate code following project conventions with 90% accuracy when provided only the documentation -- **SC-003**: API integration developers successfully implement clients without consulting source code, measured by zero source-code-related support requests -- **SC-004**: Operations teams configure new services without errors on first attempt in 80% of cases -- **SC-005**: Time to onboard new team member reduces from current baseline to under 4 hours -- **SC-006**: Documentation coverage includes 100% of public APIs, configuration options, and core modules -- **SC-007**: All configuration examples in documentation execute successfully without modification -- **SC-008**: Documentation search functionality returns relevant results for common queries within 3 seconds -- **SC-009**: Zero discrepancies between OpenAPI schema and API documentation -- **SC-010**: Architecture diagrams accurately reflect actual system structure, validated by team consensus -- **SC-011**: Developers successfully add new TSDB backend support following only integration documentation within 8 hours -- **SC-012**: Documentation maintenance burden reduces to under 2 hours per month after initial creation diff --git a/specs/001-project-documentation/tasks.md b/specs/001-project-documentation/tasks.md deleted file mode 100644 index 29cac20..0000000 --- a/specs/001-project-documentation/tasks.md +++ /dev/null @@ -1,409 +0,0 @@ -# Tasks: Comprehensive Project Documentation - -**Input**: Design documents from `/specs/001-project-documentation/` -**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/, quickstart.md - -**Tests**: Validation tests are included per research.md findings (cargo test, mdbook-linkcheck, example validation) - -**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. - -## Format: `[ID] [P?] [Story] Description` - -- **[P]**: Can run in parallel (different files, no dependencies) -- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3) -- Include exact file paths in descriptions - ---- - -## Phase 1: Setup (Shared Infrastructure) - -**Purpose**: Install tooling and configure infrastructure for documentation generation - -- [X] T001 Install mdbook and plugins: `cargo install mdbook mdbook-mermaid mdbook-linkcheck` -- [X] T002 Update doc/book.toml with preprocessor configuration for mermaid and linkcheck -- [X] T003 [P] Add schemars dependency to Cargo.toml for JSON schema generation -- [X] T004 [P] Update doc/SUMMARY.md with new documentation sections structure - ---- - -## Phase 2: Foundational (Blocking Prerequisites) - -**Purpose**: Core documentation infrastructure that MUST be complete before user story documentation can be created - -**⚠️ CRITICAL**: No user story work can begin until this phase is complete - -- [X] T006 Create build.rs script to auto-generate doc/schemas/config-schema.json from src/config.rs -- [X] T007 [P] Copy specs/001-project-documentation/contracts/patterns.json to doc/schemas/patterns.json -- [X] T008 [P] Create doc/schemas/README.md explaining schema usage for AI tools and IDEs -- [X] T009 Create validation test framework in tests/documentation_validation.rs -- [X] T010 [P] Add test function to validate all YAML examples parse correctly -- [X] T011 [P] Add test function to validate schema matches Config struct definition -- [X] T012 Enhance doc/index.md with comprehensive project overview per FR-001 - -**Checkpoint**: Foundation ready - user story documentation can now be created in parallel - ---- - -## Phase 3: User Story 1 - New Developer Onboarding (Priority: P1) 🎯 MVP - -**Goal**: New developers understand project purpose, set up environment, and locate components within 30 minutes - -**Independent Test**: A developer unfamiliar with the project follows only the onboarding documentation to set up environment, understand project purpose, and locate convertor/reporter components without external help - -### Validation for User Story 1 - -- [X] T013 [US1] Add test to validate quickstart.md examples compile and run in tests/documentation_validation.rs - -### Implementation for User Story 1 - -- [X] T014 [P] [US1] Migrate specs/001-project-documentation/quickstart.md to doc/getting-started/quickstart.md -- [X] T015 [P] [US1] Create doc/getting-started/project-structure.md documenting src/ layout and module organization -- [X] T016 [US1] Create doc/getting-started/development.md with testing, debugging, and workflow instructions -- [X] T017 [US1] Add getting-started section entries to doc/SUMMARY.md - -**Checkpoint**: Developer onboarding documentation complete - new developers can successfully set up environment independently - ---- - -## Phase 4: User Story 2 - AI-Assisted Development (Priority: P1) - -**Goal**: AI agents and IDE assistants understand project structure, conventions, and APIs to provide accurate code suggestions - -**Independent Test**: Provide only documentation to an AI agent and ask it to generate code for adding a new TSDB backend. Success means generated code follows project conventions without guidance beyond documentation. - -### Implementation for User Story 2 - -- [X] T018 [P] [US2] Run build.rs to generate doc/schemas/config-schema.json from Config struct -- [X] T019 [P] [US2] Create doc/schemas/types.json with type definitions for all core data structures -- [X] T020 [US2] Verify patterns.json includes all naming conventions and code patterns -- [X] T021 [US2] Update .github/agents/copilot-instructions.md with links to new schemas -- [X] T022 [US2] Add schema validation test ensuring schemas match Rust source in tests/documentation_validation.rs - -**Checkpoint**: AI/IDE integration complete - tools can access machine-readable schemas for accurate code generation - ---- - -## Phase 5: User Story 6 - Architecture Understanding (Priority: P2) - -**Goal**: Developers understand system design, data flow, and patterns to make safe architectural decisions - -**Independent Test**: A senior developer designs a performance optimization using only architecture documentation, respecting existing patterns and correctly identifying bottlenecks - -### Implementation for User Story 6 - -- [X] T023 [P] [US6] Create doc/architecture/overview.md describing convertor, reporter, TSDB, dashboard relationships -- [X] T024 [P] [US6] Create doc/architecture/diagrams.md with Mermaid system architecture diagram -- [X] T025 [P] [US6] Add Mermaid module dependency graph to doc/architecture/diagrams.md -- [X] T026 [US6] Create doc/architecture/data-flow.md documenting TSDB → flag metrics → health metrics flow with sequence diagram -- [X] T027 [US6] Enhance doc/convertor.md with detailed flag metric evaluation process -- [X] T028 [US6] Enhance doc/reporter.md with polling logic and dashboard integration details -- [X] T029 [US6] Add architecture section entries to doc/SUMMARY.md - -**Checkpoint**: Architecture documentation complete - developers can understand system design and data flow - ---- - -## Phase 6: User Story 3 - API Integration (Priority: P2) - -**Goal**: External teams integrate with metrics-processor API with clear endpoint documentation, examples, and error handling - -**Independent Test**: Developer unfamiliar with project builds an API client that queries health metrics using only API documentation, without consulting source code - -### Validation for User Story 3 - -- [X] T030 [US3] Add test to validate API documentation matches openapi-schema.yaml in tests/documentation_validation.rs - -### Implementation for User Story 3 - -- [X] T031 [P] [US3] Create doc/api/endpoints.md documenting /v1/health and /v1/maintenances from openapi-schema.yaml -- [X] T032 [P] [US3] Create doc/api/authentication.md documenting JWT token mechanism for status dashboard -- [X] T033 [US3] Create doc/api/examples.md with request/response samples for common scenarios -- [X] T034 [US3] Add API section entries to doc/SUMMARY.md - -**Checkpoint**: API documentation complete - external developers can integrate without source code access - ---- - -## Phase 7: User Story 4 - Configuration Management (Priority: P2) - -**Goal**: Operations teams and developers configure metrics-processor with comprehensive reference, examples, and troubleshooting - -**Independent Test**: Operations engineer configures monitoring for new service with custom flag metrics using only configuration documentation, producing valid configuration without trial-and-error - -### Validation for User Story 4 - -- [X] T035 [US4] Add test to validate all configuration examples parse and conform to schema in tests/documentation_validation.rs - -### Implementation for User Story 4 - -- [X] T036 [P] [US4] Create doc/configuration/overview.md with configuration structure introduction -- [X] T037 [P] [US4] Create doc/configuration/schema.md referencing auto-generated JSON schema with field descriptions -- [X] T038 [P] [US4] Create doc/configuration/datasource.md documenting TSDB connection configuration -- [X] T039 [P] [US4] Create doc/configuration/metric-templates.md documenting query templates and variable substitution -- [X] T040 [P] [US4] Create doc/configuration/flag-metrics.md documenting flag metric configuration and comparison operators -- [X] T041 [P] [US4] Create doc/configuration/health-metrics.md documenting health expressions and boolean operators -- [X] T042 [P] [US4] Create doc/configuration/environments.md documenting environment configuration -- [X] T043 [US4] Create doc/configuration/examples.md with working configuration samples for common scenarios -- [X] T044 [US4] Add configuration section entries to doc/SUMMARY.md - -**Checkpoint**: Configuration documentation complete - operations teams can configure without errors - ---- - -## Phase 8: User Story 5 - TSDB Backend Extension (Priority: P3) - -**Goal**: Developers add support for new TSDB backends with clear integration interfaces and testing guidance - -**Independent Test**: Developer implements Prometheus backend using only integration documentation, implementing correct traits and query translation without source code archaeology - -### Implementation for User Story 5 - -- [X] T045 [P] [US5] Create doc/integration/interface.md defining TSDB trait requirements and responsibilities -- [X] T046 [P] [US5] Create doc/integration/graphite.md documenting existing Graphite implementation as reference -- [X] T047 [US5] Create doc/integration/adding-backends.md with step-by-step guide for new backends -- [X] T048 [US5] Add integration section entries to doc/SUMMARY.md - -**Checkpoint**: Integration documentation complete - developers can extend TSDB support independently - ---- - -## Phase 9: Module Documentation (Supporting Multiple Stories) - -**Goal**: Document Rust module structure and responsibilities to support code navigation and understanding - -**Independent Test**: Developer identifies which module to edit for specific changes using only module documentation - -### Implementation for Module Documentation - -- [X] T049 [P] Create doc/modules/overview.md with module responsibility matrix -- [X] T050 [P] Create doc/modules/api.md documenting HTTP API module and axum handlers -- [X] T051 [P] Create doc/modules/config.md documenting configuration parsing and validation logic -- [X] T052 [P] Create doc/modules/types.md documenting core data structures (Config, FlagMetric, HealthMetric) -- [X] T053 [P] Create doc/modules/graphite.md documenting Graphite TSDB integration implementation -- [X] T054 [P] Create doc/modules/common.md documenting shared utilities and helper functions -- [X] T055 Create doc/modules/ section entries to doc/SUMMARY.md - -**Checkpoint**: Module documentation complete - developers understand code organization - ---- - -## Phase 10: Operational Guides (Supporting Multiple Stories) - -**Goal**: Provide troubleshooting and deployment guidance for operations teams and developers - -**Independent Test**: Operations engineer resolves configuration error using troubleshooting guide without external support - -### Implementation for Operational Guides - -- [X] T056 [P] Create doc/guides/troubleshooting.md with common issues, error messages, and solutions -- [X] T057 [P] Create doc/guides/deployment.md with deployment patterns and configuration tips -- [X] T058 Create doc/guides/ section entries to doc/SUMMARY.md - -**Checkpoint**: Operational guides complete - teams can troubleshoot and deploy independently - ---- - -## Phase 11: Validation & Testing - -**Purpose**: Ensure all documentation is accurate, links work, and examples are valid - -- [X] T059 Run cargo test --doc to validate Rust code examples compile -- [X] T060 Run cargo test --test documentation_validation to validate YAML examples and schemas -- [X] T061 Run mdbook test doc/ to validate markdown links with mdbook-linkcheck -- [X] T062 Run mdbook build doc/ to ensure documentation builds without errors -- [X] T063 Manually verify all diagrams render correctly in browser -- [X] T064 Update CI pipeline configuration to include documentation validation tests - -**Checkpoint**: All documentation validated - examples work, links resolve, schemas match code - ---- - -## Phase 12: Navigation & Polish - -**Purpose**: Final improvements for discoverability and user experience - -- [X] T065 Review and refine doc/SUMMARY.md table of contents for logical flow -- [X] T066 Add cross-references between related documentation pages -- [X] T067 Ensure consistent terminology across all documentation sections -- [X] T068 Add search keywords to page titles for better discoverability -- [X] T069 Proofread all content for clarity, grammar, and accuracy -- [X] T070 Generate final documentation with `mdbook build doc/` and review in browser -- [X] T071 Run quickstart.md validation with fresh developer environment - -**Checkpoint**: Documentation complete, polished, and ready for use - ---- - -## Dependencies & Execution Order - -### Phase Dependencies - -- **Setup (Phase 1)**: No dependencies - can start immediately -- **Foundational (Phase 2)**: Depends on Setup completion (T001-T005) - BLOCKS all user stories -- **User Stories (Phase 3-8)**: All depend on Foundational phase completion (T006-T012) - - US1 (Phase 3): Can start after Foundational - - US2 (Phase 4): Can start after Foundational - - US6 (Phase 5): Can start after Foundational - - US3 (Phase 6): Can start after Foundational - - US4 (Phase 7): Can start after Foundational - - US5 (Phase 8): Can start after Foundational -- **Module Docs (Phase 9)**: Can start after Foundational, supports all user stories -- **Operational Guides (Phase 10)**: Can start after Foundational, supports US4 -- **Validation (Phase 11)**: Depends on all documentation content being written -- **Polish (Phase 12)**: Depends on Validation completion - -### User Story Dependencies - -- **User Story 1 (P1)**: No dependencies on other stories - can proceed independently -- **User Story 2 (P1)**: No dependencies on other stories - can proceed independently -- **User Story 6 (P2)**: No dependencies on other stories - can proceed independently -- **User Story 3 (P2)**: No dependencies on other stories - can proceed independently -- **User Story 4 (P2)**: No dependencies on other stories - can proceed independently -- **User Story 5 (P3)**: No dependencies on other stories - can proceed independently - -All user stories are independently implementable after Foundational phase completion. - -### Within Each Phase - -- Setup: All [P] tasks can run in parallel (T003, T004, T005) -- Foundational: [P] tasks can run in parallel (T007-T008, T010-T011) -- User Stories: All [P] tasks within each story can run in parallel -- Module Documentation: All tasks (T049-T054) can run in parallel -- Operational Guides: Both tasks (T056-T057) can run in parallel - -### Parallel Opportunities - -**After Foundational Phase Completes, Maximum Parallelization:** - -```bash -# All user stories can proceed simultaneously: -Team Member 1: Phase 3 - US1 (Developer Onboarding) - T013-T017 -Team Member 2: Phase 4 - US2 (AI-Assisted Development) - T018-T022 -Team Member 3: Phase 5 - US6 (Architecture Understanding) - T023-T029 -Team Member 4: Phase 6 - US3 (API Integration) - T030-T034 -Team Member 5: Phase 7 - US4 (Configuration Management) - T035-T044 -Team Member 6: Phase 8 - US5 (TSDB Backend Extension) - T045-T048 -Team Member 7: Phase 9 - Module Documentation - T049-T055 -Team Member 8: Phase 10 - Operational Guides - T056-T058 -``` - ---- - -## Parallel Example: User Story 4 (Configuration Management) - -```bash -# Launch all documentation pages for US4 together (different files): -Task T036: "Create doc/configuration/overview.md with configuration structure introduction" -Task T037: "Create doc/configuration/schema.md referencing auto-generated JSON schema" -Task T038: "Create doc/configuration/datasource.md documenting TSDB connection configuration" -Task T039: "Create doc/configuration/metric-templates.md documenting query templates" -Task T040: "Create doc/configuration/flag-metrics.md documenting flag metric configuration" -Task T041: "Create doc/configuration/health-metrics.md documenting health expressions" -Task T042: "Create doc/configuration/environments.md documenting environment configuration" - -# Then complete the examples and integration tasks: -Task T043: "Create doc/configuration/examples.md with working configuration samples" -Task T044: "Add configuration section entries to doc/SUMMARY.md" -``` - ---- - -## Implementation Strategy - -### MVP First (User Stories 1 & 2 Only) - -1. Complete Phase 1: Setup (T001-T005) -2. Complete Phase 2: Foundational (T006-T012) - CRITICAL -3. Complete Phase 3: User Story 1 - Developer Onboarding (T013-T017) -4. Complete Phase 4: User Story 2 - AI-Assisted Development (T018-T022) -5. **STOP and VALIDATE**: Test that new developers can onboard in <30 min and AI tools can generate accurate code -6. Deploy documentation to team - -**Why this MVP?**: US1 and US2 are both P1 priority and deliver immediate value to current team members and AI-powered development tools. - -### Incremental Delivery - -1. **Foundation** (Phases 1-2): Setup + Schemas → Tools ready -2. **MVP** (Phases 3-4): US1 + US2 → Developer onboarding + AI assistance -3. **Architecture** (Phase 5): US6 → Architecture understanding -4. **External Integration** (Phases 6-7): US3 + US4 → API + Configuration docs -5. **Extensibility** (Phase 8): US5 → TSDB backend extension guide -6. **Supporting Docs** (Phases 9-10): Modules + Operational guides -7. **Quality** (Phases 11-12): Validation + Polish - -Each increment adds value without breaking previous deliverables. - -### Parallel Team Strategy - -With multiple developers after Foundational phase completion: - -1. **Team completes Setup + Foundational together** (2-4 hours) -2. **Once Foundational is done, parallelize by user story:** - - Developer A: US1 (Onboarding) - 2-3 hours - - Developer B: US2 (AI Integration) - 2-3 hours - - Developer C: US6 (Architecture) - 3-4 hours - - Developer D: US3 (API) - 2-3 hours - - Developer E: US4 (Configuration) - 4-5 hours - - Developer F: US5 (TSDB Integration) - 2-3 hours - - Developer G: Module Documentation - 3-4 hours - - Developer H: Operational Guides - 2-3 hours -3. **Merge all stories** → Validation phase (1-2 hours) -4. **Polish together** (1-2 hours) - -**Total Time**: -- Sequential: 24-32 hours (1 developer) -- Parallel: 8-12 hours (8 developers) - ---- - -## Task Summary - -| Phase | Task Count | Can Parallelize | Est. Time (Solo) | Est. Time (Team) | -|-------|-----------|-----------------|------------------|------------------| -| Phase 1: Setup | 5 | 3 tasks | 1-2 hours | 30 min | -| Phase 2: Foundational | 7 | 4 tasks | 2-4 hours | 1-2 hours | -| Phase 3: US1 (Onboarding) | 5 | 2 tasks | 2-3 hours | 1-2 hours | -| Phase 4: US2 (AI) | 5 | 2 tasks | 2-3 hours | 1-2 hours | -| Phase 5: US6 (Architecture) | 7 | 3 tasks | 3-4 hours | 2-3 hours | -| Phase 6: US3 (API) | 5 | 2 tasks | 2-3 hours | 1-2 hours | -| Phase 7: US4 (Configuration) | 9 | 7 tasks | 4-5 hours | 2-3 hours | -| Phase 8: US5 (Integration) | 4 | 2 tasks | 2-3 hours | 1-2 hours | -| Phase 9: Module Docs | 7 | 6 tasks | 3-4 hours | 1-2 hours | -| Phase 10: Operational Guides | 3 | 2 tasks | 2-3 hours | 1-2 hours | -| Phase 11: Validation | 6 | 0 tasks | 1-2 hours | 1-2 hours | -| Phase 12: Polish | 7 | 0 tasks | 2-3 hours | 2-3 hours | -| **TOTAL** | **71 tasks** | **33 parallel** | **26-39 hours** | **15-24 hours** | - ---- - -## Success Criteria Mapping - -Tasks are designed to achieve all success criteria from spec.md: - -- **SC-001** (30-min onboarding): Phase 3 (US1) - T014-T017 -- **SC-002** (90% AI accuracy): Phase 4 (US2) - T018-T022 -- **SC-003** (Zero source-code support): Phase 6 (US3) - T031-T034 -- **SC-004** (80% first-attempt config): Phase 7 (US4) - T036-T044 -- **SC-005** (<4 hour onboarding): Phases 3-10 combined -- **SC-006** (100% coverage): All phases 3-10 -- **SC-007** (All examples work): Phase 11 (Validation) - T035, T059-T064 -- **SC-008** (<3s search): Built-in mdbook search (no tasks needed) -- **SC-009** (Zero OpenAPI discrepancies): Phase 6 (US3) - T030 -- **SC-010** (Accurate diagrams): Phase 5 (US6) - T024-T026 + T063 -- **SC-011** (8-hour TSDB backend): Phase 8 (US5) - T045-T048 -- **SC-012** (<2hr/month maintenance): Build automation via T006 + validation via T059-T064 - ---- - -## Notes - -- **[P] marker**: Tasks that can run in parallel (different files, no dependencies on incomplete tasks) -- **[Story] label**: Maps task to specific user story for traceability (US1, US2, US3, US4, US5, US6) -- **File paths**: All paths are absolute from repository root -- **Tests**: Validation tests included per research findings (not optional for documentation feature) -- **Checkpoints**: Each phase ends with a checkpoint to validate story independently -- **Format compliance**: All tasks follow strict checklist format: `- [ ] [TaskID] [P?] [Story?] Description with file path` - -**Critical Path**: Setup → Foundational → {All User Stories in parallel} → Validation → Polish - -**Minimum Viable Product**: Phases 1-4 (Setup + Foundational + US1 + US2) = Core onboarding + AI assistance diff --git a/specs/002-functional-test-suite/IMPLEMENTATION_REPORT.md b/specs/002-functional-test-suite/IMPLEMENTATION_REPORT.md deleted file mode 100644 index 8d9f2c5..0000000 --- a/specs/002-functional-test-suite/IMPLEMENTATION_REPORT.md +++ /dev/null @@ -1,343 +0,0 @@ -# Test Suite Implementation - Final Report - -**Date**: 2025-01-21 -**Feature**: Comprehensive Functional Test Suite -**Status**: Foundation Complete, Implementation In Progress - -## Executive Summary - -Successfully established a comprehensive test infrastructure for the metrics-processor project, implementing **30 out of 80 planned tasks (37.5%)**. The foundation enables rapid development of the remaining test suite with minimal overhead. - -### Key Achievements ✅ - -1. **Robust Test Infrastructure** - - 35+ reusable test fixtures covering all scenarios - - 10+ helper functions eliminating test duplication - - Custom assertions with clear error messages - - CI/CD integration with automated coverage - -2. **Core Business Logic Coverage** - - 11 comprehensive tests for metric flag evaluation - - 100% coverage of comparison operators (Lt, Gt, Eq) - - Edge case validation (null, negative, boundary, zero) - - Regression detection validated - -3. **Fast Test Execution** - - Current suite: < 0.05 seconds - - Well below 2-minute target - - Enables rapid development feedback - -4. **Professional Documentation** - - Comprehensive testing guide (`docs/testing.md`) - - Clear implementation patterns - - CI/CD setup instructions - -## Current Coverage: 42.96% - -### Coverage by Module - -| Module | Lines Covered | Total Lines | Coverage % | Status | -|--------|--------------|-------------|------------|--------| -| `api/v1.rs` | 0 | 39 | 0.0% | ❌ Not Started | -| `common.rs` | 7 | 75 | 9.3% | 🚧 Partial | -| `config.rs` | 27 | 33 | 81.8% | ✅ Good | -| `types.rs` | 54 | 69 | 78.3% | ✅ Good | -| `graphite.rs` | 98 | 213 | 46.0% | 🚧 Partial | -| **TOTAL** | **186** | **433** | **42.96%** | 🚧 **In Progress** | - -### Gap Analysis - -**Critical Gaps** (High Impact): -1. `get_service_health()` in `common.rs` - 0% covered - - Most complex business logic - - Integrates multiple components - - **Impact**: 20-25% coverage gain when tested - -2. API endpoint handlers in `api/v1.rs` - 0% covered - - Public interface - - Error handling - - **Impact**: 10-15% coverage gain - -**Medium Gaps** (Medium Impact): -3. Graphite integration in `graphite.rs` - 46% covered - - Response parsing - - Error handling - - **Impact**: 8-12% coverage gain - -4. Additional config validation - 82% covered - - Edge cases - - Template substitution - - **Impact**: 3-5% coverage gain - -## Work Completed - -### Phase 1: Setup (4 tasks) ✅ -Created comprehensive test infrastructure: -- **3 fixture modules** with 35+ fixtures -- **Test configurations**: 11 YAML config scenarios -- **Mock responses**: 25+ Graphite response fixtures -- **Helper functions**: 10+ utilities - -**Files**: -- `tests/fixtures/mod.rs` (311 bytes) -- `tests/fixtures/configs.rs` (6.5KB) -- `tests/fixtures/graphite_responses.rs` (9.4KB) -- `tests/fixtures/helpers.rs` (11KB) - -### Phase 2: Foundational (5 tasks) ✅ -Established CI/CD and testing utilities: -- **Coverage CI**: GitHub Actions workflow -- **Test helpers**: State creation, mocking, assertions -- **Docker ignore**: Build optimization - -**Files**: -- `.github/workflows/coverage.yml` (976 bytes) -- `.dockerignore` (408 bytes) - -### Phase 3: User Story 1 (11 tasks) ✅ -**Core Metric Flag Evaluation Tests** - -Implemented 11 comprehensive unit tests in `src/common.rs`: -- ✅ Lt operator (2 tests): below and above threshold -- ✅ Gt operator (2 tests): above and below threshold -- ✅ Eq operator (2 tests): equal and not equal -- ✅ None value handling (1 test) -- ✅ Boundary conditions (1 test) -- ✅ Negative values (1 test) -- ✅ Zero threshold (1 test) -- ✅ Mixed operators (1 test) - -**Coverage**: 100% of `get_metric_flag_state()` function - -### Phase 4: User Story 6 (5 tasks) ✅ -**Regression Suite Validation** - -- ✅ Regression detection validated (8/11 tests catch operator swap) -- ✅ Zero false positives confirmed -- ✅ Fast execution verified (< 0.05 seconds) -- ✅ Documentation created (`docs/testing.md`, 7KB) -- ✅ Test patterns established - -## Work Remaining - -### Phase 5: User Story 2 (11 tasks) ⏳ -**Service Health Aggregation - NOT STARTED** - -Priority: **P1 - Critical** - -Tests needed in `src/common.rs` for `get_service_health()`: -- Expression evaluation (OR, AND logic) -- Weighted health calculations -- Error handling (unknown service/environment) -- End-to-end with mocked Graphite -- Edge cases (empty data, partial data) - -**Estimated Impact**: +20-25% coverage - -### Phase 6: User Story 4 (11 tasks) ⏳ -**Configuration Processing - PARTIALLY COMPLETE** - -Priority: **P2 - Important** - -Existing: 3 tests in `config.rs`, 1 test in `types.rs` - -Additional tests needed: -- Template variable substitution -- Multiple environment expansion -- Threshold overrides -- Dash-to-underscore conversion -- Validation errors - -**Estimated Impact**: +3-5% coverage - -### Phase 7: User Story 3 (13 tasks) ⏳ -**API Endpoints - NOT STARTED** - -Priority: **P2 - Important** - -Tests needed in `src/api/v1.rs`: -- REST endpoint handlers (/health, /info, /render, /find) -- Response format validation -- Error codes (400, 409, 500) -- Integration tests - -**Estimated Impact**: +10-15% coverage - -### Phase 8: User Story 5 (10 tasks) ⏳ -**Graphite Integration - PARTIALLY COMPLETE** - -Priority: **P3 - Nice to Have** - -Existing: 3 tests in `graphite.rs` - -Additional tests needed: -- Query building edge cases -- Response parsing (malformed, empty) -- Error handling (4xx, 5xx, timeout) -- Null/NaN value handling - -**Estimated Impact**: +8-12% coverage - -### Phase 9: Polish (10 tasks) ⏳ -**Coverage Validation - PARTIALLY COMPLETE** - -Tasks remaining: -- Gap identification and filling -- CI enforcement configuration -- HTML report generation -- Final validation -- Documentation updates - -**Estimated Impact**: +2-5% coverage (gap filling) - -## Path to 95% Coverage - -### Critical Path (Must Complete) - -**Priority 1: Service Health Tests (Phase 5)** -- Lines to cover: ~60-70 lines in `common.rs` -- Complexity: High (integration, expression evaluation) -- Time estimate: 4-6 hours -- Coverage gain: +20-25% - -**Priority 2: API Endpoint Tests (Phase 7)** -- Lines to cover: ~40 lines in `api/v1.rs` -- Complexity: Medium (handlers, error responses) -- Time estimate: 3-4 hours -- Coverage gain: +10-15% - -**Priority 3: Graphite Integration (Phase 8)** -- Lines to cover: ~50-60 lines in `graphite.rs` -- Complexity: Medium (mocking, parsing) -- Time estimate: 2-3 hours -- Coverage gain: +8-12% - -**Priority 4: Config Tests (Phase 6)** -- Lines to cover: ~10-15 lines in `config.rs` + `types.rs` -- Complexity: Low (validation, substitution) -- Time estimate: 2 hours -- Coverage gain: +3-5% - -**Total Estimated Time**: 11-15 hours to reach 95% coverage - -### Recommended Implementation Schedule - -**Week 1** (6-8 hours): -- Complete Phase 5 (Service Health) -- Expected coverage: 43% → 65-68% - -**Week 2** (5-7 hours): -- Complete Phase 7 (API Endpoints) -- Complete Phase 8 (Graphite) -- Expected coverage: 68% → 85-90% - -**Week 3** (2-3 hours): -- Complete Phase 6 (Config) -- Fill gaps identified in coverage report -- Expected coverage: 90% → 95%+ - -## Success Metrics Status - -| Metric | Target | Current | Status | Gap | -|--------|--------|---------|--------|-----| -| **Code Coverage** | ≥95% | 42.96% | 🚧 | -52% | -| **Test Count** | ≥50 | 18 | 🚧 | -32 tests | -| **Execution Time** | <2 min | <0.05s | ✅ | None | -| **Regression Detection** | 100% | 100% | ✅ | None | -| **False Positives** | 0 | 0 | ✅ | None | - -### What's Complete ✅ -- ✅ Test infrastructure (100%) -- ✅ Core metric evaluation (100% function coverage) -- ✅ Fast execution (way under target) -- ✅ Regression detection (validated) -- ✅ Documentation (comprehensive) - -### What's Remaining ⏳ -- 🚧 Service health aggregation (0% of function) -- 🚧 API endpoint tests (0%) -- 🚧 Additional Graphite tests (54% of module remaining) -- 🚧 Additional config tests (18% of module remaining) -- 🚧 Coverage gap filling - -## Risk Assessment - -### Risk Level: **LOW** - -**Rationale**: -1. **Foundation is Solid**: Infrastructure proven and working -2. **Patterns Established**: Clear examples for remaining tests -3. **Reusable Components**: Fixtures and helpers ready to use -4. **Time Estimate Reasonable**: 11-15 hours to completion - -**Known Challenges**: -1. Service health testing requires async mocking (mockito ready) -2. API endpoint testing needs request simulation (established pattern) -3. Coverage gap filling may require creative scenarios - -**Mitigation**: -- All fixtures already created for remaining phases -- Helper functions eliminate boilerplate -- Existing tests provide clear patterns - -## Recommendations - -### Immediate Next Steps - -1. **Run coverage HTML report** to visualize gaps: - ```bash - cargo tarpaulin --out Html --output-dir ./coverage - open coverage/tarpaulin-report.html - ``` - -2. **Prioritize Service Health Tests (Phase 5)**: - - Highest impact on coverage - - Most complex business logic - - Already have all fixtures needed - -3. **Use Established Patterns**: - - Copy test structure from Phase 3 - - Leverage `create_multi_metric_test_state()` helper - - Use `setup_graphite_mock()` for HTTP mocking - -### Long-term Recommendations - -1. **Maintain Test-First Approach**: Write tests before new features -2. **Enforce Coverage in CI**: Add `--fail-under 90` to coverage workflow -3. **Regular Gap Analysis**: Run `cargo tarpaulin` weekly -4. **Update Documentation**: Keep `docs/testing.md` current - -## Conclusion - -Successfully established a **professional-grade test infrastructure** for the metrics-processor project. The foundation is complete with: - -- 35+ reusable fixtures -- 10+ helper functions -- Custom assertions -- CI/CD integration -- Comprehensive documentation - -The remaining **50 tasks** follow established patterns and can be completed efficiently using the provided infrastructure. With an estimated **11-15 hours of focused work**, the project can achieve the **95% coverage target**. - -### Key Takeaways - -✅ **What Works Well**: -- Test infrastructure is excellent -- Execution performance is exceptional (< 0.05s) -- Documentation is professional -- Patterns are clear and reusable - -⚠️ **What Needs Attention**: -- Service health function (highest priority) -- API endpoint coverage (public interface) -- Remaining Graphite integration - -🎯 **Bottom Line**: The project is well-positioned to achieve 95% coverage. The hard work of infrastructure setup is complete, and the remaining tests can be implemented rapidly using established patterns. - ---- - -**Implementation Guide**: See `docs/testing.md` for detailed instructions on adding new tests. - -**Coverage Reports**: Run `cargo tarpaulin --out Html` to generate visual coverage reports. - -**Questions**: All test patterns are documented with examples in the existing test modules. diff --git a/specs/002-functional-test-suite/IMPLEMENTATION_STATUS.txt b/specs/002-functional-test-suite/IMPLEMENTATION_STATUS.txt deleted file mode 100644 index bc5f881..0000000 --- a/specs/002-functional-test-suite/IMPLEMENTATION_STATUS.txt +++ /dev/null @@ -1,206 +0,0 @@ -================================================================================ -IMPLEMENTATION STATUS: Comprehensive Functional Test Suite -================================================================================ -Date: 2025-01-24 -Status: 51.25% Complete (41/80 tasks) - -================================================================================ -COMPLETED WORK -================================================================================ - -Phase 1: Test Infrastructure Setup ✅ (4/4 tasks) --------------------------------------------------- -✓ T001-T004: Fixtures module structure with configs, graphite responses, helpers - Location: tests/fixtures/ - Files: mod.rs, configs.rs, graphite_responses.rs, helpers.rs - -Phase 2: Foundational Prerequisites ✅ (5/5 tasks) --------------------------------------------------- -✓ T005-T009: Coverage tooling and test helper functions - - cargo-tarpaulin integration - - create_test_state helpers - - Custom assertions (assert_metric_flag, assert_health_score) - -Phase 3: User Story 1 - Core Metric Flag Tests ✅ (11/11 tasks) ---------------------------------------------------------------- -✓ T010-T020: Comprehensive tests for get_metric_flag_state function - Location: src/common.rs (test module) - Coverage: - - Lt/Gt/Eq operators (6 tests) - - None value handling (1 test) - - Boundary conditions (1 test) - - Negative values (1 test) - - Zero threshold (1 test) - - Mixed operators (1 test) - -Phase 4: User Story 6 - Regression Suite ✅ (5/5 tasks) -------------------------------------------------------- -✓ T021-T025: Regression detection and test suite validation - - All tests execute successfully - - Tests catch intentional breaking changes - - Fast execution (under 2 minutes) - -Phase 5: User Story 2 - Service Health Aggregation ✅ (11/11 tasks) -------------------------------------------------------------------- -✓ T026-T033: Unit tests for health calculation logic - Location: src/common.rs (test module) - Coverage: - - Single metric OR expressions (1 test) - - Two metrics AND expressions (2 tests) - - Weighted expressions (1 test) - - All false expressions (1 test) - - Error handling: unknown service/environment (2 tests) - - Multiple datapoints time series (1 test) - -✓ T034-T036: Integration tests for end-to-end flows - Location: tests/integration_health.rs - Coverage: - - End-to-end health calculation with mocked Graphite (1 test) - - Complex weighted expression scenarios (1 test) - - Edge cases: empty datapoints and partial data (1 test) - -================================================================================ -TEST METRICS -================================================================================ - -Total Tests: 36 passing - - src/common.rs: 19 tests (Phases 3 & 5) - - src/config.rs: 3 tests (existing) - - src/graphite.rs: 3 tests (existing) - - src/types.rs: 1 test (existing) - - tests/integration_health.rs: 3 tests (Phase 5) - - tests/documentation_validation.rs: 7 tests (existing) - -Target: 50+ tests (72% achieved) -Coverage Goal: 95%+ for core business functions - -================================================================================ -REMAINING WORK -================================================================================ - -Phase 6: User Story 4 - Configuration Processing (11 tasks) ------------------------------------------------------------- -[ ] T037-T047: Configuration loading and template expansion tests - Priority: P2 - Files to test: src/types.rs, src/config.rs - Coverage needed: - - Template variable substitution ($environment, $service) - - Multiple environments expansion - - Per-environment threshold overrides - - Dash-to-underscore conversion - - Service set population - - Health metrics expression copying - - Invalid YAML syntax handling - - Missing required fields validation - - Default values application - - get_socket_addr validation - - Config loading from multiple sources - -Phase 7: User Story 3 - API Endpoint Tests (13 tasks) ------------------------------------------------------- -[ ] T048-T060: API endpoint integration tests - Priority: P2 - Files to test: src/api/v1.rs, src/graphite.rs - Coverage needed: - - /api/v1/ root endpoint - - /api/v1/info endpoint - - /api/v1/health endpoint (success, error cases) - - /render endpoint (flag/health targets) - - /metrics/find hierarchy levels - - /functions endpoint - - /tags/autoComplete/tags endpoint - - Full API integration test - - Error response format validation - -Phase 8: User Story 5 - Graphite Integration (10 tasks) --------------------------------------------------------- -[ ] T061-T070: Graphite client tests - Priority: P3 - File to test: src/graphite.rs - Coverage needed: - - Query building - - Valid JSON parsing - - Empty datapoints handling - - HTTP error handling (4xx, 5xx) - - Malformed JSON handling - - Connection timeout handling - - Metric discovery with wildcards - - Null and NaN value handling - - Partial response handling - -Phase 9: Coverage Validation & Polish (10 tasks) -------------------------------------------------- -[ ] T071-T080: Coverage validation and documentation - Priority: Final - Tasks: - - Run cargo tarpaulin - - Verify 95% coverage threshold - - Identify and fill coverage gaps - - Add coverage enforcement to CI - - Generate HTML coverage report - - Verify execution time < 2 minutes - - Document testing approach - - Add test execution commands reference - - Verify regression detection - - Final validation against all FR requirements - -================================================================================ -KEY ACHIEVEMENTS -================================================================================ - -✓ Completed 51% of implementation tasks -✓ 36 tests passing with 0 failures -✓ Core business logic fully tested (metric flags, health aggregation) -✓ Comprehensive test infrastructure in place -✓ Mockito integration working correctly for async tests -✓ Integration tests demonstrate end-to-end flows -✓ All test execution is fast (< 2 minutes total) -✓ Tests follow TDD principles and test existing code behavior - -================================================================================ -TECHNICAL NOTES -================================================================================ - -Test Infrastructure: -- Using Rust's built-in test framework with #[test] and #[tokio::test] -- Mockito for HTTP mocking with proper async/await support -- Test helpers in tests/fixtures/helpers.rs for reusable components -- Custom assertions for clear failure messages - -Mock Server Pattern: -- Use mockito::Server::new_async().await for async tests -- Always include .match_query() with at least format=json matcher -- Create mocks before calling functions that make HTTP requests -- Use .create() (not .create_async()) for immediate mock registration - -Expression Evaluation: -- Metric names like "service-name.metric" become "service_name.metric" in expressions -- The dash-to-underscore replacement happens in context building -- Expressions are evaluated with evalexpr crate -- Missing metrics default to false in expression context - -================================================================================ -NEXT STEPS FOR CONTINUATION -================================================================================ - -Immediate Priority (Phase 6): -1. Implement configuration processing tests (T037-T047) -2. Focus on AppState::process_config validation -3. Test template variable substitution -4. Verify environment expansion logic - -To Continue Implementation: -$ cd /Users/A107229207/dev/otc/stackmon/metrics-processor -$ cargo test --lib types::test # Run existing config tests -$ # Add new tests to src/types.rs test module -$ # Add new tests to src/config.rs test module - -To Check Coverage: -$ cargo install cargo-tarpaulin # If not already installed -$ cargo tarpaulin --out Html --out Lcov - -To Run All Tests: -$ cargo test --lib --tests # Run all unit and integration tests -$ cargo test -- --nocapture # Run with output visible - -================================================================================ diff --git a/specs/002-functional-test-suite/checklists/requirements.md b/specs/002-functional-test-suite/checklists/requirements.md deleted file mode 100644 index b830dd3..0000000 --- a/specs/002-functional-test-suite/checklists/requirements.md +++ /dev/null @@ -1,127 +0,0 @@ -# Specification Quality Checklist: Comprehensive Functional Test Suite - -**Purpose**: Validate specification completeness and quality before proceeding to planning -**Created**: 2025-01-24 -**Feature**: [spec.md](../spec.md) - -## Content Quality - -- [x] No implementation details (languages, frameworks, APIs) -- [x] Focused on user value and business needs -- [x] Written for non-technical stakeholders -- [x] All mandatory sections completed - -## Requirement Completeness - -- [x] No [NEEDS CLARIFICATION] markers remain -- [x] Requirements are testable and unambiguous -- [x] Success criteria are measurable -- [x] Success criteria are technology-agnostic (no implementation details) -- [x] All acceptance scenarios are defined -- [x] Edge cases are identified -- [x] Scope is clearly bounded -- [x] Dependencies and assumptions identified - -## Feature Readiness - -- [x] All functional requirements have clear acceptance criteria -- [x] User scenarios cover primary flows -- [x] Feature meets measurable outcomes defined in Success Criteria -- [x] No implementation details leak into specification - -## Validation Notes - -### Content Quality Review -✅ **PASS**: Specification focuses purely on testing requirements without mentioning specific testing frameworks, tools, or implementation approaches. Uses technology-agnostic language like "test suite", "mock", "test case" without prescribing Rust-specific tools. - -✅ **PASS**: Clearly addresses user needs (developers, QA, new team members) with focus on refactoring confidence, regression protection, and understanding business logic. Each user story has clear business value stated. - -✅ **PASS**: Written for non-technical stakeholders - describes testing needs in terms of business functions, coverage goals, and quality metrics rather than technical implementation. - -✅ **PASS**: All mandatory sections present: User Scenarios & Testing (6 prioritized stories), Requirements (25 functional requirements), Key Entities (5 entities), Success Criteria (15 measurable outcomes). - -### Requirement Completeness Review -✅ **PASS**: Zero [NEEDS CLARIFICATION] markers - all requirements are concrete and specific. - -✅ **PASS**: All 25 functional requirements are testable with clear, unambiguous criteria: -- FR-001: "minimum 95% code coverage" - measurable via coverage tools -- FR-002: "all three comparison operators" - verifiable by test case count -- FR-018: "execute in under 2 minutes" - measurable time threshold -- Each requirement uses MUST language and defines specific capabilities to verify - -✅ **PASS**: Success criteria are measurable with specific metrics: -- SC-001: "minimum 95% code coverage" (quantitative) -- SC-002: "minimum 50 test cases" (quantitative) -- SC-004: "under 2 minutes" (time-based) -- SC-005: "100% of intentional breaking changes" (percentage) -- SC-006: "Zero false positives" (count-based) - -✅ **PASS**: Success criteria are technology-agnostic - no mention of specific testing tools, frameworks, or Rust-specific constructs. Uses general terms like "test suite", "coverage report", "CI/CD pipeline". - -✅ **PASS**: All 6 user stories have detailed acceptance scenarios with Given-When-Then format. Total of 27 acceptance scenarios covering happy paths, error cases, and edge cases. - -✅ **PASS**: Edge Cases section contains 10 specific boundary conditions and error scenarios (null values, missing config, network failures, malformed data, etc.). - -✅ **PASS**: Scope is clearly bounded: -- Covers specific business functions: get_metric_flag_state, get_service_health, AppState::process_config, handler_render, get_graphite_data -- Defines specific API endpoints to test -- Specifies 95% coverage threshold for core functions (not entire codebase) -- Clear priorities (P1, P2, P3) indicating what's critical vs nice-to-have - -✅ **PASS**: Dependencies and assumptions identified implicitly: -- Assumes existing codebase has identified business functions (listed in requirements) -- Assumes mockito and tokio-test are available (mentioned in context, not prescribed) -- Assumes CI/CD pipeline exists or will be configured -- Assumes standard Rust test tooling is acceptable - -### Feature Readiness Review -✅ **PASS**: Each of 25 functional requirements maps to user stories and acceptance scenarios. For example: -- FR-001 (95% coverage) → User Story 6 (regression suite) → SC-001 -- FR-002 (comparison operators) → User Story 1 (metric flag testing) → 5 acceptance scenarios -- FR-006 (API endpoints) → User Story 3 (API testing) → 5 acceptance scenarios - -✅ **PASS**: User scenarios cover all primary flows: -- Core metric evaluation (P1) -- Health aggregation (P1) -- API testing (P2) -- Configuration processing (P2) -- Graphite integration (P3) -- Overall regression suite (P1) -Coverage is comprehensive across all layers: business logic, API, configuration, external integration. - -✅ **PASS**: Feature explicitly defines measurable outcomes across 5 categories (15 total success criteria) aligned with user needs: -- Coverage Metrics (SC-001 to SC-003) -- Quality Metrics (SC-004 to SC-006) -- Refactoring Confidence (SC-007 to SC-009) -- Documentation Value (SC-010 to SC-012) -- CI/CD Integration (SC-013 to SC-015) - -✅ **PASS**: No implementation details found: -- No mention of specific Rust testing frameworks (though project uses them) -- No code structure or module organization specified -- No test file naming conventions prescribed -- No specific assertion libraries mentioned -- One minor note: FR-012 typo "Test MUST" instead of "Tests MUST" (fixed in validation) - -## Overall Assessment - -**STATUS**: ✅ **READY FOR PLANNING** - -The specification is complete, clear, and ready for the next phase. All checklist items pass validation. - -### Strengths: -1. Comprehensive coverage of all business functions identified in codebase analysis -2. Well-prioritized user stories with clear independent value -3. Highly measurable success criteria with specific quantitative metrics -4. Technology-agnostic language throughout -5. Strong focus on user value (refactoring confidence, onboarding, regression protection) -6. Detailed acceptance scenarios for every user story - -### Minor Issue Fixed: -- FR-012: Corrected typo "Test MUST" → "Tests MUST" for consistency - -### Recommendations for Planning Phase: -1. Consider test organization strategy (by module, by business function, or by user story) -2. Define test data fixture management approach -3. Determine coverage reporting tool and thresholds -4. Plan for CI/CD pipeline integration testing diff --git a/specs/002-functional-test-suite/plan.md b/specs/002-functional-test-suite/plan.md deleted file mode 100644 index 6fb481e..0000000 --- a/specs/002-functional-test-suite/plan.md +++ /dev/null @@ -1,492 +0,0 @@ -# Implementation Plan: Comprehensive Functional Test Suite - -**Feature Branch**: `002-functional-test-suite` -**Created**: 2025-01-24 -**Status**: Ready for Implementation - ---- - -## 1. Overview - -This plan implements a comprehensive functional test suite achieving 95%+ coverage of core business functions. The approach prioritizes **bottom-up testing**: starting with pure unit tests for core logic, then building up to integration tests with mocked HTTP dependencies, and finally full API endpoint tests. - -**Key Implementation Strategy:** -- Use existing `mockito` dependency for HTTP mocking (already in dev-dependencies) -- Organize tests by module with shared fixtures per test category -- Leverage Rust's built-in parallel test execution with proper isolation -- Add `cargo-tarpaulin` for coverage measurement in CI - -**Current Test Baseline:** -- `config.rs`: 3 tests (config parsing, env merge, conf.d merge) -- `types.rs`: 1 test (AppState processing) -- `graphite.rs`: 3 tests (query building, HTTP mocking, find metrics) -- `common.rs`: 0 tests ❌ (core business logic - **highest priority**) -- `api/v1.rs`: 0 tests ❌ (API handlers) - ---- - -## 2. Design Decisions - -### 2.1 Test Framework Architecture - -**Decision**: Use Rust's built-in test framework with inline module tests + integration tests in `tests/` - -**Rationale**: -- Inline `#[cfg(test)]` modules keep tests close to implementation -- Integration tests in `tests/` directory for cross-module scenarios -- No additional test framework dependencies needed -- Follows existing codebase patterns (see `config.rs`, `graphite.rs`) - -**Structure**: -``` -src/ -├── common.rs # + #[cfg(test)] mod test { ... } -├── types.rs # + expand existing test module -├── graphite.rs # + expand existing test module -├── config.rs # existing tests - add validation tests -├── api/ -│ └── v1.rs # + #[cfg(test)] mod test { ... } -tests/ -├── fixtures/ # Shared test fixtures -│ ├── mod.rs -│ ├── configs.rs # YAML config fixtures -│ ├── graphite_responses.rs # Mock Graphite data -│ └── helpers.rs # Common test utilities -├── integration_health.rs # Health aggregation E2E -├── integration_api.rs # Full API endpoint tests -└── documentation_validation.rs # (existing) -``` - -### 2.2 Mock Server Setup - -**Decision**: Use `mockito` (already in Cargo.toml dev-dependencies) - -**Rationale**: -- Already integrated and proven working (see `graphite.rs:test_get_graphite_data`) -- Provides request matching, response mocking, and expectation verification -- Lightweight and thread-safe with `mockito::Server::new()` per test -- No need to add `wiremock-rs` or `httptest` - avoid unnecessary dependencies - -**Mock Patterns**: -```rust -// Per-test server isolation (already established pattern) -let mut server = mockito::Server::new(); -let mock = server - .mock("GET", "/render") - .match_query(Matcher::AllOf(vec![...])) - .with_body(json!([...]).to_string()) - .create(); -``` - -### 2.3 Fixture Organization Strategy - -**Decision**: Shared fixtures per module in `tests/fixtures/` - -**Rationale**: -- Centralized test data reduces duplication (spec FR-021) -- Easy to maintain consistent test scenarios -- Can be reused across unit and integration tests -- Follows DRY principle while keeping fixtures discoverable - -**Fixture Categories**: - -| Module | Fixture Type | Contents | -|--------|--------------|----------| -| `configs.rs` | YAML strings | Valid configs, minimal configs, invalid configs, edge cases | -| `graphite_responses.rs` | JSON data | Valid datapoints, empty arrays, null values, errors | -| `helpers.rs` | Test utilities | `create_test_state()`, `mock_graphite_response()`, custom assertions | - -### 2.4 Coverage Measurement Approach - -**Decision**: Use `cargo-tarpaulin` with CI integration - -**Rationale**: -- Rust-native, accurate line coverage (spec clarification) -- Supports HTML and lcov output formats -- Can enforce minimum threshold in CI -- Widely adopted in Rust ecosystem - -**CI Configuration** (for `zuul.yaml` or GitHub Actions): -```yaml -coverage: - script: - - cargo install cargo-tarpaulin - - cargo tarpaulin --out Html --out Lcov --fail-under 95 - artifacts: - - tarpaulin-report.html -``` - ---- - -## 3. Architecture - -### 3.1 Test Categories - -``` -┌─────────────────────────────────────────────────────────────────┐ -│ Test Pyramid │ -├─────────────────────────────────────────────────────────────────┤ -│ ┌───────────────────────────────────────────────────────────┐ │ -│ │ Integration Tests (tests/*.rs) │ │ -│ │ • Full API endpoint tests with mock Graphite │ │ -│ │ • Cross-module health aggregation flows │ │ -│ └───────────────────────────────────────────────────────────┘ │ -│ ┌───────────────────────────────────────────────────────────┐ │ -│ │ Unit Tests (src/**/mod test) │ │ -│ │ • get_metric_flag_state - all operators & edge cases │ │ -│ │ • get_service_health - expression evaluation │ │ -│ │ • AppState::process_config - template expansion │ │ -│ │ • Config validation & error cases │ │ -│ │ • Graphite query building & response parsing │ │ -│ └───────────────────────────────────────────────────────────┘ │ -└─────────────────────────────────────────────────────────────────┘ -``` - -### 3.2 Test Isolation Strategy - -For parallel execution safety (spec FR-018, SC-006): - -| Isolation Need | Solution | -|----------------|----------| -| Mock server ports | `mockito::Server::new()` auto-assigns unique ports | -| Shared state | Each test creates own `AppState` instance | -| Environment variables | Use `temp_env` crate or test-specific prefixes | -| File system | Use `tempfile` crate (already in dev-deps) | - -### 3.3 Custom Assertion Helpers - -For business context in failures (spec FR-019, SC-009): - -```rust -// tests/fixtures/helpers.rs -pub fn assert_metric_flag( - value: Option, - metric: &FlagMetric, - expected: bool, - context: &str, -) { - let actual = get_metric_flag_state(&value, metric); - assert_eq!( - actual, expected, - "Metric flag evaluation failed for {}: value={:?}, op={:?}, threshold={}, expected={}, got={}", - context, value, metric.op, metric.threshold, expected, actual - ); -} - -pub fn assert_health_score( - service: &str, - environment: &str, - expected_score: u8, - actual_score: u8, -) { - assert_eq!( - actual_score, expected_score, - "Health score mismatch for service '{}' in '{}': expected {}, got {}", - service, environment, expected_score, actual_score - ); -} -``` - ---- - -## 4. Implementation Phases - -### Phase 1: Test Infrastructure Setup -- [ ] **1.1** Create `tests/fixtures/mod.rs` with module structure -- [ ] **1.2** Create `tests/fixtures/configs.rs` with standard test configurations -- [ ] **1.3** Create `tests/fixtures/graphite_responses.rs` with mock response data -- [ ] **1.4** Create `tests/fixtures/helpers.rs` with custom assertions and utilities -- [ ] **1.5** Add `cargo-tarpaulin` configuration to CI pipeline - -### Phase 2: Core Function Unit Tests (P1 - Highest Priority) -- [ ] **2.1** `get_metric_flag_state` tests in `src/common.rs`: - - [ ] 2.1.1 Lt operator: value < threshold returns true - - [ ] 2.1.2 Lt operator: value >= threshold returns false - - [ ] 2.1.3 Gt operator: value > threshold returns true - - [ ] 2.1.4 Gt operator: value <= threshold returns false - - [ ] 2.1.5 Eq operator: value == threshold returns true - - [ ] 2.1.6 Eq operator: value != threshold returns false - - [ ] 2.1.7 None value always returns false - - [ ] 2.1.8 Boundary conditions (threshold ± 0.001) - - [ ] 2.1.9 Negative values - - [ ] 2.1.10 Zero threshold -- [ ] **2.2** `AppState::process_config` tests in `src/types.rs`: - - [ ] 2.2.1 Template variable substitution ($environment, $service) - - [ ] 2.2.2 Multiple environments expansion - - [ ] 2.2.3 Per-environment threshold override - - [ ] 2.2.4 Dash-to-underscore conversion in expressions - - [ ] 2.2.5 Service set population - - [ ] 2.2.6 Health metrics expression copying - -### Phase 3: Integration Tests with Mocked Graphite (P1) -- [ ] **3.1** `get_service_health` tests in `src/common.rs`: - - [ ] 3.1.1 Single metric OR expression evaluates correctly - - [ ] 3.1.2 Multiple metrics AND expression evaluates correctly - - [ ] 3.1.3 Weighted expressions return highest matching weight - - [ ] 3.1.4 All false expressions return weight 0 - - [ ] 3.1.5 Unknown service returns ServiceNotSupported error - - [ ] 3.1.6 Unknown environment returns EnvNotSupported error - - [ ] 3.1.7 Multiple datapoints across time series -- [ ] **3.2** Create `tests/integration_health.rs` for end-to-end health flows: - - [ ] 3.2.1 Full health calculation with mocked Graphite - - [ ] 3.2.2 Complex weighted expression scenarios - - [ ] 3.2.3 Edge cases: empty datapoints, partial data - -### Phase 4: API Endpoint Tests (P2) -- [ ] **4.1** Add tests to `src/api/v1.rs`: - - [ ] 4.1.1 `/api/v1/` root endpoint returns name - - [ ] 4.1.2 `/api/v1/info` returns API info - - [ ] 4.1.3 `/api/v1/health` with valid service returns 200 + JSON - - [ ] 4.1.4 `/api/v1/health` with unknown service returns 409 - - [ ] 4.1.5 `/api/v1/health` with missing params returns 400 -- [ ] **4.2** Expand Graphite route tests in `src/graphite.rs`: - - [ ] 4.2.1 `/render` with flag target returns boolean datapoints - - [ ] 4.2.2 `/render` with health target returns health scores - - [ ] 4.2.3 `/render` with invalid target returns empty array - - [ ] 4.2.4 `/metrics/find` all levels (*, flag.*, flag.env.*, etc.) - - [ ] 4.2.5 `/functions` returns empty object - - [ ] 4.2.6 `/tags/autoComplete/tags` returns empty array -- [ ] **4.3** Create `tests/integration_api.rs` for full API tests: - - [ ] 4.3.1 Health endpoint with mocked Graphite responses - - [ ] 4.3.2 Render endpoint with various targets - - [ ] 4.3.3 Error response format validation - -### Phase 5: Configuration & Error Path Tests (P2-P3) -- [ ] **5.1** Expand config tests in `src/config.rs`: - - [ ] 5.1.1 Invalid YAML syntax returns parse error - - [ ] 5.1.2 Missing required fields error - - [ ] 5.1.3 Default values applied correctly - - [ ] 5.1.4 `get_socket_addr()` produces valid address -- [ ] **5.2** Graphite client error handling in `src/graphite.rs`: - - [ ] 5.2.1 HTTP 4xx returns GraphiteError - - [ ] 5.2.2 HTTP 5xx returns GraphiteError - - [ ] 5.2.3 Malformed JSON response handling - - [ ] 5.2.4 Connection timeout handling - - [ ] 5.2.5 Empty response handling -- [ ] **5.3** Expression evaluation error tests: - - [ ] 5.3.1 Invalid expression syntax returns ExpressionError - - [ ] 5.3.2 Missing metric in context handled - -### Phase 6: Coverage Validation & CI Integration -- [ ] **6.1** Run `cargo tarpaulin` and verify 95% coverage target -- [ ] **6.2** Identify and fill any coverage gaps -- [ ] **6.3** Add coverage enforcement to CI (`--fail-under 95`) -- [ ] **6.4** Generate HTML coverage report for documentation -- [ ] **6.5** Verify all tests pass in under 2 minutes -- [ ] **6.6** Run intentional breakage tests to verify regression detection - ---- - -## 5. Dependencies - -### Task Dependency Graph - -``` -Phase 1 (Infrastructure) - │ - ├──► Phase 2 (Unit Tests) ──┐ - │ │ - └──► Phase 3 (Integration)──┼──► Phase 6 (Coverage) - │ │ - └──► Phase 4 (API)┘ - │ - └──► Phase 5 (Errors) -``` - -### Critical Path - -1. **Phase 1.1-1.4** (fixtures) → Blocks all test implementation -2. **Phase 2.1** (get_metric_flag_state) → Core logic, highest ROI -3. **Phase 3.1** (get_service_health) → Most complex business function -4. **Phase 6.1** (coverage check) → May reveal additional gaps - -### External Dependencies - -| Dependency | Version | Purpose | Status | -|------------|---------|---------|--------| -| `mockito` | ~1.0 | HTTP mocking | ✅ Already in Cargo.toml | -| `tempfile` | ~3.5 | Temp file/dir creation | ✅ Already in Cargo.toml | -| `tokio-test` | * | Async test utilities | ✅ Already in Cargo.toml | -| `hyper` | 0.14 | HTTP test client | ✅ Already in Cargo.toml | -| `cargo-tarpaulin` | latest | Coverage tool | ⚠️ Install required | - ---- - -## 6. Risk Mitigation - -### Risk 1: Test Flakiness from Parallel Execution - -**Risk**: Tests sharing mock servers or state cause intermittent failures. - -**Mitigation**: -- Each test creates its own `mockito::Server::new()` (auto-assigns port) -- Each test creates its own `AppState` instance -- Use `tokio::test` for async isolation -- Avoid global mutable state - -**Detection**: Run `cargo test -- --test-threads=1` vs default; results should match. - -### Risk 2: Coverage Target Not Achievable - -**Risk**: 95% coverage is unrealistic for error paths or edge cases. - -**Mitigation**: -- Focus coverage on core functions listed in spec (FR-001) -- Accept lower coverage for generated code, error Display impls -- Use `#[cfg(not(tarpaulin_include))]` for intentionally uncovered code -- Document any exclusions - -**Fallback**: Negotiate with stakeholders if <95% is justified. - -### Risk 3: Mock Graphite Behavior Diverges from Real - -**Risk**: Mocked responses don't match real Graphite behavior. - -**Mitigation**: -- Use real Graphite API documentation for response formats -- Record real responses as fixtures where possible -- Include malformed/error responses based on real error scenarios -- Integration test with real Graphite in CI (optional, out of scope) - -### Risk 4: Test Suite Exceeds 2-Minute Target - -**Risk**: Full test suite takes too long, reducing developer adoption. - -**Mitigation**: -- Profile test execution: `cargo test -- --nocapture 2>&1 | ts` -- Identify slow tests (usually mock server setup) -- Use `#[ignore]` for optional slow tests -- Consider test parallelization tuning - -**Target Breakdown**: -- Unit tests: <30 seconds (no I/O) -- Integration tests: <60 seconds (mock HTTP) -- API tests: <30 seconds (in-process server) - -### Risk 5: Test Maintenance Burden - -**Risk**: Tests become brittle and hard to maintain over time. - -**Mitigation**: -- Shared fixtures reduce duplication -- Custom assertion helpers provide clear failure messages -- Tests focus on behavior, not implementation details -- Documentation in test names and comments - ---- - -## 7. Test Count Targets (SC-002) - -| Category | Target | Priority | -|----------|--------|----------| -| Metric flag evaluation | 20+ tests | P1 | -| Health aggregation | 15+ tests | P1 | -| API endpoints | 10+ tests | P2 | -| Configuration | 5+ tests | P2 | -| Error handling | 10+ tests | P3 | -| **Total** | **60+ tests** | | - ---- - -## 8. Success Metrics - -| Metric | Target | Measurement | -|--------|--------|-------------| -| Code coverage | ≥95% | `cargo tarpaulin --fail-under 95` | -| Test count | ≥50 | `cargo test -- --list \| wc -l` | -| Execution time | <2 min | `time cargo test` | -| Parallel safety | 0 flaky | Run 10x with `--test-threads=8` | -| Regression detection | 100% | Intentional break tests | - ---- - -## Appendix A: Sample Test Fixtures - -### A.1 Minimal Valid Config - -```rust -// tests/fixtures/configs.rs -pub const MINIMAL_CONFIG: &str = r#" -datasource: - url: 'http://localhost:8080' -server: - port: 3000 -environments: - - name: prod -flag_metrics: [] -health_metrics: {} -"#; -``` - -### A.2 Mock Graphite Response - -```rust -// tests/fixtures/graphite_responses.rs -pub fn valid_datapoints(target: &str) -> String { - serde_json::json!([{ - "target": target, - "datapoints": [ - [85.0, 1700000000], - [90.0, 1700000060], - [95.0, 1700000120] - ] - }]).to_string() -} - -pub fn empty_datapoints(target: &str) -> String { - serde_json::json!([{ - "target": target, - "datapoints": [] - }]).to_string() -} -``` - -### A.3 Test Helper Functions - -```rust -// tests/fixtures/helpers.rs -use cloudmon_metrics::{config::Config, types::AppState}; - -pub fn create_test_state(config_yaml: &str) -> AppState { - let config = Config::from_config_str(config_yaml); - let mut state = AppState::new(config); - state.process_config(); - state -} - -pub fn create_test_state_with_mock_url(config_yaml: &str, mock_url: &str) -> AppState { - // Replace datasource URL with mock server URL - let modified = config_yaml.replace("http://localhost:8080", mock_url); - create_test_state(&modified) -} -``` - ---- - -## Appendix B: Commands Reference - -```bash -# Run all tests -cargo test - -# Run with verbose output -cargo test -- --nocapture - -# Run specific test module -cargo test common::test - -# Run tests matching pattern -cargo test metric_flag - -# Check coverage (install first: cargo install cargo-tarpaulin) -cargo tarpaulin --out Html - -# Coverage with threshold enforcement -cargo tarpaulin --fail-under 95 - -# Run tests sequentially (debugging flakiness) -cargo test -- --test-threads=1 - -# List all tests -cargo test -- --list -``` diff --git a/specs/002-functional-test-suite/spec.md b/specs/002-functional-test-suite/spec.md deleted file mode 100644 index 87bf0f3..0000000 --- a/specs/002-functional-test-suite/spec.md +++ /dev/null @@ -1,245 +0,0 @@ -# Feature Specification: Comprehensive Functional Test Suite - -**Feature Branch**: `002-functional-test-suite` -**Created**: 2025-01-24 -**Status**: Draft -**Input**: User description: "Create a feature specification for comprehensive functional tests for the metrics-processor project. - -User requirements: -- As a new developer or QA, I need to be sure that main business functionality works as expected -- I need functional tests for the whole project -- The user plans to refactor the code base and add new features -- They need confidence that the main functionality won't change during refactoring -- Minimum 95% test coverage for the main business functions is required - -The spec should cover: -1. Identifying all main business functions in the codebase -2. Creating functional/integration tests that verify business logic -3. Ensuring 95%+ coverage of core business functionality -4. Tests should serve as regression protection during refactoring" - -## User Scenarios & Testing *(mandatory)* - -### User Story 1 - Core Metric Flag Evaluation Testing (Priority: P1) - -As a developer refactoring the metrics processing logic, I need comprehensive tests that verify metric flag evaluation (comparison operators Lt/Gt/Eq) works correctly across all threshold scenarios, so I can confidently refactor without breaking core business logic. - -**Why this priority**: This is the foundation of the entire system - converting raw numeric metrics to boolean flags. If this breaks, the entire health monitoring system fails. This function (`get_metric_flag_state`) is called for every metric evaluation and is critical for accurate monitoring. - -**Independent Test**: Can be fully tested by providing numeric values and metric configurations with different comparison operators, then verifying the boolean flag output matches expected results. Delivers immediate value by preventing false positives/negatives in health monitoring. - -**Acceptance Scenarios**: - -1. **Given** a metric value of 85 and a threshold of 90 with Lt operator, **When** evaluating flag state, **Then** system returns true (85 < 90) -2. **Given** a metric value of 95 and a threshold of 90 with Gt operator, **When** evaluating flag state, **Then** system returns true (95 > 90) -3. **Given** a metric value of 90 and a threshold of 90 with Eq operator, **When** evaluating flag state, **Then** system returns true (90 == 90) -4. **Given** a metric value of 85 and a threshold of 90 with Gt operator, **When** evaluating flag state, **Then** system returns false (85 not > 90) -5. **Given** multiple metrics with mixed operators in a service, **When** evaluating all flags, **Then** each flag correctly reflects its comparison result - ---- - -### User Story 2 - Service Health Aggregation Testing (Priority: P1) - -As a QA engineer, I need tests that verify service health calculation correctly fetches metrics, evaluates boolean expressions, applies weights, and returns the highest-weighted health status, so that monitoring dashboards show accurate service health. - -**Why this priority**: This is the most complex business function (`get_service_health`) that orchestrates the entire health evaluation workflow. It combines multiple metrics using boolean expressions (AND/OR) and weighted scoring. Incorrect health status can lead to missed incidents or false alarms. - -**Independent Test**: Can be fully tested by mocking Graphite responses with known metric values, defining weighted expressions, and verifying the returned health status matches expected priority. Delivers value by ensuring monitoring accuracy. - -**Acceptance Scenarios**: - -1. **Given** a service with 2 metrics (both true) and expression "metric1 OR metric2" with weight 3, **When** calculating health, **Then** system returns impact value 3 -2. **Given** a service with 3 weighted expressions (weights: 5, 3, 1) where only weight-3 expression evaluates true, **When** calculating health, **Then** system returns impact value 3 (highest true expression) -3. **Given** a service where all expressions evaluate false, **When** calculating health, **Then** system returns impact value 0 -4. **Given** a service with AND expression "metric1 AND metric2" where metric1 is true but metric2 is false, **When** calculating health, **Then** expression evaluates false -5. **Given** a service in an unknown environment, **When** calculating health, **Then** system returns appropriate error without crashing - ---- - -### User Story 3 - API Endpoint Integration Testing (Priority: P2) - -As a developer adding new API features, I need integration tests for all REST endpoints (/api/v1/health, /render, /metrics/find) that verify request handling, response formats, error handling, and Graphite integration, so I can ensure the API contract remains stable during refactoring. - -**Why this priority**: API endpoints are the external interface used by dashboards and other services. Breaking changes affect all consumers. Currently these endpoints have no automated tests, making refactoring risky. - -**Independent Test**: Can be fully tested by starting a test server, sending HTTP requests with various parameters, and validating response structure and status codes. Delivers value by protecting the API contract. - -**Acceptance Scenarios**: - -1. **Given** a running API server, **When** GET /api/v1/health?service=myservice&environment=production, **Then** response contains valid ServiceHealthResponse JSON with status 200 -2. **Given** a running API server, **When** GET /render?target=flag.prod.myservice.metric1, **Then** response contains time-series data with boolean values -3. **Given** a running API server, **When** GET /metrics/find?query=flag.*, **Then** response contains list of matching metrics with expandable flag -4. **Given** a request for non-existent service, **When** querying health endpoint, **Then** response returns 404 or appropriate error with message -5. **Given** invalid query parameters, **When** calling any endpoint, **Then** response returns 400 with clear error description - ---- - -### User Story 4 - Configuration Processing Testing (Priority: P2) - -As a new team member, I need tests that verify configuration loading, template variable substitution ($environment, $service), and metric initialization work correctly across all configuration scenarios, so I understand how configuration changes affect system behavior. - -**Why this priority**: Configuration processing (`AppState::process_config`) is the initialization step that sets up all metrics and expressions. Errors here prevent the system from starting or cause incorrect metric mappings. This has one test but needs comprehensive coverage. - -**Independent Test**: Can be fully tested by providing various YAML configurations with templates and variables, then verifying the resulting AppState contains correctly expanded metric definitions and expression mappings. Delivers documentation value through test examples. - -**Acceptance Scenarios**: - -1. **Given** a config with template "flag.$environment.$service.cpu" and environments [prod, dev], **When** processing config, **Then** system creates metric mappings for flag.prod.*.cpu and flag.dev.*.cpu -2. **Given** a config with health expression containing dashes "api-gateway", **When** processing config, **Then** system converts to "api_gateway" for expression evaluation -3. **Given** a config file, conf.d directory with overrides, and environment variables with MP_ prefix, **When** loading config, **Then** system merges all sources with correct precedence -4. **Given** a config with invalid YAML syntax, **When** loading config, **Then** system returns clear error message with line number -5. **Given** a config with missing required fields, **When** validating config, **Then** system returns error listing all missing fields - ---- - -### User Story 5 - Graphite Integration Testing (Priority: P3) - -As a developer working on TSDB integration, I need tests that verify Graphite client query building, response parsing, and error handling for network failures or malformed data, so I can safely refactor the Graphite client without breaking monitoring. - -**Why this priority**: Graphite integration (`get_graphite_data`, `find_metrics`) is essential for data retrieval, but failures here are easier to debug and less critical than core business logic. The client already has some test coverage but needs comprehensive scenarios. - -**Independent Test**: Can be fully tested by mocking Graphite HTTP responses with various data formats and error conditions, then verifying correct parsing or error handling. Delivers value by ensuring reliable external integration. - -**Acceptance Scenarios**: - -1. **Given** a mock Graphite server returning valid JSON with datapoints, **When** querying metrics, **Then** client correctly parses values and timestamps -2. **Given** a mock Graphite server returning empty datapoints array, **When** querying metrics, **Then** client handles gracefully without errors -3. **Given** Graphite server returns HTTP 500 error, **When** querying metrics, **Then** client returns appropriate error with context -4. **Given** Graphite server times out, **When** querying metrics, **Then** client returns timeout error after configured duration -5. **Given** metric discovery query for "flag.prod.*", **When** calling find_metrics, **Then** client returns list of expandable nodes at that level - ---- - -### User Story 6 - Regression Test Suite for Refactoring (Priority: P1) - -As a developer refactoring the codebase, I need a comprehensive regression test suite that runs quickly (under 2 minutes) and fails immediately when business logic changes, so I can refactor code structure confidently without changing behavior. - -**Why this priority**: This is the primary goal - enabling safe refactoring. The test suite must cover 95%+ of business logic and serve as a safety net. Without this, refactoring is risky and slow. - -**Independent Test**: Can be fully tested by running the complete test suite after making intentional breaking changes to business logic, and verifying tests catch the breakage. Delivers immediate refactoring confidence. - -**Acceptance Scenarios**: - -1. **Given** complete test suite covering all business functions, **When** running tests, **Then** all tests pass in under 2 minutes -2. **Given** a deliberate change to metric comparison logic (swap Lt/Gt), **When** running tests, **Then** core metric tests fail with clear error messages -3. **Given** a deliberate change to health calculation weights, **When** running tests, **Then** health aggregation tests fail -4. **Given** refactored code with same behavior but different structure, **When** running tests, **Then** all tests still pass -5. **Given** test suite in CI/CD pipeline, **When** pull request is created, **Then** tests run automatically and block merge if failing - ---- - -## Clarifications - -### Session 2025-01-24 - -- Q: For Graphite integration testing, which mocking approach should the test suite use? → A: HTTP mock server (e.g., wiremock/mockito) - Realistic integration, tests full HTTP stack -- Q: How should test data fixtures (sample configs, metric values, expected outputs) be organized and managed? → A: Shared fixtures per module - Reusable, DRY, good for consistency -- Q: Should tests run in parallel or sequentially to meet the 2-minute execution goal? → A: Parallel by default (cargo test) - Fast, requires careful isolation -- Q: What approach should be used for test assertions and failure diagnostics to ensure clear error messages? → A: Balanced approach - standard assertions for unit tests, custom assertions with business context for functional/integration tests -- Q: Which code coverage tool should be used to measure and enforce the 95% coverage requirement? → A: cargo-tarpaulin - Rust-native, accurate line coverage - -### Edge Cases - -- What happens when Graphite returns null or NaN values for metrics? (Should handle gracefully, not crash) -- How does system handle services with zero health expressions configured? (Should return default/error state) -- What happens when boolean expressions contain invalid metric names? (Should return error with metric name) -- How does system handle extremely large time ranges (months of data)? (Should limit or paginate) -- What happens when configuration contains circular variable references? (Should detect and error) -- How does system behave when Graphite is completely unreachable? (Should timeout and return error) -- What happens when multiple environments have overlapping metric names? (Should namespace correctly) -- How does expression evaluation handle division by zero or math errors? (Should catch and return error) -- What happens when services list contains special characters or spaces? (Should sanitize or validate) -- How does system handle partial Graphite responses (some metrics succeed, others fail)? (Should process available data) - -## Requirements *(mandatory)* - -### Functional Requirements - -#### Test Coverage Requirements - -- **FR-001**: Test suite MUST achieve minimum 95% code coverage for all core business logic functions (get_metric_flag_state, get_service_health, AppState::process_config, handler_render) -- **FR-002**: Test suite MUST cover all three comparison operators (Lt, Gt, Eq) with boundary conditions and edge cases for metric flag evaluation -- **FR-003**: Test suite MUST verify boolean expression evaluation (AND, OR operators) with all combinations of true/false metric states -- **FR-004**: Test suite MUST validate weighted health scoring with multiple expressions at different priority levels -- **FR-005**: Test suite MUST verify configuration template variable substitution ($environment, $service) for all configured environments and services - -#### API Testing Requirements - -- **FR-006**: Test suite MUST include integration tests for all REST API endpoints: /api/v1/health, /api/v1/info, /render, /metrics/find, /functions, /tags/autoComplete/tags -- **FR-007**: API tests MUST verify correct HTTP status codes (200, 400, 404, 500) for valid and invalid requests -- **FR-008**: API tests MUST validate response JSON structure matches expected schema for each endpoint -- **FR-009**: API tests MUST verify error messages are clear and actionable when requests fail - -#### Graphite Integration Testing Requirements - -- **FR-010**: Test suite MUST mock Graphite HTTP responses using an HTTP mock server (e.g., wiremock-rs or httptest) for all query scenarios (valid data, empty data, errors, timeouts) -- **FR-011**: Tests MUST verify correct parsing of Graphite JSON response format with datapoints arrays -- **FR-012**: Tests MUST verify metric discovery (find_metrics) correctly handles hierarchical metric paths with wildcards -- **FR-013**: Tests MUST verify query building produces valid Graphite query syntax with correct time ranges and parameters - -#### Configuration Testing Requirements - -- **FR-014**: Test suite MUST verify configuration loading from multiple sources (file, conf.d directory, environment variables) with correct precedence -- **FR-015**: Tests MUST verify YAML parsing handles valid and invalid syntax with appropriate error messages -- **FR-016**: Tests MUST verify configuration validation catches missing required fields and returns all errors at once -- **FR-017**: Tests MUST verify metric template expansion creates correct mappings for all environment and service combinations - -#### Regression Protection Requirements - -- **FR-018**: Test suite MUST execute in under 2 minutes to support rapid development workflow. Tests MUST run in parallel using Rust's default test runner (cargo test) with proper isolation (unique mock server ports, isolated test data) to achieve performance goals. -- **FR-019**: Tests MUST fail immediately with clear error messages when business logic behavior changes. Functional tests, API tests, and complex integration tests MUST use custom assertion helpers that include business context (e.g., service name, metric states, expected behavior). Unit tests MAY use standard Rust assert macros for simplicity. -- **FR-020**: Test suite MUST be runnable in CI/CD pipeline with standard Rust test tools (cargo test) -- **FR-021**: Tests MUST be maintainable with clear naming, documentation, and modular structure. Test fixtures MUST be organized per module with shared fixtures for related tests to ensure consistency and reduce duplication. - -#### Error Handling Testing Requirements - -- **FR-022**: Test suite MUST verify all error paths return appropriate error types without panicking -- **FR-023**: Tests MUST verify system handles null, NaN, and missing metric values gracefully -- **FR-024**: Tests MUST verify system handles network failures (timeouts, connection refused) with retries or clear errors -- **FR-025**: Tests MUST verify system handles malformed JSON responses from Graphite without crashing - -### Key Entities - -- **Test Case**: Represents a single automated test with setup, execution, and assertion phases. Contains test name, description, mock data fixtures, expected outcomes, and cleanup logic. - -- **Mock Graphite Server**: HTTP mock server (e.g., wiremock-rs or httptest) that simulates Graphite TSDB responses for integration testing. Runs an actual HTTP server on localhost during tests, provides configurable responses for different query patterns, supports both valid data and error scenarios. Tests the complete HTTP client stack including connection handling, timeouts, and error codes. - -- **Test Fixture**: Reusable test data including sample configurations, metric values, boolean expressions, and expected health scores. Organized by scenario (happy path, edge cases, errors). Each test module (metric_evaluation_tests, health_aggregation_tests, etc.) maintains its own fixtures submodule with common test data shared across related tests, ensuring consistency while keeping fixtures close to where they're used. - -- **Coverage Report**: Generated output showing code coverage percentage for each module and function. Generated using cargo-tarpaulin with support for HTML, lcov, and JSON formats. Used to verify 95% threshold is met and identify untested code paths. Integrated into CI/CD pipeline for automated coverage tracking. - -- **Test Configuration**: YAML configuration files specifically designed for testing, including minimal valid config, maximal config with all features, and invalid configs for error testing. - -## Success Criteria *(mandatory)* - -### Measurable Outcomes - -#### Coverage Metrics - -- **SC-001**: Test suite achieves minimum 95% code coverage for core business functions (get_metric_flag_state, get_service_health, AppState::process_config, handler_render, get_graphite_data) as measured by cargo-tarpaulin -- **SC-002**: Test suite includes minimum 50 test cases covering all priority areas (20+ for metric evaluation, 15+ for health aggregation, 10+ for API endpoints, 5+ for configuration) -- **SC-003**: All 25 functional requirements (FR-001 through FR-025) have at least one passing test that validates the requirement - -#### Quality Metrics - -- **SC-004**: Test suite completes full execution in under 2 minutes on standard development hardware -- **SC-005**: Tests detect 100% of intentional breaking changes to business logic (deliberate changes to comparison operators, expression evaluation, weight calculations) -- **SC-006**: Zero false positives - tests only fail when actual business logic changes, not due to test flakiness or timing issues. Tests MUST be designed for parallel execution with proper isolation to prevent race conditions. - -#### Refactoring Confidence - -- **SC-007**: Developers can refactor code structure (rename functions, split modules, reorganize files) without any test failures as long as behavior is preserved -- **SC-008**: New developers can run test suite immediately after cloning repository with single command (cargo test) and see all tests pass -- **SC-009**: Test failures provide clear error messages identifying which business function broke and what the expected vs actual behavior was. Functional and integration test failures include business context (service names, metric states, scenarios) beyond simple value comparisons. - -#### Documentation Value - -- **SC-010**: Test cases serve as executable documentation - new team members can understand business logic by reading test scenarios -- **SC-011**: Each test case includes descriptive name and comments explaining the business scenario being tested -- **SC-012**: Test coverage report clearly identifies any untested code paths requiring additional tests. Coverage reports generated using cargo-tarpaulin in HTML and lcov formats for developer and CI integration. - -#### CI/CD Integration - -- **SC-013**: Test suite runs automatically on every pull request and blocks merge if any test fails -- **SC-014**: Test results are reported in CI/CD pipeline within 3 minutes of commit -- **SC-015**: Coverage reports are generated automatically using cargo-tarpaulin and show trend over time (no coverage decrease allowed) diff --git a/specs/002-functional-test-suite/tasks.md b/specs/002-functional-test-suite/tasks.md deleted file mode 100644 index 9f365f2..0000000 --- a/specs/002-functional-test-suite/tasks.md +++ /dev/null @@ -1,358 +0,0 @@ ---- -description: "Implementation tasks for comprehensive functional test suite" ---- - -# Tasks: Comprehensive Functional Test Suite - -**Input**: Design documents from `/specs/002-functional-test-suite/` -**Prerequisites**: plan.md, spec.md - -**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. Tests are explicitly requested in the feature specification to achieve 95% coverage and enable safe refactoring. - -## Format: `[ID] [P?] [Story] Description` - -- **[P]**: Can run in parallel (different files, no dependencies) -- **[Story]**: Which user story this task belongs to (e.g., US1, US2, US3) -- Include exact file paths in descriptions - -## Path Conventions - -- Tests use Rust's built-in test framework -- Unit tests: `#[cfg(test)]` modules in source files -- Integration tests: `tests/` directory at repository root -- Fixtures: `tests/fixtures/` for shared test data - ---- - -## Phase 1: Setup (Shared Infrastructure) - -**Purpose**: Test infrastructure and fixtures that all test phases depend on - -- [X] T001 Create fixtures module structure in tests/fixtures/mod.rs -- [X] T002 [P] Create test configuration fixtures in tests/fixtures/configs.rs -- [X] T003 [P] Create Graphite response mock data in tests/fixtures/graphite_responses.rs -- [X] T004 [P] Create custom assertion helpers in tests/fixtures/helpers.rs - ---- - -## Phase 2: Foundational (Blocking Prerequisites) - -**Purpose**: Core test utilities and CI configuration that MUST be complete before user story tests - -**⚠️ CRITICAL**: No user story testing can begin until this phase is complete - -- [X] T005 Add cargo-tarpaulin to CI pipeline configuration for coverage measurement -- [X] T006 [P] Implement create_test_state helper function in tests/fixtures/helpers.rs -- [X] T007 [P] Implement create_test_state_with_mock_url helper in tests/fixtures/helpers.rs -- [X] T008 [P] Implement assert_metric_flag custom assertion in tests/fixtures/helpers.rs -- [X] T009 [P] Implement assert_health_score custom assertion in tests/fixtures/helpers.rs - -**Checkpoint**: Foundation ready - user story testing can now begin in parallel - ---- - -## Phase 3: User Story 1 - Core Metric Flag Evaluation Testing (Priority: P1) 🎯 MVP - -**Goal**: Verify metric flag evaluation (comparison operators Lt/Gt/Eq) works correctly across all threshold scenarios to enable confident refactoring of core business logic. - -**Independent Test**: Provide numeric values and metric configurations with different comparison operators, verify boolean flag output matches expected results. - -### Tests for User Story 1 - -> **NOTE: Write these tests FIRST, ensure they FAIL before implementation (tests are testing existing code)** - -- [X] T010 [P] [US1] Test Lt operator with value < threshold returns true in src/common.rs test module -- [X] T011 [P] [US1] Test Lt operator with value >= threshold returns false in src/common.rs test module -- [X] T012 [P] [US1] Test Gt operator with value > threshold returns true in src/common.rs test module -- [X] T013 [P] [US1] Test Gt operator with value <= threshold returns false in src/common.rs test module -- [X] T014 [P] [US1] Test Eq operator with value == threshold returns true in src/common.rs test module -- [X] T015 [P] [US1] Test Eq operator with value != threshold returns false in src/common.rs test module -- [X] T016 [P] [US1] Test None value always returns false for all operators in src/common.rs test module -- [X] T017 [P] [US1] Test boundary conditions (threshold ± 0.001) in src/common.rs test module -- [X] T018 [P] [US1] Test negative values with all operators in src/common.rs test module -- [X] T019 [P] [US1] Test zero threshold edge case in src/common.rs test module -- [X] T020 [P] [US1] Test mixed operators scenario with multiple metrics in src/common.rs test module - -**Checkpoint**: At this point, User Story 1 should be fully tested - metric flag evaluation has comprehensive unit test coverage - ---- - -## Phase 4: User Story 6 - Regression Test Suite for Refactoring (Priority: P1) - -**Goal**: Comprehensive regression test suite that runs quickly (under 2 minutes) and fails immediately when business logic changes, enabling confident refactoring. - -**Independent Test**: Run complete test suite after intentional breaking changes to business logic, verify tests catch the breakage. - -**Note**: This phase depends on US1 tests being complete, as it validates the regression detection capability. - -### Tests for User Story 6 - -- [X] T021 [US6] Create intentional breakage test script to swap Lt/Gt operators in tests/ -- [X] T022 [US6] Verify all US1 tests fail with clear error messages after intentional breakage -- [X] T023 [US6] Run full test suite with parallel execution and measure timing in CI -- [X] T024 [US6] Verify zero false positives - tests only fail on actual logic changes -- [X] T025 [US6] Document test execution in README or docs/testing.md - -**Checkpoint**: Regression suite validated - safe refactoring enabled for metric flag evaluation - ---- - -## Phase 5: User Story 2 - Service Health Aggregation Testing (Priority: P1) - -**Goal**: Verify service health calculation correctly fetches metrics, evaluates boolean expressions, applies weights, and returns highest-weighted health status. - -**Independent Test**: Mock Graphite responses with known metric values, define weighted expressions, verify returned health status matches expected priority. - -### Tests for User Story 2 - -- [X] T026 [P] [US2] Test single metric OR expression evaluates correctly in src/common.rs test module -- [X] T027 [P] [US2] Test two metrics AND expression (both true) in src/common.rs test module -- [X] T028 [P] [US2] Test two metrics AND expression (one false) returns false in src/common.rs test module -- [X] T029 [P] [US2] Test weighted expressions return highest matching weight in src/common.rs test module -- [X] T030 [P] [US2] Test all false expressions return weight 0 in src/common.rs test module -- [X] T031 [P] [US2] Test unknown service returns ServiceNotSupported error in src/common.rs test module -- [X] T032 [P] [US2] Test unknown environment returns EnvNotSupported error in src/common.rs test module -- [X] T033 [P] [US2] Test multiple datapoints across time series in src/common.rs test module -- [X] T034 [US2] Create end-to-end health calculation test with mocked Graphite in tests/integration_health.rs -- [X] T035 [US2] Test complex weighted expression scenarios in tests/integration_health.rs -- [X] T036 [US2] Test edge cases: empty datapoints and partial data in tests/integration_health.rs - -**Checkpoint**: Service health aggregation fully tested - can refactor expression evaluation confidently - ---- - -## Phase 6: User Story 4 - Configuration Processing Testing (Priority: P2) - -**Goal**: Verify configuration loading, template variable substitution ($environment, $service), and metric initialization work correctly across all configuration scenarios. - -**Independent Test**: Provide various YAML configurations with templates and variables, verify resulting AppState contains correctly expanded metric definitions. - -### Tests for User Story 4 - -- [X] T037 [P] [US4] Test template variable substitution ($environment, $service) in src/types.rs test module -- [X] T038 [P] [US4] Test multiple environments expansion creates correct mappings in src/types.rs test module -- [X] T039 [P] [US4] Test per-environment threshold override in src/types.rs test module -- [X] T040 [P] [US4] Test dash-to-underscore conversion in expressions in src/types.rs test module -- [X] T041 [P] [US4] Test service set population from config in src/types.rs test module -- [X] T042 [P] [US4] Test health metrics expression copying in src/types.rs test module -- [X] T043 [P] [US4] Test invalid YAML syntax returns parse error in src/config.rs test module -- [X] T044 [P] [US4] Test missing required fields validation in src/config.rs test module -- [X] T045 [P] [US4] Test default values applied correctly in src/config.rs test module -- [X] T046 [P] [US4] Test get_socket_addr produces valid address in src/config.rs test module -- [X] T047 [P] [US4] Test config loading from multiple sources (file, conf.d, env vars) in src/config.rs test module - -**Checkpoint**: Configuration processing fully tested - can refactor config initialization safely - ---- - -## Phase 7: User Story 3 - API Endpoint Integration Testing (Priority: P2) - -**Goal**: Verify all REST endpoints handle requests correctly, return proper response formats, handle errors, and integrate with Graphite mocks. - -**Independent Test**: Start test server, send HTTP requests with various parameters, validate response structure and status codes. - -### Tests for User Story 3 - -- [X] T048 [P] [US3] Test /api/v1/ root endpoint returns name in src/api/v1.rs test module -- [X] T049 [P] [US3] Test /api/v1/info returns API info in src/api/v1.rs test module -- [X] T050 [P] [US3] Test /api/v1/health with valid service returns 200 + JSON in src/api/v1.rs test module -- [X] T051 [P] [US3] Test /api/v1/health with unknown service returns 409 in src/api/v1.rs test module -- [X] T052 [P] [US3] Test /api/v1/health with missing params returns 400 in src/api/v1.rs test module -- [X] T053 [P] [US3] Test /render with flag target returns boolean datapoints in src/graphite.rs test module -- [X] T054 [P] [US3] Test /render with health target returns health scores in src/graphite.rs test module -- [X] T055 [P] [US3] Test /render with invalid target returns empty array in src/graphite.rs test module -- [X] T056 [P] [US3] Test /metrics/find at all hierarchy levels in src/graphite.rs test module -- [X] T057 [P] [US3] Test /functions returns empty object in src/graphite.rs test module -- [X] T058 [P] [US3] Test /tags/autoComplete/tags returns empty array in src/graphite.rs test module -- [X] T059 [US3] Create full API integration test with mocked Graphite in tests/integration_api.rs -- [X] T060 [US3] Test error response format validation in tests/integration_api.rs - -**Checkpoint**: All API endpoints tested - API contract protected during refactoring - ---- - -## Phase 8: User Story 5 - Graphite Integration Testing (Priority: P3) - -**Goal**: Verify Graphite client query building, response parsing, and error handling for network failures or malformed data. - -**Independent Test**: Mock Graphite HTTP responses with various data formats and error conditions, verify correct parsing or error handling. - -### Tests for User Story 5 - -- [X] T061 [P] [US5] Test query building produces valid Graphite syntax in src/graphite.rs test module (Covered by test_get_graphite_data) -- [X] T062 [P] [US5] Test valid JSON with datapoints parses correctly in src/graphite.rs test module (Covered by test_get_graphite_data) -- [X] T063 [P] [US5] Test empty datapoints array handled gracefully in src/graphite.rs test module (Covered by integration tests) -- [X] T064 [P] [US5] Test HTTP 4xx error returns GraphiteError in src/graphite.rs test module -- [X] T065 [P] [US5] Test HTTP 5xx error returns GraphiteError in src/graphite.rs test module -- [X] T066 [P] [US5] Test malformed JSON response handling in src/graphite.rs test module -- [X] T067 [P] [US5] Test connection timeout handling in src/graphite.rs test module -- [X] T068 [P] [US5] Test metric discovery with wildcards in src/graphite.rs test module (Covered by test_get_grafana_find) -- [X] T069 [P] [US5] Test null and NaN values handled gracefully in src/graphite.rs test module (Covered by common tests) -- [X] T070 [P] [US5] Test partial response handling (some metrics succeed, others fail) in src/graphite.rs test module - -**Checkpoint**: Graphite integration fully tested - can refactor TSDB client safely - ---- - -## Phase 9: Polish & Cross-Cutting Concerns - -**Purpose**: Coverage validation, CI integration, and documentation - -- [X] T071 Run cargo tarpaulin and generate coverage report -- [X] T072 Verify 95% coverage threshold met for core business functions (Achieved: 97.18% for config+common+types) -- [X] T073 Identify and fill any coverage gaps with additional tests -- [X] T074 Add coverage enforcement to CI with --fail-under 95 flag -- [X] T075 [P] Generate HTML coverage report for documentation -- [X] T076 Verify all tests pass in under 2 minutes execution time (Achieved: < 1 second) -- [X] T077 [P] Document testing approach in README or docs/testing.md -- [X] T078 [P] Add test execution commands reference to documentation -- [X] T079 Verify tests detect 100% of intentional breaking changes (Validated via operator swap test) -- [X] T080 Final validation against all 25 functional requirements (FR-001 to FR-025) - ---- - -## Dependencies & Execution Order - -### Phase Dependencies - -- **Setup (Phase 1)**: No dependencies - can start immediately -- **Foundational (Phase 2)**: Depends on Setup completion - BLOCKS all user story tests -- **User Stories (Phase 3-8)**: All depend on Foundational phase completion - - User Story 1 (Phase 3): Core metric flag tests - highest priority - - User Story 6 (Phase 4): Regression suite - depends on US1 tests existing - - User Story 2 (Phase 5): Health aggregation - depends on US1 tests as foundation - - User Story 4 (Phase 6): Configuration processing - independent after foundational - - User Story 3 (Phase 7): API endpoints - independent after foundational - - User Story 5 (Phase 8): Graphite integration - independent after foundational -- **Polish (Phase 9)**: Depends on all user story tests being complete - -### User Story Dependencies - -- **User Story 1 (P1)**: Can start after Foundational (Phase 2) - No dependencies on other stories -- **User Story 6 (P1)**: Can start after User Story 1 complete - Validates regression detection -- **User Story 2 (P1)**: Can start after Foundational (Phase 2) - Builds on US1 tests as foundation -- **User Story 4 (P2)**: Can start after Foundational (Phase 2) - Independent of other stories -- **User Story 3 (P2)**: Can start after Foundational (Phase 2) - Independent of other stories -- **User Story 5 (P3)**: Can start after Foundational (Phase 2) - Independent of other stories - -### Within Each User Story - -- Tests marked [P] can run in parallel (different source files) -- Integration tests depend on corresponding unit tests being written first -- Each story should be complete before moving to next priority - -### Parallel Opportunities - -- All Setup tasks (T002-T004) marked [P] can run in parallel -- All Foundational tasks (T006-T009) marked [P] can run in parallel -- Once Foundational phase completes, US4, US3, and US5 can start in parallel (US1 and US2 have dependencies) -- All unit tests within a story marked [P] can be written in parallel -- Integration tests within a story can be written in parallel after unit tests exist - ---- - -## Parallel Example: User Story 1 - -```bash -# Launch all Lt operator tests for User Story 1 together: -Task T010: "Test Lt operator with value < threshold returns true" -Task T011: "Test Lt operator with value >= threshold returns false" - -# Launch all operator tests in parallel: -Task T010: "Lt operator tests" -Task T012-T013: "Gt operator tests" -Task T014-T015: "Eq operator tests" - -# All [P] marked tests within US1 can be developed simultaneously -``` - ---- - -## Implementation Strategy - -### MVP First (User Stories 1 & 6 Only) - -1. Complete Phase 1: Setup -2. Complete Phase 2: Foundational (CRITICAL - blocks all stories) -3. Complete Phase 3: User Story 1 (Core metric flag tests) -4. Complete Phase 4: User Story 6 (Regression validation) -5. **STOP and VALIDATE**: Verify regression detection works for core metrics -6. Deploy/demo if ready - -### Incremental Delivery - -1. Complete Setup + Foundational → Test infrastructure ready -2. Add User Story 1 → Test independently → 20+ metric evaluation tests complete (MVP!) -3. Add User Story 6 → Validate regression detection → Refactoring confidence achieved -4. Add User Story 2 → 15+ health aggregation tests → Complex business logic covered -5. Add User Story 4 → Configuration processing covered → Initialization safe to refactor -6. Add User Story 3 → API contract protected → External interface stable -7. Add User Story 5 → Graphite integration covered → TSDB client safe to refactor -8. Complete Coverage validation → 95% threshold achieved - -### Parallel Team Strategy - -With multiple developers: - -1. Team completes Setup + Foundational together -2. Once Foundational is done: - - Developer A: User Story 1 (T010-T020) - - Developer B: User Story 4 (T037-T047) - parallel independent work - - Developer C: User Story 5 (T061-T070) - parallel independent work -3. After US1 complete: - - Developer A: User Story 6 (T021-T025) - validates US1 - - Developer D: User Story 2 (T026-T036) - builds on US1 foundation - - Developer E: User Story 3 (T048-T060) - parallel independent work -4. Stories complete and validate independently - ---- - -## Test Count Summary - -| Category | Task Range | Test Count | Priority | -|----------|------------|------------|----------| -| **Setup** | T001-T004 | 4 fixtures | Foundation | -| **Foundational** | T005-T009 | 5 helpers | Foundation | -| **US1: Metric Flag Tests** | T010-T020 | 11 tests | P1 | -| **US6: Regression Suite** | T021-T025 | 5 tests | P1 | -| **US2: Health Aggregation** | T026-T036 | 11 tests | P1 | -| **US4: Configuration** | T037-T047 | 11 tests | P2 | -| **US3: API Endpoints** | T048-T060 | 13 tests | P2 | -| **US5: Graphite Integration** | T061-T070 | 10 tests | P3 | -| **Polish & Coverage** | T071-T080 | 10 tasks | Final | -| **Total** | 80 tasks | **61 tests** | | - -**Target Met**: 61 tests exceeds the minimum 50 required (SC-002) - -**Coverage Breakdown**: -- Metric evaluation: 11 tests (target: 20+) ✅ -- Health aggregation: 11 tests (target: 15+) ✅ -- API endpoints: 13 tests (target: 10+) ✅ -- Configuration: 11 tests (target: 5+) ✅ -- Error handling: 15+ tests (distributed across stories) ✅ - ---- - -## Success Metrics - -| Metric | Target | Measurement | Task References | -|--------|--------|-------------|-----------------| -| Code coverage | ≥95% | cargo tarpaulin --fail-under 95 | T071-T072 | -| Test count | ≥50 | 61 tests delivered | All test tasks | -| Execution time | <2 min | Verify with time cargo test | T076 | -| Regression detection | 100% | Intentional breakage tests | T021-T022, T079 | -| Functional requirements | 25/25 | All FR-001 to FR-025 validated | T080 | - ---- - -## Notes - -- [P] tasks = different files/modules, no dependencies, can run in parallel -- [Story] label maps task to specific user story for traceability (US1-US6) -- Tests are explicitly requested in spec to achieve 95% coverage goal -- All tests follow TDD principle: write tests first, verify they test existing code behavior -- Use custom assertions (assert_metric_flag, assert_health_score) for clear failure messages -- Each test module creates isolated mock servers (mockito::Server::new()) -- Commit after each user story phase completion -- Tests serve dual purpose: regression protection + executable documentation -- Avoid: shared mutable state, hardcoded ports, flaky timing dependencies diff --git a/specs/002-functional-test-suite/test-implementation-summary.md b/specs/002-functional-test-suite/test-implementation-summary.md deleted file mode 100644 index 4aa7be5..0000000 --- a/specs/002-functional-test-suite/test-implementation-summary.md +++ /dev/null @@ -1,249 +0,0 @@ -# Test Suite Implementation Summary - -## Completed Work (30/80 tasks - 37.5%) - -### Phase 1: Setup ✅ (T001-T004) -- ✅ Created fixtures module structure -- ✅ Test configuration fixtures (10+ config scenarios) -- ✅ Graphite response mock data (20+ response fixtures) -- ✅ Custom assertion helpers and test utilities - -**Files Created:** -- `tests/fixtures/mod.rs` -- `tests/fixtures/configs.rs` (6.5KB, 11 fixture functions) -- `tests/fixtures/graphite_responses.rs` (9.4KB, 25+ mock responses) -- `tests/fixtures/helpers.rs` (11KB, 10+ helper functions) - -### Phase 2: Foundational ✅ (T005-T009) -- ✅ Added cargo-tarpaulin to CI pipeline -- ✅ Implemented all test helper functions - - `create_test_state()` - - `create_test_state_with_mock_url()` - - `create_custom_test_state()` - - `create_multi_metric_test_state()` -- ✅ Custom assertions for clear error messages - - `assert_metric_flag()` - - `assert_health_score()` - - `assert_health_score_within()` - -**Files Created:** -- `.github/workflows/coverage.yml` (Coverage CI configuration) -- Updated `.dockerignore` -- Updated `.gitignore` (coverage artifacts) - -### Phase 3: User Story 1 ✅ (T010-T020) -**Core Metric Flag Evaluation - 11 Unit Tests** - -All tests passing in `src/common.rs`: -- ✅ T010: Lt operator below threshold → true -- ✅ T011: Lt operator above/equal threshold → false -- ✅ T012: Gt operator above threshold → true -- ✅ T013: Gt operator below/equal threshold → false -- ✅ T014: Eq operator equal threshold → true -- ✅ T015: Eq operator not equal threshold → false -- ✅ T016: None value returns false for all operators -- ✅ T017: Boundary conditions (threshold ± 0.001) -- ✅ T018: Negative values with all operators -- ✅ T019: Zero threshold edge case -- ✅ T020: Mixed operators scenario - -**Test Coverage:** 100% of `get_metric_flag_state()` function - -### Phase 4: User Story 6 ✅ (T021-T025) -**Regression Suite Validation** - -- ✅ T021: Documented regression validation approach -- ✅ T022: Verified tests catch breaking changes (8/11 tests failed with operator swap) -- ✅ T023: Full test suite runs in < 1 second (target: < 2 minutes) -- ✅ T024: Zero false positives confirmed -- ✅ T025: Created comprehensive `docs/testing.md` - -**Test Execution Time:** 0.02 seconds (well under 2-minute target) - -## Remaining Work (50/80 tasks - 62.5%) - -### Phase 5: User Story 2 (T026-T036) - Service Health Aggregation -**Status:** Not yet implemented - -**Required Tests (11 tests):** -- Expression evaluation (OR, AND, complex boolean) -- Weighted expression calculations -- Error handling (unknown service, environment) -- End-to-end with mocked Graphite -- Edge cases (empty data, partial data) - -**Implementation Pattern:** Add tests to `src/common.rs` test module for `get_service_health()` function - -### Phase 6: User Story 4 (T037-T047) - Configuration Processing -**Status:** Partially complete (existing tests in config.rs) - -**Existing Tests:** -- 3 tests in `src/config.rs` -- 1 test in `src/types.rs` - -**Additional Tests Needed (11 tests):** -- Template variable substitution -- Multiple environment expansion -- Threshold overrides -- Dash-to-underscore conversion -- Validation and error cases - -### Phase 7: User Story 3 (T048-T060) - API Endpoints -**Status:** Not yet implemented - -**Required Tests (13 tests):** -- REST endpoint handlers (v1/health, v1/info, render, find, functions, tags) -- Response format validation -- Error handling (400, 409, 500) -- Integration tests with mock server - -**Implementation Location:** `src/api/v1.rs` test module + `tests/integration_api.rs` - -### Phase 8: User Story 5 (T061-T070) - Graphite Integration -**Status:** Partially complete (existing tests in graphite.rs) - -**Existing Tests:** -- 3 tests in `src/graphite.rs` - -**Additional Tests Needed (10 tests):** -- Query building validation -- Response parsing (valid, empty, malformed) -- Error handling (4xx, 5xx, timeout, connection) -- Null/NaN value handling -- Partial response handling - -### Phase 9: Polish (T071-T080) - Coverage & Documentation -**Status:** Partially complete - -**Completed:** -- ✅ T077: Documentation created (`docs/testing.md`) -- ✅ T078: Test commands documented - -**Remaining (10 tasks):** -- Coverage measurement and reporting -- Gap identification -- CI enforcement -- HTML report generation -- Final validation - -## Test Infrastructure Quality - -### Strengths ✅ -1. **Comprehensive Fixtures**: 35+ test fixtures covering all scenarios -2. **Reusable Helpers**: 10+ helper functions eliminate duplication -3. **Clear Assertions**: Custom assertions provide descriptive error messages -4. **CI Integration**: Automated coverage measurement with 95% threshold -5. **Fast Execution**: Tests run in < 1 second -6. **Good Documentation**: Comprehensive testing guide created - -### Coverage Status - -| Module | Existing Tests | New Tests Added | Coverage | -|--------|---------------|-----------------|----------| -| `common.rs` | 0 | 11 | ⭐⭐⭐⭐⭐ Excellent | -| `config.rs` | 3 | 0 | ⭐⭐⭐ Good | -| `types.rs` | 1 | 0 | ⭐⭐ Fair | -| `graphite.rs` | 3 | 0 | ⭐⭐⭐ Good | -| `api/v1.rs` | 0 | 0 | ❌ Missing | - -## Critical Path to 95% Coverage - -### High Priority (Must Complete) -1. **Phase 5: Service Health Tests** (T026-T036) - - Most complex business logic - - Integration of multiple components - - Expected to add 15-20% coverage - -2. **Phase 7: API Endpoint Tests** (T048-T060) - - Public interface testing - - Error handling validation - - Expected to add 10-15% coverage - -### Medium Priority (Should Complete) -3. **Phase 6: Configuration Tests** (T037-T047) - - Build on existing 4 tests - - Validation logic coverage - - Expected to add 5-10% coverage - -4. **Phase 8: Graphite Integration** (T061-T070) - - Build on existing 3 tests - - External service mocking - - Expected to add 5-8% coverage - -### Low Priority (Nice to Have) -5. **Phase 9: Coverage Polish** (T071-T080) - - Validation and reporting - - Gap filling - - Documentation updates - -## Next Steps - -### Immediate Actions -1. Run current coverage measurement: - ```bash - cargo install cargo-tarpaulin - cargo tarpaulin --out Html --output-dir ./coverage - ``` - -2. Identify coverage gaps in critical modules - -3. Prioritize Phase 5 (Service Health) implementation - -### Recommended Implementation Order -1. **Week 1**: Phase 5 (US2) - Service health aggregation tests -2. **Week 2**: Phase 7 (US3) - API endpoint tests -3. **Week 3**: Phase 6 (US4) + Phase 8 (US5) - Configuration and Graphite -4. **Week 4**: Phase 9 - Coverage validation and polish - -## Test Metrics - -### Current State -- **Total Tests**: 18 (baseline) + 11 (new) = 29 tests -- **Test Files**: 4 modules with tests -- **Execution Time**: < 0.05 seconds -- **Coverage**: To be measured (estimated 40-50%) - -### Target State (Full Implementation) -- **Total Tests**: 60+ tests -- **Test Files**: 8+ modules with tests -- **Execution Time**: < 2 minutes (< 2 seconds expected) -- **Coverage**: ≥ 95% for core business functions - -## Success Criteria Status - -| Criterion | Target | Current Status | -|-----------|--------|----------------| -| Code Coverage | ≥95% | 🚧 In Progress (~40-50% estimated) | -| Test Count | ≥50 | ✅ On Track (29/50+) | -| Execution Time | <2 min | ✅ Excellent (<1 sec) | -| Regression Detection | 100% | ✅ Verified (8/11 failures) | -| False Positives | 0 | ✅ Verified | - -## Conclusion - -### What's Working Well ✅ -- Solid test infrastructure foundation -- Excellent test execution performance -- Clear documentation and helpers -- CI/CD integration functional -- Core business logic (metric flags) fully tested - -### What Needs Attention ⚠️ -- Health aggregation tests (highest priority) -- API endpoint tests (public interface) -- Coverage measurement and gaps -- Integration test scenarios - -### Risk Assessment -**Risk Level:** Low to Medium - -**Rationale:** -- Core metric evaluation fully tested (highest risk code) -- Test infrastructure proven and working -- Remaining tests follow established patterns -- Clear implementation roadmap - -**Mitigation:** -- Existing fixtures can be reused -- Helper functions simplify new test creation -- Test patterns established and documented diff --git a/specs/003-sd-api-v2-migration/checklists/requirements.md b/specs/003-sd-api-v2-migration/checklists/requirements.md deleted file mode 100644 index 81d8309..0000000 --- a/specs/003-sd-api-v2-migration/checklists/requirements.md +++ /dev/null @@ -1,54 +0,0 @@ -# Specification Quality Checklist: Reporter Migration to Status Dashboard API V2 - -**Purpose**: Validate specification completeness and quality before proceeding to planning -**Created**: 2025-01-22 -**Feature**: [spec.md](../spec.md) - -## Content Quality - -- [x] No implementation details (languages, frameworks, APIs) -- [x] Focused on user value and business needs -- [x] Written for non-technical stakeholders -- [x] All mandatory sections completed - -## Requirement Completeness - -- [x] No [NEEDS CLARIFICATION] markers remain -- [x] Requirements are testable and unambiguous -- [x] Success criteria are measurable -- [x] Success criteria are technology-agnostic (no implementation details) -- [x] All acceptance scenarios are defined -- [x] Edge cases are identified -- [x] Scope is clearly bounded -- [x] Dependencies and assumptions identified - -## Feature Readiness - -- [x] All functional requirements have clear acceptance criteria -- [x] User scenarios cover primary flows -- [x] Feature meets measurable outcomes defined in Success Criteria -- [x] No implementation details leak into specification - -## Validation Results - -### Initial Review (2025-01-22) - -All checklist items passed on initial review: - -✅ **Content Quality**: The specification focuses on what the reporter needs to accomplish (migrate from V1 to V2 API) without prescribing implementation details. While it references specific API endpoints and data structures, these are part of the external API contract that the reporter must integrate with, not implementation choices. - -✅ **Requirement Completeness**: All requirements are testable and unambiguous. No clarifications needed - the feature scope is well-defined based on the existing V1 implementation and the V2 API schema. - -✅ **Feature Readiness**: The three prioritized user stories cover the complete migration scope: -- P1: Core incident creation via V2 API (essential MVP) -- P2: Component cache management (enables robustness) -- P3: Authorization compatibility (confirms backward compatibility) - -Each story is independently testable and delivers standalone value. - -## Notes - -- The specification references API endpoints and data structures because these are external contracts defined by the Status Dashboard API V2, not implementation details of the reporter -- The feature scope is constrained to incident creation only; incident updates and closures are explicitly out of scope -- Authorization remains unchanged, minimizing migration risk -- Component cache management is essential for efficient operation and handling dynamic component additions diff --git a/specs/003-sd-api-v2-migration/contracts/README.md b/specs/003-sd-api-v2-migration/contracts/README.md deleted file mode 100644 index 48ee2c9..0000000 --- a/specs/003-sd-api-v2-migration/contracts/README.md +++ /dev/null @@ -1,57 +0,0 @@ -# API Contracts: Status Dashboard V2 - -**Feature**: Reporter Migration to Status Dashboard API V2 -**Date**: 2025-01-23 - -This directory contains API contract specifications for the Status Dashboard V2 endpoints used by the reporter. - -## Files - -- `components-api.yaml`: GET /v2/components endpoint contract -- `incidents-api.yaml`: POST /v2/incidents endpoint contract -- `request-examples/`: Sample request payloads -- `response-examples/`: Sample response payloads - -## Source - -All contracts are derived from the project's OpenAPI specification: -- **File**: `/openapi.yaml` (project root) -- **Version**: Status Dashboard API 1.0.0 -- **Endpoints Used**: - - `GET /v2/components` (line 138-151) - - `POST /v2/incidents` (line 254-270) - -## Testing - -Contracts can be validated using OpenAPI tooling: - -```bash -# Validate against OpenAPI schema -npx @redocly/cli lint openapi.yaml - -# Generate mock server for testing -npx @stoplight/prism mock openapi.yaml -``` - -## Usage in Reporter - -### Components Endpoint -```rust -// Fetch all components at startup -let components: Vec = - req_client.get(&format!("{}/v2/components", sdb_url)) - .send().await? - .json().await?; -``` - -### Incidents Endpoint -```rust -// Create incident -let incident = IncidentData { /* ... */ }; -let response: IncidentPostResponse = - req_client.post(&format!("{}/v2/incidents", sdb_url)) - .headers(auth_headers) - .json(&incident) - .send().await? - .json().await?; -``` diff --git a/specs/003-sd-api-v2-migration/contracts/components-api.md b/specs/003-sd-api-v2-migration/contracts/components-api.md deleted file mode 100644 index 7bf88f3..0000000 --- a/specs/003-sd-api-v2-migration/contracts/components-api.md +++ /dev/null @@ -1,296 +0,0 @@ -# GET /v2/components - -## Overview - -Fetch all components from Status Dashboard to build the component ID cache. - -**Endpoint**: `GET /v2/components` -**Authentication**: Optional (HMAC-JWT Bearer token if configured) -**Frequency**: Startup + on-demand cache refresh - -## Request - -### HTTP Method -``` -GET /v2/components HTTP/1.1 -Host: {status-dashboard-url} -Authorization: Bearer {jwt-token} -``` - -### Headers - -| Header | Required | Value | Description | -|--------|----------|-------|-------------| -| `Authorization` | Optional | `Bearer {jwt-token}` | HMAC-signed JWT if secret configured | - -### Query Parameters - -None. - -### Request Body - -None (GET request). - -## Response - -### Success Response (200 OK) - -**Content-Type**: `application/json` - -**Schema**: -```yaml -type: array -items: - type: object - required: [id, name, attributes] - properties: - id: - type: integer - format: int64 - description: Component ID (primary key) - example: 218 - name: - type: string - description: Component name - example: "Object Storage Service" - attributes: - type: array - items: - type: object - properties: - name: - type: string - enum: [category, region, type] - description: Attribute name - example: "category" - value: - type: string - description: Attribute value - example: "Storage" -``` - -**Example Response**: -```json -[ - { - "id": 218, - "name": "Object Storage Service", - "attributes": [ - { - "name": "category", - "value": "Storage" - }, - { - "name": "region", - "value": "EU-DE" - } - ] - }, - { - "id": 254, - "name": "Compute Service", - "attributes": [ - { - "name": "category", - "value": "Compute" - }, - { - "name": "region", - "value": "EU-NL" - }, - { - "name": "type", - "value": "vm" - } - ] - }, - { - "id": 312, - "name": "Database Service", - "attributes": [] - } -] -``` - -### Error Responses - -#### 401 Unauthorized -Invalid or missing authentication token (if auth required). - -```json -{ - "errMsg": "Invalid or missing authorization token" -} -``` - -#### 500 Internal Server Error -Server-side error. - -```json -{ - "errMsg": "internal server error" -} -``` - -## Rust Implementation - -### Request Struct - -```rust -// No request body struct needed (GET request) -``` - -### Response Struct - -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct StatusDashboardComponent { - pub id: u32, - pub name: String, - #[serde(default)] - pub attributes: Vec, -} - -#[derive(Clone, Deserialize, Serialize, Debug, PartialEq, Eq, Hash, Ord, PartialOrd)] -pub struct ComponentAttribute { - pub name: String, - pub value: String, -} -``` - -### Usage Example - -```rust -use reqwest::{Client, header::{HeaderMap, AUTHORIZATION}}; -use serde::Deserialize; - -async fn fetch_components( - client: &Client, - base_url: &str, - auth_headers: &HeaderMap, -) -> Result, Box> { - let url = format!("{}/v2/components", base_url); - - let response = client - .get(&url) - .headers(auth_headers.clone()) - .send() - .await?; - - response.error_for_status_ref()?; - - let components = response.json::>().await?; - - tracing::info!("Fetched {} components from Status Dashboard", components.len()); - - Ok(components) -} -``` - -## Cache Building - -Once components are fetched, build the cache: - -```rust -use std::collections::HashMap; - -fn build_component_id_cache( - components: Vec -) -> HashMap<(String, Vec), u32> { - components.into_iter().map(|c| { - let mut attrs = c.attributes; - attrs.sort(); // Ensure deterministic cache key - ((c.name, attrs), c.id) - }).collect() -} -``` - -## Error Handling - -```rust -async fn fetch_components_with_retry( - client: &Client, - base_url: &str, - auth_headers: &HeaderMap, -) -> Option> { - for attempt in 1..=3 { - match fetch_components(client, base_url, auth_headers).await { - Ok(components) => { - tracing::info!("Successfully fetched {} components", components.len()); - return Some(components); - } - Err(e) => { - tracing::error!("Failed to fetch components (attempt {}/3): {}", attempt, e); - if attempt < 3 { - tracing::info!("Retrying in 60 seconds..."); - tokio::time::sleep(Duration::from_secs(60)).await; - } else { - tracing::error!("Could not fetch components after 3 attempts"); - return None; - } - } - } - } - None -} -``` - -## Contract Validation - -### Valid Response Examples - -✅ **Complete component with attributes**: -```json -{ - "id": 218, - "name": "Object Storage Service", - "attributes": [ - {"name": "region", "value": "EU-DE"} - ] -} -``` - -✅ **Component without attributes**: -```json -{ - "id": 312, - "name": "Database Service", - "attributes": [] -} -``` - -### Invalid Response Examples - -❌ **Missing required field `id`**: -```json -{ - "name": "Storage", - "attributes": [] -} -``` -*Error*: Serde deserialization fails - -❌ **Invalid attribute structure**: -```json -{ - "id": 218, - "name": "Storage", - "attributes": [ - {"key": "region", "val": "EU-DE"} // Should be "name" and "value" - ] -} -``` -*Error*: Serde deserialization fails - -## Performance Considerations - -- **Response Size**: ~100 components × ~200 bytes = ~20 KB (small payload) -- **Frequency**: Once at startup + rare refreshes (only on cache miss) -~~- **Timeout**: Use 10-second timeout per FR-014~~ -- **Caching**: Store in memory for duration of reporter process - -## Security - -- **Authentication**: Same HMAC-JWT mechanism as V1 API (FR-008) -- **Data Exposure**: Component names and attributes are public data (Status Dashboard is public) -- **Authorization**: Reporter only needs read access to components endpoint diff --git a/specs/003-sd-api-v2-migration/contracts/incidents-api.md b/specs/003-sd-api-v2-migration/contracts/incidents-api.md deleted file mode 100644 index eb3671d..0000000 --- a/specs/003-sd-api-v2-migration/contracts/incidents-api.md +++ /dev/null @@ -1,468 +0,0 @@ -# POST /v2/incidents - -## Overview - -Create a new incident in Status Dashboard when a service health issue is detected. - -**Endpoint**: `POST /v2/incidents` -**Authentication**: Required (HMAC-JWT Bearer token) -**Frequency**: Per health issue detection (~1-10 incidents/min under normal load) - -## Request - -### HTTP Method -``` -POST /v2/incidents HTTP/1.1 -Host: {status-dashboard-url} -Authorization: Bearer {jwt-token} -Content-Type: application/json -``` - -### Headers - -| Header | Required | Value | Description | -|--------|----------|-------|-------------| -| `Authorization` | Yes | `Bearer {jwt-token}` | HMAC-signed JWT (unchanged from V1) | -| `Content-Type` | Yes | `application/json` | Request body format | - -### Request Body - -**Schema**: -```yaml -type: object -required: [title, impact, components, start_date, type] -properties: - title: - type: string - description: Incident title (static for auto-created) - example: "System incident from monitoring system" - description: - type: string - description: Generic description (optional, defaults to empty) - example: "System-wide incident affecting one or multiple components. Created automatically." - impact: - type: integer - enum: [0, 1, 2, 3] - description: "Impact level: 0=none, 1=minor, 2=major, 3=critical" - example: 2 - components: - type: array - items: - type: integer - description: Array of component IDs (resolved from cache) - example: [218] - start_date: - type: string - format: date-time - description: Incident start time (RFC3339, health metric timestamp - 1s) - example: "2025-01-22T10:30:44Z" - end_date: - type: string - format: date-time - description: Incident end time (optional, not used for auto-created) - system: - type: boolean - default: false - description: System-generated flag (true for auto-created) - example: true - type: - type: string - enum: [incident, maintenance] - description: Event type (always "incident" for auto-created) - example: "incident" -``` - -**Example Request Body** (typical auto-created incident): -```json -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 2, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} -``` - -**Example Request Body** (multi-component incident): -```json -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 3, - "components": [218, 254, 312], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} -``` - -## Response - -### Success Response (200 OK) - -**Content-Type**: `application/json` - -**Schema**: -```yaml -type: object -properties: - result: - type: array - items: - type: object - properties: - component_id: - type: integer - format: int64 - description: Component ID from request - incident_id: - type: integer - format: int64 - description: Created or existing incident ID -``` - -**Example Response** (new incident created): -```json -{ - "result": [ - { - "component_id": 218, - "incident_id": 456 - } - ] -} -``` - -**Example Response** (existing incident returned - duplicate detection): -```json -{ - "result": [ - { - "component_id": 218, - "incident_id": 123 - } - ] -} -``` - -**Duplicate Handling**: If an identical incident already exists (same component + impact + active), the API returns the existing incident ID. The reporter does not need to implement deduplication logic (FR-016). - -### Error Responses - -#### 400 Bad Request -Invalid request body (missing required fields, invalid impact value, etc.). - -```json -{ - "errMsg": "Invalid request: impact must be between 0 and 3" -} -``` - -#### 401 Unauthorized -Invalid or missing authentication token. - -```json -{ - "errMsg": "Invalid or missing authorization token" -} -``` - -#### 404 Not Found -Component ID(s) not found in Status Dashboard. - -```json -{ - "errMsg": "component does not exist" -} -``` - -#### 500 Internal Server Error -Server-side error. - -```json -{ - "errMsg": "internal server error" -} -``` - -## Rust Implementation - -### Request Struct - -```rust -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; - -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentData { - pub title: String, - #[serde(default)] - pub description: String, - pub impact: u8, - pub components: Vec, - pub start_date: DateTime, - #[serde(default)] - pub system: bool, - #[serde(rename = "type")] - pub incident_type: String, -} -``` - -### Response Struct - -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentPostResponse { - pub result: Vec, -} - -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentPostResult { - pub component_id: u32, - pub incident_id: u32, -} -``` - -### Usage Example - -```rust -use reqwest::{Client, header::HeaderMap}; - -async fn create_incident( - client: &Client, - base_url: &str, - auth_headers: &HeaderMap, - incident: &IncidentData, -) -> Result> { - let url = format!("{}/v2/incidents", base_url); - - let response = client - .post(&url) - .headers(auth_headers.clone()) - .json(incident) - .send() - .await?; - - if !response.status().is_success() { - let status = response.status(); - let body = response.text().await?; - tracing::error!("Incident creation failed [{}]: {}", status, body); - return Err(format!("API error: {} - {}", status, body).into()); - } - - let result = response.json::().await?; - tracing::info!( - "Incident created: component_id={}, incident_id={}", - result.result[0].component_id, - result.result[0].incident_id - ); - - Ok(result) -} -``` - -### Building IncidentData - -```rust -fn build_incident_data( - component_id: u32, - impact: u8, - timestamp: i64, -) -> IncidentData { - use chrono::{DateTime, Utc}; - - // Adjust timestamp by -1 second per FR-011 - let start_date = DateTime::::from_timestamp(timestamp - 1, 0) - .expect("Invalid timestamp"); - - IncidentData { - title: "System incident from monitoring system".to_string(), - description: "System-wide incident affecting one or multiple components. Created automatically.".to_string(), - impact, - components: vec![component_id], - start_date, - system: true, - incident_type: "incident".to_string(), - } -} -``` - -### Error Handling - -```rust -async fn report_incident( - client: &Client, - base_url: &str, - auth_headers: &HeaderMap, - incident: &IncidentData, -) { - match create_incident(client, base_url, auth_headers, incident).await { - Ok(response) => { - tracing::info!( - "Successfully created/updated incident {} for component {}", - response.result[0].incident_id, - response.result[0].component_id - ); - } - Err(e) => { - tracing::error!("Failed to create incident: {}", e); - // Do not retry immediately - next monitoring cycle will retry (FR-015) - } - } -} -``` - -## Contract Validation - -### Valid Request Examples - -✅ **Minimal auto-created incident**: -```json -{ - "title": "System incident from monitoring system", - "impact": 1, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "type": "incident" -} -``` - -✅ **Complete auto-created incident**: -```json -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 3, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} -``` - -### Invalid Request Examples - -❌ **Missing required field `title`**: -```json -{ - "impact": 2, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "type": "incident" -} -``` -*Error*: 400 Bad Request - -❌ **Invalid impact value**: -```json -{ - "title": "Incident", - "impact": 5, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "type": "incident" -} -``` -*Error*: 400 Bad Request (impact must be 0-3) - -❌ **Empty components array**: -```json -{ - "title": "Incident", - "impact": 2, - "components": [], - "start_date": "2025-01-22T10:30:44Z", - "type": "incident" -} -``` -*Error*: 400 Bad Request (at least one component required) - -❌ **Invalid date format**: -```json -{ - "title": "Incident", - "impact": 2, - "components": [218], - "start_date": "2025-01-22 10:30:44", - "type": "incident" -} -``` -*Error*: 400 Bad Request (must be RFC3339 format) - -## Field Constraints (FR-002) - -| Field | Value | Rationale | -|-------|-------|-----------| -| `title` | `"System incident from monitoring system"` | Static generic title (FR-002) | -| `description` | `"System-wide incident affecting one or multiple components. Created automatically."` | Static generic description (FR-017, prevents sensitive data exposure) | -| `impact` | 0-3 from health metric | Direct mapping from service health (FR-002) | -| `components` | `[component_id]` | Resolved from cache lookup (FR-004) | -| `start_date` | Health timestamp - 1s | RFC3339 format, adjusted per FR-011 | -| `system` | `true` | Always true for auto-created (FR-009) | -| `type` | `"incident"` | Always "incident" for auto-created (FR-010) | - -## Sensitive Data Separation (FR-017) - -### Data NOT Sent to API (Logged Locally Only) - -The following information MUST NOT be included in the incident payload to prevent exposing sensitive operational data on the public Status Dashboard: - -- ❌ Service name (e.g., "swift", "nova") -- ❌ Environment name (e.g., "production", "staging") -- ❌ Component name (e.g., "Object Storage Service") -- ❌ Component attributes (e.g., `region=EU-DE`) -- ❌ Triggered metric names (e.g., "latency_p95", "error_rate") -- ❌ Metric values (e.g., "latency=450ms") - -### Data Logged Locally (For Diagnostics) - -```rust -tracing::info!( - timestamp = %start_date, - service = %service_name, - environment = %env_name, - component_name = %component.name, - component_attrs = ?component.attributes, - component_id = component_id, - impact = impact, - triggered_metrics = ?triggered_metric_names, - "Creating incident for health issue" -); -``` - -### Data Sent to API (Public, Generic) - -```json -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 2, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} -``` - -## Performance Considerations - -- **Request Size**: ~300 bytes per incident (small payload) -- **Frequency**: ~1-10 incidents/min under normal load, higher during widespread issues -~~- **Timeout**: 10 seconds per FR-014 (increased from 2s)~~ -- **Retry Strategy**: No immediate retry on failure, rely on next monitoring cycle (~60s per FR-015) - -## Security - -- **Authentication**: HMAC-JWT Bearer token (unchanged from V1, FR-008) -- **Data Privacy**: Generic title/description prevent sensitive data exposure (FR-017) -- **Component IDs**: Integer IDs expose less information than names/attributes -- **Public Dashboard**: All incident data is publicly visible on Status Dashboard - -## Idempotency - -The Status Dashboard API implements built-in duplicate detection: -- If an identical incident exists (same component + impact + still active), the API returns the existing incident ID -- The reporter does NOT need to track created incidents (FR-016) -- Each health issue detection results in a new POST request, API handles deduplication diff --git a/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-multi-component.json b/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-multi-component.json deleted file mode 100644 index 49fa42b..0000000 --- a/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-multi-component.json +++ /dev/null @@ -1,9 +0,0 @@ -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 3, - "components": [218, 254, 312], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} diff --git a/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-single-component.json b/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-single-component.json deleted file mode 100644 index 48288a3..0000000 --- a/specs/003-sd-api-v2-migration/contracts/request-examples/create-incident-single-component.json +++ /dev/null @@ -1,9 +0,0 @@ -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 2, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} diff --git a/specs/003-sd-api-v2-migration/contracts/response-examples/components-list.json b/specs/003-sd-api-v2-migration/contracts/response-examples/components-list.json deleted file mode 100644 index 1756953..0000000 --- a/specs/003-sd-api-v2-migration/contracts/response-examples/components-list.json +++ /dev/null @@ -1,39 +0,0 @@ -[ - { - "id": 218, - "name": "Object Storage Service", - "attributes": [ - { - "name": "category", - "value": "Storage" - }, - { - "name": "region", - "value": "EU-DE" - } - ] - }, - { - "id": 254, - "name": "Compute Service", - "attributes": [ - { - "name": "category", - "value": "Compute" - }, - { - "name": "region", - "value": "EU-NL" - }, - { - "name": "type", - "value": "vm" - } - ] - }, - { - "id": 312, - "name": "Database Service", - "attributes": [] - } -] diff --git a/specs/003-sd-api-v2-migration/contracts/response-examples/incident-created.json b/specs/003-sd-api-v2-migration/contracts/response-examples/incident-created.json deleted file mode 100644 index 7b24bb2..0000000 --- a/specs/003-sd-api-v2-migration/contracts/response-examples/incident-created.json +++ /dev/null @@ -1,8 +0,0 @@ -{ - "result": [ - { - "component_id": 218, - "incident_id": 456 - } - ] -} diff --git a/specs/003-sd-api-v2-migration/data-model.md b/specs/003-sd-api-v2-migration/data-model.md deleted file mode 100644 index 4f99385..0000000 --- a/specs/003-sd-api-v2-migration/data-model.md +++ /dev/null @@ -1,785 +0,0 @@ -# Data Model: Status Dashboard API V2 Migration - -**Feature**: Reporter Migration to Status Dashboard API V2 -**Branch**: `003-sd-api-v2-migration` -**Date**: 2025-01-23 - -## Overview - -This document defines the data entities and their relationships for the Status Dashboard API V2 migration. The migration introduces a component ID caching layer and restructures incident data to align with the V2 API schema. - ---- - -## Core Entities - -### 1. ComponentAttribute - -**Purpose**: Represents a key-value attribute that qualifies a component (e.g., `region=EU-DE`, `category=Storage`) - -**Rust Definition**: -```rust -#[derive(Clone, Deserialize, Serialize, Debug, PartialEq, Eq, Hash, Ord, PartialOrd)] -pub struct ComponentAttribute { - pub name: String, // Attribute name (e.g., "region", "category", "type") - pub value: String, // Attribute value (e.g., "EU-DE", "Storage") -} -``` - -**JSON Representation** (Status Dashboard API V2): -```json -{ - "name": "region", - "value": "EU-DE" -} -``` - -**Validation Rules**: -- `name`: Non-empty string, typically one of `["region", "category", "type"]` (per OpenAPI enum) -- `value`: Non-empty string - -**Traits**: -- `PartialOrd`, `Ord`: Required for sorting attributes before caching -- `Hash`, `Eq`: Required for use in HashMap keys -- `Serialize`, `Deserialize`: JSON API interaction - -**Relationships**: -- **Owned by**: `Component` (from config), `StatusDashboardComponent` (from API) -- **Used in**: Component cache key construction - ---- - -### 2. Component (Config) - -**Purpose**: Represents a component definition from the reporter's configuration file. Used to look up component IDs in the cache. - -**Rust Definition**: -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct Component { - pub name: String, // Component name (e.g., "Object Storage Service") - pub attributes: Vec, // Attributes from config + environment -} -``` - -**Source**: Reporter configuration (`config.yaml`) - -**Example**: -```yaml -# In config.yaml health_metrics section -health_metrics: - swift: - component_name: "Object Storage Service" - # attributes come from environment.attributes - -environments: - - name: production - attributes: - region: "EU-DE" - category: "Storage" -``` - -**Construction Logic** (from `reporter.rs`): -```rust -// Combines component_name from health_metric + attributes from environment -let component = Component { - name: health_metric.component_name.clone(), - attributes: env.attributes.clone(), -}; -``` - -**Relationships**: -- **Created from**: Configuration file (`config.yaml`) -- **Used for**: Component ID cache lookup -- **Key construction**: `(component.name, sorted(component.attributes))` → cache key - ---- - -### 3. StatusDashboardComponent (API Response) - -**Purpose**: Represents a component as returned by the Status Dashboard API `/v2/components` endpoint. Used to build the component ID cache. - -**Rust Definition**: -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct StatusDashboardComponent { - pub id: u32, // Component ID (primary key in Status Dashboard) - pub name: String, // Component name - #[serde(default)] - pub attributes: Vec, // Component attributes (may be empty) -} -``` - -**JSON Representation** (API response from `GET /v2/components`): -```json -{ - "id": 218, - "name": "Object Storage Service", - "attributes": [ - {"name": "category", "value": "Storage"}, - {"name": "region", "value": "EU-DE"} - ] -} -``` - -**Source**: Status Dashboard API `/v2/components` endpoint - -**Validation Rules**: -- `id`: Positive integer (u32) -- `name`: Non-empty string -- `attributes`: Array (may be empty per `#[serde(default)]`) - -**Relationships**: -- **Fetched from**: Status Dashboard API -- **Used to build**: Component ID cache (`ComponentCache`) -- **Cache entry**: `(name, sorted(attributes))` → `id` - ---- - -### 4. ComponentCache - -**Purpose**: In-memory cache mapping component names and attributes to component IDs. Avoids repeated API calls during monitoring cycles. - -**Rust Definition**: -```rust -type ComponentCache = HashMap<(String, Vec), u32>; -// Key: (component_name, sorted_attributes) -// Value: component_id -``` - -**Example Cache State**: -```rust -{ - ("Object Storage Service", vec![ - ComponentAttribute { name: "category", value: "Storage" }, - ComponentAttribute { name: "region", value: "EU-DE" }, - ]): 218, - - ("Compute Service", vec![ - ComponentAttribute { name: "category", value: "Compute" }, - ComponentAttribute { name: "region", value: "EU-NL" }, - ]): 254, -} -``` - -**Cache Operations**: -1. **Build** (startup): - ```rust - fn build_component_id_cache(components: Vec) - -> ComponentCache - { - components.into_iter().map(|c| { - let mut attrs = c.attributes; - attrs.sort(); // Ensure deterministic key - ((c.name, attrs), c.id) - }).collect() - } - ``` - -2. **Lookup**: - ```rust - fn lookup_component_id( - cache: &ComponentCache, - component: &Component - ) -> Option { - let mut attrs = component.attributes.clone(); - attrs.sort(); // Match cache key format - cache.get(&(component.name.clone(), attrs)).copied() - } - ``` - -3. **Refresh** (on miss): - ```rust - async fn refresh_cache(client: &Client, url: &str) - -> Result - { - let components = fetch_components(client, url).await?; - Ok(build_component_id_cache(components)) - } - ``` - -**Lifecycle**: -- **Created**: Reporter startup (with 3 retries, 60s delays per FR-006) -- **Refreshed**: On cache miss during incident creation (1 attempt, per FR-005) -- **Invalidated**: Never (components are stable; refresh only on miss) - -**Subset Matching** (FR-012): -The cache stores full component attributes from the Status Dashboard. Config may specify fewer attributes: -```rust -// Config component -Component { - name: "Storage", - attributes: vec![region=EU-DE] -} - -// Dashboard component (in cache) -StatusDashboardComponent { - id: 218, - name: "Storage", - attributes: vec![region=EU-DE, type=block] -} - -// Lookup fails because keys don't match exactly! -// Solution: FR-012 specifies subset matching, but cache uses exact key matching. -// Implementation must iterate cache to find subset matches. -``` - -**Corrected Lookup Algorithm** (for subset matching): -```rust -fn find_component_id( - cache: &ComponentCache, - target: &Component -) -> Option { - cache.iter() - .filter(|((name, _attrs), _id)| name == &target.name) - .find(|((_name, cache_attrs), _id)| { - // Config attrs must be subset of cache attrs - target.attributes.iter().all(|target_attr| { - cache_attrs.iter().any(|cache_attr| { - cache_attr.name == target_attr.name - && cache_attr.value == target_attr.value - }) - }) - }) - .map(|((_name, _attrs), id)| *id) -} -``` - -**Performance**: O(n) worst case where n = cache size (~100 components), acceptable for 60s monitoring intervals. - ---- - -### 5. IncidentData (V2 API Request) - -**Purpose**: Represents the incident payload sent to Status Dashboard API V2 `/v2/incidents` endpoint. - -**Rust Definition**: -```rust -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentData { - pub title: String, // Static: "System incident from monitoring system" - #[serde(default)] - pub description: String, // Static generic message (FR-017) - pub impact: u8, // 0=none, 1=minor, 2=major, 3=critical - pub components: Vec, // Component IDs (resolved from cache) - pub start_date: DateTime, // Health metric timestamp - 1s (RFC3339) - #[serde(default)] - pub system: bool, // Always true for auto-created incidents - #[serde(rename = "type")] - pub incident_type: String, // Always "incident" for auto-created -} -``` - -**JSON Representation** (POST request body): -```json -{ - "title": "System incident from monitoring system", - "description": "System-wide incident affecting one or multiple components. Created automatically.", - "impact": 2, - "components": [218], - "start_date": "2025-01-22T10:30:44Z", - "system": true, - "type": "incident" -} -``` - -**Field Mapping from Health Metric**: - -| Source | Field | Value | Transformation | -|--------|-------|-------|----------------| -| Health API | `timestamp` (i64) | Epoch seconds | `DateTime::from_timestamp(ts - 1, 0)` | -| Health API | `impact` (u8) | 0-3 | Direct copy | -| Config | Component name + attrs | "Storage", {region:EU-DE} | Resolve to component ID via cache | -| Static | `title` | - | "System incident from monitoring system" | -| Static | `description` | - | "System-wide incident..." | -| Static | `system` | - | `true` | -| Static | `incident_type` | - | `"incident"` | - -**Construction Logic**: -```rust -async fn build_incident_data( - service_health: &ServiceHealthData, - component_id: u32, -) -> IncidentData { - let (timestamp, impact) = service_health.metrics.last().unwrap(); - - let start_date = DateTime::::from_timestamp( - *timestamp - 1, // -1 second per FR-011 - 0 - ).unwrap(); - - IncidentData { - title: "System incident from monitoring system".to_string(), - description: "System-wide incident affecting one or multiple components. Created automatically.".to_string(), - impact: *impact, - components: vec![component_id], - start_date, - system: true, - incident_type: "incident".to_string(), - } -} -``` - -**Validation Rules**: -- `impact`: Must be in range [0, 3] -- `components`: Non-empty array (at least one component ID) -- `start_date`: Valid RFC3339 datetime -- `type`: Must be `"incident"` (not `"maintenance"` or `"info"`) - -**API Response** (on success): -```json -{ - "result": [ - { - "component_id": 218, - "incident_id": 456 // Existing incident ID if duplicate, or new ID - } - ] -} -``` - -**Relationships**: -- **Created from**: `ServiceHealthResponse` (health API) + `Component` (config) + `ComponentCache` (ID lookup) -- **Sent to**: Status Dashboard API `/v2/incidents` -- **Security**: Generic title/description prevent exposing sensitive data (FR-017) - ---- - -### 6. ServiceHealthResponse (Existing, unchanged) - -**Purpose**: Response from the local convertor API `/api/v1/health` containing service health metrics. - -**Rust Definition** (from `src/api/v1.rs`): -```rust -#[derive(Debug, Serialize, Deserialize)] -pub struct ServiceHealthResponse { - pub name: String, // Service name (e.g., "swift") - pub service_category: String, // Category (e.g., "Storage") - pub environment: String, // Environment name (e.g., "production") - pub metrics: ServiceHealthData, // Health data points -} - -pub type ServiceHealthData = Vec<(i64, u8)>; -// Vec of (timestamp_epoch_seconds, impact_0_to_3) -``` - -**Example**: -```json -{ - "name": "swift", - "service_category": "Storage", - "environment": "production", - "metrics": [ - [1706000000, 0], - [1706000060, 0], - [1706000120, 2] // Impact level 2 (major issue) - ] -} -``` - -**Usage in Reporter**: -```rust -// Reporter queries convertor API -let response: ServiceHealthResponse = req_client - .get("http://localhost:8080/api/v1/health") - .query(&[ - ("environment", "production"), - ("service", "swift"), - ("from", "-5min"), - ("to", "-2min") - ]) - .send().await? - .json().await?; - -// Check last metric -if let Some((timestamp, impact)) = response.metrics.last() { - if *impact > 0 { - // Create incident using *impact and *timestamp - } -} -``` - -**Relationships**: -- **Source**: Local convertor API (unchanged by this migration) -- **Consumed by**: Reporter's monitoring loop -- **Used to create**: `IncidentData` when impact > 0 - ---- - -## Data Flow - -### Startup Flow - -``` -┌─────────────────────────────────────────────────────────────────┐ -│ 1. Reporter Startup │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 2. Fetch Components (with retry) │ -│ GET /v2/components → Vec │ -│ Retry: 3 attempts, 60s delay (FR-006) │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 3. Build Component Cache │ -│ ComponentCache = HashMap<(name, attrs), id> │ -│ Sort attributes before inserting │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 4. Start Monitoring Loop │ -│ Every 60 seconds │ -└─────────────────────────────────────────────────────────────────┘ -``` - -### Incident Creation Flow - -``` -┌─────────────────────────────────────────────────────────────────┐ -│ 1. Query Health API │ -│ GET /api/v1/health?env=prod&service=swift │ -│ Response: ServiceHealthResponse │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 2. Check Impact Level │ -│ if metrics.last().impact > 0 { proceed } │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 3. Lookup Component ID │ -│ component = config.get(service_name) │ -│ component_id = cache.find((component.name, component.attrs)) │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ┌─────┴─────┐ - │ │ - Found │ │ Not Found - ▼ ▼ - ┌─────────┐ ┌────────────────────────────────┐ - │ Create │ │ Refresh Cache (1 attempt) │ - │Incident │ │ Retry lookup once (FR-005) │ - └────┬────┘ └──────────┬─────────────────────┘ - │ │ - │ ┌────────┴────────┐ - │ Found │ │ Still Not Found - │ ▼ ▼ - │ ┌─────────────┐ ┌──────────────────┐ - │ │ Create │ │ Log Warning │ - │ │ Incident │ │ Skip incident │ - │ └──────┬──────┘ │ Continue loop │ - │ │ └──────────────────┘ - └─────────┴──────┐ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 4. Build IncidentData │ -│ - title: static │ -│ - description: static │ -│ - impact: from health metric │ -│ - components: [component_id] │ -│ - start_date: timestamp - 1s (RFC3339) │ -│ - system: true │ -│ - type: "incident" │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────────┐ -│ 5. POST to /v2/events │ -│ Authorization: Bearer │ -│ Body: IncidentData (JSON) │ -│ Timeout: 2s │ -└───────────────┬─────────────────────────────────────────────────┘ - │ - ┌─────┴─────┐ - │ │ - Success│ │ Error - ▼ ▼ - ┌─────────┐ ┌───────────────────────────┐ - │Log INFO │ │ Log ERROR (status + body) │ - │Continue │ │ Continue to next service │ - │to next │ │ (retry in next cycle) │ - └─────────┘ └───────────────────────────┘ -``` - ---- - -## Entity Relationships Diagram - -``` -┌──────────────────────────────────────────────────────────────────┐ -│ Configuration │ -│ (config.yaml) │ -└────────────────┬─────────────────────────────────────────────────┘ - │ - │ Defines - ▼ -┌────────────────────────────────┐ ┌───────────────────────┐ -│ Component (Config) │ │ ComponentAttribute │ -├────────────────────────────────┤ ├───────────────────────┤ -│ + name: String │◆─────│ + name: String │ -│ + attributes: Vec │1 * │ + value: String │ -└────────────────┬───────────────┘ └───────────────────────┘ - │ ▲ - │ Lookup │ - ▼ │ Uses -┌────────────────────────────────────────────┐ │ -│ ComponentCache │ │ -│ HashMap<(String, Vec), u32> │ │ -├────────────────────────────────────────────┤ │ -│ Key: (component_name, sorted_attributes) │ │ -│ Value: component_id │ │ -└────────────────┬───────────────────────────┘ │ - ▲ │ - │ Built from │ - │ │ -┌────────────────┴───────────────┐ │ -│ StatusDashboardComponent │ │ -│ (from API) │◆───────────────┘ -├────────────────────────────────┤1 * -│ + id: u32 │ -│ + name: String │ -│ + attributes: Vec │ -└────────────────────────────────┘ - ▲ - │ Fetched from - │ -┌────────┴──────────────────────────────────────────────────┐ -│ Status Dashboard API: GET /v2/components │ -└───────────────────────────────────────────────────────────┘ - - -┌────────────────────────────────┐ -│ ServiceHealthResponse │ -│ (from convertor API) │ -├────────────────────────────────┤ -│ + name: String │ -│ + service_category: String │ -│ + environment: String │ -│ + metrics: Vec<(i64, u8)> │ (timestamp, impact) -└────────────────┬───────────────┘ - │ - │ impact > 0? - ▼ - ┌───────────────┐ - │ Resolve via │ - │ Cache │ - └───────┬───────┘ - │ - ▼ component_id -┌────────────────────────────────┐ -│ IncidentData │ -│ (V2 API request) │ -├────────────────────────────────┤ -│ + title: String │ -│ + description: String │ -│ + impact: u8 │ -│ + components: Vec │──── Resolved from cache -│ + start_date: DateTime │──── From health metric (ts - 1s) -│ + system: bool │──── Always true -│ + incident_type: String │──── Always "incident" -└────────────────┬───────────────┘ - │ - │ POST - ▼ -┌───────────────────────────────────────────────────────────┐ -│ Status Dashboard API: POST /v2/incidents │ -└───────────────────────────────────────────────────────────┘ -``` - ---- - -## State Transitions - -### Component Cache States - -``` -[Uninitialized] - │ - │ Startup: fetch_components_with_retry() - │ Attempts: 3, Delay: 60s - │ - ├── Success ──→ [Loaded] - │ │ - │ │ Monitoring loop - │ │ Cache miss? - │ │ - │ ├── Yes ──→ refresh_cache() (1 attempt) - │ │ │ - │ │ ├── Success ──→ [Loaded] (updated) - │ │ │ - │ │ └── Fail ──→ [Stale] (log warning, continue) - │ │ - │ └── No ──→ [Loaded] (continue) - │ - └── Fail (after 3 retries) ──→ [Failed] (panic, reporter exits) -``` - -### Incident Creation States - -``` -[Monitoring] - │ - │ Query health API - │ - ├── impact = 0 ──→ [No Action] (continue to next service) - │ - └── impact > 0 ──→ [Resolving Component] - │ - │ Lookup component_id in cache - │ - ├── Found ──→ [Creating Incident] - │ │ - │ │ POST /v2/incidents - │ │ - │ ├── Success (200) ──→ [Incident Created] - │ │ │ - │ │ └─→ Log INFO, continue - │ │ - │ └── Fail (4xx/5xx/timeout) ──→ [Error] - │ │ - │ └─→ Log ERROR, continue - │ (retry in next cycle) - │ - └── Not Found ──→ [Refreshing Cache] - │ - │ refresh_cache() (1 attempt) - │ - ├── Found after refresh ──→ [Creating Incident] - │ - └── Still not found ──→ [Component Missing] - │ - └─→ Log WARNING, skip incident -``` - ---- - -## Data Validation - -### Input Validation - -| Entity | Field | Validation | Error Handling | -|--------|-------|------------|----------------| -| `StatusDashboardComponent` | `id` | u32 > 0 | Serde deserialization error → log + skip | -| `StatusDashboardComponent` | `name` | Non-empty string | Serde deserialization error → log + skip | -| `ComponentAttribute` | `name` | Non-empty string | Serde deserialization error → log + skip | -| `ComponentAttribute` | `value` | Non-empty string | Serde deserialization error → log + skip | -| `IncidentData` | `impact` | 0 ≤ u8 ≤ 3 | Assert in code (from health metric, already validated) | -| `IncidentData` | `components` | Non-empty Vec | Assert (only create incident if component_id found) | -| `IncidentData` | `start_date` | Valid timestamp | chrono handles validation; panic if invalid | - -### Output Validation - -| Field | Constraint | Enforcement | -|-------|-----------|-------------| -| `IncidentData.title` | Static string | Hardcoded in code | -| `IncidentData.description` | Static string | Hardcoded in code | -| `IncidentData.system` | Always `true` | Hardcoded in code | -| `IncidentData.incident_type` | Always `"incident"` | Hardcoded in code | -| `IncidentData.start_date` | RFC3339 format | `chrono::DateTime::to_rfc3339()` | - ---- - -## Security Considerations - -### Data Separation (FR-017) - -**Sensitive Data** (logged locally, NEVER sent to API): -- Service name (e.g., "swift") -- Environment name (e.g., "production") -- Component name (e.g., "Object Storage Service") -- Component attributes (e.g., `region=EU-DE`) -- Triggered metric names (e.g., "latency_p95", "error_rate") -- Metric values (e.g., "latency=450ms") - -**Public Data** (sent to Status Dashboard API): -- Static generic title: "System incident from monitoring system" -- Static generic description: "System-wide incident affecting one or multiple components. Created automatically." -- Impact level (integer 0-3, no context) -- Component IDs (integers, no names/attributes) -- Start date (timestamp only, no context) - -**Rationale**: Status Dashboard is public-facing. Exposing service names, metric details, or specific component attributes would reveal internal infrastructure details. - ---- - -## Performance Characteristics - -| Operation | Complexity | Frequency | Optimization | -|-------------------|-----------------|----------------------------------|-----------------------| -| Cache build | O(n log n) | Once at startup + rare refreshes | Acceptable; n ~100 | -| Component lookup | O(n) worst case | Per incident (~1-10/min) | Acceptable for n ~100 | -| Incident creation | O(1) | Per health issue (~1-10/min) | HTTP timeout 2s | -| Health query | O(1) | Every 60s per service | Existing, unchanged | - -**Memory Usage**: -- `ComponentCache`: ~100 entries × ~200 bytes/entry = ~20 KB -- `StatusDashboardComponent` list: ~100 × ~200 bytes = ~20 KB (transient during cache build) -- Negligible compared to reporter's base memory footprint (~10 MB) - ---- - -## Testing Strategy - -### Unit Tests - -1. **Component Cache Building**: - - Test `build_component_id_cache()` with various attribute orders - - Verify attributes are sorted in cache keys - - Test empty attributes list - -2. **Component Matching**: - - Test exact match - - Test subset matching (config has fewer attributes) - - Test no match (different attribute values) - - Test no match (different component name) - -3. **Incident Data Construction**: - - Test timestamp adjustment (-1 second) - - Test RFC3339 formatting - - Test static field values - -### Integration Tests - -1. **Cache Load & Refresh**: - - Mock `/v2/components` endpoint - - Test successful cache load - - Test retry logic (3 attempts, 60s delays) - - Test cache refresh on miss - -2. **Incident Creation**: - - Mock `/v2/incidents` endpoint - - Test successful incident creation - - Test duplicate incident handling (API returns existing ID) - - Test error handling (4xx, 5xx, timeout) - -3. **End-to-End Flow**: - - Mock both convertor and Status Dashboard APIs - - Test full flow: health query → component lookup → incident creation - - Test cache miss → refresh → retry - - Test component not found → skip incident - ---- - -## Summary - -This data model defines 6 core entities for the V2 migration: - -1. **ComponentAttribute**: Key-value pairs qualifying components -2. **Component** (config): Reporter's view of components from config -3. **StatusDashboardComponent**: API's view of components -4. **ComponentCache**: In-memory mapping for efficient lookups -5. **IncidentData**: V2 incident request payload -6. **ServiceHealthResponse**: Existing health data (unchanged) - -Key design decisions: -- **Cache structure**: HashMap with sorted attribute keys for deterministic lookups -- **Subset matching**: Iterate cache to find components where config attrs ⊆ dashboard attrs -- **Static incident fields**: Prevent exposing sensitive operational data on public dashboard -- **Timestamp handling**: RFC3339 with -1 second adjustment per FR-011 - -All entities align with OpenAPI schema and functional requirements (FR-001 through FR-017). diff --git a/specs/003-sd-api-v2-migration/plan.md b/specs/003-sd-api-v2-migration/plan.md deleted file mode 100644 index 4b76441..0000000 --- a/specs/003-sd-api-v2-migration/plan.md +++ /dev/null @@ -1,340 +0,0 @@ -# Implementation Plan: Status Dashboard API V2 Migration - -**Branch**: `003-sd-api-v2-migration` | **Date**: 2025-01-23 | **Spec**: [spec.md](spec.md) -**Input**: Feature specification from `/specs/003-sd-api-v2-migration/spec.md` - -## Summary - -Migrate the `cloudmon-metrics-reporter` from Status Dashboard API V1 (`/v1/component_status`) to V2 (`/v2/events`, `/v2/components`). The migration introduces component ID caching with retry logic, restructures incident data with static title/description for security. Authorization mechanism (HMAC-JWT) remains unchanged. All 17 functional requirements (FR-001 through FR-017) are addressed through a nested HashMap cache structure with subset attribute matching, structured diagnostic logging separate from API payloads, and graceful error handling with automatic recovery. - -## Technical Context - -**Language/Version**: Rust 2021 edition (Cargo.toml: edition = "2021", likely Rust 1.70+) -**Primary Dependencies**: -- `reqwest ~0.11` (HTTP client with rustls-tls, json features) -- `chrono ~0.4` (datetime handling, RFC3339 formatting) -- `serde ~1.0` + `serde_json ~1.0` (JSON serialization) -- `tokio ~1.42` (async runtime with full features) -- `tracing ~0.1` + `tracing-subscriber ~0.3` (structured logging) -- `jwt ~0.16`, `hmac ~0.12`, `sha2 ~0.10` (HMAC-JWT authentication) -- **NEW**: `anyhow ~1.0` (for Result error handling in cache functions) - -**Storage**: In-memory HashMap for component ID cache (~100 components × 200 bytes = ~20KB); no persistent storage - -**Testing**: -- `cargo test` with `#[cfg(test)]` unit tests in source files -- Integration tests in `tests/` directory using `mockito ~1.0`, `tokio-test`, `tower` utilities -- Existing test files: `tests/integration_api.rs`, `tests/integration_health.rs`, `tests/documentation_validation.rs` - -**Target Platform**: Linux server (primary), macOS (development); binary target `cloudmon-metrics-reporter` from `src/bin/reporter.rs` - -**Project Type**: Single Rust project (library + 2 binaries: `cloudmon-metrics-convertor`, `cloudmon-metrics-reporter`) - -**Performance Goals**: -- API response time: <200ms p95 under normal load (100 concurrent requests per Constitution IV) -- Metric conversion: <500ms for datasets up to 1000 data points -- **Reporter-specific**: HTTP timeout 10s (increased from 2s per FR-014), monitoring cycle ~60s - -**Constraints**: -- Memory footprint: <100MB RSS under normal operation (Constitution IV) -- Component cache refresh: 3 retries × 60s delays on startup (FR-006) -- Incident creation: no immediate retry on failure, rely on 60s monitoring cycle (FR-015) - -**Scale/Scope**: -- ~100 components in Status Dashboard -- ~10-20 monitored services per environment -- ~1-10 incidents/minute under normal load -- ~5000 lines of code in reporter binary (incremental change to existing) - -## Constitution Check - -*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.* - -### Principle I: Code Quality Standards - -| Requirement | Status | Compliance Notes | -|-------------|--------|------------------| -| Rust idiomatic practices | ✅ PASS | Uses standard collections (HashMap), serde for serialization, async/await patterns | -| Documentation | ✅ PASS | Existing `reporter.rs` has doc comments; new functions will follow rustdoc conventions | -| Type safety | ✅ PASS | Strong typing for all structs (ComponentAttribute, IncidentData); uses Result for errors | -| Error handling | ✅ PASS | Adding `anyhow::Result` for cache functions; no `unwrap()` in production paths (only in main setup) | -| Code review | ✅ PASS | Feature branch `003-sd-api-v2-migration` will undergo peer review before merge | - -### Principle II: Testing Excellence - -| Requirement | Status | Compliance Notes | -|-------------|--------|------------------| -| Unit test coverage | ✅ PASS | Plan includes unit tests for cache building, component matching, timestamp handling (target 95% per Constitution) | -| Integration tests | ✅ PASS | Using `mockito` to mock `/v2/components` and `/v2/incidents` endpoints; testing full flow | -| Contract testing | ✅ PASS | OpenAPI schema validation against `openapi.yaml`; contracts defined in `specs/003-sd-api-v2-migration/contracts/` | -| Mock external dependencies | ✅ PASS | `mockito ~1.0` already in dev-dependencies for mocking Status Dashboard API | -| Test organization | ✅ PASS | Unit tests in `#[cfg(test)]` modules; integration tests in `tests/reporter_v2_integration.rs` (new file) | - -### Principle III: User Experience Consistency - -| Requirement | Status | Compliance Notes | -|-------------|--------|------------------| -| API consistency | ✅ PASS | Reporter uses `serde_json` for consistent JSON handling; follows existing patterns | -| Configuration interface | ✅ PASS | No config changes required (FR-008); existing YAML config remains compatible | -| Logging standards | ✅ PASS | Uses `tracing` crate for structured logging; FR-017 adds diagnostic fields (timestamp, service, environment, component details, impact, triggered_metrics) | -| CLI consistency | ✅ PASS | Reporter binary interface unchanged; existing `--help` and exit codes maintained | -| Documentation coherence | ⚠️ DEFER | Will update `doc/` after implementation; migration documented in `specs/003-sd-api-v2-migration/quickstart.md` | - -### Principle IV: Performance Requirements - -| Requirement | Status | Compliance Notes | -|---------------------|----------|-------------------------------------------------------------------------------------------------------| -| Resource efficiency | ✅ PASS | Component cache adds ~20KB memory (negligible); no heap allocation in hot paths | -| Async operations | ✅ PASS | All I/O uses async/await with `tokio` runtime (existing pattern maintained) | -| Query optimization | ✅ PASS | Component cache eliminates repeated API calls; cached lookup is O(n) for ~100 components (acceptable) | -| Performance testing | ⚠️ DEFER | Benchmark tests for cache lookup optional (not in hot path); integration tests cover timeout behavior | - -### Development Workflow - -| Requirement | Status | Compliance Notes | -|-------------|--------|------------------| -| Branch strategy | ✅ PASS | Feature branch `003-sd-api-v2-migration` already exists | -| Pre-commit checks | ✅ PASS | `.pre-commit-config.yaml` configured; will run `cargo fmt`, `cargo clippy` | -| CI/CD gates | ✅ PASS | Existing Zuul pipeline runs `cargo build`, `cargo clippy`, `cargo test`, Docker build | -| Review requirements | ✅ PASS | PR will require maintainer approval verifying Constitution compliance | - -**Overall Assessment**: ✅ **PASS** - All critical gates pass. Two items deferred to post-implementation (documentation updates, optional benchmarks) are acceptable per Constitution governance. - -**Re-check After Phase 1**: Will verify test coverage meets 95% target for new code, structured logging includes all FR-017 diagnostic fields. - -## Project Structure - -### Documentation (this feature) - -```text -specs/003-sd-api-v2-migration/ -├── spec.md # Feature specification (17 FRs, 3 user stories) -├── plan.md # This file - implementation plan -├── research.md # Phase 0: Technology decisions, API analysis, cache design -├── data-model.md # Phase 1: Entity definitions, relationships, flows -├── quickstart.md # Phase 1: Step-by-step implementation guide -├── contracts/ # Phase 1: API contract specifications -│ ├── README.md # Contract overview and usage -│ ├── components-api.md # GET /v2/components endpoint spec -│ ├── incidents-api.md # POST /v2/incidents endpoint spec -│ ├── request-examples/ # Sample JSON request payloads -│ │ ├── create-incident-single-component.json -│ │ └── create-incident-multi-component.json -│ └── response-examples/ # Sample JSON response payloads -│ ├── components-list.json -│ └── incident-created.json -└── tasks.md # Phase 2: Generated by /speckit.tasks (NOT by /speckit.plan) -``` - -### Source Code (repository root) - -```text -src/ -├── bin/ -│ ├── convertor.rs # Unchanged - metric conversion binary -│ └── reporter.rs # ✏️ MODIFIED - V2 API migration implementation -├── api/ -│ └── v1.rs # Unchanged - health API (ServiceHealthResponse used by reporter) -├── api.rs # Unchanged - API router -├── common.rs # Unchanged - shared utilities -├── config.rs # Unchanged - config parsing (no config changes needed) -├── graphite.rs # Unchanged - TSDB backend -├── lib.rs # Unchanged - library root -└── types.rs # Unchanged - core type definitions - -tests/ -├── integration_api.rs # Existing - API integration tests -├── integration_health.rs # Existing - health endpoint tests -├── documentation_validation.rs # Existing - doc validation -└── reporter_v2_integration.rs # ✨ NEW - V2 migration integration tests - # Tests: component fetching, cache building, incident creation, error handling - -Cargo.toml # ✏️ MODIFIED - add anyhow ~1.0 dependency -openapi.yaml # Reference - Status Dashboard API V2 contract source -``` - -**Structure Decision**: Single Rust project with library + binaries. This migration only modifies `src/bin/reporter.rs` (adds ~300 lines for cache management and V2 API calls) and adds one new integration test file. No changes to project structure needed - follows existing patterns for binary implementations in `src/bin/` and integration tests in `tests/`. - -**Modified Files**: -1. **`Cargo.toml`**: Add `anyhow = "~1.0"` dependency for Result error handling -2. **`src/bin/reporter.rs`**: - - Add structs: `StatusDashboardComponent`, `IncidentData` - - Update `ComponentAttribute` derives (add `PartialOrd`, `Ord`, `Hash`) - - Add functions: `fetch_components*`, `build_component_id_cache`, `find_component_id`, `build_incident_data`, `create_incident` - - Update `metric_watcher`: load cache at startup, replace V1 endpoint with V2, add cache miss handling - -**New Files**: -3. **`tests/reporter_v2_integration.rs`**: Integration tests using `mockito` to mock Status Dashboard V2 endpoints - -## Complexity Tracking - -*No Constitution violations identified. This section intentionally left empty per template instructions.* - -**Rationale**: All implementation decisions align with Constitution principles: -- **Simple cache structure**: Standard Rust HashMap (no custom abstractions) -- **Subset matching**: O(n) iteration acceptable for n~100 components -- **Error handling**: `anyhow::Result` is idiomatic Rust pattern -- **No new architectural patterns**: Follows existing reporter structure -- **Testing strategy**: Matches project's existing mockito + tokio-test approach - ---- - -## Phase 0: Research (COMPLETED) - -**Status**: ✅ Complete - See [`research.md`](research.md) - -**Key Decisions Made**: -1. **Component Cache**: Nested HashMap with sorted attributes for deterministic keys -2. **V2 Incident Payload**: Static title/description, generic content (security per FR-017) -3. **Error Handling**: 3x retry on startup (FR-006), single refresh on miss (FR-005), no immediate retry on incident creation (FR-015) -4. **Testing**: mockito for API mocking, tokio-test for async tests -5. **Authorization**: HMAC-JWT unchanged (FR-008) -6. **Timestamp Handling**: RFC3339 with -1 second adjustment (FR-011) - -**All Technical Unknowns Resolved**: No "NEEDS CLARIFICATION" items remain. - ---- - -## Phase 1: Design & Contracts (COMPLETED) - -**Status**: ✅ Complete - See design artifacts below - -### 1. Data Model -**File**: [`data-model.md`](data-model.md) - -**Core Entities Defined**: -- `ComponentAttribute`: Key-value pairs with sorting/hashing support -- `Component` (config): Reporter's view from configuration -- `StatusDashboardComponent` (API): Status Dashboard's view from `/v2/components` -- `ComponentCache`: HashMap mapping (name, attrs) → component_id -- `IncidentData`: V2 incident request payload with static security-compliant fields -- `ServiceHealthResponse`: Existing, unchanged health metric structure - -**Key Diagrams**: -- Entity relationship diagram showing data flow -- State transition diagrams for cache and incident creation -- Startup flow: component fetch → cache build → monitoring loop -- Incident creation flow: health query → component lookup → cache refresh on miss → incident POST - -### 2. API Contracts -**Directory**: [`contracts/`](contracts/) - -**Files Created**: -- `contracts/README.md`: Contract overview and usage guide -- `contracts/components-api.md`: GET /v2/components specification with Rust implementation examples -- `contracts/incidents-api.md`: POST /v2/incidents specification with field constraints (FR-002, FR-017) -- `contracts/request-examples/*.json`: Sample incident creation payloads -- `contracts/response-examples/*.json`: Sample API responses (components list, incident created) - -**Validation**: All contracts derived from `/openapi.yaml` (project root, lines 138-270) - -### 3. Quickstart Guide -**File**: [`quickstart.md`](quickstart.md) - -**Contents**: -- Prerequisites (dependencies, Status Dashboard setup, config compatibility) -- Step-by-step implementation (6 steps: structs, fetch, cache, lookup, incident, metric_watcher update) -- Complete code examples with inline comments -- Unit test suite (cache building, component matching) -- Integration test suite (API mocking with mockito) -- Verification procedures (logs to check, expected outputs) -- Troubleshooting guide (common errors and solutions) - -### 4. Agent Context Update -**Status**: ✅ Complete - -**Updated File**: `.github/agents/copilot-instructions.md` - -**Changes**: Added project-specific context about this feature (database: N/A, project type: single) - ---- - -## Phase 2: Task Generation (DEFERRED) - -**Status**: ⏸️ Deferred to `/speckit.tasks` command - -**Rationale**: Per plan template instructions, `tasks.md` is generated by a separate command after Phase 1 design is complete. Implementation tasks will be created from: -- Data model entities → implementation tickets -- API contracts → integration test tickets -- Quickstart steps → development workflow tickets - -**Next Command**: Run `/speckit.tasks` to generate actionable task breakdown with dependency ordering. - ---- - -## Implementation Summary - -### Artifacts Generated - -| Phase | Artifact | Status | Lines | Description | -|-------|----------|--------|-------|-------------| -| 0 | `research.md` | ✅ Complete | 308 | Technology decisions, API analysis, cache design rationale | -| 1 | `data-model.md` | ✅ Complete | 650+ | Entity definitions, relationships, data flows, state machines | -| 1 | `contracts/components-api.md` | ✅ Complete | 250+ | GET /v2/components contract with Rust examples | -| 1 | `contracts/incidents-api.md` | ✅ Complete | 500+ | POST /v2/incidents contract with FR-017 security notes | -| 1 | `contracts/request-examples/` | ✅ Complete | 2 files | JSON request payload samples | -| 1 | `contracts/response-examples/` | ✅ Complete | 2 files | JSON response payload samples | -| 1 | `quickstart.md` | ✅ Complete | 800+ | Step-by-step implementation guide with code | -| 1 | `.github/agents/copilot-instructions.md` | ✅ Updated | N/A | Agent context with project details | - -**Total Documentation**: ~2,500 lines of design artifacts + code examples - -### Key Design Decisions Captured - -1. **Cache Architecture** (research.md § 1): - - Nested HashMap<(String, Vec), u32> - - Sorted attributes for deterministic keys - - Subset matching via iteration (O(n) acceptable for n~100) - -2. **Security Model** (data-model.md § 5, contracts/incidents-api.md § FR-017): - - Sensitive data (service names, environments, metric details) logged locally only - - Public data (generic title/description, impact level, component IDs) sent to API - - Clear separation documented in contracts and quickstart - -3. **Error Resilience** (research.md § 5, data-model.md § "State Transitions"): - - Startup: 3 retries × 60s delays, panic if cache load fails - - Runtime: single cache refresh on miss, log warning if still not found - - Incident creation: log error and continue, rely on next cycle (no immediate retry) - -4. **Testing Strategy** (quickstart.md § "Testing"): - - Unit tests: cache building, subset matching, timestamp handling - - Integration tests: mockito for API mocking, end-to-end flow validation - - Contract tests: OpenAPI schema validation (manual in staging) - -### Constitution Re-Check (Post-Design) - -**Status**: ✅ **PASS** - All principles maintained - -| Principle | Re-Check Result | -|---------------------|--------------------------------------------------------------------------------| -| I. Code Quality | ✅ Rust-idiomatic design, strong typing, proper error handling (anyhow::Result) | -| II. Testing | ✅ Comprehensive unit + integration tests planned (95% coverage target) | -| III. UX Consistency | ✅ Structured logging with FR-017 diagnostic fields, no config changes | -| IV. Performance | ✅ Component cache adds ~20KB, O(n) lookup acceptable | - -**No new violations introduced.** Design maintains project's existing architectural patterns. - ---- - -## Next Steps - -1. **Run** `/speckit.tasks` **command** to generate task breakdown from this plan -2. **Review** generated `tasks.md` for task dependencies and estimation -3. **Begin implementation** following `quickstart.md` step-by-step guide -4. **Reference** `data-model.md` for entity structures during coding -5. **Validate** against `contracts/*.md` during API integration -6. **Execute tests** per `quickstart.md` § Testing section -7. **Update** project documentation in `doc/` after implementation - ---- - -## References - -- **Feature Spec**: [`spec.md`](spec.md) - 17 functional requirements, 3 user stories, edge cases -- **OpenAPI Schema**: `/openapi.yaml` (project root) - Status Dashboard API V2 source of truth -- **Reference Implementation**: `sd_api_v2_migration` branch - working V2 implementation for validation -- **Constitution**: `.specify/memory/constitution.md` - CloudMon Metrics Processor principles -- **Codebase**: - - Current reporter: `src/bin/reporter.rs` - - Health API types: `src/api/v1.rs` (ServiceHealthResponse) - - Test fixtures: `tests/fixtures/` diff --git a/specs/003-sd-api-v2-migration/quickstart.md b/specs/003-sd-api-v2-migration/quickstart.md deleted file mode 100644 index 51f1389..0000000 --- a/specs/003-sd-api-v2-migration/quickstart.md +++ /dev/null @@ -1,721 +0,0 @@ -# Quickstart: Status Dashboard API V2 Migration - -**Feature**: Reporter Migration to Status Dashboard API V2 -**Branch**: `003-sd-api-v2-migration` -**Date**: 2025-01-23 - -## Overview - -This guide provides a quickstart for implementing the Status Dashboard API V2 migration. The migration replaces the V1 component status endpoint with V2 incident creation and adds component ID caching. - -**Key Changes**: -- ✅ Component ID cache at startup (with retry) -- ✅ New incident structure with static title/description -- ✅ Structured diagnostic logging (not sent to API) -- ✅ Authorization unchanged (HMAC-JWT) - ---- - -## Prerequisites - -### 1. Dependencies - -Add `anyhow` crate for error handling: - -```toml -# Cargo.toml -[dependencies] -anyhow = "~1.0" -chrono = "~0.4" # Already present -serde = { version = "~1.0", features = ["derive"] } # Already present -serde_json = "~1.0" # Already present -reqwest = { version = "~0.11", default-features = false, features = ["rustls-tls", "json"] } # Already present -``` - -### 2. Status Dashboard Requirements - -- Status Dashboard must be running with V2 API endpoints available -- All monitored components must be registered in Status Dashboard -- Component names and attributes in config must match Status Dashboard exactly (or be subsets) - -### 3. Configuration - -No configuration changes required. Existing `config.yaml` is compatible: - -```yaml -status_dashboard: - url: "https://status-dashboard.example.com" - secret: "your-hmac-secret" # Optional, for auth - -environments: - - name: production - attributes: - region: "EU-DE" - category: "Storage" - -health_metrics: - swift: - component_name: "Object Storage Service" - # ... other health metric config -``` - ---- - -## Implementation Steps - -### Step 1: Define Data Structures - -The Status Dashboard integration is consolidated in `src/sd.rs` library module. Add/update these structs: - -```rust -// src/sd.rs - Status Dashboard integration module - -use anyhow; -use hmac::{Hmac, Mac}; -use jwt::SignWithKey; -use reqwest::header::HeaderMap; -use serde::{Deserialize, Serialize}; -use sha2::Sha256; -use std::collections::{BTreeMap, HashMap}; - -// Update ComponentAttribute to support sorting and hashing -#[derive(Clone, Deserialize, Serialize, Debug, PartialEq, Eq, Hash, Ord, PartialOrd)] -pub struct ComponentAttribute { - pub name: String, - pub value: String, -} - -// Existing Component struct (no changes needed) -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct Component { - pub name: String, - pub attributes: Vec, -} - -// NEW: API response from GET /v2/components -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct StatusDashboardComponent { - pub id: u32, - pub name: String, - #[serde(default)] - pub attributes: Vec, -} - -// NEW: API request for POST /v2/incidents -#[derive(Clone, Deserialize, Serialize, Debug)] -pub struct IncidentData { - pub title: String, - #[serde(default)] - pub description: String, - pub impact: u8, - pub components: Vec, - pub start_date: DateTime, - #[serde(default)] - pub system: bool, - #[serde(rename = "type")] - pub incident_type: String, -} - -// Component ID cache type -type ComponentCache = HashMap<(String, Vec), u32>; -``` - -### Step 2: Implement Component Fetching - -```rust -/// Fetch components from Status Dashboard API -async fn fetch_components( - req_client: &reqwest::Client, - components_url: &str, -) -> Result> { - let response = req_client.get(components_url).send().await?; - response.error_for_status_ref()?; - let components = response.json::>().await?; - Ok(components) -} - -/// Fetch components with retry logic (3 attempts, 60s delays) -async fn fetch_components_with_retry( - req_client: &reqwest::Client, - components_url: &str, -) -> Option> { - let mut attempts = 0; - loop { - match fetch_components(req_client, components_url).await { - Ok(components) => { - tracing::info!("Successfully fetched {} components.", components.len()); - return Some(components); - } - Err(e) => { - attempts += 1; - tracing::error!("Failed to fetch components (attempt {}/3): {}", attempts, e); - if attempts >= 3 { - tracing::error!("Could not fetch components after 3 attempts. Giving up."); - return None; - } - tracing::info!("Retrying in 60 seconds..."); - sleep(Duration::from_secs(60)).await; - } - } - } -} -``` - -### Step 3: Implement Cache Building - -```rust -/// Build component ID cache from fetched components -fn build_component_id_cache( - components: Vec, -) -> ComponentCache { - components - .into_iter() - .map(|c| { - let mut attrs = c.attributes; - attrs.sort(); // Ensure deterministic cache keys - ((c.name, attrs), c.id) - }) - .collect() -} - -/// Update cache (with optional retry on startup) -async fn update_component_cache( - req_client: &reqwest::Client, - components_url: &str, - with_retry: bool, -) -> Result { - tracing::info!("Updating component cache..."); - - let fetch_future = if with_retry { - fetch_components_with_retry(req_client, components_url).await - } else { - fetch_components(req_client, components_url).await.ok() - }; - - match fetch_future { - Some(components) if !components.is_empty() => { - let cache = build_component_id_cache(components); - tracing::info!("Successfully updated component cache. New size: {}", cache.len()); - Ok(cache) - } - Some(_) => { - anyhow::bail!("Component list from status-dashboard is empty.") - } - None => anyhow::bail!("Failed to fetch component list from status-dashboard."), - } -} -``` - -### Step 4: Implement Component Lookup with Subset Matching - -```rust -/// Find component ID in cache with subset attribute matching -fn find_component_id( - cache: &ComponentCache, - target: &Component, -) -> Option { - // Iterate cache to find matching component - cache.iter() - .filter(|((name, _attrs), _id)| name == &target.name) - .find(|((_name, cache_attrs), _id)| { - // Config attrs must be subset of cache attrs (FR-012) - target.attributes.iter().all(|target_attr| { - cache_attrs.iter().any(|cache_attr| { - cache_attr.name == target_attr.name - && cache_attr.value == target_attr.value - }) - }) - }) - .map(|((_name, _attrs), id)| *id) -} -``` - -### Step 5: Implement Incident Creation - -```rust -/// Build incident data from health metric -fn build_incident_data( - component_id: u32, - impact: u8, - timestamp: i64, -) -> IncidentData { - // Adjust timestamp by -1 second per FR-011 - let start_date = DateTime::::from_timestamp(timestamp - 1, 0) - .expect("Invalid timestamp"); - - IncidentData { - title: "System incident from monitoring system".to_string(), - description: "System-wide incident affecting one or multiple components. Created automatically.".to_string(), - impact, - components: vec![component_id], - start_date, - system: true, - incident_type: "incident".to_string(), - } -} - -/// Create incident via API -async fn create_incident( - req_client: &reqwest::Client, - incidents_url: &str, - headers: &HeaderMap, - incident: &IncidentData, -) -> Result<()> { - let response = req_client - .post(incidents_url) - .headers(headers.clone()) - .json(incident) - .send() - .await?; - - if !response.status().is_success() { - let status = response.status(); - let body = response.text().await?; - tracing::error!("Incident creation failed [{}]: {}", status, body); - return Err(anyhow::anyhow!("API error: {} - {}", status, body)); - } - - tracing::info!("Incident created successfully"); - Ok(()) -} -``` - -### Step 6: Update metric_watcher Function - -Replace the monitoring loop in `metric_watcher()`: - -```rust -async fn metric_watcher(config: &Config) { - tracing::info!("Starting metric reporter thread"); - - let req_client: reqwest::Client = ClientBuilder::new() - .timeout(Duration::from_secs(2)) - .build() - .unwrap(); - - // Build component lookup table from config (unchanged) - let mut components_from_config: HashMap> = HashMap::new(); - for env in config.environments.iter() { - // ... existing component building logic ... - } - - // Status Dashboard configuration - let sdb_config = config - .status_dashboard - .as_ref() - .expect("Status dashboard section is missing"); - - // NEW: V2 endpoints - let components_url = format!("{}/v2/components", sdb_config.url); - let incidents_url = format!("{}/v2/incidents", sdb_config.url); - - // Setup authorization headers (unchanged) - let mut headers = HeaderMap::new(); - if let Some(ref secret) = sdb_config.secret { - let key: Hmac = Hmac::new_from_slice(secret.as_bytes()).unwrap(); - let mut claims = BTreeMap::new(); - claims.insert("stackmon", "dummy"); - let token_str = claims.sign_with_key(&key).unwrap(); - let bearer = format!("Bearer {}", token_str); - headers.insert(AUTHORIZATION, bearer.parse().unwrap()); - } - - // NEW: Load component cache at startup with retry (FR-006, FR-007) - let mut component_cache = update_component_cache(&req_client, &components_url, true) - .await - .expect("Failed to load component cache. Reporter cannot start."); - - tracing::info!("Component cache loaded with {} entries", component_cache.len()); - - // Monitoring loop - loop { - for env in config.environments.iter() { - for (service_name, _component_config) in config.health_metrics.iter() { - // Query health API (unchanged) - match req_client - .get(format!("http://localhost:{}/api/v1/health", config.server.port)) - .query(&[ - ("environment", env.name.clone()), - ("service", service_name.clone()), - ("from", "-5min".to_string()), - ("to", "-2min".to_string()), - ]) - .send() - .await - { - Ok(rsp) => { - if rsp.status().is_client_error() { - tracing::error!("Got API error {:?}", rsp.text().await); - } else { - match rsp.json::().await { - Ok(mut data) => { - if let Some((timestamp, impact)) = data.metrics.pop() { - if impact > 0 { - // Get component from config - let component = components_from_config - .get(&env.name) - .and_then(|env_map| env_map.get(service_name)) - .expect("Component not found in config"); - - // NEW: Look up component ID in cache - let component_id = match find_component_id(&component_cache, component) { - Some(id) => id, - None => { - // Cache miss: refresh and retry (FR-005) - tracing::warn!("Component not found in cache: {} {:?}", component.name, component.attributes); - tracing::info!("Refreshing component cache..."); - - match update_component_cache(&req_client, &components_url, false).await { - Ok(new_cache) => { - component_cache = new_cache; - match find_component_id(&component_cache, component) { - Some(id) => id, - None => { - tracing::warn!("Component still not found after cache refresh: {} {:?}", component.name, component.attributes); - continue; // Skip incident creation - } - } - } - Err(e) => { - tracing::error!("Failed to refresh cache: {}", e); - continue; // Skip incident creation - } - } - } - }; - - // NEW: Log diagnostic details (FR-017) - tracing::info!( - timestamp = timestamp, - service = %service_name, - environment = %env.name, - component_name = %component.name, - component_attrs = ?component.attributes, - component_id = component_id, - impact = impact, - "Creating incident for health issue" - ); - - // NEW: Build and create incident - let incident = build_incident_data(component_id, impact, timestamp); - - match create_incident(&req_client, &incidents_url, &headers, &incident).await { - Ok(_) => { - tracing::info!("Incident reported successfully"); - } - Err(e) => { - tracing::error!("Failed to create incident: {}", e); - // Continue to next service (FR-015) - } - } - } - } - } - Err(e) => { - tracing::error!("Cannot process response: {}", e); - } - } - } - } - Err(e) => { - tracing::error!("Error querying health API: {}", e); - } - } - } - } - - // Sleep between monitoring cycles - sleep(Duration::from_secs(60)).await; - } -} -``` - ---- - -## Testing - -### Unit Tests - -Add to `src/bin/reporter.rs`: - -```rust -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_build_component_id_cache() { - let components = vec![ - StatusDashboardComponent { - id: 218, - name: "Storage".to_string(), - attributes: vec![ - ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }, - ComponentAttribute { name: "category".to_string(), value: "Storage".to_string() }, - ], - }, - ]; - - let cache = build_component_id_cache(components); - - // Attributes should be sorted in cache key - let key = ( - "Storage".to_string(), - vec![ - ComponentAttribute { name: "category".to_string(), value: "Storage".to_string() }, - ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }, - ], - ); - - assert_eq!(cache.get(&key), Some(&218)); - } - - #[test] - fn test_find_component_id_exact_match() { - let mut cache = ComponentCache::new(); - cache.insert( - ( - "Storage".to_string(), - vec![ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }], - ), - 218, - ); - - let component = Component { - name: "Storage".to_string(), - attributes: vec![ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }], - }; - - assert_eq!(find_component_id(&cache, &component), Some(218)); - } - - #[test] - fn test_find_component_id_subset_match() { - let mut cache = ComponentCache::new(); - cache.insert( - ( - "Storage".to_string(), - vec![ - ComponentAttribute { name: "category".to_string(), value: "Storage".to_string() }, - ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }, - ], - ), - 218, - ); - - // Config has only region (subset of cache) - let component = Component { - name: "Storage".to_string(), - attributes: vec![ComponentAttribute { name: "region".to_string(), value: "EU-DE".to_string() }], - }; - - assert_eq!(find_component_id(&cache, &component), Some(218)); - } - - #[test] - fn test_find_component_id_no_match() { - let mut cache = ComponentCache::new(); - cache.insert( - ("Storage".to_string(), vec![]), - 218, - ); - - let component = Component { - name: "Compute".to_string(), - attributes: vec![], - }; - - assert_eq!(find_component_id(&cache, &component), None); - } -} -``` - -### Integration Tests - -Create `tests/reporter_v2_integration.rs`: - -```rust -use mockito::{Mock, Server}; -use cloudmon_metrics::config::Config; - -#[tokio::test] -async fn test_fetch_components_success() { - let mut server = Server::new_async().await; - - let mock = server.mock("GET", "/v2/components") - .with_status(200) - .with_header("content-type", "application/json") - .with_body(r#"[ - { - "id": 218, - "name": "Storage", - "attributes": [{"name": "region", "value": "EU-DE"}] - } - ]"#) - .create_async() - .await; - - let client = reqwest::Client::new(); - let url = format!("{}/v2/components", server.url()); - - let components = fetch_components(&client, &url).await.unwrap(); - - assert_eq!(components.len(), 1); - assert_eq!(components[0].id, 218); - mock.assert_async().await; -} - -#[tokio::test] -async fn test_create_incident_success() { - let mut server = Server::new_async().await; - - let mock = server.mock("POST", "/v2/incidents") - .with_status(200) - .with_header("content-type", "application/json") - .with_body(r#"{"result": [{"component_id": 218, "incident_id": 456}]}"#) - .create_async() - .await; - - let client = reqwest::Client::new(); - let url = format!("{}/v2/incidents", server.url()); - let headers = HeaderMap::new(); - - let incident = IncidentData { - title: "Test".to_string(), - description: "Test".to_string(), - impact: 2, - components: vec![218], - start_date: chrono::Utc::now(), - system: true, - incident_type: "incident".to_string(), - }; - - let result = create_incident(&client, &url, &headers, &incident).await; - - assert!(result.is_ok()); - mock.assert_async().await; -} -``` - -Run tests: -```bash -cargo test -``` - ---- - -## Verification - -### 1. Check Component Cache Loading - -Start the reporter and verify logs: - -```bash -RUST_LOG=info cargo run --bin cloudmon-metrics-reporter -``` - -Expected output: -``` -INFO Updating component cache... -INFO Successfully fetched 100 components. -INFO Successfully updated component cache. New size: 100 -INFO Component cache loaded with 100 entries -INFO Starting metric reporter thread -``` - -### 2. Trigger an Incident - -Create a health issue and check logs: - -``` -INFO Creating incident for health issue timestamp=1706000120 service="swift" environment="production" component_name="Object Storage Service" component_attrs=[ComponentAttribute { name: "region", value: "EU-DE" }] component_id=218 impact=2 -INFO Incident created successfully -INFO Incident reported successfully -``` - -### 3. Verify in Status Dashboard - -Check Status Dashboard UI: -- Incident should appear with title "System incident from monitoring system" -- `system` flag should be true -- Impact level should match health metric -- Component should be correctly associated - -### 4. Test Cache Refresh - -1. Add a new component to Status Dashboard -2. Update config to reference new component -3. Trigger health issue for new component -4. Verify logs show cache refresh: - -``` -WARN Component not found in cache: "New Service" [...] -INFO Refreshing component cache... -INFO Successfully updated component cache. New size: 101 -INFO Creating incident for health issue ... component_id=350 ... -``` - ---- - -## Troubleshooting - -### Issue: "Failed to load component cache. Reporter cannot start." - -**Cause**: Cannot fetch components from Status Dashboard (network error, auth issue, or API unavailable) - -**Solution**: -1. Check Status Dashboard URL in config -2. Verify Status Dashboard is running and `/v2/components` endpoint is accessible -3. Check authentication secret if configured -4. Review logs for specific error messages - -### Issue: "Component not found in cache" (repeated) - -**Cause**: Component name or attributes in config don't match Status Dashboard - -**Solution**: -1. Check component name spelling in config -2. Verify attributes match exactly (or are subset of) Status Dashboard -3. Check Status Dashboard API response: `curl https://status-dashboard/v2/components` -4. Ensure component is registered in Status Dashboard - -### Issue: "Incident creation failed [404]" - -**Cause**: Component ID doesn't exist in Status Dashboard - -**Solution**: -1. Verify component exists: `curl https://status-dashboard/v2/components/{id}` -2. Check cache is up-to-date -3. Manually trigger cache refresh by restarting reporter - -### Issue: "Incident creation failed [400]" - -**Cause**: Invalid incident data (impact out of range, missing required fields, invalid date format) - -**Solution**: -1. Check health metric returns valid impact (0-3) -2. Verify timestamp is valid Unix epoch seconds -3. Review incident payload in error logs -4. Validate against OpenAPI schema in `/openapi.yaml` - ---- - -## Next Steps - -After implementing the migration: - -1. **Update Documentation**: Update project docs in `doc/` to reflect V2 usage -2. **Add Monitoring**: Set up alerts for component cache failures or incident creation errors -3. **Performance Tuning**: Monitor HTTP timeout usage; adjust if needed -4. **Decommission V1**: After validation period, remove V1 endpoint usage (if not needed elsewhere) - ---- - -## Reference - -- **Feature Spec**: `specs/003-sd-api-v2-migration/spec.md` -- **Research**: `specs/003-sd-api-v2-migration/research.md` -- **Data Model**: `specs/003-sd-api-v2-migration/data-model.md` -- **API Contracts**: `specs/003-sd-api-v2-migration/contracts/` -- **OpenAPI Schema**: `/openapi.yaml` -- **Reference Implementation**: `sd_api_v2_migration` branch diff --git a/specs/003-sd-api-v2-migration/research.md b/specs/003-sd-api-v2-migration/research.md deleted file mode 100644 index b8f891a..0000000 --- a/specs/003-sd-api-v2-migration/research.md +++ /dev/null @@ -1,276 +0,0 @@ -# Research: SD API V2 Migration - -**Date**: 2025-01-22 -**Feature**: Reporter Migration to Status Dashboard API V2 -**Branch**: `003-sd-api-v2-migration` - -## Overview - -This document consolidates research findings for migrating the cloudmon-metrics-reporter from Status Dashboard API V1 to V2. All decisions are informed by: -- OpenAPI schema at `/openapi.yaml` -- Feature specification requirements (17 FRs) -- Existing V1 implementation in `src/bin/reporter.rs` -- Reference implementation in branch `sd_api_v2_migration` - ---- - -## 1. Component Cache Design - -### Decision: HashMap> (Nested Hash Maps) - -**Rationale**: -Component resolution requires matching both `name` and `attributes` as a composite key. The cache structure is: -```rust -HashMap< - String, // Component name (e.g., "Object Storage Service") - HashMap // Attributes hash -> Component ID -> -``` - -Where the inner `String` key is a deterministic hash of sorted attributes (e.g., "category=Storage,region=EU-DE"). - -**Why this approach**: -1. **Fast lookup**: O(1) for name, O(1) for attribute hash = O(1) total -2. **Subset matching support**: FR-012 requires matching where configured attributes are a subset of component's attributes. We compute the hash from configured attributes and find matches. -3. **Rust-idiomatic**: Uses standard library `HashMap` with no external dependencies -4. **Memory efficient**: ~10-100 components typical; minimal overhead - -**Alternatives considered**: -- **Option A: Vec with linear search** - O(n) lookup, too slow for 60s monitoring cycles -- **Option B: BTreeMap for sorted iteration** - Unnecessary; lookup order doesn't matter -- **Option C: Custom index struct** - Overengineering for simple cache - -**Implementation notes**: -- Attributes sorted lexicographically before hashing to ensure deterministic keys -- Cache refresh on miss (FR-005) rebuilds entire cache from `/v2/components` GET - ---- - -## 2. V2 Incident Payload Construction - -### Decision: Static struct with serde serialization - -**Rationale**: -V2 incident creation uses a fixed payload structure per OpenAPI schema: - -```rust -#[derive(Serialize)] -struct IncidentPost { - title: String, // Static: "System incident from monitoring system" - description: String, // Static: "System-wide incident affecting one or multiple components. Created automatically." - impact: u8, // From health metric (0-3) - components: Vec, // Resolved component IDs - start_date: String, // RFC3339, from health timestamp - 1s - system: bool, // Always true - #[serde(rename = "type")] - incident_type: String, // Always "incident" -} -``` - -**Why this approach**: -1. **Type safety**: Compile-time validation via Rust structs + serde derive -2. **Security compliance**: FR-002/FR-017 separation - sensitive data in logs, generic data in API -3. **OpenAPI alignment**: Fields match schema exactly (using `#[serde(rename)]` for "type" keyword) -4. **Maintainability**: Single source of truth for payload structure - -**Alternatives considered**: -- **Option A: Manual JSON construction** - Error-prone, no compile-time checks -- **Option B: Dynamic template strings** - Harder to test, type-unsafe -- **Option C: Builder pattern** - Overkill for simple static payload - -**Security implementation**: -Per FR-002 clarifications (Session 2026-01-22): -- **API fields**: Generic static messages (title, description) -- **Local logs**: Detailed diagnostic info (service, environment, component attributes, triggered metrics per FR-017) -- **Separation enforced**: Incident struct does NOT include sensitive fields; logging uses separate context variables - ---- - -## 3. Error Handling: Cache Refresh Scenarios - -### Decision: Retry with exponential backoff for initial load; single retry for cache miss - -**Rationale**: - -**Initial cache load (startup)**: -```rust -// FR-006: Retry up to 3 times with 60s delays -for attempt in 1..=3 { - match fetch_components().await { - Ok(components) => { build_cache(components); break; } - Err(e) if attempt < 3 => { - tracing::warn!("Cache load attempt {}/3 failed: {}", attempt, e); - sleep(Duration::from_secs(60)).await; - } - Err(e) => { - tracing::error!("Failed to load component cache after 3 attempts"); - return Err(e); // FR-007: Fail to start - } - } -} -``` - -**Cache miss during runtime**: -```rust -// FR-005: Refresh on miss, retry lookup once -if cache.get(name, attrs).is_none() { - tracing::info!("Component not found in cache; refreshing"); - refresh_cache().await?; // Single refresh attempt - if cache.get(name, attrs).is_none() { - tracing::warn!("Component {} still not found after refresh", name); - // FR-015: Continue to next service, don't retry incident creation - continue; - } -} -``` - -**Why this approach**: -1. **Startup reliability**: 3 retries with 60s delays handle temporary API unavailability (SC-004: starts within 3min) -2. **Runtime resilience**: Single cache refresh on miss handles new components added to Status Dashboard (FR-005) -3. **No retry on incident creation failure**: Per FR-015, log error and rely on next monitoring cycle (~60s) -4. **Constitution alignment**: Clear error messages (III. User Experience) and async operations (IV. Performance) - -**Alternatives considered**: -- **Option A: Infinite retries** - Blocks startup indefinitely; violates SC-004 -- **Option B: Exponential backoff during runtime** - Delays monitoring cycle; FR-015 says rely on next cycle -- **Option C: Circuit breaker pattern** - Overengineering; simple retry sufficient - -**Error logging**: -Per Constitution III (Logging Standards): -- Include request IDs via tower-http middleware (already configured) -- Log HTTP status codes and response bodies on errors (SC-006) -- Use structured fields: `component_name`, `attributes`, `http_status`, `response_body` - ---- - -## 4. Testing Strategies: Async HTTP with Mockito - -### Decision: mockito 1.0 for HTTP mocking + tokio-test for async assertions - -**Rationale**: -Testing async reporter logic requires: -1. **HTTP mocking**: Simulate `/v2/components` and `/v2/incidents` responses -2. **Async runtime**: Execute tokio futures in tests -3. **Deterministic timing**: Control retry delays for fast tests - -**Test structure**: -```rust -#[cfg(test)] -mod tests { - use super::*; - use mockito::{mock, Mock}; - use tokio_test::block_on; - - #[test] - fn test_cache_load_success() { - let mut server = mockito::Server::new(); - let m = server.mock("GET", "/v2/components") - .with_status(200) - .with_body(r#"[{"id":1,"name":"Service A","attributes":[]}]"#) - .create(); - - let cache = block_on(fetch_and_build_cache(&server.url())); - assert!(cache.get("Service A", &[]).is_some()); - m.assert(); - } - - #[test] - fn test_cache_refresh_on_miss() { - // Mock initial load with component A - // Mock refresh returning component A + B - // Verify lookup finds B after refresh - } - - #[test] - fn test_incident_creation_with_static_description() { - // Mock POST /v2/incidents - // Verify payload contains generic description (not service/env details) - // Verify logs contain diagnostic details (FR-017) - } -} -``` - -**Why this approach**: -1. **mockito 1.0**: Already in dev-dependencies; simple HTTP mock setup -2. **tokio-test**: Lightweight async test utilities; no heavyweight framework needed -3. **Constitution alignment**: II. Testing Excellence - integration tests in `#[cfg(test)]` modules - -**Alternatives considered**: -- **Option A: wiremock crate** - More features but heavier dependency; mockito sufficient -- **Option B: Real HTTP server in tests** - Flaky, slow, requires network -- **Option C: Trait-based mocking** - Overengineering; HTTP layer is the right boundary - -**Test coverage targets**: -Per Constitution II (Unit Test Coverage: 95%): -- Component cache: load, refresh, subset matching (FR-012) -- Incident payload: field values, serde serialization -- Error scenarios: cache failures, HTTP timeouts, malformed responses -- Retry logic: initial load retries, cache refresh - ---- - -## 5. Authorization: HMAC-JWT Token (Unchanged) - -### Decision: Reuse existing V1 authorization mechanism - -**Rationale**: -FR-008 explicitly states "continue using the existing HMAC-signed JWT authorization mechanism without changes." Current V1 code: -```rust -let key: Hmac = Hmac::new_from_slice(secret.as_bytes())?; -let mut claims = BTreeMap::new(); -claims.insert("stackmon", "dummy"); -let token_str = claims.sign_with_key(&key)?; -headers.insert(AUTHORIZATION, format!("Bearer {}", token_str).parse()?); -``` - -**No changes required**: V2 endpoints accept same Authorization header format. - -**Alternatives considered**: None - FR-008 is explicit. - ---- - -## 6. Timestamp Handling: start_date Field - -### Decision: Use health metric timestamp minus 1 second, formatted as RFC3339 - -**Rationale**: -FR-011 specifies: "Use the timestamp from the health metric as the start_date, adjusted by -1 second to align with monitoring intervals." - -Current V1 implementation gets timestamp from: -```rust -let last = data.metrics.pop(); // (timestamp, impact) tuple -// last.0 is the timestamp -``` - -V2 implementation: -```rust -use chrono::{DateTime, Utc, Duration}; - -let timestamp_secs = last.0 as i64; -let dt = DateTime::::from_timestamp(timestamp_secs, 0).unwrap(); -let start_date = (dt - Duration::seconds(1)).to_rfc3339(); -``` - -**Why this approach**: -1. **RFC3339 compliance**: OpenAPI schema specifies `format: date-time` (RFC3339) -2. **chrono crate**: Already in dependencies (v0.4); standard Rust datetime library -3. **-1 second adjustment**: Aligns with monitoring interval logic per FR-011 - -**Alternatives considered**: -- **Option A: Manual RFC3339 formatting** - Error-prone; chrono is reliable -- **Option B: Use timestamp as-is** - Violates FR-011 specification - ---- - -## Summary of Research Findings - -| Topic | Decision | Key Constraint | -|------------------|-----------------------------------------|--------------------------------------| -| Component Cache | Nested HashMap with attribute hash keys | FR-004, FR-012 (subset matching) | -| Incident Payload | Static serde struct with generic fields | FR-002, FR-017 (security separation) | -| Error Handling | 3x retry on startup, 1x refresh on miss | FR-005, FR-006, FR-007, FR-015 | -| Testing | mockito + tokio-test | Constitution II (95% coverage) | -| Authorization | Unchanged HMAC-JWT | FR-008 | -| Timestamps | RFC3339, -1 second adjustment | FR-011 | - -All decisions traceable to specific functional requirements or Constitution principles. No unknowns remaining - proceed to Phase 1 (Design). diff --git a/specs/003-sd-api-v2-migration/spec.md b/specs/003-sd-api-v2-migration/spec.md deleted file mode 100644 index d7b853e..0000000 --- a/specs/003-sd-api-v2-migration/spec.md +++ /dev/null @@ -1,202 +0,0 @@ -# Feature Specification: Reporter Migration to Status Dashboard API V2 - -**Feature Branch**: `003-sd-api-v2-migration` -**Created**: 2025-01-22 -**Status**: Draft -**Input**: User description: "Migrate the reporter from Status Dashboard API V1 to V2 for sending incidents" - -## User Scenarios & Testing *(mandatory)* - -### User Story 1 - Reporter Creates Incidents via V2 API (Priority: P1) - -The reporter monitors service health metrics and automatically creates incidents in the Status Dashboard when issues are detected. After migration, the reporter must successfully create incidents using the new V2 API endpoint while maintaining the same monitoring capabilities. - -**Why this priority**: This is the core functionality of the reporter. Without this working, no incidents can be reported to the Status Dashboard, making the entire monitoring system ineffective. - -**Independent Test**: Can be fully tested by triggering a service health issue (impact value > 0) and verifying that an incident appears in the Status Dashboard with the correct component ID, impact level, and timestamp. Delivers the fundamental value of automated incident reporting. - -**Acceptance Scenarios**: - -1. **Given** the reporter detects a service health issue with impact > 0, **When** it sends an incident to the Status Dashboard, **Then** the incident is created successfully via the `/v2/incidents` endpoint with component ID, title, description, impact, start_date, system flag, and type fields. - -2. **Given** the reporter has a valid component name and attributes from config, **When** it needs to report an incident, **Then** it successfully resolves the component name to a component ID by querying the components cache. - -3. **Given** multiple services are being monitored, **When** issues are detected in different services, **Then** each incident is created with the correct component ID matching the service's component configuration. - ---- - -### User Story 2 - Component Cache Management (Priority: P2) - -The reporter maintains a cache mapping component names and attributes to component IDs to avoid repeated lookups. When a component is not found in the cache, the reporter refreshes the cache from the Status Dashboard API. - -**Why this priority**: This enables efficient operation and handles cases where new components are added to the Status Dashboard after the reporter starts. Without this, the reporter would fail when encountering unknown components. - -**Independent Test**: Can be tested by starting the reporter, adding a new component to the Status Dashboard, triggering an issue for that component, and verifying the reporter refreshes the cache and successfully creates the incident. - -**Acceptance Scenarios**: - -1. **Given** the reporter starts up, **When** initialization occurs, **Then** the reporter fetches all components from `/v2/components` endpoint and builds a component ID cache. - -2. **Given** a component is not found in the cache, **When** the reporter needs to report an incident, **Then** it refreshes the cache from the API and retries the component lookup. - -3. **Given** the initial cache load fails, **When** the reporter starts, **Then** it retries fetching components up to 3 times with 60-second delays before giving up. - ---- - -### User Story 3 - Authorization Remains Unchanged (Priority: P3) - -The reporter continues to use the same authorization mechanism (HMAC-based JWT token) for authenticating with the Status Dashboard API, ensuring no changes to security configuration are required. - -**Why this priority**: Maintaining existing authorization reduces migration complexity and avoids requiring configuration changes or credential updates during the migration. - -**Independent Test**: Can be tested by verifying that the reporter uses the existing secret from config to generate the JWT token and successfully authenticates with the V2 endpoints using the same Authorization header format as V1. - -**Acceptance Scenarios**: - -1. **Given** the reporter has a configured secret, **When** it makes requests to V2 endpoints, **Then** it includes the same HMAC-signed JWT token in the Authorization header as used with V1. - -2. **Given** no secret is configured, **When** the reporter starts, **Then** it operates without authentication headers (for environments without auth requirements). - ---- - -### Edge Cases - -- What happens when the Status Dashboard API is unavailable during initial cache load? - - Reporter should retry up to 3 times with delays, then fail to start with clear error message - -- What happens when a component name exists but with different attributes than configured? - - Reporter should match components where the configured attributes are a subset of the component's attributes - -- What happens when the API returns an error during incident creation? - - Reporter should log the error with the response status and body, continue without retry, and rely on the next monitoring cycle (typically ~5 minutes) to re-attempt incident creation - -- What happens when multiple components match the same name and attributes? - - Reporter should use the first matching component ID found in the cache - -- What happens when the component cache refresh fails? - - Reporter should log a warning, continue using the old cache, and report that the component was not found - -- What happens when the service health response contains no datapoints? - - Reporter should skip incident creation and continue to the next service check - -## Requirements *(mandatory)* - -### Functional Requirements - -- **FR-001**: Reporter MUST send incident data to the `/v2/incidents` endpoint instead of `/v1/component_status` - -- **FR-002**: Reporter MUST use the new incident data structure containing: title (static value "System incident from monitoring system"), description (static value "System-wide incident affecting one or multiple components. Created automatically." - generic text that does not expose sensitive operational data since the Status Dashboard is public), impact (0=none, 1=minor, 2=major, 3=critical, derived directly from service health expression weight), components (array of component IDs), start_date, system flag, and type - -- **FR-003**: Reporter MUST fetch components from `/v2/components` endpoint at startup and build a cache mapping (component name, attributes) to component ID - -- **FR-004**: Reporter MUST resolve component names to component IDs using the cache before creating incidents - -- **FR-005**: Reporter MUST refresh the component cache when a component is not found and retry the lookup once - -- **FR-006**: Reporter MUST retry the initial component cache load up to 3 times with 60-second delays between attempts - -- **FR-007**: Reporter MUST fail to start if the initial component cache load fails after all retry attempts - -- **FR-008**: Reporter MUST continue using the existing HMAC-signed JWT authorization mechanism without changes - -- **FR-009**: Reporter MUST include the system flag set to true in incident data to indicate automatic creation - -- **FR-010**: Reporter MUST set the incident type to "incident" for all automatically created incidents - -- **FR-011**: Reporter MUST use the timestamp from the health metric as the start_date, adjusted by -1 second to align with monitoring intervals - -- **FR-012**: Reporter MUST match components where the configured attributes are a subset of the component's attributes in the Status Dashboard - -- **FR-013**: Reporter MUST log comprehensive incident information including timestamp, status, service, environment, component details, and triggered metrics - -~~- **FR-014**: Reporter MUST increase the HTTP timeout from 2 seconds to 10 seconds to accommodate the new endpoint's response times~~ - -- **FR-015**: Reporter MUST continue monitoring other services even if incident creation fails for one service, logging the error without immediate retry and allowing the next monitoring cycle to re-attempt - -- **FR-016**: Reporter MUST create a new incident request for every service health issue detection, relying on the Status Dashboard's built-in duplicate handling to return existing incidents when applicable - -### Logging Requirements - -- **FR-017**: Reporter MUST log structured diagnostic details for incident investigation containing: detection timestamp, service name, environment name, component name and attributes, impact value, and a list of all triggered metric names with values that contributed to the earliest health issue detection. These details MUST NOT be included in API requests to prevent exposing sensitive operational data on the public Status Dashboard. - -### Key Entities - -- **Incident (V2)**: Represents an incident in the Status Dashboard V2 API. - - **API Fields**: title (string, static value "System incident from monitoring system"), description (string, static value "System-wide incident affecting one or multiple components. Created automatically."), impact (integer 0-3 where 0=none, 1=minor, 2=major, 3=critical, derived directly from service health expression weight), components (array of component IDs), start_date (RFC3339 datetime), system (boolean, always true), type (enum: "incident", "maintenance", "info", always "incident" for auto-created). - - **Security Note**: Description uses a generic message to prevent exposing sensitive operational data (timestamps, service names, environments, component details, impact values, triggered metrics) on the public Status Dashboard. - -- **Component (V2)**: Represents a component in Status Dashboard with fields: id (integer), name (string), attributes (array of name-value pairs). Used to resolve component names to IDs. - -- **Component Cache**: In-memory mapping from (component name, sorted attributes) to component ID, used to avoid repeated API calls for component resolution - -- **Service Health Point**: Enhanced health metric data containing: timestamp, impact value, list of triggered metric names, and optional metric value for detailed logging - -### Operational Logging (Not Sent to API) - -Diagnostic details for incident investigation are logged locally and MUST NOT be included in API requests: -- Detection timestamp -- Service name and environment name -- Component name and attributes -- Impact value (0-3) -- List of triggered metric names with their values - -## Success Criteria *(mandatory)* - -### Measurable Outcomes - -- **SC-001**: Reporter successfully creates incidents in the Status Dashboard using the V2 API within 10 seconds of detecting a service health issue - -- **SC-002**: Reporter resolves component names to IDs without errors for 100% of configured components that exist in the Status Dashboard - -- **SC-003**: Reporter automatically recovers from missing component errors by refreshing the cache within one monitoring cycle (approximately 5 minutes) - -- **SC-004**: Reporter starts successfully within 3 minutes even when the Status Dashboard API is slow, thanks to retry logic - -- **SC-005**: All automatically created incidents are correctly tagged with system=true and type="incident" in the Status Dashboard - -- **SC-006**: Reporter logs provide sufficient information to troubleshoot incident creation failures, including component names, attributes, and API responses - -## Dependencies - -- **Status Dashboard API V2**: The Status Dashboard must have the `/v2/incidents` and `/v2/components` endpoints available and functional -- **Backward Compatibility**: The migration does not require changes to the reporter configuration file format or authorization mechanism -- **Component Registration**: All monitored components must be registered in the Status Dashboard with matching names and attributes - -## Assumptions - -- The Status Dashboard API V2 is stable and ready for production use -- Component IDs in the Status Dashboard are stable and do not change frequently -- The authorization mechanism (HMAC-signed JWT) is compatible with both V1 and V2 endpoints -- The reporter's monitoring logic and configuration structure remain unchanged -- The Status Dashboard will accept incidents with system=true flag for automatically generated incidents -- Component matching logic (subset attribute matching) is sufficient for all use cases -- The 10-second HTTP timeout is sufficient for the V2 API response times under normal operation - -## Clarifications - -### Session 2025-01-22 - -- Q: How should the reporter handle duplicate incident detection events within the same monitoring cycle? → A: Option A - Create incidents on every detection. The Status Dashboard ignores duplicate requests and returns the existing event, so no client-side deduplication needed. -- Q: What should the error recovery strategy be when incident creation fails? → A: Option B - Log the error and continue without retry, rely on next monitoring cycle. -- Q: How should the reporter map service health impact values to incident impact values? → A: Option B - Use service health "impact" field directly (0=none, 1=minor, 2=major, 3=critical). The current V1 implementation already passes the health expression weight directly as the impact value. -- Q: What format should the incident title use? → A: Use a generic static title "System incident from monitoring system" (as implemented in sd_api_v2_migration branch). - -### Session 2026-01-22 - -- Q: What content should be included in the incident description field? → A: Use a static generic message "System-wide incident affecting one or multiple components. Created automatically." This provides context without exposing sensitive operational data on the public Status Dashboard. -- Q: Should FR-017 (diagnostic logging) be kept? → A: Yes, FR-017 is required for incident investigations. Diagnostic details MUST be logged locally but MUST NOT be sent to the API. -- Q: How should API fields be separated from logging requirements? → A: Incident (V2) entity now clearly separates API Fields from Operational Logging requirements. API receives generic non-sensitive data; logs contain full diagnostic details for operators. -- Q: What exact wording should the description field use? → A: "System-wide incident affecting one or multiple components. Created automatically." - using "one or multiple" (not just "multiple") to accurately describe that incidents can affect a single component or several. - -## Out of Scope - -- Changes to the monitoring logic or health metric evaluation -- Modifications to the reporter configuration file format -- Updates to the authorization mechanism or secret management -- Migration of existing V1 incidents to V2 format -- Support for additional incident types beyond "incident" (e.g., "maintenance", "info") -- Batch incident creation or update operations -- Incident updates or closure operations (only creation is in scope) -- Changes to the component attribute configuration format -- Performance optimizations beyond the timeout adjustment -- Automatic component creation in the Status Dashboard if not found diff --git a/specs/003-sd-api-v2-migration/tasks.md b/specs/003-sd-api-v2-migration/tasks.md deleted file mode 100644 index 960a7b3..0000000 --- a/specs/003-sd-api-v2-migration/tasks.md +++ /dev/null @@ -1,370 +0,0 @@ -# Tasks: Status Dashboard API V2 Migration - -**Input**: Design documents from `/specs/003-sd-api-v2-migration/` -**Prerequisites**: plan.md (tech stack), spec.md (3 user stories: P1, P2, P3), research.md (cache design), data-model.md (entities), contracts/ (API endpoints) - -**Tests**: Not requested in feature specification - focusing on implementation tasks - -**Organization**: Tasks are grouped by user story to enable independent implementation and testing of each story. - -## Format: `- [ ] [ID] [P?] [Story] Description` - -- **Checkbox**: ALWAYS start with `- [ ]` -- **[ID]**: Task ID (T001, T002, etc.) -- **[P]**: Can run in parallel (different files, no dependencies) -- **[Story]**: Which user story this task belongs to (US1, US2, US3) -- Include exact file paths in descriptions - -## Path Conventions - -**Single Rust project** at repository root: -- `src/bin/reporter.rs` - main reporter implementation -- `tests/reporter_v2_integration.rs` - integration tests -- `Cargo.toml` - dependency management - ---- - -## Phase 1: Setup (Shared Infrastructure) - -**Purpose**: Project initialisation and dependency updates - -- [x] T001 Add anyhow ~1.0 dependency to Cargo.toml for Result error handling -~~- [x] T002 [P] Update reqwest client timeout from 2s to 10s in src/bin/reporter.rs per FR-014~~ - ---- - -## Phase 2: Foundational (Blocking Prerequisites) - -**Purpose**: Core data structures and utilities that ALL user stories depend on - -**⚠️ CRITICAL**: No user story work can begin until this phase is complete - -- [x] T003 [P] Add StatusDashboardComponent struct in src/bin/reporter.rs for V2 API response -- [x] T004 [P] Update ComponentAttribute with PartialOrd, Ord, Hash derives in src/bin/reporter.rs -- [x] T005 [P] Add IncidentData struct in src/bin/reporter.rs for V2 incident payload -- [x] T006 Create ComponentCache type alias HashMap> in src/bin/reporter.rs - -**Checkpoint**: Foundation ready - all user stories can now proceed - ---- - -## Phase 3: User Story 1 - Reporter Creates Incidents via V2 API (Priority: P1) 🎯 MVP - -**Goal**: Enable reporter to create incidents using the new V2 API endpoint while maintaining monitoring capabilities - -**Independent Test**: Trigger a service health issue (impact > 0) and verify incident appears in Status Dashboard with correct component ID, impact, and timestamp - -**FR Coverage**: FR-001, FR-002, FR-009, FR-010, FR-011, FR-013, FR-016, FR-017 - -### Implementation for User Story 1 - -- [x] T007 [P] [US1] Implement fetch_components() async function in src/bin/reporter.rs to call GET /v2/components -- [x] T008 [P] [US1] Implement build_component_id_cache() function in src/bin/reporter.rs to construct nested HashMap -- [x] T009 [US1] Implement find_component_id() function in src/bin/reporter.rs with subset attribute matching per FR-012 -- [x] T010 [US1] Implement build_incident_data() function in src/bin/reporter.rs with static title/description per FR-002 -- [x] T011 [US1] Add timestamp handling with RFC3339 format and -1 second adjustment in build_incident_data() per FR-011 -- [x] T012 [US1] Implement create_incident() async function in src/bin/reporter.rs to POST /v2/incidents -- [x] T013 [US1] Update metric_watcher() to replace V1 endpoint (/v1/component_status) with V2 incident creation -- [x] T014 [US1] Add structured logging with diagnostic fields (timestamp, service, environment, component details, impact) per FR-017 -- [x] T015 [US1] Add error logging for incident creation failures with status and response body per FR-015 - -**Checkpoint**: User Story 1 complete - reporter can create incidents via V2 API - ---- - -## Phase 4: User Story 2 - Component Cache Management (Priority: P2) - -**Goal**: Maintain cache mapping component names to IDs with automatic refresh when components not found - -**Independent Test**: Start reporter, add new component to Status Dashboard, trigger issue for that component, verify reporter refreshes cache and creates incident - -**FR Coverage**: FR-003, FR-004, FR-005, FR-012 - -### Implementation for User Story 2 - -- [x] T016 [US2] Add component cache initialization in metric_watcher() to fetch and build cache at startup -- [x] T017 [US2] Implement cache miss detection in metric_watcher() when component not found during lookup -- [x] T018 [US2] Implement single cache refresh attempt (call fetch_components + rebuild cache) on cache miss per FR-005 -- [x] T019 [US2] Add warning logging when component still not found after cache refresh per FR-015 -- [x] T020 [US2] Add continue to next service logic when component cannot be resolved (no retry on incident creation) - -**Checkpoint**: User Story 2 complete - cache management with automatic refresh working - ---- - -## Phase 5: User Story 3 - Authorization Remains Unchanged (Priority: P3) - -**Goal**: Verify existing HMAC-JWT authorization works with V2 endpoints without any changes - -**Independent Test**: Verify reporter uses existing secret to generate JWT token and successfully authenticates with V2 endpoints - -**FR Coverage**: FR-008 - -### Implementation for User Story 3 - -- [x] T021 [US3] Verify existing HMAC-JWT token generation in metric_watcher() is reused for V2 endpoints - - ✅ VERIFIED: Token generation at lines 268-274 uses same HMAC-SHA256 algorithm - - ✅ VERIFIED: Same `headers` variable passed to both `fetch_components()` and `create_incident()` - -- [x] T022 [US3] Verify Authorization header format remains unchanged (Bearer {jwt-token}) for V2 API calls - - ✅ VERIFIED: Line 274 uses same format: `format!("Bearer {}", token_str)` - - ✅ VERIFIED: Header inserted with same `AUTHORIZATION` constant - -- [x] T023 [US3] Test that reporter operates without auth headers when no secret configured (optional auth) - - ✅ VERIFIED: Line 268 guards with `if let Some(ref secret) = sdb_config.secret` - - ✅ VERIFIED: Empty `headers` passed to V2 endpoints when no secret configured - -**Checkpoint**: User Story 3 complete - authorization verified unchanged for V2 - -**Verification Summary**: -- ✅ No code changes required - existing auth mechanism works with V2 endpoints -- ✅ HMAC-JWT token generation unchanged (same algorithm, same claims) -- ✅ Authorization header format unchanged (Bearer token) -- ✅ Optional authentication supported (works with or without secret) -- ✅ Headers reused for both GET /v2/components and POST /v2/incidents - ---- - -## Phase 6: Startup Reliability & Error Handling - -**Goal**: Add robust error handling for startup cache loading with retry logic - -**FR Coverage**: FR-006, FR-007 - -- [x] T024 Add initial component cache load with 3 retry attempts in metric_watcher() per FR-006 -- [x] T025 Add 60-second delay between cache load retry attempts per FR-006 -- [x] T026 Add error return from metric_watcher() if cache load fails after 3 attempts per FR-007 -- [x] T027 Add warning logging for each failed cache load attempt with attempt number - -**Checkpoint**: Startup reliability complete - reporter handles API unavailability - -**Implementation Summary**: -- ✅ Retry loop with 1-3 attempts implemented (lines 285-323) -- ✅ 60-second delay using `sleep(Duration::from_secs(60))` between attempts -- ✅ Reporter exits via `return` if all attempts fail (FR-007) -- ✅ Structured logging with attempt number, max_attempts, retry_delay_seconds -- ✅ Info log on success, warning log on retry, error log on final failure -- ✅ Cache initialization broken into Option with unwrap after loop - ---- - -## Phase 7: Integration Testing - -**Purpose**: Validate end-to-end V2 migration with mocked API endpoints - -- [x] T028 [P] Create tests/reporter_v2_integration.rs test file with mockito setup -- [x] T029 [P] Add test_fetch_components_success() to verify component fetching and parsing -- [x] T030 [P] Add test_build_component_id_cache() to verify cache structure with nested HashMap -- [x] T031 [P] Add test_find_component_id_subset_matching() to verify FR-012 subset attribute matching -- [x] T032 [P] Add test_build_incident_data_structure() to verify static title/description per FR-002 -- [x] T033 [P] Add test_timestamp_rfc3339_minus_one_second() to verify FR-011 timestamp handling -- [x] T034 [P] Add test_create_incident_success() to verify POST /v2/incidents with mockito -- [x] T035 [P] Add test_cache_refresh_on_miss() to verify FR-005 single refresh attempt -- [x] T036 [P] Add test_startup_retry_logic() to verify FR-006 3 retry attempts with delays -- [x] T037 [P] Add test_error_logging_with_diagnostic_fields() to verify FR-017 structured logging - -**Checkpoint**: Integration tests complete - all V2 functionality validated - -**Implementation Summary**: -- ✅ Created src/sd.rs library module with all Status Dashboard functions -- ✅ Added sd module to lib.rs exports -- ✅ Created comprehensive test file tests/integration_sd.rs with 13 tests -- ✅ Test coverage: fetch_components, build_cache, subset matching, incident data, timestamps, retries, auth -- ✅ Additional tests: empty attributes handling, multiple components with same name -- ✅ All test logic implemented and ready for execution - ---- - -## Phase 8: Polish & Cross-Cutting Concerns - -**Purpose**: Code quality, documentation, and final validation - -- [x] T038 [P] Run cargo fmt to format all code changes -- [x] T039 [P] Run cargo clippy to check for lints and warnings -- [x] T040 Run cargo test to execute all tests including new integration tests -- [x] T041 Run cargo build to verify compilation without errors -- [x] T042 [P] Update comments and doc strings in src/bin/reporter.rs for new functions -- [x] T043 Verify quickstart.md steps match actual implementation -- [x] T044 [P] Add inline comments explaining cache structure and subset matching logic -- [x] T045 Review all error messages for clarity and actionability per Constitution III - -**Checkpoint**: Feature ready for code review and deployment - -**Implementation Summary**: -- ✅ cargo fmt applied to all files -- ✅ cargo clippy - all 24 warnings fixed, lint passing - - Fixed: redundant field names, needless returns, useless conversions, needless borrows - - Fixed: redundant pattern matching, field reassignment, clone on copy, single match - - Fixed: unnecessary casts, unwrap_or_default -- ✅ All tests passing: 27 tests (5 lib + 7 docs + 5 API + 3 health + 12 SD tests) -- ✅ Build successful without errors -- ✅ Documentation and comments updated throughout sd module -- ✅ Inline comments added for cache structure and subset matching -- ✅ Error messages use structured logging with diagnostic fields - ---- - -## Dependencies & Execution Order - -### Phase Dependencies - -- **Setup (Phase 1)**: No dependencies - can start immediately -- **Foundational (Phase 2)**: Depends on T001, T002 - BLOCKS all user stories -- **User Story 1 (Phase 3)**: Depends on Foundational (T003-T006) complete -- **User Story 2 (Phase 4)**: Depends on User Story 1 (T007-T015) complete -- **User Story 3 (Phase 5)**: Depends on User Story 1 (T007-T015) complete (verification only) -- **Startup Reliability (Phase 6)**: Depends on User Story 2 (T016-T020) complete -- **Integration Testing (Phase 7)**: Depends on all implementation phases (T003-T027) complete -- **Polish (Phase 8)**: Depends on all previous phases complete - -### User Story Dependencies - -- **User Story 1 (P1)**: Foundation → Core V2 incident creation (REQUIRED for MVP) -- **User Story 2 (P2)**: User Story 1 → Add cache refresh logic (enhances US1) -- **User Story 3 (P3)**: User Story 1 → Verification of auth (depends on US1 endpoints) - -### Within Each User Story - -**User Story 1**: -- T007, T008 can run in parallel (different functions) -- T009 depends on T008 (uses cache structure) -- T010, T011 can run after T009 (needs component resolution) -- T012 depends on T010 (uses IncidentData struct) -- T013 depends on T007-T012 (integrates all functions) -- T014, T015 can run in parallel with T013 (logging is separate) - -**User Story 2**: -- T016-T020 are sequential (modify metric_watcher flow) - -**User Story 3**: -- T021-T023 are verification tasks (can run in parallel) - -**Integration Testing**: -- All tests (T028-T037) marked [P] can run in parallel - -### Parallel Opportunities - -- **Phase 1 Setup**: T001, T002 can run in parallel (different concerns) -- **Phase 2 Foundational**: T003, T004, T005 can run in parallel (different structs) -- **Phase 3 User Story 1**: T007, T008 can run in parallel initially -- **Phase 7 Integration Testing**: T029-T037 all marked [P] can run simultaneously -- **Phase 8 Polish**: T038, T039, T042, T044, T045 marked [P] can run simultaneously - ---- - -## Parallel Example: Foundational Phase - -```bash -# Launch all struct definitions together: -Task T003: "Add StatusDashboardComponent struct in src/bin/reporter.rs" -Task T004: "Update ComponentAttribute derives in src/bin/reporter.rs" -Task T005: "Add IncidentData struct in src/bin/reporter.rs" -``` - -## Parallel Example: Integration Testing - -```bash -# Launch all integration tests together: -Task T029: "Add test_fetch_components_success()" -Task T030: "Add test_build_component_id_cache()" -Task T031: "Add test_find_component_id_subset_matching()" -Task T032: "Add test_build_incident_data_structure()" -Task T033: "Add test_timestamp_rfc3339_minus_one_second()" -# ... and so on for all test tasks -``` - ---- - -## Implementation Strategy - -### MVP First (User Story 1 Only) - -1. Complete Phase 1: Setup (T001-T002) -2. Complete Phase 2: Foundational (T003-T006) - CRITICAL -3. Complete Phase 3: User Story 1 (T007-T015) -4. **STOP and VALIDATE**: Test incident creation via V2 API manually -5. MVP READY: Reporter can create incidents using V2 endpoint - -**Estimated Tasks for MVP**: 17 tasks (T001-T015 + T001-T002 foundational) - -### Incremental Delivery - -1. MVP (US1) → Deploy/Demo → Reporter creates V2 incidents ✅ -2. Add US2 (T016-T020) → Deploy/Demo → Cache refresh on miss ✅ -3. Add US3 (T021-T023) → Deploy/Demo → Auth verified ✅ -4. Add Startup Reliability (T024-T027) → Deploy/Demo → Robust startup ✅ -5. Add Testing (T028-T037) → Full test coverage ✅ -6. Polish (T038-T045) → Production ready ✅ - -### Critical Path - -**Blocking sequence** (cannot parallelize): -1. T001-T002 (Setup) → T003-T006 (Foundation) → T009 (component lookup) → T012 (incident creation) → T013 (integration) → T024-T027 (startup reliability) - -**Total Critical Path**: ~13 tasks that MUST be sequential - -**Parallelizable**: ~32 tasks that can run in parallel (all [P] marked tasks) - ---- - -## Task Summary - -| Phase | Task Count | Parallelizable | User Story | -|-------|------------|----------------|------------| -| Phase 1: Setup | 2 | 1 | N/A | -| Phase 2: Foundational | 4 | 3 | N/A | -| Phase 3: User Story 1 | 9 | 2 | P1 (MVP) | -| Phase 4: User Story 2 | 5 | 0 | P2 | -| Phase 5: User Story 3 | 3 | 3 | P3 | -| Phase 6: Startup Reliability | 4 | 0 | N/A | -| Phase 7: Integration Testing | 10 | 10 | N/A | -| Phase 8: Polish | 8 | 4 | N/A | -| **TOTAL** | **45** | **23** | **3 stories** | - -### Task Distribution by User Story - -- **User Story 1 (P1)**: 9 tasks - Core V2 incident creation 🎯 MVP -- **User Story 2 (P2)**: 5 tasks - Cache management with refresh -- **User Story 3 (P3)**: 3 tasks - Authorization verification -- **Infrastructure**: 18 tasks - Setup, foundation, testing, polish -- **Parallel Opportunities**: 23 tasks (51%) can run simultaneously - -### Independent Test Criteria - -**User Story 1**: Manually trigger service health issue with impact > 0 → verify incident created in Status Dashboard → check component ID, impact level, timestamp, system=true flag - -**User Story 2**: Start reporter → add new component in Status Dashboard → trigger issue for new component → verify logs show cache refresh → verify incident created successfully - -**User Story 3**: Review code that no auth changes made → verify JWT token generation unchanged → verify Authorization header format unchanged → test with/without secret configuration - ---- - -## Suggested MVP Scope - -**MVP = User Story 1 only** (17 tasks: T001-T015) - -Delivers core value: -✅ Reporter creates incidents via V2 API -✅ Component ID resolution from cache -✅ Static secure incident payloads -✅ Structured diagnostic logging -✅ Error handling for incident creation - -Not in MVP (can add later): -⏸️ Automatic cache refresh on miss (US2) -⏸️ Auth verification tasks (US3) -⏸️ Startup retry logic (Phase 6) -⏸️ Integration tests (Phase 7) - -**Rationale**: US1 provides immediate business value - reporter works with V2 API. US2/US3 are enhancements that can be added incrementally. - ---- - -## Notes - -- All tasks follow checklist format: `- [ ] [ID] [P?] [Story?] Description with file path` -- [P] tasks target different files or independent functions -- [Story] labels (US1, US2, US3) map to spec.md priorities (P1, P2, P3) -- Each user story independently testable per acceptance scenarios in spec.md -- Constitution compliance: Rust idioms, anyhow::Result, structured logging, 95% test coverage target -- Reference implementations: quickstart.md (step-by-step guide), contracts/ (API specs)