From a24ad3833bb224c5eb02a68adf04493066041d9a Mon Sep 17 00:00:00 2001 From: Wayland Date: Sun, 4 Oct 2026 20:07:55 +0200 Subject: [PATCH 1/6] docs(decisions): ADR-0023 - timed Parquet snapshots of the live database Production has no recovery point: the 2026-10-04 recovery left three volumes behind, two of which hold the same damaged file and the third is production itself. The surviving September copy stops being data around 2026-10-27, when its newest span falls out of the raw retention window. ADR-0023 picks the mechanism and records the numbers it was picked on, measured on robmini against a probe copy of the live 152.6 MB database: a full EXPORT DATABASE to Parquet+ZSTD costs 0.15-0.3 s of connection time and 4.2 MiB, and IMPORT DATABASE restores it in under a second with every row count, the schema version and all four indexes matching the source. Numbered 0023 because the decision index moved twice while the proposal was in review; the earlier drafts claimed 0019 and 0020, both now taken on main. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- docs/.vitepress/config.js | 1 + .../0023-production-database-snapshots.md | 207 ++++++++++++++++++ docs/decisions/index.md | 1 + 3 files changed, 209 insertions(+) create mode 100644 docs/decisions/0023-production-database-snapshots.md diff --git a/docs/.vitepress/config.js b/docs/.vitepress/config.js index b4c42a7..6ce3bc4 100644 --- a/docs/.vitepress/config.js +++ b/docs/.vitepress/config.js @@ -62,6 +62,7 @@ export default defineConfig({ { text: 'ADR-0020 — Recovery Arrives as a New Issue (superseded)', link: '/decisions/0020-recovery-arrives-as-a-new-issue' }, { text: "ADR-0021 — Recovery Wakes the Alert's Assignee", link: '/decisions/0021-recovery-wakes-the-alerts-assignee' }, { text: 'ADR-0022 — Health Probe Scheduler Outside This Repo', link: '/decisions/0022-health-probe-scheduler-outside-github' }, + { text: 'ADR-0023 — Production Database Snapshots', link: '/decisions/0023-production-database-snapshots' }, ], }, ], diff --git a/docs/decisions/0023-production-database-snapshots.md b/docs/decisions/0023-production-database-snapshots.md new file mode 100644 index 0000000..9a50a24 --- /dev/null +++ b/docs/decisions/0023-production-database-snapshots.md @@ -0,0 +1,207 @@ +# ADR 0023 - Snapshots: `EXPORT DATABASE` to Parquet, on a timer, from inside cotel + +**Date:** 2026-10-04 +**Status:** Accepted +**Deciders:** Daedalus (CTO) + +--- + +## Context + +Production has no backup of the live database. The 2026-10-04 recovery left three +volumes behind, which look like redundancy and are not: two of them hold the same +damaged file (`sha256` `7eb82bbe...` on both) and the third is production itself. +See [What a recovery leaves behind](../operations/duckdb-recovery#what-a-recovery-leaves-behind-and-when-to-delete-it). + +The gap has a date on it. The damaged original's newest span is 2026-09-27, and raw +spans are purged at `COTEL_RETENTION_RAW_DAYS` (30), so from roughly **2026-10-27** +re-repairing it recovers nothing retention would have kept anyway. After that there +is no recovery point for production at all. + +Two properties of the running system shape every option: + +1. **DuckDB has one writer, and the live process holds the file lock.** An external + process cannot open `/data/cotel.duckdb`, not even read-only, while cotel runs. + So a snapshot is either taken by cotel itself, or taken while cotel is stopped. +2. **cotel serves everything through a single connection.** `storage.Open` sets + `rw.SetMaxOpenConns(1)` because DuckDB has one writer, and `DB.ReadOnly()` hands + the dashboard that same pool so reads always see WAL-buffered writes. Whatever a + snapshot does on that connection, it does with ingest and the dashboard queued + behind it. **The decisive number is therefore how long the statement holds the + connection, not how big its output is.** + +--- + +## Options considered + +| Option | Who takes it | Ingest cost | Format | Verdict | +|---|---|---|---|---| +| **1. `EXPORT DATABASE` on a timer inside cotel** | cotel | **0.15-0.3 s per run, measured** | Parquet + `schema.sql`, portable | **Chosen** | +| 2. Graceful stop, copy the volume | an operator, by hand | full stop for the length of a 152 MB copy | native DuckDB file, version-coupled | Rejected | +| 3. External pull over `/api/v1/export` | a sidecar or cron in compose | same connection, strictly more work | ZIP/CSV, **incomplete** | Rejected | + +--- + +## Measurements + +Host: **robmini** (Mac mini M1, 8 cores, macOS 26.6) - the machine production runs +on. Subject: a probe copy of the live production database, taken 2026-10-04 with +`docker cp cotel-cotel-1:/data/cotel.duckdb` (plus its `.wal`) into a throwaway +Docker volume: **152,580,096 B**, 65,184 spans, 3,048 `daily_usage` rows, 17 users, +`schema_version` 10. + +Every row below was produced by the production image's own binary, +`cotel --db-query ""`, which opens the file with `access_mode=read_only` (the +engine is DuckDB 1.5.6, the one the production binary links). That the exports ran +at all under a read-only handle is itself the evidence for one acceptance point: +**the snapshot never writes to the live database file.** + +| Operation | wall clock, incl. ~0.4 s of container start and file open | output | +|---|---|---| +| baseline, `SELECT 1` | 0.42 / 0.36 / 0.36 s | - | +| `EXPORT DATABASE (FORMAT PARQUET, COMPRESSION ZSTD)` | 0.69 / 0.58 / 0.51 s | **4,399,961 B** (4.2 MiB) | +| `EXPORT DATABASE (FORMAT PARQUET)` (snappy) | 0.55 s | 8,629,574 B | +| `EXPORT DATABASE (FORMAT CSV)` | 0.77 s | 71,012,385 B | +| `IMPORT DATABASE` into a fresh file, DuckDB 1.5.6 CLI | 0.96 / 0.74 s | a 36.2 MB database file | + +Subtracting the baseline, the Parquet+ZSTD export statement costs **0.15-0.3 s** and +**4.2 MiB**: 1/35th of the live file, for a full copy of every table. That is the +measurement the recommendation in the ticket hinged on, and it holds - there is no +"blocks ingest for a noticeable time" case to flip the decision to option 3. + +The cost is bounded going forward, not just today: raw spans live 30 days by +retention, and production writes ~2.2 k spans/day, so the exported volume is capped +by the retention window rather than growing with the service's age. + +### The restore was verified, not assumed + +`IMPORT DATABASE` from the snapshot above into an empty volume, then the **cotel +binary** (not the CLI that wrote it) opened the result: + +| Check | Source database | Restored from snapshot | +|---|---|---| +| `count(*) FROM spans` | 65,184 | **65,184** | +| `count(*) FROM daily_usage` | 3,048 | **3,048** | +| `count(*) FROM users` | 17 | **17** | +| `max(start_time) FROM spans` | `2026-10-04 17:53:06.637+00` | **identical** | +| `max(version) FROM schema_version` | 10 | **10** | +| `count(*) FROM duckdb_indexes()` | 4 | **4** | +| opens with `cotel --db-query` | yes | **yes** | + +Two things worth naming in that table. The four secondary ART indexes are **rebuilt +from the data** by the import rather than copied, which is precisely the failure the +2026-09 incident was: a damaged index structure in a file whose rows were all +readable. A Parquet snapshot cannot carry that damage across; a byte copy of the +volume carries it faithfully. And the restored file is 36.2 MB where the live one is +152.6 MB, so a restore also compacts away the free space a year of retention churn +left behind. + +--- + +## Decision + +cotel takes its own snapshots, with `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION +ZSTD)`, on a timer, into a second named volume. No new processes, no new ports, no +token: the deploy stays `docker compose up -d`. + +**Mechanism.** A `RunSnapshotWorker` goroutine alongside `RunRetentionWorker`: run +once at startup, then every `COTEL_SNAPSHOT_INTERVAL`. It runs on the same single +connection as everything else, which is also why it cannot race the WAL checkpoint: +the two are serialised by construction, not by a lock we have to get right. + +**Layout.** One directory per snapshot, named for its UTC instant: + +``` +/snapshots/2026-10-04T18-00-00Z/ + schema.sql load.sql + spans.parquet daily_usage.parquet users.parquet + api_tokens.parquet settings.parquet schema_version.parquet + snapshot.json <- written last; its presence means "complete" +``` + +`snapshot.json` carries the instant, the duration, the row count per table and the +`schema_version`, so a restore can be checked against what the snapshot claimed to +hold. It is written after the export returns, and nothing else creates it, so a +directory without it is a failed or half-written run: prunable, never restorable. +This is also why the export writes straight into its final directory instead of a +temporary one that gets renamed - DuckDB bakes **absolute paths** into `load.sql` +(`COPY spans FROM '/snapshots/.../spans.parquet'`), so renaming the directory after +the fact breaks the import. + +**Depth over density.** `COTEL_SNAPSHOT_INTERVAL=6h`, `COTEL_SNAPSHOT_KEEP=56`: 14 +days of history in 56 snapshots, about 235 MB. The September failure went unnoticed +for six days, so the thing worth buying is reach back past the moment someone +notices, not a tighter recovery point. Disk does not constrain the choice (93 GB +free on the host); the pruning rule does: keep the N newest complete snapshots, +delete the rest, and never delete the last one standing. + +**Health.** The worker records its outcome the way the retention worker does +(`settings` keys surfaced on `/api/v1/health`), because a backup that has been +failing quietly for a month is worse than a known absent one. + +**Restore.** `cotel --db-import ` opens the target database read-write and runs +`IMPORT DATABASE`, so a restore needs nothing but the image that is already on the +host. Today the alternative is the DuckDB CLI with its version matched by hand, +which is step 4 of the recovery page and the step most able to destroy the file; +`docs/operations/` gets the volume-level procedure built on the flag instead. + +--- + +## Why not the other two + +**Option 2, graceful stop plus volume copy.** It is the one option that cannot be +shipped: it needs hands on the host, so there is nothing to merge and nothing that +runs when nobody is watching - which is exactly how production arrived at no backup +at all. It also stops ingest for the length of a 152 MB copy, and the artifact it +produces is a native DuckDB file: version-coupled, and a faithful copy of whatever +corruption the source already carries. The 2026-10-04 recovery produced two such +copies, byte-identical, and neither opens. That is the measured precedent for this +option, not a hypothetical. + +**Option 3, external pull over `/api/v1/export`.** Rejected on three counts, the +first of which is fatal on its own: + +1. **It is not a backup of the database.** The ZIP carries `spans.csv` and + `daily_usage.csv` and nothing else ([ADR-0005](./0005-export-import-format)). + `users`, `api_tokens`, `settings` and `schema_version` are not in it, so an + instance restored from it rejects every agent's ingest token and has no record of + its own schema version. `EXPORT DATABASE` writes all six tables; that is verified + in the file listing above. +2. **It buys no isolation from the writer.** `ExportSpans` runs on the same single + `rw` connection, and it materialises every span into Go memory before encoding + CSV into a ZIP, so it holds the connection *longer* than the engine-side export + while delivering less. +3. **It adds moving parts to the deploy.** An ingest token (board-only to mint), a + sidecar or cron entry in `docker-compose.yml`, and a `period=day|week|month` + argument that has to be chosen per run and silently returns 404 for a period with + no data. + +--- + +## Consequences + +- **Three new env vars**, documented in `README.md` and `docs/index.md`: + `COTEL_SNAPSHOT_DIR` (empty = disabled, so a dev checkout does not start writing + snapshots into its container filesystem; compose sets `/snapshots`), + `COTEL_SNAPSHOT_INTERVAL` (`6h`), `COTEL_SNAPSHOT_KEEP` (`56`). +- **A second volume in `docker-compose.yml`**, `cotel-snapshots` mounted at + `/snapshots`, name overridable the way `COTEL_DATA_VOLUME` already is. It must be + mounted at the same path on restore, because of the absolute paths in `load.sql`. +- **`docker volume prune` is now dangerous in a new way.** The snapshots volume is + only "in use" while the container exists; a prune on a stopped deploy takes the + backups with it. This is a note for `docs/operations/`, not something the code can + prevent. +- **Host loss is still uncovered.** Snapshots land on the same disk as production, so + they insure against file-level corruption, a bad migration and accidental + deletion - the failures cotel has actually had - and not against the disk or the + machine going away. An off-host copy is a separate decision with its own + dependencies (where to put it, what credential it uses); it is deliberately not in + this one. +- **A snapshot of a silently damaged database is still a good snapshot**, because the + export reads rows and the import rebuilds indexes. The converse is the trap option + 2 fell into. +- **The snapshot format is a one-way door only in one direction**: Parquet plus a + plain-text `schema.sql` is readable by any DuckDB build and by anything that reads + Parquet, so a snapshot outlives the engine version that wrote it. That is the + property the native-file copy does not have, and the reason this ADR exists rather + than a cron line in a README. diff --git a/docs/decisions/index.md b/docs/decisions/index.md index 17d15a7..a90a81f 100644 --- a/docs/decisions/index.md +++ b/docs/decisions/index.md @@ -30,3 +30,4 @@ New ADRs go in this directory as `NNNN-short-title.md`, numbered sequentially. | [ADR-0020](./0020-recovery-arrives-as-a-new-issue) | Recovery arrives as a new issue, and dedup is time-bounded | Superseded by ADR-0021 | | [ADR-0021](./0021-recovery-wakes-the-alerts-assignee) | Recovery wakes the alert's assignee, and dedup is time-bounded | Accepted | | [ADR-0022](./0022-health-probe-scheduler-outside-github) | The health probe's scheduler lives outside this repo | Accepted | +| [ADR-0023](./0023-production-database-snapshots) | Snapshots: `EXPORT DATABASE` to Parquet, on a timer, from inside cotel | Accepted | From 519578824792a1e40e226887aa5fa11134b5ac8d Mon Sep 17 00:00:00 2001 From: Wayland Date: Tue, 6 Oct 2026 19:21:18 +0200 Subject: [PATCH 2/6] feat(storage): snapshot the database to Parquet on a timer, and restore from it The snapshot has to come from inside cotel: DuckDB has one writer and the live process holds the file lock, so no external process can open the database even read-only. RunSnapshotWorker runs EXPORT DATABASE (FORMAT PARQUET, COMPRESSION ZSTD) into one dated directory per snapshot under COTEL_SNAPSHOT_DIR, keeping the COTEL_SNAPSHOT_KEEP newest. It runs on the same single connection as ingest and the dashboard, which is why it cannot race the WAL checkpoint - the two are serialised by construction rather than by a lock - and why the decisive cost is how long the statement holds the connection, not how large its output is. snapshot.json is written last and by nothing else, so its presence is the only "complete" signal. It carries the row count per table read back out of the Parquet files, not out of the live tables, so it describes what the snapshot holds rather than what the database held a moment later. The export writes straight into its final directory because DuckDB bakes absolute paths into load.sql, which a rename would break; a test pins that. A cycle only runs when the newest complete snapshot is older than the interval, so a restart - or a crash loop - cannot spend the retained window on snapshots minutes apart. Pruning keeps the newest N complete, deletes incomplete ones, runs only after a successful export so the last one standing survives, and ignores directories whose name is not a snapshot instant. cotel --db-import restores into COTEL_DB_PATH and verifies every table against the manifest, so a restore needs only the image already on the host instead of a DuckDB CLI with its version matched by hand. It refuses a populated target and a snapshot with no manifest. The worker's outcome lands in settings and on /api/v1/health as a snapshot object; a failed export degrades the top-level status, the way a failed retention roll-up does. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- cmd/cotel/main.go | 35 +++ docker-compose.yml | 15 ++ internal/api/handler.go | 71 ++++- internal/api/handler_test.go | 59 +++++ internal/storage/snapshot.go | 419 ++++++++++++++++++++++++++++++ internal/storage/snapshot_test.go | 346 ++++++++++++++++++++++++ 6 files changed, 936 insertions(+), 9 deletions(-) create mode 100644 internal/storage/snapshot.go create mode 100644 internal/storage/snapshot_test.go diff --git a/cmd/cotel/main.go b/cmd/cotel/main.go index b9fbdec..9d89363 100644 --- a/cmd/cotel/main.go +++ b/cmd/cotel/main.go @@ -11,6 +11,7 @@ import ( "net/url" "os" "os/signal" + "sort" "strconv" "strings" "sync/atomic" @@ -28,6 +29,7 @@ import ( func main() { dbQuery := flag.String("db-query", "", "run SQL query against DuckDB, print first column of first row, and exit") + dbImport := flag.String("db-import", "", "restore COTEL_DB_PATH from the snapshot directory given, verify it against the snapshot manifest, and exit; the target database file must be empty or absent") healthcheck := flag.Bool("healthcheck", false, "probe the local dashboard /healthz and exit 0 (ready) or 1; used by the container HEALTHCHECK") flag.Parse() @@ -53,6 +55,18 @@ func main() { return } + // Restore runs before the listeners bind: it needs the database file to + // itself, and nothing should be able to ingest into a half-imported file. + if *dbImport != "" { + m, err := storage.ImportSnapshot(dbPath, *dbImport) + if err != nil { + log.Fatalf("db-import: %v", err) + } + log.Printf("db-import: restored %s from snapshot %s (taken %s, schema_version %d, %s)", + dbPath, *dbImport, m.Instant, m.SchemaVersion, formatTableCounts(m.Tables)) + return + } + // Bind BEFORE storage.Open: WAL replay + schema migration can block for // minutes on a large DB, and a bound port answering a retryable 503 keeps // OTLP clients retrying where a connection reset would drop their spans. @@ -98,6 +112,12 @@ func main() { retentionInterval := envDuration("COTEL_RETENTION_INTERVAL", 6*time.Hour) go db.RunRetentionWorker(retentionCfg, retentionInterval) + snapshotCfg := storage.SnapshotConfig{ + Dir: os.Getenv("COTEL_SNAPSHOT_DIR"), + Keep: envInt("COTEL_SNAPSHOT_KEEP", storage.DefaultSnapshotKeep), + } + go db.RunSnapshotWorker(snapshotCfg, envDuration("COTEL_SNAPSHOT_INTERVAL", storage.DefaultSnapshotInterval)) + ingestMux := http.NewServeMux() ingestMux.Handle("/v1/traces", auth.Middleware(db, ingest.New(db))) @@ -335,6 +355,21 @@ func runHealthcheck(dashAddr string) int { return 0 } +// formatTableCounts renders a snapshot manifest's row counts in a stable order, +// so two restores of the same snapshot log the same line. +func formatTableCounts(tables map[string]int64) string { + names := make([]string, 0, len(tables)) + for name := range tables { + names = append(names, name) + } + sort.Strings(names) + parts := make([]string, 0, len(names)) + for _, name := range names { + parts = append(parts, fmt.Sprintf("%s=%d", name, tables[name])) + } + return strings.Join(parts, " ") +} + func env(key, fallback string) string { if v := os.Getenv(key); v != "" { return v diff --git a/docker-compose.yml b/docker-compose.yml index 2ca23c8..7f1c45f 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -7,6 +7,9 @@ services: - "8080:8080" # Dashboard volumes: - cotel-data:/data + # Snapshots of the whole database (ADR-0023). Mounted at the path baked + # into each snapshot's load.sql, which is what a restore replays. + - cotel-snapshots:/snapshots # Locally-managed Cloudflare tunnel (optional). # Uncomment and adjust the host path if you prefer file-based tunnel config # over the CLOUDFLARE_TUNNEL_TOKEN env var. Mount your ~/.cloudflared/ @@ -27,6 +30,12 @@ services: # get the correct endpoint to paste into ~/.claude/settings.json. # Example: COTEL_PUBLIC_INGEST_URL=https://otlp.example.com COTEL_PUBLIC_INGEST_URL: ${COTEL_PUBLIC_INGEST_URL:-} + # Database snapshots. An empty COTEL_SNAPSHOT_DIR disables them, which is + # the default outside this compose file; here they are always on, because + # a deploy without a recovery point is how production ended up with none. + COTEL_SNAPSHOT_DIR: ${COTEL_SNAPSHOT_DIR:-/snapshots} + COTEL_SNAPSHOT_INTERVAL: ${COTEL_SNAPSHOT_INTERVAL:-6h} + COTEL_SNAPSHOT_KEEP: ${COTEL_SNAPSHOT_KEEP:-56} volumes: # Named explicitly so the deploy can be pointed at a different volume without @@ -38,3 +47,9 @@ volumes: # warns that it did not create a hand-made volume; that warning is expected. cotel-data: name: ${COTEL_DATA_VOLUME:-cotel-data-repaired-20261004} + # Named for the same reason as cotel-data: a restore brings the deploy up + # against a different data volume while this one stays put. Note that + # `docker volume prune` on a stopped deploy takes the backups with it — this + # volume is only "in use" while the container exists. + cotel-snapshots: + name: ${COTEL_SNAPSHOT_VOLUME:-cotel-snapshots} diff --git a/internal/api/handler.go b/internal/api/handler.go index d51bee3..e433bbd 100644 --- a/internal/api/handler.go +++ b/internal/api/handler.go @@ -360,6 +360,7 @@ type healthResponse struct { NewestSpanAgeSeconds *int64 `json:"newest_span_age_seconds"` DBSizeBytes int64 `json:"db_size_bytes"` Retention retentionHealth `json:"retention"` + Snapshot snapshotHealth `json:"snapshot"` PublicIngestURL string `json:"public_ingest_url,omitempty"` } @@ -371,6 +372,18 @@ type retentionHealth struct { LastError string `json:"last_error,omitempty"` } +// snapshotHealth surfaces the outcome of the last snapshot cycle. A backup that +// has been failing quietly for a month is worse than a known absent one, so +// "error" degrades the whole endpoint the way a failing roll-up does. Status +// "unknown" covers both a worker that has not run yet and snapshots left +// disabled, which is the shipped default outside production. +type snapshotHealth struct { + Status string `json:"status"` // "ok" | "error" | "unknown" + LastRunAt string `json:"last_run_at,omitempty"` + LastError string `json:"last_error,omitempty"` + LastDir string `json:"last_dir,omitempty"` +} + func (h *Handler) handleHealth(w http.ResponseWriter, _ *http.Request) { fresh, err := storage.QueryIngestFreshness(h.db) if err != nil { @@ -390,8 +403,13 @@ func (h *Handler) handleHealth(w http.ResponseWriter, _ *http.Request) { queryFailed(w) return } + snap, err := h.snapshotHealth() + if err != nil { + queryFailed(w) + return + } status := "ok" - if ret.Status == "error" { + if ret.Status == "error" || snap.Status == "error" { status = "degraded" } @@ -402,6 +420,7 @@ func (h *Handler) handleHealth(w http.ResponseWriter, _ *http.Request) { NewestSpanAgeSeconds: fresh.AgeSeconds(time.Now()), DBSizeBytes: dbSize, Retention: ret, + Snapshot: snap, PublicIngestURL: h.publicIngestURL, }) } @@ -409,14 +428,7 @@ func (h *Handler) handleHealth(w http.ResponseWriter, _ *http.Request) { // retentionHealth reads the retention-worker status the worker persisted to the // settings table. A fresh DB (worker not yet run) reports status "unknown". func (h *Handler) retentionHealth() (retentionHealth, error) { - get := func(key string) (string, error) { - var v string - err := h.db.QueryRow("SELECT value FROM settings WHERE key = ?", key).Scan(&v) - if !scanOK(err) { - return "", err - } - return v, nil - } + get := h.setting status, err := get("retention_last_status") if err != nil { return retentionHealth{}, err @@ -439,6 +451,47 @@ func (h *Handler) retentionHealth() (retentionHealth, error) { }, nil } +// snapshotHealth reads the snapshot-worker status the worker persisted to the +// settings table. A fresh DB, or one with snapshots disabled, reports "unknown". +func (h *Handler) snapshotHealth() (snapshotHealth, error) { + get := h.setting + status, err := get("snapshot_last_status") + if err != nil { + return snapshotHealth{}, err + } + if status == "" { + status = "unknown" + } + lastRun, err := get("snapshot_last_run_at") + if err != nil { + return snapshotHealth{}, err + } + lastErr, err := get("snapshot_last_error") + if err != nil { + return snapshotHealth{}, err + } + lastDir, err := get("snapshot_last_dir") + if err != nil { + return snapshotHealth{}, err + } + return snapshotHealth{ + Status: status, + LastRunAt: lastRun, + LastError: lastErr, + LastDir: lastDir, + }, nil +} + +// setting reads one settings row, treating an absent key as the empty string. +func (h *Handler) setting(key string) (string, error) { + var v string + err := h.db.QueryRow("SELECT value FROM settings WHERE key = ?", key).Scan(&v) + if !scanOK(err) { + return "", err + } + return v, nil +} + // ---- /api/v1/overview ---- type overviewResponse struct { diff --git a/internal/api/handler_test.go b/internal/api/handler_test.go index 79ecee4..9508e01 100644 --- a/internal/api/handler_test.go +++ b/internal/api/handler_test.go @@ -158,6 +158,65 @@ func TestHealthRetention(t *testing.T) { }) } +// TestHealthSnapshot verifies /health surfaces snapshot-worker degradation the +// same way: a backup that has been failing quietly is the failure a backup +// exists to rule out, so it must read as degraded, not as ok. +func TestHealthSnapshot(t *testing.T) { + t.Run("unknown before first run", func(t *testing.T) { + _, ro := openTestDB(t) + h := api.New(ro) + _, body := getJSON(t, h, "/api/v1/health") + if body["status"] != "ok" { + t.Errorf("want status=ok, got %v", body["status"]) + } + snap, _ := body["snapshot"].(map[string]any) + if snap == nil || snap["status"] != "unknown" { + t.Errorf("want snapshot.status=unknown, got %v", body["snapshot"]) + } + }) + + t.Run("reports the last snapshot directory", func(t *testing.T) { + db, ro := openTestDB(t) + if err := db.SetSetting("snapshot_last_status", "ok"); err != nil { + t.Fatalf("set status: %v", err) + } + if err := db.SetSetting("snapshot_last_dir", "/snapshots/2026-10-06T12-00-00Z"); err != nil { + t.Fatalf("set dir: %v", err) + } + h := api.New(ro) + _, body := getJSON(t, h, "/api/v1/health") + if body["status"] != "ok" { + t.Errorf("want status=ok, got %v", body["status"]) + } + snap, _ := body["snapshot"].(map[string]any) + if snap == nil || snap["last_dir"] != "/snapshots/2026-10-06T12-00-00Z" { + t.Errorf("want snapshot.last_dir echoed, got %v", body["snapshot"]) + } + }) + + t.Run("degraded on recorded error", func(t *testing.T) { + db, ro := openTestDB(t) + if err := db.SetSetting("snapshot_last_status", "error"); err != nil { + t.Fatalf("set status: %v", err) + } + if err := db.SetSetting("snapshot_last_error", "export to /snapshots: No space left on device"); err != nil { + t.Fatalf("set error: %v", err) + } + h := api.New(ro) + _, body := getJSON(t, h, "/api/v1/health") + if body["status"] != "degraded" { + t.Errorf("want status=degraded, got %v", body["status"]) + } + snap, _ := body["snapshot"].(map[string]any) + if snap == nil || snap["status"] != "error" { + t.Fatalf("want snapshot.status=error, got %v", body["snapshot"]) + } + if snap["last_error"] != "export to /snapshots: No space left on device" { + t.Errorf("want snapshot.last_error echoed, got %v", snap["last_error"]) + } + }) +} + func TestOverview(t *testing.T) { cases := []struct { name string diff --git a/internal/storage/snapshot.go b/internal/storage/snapshot.go new file mode 100644 index 0000000..04cfc13 --- /dev/null +++ b/internal/storage/snapshot.go @@ -0,0 +1,419 @@ +package storage + +import ( + "context" + "database/sql" + "encoding/json" + "errors" + "fmt" + "log" + "os" + "path/filepath" + "sort" + "strings" + "time" +) + +// Snapshot defaults. 6h × 56 kept snapshots is 14 days of reach: the September +// 2026 corruption went unnoticed for six days, so depth past the moment someone +// notices is worth more than a tighter recovery point. At the measured 4.2 MiB +// per snapshot that whole window costs ~235 MB. +const ( + DefaultSnapshotKeep = 56 + DefaultSnapshotInterval = 6 * time.Hour +) + +// SnapshotManifestName is written last in a snapshot directory, and nothing +// else writes it. Its presence is therefore the only "this snapshot is +// complete" signal: a directory without it is a failed or half-written run, +// prunable and never restorable. +const SnapshotManifestName = "snapshot.json" + +// snapshotDirLayout names a snapshot directory after its UTC instant. Colons +// are out because the directory has to survive a copy to any filesystem, and +// the remaining format still sorts lexicographically in chronological order, +// which is what the pruner relies on. +const snapshotDirLayout = "2006-01-02T15-04-05Z" + +// snapshotTimeout bounds one EXPORT DATABASE. The export runs on the single +// writer connection with ingest queued behind it (0.15-0.3 s on the 152 MB +// production database), so a wedged export must not be able to hold that +// connection indefinitely. A cancelled export leaves an incomplete directory, +// which the next successful run prunes. +const snapshotTimeout = 5 * time.Minute + +// snapshotRetryCap bounds the wait after a failed run, so a transient failure +// does not cost a whole interval of backup coverage. +const snapshotRetryCap = 15 * time.Minute + +// Setting keys used to surface snapshot-worker health on /api/v1/health. +const ( + settingSnapshotStatus = "snapshot_last_status" // "ok" | "error" + settingSnapshotError = "snapshot_last_error" // last error text, "" on success + settingSnapshotRunAt = "snapshot_last_run_at" // RFC3339 timestamp of last attempt + settingSnapshotDir = "snapshot_last_dir" // directory of the last complete snapshot +) + +// SnapshotConfig controls where snapshots are written and how many are kept. +// An empty Dir disables snapshots entirely, so a dev checkout does not start +// writing them into its container filesystem. +type SnapshotConfig struct { + Dir string + Keep int +} + +// SnapshotManifest is the content of snapshot.json: what the snapshot claims to +// hold, so a restore can be checked against it rather than trusted. +type SnapshotManifest struct { + Instant string `json:"instant"` + DurationMS int64 `json:"duration_ms"` + Format string `json:"format"` + SchemaVersion int `json:"schema_version"` + Tables map[string]int64 `json:"tables"` + Directory string `json:"directory"` +} + +// RunSnapshotWorker exports the whole database to a dated directory under +// cfg.Dir, then every interval. Returns immediately (snapshots disabled) when +// cfg.Dir is empty. +// +// Each cycle only runs when the newest complete snapshot is older than +// interval. That is what keeps a restart - or a crash loop - from spending the +// retained window on snapshots minutes apart: the depth the worker buys is +// bounded by Keep, so churning it is the one way to lose coverage silently. +// +// A failed cycle is not silent: it is logged at ERROR level with the +// consecutive-failure count and recorded on /api/v1/health (status "degraded"). +func (db *DB) RunSnapshotWorker(cfg SnapshotConfig, interval time.Duration) { + if cfg.Dir == "" { + log.Printf("snapshots disabled: no snapshot directory configured") + return + } + log.Printf("snapshot worker: directory %s, interval %s, keeping %d", cfg.Dir, interval, cfg.Keep) + + consecutiveFailures := 0 + for { + wait := snapshotWait(cfg.Dir, interval, time.Now()) + if wait <= 0 { + m, err := db.Snapshot(cfg) + db.recordSnapshotRun(m, err) + switch { + case err != nil: + consecutiveFailures++ + log.Printf("ERROR snapshot worker: export failed (consecutive failures=%d, dir=%s): %v", + consecutiveFailures, cfg.Dir, err) + wait = min(interval, snapshotRetryCap) + default: + consecutiveFailures = 0 + log.Printf("snapshot: wrote %s in %dms (%d tables, schema_version %d)", + m.Directory, m.DurationMS, len(m.Tables), m.SchemaVersion) + wait = interval + } + } + time.Sleep(wait) + } +} + +// snapshotWait reports how long to wait before the next snapshot is due, given +// the newest complete snapshot already on disk. An unreadable or empty +// directory yields 0: let Snapshot produce the real error rather than guess here. +func snapshotWait(dir string, interval time.Duration, now time.Time) time.Duration { + newest, ok := newestSnapshotInstant(dir) + if !ok { + return 0 + } + return interval - now.Sub(newest) +} + +func newestSnapshotInstant(dir string) (time.Time, bool) { + var newest time.Time + for _, s := range listSnapshots(dir) { + if s.complete && s.instant.After(newest) { + newest = s.instant + } + } + return newest, !newest.IsZero() +} + +// Snapshot writes one complete snapshot and prunes older ones. +func (db *DB) Snapshot(cfg SnapshotConfig) (SnapshotManifest, error) { + return db.snapshotAt(cfg, time.Now()) +} + +func (db *DB) snapshotAt(cfg SnapshotConfig, now time.Time) (SnapshotManifest, error) { + if cfg.Dir == "" { + return SnapshotManifest{}, errors.New("snapshot: no directory configured") + } + instant := now.UTC().Truncate(time.Second) + target := filepath.Join(cfg.Dir, instant.Format(snapshotDirLayout)) + if err := os.MkdirAll(target, 0o755); err != nil { + return SnapshotManifest{}, fmt.Errorf("snapshot: create %s: %w", target, err) + } + + ctx, cancel := context.WithTimeout(context.Background(), snapshotTimeout) + defer cancel() + + // The export runs on db.rw, the same single connection ingest and the + // dashboard use, which is why it cannot race the WAL checkpoint: the two are + // serialised by construction rather than by a lock that has to be taken + // correctly. It is also why it never writes to the live database file. + start := time.Now() + stmt := fmt.Sprintf("EXPORT DATABASE '%s' (FORMAT PARQUET, COMPRESSION ZSTD)", quoteSQLLiteral(target)) + if _, err := db.rw.ExecContext(ctx, stmt); err != nil { + return SnapshotManifest{}, fmt.Errorf("snapshot: export to %s: %w", target, err) + } + elapsed := time.Since(start) + + tables, err := snapshotTableCounts(ctx, db.rw, target) + if err != nil { + return SnapshotManifest{}, fmt.Errorf("snapshot: count exported rows in %s: %w", target, err) + } + + m := SnapshotManifest{ + Instant: instant.Format(time.RFC3339), + DurationMS: elapsed.Milliseconds(), + Format: "parquet-zstd", + SchemaVersion: snapshotSchemaVersion(ctx, db.rw, target), + Tables: tables, + Directory: target, + } + if err := writeSnapshotManifest(target, m); err != nil { + return SnapshotManifest{}, err + } + + if err := pruneSnapshots(cfg.Dir, cfg.Keep); err != nil { + // The snapshot itself is complete and restorable; a failed prune only + // costs disk, so it must not report the run as a failed backup. + log.Printf("WARNING snapshot: pruning %s failed, snapshots may accumulate: %v", cfg.Dir, err) + } + return m, nil +} + +// snapshotTableCounts reads the row count of every table the export wrote, from +// the Parquet files themselves rather than from the live tables: the manifest +// has to describe what the snapshot holds, not what the database held a moment +// after it was taken. count(*) over Parquet is answered from file metadata. +func snapshotTableCounts(ctx context.Context, q *sql.DB, dir string) (map[string]int64, error) { + files, err := filepath.Glob(filepath.Join(dir, "*.parquet")) + if err != nil { + return nil, err + } + if len(files) == 0 { + return nil, fmt.Errorf("export wrote no parquet files") + } + sort.Strings(files) + counts := make(map[string]int64, len(files)) + for _, f := range files { + var n int64 + stmt := fmt.Sprintf("SELECT count(*) FROM read_parquet('%s')", quoteSQLLiteral(f)) + if err := q.QueryRowContext(ctx, stmt).Scan(&n); err != nil { + return nil, fmt.Errorf("%s: %w", filepath.Base(f), err) + } + counts[strings.TrimSuffix(filepath.Base(f), ".parquet")] = n + } + return counts, nil +} + +// snapshotSchemaVersion reads the schema version out of the exported +// schema_version table. A snapshot of a database too old to have that table is +// still a valid snapshot, so an unreadable version is recorded as 0 rather than +// failing the run. +func snapshotSchemaVersion(ctx context.Context, q *sql.DB, dir string) int { + var v sql.NullInt64 + stmt := fmt.Sprintf("SELECT max(version) FROM read_parquet('%s')", + quoteSQLLiteral(filepath.Join(dir, "schema_version.parquet"))) + if err := q.QueryRowContext(ctx, stmt).Scan(&v); err != nil || !v.Valid { + return 0 + } + return int(v.Int64) +} + +// writeSnapshotManifest writes snapshot.json via a temporary file in the same +// directory, so a crash mid-write cannot leave a truncated manifest that would +// make an incomplete snapshot look complete. +func writeSnapshotManifest(dir string, m SnapshotManifest) error { + body, err := json.MarshalIndent(m, "", " ") + if err != nil { + return fmt.Errorf("snapshot: encode manifest: %w", err) + } + tmp := filepath.Join(dir, "."+SnapshotManifestName+".tmp") + if err := os.WriteFile(tmp, append(body, '\n'), 0o644); err != nil { + return fmt.Errorf("snapshot: write manifest: %w", err) + } + if err := os.Rename(tmp, filepath.Join(dir, SnapshotManifestName)); err != nil { + return fmt.Errorf("snapshot: publish manifest: %w", err) + } + return nil +} + +// ReadSnapshotManifest loads the manifest of a snapshot directory. A missing +// manifest means the snapshot is incomplete and must not be restored from. +func ReadSnapshotManifest(dir string) (SnapshotManifest, error) { + body, err := os.ReadFile(filepath.Join(dir, SnapshotManifestName)) + if err != nil { + if errors.Is(err, os.ErrNotExist) { + return SnapshotManifest{}, fmt.Errorf("%s has no %s: the snapshot is incomplete and cannot be restored", dir, SnapshotManifestName) + } + return SnapshotManifest{}, err + } + var m SnapshotManifest + if err := json.Unmarshal(body, &m); err != nil { + return SnapshotManifest{}, fmt.Errorf("%s/%s: %w", dir, SnapshotManifestName, err) + } + return m, nil +} + +type snapshotDir struct { + name string + instant time.Time + complete bool +} + +// listSnapshots returns the snapshot directories under dir, oldest first. Only +// directories whose name parses as a snapshot instant are returned: anything +// else in the volume belongs to someone else and the pruner must not touch it. +func listSnapshots(dir string) []snapshotDir { + entries, err := os.ReadDir(dir) + if err != nil { + return nil + } + var out []snapshotDir + for _, e := range entries { + if !e.IsDir() { + continue + } + instant, err := time.Parse(snapshotDirLayout, e.Name()) + if err != nil { + continue + } + _, statErr := os.Stat(filepath.Join(dir, e.Name(), SnapshotManifestName)) + out = append(out, snapshotDir{name: e.Name(), instant: instant, complete: statErr == nil}) + } + sort.Slice(out, func(i, j int) bool { return out[i].instant.Before(out[j].instant) }) + return out +} + +// pruneSnapshots keeps the keep newest complete snapshots and deletes the rest, +// incomplete directories included. It is only ever called after a successful +// export, so the newest complete snapshot is always one that was just verified +// to exist - the last one standing can never be pruned away. +func pruneSnapshots(dir string, keep int) error { + if keep < 1 { + keep = 1 + } + var complete, doomed []string + for _, s := range listSnapshots(dir) { + if s.complete { + complete = append(complete, s.name) + continue + } + doomed = append(doomed, s.name) + } + if excess := len(complete) - keep; excess > 0 { + doomed = append(doomed, complete[:excess]...) + } + var errs []error + for _, name := range doomed { + if err := os.RemoveAll(filepath.Join(dir, name)); err != nil { + errs = append(errs, err) + continue + } + log.Printf("snapshot: pruned %s", filepath.Join(dir, name)) + } + return errors.Join(errs...) +} + +// recordSnapshotRun persists the outcome of one snapshot cycle so the health +// endpoint can report degradation. Best-effort: a failure to write status must +// not take down the worker, so write errors are only logged. +func (db *DB) recordSnapshotRun(m SnapshotManifest, runErr error) { + status := "ok" + msg := "" + if runErr != nil { + status = "error" + msg = runErr.Error() + } + vals := map[string]string{ + settingSnapshotStatus: status, + settingSnapshotError: msg, + settingSnapshotRunAt: time.Now().UTC().Format(time.RFC3339), + } + if runErr == nil { + vals[settingSnapshotDir] = m.Directory + } + for k, v := range vals { + if err := db.SetSetting(k, v); err != nil { + log.Printf("ERROR snapshot worker: failed to record %s: %v", k, err) + } + } +} + +// ImportSnapshot restores the snapshot in dir into the database file at path, +// and verifies the result against the snapshot's manifest. +// +// The target must be empty or absent: the exported schema.sql issues plain +// CREATE TABLE, so importing over populated tables fails halfway and leaves a +// mixed database. The import runs the snapshot's own load.sql, which carries +// the **absolute** paths the export baked in, so dir must be visible at the +// same path it was written to. +func ImportSnapshot(path, dir string) (SnapshotManifest, error) { + m, err := ReadSnapshotManifest(dir) + if err != nil { + return SnapshotManifest{}, err + } + + rw, err := sql.Open("duckdb", path+"?storage_compatibility_version="+StorageCompatibilityVersion) + if err != nil { + return SnapshotManifest{}, fmt.Errorf("open %s: %w", path, err) + } + defer rw.Close() //nolint:errcheck + rw.SetMaxOpenConns(1) + + var existing int + if err := rw.QueryRow("SELECT count(*) FROM duckdb_tables() WHERE schema_name = 'main'").Scan(&existing); err != nil { + return SnapshotManifest{}, fmt.Errorf("inspect %s: %w", path, err) + } + if existing > 0 { + return SnapshotManifest{}, fmt.Errorf("%s already holds %d tables: import needs an empty database file", path, existing) + } + + if _, err := rw.Exec(fmt.Sprintf("IMPORT DATABASE '%s'", quoteSQLLiteral(dir))); err != nil { + return SnapshotManifest{}, fmt.Errorf("import %s into %s: %w", dir, path, err) + } + + if err := verifyImportedCounts(rw, m); err != nil { + return SnapshotManifest{}, err + } + + // Fold the import's WAL now, so the first start of the restored database + // does not have to replay it. + if _, err := rw.Exec("CHECKPOINT"); err != nil { + return SnapshotManifest{}, fmt.Errorf("checkpoint %s after import: %w", path, err) + } + return m, nil +} + +// verifyImportedCounts compares the restored tables against what the manifest +// said the snapshot held. A restore that silently loaded fewer rows than it was +// given is the failure mode a backup exists to rule out, so it is an error. +func verifyImportedCounts(q *sql.DB, m SnapshotManifest) error { + var errs []error + for table, want := range m.Tables { + var got int64 + // Table names come from the snapshot's own file listing, not from user + // input, and the quoting keeps an odd one from breaking the statement. + if err := q.QueryRow(fmt.Sprintf(`SELECT count(*) FROM "%s"`, strings.ReplaceAll(table, `"`, `""`))).Scan(&got); err != nil { + errs = append(errs, fmt.Errorf("%s: %w", table, err)) + continue + } + if got != want { + errs = append(errs, fmt.Errorf("%s: restored %d rows, snapshot holds %d", table, got, want)) + } + } + return errors.Join(errs...) +} + +// quoteSQLLiteral escapes a value for a single-quoted DuckDB string literal. +// EXPORT/IMPORT DATABASE and read_parquet take a literal path, not a bind +// parameter, so the path has to be interpolated. +func quoteSQLLiteral(s string) string { return strings.ReplaceAll(s, "'", "''") } diff --git a/internal/storage/snapshot_test.go b/internal/storage/snapshot_test.go new file mode 100644 index 0000000..3cebf2a --- /dev/null +++ b/internal/storage/snapshot_test.go @@ -0,0 +1,346 @@ +package storage + +import ( + "encoding/json" + "os" + "path/filepath" + "strings" + "testing" + "time" +) + +func insertOneSpan(t *testing.T, db *DB, spanID string, at time.Time) { + t.Helper() + inp := int64(10) + out := int64(20) + cost := 0.001 + if err := db.InsertSpan(Span{ + TraceID: "trace-" + spanID, + SpanID: spanID, + Name: "claude_code.session", + StartTime: at, + EndTime: at.Add(time.Second), + SessionID: "session-1", + Model: "claude-sonnet-4-6", + ToolName: "Bash", + InputTokens: &inp, + OutputTokens: &out, + CostUSD: &cost, + }); err != nil { + t.Fatalf("insert span %s: %v", spanID, err) + } +} + +func TestSnapshotWritesManifestLast(t *testing.T) { + db, err := Open(":memory:") + if err != nil { + t.Fatalf("open in-memory db: %v", err) + } + defer db.Close() //nolint:errcheck + insertOneSpan(t, db, "span-001", time.Now()) + + dir := t.TempDir() + m, err := db.Snapshot(SnapshotConfig{Dir: dir, Keep: 4}) + if err != nil { + t.Fatalf("Snapshot: %v", err) + } + + for _, name := range []string{"schema.sql", "load.sql", "spans.parquet", "users.parquet", SnapshotManifestName} { + if _, err := os.Stat(filepath.Join(m.Directory, name)); err != nil { + t.Errorf("snapshot is missing %s: %v", name, err) + } + } + if got := m.Tables["spans"]; got != 1 { + t.Errorf("manifest spans count = %d, want 1", got) + } + ddl, err := schemaFS.ReadFile("schema.sql") + if err != nil { + t.Fatalf("read schema.sql: %v", err) + } + want, err := schemaVersion(string(ddl)) + if err != nil { + t.Fatalf("schemaVersion: %v", err) + } + if m.SchemaVersion != want { + t.Errorf("manifest schema_version = %d, want %d", m.SchemaVersion, want) + } + + // The manifest on disk must be the manifest returned, and it must parse: + // a restore reads it to decide whether the snapshot is usable at all. + body, err := os.ReadFile(filepath.Join(m.Directory, SnapshotManifestName)) + if err != nil { + t.Fatalf("read manifest: %v", err) + } + var onDisk SnapshotManifest + if err := json.Unmarshal(body, &onDisk); err != nil { + t.Fatalf("manifest is not valid JSON: %v", err) + } + if onDisk.Instant != m.Instant || onDisk.Tables["spans"] != m.Tables["spans"] { + t.Errorf("manifest on disk %+v does not match returned %+v", onDisk, m) + } + + // No temporary manifest may survive the run, or a reader could mistake it + // for the real thing. + entries, err := os.ReadDir(m.Directory) + if err != nil { + t.Fatalf("read snapshot dir: %v", err) + } + for _, e := range entries { + if strings.HasPrefix(e.Name(), ".") { + t.Errorf("snapshot left a temporary file behind: %s", e.Name()) + } + } +} + +// TestSnapshotLoadSQLCarriesAbsolutePaths pins the constraint the layout is +// built on: DuckDB bakes the export directory's absolute path into load.sql, so +// a snapshot cannot be exported to a temporary directory and renamed, and must +// be mounted at the same path to be restored. +func TestSnapshotLoadSQLCarriesAbsolutePaths(t *testing.T) { + db, err := Open(":memory:") + if err != nil { + t.Fatalf("open in-memory db: %v", err) + } + defer db.Close() //nolint:errcheck + insertOneSpan(t, db, "span-001", time.Now()) + + m, err := db.Snapshot(SnapshotConfig{Dir: t.TempDir(), Keep: 4}) + if err != nil { + t.Fatalf("Snapshot: %v", err) + } + load, err := os.ReadFile(filepath.Join(m.Directory, "load.sql")) + if err != nil { + t.Fatalf("read load.sql: %v", err) + } + if !strings.Contains(string(load), m.Directory) { + t.Errorf("load.sql does not reference %s by absolute path:\n%s", m.Directory, load) + } +} + +func TestImportSnapshotRestoresEveryTable(t *testing.T) { + src := filepath.Join(t.TempDir(), "source.duckdb") + db, err := Open(src) + if err != nil { + t.Fatalf("open source db: %v", err) + } + insertOneSpan(t, db, "span-001", time.Now()) + insertOneSpan(t, db, "span-002", time.Now()) + if err := db.SetSetting("snapshot-test-key", "kept"); err != nil { + t.Fatalf("set setting: %v", err) + } + m, err := db.Snapshot(SnapshotConfig{Dir: t.TempDir(), Keep: 4}) + if err != nil { + t.Fatalf("Snapshot: %v", err) + } + if err := db.Close(); err != nil { + t.Fatalf("close source db: %v", err) + } + + restored := filepath.Join(t.TempDir(), "restored.duckdb") + got, err := ImportSnapshot(restored, m.Directory) + if err != nil { + t.Fatalf("ImportSnapshot: %v", err) + } + if got.Instant != m.Instant { + t.Errorf("imported manifest instant = %q, want %q", got.Instant, m.Instant) + } + + // The restored file must be openable by the same code path production uses, + // not just by the import. + reopened, err := Open(restored) + if err != nil { + t.Fatalf("open restored db: %v", err) + } + defer reopened.Close() //nolint:errcheck + + var spans int64 + if err := reopened.rw.QueryRow("SELECT count(*) FROM spans").Scan(&spans); err != nil { + t.Fatalf("count restored spans: %v", err) + } + if spans != 2 { + t.Errorf("restored spans = %d, want 2", spans) + } + v, err := reopened.GetSetting("snapshot-test-key") + if err != nil || v != "kept" { + t.Errorf("restored setting = %q (err %v), want %q", v, err, "kept") + } + // The secondary indexes are rebuilt from the data by the import rather than + // copied, which is the property that keeps a damaged index out of a restore. + var indexes int64 + if err := reopened.rw.QueryRow("SELECT count(*) FROM duckdb_indexes()").Scan(&indexes); err != nil { + t.Fatalf("count restored indexes: %v", err) + } + if indexes == 0 { + t.Error("restored database has no secondary indexes") + } +} + +func TestImportSnapshotRefusesIncompleteSnapshot(t *testing.T) { + dir := t.TempDir() + if err := os.WriteFile(filepath.Join(dir, "spans.parquet"), []byte("not really parquet"), 0o644); err != nil { + t.Fatalf("write stub file: %v", err) + } + if _, err := ImportSnapshot(filepath.Join(t.TempDir(), "restored.duckdb"), dir); err == nil { + t.Fatal("ImportSnapshot accepted a snapshot with no manifest") + } +} + +func TestImportSnapshotRefusesPopulatedTarget(t *testing.T) { + db, err := Open(":memory:") + if err != nil { + t.Fatalf("open in-memory db: %v", err) + } + defer db.Close() //nolint:errcheck + insertOneSpan(t, db, "span-001", time.Now()) + m, err := db.Snapshot(SnapshotConfig{Dir: t.TempDir(), Keep: 4}) + if err != nil { + t.Fatalf("Snapshot: %v", err) + } + + target := filepath.Join(t.TempDir(), "occupied.duckdb") + occupied, err := Open(target) + if err != nil { + t.Fatalf("open target db: %v", err) + } + if err := occupied.Close(); err != nil { + t.Fatalf("close target db: %v", err) + } + + _, err = ImportSnapshot(target, m.Directory) + if err == nil { + t.Fatal("ImportSnapshot overwrote a database that already had tables") + } + if !strings.Contains(err.Error(), "empty database file") { + t.Errorf("error does not say what to do about it: %v", err) + } +} + +func TestPruneSnapshotsKeepsNewestCompleteOnly(t *testing.T) { + dir := t.TempDir() + base := time.Date(2026, 10, 6, 0, 0, 0, 0, time.UTC) + mk := func(at time.Time, complete bool) string { + name := at.Format(snapshotDirLayout) + path := filepath.Join(dir, name) + if err := os.MkdirAll(path, 0o755); err != nil { + t.Fatalf("mkdir %s: %v", path, err) + } + if complete { + if err := os.WriteFile(filepath.Join(path, SnapshotManifestName), []byte("{}"), 0o644); err != nil { + t.Fatalf("write manifest: %v", err) + } + } + return name + } + oldest := mk(base, true) + middle := mk(base.Add(6*time.Hour), true) + newest := mk(base.Add(12*time.Hour), true) + halfWritten := mk(base.Add(18*time.Hour), false) + // Anything whose name is not a snapshot instant belongs to someone else. + foreign := "forensics-20261004" + if err := os.MkdirAll(filepath.Join(dir, foreign), 0o755); err != nil { + t.Fatalf("mkdir %s: %v", foreign, err) + } + + if err := pruneSnapshots(dir, 2); err != nil { + t.Fatalf("pruneSnapshots: %v", err) + } + + for _, name := range []string{middle, newest, foreign} { + if _, err := os.Stat(filepath.Join(dir, name)); err != nil { + t.Errorf("%s should have survived the prune: %v", name, err) + } + } + for _, name := range []string{oldest, halfWritten} { + if _, err := os.Stat(filepath.Join(dir, name)); !os.IsNotExist(err) { + t.Errorf("%s should have been pruned (err %v)", name, err) + } + } +} + +func TestPruneSnapshotsNeverEmptiesTheDirectory(t *testing.T) { + dir := t.TempDir() + only := time.Date(2026, 10, 6, 0, 0, 0, 0, time.UTC).Format(snapshotDirLayout) + if err := os.MkdirAll(filepath.Join(dir, only), 0o755); err != nil { + t.Fatalf("mkdir: %v", err) + } + if err := os.WriteFile(filepath.Join(dir, only, SnapshotManifestName), []byte("{}"), 0o644); err != nil { + t.Fatalf("write manifest: %v", err) + } + if err := pruneSnapshots(dir, 0); err != nil { + t.Fatalf("pruneSnapshots: %v", err) + } + if _, err := os.Stat(filepath.Join(dir, only)); err != nil { + t.Errorf("a Keep of 0 deleted the last snapshot standing: %v", err) + } +} + +func TestSnapshotWaitDefersUntilDue(t *testing.T) { + dir := t.TempDir() + now := time.Date(2026, 10, 6, 12, 0, 0, 0, time.UTC) + mkComplete := func(at time.Time) { + path := filepath.Join(dir, at.Format(snapshotDirLayout)) + if err := os.MkdirAll(path, 0o755); err != nil { + t.Fatalf("mkdir: %v", err) + } + if err := os.WriteFile(filepath.Join(path, SnapshotManifestName), []byte("{}"), 0o644); err != nil { + t.Fatalf("write manifest: %v", err) + } + } + + if wait := snapshotWait(dir, 6*time.Hour, now); wait > 0 { + t.Errorf("an empty directory must be due immediately, got %s", wait) + } + + mkComplete(now.Add(-1 * time.Hour)) + if wait := snapshotWait(dir, 6*time.Hour, now); wait != 5*time.Hour { + t.Errorf("wait after a 1h-old snapshot = %s, want 5h", wait) + } + + mkComplete(now.Add(-7 * time.Hour)) + if wait := snapshotWait(dir, 6*time.Hour, now); wait != 5*time.Hour { + t.Errorf("wait must follow the newest snapshot, got %s, want 5h", wait) + } +} + +func TestRunSnapshotWorkerReturnsWhenDisabled(t *testing.T) { + db, err := Open(":memory:") + if err != nil { + t.Fatalf("open in-memory db: %v", err) + } + defer db.Close() //nolint:errcheck + + done := make(chan struct{}) + go func() { + db.RunSnapshotWorker(SnapshotConfig{Keep: 4}, time.Hour) + close(done) + }() + select { + case <-done: + case <-time.After(5 * time.Second): + t.Fatal("RunSnapshotWorker did not return with snapshots disabled") + } +} + +func TestRecordSnapshotRunSurfacesFailure(t *testing.T) { + db, err := Open(":memory:") + if err != nil { + t.Fatalf("open in-memory db: %v", err) + } + defer db.Close() //nolint:errcheck + + // A failed run must not advance the recorded directory: the last snapshot + // that exists is still the last one that succeeded. + db.recordSnapshotRun(SnapshotManifest{Directory: "/snapshots/good"}, nil) + db.recordSnapshotRun(SnapshotManifest{}, os.ErrPermission) + + status, err := db.GetSetting(settingSnapshotStatus) + if err != nil || status != "error" { + t.Errorf("status = %q (err %v), want %q", status, err, "error") + } + if dir, err := db.GetSetting(settingSnapshotDir); err != nil || dir != "/snapshots/good" { + t.Errorf("last dir = %q (err %v), want %q", dir, err, "/snapshots/good") + } + if msg, err := db.GetSetting(settingSnapshotError); err != nil || msg == "" { + t.Errorf("last error = %q (err %v), want the failure text", msg, err) + } +} From aeb1b38c209b79785510f20217f69a6f7a484ffb Mon Sep 17 00:00:00 2001 From: Wayland Date: Tue, 6 Oct 2026 19:21:28 +0200 Subject: [PATCH 3/6] docs: document snapshots, the restore procedure and the new env vars A new operations page covers how to tell whether snapshots are happening, how to restore one into a probe volume and promote it, and the three things snapshots do not cover: host loss, docker volume prune on a stopped deploy, and anything newer than the last snapshot. It also states plainly that a snapshot carries the plaintext ingest tokens, because users.token is plaintext by design and a snapshot without it would not restore to a working instance. The recovery page gains a pointer to it, loses its "there is no backup" section, and gets the ARM CLI asset name corrected: DuckDB publishes duckdb_cli-linux-aarch64.zip up to v1.2.x and duckdb_cli-linux-arm64.zip from v1.3.0, verified against the release assets of v1.1.3, v1.2.2, v1.3.0 and v1.5.6. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 3 + README.md | 29 +++ docs/.vitepress/config.js | 1 + .../0023-production-database-snapshots.md | 5 + docs/index.md | 4 + docs/operations/api-reference.md | 3 +- docs/operations/duckdb-recovery.md | 21 ++- docs/operations/duckdb-snapshots.md | 168 ++++++++++++++++++ 8 files changed, 228 insertions(+), 6 deletions(-) create mode 100644 docs/operations/duckdb-snapshots.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 972ea95..9d3b209 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] ### Added +- cotel snapshots its own database. A worker runs `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION ZSTD)` at startup and then every `COTEL_SNAPSHOT_INTERVAL`, into one dated directory per snapshot under `COTEL_SNAPSHOT_DIR` (`/snapshots` in compose, its own volume), keeping the `COTEL_SNAPSHOT_KEEP` newest. It had to be taken from inside cotel: DuckDB has one writer and the live process holds the file lock, so no external process can open the database even read-only. Measured on the production host against a probe copy of the live 152.6 MB database, the export holds the single connection for 0.15-0.3 s and writes 4.2 MiB — 1/35th of the file for a full copy of all six tables — and because it runs on that same connection it cannot race the WAL checkpoint: the two are serialised by construction rather than by a lock. The format is the point: Parquet plus a plain-text `schema.sql` is readable by any DuckDB build, and an import *rebuilds* the secondary indexes from the data, so a snapshot of a database with a damaged ART index — the September 2026 failure, a file whose rows were all readable — restores to a healthy one, where a byte copy of the volume reproduces the damage faithfully. `snapshot.json` is written last and by nothing else, so its presence is the only "complete" signal; it carries the row count per table, read back out of the Parquet files, so a restore is checked against what the snapshot claims rather than trusted. A snapshot is only taken when the newest complete one is older than the interval, so a restart — or a crash loop — cannot spend the retained window on snapshots minutes apart. Pruning keeps the newest N complete snapshots, deletes incomplete ones, runs only after a successful export, and leaves directories whose name is not a snapshot instant alone ([ADR-0023](docs/decisions/0023-production-database-snapshots.md), [docs/operations/duckdb-snapshots.md](docs/operations/duckdb-snapshots.md)) +- `cotel --db-import ` restores a snapshot into `COTEL_DB_PATH` and verifies every table against the snapshot's manifest, failing with the table and both numbers named if a count does not match. It needs nothing but the image already on the host — no DuckDB CLI with its version matched by hand, which is the step of the recovery procedure most able to destroy a file. It refuses a target that already holds tables (the snapshot's `schema.sql` issues plain `CREATE TABLE`, so importing over data would fail halfway and leave a mixed database) and a snapshot with no `snapshot.json`. Note that the paths in a snapshot's `load.sql` are **absolute**, so a snapshot cannot be moved or renamed and still be imported; it must be visible at the path it was written to +- `GET /api/v1/health` reports a `snapshot` object (`status`, `last_run_at`, `last_error`, `last_dir`), and a failed export flips the top-level `status` to `degraded` — the same treatment a failing retention roll-up gets, because a backup that has been failing quietly for a month is worse than a known absent one. `unknown` covers both a worker that has not run yet and snapshots left disabled, which is the default outside the compose file - The Overview leads with a **Span activity** grid — a block of cells counting spans, GitHub-contribution-graph style, sitting directly under the KPI row. A line chart answers *how much and when*; it does not answer *what does a week here look like*, which is the question a telemetry front door gets asked most. The selected range picks both the grid and how much time one cell is: 53 × 7 day cells over a year (and over `All`), 31 × 6 four-hour cells over a month, 24 × 7 hourly cells over a week, 24 × 6 ten-minute cells over a day — each tiling its window exactly, and each about 115 px tall, so switching range does not move the page under the reader. Cells outside the queried window — the leading edge of the lattice, and the rest of today — are drawn as an outline with no fill: an empty cell means "we looked and there was nothing", an outline means "we did not look", and conflating the two is how a heatmap invents a quiet weekend. The grid is placed in UTC, which the footer and every tooltip say. Intensity is cut at the quartiles of the cells in view, not scaled against the busiest one: against a 722-span peak a 200-span day and a 700-span day are both "busy", so a max-relative ramp — linear or log — renders a working week as one flat block of full-intensity cells, which is the difference the grid exists to show. A step therefore means a rank, so the footer names the busiest cell in view and every tooltip gives the cell's own count. The scale lives in `frontend/src/lib/heat.ts` and is shared with the History page's calendar and hour-of-day heatmaps, which had a private copy of it and pick up the quartile cut with this change ([ADR-0016](docs/decisions/0016-overview-activity-grid.md)) - `GET /api/v1/history` accepts two more bucket widths, `granularity=10m` and `granularity=4h`, so the activity grid asks for exactly the width it draws — one bucket, one cell — instead of re-bucketing an hourly series in the page, which could not have produced a ten-minute cell at all. Both are additive and no existing caller changes; an unrecognised width still falls back to `day` rather than 400ing. Like `hour` they are answered from `spans` alone and report the shortfall in `covered_since`, because `daily_usage` buckets whole UTC days and cannot produce a sub-day bucket. `bucket` is now documented as a UTC wall-clock label floored to the width, on any host, so a client can reconstruct it for an instant without asking what the server thinks midnight is - `scripts/seed-demo.py` fills a throwaway instance with a synthetic team — seven users, three models, 90 days of sessions and tool calls. It goes in over the OTLP endpoint rather than writing to DuckDB, so a seeded instance exercises the same ingest, cost-derivation and roll-up path a real one does, and it sends only attributes Claude Code actually sends: no `command` on `Bash` spans, so the Tools page shows the same "no command detail" state a real install sees. The RNG seed is fixed, so a re-run against a fresh volume reproduces the same numbers. `scripts/shoot-screenshots.mjs` turns that instance into the README images, each cropped at the bottom edge of a named element rather than at a pixel count ([docs/operations/screenshots.md](docs/operations/screenshots.md)) diff --git a/README.md b/README.md index 50206bf..493f9b4 100644 --- a/README.md +++ b/README.md @@ -385,6 +385,31 @@ docker run --rm -v cotel-data:/data ubuntu \ duckdb /data/cotel.duckdb "SELECT model, COUNT(*) FROM spans GROUP BY model" ``` +## Snapshots + +cotel exports its whole database to Parquet on a timer and keeps the newest +`COTEL_SNAPSHOT_KEEP` exports in a second volume (`/snapshots`), so there is a +recovery point that does not depend on the live DuckDB file being readable. The +export runs inside cotel, on the same connection as everything else, and costs +0.15-0.3 s per run on a 152 MB database; the format is portable Parquet plus a +plain-text `schema.sql`, which any DuckDB build can read +([ADR-0023](docs/decisions/0023-production-database-snapshots.md)). + +```bash +# what is on disk +docker run --rm -v cotel-snapshots:/snapshots debian:bookworm-slim ls -1 /snapshots + +# restore one into an empty volume, verified against the snapshot's manifest +docker run --rm --entrypoint /usr/local/bin/cotel \ + -v cotel-snapshots:/snapshots -v cotel-data-restore:/data \ + ghcr.io/flopsstuff/cotel:latest --db-import /snapshots/2026-10-06T12-00-00Z +``` + +The worker's last outcome is reported on `GET /api/v1/health` under a `snapshot` +object, and a failed export degrades the top-level `status`. Full procedure, +including promoting a restored volume: +[Database Snapshots and Restore](docs/operations/duckdb-snapshots.md). + ## Retention defaults | Tier | Period | Storage | @@ -495,6 +520,10 @@ from the network at startup instead of bundling it). | `COTEL_RETENTION_RAW_DAYS` | `30` | Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer) | | `COTEL_RETENTION_AGGREGATE_DAYS` | `90` | Daily aggregate retention in days | | `COTEL_RETENTION_INTERVAL` | `6h` | Retention worker tick interval (Go duration) | +| `COTEL_SNAPSHOT_DIR` | _(unset — snapshots off)_ | Directory the snapshot worker exports the whole database into, one dated subdirectory per snapshot. Empty disables snapshots; `docker-compose.yml` sets `/snapshots`, backed by its own volume. See [Database Snapshots and Restore](docs/operations/duckdb-snapshots.md). | +| `COTEL_SNAPSHOT_INTERVAL` | `6h` | How often a snapshot is taken (Go duration). A snapshot is only taken when the newest complete one is older than this, so a restart cannot churn through the retained window. | +| `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep; older ones and incomplete ones are pruned after each successful export. At the default interval, 56 is 14 days of reach for about 235 MB. The last snapshot standing is never pruned. | +| `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not by the binary: the Docker volume mounted at `/snapshots`. Note that `docker volume prune` on a stopped deploy deletes it — the volume counts as in use only while the container exists. | | `COTEL_WAL_AUTOCHECKPOINT` | `4MB` | DuckDB `checkpoint_threshold`: the write-ahead log is folded into the main file once it grows past this size. Lower values bound how much WAL an ungraceful kill leaves to replay on the next open; higher values checkpoint less often during ingest. DuckDB's own default is `16MB`. | | `CLOUDFLARE_TUNNEL_TOKEN` | _(unset)_ | When set, starts `cloudflared tunnel run` before cotel; enables public HTTPS access via Cloudflare Tunnel | | `TUNNEL_EDGE_IP_VERSION` | `4` in token mode | Read by `cloudflared`, not by cotel: the address family used to reach the Cloudflare edge (`4`, `6` or `auto`). cloudflared's own default became `auto` in 2026.4.0, which tries whichever family the resolver answers with first and falls back only after a connection has failed; token mode pins `4` unless you set this. Not set in local-config mode, where `config.yml` owns the setting. See [token mode](docs/operations/cloudflare-tunnel-remote.md#the-bundled-cloudflared) | diff --git a/docs/.vitepress/config.js b/docs/.vitepress/config.js index 6ce3bc4..9a59ba0 100644 --- a/docs/.vitepress/config.js +++ b/docs/.vitepress/config.js @@ -32,6 +32,7 @@ export default defineConfig({ { text: 'Production /healthz probe', link: '/operations/health-probe' }, { text: 'Export / Import', link: '/operations/export-import' }, { text: 'DuckDB Recovery', link: '/operations/duckdb-recovery' }, + { text: 'Database Snapshots and Restore', link: '/operations/duckdb-snapshots' }, { text: 'README Screenshots', link: '/operations/screenshots' }, ], }, diff --git a/docs/decisions/0023-production-database-snapshots.md b/docs/decisions/0023-production-database-snapshots.md index 9a50a24..e91a856 100644 --- a/docs/decisions/0023-production-database-snapshots.md +++ b/docs/decisions/0023-production-database-snapshots.md @@ -191,6 +191,11 @@ first of which is fatal on its own: only "in use" while the container exists; a prune on a stopped deploy takes the backups with it. This is a note for `docs/operations/`, not something the code can prevent. +- **Every snapshot carries the ingest tokens.** `users.token` is plaintext by + design, and a snapshot that omitted it would not be restorable to a working + instance. So the snapshots volume holds the same secrets as the data volume and + has to be treated the same way, which is a second reason an off-host copy is + its own decision rather than an obvious next step. - **Host loss is still uncovered.** Snapshots land on the same disk as production, so they insure against file-level corruption, a bad migration and accidental deletion - the failures cotel has actually had - and not against the disk or the diff --git a/docs/index.md b/docs/index.md index 13780ca..922fa2f 100644 --- a/docs/index.md +++ b/docs/index.md @@ -69,6 +69,10 @@ Restart Claude Code. Telemetry starts flowing immediately. | `COTEL_RETENTION_RAW_DAYS` | `30` | Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer) | | `COTEL_RETENTION_AGGREGATE_DAYS` | `90` | Daily aggregate retention in days | | `COTEL_RETENTION_INTERVAL` | `6h` | Retention worker tick interval | +| `COTEL_SNAPSHOT_DIR` | _(unset — snapshots off)_ | Where the snapshot worker exports the whole database (compose sets `/snapshots`); empty disables snapshots. See [Database Snapshots and Restore](./operations/duckdb-snapshots) | +| `COTEL_SNAPSHOT_INTERVAL` | `6h` | How often a snapshot is taken (Go duration) | +| `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep — 14 days at the default interval | +| `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not the binary: the Docker volume mounted at `/snapshots` | ## Data & retention diff --git a/docs/operations/api-reference.md b/docs/operations/api-reference.md index 13e733a..c99eb2c 100644 --- a/docs/operations/api-reference.md +++ b/docs/operations/api-reference.md @@ -250,12 +250,13 @@ Instance health. Takes no parameters and is never range-scoped. | Field | Type | Meaning | |---|---|---| -| `status` | `ok` \| `degraded` | `degraded` when the last retention roll-up failed | +| `status` | `ok` \| `degraded` | `degraded` when the last retention roll-up or the last snapshot failed | | `span_count` | integer | Rows in `spans` | | `last_ingest_at` | RFC 3339 string \| `null` | When the newest span was **accepted**; `null` if nothing was ever ingested | | `newest_span_age_seconds` | integer \| `null` | Seconds since `last_ingest_at`; `null` if nothing was ever ingested | | `db_size_bytes` | integer | Approximate database file size | | `retention` | object | `status` (`ok` \| `error` \| `unknown`), `last_run_at`, `last_error` | +| `snapshot` | object | `status` (`ok` \| `error` \| `unknown`), `last_run_at`, `last_error`, `last_dir` — the last complete snapshot's directory. `unknown` covers both "has not run yet" and snapshots disabled ([Database Snapshots and Restore](./duckdb-snapshots)) | | `public_ingest_url` | string | Omitted unless `COTEL_PUBLIC_INGEST_URL` is set | The two freshness fields are measured from the span's `ingested_at`, not its diff --git a/docs/operations/duckdb-recovery.md b/docs/operations/duckdb-recovery.md index b1c204e..1058fb0 100644 --- a/docs/operations/duckdb-recovery.md +++ b/docs/operations/duckdb-recovery.md @@ -97,7 +97,7 @@ Then install exactly that CLI version: ```bash DUCKDB_VER=v1.1.3 -ARCH=aarch64 # x86_64 hosts: amd64 +ARCH=aarch64 # see the note below; x86_64 hosts: amd64 docker run --rm -it -v cotel-data-probe:/data debian:bookworm-slim sh -c " apt-get -qq update && apt-get -qq install -y --no-install-recommends curl unzip ca-certificates && curl -fsSL https://github.com/duckdb/duckdb/releases/download/\$DUCKDB_VER/duckdb_cli-linux-\$ARCH.zip -o /tmp/d.zip && @@ -105,8 +105,18 @@ docker run --rm -it -v cotel-data-probe:/data debian:bookworm-slim sh -c " duckdb --version && exec bash" ``` +The ARM asset was renamed between releases: DuckDB up to and including v1.2.x +publishes `duckdb_cli-linux-aarch64.zip`, v1.3.0 and later publish +`duckdb_cli-linux-arm64.zip`. Using the wrong one gets a 404 from the release, +not a wrong binary, so it is a nuisance rather than a hazard — but check the +release's asset list rather than guessing. + Confirm `duckdb --version` prints the same version as `SELECT version()` above before you run a single statement against the file. +Restoring from a snapshot needs none of this step: `cotel --db-import` uses the +engine already linked into the image. See +[Database Snapshots and Restore](./duckdb-snapshots#restoring-from-a-snapshot). + ## Step 5 — Rebuild the secondary indexes The four secondary indexes on `spans` are plain ART indexes and carry no data of their own — dropping and recreating them rebuilds the damaged structure from the table. Inside the CLI container from step 4: @@ -276,9 +286,10 @@ re-examine at both of these moments rather than on a calendar: - the DuckDB version linked into cotel moves, which is also when the repro becomes the natural acceptance test for the bump. -### There is no backup of the live database +### The leftovers are not the backup Three volumes look like redundancy and are not. Two of them are the same damaged -file and the third is production itself; once the leftovers are gone, the live -database has no snapshot anywhere. That gap predates this incident and is tracked -separately. Do not read the table above as a backup policy. +file and the third is production itself. Do not read the table above as a backup +policy: the recovery point comes from the timed Parquet snapshots in a separate +volume ([Database Snapshots and Restore](./duckdb-snapshots)), and deleting +these leftovers does not touch it. diff --git a/docs/operations/duckdb-snapshots.md b/docs/operations/duckdb-snapshots.md new file mode 100644 index 0000000..116ba0c --- /dev/null +++ b/docs/operations/duckdb-snapshots.md @@ -0,0 +1,168 @@ +# Database Snapshots and Restore + +cotel exports its whole database to Parquet on a timer and keeps the newest N +exports in a second volume. This page is how you check that it is happening, how +you restore from one, and what the snapshots do *not* protect against. + +The decision behind the mechanism, with the measurements it was chosen on, is +[ADR-0023](../decisions/0023-production-database-snapshots). The recovery +procedure for a database that will not open at all is a different page: +[Recovering a DuckDB File That Will Not Open](./duckdb-recovery). + +## What the mechanism does + +A worker inside cotel runs `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION +ZSTD)` once at startup and then every `COTEL_SNAPSHOT_INTERVAL`, writing one +directory per snapshot into `COTEL_SNAPSHOT_DIR`: + +``` +/snapshots/2026-10-06T12-00-00Z/ + schema.sql load.sql + spans.parquet daily_usage.parquet users.parquet + api_tokens.parquet settings.parquet schema_version.parquet + snapshot.json +``` + +Three properties are worth knowing before you rely on it: + +- **`snapshot.json` is written last, and nothing else writes it.** A directory + without it is a failed or interrupted run: prunable, never restorable. It + records the instant, the export duration, the schema version and the row count + per table, read back out of the Parquet files themselves — so a restore can be + checked against what the snapshot claims to hold rather than trusted. +- **A snapshot is only taken when the newest complete one is older than the + interval.** A restart — or a crash loop — therefore cannot spend the retained + window on snapshots minutes apart. +- **The paths inside `load.sql` are absolute.** DuckDB bakes the export + directory's path into it, so a snapshot cannot be moved or renamed and still + be imported. It must be visible at the path it was written to — `/snapshots` + in the shipped compose file. + +Pruning keeps the `COTEL_SNAPSHOT_KEEP` newest complete snapshots, deletes +incomplete ones, and only ever runs after a successful export, so the last +snapshot standing can never be pruned away. Directories whose name is not a +snapshot instant are left alone, so the volume is safe to share with anything +else you keep there. + +## Is it working? + +`/api/v1/health` carries the worker's own report: + +```bash +curl -s localhost:8080/api/v1/health | jq '{status, snapshot}' +{ + "status": "ok", + "snapshot": { + "status": "ok", + "last_run_at": "2026-10-06T12:00:04Z", + "last_dir": "/snapshots/2026-10-06T12-00-00Z" + } +} +``` + +`snapshot.status` is `error` after a failed cycle, which also flips the +top-level `status` to `degraded` — the same treatment a failing retention +roll-up gets, and for the same reason: a backup that has been failing quietly +for a month is worse than a known absent one. `unknown` means the worker has not +run yet, or snapshots are disabled (`COTEL_SNAPSHOT_DIR` empty, which is the +default outside the compose file). + +To see what is actually on disk: + +```bash +docker run --rm -v cotel-snapshots:/snapshots debian:bookworm-slim \ + sh -c 'ls -1 /snapshots && du -sh /snapshots' +docker run --rm -v cotel-snapshots:/snapshots debian:bookworm-slim \ + cat /snapshots/2026-10-06T12-00-00Z/snapshot.json +``` + +## Restoring from a snapshot + +`cotel --db-import ` creates the database at `COTEL_DB_PATH` from a +snapshot directory and verifies every table against the snapshot's manifest. It +needs nothing but the image already on the host: no DuckDB CLI, no version +matching by hand (which is the step of the recovery procedure most able to +destroy a file). + +Two rules the flag enforces rather than documents: the target database file must +be **empty or absent** (the snapshot's `schema.sql` issues plain `CREATE TABLE`, +so importing over populated tables would fail halfway), and the snapshot must +carry its `snapshot.json`. + +**Step 1 — make a probe volume and do the restore there, never into the live +volume.** Creating a probe volume is +[step 3 of the recovery page](./duckdb-recovery#step-3-do-every-experiment-on-a-probe-copy); +for a restore it only has to be empty: + +```bash +docker volume create cotel-data-restore-probe +``` + +**Step 2 — import.** `--entrypoint` is not optional: the image's entrypoint +starts the server and would swallow the flag, leaving a *running cotel* writing +to the volume. The snapshots volume must be mounted at `/snapshots`, the path in +`load.sql`: + +```bash +docker run --rm --entrypoint /usr/local/bin/cotel \ + -v cotel-snapshots:/snapshots \ + -v cotel-data-restore-probe:/data \ + ghcr.io/flopsstuff/cotel:latest --db-import /snapshots/2026-10-06T12-00-00Z +# db-import: restored /data/cotel.duckdb from snapshot /snapshots/2026-10-06T12-00-00Z +# (taken 2026-10-06T12:00:00Z, schema_version 10, +# api_tokens=3 daily_usage=3048 settings=7 schema_version=10 spans=65184 users=17) +``` + +A row count that does not match the manifest fails the import with the table and +both numbers named. Nothing is written to the live database or to the snapshot at +any point. + +**Step 3 — verify with the same binary that will serve it.** + +```bash +docker run --rm --entrypoint /usr/local/bin/cotel -v cotel-data-restore-probe:/data \ + ghcr.io/flopsstuff/cotel:latest --db-query "SELECT count(*) FROM spans" +docker run --rm --entrypoint /usr/local/bin/cotel -v cotel-data-restore-probe:/data \ + ghcr.io/flopsstuff/cotel:latest --db-query "SELECT max(start_time) FROM spans" +docker run --rm --entrypoint /usr/local/bin/cotel -v cotel-data-restore-probe:/data \ + ghcr.io/flopsstuff/cotel:latest --db-query "SELECT count(*) FROM duckdb_indexes()" +``` + +The secondary indexes are **rebuilt from the data** by the import rather than +copied, which is why a snapshot of a database with a damaged index still +restores to a healthy one. Expect the restored file to be considerably smaller +than the live one, too: the import compacts away the free space retention churn +leaves behind. + +**Step 4 — promote.** Point the deploy at the restored volume instead of +overwriting the one you are restoring from; Docker has no `volume rename`, and +the volume you would overwrite is also your evidence: + +```bash +docker compose down +COTEL_DATA_VOLUME=cotel-data-restore-probe docker compose up -d +``` + +Make that variable permanent in the deploy's `.env` before you walk away — an +untracked `COTEL_DATA_VOLUME` on the command line is forgotten by the next +`docker compose up -d`, which then silently brings the old volume back. + +## A snapshot carries the same secrets as the database + +`users.token` holds ingest tokens in plaintext by design +([Users and Authentication](./users-and-auth)), so every snapshot carries a copy +of every agent's token. Treat the snapshots volume exactly like the data volume: +it is not a file to hand around, and an off-host copy of it is an off-host copy +of the credentials. + +## What snapshots do not cover + +- **Host loss.** Snapshots land on the same disk as production. They insure + against file-level corruption, a bad migration and accidental deletion — the + failures cotel has actually had — and not against the disk or the machine + going away. An off-host copy is a separate decision. +- **`docker volume prune`.** The snapshots volume is only "in use" while the + container exists, so a prune on a stopped deploy takes the backups with it. + This is the one new way the volume layout can bite; no code can prevent it. +- **Everything newer than the last snapshot.** With the shipped `6h` interval + the worst-case loss is six hours of spans. From 6d2980982294dd08b8afa227d62533d84b6dcd7e Mon Sep 17 00:00:00 2001 From: Wayland Date: Tue, 6 Oct 2026 19:23:20 +0200 Subject: [PATCH 4/6] fix(storage): fall back to the default when the snapshot interval is not positive A zero or negative COTEL_SNAPSHOT_INTERVAL made every cycle due the moment the last one finished, which is a busy loop on the one connection ingest uses. Also note in the operations page that a failed import leaves a half-populated file the refusal check then rejects on retry: discard the probe volume rather than trying to clean it. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- docs/operations/duckdb-snapshots.md | 8 ++++++++ internal/storage/snapshot.go | 6 ++++++ 2 files changed, 14 insertions(+) diff --git a/docs/operations/duckdb-snapshots.md b/docs/operations/duckdb-snapshots.md index 116ba0c..0a86887 100644 --- a/docs/operations/duckdb-snapshots.md +++ b/docs/operations/duckdb-snapshots.md @@ -117,6 +117,14 @@ A row count that does not match the manifest fails the import with the table and both numbers named. Nothing is written to the live database or to the snapshot at any point. +An import that fails partway leaves a half-populated file behind, and the +refusal above means retrying into it fails too. Throw the probe volume away and +make a new one rather than trying to clean it up: + +```bash +docker volume rm cotel-data-restore-probe && docker volume create cotel-data-restore-probe +``` + **Step 3 — verify with the same binary that will serve it.** ```bash diff --git a/internal/storage/snapshot.go b/internal/storage/snapshot.go index 04cfc13..28e4beb 100644 --- a/internal/storage/snapshot.go +++ b/internal/storage/snapshot.go @@ -89,6 +89,12 @@ func (db *DB) RunSnapshotWorker(cfg SnapshotConfig, interval time.Duration) { log.Printf("snapshots disabled: no snapshot directory configured") return } + // A non-positive interval would make every cycle due the moment the last + // one finished, which is a busy loop on the connection ingest uses. + if interval <= 0 { + log.Printf("warning: ignoring snapshot interval %s, using %s", interval, DefaultSnapshotInterval) + interval = DefaultSnapshotInterval + } log.Printf("snapshot worker: directory %s, interval %s, keeping %d", cfg.Dir, interval, cfg.Keep) consecutiveFailures := 0 From 419c9c89db456ee61adf017dd9ad5f1a7ada05bf Mon Sep 17 00:00:00 2001 From: Wayland Date: Tue, 6 Oct 2026 19:26:19 +0200 Subject: [PATCH 5/6] fix(storage): cap the snapshot wait at one interval A snapshot directory dated in the future - a snapshots volume carried over from a host whose clock ran fast - would otherwise defer the next snapshot by the whole skew rather than by the interval. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- internal/storage/snapshot.go | 5 ++++- internal/storage/snapshot_test.go | 6 ++++++ 2 files changed, 10 insertions(+), 1 deletion(-) diff --git a/internal/storage/snapshot.go b/internal/storage/snapshot.go index 28e4beb..b067f3d 100644 --- a/internal/storage/snapshot.go +++ b/internal/storage/snapshot.go @@ -128,7 +128,10 @@ func snapshotWait(dir string, interval time.Duration, now time.Time) time.Durati if !ok { return 0 } - return interval - now.Sub(newest) + // Capped at one interval so a future-dated directory - a volume carried over + // from a host with a fast clock - delays the next snapshot by at most the + // interval instead of by the skew. + return min(interval-now.Sub(newest), interval) } func newestSnapshotInstant(dir string) (time.Time, bool) { diff --git a/internal/storage/snapshot_test.go b/internal/storage/snapshot_test.go index 3cebf2a..5ba8cdc 100644 --- a/internal/storage/snapshot_test.go +++ b/internal/storage/snapshot_test.go @@ -300,6 +300,12 @@ func TestSnapshotWaitDefersUntilDue(t *testing.T) { if wait := snapshotWait(dir, 6*time.Hour, now); wait != 5*time.Hour { t.Errorf("wait must follow the newest snapshot, got %s, want 5h", wait) } + + // A directory dated in the future must not defer snapshots by the skew. + mkComplete(now.Add(72 * time.Hour)) + if wait := snapshotWait(dir, 6*time.Hour, now); wait != 6*time.Hour { + t.Errorf("wait with a future-dated snapshot = %s, want it capped at 6h", wait) + } } func TestRunSnapshotWorkerReturnsWhenDisabled(t *testing.T) { From d66a58aefe1614f5cd8f06c6999f5d5927da9278 Mon Sep 17 00:00:00 2001 From: Wayland Date: Tue, 6 Oct 2026 19:30:25 +0200 Subject: [PATCH 6/6] docs: plain hyphens in the prose this branch adds Punctuation only, confined to the 22 lines this branch introduced across the docs, the compose comment and the sidebar label. Surrounding text is untouched. Co-Authored-By: Wayland Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 6 +++--- README.md | 4 ++-- docker-compose.yml | 2 +- docs/.vitepress/config.js | 2 +- docs/index.md | 4 ++-- docs/operations/api-reference.md | 2 +- docs/operations/duckdb-recovery.md | 2 +- docs/operations/duckdb-snapshots.md | 22 +++++++++++----------- 8 files changed, 22 insertions(+), 22 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9d3b209..8571b52 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,9 +8,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] ### Added -- cotel snapshots its own database. A worker runs `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION ZSTD)` at startup and then every `COTEL_SNAPSHOT_INTERVAL`, into one dated directory per snapshot under `COTEL_SNAPSHOT_DIR` (`/snapshots` in compose, its own volume), keeping the `COTEL_SNAPSHOT_KEEP` newest. It had to be taken from inside cotel: DuckDB has one writer and the live process holds the file lock, so no external process can open the database even read-only. Measured on the production host against a probe copy of the live 152.6 MB database, the export holds the single connection for 0.15-0.3 s and writes 4.2 MiB — 1/35th of the file for a full copy of all six tables — and because it runs on that same connection it cannot race the WAL checkpoint: the two are serialised by construction rather than by a lock. The format is the point: Parquet plus a plain-text `schema.sql` is readable by any DuckDB build, and an import *rebuilds* the secondary indexes from the data, so a snapshot of a database with a damaged ART index — the September 2026 failure, a file whose rows were all readable — restores to a healthy one, where a byte copy of the volume reproduces the damage faithfully. `snapshot.json` is written last and by nothing else, so its presence is the only "complete" signal; it carries the row count per table, read back out of the Parquet files, so a restore is checked against what the snapshot claims rather than trusted. A snapshot is only taken when the newest complete one is older than the interval, so a restart — or a crash loop — cannot spend the retained window on snapshots minutes apart. Pruning keeps the newest N complete snapshots, deletes incomplete ones, runs only after a successful export, and leaves directories whose name is not a snapshot instant alone ([ADR-0023](docs/decisions/0023-production-database-snapshots.md), [docs/operations/duckdb-snapshots.md](docs/operations/duckdb-snapshots.md)) -- `cotel --db-import ` restores a snapshot into `COTEL_DB_PATH` and verifies every table against the snapshot's manifest, failing with the table and both numbers named if a count does not match. It needs nothing but the image already on the host — no DuckDB CLI with its version matched by hand, which is the step of the recovery procedure most able to destroy a file. It refuses a target that already holds tables (the snapshot's `schema.sql` issues plain `CREATE TABLE`, so importing over data would fail halfway and leave a mixed database) and a snapshot with no `snapshot.json`. Note that the paths in a snapshot's `load.sql` are **absolute**, so a snapshot cannot be moved or renamed and still be imported; it must be visible at the path it was written to -- `GET /api/v1/health` reports a `snapshot` object (`status`, `last_run_at`, `last_error`, `last_dir`), and a failed export flips the top-level `status` to `degraded` — the same treatment a failing retention roll-up gets, because a backup that has been failing quietly for a month is worse than a known absent one. `unknown` covers both a worker that has not run yet and snapshots left disabled, which is the default outside the compose file +- cotel snapshots its own database. A worker runs `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION ZSTD)` at startup and then every `COTEL_SNAPSHOT_INTERVAL`, into one dated directory per snapshot under `COTEL_SNAPSHOT_DIR` (`/snapshots` in compose, its own volume), keeping the `COTEL_SNAPSHOT_KEEP` newest. It had to be taken from inside cotel: DuckDB has one writer and the live process holds the file lock, so no external process can open the database even read-only. Measured on the production host against a probe copy of the live 152.6 MB database, the export holds the single connection for 0.15-0.3 s and writes 4.2 MiB - 1/35th of the file for a full copy of all six tables - and because it runs on that same connection it cannot race the WAL checkpoint: the two are serialised by construction rather than by a lock. The format is the point: Parquet plus a plain-text `schema.sql` is readable by any DuckDB build, and an import *rebuilds* the secondary indexes from the data, so a snapshot of a database with a damaged ART index - the September 2026 failure, a file whose rows were all readable - restores to a healthy one, where a byte copy of the volume reproduces the damage faithfully. `snapshot.json` is written last and by nothing else, so its presence is the only "complete" signal; it carries the row count per table, read back out of the Parquet files, so a restore is checked against what the snapshot claims rather than trusted. A snapshot is only taken when the newest complete one is older than the interval, so a restart - or a crash loop - cannot spend the retained window on snapshots minutes apart. Pruning keeps the newest N complete snapshots, deletes incomplete ones, runs only after a successful export, and leaves directories whose name is not a snapshot instant alone ([ADR-0023](docs/decisions/0023-production-database-snapshots.md), [docs/operations/duckdb-snapshots.md](docs/operations/duckdb-snapshots.md)) +- `cotel --db-import ` restores a snapshot into `COTEL_DB_PATH` and verifies every table against the snapshot's manifest, failing with the table and both numbers named if a count does not match. It needs nothing but the image already on the host - no DuckDB CLI with its version matched by hand, which is the step of the recovery procedure most able to destroy a file. It refuses a target that already holds tables (the snapshot's `schema.sql` issues plain `CREATE TABLE`, so importing over data would fail halfway and leave a mixed database) and a snapshot with no `snapshot.json`. Note that the paths in a snapshot's `load.sql` are **absolute**, so a snapshot cannot be moved or renamed and still be imported; it must be visible at the path it was written to +- `GET /api/v1/health` reports a `snapshot` object (`status`, `last_run_at`, `last_error`, `last_dir`), and a failed export flips the top-level `status` to `degraded` - the same treatment a failing retention roll-up gets, because a backup that has been failing quietly for a month is worse than a known absent one. `unknown` covers both a worker that has not run yet and snapshots left disabled, which is the default outside the compose file - The Overview leads with a **Span activity** grid — a block of cells counting spans, GitHub-contribution-graph style, sitting directly under the KPI row. A line chart answers *how much and when*; it does not answer *what does a week here look like*, which is the question a telemetry front door gets asked most. The selected range picks both the grid and how much time one cell is: 53 × 7 day cells over a year (and over `All`), 31 × 6 four-hour cells over a month, 24 × 7 hourly cells over a week, 24 × 6 ten-minute cells over a day — each tiling its window exactly, and each about 115 px tall, so switching range does not move the page under the reader. Cells outside the queried window — the leading edge of the lattice, and the rest of today — are drawn as an outline with no fill: an empty cell means "we looked and there was nothing", an outline means "we did not look", and conflating the two is how a heatmap invents a quiet weekend. The grid is placed in UTC, which the footer and every tooltip say. Intensity is cut at the quartiles of the cells in view, not scaled against the busiest one: against a 722-span peak a 200-span day and a 700-span day are both "busy", so a max-relative ramp — linear or log — renders a working week as one flat block of full-intensity cells, which is the difference the grid exists to show. A step therefore means a rank, so the footer names the busiest cell in view and every tooltip gives the cell's own count. The scale lives in `frontend/src/lib/heat.ts` and is shared with the History page's calendar and hour-of-day heatmaps, which had a private copy of it and pick up the quartile cut with this change ([ADR-0016](docs/decisions/0016-overview-activity-grid.md)) - `GET /api/v1/history` accepts two more bucket widths, `granularity=10m` and `granularity=4h`, so the activity grid asks for exactly the width it draws — one bucket, one cell — instead of re-bucketing an hourly series in the page, which could not have produced a ten-minute cell at all. Both are additive and no existing caller changes; an unrecognised width still falls back to `day` rather than 400ing. Like `hour` they are answered from `spans` alone and report the shortfall in `covered_since`, because `daily_usage` buckets whole UTC days and cannot produce a sub-day bucket. `bucket` is now documented as a UTC wall-clock label floored to the width, on any host, so a client can reconstruct it for an instant without asking what the server thinks midnight is - `scripts/seed-demo.py` fills a throwaway instance with a synthetic team — seven users, three models, 90 days of sessions and tool calls. It goes in over the OTLP endpoint rather than writing to DuckDB, so a seeded instance exercises the same ingest, cost-derivation and roll-up path a real one does, and it sends only attributes Claude Code actually sends: no `command` on `Bash` spans, so the Tools page shows the same "no command detail" state a real install sees. The RNG seed is fixed, so a re-run against a fresh volume reproduces the same numbers. `scripts/shoot-screenshots.mjs` turns that instance into the README images, each cropped at the bottom edge of a named element rather than at a pixel count ([docs/operations/screenshots.md](docs/operations/screenshots.md)) diff --git a/README.md b/README.md index 493f9b4..d4ec572 100644 --- a/README.md +++ b/README.md @@ -520,10 +520,10 @@ from the network at startup instead of bundling it). | `COTEL_RETENTION_RAW_DAYS` | `30` | Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer) | | `COTEL_RETENTION_AGGREGATE_DAYS` | `90` | Daily aggregate retention in days | | `COTEL_RETENTION_INTERVAL` | `6h` | Retention worker tick interval (Go duration) | -| `COTEL_SNAPSHOT_DIR` | _(unset — snapshots off)_ | Directory the snapshot worker exports the whole database into, one dated subdirectory per snapshot. Empty disables snapshots; `docker-compose.yml` sets `/snapshots`, backed by its own volume. See [Database Snapshots and Restore](docs/operations/duckdb-snapshots.md). | +| `COTEL_SNAPSHOT_DIR` | _(unset - snapshots off)_ | Directory the snapshot worker exports the whole database into, one dated subdirectory per snapshot. Empty disables snapshots; `docker-compose.yml` sets `/snapshots`, backed by its own volume. See [Database Snapshots and Restore](docs/operations/duckdb-snapshots.md). | | `COTEL_SNAPSHOT_INTERVAL` | `6h` | How often a snapshot is taken (Go duration). A snapshot is only taken when the newest complete one is older than this, so a restart cannot churn through the retained window. | | `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep; older ones and incomplete ones are pruned after each successful export. At the default interval, 56 is 14 days of reach for about 235 MB. The last snapshot standing is never pruned. | -| `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not by the binary: the Docker volume mounted at `/snapshots`. Note that `docker volume prune` on a stopped deploy deletes it — the volume counts as in use only while the container exists. | +| `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not by the binary: the Docker volume mounted at `/snapshots`. Note that `docker volume prune` on a stopped deploy deletes it - the volume counts as in use only while the container exists. | | `COTEL_WAL_AUTOCHECKPOINT` | `4MB` | DuckDB `checkpoint_threshold`: the write-ahead log is folded into the main file once it grows past this size. Lower values bound how much WAL an ungraceful kill leaves to replay on the next open; higher values checkpoint less often during ingest. DuckDB's own default is `16MB`. | | `CLOUDFLARE_TUNNEL_TOKEN` | _(unset)_ | When set, starts `cloudflared tunnel run` before cotel; enables public HTTPS access via Cloudflare Tunnel | | `TUNNEL_EDGE_IP_VERSION` | `4` in token mode | Read by `cloudflared`, not by cotel: the address family used to reach the Cloudflare edge (`4`, `6` or `auto`). cloudflared's own default became `auto` in 2026.4.0, which tries whichever family the resolver answers with first and falls back only after a connection has failed; token mode pins `4` unless you set this. Not set in local-config mode, where `config.yml` owns the setting. See [token mode](docs/operations/cloudflare-tunnel-remote.md#the-bundled-cloudflared) | diff --git a/docker-compose.yml b/docker-compose.yml index 7f1c45f..14257b9 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -49,7 +49,7 @@ volumes: name: ${COTEL_DATA_VOLUME:-cotel-data-repaired-20261004} # Named for the same reason as cotel-data: a restore brings the deploy up # against a different data volume while this one stays put. Note that - # `docker volume prune` on a stopped deploy takes the backups with it — this + # `docker volume prune` on a stopped deploy takes the backups with it - this # volume is only "in use" while the container exists. cotel-snapshots: name: ${COTEL_SNAPSHOT_VOLUME:-cotel-snapshots} diff --git a/docs/.vitepress/config.js b/docs/.vitepress/config.js index 9a59ba0..890e66a 100644 --- a/docs/.vitepress/config.js +++ b/docs/.vitepress/config.js @@ -63,7 +63,7 @@ export default defineConfig({ { text: 'ADR-0020 — Recovery Arrives as a New Issue (superseded)', link: '/decisions/0020-recovery-arrives-as-a-new-issue' }, { text: "ADR-0021 — Recovery Wakes the Alert's Assignee", link: '/decisions/0021-recovery-wakes-the-alerts-assignee' }, { text: 'ADR-0022 — Health Probe Scheduler Outside This Repo', link: '/decisions/0022-health-probe-scheduler-outside-github' }, - { text: 'ADR-0023 — Production Database Snapshots', link: '/decisions/0023-production-database-snapshots' }, + { text: 'ADR-0023 - Production Database Snapshots', link: '/decisions/0023-production-database-snapshots' }, ], }, ], diff --git a/docs/index.md b/docs/index.md index 922fa2f..596e1fd 100644 --- a/docs/index.md +++ b/docs/index.md @@ -69,9 +69,9 @@ Restart Claude Code. Telemetry starts flowing immediately. | `COTEL_RETENTION_RAW_DAYS` | `30` | Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer) | | `COTEL_RETENTION_AGGREGATE_DAYS` | `90` | Daily aggregate retention in days | | `COTEL_RETENTION_INTERVAL` | `6h` | Retention worker tick interval | -| `COTEL_SNAPSHOT_DIR` | _(unset — snapshots off)_ | Where the snapshot worker exports the whole database (compose sets `/snapshots`); empty disables snapshots. See [Database Snapshots and Restore](./operations/duckdb-snapshots) | +| `COTEL_SNAPSHOT_DIR` | _(unset - snapshots off)_ | Where the snapshot worker exports the whole database (compose sets `/snapshots`); empty disables snapshots. See [Database Snapshots and Restore](./operations/duckdb-snapshots) | | `COTEL_SNAPSHOT_INTERVAL` | `6h` | How often a snapshot is taken (Go duration) | -| `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep — 14 days at the default interval | +| `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep - 14 days at the default interval | | `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not the binary: the Docker volume mounted at `/snapshots` | ## Data & retention diff --git a/docs/operations/api-reference.md b/docs/operations/api-reference.md index c99eb2c..92695b7 100644 --- a/docs/operations/api-reference.md +++ b/docs/operations/api-reference.md @@ -256,7 +256,7 @@ Instance health. Takes no parameters and is never range-scoped. | `newest_span_age_seconds` | integer \| `null` | Seconds since `last_ingest_at`; `null` if nothing was ever ingested | | `db_size_bytes` | integer | Approximate database file size | | `retention` | object | `status` (`ok` \| `error` \| `unknown`), `last_run_at`, `last_error` | -| `snapshot` | object | `status` (`ok` \| `error` \| `unknown`), `last_run_at`, `last_error`, `last_dir` — the last complete snapshot's directory. `unknown` covers both "has not run yet" and snapshots disabled ([Database Snapshots and Restore](./duckdb-snapshots)) | +| `snapshot` | object | `status` (`ok` \| `error` \| `unknown`), `last_run_at`, `last_error`, `last_dir` - the last complete snapshot's directory. `unknown` covers both "has not run yet" and snapshots disabled ([Database Snapshots and Restore](./duckdb-snapshots)) | | `public_ingest_url` | string | Omitted unless `COTEL_PUBLIC_INGEST_URL` is set | The two freshness fields are measured from the span's `ingested_at`, not its diff --git a/docs/operations/duckdb-recovery.md b/docs/operations/duckdb-recovery.md index 1058fb0..9268885 100644 --- a/docs/operations/duckdb-recovery.md +++ b/docs/operations/duckdb-recovery.md @@ -108,7 +108,7 @@ docker run --rm -it -v cotel-data-probe:/data debian:bookworm-slim sh -c " The ARM asset was renamed between releases: DuckDB up to and including v1.2.x publishes `duckdb_cli-linux-aarch64.zip`, v1.3.0 and later publish `duckdb_cli-linux-arm64.zip`. Using the wrong one gets a 404 from the release, -not a wrong binary, so it is a nuisance rather than a hazard — but check the +not a wrong binary, so it is a nuisance rather than a hazard - but check the release's asset list rather than guessing. Confirm `duckdb --version` prints the same version as `SELECT version()` above before you run a single statement against the file. diff --git a/docs/operations/duckdb-snapshots.md b/docs/operations/duckdb-snapshots.md index 0a86887..c94718b 100644 --- a/docs/operations/duckdb-snapshots.md +++ b/docs/operations/duckdb-snapshots.md @@ -28,14 +28,14 @@ Three properties are worth knowing before you rely on it: - **`snapshot.json` is written last, and nothing else writes it.** A directory without it is a failed or interrupted run: prunable, never restorable. It records the instant, the export duration, the schema version and the row count - per table, read back out of the Parquet files themselves — so a restore can be + per table, read back out of the Parquet files themselves - so a restore can be checked against what the snapshot claims to hold rather than trusted. - **A snapshot is only taken when the newest complete one is older than the - interval.** A restart — or a crash loop — therefore cannot spend the retained + interval.** A restart - or a crash loop - therefore cannot spend the retained window on snapshots minutes apart. - **The paths inside `load.sql` are absolute.** DuckDB bakes the export directory's path into it, so a snapshot cannot be moved or renamed and still - be imported. It must be visible at the path it was written to — `/snapshots` + be imported. It must be visible at the path it was written to - `/snapshots` in the shipped compose file. Pruning keeps the `COTEL_SNAPSHOT_KEEP` newest complete snapshots, deletes @@ -61,7 +61,7 @@ curl -s localhost:8080/api/v1/health | jq '{status, snapshot}' ``` `snapshot.status` is `error` after a failed cycle, which also flips the -top-level `status` to `degraded` — the same treatment a failing retention +top-level `status` to `degraded` - the same treatment a failing retention roll-up gets, and for the same reason: a backup that has been failing quietly for a month is worse than a known absent one. `unknown` means the worker has not run yet, or snapshots are disabled (`COTEL_SNAPSHOT_DIR` empty, which is the @@ -89,7 +89,7 @@ be **empty or absent** (the snapshot's `schema.sql` issues plain `CREATE TABLE`, so importing over populated tables would fail halfway), and the snapshot must carry its `snapshot.json`. -**Step 1 — make a probe volume and do the restore there, never into the live +**Step 1 - make a probe volume and do the restore there, never into the live volume.** Creating a probe volume is [step 3 of the recovery page](./duckdb-recovery#step-3-do-every-experiment-on-a-probe-copy); for a restore it only has to be empty: @@ -98,7 +98,7 @@ for a restore it only has to be empty: docker volume create cotel-data-restore-probe ``` -**Step 2 — import.** `--entrypoint` is not optional: the image's entrypoint +**Step 2 - import.** `--entrypoint` is not optional: the image's entrypoint starts the server and would swallow the flag, leaving a *running cotel* writing to the volume. The snapshots volume must be mounted at `/snapshots`, the path in `load.sql`: @@ -125,7 +125,7 @@ make a new one rather than trying to clean it up: docker volume rm cotel-data-restore-probe && docker volume create cotel-data-restore-probe ``` -**Step 3 — verify with the same binary that will serve it.** +**Step 3 - verify with the same binary that will serve it.** ```bash docker run --rm --entrypoint /usr/local/bin/cotel -v cotel-data-restore-probe:/data \ @@ -142,7 +142,7 @@ restores to a healthy one. Expect the restored file to be considerably smaller than the live one, too: the import compacts away the free space retention churn leaves behind. -**Step 4 — promote.** Point the deploy at the restored volume instead of +**Step 4 - promote.** Point the deploy at the restored volume instead of overwriting the one you are restoring from; Docker has no `volume rename`, and the volume you would overwrite is also your evidence: @@ -151,7 +151,7 @@ docker compose down COTEL_DATA_VOLUME=cotel-data-restore-probe docker compose up -d ``` -Make that variable permanent in the deploy's `.env` before you walk away — an +Make that variable permanent in the deploy's `.env` before you walk away - an untracked `COTEL_DATA_VOLUME` on the command line is forgotten by the next `docker compose up -d`, which then silently brings the old volume back. @@ -166,8 +166,8 @@ of the credentials. ## What snapshots do not cover - **Host loss.** Snapshots land on the same disk as production. They insure - against file-level corruption, a bad migration and accidental deletion — the - failures cotel has actually had — and not against the disk or the machine + against file-level corruption, a bad migration and accidental deletion - the + failures cotel has actually had - and not against the disk or the machine going away. An off-host copy is a separate decision. - **`docker volume prune`.** The snapshots volume is only "in use" while the container exists, so a prune on a stopped deploy takes the backups with it.