Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Added
- cotel snapshots its own database. A worker runs `EXPORT DATABASE ... (FORMAT PARQUET, COMPRESSION ZSTD)` at startup and then every `COTEL_SNAPSHOT_INTERVAL`, into one dated directory per snapshot under `COTEL_SNAPSHOT_DIR` (`/snapshots` in compose, its own volume), keeping the `COTEL_SNAPSHOT_KEEP` newest. It had to be taken from inside cotel: DuckDB has one writer and the live process holds the file lock, so no external process can open the database even read-only. Measured on the production host against a probe copy of the live 152.6 MB database, the export holds the single connection for 0.15-0.3 s and writes 4.2 MiB - 1/35th of the file for a full copy of all six tables - and because it runs on that same connection it cannot race the WAL checkpoint: the two are serialised by construction rather than by a lock. The format is the point: Parquet plus a plain-text `schema.sql` is readable by any DuckDB build, and an import *rebuilds* the secondary indexes from the data, so a snapshot of a database with a damaged ART index - the September 2026 failure, a file whose rows were all readable - restores to a healthy one, where a byte copy of the volume reproduces the damage faithfully. `snapshot.json` is written last and by nothing else, so its presence is the only "complete" signal; it carries the row count per table, read back out of the Parquet files, so a restore is checked against what the snapshot claims rather than trusted. A snapshot is only taken when the newest complete one is older than the interval, so a restart - or a crash loop - cannot spend the retained window on snapshots minutes apart. Pruning keeps the newest N complete snapshots, deletes incomplete ones, runs only after a successful export, and leaves directories whose name is not a snapshot instant alone ([ADR-0023](docs/decisions/0023-production-database-snapshots.md), [docs/operations/duckdb-snapshots.md](docs/operations/duckdb-snapshots.md))
- `cotel --db-import <dir>` restores a snapshot into `COTEL_DB_PATH` and verifies every table against the snapshot's manifest, failing with the table and both numbers named if a count does not match. It needs nothing but the image already on the host - no DuckDB CLI with its version matched by hand, which is the step of the recovery procedure most able to destroy a file. It refuses a target that already holds tables (the snapshot's `schema.sql` issues plain `CREATE TABLE`, so importing over data would fail halfway and leave a mixed database) and a snapshot with no `snapshot.json`. Note that the paths in a snapshot's `load.sql` are **absolute**, so a snapshot cannot be moved or renamed and still be imported; it must be visible at the path it was written to
- `GET /api/v1/health` reports a `snapshot` object (`status`, `last_run_at`, `last_error`, `last_dir`), and a failed export flips the top-level `status` to `degraded` - the same treatment a failing retention roll-up gets, because a backup that has been failing quietly for a month is worse than a known absent one. `unknown` covers both a worker that has not run yet and snapshots left disabled, which is the default outside the compose file
- The Overview leads with a **Span activity** grid — a block of cells counting spans, GitHub-contribution-graph style, sitting directly under the KPI row. A line chart answers *how much and when*; it does not answer *what does a week here look like*, which is the question a telemetry front door gets asked most. The selected range picks both the grid and how much time one cell is: 53 × 7 day cells over a year (and over `All`), 31 × 6 four-hour cells over a month, 24 × 7 hourly cells over a week, 24 × 6 ten-minute cells over a day — each tiling its window exactly, and each about 115 px tall, so switching range does not move the page under the reader. Cells outside the queried window — the leading edge of the lattice, and the rest of today — are drawn as an outline with no fill: an empty cell means "we looked and there was nothing", an outline means "we did not look", and conflating the two is how a heatmap invents a quiet weekend. The grid is placed in UTC, which the footer and every tooltip say. Intensity is cut at the quartiles of the cells in view, not scaled against the busiest one: against a 722-span peak a 200-span day and a 700-span day are both "busy", so a max-relative ramp — linear or log — renders a working week as one flat block of full-intensity cells, which is the difference the grid exists to show. A step therefore means a rank, so the footer names the busiest cell in view and every tooltip gives the cell's own count. The scale lives in `frontend/src/lib/heat.ts` and is shared with the History page's calendar and hour-of-day heatmaps, which had a private copy of it and pick up the quartile cut with this change ([ADR-0016](docs/decisions/0016-overview-activity-grid.md))
- `GET /api/v1/history` accepts two more bucket widths, `granularity=10m` and `granularity=4h`, so the activity grid asks for exactly the width it draws — one bucket, one cell — instead of re-bucketing an hourly series in the page, which could not have produced a ten-minute cell at all. Both are additive and no existing caller changes; an unrecognised width still falls back to `day` rather than 400ing. Like `hour` they are answered from `spans` alone and report the shortfall in `covered_since`, because `daily_usage` buckets whole UTC days and cannot produce a sub-day bucket. `bucket` is now documented as a UTC wall-clock label floored to the width, on any host, so a client can reconstruct it for an instant without asking what the server thinks midnight is
- `scripts/seed-demo.py` fills a throwaway instance with a synthetic team — seven users, three models, 90 days of sessions and tool calls. It goes in over the OTLP endpoint rather than writing to DuckDB, so a seeded instance exercises the same ingest, cost-derivation and roll-up path a real one does, and it sends only attributes Claude Code actually sends: no `command` on `Bash` spans, so the Tools page shows the same "no command detail" state a real install sees. The RNG seed is fixed, so a re-run against a fresh volume reproduces the same numbers. `scripts/shoot-screenshots.mjs` turns that instance into the README images, each cropped at the bottom edge of a named element rather than at a pixel count ([docs/operations/screenshots.md](docs/operations/screenshots.md))
Expand Down
29 changes: 29 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -385,6 +385,31 @@ docker run --rm -v cotel-data:/data ubuntu \
duckdb /data/cotel.duckdb "SELECT model, COUNT(*) FROM spans GROUP BY model"
```

## Snapshots

cotel exports its whole database to Parquet on a timer and keeps the newest
`COTEL_SNAPSHOT_KEEP` exports in a second volume (`/snapshots`), so there is a
recovery point that does not depend on the live DuckDB file being readable. The
export runs inside cotel, on the same connection as everything else, and costs
0.15-0.3 s per run on a 152 MB database; the format is portable Parquet plus a
plain-text `schema.sql`, which any DuckDB build can read
([ADR-0023](docs/decisions/0023-production-database-snapshots.md)).

```bash
# what is on disk
docker run --rm -v cotel-snapshots:/snapshots debian:bookworm-slim ls -1 /snapshots

# restore one into an empty volume, verified against the snapshot's manifest
docker run --rm --entrypoint /usr/local/bin/cotel \
-v cotel-snapshots:/snapshots -v cotel-data-restore:/data \
ghcr.io/flopsstuff/cotel:latest --db-import /snapshots/2026-10-06T12-00-00Z
```

The worker's last outcome is reported on `GET /api/v1/health` under a `snapshot`
object, and a failed export degrades the top-level `status`. Full procedure,
including promoting a restored volume:
[Database Snapshots and Restore](docs/operations/duckdb-snapshots.md).

## Retention defaults

| Tier | Period | Storage |
Expand Down Expand Up @@ -495,6 +520,10 @@ from the network at startup instead of bundling it).
| `COTEL_RETENTION_RAW_DAYS` | `30` | Raw span retention in days (roll-up consumes whole days, so spans survive up to a day longer) |
| `COTEL_RETENTION_AGGREGATE_DAYS` | `90` | Daily aggregate retention in days |
| `COTEL_RETENTION_INTERVAL` | `6h` | Retention worker tick interval (Go duration) |
| `COTEL_SNAPSHOT_DIR` | _(unset - snapshots off)_ | Directory the snapshot worker exports the whole database into, one dated subdirectory per snapshot. Empty disables snapshots; `docker-compose.yml` sets `/snapshots`, backed by its own volume. See [Database Snapshots and Restore](docs/operations/duckdb-snapshots.md). |
| `COTEL_SNAPSHOT_INTERVAL` | `6h` | How often a snapshot is taken (Go duration). A snapshot is only taken when the newest complete one is older than this, so a restart cannot churn through the retained window. |
| `COTEL_SNAPSHOT_KEEP` | `56` | How many complete snapshots to keep; older ones and incomplete ones are pruned after each successful export. At the default interval, 56 is 14 days of reach for about 235 MB. The last snapshot standing is never pruned. |
| `COTEL_SNAPSHOT_VOLUME` | `cotel-snapshots` | Read by `docker-compose.yml`, not by the binary: the Docker volume mounted at `/snapshots`. Note that `docker volume prune` on a stopped deploy deletes it - the volume counts as in use only while the container exists. |
| `COTEL_WAL_AUTOCHECKPOINT` | `4MB` | DuckDB `checkpoint_threshold`: the write-ahead log is folded into the main file once it grows past this size. Lower values bound how much WAL an ungraceful kill leaves to replay on the next open; higher values checkpoint less often during ingest. DuckDB's own default is `16MB`. |
| `CLOUDFLARE_TUNNEL_TOKEN` | _(unset)_ | When set, starts `cloudflared tunnel run` before cotel; enables public HTTPS access via Cloudflare Tunnel |
| `TUNNEL_EDGE_IP_VERSION` | `4` in token mode | Read by `cloudflared`, not by cotel: the address family used to reach the Cloudflare edge (`4`, `6` or `auto`). cloudflared's own default became `auto` in 2026.4.0, which tries whichever family the resolver answers with first and falls back only after a connection has failed; token mode pins `4` unless you set this. Not set in local-config mode, where `config.yml` owns the setting. See [token mode](docs/operations/cloudflare-tunnel-remote.md#the-bundled-cloudflared) |
Expand Down
35 changes: 35 additions & 0 deletions cmd/cotel/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ import (
"net/url"
"os"
"os/signal"
"sort"
"strconv"
"strings"
"sync/atomic"
Expand All @@ -28,6 +29,7 @@ import (

func main() {
dbQuery := flag.String("db-query", "", "run SQL query against DuckDB, print first column of first row, and exit")
dbImport := flag.String("db-import", "", "restore COTEL_DB_PATH from the snapshot directory given, verify it against the snapshot manifest, and exit; the target database file must be empty or absent")
healthcheck := flag.Bool("healthcheck", false, "probe the local dashboard /healthz and exit 0 (ready) or 1; used by the container HEALTHCHECK")
flag.Parse()

Expand All @@ -53,6 +55,18 @@ func main() {
return
}

// Restore runs before the listeners bind: it needs the database file to
// itself, and nothing should be able to ingest into a half-imported file.
if *dbImport != "" {
m, err := storage.ImportSnapshot(dbPath, *dbImport)
if err != nil {
log.Fatalf("db-import: %v", err)
}
log.Printf("db-import: restored %s from snapshot %s (taken %s, schema_version %d, %s)",
dbPath, *dbImport, m.Instant, m.SchemaVersion, formatTableCounts(m.Tables))
return
}

// Bind BEFORE storage.Open: WAL replay + schema migration can block for
// minutes on a large DB, and a bound port answering a retryable 503 keeps
// OTLP clients retrying where a connection reset would drop their spans.
Expand Down Expand Up @@ -98,6 +112,12 @@ func main() {
retentionInterval := envDuration("COTEL_RETENTION_INTERVAL", 6*time.Hour)
go db.RunRetentionWorker(retentionCfg, retentionInterval)

snapshotCfg := storage.SnapshotConfig{
Dir: os.Getenv("COTEL_SNAPSHOT_DIR"),
Keep: envInt("COTEL_SNAPSHOT_KEEP", storage.DefaultSnapshotKeep),
}
go db.RunSnapshotWorker(snapshotCfg, envDuration("COTEL_SNAPSHOT_INTERVAL", storage.DefaultSnapshotInterval))

ingestMux := http.NewServeMux()
ingestMux.Handle("/v1/traces", auth.Middleware(db, ingest.New(db)))

Expand Down Expand Up @@ -335,6 +355,21 @@ func runHealthcheck(dashAddr string) int {
return 0
}

// formatTableCounts renders a snapshot manifest's row counts in a stable order,
// so two restores of the same snapshot log the same line.
func formatTableCounts(tables map[string]int64) string {
names := make([]string, 0, len(tables))
for name := range tables {
names = append(names, name)
}
sort.Strings(names)
parts := make([]string, 0, len(names))
for _, name := range names {
parts = append(parts, fmt.Sprintf("%s=%d", name, tables[name]))
}
return strings.Join(parts, " ")
}

func env(key, fallback string) string {
if v := os.Getenv(key); v != "" {
return v
Expand Down
15 changes: 15 additions & 0 deletions docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@ services:
- "8080:8080" # Dashboard
volumes:
- cotel-data:/data
# Snapshots of the whole database (ADR-0023). Mounted at the path baked
# into each snapshot's load.sql, which is what a restore replays.
- cotel-snapshots:/snapshots
# Locally-managed Cloudflare tunnel (optional).
# Uncomment and adjust the host path if you prefer file-based tunnel config
# over the CLOUDFLARE_TUNNEL_TOKEN env var. Mount your ~/.cloudflared/
Expand All @@ -27,6 +30,12 @@ services:
# get the correct endpoint to paste into ~/.claude/settings.json.
# Example: COTEL_PUBLIC_INGEST_URL=https://otlp.example.com
COTEL_PUBLIC_INGEST_URL: ${COTEL_PUBLIC_INGEST_URL:-}
# Database snapshots. An empty COTEL_SNAPSHOT_DIR disables them, which is
# the default outside this compose file; here they are always on, because
# a deploy without a recovery point is how production ended up with none.
COTEL_SNAPSHOT_DIR: ${COTEL_SNAPSHOT_DIR:-/snapshots}
COTEL_SNAPSHOT_INTERVAL: ${COTEL_SNAPSHOT_INTERVAL:-6h}
COTEL_SNAPSHOT_KEEP: ${COTEL_SNAPSHOT_KEEP:-56}

volumes:
# Named explicitly so the deploy can be pointed at a different volume without
Expand All @@ -38,3 +47,9 @@ volumes:
# warns that it did not create a hand-made volume; that warning is expected.
cotel-data:
name: ${COTEL_DATA_VOLUME:-cotel-data-repaired-20261004}
# Named for the same reason as cotel-data: a restore brings the deploy up
# against a different data volume while this one stays put. Note that
# `docker volume prune` on a stopped deploy takes the backups with it - this
# volume is only "in use" while the container exists.
cotel-snapshots:
name: ${COTEL_SNAPSHOT_VOLUME:-cotel-snapshots}
2 changes: 2 additions & 0 deletions docs/.vitepress/config.js
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ export default defineConfig({
{ text: 'Production /healthz probe', link: '/operations/health-probe' },
{ text: 'Export / Import', link: '/operations/export-import' },
{ text: 'DuckDB Recovery', link: '/operations/duckdb-recovery' },
{ text: 'Database Snapshots and Restore', link: '/operations/duckdb-snapshots' },
{ text: 'README Screenshots', link: '/operations/screenshots' },
],
},
Expand Down Expand Up @@ -62,6 +63,7 @@ export default defineConfig({
{ text: 'ADR-0020 — Recovery Arrives as a New Issue (superseded)', link: '/decisions/0020-recovery-arrives-as-a-new-issue' },
{ text: "ADR-0021 — Recovery Wakes the Alert's Assignee", link: '/decisions/0021-recovery-wakes-the-alerts-assignee' },
{ text: 'ADR-0022 — Health Probe Scheduler Outside This Repo', link: '/decisions/0022-health-probe-scheduler-outside-github' },
{ text: 'ADR-0023 - Production Database Snapshots', link: '/decisions/0023-production-database-snapshots' },
],
},
],
Expand Down
Loading
Loading