Skip to content

Latest commit

 

History

History
297 lines (246 loc) · 15.3 KB

File metadata and controls

297 lines (246 loc) · 15.3 KB

Upgrade and rollback guide

This guide covers upgrades of a SQLite/libSQL database already containing Timeless virtual tables and upgrades of the three Rust signal servers. It does not turn the server binaries into readers for an unrelated Rust block-store directory; that conversion belongs to the higher-order product that owns the legacy format and must write through the public Timeless batch/SQL contracts.

Use the publication status to select a published bundle and the compatibility contract for the current source versions and pairing floors. Never mix archives from different versions. Changes marked unreleased below require builds from main; they are not claims about the latest downloadable binaries. Historical versions in this guide describe migration boundaries, not installation recommendations.

Invariants

  • Upgrade the telemetry extension and all deployed Rust signal binaries as one compatibility set.
  • Stop the current owner before replacing a database or extension artifact.
  • Keep the pre-upgrade database and its coordinated backup until rollback is no longer required.
  • Never copy, delete, rename, or selectively import private shadow tables.
  • Never test a downgrade against the only copy of production data.
  • A green --version string is not enough; require the complete capability, database-ledger, readiness, and semantic checks below.

Opening an older timeless_traces table with this extension adds a private duration-extrema side table. Existing payloads and indexes are not rewritten during startup. Unknown blocks remain exactly queryable through the decode fallback; schedule INSERT INTO traces(traces) VALUES ('optimize') (or the bounded optimize:<max source spans> form) to populate them. Observe duration_unknown_blocks reaching zero through timeless_stats('traces'). Keep the pre-upgrade backup: an older extension is not expected to understand the additive private schema even though stored block payloads retain their codec and data-ABI meaning.

Starting with 0.5.0, rich-logs optimize() writes template-compressed message blocks (codec byte 8). Nothing is rewritten at open — the new codec appears only as background or scheduled optimize re-encodes rich blocks — but once it has, a pre-0.5.0 extension refuses those blocks with a loud decode error naming the unknown codec; it never returns partial or flattened rows. Rolling back the extension after a 0.5.0 optimize therefore requires restoring the pre-upgrade backup (or re-importing through the public batch contracts with the older binary). Upgrade every binary that opens the same database file as one set before allowing optimize to run, exactly as the invariants above require.

Log-schema maintenance in 0.8.4 records configuration freshness in the inventory. Reindexing refreshes the owned discovery companions transactionally; reopen and run the explicit schema command to repair stale companions from an older extension. Merely opening a database never repairs or replaces views.

Logs binaries from 0.8.4 support TIMELESS_LOGS_TIMESTAMP_UNIT=ms|us (default us). Select ms to serve a default SQL-created logs database without converting its data. Unit mismatches now stop startup before policy or ledger changes. The revised SQL quickstart explicitly declares microseconds for new tables.

1. Inventory the current installation

Metrics servers from 0.8.4 enforce the configured TIMELESS_METRICS_PROMQL_MAX_* budgets on native latest, range, export, and discovery routes, including Prometheus discovery aliases. Previously unbounded requests may now fail with an explicit budget error. Narrow the selection, increase the range step, split exports by time, or configure a larger deployment budget as needed. The server requires bounded catalog and latest/raw-frame capabilities at startup; upgrade the extension and server together. See the metrics query limits.

Record the current artifact identities and complete database paths. For a current server binary:

timeless-metrics-api --version
timeless-logs-api --version
timeless-traces-api --version

For an extension that already exposes the handshake:

sqlite3 /absolute/path/telemetry.db <<'SQL'
.load /absolute/path/libtimeless_ext
.mode json
SELECT timeless_capabilities() AS capabilities;
SELECT name, sql
  FROM sqlite_schema
 WHERE name = '_timeless_schema_migrations';
SQL

An old extension may fail the first statement with “no such function.” That is an identified old artifact, not permission to skip preflight in the new deployment. If the inventory query reports the ledger table, record its rows with SELECT * FROM _timeless_schema_migrations ORDER BY signal, version.

Also record:

  • database, -wal, and -shm paths and sizes;
  • signal table names and creation SQL from sqlite_schema;
  • configured timestamp units, retention, and log index keys;
  • current row/series/time-range counts obtained through public tables/TVFs;
  • artifact checksums and the deployed configuration/policy files.

Auth default change — read before upgrading the servers

Signal server binaries now start with authentication disabled unless TIMELESS_AUTH_MODE=required is set explicitly. Before this line, an unset TIMELESS_AUTH_MODE demanded a policy file and refused to start without one, so any deployment that ran at all either set required (still enforced, nothing changes — this includes every timeless_stack deployment) or set disabled (already open, nothing changes). The deployment that must act is one that relied on the refusal-to-start as a safety net: after upgrading, such a binary starts open instead of failing. If you want token verification, set TIMELESS_AUTH_MODE=required with TIMELESS_AUTH_POLICY_FILE — one line, and the enforcement is exactly as before.

Metrics retention default change — 0.8.4

The metrics server no longer applies an implicit seven-day wall-clock cutoff. By default it preserves the table's declared retention, which the extension measures from the newest queryable data timestamp. This keeps historical imports replayable and honors windows such as retention='30d'.

Deployments that want the previous additional seven-day wall-clock expiry must set TIMELESS_METRICS_RAW_RETENTION_SECS=604800. Unset or 0 disables that additional policy; a table without declared retention then keeps data until explicitly pruned. A positive override can shorten the table's window, but cannot lengthen it. Startup logs and /select/metrics/stats expose both policies (table_retention_seconds and raw_retention_seconds). Existing table definitions are not rewritten.

Companion view upgrades — 0.8.4

Opening or reading a signal table no longer refreshes companion views. On a writable connection, run INSERT INTO metrics(metrics) VALUES ('schema') for each source that needs its companions installed or upgraded (substitute the actual table name). Fresh CREATE VIRTUAL TABLE statements still install companions automatically. The command preserves user objects, fails on unowned name collisions, and rolls back both DDL and inventory on error. Run it before deploying updated companion queries to read-only replicas. The full lifecycle and attached-database examples are in the observability schema reference.

Schema version 2 changes the metrics series companion to a read-only catalog table with the same visible columns. Run the schema command for each metrics source to replace its version-1 view; other companions remain at version 1. This fixes attached-database queries and preserves binding through standalone opens and backup copies. Load the updated extension wherever these catalogs are queried before upgrading them.

2. Drain and create the rollback point

For a Rust signal server, stop producers, call its flush route (authenticated only if you enabled auth), then use its verified backup route while it still owns the database. The backup operation flushes, performs signal maintenance, requires a complete WAL checkpoint, uses SQLite's online-backup API, validates the result, fsyncs, and publishes without overwrite. Exact request/response behavior is in the server API reference.

Then stop the server normally and wait for it to exit. SIGINT/SIGTERM drains accepted HTTP requests, places a final flush behind admitted writes, checkpoints WAL with TRUNCATE, joins workers, and releases its owner lease. Do not use SIGKILL as an upgrade procedure.

For a directly embedded host, stop every writer and either use SQLite's online-backup API before shutdown or perform a coordinated complete WAL checkpoint before copying the entire database. A bare copy of the main file while WAL frames are outstanding is not a rollback point.

3. Preflight the new artifacts on a copy

Verify checksums, place the new extension and matching signal binaries beside the old artifacts, and retain the old files for rollback. Do not overwrite the only known-good copy.

Load the new extension against a copy of the database:

sqlite3 /absolute/path/telemetry-upgrade-check.db <<'SQL'
.load /absolute/path/libtimeless_ext
.mode json
SELECT timeless_capabilities() AS capabilities;
PRAGMA quick_check;
SELECT name, sql
  FROM sqlite_schema
 WHERE type = 'table' AND sql LIKE '%USING timeless_%'
 ORDER BY name;
SQL

Require the extension identity to match the selected artifact manifest and both binaries to satisfy the pairing floors in the compatibility contract. Check data_abi, sql_surface_version, the expected signal batch generations, and every query work guard required by the intended server. The canonical field inventory is in the SQL API reference.

Run representative public queries on the copy for every signal and compare counts, identities, timestamp extrema, float bits where relevant, labels, severity/metadata, trace relationships, and rich-span fields with the pre-upgrade inventory. Flush/maintain/checkpoint, close SQLite, reopen cold, and repeat the semantic checks. Do not treat quick_check alone as semantic parity.

4. Replace and start in dependency order

  1. Confirm the old owner is stopped and no process holds the signal lease.
  2. Install the new extension and matching server binaries atomically or under a deployment-level maintenance lock.
  3. Start one signal owner with its existing database path and policy.
  4. Let the writer preflight the extension and database, initialize/connect the public virtual table, and add the idempotent schema-ledger v1 row when the database is pre-ledger schema 0.
  5. Require /live, /ready (unauthenticated probe endpoints), and signal stats (authenticated when auth is enabled) to report the expected build, data ABI, table, queue, and storage state.
  6. Resume producers only after readiness succeeds.
  7. Repeat for each independently owned signal database/process.

Startup must fail closed for a missing capability function, extension below the server's floor, server below the extension's floor, wrong data ABI, missing rich batch generation, required query guard absence, future schema ledger, incompatible timestamp/retention policy, corruption, or another owner lease.

5. Post-upgrade verification

After admitting a bounded test batch:

  • call the ordered flush barrier and require completed watermarks;
  • execute exact latest/range metrics queries and at least one applicable PromQL expression;
  • execute a bounded log read/count and verify all expected typed fields;
  • fetch a complete trace and verify parent relationships and rich-span fields;
  • inspect public stats for failures, queue growth, decoded work, WAL/checkpoint state, and storage accounting;
  • for upgraded traces, require duration_unknown_blocks=0 after the planned optimize backfill and verify an impossible duration filter reports zero candidate blocks and decoded spans;
  • when a newly created trace table declares attribute_indexes, verify attribute_index_fields, attribute_bloom_rows, and attribute_bloom_bytes, then compare one hidden-filter result with its public JSON1 control. Existing tables do not acquire an allowlist merely by opening them; changing fields is a side-by-side public row/batch migration;
  • restart normally and repeat the reads cold; and
  • retain the rollback backup and old artifacts for the documented support window.

Do not silently route an unsupported query to another process during this verification.

Rollback

Rollback is backup restoration, not an in-place binary downgrade:

  1. Stop producers and the new signal owner normally.
  2. Preserve the failed/new database, WAL, SHM, logs, and build identities for diagnosis.
  3. Restore the coordinated pre-upgrade backup to a new path or replace the stopped database using an atomic operator-controlled operation.
  4. Restore the previous extension and server binaries as the same known-good set.
  5. Start against the restored database, require its normal compatibility checks/readiness, and run the pre-upgrade semantic oracle.
  6. Resume producers only after verification.

Never point an older server at a database already mutated by a newer server and call that rollback, even when data_abi has not changed. The older server may not understand additive schema, batch, query, limit, or lifecycle contracts. Restore the matching backup.

Source-state table

Detected source Action
Fresh database Create the public signal vtab with the desired timestamp/index/retention options; record schema ledger v1 through the writer.
Current SQLite/libSQL database and compatible extension/server bundle Start normally after backup; full handshake and schema preflight remain mandatory.
Pre-ledger SQLite telemetry database Back up, preflight with the new extension on a copy, then allow the new writer to add ledger v1 idempotently.
Database created by tagged v0.3.0 Replace the extension with the current line before starting a release server; validate all signals on a copy because the old tag has no capability document.
Future ledger or different data ABI Stop. Use a compatible newer binary or an explicit versioned migration; do not mutate/downgrade.
Corrupt database or failed semantic parity Stop and retain both source and candidate; restore/continue the old deployment.
External legacy Rust block store Use the owning higher-order library's explicit converter. These signal binaries do not read it and must not recreate its storage format.
Ambiguous multiple candidate/production databases Stop and require operator selection; never choose by newest mtime or largest file.

SQLite/libSQL, WAL, backup, and replication

All durable Timeless state, including compressed blocks, indexes, rollups, and private metadata, is stored in the containing SQLite database and WAL. Use a whole-database SQLite/libSQL backup or replication mechanism. Timeless does not add a second replication protocol and does not claim that an arbitrary file copy made during active WAL writes is valid.

Replication compatibility does not replace application-level verification: open a replica with the same extension generation, require the capability handshake, and compare public semantic queries after the host reports the expected replication/checkpoint boundary.