Skip to content

feat: split raw data storage into per-dataset SQLite files - #8219

Merged
Mikhail Astafev (astafan8) merged 28 commits into
microsoft:mainfrom
astafan8:feature/split-raw-data-sqlite
Sep 28, 2026
Merged

Mikhail Astafev (astafan8) merged 28 commits into
microsoft:mainfrom
astafan8:feature/split-raw-data-sqlite

Conversation

@astafan8

@astafan8 Mikhail Astafev (astafan8) commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Opt-in split raw data storage: each dataset's raw measurement data (its results table) can be written to an individual per-dataset SQLite file, while all metadata stays in the main database. This keeps the main DB small instead of growing to many gigabytes as datasets accumulate. The public DataSet API is unchanged — every method behaves identically whichever backend is used.

Selected via dataset config options in qcodesrc.json:

  • raw_data_backend (string, default "sqlite_main_db") — the results backend; set to "sqlite_per_dataset_db" to enable split storage.
  • raw_data_backend_config (object) — per-backend settings keyed by backend name. The per-dataset backend reads raw_data_backend_config.sqlite_per_dataset_db.raw_data_path (default "{db_location}", expanded relative to the main DB path), the folder for the per-dataset files.

How it works

  • Pluggable storage backend. A ResultsBackend strategy (src/qcodes/dataset/_results_backend.py) encapsulates where/how a dataset's results are stored, so DataSet carries no storage-specific conditionals. Backends are registered by their backend_name (the raw_data_backend config value): MainDatabaseResultsBackend (sqlite_main_db, default) and SeparateSqliteFileResultsBackend (sqlite_per_dataset_db). The backend is chosen inside DataSet.__init__ — from config for new runs, and from the run's recorded state for existing runs, so a plain DataSet(run_id=...) auto-detects. Everything routes through the single results_conn; the backend also owns the results-table operations (exists/count/length/insert/read). Adding a backend is a new subclass plus a raw_data_backend enum value and a raw_data_backend_config section.
  • Generic lifecycle. The backend exposes backend-agnostic hooks — setup_on_load, setup_on_new_run, setup_on_start, close — so the abstraction never presumes a "results table"; a future backend that stores data in another format just implements the hooks.
  • Per-dataset files. Named <guid>.db, containing only the results table + numpy type adapters (no metadata schema). Created via shared helpers _connect_to_sqlite_file() / _create_run_table() so schema and connection settings match the main DB (also used to dedup connect()).
  • No empty results table in the main DB. With sqlite_per_dataset_db, no results table is created there — only run metadata. A split run is marked by a dedicated raw_data_db_path column in runs (recorded at dataset creation), which also distinguishes it from a DataSetInMem run when loading. That column is an internal detail: it is excluded from the user-facing metadata dict (accessed via get/set_raw_data_db_path_for_run), so it never leaks into the_same_dataset_as comparisons or extracted/exported DBs.
  • Edge cases. number_of_results/__len__ return 0 when the results table doesn't exist yet (pristine dataset). Subscriptions requested before start are deferred and materialised on start; read_only is threaded through the public loaders so a read-only load opens the per-dataset file read-only too. Extract/export work unchanged (extract_runs_into_db reads via the backend and doesn't carry over raw_data_db_path). Per-dataset setup writes are batched into single transactions, keeping measurement throughput on par with the default backend.

Management helpers

Exposed via qcodes.dataset (both destructive ones default to dry_run=True, show a tqdm progress bar, and return a result object; deletion reuses the existing remove_dataset_from_db):

  • update_raw_data_paths(db_path, new_raw_data_folder) — fix stored paths after moving the per-dataset files.
  • purge_orphaned_datasets(db_path, *, dry_run=True) — remove main-DB records whose raw data file is gone.
  • cleanup_datasets(db_path, *, older_than_days=None, sample_name=None, larger_than_mb=None, dry_run=True) — remove datasets (records + files) matching the given criteria (AND).

Files

  • New: _results_backend.py (the backend strategy + registry); _raw_data_storage.py (storage + config accessors + management helpers); tests/dataset/test_raw_data_storage.py; docs/changes/newsfragments/8219.new.
  • Modified: data_set.py (delegate to the backend, no-empty-table, deferred subscribers, disambiguation); sqlite/queries.py (get/set_raw_data_db_path_for_run, get_datasets_with_raw_data_path + RawDataDatasetRecord, remove_dataset_from_db, column excluded from metadata); sqlite/database.py (shared connect helpers); data_set_cache.py, subscriber.py, database_extract_runs.py (use results_conn); __init__.py (exports); qcodesrc.json + qcodesrc_schema.json (config); docs.

Verification

test_raw_data_storage.py covers unit, integration, management, extract/export, config/backend selection + DataSet(run_id=...) auto-detection, and the no-empty-table / count-before-start / subscribe-before-start / background-write cases. Full tests/dataset suite passes with no regressions in default mode. Pyright, ruff and pre-commit all clean.

@codecov

codecov Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.67100% with 20 lines in your changes missing coverage. Please review.
✅ Project coverage is 72.28%. Comparing base (dccb1d7) to head (cd2ec34).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/qcodes/dataset/_raw_data_storage.py 95.52% 9 Missing ⚠️
src/qcodes/dataset/_results_backend.py 95.28% 5 Missing ⚠️
src/qcodes/dataset/data_set.py 96.77% 2 Missing ⚠️
src/qcodes/dataset/sqlite/db_overview.py 93.33% 2 Missing ⚠️
src/qcodes/dataset/sqlite/queries.py 95.12% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8219      +/-   ##
==========================================
+ Coverage   71.99%   72.28%   +0.28%     
==========================================
  Files         305      307       +2     
  Lines       32015    32424     +409     
==========================================
+ Hits        23050    23438     +388     
- Misses       8965     8986      +21     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread src/qcodes/dataset/data_set.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

Comment thread src/qcodes/dataset/_raw_data_storage.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 10 out of 10 changed files in this pull request and generated 3 comments.

Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/data_set.py Outdated
Comment thread tests/dataset/test_raw_data_storage.py
@astafan8

Copy link
Copy Markdown
Contributor Author

Added: update_raw_data_paths helper

Added a utility function for users who move their per-dataset raw data files to a new location. This mirrors the pattern used for exported netCDF files.

Usage:

`python
from qcodes.dataset import update_raw_data_paths

update_raw_data_paths(
db_path="/path/to/main_database.db",
new_raw_data_folder="/new/location/of/raw_files/"
)
`

The function:

  • Scans all datasets in the main DB that have raw_data_db_path metadata
  • For each, checks if the corresponding .db file exists in the new folder
  • Updates the stored path to point to the new location
  • Logs warnings for any files not found in the new folder

4 tests added, documentation updated in introduction.rst and Database.ipynb.

Comment thread docs/dataset/introduction.rst Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread docs/changes/newsfragments/8219.new Outdated
Comment thread docs/changes/newsfragments/8219.new Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 14 out of 14 changed files in this pull request and generated 8 comments.

Comment thread src/qcodes/dataset/_raw_data_storage.py Outdated
Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread docs/dataset/dataset_design.rst Outdated
Comment thread docs/dataset/introduction.rst Outdated
Comment thread src/qcodes/dataset/sqlite/queries.py
Mikhail Astafev and others added 17 commits September 27, 2026 08:22
- Change per-dataset log.info to log.debug in purge/cleanup/update_paths
- Replace os.remove with Path.unlink() for robustness
- Remove unused os import

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- _extract_single_dataset_into_db: use dataset._data_conn instead of
  dataset.conn so raw data is read from the per-dataset file
- unsubscribe/unsubscribe_all: use _data_conn to remove triggers that
  were created on the raw data connection
- Add tests for extract_runs_into_db and
  export_datasets_and_create_metadata_db with split raw data

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Mikhail Astafev <astafan8@gmail.com>
- Extract connect_to_sqlite_file() and _register_numpy_sqlite_adapters_and_converters()
  in sqlite/database.py; use them in both connect() and connect_to_raw_data_db()
  so the numpy adapters and connection settings are defined once
- Reuse _create_run_table() in create_raw_data_db() instead of duplicating
  the results-table CREATE TABLE logic (also gains table-name validation)
- Validate table name via _validate_table_name() before DROP TABLE in
  remove_dataset_from_db()
- Apply doc review suggestions and fix module path in dataset_design.rst

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previously a results table was still created (empty) in the main database
for split-storage datasets, purely to satisfy code paths that assumed the
table exists. This removes that table entirely so the main DB holds only
metadata - cleaner and future-proof for non-sqlite raw data backends.

Changes:
- Do not create the results table in the main DB when raw storage is
  enabled (create_run_table=False, insert_into_results_table=False),
  mirroring how DataSetInMem records runs.
- Identify split-storage runs via a dedicated 'raw_data_db_path' column in
  the runs table, recorded at dataset creation. This disambiguates them
  from DataSetInMem runs (which also have no results table) in
  _get_datasetprotocol_from_guid, replacing the old reliance on the empty
  table's presence.
- Keep 'raw_data_db_path' out of the user-facing metadata dict (read/write
  via get/set_raw_data_db_path_for_run); this also stops it leaking into
  metadata comparisons and extract/export copies.
- number_of_results/__len__ return 0 when the results table does not exist
  (e.g. a pristine, not-yet-started split dataset).
- Defer subscriptions requested before the dataset is started until start
  time, when the raw data table exists (fixes subscriber trigger creation).
- Restrict management queries (purge/cleanup) to started runs only.
- Docs + tests updated; add tests for subscribe-before-start, count before
  start, no-main-table, and extract not carrying over raw_data_db_path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address review: factor the subscriber creation into
_create_and_start_subscriber (reused by subscribe and
_start_pending_subscribers) and the pending-subscription bookkeeping into
_queue_pending_subscriber.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…s for query result

- Rename connect_to_sqlite_file -> _connect_to_sqlite_file (internal helper,
  not part of the public API)
- Return a RawDataDatasetRecord dataclass from get_datasets_with_raw_data_path
  instead of an opaque 8-tuple

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add noqa: BLE001 with rationale for the intentional broad excepts in
  purge_orphaned_datasets / cleanup_datasets that collect per-dataset errors
- Collapse nested/duplicate branches flagged by SIM in cleanup_datasets and
  _get_datasetprotocol_from_guid

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- remove_dataset_from_db: drop the redundant result_table_name argument and
  look it up from the runs table internally.
- Convert the record types RawDataDatasetRecord and DatasetInfo from
  dataclasses to NamedTuples (lighter immutable value records; the mutable
  result aggregates PurgeResult/CleanupResult stay dataclasses).
- update_raw_data_paths: use closing(conn) so the connection is always closed;
  fix a docstring typo.
- Add tqdm progress bars to the I/O-bound loops: dataset scanning
  (_build_dataset_info_list), path updates, and the purge/cleanup removals.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Move the split-storage behaviour out of DataSet's is_raw_data_storage_enabled()
conditionals into a composed ResultsBackend strategy, chosen inside
DataSet.__init__ (from config for new runs, from the run's recorded
raw_data_db_path for existing runs, so DataSet(run_id=...) auto-detects).

- New _results_backend.py: ResultsBackend base (MainDatabaseResultsBackend)
  plus SeparateSqliteFileResultsBackend. Backends expose a single results_conn
  and encapsulate the results-table operations (exists/count/length/insert/
  read) and a results_db_path used to route background writes.
- DataSet delegates to self._results_backend; the confusing _data_conn /
  _raw_data_conn / _raw_data_db_path trio is replaced by a single _results_conn.
  data_set_cache.py, subscriber.py and database_extract_runs.py use it too.
- Tests assert backend selection and that DataSet(run_id=...) auto-detects, and
  that split-storage background writes land in the per-dataset file.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…, robustness

- Replace the `creates_results_table_in_main_db` boolean with a polymorphic
  `ResultsBackend.create_results_table` (and `setup_on_new_run`) so each
  backend owns where its results table is created; the main backend still
  creates the (empty) table up front to keep DataSet/DataSetInMem
  disambiguation and behaviour unchanged.
- Thread `read_only` through the public loaders into `DataSet` so split
  datasets loaded read-only open their per-dataset file read-only too.
- `unsubscribe_all` now also clears deferred pending subscribers.
- `update_raw_data_paths` resolves the new folder to an absolute path.
- cleanup removes the DB record before unlinking the raw file (no data loss on
  DB failure).
- Restore the lost `test_get_parameter_data_from_raw_data` method boundary.
- Docs/docstrings: describe raw_data_db_path as an internal runs column (not
  metadata); fix _results_conn name and design section; remove a stale comment.
The results-backend refactor split results-table creation across
setup_on_new_run and create_results_table, roughly doubling the number of
SQLite commits (fsyncs) in the per-dataset start path versus main. Batch the
add_parameter registration loop and the backend column creation each into a
single transaction so dataset setup performs one commit per phase instead of
one per parameter. This restores measurement throughput to parity with main
(realistic 1500-point measurements: was ~+8%, now within noise).
Config: replace the dataset.raw_data_to_separate_db boolean and top-level
dataset.raw_data_path with a dataset.raw_data_backend enum
('sqlite_main_db' | 'sqlite_per_dataset_db') plus a dataset.raw_data_backend_config
object holding per-backend settings (the per-dataset backend's raw_data_path
lives there). Backends are registered by their backend_name, so adding a
backend is a matter of a new ResultsBackend subclass plus a raw_data_backend
enum value and raw_data_backend_config section.

Lifecycle: rename the backend's create_results_table hook to the generic
setup_on_start (paired with setup_on_new_run), so the abstraction does not
presume a 'results table' - a future backend that stores data in another
format just implements the lifecycle hooks. Behaviour is unchanged.

Also rename the internal background-writer queue key raw_data_path to
results_db_path to avoid confusion with the config path. Updates tests,
docs, docstrings and the newsfragment.
…rom docs

- create_raw_data_db now creates an empty results table and adds columns via
  insert_column (which quotes identifiers), matching the main-database path, so
  a parameter whose name is a SQL keyword (e.g. 'from') no longer breaks
  split-storage startup at CREATE TABLE. Adds a regression test.
- Database.ipynb: replace the manual qcodesrc.json editing block with a pointer
  to the Configuring QCoDeS notebook (config is best changed via qc.config).
- introduction.rst: describe the option via the config system and reference the
  Configuring QCoDeS notebook instead of qcodesrc.json.
…lve {db_location} from the dataset's own DB

- cleanup_datasets now rejects negative older_than_days / larger_than_mb, which
  would otherwise match (and, with dry_run=False, delete) every split dataset.
- The per-dataset raw-data folder now resolves {db_location} relative to the
  owning dataset's database (DataSet.path_to_db) rather than the global
  core.db_location, so datasets created against an explicit connection/path
  (e.g. via Measurement) place their raw files beside the correct database.
  _expand_export_path gains an optional db_location argument for this.
- Adds tests for both.

Copilot AI commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

One or more custom setup steps configured for this repository failed during this Copilot code review run:

Fetch main branch

Setup steps run before each review. If the review above is missing context, or no review was posted at all, the failing step above may be the cause. See the workflow run for failure details, fix your setup steps configuration, and re-request a review.

Note

You can configure setup steps for Copilot code review separately from Copilot cloud agent with a copilot-code-review.yml file. Read the docs for details.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Backend-marker collisions, unsafe active-run cleanup, inaccurate overview counts, and background subscriber failures remain unresolved.

Review effort: Balanced
Findings: 2 High severity

Open (2)
Resolved since last review (4)

Comment thread src/qcodes/dataset/_raw_data_storage.py
Comment thread src/qcodes/dataset/data_set.py
Comment thread src/qcodes/dataset/sqlite/queries.py
Comment thread src/qcodes/dataset/sqlite/queries.py Outdated

@jenshnielsen Jens Hedegaard Nielsen (jenshnielsen) left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good after addressing open pr comments. It might be useful to allow compression into a zip file or similar for the raw data but that can be added in a follow up pr.

…moval helper

- cleanup_datasets never removes in-progress (started-but-not-completed) runs,
  whose per-dataset file may still be actively written; adds a test.
- get_db_overview now counts records from the recorded per-dataset raw-data
  file for split runs (results table is absent from the main DB), fixing
  records==0 for split runs; adds a test.
- Make remove_dataset_from_db private (_remove_dataset_from_db): it is an
  internal helper for the management functions, not public API.
@astafan8
Mikhail Astafev (astafan8) added this pull request to the merge queue Sep 28, 2026
Merged via the queue into microsoft:main with commit ffafd35 Sep 28, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants