Skip to content

feat(artifacts): add store_dataframe and delete_dataframe - #135

Draft
tkislan wants to merge 1 commit into
mainfrom
feat/artifacts-store-dataframe
Draft

tkislan wants to merge 1 commit into
mainfrom
feat/artifacts-store-dataframe

Conversation

@tkislan

@tkislan tkislan commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

@coderabbitai ignore

Summary

Toolkit side of the "Full Dataframe storage" RFC (revision 10, §3.4): a function that notebook code calls to store a whole DataFrame under a name, and one that deletes it. It does what the user's own df.to_feather(...) could do; the toolkit only decides where the file goes and replaces it safely.

from deepnote_toolkit import artifacts

artifacts.store_dataframe(df, "revenue")   # replaces what was stored as "revenue" before
artifacts.delete_dataframe("revenue")
  • New deepnote_toolkit/artifacts.py: store_dataframe and delete_dataframe, exported from the package like ocelots (_IMPORT_MAPPINGS).
  • New deepnote_toolkit/dataframe_storage.py: the writer both phases of the RFC share. Project root, read-only detection, codec choice, the write, the manifest commit and cleanup of replaced files.
  • New deepnote_toolkit/dataframe_storage_manifest.py: the manifest model, reading, writing and <uuid>.arrow validation.
  • Layout: <project root>/deepnote_dataframes/<name>/<uuid>.arrow plus manifest.json, which names the current file and the one it replaced. The frame is written with df.to_feather() (pandas) or df.write_ipc() (polars), zstd, falling back to lz4, then uncompressed. Write order: frame, manifest, delete.

Out of scope (other repos, or the follow-up): the read endpoint, cleanup on project deletion, and phase 2's automatic storage of a block's result. That is the post_run_cell hook in #131, which builds on this PR.

Behaviour worth reviewing

  • Raises like the user's own call. name outside [A-Za-z0-9_-]{1,128} raises ValueError. Anything but a pandas or polars DataFrame raises TypeError: PySpark and pandas-on-Spark frames, polars LazyFrames and SQL query previews included. Both are checked before the mount is, so they raise on a read-only mount too. A frame to_feather can't serialize, or a failed write, raises its own exception and leaves the name's previous frame current.
  • Read-only mount: store_dataframe returns None before converting anything (os.statvfs flag, and EROFS from the write); delete_dataframe returns False and deletes nothing.
  • Returns the new frame's reference ID, the UUID in the file name without .arrow.
  • Stop during a write. The partial file is deleted and the KeyboardInterrupt propagates. The manifest commit is the exception: an interrupt there is held until the manifest is final, so it can neither leave an empty manifest nor delete the frame the manifest names.
  • Untrusted manifest. Anyone who can edit project files can write it. A manifest over 4 KiB, not JSON, or naming anything but <uuid>.arrow counts as absent, and only <uuid>.arrow files inside the name's own directory are ever deleted. Besides the replaced previous file, a write deletes other unreferenced <uuid>.arrow files older than a day.
  • A failure after the commit raises too. If deleting the replaced files fails (say EIO), the call raises but the new frame is already current. The RFC says every failure other than a read-only mount raises, so I kept that.
  • delete_dataframe uses shutil.rmtree, which refuses a symlinked name directory instead of following it out of deepnote_dataframes/.

Not in this PR

No feature flag, no environment variable, no execute_request metadata and no error reports to the webapp. A call is the user's own code, so none of them applies (RFC §3.4.1, §3.4.3).

No documentation change: the repo's docs cover configuration, and the function has no user-visible effect until the read endpoint exists. The product docs belong with that change.

Not verified here

  • Only Python 3.13 / pyarrow 24 / pandas 2.2.3 / polars 1.39 ran locally, so the pyarrow 16.1 failure branches (UUID column, dictionary pd.concat) are covered by "whatever pandas does" tests, not by 16.1 itself.
  • A real s3fs mounter, and whether polars writing to a Python file object surfaces a failed upload from close() (fallback: flush plus fsync).
  • The project root copies set_notebook_path()'s precedence, which uses home_dir without a /work suffix. Confirm on a pod that it equals the project mount.

Test plan

  • poetry run python -m pytest tests/unit/dataframe_storage -p no:randomly: 141 passed
  • Broke the code on purpose in 25 places (name regex, deferred SIGINT, read-only check, partial-file cleanup, stale-file age, type check, delete scope, manifest size cap and more); every break failed a test
  • black, isort, flake8 clean on the touched files; mypy deepnote_toolkit/ clean
  • Full tests/unit: 1424 passed, 4 skipped, 1 failed. The failure is test_redshift_dialect.py::test_redshift_distribution_matches_python_version: the venv has two redshift packages installed, and it fails identically on a clean origin/main.
  • CI across the Python / pyarrow matrix
  • Pod test: mount upload on close, project root

🤖 Generated with Claude Code

https://claude.ai/code/session_01B441SpUqRn7QTzcnWxjNa5

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

📦 Python package built successfully!

  • Version: 2.8.0.dev6+3929984
  • Wheel: deepnote_toolkit-2.8.0.dev6+3929984-py3-none-any.whl
  • Install:
    pip install "deepnote-toolkit @ https://deepnote-staging-runtime-artifactory.s3.amazonaws.com/deepnote-toolkit-packages/2.8.0.dev6%2B3929984/deepnote_toolkit-2.8.0.dev6%2B3929984-py3-none-any.whl"

`artifacts.store_dataframe(df, name)` writes a whole pandas or polars
DataFrame to `deepnote_dataframes/<name>/<uuid>.arrow` in the project's
filesystem, with `to_feather` / `write_ipc`, zstd falling back to lz4 and then
uncompressed. A `manifest.json` names the current file and the one it replaced,
and every write rewrites it. `artifacts.delete_dataframe(name)` deletes a name's
directory.

The function raises like the user's own `df.to_feather()` would, and returns
None on a read-only mount. An interrupt stops the write and removes the partial
file, except during the manifest update, which finishes first.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B441SpUqRn7QTzcnWxjNa5
@tkislan
tkislan force-pushed the feat/artifacts-store-dataframe branch from 3d68fb4 to 62a4d0e Compare October 8, 2026 12:37
@codecov

codecov Bot commented Oct 8, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 77.57%. Comparing base (f07a89f) to head (62a4d0e).
✅ All tests successful. No failed tests found.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #135      +/-   ##
==========================================
+ Coverage   77.26%   77.57%   +0.31%     
==========================================
  Files         115      117       +2     
  Lines        6589     6663      +74     
  Branches      961      972      +11     
==========================================
+ Hits         5091     5169      +78     
+ Misses       1186     1184       -2     
+ Partials      312      310       -2     
Flag Coverage Δ
combined 77.57% <100.00%> (+0.31%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant