Repository navigation
Conversation
|
📦 Python package built successfully!
|
`artifacts.store_dataframe(df, name)` writes a whole pandas or polars DataFrame to `deepnote_dataframes/<name>/<uuid>.arrow` in the project's filesystem, with `to_feather` / `write_ipc`, zstd falling back to lz4 and then uncompressed. A `manifest.json` names the current file and the one it replaced, and every write rewrites it. `artifacts.delete_dataframe(name)` deletes a name's directory. The function raises like the user's own `df.to_feather()` would, and returns None on a read-only mount. An interrupt stops the write and removes the partial file, except during the manifest update, which finishes first. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B441SpUqRn7QTzcnWxjNa5
tkislan
force-pushed
the
feat/artifacts-store-dataframe
branch
from
October 8, 2026 12:37
3d68fb4 to
62a4d0e
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #135 +/- ##
==========================================
+ Coverage 77.26% 77.57% +0.31%
==========================================
Files 115 117 +2
Lines 6589 6663 +74
Branches 961 972 +11
==========================================
+ Hits 5091 5169 +78
+ Misses 1186 1184 -2
+ Partials 312 310 -2
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. |
5 of 7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@coderabbitai ignore
Summary
Toolkit side of the "Full Dataframe storage" RFC (revision 10, §3.4): a function that notebook code calls to store a whole DataFrame under a name, and one that deletes it. It does what the user's own
df.to_feather(...)could do; the toolkit only decides where the file goes and replaces it safely.deepnote_toolkit/artifacts.py:store_dataframeanddelete_dataframe, exported from the package likeocelots(_IMPORT_MAPPINGS).deepnote_toolkit/dataframe_storage.py: the writer both phases of the RFC share. Project root, read-only detection, codec choice, the write, the manifest commit and cleanup of replaced files.deepnote_toolkit/dataframe_storage_manifest.py: the manifest model, reading, writing and<uuid>.arrowvalidation.<project root>/deepnote_dataframes/<name>/<uuid>.arrowplusmanifest.json, which names the current file and the one it replaced. The frame is written withdf.to_feather()(pandas) ordf.write_ipc()(polars), zstd, falling back to lz4, then uncompressed. Write order: frame, manifest, delete.Out of scope (other repos, or the follow-up): the read endpoint, cleanup on project deletion, and phase 2's automatic storage of a block's result. That is the
post_run_cellhook in #131, which builds on this PR.Behaviour worth reviewing
nameoutside[A-Za-z0-9_-]{1,128}raisesValueError. Anything but a pandas or polars DataFrame raisesTypeError: PySpark and pandas-on-Spark frames, polarsLazyFrames and SQL query previews included. Both are checked before the mount is, so they raise on a read-only mount too. A frameto_feathercan't serialize, or a failed write, raises its own exception and leaves the name's previous frame current.store_dataframereturnsNonebefore converting anything (os.statvfsflag, andEROFSfrom the write);delete_dataframereturnsFalseand deletes nothing..arrow.KeyboardInterruptpropagates. The manifest commit is the exception: an interrupt there is held until the manifest is final, so it can neither leave an empty manifest nor delete the frame the manifest names.<uuid>.arrowcounts as absent, and only<uuid>.arrowfiles inside the name's own directory are ever deleted. Besides the replacedpreviousfile, a write deletes other unreferenced<uuid>.arrowfiles older than a day.EIO), the call raises but the new frame is already current. The RFC says every failure other than a read-only mount raises, so I kept that.delete_dataframeusesshutil.rmtree, which refuses a symlinked name directory instead of following it out ofdeepnote_dataframes/.Not in this PR
No feature flag, no environment variable, no
execute_requestmetadata and no error reports to the webapp. A call is the user's own code, so none of them applies (RFC §3.4.1, §3.4.3).No documentation change: the repo's docs cover configuration, and the function has no user-visible effect until the read endpoint exists. The product docs belong with that change.
Not verified here
pd.concat) are covered by "whatever pandas does" tests, not by 16.1 itself.close()(fallback: flush plus fsync).set_notebook_path()'s precedence, which useshome_dirwithout a/worksuffix. Confirm on a pod that it equals the project mount.Test plan
poetry run python -m pytest tests/unit/dataframe_storage -p no:randomly: 141 passedblack,isort,flake8clean on the touched files;mypy deepnote_toolkit/cleantests/unit: 1424 passed, 4 skipped, 1 failed. The failure istest_redshift_dialect.py::test_redshift_distribution_matches_python_version: the venv has two redshift packages installed, and it fails identically on a cleanorigin/main.🤖 Generated with Claude Code
https://claude.ai/code/session_01B441SpUqRn7QTzcnWxjNa5