Repository navigation
Memory is not reclaimed with unregister() when inside a transaction #469
Description
Activity
The underlying issue is that when the object is
register()ed, to make it available for the query planner a view is created with the same name. This view is added to the catalog, and isDROPed whenunregister()is invoked. The issue is that when inside a transaction theDROP VIEWdoesn't clear the view object from the catalog because it is required to maintain transactionality. That view holds the reference to the data object which prevents its garbage collection.In contrast, the approach when
register()is bypassed doesn't suffer from this. It is PythonReplacementScan that reads the variable from python's variable stack directly.I can see some alternative solutions for this problem but would love some guidance on which option is more preferable:
- We keep the
CREATE VIEW, but make it so that the view instead stores only a weakref. If the view is re-read after the python object is deleted then an duckdb runtime error is raised. - We remove the
CREATE VIEW, and makeregister()/unregister()behave just like an explicit version ofPythonReplacementScan. In this format, the two functions simply modify theClientContextto maintain the a map of name -> python object (outside of transaction semantics). Then PythonReplacementScan is modified to first check this map of names instead of checking python locals.
I'm more inclined on the second option (my mental model of register/unregister was simply an explicit version of
python_scan_all_frames), but maybe there are some use cases that require the VIEWs to be created. I didn't find any unit test that makes such assumption explicit if it exists.- We keep the
Thanks for the report! I agree that this is not behavior you might expect and we should update our docs. But I do think we should keep
register()the way it is. It creates a view, and the mental model should be more or less the same as when you would create and drop a (materialized) view repeatedly in a transaction. This would result in accumulating memory in the same way, as long as the transaction is not committed or rolled back. The only real difference is that a Python object lives outside of DuckDB's managed heap so it can't spill etc. But otherwise the semantics line up, and I think they are correct.For this sort of bulk loading it is indeed better to use a replacement scan rather than register.
The gripe I have with that is I'd rather not have duckdb magically have access to all python objects in scope and would prefer the option to have explicit/whitelisted access only. Is there a way we can keep
register()with its current semantics but also have the option to not create the view? e.g.register(.., create_view=False)orregister_unsafe(...)
What happens?
Repeatedly inserting pandas DataFrames via
connection.register()inside a single open transaction causes process RSS to grow linearly with each chunk, even afterunregister(),del, andgc.collect().Using
SET python_scan_all_frames=trueand referencing a local variable instead of explicit use ofregister()/unregister()avoids the problem.To Reproduce
The following unmodified script shows the memory growing with each chunk until the process hits >10GB.
Change the
USE_REGISTERflag toFalseto avoid the use ofregister()which bypasses the problem.OS:
Ubuntu 24.04LTS
DuckDB Package Version:
1.5.2
Python Version:
3.14
Full Name:
João Pedro Maia Rafael
Affiliation:
Upper Delta, Unipessoal LDA
What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.
I have tested with a nightly build
Did you include all relevant data sets for reproducing the issue?
Not applicable - the reproduction does not require a data set
Did you include all code required to reproduce the issue?
Did you include all relevant configuration to reproduce the issue?