Skip to content

GH-51043: [Python] Reject null required Arrow objects - #51161

Open
1fanwang wants to merge 3 commits into
apache:mainfrom
1fanwang:1fannnw/fix-pyarrow-none-segfaults
Open

GH-51043: [Python] Reject null required Arrow objects#51161
1fanwang wants to merge 3 commits into
apache:mainfrom
1fanwang:1fannnw/fix-pyarrow-none-segfaults

Conversation

@1fanwang

@1fanwang 1fanwang commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Rationale for this change

Passing null values for required Arrow objects can terminate Python or raise an unrelated AttributeError. Array buffer constructors should reject missing types and child arrays with TypeError.

Fixes #51043.

What changes are included in this PR?

Required schema, file-format, data-type, dictionary, fragment, and child-array arguments reject null before use. The base Array and RunEndEncodedArray buffer constructors require a data type. String constructors retain their existing buffer validation, including null offsets for empty arrays.

Are these changes tested?

Native Arrow 26.0.0-SNAPSHOT reproduced the remaining AttributeErrors before the fix and TypeErrors afterward. The focused buffer tests include dictionary construction and both string variants.

Raw logs
$ python -c 'import pyarrow as pa; pa.Array.from_buffers(type=None, length=0, buffers=[])'
Before:
AttributeError: 'NoneType' object has no attribute 'num_fields'
After:
TypeError: Argument 'type' has incorrect type (expected pyarrow.lib.DataType, got NoneType)

$ python -c 'import pyarrow as pa; pa.RunEndEncodedArray.from_buffers(type=None, length=0, buffers=[])'
Before:
AttributeError: 'NoneType' object has no attribute 'num_fields'
After:
TypeError: Argument 'type' has incorrect type (expected pyarrow.lib.DataType, got NoneType)

$ python -c 'import pyarrow as pa; pa.Array.from_buffers(type=pa.list_(pa.int8()), length=0, buffers=[None, pa.py_buffer(bytes(4))], children=[None])'
Before:
AttributeError: 'NoneType' object has no attribute 'ap'
After:
TypeError: Array child must not be None

$ cd python
$ python -m pytest pyarrow/tests/test_array.py -k from_buffers -q
.............                                                            [100%]
13 passed, 317 deselected, 1 warning in 0.45s

Are there any user-facing changes?

Missing required objects raise TypeError. Valid null buffers keep their existing behavior.

New Contributor's Guide |
Contributing Overview |
AI-generated Code Guidance

Generated-by: GitHub Copilot CLI (GPT-5.6 Sol)
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Copilot AI lite review requested due to automatic review settings September 4, 2026 21:47
@1fanwang
1fanwang requested a review from rok as a code owner September 4, 2026 21:47
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

⚠️ GitHub issue #51043 has been automatically assigned in GitHub to PR creator.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

DictionaryArray.from_buffers() still allows dictionary=None while dereferencing it unconditionally, so a None argument can still trigger a crash.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR hardens several PyArrow Cython entry points by rejecting None for required Arrow extension objects, preventing null C++ pointer dereferences that can segfault the Python interpreter (GH-51043).

Changes:

  • Mark required Schema, FileFormat, and DataType parameters as non-nullable (not None) at the Cython boundary to raise TypeError instead of crashing.
  • Add an explicit None check for FileSystemDataset fragments before unwrapping.
  • Add focused regression tests covering the newly rejected None arguments.
File summaries
File Description
python/pyarrow/_dataset.pyx Reject None fragments and make schema/format non-nullable in FileSystemDataset.
python/pyarrow/_parquet.pyx Make SortingColumn conversion helpers reject schema=None at the boundary.
python/pyarrow/array.pxi Make DictionaryArray.from_buffers reject type=None at the boundary.
python/pyarrow/tests/test_dataset.py Add regression assertions for FileSystemDataset(..., schema=None/format=None) and [None] fragments.
python/pyarrow/tests/parquet/test_metadata.py Add regression assertions for SortingColumn.* with schema=None.
python/pyarrow/tests/test_array.py Add regression assertion for DictionaryArray.from_buffers(type=None, ...).
Review details
  • Files reviewed: 6/6 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/pyarrow/array.pxi
Generated-by: GitHub Copilot CLI (GPT-5.6 Sol)
Signed-off-by: 1fanwang <1fannnw@gmail.com>
@github-actions

github-actions Bot commented Sep 5, 2026

Copy link
Copy Markdown

⚠️ GitHub issue #51043 has been automatically assigned in GitHub to PR creator.

Copilot AI review requested due to automatic review settings September 5, 2026 13:10

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The changes directly address the null-dereference crash at the Cython boundary and are covered by targeted regression tests for the reported cases.

Review details
  • Files reviewed: 6/6 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@AlenkaF AlenkaF left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there any other instances of this pattern anywhere else in the Cython bindings? They might produce other errors (like AttributeError) but would still benefit from similar change.

Comment thread python/pyarrow/array.pxi
@staticmethod
def from_buffers(DataType type, int64_t length, buffers, Array dictionary,
int64_t null_count=-1, int64_t offset=0):
def from_buffers(DataType type not None, int64_t length, buffers,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense to also update other from_buffer() methods (base Array class too)?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in e031131.

@AlenkaF

AlenkaF commented Sep 10, 2026

Copy link
Copy Markdown
Member

Are there any other instances of this pattern anywhere else in the Cython bindings?

Ah, ok, there is another PR aiming at other files in PyArrow: #51163

For future work we could try keeping the number of PRs down and tackle similar issues in one.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
Copilot AI review requested due to automatic review settings September 12, 2026 04:57

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

No unresolved issues were identified, and focused regression tests cover the changes.

Review details
  • Files reviewed: 6/6 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@github-actions

Copy link
Copy Markdown

⚠️ GitHub issue #51043 has been automatically assigned in GitHub to PR creator.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Python] Several APIs segfault when required Arrow object arguments are None

3 participants