Skip to content

(#765) Restore Hadoop ByteBuffer compression compatibility - #766

Merged
xerial merged 1 commit into
xerial:mainfrom
vjanelle:fix-hadoop-output-buffer-765
Oct 5, 2026
Merged

xerial merged 1 commit into
xerial:mainfrom
vjanelle:fix-hadoop-output-buffer-765

Conversation

@vjanelle

@vjanelle vjanelle commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Fixes #765.

Hadoop 3.5.0's SnappyCompressor sets its output buffer limit to zero before calling Snappy.compress. The bounds check added in #739 uses remaining(), so it rejects that buffer even when its allocated capacity is sufficient and causes Snappy-compressed SequenceFile writes to fail.

Validate capacity() - position() against maxCompressedLength(input.remaining()) before invoking JNI. This restores the existing output-buffer behavior while preserving the check against insufficient allocated space, including output offsets and slice boundaries. The decompression bounds check protecting against #728 remains unchanged. Update the Javadoc to make the compression output-buffer semantics explicit.

Add five regression tests covering empty output limits, nonzero positions, slices, insufficient capacity after the output position, and undersized slices backed by larger allocations.

Validation:

  • ./sbt -batch testFull: 144 passed, 1 ignored, 0 failures.
  • Direct Hadoop 3.5.0 SnappyCompressor reproduction: successful compression and roundtrip with the fix; the prior check produces the exact reported exception.
  • New empty-limit regression cases fail before the fix; all 39 bounds-check tests pass afterward.

The complete Spark suite has not been rerun with this change.

@xerial
xerial merged commit 2dc4622 into xerial:main Oct 5, 2026
13 checks passed
@vjanelle
vjanelle deleted the fix-hadoop-output-buffer-765 branch October 5, 2026 20:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ByteBuffer compression bounds check breaks Hadoop SnappyCompressor with an empty output limit

2 participants