You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
ByteBuffer compression bounds check breaks Hadoop SnappyCompressor with an empty output limit #765
Snappy's new compress(ByteBuffer, ByteBuffer) bounds check rejects Hadoop's normal output-buffer setup, causing Snappy-compressed SequenceFile writes to fail with snappy-java 1.1.10.10 and Hadoop 3.5.0.
The security checks should remain. The compatibility problem is that the compression check uses the output buffer's current limit rather than its allocated capacity after the output position.
Observed failure and impact
Apache Spark's org.apache.spark.FileSuite test SequenceFile (compressed) - snappy fails while closing a block-compressed SequenceFile writer. The test writes 100 ordinary ("abc", "abc") records using saveAsSequenceFile with Hadoop's SnappyCodec.
The relevant exception chain is:
java.lang.IllegalArgumentException:
not enough space for output: need 148 bytes, but only 0 remaining
at org.xerial.snappy.Snappy.checkOutputSpace(Snappy.java:622)
at org.xerial.snappy.Snappy.compress(Snappy.java:160)
at org.apache.hadoop.io.compress.snappy.SnappyCompressor.compressDirectBuf(SnappyCompressor.java:282)
at org.apache.hadoop.io.compress.snappy.SnappyCompressor.compress(SnappyCompressor.java:210)
at org.apache.hadoop.io.compress.BlockCompressorStream.compress(BlockCompressorStream.java:153)
at org.apache.hadoop.io.compress.BlockCompressorStream.finish(BlockCompressorStream.java:146)
at org.apache.hadoop.io.SequenceFile$BlockCompressWriter.writeBuffer(SequenceFile.java:1631)
at org.apache.hadoop.io.SequenceFile$BlockCompressWriter.sync(SequenceFile.java:1647)
at org.apache.hadoop.io.SequenceFile$BlockCompressWriter.close(SequenceFile.java:1671)
This is a failure in the production write path, rather than an assertion about compression ratios or a test that intentionally supplies an undersized allocation. Spark aborts the write job. The observed run used Linux aarch64 and Java 17; the behavior was also reproduced independently on macOS aarch64 with Java 17.
Minimal reproduction without Spark or Hadoop dependencies
The output has 65,536 bytes of allocated capacity, but zero remaining() bytes because its limit is zero. Compression requires a worst-case output bound of 148 bytes for this input. The new check throws the exception above despite the sufficiently sized allocation.
Why the buffer state occurs
Hadoop 3.5.0's SnappyCompressor.compress() resets its direct output buffer with:
It then calls Snappy.compress() from compressDirectBuf() and sets the output limit to the returned compressed size.
Historically, Snappy's ByteBuffer compression API wrote starting at the output position and set the output limit after compression. Its output parameter documentation described the range as [pos()..]. The JNI implementation uses GetDirectBufferAddress() plus the output position; it does not use the buffer's existing limit as a native write boundary.
PR #739, commit 943c6043a292568575e8347af4507d015e802c75, added a check against compressed.remaining(). The check in the v1.1.10.10 tag is:
This adds an output-limit requirement that the existing Hadoop client does not satisfy. PR #741 retained that requirement while refactoring the compressed-length validation. The bounds-check tests identify compression protection with #732; #728 concerns the separate decompression overload.
Proposed compatibility-preserving correction
Check the destination buffer view's capacity after the actual output position:
This accepts Hadoop's empty-limit buffer while still rejecting insufficient allocated space before JNI compression. Accounting for cPos prevents accepting a buffer whose total capacity is sufficient but whose available capacity after its position is not. Using the destination view's capacity also respects a slice's boundaries, rather than the larger parent allocation.
The output position remains unchanged, and the output limit is still set to cPos + compressedSize. The Javadoc should explicitly describe these output-buffer semantics.
The Snappy.uncompress(ByteBuffer, ByteBuffer) check introduced for #728 should remain unchanged by this compression fix. Removing bounds checks or reverting the security fixes is unnecessary.
Validation performed locally
I compared the cached 1.1.10.5 library with Java sources compiled from the v1.1.10.10 tag, using existing prebuilt native libraries. I also compiled a temporary variant with only the capacity-based compression check changed; the native code was not rebuilt.
1.1.10.5 successfully compresses and roundtrips through Hadoop 3.5.0's SnappyCompressor.
The v1.1.10.10 Java source fails through that same compressor with the exact need 148 bytes, but only 0 remaining error.
The capacity-based variant successfully compresses and roundtrips through the same Hadoop compressor.
Five additional regression tests cover an empty output limit, an empty limit at a nonzero output position, an empty-limit output slice, insufficient capacity after the output position, and an undersized slice backed by a larger allocation.
Before the correction, the three empty-limit cases fail. Afterward, all 39 bounds-check tests pass. Undersized allocations, slices, and available capacity after an output position are still rejected.
The complete Spark suite has not been rerun with this correction. The focused reproduction isolates the compatibility issue without requiring a Spark build.
Coverage gap
The existing SnappyHadoopCompatibleOutputStreamTest.testXerialCompressionHadoopDecompressionCodec is annotated @Ignore and only checks Hadoop decompression of Xerial-compressed data. It does not exercise Hadoop's compression output-buffer setup. A regression test for the empty-limit output behavior would prevent this particular compatibility break from recurring.
Snappy's new
compress(ByteBuffer, ByteBuffer)bounds check rejects Hadoop's normal output-buffer setup, causing Snappy-compressed SequenceFile writes to fail with snappy-java 1.1.10.10 and Hadoop 3.5.0.The security checks should remain. The compatibility problem is that the compression check uses the output buffer's current limit rather than its allocated capacity after the output position.
Observed failure and impact
Apache Spark's
org.apache.spark.FileSuitetestSequenceFile (compressed) - snappyfails while closing a block-compressed SequenceFile writer. The test writes 100 ordinary("abc", "abc")records usingsaveAsSequenceFilewith Hadoop'sSnappyCodec.The relevant exception chain is:
This is a failure in the production write path, rather than an assertion about compression ratios or a test that intentionally supplies an undersized allocation. Spark aborts the write job. The observed run used Linux aarch64 and Java 17; the behavior was also reproduced independently on macOS aarch64 with Java 17.
Minimal reproduction without Spark or Hadoop dependencies
The output has 65,536 bytes of allocated capacity, but zero
remaining()bytes because its limit is zero. Compression requires a worst-case output bound of 148 bytes for this input. The new check throws the exception above despite the sufficiently sized allocation.Why the buffer state occurs
Hadoop 3.5.0's
SnappyCompressor.compress()resets its direct output buffer with:It then calls
Snappy.compress()fromcompressDirectBuf()and sets the output limit to the returned compressed size.Historically, Snappy's ByteBuffer compression API wrote starting at the output position and set the output limit after compression. Its output parameter documentation described the range as
[pos()..]. The JNI implementation usesGetDirectBufferAddress()plus the output position; it does not use the buffer's existing limit as a native write boundary.PR #739, commit
943c6043a292568575e8347af4507d015e802c75, added a check againstcompressed.remaining(). The check in the v1.1.10.10 tag is:This adds an output-limit requirement that the existing Hadoop client does not satisfy. PR #741 retained that requirement while refactoring the compressed-length validation. The bounds-check tests identify compression protection with #732; #728 concerns the separate decompression overload.
Proposed compatibility-preserving correction
Check the destination buffer view's capacity after the actual output position:
The enforced invariant remains:
This accepts Hadoop's empty-limit buffer while still rejecting insufficient allocated space before JNI compression. Accounting for
cPosprevents accepting a buffer whose total capacity is sufficient but whose available capacity after its position is not. Using the destination view's capacity also respects a slice's boundaries, rather than the larger parent allocation.The output position remains unchanged, and the output limit is still set to
cPos + compressedSize. The Javadoc should explicitly describe these output-buffer semantics.The
Snappy.uncompress(ByteBuffer, ByteBuffer)check introduced for #728 should remain unchanged by this compression fix. Removing bounds checks or reverting the security fixes is unnecessary.Validation performed locally
I compared the cached 1.1.10.5 library with Java sources compiled from the v1.1.10.10 tag, using existing prebuilt native libraries. I also compiled a temporary variant with only the capacity-based compression check changed; the native code was not rebuilt.
SnappyCompressor.need 148 bytes, but only 0 remainingerror.SnappyBoundsCheckTesttests pass with the proposed correction, including the Off heap OOB Write in ByteBuffer Uncompress #728 decompression test.The complete Spark suite has not been rerun with this correction. The focused reproduction isolates the compatibility issue without requiring a Spark build.
Coverage gap
The existing
SnappyHadoopCompatibleOutputStreamTest.testXerialCompressionHadoopDecompressionCodecis annotated@Ignoreand only checks Hadoop decompression of Xerial-compressed data. It does not exercise Hadoop's compression output-buffer setup. A regression test for the empty-limit output behavior would prevent this particular compatibility break from recurring.