Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
d7ae3bb
perf: project cached batches by buffer selection, prune on collated s…
andygrove Aug 28, 2026
f59c9dc
refactor: use Spark's interpreted ordering for bounds, hoist projecti…
andygrove Aug 29, 2026
ccd469e
fix: relocate the arrow-compression service file when shading
andygrove Aug 29, 2026
e71d802
test: drop the cache leak test that depends on zstd corruption detection
andygrove Aug 29, 2026
d8d4196
feat: enable Comet's in-memory cache by default
andygrove Sep 2, 2026
4010f78
test: cover nested columns in the cached-batch projection tests and b…
andygrove Sep 5, 2026
8c19267
fix: drop a redundant string interpolator flagged by scalafix Redunda…
andygrove Sep 5, 2026
1699665
Merge remote-tracking branch 'apache/main' into feat/cache-buffer-sel…
andygrove Sep 7, 2026
792a465
Merge branch 'main' into feat/cache-enabled-by-default
andygrove Sep 8, 2026
bac454e
Merge remote-tracking branch 'apache/main' into feat/cache-buffer-sel…
cincrement Sep 10, 2026
f05c204
Merge branch 'main' into feat/cache-buffer-selection-projection
andygrove Sep 14, 2026
3e74e9d
review: check the cached layout, and address the rest of the review
andygrove Sep 15, 2026
b978e54
review: own the write-side compression buffers, fix the activation ex…
andygrove Sep 15, 2026
dbf487b
Merge remote-tracking branch 'origin/feat/cache-buffer-selection-proj…
andygrove Sep 21, 2026
a71e8cb
Merge remote-tracking branch 'apache/main' into feat/cache-enabled-by…
andygrove Sep 21, 2026
e483d08
fix: write cached batches to the schema width, not the batch width
andygrove Sep 21, 2026
3cf15ac
Merge branch 'fix/cache-wide-columnar-batch' into feat/cache-enabled-…
andygrove Sep 21, 2026
c7f1ce3
Merge branch 'main' into feat/cache-enabled-by-default
andygrove Sep 22, 2026
89cd107
Merge remote-tracking branch 'apache/main' into HEAD
andygrove Sep 24, 2026
d5983c2
Merge remote-tracking branch 'apache/main' into feat/cache-enabled-by…
andygrove Sep 24, 2026
299d381
Merge remote-tracking branch 'apache/main' into feat/cache-enabled-by…
andygrove Sep 25, 2026
a371d64
fix: report a cached relation's decoded size to the planner, not its …
andygrove Sep 29, 2026
6899380
fix: keep Spark's cache scan for a relation whose cached plan records…
andygrove Sep 29, 2026
d84b1f0
Merge remote-tracking branch 'apache/main' into feat/cache-enabled-by…
andygrove Sep 30, 2026
0ba6716
test: install Comet's cache serializer in the Spark SQL test sessions
andygrove Sep 30, 2026
37560b8
test: fix three Spark SQL test adaptations for Comet's cache format
andygrove Sep 30, 2026
c57eb9d
Merge branch 'fix/cache-decoded-size-stats' into feat/cache-enabled-b…
andygrove Oct 1, 2026
75b714b
Merge branch 'fix/cache-observed-metrics' into feat/cache-enabled-by-…
andygrove Oct 1, 2026
45dcd0e
Merge branch 'main' into feat/cache-enabled-by-default
andygrove Oct 1, 2026
6b6b90f
Merge apache/main into feat/cache-enabled-by-default
andygrove Oct 2, 2026
29edf93
fix: keep Spark's cache format when Kryo would reject Comet's or Come…
andygrove Oct 2, 2026
b53ebdd
Merge branch 'fix/cache-serializer-kryo-gate' into andygrove/5634-fee…
andygrove Oct 2, 2026
567956c
docs: describe the in-memory cache default in the upgrade guide
andygrove Oct 2, 2026
6cf61b6
test: benchmark the in-memory cache against Spark's format with AQE on
andygrove Oct 2, 2026
d124b1f
feat: record why Spark scans a relation cached in Comet's format when…
andygrove Oct 2, 2026
a09127b
Merge branch 'feat/cache-spark-read-fallback-reason' into andygrove/5…
andygrove Oct 2, 2026
73326f8
docs: say an application keeps the cache format it started with
andygrove Oct 2, 2026
5a96d6d
Merge branch 'fix/cache-serializer-kryo-gate' into andygrove/5634-fee…
andygrove Oct 2, 2026
972c47f
test: format the in-memory cache benchmark
andygrove Oct 2, 2026
5528fc4
fix: bind the fallback reason result for the strict Scala warnings build
andygrove Oct 2, 2026
3fb6673
Merge branch 'feat/cache-spark-read-fallback-reason' into andygrove/5…
andygrove Oct 2, 2026
b9ab592
docs: compare the in-memory cache with Spark's format as an applicati…
andygrove Oct 2, 2026
667ea36
fix: count Kryo registrations however the application made them
andygrove Oct 2, 2026
15dc3a9
Merge branch 'fix/cache-serializer-kryo-gate' into andygrove/5634-fee…
andygrove Oct 2, 2026
1330b75
docs: describe the Kryo condition by registration in the upgrade guide
andygrove Oct 2, 2026
5fd331e
Merge apache/main into feat/cache-enabled-by-default
andygrove Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/pr_build_linux.yml
Original file line number Diff line number Diff line change
Expand Up @@ -516,6 +516,8 @@ jobs:
org.apache.comet.exec.CometInMemoryCacheSuite
org.apache.comet.exec.CometInMemoryCachePruningSuite
org.apache.comet.exec.CometInMemoryCacheKryoSuite
org.apache.comet.exec.CometInMemoryCacheKryoUnregisteredSuite
org.apache.comet.exec.CometInMemoryCacheKryoClassesToRegisterSuite
org.apache.comet.exec.CometGenerateExecSuite
org.apache.comet.exec.CometWindowExecSuite
org.apache.comet.exec.CometJoinSuite
Expand Down
2 changes: 2 additions & 0 deletions .github/workflows/pr_build_macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -220,6 +220,8 @@ jobs:
org.apache.comet.exec.CometInMemoryCacheSuite
org.apache.comet.exec.CometInMemoryCachePruningSuite
org.apache.comet.exec.CometInMemoryCacheKryoSuite
org.apache.comet.exec.CometInMemoryCacheKryoUnregisteredSuite
org.apache.comet.exec.CometInMemoryCacheKryoClassesToRegisterSuite
org.apache.comet.exec.CometGenerateExecSuite
org.apache.comet.exec.CometWindowExecSuite
org.apache.comet.exec.CometJoinSuite
Expand Down
331 changes: 316 additions & 15 deletions dev/diffs/3.4.3.diff

Large diffs are not rendered by default.

374 changes: 356 additions & 18 deletions dev/diffs/3.5.9.diff

Large diffs are not rendered by default.

420 changes: 400 additions & 20 deletions dev/diffs/4.0.4.diff

Large diffs are not rendered by default.

493 changes: 466 additions & 27 deletions dev/diffs/4.1.3.diff

Large diffs are not rendered by default.

493 changes: 466 additions & 27 deletions dev/diffs/4.2.0.diff

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions docs/source/contributor-guide/plugin_overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,8 @@ and skips the remaining steps. Otherwise it:
- Appends `CometSparkSessionExtensions` to `spark.sql.extensions`, unless it is already listed.
- Sets `spark.sql.cache.serializer` to Comet's `ArrowCachedBatchSerializer` when
`spark.comet.exec.inMemoryCache.enabled=true`, unless the application has chosen a different serializer.
It leaves Spark's serializer in place where Comet could not scan its format natively or Kryo would
reject it; see [In-Memory Cache](../user-guide/latest/in-memory-cache.md).
- Registers `CometSource` with Spark's metrics system and adds `CometMetricsListener` to
`spark.sql.queryExecutionListeners` when `spark.comet.metrics.enabled=true`.
- Logs a warning for settings that are likely to cause problems, such as an unset `spark.executor.memoryOverhead`.
Expand Down
4 changes: 2 additions & 2 deletions docs/source/user-guide/latest/compatibility/operators.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,8 +35,8 @@ readable empty output files and their schema metadata.
## In-Memory Cache

Comet can store cached relations (`df.cache()`, `CACHE TABLE`) in Arrow format and scan them
natively. This is experimental and disabled by default; see [In-Memory Cache](../in-memory-cache.md)
for how to enable it. Comet does not replace a `spark.sql.cache.serializer` that the application
natively. This is experimental and enabled by default; see [In-Memory Cache](../in-memory-cache.md)
for how to turn it off. Comet does not replace a `spark.sql.cache.serializer` that the application
has already set. Relations whose schema Comet's Arrow writer does not support are cached in
Spark's default format, and their scans fall back to Spark. Reads that feed Spark operators rather
than Comet operators can be slower than Spark's cache.
Expand Down
90 changes: 69 additions & 21 deletions docs/source/user-guide/latest/in-memory-cache.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,24 +21,27 @@

Comet can store Spark's in-memory cache (`CACHE TABLE`, `df.cache()`, `df.persist()`) in an Arrow
format that Comet operators read directly. Without it, a cached table is stored in Spark's own
format and every scan of it has to convert each batch before Comet can continue, which shows up in
the plan as a `CometSparkColumnarToColumnar` above the cache scan.
format, which Comet operators cannot read. Under Comet's default settings the operators above the
cache scan then run on Spark. With `spark.comet.sparkToColumnar.enabled`, a
`CometSparkColumnarToColumnar` above the scan converts each batch for Comet operators instead.

This feature is **experimental and disabled by default**. Turn it on at startup, alongside the rest
of Comet's configuration:
This feature is **experimental and enabled by default**. To turn it off, set the config at startup,
alongside the rest of Comet's configuration:
Comment on lines +28 to +29

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This key did not exist in 1.0.0, so a user upgrading from 1.0.0 goes from Spark's cache format to Comet's without setting anything. With spark.kryo.registrationRequired=true and no CometKryoRegistrator, a df.cache() that spills to disk now fails with "Class is not registered" where it did not before. The plugin only logs a warning for that.

The versioning policy counts a new error under the same explicit configuration as a behavior change. Could you add an entry to the upgrade guide under the next release that covers the format change and the Kryo requirement? The policy asks for a spark.comet.legacy.* key, but spark.comet.exec.inMemoryCache.enabled=false already restores the old behavior, so naming that key in the entry seems enough. If you read the policy differently, it would be good to settle that here, since this is one of the first behavior changes since 1.0.0.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed on the upgrade guide entry, covering both the format change and the Kryo requirement. Moving the flip past 1.1.0 changes one premise, though. 1.1.0 ships this key with a default of false, so turning it on in the next release is a change to an existing key's default, which is the first case the policy lists. Let's settle the legacy-key question when this comes out of draft.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The upgrade guide has a 1.2.0 entry now. It covers the format change, how to keep Spark's format, and the cases where Comet keeps Spark's format without being asked.

For Kryo, #6537 (merged into this branch) does more than document the requirement. When Kryo requires registration and has not registered Comet's cached batch, the plugin now keeps Spark's format rather than installing one that Kryo would reject, and its startup warning says so. It asks a Kryo instance built from the application's conf, so registrations made through CometKryoRegistrator, another registrator or spark.kryo.classesToRegister all count. Two new suites run this end to end with a DISK_ONLY cache. In one, the application registers only Spark's cached batch. Without the change, its cache fails with Class is not registered: org.apache.spark.sql.comet.execution.arrow.CometCachedBatch, and with it the cache is stored as DefaultCachedBatch and reads back correctly. In the other, the application registers Comet's classes through spark.kryo.classesToRegister, and the cache keeps Comet's format.

That also answers the legacy-key question for me. With the gate in place, the flip changes no result and raises no new error under the same explicit configuration. What it changes is which operators run natively and the storage format underneath them. The config conventions exempt changes to which expressions and operators run natively, and spark.comet.exec.inMemoryCache.enabled=false restores the old format exactly. So the entry sits with the changes that need no legacy key, as the guide's introduction allows, and it names that setting. The versioning policy does list default changes among its examples, though. If you read that as applying even when results don't change, I can add a spark.comet.legacy.* key, but it would do exactly what spark.comet.exec.inMemoryCache.enabled=false already does.


```shell
$SPARK_HOME/bin/spark-shell \
... \
--conf spark.comet.exec.inMemoryCache.enabled=true
--conf spark.comet.exec.inMemoryCache.enabled=false
```

It has to be set before the `SparkContext` starts. Comet's driver plugin chooses
`spark.sql.cache.serializer` once, while the context is initializing, so a session that started
with the default goes on using Spark's cache format however the config is set afterwards. The
plugin installs Comet's serializer only if `spark.comet.enabled` and `spark.comet.exec.enabled`
are enabled at that point too, because an application that starts without native execution could
not scan Comet's format natively.
`spark.sql.cache.serializer` once, while the context is initializing, so an application keeps the
cache format it started with however the config is set afterwards. The plugin installs Comet's
serializer only if `spark.comet.enabled` and `spark.comet.exec.enabled` are enabled at that point
too, because an application that starts without native execution could not scan Comet's format
natively. It also keeps Spark's format when Comet shuffle is enabled but `spark.shuffle.manager` is
not one of Comet's shuffle managers, since Comet then disables itself, and when Kryo would reject
Comet's format; see [Kryo](#kryo).

## What changes when it is enabled

Expand Down Expand Up @@ -109,7 +112,7 @@ nowhere to record either that a column is dictionary encoded or the dictionary i

| Config | Default | Description |
| ------------------------------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| `spark.comet.exec.inMemoryCache.enabled` | `false` | Whether to store and scan Spark's in-memory cache in Comet's format. Read at startup. |
| `spark.comet.exec.inMemoryCache.enabled` | `true` | Whether to store and scan Spark's in-memory cache in Comet's format. Read at startup. |
| `spark.comet.exec.inMemoryCache.compression.codec` | `zstd` | Arrow IPC compression codec for cached data: `zstd` or `none`. Affects newly cached data only — a batch records the codec it was written with. |
| `spark.comet.exec.inMemoryCache.compression.zstd.level` | `1` | Compression level when the codec is `zstd`. Ignored otherwise. |

Expand Down Expand Up @@ -158,10 +161,43 @@ back to Spark row execution above the scan and the two columns stop measuring th
Read what this compares carefully. Comet execution is on in both columns, so the aggregation runs
on Comet either way and only the cache-scan boundary moves: on the left, Spark's
`InMemoryTableScanExec` feeds those same Comet operators through a `CometSparkColumnarToColumnar`
bridge; on the right, `CometInMemoryTableScan` feeds them directly. Both columns read the same
bridge, which the benchmark turns on with `spark.comet.sparkToColumnar.enabled`; on the right,
`CometInMemoryTableScan` feeds them directly. Both columns read the same
Comet-written `CometCachedBatch`. These numbers are therefore "keep the cached scan native" against
"fall back to a Spark cache scan and convert", not Comet against Spark execution, and not a
comparison with Spark's own cache format. That comparison is under [Limitations](#limitations).
comparison with Spark's own cache format, which follows.

### Against Spark's cache format

What turning the feature on changes for a query that Comet runs is measured against Spark's own
cache format by the benchmark's adaptive cases. Comet and AQE are on, Comet's other settings are at
their defaults, and the same 5M-row relation is cached in each format. The defaults leave
`spark.comet.sparkToColumnar.enabled` off, so Comet operators cannot read Spark's cache scan, and
with Spark's format the operators directly above the scan run on Spark. Measured on an AMD Ryzen 9
7950X3D (JDK 17, Spark 4.1, release build):

| Query shape | Spark's cache format | Comet's cache format | Relative |
| -------------------------- | -------------------: | -------------------: | -------: |
| Row count only (0 of 6) | 29 ms | 24 ms | 1.2x |
| Narrow projection (1 of 6) | 52 ms | 34 ms | 1.5x |
| 3 of 6 columns | 102 ms | 112 ms | 0.9x |
| Full projection (6 of 6) | 299 ms | 224 ms | 1.3x |

A Spark operator above the cache scan, standing in for any operator Comet does not support, is
measured the same way, with Comet's aggregate turned off. With Comet's format, the native scan feeds
that operator through a columnar-to-row transition:

| Query shape | Spark's cache format | Comet's cache format | Relative |
| -------------------------- | -------------------: | -------------------: | -------: |
| Row count only (0 of 6) | 39 ms | 16 ms | 2.4x |
| Narrow projection (1 of 6) | 50 ms | 27 ms | 1.8x |
| 3 of 6 columns | 97 ms | 112 ms | 0.9x |
| Full projection (6 of 6) | 303 ms | 299 ms | 1.0x |

Comet's format is as fast or faster in every shape but one: the read of three of the six columns,
all of them longs, is about 10% slower under either kind of operator. That cost is `zstd`
decompression. With the `none` codec, the same read is 2.7x faster than Spark's format with Comet
operators above the scan, and 1.6x faster with a Spark operator above it.

## Kryo

Expand All @@ -179,16 +215,27 @@ spark.kryo.registrator=org.apache.comet.CometKryoRegistrator

Comet cannot set `spark.kryo.registrator` for you the way it sets `spark.sql.cache.serializer`:
`KryoSerializer` reads it when `SparkEnv` builds the serializer, which happens before any plugin
runs. Without it, caching fails with a "Class is not registered" error that does not name this
feature. Comet's driver plugin warns at startup when it sees Kryo, `registrationRequired`, and no
registrator. Native broadcast needs the same registrator even when the cache is disabled; see
runs. Without it, Kryo would reject Comet's cached batch with a "Class is not registered" error
that does not name this feature. So when Kryo requires registration and has not registered
Comet's cached batch, Comet's driver plugin does not install Comet's serializer, and caches stay in
Spark's format. Registrations made another way, through a registrator of the application's own or
`spark.kryo.classesToRegister`, count as well. The plugin warns at startup when Kryo requires
registration and has not registered every class `CometKryoRegistrator` registers. An application
that sets `spark.sql.cache.serializer` to Comet's serializer itself gets the error instead. Native
broadcast needs the same registrator even when the cache is disabled; see
[Kryo serialization](installation.md#kryo-serialization).

Spark registers its own cached batch with Kryo only from Spark 4.1, so on earlier versions caching
in either format under `registrationRequired` needs a registrator. `CometKryoRegistrator` registers
Spark's cached batch too.

## Limitations

Reads that feed **Spark** operators rather than Comet ones are slower than Spark's own cache
format, and the narrower the read, the wider the gap. Measured by the same benchmark over the same
5M-row relation, with Comet off so that Spark operators consume the cached data:
Spark's own cache scan, `InMemoryTableScanExec`, reads Comet's format more slowly than Spark's, and
the narrower the read, the wider the gap. Spark's scan reads a cached relation when a session turns
Comet or its native execution off after caching, and when the relation's cached plan records
`Dataset.observe` metrics, and Comet records a fallback reason on the scan in either case. Measured
by the same benchmark over the same 5M-row relation, with Comet off:

| Read shape | Spark's cache format | Comet's cache format | Slowdown |
| ----------------------- | -------------------: | -------------------: | -------: |
Expand All @@ -197,8 +244,9 @@ format, and the narrower the read, the wider the gap. Measured by the same bench
| 3 of 6 columns | 98 ms | 331 ms | 3.4x |
| 6 of 6 columns | 410 ms | 623 ms | 1.5x |

This is why the feature is off by default. The cause is not yet established;
[#5485](https://github.com/apache/datafusion-comet/issues/5485) tracks it.
A Spark operator above Comet's native cache scan does not pay this; see [Performance](#performance).
This gap is the main reason the feature is still described as experimental. The cause is not yet
established; [#5485](https://github.com/apache/datafusion-comet/issues/5485) tracks it.

Comet's serializer exists because Spark's own Arrow cache format
([SPARK-57268](https://issues.apache.org/jira/browse/SPARK-57268)) is only available from Spark
Expand Down
11 changes: 6 additions & 5 deletions docs/source/user-guide/latest/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -269,8 +269,9 @@ If the application uses Kryo (`spark.serializer=org.apache.spark.serializer.Kryo

Without it, any query that uses Comet's native broadcast exchange, which is enabled by default,
fails with Kryo's "Class is not registered" error, for example on the first broadcast hash join.
The [in-memory cache](in-memory-cache.md#kryo) needs the same registrator. Set it before the
`SparkContext` is created: `KryoSerializer` reads it before Comet's plugin runs, so Comet cannot
add it for you. `spark.kryo.registrator` accepts a comma-separated list, so an application with
its own registrator can list both. Comet logs a warning at startup when Kryo requires registration
and this registrator is missing.
Comet's [in-memory cache](in-memory-cache.md#kryo) format needs the same registrations, and while
Kryo has not registered Comet's cached batch, Comet's plugin keeps caches in Spark's format. Set it
before the `SparkContext` is created: `KryoSerializer` reads it before Comet's plugin runs, so
Comet cannot add it for you. `spark.kryo.registrator` accepts a comma-separated list, so an
application with its own registrator can list both. Comet logs a warning at startup when Kryo
requires registration and has not registered the classes this registrator covers.
29 changes: 29 additions & 0 deletions docs/source/user-guide/latest/migration-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,35 @@ Treat setting one of these keys as a temporary measure. If you find you cannot s
legacy behavior, please open an issue describing your use case so it can be considered before the
key is removed.

## Upgrading to Comet 1.2.0

Comet `1.2.0` makes no behavior changes that need a `spark.comet.legacy.*` key. The changes below
need none either, but check whether any of them applies to your deployment.

### In-Memory Cache Enabled by Default

`spark.comet.exec.inMemoryCache.enabled` now defaults to `true`. An application that loads
`CometPlugin` now stores what it caches with `CACHE TABLE`, `df.cache()` or `df.persist()` in
Comet's Arrow format instead of Spark's, and Comet scans it natively. The format does not change
query results, but it can change performance. Spark's own cache scan reads Comet's format more
slowly than Spark's, which matters when a session turns Comet or its native execution off after
caching, and Comet records a fallback reason on such a scan. See
[In-Memory Cache](in-memory-cache.md#limitations).

The format is chosen once, when the application starts. To keep Spark's format, set
`spark.comet.exec.inMemoryCache.enabled=false` then. Comet also keeps Spark's format without that
setting when the application:

- starts with `spark.comet.enabled` or `spark.comet.exec.enabled` set to `false`.
- leaves Comet shuffle enabled without one of Comet's shuffle managers, so that Comet disables
itself.
- uses Kryo with `spark.kryo.registrationRequired=true` and has not registered Comet's cached
batch, because Kryo would reject it. To use Comet's format, register Comet's classes with
`spark.kryo.registrator=org.apache.comet.CometKryoRegistrator`; see
[Kryo](in-memory-cache.md#kryo).

An application that sets `spark.sql.cache.serializer` itself keeps the serializer it chose.

## Upgrading to Comet 1.1.0

Comet `1.1.0` makes no behavior changes that need a `spark.comet.legacy.*` key. The changes below
Expand Down
2 changes: 1 addition & 1 deletion docs/source/user-guide/latest/operators.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ omitted from the tables below and may be reconsidered based on demand:
| `LocalTableScanExec` | ⚠️ | Disabled by default; there is no acceleration advantage and this operator is typically only used in test code. Can be opted into via config ([#4393](https://github.com/apache/datafusion-comet/pull/4393)). |
| `EmptyRelationExec` | ✅ | Spark 4.0 and later. See [Empty Relations](compatibility/operators.md#empty-relations) for native-input support and writer fallback. |
| `RangeExec` | ⚠️ | Disabled by default. Set `spark.comet.exec.range.enabled=true` to generate the rows of `spark.range` and SQL `range()` in native code, so the operators above them run natively. It can be slower than Spark when those operators are only cheap expressions, such as a filter, which Spark compiles together with the range into one loop. Ranges whose arithmetic overflows the `Long` range fall back to Spark. |
| `InMemoryTableScanExec` | ⚠️ | Experimental, disabled by default. Set `spark.comet.exec.inMemoryCache.enabled=true` before the application starts so Comet installs its Arrow cache serializer. Relations with unsupported column types stay in Spark's cache format and fall back. See [In-Memory Cache](in-memory-cache.md). |
| `InMemoryTableScanExec` | ⚠️ | Experimental, enabled by default. Comet installs its Arrow cache serializer as the application starts, unless `spark.comet.exec.inMemoryCache.enabled` is false. Relations with unsupported column types stay in Spark's cache format and fall back. See [In-Memory Cache](in-memory-cache.md). |

## Projection and filtering

Expand Down
Loading
Loading