Apache Spark vs VecRuntime vs DataFusion Comet on TPC-DS 1 TB

VecRuntime runs Spark SQL's filters, projections, aggregates, sorts and joins on Arrow-layout batches with the Java Vector API -- on the JVM, no native code -- and moves batches between executors over its own Arrow Flight shuffle. Apache DataFusion Comet offloads the same operators to a native Rust engine. This page compares both against plain Apache Spark on the TPC-DS 1 TB workload on Amazon EKS, all three on identical hardware, data and Spark settings.

TL;DR. Over the 103 TPC-DS queries at 1 TB, VecRuntime finished in 2,420 s against Spark's 3,313 s -- 1.37x, 27% less runtime, faster than Spark on 88 of 103 queries (best q29: 3.49x; largest regression q18: 35%). Comet finished in 2,178 s -- 1.52x, 34% less, faster on 95 of 103 (best q6: 4.43x; largest regression q32: 63%). Between the two, VecRuntime is faster on 25 of 103 queries and leads on the heavy joins (q93, q64, q29, q50, q23b); Comet leads on the scan-bound and aggregate-heavy ones (q9, q28, q88, q4, q11, q67). Every engine returned Spark's results (one footnoted exception, q65). Comet Native Scan + VecRuntime finished in 2,324 s -- 1.43x, 30% less, faster than Spark on 87 of 103 (best q29: 3.72x; largest regression q83: 163%), and faster than VecRuntime on Spark's reader on 56 of 103. It is VecRuntime's operators and shuffle on Comet's native Parquet reader: the scan-bound queries drop to Comet's level (q28 82.5 s, q9 65.4 s, q88 97.1 s), while small queries pay a fixed cost per scan (q83 1.4 to 4.5 s). It ran in a separate cluster session, see the methodology.
Apache Spark 4.1.3
3,313 s
baseline, 103 queries
VecRuntime
2,420 s
1.37x · 27% less runtime
DataFusion Comet 1.2.0-SNAPSHOT (+ #6268, #6270)
2,178 s
1.52x · 34% less runtime
Comet Native Scan + VecRuntime
2,324 s
1.43x · 30% less runtime

Summary

EngineCompletion time (s)SpeedupFaster than Spark onExecutor time (h)GC (h)Shuffle read (TB)
Apache Spark 4.1.33,313.2baseline--80.60.600.94
VecRuntime2,419.71.37x (27% less)88 / 10358.50.430.47
DataFusion Comet 1.2.0-SNAPSHOT (+ #6268, #6270)2,177.91.52x (34% less)95 / 10353.00.010.66
Comet Native Scan + VecRuntime2,324.21.43x (30% less)87 / 10355.10.130.48

Benchmark infrastructure

Methodology. Spark, VecRuntime and Comet ran one after another in one cluster session on 2026-09-30 (morning), on the same nine nodes, each alone on the cluster, over the same S3 data with the same Spark settings; only the execution engine and its own memory split differ (every engine has 50 GB per executor; where it puts them follows where it allocates). Comet Native Scan + VecRuntime ran the same afternoon in a second cluster session on the same node group, availability zone, image and settings; S3 throughput varies between sessions, so its comparison with the other three carries that uncertainty. Each query ran once after the plan was compiled; the time is the wall-clock of the query's execution as the runner measures it.

Test environment

ComponentConfiguration
DatasetTPC-DS scale factor 1000 (1 TB), Parquet on Amazon S3 (103 query variants, one measured iteration each, no warm-up)
ClusterAmazon EKS 1.36; 9 x m5.4xlarge (16 vCPU, 64 GB, x86-64 with AVX-512), 300 GB root volume; one node group in one availability zone (us-east-1b), the driver on the ninth node
Executors8 executors x 13 cores x 50 GB each (Spark: 20 GB heap / 30 GB overhead; VecRuntime: 30 GB heap / 20 GB overhead; Comet: 20 GB heap / 6 GB overhead / 24 GB off-heap; Comet Native Scan + VecRuntime: 22 GB heap / 18 GB overhead / 10 GB off-heap); driver 2 cores x 4 GB
StorageAmazon S3 through an S3 gateway VPC endpoint, Hadoop 3.4.3 S3A with the Analytics Accelerator input stream (the default in 3.4.3)

Versions

ComponentVersion
Spark4.1.3
Scala2.13
JDKAmazon Corretto 25
VecRuntimemain at cb755d1 (#554 in-place dictionary decode)
Comet1.2.0-SNAPSHOT, a source build: apache/datafusion-comet main b58b2f3a with the unmerged fixes #6268 (q5) and #6270 (q64) applied, native library for x86-64-v3
Hadoop3.4.3

Configuration

Common to all three engines:

spark.sql.shuffle.partitions=300
# spark.sql.adaptive.advisoryPartitionSizeInBytes left at Spark's default (64 MB); no coalescePartitions.minPartitionNum
spark.eventLog.enabled=true

VecRuntime:

spark.plugins=io.vecruntime.spark.VectorPlugin
spark.shuffle.manager=org.apache.spark.sql.vecruntime.shuffle.VectorShuffleManager
spark.vecruntime.exec.strictFloatingPoint=false   # Comet's default too
spark.vecruntime.shuffle.aqe.mapSizeScaling=true   # AQE sees Spark-scale map output sizes (#514)
spark.vecruntime.shuffle.aqe.sparkCompressionRatio=0   # the uncompressed-bytes ratio (#514)
AOT class-data cache off
--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED (driver and executors)

Comet:

spark.plugins=org.apache.spark.CometPlugin
spark.shuffle.manager=org.apache.spark.sql.comet.execution.shuffle.CometShuffleManager
spark.memory.offHeap.enabled=true
spark.memory.offHeap.size=24g
spark.comet.exec.enabled=true
spark.comet.scan.enabled=true
spark.comet.exec.shuffle.enabled=true
spark.comet.exec.shuffle.mode=auto
spark.comet.cast.allowIncompatible=true
spark.comet.explainFallback.enabled=true

Comet Native Scan + VecRuntime:

spark.plugins=org.apache.spark.CometPlugin,io.vecruntime.spark.VectorPlugin
spark.shuffle.manager=org.apache.spark.sql.vecruntime.shuffle.VectorShuffleManager
spark.vecruntime.shuffle.enabled=true
spark.comet.enabled=true
spark.comet.scan.enabled=true   # CometNativeScanExec: DataFusion's Rust Parquet reader and object_store S3 I/O
spark.comet.exec.enabled=true   # with every Comet operator switched off: filter, project, aggregate, joins, sort, window, ...
spark.comet.exec.shuffle.enabled=false
spark.memory.offHeap.enabled=true
spark.memory.offHeap.size=10g   # executors 22 GB heap / 18 GB overhead (16 GB direct) / 10 GB off-heap = 50 GB
# plus VecRuntime's keys above (strictFloatingPoint=false, mapSizeScaling=true, sparkCompressionRatio=0)

Performance results

Per query

Seconds per query, all 103, the three engines side by side (hover for values; click a legend entry to hide an engine).

Speedup over Spark, per query

Spark's time divided by the engine's; above 1 is faster than Spark. Log scale.

Performance distribution

VecRuntime vs Spark

RangeQueriesShare
20%+ improvement5049%
10-20% improvement1817%
within ±10%2928%
10-20% degradation22%
20%+ degradation44%

Comet vs Spark

RangeQueriesShare
20%+ improvement7472%
10-20% improvement1414%
within ±10%1313%
10-20% degradation11%
20%+ degradation11%

Comet Native Scan + VecRuntime vs Spark

RangeQueriesShare
20%+ improvement5452%
10-20% improvement2221%
within ±10%2019%
10-20% degradation44%
20%+ degradation33%

Top 10 improvements -- VecRuntime

QuerySpark (s)VecRuntime (s)Speedup
q2934.39.83.49x (+71%)
q9735.712.72.81x (+64%)
q610.13.92.60x (+62%)
q93139.357.02.44x (+59%)
q23b281.7117.72.39x (+58%)
q158.63.62.35x (+57%)
q8120.810.12.05x (+51%)
q548.04.02.02x (+50%)
q686.93.61.93x (+48%)
q457.64.01.92x (+48%)

Regressions -- VecRuntime

QuerySpark (s)VecRuntime (s)Degradation
q189.212.435% slower (0.74x)
q122.22.931% slower (0.76x)
q365.46.724% slower (0.81x)
q9910.112.322% slower (0.82x)
q33.64.216% slower (0.86x)
q1142.248.214% slower (0.88x)
q211.81.98% slower (0.93x)
q76.67.17% slower (0.93x)
q24b100.6105.65% slower (0.95x)
q28110.6115.65% slower (0.96x)

Top 10 improvements -- Comet

QuerySpark (s)Comet (s)Speedup
q610.12.34.43x (+77%)
q9735.712.42.88x (+65%)
q8732.112.32.60x (+62%)
q8120.88.52.45x (+59%)
q158.63.62.38x (+58%)
q67126.455.02.30x (+56%)
q23b281.7123.52.28x (+56%)
q228.43.72.25x (+56%)
q211.80.82.22x (+55%)
q735.72.62.22x (+55%)

Regressions -- Comet

QuerySpark (s)Comet (s)Degradation
q321.52.463% slower (0.61x)
q713.23.713% slower (0.88x)
q772.42.68% slower (0.93x)
q6224.426.17% slower (0.94x)
q552.12.25% slower (0.95x)
q4945.147.04% slower (0.96x)
q7642.443.94% slower (0.96x)
q6610.010.00% slower (1.00x)

Top 10 improvements -- Comet Native Scan + VecRuntime

QuerySpark (s)Mixed (s)Speedup
q2934.39.23.72x (+73%)
q93139.352.42.66x (+62%)
q9735.713.82.59x (+61%)
q5066.827.92.39x (+58%)
q23b281.7119.42.36x (+58%)
q735.72.52.27x (+56%)
q686.93.12.22x (+55%)
q8120.810.02.09x (+52%)
q95115.457.12.02x (+50%)
q457.63.82.01x (+50%)

Regressions -- Comet Native Scan + VecRuntime

QuerySpark (s)Mixed (s)Degradation
q831.74.5163% slower (0.38x)
q6610.013.131% slower (0.76x)
q33.64.628% slower (0.78x)
q365.46.418% slower (0.85x)
q9910.111.918% slower (0.85x)
q39a5.96.715% slower (0.87x)
q76.67.514% slower (0.88x)
q1142.246.310% slower (0.91x)
q122.22.49% slower (0.91x)
q6224.426.38% slower (0.93x)

Where each engine wins

The heavy joins are VecRuntime's: q93 57.0 s against Spark's 139.3 and Comet's 85.9; q64 48.6 against 92.7 and 63.8; q50 38.4 against 66.8 and 47.4; q29 9.8 against 34.3 and 15.6; q23b 117.7 against 281.7 and 123.5. Its shuffle moves 0.47 TB where Spark's moves 0.94 and Comet's 0.66. Comet leads where the scan dominates -- q9 60.1 s against VecRuntime's 87.1, q28 84.6 against 115.6, q88 96.5 against 111.2 -- and on the heavy aggregates: q4 46.4 against 86.4, q11 27.4 against 48.2, q67 55.0 against 66.8. Comet Native Scan + VecRuntime separates the two: on Comet's Parquet reader VecRuntime's operators take q28 to 82.5 s, q9 to 65.4, q88 to 97.1 and q44 to 26.5 (Comet 26.9), and q50 to 27.9 and q93 to 52.4, so the scan-bound gap is the reader's; q4 (84.1), q11 (46.3) and q67 (68.2) barely move, so that gap is in the aggregates. VecRuntime's losses to Spark are q18 (12.4 s against 9.2, a cold query; the AOT cache is off here, #558), q99, q11, and the scan-bound q28 and q24b.

Notes

  • Every engine returned Spark's row counts on every query, and Spark's checksums on every query except q65, whose result has ties that every engine orders differently.
  • The three legs ran on 2026-09-30 in one cluster session on the same nine nodes, one after another (Spark, VecRuntime, Comet), each alone on the cluster. S3 throughput varies between sessions, so only legs from one session are compared.
  • Comet is a source build of main with two unmerged fixes: #6268, without which q5 fails at 1 TB with native scans, and #6270, without which q64 loses rows (apache/datafusion-comet#6264; 0 rows against 12,185 in the earlier run, #6133). With both, q5 and q64 return Spark's rows and checksums.
  • AQE is at its defaults for every engine. The earlier x86 run used a 128m advisory size and coalescePartitions.minPartitionNum=208 (deprecated in Spark 3.2+); its numbers are in docs/results.md.
  • Comet Native Scan + VecRuntime returned Spark's row counts on all 103 queries and Spark's checksums on all but q65 (ties). It uses the same patched Comet build.

All queries

Seconds per query, the fastest engine in bold
QuerySparkVecRuntimeCometMixedVecRuntime speedupComet speedupMixed speedupNote
q113.412.912.313.81.04x1.09x0.97x
q254.952.041.449.51.06x1.33x1.11x
q33.64.23.44.60.86x1.05x0.78x
q492.386.446.484.11.07x1.99x1.10x
q547.027.423.130.51.72x2.04x1.54x
q610.13.92.36.32.60x4.43x1.60x
q76.67.14.87.50.93x1.39x0.88x
q87.03.83.44.31.84x2.08x1.64x
q988.187.160.165.41.01x1.47x1.35x
q108.35.95.06.81.41x1.64x1.21x
q1142.248.227.446.30.88x1.54x0.91x
q122.22.91.62.40.76x1.42x0.91x
q138.38.05.36.31.04x1.57x1.33x
q14a99.874.664.269.71.34x1.55x1.43x
q14b93.767.660.563.61.39x1.55x1.47x
q158.63.63.65.32.35x2.38x1.61x
q1635.126.622.521.11.32x1.56x1.66x
q1712.77.38.87.21.73x1.44x1.76x
q189.212.44.79.40.74x1.95x0.98x
q194.73.22.32.61.45x2.06x1.79x
q202.42.21.62.01.05x1.51x1.18x
q211.81.90.81.50.93x2.22x1.21x
q228.47.93.76.81.07x2.25x1.24x
q23a201.2106.9104.2111.31.88x1.93x1.81x
q23b281.7117.7123.5119.42.39x2.28x2.36x
q24a104.5102.692.3100.11.02x1.13x1.04x
q24b100.6105.686.6103.70.95x1.16x0.97x
q259.79.76.26.21.00x1.57x1.58x
q264.03.63.13.81.13x1.31x1.05x
q276.45.85.05.91.12x1.29x1.09x
q28110.6115.684.682.50.96x1.31x1.34x
q2934.39.815.69.23.49x2.20x3.72x
q3018.314.711.915.11.25x1.54x1.21x
q3114.411.211.810.81.28x1.22x1.33x
q321.51.32.41.31.16x0.61x1.15x
q334.12.72.22.71.52x1.83x1.49x
q346.64.63.74.01.44x1.78x1.63x
q3520.711.611.211.81.78x1.85x1.76x
q365.46.75.16.40.81x1.07x0.85x
q378.58.06.98.51.06x1.23x1.00x
q3830.722.315.123.31.38x2.03x1.32x
q39a5.94.73.56.71.25x1.68x0.87x
q39b5.34.33.34.61.24x1.61x1.17x
q4010.810.610.19.21.03x1.07x1.18x
q410.90.50.40.51.60x1.93x1.75x
q421.61.31.51.61.23x1.08x1.00x
q435.25.23.74.20.99x1.39x1.22x
q4434.633.926.926.51.02x1.29x1.31x
q457.64.05.13.81.92x1.49x2.01x
q467.77.76.16.71.00x1.25x1.14x
q4712.012.19.510.90.99x1.26x1.11x
q487.96.24.84.61.27x1.64x1.71x
q4945.143.847.048.31.03x0.96x0.93x
q5066.838.447.427.91.74x1.41x2.39x
q5124.615.212.114.81.62x2.03x1.67x
q521.51.11.11.11.38x1.39x1.39x
q534.54.33.73.71.04x1.21x1.23x
q548.04.04.34.52.02x1.85x1.76x
q552.11.62.21.71.27x0.95x1.21x
q563.92.21.92.11.79x2.05x1.90x
q577.46.54.76.51.14x1.59x1.15x
q583.72.42.02.41.51x1.81x1.50x
q5932.130.028.228.41.07x1.14x1.13x
q604.12.72.22.81.50x1.84x1.46x
q615.23.32.63.11.57x2.01x1.68x
q6224.424.926.126.30.98x0.94x0.93x
q634.54.63.84.20.99x1.19x1.08x
q6492.748.663.850.71.91x1.45x1.83x
q6528.622.913.921.31.25x2.06x1.35xties
q6610.09.810.013.11.02x1.00x0.76x
q67126.466.855.068.21.89x2.30x1.85x
q686.93.63.83.11.93x1.83x2.22x
q696.83.84.43.91.82x1.54x1.77x
q7011.18.37.78.71.34x1.44x1.28x
q713.22.63.72.81.26x0.88x1.17x
q7237.331.135.332.11.20x1.06x1.16x
q735.73.12.62.51.83x2.22x2.27x
q7445.332.226.933.81.41x1.68x1.34x
q7575.167.973.072.61.11x1.03x1.03x
q7642.442.243.942.51.00x0.96x1.00x
q772.42.02.62.21.22x0.93x1.10x
q78117.882.975.685.31.42x1.56x1.38x
q796.04.93.75.21.22x1.60x1.15x
q8050.038.842.039.31.29x1.19x1.27x
q8120.810.18.510.02.05x2.45x2.09x
q8219.116.414.615.31.16x1.30x1.24x
q831.71.41.34.51.22x1.35x0.38x
q8420.617.417.320.31.18x1.19x1.01x
q8522.519.517.016.81.15x1.32x1.34x
q865.84.94.34.11.18x1.36x1.40x
q8732.117.712.320.71.81x2.60x1.55x
q88119.1111.296.597.11.07x1.23x1.23x
q895.55.14.54.51.08x1.21x1.20x
q9042.034.933.535.21.20x1.26x1.19x
q913.82.42.32.01.59x1.68x1.86x
q922.01.51.71.51.35x1.19x1.36x
q93139.357.085.952.42.44x1.62x2.66x
q9466.350.756.253.91.31x1.18x1.23x
q95115.462.352.957.11.85x2.18x2.02x
q9619.815.113.713.41.31x1.45x1.48x
q9735.712.712.413.82.81x2.88x2.59x
q982.92.41.72.11.19x1.73x1.37x
q9910.112.38.511.90.82x1.18x0.85x

Running the benchmark

The cluster runner, the Spark-on-Kubernetes manifests, the image and the data generation are in the benchmark runner's README: run-matrix.sh renders a SparkApplication per engine configuration (spark, vector-shuffle, comet, ...) and writes one JSON-lines result file per run; this page is rendered from three of them by benchmarks/scripts/render-benchmark-page.py.

TPC-DS is a benchmark of the Transaction Processing Performance Council; these results are not audited TPC results and are not comparable to published TPC-DS results. Times are the median of one measured iteration per query on the cluster described above; run-to-run variation on the heavy queries is a few percent.