Apache Spark vs VecRuntime on TPC-DS 1 TB, AWS Graviton4
VecRuntime runs Spark SQL's filters, projections, aggregates, sorts and joins on Arrow-layout batches with the Java Vector API -- on the JVM, no native code -- and moves batches between executors over its own Arrow Flight shuffle. This page compares it against plain Apache Spark on the TPC-DS 1 TB workload on Amazon EKS with AWS Graviton4 (arm64) nodes, and sets both against the published x86 run from the same day, with the same settings and the same single-AZ layout (its own cluster session) (#253).
Summary
| Engine | Completion time (s) | Speedup | Faster than Spark on | Executor time (h) | GC (h) | Shuffle read (TB) |
|---|---|---|---|---|---|---|
| Apache Spark 4.1.3 | 2,470.2 | baseline | -- | 60.5 | 0.28 | 0.94 |
| VecRuntime | 1,892.7 | 1.31x (23% less) | 96 / 103 | 46.5 | 0.17 | 0.47 |
Graviton4 against x86
The x86 numbers are the published run: 9 x m5.4xlarge (16 vCPU, 64 GB, AVX-512), the same data, executors, Spark settings and single-AZ layout, measured the same day in its own cluster session.
| Spark (s) | VecRuntime (s) | VecRuntime speedup | geometric mean | |
|---|---|---|---|---|
| Graviton4 (m8g.4xlarge) | 2,470.2 | 1,892.7 | 1.31x | 1.31x |
| x86 (m5.4xlarge) | 3,313.3 | 2,419.6 | 1.37x | 1.31x |
| Graviton4 / x86, per query (geomean) | 0.76 | 0.76 |
Where VecRuntime's lead over Spark shrinks most on Graviton4 -- mostly queries where Spark itself gains more from Graviton4:
| Query | Graviton4 | x86 | Spark G/x | VecRuntime G/x |
|---|---|---|---|---|
| q92 | 0.49x | 1.36x | 0.73 | 2.01 |
| q58 | 0.90x | 1.51x | 0.59 | 0.99 |
| q29 | 2.22x | 3.49x | 0.48 | 0.75 |
| q45 | 1.25x | 1.92x | 0.62 | 0.94 |
| q54 | 1.35x | 2.02x | 0.56 | 0.83 |
| q41 | 1.12x | 1.59x | 0.61 | 0.87 |
| q95 | 1.39x | 1.85x | 0.56 | 0.75 |
| q93 | 1.87x | 2.44x | 0.71 | 0.93 |
Where it grows most -- mostly queries at parity or behind on x86:
| Query | Graviton4 | x86 | Spark G/x | VecRuntime G/x |
|---|---|---|---|---|
| q12 | 1.35x | 0.76x | 0.92 | 0.52 |
| q25 | 1.58x | 1.00x | 0.76 | 0.49 |
| q18 | 1.14x | 0.74x | 0.57 | 0.37 |
| q22 | 1.55x | 1.07x | 0.73 | 0.50 |
| q11 | 1.19x | 0.88x | 0.77 | 0.57 |
| q2 | 1.43x | 1.06x | 0.98 | 0.72 |
| q21 | 1.24x | 0.93x | 0.71 | 0.53 |
| q5 | 2.28x | 1.72x | 0.88 | 0.66 |
"G/x" is the query's time on Graviton4 over its time on x86 (below 1 = faster on Graviton4).
Benchmark infrastructure
Methodology. The two engines ran one after another on the same nodes, each alone on the cluster, over the same S3 data with the same Spark settings; only the execution engine and its own memory split differ (every engine has 50 GB per executor). Each query ran once after the plan was compiled; the time is the wall-clock of the query's execution as the runner measures it.
Test environment
| Component | Configuration |
|---|---|
| Dataset | TPC-DS scale factor 1000 (1 TB), Parquet on Amazon S3 (103 query variants, one measured iteration each, no warm-up) |
| Cluster | Amazon EKS 1.36; 9 x m8g.4xlarge (16 vCPU AWS Graviton4 / Neoverse V2 with SVE2, 64 GB, arm64), 300 GB root volume; one node group in one availability zone (us-east-1b), the driver on the ninth node |
| Executors | 8 executors x 13 cores x 50 GB each (Spark: 20 GB heap / 30 GB overhead; VecRuntime: 30 GB heap / 20 GB overhead); driver 2 cores x 4 GB |
| Storage | Amazon S3 through an S3 gateway VPC endpoint, Hadoop 3.4.3 S3A with the Analytics Accelerator input stream (the default in 3.4.3) |
Versions
| Component | Version |
|---|---|
| Spark | 4.1.3 |
| Scala | 2.13 |
| JDK | Amazon Corretto 25.0.4.1 (aarch64) |
| VecRuntime | main at cb755d1 (#554 in-place dictionary decode; arm64 image main-cb755d1-arm64) |
| Hadoop | 3.4.3 |
Configuration
Common to both engines:
spark.sql.shuffle.partitions=300 # spark.sql.adaptive.advisoryPartitionSizeInBytes left at Spark's default (64 MB); no coalescePartitions.minPartitionNum spark.eventLog.enabled=true
VecRuntime:
spark.plugins=io.vecruntime.spark.VectorPlugin spark.shuffle.manager=org.apache.spark.sql.vecruntime.shuffle.VectorShuffleManager spark.vecruntime.exec.strictFloatingPoint=false # Comet's default too spark.vecruntime.shuffle.aqe.mapSizeScaling=true # AQE sees Spark-scale map output sizes (#514) spark.vecruntime.shuffle.aqe.sparkCompressionRatio=0 # the uncompressed-bytes ratio (#514) AOT class-data cache off --add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED (driver and executors)
The arm64 platform
- HotSpot on Graviton4: UseSVE=2, MaxVectorSize=16 -- the Vector API's species are 128 bits wide, as on NEON.
- VecRuntime's platform probe picks its SVE paths: native compress (SVE COMPACT) for selection, and the broadcast-AND-compare lane masks, because VectorMask.fromLong is not intrinsified at 128-bit SVE on JDK 25 (#253, #484).
- The native codecs (snappy, zstd, lz4) and Netty's epoll transport load their aarch64 libraries; neither architecture has libhadoop, as on x86.
Performance results
Per query
Seconds per query on Graviton4, both engines side by side (hover for values; click a legend entry to hide an engine).
Performance distribution
| VecRuntime vs Spark | Queries | Share |
|---|---|---|
| 20%+ improvement | 51 | 50% |
| 10-20% improvement | 22 | 21% |
| within ±10% | 26 | 25% |
| 10-20% degradation | 2 | 2% |
| 20%+ degradation | 2 | 2% |
Top 10 improvements
| Query | Spark (s) | VecRuntime (s) | Speedup | x86 speedup |
|---|---|---|---|---|
| q6 | 5.5 | 1.8 | 3.06x (+67%) | 2.60x |
| q73 | 4.2 | 1.8 | 2.36x (+58%) | 1.82x |
| q15 | 5.3 | 2.3 | 2.34x (+57%) | 2.35x |
| q97 | 21.1 | 9.1 | 2.31x (+57%) | 2.81x |
| q5 | 41.5 | 18.2 | 2.28x (+56%) | 1.72x |
| q29 | 16.4 | 7.4 | 2.22x (+55%) | 3.49x |
| q68 | 5.2 | 2.5 | 2.06x (+52%) | 1.93x |
| q69 | 6.1 | 3.1 | 1.99x (+50%) | 1.82x |
| q23b | 159.7 | 80.6 | 1.98x (+50%) | 2.39x |
| q81 | 14.4 | 7.6 | 1.89x (+47%) | 2.06x |
Regressions
| Query | Spark (s) | VecRuntime (s) | Degradation | x86 speedup |
|---|---|---|---|---|
| q92 | 1.5 | 3.0 | 103% slower (0.49x) | 1.36x |
| q99 | 7.6 | 9.5 | 25% slower (0.80x) | 0.82x |
| q36 | 4.7 | 5.3 | 12% slower (0.89x) | 0.81x |
| q58 | 2.2 | 2.4 | 11% slower (0.90x) | 1.51x |
| q89 | 4.2 | 4.4 | 6% slower (0.94x) | 1.08x |
| q90 | 35.8 | 37.4 | 5% slower (0.96x) | 1.20x |
| q88 | 101.9 | 104.6 | 3% slower (0.97x) | 1.07x |
Notes
- Both engines returned the same row counts on every query and the same checksums on every query except q65, whose result has ties that each engine orders differently.
- Both legs ran on 2026-09-30 in one cluster session, Spark first and VecRuntime straight after, on the same nine nodes in one availability zone, each alone on the cluster. S3 throughput varies between sessions, so only runs from the same session are compared.
- The earlier Graviton4 runs (2026-09-26 at 1.23x, 2026-09-29 at 1.26x) had their nodes split across two availability zones; this run keeps all nine in one and adds #554.
- VecRuntime reports its map output sizes to AQE on Spark's scale (#511, #514): its columnar shuffle is 1.6-4.3x smaller than Spark's for the same rows, and without the scaling AQE packs up to twice Spark's rows into a task.
- The x86 reference ran the same day with the same image version, settings and single-AZ layout, in its own cluster session; the architecture comparison below therefore compares two sessions.
All queries
Seconds per query on Graviton4, the faster engine in bold, with the x86 run alongside
| Query | Spark | VecRuntime | Speedup | x86 Spark | x86 VecRuntime | x86 speedup | Note |
|---|---|---|---|---|---|---|---|
| q1 | 11.5 | 9.7 | 1.19x | 13.4 | 12.9 | 1.04x | |
| q2 | 53.7 | 37.6 | 1.43x | 54.9 | 52.0 | 1.06x | |
| q3 | 3.5 | 3.4 | 1.02x | 3.6 | 4.2 | 0.86x | |
| q4 | 50.1 | 49.7 | 1.01x | 92.3 | 86.4 | 1.07x | |
| q5 | 41.5 | 18.2 | 2.28x | 47.0 | 27.4 | 1.72x | |
| q6 | 5.5 | 1.8 | 3.06x | 10.1 | 3.9 | 2.60x | |
| q7 | 5.0 | 4.6 | 1.07x | 6.6 | 7.1 | 0.94x | |
| q8 | 5.1 | 2.9 | 1.73x | 7.0 | 3.8 | 1.84x | |
| q9 | 76.7 | 70.2 | 1.09x | 88.1 | 87.1 | 1.01x | |
| q10 | 5.3 | 3.9 | 1.34x | 8.3 | 5.9 | 1.41x | |
| q11 | 32.5 | 27.3 | 1.19x | 42.2 | 48.2 | 0.88x | |
| q12 | 2.1 | 1.5 | 1.35x | 2.2 | 2.9 | 0.76x | |
| q13 | 7.2 | 5.6 | 1.27x | 8.3 | 8.0 | 1.04x | |
| q14a | 83.1 | 55.1 | 1.51x | 99.8 | 74.6 | 1.34x | |
| q14b | 74.9 | 51.6 | 1.45x | 93.7 | 67.6 | 1.39x | |
| q15 | 5.3 | 2.3 | 2.34x | 8.6 | 3.6 | 2.35x | |
| q16 | 26.2 | 17.9 | 1.46x | 35.1 | 26.6 | 1.32x | |
| q17 | 10.0 | 5.5 | 1.81x | 12.7 | 7.3 | 1.73x | |
| q18 | 5.2 | 4.6 | 1.14x | 9.2 | 12.4 | 0.74x | |
| q19 | 3.3 | 2.0 | 1.67x | 4.7 | 3.2 | 1.45x | |
| q20 | 1.7 | 1.5 | 1.09x | 2.4 | 2.2 | 1.06x | |
| q21 | 1.3 | 1.0 | 1.24x | 1.8 | 1.9 | 0.93x | |
| q22 | 6.1 | 4.0 | 1.55x | 8.4 | 7.9 | 1.07x | |
| q23a | 127.9 | 76.3 | 1.68x | 201.2 | 106.9 | 1.88x | |
| q23b | 159.7 | 80.6 | 1.98x | 281.7 | 117.7 | 2.39x | |
| q24a | 79.9 | 76.6 | 1.04x | 104.5 | 102.6 | 1.02x | |
| q24b | 82.7 | 77.0 | 1.07x | 100.6 | 105.6 | 0.95x | |
| q25 | 7.5 | 4.7 | 1.58x | 9.7 | 9.7 | 1.00x | |
| q26 | 3.5 | 2.7 | 1.33x | 4.0 | 3.6 | 1.13x | |
| q27 | 4.9 | 4.6 | 1.05x | 6.5 | 5.8 | 1.12x | |
| q28 | 100.3 | 98.5 | 1.02x | 110.6 | 115.6 | 0.96x | |
| q29 | 16.4 | 7.4 | 2.22x | 34.3 | 9.8 | 3.49x | |
| q30 | 14.0 | 11.5 | 1.22x | 18.3 | 14.7 | 1.25x | |
| q31 | 11.6 | 9.2 | 1.26x | 14.4 | 11.2 | 1.29x | |
| q32 | 1.1 | 0.8 | 1.44x | 1.5 | 1.3 | 1.16x | |
| q33 | 2.9 | 2.0 | 1.51x | 4.1 | 2.7 | 1.53x | |
| q34 | 4.7 | 3.4 | 1.38x | 6.6 | 4.6 | 1.44x | |
| q35 | 12.6 | 8.2 | 1.55x | 20.7 | 11.6 | 1.78x | |
| q36 | 4.7 | 5.3 | 0.89x | 5.4 | 6.7 | 0.81x | |
| q37 | 7.7 | 6.4 | 1.21x | 8.5 | 8.0 | 1.06x | |
| q38 | 22.6 | 13.6 | 1.66x | 30.7 | 22.3 | 1.38x | |
| q39a | 4.2 | 3.3 | 1.26x | 5.8 | 4.7 | 1.25x | |
| q39b | 3.8 | 3.2 | 1.20x | 5.3 | 4.3 | 1.24x | |
| q40 | 9.7 | 7.3 | 1.34x | 10.8 | 10.6 | 1.03x | |
| q41 | 0.5 | 0.5 | 1.12x | 0.9 | 0.5 | 1.59x | |
| q42 | 1.5 | 1.2 | 1.27x | 1.6 | 1.3 | 1.22x | |
| q43 | 4.1 | 3.9 | 1.06x | 5.2 | 5.2 | 0.99x | |
| q44 | 32.8 | 30.1 | 1.09x | 34.6 | 33.9 | 1.02x | |
| q45 | 4.7 | 3.7 | 1.25x | 7.6 | 4.0 | 1.92x | |
| q46 | 5.4 | 4.9 | 1.08x | 7.7 | 7.7 | 1.00x | |
| q47 | 9.0 | 7.7 | 1.16x | 12.0 | 12.1 | 0.99x | |
| q48 | 6.2 | 5.2 | 1.21x | 7.9 | 6.2 | 1.27x | |
| q49 | 48.9 | 43.1 | 1.14x | 45.1 | 43.8 | 1.03x | |
| q50 | 46.8 | 30.9 | 1.51x | 66.8 | 38.4 | 1.74x | |
| q51 | 17.4 | 10.4 | 1.66x | 24.6 | 15.2 | 1.62x | |
| q52 | 1.2 | 0.9 | 1.24x | 1.5 | 1.1 | 1.38x | |
| q53 | 3.8 | 3.4 | 1.09x | 4.5 | 4.3 | 1.04x | |
| q54 | 4.5 | 3.3 | 1.35x | 8.0 | 4.0 | 2.02x | |
| q55 | 2.0 | 1.3 | 1.53x | 2.1 | 1.6 | 1.27x | |
| q56 | 2.6 | 1.7 | 1.57x | 3.9 | 2.2 | 1.79x | |
| q57 | 4.9 | 4.3 | 1.15x | 7.4 | 6.5 | 1.13x | |
| q58 | 2.2 | 2.4 | 0.90x | 3.7 | 2.4 | 1.51x | |
| q59 | 28.7 | 25.6 | 1.12x | 32.1 | 30.0 | 1.07x | |
| q60 | 3.1 | 2.0 | 1.58x | 4.1 | 2.7 | 1.50x | |
| q61 | 3.3 | 2.2 | 1.49x | 5.2 | 3.3 | 1.57x | |
| q62 | 26.5 | 24.7 | 1.07x | 24.4 | 24.9 | 0.98x | |
| q63 | 4.3 | 4.0 | 1.07x | 4.5 | 4.6 | 0.99x | |
| q64 | 72.0 | 41.1 | 1.75x | 92.7 | 48.6 | 1.91x | |
| q65 | 20.2 | 14.2 | 1.43x | 28.6 | 22.9 | 1.25x | ties |
| q66 | 9.8 | 8.6 | 1.14x | 10.0 | 9.8 | 1.02x | |
| q67 | 67.5 | 38.2 | 1.77x | 126.4 | 66.8 | 1.89x | |
| q68 | 5.2 | 2.5 | 2.06x | 6.9 | 3.6 | 1.93x | |
| q69 | 6.1 | 3.1 | 1.99x | 6.8 | 3.8 | 1.82x | |
| q70 | 8.4 | 6.9 | 1.21x | 11.1 | 8.3 | 1.34x | |
| q71 | 3.4 | 2.3 | 1.45x | 3.2 | 2.6 | 1.26x | |
| q72 | 20.9 | 17.8 | 1.17x | 37.3 | 31.1 | 1.20x | |
| q73 | 4.2 | 1.8 | 2.36x | 5.7 | 3.1 | 1.82x | |
| q74 | 29.6 | 24.2 | 1.22x | 45.3 | 32.2 | 1.41x | |
| q75 | 66.8 | 60.8 | 1.10x | 75.1 | 67.9 | 1.11x | |
| q76 | 44.9 | 41.4 | 1.08x | 42.4 | 42.2 | 1.00x | |
| q77 | 2.1 | 2.0 | 1.05x | 2.4 | 2.0 | 1.22x | |
| q78 | 85.3 | 64.4 | 1.32x | 117.8 | 82.9 | 1.42x | |
| q79 | 4.0 | 3.6 | 1.13x | 6.0 | 4.9 | 1.22x | |
| q80 | 44.1 | 38.6 | 1.15x | 50.0 | 38.8 | 1.29x | |
| q81 | 14.4 | 7.6 | 1.89x | 20.8 | 10.1 | 2.06x | |
| q82 | 17.2 | 14.9 | 1.15x | 19.1 | 16.4 | 1.16x | |
| q83 | 1.2 | 0.9 | 1.32x | 1.7 | 1.4 | 1.22x | |
| q84 | 19.0 | 17.8 | 1.07x | 20.6 | 17.4 | 1.18x | |
| q85 | 20.0 | 19.5 | 1.03x | 22.5 | 19.5 | 1.15x | |
| q86 | 5.0 | 4.8 | 1.05x | 5.8 | 4.9 | 1.19x | |
| q87 | 24.5 | 15.8 | 1.55x | 32.1 | 17.7 | 1.81x | |
| q88 | 101.9 | 104.6 | 0.97x | 119.1 | 111.2 | 1.07x | |
| q89 | 4.2 | 4.4 | 0.94x | 5.5 | 5.1 | 1.08x | |
| q90 | 35.8 | 37.4 | 0.96x | 42.0 | 34.9 | 1.20x | |
| q91 | 2.6 | 1.4 | 1.77x | 3.8 | 2.4 | 1.59x | |
| q92 | 1.5 | 3.0 | 0.49x | 2.0 | 1.5 | 1.36x | |
| q93 | 98.9 | 52.8 | 1.87x | 139.3 | 57.0 | 2.44x | |
| q94 | 55.4 | 52.9 | 1.05x | 66.3 | 50.7 | 1.31x | |
| q95 | 64.8 | 46.8 | 1.39x | 115.3 | 62.4 | 1.85x | |
| q96 | 15.0 | 14.1 | 1.06x | 19.8 | 15.1 | 1.31x | |
| q97 | 21.1 | 9.1 | 2.31x | 35.7 | 12.7 | 2.81x | |
| q98 | 1.9 | 1.5 | 1.23x | 2.9 | 2.4 | 1.19x | |
| q99 | 7.6 | 9.5 | 0.80x | 10.1 | 12.3 | 0.82x |
Running the benchmark
The cluster runner, the Spark-on-Kubernetes manifests, the image and the data generation are in
the benchmark runner's README; on Graviton the image builds for arm64 on an arm64 node and the runs select the arm64 node group.
run-matrix.sh writes one JSON-lines result file per run; this page is rendered from two of them, with the x86 page as the reference,
by benchmarks/scripts/render-graviton-page.py.
TPC-DS is a benchmark of the Transaction Processing Performance Council; these results are not audited TPC results and are not comparable to published TPC-DS results. Times are one measured iteration per query on the cluster described above; run-to-run variation on the heavy queries is a few percent.