Apache Spark vs VecRuntime on TPC-DS 1 TB, AWS Graviton4

VecRuntime runs Spark SQL's filters, projections, aggregates, sorts and joins on Arrow-layout batches with the Java Vector API -- on the JVM, no native code -- and moves batches between executors over its own Arrow Flight shuffle. This page compares it against plain Apache Spark on the TPC-DS 1 TB workload on Amazon EKS with AWS Graviton4 (arm64) nodes, and sets both against the published x86 run from the same day, with the same settings and the same single-AZ layout (its own cluster session) (#253).

TL;DR. On Graviton4, over the 103 TPC-DS queries at 1 TB, VecRuntime finished in 1,893 s against Spark's 2,470 s -- 1.31x, 23% less runtime (geometric mean 1.31x), faster than Spark on 96 of 103 queries (best q6: 3.06x; largest regression q92: 103%). On the x86 nodes the same comparison is 1.37x (geometric mean 1.31x): the lead carries over to arm64. Both engines run faster on Graviton4 than on the m5.4xlarge nodes -- Spark's time is 0.76 of its x86 time and VecRuntime's 0.76, per query (geometric mean). Row counts equal Spark's on 103 of 103 queries; checksums on every query except q65 (ties in the result, ordered differently by each engine).
Apache Spark 4.1.3 on Graviton4
2,470 s
baseline, 103 queries (x86: 3,313 s)
VecRuntime on Graviton4
1,893 s
1.31x · 23% less runtime (x86: 2,420 s)
Graviton4 against x86
0.76 / 0.76
time per query, Spark / VecRuntime (geometric mean)

Summary

EngineCompletion time (s)SpeedupFaster than Spark onExecutor time (h)GC (h)Shuffle read (TB)
Apache Spark 4.1.32,470.2baseline--60.50.280.94
VecRuntime1,892.71.31x (23% less)96 / 10346.50.170.47

Graviton4 against x86

The x86 numbers are the published run: 9 x m5.4xlarge (16 vCPU, 64 GB, AVX-512), the same data, executors, Spark settings and single-AZ layout, measured the same day in its own cluster session.

Spark (s)VecRuntime (s)VecRuntime speedupgeometric mean
Graviton4 (m8g.4xlarge)2,470.21,892.71.31x1.31x
x86 (m5.4xlarge)3,313.32,419.61.37x1.31x
Graviton4 / x86, per query (geomean)0.760.76

Where VecRuntime's lead over Spark shrinks most on Graviton4 -- mostly queries where Spark itself gains more from Graviton4:

QueryGraviton4x86Spark G/xVecRuntime G/x
q920.49x1.36x0.732.01
q580.90x1.51x0.590.99
q292.22x3.49x0.480.75
q451.25x1.92x0.620.94
q541.35x2.02x0.560.83
q411.12x1.59x0.610.87
q951.39x1.85x0.560.75
q931.87x2.44x0.710.93

Where it grows most -- mostly queries at parity or behind on x86:

QueryGraviton4x86Spark G/xVecRuntime G/x
q121.35x0.76x0.920.52
q251.58x1.00x0.760.49
q181.14x0.74x0.570.37
q221.55x1.07x0.730.50
q111.19x0.88x0.770.57
q21.43x1.06x0.980.72
q211.24x0.93x0.710.53
q52.28x1.72x0.880.66

"G/x" is the query's time on Graviton4 over its time on x86 (below 1 = faster on Graviton4).

Benchmark infrastructure

Methodology. The two engines ran one after another on the same nodes, each alone on the cluster, over the same S3 data with the same Spark settings; only the execution engine and its own memory split differ (every engine has 50 GB per executor). Each query ran once after the plan was compiled; the time is the wall-clock of the query's execution as the runner measures it.

Test environment

ComponentConfiguration
DatasetTPC-DS scale factor 1000 (1 TB), Parquet on Amazon S3 (103 query variants, one measured iteration each, no warm-up)
ClusterAmazon EKS 1.36; 9 x m8g.4xlarge (16 vCPU AWS Graviton4 / Neoverse V2 with SVE2, 64 GB, arm64), 300 GB root volume; one node group in one availability zone (us-east-1b), the driver on the ninth node
Executors8 executors x 13 cores x 50 GB each (Spark: 20 GB heap / 30 GB overhead; VecRuntime: 30 GB heap / 20 GB overhead); driver 2 cores x 4 GB
StorageAmazon S3 through an S3 gateway VPC endpoint, Hadoop 3.4.3 S3A with the Analytics Accelerator input stream (the default in 3.4.3)

Versions

ComponentVersion
Spark4.1.3
Scala2.13
JDKAmazon Corretto 25.0.4.1 (aarch64)
VecRuntimemain at cb755d1 (#554 in-place dictionary decode; arm64 image main-cb755d1-arm64)
Hadoop3.4.3

Configuration

Common to both engines:

spark.sql.shuffle.partitions=300
# spark.sql.adaptive.advisoryPartitionSizeInBytes left at Spark's default (64 MB); no coalescePartitions.minPartitionNum
spark.eventLog.enabled=true

VecRuntime:

spark.plugins=io.vecruntime.spark.VectorPlugin
spark.shuffle.manager=org.apache.spark.sql.vecruntime.shuffle.VectorShuffleManager
spark.vecruntime.exec.strictFloatingPoint=false   # Comet's default too
spark.vecruntime.shuffle.aqe.mapSizeScaling=true   # AQE sees Spark-scale map output sizes (#514)
spark.vecruntime.shuffle.aqe.sparkCompressionRatio=0   # the uncompressed-bytes ratio (#514)
AOT class-data cache off
--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED (driver and executors)

The arm64 platform

  • HotSpot on Graviton4: UseSVE=2, MaxVectorSize=16 -- the Vector API's species are 128 bits wide, as on NEON.
  • VecRuntime's platform probe picks its SVE paths: native compress (SVE COMPACT) for selection, and the broadcast-AND-compare lane masks, because VectorMask.fromLong is not intrinsified at 128-bit SVE on JDK 25 (#253, #484).
  • The native codecs (snappy, zstd, lz4) and Netty's epoll transport load their aarch64 libraries; neither architecture has libhadoop, as on x86.

Performance results

Per query

Seconds per query on Graviton4, both engines side by side (hover for values; click a legend entry to hide an engine).

Performance distribution

VecRuntime vs SparkQueriesShare
20%+ improvement5150%
10-20% improvement2221%
within ±10%2625%
10-20% degradation22%
20%+ degradation22%

Top 10 improvements

QuerySpark (s)VecRuntime (s)Speedupx86 speedup
q65.51.83.06x (+67%)2.60x
q734.21.82.36x (+58%)1.82x
q155.32.32.34x (+57%)2.35x
q9721.19.12.31x (+57%)2.81x
q541.518.22.28x (+56%)1.72x
q2916.47.42.22x (+55%)3.49x
q685.22.52.06x (+52%)1.93x
q696.13.11.99x (+50%)1.82x
q23b159.780.61.98x (+50%)2.39x
q8114.47.61.89x (+47%)2.06x

Regressions

QuerySpark (s)VecRuntime (s)Degradationx86 speedup
q921.53.0103% slower (0.49x)1.36x
q997.69.525% slower (0.80x)0.82x
q364.75.312% slower (0.89x)0.81x
q582.22.411% slower (0.90x)1.51x
q894.24.46% slower (0.94x)1.08x
q9035.837.45% slower (0.96x)1.20x
q88101.9104.63% slower (0.97x)1.07x

Notes

  • Both engines returned the same row counts on every query and the same checksums on every query except q65, whose result has ties that each engine orders differently.
  • Both legs ran on 2026-09-30 in one cluster session, Spark first and VecRuntime straight after, on the same nine nodes in one availability zone, each alone on the cluster. S3 throughput varies between sessions, so only runs from the same session are compared.
  • The earlier Graviton4 runs (2026-09-26 at 1.23x, 2026-09-29 at 1.26x) had their nodes split across two availability zones; this run keeps all nine in one and adds #554.
  • VecRuntime reports its map output sizes to AQE on Spark's scale (#511, #514): its columnar shuffle is 1.6-4.3x smaller than Spark's for the same rows, and without the scaling AQE packs up to twice Spark's rows into a task.
  • The x86 reference ran the same day with the same image version, settings and single-AZ layout, in its own cluster session; the architecture comparison below therefore compares two sessions.

All queries

Seconds per query on Graviton4, the faster engine in bold, with the x86 run alongside
QuerySparkVecRuntimeSpeedupx86 Sparkx86 VecRuntimex86 speedupNote
q111.59.71.19x13.412.91.04x
q253.737.61.43x54.952.01.06x
q33.53.41.02x3.64.20.86x
q450.149.71.01x92.386.41.07x
q541.518.22.28x47.027.41.72x
q65.51.83.06x10.13.92.60x
q75.04.61.07x6.67.10.94x
q85.12.91.73x7.03.81.84x
q976.770.21.09x88.187.11.01x
q105.33.91.34x8.35.91.41x
q1132.527.31.19x42.248.20.88x
q122.11.51.35x2.22.90.76x
q137.25.61.27x8.38.01.04x
q14a83.155.11.51x99.874.61.34x
q14b74.951.61.45x93.767.61.39x
q155.32.32.34x8.63.62.35x
q1626.217.91.46x35.126.61.32x
q1710.05.51.81x12.77.31.73x
q185.24.61.14x9.212.40.74x
q193.32.01.67x4.73.21.45x
q201.71.51.09x2.42.21.06x
q211.31.01.24x1.81.90.93x
q226.14.01.55x8.47.91.07x
q23a127.976.31.68x201.2106.91.88x
q23b159.780.61.98x281.7117.72.39x
q24a79.976.61.04x104.5102.61.02x
q24b82.777.01.07x100.6105.60.95x
q257.54.71.58x9.79.71.00x
q263.52.71.33x4.03.61.13x
q274.94.61.05x6.55.81.12x
q28100.398.51.02x110.6115.60.96x
q2916.47.42.22x34.39.83.49x
q3014.011.51.22x18.314.71.25x
q3111.69.21.26x14.411.21.29x
q321.10.81.44x1.51.31.16x
q332.92.01.51x4.12.71.53x
q344.73.41.38x6.64.61.44x
q3512.68.21.55x20.711.61.78x
q364.75.30.89x5.46.70.81x
q377.76.41.21x8.58.01.06x
q3822.613.61.66x30.722.31.38x
q39a4.23.31.26x5.84.71.25x
q39b3.83.21.20x5.34.31.24x
q409.77.31.34x10.810.61.03x
q410.50.51.12x0.90.51.59x
q421.51.21.27x1.61.31.22x
q434.13.91.06x5.25.20.99x
q4432.830.11.09x34.633.91.02x
q454.73.71.25x7.64.01.92x
q465.44.91.08x7.77.71.00x
q479.07.71.16x12.012.10.99x
q486.25.21.21x7.96.21.27x
q4948.943.11.14x45.143.81.03x
q5046.830.91.51x66.838.41.74x
q5117.410.41.66x24.615.21.62x
q521.20.91.24x1.51.11.38x
q533.83.41.09x4.54.31.04x
q544.53.31.35x8.04.02.02x
q552.01.31.53x2.11.61.27x
q562.61.71.57x3.92.21.79x
q574.94.31.15x7.46.51.13x
q582.22.40.90x3.72.41.51x
q5928.725.61.12x32.130.01.07x
q603.12.01.58x4.12.71.50x
q613.32.21.49x5.23.31.57x
q6226.524.71.07x24.424.90.98x
q634.34.01.07x4.54.60.99x
q6472.041.11.75x92.748.61.91x
q6520.214.21.43x28.622.91.25xties
q669.88.61.14x10.09.81.02x
q6767.538.21.77x126.466.81.89x
q685.22.52.06x6.93.61.93x
q696.13.11.99x6.83.81.82x
q708.46.91.21x11.18.31.34x
q713.42.31.45x3.22.61.26x
q7220.917.81.17x37.331.11.20x
q734.21.82.36x5.73.11.82x
q7429.624.21.22x45.332.21.41x
q7566.860.81.10x75.167.91.11x
q7644.941.41.08x42.442.21.00x
q772.12.01.05x2.42.01.22x
q7885.364.41.32x117.882.91.42x
q794.03.61.13x6.04.91.22x
q8044.138.61.15x50.038.81.29x
q8114.47.61.89x20.810.12.06x
q8217.214.91.15x19.116.41.16x
q831.20.91.32x1.71.41.22x
q8419.017.81.07x20.617.41.18x
q8520.019.51.03x22.519.51.15x
q865.04.81.05x5.84.91.19x
q8724.515.81.55x32.117.71.81x
q88101.9104.60.97x119.1111.21.07x
q894.24.40.94x5.55.11.08x
q9035.837.40.96x42.034.91.20x
q912.61.41.77x3.82.41.59x
q921.53.00.49x2.01.51.36x
q9398.952.81.87x139.357.02.44x
q9455.452.91.05x66.350.71.31x
q9564.846.81.39x115.362.41.85x
q9615.014.11.06x19.815.11.31x
q9721.19.12.31x35.712.72.81x
q981.91.51.23x2.92.41.19x
q997.69.50.80x10.112.30.82x

Running the benchmark

The cluster runner, the Spark-on-Kubernetes manifests, the image and the data generation are in the benchmark runner's README; on Graviton the image builds for arm64 on an arm64 node and the runs select the arm64 node group. run-matrix.sh writes one JSON-lines result file per run; this page is rendered from two of them, with the x86 page as the reference, by benchmarks/scripts/render-graviton-page.py.

TPC-DS is a benchmark of the Transaction Processing Performance Council; these results are not audited TPC results and are not comparable to published TPC-DS results. Times are one measured iteration per query on the cluster described above; run-to-run variation on the heavy queries is a few percent.