VecRuntime

A vectorized execution runtime for Apache Spark using Java

A Spark SQL plugin that runs Filter, Project, HashAggregate, Sort, Window (ROWS and RANGE frames), Range, Expand, Generate, the limits, Union and the hash, sort-merge and nested-loop joins on Arrow-layout batches with the Java Vector API -- on the JVM, no native code -- with its own columnar shuffle over Arrow Flight. Source, releases and the README: github.com/vecruntime/vecruntime.

VecRuntime accelerates Spark SQL workloads by executing core operators directly on Arrow-layout columnar batches using the Java Vector API (jdk.incubator.vector), bringing SIMD-optimized execution to the JVM without native libraries, JNI, or serialization boundaries. Unsupported operators, expressions, and types transparently fall back to Spark, always with a recorded reason.

Key characteristics

  • JVM-native: no native runtime, JNI, or external execution engine.
  • SIMD-accelerated: uses the Java Vector API for hardware-vectorized execution.
  • Columnar by design: operators consume and produce Arrow-layout batches.
  • Spark-compatible: unsupported operators, expressions, and types fall back to Spark.
  • Vectorized operators: Filter, Project, HashAggregate, Sort, Window (including RANGE frames with value offsets), Range, Expand, Generate, the limits, Union, and the hash, sort-merge and nested-loop joins -- the full list is on the Operators page.
  • Zero-copy integration: interoperates with columnar native components such as Apache DataFusion Comet without serialization between execution stages.
  • Incremental adoption: operators can be accelerated individually while the rest of the Spark plan continues to execute normally.

Getting started

VecRuntime needs JDK 25 (the Java Vector API) and Spark 4.1. On JDK 25, Spark 4.1.3's bundled Hadoop 3.4.2 fails at start-up (Subject.getSubject, HADOOP-19212): replace hadoop-client-api and hadoop-client-runtime in $SPARK_HOME/jars with their 3.4.3 versions (drop-in jars).

Let Spark download the plugin: --packages takes the Maven coordinates and --repositories points at the Maven repository served from this project's maven-repo branch. No code changes:

spark-submit \
  --repositories https://raw.githubusercontent.com/vecruntime/vecruntime/maven-repo/ \
  --packages io.github.vecruntime:vecruntime-spark_2.13:0.0.4 \
  --conf spark.plugins=io.vecruntime.spark.VectorPlugin \
  --conf spark.driver.extraJavaOptions="--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED --sun-misc-unsafe-memory-access=allow" \
  --conf spark.executor.extraJavaOptions="--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED --sun-misc-unsafe-memory-access=allow" \
  ...

With the columnar shuffle (Arrow IPC map outputs, fetched over Arrow Flight) add its artifact and its shuffle manager. Spark's own shuffle keeps serving every exchange the plugin does not produce:

spark-submit \
  --repositories https://raw.githubusercontent.com/vecruntime/vecruntime/maven-repo/ \
  --packages io.github.vecruntime:vecruntime-spark_2.13:0.0.4,io.github.vecruntime:vecruntime-shuffle_2.13:0.0.4 \
  --conf spark.plugins=io.vecruntime.spark.VectorPlugin \
  --conf spark.shuffle.manager=org.apache.spark.sql.vecruntime.shuffle.VectorShuffleManager \
  --conf spark.driver.extraJavaOptions="--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED --sun-misc-unsafe-memory-access=allow" \
  --conf spark.executor.extraJavaOptions="--add-modules=jdk.incubator.vector --enable-native-access=ALL-UNNAMED --sun-misc-unsafe-memory-access=allow" \
  ...

The same flags work on spark-shell, pyspark and spark-sql.

  • --sun-misc-unsafe-memory-access=allow is required on JDK 25: without it Arrow's Netty allocator cannot address direct memory and the first columnar operator fails. --add-modules enables the Vector API; --enable-native-access is for the zstd and Netty native code.
  • The plugin's Spark, Arrow and Scala dependencies are provided, so --packages downloads the plugin jar alone. The shuffle also pulls Arrow Flight, gRPC and protobuf-java (about 40 jars, none of which Spark bundles apart from versions it already ships).
  • The Flight server has no TLS: where RPC encryption is on, add --conf spark.vecruntime.shuffle.backend=block (Spark's block transfer over the same files). See the Configuration reference.

Without network access, download the jars from the releases page (plugin jar, shuffle jar, SHA256SUMS) and pass them with --jars vecruntime-spark_2.13-0.0.4.jar (the shuffle's Flight and gRPC jars must then be on the classpath too; --packages resolves them for you).

spark.plugins registers the session extension automatically; alternatively set spark.sql.extensions=io.vecruntime.spark.VectorSparkSessionExtensions. See the README for the Maven coordinates, and the Configuration reference for every spark.vecruntime.* key.

Status

Version 0.0.4, a preview release under the Apache License 2.0. The plugin runs the whole of TPC-DS (103 queries) and TPC-H (22) with every operator accelerated and returns Spark's results. On the 1 TB TPC-DS Parquet dataset on EKS, VecRuntime finished the 103 queries in 2,557 s against Spark's 3,309 s (23% less runtime, faster on 82 of 103 queries), within 2% of Apache DataFusion Comet. The per-query tables and configurations are on the TPC-DS 1 TB page.

Benchmarks

Reference