mattone
dopo mattoneTHE BOOK SERIES
IT/EN
← Java guide

Java 25 · 23/39

23. Further reading: Vector API, processing data in groups

Java 25 · Complete guide · Draft under review

This guide retains the book’s draft status. Editorial review and comprehension checks with independent readers remain to be completed.

Search the whole guide →

The vector API in this chapter addresses computation on several elements at once. The name “vector” and the idea of parallelism may recall threads, but the problem differs from that of chapter 20A: lane-wise addition does not make an update shared between threads atomic. Before choosing this tool, ask whether the work concerns independent numerical data, rather than a concurrent decision on a single state.

This is a self-contained further reading chapter between the chapter on classloaders and the one on reflection. If you want to follow the thread “loading a type → inspecting the type” straight away, you can move to chapter 24 and return here later: vectors are not a prerequisite for reflection.

Java version – Incubator API. The Vector API belongs to the jdk.incubator.vector module in Java 25: it requires --add-modules jdk.incubator.vector and may change or be removed in a future version. It is not an ordinary stable Java SE 25 API. The examples in this further reading chapter follow the module documentation in JDK 25.

Why discuss vectors

In many scientific, graphics or signal-processing problems, the same operation is performed on many independent numbers. To add two float[] arrays, an ordinary loop reads one pair of elements at a time, adds them and writes the result. A CPU may have SIMD instructions, capable of applying the same operation to several data lanes in one instruction. Automatic compiler optimization can sometimes exploit them, but not all loops have an easily recognizable form. The Vector API provides an explicit way to describe operations on groups of values, leaving the JVM to adapt them to the available hardware.

Definition – Vector and lane. A vector in this API is a group of values of the same primitive type, for example several float values. Each position in the group is a lane. Adding two vectors adds corresponding lanes: it does not calculate the sum of all elements of a single vector. This distinction avoids confusing an element-wise operation with a reduction.

In the complete program, FloatVector.SPECIES_PREFERRED chooses a species suitable for the platform for float values. Among other things, the species describes how many lanes the vector contains. The number printed by SPECIES.length() may vary between machines and JVMs; the example therefore does not promise a fixed width. The numerical result, on the other hand, is predictable for the small values used: [11.0, 22.0, 33.0, 44.0, 55.0].

FloatVector a = FloatVector.fromArray(SPECIES, left, i);
FloatVector b = FloatVector.fromArray(SPECIES, right, i);
a.add(b).intoArray(result, i);

The first two lines load groups of numbers from the arrays, starting at index i. The third adds lane by lane and writes the group into the result. The loop advances by SPECIES.length() elements. If the array length is not a multiple of the lanes, a tail remains: the program processes it with an ordinary scalar loop. SPECIES.loopBound(left.length) computes the boundary up to which complete groups can be read without going beyond the array. This is why the example works even when the machine uses more lanes than the five available elements: in that case the vector loop may execute no iterations, and the scalar tail does all the work.

Lane-wise addition and scalar tail
Figure 23.1 – Illustrated example with four lanes. The actual width of SPECIES_PREFERRED depends on the platform.

Before starting, the method checks that the two arrays have the same length. Without that check, an out-of-bounds access could cause an error midway through processing and leave a partial result. Validation clearly establishes the contract: element-wise addition requires two comparable sequences. There is no need to introduce implicit behavior for arrays of different lengths.

Compiling and interpreting the test

Compile the file with javac --release 25 --add-modules jdk.incubator.vector VectorSum.java and run it with java --add-modules jdk.incubator.vector VectorSum. The JDK issues a warning about use of the incubator module: this is expected and confirms the API's experimental nature. If you omit the option, the module may be unavailable during compilation or execution. The examples in the stable chapters continue to compile without this option.

The presence of a vector loop does not by itself prove a performance improvement. Measurement depends on data size, hardware, JVM configuration, JIT compiler warmup and the cost of moving data. For five numbers, the program demonstrates semantics, rather than serving as a benchmark. Even a comparison with a scalar loop requires repeated trials and a suitable measurement method; timing a single execution from main would confuse JVM startup, compilation and actual work.

A useful exercise is to change the array length so that zero, one or many elements remain after the last complete vector. Readers can verify that the result never loses the tail. Only after understanding this boundary does it make sense to explore Vector API masks, which allow operations on selected lanes to be expressed. They still belong to the same incubator API, however: every additional example must be verified on the JDK version used by the volume.

When the length does not match the lanes

The first program uses two loops: one for complete groups, one for the remainder. This is a readable solution and useful for getting started, because the scalar loop makes it clear that no element must be lost. The Vector API also provides masks. A mask has a Boolean condition for each lane: it indicates which lanes are valid for a given operation. In the second program, SPECIES.indexInRange(i, left.length) builds the mask for the positions still within the array. The masked forms of fromArray and intoArray read and write only the selected lanes.

VectorMask<Float> valid = SPECIES.indexInRange(i, left.length);
FloatVector a = FloatVector.fromArray(SPECIES, left, i, valid);
FloatVector b = FloatVector.fromArray(SPECIES, right, i, valid);
a.add(b).intoArray(result, i, valid);

To understand these four lines, take the figure with four lanes and five numbers. The first iteration uses indices zero through three: all lanes are valid. The second iteration starts at index four: only the first lane corresponds to the fifth element; the other three would fall beyond the end of the array and the mask excludes them. The written result still contains five values, not eight. If the platform prefers a different width, the same method computes the mask from the actual species length. This is why the program does not hard-code the number four as the loop boundary.

The masked version is not automatically faster than the one with a final scalar loop. The chosen form can affect machine-code generation and the outcome depends on data and hardware. Here the mask teaches an API mechanism: it lets us express “operate only on these lanes” and handle a partial group uniformly. The first program, with its scalar tail, may remain preferable as an initial explanation. Both must be verified with empty arrays, arrays shorter than a species, arrays exactly one species long and arrays one species plus one element long.

Further reading – Mask and Java condition. A VectorMask<Float> is not a single boolean deciding whether to execute the whole loop. It contains a choice for each lane. If the group has four lanes and only the first is still within the array, the mask conceptually represents true, false, false, false. The object's exact form and the number of lanes depend on the species. This distinction avoids thinking of the mask as a simple if around loading the whole vector.

Species, element type and portability

VectorSpecies<Float> describes the combination of element type and vector shape. Float in the generic parameter identifies the type of the lanes in the API, while the array data remains primitive float. The length() method returns the number of lanes; loopBound(length) computes the greatest multiple of that number that does not exceed the supplied length. If a species has four lanes and the array contains eleven elements, the boundary is eight: complete groups start at indices zero and four; three elements remain for the scalar loop. If the species has eight lanes, the boundary is still eight, but there is only one complete group. The formula is simple, but it must be connected to the actual loop index to avoid out-of-bounds accesses.

The preferred species leaves a choice to the platform. This favors source portability, rather than identical performance behavior. The program must not promise that SPECIES.length() equals a specific number. Tests check the operation's result and, if desired, print the width as diagnostic information. A teaching figure may draw four lanes to make the reasoning visible, provided its caption states that this is an illustrated example and not an API constraint.

A vector algorithm must also consider the data type. Adding float values follows the rules of floating-point numbers: very large values, very small values, infinities and NaN do not behave like exact mathematical integers. In our example we use small numbers with an intuitive result. If the program were transformed into a reduction, such as summing all the elements, the order of additions could change, and with it some rounded results. Before promising bit-for-bit equivalence with a scalar loop, we would need to establish the required numerical contract and check the concrete operation.

From the example to a real problem

Element-wise addition is a first brick. A real application might process audio samples, transform image channels or apply a function to large blocks of numerical data. The potential advantage arises when the same simple computation repeats over enough elements to make work on lanes significant. If every element requires a complex decision or irregular access to scattered objects, vector code may lose clarity without offering a benefit. First identify a loop that really contributes to total time, then try a vector form and measure it.

Measurement must compare correct results and comparable conditions. A benchmark that includes array creation in the time for one variant and excludes it for the other does not compare the same work. A single JVM execution includes startup and JIT compiler activity. Array size, warmup, memory, architecture and JDK version must be recorded. If the data is already cached in one trial and not in the other, the difference may depend on memory rather than SIMD instructions. The Vector API supplies a language for expressing computation; proof of an advantage belongs to a measured experiment.

This further reading chapter lets us experiment with SIMD and masks while knowing which dependency we introduce. If the module changes in a later version, examples and commands will need to be verified again before reuse. For now, the point to take away is the method: start with a concrete computation, understand what each lane does, check the array tail as well and measure only after verifying that the result is correct.

To verify

Run the two programs with arrays of length zero, one, five and ten. For each length, indicate how many complete groups the first program processes on the machine used and which mask applies in the second program's final iteration. Add a test with arrays of different lengths and verify that both throw IllegalArgumentException before producing a result. Distinguish the numerical output, which must be equal in the two test cases, from the number of lanes and execution time, which are not universal results of the book. If you transfer the method to audio samples or pixels, establish first which numerical result must remain invariant.

Try the examples

Running the programs requires JDK 25. Download the individual Java files linked in the chapter or the complete example project, which includes instructions and a launcher script. The explanations also compare expected output: predict it before running the program.

Massimiliano Tarquini · CC BY-NC 4.0

Back to the top ↑