mattone
dopo mattoneTHE BOOK SERIES
IT/EN
← Java guide

Java 25 · 11/39

11. Strings and Unicode text

Java 25 · Complete guide · Draft under review

This guide retains the book’s draft status. Editorial review and comprehension checks with independent readers remain to be completed.

Search the whole guide →

Text is a value, not a container to modify

When we write String title = "Java", we obtain a reference to a String object. A string is immutable: no method changes its contents after creation. An operation such as strip() returns a string with leading and trailing whitespace removed; the original variable continues to refer to the previous text. This is useful when several program parts share a string, but requires assigning the result if we want to keep it.

In the complete program we start from " Hello ". Printing the result of strip() produces [Hello], while a later print of the original produces [ Hello ]. The square brackets only make spaces visible. They are not part of the cleaned text.

The most frequent comparison concerns content. first.equals(second) is true if the two strings have the same characters in the same sequence; first == second instead asks whether they are the very same object. The program deliberately creates a second instance with new String("Java"), so the results are false and true. Ordinary code does not need to create a string copy this way: here we do it to show the distinction.

Important note – The string pool. Java can share the object associated with equal literals and provides the intern() method. This explains some results of == between literals, but does not turn == into a reliable content comparison. For text value we use equals; when a variable can be null, we explicitly decide how to treat it, for example with Objects.equals(a, b).

Figure 11.1 compares literals with objects created using new. The arrows describe references' logical identity; they do not indicate the memory area in which the JVM must store the pool.

Shared literals and a distinct string
Figure 11.1 – Two equal literals can indicate the canonical instance; new String(a) creates a distinct instance with the same contents. equals compares text, while == compares reference identity.

From literals to objects: four comparisons

A string literal is text enclosed in quotation marks, for example "Java"; "" is also a valid literal and represents an empty string. We can call a method directly on a literal: "Java".length() is 4. The form String title = "Java"; is the ordinary way to obtain that value. Writing new String("Java") deliberately creates another instance and usually brings no advantage to application code. We use it here only to make the content/identity distinction visible.

Compare four cases. String a = "Hello world"; String b = "Hello world"; uses two equal literals referring to the same canonical instance. String c = "Hello" + " world"; concatenates constant expressions: the result is a compile-time constant and can share the complete literal's instance. String d = new String(a); is instead a new instance, although its contents are the same. Thus a == b and a == c are true, a == d is false, while a.equals(d) is true. The rule to take into everyday code remains equals for content: text calculated at runtime must not be treated as an interned literal.

The variable and object have different histories. After String message = "first"; message = "second";, the variable indicates another value; the string "first" was not modified. If we declare final String message = "first";, we prevent the second assignment, but do not make String “more immutable” than it already is. This is the same principle encountered with arrays, viewed from the other side: final regulates variable reassignment; mutability depends on the object's type.

Following a transformation to its result

Try reading three lines without executing them: String original = " Java "; String clean = original.strip(); String uppercase = clean.toUpperCase();. After the first line, original indicates text with a space before and after. After strip(), clean indicates the result without those spaces, but original keeps the initial value. The third line produces "JAVA" and changes neither original nor clean. If we print the three values between square brackets, we see [ Java ], [Java] and [JAVA]. This small path is the best way to avoid calling a method returning a new string and ignoring its result.

Comparison with C character arrays helps us focus on a difference. The useful distinction for a Java reader is that a String is not a char[] array to which we can assign a new letter at a position. We can obtain a character or part of the text through class methods, but operations produce new results when they change content. The internal representation chosen by the JDK is not String's public contract; we must not infer a particular physical memory format from the word “immutable”.

Length in Java and the characters we see

A String uses indices of UTF-16 code units. length() counts these units, not always the characters a person sees on the screen. The smiling emoji U+1F600 is represented by two UTF-16 units and a single Unicode code point: this is why the program prints 2 and then 1. Figure 11.2 visualizes the two units as a pair representing a single code point.

UTF-16 units and code points
Figure 11.2 – The example's symbol occupies two UTF-16 positions. length() returns 2; codePointCount(0, length()) returns 1.
String symbol = "\uD83D\uDE00";
System.out.println(symbol.length());
System.out.println(symbol.codePointCount(0, symbol.length()));

The sequences \uD83D and \uDE00 write the two UTF-16 units in hexadecimal form in the source: D83D is the high surrogate, DE00 the low one. The character u introduces the Unicode escape; every following digit represents four bits. The combined value corresponds to code point U+1F600. Writing escapes in the listing also avoids a font without the emoji glyph hiding the example on the printed page.

A code point is a value in the Unicode set. It does not always coincide with what we perceive as a character: some visible symbols are sequences of several code points, for example a letter with a combining mark or certain compound emoji. An interface that must count user-perceived characters needs further rules; codePointCount solves only the passage from UTF-16 units to code points. The String specification defines both methods.

charAt(i) returns the individual char unit at index i; on a supplementary symbol, the two calls separately return the pair's parts. If we want to traverse code points, we use APIs such as codePoints() or codePointAt, paying attention to indices. An index outside the range causes an exception, not an empty character.

Further reading – Three ways to count text. For "A", length(), codePointCount(0, length()) and the number of symbols perceived by the reader coincide: they are one. For the example's U+1F600 symbol, UTF-16 units are two and code points one. For a letter followed by a combining accent, code points can be two even if the reader sees one accented letter. The question “how long is the string?” must therefore indicate what we want to count. substring and charAt indices remain expressed in UTF-16 units; merely replacing every length() with codePointCount() without reconsidering indices is not enough.

Searching and transforming without changing the original

Before transforming a string, we often want to find a piece of text. title.contains("Java") answers with a boolean; title.indexOf("Java") returns the first occurrence's UTF-16 index or -1 if it finds none. The number -1 is not a valid index, but an absence signal. substring(start, end) produces a portion whose start is included and end excluded: "Java".substring(0, 2) is "Ja". The second bound is not the portion's length, but the position at which to stop. An index cutting between the two char units of a surrogate pair can produce a string difficult to interpret as complete Unicode text; this is why we do not equate UTF-16 indices with “letters”.

To remove surrounding whitespace we can use strip(), which considers whitespace characters according to the API's Unicode rules. The older trim() has a different criterion, based on character values up to U+0020. If text comes from a form or file, choose the rule matching expected data, and test concrete input. toUpperCase() and toLowerCase() return new strings; the result can depend on the locale if we use forms without a parameter. For technical keys that must behave identically regardless of user language, explicitly choosing a locale is more reliable. Furthermore, conversion does not always mean replacing a single letter with another: some transformations can change text length.

To remove simple spaces we might use replaceAll. replaceAll interprets the first argument as a regular expression, a pattern for finding text sequences. If we want to replace a literal space character, replace(" ", "") communicates the intention better and introduces no new grammar. replaceFirst instead uses a regular expression and changes only the first match. Java 25's String class has no replaceLast method: for the last match another strategy is needed, for example finding the position with lastIndexOf and building the result. This is a case where a table of plausible names would be harmful: the reader would try code that does not compile.

Definition – Regular expression. It is a language for describing sets of textual sequences. The pattern [A-Za-z0-9]+ indicates one or more unaccented Latin letters or digits; it does not indicate “any word in any language”. When you use replaceAll, special pattern characters have a meaning that replace does not give them. Before adopting a regex, verify examples that should match and examples that should not.

A guided search, from the result to method choice

Take the text "Java, Brick by Brick" and search for "Brick". indexOf returns the first occurrence's position, measured in UTF-16 units from zero; lastIndexOf finds the last. Here the first begins at 6 and the last at 15. contains answers only “does at least one occurrence exist?”, without giving its position. If we need the position to extract a portion, first check that it is not -1: directly using substring(-1) does not mean “no result”, but causes an exception. Method choice thus arises from the question we must solve, not the similarity of names.

To extract the first word, int separator = text.indexOf(' '); locates the first space and text.substring(0, separator) returns "Java,". This works for the chosen data; if the string contains no spaces, separator is -1 and we must decide whether to return the whole text, nothing or an error. If the text contains a symbol outside the BMP before the space, the obtained index remains correct for substring because both use UTF-16 units. Mixing that index with a code-point count without conversion would produce a wrong bound.

Distinguish replace, replaceFirst and replaceAll on the text "one one one". replace("one", "two") changes all literal occurrences and produces "two two two". replaceFirst("one", "two") changes the first and produces "two one one"; its first argument, however, is a regex. replaceAll("one", "two") again produces "two two two", but also interprets a regex. With the word one the literal/pattern distinction is invisible; with a dot it appears immediately: replace(".", "!") changes only actual dots, while replaceAll(".", "!") uses . as a pattern matching many characters. If we really want a regex for a literal dot, we must escape the pattern. Until regex power is needed, replace is the choice a reader can understand without another grammar.

To verify – Output before code. Write the expected result of " Java ".strip(), "Java".substring(1, 3) and "one one".replaceFirst("one", "two") on paper. Then execute the three expressions and check respectively whitespace, excluded end bound and replacement count. If a prediction is wrong, correct the rule you used, not just the answer.

Building text in steps

The operator + joins strings and is convenient for short expressions. But if an algorithm adds many fragments in a loop, StringBuilder makes explicit a mutable container that then produces a final String. In our example the two append calls form Java.

StringBuilder builder = new StringBuilder();
builder.append("Ja").append("va");
System.out.println(builder.toString());

There is no need to turn every small concatenation into a builder: choose the tool according to what the code is doing, and measure if performance matters. StringBuilder is not designed for simultaneous modification by several threads without coordination, a subject we will encounter in the concurrency chapter.

The operator + can concatenate text and values of other types. In "Grades: " + 3, the number enters the textual representation and the result is "Grades: 3". But expression order matters: "Total: " + 2 + 3 produces "Total: 23", while "Total: " + (2 + 3) produces "Total: 5". Parentheses declare that we first want numeric addition. concat(String) is another method producing a new string; it does not modify the original, as in the example. For repeated construction in a loop, StringBuilder avoids using a variable as if it were a mutable string, without forcing us to claim every concatenation with + is slow.

Java also offers text blocks, delimited by three quotation marks, for strings spanning several source lines. They remain String objects: they change how we write the literal, not the value model. Formatting depends on indentation and line-break rules, so must be verified with actual output when text must be exact. The string templates that appeared as preview features in earlier versions were withdrawn: we do not present them as syntax available in Java 25. The language-change summary documents feature status.

The text-block program contains two lines and prints Hello, and Java! on separate lines. The line containing the closing delimiter also determines whether a final line break remains. Common spaces before the two source lines are treated as incidental indentation, not printed before Hello and Java!.

An HTML example clearly shows the readability gain: without text blocks we needed several concatenations, \n and escaped quotes; with three quotes the text looks more like the document we want to produce. The opening delimiter must be followed by a line terminator before the content. The closing delimiter decides where text ends and affects common spaces removed. A text block is still a String: if HTML requires variable data, we must build it with normal APIs, also accounting for the target format's escaping rules. Text blocks have been stable since Java 15.

String message = """
        Hello,
        Java!
        """;
System.out.print(message);

To verify

Compile the complete source with javac --release 25 -Xlint:all -d build DemoStrings.java and run java -cp build DemoStrings. Try replacing symbol with a simple letter and then with a sequence of several code points: note length() and codePointCount each time. Activity C11 helps interpret the numbers without generically calling them “character count”.

Try the examples

Running the programs requires JDK 25. Download the individual Java files linked in the chapter or the complete example project, which includes instructions and a launcher script. The explanations also compare expected output: predict it before running the program.

Massimiliano Tarquini · CC BY-NC 4.0

Back to the top ↑