An archive that survives execution
So far, many examples have created data in memory and lost it at the end of main. A library, however, must be able to read tomorrow the titles recorded today. This change introduces a boundary: the program controls its objects, but does not completely control the file system. A file may be missing, unreadable or change while we consult it; a write may exhaust available space. The first question is not “which Files method must I remember?”, but “what am I transferring, between which places, and what must remain true if the transfer fails?”.
We call the set of input and output operations I/O. A source provides data, a destination receives it. A file is one possible source or destination; so is a network connection. A resource is an object that keeps something open outside the ordinary memory of Java objects, such as a file descriptor. Closing the resource is not a cosmetic detail: it releases that capacity and concludes the operation according to the API's contract.
Figure 25.1 follows text from the archive into the program. Bytes on disk are interpreted using an encoding; the result is characters and strings. In the reverse direction, characters must be encoded into bytes. If the archive contains an image, that textual conversion would be a mistake: the image must be treated as bytes. This distinction connects chapter 11 on Unicode to chapter 12 on exceptions.
Bytes, characters and encoding
InputStream reads bytes, OutputStream writes them. The value returned by read() is an integer from 0 to 255, or -1 when the stream has ended: -1 is not a byte from the file. A call reading into an array may obtain fewer bytes than requested, even when more data will arrive. Code that needs all the content must repeat reading until end of stream; transferTo expresses a simple copy, but does not make the destination automatically atomic or recoverable after a failure.
Reader and Writer instead work with Java characters, represented in their elementary methods by UTF-16 units. A Reader is not equivalent to a Unicode code-point reader: a symbol outside the basic multilingual plane uses two units. Moving between bytes and characters requires a decoder and an encoder. InputStreamReader and OutputStreamWriter provide the bridge; Files.newBufferedReader and Files.newBufferedWriter are convenient choices when the source is a Path. In the complete program, we specify StandardCharsets.UTF_8 in both directions, so the archive format is declared by the code.
Key concept – The charset is part of the format. UTF-8 establishes how characters are represented in bytes. A Java
Stringis not already “a UTF-8 file”: it is text in memory. If producer and reader choose different encodings, text may be corrupted or decoding may fail. When the format is known, do not rely on the machine's default charset; any byte order mark and the treatment of malformed inputs must be decided by the file's contract.
A buffer temporarily holds data to reduce the number of accesses to the source. Wrapping a stream in BufferedInputStream or a reader in BufferedReader can help when we read many small fragments. It does not, however, solve the problem of an algorithm loading an enormous archive entirely into memory. Files.readString suits a small file whose memory cost we can accept; for a long archive, use a reader and process one line at a time. The analogous output choice is between Files.writeString and progressive writing with BufferedWriter.
A concrete case helps us choose. If we must copy an image without modifying it, Files.copy already expresses the operation between two paths. If the source is a connection and the destination a file, InputStream.transferTo(OutputStream) transfers bytes without introducing a textual encoding. If instead we must count titles starting with a letter, we need to interpret lines and characters, so the text reader is part of the solution. The fact that all three operations “read a file” does not make them interchangeable: the meaning of the data changes.
With manual block reading, the value returned by read(byte[]) indicates how many array elements have been filled; only those must be written. A program that always writes the whole array may add old bytes to the output file on the final iteration. 0 and -1 also have different meanings when a particular API can return zero: end of stream is signaled by -1. To learn this discipline, copy a file whose length is not a multiple of the buffer size and compare the result byte for byte.
Decorating a stream: every layer adds a task
Suppose we read from a source providing a few bytes at a time. The format does not change, but the cost of each access does. We can wrap the source with a buffer and then a text reader: byte source → buffer → decoder → line reader. Decoration preserves the base contract and adds behavior. A BufferedReader does not by itself know the encoding of the bytes before the decoder; a BufferedInputStream does not know where a line ends. Every layer answers a different question.
When we create the entire chain ourselves, close its outer object with try-with-resources: ordinary wrappers propagate closure to the underlying resource. If a method receives a stream from the caller, however, the contract must say who owns it. A copy method should not close System.out without warning, or a stream the caller must reuse. transferTo transfers content, but does not close either end. Resource ownership is a design responsibility, rather than something the compiler deduces from the type.
Important note – End of stream and availability.
available()does not measure the total input length and does not prove end of stream: it estimates bytes readable without blocking.readAllBytes()reads everything and consumes memory proportional to the data;readNBytes(n)may return fewer thannbytes if the input ends. If the format requires exactly a certain amount, check the length obtained.markandresetallow returning to a point only in streams supporting them and within the declared limits; they do not turn any source into a random-access file.
In the I/O contract program we use an in-memory input that deliberately limits every read to three bytes. The loop writes only the returned count and the final comparison checks the entire result. It is a small experiment, but removes the illusion that “I requested eight bytes” means “I received eight”. The variant replaces the in-memory array with a network source: the principle remains valid even if a read can now wait for new data.
For a binary format with primitive numbers, DataInputStream and DataOutputStream add operations such as readInt and writeInt. The contract must fix field order, sizes, version and response to truncation; readInt reports premature end with EOFException. writeUTF uses a modified UTF-8 form and a limited length: it is not a shortcut for a general UTF-8 text file. For compressed data we can add streams from java.util.zip, but decompressed size needs its own limits: a small input can produce much more content. The criterion remains choosing the layer representing the format, rather than adding layers until the example works.
Decoding that rejects ambiguous data
The catalog comes from an external system declaring UTF-8. We want to reject a malformed byte sequence rather than replace it with a character and publish a different title. A CharsetDecoder allows choosing CodingErrorAction.REPORT for malformed input and unmappable characters. We can pass it to InputStreamReader, then wrap that in BufferedReader. Do not rely on supposed uniformity among constructors: a bridge built with only a charset may use a replacement policy, while the contract of Files.newBufferedReader requires errors for malformed or unmappable sequences.
To construct this reader, the following fragment assumes an already available InputStream input. The method performing the reading propagates or handles IOException; closing reader also closes input, according to the ownership contract just established.
var decoder = StandardCharsets.UTF_8.newDecoder()
.onMalformedInput(CodingErrorAction.REPORT)
.onUnmappableCharacter(CodingErrorAction.REPORT);
try (var reader = new BufferedReader(new InputStreamReader(input, decoder))) {
String line;
while ((line = reader.readLine()) != null) {
System.out.println(line);
}
}
Import CodingErrorAction from java.nio.charset and the two readers from java.io. The decoder is created for this stream: do not share it between concurrent reads. Printing is only the demonstration action; in the importer, the line must be validated before producing the result.
The check uses bytes C3 28: the first starts a UTF-8 sequence that the second does not complete. The reader configured to reject must throw MalformedInputException, rather than produce a title. Then test a valid title with an accented character and one outside the basic multilingual plane. Success on ASCII characters alone would not have tested the interesting part of encoding.
Further reading – A sequence can cross two blocks. A UTF-8 character may occupy several bytes. If we decode each block in isolation with
new String(block, UTF_8), the buffer boundary may split the sequence and corrupt it. A reader or decoder maintained throughout the stream retains the necessary state. Block I/O divides transport; it does not define the boundaries of characters or format records.
Even reading “one line at a time” has a limit: readLine() builds the entire line in memory and removes its terminator. A file with no line breaks can therefore be one enormous line. If input is uncontrolled, define maxima for bytes, characters and line length, enforcing them during reading rather than after allocating everything. newLine() uses the platform separator; if the protocol requires LF, explicitly write \n. Our catalog transformation chooses a local archive: for an interchange format, make that choice explicit too.
A file's name is not its content
Path represents a path in the file system; Files performs operations on that path. With Path.of("data", "books.txt"), the program constructs a relative path: its meaning depends on the process's working directory. An absolute path, such as the one returned by toAbsolutePath(), indicates a location relative to the file system root. Neither guarantees that the file exists. resolve joins path parts; relativize describes how to reach one path from another, when the provider allows it and the roots are compatible.
normalize() removes lexical elements such as . and name/.. pairs. It does not read the file system and is not a security check. In the presence of symbolic links, the normalized path and the one actually visited can differ. toRealPath() queries the file system and resolves links according to the specified options, but requires the path to be accessible. If an application accepts filenames from a user, a simple textual prefix check is not enough to guarantee that all future openings remain inside a directory: other processes may change the file system between checking and use.
Take Path.of("archives").resolve("2026").resolve("books.txt"). The API preserves separators appropriate to the provider: there is no need to manually insert / or \\ into the string. On a relative path, toAbsolutePath uses the working directory; it does not turn the path into a fixed application resource. If the user starts the program from another folder, that same relative name may designate another file. The program must therefore receive the working directory as a parameter or obtain it from explicit configuration. A path used in tests can be created with Files.createTempDirectory, as our example does, to avoid dependence on the author's computer.
A call to Files.exists may return false both when the file is missing and when its existence cannot be determined. In any case, if (Files.exists(p)) Files.readString(p) does not eliminate the exception: the file may disappear between the two calls. It is better to attempt the operation, then handle NoSuchFileException, AccessDeniedException or a more general IOException according to the decision the application must make. The exceptions chapter prepared us precisely to distinguish expected failure from a violated invariant.
Opening for writing means choosing what may be lost
Before modifying an archive, we must answer a question: must the file be new, replaced or extended? Opening options make the answer concrete. With the ordinary settings of Files.newBufferedWriter, the file is created if missing and truncated if it exists. This is convenient for a regenerable result; it is dangerous if we thought we were adding a title at the end. APPEND requests writing at the end; CREATE_NEW requires the name not to exist and makes opening fail if it is already occupied. CREATE, on the other hand, also permits an existing file.
Imagine creating a numbered receipt. The check “if it does not exist, create it” leaves an interval in which another process can use the same name. Opening with CREATE_NEW expresses the constraint directly. The test program creates the receipt and tries to create it a second time: we expect FileAlreadyExistsException and verify that the original content remains intact. If the requirement changes to “add lines to the log”, the choice becomes APPEND, but we must still define coordination and recovery for concurrent writers. Do not assume that a sequence of several writes forms an indivisible record on every provider.
flush() pushes data from the buffer toward the next layer. It does not certify that it is already on persistent storage; close() may fail while completing a write. This is why our importer publishes after leaving the resource block. FileChannel.force can request persistence of file updates under the documented conditions, but the durability of a replacement also involves the file system and directory metadata. We must not derive a universal recovery promise from a single method.
An import that does not publish half-finished results
Suppose we transform an archive of titles. We want to remove blank lines and spaces at the edges, then write the titles in uppercase. If we write directly to the final file and line one hundred contains invalid data, readers will see a truncated archive. The worked solution uses a temporary file in the same directory as the destination. Only after reading, validation and closing the writer does it request an atomic move. Figure 25.2 makes the difference between “I am building it” and “it is ready for readers” visible.
The code is complete and can be compiled with javac --release 25 ImportArchive.java and run with java ImportArchive. The test program's paths are created in a temporary directory: the example does not depend on a folder name on the reader's computer.
import java.io.BufferedReader;
import java.io.BufferedWriter;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.AtomicMoveNotSupportedException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
public class ImportArchive {
static int importArchive(Path source, Path destination) throws IOException {
Path folder = destination.toAbsolutePath().getParent();
Path temporary = Files.createTempFile(folder, "archive-", ".tmp");
boolean published = false;
Throwable failure = null;
try {
int lines = 0;
int lineNumber = 0;
try (BufferedReader reader = Files.newBufferedReader(
source, StandardCharsets.UTF_8);
BufferedWriter writer = Files.newBufferedWriter(
temporary, StandardCharsets.UTF_8)) {
String line;
while ((line = reader.readLine()) != null) {
lineNumber++;
String value = line.strip();
if (value.isEmpty()) {
continue;
}
if (value.contains(";")) {
throw new IOException("Unexpected separator at line " + lineNumber);
}
writer.write(value.toUpperCase(java.util.Locale.ROOT));
writer.newLine();
lines++;
}
}
try {
Files.move(temporary, destination,
StandardCopyOption.ATOMIC_MOVE,
StandardCopyOption.REPLACE_EXISTING);
} catch (AtomicMoveNotSupportedException ex) {
throw new IOException("Atomic publication unavailable", ex);
}
published = true;
return lines;
} catch (IOException | RuntimeException | Error ex) {
failure = ex;
throw ex;
} finally {
if (!published) {
try {
Files.deleteIfExists(temporary);
} catch (IOException | RuntimeException cleanup) {
if (failure != null) {
failure.addSuppressed(cleanup);
} else {
throw cleanup;
}
}
}
}
}
public static void main(String[] args) throws IOException {
Path folder = Files.createTempDirectory("example-c25-");
Path source = folder.resolve("input.txt");
Path destination = folder.resolve("output.txt");
try {
Files.writeString(source, " book \n\njava\n", StandardCharsets.UTF_8);
System.out.println("Imported lines: " + importArchive(source, destination));
System.out.print(Files.readString(destination, StandardCharsets.UTF_8));
} finally {
Files.deleteIfExists(destination);
Files.deleteIfExists(source);
Files.deleteIfExists(folder);
}
}
}
createTempFile really creates the file, rather than inventing a name and hoping it is free. The file is created next to the destination because an atomic move between different file systems is generally unavailable. The published variable distinguishes success from an exceptional exit; in the second case, finally tries to delete the temporary file. If cleanup fails while handling an error, addSuppressed preserves that second problem without hiding the first. An undeleted temporary file will require later cleanup: we cannot promise that every failure leaves the directory empty. The two resources in the try are closed in reverse order. If reading or writing fails and closing also fails, the closing error is retained as a suppressed exception: it does not erase the cause that interrupted the work.
In the loop, readLine() returns null at the end. Every nonempty line is validated before being written. The semicolon condition is a constraint invented for the example, rather than a rule of text files. The lineNumber counter advances even on blank lines, while lines counts only imported titles: the error message therefore indicates the physical position in the input. A real archive would also have a declared grammar and a duplicate policy. Running the program, we observe Imported lines: 2, then BOOK and JAVA on separate lines. The result demonstrates the positive path; to test the negative one, insert invalid;title into the source and verify that the previous final file has not been replaced.
The failure test performs exactly this last experiment on temporary files: it creates a destination containing PREVIOUS VERSION, encounters a malformed line, then checks that the previous content is intact and no temporary file remains in the folder. It prints Previous destination intact; temporary file removed. The test makes observable the property for which we introduced the temporary file; a simple positive execution would not have demonstrated it. It does not yet prove resistance to a sudden crash or behavior on a different provider.
Warning – Atomic does not mean durable.
ATOMIC_MOVEasks the provider for a move indivisible to observers; if unsupported, the program fails explicitly. The combination of options and replacement of an existing destination have provider-dependent details: a real system must test them on its own file system. An atomic move does not by itself promise that bytes survive a sudden power loss: that is a different persistence guarantee. If the application accepts non-atomic publication, it must define and document a recovery plan, rather than hide the fallback inside acatch.
The term “atomic” appeared in chapter 20A for a decision on shared state in the JVM. Here the boundary differs: file system observers must see either the old file or the new one, according to the provider's guarantees. An atomic Java counter does not make file publication indivisible; ATOMIC_MOVE does not make Java fields updated before or after the move atomic. When an operation requires both things, coordination and failure handling between the two worlds must be designed explicitly.
Exploring directories without retaining resources
Files.createDirectories creates missing directories along a path, while createDirectory requires the parent to exist already. Files.list returns a Stream<Path> that keeps the directory open: it must be closed with try-with-resources, even though the type resembles the streams of chapter 18. Files.walk visits several levels and requires the same discipline. DirectoryStream is another choice for iterating through a directory without immediately constructing a list. A FileVisitor makes explicit the moments when we enter a directory, see a file, encounter an error and leave: it is especially useful when copying, cleanup or searching must react differently to each event.
To delete a tree, calling Files.delete on its nonempty root is not enough. A traversal is needed in which children are handled before the directory. Deletion can fail after already removing some children: the program must decide whether to continue, stop or record partial results. The same applies to copying or moving across several files. Files.copy and Files.move operate on specified paths; they are not a general transaction on the entire tree. Symbolic links require an explicit choice: copying them as links or following them changes both content and the boundaries crossed.
For a simple local search, we can write a fragment that closes the stream even if the filter throws an exception:
try (var paths = Files.list(folder)) {
long texts = paths.filter(p -> p.getFileName().toString()
.endsWith(".txt")).count();
System.out.println("Text files: " + texts);
}
This fragment assumes that folder is an already defined Path. Files.list observes only direct children, rather than the entire tree; Files.walk would change the problem and traversal cost. Filtering names does not demonstrate that a file really is text: an extension is a convention, rather than a verified property of content. If another process creates or removes files while we iterate, we must not treat the result as an atomic snapshot of the directory.
Temporary files, attributes and permissions are other questions about the boundary. Files.createTempFile lets us choose a directory and creates an unoccupied name; initial permissions and available attributes depend on the provider. A POSIX permission view is not guaranteed on every system. Files.readAttributes and methods such as size provide a snapshot: the file may change after reading. Do not confuse an observed measurement with a reservation of the resource.
Paths, links and traversals: choosing what we are crossing
The library receives the relative name inputs/../catalog.txt. resolve and normalize help construct and simplify the name. If the part being resolved is absolute, however, resolve returns that part: it does not force it inside the base directory. If the name comes from outside, the requirement “remain in the authorized folder” needs further checks. Lexical comparison after normalization is an initial check, but does not alone handle symbolic links or a concurrent replacement of the path.
A symbolic link contains a reference to another path. The name we observe and the file reached may therefore belong to different places. NOFOLLOW_LINKS changes the behavior of operations accepting it; it is not a global mode enabled once for the whole program. For a copy, decide whether the link or its content is of interest. For recursive traversal, Files.walk does not follow symbolic links by default; enabling FOLLOW_LINKS also introduces the problem of cycles and the boundaries crossed.
Key concept – Security and portability are contracts. A
Pathbelongs to a file system provider, which translates Java operations to the concrete system. Roots, case sensitivity, separators and attribute support may differ. A valid path is not authorization; a normalized path is not a file reservation. Sensitive operations on directories modifiable by other actors require permission design and, where available, tools such asSecureDirectoryStream, which operate relative to an open directory.
For catalog scanning, walkFileTree with a SimpleFileVisitor allows declaring what to do in visitFile, visitFileFailed and postVisitDirectory. The latter also receives any traversal error: deleting without checking it may hide why we did not complete the tree. The results CONTINUE, SKIP_SUBTREE and TERMINATE express a policy, rather than repairing an error. A traversal may stop after effects have already occurred; retain a report when those results must be recoverable.
Files.lines returns a lazy text stream that must be closed. An error may arise during consumption as UncheckedIOException, rather than only during opening as IOException. A pipeline from chapter 18 does not automatically make the source pure or resource-free. The useful variant is a file disappearing or becoming unreadable during scanning: first choose whether the report must be complete or explicitly partial, then write error handling.
Reading a packaged resource
An image distributed inside a JAR is not automatically a Path in the ordinary file system. Class.getResource resolves an absolute name with / from the class path or module root, or a name relative to the class's package; ClassLoader.getResource uses names without an initial / and has different search rules. If we need to read a packaged resource, getResourceAsStream avoids arbitrarily turning its URL into a path. Chapter 21 on classloaders explains why the physical location of a class and its resources must not be assumed.
Imagine a config/example.txt file distributed alongside the classes. ImportArchive.class.getResourceAsStream("/config/example.txt") requests the name from the root and returns a stream or null if the resource is not found. Before passing it to an InputStreamReader, check the missing case and declare UTF-8 if the content is text. The stream must be closed. If the user must modify the file, it should instead not be treated as an immutable internal resource: choose a configurable external Path. The transfer test is to actually package the resource in a JAR and repeat reading without assuming that the JAR has been extracted to disk.
Channels and buffer state
A Channel offers another interface for data and devices. ByteBuffer distinguishes capacity, position and limit: capacity is the buffer's space, position indicates where the next read or write will occur, and limit bounds the active part. flip() prepares the bytes just written into the buffer for reading; clear() prepares it to be filled again without physically erasing the bytes. This is a useful model for advanced I/O, but a linear text archive does not require automatically replacing a BufferedReader with channels and buffers.
Imagine a buffer of capacity eight. After writing three bytes, position is three and limit remains eight. flip() moves limit to three and position to zero: a subsequent read sees only the three valid bytes. If we skipped flip(), we would attempt to read from position three up to limit eight, meaning from the part not prepared for reading. If we called clear() thinking it zeroed the data, we might instead still observe old bytes through an inappropriate view. These numbers are a state model, rather than a recommendation to use an eight-byte buffer in production.
From bytes to the channel: a copy with explicit state
If the requirement is to reach a precise file position or integrate processing already based on buffers, FileChannel may be suitable. Keep the channel's state, which includes the position in the file, separate from the ByteBuffer state, which includes the active region in memory. The channel fills the buffer; flip() exposes the data just obtained; writing consumes that region; clear() prepares the next iteration.
In the IOContracts program, we copy eleven bytes with a buffer of capacity four. We will have two full blocks and one partial block. Writing is repeated until hasRemaining() becomes false: a single write does not promise to consume the whole buffer. For this example we use ordinary file channels; with a non-blocking channel, a return of zero would require a waiting strategy, otherwise a loop could occupy the CPU without making progress.
The core of the copy makes the four steps visible. Here source and destination are FileChannel objects opened in the complete program's try-with-resources; ByteBuffer belongs to java.nio.
ByteBuffer buffer = ByteBuffer.allocate(4);
while (source.read(buffer) != -1) {
buffer.flip();
while (buffer.hasRemaining()) destination.write(buffer);
buffer.clear();
}
On the final iteration, three bytes are read: flip() limits the region to write to three and prevents recopying the fourth byte left from the previous iteration. Comparing all bytes between input and output verifies this detail, as well as the final length.
The following diagram shows the first iteration, with four valid bytes. We do not erase the data when changing phase: we change the indices authorizing reading or writing them.
compact() serves a different problem: we have consumed only part of the data and must preserve the remainder to complete, for example, a message. It moves them to the beginning and prepares space for more input. Using clear() at that point would lose the remaining region from our reading protocol. rewind() instead allows rereading the active region while retaining its limit. Verification means following position and limit after every step, rather than memorizing three similar names.
Further reading – NIO does not always mean non-blocking. A
FileChanneldoes not become selectable simply because it belongs to NIO.Selectorcoordinates selectable channels, such as certain network channels, registered in non-blocking mode.AsynchronousFileChanneloffers another contract, based on operation completion: the buffer must not be reused while the operation using it is still underway. A direct buffer or memory-mapped file may be useful for specific workloads, but involves costs and a life cycle different from a simple array. Choose them after measurement and a concrete need.
If we read numbers with ByteBuffer.getInt, the format must establish byte order and the buffer must have enough remaining bytes. The fact that a Java integer has the same value on two machines does not mean every protocol encodes it in the same order. Our example's copy avoids this problem because it preserves bytes without interpreting them. The variant is a binary format with a header: first define length and order, then add field reading.
Watching a directory and rereading the state
WatchService can signal directory changes. It is useful for updating a view or starting a new scan, rather than treating every notification as a reliable and complete record: events may be coalesced, lost or delivered with system-dependent semantics. When the service signals overflow, a sensible strategy is to reconcile the directory's actual state. The same principle seen in C20 returns here: a notification invites checking a condition, rather than replacing the condition.
Suppose the library updates its archive when another program creates a file in a folder. Periodic scanning might be sufficient; WatchService becomes interesting when we want to react promptly. Register the directory for relevant events, then a loop reads the keys and their events. The name reported in an event is relative to the registered directory: resolve it against that directory before reading the file. After consuming the events, reset the key, checking whether it is still valid. If events are lost, the correct solution is a new state scan, rather than assuming the notification stream alone reconstructs everything that happened. Polling may be simpler when timeliness is not critical; this is a problem-driven choice.
Notifications, completion and state verification
A creation event does not guarantee that the writer has finished the file. If we import immediately, we may read an archive that is still incomplete. A useful protocol has the producer write a temporary file and publish the definitive name only upon completion; the consumer rereads the state and still applies validation. WatchService signals that checking is worthwhile, while format and publication establish what to accept.
In the loop, a WatchKey gathers events; after pollEvents() we call reset() and, if it returns false, registration is no longer valid. OVERFLOW requires a reconciliation scan. Registering a directory does not automatically register all its descendants. Closing the service and interrupting the thread waiting for a key are part of shutdown; do not confuse a requested stop with a failure to retry forever.
The transfer test rapidly creates several files, renames them and then compares the catalog with a complete scan. We do not impose a particular number of notifications as the result: we verify that reconciliation reaches the correct state. If the system requires a complete, ordered history of all modifications, this service is not the event log needed. An application protocol retaining that history is required.
Historical serialization and external formats
Native Java object serialization belongs to the heritage of existing systems, but binds the format to classes and introduces compatibility and security problems. For a new external archive, choose a format with an explicit contract and a suitable parser; Java SE does not provide a single general parser for JSON, CSV and XML. The example file is deliberately simpler and declares its own small grammar.
Imagine a colleague proposing to replace the title file with an ObjectOutputStream: it would save catalog objects directly and seem to avoid the work of defining lines. The initial advantage has a cost. When classes change, previously saved data may no longer be readable under the expected contract; when the file comes from outside, ObjectInputStream reconstructs object graphs and must not be treated as a harmless parser for untrusted input. The deserialization filter documented for Java 25 can restrict permitted classes and sizes in a legacy system, but does not by itself create a simple, stable public format.
If the catalog must be exchanged with another program, first define the necessary data, encoding, size limits, format version and response to invalid fields. Then choose a parser respecting that contract and test both a valid file and a truncated one. If instead we must maintain an already distributed Java archive, start from an inventory of the classes actually present and a controlled migration; do not convert it blindly. The reusable brick is recognizing who will read the file and for how long: this question guides the format more than the number of lines needed to write the first example.
Verifying the brick and transferring it
Change the problem: titles must be unique, input may contain millions of lines and the final file must not change if a line is malformed. Progressive reading remains suitable for available memory; however, we must decide how to recognize duplicates, using an in-memory Set if the size is manageable or an external strategy if it is not. Publishing the temporary file remains the final step. If the consumer must see all titles or none, direct writing loses the required property even if every line is correct.
A second variant replaces the local archive with a resource in the JAR. Now we cannot publish it with Files.move: we must distinguish the packaged read-only resource from the external working file. Being able to formulate this difference is the chapter's test: associate the problem with the correct abstraction, recognize its limit and change solution when the boundary changes.
Essential references. The Java 25
Filesspecification describes file system operations, options and possible errors; it does not guarantee that every provider offers the same atomic operations. The charset documentation defines conversion between bytes and characters. These sources support the API contracts; archive policy, including the choice to fail without a fallback, belongs to our example.References for further reading. InputStream defines partial reads and transfers; CharsetDecoder makes error policy explicit. ByteBuffer describes buffer indices and FileChannel the file channel contract. These references let us check the boundaries of the small experiments without turning a local result into a guarantee for every device.