Reading a CSV file in Java is a common task for developers who need to import data from spreadsheets, logs, or external systems into their applications. CSV (Comma‑Separated Values) format remains popular because it is simple, human‑readable, and supported by almost every data‑processing tool. This guide walks you through the most reliable ways to parse CSV files in Java, explains the underlying mechanics, and provides practical code samples you can adapt to your own projects.
Why Choose CSV for Data Exchange in Java Applications
CSV files store tabular data as plain text, with each line representing a record and fields separated by a delimiter—most often a comma. The format’s simplicity makes it ideal for:
- Interoperability – Excel, Google Sheets, databases, and many ETL tools can read or write CSV without extra libraries.
- Low overhead – No binary parsing or schema definition is required, which keeps memory usage predictable.
- Human readability – Developers can open a CSV in any text editor to inspect or modify data quickly.
Despite its advantages, CSV has pitfalls: embedded commas, quoted fields, varying line endings, and inconsistent headers can break naïve parsers. Understanding these challenges helps you pick the right approach for reading a CSV file in Java.
Core Approaches to Parse CSV in Java
When you need to read a CSV file in Java, you have three main options:
- Manual parsing with
java.ioclasses – UsingBufferedReaderorScannerto read lines and split them withString.split(). - Dedicated third‑party libraries – Such as OpenCSV or Apache Commons CSV, which handle quoting, escaping, and custom delimiters out of the box.
- Java Streams API – Combining
Files.lines()withmapandcollectfor functional‑style processing (still relies on manual splitting unless you wrap a library).
Each method trades off control, convenience, and dependency size. For quick prototypes or environments where adding JARs is undesirable, the manual method works fine. For production‑grade applications, a library saves time and reduces bugs related to edge cases Worth keeping that in mind. And it works..
Manual CSV Reading with BufferedReader – Step‑by‑Step
Below is a complete example that demonstrates how to read a CSV file in Java using only the standard library. The code assumes a simple CSV where fields are not quoted and do not contain the delimiter Less friction, more output..
import java.io.BufferedReader;
import java.io.FileReader;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;
public class SimpleCsvReader {
public static List readCsv(String filePath) throws IOException {
List records = new ArrayList<>();
try (BufferedReader br = new BufferedReader(new FileReader(filePath))) {
String line;
while ((line = br.Practically speaking, readLine()) ! = null) {
// Split on comma; trim whitespace if needed
String[] values = line.split(",");
records.
public static void main(String[] args) {
String csvFile = "data/sample.On the flip side, csv";
try {
List rows = readCsv(csvFile);
for (String[] row : rows) {
System. out.println(java.util.Arrays.toString(row));
}
} catch (IOException e) {
System.On the flip side, err. println("Failed to read CSV: " + e.
**Explanation of key parts**
* `BufferedReader` wraps a `FileReader` to read the file line by line efficiently.
* The `try‑with‑resources` block guarantees the stream closes even if an exception occurs.
* `line.split(",")` creates an array of strings for each record. If your data may contain spaces after commas, you can use `line.split("\\s*,\\s*")` to trim them.
* The method returns a `List` where each inner array corresponds to a row.
**Limitations** – This approach fails when:
* A field contains a comma inside quotes (e.g., `"Doe, John"`).
* Fields are enclosed in double quotes to allow embedded newlines.
* The file uses a different delimiter like a semicolon or tab.
When any of these situations arise, switching to a library is the safer route.
## Using OpenCSV – A reliable Alternative
OpenCSV is a lightweight, widely adopted library that simplifies CSV parsing. Add the following Maven dependency to your project:
```xml
com.opencsv
opencsv
5.8
Basic OpenCSV Example
import com.opencsv.CSVReader;
import com.opencsv.exceptions.CsvValidationException;
import java.io.FileReader;
import java.io.IOException;
import java.util.List;
public class OpenCsvDemo {
public static void main(String[] args) {
String csvFile = "data/sample.Consider this: csv";
try (CSVReader reader = new CSVReader(new FileReader(csvFile))) {
List allRows = reader. Plus, readAll();
for (String[] row : allRows) {
System. Plus, out. println(java.Still, util. Day to day, arrays. toString(row));
}
} catch (IOException | CsvValidationException e) {
System.In practice, err. println("Error while reading CSV: " + e.
**Why OpenCSV helps**
* It automatically handles quoted fields, escaped quotes (`""`), and varying line endings.
* You can specify a custom delimiter, quote character, and escape character via the constructor.
* The library provides mapping to Java beans, letting you convert each row directly into a POJO.
### Mapping CSV to Java Beans
Suppose you have a CSV with columns `id,name,age`. Define a bean:
```java
public class Person {
private int id;
private String name;
private int age;
// getters and setters omitted for brevity
}
Then read the file with a header mapping:
import com.opencsv.bean.CsvToBean;
import com.opencsv.bean.CsvToBeanBuilder;
import com.opencsv.CSVReader;
import java.io.FileReader;
import java.io.IOException;
import java.util.List;
public class BeanCsvReader {
public static List readPersons(String filePath) throws IOException {
try (CSVReader reader = new CSVReader(new FileReader(filePath))) {
CsvToBean csvToBean = new CsvToBeanBuilder(reader)
.withType(Person.class)
.withIgnoreLeadingWhiteSpace(true)
.build();
return csvToBean.
This approach eliminates
manual field-by-field parsing and type conversion, reducing boilerplate code and the risk of index-mismatch bugs. The bean-mapping strategy also makes the code resilient to column reordering—provided the header names match the bean property names (or are annotated with `@CsvBindByName`).
## Writing CSV Files with OpenCSV
Reading is only half the story. OpenCSV’s `CSVWriter` and `StatefulBeanToCsv` make exporting data just as straightforward.
### Low-Level Writing
```java
import com.opencsv.CSVWriter;
import java.io.FileWriter;
import java.io.IOException;
public class CsvWriterDemo {
public static void main(String[] args) {
String[] header = {"id", "name", "age"};
String[] row1 = {"1", "Alice", "30"};
String[] row2 = {"2", "Bob", "25"};
try (CSVWriter writer = new CSVWriter(new FileWriter("output/people.NO_QUOTE_CHARACTER,
CSVWriter.csv"),
CSVWriter.writeNext(header);
writer.Day to day, dEFAULT_LINE_END)) {
writer. But dEFAULT_SEPARATOR,
CSVWriter. writeNext(row2);
} catch (IOException e) {
System.writeNext(row1);
writer.Practically speaking, dEFAULT_ESCAPE_CHARACTER,
CSVWriter. err.println("Failed to write CSV: " + e.
### Bean-Based Writing
```java
import com.opencsv.bean.StatefulBeanToCsv;
import com.opencsv.bean.StatefulBeanToCsvBuilder;
import com.opencsv.CSVWriter;
import java.io.FileWriter;
import java.io.IOException;
import java.util.List;
public class BeanCsvWriter {
public static void writePersons(List persons, String filePath) throws IOException {
try (FileWriter writer = new FileWriter(filePath)) {
StatefulBeanToCsv beanToCsv = new StatefulBeanToCsvBuilder(writer)
.Because of that, withSeparator(CSVWriter. DEFAULT_SEPARATOR)
.Day to day, withOrderedResults(false) // order follows @CsvBindByPosition or header annotation
. NO_QUOTE_CHARACTER)
.withQuotechar(CSVWriter.build();
beanToCsv.
Annotate the bean for explicit control:
```java
import com.opencsv.bean.CsvBindByName;
import com.opencsv.bean.CsvBindByPosition;
public class Person {
@CsvBindByPosition(position = 0)
@CsvBindByName(column = "id")
private int id;
@CsvBindByPosition(position = 1)
@CsvBindByName(column = "name")
private String name;
@CsvBindByPosition(position = 2)
@CsvBindByName(column = "age")
private int age;
// getters / setters
}
Handling Large Files – Streaming vs. In-Memory
The readAll() and parse() methods load the entire dataset into memory. For files larger than a few hundred megabytes, switch to streaming:
try (CSVReader reader = new CSVReader(new FileReader(largeFile))) {
String[] line;
while ((line = reader.readNext()) != null) {
// process one row at a time
}
}
For bean streaming, use CsvToBean’s iterator:
CsvToBean csvToBean = new CsvToBeanBuilder(reader)
.withType(Person.class)
.build();
for (Person p : csvToBean) {
// handle each Person without materialising the full list
}
Streaming keeps the memory footprint constant regardless of file size.
Performance Tips
| Concern | Recommendation |
|---|---|
| GC pressure | Reuse String[] buffers when parsing millions of rows in a tight loop (OpenCSV’s readNext() already reuses an internal buffer). |
| Charset | Explicitly specify StandardCharsets.UTF_8 (or the known encoding) in FileReader/FileWriter constructors to avoid platform-default surprises. |
| Parallelism | For CPU-bound post-processing, hand rows off to a CompletableFuture pipeline or a fixed thread pool after the single-threaded CSV read. On the flip side, |
| Validation | Enable CsvToBeanBuilder. That said, withThrowExceptions(false) and collect CsvToBean. getCapturedExceptions() for batch error reporting instead of failing fast. |
Common Pitfalls & Quick Fixes
| Symptom | Likely Cause | Fix |
|---|---|---|
CsvValidationException: Unterminated quoted field |
Multi-line value without closing quote | Ensure the source escapes newlines ("" inside quotes) or use a reader that supports multi-line (CSVReaderBuilder.withFieldAsNull(CSVReaderNullFieldIndicator.BOTH)). |
|---------|--------------|-----|
| CsvValidationException: Unterminated quoted field | Multi-line value without closing quote | Ensure the source escapes newlines ("" inside quotes) or use a reader that supports multi-line (CSVReaderBuilder.In real terms, withFieldAsNull(CSVReaderNullFieldIndicator. Still, bOTH)). |
| Leading/trailing spaces in fields | withIgnoreLeadingWhiteSpace(true) not set | Chain .withIgnoreLeadingWhiteSpace(true) onto the CSVReaderBuilder. |
| NullPointerException on bean mapping | Missing no-arg constructor or mismatched column names | Add a public no-arg constructor; verify column headers match @CsvBindByName values or supply a custom HeaderColumnNameMappingStrategy. That said, |
| Garbled characters in output | Default platform encoding used by FileWriter | Wrap FileWriter in an OutputStreamWriter with StandardCharsets. UTF_8. |
| Slow writes with millions of rows | StatefulBeanToCsv re-validates annotations on every call | Use CsvWriter directly for raw throughput, or batch writes in chunks of ~10 000 rows.
Advanced: Custom Separators, Quoting & Escaping
Not all CSV files use commas. OpenCSV lets you configure the entire dialect:
CSVWriterBuilder writerBuilder = new CSVWriterBuilder(new FileWriter("output.tsv"))
.withSeparator('\t')
.withQuoteChar(CSVWriter.NO_QUOTE_CHARACTER)
.withEscapeChar(CSVWriter.DEFAULT_ESCAPE_CHARACTER);
try (CSVWriter writer = writerBuilder.build()) {
writer.writeAll(rows);
}
For RFC 4180–compliant quoting (double-quote around every field), use:
.withQuoteChar('"')
.withStrictQuotedStrings(true)
This is essential when fields themselves contain commas, quotes, or newlines.
Integrating with Spring Boot
If you are building a REST service that imports CSV uploads, the typical flow is:
- Accept a
MultipartFilein a controller. - Wrap it in an
InputStreamReaderwith the correct charset. - Pipe it through
CsvToBeanBuilderand persist the resulting beans via a repository.
@PostMapping("/import")
public ResponseEntity importCsv(@RequestParam("file") MultipartFile file) {
try (InputStreamReader isr = new InputStreamReader(file.getInputStream(), StandardCharsets.UTF_8)) {
CsvToBean csvToBean = new CsvToBeanBuilder(isr)
.withType(Person.class)
.withIgnoreLeadingWhiteSpace(true)
.build();
List persons = csvToBean.parse();
personRepository.size() + " records.Also, badRequest(). ok("Imported " + persons.");
} catch (Exception e) {
return ResponseEntity.saveAll(persons);
return ResponseEntity.body("Import failed: " + e.
For very large uploads, push the parsing into a `@Async` method or a message queue to avoid blocking the HTTP thread.
---
## Conclusion
OpenCSV provides a pragmatic, well-documented toolkit for reading and writing CSV data in Java. Whether you are working with small configuration files or multi-gigabyte data dumps, the library offers both high-level abstractions—`readAll()`, `StatefulBeanToCsv`, and annotation-driven mapping—and low-level streaming APIs that keep memory usage predictable.
Most guides skip this. Don't.
The key takeaways are straightforward: **annotate your beans** for clean object mapping, **stream when dealing with large files** to avoid `OutOfMemoryError`, **always specify a charset** to prevent silent data corruption, and **configure the dialect** (separator, quoting, escape) to match the source format. By combining these practices with the validation and error-collection strategies outlined above, you can build strong CSV ingestion pipelines that are both performant and maintainable.
As data interchange formats evolve, CSV remains one of the most widely supported standards, and OpenCSV ensures that working with it in Java is never more complicated than it needs to be.