I have a method that reads the content of a csv file. This works.
CSVParser parser = new CSVParserBuilder().withSeparator(';').build();
try (CSVReader reader = new CSVReaderBuilder(
new FileReader(getFilePath()))
.withCSVParser(parser)
.build()) {
String[] lineArray;
while ((lineArray = reader.readNext()) != null) {
for (int i = 0; i < lineArray.length; i++) {
String s = lineArray[i];
lineArray[i] = s.replace("\n", "");
}
sheetContent.add(new ArrayList<>(List.of(lineArray)));
}
} catch (IOException e) {
e.printStackTrace();
}
But some Strings are not read out as I imagined. e.g.
expected output : Geschäft
current output : Gesch�fts
So I think I need to change the encoding. Therefore I check the encoding before reading.
This looks like this:
try (InputStream inputStream = new FileInputStream(getFilePath())) {
Charset charSet = Charset.forName(new TikaEncodingDetector().guessEncoding(inputStream));
if (charSet.equals(StandardCharsets.ISO_8859_1)) {
}
} catch (IOException e) {
e.printStackTrace();
}
In this example, the file has the StandardCharsets.ISO_8859_1, so I check for it in the if statement.
I had the idea to store the content of the file in a byte[]
byte[] fileContent;
File file = new File(getFilePath());
FileInputStream fis = new FileInputStream(getFilePath());
fileContent = new byte[(int) file.length()];
fis.read(fileContent);
byte[] data = fileContent;
and convert it via newString() to UTF-8.
byte[] utf8 = new String(data, StandardCharsets.ISO_8859_1).getBytes(StandardCharsets.UTF_8);
However, with this approach I don't get further.
Does anyone know how I can convert the encoding to UTF-8 before reading it out?