How to change text encoding to utf-8 while using apache tika text parsing (most specifically for .txt files)

Viewed 763

I'm using apache tika for text extraction. It was working fine over almost all filetypes unless I tried testing it over a Chinese machine with a .txt document written in Chinese. I did not save the file in utf-8 encoding format. Tika started parsing wrong string characters. This seems to be an encoding issue, I tried setting encoding type like this metadata.add(Metadata.CONTENT_ENCODING, "UTF_8") still no luck. I've seen some methods in java that convert text from one encoding type to another but only if the source encoding type is known. In my case, I'm not sure about the client's encoding type and can't force him to use utf-8. kindly help me with this!! Thanks in advance:)

1 Answers

I had the same issue but when converting Powerpoint to text and I found out that by using the correct OutputStream which you can specify the encoding, the encoding is working well. The metadata you try to add changes nothing for the conversion but just add the line in the headers of the html file.

Here is my code:

public String tranformPowerpointToText(File file) throws IOException, TikaException {
        ByteArrayOutputStream byteArrayOutputStream = new ByteArrayOutputStream();
        ToTextContentHandler toTextContentHandler= new ToTextContentHandler(byteArrayOutputStream, "UTF-8");

        AutoDetectParser parser = new AutoDetectParser();

        Metadata metadata = new Metadata();
        try (InputStream stream = new FileInputStream(file)) {
            parser.parse(stream, toTextContentHandler, metadata);
            return byteArrayOutputStream.toString();
        } catch (SAXException e) {
            e.printStackTrace();
        }
}
Related