Tika-app extracts pages (numbers/keynotes) as zip File and prints only the file names inside it. It doesn't return the exact content inside the document file.
I tried using autodetect parser ,which tried to parse it using IWorkDocument but fails to get content inside of it. Using tika-app-1.22 to extract the contents.
BodyContentHandler handler = new BodyContentHandler();
AutoDetectParser parser = new AutoDetectParser();
Metadata metadata = new Metadata();
try (InputStream stream = AutoDetectParseExample.class.getResourceAsStream("Hello.pages")) {
parser.parse(stream, handler, metadata);
return handler.toString();
}
expected result:
Lorem ipsum dolor sit amet,....
actual result:
Data/92317989_242x291px-small-17.jpeg
Data/108151441_276x185px-small-13.jpeg
Data/125144832_750x539px-small-11.jpeg
Data/200250285_276x185px-small-15.jpeg
Index/Document.iwa
Index/ViewState.iwa
Index/CalculationEngine-4759.iwa
Index/AnnotationAuthorStorage-4758.iwa
Index/DocumentStylesheet-4762.iwa
Index/DocumentMetadata.iwa
Index/Metadata.iwa
Metadata/Properties.plist
Metadata/DocumentIdentifier D45D90E8-2C22-4115-98BA-1EDBA675DD55
Metadata/BuildVersionHistory.plist
Template: 09_School_Report (2018-07-03 15:42) M7.3-5989-2
preview.jpg
preview-micro.jpg
preview-web.jpg