I have a piece of Java code which pretty prints xml. When using a LSSerializer to pretty print the output, it is formatted nicely and indented but elements which contain CDATA behave strangely. The XML
<root><outer><inner><text><![CDATA[Content of the CDATA block]]></text></inner></outer></root>
gets transformed into the following xml
<?xml version="1.0" encoding="UTF-8"?><root>
<outer>
<inner>
<text>
<![CDATA[Content of the CDATA block]]>
</text>
</inner>
</outer>
</root>
and has the CDATA element in a separate line. This causes issues when extracting the content later on with xpath expressions.
The code
@Test
public void testOutputXML() throws Exception {
final Document document = loadXMLFromString( "<root><outer><inner><text><![CDATA[Content of the CDATA block]]></text></inner></outer></root>" );
final String formattedXml = toXmlPrettyLS( document );
final Document formattedDocument = loadXMLFromString( formattedXml );
XPathFactory xPathfactory = XPathFactory.newInstance();
XPath xpath = xPathfactory.newXPath();
XPathExpression expr = xpath.compile("//text/text()");
final String evaluate = expr.evaluate( formattedDocument );
assertThat( evaluate ).isEqualTo( "Content of the CDATA block" );
}
private String toXmlPrettyLS( final Document document ) throws Exception {
final ByteArrayOutputStream bos = new ByteArrayOutputStream();
final DOMImplementationRegistry registry = DOMImplementationRegistry.newInstance();
final DOMImplementationLS loadSave = ( DOMImplementationLS ) registry.getDOMImplementation( "LS" );
final LSOutput output = loadSave.createLSOutput();
output.setByteStream( bos );
final LSSerializer serializer = loadSave.createLSSerializer();
final DOMConfiguration config = serializer.getDomConfig();
config.setParameter( "format-pretty-print", true );
serializer.write( document, output );
return String.valueOf( bos );
}
private Document loadXMLFromString( final String xml ) throws Exception {
final DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
factory.setNamespaceAware( true );
final DocumentBuilder builder = factory.newDocumentBuilder();
return builder.parse( new ByteArrayInputStream( xml.getBytes() ) );
}
is used to transform the xml and extract the content, the environment is Java 11.
How can I adjust the formatting to get
<text>![CDATA[Content of the CDATA block]]></text>
instead?