XML pretty print add unnecessary whitespace element content containing CDATA

Viewed 88

I have a piece of Java code which pretty prints xml. When using a LSSerializer to pretty print the output, it is formatted nicely and indented but elements which contain CDATA behave strangely. The XML

<root><outer><inner><text><![CDATA[Content of the CDATA block]]></text></inner></outer></root>

gets transformed into the following xml

<?xml version="1.0" encoding="UTF-8"?><root>
    <outer>
        <inner>
            <text>
                <![CDATA[Content of the CDATA block]]>
            </text>
        </inner>
    </outer>
</root>

and has the CDATA element in a separate line. This causes issues when extracting the content later on with xpath expressions.

The code

@Test
public void testOutputXML() throws Exception {
    final Document document = loadXMLFromString( "<root><outer><inner><text><![CDATA[Content of the CDATA block]]></text></inner></outer></root>" );
    final String formattedXml = toXmlPrettyLS( document );
    final Document formattedDocument = loadXMLFromString( formattedXml );
    XPathFactory xPathfactory = XPathFactory.newInstance();
    XPath xpath = xPathfactory.newXPath();
    XPathExpression expr = xpath.compile("//text/text()");
    final String evaluate = expr.evaluate( formattedDocument );
    assertThat( evaluate ).isEqualTo( "Content of the CDATA block" );
}

private String toXmlPrettyLS( final Document document ) throws Exception {
    final ByteArrayOutputStream bos = new ByteArrayOutputStream();
    final DOMImplementationRegistry registry = DOMImplementationRegistry.newInstance();
    final DOMImplementationLS loadSave = ( DOMImplementationLS ) registry.getDOMImplementation( "LS" );
    final LSOutput output = loadSave.createLSOutput();
    output.setByteStream( bos );
    final LSSerializer serializer = loadSave.createLSSerializer();
    final DOMConfiguration config = serializer.getDomConfig();
    config.setParameter( "format-pretty-print", true );
    serializer.write( document, output );
    return String.valueOf( bos );
}

private Document loadXMLFromString( final String xml ) throws Exception {
    final DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
    factory.setNamespaceAware( true );
    final DocumentBuilder builder = factory.newDocumentBuilder();
    return builder.parse( new ByteArrayInputStream( xml.getBytes() ) );
}

is used to transform the xml and extract the content, the environment is Java 11.

How can I adjust the formatting to get

<text>![CDATA[Content of the CDATA block]]></text>

instead?

0 Answers
Related