Update PDF using pdfBox

Viewed 36

I would like to ask if having a PDF it is possible, using pdfbox libraries, to update it at a specific point.

I am trying to use a solution already online but seems the gettoken() method does not enter code heresection the words properly to allow me to find the part I would like to modify.

This is the code(Groovy):

for( int i = 0; i < dataContext.getDataCount(); i++ ) {
InputStream is = dataContext.getStream(i);
Properties props = dataContext.getProperties(i);
String searchString= "Hours worked";
String replacement = "Hours worked: 2";
File file = new File("\\\\****\\UKDC\\GFS\\PRE\\PREPROD\\Alchemer\\Template\\***.pdf"); 
PDDocument doc = PDDocument.load(file);  
for ( PDPage page : doc.getPages() )
    {
        PDFStreamParser parser = new PDFStreamParser(page);
        parser.parse();
        List tokens = parser.getTokens();
        logger.info("in Page");
        for (int j = 0; j < tokens.size(); j++) 
        {
            logger.info("tokens:"+tokens[j]);
            Object next = tokens.get(j);
             //logger.info("in Object");
             if (next instanceof Operator) 
             {
                Operator op = (Operator) next;
                String pstring = "";
                int prej = 0;
                
                //Tj and TJ are the two operators that display strings in a PDF
                if (op.getName().equals("Tj")) 
                {
                    logger.info("in Tj");
                    // Tj takes one operator and that is the string to display so lets update that operator
                    COSString previous = (COSString) tokens.get(j - 1);
                    String string = previous.getString();
                    logger.info("previousString:"+string);
                    string = string.replaceFirst(searchString, replacement);
                    previous.setValue(string.getBytes());
                } else 
                if (op.getName().equals("TJ")) 
                {
                    logger.info("in TJ:"+ op.getName());
                    COSArray previous = (COSArray) tokens.get(j - 1);
                    logger.info("previous:"+previous);
                    for (int k = 0; k < previous.size(); k++) 
                    {
                        Object arrElement = previous.getObject(k);
                        if (arrElement instanceof COSString) 
                        {
                            COSString cosString = (COSString) arrElement;
                            String string = cosString.getString();
                             logger.info("string:"+string);
                            if (j == prej || string.equals(" ") || string.equals(":") || string.equals("-")) {
                                pstring += string;
                            } else {
                                prej = j;
                                pstring = string;
                            }
                        }                       
                    }                        
                    logger.info("pstring:"+pstring);
                    if (searchString.equals(pstring.trim())) 
                    {  
                        logger.info("in searchString");
                        COSString cosString2 = (COSString) previous.getObject(0);
                        cosString2.setValue(replacement.getBytes());                           

                        int total = previous.size()-1;    
                        for (int k = total; k > 0; k--) {
                            previous.remove(k);
                        }                            
                    }
                }
            }
        }
        logger.info("in updatedStream");
        // now that the tokens are updated we will replace the page content stream.
        PDStream updatedStream = new PDStream(doc);
        OutputStream out = updatedStream.createOutputStream(COSName.FLATE_DECODE);
        ContentStreamWriter tokenWriter = new ContentStreamWriter(out);
        tokenWriter.writeTokens(tokens);            
        logger.info("in tokenWriter");
        out.close();
        page.setContents(updatedStream);
        
        doc.save("\\\\***\\UKDC\\GFS\\PRE\\PREPROD\\Alchemer\\***1.pdf");
    }

Executing the code I am trying to search "Hours worked" String and update with "Hours worked: 2" There are 2 questions: 1.When I execute and check the logs can see the Tokens are not created properly: enter image description here

enter image description here

So are created two different COSArrays meantime I have all in one Line:

enter image description here and this can be a problem if I have to search a specific word.

  1. When it find the word it seems it is working but it apply a strange char:

enter image description here

So Here 2 questions:

  1. How to manage to specify the token behaviour (or maybe for the parser) to get an entire phrase in the same token until a special char happen?
  2. Hot to format the new char in the new PDF?

Hope you can help me, thanks for your support.

0 Answers
Related