Google Script - Read PDF file, recognised as text/html

Viewed 554

Do you have any idea how to read PDF files, which mimetype is text/html?

I have tried the snippet below, but OCR doesn't work, resulting in this issue "API call to drive.files.insert failed with error: OCR is not supported for files of type text/html"

function extractTextFromPDF(pdfID) {
      // PDF File URL
      // You can also pull PDFs from Google Drive
      var url =  "https://drive.google.com/file/d/"+pdfID
      var blob = UrlFetchApp.fetch(url).getBlob();
      var resource = {
        title: blob.getName(),
        mimeType: blob.getContentType(),
      };
    
      // Enable the Advanced Drive API Service
      var file = Drive.Files.insert(resource, blob, { ocr: true, ocrLanguage: 'en' });
    
      // Extract Text from PDF file
      var doc = DocumentApp.openById(file.id);
      var text = doc.getBody().getText();
    
      return text;
    }

Also, I have tried to convert files to any other format like .csv .css or text, but when did it the text is horrible, long HTML, with content encrypted I think. I considered splitting data from extracted HTML, but unfortunately, content is not there or is encrypted somehow.

What I want to do is to print the text from this wired pdf, so I can later write it to Google Sheets. Do you have any idea how I can read this file? File I am attaching a pdf here, so you can see what I am fighting with. https://drive.google.com/file/d/1HXQk6PU9hzBb26EwoFQ0840W6ZihDUIX/view?usp=sharing

1 Answers

I used your sample file, see how I did it below:

function myFunction() {
  var pdfFile = DriveApp.getFilesByName("222-1522118.pdf").next();
  var blob = pdfFile.getBlob();

  // Get the text from pdf
  var filetext = pdfToText( blob, {keepTextfile: false} );

  console.log(filetext)
}

Output:

output

I used Mogsdad's library pdfToText

Reference: Get text from PDF in Google

Related