i am working with tesseract and using the following command to convert image into searchable pdf form.
tesseract test.png -l urd -psm 3 result pdf
this is the image which i converted in pdf.
after conversion, when i copy the text in pdf file and paste in any text editor (word, notepad etc) 4, i get the following result.
Lf ELINOR BI LF ERE I LPM DAT? MON IVAN DEBI OE SI D7 Pipips FEIN AAASQE PIAA IG or esddspp- PLDI AOL ko26RDLT HOY
i have tried both ways (opening pdf file in acrobat and opening the file in browser and the copy/paste the data in text editor, both did not work for me, i also tried all solution given on following two links, not a single solution worked for me.
https://stackoverflow.com/questions/9143154/how-to-cut-paste-from-pdf-with-non-ascii-encoding
And
https://stackoverflow.com/questions/12703387/pdf-font-encoding-why-cant-i-copy-text-from-a-pdf
any help will be greatly appreciated. thanks in advance.
