unable to copy paste urdu text from pdf file (getting weired english text instead urdu in my editor)

Viewed 2394

i am working with tesseract and using the following command to convert image into searchable pdf form.

tesseract test.png -l urd -psm 3 result pdf

this is the image which i converted in pdf.

enter image description here

after conversion, when i copy the text in pdf file and paste in any text editor (word, notepad etc) 4, i get the following result.

Lf ELINOR BI LF ERE I LPM DAT? MON IVAN DEBI OE SI D7 Pipips FEIN AAASQE PIAA IG or esddspp- PLDI AOL ko26RDLT HOY

i have tried both ways (opening pdf file in acrobat and opening the file in browser and the copy/paste the data in text editor, both did not work for me, i also tried all solution given on following two links, not a single solution worked for me.

https://stackoverflow.com/questions/9143154/how-to-cut-paste-from-pdf-with-non-ascii-encoding

And

https://stackoverflow.com/questions/12703387/pdf-font-encoding-why-cant-i-copy-text-from-a-pdf

any help will be greatly appreciated. thanks in advance.

0 Answers
Related