I just copied a piece of text from a PDF, the text contained a hyphenated word broken across two lines. The pasted text didn't contain the hyphen character and I'm wondering if someone can explain why that's the case.
I checked by opening the PDF in both Chrome, Firefox, and Acrobat Reader DC and had the same behaviour so I'm assuming it's either a feature of the Window's copy-pasting mechanism or something to do with the underlying encoding of the text. I would have guessed that the PDF file-format would hard-code the breaks into the text rather than dynamically (but consistently) calculating them?
If it's useful to know, I believe the document was generated in some form of TeX program, probably in LaTeX.
How does this work? Is it related to "Soft Hyphens"?