I would like to build a corpus for movie scripts to analyze gender representation in film discourse. I would like to know a way to process .pdf or .txt format of movie scripts so that I can:
- separate dialogues from scenes
- extract only the dialogues of the main character (providing all names are in caps, and lines are separated by linebreak)

Thank you!