You can use
import os, re
matches = []
filename = r'C:\Users\Documents\romeo.txt'
with open(filename, 'r') as f:
for line in f:
matches.extend([x for x in re.findall(r'\w+', line) if x[0].isupper()])
print(matches)
The idea is to extract all words with a simple \w+ regex and add only those to the final matches list that start with an uppercase letter.
See the Python demo.
NOTE: If you want to only match letter words use r'\b[^\W\d_]+\b' regex.
This approach is Unicode friendly, that is, any Unicode word with the first capitalized letter will be found.
You also ask:
Is there a way to limit this to only words that start with an upper case letter and not all uppercase words
You can extend the previous code to
[x for x in re.findall(r'\w+', line) if x[0].isupper() and not x.isupper()]
See this Python demo, "Hi, How ARE You?" yields ['Hi', 'How', 'You'].
Or, to avoid getting CaMeL words in the output, use
matches.extend([x for x in re.findall(r'\w+', line) if x[0].isupper() and all(i.islower() for i in x[1:])])
See this Python demo where all(i.islower() for i in x[1:]) makes sure all letters after the first one are all lowercase.
Fully regex approach
You can use PyPi regex module that has support for both Unicode property and POSIX character classes, \p{Lu}/\p{Ll} and [:upper:]/[:lower:]. So, the solution will look like
import regex
text = "Hi, How ARE You?"
# Word starting with an uppercase letter:
print( regex.findall(r'\b\p{Lu}\p{L}*\b', text) )
## => ['Hi', 'How', 'ARE', 'You']
# Word starting with an uppercase letter but not ALLCAPS:
print( regex.findall(r'\b\p{Lu}\p{Ll}*\b', text) )
## => ['Hi', 'How', 'You']
See the Python demo online where
\b - a word boundary
\p{Lu} - any uppercase letter
\p{L}* - any zero or more letters
\p{Ll}* - any zero or more lowercase letters