I am trying to write a regular expression to select the text I want from a corpus, and then write the extracted text into a dataframe in CSV format.
Here is the code that I used:
import re
import pandas as pd
def main():
pattern = re.compile(r'(case).(reason)(.+)(})')
with open('/Users/cleantext.txt', 'r') as f:
content = f.read()
matches = pattern.finditer(content)
for match in matches:
print(tuple(match.groups()))
# Create a DF for the expenses
df = pd.DataFrame(data=[tuple(match.groups())])
df.to_csv("judgement.csv", index=True)
if __name__ == '__main__':
main()
However the CSV would only return one line of output:
,0,1,2,3
0,xxx,yyy,zzz,}
where I was expecting multiple lines since the corpus contained at least 100 judicial judgements.
the orginal corpus looks something like this:
{mID a9d50454f624 case xxx reason yyy judgement zzz}
{mID a9d5049e34e934bff9b case xxx reason yyy judgement zzz}
{mID a67c9e34e934bff9b case xxx reason yyy judgement zzz}
Thank you so much for your help.