I am using Enron email data from kaggle. https://www.kaggle.com/wcukierski/enron-email-dataset I am reading emails.csv file.I am using BeautifulSoup to parse the message column.
import pandas as pd
train = pd.read_csv( "C:\Users\JAYASHREE\Documents\NLP\enron-email-dataset (1)\emails.csv")
from bs4 import BeautifulSoup
message=train["message"]
message[0]
soup = BeautifulSoup(message[0],"lxml")
message=soup.body.p
print message
First line parsed by beautifulsoup prints the following output
<p>Message-ID: <18782981.1075855378110.JavaMail.evans@thyme>
Date: Mon, 14 May 2001 16:39:00 -0700 (PDT)
From: phillip.allen@enron.com
To: tim.belden@enron.com
Subject:
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-From: Phillip K Allen
X-To: Tim Belden <tim belden="">
X-cc:
X-bcc:
X-Folder: \Phillip_Allen_Jan2002_1\Allen, Phillip K.\'Sent Mail
X-Origin: Allen-P
X-FileName: pallen (Non-Privileged).pst
Here is our forecast
</tim></p>
I need to extract only this line Here is our forecast
The line followed by X-FileName
How to parse the text and retrieve the specific portion.