Parsing text in python with BeautifulSoup

Viewed 2038

I am using Enron email data from kaggle. https://www.kaggle.com/wcukierski/enron-email-dataset I am reading emails.csv file.I am using BeautifulSoup to parse the message column.

import pandas as pd
train = pd.read_csv( "C:\Users\JAYASHREE\Documents\NLP\enron-email-dataset (1)\emails.csv")
from bs4 import BeautifulSoup
message=train["message"]
message[0]
soup = BeautifulSoup(message[0],"lxml")
message=soup.body.p
print message

First line parsed by beautifulsoup prints the following output

<p>Message-ID: &lt;18782981.1075855378110.JavaMail.evans@thyme&gt;
Date: Mon, 14 May 2001 16:39:00 -0700 (PDT)
From: phillip.allen@enron.com
To: tim.belden@enron.com
Subject: 
Mime-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
X-From: Phillip K Allen
X-To: Tim Belden <tim belden="">
X-cc: 
X-bcc: 
X-Folder: \Phillip_Allen_Jan2002_1\Allen, Phillip K.\'Sent Mail
X-Origin: Allen-P
X-FileName: pallen (Non-Privileged).pst

Here is our forecast

 </tim></p>

I need to extract only this line Here is our forecast

The line followed by X-FileName

How to parse the text and retrieve the specific portion.

1 Answers
Related