Python String Manipulation Extracting HTML Data

Viewed 48

Using Python, I am trying to extract from a html page data that changes constantly. I know that the data that I want is between a tag that looks like, 'abcd>' and a tag. EX: abcd>MyData... remaining html...

I can replace the html up to and including the abcd> tag by finding the unique occurrence of abcd> and using the replace method. That leaves me with MyData... remaining html. I can find the position of the tag in the remaining html.

Can anyone tell me how to replace the html starting with the tag along with the rest of the trailing html and assign 'MyData' to a variable?

In short, it looks like I can only remove characters from the left unless I know exactly what the data is that I want to extract. If I knew what the data was that I wanted to extract I wouldn't need to parse through the html to get it.

Thank you for your assistance.

Tom

1 Answers

Not sure I understand the question.
If you have an html string like :

string = '<html class="html__responsive " lang="en"><head><title>Python String Manipulation Extracting HTML Data - Stack Overflow</title></head><body>mybody</body></html>'

Let's say the target tags are < head > and < /head >.
You can use the split() method which returns a list with two elements.

split1 = string.split("<head>")
split2 = split1[1].split("</head>")
left = split1[0]
right = split2[1]
middle = split2[0]

Printing :

left =  <html class="html__responsive " lang="en">
right =  <body>mybody</body></html>
middle = <title>Python String Manipulation Extracting HTML Data - Stack Overflow</title>

Is this the answer you were expecting???

Related