I've an html table like this:
<TABLE>
<TR>
<TD><P>Name</P></TD>
<TD><P>Fees</P></TD>
<TD><P>Awards</P></TD>
<TD><P>Total</P></TD>
</TR>
<TR>
<TD><P>Tony</P></TD>
<TD >7,800</TD>
<TD >7</TD>
<TD>15,400</TD>
</TR>
<TR>
<TD><P>Paul</FONT></P></TD>
<TD >7,800</TD>
<TD >7</TD>
<TD>15,400</TD>
</TR>
<TR>
<TD><P>Richard</P></TD>
<TD >7,800</TD>
<TD >7</TD>
<TD>15,400</TD>
</TR>
</TR>
</TABLE>
I want to extract the values of table. I'd tried the following.
import lxml.html
html = lxml.html.parse(''html_table)
text_value = html.xpath('//tr/td/text()')
packages = html.xpath('//tr/td/p')
p_content = [p.text_content() for p in packages]
is there any way to extract both the <p> text and the text of <td> to a single list ?