How to parse an HTML table with headers in in rows

Viewed 150

I have an HTML table similar to following where table headers are also within a row. How can I extract it in a single pass using a third party python package? (should be either a list or a dict)

<table>
<tr>
<th>Header 1</th><td>Value 1</td>
</tr>
<tr>
<th>Header 2</th><td>Value 2</td>
</tr>
<tr>
<th>Header 3</th><td>Value 3</td>
</tr>
</table>
1 Answers

I'm assuming, that you want a dictionary out of the table:

from bs4 import BeautifulSoup


txt = '''<table>
<tr>
<th>Header 1</th><td>Value 1</td>
</tr>
<tr>
<th>Header 2</th><td>Value 2</td>
</tr>
<tr>
<th>Header 3</th><td>Value 3</td>
</tr>
</table>
'''

soup = BeautifulSoup(txt, 'html.parser')

out = {}
for tr in soup.select('tr'):
    out[tr.select_one('th').get_text(strip=True)] = [td.get_text(strip=True) for td in tr.select('td')]

print(out)

Prints:

{'Header 1': ['Value 1'], 'Header 2': ['Value 2'], 'Header 3': ['Value 3']}

Or:

out = {}
for tr in soup.select('tr'):
    out[tr.select_one('th').get_text(strip=True)] = tr.select_one('td').get_text(strip=True)

print(out)

Prints:

{'Header 1': 'Value 1', 'Header 2': 'Value 2', 'Header 3': 'Value 3'}
Related