How to access a GCS Blob that contains an xml file in a bucket with the pandas.read_xml() function in python?

Viewed 56

I would like to access a blob file via the pandas.read_xml() function. Like this:

pandas.read_xml(blob.open())

When printing the blob it looks like this:

<Blob: Bucket, filename.0.xml.gz, 1612169959288959>

the blob.open()function gives this:

<_io.TextIOWrapper encoding='iso-8859-1'>

and I get the error UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8b in position 1: invalid start byte. When I change the code to: blob.open(mode='rt', encoding='iso-8859-1') I get ther error lxml.etree.XMLSyntaxError: Start tag expected, '<' not found, line 1, column 1.

Is there even a way to read in a xml file from a bucket on gcs?

1 Answers

read_xml() can directly read GCS files. Just provide the GCS URI and it can transform it to a dataframe. See sample code below and testing:

Sample file stored in GCS:

<?xml version="1.0" encoding="UTF-8"?>
<root xmlns="http://example.com">
    <bathrooms>
        <n35237 type="number">1.0</n35237>
        <n32238 type="number">3.0</n32238>
        <n44699 type="number">nan</n44699>
    </bathrooms>
    <price>
        <n35237 type="number">7020000.0</n35237>
        <n32238 type="number">10000000.0</n32238>
        <n44699 type="number">4128000.0</n44699>
    </price>
    <property_id>
        <n35237 type="number">35237.0</n35237>
        <n32238 type="number">32238.0</n32238>
        <n44699 type="number">44699.0</n44699>
    </property_id>
</root>

Code:

import pandas as pd

df = pd.read_xml("gs://my-bucket/note.xml.gz",compression="gzip")

print(df)

Output:

enter image description here

Related