Read a file in Azure data lake storage using pandas

Viewed 132

I am trying to read a parquet file which is stored in adls:

import pandas as pd
parquet_file = 'abfss://<>abc.parquet'
pd.read_parquet(parquet_file, engine='pyarrow')

But it gives the below error:

ValueError: Protocol not known: abfss

Is the only way to make it work is to read the file through pyspark and then convert it into pandas dataframe?

I installed adlfs

pip install adlfs

But now I am getting the following error -

ClientAuthenticationError: Server failed to authenticate the request. Please refer to the information in the www-authenticate header.
2 Answers

There are 2 approaches to make it work.

  1. Downgrade Pandas version, as suggested here

      dbutils.library.installPyPI("pandas", version="version_number")
    
      dbutils.library.restartPython()
    
  2. If first approach does not work, you need to get data in rdd first and then create Pandas dataframe.

    df = spark.read.parquet('<the path of your parquet file>')
    pandas_df = df.toPandas()
    

If possible, I would recommend mounting the storage and query the data using this mount point using :

df = spark.read.parquet('dbfs://mnt/<mounting_path>abc.parquet').toPandas()

Also, I would suggest using the Pandas API on Spark to leverage the benefits of using Databricks. You'd keep the syntax of Pandas but would call the Spark API to do your computations. It's just as simple as :

import pyspark.pandas as pd
Related