I use pyarrow 2.0.0 to interact Hadoop 3.3 on CentOS 8. Installations of Hadoop and pyarrow module are successful. So I copied some local csv files into Hadoop file system. I try to read CSV files from Hadoop file system and convert the csv lines to list of string.
Below is my first attempt:
from pyarrow import fs
hdfs = fs.HadoopFileSystem('localhost', port=9000)
def readHdFile(filename):
with hdfs.open_input_file(filename) as inf:
read_data = inf.read().decode('utf-8')
return read_data
data = readHdFile('test.csv')
print(data)
Above codes work without errors. It print the data successfully. For example:
date,values,
2007-01-01,6.3
2008-01-01,6.7
2009-01-01,7.7
But the type of these lines are not list of string but just big size of string itself. So the next step is blocked because of the big size of returned string. Then I change the pyarrow method to CSV like below:
from pyarrow import csv
from pyarrow import fs
def readHdFile(filename):
with hdfs.open_input_file(filename) as inf:
read_data = csv.read_csv(inf)
return read_data
data = readHdFile('test.csv')
print(data)
But the returned value of pyarrow table type are not what I expect.
pyarrow.Table
date: timestamp[s]
values: double
How can I convert the CSV file stored in Hadoop file system into list of string type using pyarrow?