I need to use python to connect to a remote HDFS. One of the ways I found from google is to use pyarrow. Unfortunately, I got this error for this simple example. And I have no idea why this error occurs and how to solve it.
from pyarrow import fs
hdfs = fs.HadoopFileSystem('http://172.16.1.7', 8020)
Traceback (most recent call last):
File "/notebooks/test/0.1-hdfs-test.py", line 2, in <module>
hdfs = fs.HadoopFileSystem('http://172.16.1.7', 8020)
File "pyarrow/_hdfs.pyx", line 83, in pyarrow._hdfs.HadoopFileSystem.__init__
File "pyarrow/error.pxi", line 122, in pyarrow.lib.pyarrow_internal_check_status
File "pyarrow/error.pxi", line 99, in pyarrow.lib.check_status
OSError: Unable to load libjvm: /usr/java/latest//lib/server/libjvm.so: cannot open shared
object file: No such file or directory
There are two machines in my environment that one of the machines is my working station where my code stay and the other one contains the HDFS that I would like to read/write data. I use anaconda as my python environment.
Python=3.6.13,hdfs3=0.3.1,pyarrow=3.0.0. Java installed on my Centos7 machine is jdk1.8.0_144.
I see someone solved their issue by setting HADOOP_HOME. However, I did not install Hadoop on my working machine, do I need to also install it?