How to read data from Cassandra to Pyspark that are running on docker?

Viewed 285

I am running the Cassandra and Pyspark container on a same network(network name: test-cs) with these commands on the Docker :

docker run --name cassandra -v $HOME/Documents/datastax/cassandra:/var/lib/cassandra --network test-cs -d datastax/cassandra:4:0

docker run --name pyspark -p 8888:8888 -p 4040:4040 -p 4041:4041 -p 4042:4042 -e CHOWN_HOME=yes -e GRANT_SUDO=yse -e NB_GID=1000 -e NB_GID=100 -v $HOME/Documents/spark:/home/jovyan/work --network test-cs jupyter/pyspark-notebook

I want to read data from tables that are exist on The cassandra table, so I use these pyspark codes on the Jupyter notebook for connecting spark to Cassandra:

# Configuratins related to Cassandra connector & Cluster
import os
os.environ['PYSPARK_SUBMIT_ARGS'] = '--packages com.datastax.spark:spark-cassandra-connector_2.11:2.3.0 --conf spark.cassandra.connection.host=cassandra pyspark-shell' 

pay attention that I use cassandra(cassandra container name) value for "spark.cassandra.connection.host" argument of these code instead of Ip(127.0.0.1).

# Creating PySpark Context
from pyspark import SparkContext
sc = SparkContext("local", "movie lens app")

# Creating PySpark SQL Context
from pyspark.sql import SQLContext
sqlContext = SQLContext(sc)

# Loads and returns data frame for a table including key space given
def load_and_get_table_df(keys_space_name, table_name):
    table_df = sqlContext.read\
        .format("org.apache.spark.sql.cassandra")\
        .options(table=table_name, keyspace=keys_space_name)\
        .load()
    return table_df

# Loading movies & ratings table data frames
movies = load_and_get_table_df("movie_lens", "movies")
ratings = load_and_get_table_df("movie_lens", "ratings")

after run the above codes, I see errors, and I can't read data from the Cassandra and connect to it. please help me because I am very beginner to communicating Pyspark and cassandra.

0 Answers
Related