Encoding option in Spark JDBC

Viewed 525

I want to read data from an Oracle DB using Spark JDBC in a specific charset encoding like us-ascii but I am unable to.

The code I tried as per this answer:

val res=spark.read.format("jdbc")
  .option("url", url)
  .option("user", "userid")
  .option("password", "pwd")
  .option("driver","oracle.jdbc.OracleDriver")
  .option("encoding", "us-ascii")
  .option("characterEncoding", "us-ascii")
  .option("query", tableQuery).option("fetchsize","10000")
  .load()

This always returns the data in utf-8 encoding.

Is there a way to achieve this?

1 Answers

As per Oracle docs "The JDBC drivers perform all character set conversions transparently. No user intervention is necessary for the conversions to occur".

It appears the Oracle JDBC driver does not support the connection params characterEncoding or encoding.

Better you can try out the below steps to understand the issue better

  1. Validate encoding works with Spark - Extract data into a delimited file with the same encoding and read the file by providing encoding detail then display data frame
  2. Ensure proper encoding is used in Oracle Database - Oracle JDBC drivers use to perform character set conversion for Java applications to depend on the character set the database uses and the equivalent Oracle Character Set Name is US7ASCII.
Related