Can Apache Hudi be used to upsert a row from Apache Spark dataframe into Postgres database?

Viewed 33

Problem Statement: There is no upsert to database feature in Apache Spark, instead we have to overwrite the entire table. But Apache Hudi can be used to upsert one or more rows to table without overwriting the entire table.

I understand Apache Hudi is table/file format that can used along with S3. But can it also be used with Postgresdb or MySql or Oracledb?

1 Answers

Hudi manages the storage layer of the datasets on HCFS (Hadoop Compatible File System), the answer is no, Hudi cannot manage Postgresdb, MySql and Oracledb tables because they are not HCFS, and they will never be.

Instead, and where the jdbc DataFrameWriter can only append to existing table or overwrite it, you can use foreach or foreachPartition to call a function which create a jdbc connection (based on the language you want to use), and upsert the data in the table.

Related