SPARK SQL - update MySql table using DataFrames and JDBC

Question

1 Answer

Amit Rawat · Answer 1 · 2019-07-28T06:39:24+0000

It is not actually possible. As for Spark 1.6.0+ Spark DataFrameWriter supports only four writing modes:

SaveMode.Overwrite: overwrite the existing data.
SaveMode.Append: append the data.
SaveMode.Ignore: ignore the operation (i.e. no-op).
SaveMode.ErrorIfExists: default option, throw an exception at runtime.

You can insert manually for example using mapPartitions (since you want an UPSERT operation should be idempotent and as such easy to implement), write to temporary table and execute upsert manually, or use triggers.

In general achieving upsert behavior for batch operations and keeping decent performance is far from trivial. You have to remember that in general case there will be multiple concurrent transactions in place (one per each partition) so you have to ensure that there will be no write conflicts (typically when application specific partitioning is used) or provide appropriate recovery procedures. In practice it may be better to perform and batch writes to a temporary table and resolve upsert part directly in the database.

SPARK SQL - update MySql table using DataFrames and JDBC

1 Answer

Related questions

Browse Categories

Browse By Domains

Popular Courses

Popular Tutorials

Popular Resources