I have noticed a very low throughput in our pipeline when inserting data from spark to cassandra (less than 1 MB/s per core).
When trying to tune write conf (spark.cassandra.output.concurrent.writes, spark.cassandra.output.batch.grouping.key and spark.cassandra.output.batch.size.rows) I am getting quickly a write timeout.
My questions:
- Is it recommended/normal to increase the cassandra write timeout when writing data in batch (via spark)?
- Is it possible to increase it just for spark workloads? or just for batch writes?
- The default value for
spark.cassandra.output.batch.size.bytesis 1024, I find that too low as a default value, I guess in most times that would correspond to 1 or 2 rows, am I missing something ?
I am using spark-cassandra-connector 2.4.3