Is it possible to configure batch size for SingleStore Kafka pipeline?

Viewed 164

I'm using SingleStore to load events from Kafka. I created a Kafka pipeline with the following script:

create pipeline `events_stream`
as load data kafka 'kafka-all-broker:29092/events_stream'
batch_interval 10000
max_partitions_per_batch 6
into procedure `proc_events_stream`
fields terminated by '\t' enclosed by '' escaped by '\\'
lines terminated by '\n' starting by '';

And SingleStore failing with OOM error like the following:

Memory used by MemSQL (4537.88 Mb) has reached the 'maximum_memory' setting (4915 Mb) on this node. Possible causes include (1) available query execution memory has been used up for table memory (in use table memory: 71.50 Mb) and (2) the query is large and complex and requires more query execution memory than is available

I'm a quite confused why 4Gb is not enough to read Kafka by batches....

Is it possible to configure batch_size for the pipeline to avoid memory issues and make the pipeline more predictable?

1 Answers

unfortunately, current version of Singlestore's pipelines only has a global batch size and can not be set individually in stored procedures.

However, In general, each pipeline batch has some overhead, so running 1 batch for 10000 messages should be better in terms of total resources than 100 or even 10 batches for the same number of messages. If the stored procedure is relatively light, the time delay you are experiencing is likely dominated by the network download of 500mb.

It is worth checking if 10000 large messages are arriving on the same partition or several. In Singlestore's pipelines, each database partition downloads messages from a different kafka partition, and that helps parallelize the workload. However, if the messages arrive on just one partition, then you are not getting the benefits of parallel execution.

Related