Spark streaming from eventhub: how to stop stream once there is no more data?

Viewed 747

What I am trying to do, is to read some data from my event hub, and save it in azure data lake. However, the issue is, that the stream doesn't stop, and the writeStream step is not triggered. I am not able to find any setting to identify when the input rate reaches 0 in order to stop the stream then.

enter image description here

1 Answers

There is a special trigger in Apache Spark often called Trigger.Once - it will process all available data, and then shutdown the stream. Just add the .trigger(once=True) after .writeStream to enable it.

The only problem with it is that in Spark 3.x (DBR >= 7.x), it completely ignore options like maxFilesPerTrigger, etc. that are limiting an amount of data pulled for processing - in this case it will try to process all data in one go, and sometimes it may lead to a performance problems. To workaround that you may do following hack - assign result of raw_data.writeStream.....start(), like query = raw_data.writeStream.... to a variable - and check periodically the value of query.get('numInputRows'), and if it's equal to 0 for a some period of time, issue query.stop()

Related