I have started a spark streaming job which streams data from kafka.I have assigned only 2 worker nodes with 15gb disk for testing.Within 2 hours the disk is full and the status of these nodes is showing as unhealthy on YARN Resource Manager web interface, and I have checked HDFS web interface which shows the Block Pool has used 95% of disk space. The problem is I am not storing any data on the nodes, just reading from kafka, processing and storing to MongoDB.