I have a table - table_1 -- 1.9Billion records -- orc format -- 229G size
I have a script to create table2 by joining table1 and <some other table>.
table2 - 30K records -- orc format -- 229G size
when I check the hdfs path for table2, I could see dir name - scratchdir and it occupies 220G of space.
[root@host ~]$ hadoop fs -du -s -h /db_name/data/db/db_name.db/managed/table2
220.8 G /db_name/data/db/db_name.db/managed/table2
[root@lnx1397 ~]$ ls -ltr /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2
total 2943
drwxr-xr-x 4 root root 2 Sep 16 23:29 _scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-17
drwxr-xr-x 4 root root 2 Sep 16 23:29 _scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-18
drwxr-xr-x 4 root root 2 Sep 16 23:30 _scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-19
drwxr-xr-x 4 root root 2 Sep 16 23:42 _scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-20
drwxr-xr-x 7 root root 5 Sep 16 23:45 _scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-1
-rwxr-xr-x 1 root root 3010663 Sep 17 00:25 000000_0
[root@host ~]$ du -sh /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/*
2.9M /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/000000_0
221G /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/_scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-1
1.0K /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/_scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-17
1.0K /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/_scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-18
1.0K /mapr/mapre04p//db_name/data/db/db_name.db/managed/table2/_scratchdir_hive_2021-09-16_23-29-42_576_3081873685168820741-19
1.0K /mapr/mapre04p//db_name/data/db/receipt
database volume is 1TB, i have script to drop table1 but the jobs are failing with disk quota exceeded error due to high volume of scratch dir
is there an hive property or configuration to remove these scratchdir