I have structured base text files in HDF which have data like this (in file.txt):
OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|
4295877341|^|136|^|4|^|1|^|I|!|
4295877346|^|136|^|4|^|1|^|I|!|
4295877341|^|138|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877341|^|145|^|14|^|1|^|I|!|
123456789|^|145|^|14|^|1|^|I|!|
The size of the file.txt is 30 GB.
I have incremental data file1.txt of size approx 2 GB coming up in the same format in HFDS like below:
OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|
4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|215|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877343|^|149|^|14|^|2|^|I|!|
123456789|^|145|^|14|^|1|^|D|!|
Now i have to combine file.txt and file1.txt and create a final text file that has all unique records.
The key in both files are OrgId. If the same OrgId is found in the first file then I have to replace with the new OrgId and if not then then I have to insert the new OrgId.
The Final Output is like this .
OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|
4295877346|^|136|^|4|^|1|^|I|!|
4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|215|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877343|^|149|^|14|^|2|^|I|!|
How can i do it in mapreduce?
I am not going for the HIVE solution because I have so many distinct file like this, approx 10.000 and so I have to create 10.000 partitions in HIVE.
Any suggestion to use Spark for this use case ?