How to handle Incremental Update in HDFS hadoop Map-Reduce

Viewed 4038

I have structured base text files in HDF which have data like this (in file.txt):

OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|

4295877341|^|136|^|4|^|1|^|I|!|
4295877346|^|136|^|4|^|1|^|I|!|
4295877341|^|138|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877341|^|145|^|14|^|1|^|I|!|
123456789|^|145|^|14|^|1|^|I|!|

The size of the file.txt is 30 GB.

I have incremental data file1.txt of size approx 2 GB coming up in the same format in HFDS like below:

OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|

4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|215|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877343|^|149|^|14|^|2|^|I|!|
123456789|^|145|^|14|^|1|^|D|!|

Now i have to combine file.txt and file1.txt and create a final text file that has all unique records.

The key in both files are OrgId. If the same OrgId is found in the first file then I have to replace with the new OrgId and if not then then I have to insert the new OrgId.

The Final Output is like this .

OgId|^|ItemId|^|segmentId|^|Sequence|^|Action|!|

4295877346|^|136|^|4|^|1|^|I|!|
4295877341|^|213|^|4|^|1|^|I|!|
4295877341|^|215|^|2|^|1|^|I|!|
4295877341|^|141|^|4|^|1|^|I|!|
4295877341|^|143|^|2|^|1|^|I|!|
4295877343|^|149|^|14|^|2|^|I|!|

How can i do it in mapreduce?

I am not going for the HIVE solution because I have so many distinct file like this, approx 10.000 and so I have to create 10.000 partitions in HIVE.

Any suggestion to use Spark for this use case ?

1 Answers
Related