Best approach for Cartesian product or 2 large text files

Viewed 186

I have problem where I want to merge 2 large text files together and generate new file with cartesian product of 2 input files. I do know how code would look but not sure in which language to build such a utility. I have windows server and I'm familiar with C#, Shell script.

Note : File1 can be around 20 MB and File2 can contain around 6000 records. So what I want to achieve is Copy 20MB data 6000 times in new file.

Below are smaller examples of how my files would look like

File1

Head-A-AA-AAA
Child-A1-AA1-AAA1
Child-A2-AA2-AAA2
Child-A3-AA3-AAA3
Head-B-BB-BBB
Child-B1-BB1-BBB1
Child-B2-BB2-BBB2
Child-B3-BB3-BBB3

File2

Store1
Store2
Store3

Expected output file

Store1
Head-A-AA-AAA
Child-A1-AA1-AAA1
Child-A2-AA2-AAA2
Child-A3-AA3-AAA3
Head-B-BB-BBB
Child-B1-BB1-BBB1
Child-B2-BB2-BBB2
Child-B3-BB3-BBB3
Store2
Head-A-AA-AAA
Child-A1-AA1-AAA1
Child-A2-AA2-AAA2
Child-A3-AA3-AAA3
Head-B-BB-BBB
Child-B1-BB1-BBB1
Child-B2-BB2-BBB2
Child-B3-BB3-BBB3
Store3
Head-A-AA-AAA
Child-A1-AA1-AAA1
Child-A2-AA2-AAA2
Child-A3-AA3-AAA3
Head-B-BB-BBB
Child-B1-BB1-BBB1
Child-B2-BB2-BBB2
Child-B3-BB3-BBB3

Looking for suggestion if C# code with windows service will serve purpose or I need to use any other tool/utility/scripting?

EDIT : Created below c# code. But it's taking hours to generate 150 GB output file. I'm looking for faster way. I'm taking content from file 1 and copying it for each record in second file

FileInfo[] fi;
            List<FileInfo> TodaysFiles = new List<FileInfo>();
            string PublishId;
            DirectoryInfo di = new DirectoryInfo(@"\\InputPath");

            fi = di.GetFiles().Where(file => file.FullName.Contains("TRANSMIT_MASS")).ToArray();

            foreach (FileInfo f in fi)
            {
                string[] tokens = f.Name.Split('_');
                if(tokens[2] == DateTime.Now.AddDays(1).ToString("MMddyyyy"))
                {
                    PublishId = tokens[0];
                    string MACSFile = @"\\OutputPath\\" + PublishId + ".txt";
                    string path =f.FullName;

                    string StoreFile = di.GetFiles().Where(file => file.Name.StartsWith(PublishId) && file.Name.Contains("SUBS")).Single().FullName;

                    using (FileStream fs = File.Open(StoreFile, FileMode.Open, FileAccess.Read, FileShare.ReadWrite))
                    using (BufferedStream bs = new BufferedStream(fs))
                    using (StreamReader sr = new StreamReader(bs))
                    {
                        using (StreamWriter outfile = new StreamWriter(MACSFile))
                        {
                            String StoreNumber;
                            while ((StoreNumber = sr.ReadLine()) != null)
                            {
                                Console.WriteLine(StoreNumber);
                                if (StoreNumber.Length > 5)
                                {
                                    using (FileStream fsProfile = File.Open(path, FileMode.Open, FileAccess.Read, FileShare.ReadWrite))
                                    using (BufferedStream bsProfile = new BufferedStream(fsProfile))
                                    using (StreamReader srProfile = new StreamReader(bsProfile))
                                    {
                                        outfile.WriteLine(srProfile.ReadToEnd().TrimEnd());
                                        
                                    }

                                }

                            }
                        }
                    }

                }
            }
3 Answers

You mention shell script. Here's a working shell example:

while read line; do
  echo "$line" >> Output
  cat File1 >> Output
done < File2

Here the lines of File2 are being looped over and written along with the entirety of File1 into an arbitrary output file Output.

Easily run by saving it in a local file something.sh and running sh something.sh.

We could further optimise the code for performance, at the cost of memory. All refactor it to make it cleaner.

File1 : 6000 lines

File2 : 20Mb

As File 1, (smaller file) just contains a few number of lines, would read the entire file into memory and loop over it.

foreach (string line in File.ReadAllLines(File1))

If you still have memory capacity, you can read the entire second file into memory as well

var file2 = File.ReadAllText(File2)

Now all you have to do is append everything to a 3rd file. Which we will not store in memory because of size.

So the entire code will be

var file2 = File.ReadAllText(File2);
var destinationFile = "destination/file/path";

foreach (string line in File.ReadAllLines(File1)){
File.AppendAllText(destinationFile, line);
File.AppendAllText(destinationFile, file2);
}

Further Optimisation: Skipped to keep code simple

File.AppendAllText is called twice, because we don't want to do line + file2 in code. It will allocate more memory.

To optimise this further you can use StringBuilder, load file2 into it.

var file2 = new StringBuilder(File.ReadAllText(File2));

And mutate it. This should, prevent the 2 calls to File.AppendAllText and give more performance.

It is difficult to reduce I/O time. You can try the case with a reading/writing in large portions (I think It is more efficient because I/O operations require to allocate/release resources of OS). So if you read all, aggregate the result in-memory, write to file, then it will spend less time on I/O. A higher speed here is reached by in-memory operations, because RAM and processor operations are very rapid to process in comparison with an IO operation.

  1. File 1 - is small - read it once and keep the results in memory.
  2. File 2 - is large - read it in chunks. For example, you can use streamReader.ReadLine() N times
  3. Combine in-memory data of the first file with each chunk of the second one parallelly if possible.
  4. Output - open/close stream only once, write after each chuck is processed.

PS: no need in buffered streams here because file streams are already buffered. Buffered streams are useful for network IO operations.

Related