To preface this, I am not the greatest in C and have learnt a lot in order to improve this code and thus many be incorrect about some of the things which I am talking about.
Background
I inherited some code which works on smaller datasets. However, my task is to adapt this code to run on datasets which are significantly larger (~3 billion values vs 4 million) using the same techniques.
However, what I found was that I needed to change some of the array declarations to malloc's because otherwise the allocations with a predetermined size would give me a segfault as this would happen on the larger dataset but it would work correctly on the smaller dataset.
I soon ran into the issue that the memory required was going to be significantly larger than the amount of memory on my machine (my machine has 40 GB while it would have needed ~300 GB). As such, I used mmap to map these arrays to files on my computer. A few changes from ints to ssize_ts later to account for the index positions which were greater than the int range, the first two parts of the three part program was was working. However, I am having significant issues with trying to get the third part working. I recently asked a question here however although I accepted the answer, after running the code (which took ~20 minutes, it errored out again.)
I believe however that this may not have been the root of the issue but still am not sure what to do and have been working at this specific problem for about a month now.
The Problem
Although I cannot be completely sure that this is the correct problem, what seems to be the issue is that I need to allocate a 1,000,000 x 1,000,000 sparse array (which at the moment is integers but most likely will have to be ssize_t) [NOTE: This is running a smaller versions of true dataset. The true dataset will most likely have to allocate an even larger (still sparse) matrix]. Although this array is too large to fit on my 1TB SSD or in memory, I assume that since it is a sparse array that over promising of the kernel would be able to go into effect and allow for this to successfully play out. However, when I do this, I get the same result---the resident memory starts to climb from taking up approximately 80% to now taking 90% before getting killed.
There are two strange things which occur that I cannot understand:
As elucidated in the question I asked prior, my superior is able to run the program on his Macbook pro with 16GB of ram while my linux box with 40GB get killed (i assume for out of memory...).
When I reduce that array size to be 1,000,000 x 100,000, still not able to be represented in memory or on my drive since it only has around 600 GB left, it seems to pass this point but then later segfaults even though my superior is able to run this just fine. By contrast, but superior somehow gets a segfault in an earlier part of my code which works on my computer. The only difference I believe would be
I believe that it has to do something with memory because when I run the adapted code on a smaller file it is able to work.
Attempts at a solution
A friend suggested that I run it with valgrind, which although I was unfamiliar with tried to get it to run. When I did, it seems to work on the smaller dataset (it says that I have some memory leaks but they are all after where the possible memory issue would be.) but it gives me a mmap error on the larger dataset. I looked into this but was not able to resolve it.
Considering Trying
I am starting to suspect that it has to do with the way that memory is over promised in linux in comparison to macos and was considering trying to instal a hackintosh in order to try and get it to run. However, I have never done that before and I think that this would take a bit of time and thus be very unfortunate if I am not able to get it to run.
Relevant Code Segments
Allocating & freeing for mmap:
double* newArrayDouble(ssize_t size, char* backupFile, int *filedestination) {
printf("Trying to create %s with size %zu\n", backupFile, size);
*filedestination = open(backupFile, O_RDWR | O_CREAT | O_TRUNC, 0644);
if (*filedestination < 0) {
perror("open failed");
exit(1);
}
// make sure file is big enough
if (lseek(*filedestination,size*sizeof(double), SEEK_SET) == -1) {
perror("seek to len failed");
exit(1);
}
if (write(*filedestination, "x", 1) == -1) {
perror("write at end failed");
exit(1);
}
if (lseek(*filedestination, 0, SEEK_SET) == -1) {
perror("seek to 0 failed");
exit(1);
}
double *array1 = mmap(NULL, size*sizeof(double), PROT_READ | PROT_WRITE, MAP_SHARED, *filedestination, 0);
if (array1 == MAP_FAILED) {
perror("mmap failed");
exit(1);
}
return array1;
}
void freeArrayDouble(ssize_t size, double* array, int filedestination, char* backupFileName) {
munmap(array, size*sizeof(double));
close(filedestination);
char filename[500];
sprintf(filename, "rm %s", backupFileName);
system(filename);
}
Reading from an MMAPed File:
double* loadArrayDouble(ssize_t size, char* backupFile, int *filedestination) {
*filedestination = open(backupFile, O_RDWR | O_CREAT, 0644);
if (*filedestination < 0) {
perror("open failed");
exit(1);
}
if (lseek(*filedestination,size*sizeof(double), SEEK_SET) == -1) {
perror("seek to len failed");
exit(1);
}
if (lseek(*filedestination, 0, SEEK_SET) == -1) {
perror("seek to 0 failed");
exit(1);
}
double *array1 = mmap(NULL, size*sizeof(double), PROT_READ, MAP_SHARED, *filedestination, 0);
if (array1 == MAP_FAILED) {
perror("mmap failed");
exit(1);
}
return array1;
}
Matrix Allocation:
IdxGroup = (ssize_t **) malloc(Num_Idx*sizeof(ssize_t*));
for (i = 0; i < Num_Idx; i++) {
IdxGroup[i] = (ssize_t *)malloc(Num_Idx*sizeof(ssize_t));
}
^ I changed this to ssize_t because I realized that it will have to be ssize_t later on but it does not work regardless.
NOTE: I believe that over-promising is enabled because I can allocate a really large array (granted its just main and then malloc so I am not sure if the compiler removes it?) and I ran the command in the question asked previously.
Thank you for your help! Please let me know if there is anymore information I should include. It is hard because there is so much code and I don't really completely know why it is broken.