I'm writing a simple clone of GNU coreutils in order to better understand the GNU/Linux system and POSIX systems in general. In my "cat" clone, I used fseek() and ftell() to calculate the size of an input buffer size (no support for stdin), and allocate this to a buffer. I would then fread() size bytes from stream fp into the buffer buff, which I would print to stdout.
Here's a snippet of code:
#include <stdio.h>
#include <stdlib.h>
int main() {
FILE* stream = fopen("file.txt", "r");
fseek(stream, 0L, SEEK_END);
long sz = ftell(stream);
rewind(stream);
char* buf = malloc(sz);
fread(buf, sz, 1, stream);
printf("%s", buf);
}
This implementation has two major problems:
- If the size in bytes of a stream is greater than the largest value the integer data type you use to hold it (
int,size_t) the value will overflow and you will be left with an incorrect number. For the most part, this is solved by using asize_trather than anintassize_tis 64 bit and can thus handle larger numbers - but ultimately, it can't handle arbitrarily sized files without crashing and burning. - With a very large file (several gigabytes) the memory allocated would be too large and the kernel would send a SIGKILL for the sake of freeing up major amounts of resources (in my case). This would mean the program would be wholly unable to handle very large files and would not be very efficient.
I would rather not use mmap because it adds to the complexity of the program - I'm trying to keep the program as simple as possible. I've decided to buffer the input, ie. fread() x bytes, print them to stdout, free the memory and rinse and repeat until the entire file has been printed.
My question is, how large should x be? Is there a standard for how many bytes should be read? Too large a number and the program becomes incredibly memory-intensive, too little a number and the program would call fread() far too many times, slowing it down as a whole. How many bytes should be read in one go for the best compromise of speed and resource usage?