Size of buffered input in C

Viewed 131

I'm writing a simple clone of GNU coreutils in order to better understand the GNU/Linux system and POSIX systems in general. In my "cat" clone, I used fseek() and ftell() to calculate the size of an input buffer size (no support for stdin), and allocate this to a buffer. I would then fread() size bytes from stream fp into the buffer buff, which I would print to stdout. Here's a snippet of code:

#include <stdio.h>
#include <stdlib.h>

int main() {
    FILE* stream = fopen("file.txt", "r");
    fseek(stream, 0L, SEEK_END);
    long sz = ftell(stream);
    rewind(stream);
    char* buf = malloc(sz);
    fread(buf, sz, 1, stream);
    printf("%s", buf);
}

This implementation has two major problems:

  1. If the size in bytes of a stream is greater than the largest value the integer data type you use to hold it (int, size_t) the value will overflow and you will be left with an incorrect number. For the most part, this is solved by using a size_t rather than an int as size_t is 64 bit and can thus handle larger numbers - but ultimately, it can't handle arbitrarily sized files without crashing and burning.
  2. With a very large file (several gigabytes) the memory allocated would be too large and the kernel would send a SIGKILL for the sake of freeing up major amounts of resources (in my case). This would mean the program would be wholly unable to handle very large files and would not be very efficient.

I would rather not use mmap because it adds to the complexity of the program - I'm trying to keep the program as simple as possible. I've decided to buffer the input, ie. fread() x bytes, print them to stdout, free the memory and rinse and repeat until the entire file has been printed.

My question is, how large should x be? Is there a standard for how many bytes should be read? Too large a number and the program becomes incredibly memory-intensive, too little a number and the program would call fread() far too many times, slowing it down as a whole. How many bytes should be read in one go for the best compromise of speed and resource usage?

2 Answers

Two things -- the underlying *nix file system will employ some pretty sophisticated buffering, so not there is no point in being too clever about it in your implementation.

Just allocate a decent size buffer (some multiple of 4K) and keep reusing it by looping round "fread" and "print". You will to capture the number of bytes actually read by fread as the last read will be smaller than the buffer.

When in rome building a coreutils clone, do as the romans GNU coreutils does (or at least have a look at how they do it).

According to this, the block size used by GNU cat is defined by the maximum block size of the input file handle and the output file handle reported by stat()/fstat(), but at minumum 128 kiB (this value, according to comments, is based on benchmarks).

Related