Same program is 10 times slower on windows

Viewed 159

I wrote a simple c program that copies 10 million bytes from a file and pastes them in reverse order on another file (this is done one byte at a time, I know it's not efficient but it's just to make some tests), I don't understand why on linux it takes 2.5 seconds while on windows it takes more than 20 seconds. I run the same program changing only the paths. I use windows 10 and archlinux, the files are on an ntfs partition.

code on windows

#include <stdio.h>
#include <time.h>

void get_nth_byte(FILE *fp, int nth_index,unsigned char* output){
    fseek(fp,nth_index,SEEK_SET);
    fread(output, sizeof(unsigned char), 1,fp);
}

int main() {
    clock_t begin = clock();
    //
    FILE* input = fopen( "C:\\Users\\piero\\Desktop\\input.txt","rb");
    FILE* output = fopen("C:\\Users\\piero\\Desktop\\output.txt","wb");
    unsigned char byte;
    for (int i = 10000000; i > 0; i--) {
        get_nth_byte(input,i,&byte);
        fwrite(&byte, sizeof(unsigned char),1,output);
    }
    //
    clock_t end = clock();
    double result = (double) (end - begin)/CLOCKS_PER_SEC;
    printf("%f",result);
    return 0;
}

code on linux

#include <stdio.h>
#include <time.h>

void get_nth_byte(FILE *fp, int nth_index,unsigned char* output){
    fseek(fp,nth_index,SEEK_SET);
    fread(output, sizeof(unsigned char), 1,fp);
}

int main() {
    clock_t begin = clock();
    //
    FILE* input = fopen( "/run/media/piero/Windows/Users/piero/Desktop/input.txt","rb");
    FILE* output = fopen("/run/media/piero/Windows/Users/piero/Desktop/output.txt","wb");
    unsigned char byte;
    for (int i = 10000000; i > 0; i--) {
        get_nth_byte(input,i,&byte);
        fwrite(&byte, sizeof(unsigned char),1,output);
    }
    //
    clock_t end = clock();
    double result = (double) (end - begin)/CLOCKS_PER_SEC;
    printf("%f",result);
    return 0;
}

output on linux : 2.224549

output on windows : 25.349647

UPDATE

I solved the problem by using cygwin rather than mingwin, now it takes about 4.3 seconds

2 Answers

This is a great demonstration of how it's not the code we write that runs, it's the executable that the compiler makes from the code that runs.

It is possible that your Windows C compiler is not as advanced as your Linux C compiler, and is not optimizing your code as well as it could, or it's possible that the libraries that the Windows compiler is linking to for fread() and fwrite() are slower than the equivalent libraries in the Linux system.

If I had to put up my best guess, the Linux C compiler probably noticed that it would be more efficient to read more than one byte at a time, and it could do that without affecting the semantics of your program, and the Windows compiler either didn't infer the same, or wasn't able to optimize in the same way due to some underlying proprietary filesystem thing that only Microsoft engineers understand.

I can't say for sure without a peek at the disassembled binaries

One of the strengths of Unix/Linux is that files are designed to be treated as streams of bytes, with it being maximally easy and efficient to seek to the n'th byte using fseek or lseek.

Non-Unix operating systems, such as Windows, tend to have to work much harder to implement those seek operations. In the worst case, they may actually need to read through the file, counting characters as they go.

Your code opens both files in binary mode, and this should reduce the need for the fseek implementation to perform any expensive emulations. In text mode, a 10x performance penalty for heavy fseek use wouldn't surprise me. I'm much more surprised you're seeing it in binary mode.

[Disclaimer: strictly speaking, in text mode fseek is not defined as seeking to an arbitrary byte offset at all, but rather, only to a position defined by the number returned by a previous call to ftell. If an implementation takes advantage of that freedom, it can reduce the performance penalty for text-mode fseek operations, also, but it then means that code like yours, that constructs positions to seek to on the assumption that they're pure byte offsets, may not work at all.]

Related