How does git detect directory changes so fast?

Viewed 179

I've read a number of articles and answers here, but they weren't helpful.

I know that git uses mtime and ctime do detect that file was changed without reading it, that makes sense, but:

  • Running lstat on each file in my repo takes 79 seconds, but git does that in less than a second
  • How does it detect added or removed files without scanning the whole directory tree?

I tried looking into sources of diff-index but they seem to be quite complicated.

Please note that it's not a duplicate of How does git detect that a file has been modified?. I get that git uses mtime and ctime. I wonder how git can get them so fast. Or may be git doesn't compute them each time you run git diff? That's the point of this question.

1 Answers

Long answer short:

strace -fostrace.log git diff-index --quiet @
vi strace.log

At least when there's a lot to do it fires off a big-batch-o'-threads issuing stat's in parallel so the filesystem's got a lot of pending requests and has the opportunity to prioritize for throughput.

Also:

git still has to traverse the whole directory tree

no, it doesn't. tttt, that's the reason the index is called "the cache". All the names (and last-it-looked data) it cares about it reads in in one big fat read right up front, .git/index is 5MB for a full linux checkout, that's going to be like two seeks, very few ms even the first time on a hdd, when that means hunting it up and siphoning its wiggly bits off a platter.

Related