I'm fighting what is probably an encoding issue, just can't find it. The GitHub API is giving me three bytes instead of one for a file containing only 0xC4. Illustration:
Creating the file:
~/github-binary-api-problem(master*) » echo -n -e '\xc4' > c4-createdfromfilesystem
~/github-binary-api-problem(master*) » hexdump c4-createdfromfilesystem
0000000 c4
0000001
I committed that file to GitHub as usual - go take a look - and GitHub thinks it's a single byte:
So far so good. Now I try to download it, using the Contents API (GET /repos/{owner}/{repo}/contents/{path}):
~/github-binary-api-problem(master*) » curl \
-H "Accept: application/vnd.github.v3.raw" \
https://api.github.com/repos/Undo1/github-binary-api-problem/contents/c4-createdfromfilesystem | hexdump
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 3 100 3 0 0 8 0 --:--:-- --:--:-- --:--:-- 8
0000000 ef bf bd
0000003
~/github-binary-api-problem(master*) »
And I get three bytes back! This example is in a macOS environment, but I first saw it on Windows. I'm sure it's an encoding issue somewhere in the stack, but I can't find it. What do I need to do to fetch an accurate representation of a binary file from the GitHub API?
Update - I've found that 0xef 0xbf 0xbd is the UTF-8 replacement character, so I'm guessing GitHub's API is trying to UTF-8 encode the file before sending it, even though raw is specified. I've sent GitHub a support ticket.
