Second edit: as it turns out, the problem described below occurs when the program ("P1" in the first edit) is compiled with Dev-C++ (Orwell's version 5.11, if this matters) and run on Windows 10 (at least, on the PC that I use in the office, with that OS installed). If the same program is compiled with the MS compiler, it works as expected (also in the office). The program compiled with Dev-C++ works or worked as expected on at least two Windows 7 systems, since a few years.
So, I have this easy empirical solution: use MS compiler when Dev-C++ "fails".
But I still couldn't understand. What is actually happening? Is the code built by Dev-C++ "wrong"? (It would be the first time I observe such a phenomenon). But it used to work... Or is it just incompatible with Windows 10? (Other programs compiled with Dev-C++ are working in Windows 10 exactly as they did in Windows 7). Or did I write a wrong source code, that (in some combinations of compiler and OS) works just out of luck?
This was the first edit:
Hypotesis: write and compile a program (say: "P1.cpp") like this one:
include <stdio.h>
int main(int argc, char *argv[])
{
unsigned char c;
while(1)
{
c=(unsigned char)fgetc(stdin); if(feof(stdin)) break;
if(c<=' ') fputc(c,stdout);
else if(c&128) fputc('N',stdout); // Non-ASCII
else fputc('a',stdout); // ASCII
}
return 0;
}
Now, make an input file (say: "A"):
aÑa
cËc
zzz
(Encoded as codepage 850, DOS-style line endings)
At this point, run CMD.EXE in Windows 10 (if needed, change the codepage to 850) and some commands inside CMD...
D>type A
aÑa
cËc
zzz
D>rffo A
FILE: "A" [cp:850]...
Offs:dec Offs:hex dec hex chr bin U-8
0 0 97 61 'a' 01100001 'a'
1 1 165 A5 'Ñ' 10100101 ?!?
2 2 97 61 'a' 01100001 'a'
3 3 13 D CR 00001101 CR
4 4 10 A LF 00001010 LF
5 5 99 63 'c' 01100011 'c'
6 6 211 D3 'Ë' 11010011 . 2
7 7 99 63 'c' 01100011 'c'
8 8 13 D CR 00001101 CR
9 9 10 A LF 00001010 LF
10 A 122 7A 'z' 01111010 'z'
11 B 122 7A 'z' 01111010 'z'
12 C 122 7A 'z' 01111010 'z'
13 D 13 D CR 00001101 CR
14 E 10 A LF 00001010 LF
D>type A | sort
aÑa
cËc
zzz
D>type A | P1
aNa
aNa
aaa
D>
Note: rffo is a program to dump the contents of the file, used just to check: the two non-ASCII characters are actually a byte 165/A5 ('Ñ' in CP850) and a byte 211/D3 ('Ë'). [The "U-8" column doesn't matter in this case].
The program P1 has written 'a's for the ASCII letters, and 'N's for the non-ASCII letters - as intended.
Now, replace the 'Ë' with a 'Ä' (byte 142/8E), and run the same commands...
D>type A
aÑa
cÄc
zzz
D>rffo A
FILE: "A" [cp:850]...
Offs:dec Offs:hex dec hex chr bin U-8
0 0 97 61 'a' 01100001 'a'
1 1 165 A5 'Ñ' 10100101 ?!?
2 2 97 61 'a' 01100001 'a'
3 3 13 D CR 00001101 CR
4 4 10 A LF 00001010 LF
5 5 99 63 'c' 01100011 'c'
6 6 142 8E 'Ä' 10001110 ?!?
7 7 99 63 'c' 01100011 'c'
8 8 13 D CR 00001101 CR
9 9 10 A LF 00001010 LF
10 A 122 7A 'z' 01111010 'z'
11 B 122 7A 'z' 01111010 'z'
12 C 122 7A 'z' 01111010 'z'
13 D 13 D CR 00001101 CR
14 E 10 A LF 00001010 LF
D>type A | sort
aÑa
cÄc
zzz
D>type A | P1
aNa
aNa^C
D>
That's it: P1 didn't work. After writing the second line (excluding the newline), it started waiting indefinitely, and I had to terminate it with [Ctrl]+[C].
By long and boring tests, the outcome is: there are "bad" symbols, and "safe" symbols. If a "bad" symbol is read, the program will enter an infinite loop at the end of the line containing it (regardless the lenght of the line).
More precisely: when the newline after a "bad" symbol is to be read, this instruction "hangs" (never ends): c=(unsigned char)fgetc(stdin);.
(So, by the way, if "bad" symbols are in the last line, without CR+LF following, P1 does it work without any trouble).
"Safe" symbols don't harm.
In Windows 7, all symbols were "safe", i.e. a program like P1 never failed. In Windows 10, "bad" symbols make P1 fail.
The question is: why?!? What is happening?
Some "safe" symbols are Ñ, ñ, Ï (uppercase only), Ë (uppercase only), ß, ü (lowercase only), and the box-drawing characters with single lines like │ ┤ ┬ etc. Some "bad" symbols are ä, ë, ï, ö, Ä, Ö, Ü
The problem is not in the pipe: type A | sort works correctly.
As pointed out in a comment, stdin is not being read as a binary file. (How can I read it as binary?) (Why does this cause troubles only with some symbols?) (Why didn't it on Windows 7?) (Why....?)
ORIGINAL QUESTION
A common task: sometimes, a set of four files encoded in "codepage 850" should be converted to UTF-8.
Solution: a simple program (named "850U8" with a big effort of fantasy) that reads bytes from stdin, converts the non-ASCII ones, and writes to stdout.
It was intended to be used like this: type input_file | 850U8 > output_file.
This program has always worked properly under Windows 7, since 2016 if the file's date is reliable. Some time ago, unfortunately, the PC was replaced with another, with Windows 10 installed. Then, when I tried the conversion 850→UTF-8, I had this surprise: it works on three files out of four. With the fourth file, it "hangs", like a program in infinite loop or waiting for an input that doesn't arrive.
For a quick & dirty fix, I modified the program "850U8" to let it read directly the input file, and write directly the output file. It works correctly.
So, I could call it "problem solved", but I can't understand what's going on. Does Windows 10 fail to pipe the output of type into 850U8? Why, in case? And why it (whatever is the cause) happens only with one file out of four? I'm puzzled. Any ideas will be higly appreciated.
Some details...
First of all, about when I say "one out of four". Actually, for each "job" (set of outputs) there are four files to be converted: now, the conversion fails regularly with one of those. Also using the old files, that never gave a trouble at the times of Windows 7.
The alternative way to try to run the program reading from stdin: 850U8 < input_file > output_file doesn't work either. The outcome is exactly the same as with typeand pipe.
The same executable, with the same input, works in every possible way on my PC with Windows 7 (at home).
The file on which the conversion fails starts like this:
┌──────┬──┬─────┬─┬──────────────────────────────────────┬┬────┬─────┬────┐
│Xxxxxx│ │ │ │7 77-77-77 XXX:7777││X/Ä:│Xxxxx│ │
├──────┤X │XXxx │ └───────────────────────┬──────────────┘├────┼─────┤ ││││
│Xxxxx │ │ │Xxxx │Xxxxxxxxxxx │ │ │ ││ │ │
├──────┼──┼─────┼─────────────────────────┼───────────────┼────┼─────┼────┤
│ X7777│ │ │XXXXX XXXXXXX │XX77777777X7777│ │ │ │
├──────┼──┼─────┼─────────────────────────┼───────────────┼────┼─────┼────┤
│------│--│-----│-------------------------│---------------│----│-----│----│
├──────┼──┼─────┼─────────────────────────┼───────────────┼────┼─────┼────┤
│ +│ │X7777│XXXX/XXXX.XXXXX XXXXXX │XXXX777777X7777│ │ │ │
├──────┼──┼─────┼─────────────────────────┼───────────────┼────┼─────┼────┤
│ +│ │X7777│XXXX.XXXXX XXXXXX │XXXX777777X7777│ │ │ │
├──────┼──┼─────┼─────────────────────────┼───────────────┼────┼─────┼────┤
It has many box-drawing characters, and some non-ASCII letters like the 'Ä' in the second line (all encoded as codepage 850, of course - probably unlike it is in this page). This same mixture of characters is also in the other three files, that are converted by piping. Each line ends with 36 spaces + CR + LF (for a total of 111 printable characters + CR + LF per line). Also the other three files have trailing spaces + CR + LF. The width is different, anyway.
The file shown is a bit "censored"... the conversion hangs with this censored version exactly at the same point as with the original file (and the similar files).
This point is after the trailing spaces of the second line, when it is going to read the newline (actually only the LF: as it's turning out, Windows' type ouputs only LF when it reads CR+LF).
The source of the program "850U8" (sorry for the Italian inside... and sorry for the long, boring table of data):
// "Transcodifica" un file di testo: da CodePage 850 a UTF-8
// [v2]
#include <stdio.h>
unsigned char t850[]=
{
// ===UTF-8===== // CP.850
// dec, dec, dec // dec hx
0, 000, 195, 135, // 128 80
0, 000, 195, 188, // 129 81
0, 000, 195, 169, // 130 82
0, 000, 195, 162, // 131 83
0, 000, 195, 164, // 132 84
0, 000, 195, 160, // 133 85
0, 000, 195, 165, // 134 86
0, 000, 195, 167, // 135 87
0, 000, 195, 170, // 136 88
0, 000, 195, 171, // 137 89
0, 000, 195, 168, // 138 8A
0, 000, 195, 175, // 139 8B
0, 000, 195, 174, // 140 8C
0, 000, 195, 172, // 141 8D
0, 000, 195, 132, // 142 8E
0, 000, 195, 133, // 143 8F
0, 000, 195, 137, // 144 90
0, 000, 195, 166, // 145 91
0, 000, 195, 134, // 146 92
0, 000, 195, 180, // 147 93
0, 000, 195, 182, // 148 94
0, 000, 195, 178, // 149 95
0, 000, 195, 187, // 150 96
0, 000, 195, 185, // 151 97
0, 000, 195, 191, // 152 98
0, 000, 195, 150, // 153 99
0, 000, 195, 156, // 154 9A
0, 000, 195, 184, // 155 9B
0, 000, 194, 163, // 156 9C
0, 000, 195, 152, // 157 9D
0, 000, 195, 151, // 158 9E
0, 000, 198, 146, // 159 9F
0, 000, 195, 161, // 160 A0
0, 000, 195, 173, // 161 A1
0, 000, 195, 179, // 162 A2
0, 000, 195, 186, // 163 A3
0, 000, 195, 177, // 164 A4
0, 000, 195, 145, // 165 A5
0, 000, 194, 170, // 166 A6
0, 000, 194, 186, // 167 A7
0, 000, 194, 191, // 168 A8
0, 000, 194, 174, // 169 A9
0, 000, 194, 172, // 170 AA
0, 000, 194, 189, // 171 AB
0, 000, 194, 188, // 172 AC
0, 000, 194, 161, // 173 AD
0, 000, 194, 171, // 174 AE
0, 000, 194, 187, // 175 AF
0, 226, 150, 145, // 176 B0
0, 226, 150, 146, // 177 B1
0, 226, 150, 147, // 178 B2
0, 226, 148, 130, // 179 B3
0, 226, 148, 164, // 180 B4
0, 000, 195, 129, // 181 B5
0, 000, 195, 130, // 182 B6
0, 000, 195, 128, // 183 B7
0, 000, 194, 169, // 184 B8
0, 226, 149, 163, // 185 B9
0, 226, 149, 145, // 186 BA
0, 226, 149, 151, // 187 BB
0, 226, 149, 157, // 188 BC
0, 000, 194, 162, // 189 BD
0, 000, 194, 165, // 190 BE
0, 226, 148, 144, // 191 BF
0, 226, 148, 148, // 192 C0
0, 226, 148, 180, // 193 C1
0, 226, 148, 172, // 194 C2
0, 226, 148, 156, // 195 C3
0, 226, 148, 128, // 196 C4
0, 226, 148, 188, // 197 C5
0, 000, 195, 163, // 198 C6
0, 000, 195, 131, // 199 C7
0, 226, 149, 154, // 200 C8
0, 226, 149, 148, // 201 C9
0, 226, 149, 169, // 202 CA
0, 226, 149, 166, // 203 CB
0, 226, 149, 160, // 204 CC
0, 226, 149, 144, // 205 CD
0, 226, 149, 172, // 206 CE
0, 000, 194, 164, // 207 CF
0, 000, 195, 176, // 208 D0
0, 000, 195, 144, // 209 D1
0, 000, 195, 138, // 210 D2
0, 000, 195, 139, // 211 D3
0, 000, 195, 136, // 212 D4
0, 000, 196, 177, // 213 D5
0, 000, 195, 141, // 214 D6
0, 000, 195, 142, // 215 D7
0, 000, 195, 143, // 216 D8
0, 226, 148, 152, // 217 D9
0, 226, 148, 140, // 218 DA
0, 226, 150, 136, // 219 DB
0, 226, 150, 132, // 220 DC
0, 000, 194, 166, // 221 DD
0, 000, 195, 140, // 222 DE
0, 226, 150, 128, // 223 DF
0, 000, 195, 147, // 224 E0
0, 000, 195, 159, // 225 E1
0, 000, 195, 148, // 226 E2
0, 000, 195, 146, // 227 E3
0, 000, 195, 181, // 228 E4
0, 000, 195, 149, // 229 E5
0, 000, 194, 181, // 230 E6
0, 000, 195, 190, // 231 E7
0, 000, 195, 158, // 232 E8
0, 000, 195, 154, // 233 E9
0, 000, 195, 155, // 234 EA
0, 000, 195, 153, // 235 EB
0, 000, 195, 189, // 236 EC
0, 000, 195, 157, // 237 ED
0, 000, 194, 175, // 238 EE
0, 000, 194, 180, // 239 EF
0, 000, 194, 173, // 240 F0
0, 000, 194, 177, // 241 F1
0, 226, 128, 151, // 242 F2
0, 000, 194, 190, // 243 F3
0, 000, 194, 182, // 244 F4
0, 000, 194, 167, // 245 F5
0, 000, 195, 183, // 246 F6
0, 000, 194, 184, // 247 F7
0, 000, 194, 176, // 248 F8
0, 000, 194, 168, // 249 F9
0, 000, 194, 183, // 250 FA
0, 000, 194, 185, // 251 FB
0, 000, 194, 179, // 252 FC
0, 000, 194, 178, // 253 FD
0, 226, 150, 160, // 254 FE
0, 000, 194, 160 // 255 FF
};
int main(int argc, char *argv[])
{
unsigned char c;
int i;
unsigned int xx; //xx
FILE *fi=stdin;
FILE *fo=stdout;
for(i=1; i<argc; i++)
if(argv[i][0]=='/') ; // TO DO: opzioni?
else if(fi==stdin) {fi=fopen(argv[i],"rb"); if(!fi) {fprintf(stderr,"ERRORE nell'apertura del file \"%s\" (rb).\n", argv[i]); return 1;}}
else if(fo==stdout) {fo=fopen(argv[i],"wb"); if(!fo) {fprintf(stderr,"ERRORE nell'apertura del file \"%s\" (wb).\n", argv[i]); return 1;}}
else {fprintf(stderr,"Parametro \"%s\" di troppo...\n"); return 2;}
xx=0; //xx
while(1)
{
fprintf(stderr,"@%4u",xx++); //xx
c=(unsigned char)fgetc(fi); if(feof(fi)) break;
fprintf(stderr," : %3d : %2X\n", (int)c, (int)c);//xx
if(c&128) // Non-ASCII
{
i=((c&127)<<2)|1; // Indice del numero (in seconda colonna sopra, dopo lo "0" - potrebbe essere "000".
if(c=t850[i])fputc(c,fo); // 1° numero: output se non è "000".
fputc(t850[i+1],fo); // 2° numero
fputc(t850[i+2],fo); // 3° numero
}
else fputc(c,fo); // ASCII
}
if(fi!=stdin) fclose(fi);
if(fo!=stdout) fclose(fo);
return 0;
}
It is "instrumented" with some extra lines (marked with the comment //xx) to see where it stops working: those lines are normally not there.
Also, I have the doubt: maybe it's not Windows 10's (exclusive) fault? Maybe some other thing is interfering? There is a Kaspersky running (and heavily slowing down things), but the same (probably an earlier version) was running on the Window 7 machine.
Of course I have all the authorizations to read and write the files (indeed, I do when the converter program is run reading directly the input).