We have a browser extension that must add javascript to the page body (outside the extension context) via tags in order for webcomponent functionality to work properly. The source files are encoded in UTF-8 (and must be), with no BOM, so on pages where the doctype is not UTF-8, the javascript is not interpreted correctly, leading to errors. (An example is http://info.cern.ch, which has its charset detected as "windows-1252" in Chrome.) I've seen that adding a BOM can cause issues, so we haven't explored that avenue.
We can fix the issue by adding charset="UTF-8" on the script tags. However, MDN says this attribute is deprecated and not necessary:
It’s unnecessary to specify the charset attribute, because documents must use UTF-8, and the script element inherits its character encoding from the document.
How can it be "unnecessary" if it clearly causes an issue for us on some pages? Is there standards information somewhere upon which MDN is basing their claim that it is deprecated?
WHATWG doesn't mention this parameter being deprecated.
They do mention that the parameter is ignored if you set type="module" on the script tag. If we do this (without setting charset), it works fine. Does it just assume UTF-8 encoding in this case? Does it autodetect by some other means?
Clearly most pages don't have this issue, but our extension gets used on all kinds of outdated pages, so it is important that we can trust that our javascript will load correctly regardless of the document encoding. It seems like we can probably just use type="module", but after digging through this I have concerns because there seems to be a general lack of good information about this topic out there.