How to detect character set of a file in Angular 8?

Viewed 3004

I am wondering how to detect the character set of a file before I read it in using the FileReader Web API. I need to know what the file charset is before I read it in using fileReader.readAsText(file, "UTF-8") where "UTF-8" for me at the moment is unknown.

Is there any npm packages that I can use with Angular or any manual Vanilla way of detecting a character set without looking at the signatures or using a BOM code (files on my PC saved in either ISO-8859-1 or UTF-8 have same signature and no BOM code).

The packages I have tried to use so far are 'encoding', 'chardet' and 'encoding-japanese'. These don't work with Angular 8 as they are made for use with Node.

Back story: I have a CSV and as soon as it saves in Excel, it saves with the encoding of ISO-8859-1 and I cannot expect all my clients to save their files with a specific encoding (non-technically minded folk). However, other clients may use Notepad++ which will save these files in UTF-8. I need a way of determining the encoding used to stop characters like: "�" appearing.

3 Answers

Unless your input files are really really small, I think you should check out detect-file-encoding-and-language!

I'm using it in my React application to detect the charset of subtitle files before loading them via the FileReader Web API.

This is how I do it:

import languageEncoding from "detect-file-encoding-and-language";

function inputHandler(e) {
  const file = e.target.files[0];
  languageEncoding(file).then(fileInfo => console.log(fileInfo.encoding));  // UTF-8
}

And of course, you'll have to install it:

$ npm i detect-file-encoding-and-language

You might need to use detect-character-encoding which is an external npm module which will do the job for you like this.

const fs = require('fs');
const detectCharacterEncoding = require('detect-character-encoding');

const fileBuffer = fs.readFileSync('file.txt');
const charsetMatch = detectCharacterEncoding(fileBuffer);

console.log(charsetMatch);
// {
//   encoding: 'UTF-8',
//   confidence: 60
// }

You could use the encoding-japanese package in the Angular app. Try the following

  1. Add the package to package.json and run npm install
{
  "dependencies": {
    ...,
    "encoding-japanese": "^1.0.30",
  }
}
  1. Use the package in the app.

Controller

import { Component } from '@angular/core';
import { Observable, Subject } from 'rxjs';

declare const require: any;
export const Encoding = require('encoding-japanese');

@Component({
  selector: 'my-app',
  templateUrl: './app.component.html',
  styleUrls: [ './app.component.css' ]
})
export class AppComponent  {
  encoding: string;

  constructor() { }

  onUpload(event: any) {
    this.detectEncoding(event.currentTarget.files[0]).subscribe(
      encoding => {
        console.log('File encoding is: ' + encoding);
        this.encoding = encoding;
      }
    );
  }

  private detectEncoding(file): Observable<string> {
    let result = new Subject<string>();

    const reader = new FileReader();
    reader.onload = (e) => {
      const codes = new Uint8Array(e.target.result as ArrayBuffer);
      const detectedEncoding = Encoding.detect(codes);
      result.next(detectedEncoding);
    };
    reader.readAsArrayBuffer(file);

    return result.asObservable();
  }
}

Template

<input type="file" (change)="onUpload($event)"/>
<ng-container *ngIf="encoding">
  <p>File encoding is: {{ encoding }}</p>
</ng-container>

The encoding detection mechanism was sourced from the encoding-japanese example here.

  1. You could then verify the encoding inside the subscription
this.detectEncoding(event.currentTarget.files[0]).subscribe(
  encoding => {
    if (encoding === 'UTF8') {
      // encoding is UTF-8
    } else {
      // encoding isn't UTF-8
    }
  }
);
  1. You could check for the following encoding strings.
    • UTF32
    • UTF16
    • UTF16BE
    • UTF16LE
    • BINARY
    • ASCII
    • JIS
    • UTF8
    • EUCJP
    • SJIS
    • UNICODE

Working example: Stackblitz

Related