I have a school project which consists of identifying each language of a tweet from a dataset of tweets. The dataset contains tweets in Spanish, Portuguese, English, Basque, Galician and Catalan. The task is to implement a language identification model using unigrams, bigrams and trigrams and to analyze the efficiency of each model.
I understand the concepts of ngrams and I understand that the languages are somewhat similar (hence it's not that trivial of a task), but what I don't understand is that I'm getting better results for unigrams than bigrams and I'm getting better results for bigrams than trigrams.
I can't comprehend how is that possible since I expected a better efficiency for bigrams and trigrams.
Could you help me shed some light on why is this happening?
Thank you for your time.