I have two publicly available word embeddings such as Glove and Google Word2vec. However, in their vocabulary, there are too many misspelling words or garbage words(e.g., ##AA##, adirty, etc). To avoid this words, I would like to extract frequent word(e.g., top 50000 words) since I think relatively high frequent words has normal forms.
So, I wonder if there is a way to find word frequency in above two pretrained word embeddings. If not, I want to know if there are some techniques to exclude this words.