how to embed N-grams

Viewed 120

for improving my model I use to give character based 3- Gram instead of word :) the code snippet is in below:


def MakeNGram(sent_list, N, vocab_size, seq_size):
    NGramList = []

    for sent in sent_list:
        # ---------------------------- حذف فاصله ----------------------------
        sent = sent.replace(" ", "")
        # ------------------------- استخراج ان تایی ها --------------------------
        NGram = [sent[i:i + N] for i in range(len(sent) - N + 1)]
        # ----------------------- تبدیل به ان تایی با فاصله ------------------------
        new_string = " ".join(NGram)
        # ------------------------- رمزگذاری وان هات --------------------------
        OneHot = one_hot(new_string, round(vocab_size * 1.3))
        # --------------------------- padding ------------------------------
        HotLen = len(OneHot)
        if HotLen >= seq_size:
            OneHot = OneHot[0:seq_size]
        else:
            diff = seq_size - HotLen
            extra = [0] * diff
            OneHot = OneHot + extra
        NGramList.append(OneHot)
    NGramArray = np.array(NGramList)
    return NGramArray


up to here there is no problem, but I want vectorize N-grams without onehot, and ofcourse there is no Ngram2vec model for my language(persian), so please help me to change the code with best embedding function for Ngrams :)
note : I use the keras embeddings but there were many problems, I thing I had mistake in replacing keras embedding with onehot...
what is the best way that is suitable for vriable languages?

0 Answers
Related