for improving my model I use to give character based 3- Gram instead of word :) the code snippet is in below:
def MakeNGram(sent_list, N, vocab_size, seq_size):
NGramList = []
for sent in sent_list:
# ---------------------------- حذف فاصله ----------------------------
sent = sent.replace(" ", "")
# ------------------------- استخراج ان تایی ها --------------------------
NGram = [sent[i:i + N] for i in range(len(sent) - N + 1)]
# ----------------------- تبدیل به ان تایی با فاصله ------------------------
new_string = " ".join(NGram)
# ------------------------- رمزگذاری وان هات --------------------------
OneHot = one_hot(new_string, round(vocab_size * 1.3))
# --------------------------- padding ------------------------------
HotLen = len(OneHot)
if HotLen >= seq_size:
OneHot = OneHot[0:seq_size]
else:
diff = seq_size - HotLen
extra = [0] * diff
OneHot = OneHot + extra
NGramList.append(OneHot)
NGramArray = np.array(NGramList)
return NGramArray
up to here there is no problem, but I want vectorize N-grams without onehot, and ofcourse there is no Ngram2vec model for my language(persian), so please help me to change the code with best embedding function for Ngrams :)
note : I use the keras embeddings but there were many problems, I thing I had mistake in replacing keras embedding with onehot...
what is the best way that is suitable for vriable languages?