So far I managed to structure the dataset like this:

I want to predict the diagnosis field based on the all other fields with python and fastxml. I assigned integers to all diagnosis terms.
Ex. When I find 'hypertensive' this word I replaced that with 1, when I find 'diabetes' i replaced with 2 and so on...
then the diagnosis column becomes like:

but when I'm trying to predict the diagnosis it gives the same result for all prediction. Here is my snippet code:
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler().fit(X_train)
X_train = scaler.transform(X_train)
X_test = scaler.transform(X_test)
from fastxml import Trainer, Inferencer
from fastxml.weights import propensity
from scipy.sparse import csr_matrix
trainer = Trainer(n_trees=8, n_jobs=1, leaf_classifiers=True)
trainer.fit([csr_matrix(X_train)], lt)
#trainer.save('fastxml_model.h5')
clf = Inferencer('fastxml_model.h5')
predict_mat=np.array([[14,1,78,70,140,80,178,80]])
scaler = StandardScaler().fit_transform(predict_mat)
res=clf.predict(csr_matrix(scaler.astype(np.float32)))
print(res)
Ouput is comming:
(0, 14207) -5.445482
(0, 15876) -0.13862942
(0, 21138) -0.13862942
(0, 33919) -0.13862942
(0, 43647) -0.13862942
(0, 46255) -0.13862942
Which is same for all..
Expected output:
hypertension,diabetes mellitus type 2,chronic kidney disease,parkinsonism
How do I resolve it?