How do I use Fastxml algorithm in python to predict the value of the output column?

Viewed 262

So far I managed to structure the dataset like this: init dataset

I want to predict the diagnosis field based on the all other fields with python and fastxml. I assigned integers to all diagnosis terms.

Ex. When I find 'hypertensive' this word I replaced that with 1, when I find 'diabetes' i replaced with 2 and so on...

then the diagnosis column becomes like: after categorizing

but when I'm trying to predict the diagnosis it gives the same result for all prediction. Here is my snippet code:

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler().fit(X_train)
X_train = scaler.transform(X_train)
X_test = scaler.transform(X_test)

from fastxml import Trainer, Inferencer
from fastxml.weights import propensity
from scipy.sparse import csr_matrix

trainer = Trainer(n_trees=8, n_jobs=1, leaf_classifiers=True)
trainer.fit([csr_matrix(X_train)], lt)
#trainer.save('fastxml_model.h5')
clf = Inferencer('fastxml_model.h5')
predict_mat=np.array([[14,1,78,70,140,80,178,80]])
scaler = StandardScaler().fit_transform(predict_mat)

res=clf.predict(csr_matrix(scaler.astype(np.float32)))
print(res)

Ouput is comming:

  (0, 14207)    -5.445482
  (0, 15876)    -0.13862942
  (0, 21138)    -0.13862942
  (0, 33919)    -0.13862942
  (0, 43647)    -0.13862942
  (0, 46255)    -0.13862942

Which is same for all..

Expected output:

hypertension,diabetes mellitus type 2,chronic kidney disease,parkinsonism

How do I resolve it?

0 Answers
Related