Decision tree: how to manually setting a decision node

Viewed 599

I am trying to build a decision tree from scratch. This decision tree should take one email from a list and, based on some factors, assign a score. For example: let's say that the email that I want to test is

email="I transferred £20,000 into our pension plan, but the money never arrived. I had asked my financial adviser at Brewin Dolphin for the relevant bank details and he sent them by email."

Root node: I would like to consider as root node the variable "contain_number".If the email contains a number than assign -0.1 otherwise 0. These would be the two split. I would like to consider only binary split.

Child node: if the email contains a word in the list ['financial','bank', 'plan', 'money'], then assign -0.1 otherwise 0.

and so on. I know that there is a built-in classifier in Python:

from sklearn.tree import DecisionTreeClassifier # Import Decision Tree Classifier
from sklearn.model_selection import train_test_split # Import train_test_split function
from sklearn import metrics #Import scikit-learn metrics module for accuracy calculation
#split dataset in features and target variable
feature_cols = ['email', 'date', 'author', 'server']
X = pima[feature_cols] # Features
y = pima.label # Target variable

# Split dataset into training set and test set
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=1) # 70% training and 30% test

clf = DecisionTreeClassifier()

# Train Decision Tree Classifer
clf = clf.fit(X_train,y_train)

#Predict the response for test dataset
y_pred = clf.predict(X_test)

However I was wondering how to build a decisione tree based on some criteria. Thank you for your time.

For clarifying what I would like to do I have several columns like the following:

Sender         Server       Subject       Corpus       Date         % of xsuks@libero.it  libero    DEAL!!!     I want to propose a deal 29/10/2020   punctuation          % of Upper case
60                         40

I would like to build a decision tree which takes them into account based on some 'weight' that I give you to each variable. For example:

  • if the email does not contain any word in a spam_list, then assign 0;
  • if the email contains words in a spam list, then assign -0.1 per each word included in the spam list;
  • if Sender includes misspelled bad words, then assign -0.1;
  • if Sender does not include misspelled bad words, then assign 0;
  • if the percentage of punctuation is > 50%, then assign -0.1; otherwise 0; ...

I have also a labelled column which includes labels that I assigned manually for automatically build a Decision Tree in Python (using the code I shown above); however I do not know how to include all the fields of my interest because some of them are categorical (like senders, subject and corpus; others are numerical, like %). At the end, I would like to compare the results that I get from my own decision tree for scoring to that got from a classifier in Python. For checking the accuracy, I would use confusion matrix or ROC.

This is pretty much what I would like to achieve.

0 Answers
Related