Training a model with List of points column

Viewed 86

I want to classify cracks by their depths. To do it, I store in a data frame the following features:

WindowsDf = pd.DataFrame(dataForWindowsDf, columns=['IsCrack', 'CheckTypeEncode', 'DepthCrack',
                                                    'WindowOfInterest'])
#dataForWindowsDf is a list which iteratively built from csv files.
#Windows data frame taking this list and build a data frame from it.

So, my target column is 'DepthCrack' and the other columns are part of feature vector. WindowOfInterest is a column of 2d list - list of points - a graph that represents a test that is done (based on electro-magnetic waves returned from a surface as a function of time) :

[[0.9561600000000001, 0.10913097635410397], [0.95621,0.1100000]...]

The problem I faced is how to train a model - using a column of 2d list(I tried to push that as it is and it didn't work)? What way do you suggest to deal with this problem?

I thought about extracting features from the 2d-list - to get one dimensional features(integral and etc.)

1 Answers

You might transform this one feature in two, like WindowOfInterest can become :

WindowOfInterest_x1 and WindowOfInterest_x2

For example from your DataFrame :

>>> import pandas as pd

>>> df = pd.DataFrame({'IsCrack': [1, 1, 1, 1, 1], 
...                    'CheckTypeEncode': [0, 1, 0, 0, 0], 
...                    'DepthCrack': [0.4, 0.2, 1.4, 0.7, 0.1], 
...                    'WindowOfInterest': [[0.9561600000000001, 0.10913097635410397], [0.95621,0.1100000], [0.459561, 0.635410397], [0.4495621,0.32], [0.621,0.2432]]}, 
...                   index = [0, 1, 2, 3, 4])
>>> df
    IsCrack CheckTypeEncode DepthCrack  WindowOfInterest
0   1       0               0.4         [0.9561600000000001, 0.10913097635410397]
1   1       1               0.2         [0.95621, 0.11]
2   1       0               1.4         [0.459561, 0.635410397]
3   1       0               0.7         [0.4495621, 0.32]
4   1       0               0.1         [0.621, 0.2432]

We can split the list like so :

>>> df[['WindowOfInterest_x1','WindowOfInterest_x2']] = pd.DataFrame(df['WindowOfInterest'].tolist(), index=df.index)
>>> df

        IsCrack  CheckTypeEncode    DepthCrack          WindowOfInterest                           WindowOfInterest_x1  WindowOfInterest_x2
0       1        0                  0.4                 [0.9561600000000001, 0.10913097635410397]  0.956160             0.109131
1       1        1                  0.2                 [0.95621, 0.11]                            0.956210             0.110000
2       1        0                  1.4                 [0.459561, 0.635410397]                    0.459561             0.635410
3       1        0                  0.7                 [0.4495621, 0.32]                          0.449562             0.320000
4       1        0                  0.1                 [0.621, 0.2432]                            0.621000             0.243200

To finish, we can drop the WindowOfInterest column :

>>> df = df.drop(['WindowOfInterest'], axis=1)
>>> df
    IsCrack CheckTypeEncode DepthCrack  WindowOfInterest_x1 WindowOfInterest_x2
0   1       0               0.4         0.956160            0.109131
1   1       1               0.2         0.956210            0.110000
2   1       0               1.4         0.459561            0.635410
3   1       0               0.7         0.449562            0.320000
4   1       0               0.1         0.621000            0.243200

Now you can pass WindowOfInterest_x1 and WindowOfInterest_x2 as features for you model.

Related