How indexing of sliced data frame works in pandas

Viewed 44

How to get the correct row from a datfarme which is sliced?

To show what I mean, look at this code sample:

import lightgbm as lgb
from sklearn.model_selection import train_test_split
import numpy as np
data=pd.DataFrame()
data['one']=range(0,1000)
data['p1']=data['one']+1
data['p2']=data['one']+2
label=data['p1']%2==0
X_train, X_test, y_train, y_test = train_test_split(data, label, test_size=0.2, random_state=100)
lgb_model = lgb.LGBMClassifier(objective = 'binary')
lgb_fitted = lgb_model.fit(X_train, y_train, verbose = False)
y_prob=lgb_fitted.predict_proba(X_test)
y_prob= pd.DataFrame(y_prob,columns = ['No','Yes'])
model_uncertain=y_prob.loc[(y_prob['Yes'] >= .5) & (y_prob['Yes'] <= .52)]
model_uncertain

enter image description here

My question:

How can I get the row in the X_test dataframe which is related to the first raw in model_uncertain data frame?

To make sure that I am getting the right row, I test it using passing the same row to predict_proba using the following code as I should get the same result:

y_prob_3=lgb_fitted.predict_proba([X_test.iloc[3]])
y_prob_3

enter image description here

But the result is not the same.

I think I am not sending the correct row to predict_proba, as it should return the same value for a row.

What is the correct way to find the n row in model_uncertain and find the corresponding row in X_test data frame?

1 Answers

How can I get the row in the X_test dataframe which is related to the first raw in model_uncertain data frame?

You're on the right track:

>>> idx_of_first_uncertainty_row = model_uncertain.iloc[0].index
>>> row_in_test_data = X.loc[idx_of_first_uncertainty_row]

Yes, indexes are preserved between the original dataframe and its slices (unless you reset the index somewhere in between).

To make sure that I am getting the right row, I test it using passing the same row to predict_proba using the following code as I should get the same result (...) But the result is not the same.

Why do you think they're not the same? In the dataframe image you can't see all of the decimals. A better way to confirm if they're the same (well, really really similar) would be to use something like np.isclose to compare model_uncertain.iloc[0] (first row of dataframe) and X_train.loc[3] (row where index is 3):

>>> np.isclose(model_uncertain.iloc[0].values, X_train.loc[3].values)
Related