Shuffling rows in a Pandas DataFrame while retaining the index

Viewed 277

I am currently trying to find a way to randomize items in a dataframe row-wise. I want to preserve the column names as well as the index. I just want to change the order of entries in my dataframe.

Currently, I was using

data = data.sample(frac=1).reset_index(drop=True)

However, this is causing some issues in terms of output. I don't think the rows are being shuffled properly. Is there another way to achieve that?

The issue is that I am doing text analysis and when I am looking at the most correlated unigrams and bigrams with each class, I am getting different answers for shuffled and original data.

This is the code I am using for monograms and bigrams

tfidf = TfidfVectorizer(sublinear_tf=True, 
                    min_df=5, 
                    stop_words=STOPWORDS, 
                    norm = 'l2', 
                    encoding='latin-1', 
                    ngram_range=(1, 2))

feat = tfidf.fit_transform(data['Combine']).toarray()

N = 5    # Number of examples to be listed
for f, i in sorted(category_labels.items()):
    chi2_feat = chi2(feat, labels == i)
    indices = np.argsort(chi2_feat[0])
    feat_names = np.array(tfidf.get_feature_names())[indices]
    unigrams = [w for w in feat_names if len(w.split(' ')) == 1]
    bigrams = [w for w in feat_names if len(w.split(' ')) == 2]
    print("\nFlair '{}':".format(f))
    print("Most correlated unigrams:\n\t. {}".format('\n\t. '.join(unigrams[-N:])))
    print("Most correlated bigrams:\n\t. {}".format('\n\t. '.join(bigrams[-N:])))
1 Answers

Just using data = data.sample(frac=1) samples the index as well and that is problematic. You can see the output below. We just need to change the values.

enter image description here

The correct method to achieve this is by just sampling the values. I just figured it out. We can do it this way. Thank you everybody who tried to help.

data[:] = data.sample(frac=1).values

I was getting the correct output from this.

Related