Pearson correlation using scipy, ways to improve speed

Viewed 163

I have a data frame of 1222 rows and 33000 columns which I need to run a correlation between 16000 versus remaining columns within the data frame. Currently, I am using scipy.stats library from python using the Pearson correlation method. Here is the function I am trying:

from scipy.stats import pearsonr


def correlation_analysis(lncRNA_PC_T):
    """Function for correlation analysis"""
    correlations = pd.DataFrame()
    for PC in [column for column in lncRNA_PC_T.columns if "_PC" in column]:
        for lncRNA in [
            column for column in lncRNA_PC_T.columns if "_lncRNAs" in column
        ]:
            correlations = correlations.append(
                pd.Series(
                    pearsonr(lncRNA_PC_T[PC], lncRNA_PC_T[lncRNA]),
                    index=["PCC", "p-value"],
                    name=PC + "_" + lncRNA,
                )
            )

    return correlations

The above code is doing its job, however, for my data frame of size 1222 X 33000 is taking really more than 30 minutes to finish the job. It would be really great if someone could suggest ways to improve the speed of this function for big data frames. Thanks

0 Answers
Related