My question is similar to this but i'm not sure how to modify the last part for elements in a list only.
I want to split a dataframe into smaller dataframes based on how the column name starts.
For example, column names are in the format of:
df = pd.DataFrame(np.random.randint(0,100,size=(10, 4)))
df.columns = ['P1_ATGC', 'P1_GCTA', 'P2_AACT', 'P2_CGAT']
df
P1_ATGC P1_GCTA P2_AACT P2_CGAT
0 78 86 47 78
1 22 48 22 43
2 91 12 45 10
3 83 85 9 20
4 82 26 25 71
5 13 36 53 19
6 93 15 30 28
7 24 13 55 23
8 10 49 98 45
9 85 35 77 89
and want to end up with separate df's for each PX. For example something like:
df[0]
P1_ATGC P1_GCTA
0 78 86
1 22 48
2 91 12
3 83 85
4 82 26
5 13 36
6 93 15
7 24 13
8 10 49
9 85 35
df[1]
P2_AACT P2_CGAT
0 47 78
1 22 43
2 45 10
3 9 20
4 25 71
5 53 19
6 30 28
7 55 23
8 98 45
9 77 89
i'm able to get the unique PXs with: np.unique([x.split('_')[0] for x in df.columns])
it returns:
array(['P1', 'P2'], dtype='<U2')
But how do I split the dataframes by columns ones based on the PX it belongs to?