DataFrameGroupBy select subset using multiindex

Viewed 131

2 exactly similar structured dataframes. I group them each by columns A and B.

dfgrouby1=df1.groupby(['A','B'])
dfgrouby2=df2.groupby(['A','B'])

I iterate through dfgrouby1 by subgroup (which are each dataframes), and want to get subgroups (dataframes) with the same indices (iA,iB) from dfgrouby2.

2 Questions

  1. how to retrieve the corresponding subgroup in dfgrouby2;
  2. how to catch if the (iA,iB) index doesn't exist in dfgrouby2.

The loop works fine, and documentation shows dataframes with multiindices use .loc[(index tuple)], but apparently not DataFrameGroupBy objects.

Searched extensively. Maybe not using the correct descriptors.

for (iA,iB),eachgroup1 in dfgrouby1:
    eachgroup2 =dfgrouby2.loc[(iA,iB)]
    #do things with eachgroup1['C':'Q'] vs. eachgroup2['C':'Q'] 

AttributeError: 'DataFrameGroupBy' object has no attribute 'loc'

Also tried:

    eachgroup2 =dfgrouby2[[iA,iB]]
KeyError: "Columns not found: 204, 34"
OR
    eachgroup2 =dfgrouby2[(iA,iB)]
KeyError: "Columns not found: 204, 34"

note: 204, 34 are the first values of iA,iB

1 Answers

get_group is the statement I couldn't find. This pulls the corresponding group from the 2nd groupby. And a simple try/except will suffice.

for (iA,iB),eachgroup1 in dfgrouby1:
     try:
          eachgroup2 =dfgrouby2.get_group(iA,iB)
          #comparison code for eachgroup1 and eachgroup2
     except:
          #missing statement/or exception code
     
Related