Pandas - Get unique values from column along with lists of row indices where they appear

Viewed 5034

My dataframe has a string column that can contain long strings. I want to get a list of unique strings, and also a list for each unique string containing row indices where it appears.

I can think of two ways of doing this.

  1. First get the unique list using .unique() and then iterate over the dataframe to build up lists of indices where each unique value shows up
  2. Use .groupBy() to create groups and get the lists of row indices in each group

But I am not quite sure which one is more efficient (or if there are other ways to do this more efficiently). The reason I am thinking about efficiency is that the field I want to uniquify and groupBy is a string field possibly having long strings!

Thanks!

2 Answers
Related