My dataframe has a string column that can contain long strings. I want to get a list of unique strings, and also a list for each unique string containing row indices where it appears.
I can think of two ways of doing this.
- First get the unique list using
.unique()and then iterate over the dataframe to build up lists of indices where each unique value shows up - Use
.groupBy()to create groups and get the lists of row indices in each group
But I am not quite sure which one is more efficient (or if there are other ways to do this more efficiently). The reason I am thinking about efficiency is that the field I want to uniquify and groupBy is a string field possibly having long strings!
Thanks!