Say I have a DataFrame like below
UUID domains
0 asd [foo.com, foo.ca]
1 jkl [foo.ca, foo.fr]
2 xyz [foo.fr]
3 iek [bar.com, bar.org]
4 qkr [bar.org]
5 kij [buzz.net]
How can I turn it in to something like this?
UUID
0 [asd, jkl, xyz]
1 [iek, qkr]
2 [kij]
I want to group all the UUIDs where any domain is present in any other domains column. For example, rows 0 and 1 both contain foo.ca and rows 1 and 2 both contain foo.fr so should be grouped together.
The size of my data set is millions of rows so I can't brute force it.
