I'm using PySpark 2.4.
I have a dataframe like below as input:
ceci_p| ceci_l|ceci_stok|
-------+-------+---------+
SFIL401| BPI202| BPI202|
BPI202| CDC111| BPI202|
LBP347|SFIL402| SFIL402|
LBP347|SFIL402| LBP347|
-------+-------+---------+
I want to detect which ceci_stok values exist in both ceci_l and ceci_p columns using a join (maybe a self join).
For example: ceci_stok = BPI202 exists in both ceci_l and ceci_p.
I want to create a new dataframe as a result that contains ceci_stok which exist in both ceci_l and ceci_p.