Let's say if I have a Pandas df called df_1 where one of the rows looks like this:
| id | rank_url_agg | url_list |
|---|---|---|
| 2223 | ['gtech.com','gm.com', 'ford.com'] | ['google.com','gtech.com','autoblog.com','gm.com', 'ford.com'] |
I want to create a new column called url_list_agg which does the following things for each row:
- Iterate through the URLs in
url_list - If URL doesn't exist in
rank_url_aggin the same row, assign a value of 0. - If URL exists in
rank_url_agg, then assign the value that corresponds to the difference between the length of therank_url_agglist and the index of that URL inrank_url_agg. - Once done iterating through all URLs in
url_list, wrap the results into a list.
So at the end, the first row in the new url_list_agg column will become [0,3,0,2,1].
I've tried running the following script (only to test the 1st row and not entire dataframe):
for item in agg_report['url_list'][0]:
if item in agg_report['rank_url_agg'][0]:
item=len(rank_url_agg[0]) - agg_report['rank_url_agg'][0].index(item)
else:
item=0
But when I checked agg_report['url_list'][0], it still returned just this list: ['google.com','gtech.com','autoblog.com','gm.com', 'ford.com']. So my code didn't work.
Any advice on how to achieve this goal for every row in the dataframe will be greatly appreciated!