I'm trying to extract items from column gen in a dataframe (sample below). My goal is to iterate through every line in gen into a new dataframe column with items matching the predefined list genre_code.
df = pd.DataFrame({'id': [620, 843, 986], 'tit': ['AAA', 'BBB', 'CCC'], 'gen': [['Romance', 'Satire', 'Fiction'], ['Science Fiction', 'Novel'], ['Mystery', 'Novel']]})
genre_code = ['Science Fiction', 'Mystery', 'Non-fiction']
So far I was able to come up with the following:
new_gen = []
for i in df['gen']:
for j in i:
if j in genre_code:
new_gen.append(j)
else:
new_gen.append('NA')
df['gen'] = new_gen
which does iterate through the column but the lenght of resulting new_gen does not match the original dataframe row length.
/usr/local/lib/python3.7/dist-packages/pandas/core/internals/construction.py in sanitize_index(data, index)
746 if len(data) != len(index):
747 raise ValueError(
--> 748 "Length of values "
749 f"({len(data)}) "
750 "does not match length of index "
ValueError: Length of values (30004) does not match length of index (12841)
I know this must be something very basic, but could someone please point me what I missing?