From the documentation of DataFrame.assign:
DataFrame.assign(**kwargs)
(...)
Parameters **kwargs : dict of {str: callable or Series}
The column names are keywords. If the values are callable, they are computed on the DataFrame and assigned to the new columns. The callable must not change input DataFrame (though pandas doesn’t check it). If the values are not callable, (e.g. a Series, scalar, or array), they are simply assigned.
This means that in
dataframe = dataframe.assign(gender=lambda ref: add_gender(ref))
ref stands for the calling DataFrame, i.e. dataframe, and thus you are passing the whole dataframe to the function add_gender. However, according to how it's defined, add_gender expects a single row (Series object) to be passed as the argument x, not the whole DataFrame.
if re.search("(womens?)", x.heading, re.IGNORECASE):
In the case of assign, x.heading stands for the whole column heading of dataframe (x), which is a Series object. However, re.search only works with string or bytes-like objects, so the error is raised. While in the case of apply, x.heading corresponds to the field heading of each individual row x of dataframe, which are string values.
To solve this just use assign with apply. Note that the lambda in lambda ref: add_gender(ref) is redundant, it's equivalent to just passing add_gender.
dataframe = dataframe.assign(gender=lambda df: df.apply(add_gender, axis=1))
As a suggestion, here is a more concise way of defining add_gender, using Series.str.extract and Series.fillna.
def add_gender(df):
pat = r'\b(men|women)s?\b'
return df['heading'].str.extract(pat, flags=re.IGNORECASE).fillna('unisex')
Regarding the regex pattern '\b(men|women)s?\b':
\b matches a word boundary
(men|women) matches men or women literally and captures the group
s? matches s zero or one times
Series.str.extract extract the capture group of each string value of the column heading. Non-matches are set to NaN. Then, Series.fillna replaces the NaNs with 'unisex'.
In this case, add_gender expects the whole DataFrame to be passed. With this definition, you can simply do
dataframe = dataframe.assign(gender=add_gender)
Setup:
import pandas as pd
import re
data = {'heading': ['some men', 'some men', 'some women', 'x mens', 'y womens', 'other', 'blahmenblah', 'blahwomenblah']}
dataframe = pd.DataFrame(data=data)
def add_gender(df):
pat = r'\b(men|women)s?\b'
return df['heading'].str.extract(pat, flags=re.IGNORECASE).fillna('unisex')
Output:
>>> dataframe
heading
0 some men
1 some men
2 some women
3 x mens
4 y womens
5 other
6 blahmenblah
7 blahwomenblah
>>> dataframe = dataframe.assign(gender = add_gender)
>>> dataframe
heading gender
0 some men men
1 some men men
2 some women women
3 x mens men
4 y womens women
5 other unisex
6 blahmenblah unisex
7 blahwomenblah unisex