Regex and Pandas Dataframe error: 'NoneType' object has no attribute 'group'

Viewed 59

there!

I have a dataframe and I want to extract cellphone numbers.

VARIAVEL
Telefone:(11) 95262-7297
Celular:(31) 97250-8639
Não possui

Below is the code which I'm using to do it:

df['TELEFONE'] = df['VARIAVEL'].apply(lambda x: re.search('\(\d\d\)\s\d\d\d\d\d-\d\d\d\d', x).group(0) if pd.notnull(x) else x )

The erros which is returning is:

AttributeError: 'NoneType' object has no attribute 'group'

I know the error is because 'Não possui' returns no match, and due to it I can't use 'group(0)'.

How can I fix it? How can I apply regex and group(0) just on matched cases?

1 Answers

You can use

df['TELEFONE'] = df['VARIAVEL'].str.extract(r'(\(\d{2}\)\s\d{5}-\d{4})', expand=False)

See a Pandas test:

import pandas as pd
df = pd.DataFrame({'VARIAVEL':['(11) 95262-7297', '(31) 97250-8639', 'Não possui']})
df['TELEFONE'] = df['VARIAVEL'].str.extract(r'(\(\d{2}\)\s\d{5}-\d{4})', expand=False)

Output:

>>> df
          VARIAVEL         TELEFONE
0  (11) 95262-7297  (11) 95262-7297
1  (31) 97250-8639  (31) 97250-8639
2       Não possui              NaN 

If you need to extract a multiple matches, you would need a Series.str.findall:

df['TELEFONE'] = df['VARIAVEL'].str.findall(r'\(\d{2}\)\s\d{5}-\d{4}').str.join(', ')
Related