I have the following string in a pandas dataframe:
app:fb,msg:something1, somethin2..
app:fdsfe,msg:Service Temporarily Unavailable
app:unknown,size:416746,msg:blabla
I want to parse this columns to have three distinct colum (app, msg, size); for now I do the following:
df["message_soft"] = df['message'].str.extract(".*app:([\.'A-Za-z\s0123456789déèàç^êâôîè-]*),.*",expand=False)
df["message_content"] = df['message'].str.extract('.*msg\:(.*).*', expand=False)
df["message_size"] = df['message'].str.extract('.*size\:(\d{0,10})\,.*', expand=False)
My partern inside 'message', is often:
'app: [msg: [\.'A-Za-z\s0123456789déèàç^êâôîè-] ' --> So in message I can have a ","
but sometime I could have other fields in between:
'app: .., size: /d ,msg: .... ' --> So in message I can have a ","
My solution works but it's way to long (regex hurts !), and to not take into account the partern. Anyidea?