How to extract age and gender from reddit post titles?

Viewed 519

I am trying to scrape Reddit posts of subreddits where a lot of questions are in the form:

s1 = "I [22M] and my partner (21F) are foo and bar"

s2 = "My (22m) and my partner (21m) are bar and foo"

I want to make a function that can parse each string and then return age and gender pairs. So:

def parse(s1):
 ....
 return [(22, "male"), (21, "female")]

Essentially, each age/gender tag is a two-digit number followed by either f, F, m, M.

3 Answers

We can try using re.findall here:

s1 = "I [22m] and my partner (21F) are foo and bar"
matches = re.findall(r'(?:[\[(](\d+[MF])[\])])', s1, re.IGNORECASE)
print(matches)

[('22', 'm'), ('21', 'F')]

You could try to extract the matches using this Regex:

(?:[\[\(])(\d{1,2})([MF])(?:[\]\)]) /i

Demo

For the python part of things I would recommend re's findall method:

import re

def parse(title):
    return re.findall(r'(?:\[|\()(\d{1,2})([MF])(?:\]|\))', title, re.IGNORECASE)

title = 'I [22M] and my partner (21F) are foo and bar'
matches = parse(title)

print(matches)

Demo

EDIT:

You could try to modify your Regex to this, in order to fit the new requirement you mentioned in your comment:

(?:[\[\(])(\d{1,2})\s?([MF]|male|female)(?:[\]\)]) /i

Demo

You can use Regex with re :

import re
>>> re.findall(r'(?<=\[|\()[^\)\]]+', s1)  # find text within () or []
['22M', '21F']
>>> re.findall(r'\d+', '22M') # find age
['22']
>>> re.findall(r'[fFmM]+', '22M') # find gender
['M']

This website is really nice to learn and practice on Regex: https://regex101.com/

Related