Defining a window size in a regular expression - Python

Viewed 77

We are trying to locate two words within a limited window size in a text file, e.g., 5-10 words, using regular expression. We have imported the text using PyPDF2 and then tokenized. We are able to locate single words in the text (see code below), however, we want to see if e.g. "GHG | CO2 | carbon" and "tonnes | tonne | ton | tons" appear together in the text within a window size of 5 or 10 words. Is there a function in nltk we have to import? And do you guys have a suggestion for how we can reformulate the code below to check if the words above appear within a window?

match = re.compile(r"climate", flags=re.I | re.X)
match.findall(text)

We are new to doing textual analysis in Python, so all help is highly appreciated!

1 Answers

You could use the ngram function from NLTK to create the 5-10 words window. Use word_tokenize follow by a filter for alphanumeric content to cleanup the text. Then, use a for loop to apply the appropriate ngram window size of 5 or 10 (or, for multiple window sizes within the range, remove the for loop and use everygrams(words, min_len=5, max_len=10), if applicable).

With the list of ngrams inside the required window, you can now check if the expected terms appear within the ngram window using the regex \s(GHG|CO2|carbon)\ston[nes]*. As expected, some overlap occurs in the final result, not only per the creation of neighboring sequences of items by the ngram but also because the funcion is executed for two different window sizes.

import nltk
nltk.download('punkt')
from nltk.tokenize import word_tokenize
from nltk.util import ngrams
import re

# original text: https://ghgprotocol.org/sites/default/files/Wood_Products.pdf (page 2)
text = """These tools estimate CO2 tonnes from fossil fuel combustion based on the carbon content of the fuel (or a comparable emission factor) and the amount burned. Carbon dioxide
emissions from biomass combustion are not counted as GHG tons. Companies that wish to comply with the WRI/WBCSD GHG ton Protocol should include these biomass combustion CO2 tons, and they should be
reported separately from direct GHG emissions. Regardless of the reporting approach chosen, it is important to clearly separate estimates of CO2 emissions from fossil fuel
combustion from those of CO2 tonne from biomass combustion. Methods for estimating CO2 emissions originating from resins contained in wood residual fuels are also provided.
In all cases, however, companies may use site-specific information where it yields more accurate estimates of GHG ton than the tools outlined in this report.
Annex H contains tables populated with the recommended GHG emission factors discussed throughout the body of the report."""

tokens = word_tokenize(text)
words = [word for word in tokens if word.isalnum()]

def get_ngrams(words, n):
    n_grams = ngrams(words, n)
    return [' '.join(grams) for grams in n_grams]

ngram_list = []
for ngram in [5,10]:
    ngram_list.extend(get_ngrams(words, ngram))

for idx, gram in enumerate(ngram_list):
    r = re.search(r'\s(GHG|CO2|carbon)\ston[nes]*', gram)
    if r:
        print(idx, gram)

Output

0 These tools estimate CO2 tonnes
1 tools estimate CO2 tonnes from
2 estimate CO2 tonnes from fossil
33 not counted as GHG tons
34 counted as GHG tons Companies
35 as GHG tons Companies that
...
...
267 more accurate estimates of GHG ton than the tools outlined
268 accurate estimates of GHG ton than the tools outlined in
269 estimates of GHG ton than the tools outlined in this
270 of GHG ton than the tools outlined in this report

Related