I'm extending the standard TfidfVectorizer to add an additional mandatory preprocessing step. Looking at the source code for build_preprocessor, their code has a "branching" logic, but I want an additional function tacked on top of whichever function gets returned by the super call. Here is what I have so far:
from functools import partial
from sklearn.feature_extraction.text import TfidfVectorizer
class MyTfidfVectorizer(TfidfVectorizer):
def __init__(self, some_file):
super().__init__()
f = open(some_file)
self.suffixes = f.readlines()
def build_preprocessor(self):
existing_preprocessor = super(MyTfidfVectorizer, self).build_preprocessor()
def additional_preprocessor(doc, suffixes, threshold=2):
# Call existing_preprocessor here?
processed_doc=doc
return processed_doc
return partial(additional_preprocessor, suffixes=self.suffixes)
I'm not sure how to chain the call hierarchy so when the preprocessing happens first the parent class' preprocessor is called and then the derived class' preprocessor gets called?