I am looking at sklearn's TfidfVectorizer, specifically at the preprocessor input parameter which has the following documentation:
"Override the preprocessing (string transformation) stage while preserving the tokenizing and n-grams generation steps."
I am trying to figure out exactly what (if anything) does the preprocessing stage do when I don't override it?
I have an experiment where I am looking at the number of stored elements in the resulting sparse matrix using the following code:
vectorizer = TfidfVectorizer(stop_words=words, preprocessor=process, ngram_range=(1,1), strip_accents='unicode')
vect = vectorizer.fit_transform(twenty_train.data)
items_stored = vect.nnz
- When I do not override the preprocessor, the resulting matrix stores 1278323 elements.
- When I override the preprocessor with an empty method, the resulting matrix stores 1441372 elements.
- When I override the preprocessor with a method including
s = re.sub("[^a-zA-Z]", " ", s), the resulting matrix stores 1331597 elements. - I have not been able to influence the sparse matrix's size (or accuracy when used in classification) with any other processing steps.
Clearly there are differences from the deafult sklearn result, no preprocessing and my attempt at replicating the preprocessing step. I am struggling to find documentation on what specifically the preprocessor does by default.
I have also checked the source code for the TfidfVectorizer - however I wasn't able to figure out what preprocessor was doing from here either.
Does anyone happen to know what code is executed or what preprocessing steps are taken by sklearn's default preprocessor?