I am using the following code:
pipeline = Pipeline([('vect',
TfidfVectorizer( ngram_range=(1,2),
stop_words="english",
sublinear_tf=True ,
use_idf=True,
norm='l2' )),
('reduce_dim',
SelectPercentile(f_classif, 90)),
('clf',
SVC(kernel='linear',C=1.0,
probability=True, max_iter=70000,
class_weight='balanced'))])
model = pipeline.fit(X_train,y_train)
model.predict(X_test)
x=vectorizer.fit_transform(X_train_text)
y=vectorizer.transform(X_test_text)
As per my understanding, pipeline.fit() fits tfidf to the train data and when model.predict() is called on X_test, it only does a tfidf transformation based on the fitted train data.
Since tf idf works by getting frequency of words in the document and corpus, I am wondering what happens underneath in the .fit_transform and .transform functions.