I need to calculate my own tf-idf instead of using the TfidfVectorizer built into scikit-learn, but I want my output to be in the same format as I would get using scikit-learn's TfidfVectorizer when I call the fit_transform() function.
My original code was as follows, and it gave me the ouput style that I wanted.
doc1 = "These are some words I'm putting in a document."
doc2 = "This document is comprised of a number of words."
doc3 = "Some of the words in this document are found in other documents also."
doc4 = "We are now writing a piece of text which is entirely separate from the others and, hence, dissimilar."
doc5 = "Another piece of writing in which we are interested for its dissimilarity to its precursors is this."
docs = [doc1, doc2, doc3, doc4, doc5]
def tfidf(documents, reduced_documents=None):
vectorizer = TfidfVectorizer(analyzer='word')
vectors = vectorizer.fit_transform(documents)
return vectors
print(tfidf(docs)
The output of this is a sparse matrix, which I'm not entirely sure how to recreate. It looks like this:
(0, 7) 0.322444504927649
(0, 14) 0.322444504927649
(0, 25) 0.481467662591444
(0, 35) 0.322444504927649
...
Because I want to only calculate the tf-idf score for some words in each document I've had to rewrite the tfidf() function as follows:
reduced_doc1 = "These I'm in a"
reduced_doc2 = "This of a of"
reduced_doc3 = "Some of the in this in"
reduced_doc4 = "We a of which from the and"
reduced_doc5 = "Another of in which we for its to its this"
reduced_docs = [reduced_doc1, reduced_doc2, reduced_doc3, reduced_doc4, reduced_doc5]
def tfidf(documents, reduced_documents=None):
documents = [i.split(" ") for i in documents]
reduced_documents = [i.split(" ") for i in reduced_documents]
vectors = []
if not reduced_documents:
reduced_documents = documents
N = len(documents)
for docnum, d in enumerate(documents):
rd = reduced_documents[docnum]
maxrf = max([d.count(i) for i in d])
for t in rd:
tfd = rd.count(t)
tf = tfd/maxrf
df = 1 # Start count at 1 for smoothing
for doc in reduced_documents:
if t in doc:
df += 1
idf = math.log(N/df)
vectors.append(tf*idf)
return vectors
print(tfidf(docs, reduced_docs)
This, however, only gives me a list of the tf-idf score for each term:
[0.9162907318741551, 0.9162907318741551, 0.22314355131420976, 0.22314355131420976, ...]
I'm not sure how to create the sparse matrix, as is done in scikit-learn. If what I'm trying to do can be done using scikit-learn alone I'd probably just do that, but I couldn't make it exclude the words I don't want tf-idf calculations for.
How do I render my output in the same format as I would get by running the scikit-learn fit_transform() function?
EDIT: to include tokenisation for documents in the second function. It shouldn't affect the question.