Can I calculate my own TF-IDF scores and still output them in the same format as scikit-learn's TfidfVectorizer does?

Viewed 86

I need to calculate my own tf-idf instead of using the TfidfVectorizer built into scikit-learn, but I want my output to be in the same format as I would get using scikit-learn's TfidfVectorizer when I call the fit_transform() function.

My original code was as follows, and it gave me the ouput style that I wanted.

doc1 = "These are some words I'm putting in a document."
doc2 = "This document is comprised of a number of words."
doc3 = "Some of the words in this document are found in other documents also."
doc4 = "We are now writing a piece of text which is entirely separate from the others and, hence, dissimilar."
doc5 = "Another piece of writing in which we are interested for its dissimilarity to its precursors is this."
docs = [doc1, doc2, doc3, doc4, doc5]

def tfidf(documents, reduced_documents=None):
    vectorizer = TfidfVectorizer(analyzer='word')
    vectors = vectorizer.fit_transform(documents)
    return vectors

print(tfidf(docs)

The output of this is a sparse matrix, which I'm not entirely sure how to recreate. It looks like this:

(0, 7)  0.322444504927649
(0, 14) 0.322444504927649
(0, 25) 0.481467662591444
(0, 35) 0.322444504927649
...

Because I want to only calculate the tf-idf score for some words in each document I've had to rewrite the tfidf() function as follows:

reduced_doc1 = "These I'm in a"
reduced_doc2 = "This of a of"
reduced_doc3 = "Some of the in this in"
reduced_doc4 = "We a of which from the and"
reduced_doc5 = "Another of in which we for its to its this"
reduced_docs = [reduced_doc1, reduced_doc2, reduced_doc3, reduced_doc4, reduced_doc5]

def tfidf(documents, reduced_documents=None):

    documents = [i.split(" ") for i in documents]
    reduced_documents = [i.split(" ") for i in reduced_documents]
    vectors = []

    if not reduced_documents:
        reduced_documents = documents

    N = len(documents)

    for docnum, d in enumerate(documents):
        rd = reduced_documents[docnum]
        maxrf = max([d.count(i) for i in d])
        for t in rd:
            tfd = rd.count(t)
            tf = tfd/maxrf
            df = 1  # Start count at 1 for smoothing
            for doc in reduced_documents:
                if t in doc:
                    df += 1
            idf = math.log(N/df)
            vectors.append(tf*idf)

    return vectors

print(tfidf(docs, reduced_docs)

This, however, only gives me a list of the tf-idf score for each term:

[0.9162907318741551, 0.9162907318741551, 0.22314355131420976, 0.22314355131420976, ...]

I'm not sure how to create the sparse matrix, as is done in scikit-learn. If what I'm trying to do can be done using scikit-learn alone I'd probably just do that, but I couldn't make it exclude the words I don't want tf-idf calculations for.

How do I render my output in the same format as I would get by running the scikit-learn fit_transform() function?

EDIT: to include tokenisation for documents in the second function. It shouldn't affect the question.

0 Answers
Related