Can't get metadata from dataframe using DataframeSource in tm for R

Viewed 1548

I have a dataframe with the following variables:

doc_id  text  URL  author  date  forum 

When I run

samplecorpus <- Corpus(DataframeSource(sampledataframe))

the documentation says I should get a corpus with all of the extra variables added as document-level metadata. https://rdrr.io/rforge/tm/man/DataframeSource.html http://finzi.psych.upenn.edu/R/library/tm/html/DataframeSource.html

Instead, I get a corpus that has all of the right documents in the right order, but all of their metadata is blank. I need this metadata to filter the documents for future analysis.

Someone else asked a similar question, but it never got answered... In tm version a readTabular() replacement tm package DataframeSource () ignores my other columns as metadata

Does anyone have any ideas on how to fix this?

Thanks!

2 Answers

The documentation for tm explains this if you dig down (see ??tm::DublicCore). From the docs:

A corpus has two types of metadata. Corpus metadata ("corpus") contains corpus specific metadata in form of tag-value pairs. Document level metadata ("indexed") contains document specific metadata but is stored in the corpus as a data frame. Document level metadata is typically used for semantic reasons (e.g., classifications of documents form an own entity due to some high-level information like the range of possible values) or for performance reasons (single access instead of extracting metadata of each document). The latter can be seen as a from of indexing, hence the name "indexed". Document metadata ("local") are tag-value pairs directly stored locally at the individual documents.

DataframeSource automatically assigns only the the corpus metadata*. For example, see what the following prints:

library(tm)
data <- data.frame(doc_id = c(234345345, 1299),
                   text = c("The Prince and the Pauper", 
                            "Little Women"),
                   author = c('Mark Twain', 'Louisa May Alcott'),
                   date = c(1881, 1868),
                   stringsAsFactors = FALSE)

samplecorpus <- Corpus(DataframeSource(data))
meta(samplecorpus) 
# Or even
meta(samplecorpus[1], tag = 'author')

In order to assign metadata at the document level, you can work with meta to change tags. Bizarrely, this only works if you use VCorpus. So changing the above slightly, you can do:

samplecorpus <- VCorpus(DataframeSource(data))
# Can now set document metadata tags
meta(samplecorpus[[1]], tag = 'author') <- 'Mark Twain'

*EDIT:

Contemplating further (and responding to OP's comment), I agree that the documentation is not a completely accurate description of the package's observed behavior. The quoted documentation above refers to three levels (Corpus, indexed document level, and local document level), which in my example appear to correspond to samplecorpus, samplecorpus[1], and samplecorpus[[1]], respectively. If this correct, then the metadata is being assigned by DataframeSource at the promised level (if somewhat vaguely, as they never specified which document-level). However, the docs also claims the indexed document level is stored as a data frame and local as tag-value pairs, but both are stored as lists. Confusing. I can only conclude that this is either a bug in the package implementation or an error in the docs.

Barring contacting the package authors to clear this up (not a bad idea), I would propose the following workaround:

samplecorpus <- VCorpus(DataframeSource(data))
transfer_metadata <- function(x, i, tag){
  return(meta(x[i], tag=tag)[[tag]])
}

tags <- colnames(data)
tags <- tags[! tags %in% c('doc_id', 'text')]

for(i in 1:length(samplecorpus)){
  for (tag in tags){
    meta(samplecorpus[[i]], tag=tag) <- transfer_metadata(samplecorpus, i=i, tag=tag)
    }
}

You have to check if everything is loaded correctly. I made an example docs data.frame so you can see how it works. I used the same column names you have and added 1 extra (tags). Based on this example you might check if you have an issue somewhere.

docs <- data.frame(doc_id = c("doc_1", "doc_2"),
                   text = c("This is a text.", "This another one."),
                   url = c("https://stackoverflow.com/questions/52433344/cant-get-metadata-from-dataframe-using-dataframesource-in-tm-for-r",
                           "https://stackoverflow.com/questions/52433344/cant-get-metadata-from-dataframe-using-dataframesource-in-tm-for-r"), 
                   author = c("Emi", "Emi"),
                   date = as.Date(c("2018-09-20", "2018-09-21")),
                   forum = c("stackoverflow", "stackoverflow"),
                   tags = c("r", "tm"),
                   stringsAsFactors = T)

# use Corpus or VCorpus
my_corpus <- Corpus(DataframeSource(docs))
meta(my_corpus)

    url author       date
1 https://stackoverflow.com/questions/52433344/cant-get-metadata-from-dataframe-using-dataframesource-in-tm-for-r    Emi 2018-09-20
2 https://stackoverflow.com/questions/52433344/cant-get-metadata-from-dataframe-using-dataframesource-in-tm-for-r    Emi 2018-09-21
          forum tags
1 stackoverflow    r
2 stackoverflow   tm

my_index <- meta(my_corpus, "tags") == "r"

inspect(my_corpus[my_index])
<<SimpleCorpus>>
Metadata:  corpus specific: 1, document level (indexed): 5
Content:  documents: 1

          doc_1 
This is a text. 

Now beware there is a difference in how meta is treated. If you do str(my_corpus) you will see the following:

List of 2
 $ doc_1:List of 2
  ..$ content: chr "This is a text."
  ..$ meta   :List of 7
  .. ..$ author       : chr(0) 
  .. ..$ datetimestamp: POSIXlt[1:1], format: "2018-09-21 08:55:44"
  .. ..$ description  : chr(0) 
  .. ..$ heading      : chr(0) 
  .. ..$ id           : chr "doc_1"
  .. ..$ language     : chr "en"
  .. ..$ origin       : chr(0) 
  .. ..- attr(*, "class")= chr "TextDocumentMeta"
  ..- attr(*, "class")= chr [1:2] "PlainTextDocument" "TextDocument"
 $ doc_2:List of 2
......

The meta info you see here is from meta(my_corpus, type = "local"). The metadata loaded with DataframeSource is of type indexed, meta(my_corpus, type = "indexed")

Page 5 of the vignette is important to read and experiment with to see all the different options that meta and DublinCore.

Related