How do I use Solr "relatedness()" function to measure relatedness of two sets of documents?

Viewed 376

I'd like to use the new Semantic Knowledge Graph capability in Solr to answer this question: Given a set of documents from several different publishers, compute a "relatedness" metric between a given publisher and every other publisher, based on the text content of their respective documents.

I've watched several of Trey Grainger's talks regarding the Semantic Knowledge Graph functionality in Solr (this is a great recent one: https://www.youtube.com/watch?v=lLjICpFwbjQ) I have a reasonably good understanding of Solr faceted search functionality, and I have a working Solr engine with my dataset indexed and searchable. So far I've been unable to construct a facet query to do what I want.

Here is an example curl command which I thought might get me what I want

curl -sS -X POST http://localhost:8983/solr/plans/query -d '
{
  params: {
    fore:"publisher_url:life.church"
    back:"*:*",
  },
  query:"*:*",
  limit: 0,
  facet:{
      pub_type: {
        type: terms,
        field: "publisher_url",
        limit: 5,
        sort: { "r1": "desc" },
        facet: {
          r1: "relatedness($fore,$back)"
        }
      }
    }
  }
}'

Below are the result facets. Notice that after the first bucket (which matches the foreground query), the others all have exactly the same relatedness. Which leads me to believe that the "relatedness" is only based on the publisher_url field rather than the entire text content of the documents.

{
  "facets":{
    "count":2152,
    "pub_type":{
      "buckets":[{
          "val":"life.church",
          "count":141,
          "r1":{
            "relatedness":0.38905,
            "foreground_popularity":0.06552,
            "background_popularity":0.06552}},
        {
          "val":"10ofthose.com/us/products/1039/colossians",
          "count":1,
          "r1":{
            "relatedness":-0.00285,
            "foreground_popularity":0.0,
            "background_popularity":4.6E-4}},
        {
          "val":"14DAYMARRIAGECHALLENGE.COM",
          "count":1,
          "r1":{
            "relatedness":-0.00285,
            "foreground_popularity":0.0,
            "background_popularity":4.6E-4}},
        {
          "val":"23blast.com",
          "count":1,
          "r1":{
            "relatedness":-0.00285,
            "foreground_popularity":0.0,
            "background_popularity":4.6E-4}},
        {
          "val":"2911worship.com",
          "count":1,
          "r1":{
            "relatedness":-0.00285,
            "foreground_popularity":0.0,
            "background_popularity":4.6E-4}}]}}}
1 Answers

I'm not very familiar with the relatedness function, but as far as I understand, the relatedness score is generated from the similarity between your foreground and background set of documents for that facet bucket.

Since your foreground set only contain that single value (and none of the other), the first bucket is the only one that will generate a different similarity score when you're faceting for the same field as you use for selecting documents.

I'm not sure if your use case is a good match for what you're trying to use, as relatedness would indicate that single terms in a field is related between the two sets you're using, and not a similarity score across a different field for the two comparison operators.

You probably want something more structured than a text field to generate relatedness() scores, as that's usually more useful for finding single values that generate statistical insight into the structure of your query set.

The More Like This functionality might actually be a better match for getting the most similar other sites instead.

Again, this is based on my understanding of the functionality at the moment, so someone else can hopefully add more details and correct me as necessary.

Related