Django taggit: why annotate(same_tags=Count('tags')) counts the number of common tags instead of the total number of tags?

Viewed 398

From the tutorials in Django 2 by Example, I don't understand:

step (2): Why is `Count('tags')` **not** counting 
the total number of tags possessed by that post?

This code:

# List of similar posts
post_tags_ids = post.tags.values_list('id', flat=True)
similar_posts = Post.published.filter(tags__in=post_tags_ids)\
                              .exclude(id=post.id)
similar_posts = similar_posts.annotate(same_tags=Count('tags'))\
                             .order_by('-same_tags','-publish')[:4]

does this:

  1. searches similar posts by looking at their common tags.
  2. uses Count aggregation function to generate a calculated field same_tags.
  3. orders the result by the number of shared tags in descending order etc...

I searched Taggit's API reference but it seems irrelevant.

1 Answers

I don't understand step (2): Why is Count('tags') not counting the total number of tags possessed by that post?

Because the .annotate(..) clause is used after the .filter(..) clause. You thus first filter the joined model, and then you count the elements that are still retained.

As described in the aggregation section of the documentation:

When used with an annotate() clause, a filter has the effect of constraining the objects for which an annotation is calculated. For example, you can generate an annotated list of all books that have a title starting with “Django” using the query:

>>> from django.db.models import Avg, Count
>>> Book.objects.filter(name__startswith="Django").annotate(num_authors=Count('authors'))

You thus create a query that looks like:

SELECT post.*
       COUNT(tag.id) AS same_tags
FROM post
INNER JOIN tag
WHERE tag.id IN list_of_tag_ids
  AND post.id != id_of_post
ORDER BY same_tags DESC, post.publish DESC
Related