Firebase Firestore Structure for getting un-seen trending posts - Social

Viewed 1146

This is the current sample structure

Posts(Collection)
    - post1Id : {
          viewCount : 100,
          likes     : 45,
          points    : 190,
          title     : "Title",
          postType  : image/video
          url       : FileUrl,
          createdOn : Timestamp,
          createdBy : user20Id,
          userName  : name,
          profilePic: url
      }
Users(Collection)
    - user1Id(Document):{
          postsCount : 10,
          userName  : name,
          profilePic : url
      }
        viewed(Collection)
            - post1Id(Document):{
                  viewedTime : ""
              }
                 
    - user2Id(Document)

The End goal is

  • I need to getPosts that the current user did not view and in points field descending order with paging.

What are the possible optimal solutions(like changing structure, cloud functions, multiple queries from client-side)?

2 Answers

I'm working on a solution to show trending posts and eliminate posts that are already seen by users or poor content. It's really painful to deal with two queries especially when the user base is increasing. It's difficult to maintain the "viewed" collection and filter the new posts. Imagine having 1 million viewed posts and then filter for the un-seen posts.

So I figured a solution, which is not that great, but still cool.

So here is our data structure

posts(Collection) --postid(document)

  1. Title.
  2. Description.
  3. Image.
  4. timestamp.
  5. priority

This is a simple post structure with basic details. You can see I have added a Priority field. This field will do the magic.

How to use Priority.

  1. We should query the posts that start with the higher priority and ends with lower priority.
  2. When a user posts a new Post. Assign the current timestamp as the default priority.
  3. When the user upvotes (Likes) a post increase the priority by 1 minute(60000 milliseconds)
  4. When the user downvotes (Dislike) a post decrease the priority by 1 minute (60000 ms)
  5. You can reset the priority every 24 hours. If you start browsing the feed today morning you will see posts with the last 24 hours in past. Once the 24-hour duration reached you can reset the priority to the present time. The 24-hour limit can be changed according to your needs. You may want to reset the limit every 15 min. because in every 15 min 100s of new posts might have added. This limit will ensure the repetition of content in the feed.

So when you start scrolling the feed you will get all the trending posts first then lower priority posts later. If you post a post today and people start upvoting it. It will get an increased lifetime, thus overpowers the poor content and when you downvote it, it will push down the post as long as users will not reach it.

Using timestamp as a priority because the old posts should lose priority with time. Even the trending posts today should lose the priority tomorrow.

Things to consider:

The lifetime can vary according to your needs. The bigger the user base. You should lower the lifetime value. because if a post posted today is upvoted by 10,000 users it trends 6.9 days in the future. And if there are more than 100 posts that have been upvoted by more than 10,000 users then you will never get to see a new post in those 6.9 days. So a trending post should hardly last a day or two.

So in this case you can give 10 seconds lifetime, it will give 1.1 day lifetime for 10,000 upvotes.

This is not a perfect solution but it may help you get started.

Edit: 11th June 2021

Nowadays, there are two more options that can help you solve such a problem. The first one would be the whereNotEqualTo method and the second one would be whereNotIn. You might choose one, or the other according to your needs.


Seeing your database structure, I can say you're almost there. According to your comment, you are hosting under the following reference:

Users(Collection) -> userId(Document) -> viewed(Collection)

As documents, all the posts a user has seen and you want to get all the post that the user hasn't seen. Because there is no != (not equal to) operator in Firestore nor a arrayNotContains() function, the only option that you have is to create an extra database call for each post that you want to display and check if that particular post is already seen or not.

To achieve this, first you need to add another property under your post object named postId, which will hold as String the actual post id. Now everytime you want to display the new posts, you should check if the post id already exist in viewed collection or not. If it dons't exist, display that post in your desired view, otherwise don't. That's it.


Edit: According to your comments:

So, for the first post to appear, it needs two Server calls.

Yes, for the first post to appear, two database calls are need, one to get post and second to see if it was or not seen.

large number of server calls to get the first post.

No, only two calls, as explained above.

Am I seeing it the wrong way

No, this is how NoSQL database work.

or there is no other efficient way?

Not I'm aware of. There is another option that will work but only for apps that have limited number of users and limited number of post views. This option would be to store the user id within an array in each post object and everytime you want to display a post, you only need to check if that user id exist or not in that array.

But if a post can be viewd by millions of users, storing millions of ids within an array is not a good option because the problem in this case is that the documents have limits. So there are some limits when it comes to how much data you can put into a document. According to the official documentation regarding usage and limits:

Maximum size for a document: 1 MiB (1,048,576 bytes)

As you can see, you are limited to 1 MiB total of data in a single document. So you cannot store pretty much everything in a document.

Related