Swift Firebase "Fan Out" Technique vs queryLimited Efficiency

Viewed 86

I have a group chat feature in my app that has its messages node structured like this. Currently, it doesn't use the fan-out technique. It just lists all of the messages under the group name e.g. "group1"

groups: {
    group1: {
      -MEt4K5xhsYL33anhXpP: {
          fromUid: "diidssm......."
          userImage: "https://firebasestorage..."
          text: "hello"
          date: 1617919946
          emojis: {
              "heart": 2
              "like": 1
          }
      }
      -MEt8BLP2yMEUMPbG2zV: {
          ...
      }
      -MF-Grpl8Jchxpbn2mxH: {
          ...
      }
      -MF-OUjWXsFh7lBPosMf: {
          ...
      }
    }
}

I first observe the most recent 40 messages and observe whether new children get added as such

ref = Database.database().reference().child("groups").child("group1")
ref.queryLimited(toLast: 40).observe(.childAdded, with: { (snapshot) in
    ...
    //add to messages array to load collection view
    //for each message observe emojis and update emojis to reflect changes e.g. +1 like

    ref.child("emojis").observe(.value, with: { (snapshot) in
        ...
    })
})

Every time the user scrolls up I load another 40 messages (and observe the emojis child under each of those message nodes) using the last date (and index by date in security rules) as such

ref.queryOrdered(byChild: "date").queryEnding(beforeValue: prevdate, childKey: messageId).queryLimited(toLast: 40).observeSingleEvent(of: .value, with: { (snapshot) in

I understand the fan-out technique is used to get less information per synchronization. If I attach a listener to the groups/groupname/ to get a list of all messages for that group, I will also ask for all the info of each and every message under that node. With the fan out approach I can also just ask for the message information of the 40 most recent messages and the next 40 per scroll up using the keys of the messages from another node like this.

allGroups: {
    group1: {
      -MEt4K5xhsYL33anhXpP: 1
      -MEt8BLP2yMEUMPbG2zV: 1
      -MF-Grpl8Jchxpbn2mxH: 1
      -MF-OUjWXsFh7lBPosMf: 1
    }
}

However, if I am using queryLimited(toLast: 40) is the fan-out approach beneficial or even necessary? Wouldn't this fix the problem of "I will also ask for all the info of each and every message under that node"?

In terms of checking for new messages, I just check using .childAdded in the first code above (ref.queryLimited(toLast: 40).observe(.childAdded)). According to the post below, queryLimited(toLast: 40) will sync the last 40 child nodes, and keep synchronizing those (removing previous ones as new ones are added).

Some questions about keepSynced(true) on a limited query reference

I'm assuming if group1 had 1000 messages, with this approach I am just reading the 40 most recent messages I need and the next 40 per scroll, thus ignoring the other several hundred. Why would I use the fan-out technique then? May be I'm not understanding something fundamental about limited queries.

Side Question: Should I be including references to profile images under each message node? Is it bad to do this in terms of cloud storage and realtime database storage? Ideally there would be hundreds of groupchats.

1 Answers

There's a lot of comments to the question so I thought I would condense all of that into an answer.

The intention of the 'fan out technique' in the question was to maximize query performance.

In this use case the query only returns the last 40 results

ref.queryLimited(toLast: 40)

The assumption in the question was that Firebase had to 'go through' all of the nodes before those 40 to get to the 40, therefore affecting performance. That's not the case with Firebase so whether it be the first 40 or the last 40, the performance is 'the same'.

Because of that, no 'fan-out' is really needed in this situation. For clarity

Fan-out is the process duplicating data in the database. When data is duplicated it eliminates slow joins and increases read performance.

I am going to steal a fan out example from an old Firebase Blog. Here's a fan out to update multiple nodes at once, and since it's an atomic operation it either all passes or all fails.

let updatedUser = ["name": "Shannon", "username": "shannonrules"]
let ref = Firebase(url: "https://<YOUR-FIREBASE-APP>.firebaseio.com")

let fanoutObject = ["/users/1": updatedUser, 
                    "/usersWhoAreCool/1": updatedUser, 
                    "/usersToGiveFreeStuffTo/1", updatedUser]

ref.updateChildValues(updatedUser) // atomic updating goodness

I will also include a link to Introducing multi-location updates and more as well as suggesting a read on the topic of denormalization.

In the question, there isn't really any data to 'fan out' so it would not be applicable as there isn't an attempt to join (pull data from multiple nodes) or to update multiple nodes.

The one change I would suggest would be to remove the emoji's node from the message node.

As is, every one of those has an observer which results in thousands of observers which can be difficult to manage. I would create a separate high-level node just for those emojis

emojis
   -MEt4K5xhsYL33anhXpP: //the message id
      "heart": 2  //or however you want to store them
      "like": 1

Then add a single observer (much easier to manage!) to the emoji node. When an emoji changes, that one observer will notify the app of which message it was for, and what the change was. It will also cut down on reads and overall cost.

Related