Dating app Schema pattern nosql, is array extraction viable?

Viewed 156

I'm trying to build a dating app, and for my backend, I'm using a nosql database. When it comes to the user's collection, some relations are happening between documents of the same collection. For example, a user A can like, dislike, or may haven’t had the choice yet. A simple schema for this scenario is the following:

database = {
"users": {
    "UserA": {
        "_id": "jhas-d01j-ka23-909a",
        "name": "userA",
        "geo": {
            "lat": "",
            "log": "",
            "perimeter": ""
        },
        "session": {
            "lat": "",
            "log": ""
        },
        "users_accepted": [
            "j2jl-564s-po8a-oej2",
            "soo2-ap23-d003-dkk2"

        ],
        "users_rejected": [
            "jdhs-54sd-sdio-iuiu",
            "mbb0-12md-fl23-sdm2",
        ],
    },
    "UserB": {...},
    "UserC": {...},
    "UserD": {...},
    "UserE": {...},
    "UserF": {...},
    "UserG": {...},

},

}

Here userA has a reference from the users it has seen and made a decision, and stores them either in “users_accepted” or “users_rejected”. If User C hasn’t been seen (either liked or disliked) by userA, then it is clear that it won’t appear in both of the arrays. However, these arrays are unbounded and may exceed the max size that a document can handle. One of the approaches may be to extract both of these arrays and create the following schema:

database = {
"users": {
    "UserA": {
        "_id": "jhas-d01j-ka23-909a",
        "name": "userA",
        "geo": {
            "lat": "",
            "log": "",
            "perimeter": ""
        },
        "session": {
            "lat": "",
            "log": ""
        },
    },
    "UserB": {...},
    "UserC": {...},
    "UserD": {...},
    "UserE": {...},
    "UserF": {...},
    "UserG": {...},


},
"likes": {
    "id_27-82" : {
        "user_give_like" : "userB",
        "user_receive_like" : "userA"
    },
    "id_27-83" : {
        "user_give_like" : "userA",
        "user_receive_like" : "userC"
    },
},
"dislikes": {
    "id_23-82" : {
        "user_give_dislike" : "userA",
        "user_receive_dislike" : "userD"
    },
    "id_23-83" : {
        "user_give_dislike" : "userA",
        "user_receive_dislike" : "userE"
    },
}

}

I need 4 basic queries

  1. Get the users that have liked UserA (Show who is interested in userA)
  2. Get the users that UserA has liked
  3. Get the users that UserA has disliked
  4. Get the matches that UserA has

The query 1. is fairly simple, just query the likes collection and get the users where "user_receive_like" is "userA".

Query 2. and 3. are used to get the users that userA has not seen yet, get the users that are not in query 2. or query 3.

Finally query 4. may be another collection

"matches": {
    "match_id_1": {
        "user_1": "referece_user1",
        "user_2": "referece_user2"
    },
    "match_id_2": {
        "user_1": "referece_user3",
        "user_2": "referece_user4"
    }
}

Is this approach viable and efficient?

2 Answers

You are right to notice, that these arrays are unbounded and pose a serious scalability problem for your application. If you were to assign 2-3 user roles to a user with the 1st approach it would be totally fine, but it is not the case for you. The official MongoDB documentation suggests that you should not use unbounded arrays: https://www.mongodb.com/docs/atlas/schema-suggestions/avoid-unbounded-arrays/

Your second approach is the superior implementation choice for you, because:

  • you can build indices of form (user_give_dislike, user_receive_like) which will improve your query performance even in case when you have 1M+ documents
  • you can store additional metadata, (like timestamps etc) on the likes collection without affecting the design of the user collection
  • the query for "matches" will be much simpler to write with this approach: https://mongoplayground.net/p/sFRvUniHKn8

More about NoSQL data modelling: https://www.mongodb.com/docs/manual/data-modeling/ and https://www.mongodb.com/docs/manual/tutorial/model-referenced-one-to-many-relationships-between-documents/

To answer your question , let me write some more assumptions about the domain and then lets try to answer it.

Assumptions:

  1. System should support scale for 100 million users
  2. A single user might like or dislike ~100k users in its lifetime

Also some thoery about nosql , if our queries go to all the shards of the collection then max scale of the system depends on the scale of the single shard

Now with these assumptions see the query performance of the question that you asked :

  1. Get the users that have liked UserA (Show who is interested in userA) -

Assuming we are doing sharding or user_give_like column then if we filter on user_receive_like then it will do query on all shards , which is not the right thing for scalability

  1. Get the users that UserA has liked This will work fine as we have created shard based on user_give_like

  2. Get the users that UserA has disliked This will work fine as we have created shard based on user_give_dislike

  3. Get the matches that UserA has In this case if we do a join between existing users and all users which UserA has liked and disliked this will create a parallel query on all shard and is not scalable when UserA like or dislike has huge count

Now to conclude this dosen't look like a reasonable approach to me.

Related