Is it OK to have millions of directed relationships between two nodes in Neo4j? Will it add to latency in fetching the data?

Viewed 117

We have millions of users of can make millions of transactions between them.

Lets say there is a user_A who pays money to user_B on the daily basis. We have a relationship SEND_MONEY_TO between A and B nodes.

A ---SENDS_MONEY_TO---> B

What will be the better design to accommodate this data in Neo4j.

Option A: We will create a new relationship edge every time a transaction happens.

Option B: We will keep a list of transactions as property of a same relationship edge and append the transaction details to existing list whenever a new transaction happens.

Our queries will look like:

a.) Find number of transaction between user_A and user_B in month of April 2021 where HDFC credit card is used.

b.) Find the total amount of transactions which involves user_A

We are open to any new approach as well.

2 Answers

I assume in the future, you may want to add other data to transactions, like the :Card :Device or :Platform on which it was triggered/Executed. As you probably know, you cannot create an edge between an edge and a vertex.

I would therefore recommend to use (:Transaction) vertices, with edges to the (:Account) vertices. You could use [:SENDER] and [:RECEIVER] types for the edges.

Here is how I've modeled this in the past. You might need to make a few adjustments to model it in Neo4j, but the general gist is there. Sample projects attached.


Because a customer might have multiple accounts

Customer -(CUSTOMER_ACCOUNT)- Account

Capture the directionality of the transactions

Account - (SEND_TRANSACTION)-> Transaction

Account <-(RECEIVE_TRANSACTION)-Transaction

Here I capture a running total of min_Send, min_receive, max_send, max receive, transaction_send_count, transaction_recive_count, etc

Account -(SEND_TO)-> Account


I've open-sourced the project here:

Sample Google Colab:


Related