Are client side joins permissable in Cassandra if client drills down on datapoint?

Viewed 27

I have this structure with about 1000 data points in a list on the website:

  • Datapoint1:
  • Datapoint2:

...

  • Datapoint1000:

With each datapoint containing 6 fields of information. Each datapoint can be opened to reveal an additional 2-3x of information in sublist. Would making a new request upon the user clicking on one of my datapoints be considered bad practice in Cassandra? Should I just go ahead and get it all in one go?

2 Answers

Should I just go ahead and get it all in one go?

Definitely not.

Would making a new request upon the user clicking on one of my datapoints be considered bad practice in Cassandra?

That's absolutely the way you should do it. Cassandra is great at writing large amounts of data, but not so great a returning large amounts of data. More, small key-based queries are definitely the way to go.

It is possible to do the JOINs on the client side but as a general proposition, queries which require joins indicate that you possibly didn't design the data model correctly.

You need to model your data such that (a) each application query (b) maps to a single table. If you need to do a client-side JOIN then you need to query the database multiple times to get the data required by your app. It will work but it's not efficient so affects the performance of the app and the database.

To illustrate with an example, let's say you app needs to display a customer's list of orders. The table design would need to be partitioned by customer with (clustered) multiple rows of orders:

CREATE TABLE orders_by_customerid (
    customerid text,
    orderid text,
    orderdate timestamp,
    ordertotal decimal,
    ...
    PRIMARY KEY (customerid, orderid)
)

You would retrieve the list of orders for a customer with:

SELECT ... FROM orders_by_customerid WHERE customerid = ?

By default, the driver or Stargate API your app is using would page the results so only the first 100 rows (for example) will be returned instead of retrieving thousands of rows in a single pass. Note that the page size is configurable. Cheers!

Related