Is a unique column good as partition key in Cassandra?

Viewed 362

I have a table user with multiple columns, every user has a unique userid. Because it is unique, I dont have to specify a clustering key unless I want to use the column in queries. Is this bad, because every partition consists of a single row? If it is bad for whatever reason, what is the best practise to do in this case? Thank you for your help!

Edit: If I have a query that needs to return all usernames, how can I do that with a good performance? Doing it from this table seems not very efficient for me, should I make another table where I simply duplicate all usernames in a Collection? Then they are all in one place and the read doesn't have to jump over multiple nodes.

2 Answers

I just answered the similar question. Short story - it really depends on the access patterns, and table settings. You may need to tune the table parameters to get best performance, but the settings may depend on the amount of data, and other requirements.

There are always two (main) considerations when defining your primary keys in Cassandra:

  • Data distribution
  • Query pattern match

From a data distribution standpoint, you can't get much better than using a unique key as the partition key. The more of them, the more evenly they should hash-out and thus be evenly distributed.

However, a key which distributes well but doesn't fit the desired query pattern, is pretty useless.

tl;dr;

If that unique key is all you'll ever query the table by, then it makes a great choice for a partition key.

Related