What is difference between DynamoDB and ScyllaDB

Viewed 81

Looks like DynamoDB and ScyllaDB are exactly similar in functionality where they have just used different names for keys, secondary indexes etc.

Only difference I am aware of is costing. DynamoDB charges for throughput whereas ScyllaDB charges for storage size.

So wanted to know when to use which DB.

3 Answers

It depends on the bunch of factors e.g. in some projects, Team chose DynamoDB over ScyllaDB since they're using all other services from the same cloud provider and integration/support/cost was great when they picked up DynamoDB over ScyllaDB.

Following are few things to consider (at high-level before choosing between DynamoDB and ScyllaDB)

DynamoDB

  • Excellent for projects where you need to store a large amount of data, but you do not know how many will be so you need the database to increase its storage capacity together with the number of users, without having to spend extra money.

ScyllaDB

  • Scylla is well suited for high-throughput scenarios where keyed data must be read or written with consistently low latency.

Both DynamoDB and ScyllaDB were inspired by Cassandra, so you're right about the "similar in functionality" and that they, indeed, "used different names for keys" (e.g., what Cassandra and Scylla calls "clustering keys", are calls "sort keys" (or sometimes, "range keys") in DynamoDB).

However, their capabilities, and their performance tradeoffs, are not really 100% identical. A couple of years ago I wrote a blog post, Comparing CQL and the DynamoDB API, which compares some of the more interesting differences between the capabilities and performance tradeoffs that CQL (the Cassandra Query Language, also used natively by Scylla) took, compared to DynamoDB's API. Some example differences explained in that blog post are a different network protocol (with different advantages and disadvantages), topology-aware vs. "dumb" clients, and perhaps most interestingly - a very different write model: Scylla focuses on very efficient CRDT (write-only) operations, while in DynamoDB, every write can involve a read as well - more powerful but slower (Scylla also has this power, through "LWT" (lightweight transacations)).

Because of the similarities between Scylla's and DynamoDB's APIs, we were actually able to fully (or almost fully) support the DynamoDB API in ScyllaDB - so ScyllaDB now supports the DynamoDB API as well (see ScyllaDB Alternator).

Besides the above differences in functionality, the most obvious difference between the two products is in how it is deployed and used in practice: DynamoDB is, like other Amazon products, a service on AWS where you pay per request, whereas ScyllaDB is software which you either install yourself, or get pre-deployed but in either case you get a cluster of your own (it's not shared with other customers) and you need to choose its size explicitly - by the number of nodes, not the number of requests.

Disclosure: I work for ScyllaDB.

DynamoDB is a key-value NoSQL store. ScyllaDB's Alternator interface is an API-compatible implementation of DynamoDB. The advantage of ScyllaDB is you can run it on any cloud or on-premises; DynamoDB only works in AWS.

ScyllaDB also has a CQL interface, which is technically a wide column NoSQL store.

EDIT In fact, both DynamoDB, ScyllaDB and Cassandra should be technically described as "wide column NoSQL stores." Or, as my colleague Nadav describes, a "key-key-value" store. Both ScyllaDB and DynamoDB use the term "partition key." ScyllaDB refers to the second part of the key as a the "clustering key," whereas DynamoDB calls it the "sort key." END EDIT

We've also heard from customers that DynamoDB was a great place to start, but affordability suffered as they reached scale. Moving to ScyllaDB meant they were not paying transactional costs against their own data. i.e., with DynamoDB, the more you query the more you pay. Which makes heavy read/write workloads prohibitively expensive.

So a lot may depend on your use case. Where do you need to deploy? How much data are you managing? How hard are you hitting that data? How many operations per second do you need to maintain?

Then, with ScyllaDB you have the choice of which interface: DynamoDB API or CQL. Generally unless you need to remain compatible with a current DynamoDB Workloads, we generally recommend the CQL interface. It provides some greater flexibility and performance.

Related