I am writing horse racing odds aggregation application which will get data from difference bookies website. For a start I will get data from 3 websites(it can go up to more than 10 later) every 10 seconds. So in 3 websites case, there will be about 10,000 records(runners) each day and each record could be read 3 times every 10 seconds and updated if there are changes in odds.
- Is DynamoDB suitable that kinds of application or should I stick to RDBMS?
- Will there be any consistency issue with DyanmoDB when I am updating the odds(from different websites) for the same runner(record) simultaneously?
- The application may grow into other sports and racings and get data from more websites, will that incurred a huge cost with DyamoDB as it will do more read and write?
UPDATE - 21/07/2020 9:30AM Record structure will be something like below. There will be a few scheduled services running with each service taking care for a bookie. There are chances that a record will be read by services and updated simultaneously. Calculated column value will be based on the value of Bookies column. Hence, I want to be able to read the most recent value of Bookies column consistently.
RUNNER EVENTID BOOKIE1 BOOKIE2 BOOKIE3 BOOKIE... CALCULATED
Runner 1 12345 Odds1 Odds2 Odds3 Odds... Value
Runner 2 67890 Odds1 Odds2 Odds3 Odds... Value
UPDATE - 21/07/2020 12:20PM
After updating my post i get some numbers pop up in my head and DynamoDB seems to be very expensive. Here are my numbers, please let me know if anything is incorrect.
Assumptions:
- 10,000 Runner
- Every 10 seconds for a month roughly round up to 270,000 calls
- 3 bookies
- Assuming each record/item is under 4KB.
- One RCU can read 5.2million read per month (found somewhere)
- One WCU can read 2.5million read per month
RCU Required per month: (3 * 1,0000 * 270,000)/5.2 mill = 1558 RCU
WCU Required per month: (3 * 1,0000 * 270,000)/2.5 mill = 3240 WCU