What is the best practice for handling load when collecting billions of datasets in a .NET No-SQL environment?

Viewed 26

I have a new project requirement that will see the collection billions of datasets from different web-sources every week. I have established that a No-SQL storage solution is most appropriate (probably using Apache Spark) but I need to consider the performance and load when receiving and storing these billions of records. All literature and examples I can find for working with No-SQL on large datasets is utilising pre-existing large datasets, so ignoring the problem of collecting the data to begin with.

My previous experience has been in SQL Server supported solutions, interfacing through either .NET Web-API or through Signalr. I've not had to deal with anything like the amount of load this will require before.

Please advise as to best practices for collecting large datasets and link to any useful resources.

0 Answers
Related