Efficient/Performant way to visualise a lot of data in javascript + D3/mapbox

Viewed 96

I am currently looking at an efficient way to visualise a lot of data in javascript. The data is geospatial and I have approximately 2 million data points.

Now I know that I cannot give that many datapoint to the browser directly otherwise it would just crash most of the time (or the response time will be very slow anyway).

I was thinking of having a javascript window communicating with a python which would do all the operations on the data and stream json data back to the javascript app.

My idea was to have the javascript window send in real time the bounding box of the map (lat and lng of north east and south west Point) so that the python script could go through all the entries before sending the json of only viewable objects.

I just did a very simple script that could do that which basically

  1. Reads the whole CSV and store data in a list with lat, lng, and other attributes (2 or 3)
  2. A naive implementation to check whether points are within the bounding box sent by the javascript.

Currently, going through all the datapoints takes approximately 15 seconds... Which is way too long, since I also have to then transform them into a geojson object before streaming them to my javascript application.

Now of course, I could first of all sort my points in ascending order of lat and lng so that the function checking if a point is within the javascript sent bounding box would be an order of magnitude faster. However, the processing time would still be too slow.

But even admitting that it is not, I still have the problem that at very low zoom levels, I would get too many points. Constraining the min_zoom_level is not really an option for me. So I was thinking that I should probably try and cluster data points.

My question is therefore do you think that this approach is the right one? If so, how does one compute the clusters... It seems to me that I would have to generate a lot of possible clusters (different zoom levels, different places on the map...) and I am not sure if this is an efficient and smart way to do that.

I would very much like to have your input on that, with possible adjustments or completely different solutions if you have some.

This is almost language agnostic, but I will tag as python since currently my server is running python script and I believe that python is quite efficient for large datasets.

Final note:

I know that it is possible to pre-compute tiles that I could just feed my javascript visualization but as I want to have interactive control over what is being displayed, this is not really an option for me.


Edit:

I know that, for instance, mapbox provides the clustering of data point to facilitate displaying something like a million data point.

However, I think (and this is related to an open question here ) while I can easily display clusters of points, I cannot possibly make a data-driven style for my cluster.

For instance, if we take the now famous example of ethnicity maps, if I use mapbox to cluster data points and a cluster is giving me 50 people per cluster, I cannot make the cluster the color of the most represented ethnicity in the sample of 50 people that it gathers.

Edit 2:

Also learned about supercluster, but I am quite unsure whether this tool could support multiple million data points without crashing either.

0 Answers
Related