Elasticsearch: Modeling product data with frequent updates

Viewed 33

We're struggling with modeling our data in Elasticsearch, and decided to change it.

What we have today: single index to store product data, which holds data of 2 types -

[1] Some product data that changes rarely -

* `name, category, URL, product attributes(e.g. color,price) etc...`

[2] Product data that might change frequentley for past documents, and indexed on a daily level - [KPIs]

* `product-family, daily sales, daily price, daily views...`

Our requirements are -

  • Store product-related data (for millions of products)
  • Index KPIs on a daily level, and store those KPIs for a period of 2 years.
  • Update "product-family" on a daily level, for thousands of products. (no need to index it daily)
  • Query and aggregate the data with low latency, to display it in our UI. aggregation examples -
    1. Sum all product sales in the last 3 months, from category 'A' and sort by total sales.
    2. Same as the above, but in-addition aggregate based on product-family field.
  • Keep efficient indexing rate.

Currently, we're storing everything on the same index, daily, meaning we store repetitive data such as name, category and URL over and over again. This approach is very problematic for multiple reasons-

  • We're holding duplicates for data of type [1], which hardly changes and causes the index to be very large.
  • when data of type [2] changes, specifically the product-family field(this happens daily), it requires updating tens of millions of documents (from more than a year ago), which causes the system to be very slow and timeout on queries.

Splitting this data into 2 different indices won't work for us since we have to filter data of type [2] by data of type [1] (e.g. all sales from category 'A'), moreover, we'll have to join that data somehow, and our backend server won't handle this load.

We're not sure how to model this data properly, our thoughts are -

  • Using parent-child relations - parent is product data of type [1] and children are KPIs of type [2]
  • Using nested fields to store KPIs (data of type [2]).

Both of these methods allow us to reduce the current index size by eliminating the duplicated data of type [1], and efficiently updating data of type [2] for very old documents.

Specifically, both methods allow us to store product-family for each product once in the parent/non-nested fields, which implies we can only update a single document per product. (these updates are daily)

We think parent-child relation is more suitable, due to the fact that we're adding KPIs on a daily level, which per our understanding - will cause re-indexing for documents with new KPIs when using nested fields. On the other side, we're afraid that parent-child relations will increase query latency dramatically, hence will cause our UI to be very slow.

We're not sure what is the proper way to model the data, and if our solutions are on the right path, we would appreciate any help since we're struggling with it for a long time.

1 Answers

First off, I would recommend against indexing data that changes frequently in Elasticsearch. It is not designed for this and you will get poor performance as well as encounter difficulties when cleaning up old data.

Elasticsearch is best used for immutable data (once you insert it, you don't modify it). For time based data, I would recommend inserting measurements once with their timestamp, in e.g. daily indices (see: index templates), and leaving them alone. Each measurement document would look something like

{"product_family": "widget", # keyword
 "timestamp": "2022-08-23",  # date
 "sales": 798137,
 "price": "and so on"}

This document would be inserted into the index yourindex_20220823.

You can have Elasticsearch run roll-up jobs for aggregating historical data, and set up index lifecycle management so that indices older than your retention period get deleted. This is very fast, way faster than running delete-by-query requests to remove all documents with insertionDate > -2yrs.

Now, we have the issue of storing the product category metadata. As you might have found out, ES is better at denormalized data, but it does lead to repetition and you might find your index size blowing up.

For minimizing disk usage, the trick is to tweak individual field mappings (and no, you can't rely on dynamic mapping). You can avoid storing a lot of stuff in the inverted index. See https://www.elastic.co/guide/en/elasticsearch/reference/current/tune-for-disk-usage.html. I'd need to see your current mapping to check if there are any obvious gains to be made here.

Lastly, a feature that I've never tried out is to move older data (again, having daily indices helps here) to slower storage modes. See cold/frozen storage tiers.

Related