What is the best practice for placement of incremental models in DBT pipelines?

Viewed 235

I am currently getting to grips with DBT to populate and keep our data warehouse up to date and am starting to look at where we would benefit from incremental load to reduce compute resource and processing time.

One thing I can't really grasp from the DBT documentation is where incremental models should sit in the pipeline and I'm wondering if anyone out there has any real world examples they can share.

Right now the pipeline consists of source (materialized as views) -> staging (table) -> production (table) stages. Things I read online seem to suggest that incremental models should sit as close to the source as possible, which makes sense, but this means that we'll be storing table data at every stage in the pipeline which almost feels unnecessary. If I move the incremental load to being in the staging models, this removes this problem, but means the views are always returning the full dataset. If I change the source stage to use an incremental load strategy, should I also be making the other stages in the pipeline incremental?

I realise this is a bit of an open ended question so apologies for that - I'm just looking for some pointers.

0 Answers
Related