Currently Python SDK does not support BigQuery Storage API.
We need some features of this API in our pipeline, specifically:
- Committed mode (records are available for read immediately after they are written)
- Exactly-once delivery semantics (no duplicated records)
Both of these features are not available in WriteToBigQuery PTransform.
Instead of waiting for official implementation, I am thinking about a simple workaround: custom DoFn that creates BQ Storage API Client in its setup method and then simply commits record to BQ in the process method. Records can be batched or grouped in a single commit.
Such naive approach looks like a pretty straightforward thing to do, however I am thinking what the Cons of it would be. One thing that comes to mind, is that such DoFn would become a bottleneck of a pipeline and could make scaling of it a very difficult task for the runner.
Is my thinking correct? Any other considerations?