Apache Beam Python SDK support for BigQuery Storage API

Viewed 126

Currently Python SDK does not support BigQuery Storage API.

We need some features of this API in our pipeline, specifically:

  • Committed mode (records are available for read immediately after they are written)
  • Exactly-once delivery semantics (no duplicated records)

Both of these features are not available in WriteToBigQuery PTransform.

Instead of waiting for official implementation, I am thinking about a simple workaround: custom DoFn that creates BQ Storage API Client in its setup method and then simply commits record to BQ in the process method. Records can be batched or grouped in a single commit.

Such naive approach looks like a pretty straightforward thing to do, however I am thinking what the Cons of it would be. One thing that comes to mind, is that such DoFn would become a bottleneck of a pipeline and could make scaling of it a very difficult task for the runner.

Is my thinking correct? Any other considerations?

0 Answers
Related