How can Druid be configured in Dataproc?

Viewed 188

Now that Druid is made an optional component of Google Cloud Dataproc (https://cloud.google.com/dataproc/docs/concepts/components/druid), I am wondering how Druid configuration can be performed from the Dataproc cluster creation? I have tried the following gcloud command:

%gcloud dataproc clusters create test1 --region=us-east1 --zone=us-east1-b -- 
num-masters=1 --num-workers=2 --optional-components=ZOOKEEPER,DRUID -- 
properties=druid:druid.storage.type=google,...

But it returns an error:

Property 'druid:druid.storage.type' has an unsupported prefix

Apparently druid is not a valid prefix. Then how can I configure Druid in Dataproc?

Thanks.

2 Answers

Druid is still in alpha stage and does not support Deep Storage or Metadata storage configuration. Only JVM properties and Druid's component's (Broker, historical etc) runtime properties are supported.

That also means only HDFS is supported as deep storage and MySql as metadata storage.

To configure Druid you can use next cluster properties prefixes when creating a Dataproc cluster with Druid:

druid-broker:<property-name>=<value>
druid-broker-jvm:<property-name>=<value>
druid-broker-runtime:<property-name>=<value>
druid-coordinator:<property-name>=<value>
druid-historical:<property-name>=<value>
druid-historical-jvm:<property-name>=<value>
druid-historical-runtime:<property-name>=<value>
druid-middleManager:<property-name>=<value>
druid-overlord:<property-name>=<value>
druid-router:<property-name>=<value>
Related