We currently use a ElasticsearchSinkConnector within our kafka-connect-cluster to move data from Kafka to Elastic. We utilize JSON-Schemas and the SchemaRegistry in our setup.
In this use-case our data is very dynamic and thus can only be defined as a type: object in our JSON-Schema. When sending the following message,
{
data: {
propOne: "hi"
āā propTwo: "hi"
},
type: "Person",
}
with the following schema,
{
required: ["data", "type"],
properties: {
"type": { type: "String" }
"data": { type: "Object" }
},
}
we run into the following problem:
Our ElasticSinkConnector only gets a connect-struct that contains a type with the corresponding string-value, but always an empty data-object. This makes sense, as based on the schema, this is the maximum structure it can infer. However, we would of course like to move the whole original data-object to our elasticsearch instance and not only an empty object.
The Elasticsearch Sink Connector supports JSON (schemaless) data output from Apache Kafka topics. But if we use the org.apache.kafka.connect.json.JsonConverter converter to ignore the schema we, of course, get the issue, that the subjectid-byte cannot be interpreted. Thus I'm looking for something like org.apache.kafka.connect.json.JsonConverter that ignores the subjectid and magicbyte, takes the JSON however it comes and dumps it into elastic.
Alternatives we are aware of: we could of course use a schemaless JSON-event for this use-case/topic. We still would like to avoid this, as our producer does currently not support it.
