Logstash - csv output headers

Viewed 5444

I'm trying to request database with logstash jdbc plugins and returns a csv output file with headers with logstash csv plugin.

I spent a lot of time on logstash documentation but I'm still missing a point.

With the following logstash configuration, the results give me a file with headers for each row. I couldn't find a way to add the headers for only the first row in the logstash configuration.

Helps very much appreciated.

Output file

_object$id;_object$name;_object$type;nb_surveys;csat_score
2;Jeff Karas;Agent;2;2  
_object$id;_object$name;_object$type;nb_surveys;csat_score
3;John Lafer;Agent;2;2;2;2;$2;2
_object$id;_object$name;_object$type;nb_surveys;csat_score
4;Michele Fisher;Agent;2;2
_object$id;_object$name;_object$type;nb_surveys;csat_score
5;Chad Hendren;Agent;2;78

file: simple-out.conf

input {
    jdbc {
        jdbc_connection_string => "jdbc:postgresql://localhost:5432/postgres"
        jdbc_user => "postgres"
        jdbc_password => "postgres"
        jdbc_driver_library => "/tmp/drivers/postgresql/postgresql_jdbc.jar"
        jdbc_driver_class => "org.postgresql.Driver"
        statement_filepath => "query.sql"
    }
}
output {
    csv {
        fields => ["_object$id","_object$name","_object$type","nb_surveys","csat_score"]
        path => "output/%{team}/output-%{team}.%{+yyyy.MM.dd}.csv"
        csv_options => {
        "write_headers" => true
        "headers" =>["_object$id","_object$name","_object$type","nb_surveys","csat_score"]
        "col_sep" => ";"
        }
    }
}

Thanks

2 Answers

I am using dynamic file names that leverage the date of the event (index-YYYY-MM-DD.csv) so writing the headers on pipeline start was not a viable option for me.

Instead, I allowed the duplicate headers to be written and set up a cron job to run every few minutes and remove all duplicate rows and write the result back into the same file.

#!/bin/bash -xe
 for filename in /tmp/logstash/*.csv; do awk '!v[$1]++' $filename > $filename.tmp && mv -f $filename.tmp $filename; done

NOTE: This is only tested on an instance where I am pulling a couple hundred MB of data - this may not be a viable option if your data pipeline is ingesting GB per minute.

Related