Is it possible to pass a group of rows together, instead of each row individually, to a UDF? Assuming that I am able to, I want to run custom code inside MyFunction which writes to the database after performing some manipulation.
Below is what I have implemented for passing each row individually.
static void Main(string[] args)
{
spark.Udf().Register<Row, int>("MyUDF", (text) => MyFunction(text));
spark.Read()
.Option("header", true)
.Schema(schema)
.Csv(@"C:myfile.csv")
.CreateOrReplaceTempView("Tweets");
var sqlDf = spark.Sql("SELECT MyUDF(Struct(*)) as c12 FROM Tweets");
}
public static int MyFunction(Row text)
{
//Process data complex logic
}
What I am trying to achieve is that there is already an ETL job written in C# which has loads of transformation logic after which it saves to destination database. It takes a huge collection of object and performs this transformation. How can I parallelize this using spark?