Hi everyone I am using Spark SVD in a dataset which has a column called "value" that contains some sentences and a column called "features" that contains these sentences transformed into vectors through TF-IDF. RowMatrix accepts as input only an RDD of these vectors so I need to split these two columns. My concern is: after I compute SVD I need to zip SVD results with the original sentences, but due to RDD partitioning and distribution, how can I be sure that the order of the rows is correct?
//Selecting only features column
JavaRDD<org.apache.spark.ml.linalg.Vector> features =
dataset.select("features").javaRDD().map((Function<Row,
org.apache.spark.ml.linalg.Vector>) row -> (org.apache.spark.ml.linalg.Vector)
row.get(0));
//Converting from ml.linalg.Vector to mllib.linalg.Vector in order to be
compatible with RowMatrix
JavaRDD<Vector> vectorJavaRDD = features.map(RDDUtils::conversionFromMlToMlLib);
//Converting into rdd
RDD<Vector> rdd = vectorJavaRDD.rdd();
RowMatrix rowMatrix = new RowMatrix(rdd);
SingularValueDecomposition<RowMatrix, Matrix> svd =
rowMatrix.computeSVD(numTopics, true,1.0E-9d);
RowMatrix u = svd.U();
Vector s = svd.s();
Matrix v = svd.V();
//Creating JavaRDD in order to join with original sentences
JavaRDD<Vector> uRowsRDD = u.rows().toJavaRDD();
JavaRDD<Row> value = dataset.select("value").toJavaRDD();
//Zipping U Rows with original sentences
JavaPairRDD<Row, Vector> zip = value.zip(uRowsRDD);
Also from what I read online also the "select" operations could not preserver the order of the original dataset.