RDD order after SVD

Viewed 30

Hi everyone I am using Spark SVD in a dataset which has a column called "value" that contains some sentences and a column called "features" that contains these sentences transformed into vectors through TF-IDF. RowMatrix accepts as input only an RDD of these vectors so I need to split these two columns. My concern is: after I compute SVD I need to zip SVD results with the original sentences, but due to RDD partitioning and distribution, how can I be sure that the order of the rows is correct?

//Selecting only features column
JavaRDD<org.apache.spark.ml.linalg.Vector> features = 
dataset.select("features").javaRDD().map((Function<Row, 
org.apache.spark.ml.linalg.Vector>) row -> (org.apache.spark.ml.linalg.Vector) 
row.get(0));

//Converting from ml.linalg.Vector to mllib.linalg.Vector in order to be 
compatible with RowMatrix
JavaRDD<Vector> vectorJavaRDD = features.map(RDDUtils::conversionFromMlToMlLib);

//Converting into rdd
RDD<Vector> rdd = vectorJavaRDD.rdd();

RowMatrix rowMatrix = new RowMatrix(rdd);

SingularValueDecomposition<RowMatrix, Matrix> svd = 
rowMatrix.computeSVD(numTopics, true,1.0E-9d);

RowMatrix u = svd.U();
Vector s = svd.s();
Matrix v = svd.V();

//Creating JavaRDD in order to join with original sentences
JavaRDD<Vector> uRowsRDD = u.rows().toJavaRDD();
JavaRDD<Row> value = dataset.select("value").toJavaRDD();
//Zipping U Rows with original sentences
JavaPairRDD<Row, Vector> zip = value.zip(uRowsRDD);

Also from what I read online also the "select" operations could not preserver the order of the original dataset.

0 Answers
Related