Currently I have been working on code snippet classification using Code2vec models which I have trained on a set of python code snippet, my idea was to produce embeddings for each code snippet and attach the label to it and use it further for the final classification for instance the arff file for weka will look like the following:
relation XYZ
@Attributes @Class@ {buggy,non_buggy}
@Attributes index1 real
.........
.........
.........
.........
.........
@Attributes index380 real
@data
buggy, 0.28600096702575684, -0.03643874451518059, -0.06801733374595642,.......
..................
..................
..................
non_buggy, 0.4966501295566559, -0.38083720207214355, -0.378182053565979,.......
For the classification, I split my full dataset into 80% for training and 20% for testing using the percentage split option provided by WEKA. I got a precision of 99% I was surprised though I tried to do other splits for instance 1% for training and 99% for testing however the performance is still good almost 99% precision which I found not logic in this case.
do I have to change anything before the second split ? does anybody have experienced this issue while working with embeddings in WEKA?