Piggybacking on the question I posted here, I would like to ask if it is possible to rule out certain phrase-level tags when parsing. Specifically, I am using the Stanford CorenNLP version 3.9.2 Shift-Reduce parser (for its constituency-style output) and now have experience adding ParserConstraint constraints to a ParserQuery. However, it is not immediately apparent whether a ParserConstraint can be used to (efficiently) do what I want to do.
I have the luxury of knowing that my input text is homogeneously grammatical, in complete sentences with finite matrix clauses. Consequently, any time the parser output contains a FRAG or UCP label, the parse is almost certainly inaccurate. I would like to be able to tell the parser beforehand "do not use FRAG or UCP", in an attempt to improve the quality of the output by restricting the solution set.
Is that possible? How, if so?
Snippet:
import java.io.*;
import java.util.*;
import java.text.SimpleDateFormat;
import edu.stanford.nlp.io.*;
import edu.stanford.nlp.trees.*;
import edu.stanford.nlp.ling.HasWord;
import edu.stanford.nlp.ling.TaggedWord;
import edu.stanford.nlp.process.DocumentPreprocessor;
import edu.stanford.nlp.tagger.maxent.MaxentTagger;
import edu.stanford.nlp.parser.shiftreduce.ShiftReduceParser;
import edu.stanford.nlp.parser.common.ParserQuery;
import edu.stanford.nlp.parser.common.ParserConstraint;
public class constraintTest {
// Initialize POS tagger and parser.
private static MaxentTagger meTagger = new MaxentTagger("edu/stanford/nlp/models/pos-tagger/english-left3words/english-left3words-distsim.tagger");
private static ShiftReduceParser srParser = ShiftReduceParser.loadModel("edu/stanford/nlp/models/srparser/englishSR.ser.gz");
public static void main(String[] args) throws IOException {
String text = "";
// If user passes in the name of an input file, use that.
if (args.length > 0) {
text = IOUtils.slurpFileNoExceptions(args[0]);
System.out.println(text);
// If user does not pass in a file, ask for some sentences.
} else {
System.out.println("Please enter a sentence for parsing:");
Scanner input = new Scanner(System.in);
text = input.nextLine();
}
// Create output filename and file.
String fileName = new SimpleDateFormat("'output/trees'yyyyMMdd'_'HHmmss'.txt'").format(new Date());
PrintWriter writer = new PrintWriter(fileName, "UTF-8");
// Prepare document for reading.
DocumentPreprocessor tokenizedText = new DocumentPreprocessor(new StringReader(text));
int i = 1;
for (List<HasWord> sentence : tokenizedText) {
List<TaggedWord> taggedSentence = meTagger.tagSentence(sentence);
// To parse sentences WITHOUT parser constraints:
// Tree tree = srParser.apply(taggedSentence);
// To parse sentences WITH parser constraints:
Tree tree = constrainedTree(taggedSentence);
// Print to file.
writer.println(tree);
// Print to standard out while you're at it.
System.out.println("/-/-/-/ Sentence #" + i + " /-/-/-/");
tree.pennPrint();
System.out.println();
i += 1;
}
writer.close();
}
// Takes a list of TaggedWord objects and outputs a parse tree with
// a constraint that the topmost label (below ROOT) be S.
public static Tree constrainedTree(List<TaggedWord> taggedSentence) {
int sentenceLength = taggedSentence.size();
ParserConstraint constraint = new ParserConstraint(0, sentenceLength, "S");
List<ParserConstraint> constraints = Collections.singletonList(constraint);
ParserQuery pq = srParser.parserQuery();
pq.setConstraints(constraints);
pq.parse(taggedSentence);
Tree tree = pq.getBestParse();
return tree;
}
}
As always, thanks!