Let's say I have a tree string for a sentence:
s = "(TOP (S (NP-TMP (NP (DT This) (NN time)) (ADVP (RP around))) (NP-SBJ (PRP they)) (VP (VBP 're) (VP (VBG moving) (ADVP (RB even) (RBR faster))))))"
I want to convert it into a bracketed structure like this:
"(((This time)(around))(they)(('re)((moving)(even faster))))"
I tried to do the following:
import nltk
s = "(TOP (S (NP-TMP (NP (DT This) (NN time)) (ADVP (RP around))) (NP-SBJ (PRP they)) (VP (VBP 're) (VP (VBG moving) (ADVP (RB even) (RBR faster))))))"
tree = nltk.Tree.fromstring(s)
out = "("
for subtrees in tree:
# there are threee subtrees
# print(len(subtree))
for i, subtree in enumerate(subtrees):
if len(subtree) > 1:
out += "("
for bracketing in range(len(subtree)):
# print(subtree[bracketing])
flattened_tree = subtree[bracketing].flatten()
flattened_string = str(flattened_tree)
flattened_string = flattened_string.replace(flattened_tree.label() + " ", "")
print(flattened_string)
out += flattened_string
if len(subtree) > 1:
out += ")"
# break
out += ")"
print(out)
# (((This time)(around))(they)(('re)(moving even faster)))
Edit:
if you see, "This" and "time" are part of the same parent, "NP". So, they become contiguous constituents, i.e. (This time).
Whereas, "around" is a single word constituent although a part of the same left sub-tree. So, it becomes ((This time)(around)).
Similarly, for the case of the right-subtree - "'re" and "'moving even faster", we see that "moving" as well as "even faster" share the same parent, "VP".
So, it becomes, (('re)((moving)(even faster)).
