I am trying to build a matrix where the first row will be a part of speech, first column a sentence. values in the matrix should show the number of such POS in a sentence.
So I am creating POS tags in this way:
data = pd.read_csv(open('myfile.csv'),sep=';')
target = data["label"]
del data["label"]
data.sentence = data.sentence.str.lower() # All strings in data frame to lowercase
for line in data.sentence:
Line_new= nltk.pos_tag(nltk.word_tokenize(line))
print(Line_new)
The output is:
[('together', 'RB'), ('with', 'IN'), ('the', 'DT'), ('6th', 'CD'), ('battalion', 'NN'), ('of', 'IN'), ('the', 'DT')]
How can I create a matrix which I have described above from such output?
UPDATE: The desired output is
NN VB IN VBZ DT
I was there 1 1 1 0 0
He came there 0 0 1 1 1
myfile.csv:
"A child who is exclusively or predominantly oral (using speech for communication) can experience social isolation from his or her hearing peers, particularly if no one takes the time to explicitly teach them social skills that other children acquire independently by virtue of having normal hearing.";"certain"
"Preliminary Discourse to the Encyclopedia of Diderot";"certain"
"d'Alembert claims that it would be ignorant to perceive that everything could be known about a particular subject.";"certain"
"However, as the overemphasis on parental influence of psychodynamics theory has been strongly criticized in the previous century, modern psychologists adopted interracial contact as a more important determinant than childhood experience on shaping people’s prejudice traits (Stephan & Rosenfield, 1978).";"uncertain"
"this can also be summarized as a distinguish behaviour on the peronnel level";"uncertain"