I want to adjust my script so that I can use a .csv file with multiple rows as input instead of using one example row. I've already did some thinking and I guess this can be solved using a for-loop but since I'm not really adept in using for-loops. I've added my Jupyter Notebook script below. Hopefully somebody can help me out with this one!
Context I'm currently working on a pipeline of code that extracts surgery entities out of medical notes by using Named Entity Recognition (NER) based on a pre-trained BERT model, and then using the Entity Linker to fetch a surgery type (T061) from the Unified Medical Language System (UMLS).
Current situation Currently I have a working script in Jupyter Notebook that works for one example sentence, but can not parse .csv file for example with multiple rows in it.
Desired situation I want this script to be able to parse a .csv file with multiple rows and each rows being a medical note. The needed output is the original row (medical note) and a added column that shows 'Yes'/'No' based on the presence of the surgery entity in that row.
Script
%%capture
pip install https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.4.0/en_core_sci_scibert-0.4.0.tar.gz
import pandas as pd
import spacy
from scispacy.abbreviation import AbbreviationDetector
from scispacy.umls_linking import UmlsEntityLinker
from spacy import displacy
%%capture
nlp = spacy.load("en_core_sci_scibert")
nlp.add_pipe("scispacy_linker", config={"resolve_abbreviations": True, "linker_name": "umls"})
doc = nlp("Spinal and bulbar muscular atrophy (SBMA) is an \
inherited motor neuron disease caused by the expansion \
of a polyglutamine tract within the androgen receptor (AR).\
SBMA can be caused this easily, and also solved with excision. Also, the natalizumab \
caused a very strong delirium. The patient also has decreased cognition, where the\
MMSE score is 20. The patient also is not native speaker Dutch, and has a languare barrier.\
Patient is living alone.")
linker = nlp.get_pipe("scispacy_linker")
entity = doc.ents[0]
concept_entity = []
all_umls_data = []
for entity in doc.ents:
print("Name: ",entity)
highest_umls_ent = entity._.umls_ents[0]
concept_entity.append((highest_umls_ent[0],entity))
umls_data = linker.umls.cui_to_entity[highest_umls_ent[0]]
print(umls_data)
all_umls_data.append(umls_data)
print('\n')
umls_df = pd.DataFrame(all_umls_data)
umls_df['types'] = umls_df['types'].str[0]
df_surgery = umls_df.loc[umls_df['types'] == 'T061']
df_surgery