I am working on an nlp problem where I have to analyze strangely formatted excel files.
There is one column with text, where each document spans multiple cells. Documents themselves are separated by empty cells. There are other columns with scores that I want to predict from the text data.
I have imported the sheets to a pandas dataframe and now I am trying to aggregate the cells belonging to each document while preserving the scores.
I have started to play around with nested loops, but I feel like it is much more complicated than necessary.
How would you approach this? Each document covers a different number of cells and documents are separated by different numbers of empty cells. To make it more complicated the scores in the columns to the right are sometimes in the same row as the first and sometimes in the same row as the last cell of the corresponding document.
I would greatly appreciate your help! There must be a simple solution.