I working on a Python project that has a DataFrame like this:
row1 = ['BBB', 'AAA', 'CCC']
row2 = ['CCC']
row3 = ['DDD', 'AAA']
row4 = ['DDD', 'BBB', 'AAA', 'EEE', 'CCC']
row5 = ['EEE', 'AAA', 'EEE', 'CCC']
data = {'List': [row1, row2, row3, row4, row5],
'Path_length': [3, 1, 2, 5, 4]}
df = pd.DataFrame(data)
which leads to:
| | List | Path_length |
| 0 | ['BBB', 'AAA', 'CCC'] | 3 |
| 1 | ['CCC'] | 1 |
| 2 | ['DDD','AAA'] | 2 |
| 3 | ['DDD','BBB', 'AAA', 'EEE', 'CCC'] | 5 |
| 4 | ['EEE', 'AAA', 'EEE', 'CCC'] | 4 |
And the task consists of generating the following DataFrame:
| | Content | Unique | Started | Middleway | Finished |
| 0 | AAA | 0 | 0 | 3 | 1 |
| 1 | BBB | 0 | 1 | 1 | 0 |
| 2 | CCC | 1 | 0 | 0 | 3 |
| 3 | DDD | 0 | 2 | 0 | 0 |
| 4 | EEE | 0 | 1 | 2 | 0 |
where the columns contain the following:
- Content: the elements found in the List
- Unique: the number of times that the element appears alone in the list
- Started: the number of times that the element appears at the beginning
- Finished: the number of times that the element appears at the end
- Middleway: the number of times that the element appears between the beginning and the end.
I kind of got a solution for this task, but since I used a lot of loop functions, I can't apply the algorithm to a larger database because the time processing is too high. Could you help me by suggesting a code that solves this task?