I have the following pandas DataFrame:
df = pd.DataFrame({"id": [0, 1, 2, 3, 4, 5, 6],
"from": ["A", "B", "B", "D", "B", "C", "B"],
"to": ["B", "C", "D", "F", "G", "F", "E"],
"cases": [[1, 2, 44], [2, 4, 3], [5, 2], [5], [1, 7], [4], [44, 7]]
"start1": [1, 5, 4, 4, 23, 12, 8],
"start2": [4, 7, 9, 30, 26, 15, 18],
"end1": [5, 7, 11, 32, 15, 17, 21],
"end2": [9, 12, 15, 35, 17, 20, 25],})
which looks like:
id from to cases start1 start2 end1 end2
0 0 A B [1, 2, 44] 1 4 5 9
1 1 B C [2, 4, 3] 5 7 7 12
2 2 B D [5, 2] 4 9 11 15
3 3 D F [5] 4 30 32 35
4 4 B G [1, 7] 23 26 15 17
5 5 C F [4] 12 15 17 20
6 6 B E [44, 7] 8 18 21 25
I am trying to create a column adjacency_list which contains for row i the id values of rows j for which:
i["to"] == j["from"]i["cases"]overlaps withj["cases"]- the intervals (
i["end1"],i["end2"]) and (j["start1"],j["start2"]) overlap
I am trying to execute the following code to achieve this:
data["adjacency_list"] = data.apply(
lambda x: [
row["id"]
for i, row in data[(x["to"] == data["from"])].iterrows()
if ((not set(row["cases"]).isdisjoint(x["cases"])) and ((x["end1"] <= test["start1"] <= x["end2"]) or (test["start1"] <= x["end1"] <= test["start2"])))
],
axis=1,
)
The output should look like this:
id from to cases start1 start2 end1 end2 adjacency_list
0 0 A B [1, 2, 44] 1 4 5 9 [1, 2, 6]
1 1 B C [2, 4, 3] 5 7 7 12 [5]
2 2 B D [5, 2] 4 9 11 15 [3]
3 3 D F [5] 4 30 32 35 []
4 4 B G [1, 7] 23 26 15 17 []
5 5 C F [4] 12 15 17 20 []
6 6 B E [44, 7] 8 18 21 25 []
But it gives me the following error:
ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
I read a lot of other answers from users who got this error in a different context and tried replacing the and and or with & and |, but this did not work. Also, replacing the double <= comparisons with two single <='s did not help.
How to solve this?