Given a table below
| X | Y | pr |
|---|---|---|
| 0 | 1 | 0.30 |
| 0 | 2 | 0.25 |
| 1 | 1 | 0.15 |
| 1 | 2 | 0.30 |
I intended to create a function to check the independence between the two variables X and Y. Note that the third column pr in the table is probability. For example P(X=0 ^ Y=1) = 0.3. Similarly, P(Y=1) = 0.3+0.15 = 0.45.
Two random variables are independent if for each possible value of x for X and for each possible value y for Y
P(X =x ^ Y = y) = P(X = x)*P(Y = y).
I understand that we can use iterrows() or itertuples() to iterate over the DataFrame. But I am getting issues to get the marginal probabilities within the for loop.
Note: Marginal probabilities are P(X = x) and P(Y = y) .
Here is my basic code
import pandas as pd
#you can use this table as an example
distr_table = pd.DataFrame({'X': [0, 0, 1, 1], 'Y': [1, 2, 1, 2], 'pr': [0.3, 0.25, 0.15, 0.3]})
x_0,x_1 = distr_table.groupby('X').pr.sum()
y_1,y_2 = distr_table.groupby('Y').pr.sum()
x_u = distr_table.X.unique()
y_u = distr_table.Y.unique()
for index, row in distr_table.iterrows():
print(row['X'], row['Y'], row['pr'])