Data Preparation for Graph based CNN (Cora like)

Viewed 75

I'm new to Graph based CNNs and hence tried to explore the area with a side project but with preparing a dataset from credit-streams. I'm trying to replicate the dataset that is been curated in the Cora Dataset.

I have prepared a dataframe post pre-processing which is similar to the cora.content part of the dataset as follows:

dict form of df:

{'New Context': {'900': 'Settlements', '427': 'Settlements', '219': 'MFA', '1101': 'Settlements', '748': 'Settlements'}, 'SETTLEMENT DATE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CASH ACCOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'ISIN': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'TRADE DATE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'PRICE CFA': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SECURITY NAME': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CLEARING BIC': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SCA': {'900': 0, '427': 0, '219': 0, '1101': 1, '748': 0}, 'TRADE TYPE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CLIENT NAME': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'MARKET': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SAFEKEEP ACCOUNT': {'900': 1, '427': 1, '219': 0, '1101': 0, '748': 0}, 'PORTFOLIO ACCOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SETTLEMENT AMOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}}




    New Context SETTLEMENT DATE CASH ACCOUNT    ISIN    TRADE DATE  PRICE CFA   SECURITY NAME   CLEARING BIC    SCA TRADE TYPE  CLIENT NAME MARKET  SAFEKEEP ACCOUNT    PORTFOLIO ACCOUNT   SETTLEMENT AMOUNT
900 Settlements 0   0   0   0   0   0   0   0   0   0   0   1   0   0
427 Settlements 0   0   0   0   0   0   0   0   0   0   0   1   0   0
219 MFA         0   0   0   0   0   0   0   0   0   0   0   0   0   0

I want to solve a similar problem, classify 'New Context' based on the entries in the other 14 columns which has binary values.

In order to prepare the edges set of the data, I have stacked the dataframe based on entries:

stacked = df.set_index('New Context').stack()
edges = stacked.index.tolist()
square_edges = pd.DataFrame(edges)
square_edges.columns =['source', 'target']

square_edges=

      source    target
0   Settlements SETTLEMENT DATE
1   Settlements CASH ACCOUNT
2   Settlements ISIN
3   Settlements TRADE DATE
4   Settlements PRICE CFA

cora.cites has prepared similar data where the first column identifies the cited paper, and the second column identifies the paper that cites it. The first three lines of the file look like:

target  source
0   35  1033
1   35  103482
2   35  103515
3   35  1050679

I'm failing to logically imbibe this part, makes sense for the cora dataset, however, in order to solve the similar problem that I want to solve, am I going the right way? How do i transform my 'square_edges' into more of a 'cora.cites'.

Any input is highly appreciated.

0 Answers
Related