I'm new to Graph based CNNs and hence tried to explore the area with a side project but with preparing a dataset from credit-streams. I'm trying to replicate the dataset that is been curated in the Cora Dataset.
I have prepared a dataframe post pre-processing which is similar to the cora.content part of the dataset as follows:
dict form of df:
{'New Context': {'900': 'Settlements', '427': 'Settlements', '219': 'MFA', '1101': 'Settlements', '748': 'Settlements'}, 'SETTLEMENT DATE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CASH ACCOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'ISIN': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'TRADE DATE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'PRICE CFA': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SECURITY NAME': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CLEARING BIC': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SCA': {'900': 0, '427': 0, '219': 0, '1101': 1, '748': 0}, 'TRADE TYPE': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'CLIENT NAME': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'MARKET': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SAFEKEEP ACCOUNT': {'900': 1, '427': 1, '219': 0, '1101': 0, '748': 0}, 'PORTFOLIO ACCOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}, 'SETTLEMENT AMOUNT': {'900': 0, '427': 0, '219': 0, '1101': 0, '748': 0}}
New Context SETTLEMENT DATE CASH ACCOUNT ISIN TRADE DATE PRICE CFA SECURITY NAME CLEARING BIC SCA TRADE TYPE CLIENT NAME MARKET SAFEKEEP ACCOUNT PORTFOLIO ACCOUNT SETTLEMENT AMOUNT
900 Settlements 0 0 0 0 0 0 0 0 0 0 0 1 0 0
427 Settlements 0 0 0 0 0 0 0 0 0 0 0 1 0 0
219 MFA 0 0 0 0 0 0 0 0 0 0 0 0 0 0
I want to solve a similar problem, classify 'New Context' based on the entries in the other 14 columns which has binary values.
In order to prepare the edges set of the data, I have stacked the dataframe based on entries:
stacked = df.set_index('New Context').stack()
edges = stacked.index.tolist()
square_edges = pd.DataFrame(edges)
square_edges.columns =['source', 'target']
square_edges=
source target
0 Settlements SETTLEMENT DATE
1 Settlements CASH ACCOUNT
2 Settlements ISIN
3 Settlements TRADE DATE
4 Settlements PRICE CFA
cora.cites has prepared similar data where the first column identifies the cited paper, and the second column identifies the paper that cites it. The first three lines of the file look like:
target source
0 35 1033
1 35 103482
2 35 103515
3 35 1050679
I'm failing to logically imbibe this part, makes sense for the cora dataset, however, in order to solve the similar problem that I want to solve, am I going the right way? How do i transform my 'square_edges' into more of a 'cora.cites'.
Any input is highly appreciated.