Want to run Specific node or group of nodes and capture the output into a variable in kedro jupyter lab

Viewed 229

I am new to kedro, I am trying to run Spaceflights tutorial. I want to run the complete data_processing_pipeline 'dp', and capture the output in a dataframe. I am running it on Jupyter Lab. I used the following command: model_input_table = session.run(pipeline_name='dp') or model_input_table = context.run(pipeline_name='dp')

I even tried running a specific node to capture the output returned into a variable.

Nothing seems working! Please help!

1 Answers

So the way to do this is to use the catalog object which (should also be in scope) and read the results of the files persisted on disk. So you can do catalog.load('companies') and inspect the DataFrame right there in your notebook.

You can also do catalog.list() to inspect which items in your catalog are available for inspection. Kedro is fundamentally dataset-centric so we always interact with data via the DataCatalog, the pipeline is just something that organises the execution order at runtime.

Related