I'm refactoring a Python script for a data science project, and have made it fairly modularised; each function "does one thing". In order to run the script I have a "top-level" function, and my question regards what is considered the best practice for how the functions call each other.
Is it better to have:
(a) a "single trunk with branches" structure where the "top-level" function contains ~12 function calls, each of which does not call any further functions itself
def main_function():
connect_to_database()
get_data()
clean_data()
....
train_model()
test_model()
analyse_results()
(b) a "branching tree" structure where the "top-level" function has 3-4 function calls, which in turn may call functions, which may call a third level of hierarchy.
def get_and_clean_data():
connect_to_database()
get_data()
clean_data()
def train_and_deploy_model():
train_model()
test_model()
def main_function():
get_and_clean_data()
train_and_deploy_model()
analyse_results()
Function names in both examples are for illustrative purposes, not actual code.
Approach (a) is relatively easy to follow through sequentially, but either the top-level function can get very long, or the functions within it end up doing more than one thing. Approach (b) allows groupings of smaller functions that each have a clearly defined purpose, but I've found with larger projects, it can be trickier to chase bugs through a stack of callbacks.
I understand that there are no hard and fast rules but I'd like some intuition or rules of thumb about how others approach this.
The script in question is a single .py file ~300 lines long, though I'm interested in the question more broadly too, for example if the code is spread across multiple .py files. TIA