I am struggling with my latest Python project and require good ideas.
I have dozens of webcrawlers written with Python that all produce similar type of output, but the code itself differs a lot by crawler. I run all these crawlers dynamically from one "main_runner" script that imports them all from a folder called "crawlers" like this:
main_runner.py:
import crawlers
The "main runner" script does a lot of database stuff and checking of data, sending emails etc.
I am constantly adding new webcrawlers and deleting some of them. That is why I would like my "main runner" script to import everything dynamically so I wouldnt have to change any code in the main script ever and that is where my frustration starts.
In the "crawlers" folder I have modules: "crawler1.py, crawler2.py" etc then of course the __init__.py file. To get the __init__.py file to populate __all__ I had to use this monstrosity:
__init__.py file:
from os.path import dirname, basename, isfile, join
import glob
modules = glob.glob(join(dirname(__file__), "*.py"))
__all__ = [ basename(f)[:-3] for f in modules if isfile(f) and not f.endswith('__init__.py')]
from . import *
And to actually run all my webcrawler modules in a loop inside my "main runner" script the only way I have discovered how to run them is this way:
main_runner.py file:
import crawlers
all_modules = crawlers.__all__
for module in all_modules:
#run a function called "crawl_data" and pass url to it.
data = getattr(getattr(crawlers,module),"crawl_data")("www.google.com")
This works, but I feel it is very messy.
So am I approaching this in a completely wrong way in the first place? Anyone have any better ideas? I would love for the ability to split my crawlers in to separate files and to run them dynamically. This would give me the ability to delete/add/edit the crawlers without editing the main script.