Disclaimer: Just starting to learn python, forgive me for noobish questions or explanations.
EDIT- I just hacked around with the core issue I was facing with something like:
def to():
print('To...')
data = ' '.join(req(urls))
p2 = subprocess.run(['bash-tool'], text=True, capture_output=True, input=data)
p3 = subprocess.run(["another-tool"], text=True, capture_output=True, input=p2.stdout)
with open(test + '-check', 'w') as f:
subprocess.run(['sort', '-u'], text=True, stdout=f, input=p3.stdout)
another-func()
Not sure if this is the right approach. As I think it'd break if there are large number of URLs(>10k). And also would probably be quite slow?
What am I trying to achieve?
To request some urls(in a fast asynchronous way or using threading/multiprocessing), save the responses, grep certain things from them and save it to a text file. Can be easily done using bash, but I want it to be scalable for future additions and to learn the concepts of python, python networking and multithreading/processing.
The Problem:
This following function performs get requests and returns the body of pages.
import subprocess
from multiprocessing.dummy import Pool as ThreadPool
def req(urls):
print('\nRequesting...')
results = pool.map(requests.get, urls)
return [(result.text) for result in results]
I want to pass the results(or pipe to be precise) or the return value to another function which has subprocess calls to a bash command/tool.(Something like echo 'abc' | bash-tool | another-tool | sort -u >> out.txt:
def to():
print('To...')
<Cat the return value here and pass it to below subprocesses>
p2 = subprocess.run(['bash-tool'], text=True, capture_output=True, input=p1.stdout)
p3 = subprocess.run(["another-tool"], text=True, capture_output=True, input=p2.stdout)
with open(test + '-check', 'w') as f:
subprocess.run(['sort', '-u'], text=True, stdout=f, input=p3.stdout)
another-func()
So the whole program will be something like:
import os
import argparse
from multiprocessing.dummy import Pool as ThreadPool
import subprocess
import requests
def req(urls):
print('\nRequesting...')
results = pool.map(requests.get, urls)
return [(result.text) for result in results]
def to():
print('To...')
<Cat the return value here and pass it to below subprocesses>
p2 = subprocess.run(['bash-tool'], text=True, capture_output=True, input=p1.stdout)
p3 = subprocess.run(["another-tool"], text=True, capture_output=True, input=p2.stdout)
with open(test + '-check', 'w') as f:
subprocess.run(['sort', '-u'], text=True, stdout=f, input=p3.stdout)
another-func()
def another-func():
<Some more bash-fu>
def main():
with open(dom,'r') as fl:
urls=fl.read().splitlines()
req(urls)
if __name__ == '__main__':
pool = ThreadPool(int(threads))
main()
I can store the return value of req() to a variable in To(), something like data = req(urls) but how can I echo the value of "data" variable to be piped using subprocess?
Also let me know if I'm maybe approaching it with a somewhat bad logic maybe? Any help would be much appreciated!