How to prevent gunicorn eventlet workers timeout exiting in Flask-APScheduler long running tasks?

Viewed 613

I have a small Flask app, running on a server intended for disk analysis. It shows s.m.a.r.t. parameters of disks and Im trying to add disk formatting feature to it. Doesnt matter what kind of operation it is (secure erace or just "shred"), but it takes quite a long time. At first I tried sync workers with command:

./venv/bin/gunicorn -b :5000 --access-logfile - --error-logfile - --log-file - -p myapp.pid -w 5 run:app

With scheduler tasks creating part:

if disk_type == 'NVMe':
    scheduler.add_job(id=f'{serial}_formatting',
                      func=format_nvme,
                      args=[name, serial])
elif disk_type == 'HDD':
    scheduler.add_job(id=f'{serial}_formatting',
                      func=format_hdd,
                      args=[name, serial])
...

And formatting function:

def format_hdd(name, serial):
    from run import app
    with app.app_context():
        print(f'[{datetime.now()}] initiated ssd formatting. name={name}, serial={serial}')
        record_in_db = Disk.query.filter(Disk.serial == serial).first()
        record_in_db.formatting_status = 1
        db.session.commit()
        info, error = Popen(f'wipefs -a {name}', shell=True, stdout=PIPE,
                        stderr=PIPE).communicate()
        process = Popen(f'shred -v {name}', shell=True, stdout=PIPE, stderr=PIPE,
                    encoding='utf-8', errors='replace', bufsize=1)
        while True:
            realtime_output = process.stdout.readline()
            if realtime_output == '' and process.poll() is not None:
                break
            if realtime_output:
                record_in_db.formatting_last_msg = realtime_output
                db.session.commit()

As you can see, i want to store process messages to be able to see the progress of formatting, its an important part.

run.py is:

from flaskapp import create_app

app = create_app()
from flaskapp import routes

if __name__ == '__main__':
    app.run(debug=True, host='0.0.0.0', use_reloader=True)

And in create_app() it just initializes scheduler/db like that:

    if not scheduler.running:
        scheduler.init_app(app)
        scheduler.start()
    ...
    return app

Starting flaskapp with gunicorn command like above in most cases results with one/few workers timeout exiting, because sometimes there is no any command output at console, while it is executing:

[11209] [CRITICAL] WORKER TIMEOUT (pid:11214)

Simple decision - bigger timeout, but formatting is really long operation... This stackoverflow question answer gave me an idea to use eventlet workers, but it doesnt help and I still getting mistakes:

[2020-03-23 14:13:33 +0800] [29136] [CRITICAL] WORKER TIMEOUT (pid:29142)
Exception in worker
Traceback (most recent call last):
  File "/usr/lib/python3.7/concurrent/futures/thread.py", line 78, in _worker
    work_item = work_queue.get(block=True)
  File "/home/user/flaskapp/venv/lib/python3.7/site-packages/gunicorn/workers/base.py", line 201, in handle_abort
    sys.exit(1)
SystemExit: 1

How can I fix it and always be able to see, that operation has ended? (All these commands finishing in any way, but I just cant see when, because of died workers) Or maybe there are more suitable ways to implement this functionality?

0 Answers
Related