How to Increase Flask RestAPI Concurrent Request Performance

Viewed 37

I'm using "gunicorn" and "gevent" to serve flask APIs with 4 workers. When I do AB (apache-benchmark) test to a single API response time per request is 60ms. However, when I call the second API from the first one response time per request is going up to 358ms. As I share the code sample of APIs, they only return a "Hi there! " response. What could be the reason for this response time increase?

First API

import requests
from flask import Flask, request

api_url = f'http://127.0.0.1:4000/'

app = Flask(__name__)


@app.route('/', methods=['GET'])
def index():
    return 'Hi there! '


@app.route('/test', methods=['GET'])
def test():
    resp = requests.get(f'{api_url}')
    response = resp.text
    return response

Second API

from flask import Flask, request

app = Flask(__name__)


@app.route('/', methods=['GET'])
def index():
    return 'Hi there!'

ab test on "/" path with 1000 concurrent 10000 request

Benchmarking 0.0.0.0 (be patient)
Completed 1000 requests
Completed 2000 requests
Completed 3000 requests
Completed 4000 requests
Completed 5000 requests
Completed 6000 requests
Completed 7000 requests
Completed 8000 requests
Completed 9000 requests
Completed 10000 requests
Finished 10000 requests


Server Software:        gunicorn
Server Hostname:        0.0.0.0
Server Port:            3000

Document Path:          /
Document Length:        10 bytes

Concurrency Level:      1000
Time taken for tests:   0.602 seconds
Complete requests:      10000
Failed requests:        0
Total transferred:      1630000 bytes
HTML transferred:       100000 bytes
Requests per second:    16604.12 [#/sec] (mean)
Time per request:       60.226 [ms] (mean)
Time per request:       0.060 [ms] (mean, across all concurrent requests)
Transfer rate:          2643.04 [Kbytes/sec] received

Connection Times (ms)
              min  mean[+/-sd] median   max
Connect:        0    1   3.2      0      18
Processing:     2   55  11.7     59      60
Waiting:        1   55  11.7     59      60
Total:          8   56   9.1     59      67

Percentage of the requests served within a certain time (ms)
  50%     59
  66%     59
  75%     59
  80%     59
  90%     60
  95%     60
  98%     60
  99%     60
 100%     67 (longest request)

ab test on "/test" path by calling the second API with 1000 concurrent 10000 request response result per request -> 358.406 ms

Benchmarking 0.0.0.0 (be patient)
Completed 1000 requests
Completed 2000 requests
Completed 3000 requests
Completed 4000 requests
Completed 5000 requests
Completed 6000 requests
Completed 7000 requests
Completed 8000 requests
Completed 9000 requests
Completed 10000 requests
Finished 10000 requests


Server Software:        gunicorn
Server Hostname:        0.0.0.0
Server Port:            3000

Document Path:          /test
Document Length:        9 bytes

Concurrency Level:      1000
Time taken for tests:   3.584 seconds
Complete requests:      10000
Failed requests:        0
Total transferred:      1610000 bytes
HTML transferred:       90000 bytes
Requests per second:    2790.13 [#/sec] (mean)
Time per request:       358.406 [ms] (mean)
Time per request:       0.358 [ms] (mean, across all concurrent requests)
Transfer rate:          438.68 [Kbytes/sec] received

Connection Times (ms)
              min  mean[+/-sd] median   max
Connect:        0    1   3.1      0      13
Processing:    16  339  54.2    358     446
Waiting:        3  339  54.2    358     446
Total:         16  340  53.4    359     458

Percentage of the requests served within a certain time (ms)
  50%    359
  66%    364
  75%    367
  80%    369
  90%    372
  95%    376
  98%    379
  99%    401
 100%    458 (longest request)

0 Answers
Related