I'm using "gunicorn" and "gevent" to serve flask APIs with 4 workers. When I do AB (apache-benchmark) test to a single API response time per request is 60ms. However, when I call the second API from the first one response time per request is going up to 358ms. As I share the code sample of APIs, they only return a "Hi there! " response. What could be the reason for this response time increase?
First API
import requests
from flask import Flask, request
api_url = f'http://127.0.0.1:4000/'
app = Flask(__name__)
@app.route('/', methods=['GET'])
def index():
return 'Hi there! '
@app.route('/test', methods=['GET'])
def test():
resp = requests.get(f'{api_url}')
response = resp.text
return response
Second API
from flask import Flask, request
app = Flask(__name__)
@app.route('/', methods=['GET'])
def index():
return 'Hi there!'
ab test on "/" path with 1000 concurrent 10000 request
Benchmarking 0.0.0.0 (be patient)
Completed 1000 requests
Completed 2000 requests
Completed 3000 requests
Completed 4000 requests
Completed 5000 requests
Completed 6000 requests
Completed 7000 requests
Completed 8000 requests
Completed 9000 requests
Completed 10000 requests
Finished 10000 requests
Server Software: gunicorn
Server Hostname: 0.0.0.0
Server Port: 3000
Document Path: /
Document Length: 10 bytes
Concurrency Level: 1000
Time taken for tests: 0.602 seconds
Complete requests: 10000
Failed requests: 0
Total transferred: 1630000 bytes
HTML transferred: 100000 bytes
Requests per second: 16604.12 [#/sec] (mean)
Time per request: 60.226 [ms] (mean)
Time per request: 0.060 [ms] (mean, across all concurrent requests)
Transfer rate: 2643.04 [Kbytes/sec] received
Connection Times (ms)
min mean[+/-sd] median max
Connect: 0 1 3.2 0 18
Processing: 2 55 11.7 59 60
Waiting: 1 55 11.7 59 60
Total: 8 56 9.1 59 67
Percentage of the requests served within a certain time (ms)
50% 59
66% 59
75% 59
80% 59
90% 60
95% 60
98% 60
99% 60
100% 67 (longest request)
ab test on "/test" path by calling the second API with 1000 concurrent 10000 request response result per request -> 358.406 ms
Benchmarking 0.0.0.0 (be patient)
Completed 1000 requests
Completed 2000 requests
Completed 3000 requests
Completed 4000 requests
Completed 5000 requests
Completed 6000 requests
Completed 7000 requests
Completed 8000 requests
Completed 9000 requests
Completed 10000 requests
Finished 10000 requests
Server Software: gunicorn
Server Hostname: 0.0.0.0
Server Port: 3000
Document Path: /test
Document Length: 9 bytes
Concurrency Level: 1000
Time taken for tests: 3.584 seconds
Complete requests: 10000
Failed requests: 0
Total transferred: 1610000 bytes
HTML transferred: 90000 bytes
Requests per second: 2790.13 [#/sec] (mean)
Time per request: 358.406 [ms] (mean)
Time per request: 0.358 [ms] (mean, across all concurrent requests)
Transfer rate: 438.68 [Kbytes/sec] received
Connection Times (ms)
min mean[+/-sd] median max
Connect: 0 1 3.1 0 13
Processing: 16 339 54.2 358 446
Waiting: 3 339 54.2 358 446
Total: 16 340 53.4 359 458
Percentage of the requests served within a certain time (ms)
50% 359
66% 364
75% 367
80% 369
90% 372
95% 376
98% 379
99% 401
100% 458 (longest request)