I wrote simple fastapi server in which I want to do performance testing. Which means that I would like to test async vs. sync endpoint, but I got always the same results. Which concept I don't understand or what I am doing wrong?
from fastapi import FastAPI
import time
import asyncio
app = FastAPI()
@app.get("/hello")
def say_hello():
time.sleep(5)
@app.get("/hello-async")
async def say_hello():
await asyncio.sleep(5)
I run server like this: gunicorn main:app --workers=4 --worker-class uvicorn.workers.UvicornWorker --bind 0.0.0.0:8000
And do performance testing:
ab -n 80 -c 8 http://localhost:8000/hello-async
ab -n 80 -c 8 http://localhost:8000/hello
And got this kind of results:
This is ApacheBench, Version 2.3 <$Revision: 1879490 $>
Copyright 1996 Adam Twiss, Zeus Technology Ltd, http://www.zeustech.net/
Licensed to The Apache Software Foundation, http://www.apache.org/
Benchmarking localhost (be patient).....done
Server Software: uvicorn
Server Hostname: localhost
Server Port: 8000
Document Path: /hello-async
Document Length: 4 bytes
Concurrency Level: 8
Time taken for tests: 55.016 seconds
Complete requests: 80
Failed requests: 0
Total transferred: 10240 bytes
HTML transferred: 320 bytes
Requests per second: 1.45 [#/sec] (mean)
Time per request: 5501.634 [ms] (mean)
Time per request: 687.704 [ms] (mean, across all concurrent requests)
Transfer rate: 0.18 [Kbytes/sec] received
Connection Times (ms)
min mean[+/-sd] median max
Connect: 0 0 0.1 0 1
Processing: 5000 5001 0.6 5001 5003
Waiting: 5000 5001 0.6 5001 5003
Total: 5000 5001 0.6 5001 5003
Percentage of the requests served within a certain time (ms)
50% 5001
66% 5002
75% 5002
80% 5002
90% 5002
95% 5002
98% 5003
99% 5003
100% 5003 (longest request)
I wanted to simulate I/O operation, if I call API or database I want to se how much faster is async approach. What happened now I got the same time taken for tests.
If I would do CPU intensive operation I would understand that I will got the same result. What I am doing wrong?
Also in syncing approach the web server is limited to number threads which are spawned, but then what is theoretical limit of async approach? Also how can I test limits number of spawned threads in syncing approach and in async approach what is the limit of event loop.