How to vary test data running Locust performance test from huge .txt file

Viewed 565

I'm trying to resolve 2 issues, and check the performance of my API.

Now, I have a file with a few million records (1 line, 1 record), and I need to fire all of them to my API. Loading the whole file in memory could work, even though I'd rather load batches of it. (That's 1 problem)

Second, I need to send X requests in and each request needs to have a different test data due to data being coordinates, and path recording (e.g. similar to strava when you are tracking a ride).

I have tried reading file in chunks, but not sure how to pass the data then to the method. Tried to loop within the method, but, that failed as well as Locust execute task automagically.

I cannot find anywhere in the documentation how to vary the data.

My idea was - read 1000 lines from file. Perform 1 request for each 1000 lines. Read next 1000 lines...

And do that as long as there are lines in the file.

I have this:

from locust import HttpUser, task, between

test_data_filename = '/Users/i/IdeaProjects/TestIntelliJ/test10.txt'

#loading test data set from .txt file - trying to read here in chunks, but I need single line from a chunk
with open(test_data_filename) as f:
    test_data = f.readlines(100)

class HtpAPIUser(HttpUser):
        host = 'myserver'
        wait_time = between(3,5)
        i = 0

        # Was thinking of running here the task, but not sure how to do it?
        for line in test_data:
            test_data_line = line

        @task()
        # Thought of passing different test data line into the method, but not sure is that the way?
        def tele_endpoint(self, test_data_line):
            headers = {'Content-Type': 'application/xml'}
            response = self.client.post('/push/data/points', data=test_data_line, headers=headers, verify=False)
            print(response)
2 Answers

Reading a new line from the same file handle in the task should be fine, you dont need to do your own buffering. Files are buffered by the OS and/or python so reading another line is very fast. You might want to have a look at locus-plugins CSVReader https://github.com/SvenskaSpel/locust-plugins/blob/master/examples/csvreader_ex.py (just as an example, not that it is faster or anything)

I think you should be able to do 1000 lines/s at least (probably a lot more), and if that is not enough you might look in to splitting the file and running one Locust worker process per file.

The challenge that you will have is that you will have a great number of users which need access to the same file and data. You also might need data to be exclusive to a single user (no reuse). Using a service, specifically a queue system works well for this.

if you are in a cloud provider for your tests, all of them offer a queue service with an HTTP interface. Look at Amazon Simple Queue Service (SQS) as an example of this. You can offload the file from your load generators and have it managed inside of the service. If my memory is correct then you won't even incur a cost with Amazon until you hit something on the order of one million queue requests.

If you are doing this in house, then a solution like RabbitMQ works exceptionally well to service your load generators. You want this queue service to be independent of your application under test, so install it on a dedicated host. You want to avoid the in-house generators with an in cloud queue as you will incur byte costs in/out of the cloud provider.

By default the queue service will provide you uniqueness across your virtual users, lower overhead without the file having to be loaded and accessed to the local file system (also a drag on the load generator with hundreds vying for a file access), with a solution designed for multiple user access.

If you need to re-use the data then you can simply push it to the back of the queue. If you want a self terminating tests then as soon as your user is not able to pull tests data because of an empty queue, then after that iteration ends terminate the virtual user.

Related