Our application makes batch requests to PageSpeed Insights on behalf of the user. We queue up one request per webpage and send them off at regular intervals determined by the maximum "Queries per minute" advertised by PageSpeed, which currently is 240, and may have many requests pending at any one time. We use a timer of 250 milliseconds between requests to enforce this.
As far as I can tell, this should all be within the allowed parameters of the API, yet the API frequently seems to trip over itself and periodically return batch responses of PageSpeedApi request error 500: Unable to process request. Please wait a while and try again. The requests are valid, so this appears to be the API refusing to serve responses as fast as we are sending requests.
In response to this, we implemented a finite "bucket" of pending requests, which we arbitrarily set to 50. That is, we can only have 50 outstanding requests at any one time, and when we hit that limit we will stop sending further requests until any in the bucket have been resolved, making room for more to go out. In practice, our application is limited more by this bucket than by the 240 requests per second timer. This causes our "real" rate of requests to go down to approximately 150 per minute, and does alleviate the issue somewhat, but we still occasionally get batches of requests come back with the 500 error.
Here is a screenshot of my API dashboard.

The spikes towards the left side of the graph are the result of me running the application at full speed. According to the API, I am requesting at a maximum rate of around 220, lower than the limit of 240. This proves that even the server recognizes that I'm within the limits.
Even when tightening the variables to 40 requests per minute (1500ms between requests) and 10 outstanding requests at a time, the same occurs.
We can't figure out why the server would be responding in error with these vague 500 responses. We've not seen anything explicit in the documentation. Is this a known issue, or is there some other limit such as concurrency that we're exceeding? What is the expected implementation for solutions involving batch requests?