I want to download data using 3 API codes, and I would like to multi-thread the APIs (one thread per API). Would something like this work and be safe from data races?
# DataKeys: a DataFrame object with keys to search for
len_data=size(DataKeys)[1]
array_split=[1,floor(Int, len_data/3),floor(Int,2*len_data/3), floor(Int,len_data)]
api_keys=[api1, api2, api3]
data_slices=[DataFrame(), DataFrame(), DataFrame()]
Threads.@threads for k in 1:3
key=api_keys[k]
for i in array[k]:array[k+1]
loc=DataKeys[i,:index]
r=HTTP.get(url(loc),headers)
json_r=JSON3.read(String(r.body))
temp=DataFrame(json_r[:data])
global data_slices[k]=vcat(data_slices[k], temp)
end
end
On the one hand, I feel I should be safe since every thread works on a different element of data_slices, but OTOH they're all part of the same vector.
Updated with more details on the function within the threaded for.
Things do seem to work fine on this simple example:
using DataFrames
len=120
array_split=[1,floor(Int, len/3),floor(Int,2*len/3), floor(Int,len)]
data_slices=[DataFrame(), DataFrame(), DataFrame()]
Threads.@threads for k in 1:3
for loc in array_split[k]:array_split[k+1]
temp=DataFrame(k=k,squared=loc^2,half=loc/2, thrd=Threads.threadid())
global data_slices[k]=vcat(data_slices[k],temp)
end
end
data_fin=vcat(data_slices...)
although the order of the threads is [1,3,2] which is a bit odd. Also odd is that for this simple example, I ran a threaded and a non-threaded loop and the non-threaded one is faster (len=120000):
Threaded: 14.789920 seconds (20.25 M allocations: 72.492 GiB, 28.02% gc time, 3.78% compilation time)
Non-threaded 9.614164 seconds (19.14 M allocations: 72.434 GiB, 11.00% gc time)