I am working on a machine learning system using the C++ API of PyTorch (libtorch).
One thing that I have been recently working on is researching performance, CPU utilization and GPU usage of libtorch. Trough my research I understand that Torch utilizes two ways of parallelization on CPUs:
inter-opparallelizationintra-opparallelization
My main questions are:
- difference between these two
- how can I utilize
inter-opparallelism
I know that I can specify the number of threads used for intra-op parallelism (which from my understanding is performed using the openmp backend) using the torch::set_num_threads() function, as I monitor the performance of my models, I can see clearly that it utilizes the number of threads I specify using this function, and I can see clear performance difference by changing the number of intra-op threads.
There is also another function torch::set_num_interop_threads(), but it seems that no matter how many interop threads I specify, I never see any difference in performance.
Now I have read this PyTorch documentation article but it is still is unclear to me how to utilize the inter op thread pool.
The docs say:
PyTorch uses a single thread pool for the inter-op parallelism, this thread pool is shared by all inference tasks that are forked within the application process.
I have two questions to this part:
- do I need to create new threads myself to utilize the
interopthreads, or does torch do it somehow for me internally? - If I need to create new threads myself, how do I do it in C++, so that I create a new thread form the
interopthread pool?
In python example they use a fork function from torch.jit module, but I cant find anything similar in the C++ API.
