The OpenMP API does not have a specific feature to deal with this situation. If you know from your system that the P-cores are are cores 0-7 and the E-cores are 8-15, then you can do the following to restrict your OpenMP threads to run only on the P-cores:
In the shell (bash-like):
export OMP_PLACES=0-7
export OMP_PROC_BIND=true
Then, in your code do something like this (actually no change :-)):
!$omp parallel do
do ...
...
end do
!$omp end parallel do
Or in C/C++ syntax:
#pragma omp parallel for
for(...) {...}
If you want to span all P and E cores at the same code, you will have to accept some sort of load imbalance, but you could still make good use of them.
In the shell (bash-like):
export OMP_PLACES=cores
export OMP_PROC_BIND=true
Then, in your Fortran code:
!$omp parallel do schedule(nonmonotonic:dynamic,chunkz)
do ...
...
end do
!$omp end parallel do
Or in C/C++ syntax:
#pragma omp parallel for schedule(nonmonotonic:dynamic,chunksz)
for(...) {...}
In that case, I would anticipate that a dynamic schedule with a chunk size of chunksz would a good solution so that the (faster) P-cores get some more work compared to E-cores.
If you use OpenMP tasks, then you might still want to pin the OpenMP threads to cores, but since OpenMP tasks are dynamically scheduled to idling OpenMP threads, you get automatic load balancing. As a rough rule of thumb you should make sure that you create 10x more tasks than you have OpenMP threads.