Based on AWS documentation, docs, I've set up a batch inference job. however, once we choose the instance type and instance count, bare minimum , does sagemaker choose optimal plan to process jobs, say if there are more than one files , and if resource are available, can those files in parallel?
from sagemaker.transformer import Transformer
tr = Transformer(model_name='custom_model',instance_count=2, instance_type='ml.m4.xlarge')