I have the following batch configuration:
The python script (image) along with job definition is deployed via gitlab to AWS ECR and used as a task in the step function. The whole script is surrounded with try except blocks with custom error classes.
Ex of error class
class CustomException(Exception):
def __init__(self, message):
self.message = message
Ex of main function
try:
...
except Exception as ex
raise CustomException(f"An exception occurred. Exception {str(ex)}")
The step function definition for this task is:
"Batch": {
"Type": "Task",
"Resource": "arn:aws:states:::batch:submitJob.sync",
"Parameters": {
"JobDefinition": "...",
"JobName": ".....",
"JobQueue": ".....",
"Parameters": {
"parameters.$": "States.JsonToString($)"
},
"ContainerOverrides": {
"Command": [
"Ref::parameters"
]
}
},
"Retry": [
{
"ErrorEquals": [
"CustomException"
],
"IntervalSeconds": 60,
"MaxAttempts": 3
}
],
"Catch": [
{
"ErrorEquals": [
"CustomException"
],
"ResultPath": "$.errorMessage",
"Next": "Failed or Success?"
},
{
"ErrorEquals": [
"States.ALL"
],
"ResultPath": "$.errorMessage",
"Next": "Failed or Success?"
}
],
"ResultPath": null,
"Next": "...."
}
The main reason for error handle is of course an appropriate message that is being sent to custom location, and the retry strategy for custom exceptions.
The issue with this is that no matter how I try to raise the exception the step function marks is as:
"errorMessage": {
"Error": "States.TaskFailed",
"Cause": "{\"Attempts\":[{\"Container\":{\"ContainerInstanceArn\":\".....\",\"ExitCode\":1,\"LogStreamName ........}
Why is the exception override with the following Error. Is this behavior expected or is there a way to handle this?
How to overcome this States.TaskFailed