I am using AWS Textract for Form and Table extraction using following code. For some pdf it extracts forms from all the pages but for some pdf is extracts only first page. While using the textract user interface it extracts all the pages. What could be the reason for this??
I am using following code which is available on aws.
def create_client(access_key, secret_key):
return boto3.client('textract',region_name='us-east-2',
aws_access_key_id=access_key,
aws_secret_access_key=secret_key)
def isJobComplete(jobId):
client = create_client(access_key, secret_key)
response = client.get_document_analysis(JobId=jobId)
status = response["JobStatus"]
print("Job status: {}".format(status))
while(status == "IN_PROGRESS"):
time.sleep(2)
response = client.get_document_analysis(JobId=jobId)
status = response["JobStatus"]
print("Job status: {}".format(status))
return status
def getJobResults(jobId):
client = create_client(access_key, secret_key)
response = client.get_document_analysis(JobId=jobId)
return response
Edited : It looks like its related to response size. The size is almost fixed.
Can anyone help me with this ?