How to reduce wand memory usage?

Viewed 2391

I am using wand and pytesseract to get the text of pdfs uploaded to a django website like so:

image_pdf = Image(blob=read_pdf_file, resolution=300)
image_png = image_pdf.convert('png')

req_image = []
final_text = []

for img in image_png.sequence:
    img_page = Image(image=img)
    req_image.append(img_page.make_blob('png'))

for img in req_image:
    txt = pytesseract.image_to_string(PI.open(io.BytesIO(img)).convert('RGB'))
    final_text.append(txt)

return " ".join(final_text)

I have it running in celery in a separate ec2 server. However, because the image_pdf grows to approximately 4gb for even a 13.7 mb pdf file, it is being stopped by the oom killer. Instead of paying for higher ram, I want to try to reduce the memory used by wand and ImageMagick. Since it is already async I don't mind increased computation times. I have skimmed this: http://www.imagemagick.org/Usage/files/#massive, but am not sure if it can be implemented with wand. Another possible fix is a way to open a pdf in wand one page at a time rather than putting the full image into RAM at once. Alternatively, how could I interface with ImageMagick directly using python so that I could use these memory limiting techniques?

4 Answers

I was also suffering from memory leaks issues. After some research and tweaking the code implementation, my issues were resolved. I basically worked correctly using with and destroy() function.

In some cases I could use with to open and read the files, as in the example below:

with Image(filename = pdf_file, resolution = 300) as pdf:

This case, using with, the memory and tmp files are correctly managed.

And in another case I had to use the destroy() function, preferably inside a try / finally block, as below:

try:
    for img in pdfImg.sequence:
    # your code
finally:
    pdfImg.destroy()

The second case, is an example where I cann't use with because I had to iterate the pages through the sequence, so, I already had the file open and was iterating your pages.

This conbination of solution resolved my problems with memory leaks.

The code from @emcconville works, and my temp folder is not filling up with magick-* files anymore

I needed to Import ctypes and not cstyles

I also got the error mentioned by @kerthik

solved it by saving the image and loading it again, it is properly also possible to save it to memory

from PIL import Image as PILImage

...
context.save(filename="temp.jpg")
text = pytesseract.image_to_string(PILImage.open("temp.jpg"))`

EDIT I found the in memory conversion on How to convert wand.image.Image to PIL.Image?

img_buffer = np.asarray(bytearray(context.make_blob(format='png')),dtype='uint8')
bytesio = io.BytesIO(img_buffer)
text = ytesseract.image_to_string(PILImage.open(bytesio),lang="dan")

I run into a similar issue.

Found this page interesting: http://www.imagemagick.org/script/architecture.php#tera-pixel

And how to limit the amount of memory used by ImageMagick through wand: http://docs.wand-py.org/en/latest/wand/resource.html

Just adding something like:

from wand.resource import limits

# Use 100MB of ram before writing temp data to disk.
limits['memory'] = 1024 * 1024 * 100

It may increase the computation time (but like you, I don't mind too much) and I actually did not notice so much difference.

I confirmed using Python's memory-profiler that it is working as expected.

Related