How to work with OpenAI maximum context length is 2049 tokens?

Viewed 599

I'd like to send the text from various PDF's to OpenAI's API. Specifically the Summarize for a 2nd grader or the TL;DR summarization API's.

I can extract the text from PDF's using PyMuPDF and prepare the OpenAI prompt.

Question: How best to prepare the prompt when the token count is longer than the allowed 2049?

  • Do I just truncate the text then send multiple requests?
  • Or is there a way to sample the text to "compress" it to lose key points?
1 Answers

You have to make sure the context length is within the 2049 tokens. So for the prompt, you need to reduce the size.

OpenAI uses GPT-3 which has a context length of 2049, and text needs to fit within that context length.

I am not sure what you meant to sample the text and compress it. But if you meant how to summarize a longer text, then I would suggest you to chunk the text so that it fits within the 2049 tokens and query OpenAI that way.

Related