python, requests: pass captcha with post requests

Viewed 2390

I have the following list of urls that i am tring to scrape:

base_url='https://misc.interactivebrokers.com/cstools/contract_info/index2.php?action=Details&site=GEN&conid='
conId=[280313912,289230167,285817885,289229956,256019341,289230102,289128542,289128563,289230371,289230337,287565355,285578089,287695374,285714991,287565358]
url_list=[base_url+str(x) for x in conId]

If I try to get the page for only one of the urls, it does work fine:

import requests
import lxml.html
s = requests.session()
page= s.get(url_list[0])
html = lxml.html.fromstring(page.text)
if html.xpath("//*[@id='contractSpecs']/table[5]/tr[2]/td[2]")!=[]:
    print("get_result")
else:
    print("missing")

However if I try to get multiple pages the site throws in a catcha:

for url in url_list:
    page= s.get(url)
    html = lxml.html.fromstring(page.text)
    if html.xpath("//*[@id='contractSpecs']/table[5]/tr[2]/td[2]")!=[]:
        print("get_result")
    else:
        print("missing")

Upon inspecting the captcha result i can see that the source page is as follow:

<br>
 <form type="post">
  To continue please enter the text from the image below
  <br>
   <img src="image.php?str=1wVfP3">
    <!--img src="https://chatsrv1.interactivebrokers.com/cstools/contract_info/v3.8/image.php?str=1wVfP3"-->
    <br>
     Text:
     <input name="filter" type="text">
      <input name="action" type="hidden" value="Details"/>
      <input name="conid" type="hidden" value="280313912"/>
      <input name="contract_id" type="hidden" value=""/>
      <input name="noBanner" type="hidden" value=""/>
      <input name="rescnt" type="hidden" value="100"/>
      ..............
     </input>
    </br>
   </img>
  </br>
 </form>
</br>

The captcha answer is in the tag <img src="image.php?str=1wVfP3"> ! (after str=)

I know this because i had a third person do that code for me in the past, but he did it with a module called "Grab". I am now trying to refactor that code using requests.

In order to pas that captcha, that other coder used in Grab:

try:
    text = r.xpath_text('//form')
    standart_text = 'To continue please enter the text from the image below Text:'
    if text == standart_text:
        print('captcha detected')
        cupcha = r.xpath('//img').get('src')
        cupcha_text = str(cupcha[14:])
        r.doc.set_input('filter', cupcha_text)
        r.doc.submit()

The set_input method correspond to a post request.. It seems that grab can understand what are the form fields and fill them up automatically.

My problem is that the captcha appears now only when scraping - which means i cannot use google chrome's inspect>network to check the filed names for the form.

So when i run the following (see below) on the html var I still receive the captcha problem:

def captcha_test(html):
    try:
        text=html.xpath("//form")[0].text
        standart_text = '\nTo continue please enter the text from the image below\n'
        if text == standart_text:
            print('captcha detected')
            cupcha = html.xpath('//img')[0].attrib['src']
            cupcha_text = str(cupcha[14:])
            s.post(fut_links.loc[0,"url"],data={'filter':cupcha_text})
            page=s.get(fut_links.loc[0,"url"])
            html = lxml.html.fromstring(page.text)
    except:
        print("no captcha")
    return html

Anyone know how to inspect the form's field properly in the captcha page that i get at first, in order to do the post properly, knowing all the right field names and field to fill?

0 Answers
Related