I have the following list of urls that i am tring to scrape:
base_url='https://misc.interactivebrokers.com/cstools/contract_info/index2.php?action=Details&site=GEN&conid='
conId=[280313912,289230167,285817885,289229956,256019341,289230102,289128542,289128563,289230371,289230337,287565355,285578089,287695374,285714991,287565358]
url_list=[base_url+str(x) for x in conId]
If I try to get the page for only one of the urls, it does work fine:
import requests
import lxml.html
s = requests.session()
page= s.get(url_list[0])
html = lxml.html.fromstring(page.text)
if html.xpath("//*[@id='contractSpecs']/table[5]/tr[2]/td[2]")!=[]:
print("get_result")
else:
print("missing")
However if I try to get multiple pages the site throws in a catcha:
for url in url_list:
page= s.get(url)
html = lxml.html.fromstring(page.text)
if html.xpath("//*[@id='contractSpecs']/table[5]/tr[2]/td[2]")!=[]:
print("get_result")
else:
print("missing")
Upon inspecting the captcha result i can see that the source page is as follow:
<br>
<form type="post">
To continue please enter the text from the image below
<br>
<img src="image.php?str=1wVfP3">
<!--img src="https://chatsrv1.interactivebrokers.com/cstools/contract_info/v3.8/image.php?str=1wVfP3"-->
<br>
Text:
<input name="filter" type="text">
<input name="action" type="hidden" value="Details"/>
<input name="conid" type="hidden" value="280313912"/>
<input name="contract_id" type="hidden" value=""/>
<input name="noBanner" type="hidden" value=""/>
<input name="rescnt" type="hidden" value="100"/>
..............
</input>
</br>
</img>
</br>
</form>
</br>
The captcha answer is in the tag <img src="image.php?str=1wVfP3"> ! (after str=)
I know this because i had a third person do that code for me in the past, but he did it with a module called "Grab". I am now trying to refactor that code using requests.
In order to pas that captcha, that other coder used in Grab:
try:
text = r.xpath_text('//form')
standart_text = 'To continue please enter the text from the image below Text:'
if text == standart_text:
print('captcha detected')
cupcha = r.xpath('//img').get('src')
cupcha_text = str(cupcha[14:])
r.doc.set_input('filter', cupcha_text)
r.doc.submit()
The set_input method correspond to a post request.. It seems that grab can understand what are the form fields and fill them up automatically.
My problem is that the captcha appears now only when scraping - which means i cannot use google chrome's inspect>network to check the filed names for the form.
So when i run the following (see below) on the html var I still receive the captcha problem:
def captcha_test(html):
try:
text=html.xpath("//form")[0].text
standart_text = '\nTo continue please enter the text from the image below\n'
if text == standart_text:
print('captcha detected')
cupcha = html.xpath('//img')[0].attrib['src']
cupcha_text = str(cupcha[14:])
s.post(fut_links.loc[0,"url"],data={'filter':cupcha_text})
page=s.get(fut_links.loc[0,"url"])
html = lxml.html.fromstring(page.text)
except:
print("no captcha")
return html
Anyone know how to inspect the form's field properly in the captcha page that i get at first, in order to do the post properly, knowing all the right field names and field to fill?