Why do I get bad output when using Nokogiri "search"?

Viewed 152

I want to scrape data from specific divs on a CarFax report. However, when I search for divs, I always get this weird garbage output.

I tried search(#divId) , search(.divClass), and even tried to grab all divs with search('div'). Each time I get similar results: the div's content is partially truncated and the tags are all messed up.

This is the URL I am loading into my agent: https://gist.github.com/atkolkma/8024287

This is the code (user and pass ommited):

require "rubygems"
require "mechanize"

scraper = Mechanize.new
scraper.user_agent_alias = 'Mac Safari'
scraper.follow_meta_refresh = true
scraper.redirect_ok = true

scraper.get("http://www.carfaxonline.com")
form = scraper.page.forms.first
form.j_username = "******"
form.j_password = "*****"
scraper.submit(form)

scraper.get("http://www.carfaxonline.com/api/report?vin=1G1AT58H697144202&track=true")

puts scraper.page.search("#headerBodyType")

This is what the file returns when I run it:

</div>4 DRderBodyType">

What I expect is:

<div id="headerBodyType"> SEDAN 4 DR </div>

The strangest thing is, if I copy the HTML source, save it as a new file, upload it and search it, I get the correct output! I've uploaded the copied HTML to my chevy-pics dot com domain and run the following code:

scraper2 = Mechanize.new

scraper2.get("http://www.chevy-pics.com/test.html")

puts scraper2.page.search("#headerBodyType")

I get this as output, as expected:

<div id="headerBodyType"> SEDAN 4 DR </div>
1 Answers
Related