https://www.google.com/search?q=LAPTOP+ACER+I3/4/1TB/8GEN+full+specs
For example: I want to search this product and scrape the specifications of it directly from the featured snippets. How do I get everything inside of that box??
https://www.google.com/search?q=LAPTOP+ACER+I3/4/1TB/8GEN+full+specs
For example: I want to search this product and scrape the specifications of it directly from the featured snippets. How do I get everything inside of that box??
According to Google featured snippets,
Featured snippets come from web search listings. Google's automated systems determine whether a page would make a good featured snippet to highlight for a specific search request.
So this won't be a reliable way if you want to scrape for multiple searches, as they will vary wildly.
However, for this particular search, you can redirect your scraper to follow that link and then you have to write code to scrape info for that link.
How do I get everything inside of that box??
That box only contains the info that you can see, maybe not all the info that you want. And if you want to scrape only that info, that's quite straight-forward.
import requests
from bs4 import BeautifulSoup
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
div = soup.find("div", {"class": "webanswers-webanswers_table__webanswers-
table"})
tr = div.findAll("tr")
for row in tr:
td = row.findAll("td")
print(td[0].text.strip(), ": " ,td[1].text.strip())
If above code doesn't work or returns 429 or other status code, Google might be blocking scraping scripts/spiders. Try adding user agents like:
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.129 Safari/537.36'}
response = requests.get(url, headers=headers)
If that also fails, try using selenium
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait as wait
url = "https://www.google.com/search?q=acer+i3+8th+gen+1tb+laptop+full+specs"
driver = webdriver.Firefox("path/to/geckodriver")
driver.get(url)
snippet = wait(driver, 60).until(lambda driver:
driver.find_element_by_css_selector("div.webanswers-webanswers_table__webanswers-table"))
print(snippet.text)
You can achieve this via:
requests-htmlbeautifulsoupUsing Requests-html and example in online IDE:
from requests_html import HTMLSession
session = HTMLSession()
response = session.get('https://www.google.com/search?q=Acer+Aspire+3+A315-53+Laptop++specs')
specs_no_split = response.html.find('.Crs1tb', first=True).text
# splitting by a new line and grabbing every second value
specs_split = response.html.find('.Crs1tb', first=True).text.split('\n')[0::2]
# converting from list to a string
specs_filtered = ''.join(specs_split)
print(specs_no_split)
print(specs_split)
print(specs_filtered)
# output:
'''
Acer Aspire 3 A315-53 (NX.H38SI.002) Laptop (Core i3 8th Gen/4 GB/1 TB/Windows 10) Specifications
display type
LED
display size
15.6 Inches (39.62 cm)
display resolution
1920 x 1080 Pixels
display touchscreen
No
display features
Full HD LED Backlit IPS Display
ще 1 рядок
['Acer Aspire 3 A315-53 (NX.H38SI.002) Laptop (Core i3 8th Gen/4 GB/1 TB/Windows 10) Specifications', 'LED', '15.6 Inches (39.62 cm)', '1920 x 1080 Pixels', 'No', 'Full HD LED Backlit IPS Display']
Acer Aspire 3 A315-53 (NX.H38SI.002) Laptop (Core i3 8th Gen/4 GB/1 TB/Windows 10) SpecificationsLED15.6 Inches (39.62 cm)1920 x 1080 PixelsNoFull HD LED Backlit IPS Display
'''
Using BeautifulSoup and example in online IDE:
from bs4 import BeautifulSoup
import requests, lxml
headers = {
'User-agent':
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.102 Safari/537.36 Edge/18.19582"
}
html = requests.get('https://www.google.com/search?q=https://www.google.com/search?q=Acer+Aspire+3+A315-53+Laptop+specs, headers=headers).text
soup = BeautifulSoup(html, 'lxml')
specifications = soup.find('div', class_='wDYxhc').text
print(specifications)
# Output:
'''
Acer Aspire 3 A315-53-35ZY 15.6" Notebook, Intel i3, 4GB Memory, Windows 10Operating System. Windows 10.Hard Drive. HDD.Memory. 4GB RAM DDR4.Graphics card. Intel UHD Graphics 620.Processor. 2.2 Ghz Intel i3 processor.Display. 15.6-inch 1920 x 1080 resolution.
'''
Alternatively, you can use Google Direct Answer Box API from SerpApi. It's a paid API with a free plan.
The difference is that you only need to think about the data you want to extract, and what query parameters to use, rather than figuring out how to bypass blocks or extract something.
Code to integrate (example in online IDE):
from serpapi import GoogleSearch
import os
params = {
"api_key": os.getenv("API_KEY"),
"engine": "google",
"q": "Acer Aspire 3 A315-53 Laptop specs",
"google_domain": "google.com",
}
search = GoogleSearch(params)
results = search.get_dict()
specs = results['answer_box']['contents']['formatted']
print(specs)
# output:
'''
[{'display_type': 'display size', 'led': '15.6 Inches (39.62 cm)'}, {'display_type': 'display resolution', 'led': '1920 x 1080 Pixels'}, {'display_type': 'display touchscreen', 'led': 'No'}, {'display_type': 'display features', 'led': 'Full HD LED Backlit IPS Display'}]
'''
Disclaimer, I work for SerpApi.