In order to access data such as title and others, you first need to collect all the news in a list. Each news item is located in the iter tag, and they are in the channel tag. So let's use this sample:
soup.channel.find_all('item')
After that, you can extract the necessary data for each news.
for result in soup.channel.find_all('item'):
title = result.title.text
link = result.link.text
date = result.pubDate.text
source = result.source.get("url")
print(title, link, date, source, sep='\n', end='\n\n')
Also, make sure you're using request headers user-agent to act as a "real" user visit. Because default requests user-agent is python-requests and websites understand that it's most likely a script that sends a request. Check what's your user-agent.
Code and full example in online IDE:
from bs4 import BeautifulSoup
import requests
# https://docs.python-requests.org/en/master/user/quickstart/#passing-parameters-in-urls
params = {
"hl": "en-US", # language
"gl": "US", # country of the search, US -> USA
"ceid": "US:en",
}
# https://docs.python-requests.org/en/master/user/quickstart/#custom-headers
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/100.0.4896.88 Safari/537.36",
}
html = requests.get("https://news.google.com/rss", params=params, headers=headers, timeout=30)
soup = BeautifulSoup(html.text, "xml")
for result in soup.channel.find_all('item'):
title = result.title.text
link = result.link.text
date = result.pubDate.text
source = result.source.get("url")
print(title, link, date, source, sep='\n', end='\n\n')
Output:
UK and Europe Heat Wave News: Live Updates - The New York Times
https://news.google.com/__i/rss/rd/articles/CBMiRGh0dHBzOi8vd3d3Lm55dGltZXMuY29tL2xpdmUvMjAyMi8wNy8xOS93b3JsZC91ay1ldXJvcGUtaGVhdC13ZWF0aGVy0gEA?oc=5
Tue, 19 Jul 2022 11:56:58 GMT
https://www.nytimes.com
... other results
Another way to achieve the same thing is to scrape Google News from the HTML instead.
I want to demonstrate how to scrape Google News using pagination. Оne of the ways is to use the start URL parameter which is equal to 0 by default. 0 means the first page, 10 is for the second, and so on.
Also, default search results return about ~10-15 pages. To increase the number of returned pages, you need to set the filter parameter to 0 and pass it to the URL which will return 10+ pages. Basically, this parameter defines the filters for Similar Results and Omitted Results.
While the next button exists, you need to increment the ["start"] parameter value by 10 to access the next page if it's present, otherwise we need to break out of the while loop.
And here is the code:
from bs4 import BeautifulSoup
import requests, lxml
# https://docs.python-requests.org/en/master/user/quickstart/#passing-parameters-in-urls
params = {
"q": "Elon Musk",
"hl": "en-US", # language
"gl": "US", # country of the search, US -> USA
"tbm": "nws", # google news
"start": 0, # number page by default up to 0
# "filter": 0 # shows more than 10 pages. By default up to ~10-15 if filter = 1.
}
# https://docs.python-requests.org/en/master/user/quickstart/#custom-headers
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/100.0.4896.88 Safari/537.36",
}
page_num = 0
while True:
page_num += 1
print(f"{page_num} page:")
html = requests.get("https://www.google.com/search", params=params, headers=headers, timeout=30)
soup = BeautifulSoup(html.text, "lxml")
for result in soup.select(".WlydOe"):
source = result.select_one(".NUnG9d").text
title = result.select_one(".mCBkyc").text
link = result.get("href")
try:
snippet = result.select_one(".GI74Re").text
except AttributeError:
snippet = None
date = result.select_one(".ZE0LJd").text
print(source, title, link, snippet, date, sep='\n', end='\n\n')
if soup.select_one('.d6cvqb a[id=pnnext]'):
params["start"] += 10
else:
break
Output:
1 page:
BuzzFeed News
Elon Musk’s Viral Shirtless Photos Have Sparked A Conversation Around
Body-Shaming After Some People Argued That He “Deserves” To See The Memes
Mocking His Physique
https://www.buzzfeednews.com/article/leylamohammed/elon-musk-shirtless-yacht-photos-memes-body-shaming
None
18 hours ago
People
Elon Musk Soaks Up Sun While Spending Time with Pals Aboard Luxury Yacht in
Greece
https://people.com/human-interest/elon-musk-spends-time-with-friends-aboard-luxury-yacht-in-greece/
None
2 days ago
New York Post
Elon Musk jokes shirtless pictures in Mykonos are 'good motivation' to hit
gym
https://nypost.com/2022/07/21/elon-musk-jokes-shirtless-pics-in-mykonos-are-good-motivation/
None
14 hours ago
... other results from the 1st and subsequent pages.
10 page:
Vanity Fair
A Reminder of Just Some of the Terrible Things Elon Musk Has Said and Done
https://www.vanityfair.com/news/2022/04/elon-musk-twitter-terrible-things-hes-said-and-done
... yesterday's news with “shock and dismay,” a lot of people are not
enthused about the idea of Elon Musk buying the social media network.
Apr 26, 2022
CNBC
Elon Musk is buying Twitter. Now what?
https://www.cnbc.com/2022/04/27/elon-musk-just-bought-twitter-now-what.html
Elon Musk has finally acquired Twitter after a weekslong saga during which
he first became the company's largest shareholder, then offered...
Apr 27, 2022
New York Magazine
11 Weird and Upsetting Facts About Elon Musk
https://nymag.com/intelligencer/2022/04/11-weird-and-upsetting-facts-about-elon-musk.html
3. Elon allegedly said some pretty awful things to his first wife · While
dancing at their wedding reception, Musk told Justine, “I am the alpha...
Apr 30, 2022
... other results from 10th page.
If you need more information about Google News, have a look at Web Scraping Google News with Python blog post.