Google Scholar blocked me from using search_pubs

Viewed 808

I am using Pycharm Community Edition 2020.3.2, Scholarly version 1.0.2, Tor version 1.0.0. I tried to scrape 700 articles to find their numbers of citations. Google Scholar blocked me from using search_pubs (a function of Scholarly). However, another function of Scholarly, which is search_author, is still working well. In the beginning, search_pubs function worked properly. I tried these codes.

from scholarly import scholarly
scholarly.search_pubs('Large Batch Optimization for Deep Learning: Training BERT in 76 minutes')

After a few trials, it shows the below error.

Traceback (most recent call last):
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\IPython\core\interactiveshell.py", line 3343, in run_code
    exec(code_obj, self.user_global_ns, self.user_ns)
  File "<ipython-input-9-3bbcfb742cb5>", line 1, in <module>
    scholarly.search_pubs('Large Batch Optimization for Deep Learning: Training BERT in 76 minutes')
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\_scholarly.py", line 121, in search_pubs
    return self.__nav.search_publications(url)
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\_navigator.py", line 256, in search_publications
    return _SearchScholarIterator(self, url)
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\publication_parser.py", line 53, in __init__
    self._load_url(url)
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\publication_parser.py", line 58, in _load_url
    self._soup = self._nav._get_soup(url)
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\_navigator.py", line 200, in _get_soup
    html = self._get_page('https://scholar.google.com{0}'.format(url))
  File "C:\Users\binhd\anaconda3\envs\t2\lib\site-packages\scholarly\_navigator.py", line 152, in _get_page
    raise Exception("Cannot fetch the page from Google Scholar.")
Exception: Cannot fetch the page from Google Scholar.

Then, I figured out that the reason is I need to pass the CAPTCHA from Google in order to continue to fetch the info from Google Scholar. Many people suggest that I need to use Proxy since my IP was blocked by Google. I tried to change Proxy using FreeProxies()

from scholarly import scholarly, ProxyGenerator

pg = ProxyGenerator()
pg.FreeProxies()
scholarly.use_proxy(pg)
scholarly.search_pubs('Large Batch Optimization for Deep Learning: Training BERT in 76 minutes')

It does not work and Pycharm is frozen for a long time. Then, I installed Tor (pip install Tor) and tried again:

from scholarly import scholarly, ProxyGenerator
pg = ProxyGenerator()
pg.Tor_External(tor_sock_port=9050, tor_control_port=9051, tor_password="scholarly_password")
scholarly.use_proxy(pg)
scholarly.search_pubs('Large Batch Optimization for Deep Learning: Training BERT in 76 minutes')

It does not work. Then, I tried with SingleProxy()

from scholarly import scholarly, ProxyGenerator
pg = ProxyGenerator()
pg.SingleProxy(https='socks5://127.0.0.1:9050',http='socks5://127.0.0.1:9050')
scholarly.use_proxy(pg)
scholarly.search_pubs('Large Batch Optimization for Deep Learning: Training BERT in 76 minutes')

It also does not work. I have never tried Luminati since I am not familiar with it. If anyone knows the solution, please help!

1 Answers

As an alternative to a scholarly solution, you can try to use Google Scholar Organic Results API from SerpApi.

It's a paid API with a free plan which handles bypassing blocks from Google or other search engines on their back end by solving CAPTCHA and rotating proxies so you don't have to.

Code and example in the online IDE:

import os, json
from serpapi import GoogleSearch
from urllib.parse import urlsplit, parse_qsl

params = {
    # os.getenv(): https://docs.python.org/3/library/os.html#os.getenv
    "api_key": os.getenv("API_KEY"),                 # your Serpapi API key
    "engine": "google_scholar",                      # search engine
    "q": "blizzard",                                 # search query
    "hl": "en",                                      # language
    # "as_ylo": "2017",                              # from 2017
    # "as_yhi": "2021",                              # to 2021
    "start": "0"                                     # first page
}

search = GoogleSearch(params)         # where data extraction happens

organic_results_data = []

papers_is_present = True
while papers_is_present:
    results = search.get_dict()      # JSON -> Python dictionary

    for publication in results["organic_results"]:
        organic_results_data.append({
            "page_number": results.get("serpapi_pagination", {}).get("current"),
            "result_type": publication.get("type"),
            "title": publication.get("title"),
            "link": publication.get("link"),
            "result_id": publication.get("result_id"),
            "summary": publication.get("publication_info").get("summary"),
            "snippet": publication.get("snippet"),
            })
             
        # paginates to the next page if the next page is present
        if "next" in results.get("serpapi_pagination", {}):
            search.params_dict.update(dict(parse_qsl(urlsplit(results["serpapi_pagination"]["next"]).query)))
        else:
            papers_is_present = False

print(json.dumps(organic_results_data, indent=2, ensure_ascii=False))

Outputs in this case up to the 100th page:

]
 {
    "page_number": 1,
    "result_type": null,
    "title": "Base catalyzed ring opening reactions of erythromycin A",
    "link": "https://www.sciencedirect.com/science/article/pii/S004040390074754X",
    "result_id": "RRGl5nJoIi4J",
    "summary": "ST Waddell, TA Blizzard - Tetrahedron letters, 1992 - Elsevier",
    "snippet": "While the direct opening of the lactone of erythromycin A by hydroxide to give the seco acid has so far proved elusive, two types of base catalyzed reactions which lead to rupture of the …"
  }, ... other results
  {
    "page_number": 100,
    "result_type": null,
    "title": "Síndrome de Johanson-Blizzard: importância do diagnóstico diferencial em pediatria",
    "link": "https://www.scielo.br/j/jped/a/Mc3X8DGcZSYQVqnL99kTtBH/abstract/?lang=pt",
    "result_id": "z_CmhVgEW2oJ",
    "summary": "MW Vieira, VLGS Lopes, H Teruya… - Jornal de …, 2002 - SciELO Brasil",
    "snippet": "… Description: we describe a Brazilian girl affected by Johanson-blizzard syndrome and review the literature. Comments: Johanson-Blizzard syndrome is an autosomal recessive …"
  }
]

Disclaimer, I work for SerpApi.

Related