Bypass Cloudflare with puppeteer

Viewed 14002

I am trying to scrape some startups data of a site with puppeteer and when I try to navigate to the next page the cloudflare waiting screen comes in and disrupts the scraper. I tried changing the IP but its still the same. Is there a way to bypass it with puppeteer.

(async () => {

  const browser = await puppeteer.launch({
    headless: false,
    defaultViewport: null,
  });

  const page = await browser.newPage();

  page.setDefaultNavigationTimeout(0);

  let links = [];

  // initial page

  await page.goto(`https://www.startupranking.com/top/india`, {
    waitUntil: "networkidle0",
  });

  // looping through the url to different pages

  for (let i = 2; i <= 7; i++) {
    if (i === 3) {
      console.log("waiting");

      await page.waitFor(20000);

      console.log("waited");
    }

    const onPageLinks = await page.$$eval("tr .name a", (arr) =>
      arr.map((cur) => cur.href)
    );

    links = links.concat(onPageLinks);

    console.log(onPageLinks, "inside loop");

    await page.goto(`https://www.startupranking.com/top/india/${i}`, {
      waitUntil: "networkidle0",
    });
  }

  console.log(links, links.length, "outside loop");
})();

As it is only checking for the first loop i put in a waitFor to bypass the time it takes to check, it works fine on some IP's but on others it gives challenges to solve, I have to run this on a server so I am thinking of bypassing it completely.

1 Answers

Using browser-based automation tools like Puppeteer or Playwright alone is usually insufficient to get around the WAF. To avoid being ignored by the WAF, you should try to imitate a real user as much as possible. It's critical to collect valid and proper cookies, and to replicate the original browser properties as much as possible.

In your case, puppeteer-extra-plugin-stealth without headless mode is sufficient to get around the WAF**. If you were to scale this and run it in headless mode, sending correct headers would be crucial.

I copied your code, added stealth and a few lines in the body (with a comment for you to locate them).

I believe the main problem is that the browser wasn't waiting for the right content to load. Once puppeteer hits the page, you might want to wait for some time till you bypass the waiting room. In this case, I run into some errors when the code tried to execute on elements that didn't exist yet.

You can simply wait for specific CSS selectors to render and then continue with the rest of the code. I added await page.waitForSelector(".top-100 .table-striped") to avoid those errors.

const puppeteer = require("puppeteer-extra");

const StealthPlugin = require("puppeteer-extra-plugin-stealth")();
puppeteer.use(StealthPlugin);

(async () => {
  const browser = await puppeteer.launch({
    headless: false,
    defaultViewport: null,
  });

  const page = await browser.newPage();

  page.setDefaultNavigationTimeout(0);

  let links = [];

  // initial page
  await page.goto(`https://www.startupranking.com/top/india`, {
    waitUntil: "networkidle0",
  });
  await page.waitForSelector(".top-100 .table-striped"); // NEW! wait for data

  // looping through the url to different pages

  for (let i = 2; i <= 7; i++) {
    if (i === 3) {
      console.log("waiting");

      await page.waitFor(2_000);

      console.log("waited");
    }

    const onPageLinks = await page.$$eval("tr .name a", (arr) =>
      arr.map((cur) => cur.href)
    );

    links = links.concat(onPageLinks);

    console.log(onPageLinks, "inside loop");

    await page.goto(`https://www.startupranking.com/top/india/${i}`, {
      waitUntil: "networkidle0",
    });
    await page.waitForSelector('.top-100 .table-striped'); // NEW! wait for data
  }

  console.log(links, links.length, "outside loop");

  await browser.close(); // NEW! close the browser when done
})();

You can learn more about Cloudflare from here: bypassing Cloudflare.

Related