scrape ASIN from amazon URL using javascript

Viewed 20012

Assuming I have an Amazon product URL like so

http://www.amazon.com/Kindle-Wireless-Reading-Display-Generation/dp/B0015T963C/ref=amb_link_86123711_2?pf_rd_m=ATVPDKIKX0DER&pf_rd_s=center-1&pf_rd_r=0AY9N5GXRYHCADJP5P0V&pf_rd_t=101&pf_rd_p=500528151&pf_rd_i=507846

How could I scrape just the ASIN using javascript? Thanks!

17 Answers

This worked perfectly for me, I tried all the links on this page and some other links:

function ExtractASIN(url){
    var ASINreg = new RegExp(/(?:\/)([A-Z0-9]{10})(?:$|\/|\?)/);
    var  cMatch = url.match(ASINreg);
    if(cMatch == null){
        return null;
    }
    return cMatch[1];
}
ExtractASIN('http://www.amazon.com/Kindle-Wireless-Reading-Display-Generation/dp/B0015T963C/ref=amb_link_86123711_2?pf_rd_m=ATVPDKIKX0DER&pf_rd_s=center-1&pf_rd_r=0AY9N5GXRYHCADJP5P0V&pf_rd_t=101&pf_rd_p=500528151&pf_rd_i=507846');
  • I assumed that the ASIN is a 10-length with capital letters and numbers
  • I assumed that after the ASIN must be: end of the link, question mark or slash
  • I assumed that before the ASIN must be a slash

Inspired by many of the answers here, I found that

(?:[/])([A-Z0-9]{10})(?:[\/|\?|\&|\s|$])

let url="https://www.amazon.com/Why-We-Sleep-Science-Dreams-ebook/dp/B06Y649387/ref=pd_sim_351_4/131-0417603-5732106?_encoding=UTF8&pd_rd_i=B06Y649387&pd_rd_r=5ebbfdd5-a2f6-4ee3-ad13-5036b5e20827&pd_rd_w=LBo2H&pd_rd_wg=OBomS&pf_rd_p=3c412f72-0ba4-4e48-ac1a-8867997981bd&pf_rd_r=TN0WDV3AC7ED4Y7EKNVP&psc=1&refRID=TN0WDV3AC7ED4Y7EKNVP"
url.match("(?:[/])([A-Z0-9]{10})(?:[\/|\?|\&|\s])")

>> Array [ "/B06Y649387/", "B06Y649387" ]

works really well for extracting asin from anywhere in the url. You can try it out here. https://regexr.com/56jm7

edit: Added end-of-string as one of the stopping checks. This is needed when the regex is used in python

You can scrape ASIN codes from the data-asin attribute in the search results using XPath.

For example $x('//@data-asin').map(function(v,i){return v.nodeValue}) can be ran in Chrome's console.

Used both methods in a single function:

const extractASIN = (url: string) => {
  var regex = RegExp('(?:[/])([A-Z0-9]{10})(?:[/|?|&|s])');
  const m = url.match(regex);
  if (m) {
    return m[1];
  }
  return url.split('/ref')[0].split('/dp/')[1];
};
// function to find the nth instance of character in string
function nthIndex(str, pat, n) {
  var L = str.length,
    i = -1;
  while (n-- && i++ < L) {
    i = str.indexOf(pat, i);
    if (i < 0) break;
  }
  return i;
}
// this function takes a string and split string list as parameters and slices off entirely after that character is found
function splitSliceFunc(splitStr, splitStrList) {
  for (i = 0; i < splitStrList.length; i++) {
    splitStr = splitStr.split(splitStrList[i])[0];
  }
  return splitStr;
}
try {
  const amzUrl = 'https://www.amazon.com/Encyclopedia-Country-Living-50th-Anniversary/dp/1632172895/ref=sr_1_1?keywords=survival+encyclopedia&pd_rd_r=8e62738c-ae2b-46c0-b477-db5cf23a6b0a&pd_rd_w=0Eazc&pd_rd_wg=E51TF&pf_rd_p=54cea6b7-0efb-45a3-b68b-8c1ccfbfa553&pf_rd_r=EE9X3J3QBPCDQAVQJ9FQ&qid=1651929404&sr=8-1';
  const sliceUptoAsinList = ["/dp/", "/gp/product/"]; // list for slice occurrences before asin
  const sliceAfterAsinList = ["/", "?"]; // list for slice occurrences after the asin
  let sliceUptoAsin; // variable to store index of slice occurrence before asin
  let shortenedUrl;
  // if else statements for all the possible slice occurrences before asin
  if (amzUrl.includes(sliceUptoAsinList[0])) {
    sliceUptoAsin = nthIndex(amzUrl, "/dp/", 1);
    shortenedUrl = amzUrl.slice(sliceUptoAsin + 4); // + 4 to remove /dp/ also
    console.log(sliceUptoAsin, shortenedUrl);
  } else if (amzUrl.includes(sliceUptoAsinList[1])) {
    sliceUptoAsin = nthIndex(amzUrl, "/gp/product/", 1);
    shortenedUrl = amzUrl.slice(sliceUptoAsin + 12); // + 12 to remove /gp/product/ also
    console.log(sliceUptoAsin, shortenedUrl);
  } else {
    throw "url format not supported";
  }
  // removes everything after the asin following 'sliceAfterAsinList'
  shortenedUrl = splitSliceFunc(shortenedUrl, sliceAfterAsinList);
  console.log(shortenedUrl);
} catch (error) {
  console.log(error)
}

I opted for a non regex approach because they become harder to maintain. IF you are treating the urls as simple strings, split and slice can also do the job.

The above code assumes that the ASIN is followed by "/dp/" or "/gp/product/" (*but is not limited to these occurrences only because 'sliceUptoAsinList' array can have as many as slice occurrences before ASIN as many you want, followed by an added else-if condition).

The code will work irrespective of whether there are 10 or more characters in ASIN because it will only look for the first occurrence of any character found in the 'sliceAfterAsinList' array in the url and will remove everything next to that character (including the character also).

Related