How to fully scrape the youtube comments i.e to click the "Read more" button under some comments using selenium and python?

Viewed 70

I have been trying to scrape the youtube comments. Although I am successful at scraping one-liner comments but long comments includes a "Read more" button to read it.I am not able to intercept these buttons using selenium and python.here is my code

from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium import webdriver
import time

driver_loc= r"C:\Users\dnsingh\Downloads\Compressed\chromedriver_win32\chromedriver.exe"
options = webdriver.ChromeOptions()
options.add_argument("start-maximized")
#options.add_argument("--headless")
service = Service(executable_path=driver_loc)
driver = webdriver.Chrome(service=service,options=options)

driver.get('https://www.youtube.com/watch?v=VC2BL_ChGeg')
driver.execute_script("document.body.style.zoom = '0.25'")
time.sleep(2)
last_height = driver.execute_script("return document.documentElement.scrollHeight")
while True:
        # Scroll down 'til "next load".
        driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
        # Wait to load everything thus far.
        time.sleep(2)
        # Calculate new scroll height and compare with last scroll height.
        new_height = driver.execute_script("return document.documentElement.scrollHeight")
        if new_height == last_height:
            break
        last_height = new_height
k=driver.find_elements(By.XPATH,"//*[@id='more']/span")#finding "Read more" buttons
for i in k:
    if i.text=='Read more':#clicking the read more  buttons only
        print(i.text)
        time.sleep(2)
        i.click()
comment_elems = driver.find_elements(By.XPATH,'//*[@id="content-text"]')
print([i.text for i in comment_elems])#printing comments but not the full ones

driver.quit()

error that i am getting:

ElementClickInterceptedException: Message: element click intercepted: Element ... is not clickable at point (770, 152). Other element would receive the click: ...

1 Answers

I just checked and for me each comment is fully loaded in the HTML even if there is a read more button. Each line of a multi-line comment is stored within the comment, for example like this:

<yt-formatted-string id="content-text" slot="content" split-lines="" class="style-scope ytd-comment-renderer">
    <span dir="auto" class="style-scope yt-formatted-string">START OF COMMENT</span>
    <span dir="auto" class="style-scope yt-formatted-string"> MIDDLE OF COMMENT</span>
    <span dir="auto" class="style-scope yt-formatted-string"></span>
    <span dir="auto" class="style-scope yt-formatted-string">COMMENT THAT COMES AFTER READ MORE</span>
</yt-formatted-string>

Therefore you could get each of the lines and concatenate them to get the entire comment, without having to interact with the read more button.

I think something like this might work.

comment_elems = driver.find_elements(By.XPATH,'//*[@id="content-text"]')


for comment in comment_elems:
    comment_lines = comment.find_elements(By.CLASS_NAME, "yt-formatted-string")

    full_comment = " ".join([line.text for line in comment_lines])
    print(full_comment)

I haven't had time to check that works exactly, but I hope the idea helps you.

Let me know if it works, so that I can update my answer accordingly.

Related