Extracting web address and using for loop

Viewed 18

I am trying to extract websites of the members from https://www.mhi.org/members. So, I wrote a code to visit the member page one by one and extract the web addresses. I am using BeautifulSoup Library to extract. However my problem is not implementation of BeautifulSoup but in the for loop. The above link has 15 members per page. When I try to run the code for one page it just returns the web address of only one member. I have imported all the desired libraries.

url = input('Enter URL:')

#position = int(input('Enter position:'))-1
html = urlopen(url).read()
lst1 = list()
lst = list()
lst2 = list()
lst3 = list()
lst4 = list()
soup = BeautifulSoup(html,"html.parser")
url1 = "https://www.mhi.org"

conn = sqlite3.connect('list.sqlite')
cur = conn.cursor()

cur.execute('''CREATE TABLE IF NOT EXISTS weblinks (URL TEXT UNIQUE)''')
cur.execute('''CREATE TABLE IF NOT EXISTS weblinks1 (website TEXT UNIQUE)''')
cur.execute('''CREATE TABLE IF NOT EXISTS weblinks2 (website1 TEXT UNIQUE)''')

#print(soup)
for link in soup.find_all('a'):
    lst.append(link.get('href'))

for links in lst:
    if 'members' in links:
        url2 = urllib.parse.urljoin(url1, links)
        lst2.append(url2)
#print(lst2)

url3 = lst2[4:18]
print(url3)
for x in url3:
    html1 = urlopen(x).read()
    soup1 = BeautifulSoup(html1,"html.parser")
    for url4 in soup1.find_all('a'):
        url4 = url4.get('href')
        if url4 not in lst3:
            lst3.append(url4)
    url6 = lst3.pop(108)
    print(url6)

So, the last "for loop" where url6 is the desired output. It just prints output one web address. Please advise what am I missing here.

0 Answers
Related