We are trying to scrape general info on college basketball coaches. Here are two example pages that I am trying to scrape:
- https://gobonnies.sbu.edu/sports/m-baskbl/coaches/index
- https://uofoathletics.com/sports/wbkb/coaches/index
Our ideal output is:
data.frame(
name = c('Mark Schmidt', 'Sean Neal', 'Matt Pappano', 'Steve Curran', 'Tray Woodall', NA, 'Dominique Broadus'),
title = c("Head Men's Basketball Coach", "Assistant Men's Basketball Coach", "Director Of Basketball Operations", "Associate Head Coach, Men's Basketball", "Assistant Men's Basketball Coach", "Head Women's Basketball Coach", "Assistant Women's Basketball Coach"),
email = c(NA, 'sneal@sbu.edu', 'mpappano@sbu.edu', 'scurran@sbu.edu', 'twoodall@sbu.edu', NA, 'dbroadus@ozarks.edu'),
phone = c('716-375-2207', '716-375-2257', '716-375-2218', '716-375-2258', '716-375-2259', '479-979-1325', '479-979-1325'),
stringsAsFactors = FALSE
)
name title email phone
1 Mark Schmidt Head Men's Basketball Coach <NA> 716-375-2207
2 Sean Neal Assistant Men's Basketball Coach sneal@sbu.edu 716-375-2257
3 Matt Pappano Director Of Basketball Operations mpappano@sbu.edu 716-375-2218
4 Steve Curran Associate Head Coach, Men's Basketball scurran@sbu.edu 716-375-2258
5 Tray Woodall Assistant Men's Basketball Coach twoodall@sbu.edu 716-375-2259
6 <NA> Head Women's Basketball Coach <NA> 479-979-1325
7 Dominique Broadus Assistant Women's Basketball Coach dbroadus@ozarks.edu 479-979-1325
This is causing us issues for a few reasons:
- on both pages, the data isn't saved in a table, but rather is saved in individual
divsfor each person. - there is some missing data. there are 2 missing emails and a missing name as well.
Here's what we got so far:
# go to pages, grab person bios
page1 <- 'https://gobonnies.sbu.edu/sports/m-baskbl/coaches/index' %>% read_html()
page1_bios <- page1 %>% html_nodes('div.coach-bios .coach-bio .info')
page2 <- 'https://uofoathletics.com/sports/wbkb/coaches/index' %>% read_html()
page2_bios <- page2 %>% html_nodes('div.coach-bios .coach-bio .info')
# turn bios into 1-column dataframes (not really what we need)
page1_list <- lapply(page1_bios, function(x) paste(x %>% html_children() %>% html_text(), collapse = " "))
page1_bios_df <- unlist(page1_list) %>% as.data.frame()
page2_list <- lapply(page2_bios, function(x) paste(x %>% html_children() %>% html_text(), collapse = " "))
page2_bios_df <- unlist(page2_list) %>% as.data.frame()
We are not all that close, and in fact we're not quite certain if this is even possible to do. I think we need to first get the data into a dataframe even if the columns names are wrong, and then examine the contents of the columns (e.g. look for @ symbols for emails, for #s for phone numbers, for the word "coach" for titles, etc.) to try to name them correctly.