I am using an R package, edgarWebR, to parse SEC filings, such as https://www.sec.gov/Archives/edgar/data/1060224/000090480206000008/sa10k306.htm. It returns a dataframe, of which one column - called "raw" - is HTML. It breaks up the HTML page into paragraphs, one row per paragraph:
| other columns | raw | text |
|---|---|---|
| First row | <p id="PARA339" style="TEXT-ALIGN: left; MARGIN: 0pt; LINE-HEIGHT: 1.25"><font style="FONT-SIZE: 10pt; FONT-FAMILY: Times New Roman, Times, serif"><i>We had a net loss of $1.</i><i><b>55</b></i><i> million for the year ended December 31, 201</i><i>6</i><i> and have an accumulated deficit of $</i><i>61.5</i><i> million as of December 31, 201</i><i>6</i><i>. To achieve sustainable profitability, we must generate increased revenue.</i></font></p> |
We had a net loss of $1.55 million for the year ended December 31, 2016 and have an accumulated deficit of $61.5 million as of December 31, 2016. To achieve sustainable profitability, we must generate increased revenue. |
| Second row | <div style="line-height:174%;text-align:left;font-size:9pt;"><font style="font-family:inherit;font-size:9pt;font-style:italic;font-weight:bold;">We have a history of losses, and we cannot assure you that we will achieve profitability.</font></div> |
We have a history of losses, and we cannot assure you that we will achieve profitability. |
You can easily replicate an example dataframe by running
library(edgarWebR)
df <- parse_filing("https://www.sec.gov/Archives/edgar/data/1060224/000090480206000008/sa10k306.htm",include.raw=TRUE)
My goal is to parse the HTML to determine which of the paragraphs represent headings, by calculating which formats (e.g. bold italics) are appearing less frequently throughout the document.
One problem is that some paragraphs will be descriptive (i.e. not a heading level), but contain one or a few words emboldened in the middle of the paragraph, for emphasis. To this end, I have a function which, for each paragraph, will take a list of css selectors (e.g. those for bold or italic) and tell me what proportion of the characters are of that style, thanks to @QHarr:
## the function
paragraph_style_proportion <- Vectorize(function(html, css_selector) {
html_content <- tryCatch(read_html(html), error=function(err) "NOT HTML")
if (html_content == "NOT HTML") {
style_proportion <- -1
}
else {
whole_paragraph_length <- nchar(str_squish(html_content %>% html_text()))
style_text_length <- sum(nchar(str_squish(html_content %>% html_nodes(css_selector) %>% html_text())))
style_proportion <- round(style_text_length/whole_paragraph_length, 2)
}
return(style_proportion)
})
## to apply the function to the html-containing column "raw" of a dataframe "df"
df <- df %>%
mutate(bold_proportion = paragraph_style_proportion(raw, 'b, strong, [style*="font-weight:bold"]'))
From there, I can easily apply a rule such as to record a paragraph as, say, bold, if at least 40% of the characters are bold. From there, I proceed to calculate frequencies of each combination of styles (e.g. italic capitalised text) and assign numerical heading levels.
However, the problem I now face is cases such as the below - in which the first few words of a paragraph are of a different style:
In the example above, my algorithm would class these as "normal" paragraphs - i.e. not bold or italic or anything - as it is well under 40% italic. But clearly it represents a heading of some kind - because the different style is at the beginning of the paragraph.
Some examples with this problem can be found at https://www.sec.gov/Archives/edgar/data/1376067/000137606711000002/vll51231201010k.htm and https://www.sec.gov/Archives/edgar/data/1060224/000090480206000008/sa10k306.htm
How would I start to solve this? The previous function cannot tackle this; I would need some way of parsing each text-containing HTML part of the paragraph one-by-one, rather than just pulling out all bold text and dividing by the total length of the text of the paragraph. Especially difficult is that the length of the 'first part of the paragraph' will vary (and usually won't even exist), as will the styles of this first part, and the style of the rest of the paragraph.
Ideally, in cases where the first part of the paragraph is a different format, I'd want to split the paragraph into two rows - making the different-format part its own row, which can then have its heading level classified in the way I've described. If not, I'd either want to flag the whole paragraph as the style of the first few words, or maybe create a column that flags paragraphs where the first few words are of a different style.
Thank you
