HTML tokenizer algorithm

Viewed 1700

I'm trying to write a basic html parser which doesn't tolerate errors and was reading HTML5 parsing algorithm but it's just too much information for a simple parser. I was wondering if someone had an idea on the logic for a basic tokenizer which would simply turn a small html into a list of significant tokens. I'm more of interested in the logic than the code..

std::string html = "<div id='test'> Hello <span>World</span></div>";

Tokenizer t;
t.tokenize(html);

So for the above html, I want to convert it to a list of something like this:

["<","div","id", "=", "test", ">", "Hello", "<", "span", ">", "world", "</", "span", ">", "<", "div", ">"]

I don't have anything for the tokenize method but was wondering if iterating over the html character by character is the best way to build the list..

void Tokenizer::tokenize(std::string html){
    std::list<std::string> tokens;

    for(int i = 0; i < html.length();i++){
        char c = html[i];
        if(...){
            ...
        }
    }
}
2 Answers
Related