Skipping nested comments recursively in a scanner

Viewed 207

I'm writing a scanner for a compiler and have a function to skip comments whenever it sees them, I wanted to know how one would skip nested comments, i.e something like " hello "w" world", recursively, So far I have something like:

while ( current_char != '#' ){ // comments in this language starts with $# and end with #$)
      next_char(); // gets the next character in our file

}

1 Answers

Nested comments using a counter for the depth

The easiest method is to maintain a counter of the "comment depth":

  • Every time you encounter $#, increment the counter;
  • Every time you encounter #$, decrement the counter.

When the counter is 0, you're reading code; when the counter is 1 or more, you're reading a comment.

While reading a comment, ignore everything except $# and #$.

Example:

Code $# comment depth 1 $# comment depth 2 #$
comment depth 1 #$ code $# comment depth 1 #$ code

Escaping characters in strings

You mentioned the following example:

" hello "w" world"

Let me put some emphasis on the following advice:

Do not allow nested comments if the begin-comment and the end-comment symbols are identical.

Otherwise, there would be no way to distinguish between the two following situations:

Situation 1:  "comment" code "comment"
Situation 2:  "comment "nested comment" comment"

Note that the symbol " is usually used for strings, not for comments. There is no such thing as a "nested string" (what would that mean??). However, there is such a thing as "escaped characters in a string". Indeed, what if you want a string to contain the character "? The usual approach is to reserve an escaping character; characters directly following the escaped character are not interpreted. So you could write the following string:

" hello \"w\" world"

Satisfyingly, you can note that StackOverflow's automatic syntax-colouring correctly coloured all the string in green; whereas the previous string without the \ was not correctly coloured.

Related