How would I make a VSCode syntax highlighter that is not extremely hard to understand?

Viewed 292

I'm trying to make a custom syntax highlighter for my own markup language. All the examples are complicated, missing steps and are very, very hard to understand.

Is there anything that fully documents how to make a syntax highlighter?

(for VSCode, by the way)


For example, this video https://www.youtube.com/watch?v=5msZv-nKebI which has an extremely large skip in the middle and doesn't really explain much.

My current code, made with Yeoman generator is:

{
    "$schema": "https://raw.githubusercontent.com/martinring/tmlanguage/master/tmlanguage.json",
    "name": "BetterMarkupLanguage",
    "patterns": [
        {
            "include": "#keywords"
        },
        {
            "include": "#strings"
        }
    ],
    "repository": {
        "keywords": {
            "patterns": [{
                "name": "entity.other.bml",
                "match": "\\b({|}|\\\\|//)\\b"
            }]
        },
        "strings": {
            "name": "string.quoted.double.bml",
            "begin": "`",
            "end": "`"
        }
    },
    "scopeName": "source.bml"
}
1 Answers

Synopsis

I'm not sure at what level you're approaching the problem, but there are basically two kinds of syntax-highlighting:

  • Just identify little nuggets of canonically identifiable tokens (strings, numbers, maybe operators, reserved words, comments) and highlight those, OR
  • Do the former, and also add in context awareness.

tmLanguage engines basically have two jobs:

  • Assign scopes.
  • Maintain a stack of contexts.

Samples

Lets say you make a definition for integers with this pattern:

"integers": {
    "patterns": [{
        "name": "constant.numeric.integer.bml",
        "match": "[+-]\\d+"
    }]
},

When the engine matches an integer like that, it will match to the end of the regex, assign the scope from "name", and then continue matching things in this same context.

Compare that to your "strings" definition:

"strings": {
    "name": "string.quoted.double.bml", // should be string.quoted.backtick.bml
    "begin": "`",
    "end": "`"
},

Those "begin" and "end" markers denote a change in the tmLanguage stack. You have pushed into a new context inside of a string.

Right now, there are no matches configured in this context, but you could do that by adding a "patterns" key with some "match"es or "include"s. "include"s are other sets of matches like "integers" that you've defined elsewhere. You can add it to the "strings" patterns to match integers inside strings. Matching integers might be silly, but think about escaped backticks: You want to scope those and stay in the same context within "strings". You don't want those popping back out prematurely.

Order of operations

You'll eventually notice that the first pattern encountered is matched. Remember the integers set? What happens when you have 45.125? It will decide to match the 45 and the 125 as integers and ignore the . entirely. If you have a "floats" pattern, you want to include that before your naïve integer pattern. Both these "numbers" definitions below are equivalent, but one lets you re-use floats and integers independently (if that's useful for your language):

  • "numbers": {
        "patterns": [
            {"include": "#floats"},
            {"include": "#integers"}
        ]
    },
    "integers": {
        "patterns": [{
            "name": "constant.numeric.integer.bml",
            "match": "[+-]\\d+"
        }]
    },
    "floats": {
        "patterns": [{
            "name": "constant.numeric.float.bml",
            "match": "[+-]\\d+\\.\\d*"
        }]
    },
    
  • "numbers": {
        "patterns": [{
                "name": "constant.numeric.float.bml",
                "match": "[+-]\\d+\\.\\d*"
            }, {
                "name": "constant.numeric.integer.bml",
                "match": "[+-]\\d+"
        }]
    },
    

Doing it right

The "numbers"/"integers"/"floats" thing was trivial, but well-designed syntax definitions will define utility groups that "include" equivalent things together for re-usability:

  • A normal programming language will have things like

    • A "statements" group of all things that can be directly executed. This then may or may not (language-dependent) include...
    • An "expressions" group of things you can put on the right-hand-side of an assignment, which will definitely include...
    • An "atoms" group of strings, numbers, chars, etc. that might also be valid statements, but that also depends on your language.
    • "function-definitions" probably won't be in "expressions" (unless they are lambdas) but probably would be in "statements." Function definitions might push into a context that lets you return and so on.
  • A markup language like yours might have

    • An "inline" group to keep track of all the markup one can have within a block.
    • A "block" group to hold lists, quotes, paragraphs, headers.
    • ...

Though there is more you could learn (capture groups, injections, scope conventions, etc.), this is hopefully a practical overview for getting started.

Conclusion

When you write your syntax highlighting, think to yourself: Does matching this token put me in a place where things like it can be matched again? Or does it put me in a different place where different things (more or fewer) ought to be matched? If the latter, what returns me to the original set of matches?

Related