In Compiler Construction by Aho Ullman and Sethi, it is given that the input string of characters of the source are read by scanner(lexical analysis) and groups characters into meaningful sequences called lexems,and for each lexeme scanner produces output as a token of the form. like below
<token-name, attribute-value>
e.g position = initial + rate * 60
these characters are group grouped into lexemes and mapped into tokens like
- position is lexeme and mapped into token as <id, 1> where id is an abstract symbol for identifier and 1 points to the symbol table entry for position.
- initial is lexeme and mapped into token <id, 2>, where 2 points to symbol table entry for initial
my question is, how these tokens are stored into symbol table? as we are only mapping lexemes into tokens like <id , 1>, <id, 2>..etc. where are we storing values corresponding to these tokens in symbol table? I am aware of the symbol table but, can somebody please tell me the signature of ST which is used here? Is it something like <id, map<token-name, attribute-value>> ??
also for all id fields(identifiers) which data-structure is being used to store information related to identifiers like name, scope, size, dataType.
And which state ST is generated? because all stages(scanner, parser, semantic analyzer etc) in compiler design uses ST for reference
Another question is when parser asks for next input token then does the scanner reads input token from ST or from input data? Please help me to understand or attribute-value is simply contains the pointer to the symbol table?