From: Eric Mahurin Date: 2005-11-10T00:35:17+09:00 Subject: Re: RUBY GRAMMAR ------=_Part_31706_30342465.1131550512345 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Content-Disposition: inline On 11/9/05, puellula@gmail.com wrote: > > Hi! > Can you help me? I have to write a Ruby Editor + Parser with Java > Language. > I use ANTLR Parser Generator. > The grammar (a little part of Ruby grammar) I use is my hands made. I > had use ANTLR and I had write a main in Java Project. Now the problem > is: if I write a correct ruby code or not correct ruby code the result > is the same. Therefore I had think the problem is on the grammar. > Please, can you validate my Grammar. For me it's correct, but I think > the problem is this for my project. > This is the file.g how I had write my Grammar for ANTLR. > > class P extends Parser; > options { > k =3D 2; // pone a 2 il lookahead di token > exportVocab=3DRuby; // chiama il suo vocabolario "Ruby" > defaultErrorHandler =3D false; // non genera gestori di errori di > parser > buildAST=3Dtrue; // il parser costruisce un AST (Abstract Syntax > Tree) > } > > // l'intera classe > program : (CLASS IDENTIFIER body_class)*; > > // il corpo della classe > body_class : (DEF (IDENTIFIER)* body)*; > > // il corpo > body : (statement)+; > > // istruzioni che si possono trovare nel corpo del programma > statement : (IF | WHILE) condition > | LOOP_DO > | YIELD LPAREN IDENTIFIER RPAREN > | RETURN ret > | END > | IDENTIFIER ASSIGN (LPAREN)+ (IDENTIFIER | NUMBER | STRING) > (RPAREN)+ > | IDENTIFIER ASSIGN (LPAREN)+ IDENTIFIER (RPAREN)+ > | IDENTIFIER ASSIGN (LPAREN)+ instruction (RPAREN)+; > > // condizioni dei costrutti > condition : (LPAREN)+ IDENTIFIER booleano IDENTIFIER (RPAREN)+ > | (LPAREN)+ IDENTIFIER booleano NUMBER (RPAREN)+ > | (LPAREN)+ NUMBER booleano IDENTIFIER (RPAREN)+ > | (LPAREN)+ NUMBER booleano NUMBER (RPAREN)+ > | (LPAREN)+ instruction (RPAREN)+ > | IDENTIFIER booleano (LPAREN)+ instruction (RPAREN)+; > > instruction : NUMBER operator NUMBER > | IDENTIFIER operator NUMBER > | NUMBER operator IDENTIFIER; > > // operatori che si possono trovare nelle condizioni > booleano : LT > | LE > | GE > | GT > | EGUAL > | MOD > | AND > | OR > | LPAREN > | RPAREN > | DIV; > > // ci=F2 che pu=F2 essere scritto dopo il "return" > ret : (LPAREN)+ IDENTIFIER (RPAREN)+ > | (LPAREN)+ NUMBER (RPAREN)+ > | (LPAREN)+ TRUE (RPAREN)+ > | (LPAREN)+ FALSE (RPAREN)+; > > operator : DIV > | MUL > | PLUS > | SUB > | MOD; > > > > //-----------------------------------------------------------------------= ------- > // LEXER > > //-----------------------------------------------------------------------= ------- > class RubyLexer extends Lexer; > > options { > charVocabulary =3D '\0'..'\377'; > exportVocab=3DRuby; // chiama il suo vocabolario "Ruby" > testLiterals =3D false; // don't automatically test for literals > k =3D 4; // four characters of lookahead > caseSensitive =3D true; > caseSensitiveLiterals =3D false; > filter =3D true; > } > > tokens { > CLASS =3D "class"; > DEF =3D "def"; > IF =3D "if"; > RETURN =3D "return"; > END =3D "end"; > LOOP_DO =3D "loop do"; > YIELD =3D "yield"; > DO =3D "do"; > FALSE =3D "false"; > TRUE =3D "true"; > } > > // Operatori > LT : '<'; > LE : "<=3D"; > GE : ">=3D"; > GT : '>'; > EGUAL : "=3D=3D"; > DIV : '/'; > MUL : '*'; > ASSIGN : '=3D'; > LPAREN : '('; > RPAREN : ')'; > PLUS : '+'; > POINT : '.'; > AT : '@'; > OR : '|'; // questo simbolo non ha solo la funzionalit=E0 > dell'OR! > AND : '&'; > SUB : '-'; > MOD : '%'; > > NUMBER : ('0'..'9')+; > > // Identificatori > IDENTIFIER : ('a'..'z'|'A'..'Z')+ (NUMBER)?; > > // Stringhe > STRING : '"' (('a'..'z'|'A'..'Z')+)* '"' > | '\'' (('a'..'z'|'A'..'Z')+)* '\''; > > // regola per il ritorno a capo > NEWLINE : ( "\r\n" // DOS > | '\r' // MAC > | '\n' // Unix > ); > > > > > If you can help me for me it's very important. > Thank you very much! > bye, > puellula > > > Why are you posting this to ruby-talk? Wouldn't antlr-interest be better? And I think you need to do a little more debugging and narrow your question= . Asking about whether a complete grammar for a language you created (you picked a certain subset of ruby and modified it from the looks of it) seems a little silly. Although I feel like I'm doing your homework, here are a fe= w problems I see: - what's the + after the LPAREN and RPAREN for? You don't want an arbitrary and unbalanced number of them do you? - you have a strange definition for an IDENTIFIER - I haven't used antlr in a while, but I doubt "loop do" will work since it has a space in it. I think keywords are determined when an identifier is matched (looks up the identifier in a hash to see if it is a keyword instead). - in LL parsers, you typically need separate rules for each level of precedence. I don't see that above. ------=_Part_31706_30342465.1131550512345--