version 0.8.2

This commit is contained in:
rumpf_a@web.de 2009-10-21 10:20:15 +02:00
commit 053309e60a
125 changed files with 6564 additions and 1308 deletions

3
doc/apis.txt Normal file → Executable file
View file

@ -66,4 +66,7 @@ function func
coordinate coord
rectangle rect
point point
symbol sym
identifier ident
indentation indent
------------------- ------------ --------------------------------------

41
doc/effects.txt Executable file
View file

@ -0,0 +1,41 @@
=====================================================================
Side effects in Nimrod
=====================================================================
Note: Side effects are implicit produced values! Maybe they should be
explicit like in Haskell?
The idea is that side effects and partial evaluation belong together:
Iff a proc is side effect free and all its argument are evaluable at
compile time, it can be evaluated by the compiler. However, really
difficult is the ``newString`` proc: If it is simply wrapped, it
should not be evaluated at compile time! On other occasions it can
and should be evaluted:
.. code-block:: nimrod
proc toUpper(s: string): string =
result = newString(len(s))
for i in 0..len(s) - 1:
result[i] = toUpper(s[i])
No, it really can always be evaluated. The code generator should transform
``s = "\0\0\0..."`` back into ``s = newString(...)``.
``new`` cannot be evaluated at compile time either.
Raise statement
===============
It is impractical to consider ``raise`` a statement with side effects.
Solution
========
Being side effect free does not suffice for compile time evaluation. However,
the evaluator can attempt to evaluate at compile time.

View file

@ -8,7 +8,8 @@ nimrod main module: parses the command line and calls
``main.MainCommand``
main implements the top-level command dispatching
nimconf implements the config file reader
syntaxes dispatcher for the different parsers and filters
ptmplsyn standard template filter (``#! stdtempl``)
lexbase buffer handling of the lexical analyser
scanner lexical analyser
pnimsyn Nimrod's parser

232
doc/filters.txt Executable file
View file

@ -0,0 +1,232 @@
===================
Parsers and Filters
===================
.. contents::
The Nimrod compiler contains multiple parsers. (The standard is
indentation-based.) Two others are available: The `braces`:idx: parser and the
`endX`:idx: parser. Both parsers use the same lexer as the standard parser.
To use a different parser for a source file the *shebang* notation is used:
.. code-block:: nimrod
#! braces
if (x == 10) {
echo "x is ten"
} else {
echo "x isn't ten"
}
The special ``#!`` comment for specifying a parser needs to be in the first
line with no leading whitespace, unless an UNIX shebang line is used. Then the
parser shebang can occur in the second line:
.. code-block:: nimrod
#! /usr/bin/env nimrod c -r
#! braces
if (x == 10) {
echo "x is ten"
} else {
echo "x isn't ten"
}
An UNIX shebang line is defined by the pattern ``'#!' \s* '/' .*``
(``#!`` followed by optional whitespace followed by ``/``).
Filters
=======
Nimrod's shebang also supports the invokation of `source filters`:idx: before
the source code file is passed to the parser::
#! stdtmpl(subsChar = '$', metaChar = '#')
#proc generateXML(name, age: string): string =
# result = ""
<xml>
<name>$name</name>
<age>$age</age>
</xml>
Filters transform the input character stream to an in-memory output stream.
They are used to provide templating systems or preprocessors.
As the example shows, passing arguments to a filter (or parser) can be done
just like an ordinary procedure call with named or positional arguments. The
available parameters depend on the invoked filter/parser.
Pipe operator
-------------
Filters and parsers can be combined with the ``|`` `pipe operator`:idx:. Only
the last operand can be a parser because a parser returns an abstract syntax
tree which a filter cannot process::
#! strip(startswith="<") | stdtmpl | standard
#proc generateXML(name, age: string): string =
# result = ""
<xml>
<name>$name</name>
<age>$age</age>
</xml>
Available filters
=================
**Hint:** With ``--verbosity:2`` (or higher) Nimrod lists the processed code
after each filter application.
Replace filter
--------------
The `replace`:idx: filter replaces substrings in each line.
Parameters and their defaults:
``sub: string = ""``
the substring that is searched for
``by: string = ""``
the string the substring is replaced with
Strip filter
------------
The `strip`:idx: filter simply removes leading and trailing whitespace from
each line.
Parameters and their defaults:
``startswith: string = ""``
strip only the lines that start with *startswith* (ignoring leading
whitespace). If empty every line is stripped.
``leading: bool = true``
strip leading whitespace
``trailing: bool = true``
strip trailing whitespace
StdTmpl filter
--------------
The `stdtmpl`:idx: filter provides a simple templating engine for Nimrod. The
filter uses a line based parser: Lines prefixed with a *meta character*
(default: ``#``) contain Nimrod code, other lines are verbatim. Because
indentation-based parsing is not suited for a templating engine, control flow
statements need ``end X`` delimiters.
Parameters and their defaults:
``metaChar: char = '#'``
prefix for a line that contains Nimrod code
``subsChar: char = '$'``
prefix for a Nimrod expression within a template line
``conc: string = " & "``
the operation for concatenation
``emit: string = "result.add"``
the operation to emit a string literal
``toString: string = "$"``
the operation that is applied to each expression
Example::
#! stdtmpl | standard
#proc generateHTMLPage(title, currentTab, content: string,
# tabs: openArray[string]): string =
# result = ""
<head><title>$title</title></head>
<body>
<div id="menu">
<ul>
#for tab in items(tabs):
#if currentTab == tab:
<li><a id="selected"
#else:
<li><a
#end if
href="${tab}.html">$tab</a></li>
#end for
</ul>
</div>
<div id="content">
$content
A dollar: $$.
</div>
</body>
The filter transforms this into:
.. code-block:: nimrod
proc generateHTMLPage(title, currentTab, content: string,
tabs: openArray[string]): string =
result = ""
result.add("<head><title>" & $(title) & "</title></head>\n" &
"<body>\n" &
" <div id=\"menu\">\n" &
" <ul>\n")
for tab in items(tabs):
if currentTab == tab:
result.add(" <li><a id=\"selected\" \n")
else:
result.add(" <li><a\n")
#end
result.add(" href=\"" & $(tab) & ".html\">" & $(tab) & "</a></li>\n")
#end
result.add(" </ul>\n" &
" </div>\n" &
" <div id=\"content\">\n" &
" " & $(content) & "\n" &
" A dollar: $.\n" &
" </div>\n" &
"</body>\n")
Each line that does not start with the meta character (ignoring leading
whitespace) is converted to a string literal that is added to ``result``.
The substitution character introduces a Nimrod expression *e* within the
string literal. *e* is converted to a string with the *toString* operation
which defaults to ``$``. For strong type checking, set ``toString`` to the
empty string. *e* must match this PEG pattern::
e <- [a-zA-Z\128-\255][a-zA-Z0-9\128-\255_.]* / '{' x '}'
x <- '{' x+ '}' / [^}]*
To produce a single substitution character it has to be doubled: ``$$``
produces ``$``.
The template engine is quite flexible. It is easy to produce a procedure that
writes the template code directly to a file::
#! stdtmpl(emit="f.write") | standard
#proc writeHTMLPage(f: TFile, title, currentTab, content: string,
# tabs: openArray[string]) =
<head><title>$title</title></head>
<body>
<div id="menu">
<ul>
#for tab in items(tabs):
#if currentTab == tab:
<li><a id="selected"
#else:
<li><a
#end if
href="${tab}.html" title = "$title - $tab">$tab</a></li>
#end for
</ul>
</div>
<div id="content">
$content
A dollar: $$.
</div>
</body>

179
doc/gramcurl.txt Executable file
View file

@ -0,0 +1,179 @@
module ::= stmt*
comma ::= ',' [COMMENT] [IND]
operator ::= OP0 | OR | XOR | AND | OP3 | OP4 | OP5 | OP6 | OP7
| 'is' | 'isnot' | 'in' | 'notin'
| 'div' | 'mod' | 'shl' | 'shr' | 'not'
prefixOperator ::= OP0 | OP3 | OP4 | OP5 | OP6 | OP7 | 'not'
optInd ::= [COMMENT] [IND]
lowestExpr ::= orExpr (OP0 optInd orExpr)*
orExpr ::= andExpr (OR | 'xor' optInd andExpr)*
andExpr ::= cmpExpr ('and' optInd cmpExpr)*
cmpExpr ::= ampExpr (OP3 | 'is' | 'isnot' | 'in' | 'notin' optInd ampExpr)*
ampExpr ::= plusExpr (OP4 optInd plusExpr)*
plusExpr ::= mulExpr (OP5 optInd mulExpr)*
mulExpr ::= dollarExpr (OP6 | 'div' | 'mod' | 'shl' | 'shr' optInd dollarExpr)*
dollarExpr ::= primary (OP7 optInd primary)*
indexExpr ::= '..' [expr] | expr ['=' expr | '..' expr]
castExpr ::= 'cast' '[' optInd typeDesc [SAD] ']' '(' optInd expr [SAD] ')'
addrExpr ::= 'addr' '(' optInd expr ')'
symbol ::= '`' (KEYWORD | IDENT | operator | '(' ')'
| '[' ']' | '=' | literal)+ '`'
| IDENT
primaryPrefix ::= (prefixOperator | 'bind') optInd
primarySuffix ::= '.' optInd symbol
| '(' optInd namedExprList [SAD] ')'
| '[' optInd [indexExpr (comma indexExpr)* [comma]] [SAD] ']'
| '^'
| pragma
primary ::= primaryPrefix* (symbol | constructor | castExpr | addrExpr)
primarySuffix*
literal ::= INT_LIT | INT8_LIT | INT16_LIT | INT32_LIT | INT64_LIT
| FLOAT_LIT | FLOAT32_LIT | FLOAT64_LIT
| STR_LIT | RSTR_LIT | TRIPLESTR_LIT
| CHAR_LIT
| NIL
constructor ::= literal
| '[' optInd colonExprList [SAD] ']'
| '{' optInd sliceExprList [SAD] '}'
| '(' optInd colonExprList [SAD] ')'
colonExpr ::= expr [':' expr]
colonExprList ::= [colonExpr (comma colonExpr)* [comma]]
namedExpr ::= expr ['=' expr]
namedExprList ::= [namedExpr (comma namedExpr)* [comma]]
sliceExpr ::= expr ['..' expr]
sliceExprList ::= [sliceExpr (comma sliceExpr)* [comma]]
exprOrType ::= lowestExpr
| 'if' '(' expr ')' expr ('elif' '(' expr ')' expr)* 'else' expr
| 'var' exprOrType
| 'ref' exprOrType
| 'ptr' exprOrType
| 'type' exprOrType
| 'tuple' tupleDesc
expr ::= exprOrType
| 'proc' paramList [pragma] ['=' stmt]
qualifiedIdent ::= symbol ['.' symbol]
typeDesc ::= exprOrType
| 'proc' paramList [pragma]
macroStmt ::= '{' [stmt] '}' ('of' [sliceExprList] stmt
|'elif' '(' expr ')' stmt
|'except' '(' exceptList ')' stmt )*
['else' stmt]
simpleStmt ::= returnStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
complexStmt ::= ifStmt | whileStmt | caseStmt | tryStmt | forStmt
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl | methodDecl
| constSection | typeSection | whenStmt | varSection
stmt ::= simpleStmt
| indPush (complexStmt | simpleStmt) (';' (complexStmt | simpleStmt))*
DED indPop
exprStmt ::= lowestExpr ['=' expr | [expr (comma expr)*] [macroStmt]]
returnStmt ::= 'return' [expr]
yieldStmt ::= 'yield' expr
discardStmt ::= 'discard' expr
raiseStmt ::= 'raise' [expr]
breakStmt ::= 'break' [symbol]
continueStmt ::= 'continue'
ifStmt ::= 'if' '(' expr ')' stmt ('elif' '(' expr ')' stmt)* ['else' stmt]
whenStmt ::= 'when' '(' expr ')' stmt ('elif' '(' expr ')' stmt)* ['else' stmt]
caseStmt ::= 'case' '(' expr ')' ('of' sliceExprList ':' stmt)*
('elif' '(' expr ')' stmt)*
['else' stmt]
whileStmt ::= 'while' '(' expr ')' stmt
forStmt ::= 'for' '(' symbol (comma symbol)* 'in' expr ['..' expr] ')' stmt
exceptList ::= [qualifiedIdent (comma qualifiedIdent)*]
tryStmt ::= 'try' stmt
('except' '(' exceptList ')' stmt)*
['finally' stmt]
asmStmt ::= 'asm' [pragma] (STR_LIT | RSTR_LIT | TRIPLESTR_LIT)
blockStmt ::= 'block' [symbol] stmt
filename ::= symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT
importStmt ::= 'import' filename (comma filename)*
includeStmt ::= 'include' filename (comma filename)*
fromStmt ::= 'from' filename 'import' symbol (comma symbol)*
pragma ::= '{.' optInd (colonExpr [comma])* [SAD] ('.}' | '}')
param ::= symbol (comma symbol)* (':' typeDesc ['=' expr] | '=' expr)
paramList ::= ['(' [param (comma param)*] [SAD] ')'] [':' typeDesc]
genericParam ::= symbol [':' typeDesc] ['=' expr]
genericParams ::= '[' genericParam (comma genericParam)* [SAD] ']'
routineDecl := symbol ['*'] [genericParams] paramList [pragma] ['=' stmt]
procDecl ::= 'proc' routineDecl
macroDecl ::= 'macro' routineDecl
iteratorDecl ::= 'iterator' routineDecl
templateDecl ::= 'template' routineDecl
methodDecl ::= 'method' routineDecl
colonAndEquals ::= [':' typeDesc] '=' expr
constDecl ::= symbol ['*'] [pragma] colonAndEquals ';' [COMMENT]
constSection ::= 'const' [COMMENT] (constDecl | '{' constDecl+ '}')
typeDef ::= typeDesc | objectDef | enumDef | 'distinct' typeDesc
objectField ::= symbol ['*'] [pragma]
objectIdentPart ::= objectField (comma objectField)* ':' typeDesc
[COMMENT|IND COMMENT]
objectWhen ::= 'when' expr ':' [COMMENT] objectPart
('elif' expr ':' [COMMENT] objectPart)*
['else' ':' [COMMENT] objectPart]
objectCase ::= 'case' expr ':' typeDesc [COMMENT]
('of' sliceExprList ':' [COMMENT] objectPart)*
['else' ':' [COMMENT] objectPart]
objectPart ::= objectWhen | objectCase | objectIdentPart | 'nil'
| indPush objectPart (SAD objectPart)* DED indPop
tupleDesc ::= '[' optInd [param (comma param)*] [SAD] ']'
objectDef ::= 'object' [pragma] ['of' typeDesc] objectPart
enumField ::= symbol ['=' expr]
enumDef ::= 'enum' ['of' typeDesc] (enumField [comma] [COMMENT | IND COMMENT])+
typeDecl ::= COMMENT
| symbol ['*'] [genericParams] ['=' typeDef] [COMMENT | IND COMMENT]
typeSection ::= 'type' indPush typeDecl (SAD typeDecl)* DED indPop
colonOrEquals ::= ':' typeDesc ['=' expr] | '=' expr
varField ::= symbol ['*'] [pragma]
varPart ::= symbol (comma symbol)* colonOrEquals [COMMENT | IND COMMENT]
varSection ::= 'var' (varPart
| indPush (COMMENT|varPart)
(SAD (COMMENT|varPart))* DED indPop)

View file

@ -19,25 +19,23 @@ plusExpr ::= mulExpr (OP5 optInd mulExpr)*
mulExpr ::= dollarExpr (OP6 | 'div' | 'mod' | 'shl' | 'shr' optInd dollarExpr)*
dollarExpr ::= primary (OP7 optInd primary)*
namedTypeOrExpr ::=
'..' [expr]
| expr ['=' (expr ['..' expr] | typeDescK | '..' [expr]) | '..' [expr]]
| typeDescK
indexExpr ::= '..' [expr] | expr ['=' expr | '..' expr]
castExpr ::= 'cast' '[' optInd typeDesc [SAD] ']' '(' optInd expr [SAD] ')'
addrExpr ::= 'addr' '(' optInd expr ')'
symbol ::= '`' (KEYWORD | IDENT | operator | '(' ')'
| '[' ']' | '=' | literal)+ '`'
| IDENT
primary ::= ((prefixOperator | 'bind') optInd)* (symbol | constructor |
castExpr | addrExpr) (
'.' optInd symbol
| '(' optInd namedExprList [SAD] ')'
| '[' optInd
[namedTypeOrExpr (comma namedTypeOrExpr)* [comma]]
[SAD] ']'
| '^'
| pragma)*
primaryPrefix ::= (prefixOperator | 'bind') optInd
primarySuffix ::= '.' optInd symbol
| '(' optInd namedExprList [SAD] ')'
| '[' optInd [indexExpr (comma indexExpr)* [comma]] [SAD] ']'
| '^'
| pragma
primary ::= primaryPrefix* (symbol | constructor | castExpr | addrExpr)
primarySuffix*
literal ::= INT_LIT | INT8_LIT | INT16_LIT | INT32_LIT | INT64_LIT
| FLOAT_LIT | FLOAT32_LIT | FLOAT64_LIT
@ -59,24 +57,21 @@ namedExprList ::= [namedExpr (comma namedExpr)* [comma]]
sliceExpr ::= expr ['..' expr]
sliceExprList ::= [sliceExpr (comma sliceExpr)* [comma]]
anonymousProc ::= 'lambda' paramList [pragma] '=' stmt
expr ::= lowestExpr
| anonymousProc
| 'if' expr ':' expr ('elif' expr ':' expr)* 'else' ':' expr
exprOrType ::= lowestExpr
| 'if' expr ':' expr ('elif' expr ':' expr)* 'else' ':' expr
| 'var' exprOrType
| 'ref' exprOrType
| 'ptr' exprOrType
| 'type' exprOrType
| 'tuple' tupleDesc
namedTypeDesc ::= typeDescK | expr ['=' (typeDescK | expr)]
namedTypeDescList ::= [namedTypeDesc (comma namedTypeDesc)* [comma]]
expr ::= exprOrType
| 'proc' paramList [pragma] ['=' stmt]
qualifiedIdent ::= symbol ['.' symbol]
typeDescK ::= 'var' typeDesc
| 'ref' typeDesc
| 'ptr' typeDesc
| 'type' expr
| 'tuple' tupleDesc
| 'proc' paramList [pragma]
typeDesc ::= typeDescK | primary
typeDesc ::= exprOrType
| 'proc' paramList [pragma]
macroStmt ::= ':' [stmt] ('of' [sliceExprList] ':' stmt
|'elif' expr ':' stmt
@ -84,20 +79,20 @@ macroStmt ::= ':' [stmt] ('of' [sliceExprList] ':' stmt
['else' ':' stmt]
simpleStmt ::= returnStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
complexStmt ::= ifStmt | whileStmt | caseStmt | tryStmt | forStmt
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl
| constSection | typeSection | whenStmt | varSection
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl | methodDecl
| constSection | typeSection | whenStmt | varSection
indPush ::= IND # and push indentation onto the stack
indPop ::= # pop indentation from the stack
@ -141,25 +136,24 @@ paramList ::= ['(' [param (comma param)*] [SAD] ')'] [':' typeDesc]
genericParam ::= symbol [':' typeDesc] ['=' expr]
genericParams ::= '[' genericParam (comma genericParam)* [SAD] ']'
procDecl ::= 'proc' symbol ['*'] [genericParams] paramList [pragma]
['=' stmt]
macroDecl ::= 'macro' symbol ['*'] [genericParams] paramList [pragma]
['=' stmt]
iteratorDecl ::= 'iterator' symbol ['*'] [genericParams] paramList [pragma]
['=' stmt]
templateDecl ::= 'template' symbol ['*'] [genericParams] paramList [pragma]
['=' stmt]
routineDecl := symbol ['*'] [genericParams] paramList [pragma] ['=' stmt]
procDecl ::= 'proc' routineDecl
macroDecl ::= 'macro' routineDecl
iteratorDecl ::= 'iterator' routineDecl
templateDecl ::= 'template' routineDecl
methodDecl ::= 'method' routineDecl
colonAndEquals ::= [':' typeDesc] '=' expr
constDecl ::= symbol ['*'] [pragma] colonAndEquals [COMMENT | IND COMMENT]
| COMMENT
constSection ::= 'const' indPush constDecl (SAD constDecl)* DED indPop
typeDef ::= typeDesc | objectDef | enumDef | 'abstract' typeDesc
typeDef ::= typeDesc | objectDef | enumDef | 'distinct' typeDesc
objectField ::= symbol ['*'] [pragma]
objectIdentPart ::=
objectField (comma objectField)* ':' typeDesc [COMMENT|IND COMMENT]
objectIdentPart ::= objectField (comma objectField)* ':' typeDesc
[COMMENT|IND COMMENT]
objectWhen ::= 'when' expr ':' [COMMENT] objectPart
('elif' expr ':' [COMMENT] objectPart)*

View file

@ -66,8 +66,7 @@ Generic Operating System Services
* `os <os.html>`_
Basic operating system facilities like retrieving environment variables,
reading command line arguments, working with directories, running shell
commands, etc. This module is -- like any other basic library --
platform independant.
commands, etc.
* `osproc <osproc.html>`_
Module for process communication beyond ``os.execShellCmd``.

View file

@ -342,7 +342,7 @@ Syntax
======
This section lists Nimrod's standard syntax in ENBF. How the parser receives
indentation tokens is already described in the Lexical Analysis section.
indentation tokens is already described in the `Lexical Analysis`_ section.
Nimrod allows user-definable operators.
Binary operators have 8 different levels of precedence. For user-defined
@ -363,8 +363,7 @@ Precedence level Operators First characte
================ ============================================== ================== ===============
The grammar's start symbol is ``module``. The grammar is LL(1) and therefore
not ambiguous.
The grammar's start symbol is ``module``.
.. include:: grammar.txt
:literal:
@ -875,8 +874,7 @@ Procedural type
~~~~~~~~~~~~~~~
A `procedural type`:idx: is internally a pointer to a procedure. ``nil`` is
an allowed value for variables of a procedural type. Nimrod uses procedural
types to achieve `functional`:idx: programming techniques. Dynamic dispatch
for OOP constructs can also be implemented with procedural types.
types to achieve `functional`:idx: programming techniques.
Example:
@ -946,6 +944,16 @@ each other:
Most calling conventions exist only for the Windows 32-bit platform.
Assigning/passing a procedure to a procedural variable is only allowed if one
of the following conditions hold:
1) The procedure that is accessed resists in the current module.
2) The procedure is marked with the ``procvar`` pragma (see `procvar pragma`_).
3) The procedure has a calling convention that differs from ``nimcall``.
4) The procedure is anonymous.
These rules should prevent the case that extending a non-``procvar``
procedure with default parameters breaks client code.
Distinct type
~~~~~~~~~~~~~
@ -1054,7 +1062,7 @@ describe the type checking done by the compiler.
Type equality
~~~~~~~~~~~~~
Nimrod uses structural type equivalence for most types. Only for objects,
enumerations and abstract types name equivalence is used. The following
enumerations and distinct types name equivalence is used. The following
algorithm determines type equality:
.. code-block:: nimrod
@ -1800,6 +1808,77 @@ Even more elegant is to use `tuple unpacking`:idx: to access the tuple's fields:
assert y == 3
Multi-methods
~~~~~~~~~~~~~
Procedures always use static dispatch. Dynamic dispatch is achieved by
`multi-methods`:idx:.
.. code-block:: nimrod
type
TExpr = object ## abstract base class for an expression
TLiteral = object of TExpr
x: int
TPlusExpr = object of TExpr
a, b: ref TExpr
method eval(e: ref TExpr): int =
# override this base method
quit "to override!"
method eval(e: ref TLiteral): int = return e.x
method eval(e: ref TPlusExpr): int =
# watch out: relies on dynamic binding
return eval(e.a) + eval(e.b)
proc newLit(x: int): ref TLiteral =
new(result)
result.x = x
proc newPlus(a, b: ref TExpr): ref TPlusExpr =
new(result)
result.a = a
result.b = b
echo eval(newPlus(newPlus(newLit(1), newLit(2)), newLit(4)))
In the example the constructors ``newLit`` and ``newPlus`` are procs
because they should use static binding, but ``eval`` is a method because it
requires dynamic binding.
In a multi-method all parameters that have an object type are used for the
dispatching:
.. code-block:: nimrod
type
TThing = object
TUnit = object of TThing
x: int
method collide(a, b: TThing) {.inline.} =
quit "to override!"
method collide(a: TThing, b: TUnit) {.inline.} =
echo "1"
method collide(a: TUnit, b: TThing) {.inline.} =
echo "2"
var
a, b: TUnit
collide(a, b) # output: 2
Invokation of a multi-method cannot be ambiguous: Collide 2 is prefered over
collide 1 because the resolution works from left to right.
Thus ``TUnit, TThing`` is prefered over ``TThing, TUnit``.
**Perfomance note**: Nimrod does not produce a virtual method table, but
generates dispatch trees. This avoids the expensive indirect branch for method
calls and enables inlining. However, other optimizations like compile time
evaluation or dead code elimination do not work with methods.
Iterators and the for statement
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
@ -2275,6 +2354,12 @@ error to mark a proc/iterator to have no side effect if the compiler cannot
verify this.
procvar pragma
--------------
The `procvar`:idx: pragma is used to mark a proc that it can be passed to a
procedural variable.
compileTime pragma
------------------
The `compileTime`:idx: pragma is used to mark a proc to be used at compile

View file

@ -72,8 +72,7 @@ New Pragmas and Options
-----------------------
Because Nimrod generates C code it needs some "red tape" to work properly.
Thus lots of options and pragmas for tweaking the generated C code are
available.
Lots of options and pragmas for tweaking the generated C code are available.
Importc Pragma
~~~~~~~~~~~~~~
@ -137,7 +136,7 @@ and instead the generated code should contain an ``#include``:
PFile {.importc: "FILE*", header: "<stdio.h>".} = distinct pointer
# import C's FILE* type; Nimrod will treat it as a new pointer type
The ``header`` pragma expects always a string constant. The string contant
The ``header`` pragma always expects a string constant. The string contant
contains the header file: As usual for C, a system header file is enclosed
in angle brackets: ``<>``. If no angle brackets are given, Nimrod
encloses the header file in ``""`` in the generated C code.
@ -145,9 +144,9 @@ encloses the header file in ``""`` in the generated C code.
Varargs Pragma
~~~~~~~~~~~~~~
The `varargs`:idx: pragma can be applied to procedures only. It tells Nimrod
that the proc can take a variable number of parameters after the last
specified parameter. Nimrod string values will be converted to C
The `varargs`:idx: pragma can be applied to procedures only (and procedure
types). It tells Nimrod that the proc can take a variable number of parameters
after the last specified parameter. Nimrod string values will be converted to C
strings automatically:
.. code-block:: Nimrod
@ -218,7 +217,7 @@ collector to not consider objects of this type as part of a cycle:
data: string
In the example a tree structure is declared with the ``TNode`` type. Note that
the type definition is recursive thus the GC has to assume that objects of
the type definition is recursive and the GC has to assume that objects of
this type may form a cyclic graph. The ``acyclic`` pragma passes the
information that this cannot happen to the GC. If the programmer uses the
``acyclic`` pragma for data types that are in reality cyclic, the GC may leak

180
doc/pegdocs.txt Executable file
View file

@ -0,0 +1,180 @@
PEG syntax and semantics
========================
A PEG (Parsing expression grammar) is a simple deterministic grammar, that can
be directly used for parsing. The current implementation has been designed as
a more powerful replacement for regular expressions. UTF-8 is supported.
The notation used for a PEG is similar to that of EBNF:
=============== ============================================================
notation meaning
=============== ============================================================
``A / ... / Z`` Ordered choice: Apply expressions `A`, ..., `Z`, in this
order, to the text ahead, until one of them succeeds and
possibly consumes some text. Indicate success if one of
expressions succeeded. Otherwise do not consume any text
and indicate failure.
``A ... Z`` Sequence: Apply expressions `A`, ..., `Z`, in this order,
to consume consecutive portions of the text ahead, as long
as they succeed. Indicate success if all succeeded.
Otherwise do not consume any text and indicate failure.
The sequence's precedence is higher than that of ordered
choice: ``A B / C`` means ``(A B) / Z`` and
not ``A (B / Z)``.
``(E)`` Grouping: Parenthesis can be used to change
operator priority.
``{E}`` Capture: Apply expression `E` and store the substring
that matched `E` into a *capture* that can be accessed
after the matching process.
``&E`` And predicate: Indicate success if expression `E` matches
the text ahead; otherwise indicate failure. Do not consume
any text.
``!E`` Not predicate: Indicate failure if expression E matches the
text ahead; otherwise indicate success. Do not consume any
text.
``E+`` One or more: Apply expression `E` repeatedly to match
the text ahead, as long as it succeeds. Consume the matched
text (if any) and indicate success if there was at least
one match. Otherwise indicate failure.
``E*`` Zero or more: Apply expression `E` repeatedly to match
the text ahead, as long as it succeeds. Consume the matched
text (if any). Always indicate success.
``E?`` Zero or one: If expression `E` matches the text ahead,
consume it. Always indicate success.
``[s]`` Character class: If the character ahead appears in the
string `s`, consume it and indicate success. Otherwise
indicate failure.
``[a-b]`` Character range: If the character ahead is one from the
range `a` through `b`, consume it and indicate success.
Otherwise indicate failure.
``'s'`` String: If the text ahead is the string `s`, consume it
and indicate success. Otherwise indicate failure.
``i's'`` String match ignoring case.
``y's'`` String match ignoring style.
``v's'`` Verbatim string match: Use this to override a global
``\i`` or ``\y`` modifier.
``.`` Any character: If there is a character ahead, consume it
and indicate success. Otherwise (that is, at the end of
input) indicate failure.
``_`` Any Unicode character: If there is an UTF-8 character
ahead, consume it and indicate success. Otherwise indicate
failure.
``A <- E`` Rule: Bind the expression `E` to the *nonterminal symbol*
`A`. **Left recursive rules are not possible and crash the
matching engine.**
``\identifier`` Built-in macro for a longer expression.
``\ddd`` Character with decimal code *ddd*.
``\"``, etc Literal ``"``, etc.
=============== ============================================================
Built-in macros
---------------
============== ============================================================
macro meaning
============== ============================================================
``\d`` any decimal digit: ``[0-9]``
``\D`` any character that is not a decimal digit: ``[^0-9]``
``\s`` any whitespace character: ``[ \9-\13]``
``\S`` any character that is not a whitespace character:
``[^ \9-\13]``
``\w`` any "word" character: ``[a-zA-Z_]``
``\W`` any "non-word" character: ``[^a-zA-Z_]``
``\n`` any newline combination: ``\10 / \13\10 / \13``
``\i`` ignore case for matching; use this at the start of the PEG
``\y`` ignore style for matching; use this at the start of the PEG
``\ident`` a standard ASCII identifier: ``[a-zA-Z_][a-zA-Z_0-9]*``
============== ============================================================
A backslash followed by a letter is a built-in macro, otherwise it
is used for ordinary escaping:
============== ============================================================
notation meaning
============== ============================================================
``\\`` a single backslash
``\*`` same as ``'*'``
``\t`` not a tabulator, but an (unknown) built-in
============== ============================================================
Supported PEG grammar
---------------------
The PEG parser implements this grammar (written in PEG syntax)::
# Example grammar of PEG in PEG syntax.
# Comments start with '#'.
# First symbol is the start symbol.
grammar <- rule* / expr
identifier <- [A-Za-z][A-Za-z0-9_]*
charsetchar <- "\\" . / [^\]]
charset <- "[" "^"? (charsetchar ("-" charsetchar)?)+ "]"
stringlit <- identifier? ("\"" ("\\" . / [^"])* "\"" /
"'" ("\\" . / [^'])* "'")
builtin <- "\\" identifier / [^\13\10]
comment <- '#' !\n* \n
ig <- (\s / comment)* # things to ignore
rule <- identifier \s* "<-" expr ig
identNoArrow <- identifier !(\s* "<-")
primary <- (ig '&' / ig '!')* ((ig identNoArrow / ig charset / ig stringlit
/ ig builtin / ig '.' / ig '_'
/ (ig "(" expr ig ")"))
(ig '?' / ig '*' / ig '+')*)
# Concatenation has higher priority than choice:
# ``a b / c`` means ``(a b) / c``
seqExpr <- primary+
expr <- seqExpr (ig "/" expr)*
Examples
--------
Check if `s` matches Nimrod's "while" keyword:
.. code-block:: nimrod
s =~ peg" y'while'"
Exchange (key, val)-pairs:
.. code-block:: nimrod
"key: val; key2: val2".replace(peg"{\ident} \s* ':' \s* {\ident}", "$2: $1")
Determine the ``#include``'ed files of a C file:
.. code-block:: nimrod
for line in lines("myfile.c"):
if line =~ peg"""s <- ws '#include' ws '"' {[^"]+} '"' ws
comment <- '/*' (!'*/' . )* '*/' / '//' .*
ws <- (comment / \s+)* """:
echo matches[0]
PEG vs regular expression
-------------------------
As a regular expression ``\[.*\]`` maches longest possible text between ``'['``
and ``']'``. As a PEG it never matches anything, because a PEG is
deterministic: ``.*`` consumes the rest of the input, so ``\]`` never matches.
As a PEG this needs to be written as: ``\[ ( !\] . )* \]``
Note that the regular expression does not behave as intended either:
``*`` should not be greedy, so ``\[.*?\]`` should be used.
PEG construction
----------------
There are two ways to construct a PEG in Nimrod code:
(1) Parsing a string into an AST which consists of `TPeg` nodes with the
`peg` proc.
(2) Constructing the AST directly with proc calls. This method does not
support constructing rules, only simple expressions and is not as
convenient. It's only advantage is that it does not pull in the whole PEG
parser into your executable.

File diff suppressed because it is too large Load diff

View file

@ -1,4 +1,4 @@
========================
========================
Nimrod Tutorial (Part I)
========================
@ -10,7 +10,7 @@ Nimrod Tutorial (Part I)
Introduction
============
"Before you run you must learn to walk."
"Der Mensch ist doch ein Augentier -- schöne Dinge wünsch ich mir."
This document is a tutorial for the programming language *Nimrod*. After this
tutorial you will have a decent knowledge about Nimrod. This tutorial assumes
@ -34,8 +34,8 @@ Save this code to the file "greetings.nim". Now compile and run it::
nimrod compile --run greetings.nim
As you see, with the ``--run`` switch Nimrod executes the file automatically
after compilation. You can even give your program command line arguments by
With the ``--run`` switch Nimrod executes the file automatically
after compilation. You can give your program command line arguments by
appending them after the filename::
nimrod compile --run greetings.nim arg1 arg2

View file

@ -177,8 +177,7 @@ bound to a class. This has disadvantages:
* Adding a method to a class the programmer has no control over is
impossible or needs ugly workarounds.
* Often it is unclear where the method should belong to: Is
``join`` a string method or an array method? Should the complex
``vertexCover`` algorithm really be a method of the ``graph`` class?
``join`` a string method or an array method?
Nimrod avoids these problems by not assigning methods to a class. All methods
in Nimrod are `multi-methods`:idx:. As we will see later, multi-methods are
@ -206,7 +205,7 @@ for any type:
(Another way to look at the method call syntax is that it provides the missing
postfix notation.)
So code that looks "pure object oriented" is easy to write:
So "pure object oriented" code is easy to write:
.. code-block:: nimrod
import strutils
@ -277,7 +276,7 @@ already provides ``v[]`` access.
Dynamic dispatch
----------------
Procedures always use static dispatch. To get dynamic dispatch, replace the
Procedures always use static dispatch. For dynamic dispatch replace the
``proc`` keyword by ``method``:
.. code-block:: nimrod