version 0.7.0

This commit is contained in:
Andreas Rumpf 2008-11-16 22:08:15 +01:00
commit 8b2a9401a1
185 changed files with 21451 additions and 24296 deletions

View file

@ -11,6 +11,7 @@ ast type definitions of the abstract syntax tree (AST) and
node constructors
astalgo algorithms for containers of AST nodes; converting the
AST to YAML; the symbol table
passes implement the passes managemer for passes over the AST
trees few algorithms for nodes; this module is less important
types module for traversing type graphs; also contain several
helpers for dealing with types
@ -23,19 +24,15 @@ semtypes contains the semantic checking phase for types
idents implements a general mapping from identifiers to an internal
representation (``PIdent``) that is used, so that a simple
pointer comparison suffices to say whether two Nimrod
identifiers are equivalent
id-comparison suffices to say whether two Nimrod identifiers
are equivalent
ropes implements long strings using represented as trees for
lazy evaluation; used mainly by the code generators
ccgobj contains type definitions neeeded for C code generation
and some helpers
ccgmangl contains the name mangler for converting Nimrod
identifiers to their C counterparts
ccgutils contains helpers for the C code generator
ccgtemps contains the handling of temporary variables for the
C code generator
ccgtypes the generator for C types
ccgstmts the generator for statements
ccgexprs the generator for expressions

View file

@ -1,187 +1,187 @@
module ::= ([COMMENT] [SAD] stmt)*
optComma ::= [ ',' ] [COMMENT] [IND]
operator ::= OP0 | OR | XOR | AND | OP3 | OP4 | OP5 | IS | ISNOT | IN | NOTIN
| OP6 | DIV | MOD | SHL | SHR | OP7 | NOT
prefixOperator ::= OP0 | OP3 | OP4 | OP5 | OP6 | OP7 | NOT
optInd ::= [COMMENT] [IND]
lowestExpr ::= orExpr ( OP0 optInd orExpr )*
orExpr ::= andExpr ( OR | XOR optInd andExpr )*
andExpr ::= cmpExpr ( AND optInd cmpExpr )*
cmpExpr ::= ampExpr ( OP3 | IS | ISNOT | IN | NOTIN optInd ampExpr )*
ampExpr ::= plusExpr ( OP4 optInd plusExpr )*
plusExpr ::= mulExpr ( OP5 optInd mulExpr )*
mulExpr ::= dollarExpr ( OP6 | DIV | MOD | SHL | SHR optInd dollarExpr )*
dollarExpr ::= primary ( OP7 optInd primary )*
namedTypeOrExpr ::=
DOTDOT [expr]
| expr [EQUALS (expr [DOTDOT expr] | typeDescK | DOTDOT [expr] )
| DOTDOT [expr]]
| typeDescK
castExpr ::= CAST BRACKET_LE optInd typeDesc BRACKERT_RI
PAR_LE optInd expr PAR_RI
addrExpr ::= ADDR PAR_LE optInd expr PAR_RI
symbol ::= ACC (KEYWORD | IDENT | operator | PAR_LE PAR_RI
| BRACKET_LE BRACKET_RI) ACC | IDENT
accExpr ::= KEYWORD | IDENT | operator [DOT KEYWORD | IDENT | operator]
paramList
primary ::= ( prefixOperator optInd )* ( IDENT | literal | ACC accExpr ACC
| castExpr | addrExpr ) (
DOT optInd symbol
#| CURLY_LE namedTypeDescList CURLY_RI
| PAR_LE optInd
namedExprList
PAR_RI
| BRACKET_LE optInd
(namedTypeOrExpr optComma)*
BRACKET_RI
| CIRCUM
| pragma )*
literal ::= INT_LIT | INT8_LIT | INT16_LIT | INT32_LIT | INT64_LIT
| FLOAT_LIT | FLOAT32_LIT | FLOAT64_LIT
| STR_LIT | RSTR_LIT | TRIPLESTR_LIT
| CHAR_LIT | RCHAR_LIT
| NIL
| BRACKET_LE optInd (expr [COLON expr] optComma )* BRACKET_RI # []-Constructor
| CURLY_LE optInd (expr [DOTDOT expr] optComma )* CURLY_RI # {}-Constructor
| PAR_LE optInd (expr [COLON expr] optComma )* PAR_RI # ()-Constructor
exprList ::= ( expr optComma )*
namedExpr ::= expr [EQUALS expr] # actually this is symbol EQUALS expr|expr
namedExprList ::= ( namedExpr optComma )*
exprOrSlice ::= expr [ DOTDOT expr ]
sliceList ::= ( exprOrSlice optComma )+
anonymousProc ::= LAMBDA paramList [pragma] EQUALS stmt
expr ::= lowestExpr
| anonymousProc
| IF expr COLON expr
(ELIF expr COLON expr)*
ELSE COLON expr
namedTypeDesc ::= typeDescK | expr [EQUALS (typeDescK | expr)]
namedTypeDescList ::= ( namedTypeDesc optComma )*
qualifiedIdent ::= symbol [ DOT symbol ]
typeDescK ::= VAR typeDesc
| REF typeDesc
| PTR typeDesc
| TYPE expr
| TUPLE tupleDesc
| PROC paramList [pragma]
typeDesc ::= typeDescK | primary
optSemicolon ::= [SEMICOLON]
macroStmt ::= COLON [stmt] (OF [sliceList] COLON stmt
| ELIF expr COLON stmt
| EXCEPT exceptList COLON stmt )*
[ELSE COLON stmt]
simpleStmt ::= returnStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
complexStmt ::= ifStmt | whileStmt | caseStmt | tryStmt | forStmt
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl
| constSection | typeSection | whenStmt | varSection
indPush ::= IND # push
stmt ::= simpleStmt [SAD]
| indPush (complexStmt | simpleStmt)
([SAD] (complexStmt | simpleStmt) )*
DED
exprStmt ::= lowestExpr [EQUALS expr | (expr optComma)* [macroStmt]]
returnStmt ::= RETURN [expr]
yieldStmt ::= YIELD expr
discardStmt ::= DISCARD expr
raiseStmt ::= RAISE [expr]
breakStmt ::= BREAK [symbol]
continueStmt ::= CONTINUE
ifStmt ::= IF expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
whenStmt ::= WHEN expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
caseStmt ::= CASE expr (OF sliceList COLON stmt)*
(ELIF expr COLON stmt)*
[ELSE COLON stmt]
whileStmt ::= WHILE expr COLON stmt
forStmt ::= FOR (symbol optComma)+ IN expr [DOTDOT expr] COLON stmt
exceptList ::= (qualifiedIdent optComma)*
tryStmt ::= TRY COLON stmt
(EXCEPT exceptList COLON stmt)*
[FINALLY COLON stmt]
asmStmt ::= ASM [pragma] (STR_LIT | RSTR_LIT | TRIPLESTR_LIT)
blockStmt ::= BLOCK [symbol] COLON stmt
importStmt ::= IMPORT ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) [AS symbol] optComma)+
includeStmt ::= INCLUDE ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) optComma)+
fromStmt ::= FROM (symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) IMPORT (symbol optComma)+
pragma ::= CURLYDOT_LE (expr [COLON expr] optComma)+ (CURLYDOT_RI | CURLY_RI)
paramList ::= [PAR_LE ((symbol optComma)+ COLON typeDesc optComma)* PAR_RI] [COLON typeDesc]
genericParams ::= BRACKET_LE (symbol [EQUALS typeDesc] )* BRACKET_RI
procDecl ::= PROC symbol ["*"] [genericParams]
paramList [pragma]
[EQUALS stmt]
macroDecl ::= MACRO symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
iteratorDecl ::= ITERATOR symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
templateDecl ::= TEMPLATE symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
colonAndEquals ::= [COLON typeDesc] EQUALS expr
constDecl ::= symbol ["*"] [pragma] colonAndEquals [COMMENT | IND COMMENT]
| COMMENT
constSection ::= CONST indPush constDecl (SAD constDecl)* DED
typeDef ::= typeDesc | objectDef | enumDef
objectIdentPart ::=
(symbol ["*" | "-"] [pragma] optComma)+ COLON typeDesc [COMMENT | IND COMMENT]
objectWhen ::= WHEN expr COLON [COMMENT] objectPart
(ELIF expr COLON [COMMENT] objectPart)*
[ELSE COLON [COMMENT] objectPart]
objectCase ::= CASE expr COLON typeDesc [COMMENT]
(OF sliceList COLON [COMMENT] objectPart)*
[ELSE COLON [COMMENT] objectPart]
objectPart ::= objectWhen | objectCase | objectIdentPart
| indPush objectPart (SAD objectPart)* DED
tupleDesc ::= BRACKET_LE optInd ((symbol optComma)+ COLON typeDesc optComma)* BRACKET_RI
objectDef ::= OBJECT [pragma] [OF typeDesc] objectPart
enumDef ::= ENUM [OF typeDesc] (symbol [EQUALS expr] optComma [COMMENT | IND COMMENT])+
typeDecl ::= COMMENT
| symbol ["*"] [genericParams] [EQUALS typeDef] [COMMENT | IND COMMENT]
typeSection ::= TYPE indPush typeDecl (SAD typeDecl)* DED
colonOrEquals ::= COLON typeDesc [EQUALS expr] | EQUALS expr
varPart ::= (symbol ["*" | "-"] [pragma] optComma)+ colonOrEquals [COMMENT | IND COMMENT]
varSection ::= VAR (varPart | indPush (COMMENT|varPart) (SAD (COMMENT|varPart))* DED)
module ::= ([COMMENT] [SAD] stmt)*
optComma ::= [ ',' ] [COMMENT] [IND]
operator ::= OP0 | OR | XOR | AND | OP3 | OP4 | OP5 | IS | ISNOT | IN | NOTIN
| OP6 | DIV | MOD | SHL | SHR | OP7 | NOT
prefixOperator ::= OP0 | OP3 | OP4 | OP5 | OP6 | OP7 | NOT
optInd ::= [COMMENT] [IND]
lowestExpr ::= orExpr ( OP0 optInd orExpr )*
orExpr ::= andExpr ( OR | XOR optInd andExpr )*
andExpr ::= cmpExpr ( AND optInd cmpExpr )*
cmpExpr ::= ampExpr ( OP3 | IS | ISNOT | IN | NOTIN optInd ampExpr )*
ampExpr ::= plusExpr ( OP4 optInd plusExpr )*
plusExpr ::= mulExpr ( OP5 optInd mulExpr )*
mulExpr ::= dollarExpr ( OP6 | DIV | MOD | SHL | SHR optInd dollarExpr )*
dollarExpr ::= primary ( OP7 optInd primary )*
namedTypeOrExpr ::=
DOTDOT [expr]
| expr [EQUALS (expr [DOTDOT expr] | typeDescK | DOTDOT [expr] )
| DOTDOT [expr]]
| typeDescK
castExpr ::= CAST BRACKET_LE optInd typeDesc BRACKERT_RI
PAR_LE optInd expr PAR_RI
addrExpr ::= ADDR PAR_LE optInd expr PAR_RI
symbol ::= ACC (KEYWORD | IDENT | operator | PAR_LE PAR_RI
| BRACKET_LE BRACKET_RI | EQUALS | literal )+ ACC
| IDENT
primary ::= ( prefixOperator optInd )* ( symbol | constructor |
| castExpr | addrExpr ) (
DOT optInd symbol
#| CURLY_LE namedTypeDescList CURLY_RI
| PAR_LE optInd
namedExprList
PAR_RI
| BRACKET_LE optInd
(namedTypeOrExpr optComma)*
BRACKET_RI
| CIRCUM
| pragma )*
literal ::= INT_LIT | INT8_LIT | INT16_LIT | INT32_LIT | INT64_LIT
| FLOAT_LIT | FLOAT32_LIT | FLOAT64_LIT
| STR_LIT | RSTR_LIT | TRIPLESTR_LIT
| CHAR_LIT | RCHAR_LIT
| NIL
constructor ::= literal
| BRACKET_LE optInd (expr [COLON expr] optComma )* BRACKET_RI # []-Constructor
| CURLY_LE optInd (expr [DOTDOT expr] optComma )* CURLY_RI # {}-Constructor
| PAR_LE optInd (expr [COLON expr] optComma )* PAR_RI # ()-Constructor
exprList ::= ( expr optComma )*
namedExpr ::= expr [EQUALS expr] # actually this is symbol EQUALS expr|expr
namedExprList ::= ( namedExpr optComma )*
exprOrSlice ::= expr [ DOTDOT expr ]
sliceList ::= ( exprOrSlice optComma )+
anonymousProc ::= LAMBDA paramList [pragma] EQUALS stmt
expr ::= lowestExpr
| anonymousProc
| IF expr COLON expr
(ELIF expr COLON expr)*
ELSE COLON expr
namedTypeDesc ::= typeDescK | expr [EQUALS (typeDescK | expr)]
namedTypeDescList ::= ( namedTypeDesc optComma )*
qualifiedIdent ::= symbol [ DOT symbol ]
typeDescK ::= VAR typeDesc
| REF typeDesc
| PTR typeDesc
| TYPE expr
| TUPLE tupleDesc
| PROC paramList [pragma]
typeDesc ::= typeDescK | primary
optSemicolon ::= [SEMICOLON]
macroStmt ::= COLON [stmt] (OF [sliceList] COLON stmt
| ELIF expr COLON stmt
| EXCEPT exceptList COLON stmt )*
[ELSE COLON stmt]
simpleStmt ::= returnStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
complexStmt ::= ifStmt | whileStmt | caseStmt | tryStmt | forStmt
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl
| constSection | typeSection | whenStmt | varSection
indPush ::= IND # push
stmt ::= simpleStmt [SAD]
| indPush (complexStmt | simpleStmt)
([SAD] (complexStmt | simpleStmt) )*
DED
exprStmt ::= lowestExpr [EQUALS expr | (expr optComma)* [macroStmt]]
returnStmt ::= RETURN [expr]
yieldStmt ::= YIELD expr
discardStmt ::= DISCARD expr
raiseStmt ::= RAISE [expr]
breakStmt ::= BREAK [symbol]
continueStmt ::= CONTINUE
ifStmt ::= IF expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
whenStmt ::= WHEN expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
caseStmt ::= CASE expr (OF sliceList COLON stmt)*
(ELIF expr COLON stmt)*
[ELSE COLON stmt]
whileStmt ::= WHILE expr COLON stmt
forStmt ::= FOR (symbol optComma)+ IN expr [DOTDOT expr] COLON stmt
exceptList ::= (qualifiedIdent optComma)*
tryStmt ::= TRY COLON stmt
(EXCEPT exceptList COLON stmt)*
[FINALLY COLON stmt]
asmStmt ::= ASM [pragma] (STR_LIT | RSTR_LIT | TRIPLESTR_LIT)
blockStmt ::= BLOCK [symbol] COLON stmt
importStmt ::= IMPORT ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) [AS symbol] optComma)+
includeStmt ::= INCLUDE ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) optComma)+
fromStmt ::= FROM (symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) IMPORT (symbol optComma)+
pragma ::= CURLYDOT_LE (expr [COLON expr] optComma)+ (CURLYDOT_RI | CURLY_RI)
paramList ::= [PAR_LE ((symbol optComma)+ COLON typeDesc optComma)* PAR_RI] [COLON typeDesc]
genericParams ::= BRACKET_LE (symbol [EQUALS typeDesc] )* BRACKET_RI
procDecl ::= PROC symbol ["*"] [genericParams]
paramList [pragma]
[EQUALS stmt]
macroDecl ::= MACRO symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
iteratorDecl ::= ITERATOR symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
templateDecl ::= TEMPLATE symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
colonAndEquals ::= [COLON typeDesc] EQUALS expr
constDecl ::= symbol ["*"] [pragma] colonAndEquals [COMMENT | IND COMMENT]
| COMMENT
constSection ::= CONST indPush constDecl (SAD constDecl)* DED
typeDef ::= typeDesc | objectDef | enumDef
objectIdentPart ::=
(symbol ["*" | "-"] [pragma] optComma)+ COLON typeDesc [COMMENT | IND COMMENT]
objectWhen ::= WHEN expr COLON [COMMENT] objectPart
(ELIF expr COLON [COMMENT] objectPart)*
[ELSE COLON [COMMENT] objectPart]
objectCase ::= CASE expr COLON typeDesc [COMMENT]
(OF sliceList COLON [COMMENT] objectPart)*
[ELSE COLON [COMMENT] objectPart]
objectPart ::= objectWhen | objectCase | objectIdentPart | NIL
| indPush objectPart (SAD objectPart)* DED
tupleDesc ::= BRACKET_LE optInd ((symbol optComma)+ COLON typeDesc optComma)* BRACKET_RI
objectDef ::= OBJECT [pragma] [OF typeDesc] objectPart
enumDef ::= ENUM [OF typeDesc] (symbol [EQUALS expr] optComma [COMMENT | IND COMMENT])+
typeDecl ::= COMMENT
| symbol ["*"] [genericParams] [EQUALS typeDef] [COMMENT | IND COMMENT]
typeSection ::= TYPE indPush typeDecl (SAD typeDecl)* DED
colonOrEquals ::= COLON typeDesc [EQUALS expr] | EQUALS expr
varPart ::= (symbol ["*" | "-"] [pragma] optComma)+ colonOrEquals [COMMENT | IND COMMENT]
varSection ::= VAR (varPart | indPush (COMMENT|varPart) (SAD (COMMENT|varPart))* DED)

View file

@ -32,7 +32,7 @@ Path Purpose
``config`` configuration files for Nimrod go into here
``lib`` the Nimrod library lives here; ``rod`` depends
on it!
``web`` website of Nimrod; generated by ``genweb.py``
``web`` website of Nimrod; generated by ``koch.py``
from the ``*.txt`` and ``*.tmpl`` files
``koch`` the Koch Build System (written for Nimrod)
``obj`` generated ``*.obj`` files go into here
@ -49,20 +49,19 @@ and Free Pascal can compile the Nimrod compiler.
Requirements for bootstrapping:
- Free Pascal (I used version 2.2); it may not be needed
- Python (should work with 2.4 or higher) and the code generator *cog*
(included in this distribution!)
- Free Pascal (I used version 2.2) [optional]
- Python (should work with version 1.5 or higher)
- C compiler -- one of:
* win32-lcc
* Borland C++ (tested with 5.5)
* win32-lcc (currently broken)
* Borland C++ (tested with 5.5; currently broken)
* Microsoft C++
* Digital Mars C++
* Watcom C++ (currently broken; a fix is welcome!)
* Watcom C++ (currently broken)
* GCC
* Intel C++
* Pelles C
* Pelles C (currently broken)
* llvm-gcc
| Compiling the compiler is a simple matter of running:
@ -75,7 +74,7 @@ If you want to debug the compiler, use the command::
The ``koch.py`` script is Nimrod's maintainance script: Everything that has
been automated is accessible with it. It is a replacement for make and shell
scripting with the advantage that it is more portable and is easier to read.
scripting with the advantage that it is more portable.
Coding standards
@ -124,7 +123,7 @@ Complex assignments
We already knew the type information as a graph in the compiler.
Thus we need to serialize this graph as RTTI for C code generation.
Look at the files ``lib/typeinfo.nim``, ``lib/hti.nim`` for more information.
Look at the file ``lib/hti.nim`` for more information.
The Garbage Collector
@ -135,18 +134,11 @@ Introduction
We use the term *cell* here to refer to everything that is traced
(sequences, refs, strings).
This section describes how the new GC works. The old algorithms
all had the same problem: Too complex to get them right. This one
tries to find the right compromise.
This section describes how the new GC works.
The basic algorithm is *Deferrent reference counting* with cycle detection.
References in the stack are not counted for better performance and easier C
code generation. The GC starts by traversing the hardware stack and increments
the reference count (RC) of every cell that it encounters. After the GC has
done its work the stack is traversed again and the RC of every cell
that it encounters are decremented again. Thus no marking bits in the RC are
needed. Between these stack traversals the GC has a complete accurate view over
the RCs.
code generation.
Each cell has a header consisting of a RC and a pointer to its type
descriptor. However the program does not know about these, so they are placed at
@ -156,72 +148,44 @@ is extremely important that ``pointer`` is not confused with a ``PCell``
as this would lead to a memory corruption.
When to trigger a collection
----------------------------
Since there are really two different garbage collectors (reference counting
and mark and sweep) we use two different heuristics when to run the passes.
The RC-GC pass is fairly cheap: Thus we use an additive increase (7 pages)
for the RC_Threshold and a multiple increase for the CycleThreshold.
The AT and ZCT sets
-------------------
The GC maintains two sets throughout the lifetime of
the program (plus two temporary ones). The AT (*any table*) simply contains
every cell. The ZCT (*zero count table*) contains every cell whose RC is
zero. This is used to reclaim most cells fast.
The ZCT contains redundant information -- the AT alone would suffice.
However, traversing the AT and look if the RC is zero would touch every living
cell in the heap! That's why the ZCT is updated whenever a RC drops to zero.
The ZCT is not updated when a RC is incremented from zero to one, as
this would be too costly.
The CellSet data structure
--------------------------
The AT and ZCT depend on an extremely efficient datastructure for storing a
set of pointers - this is called a ``PCellSet`` in the source code.
The GC depends on an extremely efficient datastructure for storing a
set of pointers - this is called a ``TCellSet`` in the source code.
Inserting, deleting and searching are done in constant time. However,
modifying a ``PCellSet`` during traversation leads to undefined behaviour.
modifying a ``TCellSet`` during traversation leads to undefined behaviour.
.. code-block:: Nimrod
type
PCellSet # hidden
TCellSet # hidden
proc allocCellSet: PCellSet # make a new set
proc deallocCellSet(s: PCellSet) # empty the set and free its memory
proc incl(s: PCellSet, elem: PCell) # include an element
proc excl(s: PCellSet, elem: PCell) # exclude an element
proc CellSetInit(s: var TCellSet) # initialize a new set
proc CellSetDeinit(s: var TCellSet) # empty the set and free its memory
proc incl(s: var TCellSet, elem: PCell) # include an element
proc excl(s: var TCellSet, elem: PCell) # exclude an element
proc `in`(elem: PCell, s: PCellSet): bool
proc `in`(elem: PCell, s: TCellSet): bool # tests membership
iterator elements(s: PCellSet): (elem: PCell)
iterator elements(s: TCellSet): (elem: PCell)
All the operations have to be performed efficiently. Because a Cellset can
become huge (the AT contains every allocated cell!) a hash table is not
suitable for this.
All the operations have to be perform efficiently. Because a Cellset can
become huge a hash table alone is not suitable for this.
We use a mixture of bitset and patricia tree for this. One node in the
patricia tree contains a bitset that decribes a page of the operating system
(not always, but that doesn't matter).
So including a cell is done as follows:
We use a mixture of bitset and hash table for this. The hash table maps *pages*
to a page descriptor. The page descriptor contains a bit for any possible cell
address within this page. So including a cell is done as follows:
- Find the page descriptor for the page the cell belongs to.
- Set the appropriate bit in the page descriptor indicating that the
cell points to the start of a memory block.
Removing a cell is analogous - the bit has to be set to zero.
Single page descriptors are never deleted from the tree. Typically a page
descriptor is only 19 words big, so it does not waste much by not deleting
it. Apart from that the AT and ZCT are rebuilt frequently, so removing a
single page descriptor from the tree is never necessary.
Single page descriptors are never deleted from the hash table. This is not
needed as the data structures need to be periodically rebuilt anyway.
Complete traversal is done like so::
Complete traversal is done in this way::
for each page decriptor d:
for each bit in d:
@ -265,103 +229,12 @@ roughly like this:
Note that for systems with a continous stack (which most systems have)
the check whether the ref is on the stack is very cheap (only two
comparisons). Another advantage of this scheme is that the code produced is
a tiny bit smaller.
smaller.
The algorithm in pseudo-code
----------------------------
Now we come to the nitty-gritty. The algorithm works in several phases.
Phase 1 - Consider references from stack
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
::
for each pointer p in the stack: incRef(p)
This is necessary because references in the hardware stack are not traced for
better performance. After Phase 1 the RCs are accurate.
Phase 2 - Free the ZCT
~~~~~~~~~~~~~~~~~~~~~~
This is how things used to (not) work::
for p in elements(ZCT):
if RC(p) == 0:
call finalizer of p
for c in children(p): decRef(c) # free its children recursively
# if necessary; the childrens RC >= 1, BUT they may still be in the ZCT!
free(p)
else:
remove p from the ZCT
Instead we do it this way. Note that the recursion is gone too!
::
newZCT = nil
for p in elements(ZCT):
if RC(p) == 0:
call finalizer of p
for c in children(p):
assert(RC(c) > 0)
dec(RC(c))
if RC(c) == 0:
if newZCT == nil: newZCT = allocCellSet()
incl(newZCT, c)
free(p)
else:
# nothing to do! We will use the newZCS
deallocCellSet(ZCT)
ZCT = newZCT
This phase is repeated until enough memory is available or the ZCT is nil.
If still not enough memory is available the cyclic detector gets its chance
to do something.
Phase 3 - Cycle detection
~~~~~~~~~~~~~~~~~~~~~~~~~
Cycle detection works by subtracting internal reference counts::
newAT = allocCellSet()
for y in elements(AT):
# pretend that y is dead:
for c in children(y):
dec(RC(c))
# note that this should not be done recursively as we have all needed
# pointers in the AT! This makes it more efficient too!
proc restore(y: PCell) =
# unfortunately, the recursion here cannot be eliminated easily
if y not_in newAT:
incl(newAT, y)
for c in children(y):
inc(RC(c)) # restore proper reference counts!
restore(c)
for y in elements(AT) with rc > 0:
restore(y)
for y in elements(AT) with rc == 0:
free(y) # if pretending worked, it was part of a cycle
AT = newAT
Phase 4 - Ignore references from stack again
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
::
for each pointer p in the stack:
dec(RC(p))
if RC(p) == 0: incl(ZCT, p)
Now the RCs correctly discard any references from the stack. One can also see
this as a temporary marking operation. Things that are referenced from stack
are marked during the GC's operarion and now have to be unmarked.
To be written.
The compiler's architecture

View file

@ -1,131 +1,216 @@
=======================
Nimrod Standard Library
=======================
:Author: Andreas Rumpf
:Version: |nimrodversion|
Though the Nimrod Standard Library is still evolving, it is already quite
usable. It is divided into basic libraries that contains modules that virtually
every program will need and advanced libraries which are more heavy weight.
Advanced libraries are in the ``lib/base`` directory.
Basic libraries
===============
* `system <system.html>`_
Basic procs and operators that every program needs. It also provides IO
facilities for reading and writing text and binary files. It is imported
implicitly by the compiler. Do not import it directly. It relies on compiler
magic to work.
* `strutils <strutils.html>`_
This module contains common string handling operations like converting a
string into uppercase, splitting a string into substrings, searching for
substrings, replacing substrings.
* `os <os.html>`_
Basic operating system facilities like retrieving environment variables,
reading command line arguments, working with directories, running shell
commands, etc. This module is -- like any other basic library --
platform independant.
* `math <math.html>`_
Mathematical operations like cosine, square root.
* `complex <complex.html>`_
This module implements complex numbers and their mathematical operations.
* `times <times.html>`_
The ``times`` module contains basic support for working with time.
* `parseopt <parseopt.html>`_
The ``parseopt`` module implements a command line option parser. This
supports long and short command options with optional values and command line
arguments.
* `parsecfg <parsecfg.html>`_
The ``parsecfg`` module implements a high performance configuration file
parser. The configuration file's syntax is similar to the Windows ``.ini``
format, but much more powerful, as it is not a line based parser. String
literals, raw string literals and triple quote string literals are supported
as in the Nimrod programming language.
* `strtabs <strtabs.html>`_
The ``strtabs`` module implements an efficient hash table that is a mapping
from strings to strings. Supports a case-sensitive, case-insensitive and
style-insensitive mode. An efficient string substitution operator ``%``
for the string table is also provided.
* `hashes <hashes.html>`_
This module implements efficient computations of hash values for diverse
Nimrod types.
* `lexbase <lexbase.html>`_
This is a low leve module that implements an extremely efficent buffering
scheme for lexers and parsers. This is used by the ``parsecfg`` module.
Advanced libaries
=================
* `regexprs <regexprs.html>`_
This module contains procedures and operators for handling regular
expressions.
* `dialogs <dialogs.html>`_
This module implements portable dialogs for Nimrod; the implementation
builds on the GTK interface. On Windows, native dialogs are shown if
appropriate.
Wrappers
========
Note that the generated HTML for some of these wrappers is so huge, that it is
not contained in the distribution. You can then find them on the website.
* `posix <posix.html>`_
Contains a wrapper for the POSIX standard.
* `windows <windows.html>`_
Contains a wrapper for the Win32 API.
* `shellapi <shellapi.html>`_
Contains a wrapper for the ``shellapi.h`` header.
* `shfolder <shfolder.html>`_
Contains a wrapper for the ``shfolder.h`` header.
* `mmsystem <mmsystem.html>`_
Contains a wrapper for the ``mmsystem.h`` header.
* `ole2 <ole2.html>`_
Contains GUIDs for OLE2 automation support.
* `nb30 <nb30.html>`_
This module contains the definitions for portable NetBIOS 3.0 support.
* `cairo <cairo.html>`_
Wrapper for the cairo library.
* `cairoft <cairoft.html>`_
Wrapper for the cairoft library.
* `cairowin32 <cairowin32.html>`_
Wrapper for the cairowin32 library.
* `cairoxlib <cairoxlib.html>`_
Wrapper for the cairoxlib library.
* `atk <atk.html>`_
Wrapper for the atk library.
* `gdk2 <gdk2.html>`_
Wrapper for the gdk2 library.
* `gdk2pixbuf <gdk2pixbuf.html>`_
Wrapper for the gdk2pixbuf library.
* `gdkglext <gdkglext.html>`_
Wrapper for the gdkglext library.
* `glib2 <glib2.html>`_
Wrapper for the glib2 library.
* `gtk2 <gtk2.html>`_
Wrapper for the gtk2 library.
* `gtkglext <gtkglext.html>`_
Wrapper for the gtkglext library.
* `gtkhtml <gtkhtml.html>`_
Wrapper for the gtkhtml library.
* `libglade2 <libglade2.html>`_
Wrapper for the libglade2 library.
* `pango <pango.html>`_
Wrapper for the pango library.
* `pangoutils <pangoutils.html>`_
Wrapper for the pangoutils library.
=======================
Nimrod Standard Library
=======================
:Author: Andreas Rumpf
:Version: |nimrodversion|
Though the Nimrod Standard Library is still evolving, it is already quite
usable. It is divided into basic libraries that contains modules that virtually
every program will need and advanced libraries which are more heavy weight.
Advanced libraries are in the ``lib/base`` directory.
Basic libraries
===============
* `system <system.html>`_
Basic procs and operators that every program needs. It also provides IO
facilities for reading and writing text and binary files. It is imported
implicitly by the compiler. Do not import it directly. It relies on compiler
magic to work.
* `strutils <strutils.html>`_
This module contains common string handling operations like converting a
string into uppercase, splitting a string into substrings, searching for
substrings, replacing substrings.
* `os <os.html>`_
Basic operating system facilities like retrieving environment variables,
reading command line arguments, working with directories, running shell
commands, etc. This module is -- like any other basic library --
platform independant.
* `math <math.html>`_
Mathematical operations like cosine, square root.
* `complex <complex.html>`_
This module implements complex numbers and their mathematical operations.
* `times <times.html>`_
The ``times`` module contains basic support for working with time.
* `parseopt <parseopt.html>`_
The ``parseopt`` module implements a command line option parser. This
supports long and short command options with optional values and command line
arguments.
* `parsecfg <parsecfg.html>`_
The ``parsecfg`` module implements a high performance configuration file
parser. The configuration file's syntax is similar to the Windows ``.ini``
format, but much more powerful, as it is not a line based parser. String
literals, raw string literals and triple quote string literals are supported
as in the Nimrod programming language.
* `strtabs <strtabs.html>`_
The ``strtabs`` module implements an efficient hash table that is a mapping
from strings to strings. Supports a case-sensitive, case-insensitive and
style-insensitive mode. An efficient string substitution operator ``%``
for the string table is also provided.
* `streams <streams.html>`_
This module provides a stream interface and two implementations thereof:
the `PFileStream` and the `PStringStream` which implement the stream
interface for Nimrod file objects (`TFile`) and strings. Other modules
may provide other implementations for this standard stream interface.
* `hashes <hashes.html>`_
This module implements efficient computations of hash values for diverse
Nimrod types.
* `lexbase <lexbase.html>`_
This is a low level module that implements an extremely efficent buffering
scheme for lexers and parsers. This is used by the ``parsecfg`` module.
Advanced libaries
=================
* `regexprs <regexprs.html>`_
This module contains procedures and operators for handling regular
expressions.
* `dialogs <dialogs.html>`_
This module implements portable dialogs for Nimrod; the implementation
builds on the GTK interface. On Windows, native dialogs are shown if
appropriate.
* `zipfiles <zipfiles.html>`_
This module implements a zip archive creator/reader/modifier.
Wrappers
========
Note that the generated HTML for some of these wrappers is so huge, that it is
not contained in the distribution. You can then find them on the website.
* `posix <posix.html>`_
Contains a wrapper for the POSIX standard.
* `windows <windows.html>`_
Contains a wrapper for the Win32 API.
* `shellapi <shellapi.html>`_
Contains a wrapper for the ``shellapi.h`` header.
* `shfolder <shfolder.html>`_
Contains a wrapper for the ``shfolder.h`` header.
* `mmsystem <mmsystem.html>`_
Contains a wrapper for the ``mmsystem.h`` header.
* `ole2 <ole2.html>`_
Contains GUIDs for OLE2 automation support.
* `nb30 <nb30.html>`_
This module contains the definitions for portable NetBIOS 3.0 support.
* `cairo <cairo.html>`_
Wrapper for the cairo library.
* `cairoft <cairoft.html>`_
Wrapper for the cairoft library.
* `cairowin32 <cairowin32.html>`_
Wrapper for the cairowin32 library.
* `cairoxlib <cairoxlib.html>`_
Wrapper for the cairoxlib library.
* `atk <atk.html>`_
Wrapper for the atk library.
* `gdk2 <gdk2.html>`_
Wrapper for the gdk2 library.
* `gdk2pixbuf <gdk2pixbuf.html>`_
Wrapper for the gdk2pixbuf library.
* `gdkglext <gdkglext.html>`_
Wrapper for the gdkglext library.
* `glib2 <glib2.html>`_
Wrapper for the glib2 library.
* `gtk2 <gtk2.html>`_
Wrapper for the gtk2 library.
* `gtkglext <gtkglext.html>`_
Wrapper for the gtkglext library.
* `gtkhtml <gtkhtml.html>`_
Wrapper for the gtkhtml library.
* `libglade2 <libglade2.html>`_
Wrapper for the libglade2 library.
* `pango <pango.html>`_
Wrapper for the pango library.
* `pangoutils <pangoutils.html>`_
Wrapper for the pangoutils library.
* `gl <gl.html>`_
Part of the wrapper for OpenGL.
* `glext <glext.html>`_
Part of the wrapper for OpenGL.
* `glu <glu.html>`_
Part of the wrapper for OpenGL.
* `glut <glut.html>`_
Part of the wrapper for OpenGL.
* `glx <glx.html>`_
Part of the wrapper for OpenGL.
* `wingl <wingl.html>`_
Part of the wrapper for OpenGL.
* `lua <lua.html>`_
Part of the wrapper for Lua.
* `lualib <lualib.html>`_
Part of the wrapper for Lua.
* `lauxlib <lauxlib.html>`_
Part of the wrapper for Lua.
* `odbcsql <odbcsql.nim>`_
interface to the ODBC driver.
* `zlib <zlib.nim>`_
Wrapper for the zlib library.
* `sdl <sdl.html>`_
Part of the wrapper for SDL.
* `sdl_gfx <sdl_gfx.html>`_
Part of the wrapper for SDL.
* `sdl_image <sdl_image.html>`_
Part of the wrapper for SDL.
* `sdl_mixer <sdl_mixer.html>`_
Part of the wrapper for SDL.
* `sdl_net <sdl_net.html>`_
Part of the wrapper for SDL.
* `sdl_ttf <sdl_ttf.html>`_
Part of the wrapper for SDL.
* `smpeg <smpeg.html>`_
Part of the wrapper for SDL.
* `cursorfont <cursorfont.html>`_
Part of the wrapper for X11.
* `keysym <keysym.html>`_
Part of the wrapper for X11.
* `x <x.html>`_
Part of the wrapper for X11.
* `xatom <xatom.html>`_
Part of the wrapper for X11.
* `xcms <xcms.html>`_
Part of the wrapper for X11.
* `xf86dga <xf86dga.html>`_
Part of the wrapper for X11.
* `xf86vmode <xf86vmode.html>`_
Part of the wrapper for X11.
* `xi <xi.html>`_
Part of the wrapper for X11.
* `xinerama <xinerama.html>`_
Part of the wrapper for X11.
* `xkb <xkb.html>`_
Part of the wrapper for X11.
* `xkblib <xkblib.html>`_
Part of the wrapper for X11.
* `xlib <xlib.html>`_
Part of the wrapper for X11.
* `xrandr <xrandr.html>`_
Part of the wrapper for X11.
* `xrender <xrender.html>`_
Part of the wrapper for X11.
* `xresource <xresource.html>`_
Part of the wrapper for X11.
* `xshm <xshm.html>`_
Part of the wrapper for X11.
* `xutil <xutil.html>`_
Part of the wrapper for X11.
* `xv <xv.html>`_
Part of the wrapper for X11.
* `xvlib <xvlib.html>`_
Part of the wrapper for X11.
* `libzip <libzip.html>`_
Interface to the `lib zip <http://www.nih.at/libzip/index.html>`_ library by
Dieter Baron and Thomas Klausner.

View file

@ -160,8 +160,8 @@ case-sensitive and even underscores are ignored:
this is that this allows programmers to use their own prefered spelling style
and libraries written by different programmers cannot use incompatible
conventions. The editors or IDE can show the identifiers as preferred. Another
advantage is that it frees the programmer from remembering the spelling of an
identifier.
advantage is that it frees the programmer from remembering the exact spelling
of an identifier.
Literal strings
@ -601,20 +601,20 @@ Array and sequence types
has the same type. Arrays always have a fixed length which is specified at
compile time (except for open arrays). They can be indexed by any ordinal type.
A parameter ``A`` may be an *open array*, in which case it is indexed by
integers from 0 to ``len(A)-1``.
integers from 0 to ``len(A)-1``. An array expression may be constructed by the
array constructor ``[]``.
`Sequences`:idx: are similar to arrays but of dynamic length which may change
during runtime (like strings). A sequence ``S`` is always indexed by integers
from 0 to ``len(S)-1`` and its bounds are checked. Sequences can also be
constructed by the array constructor ``[]``.
from 0 to ``len(S)-1`` and its bounds are checked. Sequences can be
constructed by the array constructor ``[]`` in conjunction with the array to
sequence operator ``@``. Another way to allocate space for a sequence is to
call the built-in ``newSeq`` procedure.
A sequence may be passed to a parameter that is of type *open array*, but
not to a multi-dimensional open array, because it is impossible to do so in an
efficient manner.
An array expression may be constructed by the array constructor ``[]``.
A constructed array is assignment compatible to a sequence.
Example:
.. code-block:: nimrod
@ -625,13 +625,13 @@ Example:
var
x: TIntArray
y: TIntSeq
x = [1, 2, 3, 4, 5, 6] # [] this is the array constructor that is compatible
# with arrays, open arrays and
y = [1, 2, 3, 4, 5, 6] # sequences
x = [1, 2, 3, 4, 5, 6] # [] this is the array constructor
y = @[1, 2, 3, 4, 5, 6] # the @ turns the array into a sequence
The lower bound of an array may be received by the built-in proc
The lower bound of an array or sequence may be received by the built-in proc
``low()``, the higher bound by ``high()``. The length may be
received by ``len()``.
received by ``len()``. ``low()`` for a sequence or an open array always returns
0, as this is the first valid index.
Arrays are always bounds checked (at compile-time or at runtime). These
checks can be disabled via pragmas or invoking the compiler with the
@ -644,15 +644,15 @@ A variable of a `tuple`:idx: or `object`:idx: type is a heterogenous storage
container.
A tuple or object defines various named *fields* of a type. A tuple defines an
*order* of the fields additionally. Tuples are meant for heterogenous storage
types with no overhead and few abstraction possibilities. The constructor ``()``
can be used to construct tuples. The order of the fields in the constructor
must match the order of the tuple's definition. Different tuple-types are
*equivalent* if they specify the same fields of the same type in the same
order.
types with no overhead and few abstraction possibilities. The constructor ``()``
can be used to construct tuples. The order of the fields in the constructor
must match the order of the tuple's definition. Different tuple-types are
*equivalent* if they specify the same fields of the same type in the same
order.
The assignment operator for tuples copies each component.
The default assignment operator for objects is not defined. The programmer may
provide one, however.
The assignment operator for tuples copies each component.
The default assignment operator for objects is not defined. The programmer may
provide one, however.
.. code-block:: nimrod
@ -662,7 +662,7 @@ provide one, however.
# and an age
var
person: TPerson
person = (name: "Peter", age: 30)
person = (name: "Peter", age: 30)
# the same, but less readable:
person = ("Peter", 30)
@ -670,8 +670,8 @@ The implementation aligns the fields for best access performance. The alignment
is done in a way that is compatible the way the C compiler does it.
Objects provide many features that tuples do not. Object provide inheritance
and information hiding. Objects have access to their type at runtime, so that
the ``is`` operator can be used to determine the object's type.
and information hiding. Objects have access to their type at runtime, so that
the ``is`` operator can be used to determine the object's type.
.. code-block:: nimrod
@ -689,9 +689,51 @@ the ``is`` operator can be used to determine the object's type.
assert(student is TStudent) # is true
Object fields that should be visible outside from the defining module, have to
marked by ``*``. In contrast to tuples, different object types are
marked by ``*``. In contrast to tuples, different object types are
never *equivalent*.
Object variants
~~~~~~~~~~~~~~~
Often an object hierarchy is overkill in certain situations where simple
`variant`:idx: types are needed.
An example:
.. code-block:: nimrod
# This is an example how an abstract syntax tree could be modelled in Nimrod
type
TNodeKind = enum # the different node types
nkInt, # a leaf with an integer value
nkFloat, # a leaf with a float value
nkString, # a leaf with a string value
nkAdd, # an addition
nkSub, # a subtraction
nkIf # an if statement
PNode = ref TNode
TNode = object
case kind: TNodeKind # the ``kind`` field is the discriminator
of nkInt: intVal: int
of nkFloat: floavVal: float
of nkString: strVal: string
of nkAdd, nkSub:
leftOp, rightOp: PNode
of nkIf:
condition, thenPart, elsePart: PNode
var
n: PNode
new(n) # creates a new node
n.kind = nkFloat
n.floatVal = 0.0 # valid, because ``n.kind==nkFloat``, so that it fits
# the following statement raises an `EInvalidField` exception, because
# n.kind's value does not fit:
n.strVal = ""
As can been seen from the example, an advantage to an object hierarchy is that
no casting between different object types is needed. Yet, access to invalid
object fields raises an exception.
Set type
~~~~~~~~
@ -749,7 +791,7 @@ The ``^`` operator can be used to derefer a reference, the ``addr`` procedure
returns the address of an item. An address is always an untraced reference.
Thus the usage of ``addr`` is an *unsafe* feature.
The ``.`` (access a tuple/object field operator)
The ``.`` (access a tuple/object field operator)
and ``[]`` (array/string/sequence index operator) operators perform implicit
dereferencing operations for reference types:
@ -773,7 +815,7 @@ further information.
Special care has to be taken if an untraced object contains traced objects like
traced references, strings or sequences: In order to free everything properly,
the built-in procedure ``finalize`` has to be called before freeing the
the built-in procedure ``GCunref`` has to be called before freeing the
untraced memory manually!
.. XXX finalizers for traced objects
@ -867,7 +909,7 @@ statement.
Statements are separated into `simple statements`:idx: and
`complex statements`:idx:.
Simple statements are statements that cannot contain other statements, like
Simple statements are statements that cannot contain other statements like
assignments, calls or the ``return`` statement; complex statements can
contain other statements. To avoid the `dangling else problem`:idx:, complex
statements always have to be intended::
@ -1028,10 +1070,10 @@ Example:
The `case`:idx: statement is similar to the if statement, but it represents
a multi-branch selection. The expression after the keyword ``case`` is
evaluated and if its value is in a *vallist* the corresponding statements
(after the ``of`` keyword) are executed. If the value is no given *vallist*
the ``else`` part is executed. If there is no ``else`` part and not all
possible values that ``expr`` can hold occur in a ``vallist``, a static
error is given. This holds only for expressions of ordinal types.
(after the ``of`` keyword) are executed. If the value is not in any
given *slicelist* the ``else`` part is executed. If there is no ``else``
part and not all possible values that ``expr`` can hold occur in a ``vallist``,
a static error is given. This holds only for expressions of ordinal types.
If the expression is not of an ordinal type, and no ``else`` part is
given, control just passes after the ``case`` statement.
@ -1331,10 +1373,8 @@ is used if the caller does not provide a value for this parameter. Example:
`Operators`:idx: are procedures with a special operator symbol as identifier:
.. code-block:: nimrod
proc `$` (x: int): string = # converts an integer to a string;
# since it has one parameter this is a prefix
# operator. With two parameters it would be
# an infix operator.
proc `$` (x: int): string =
# converts an integer to a string; this is a prefix operator.
return intToStr(x)
Calling a procedure can be done in many different ways:
@ -1544,7 +1584,7 @@ own file. Modules enable `information hiding`:idx: and
`separate compilation`:idx:. A module may gain access to symbols of another
module by the `import`:idx: statement. `Recursive module dependancies`:idx: are
allowed, but slightly subtle. Only top-level symbols that are marked with an
asterisk (``*``) are exported.
asterisk (``*``) are exported.
The algorithm for compiling modules is:
@ -1557,8 +1597,8 @@ This is best illustrated by an example:
.. code-block:: nimrod
# Module A
type
T1* = int
import B # the compiler starts parsing B
T1* = int # Module A exports the type ``T1``
import B # the compiler starts parsing B
proc main() =
var i = p(3) # works because B has been parsed completely here
@ -1660,12 +1700,19 @@ Nimrod source code. The conditional symbols go into a special symbol table.
The compiler defines the target processor and the target operating
system as conditional symbols.
Warning: The ``define`` pragma is deprecated as it conflicts with separate
compilation! One should use boolean constants as a replacement - this is
cleaner anyway.
undef pragma
------------
The `undef`:idx: pragma the counterpart to the define pragma. It undefines a
conditional symbol.
Warning: The ``undef`` pragma is deprecated as it conflicts with separate
compilation!
error pragma
------------

View file

@ -34,8 +34,15 @@ Advanced command line switches are:
Configuration file
------------------
The ``nimrod`` executable loads the configuration file ``config/nimrod.cfg``
unless this is suppressed by the ``--skip_cfg`` command line option.
The default configuration file is ``nimrod.cfg``. The ``nimrod`` executable
looks for it in the following directories (in this order):
1. ``/home/$user/.config/nimrod.cfg`` (UNIX) or ``$APPDATA/nimrod.cfg`` (Windows)
2. ``$nimrod/config/nimrod.cfg`` (UNIX, Windows)
3. ``/etc/nimrod.cfg`` (UNIX)
The search stops as soon as a configuration file has been found. The reading
of ``nimrod.cfg`` can be suppressed by the ``--skip_cfg`` command line option.
Configuration settings can be overwritten in a project specific
configuration file that is read automatically. This specific file has to
be in the same directory as the project and be of the same name, except
@ -47,11 +54,10 @@ Command line settings have priority over configuration file settings.
Nimrod's directory structure
----------------------------
The generated files that Nimrod produces all go into a subdirectory called
``rod_gen``. This makes it easy to write a script that deletes all generated
files. For example the generated C code for the module ``path/modA.nim``
will become ``path/rod_gen/modA.c``.
``nimcache`` in your project directory. This makes it easy to delete all
generated files.
However, the generated C code is not platform independant! C code generated for
However, the generated C code is not platform independant. C code generated for
Linux does not compile on Windows, for instance. The comment on top of the
C file lists the OS, CPU and CC the file has been compiled for.
@ -77,10 +83,46 @@ Because Nimrod generates C code it needs some "red tape" to work properly.
Thus lots of options and pragmas for tweaking the generated C code are
available.
Importc Pragma
~~~~~~~~~~~~~~
The `importc`:idx: pragma provides a means to import a type, a variable, or a
procedure from C. The optional argument is a string containing the C
identifier. If the argument is missing, the C name is the Nimrod
identifier *exactly as spelled*:
.. code-block::
proc printf(formatstr: cstring) {.importc: "printf", varargs.}
Exportc Pragma
~~~~~~~~~~~~~~
The `exportc`:idx: pragma provides a means to export a type, a variable, or a
procedure to C. The optional argument is a string containing the C
identifier. If the argument is missing, the C name is the Nimrod
identifier *exactly as spelled*:
.. code-block:: Nimrod
proc callme(formatstr: cstring) {.exportc: "callMe", varargs.}
Dynlib Pragma
~~~~~~~~~~~~~
With the `dynlib`:idx: pragma a procedure or a variable can be imported from
a dynamic library (``.dll`` files for Windows, ``lib*.so`` files for UNIX). The
non-optional argument has to be the name of the dynamic library:
.. code-block:: Nimrod
proc gtk_image_new(): PGtkWidget {.cdecl, dynlib: "libgtk-x11-2.0.so", importc.}
In general, importing a dynamic library does not require any special linker
options or linking with import libraries. This also
implies that no *devel* packages need to be installed.
No_decl Pragma
~~~~~~~~~~~~~~
The `no_decl`:idx: pragma can be applied to almost any symbol (variable, proc,
type, etc.) and is one of the most important for interoperability with C:
type, etc.) and is sometimes useful for interoperability with C:
It tells Nimrod that it should not generate a declaration for the symbol in
the C code. Thus it makes the following possible, for example:
@ -89,17 +131,7 @@ the C code. Thus it makes the following possible, for example:
EOF {.importc: "EOF", no_decl.}: cint # pretend EOF was a variable, as
# Nimrod does not know its value
Varargs Pragma
~~~~~~~~~~~~~~
The `varargs`:idx: pragma can be applied to procedures only. It tells Nimrod
that the proc can take a variable number of parameters after the last
specified parameter. Nimrod string values will be converted to C
strings automatically:
.. code-block:: Nimrod
proc printf(formatstr: cstring) {.nodecl, varargs.}
printf("hallo %s", "world") # "world" will be passed as C string
However, the ``header`` pragma is often the better alternative.
Header Pragma
@ -119,12 +151,25 @@ in angle brackets: ``<>``. If no angle brackets are given, Nimrod
encloses the header file in ``""`` in the generated C code.
Varargs Pragma
~~~~~~~~~~~~~~
The `varargs`:idx: pragma can be applied to procedures only. It tells Nimrod
that the proc can take a variable number of parameters after the last
specified parameter. Nimrod string values will be converted to C
strings automatically:
.. code-block:: Nimrod
proc printf(formatstr: cstring) {.nodecl, varargs.}
printf("hallo %s", "world") # "world" will be passed as C string
No_static Pragma
~~~~~~~~~~~~~~~~
The `no_static`:idx: pragma can be applied to almost any symbol and specifies
that it shall not be declared ``static`` in the generated C code. Note that
symbols in the interface part of a module never get declared ``static``, so
only in special cases is this pragma necessary.
only in very special cases this pragma is necessary.
Line_dir Option
@ -171,8 +216,29 @@ The `register`:idx: pragma is for variables only. It declares the variable as
in a hardware register for faster access. C compilers usually ignore this
though and for good reason: Often they do a better job without it anyway.
In highly specific cases (a dispatch loop of interpreters for example) it
may provide benefits, though.
In highly specific cases (a dispatch loop of an bytecode interpreter for
example) it may provide benefits, though.
Acyclic Pragma
~~~~~~~~~~~~~~
The `acyclic`:idx: pragma can be used for object types to mark them as acyclic
even though they seem to be cyclic. This is an **optimization** for the garbage
collector to not consider objects of this type as part of a cycle::
type
PNode = ref TNode
TNode {.acyclic, final.} = object
left, right: PNode
data: string
In the example a tree structure is declared with the ``TNode`` type. Note that
the type definition is recursive thus the GC has to assume that objects of
this type may form a cyclic graph. The ``acyclic`` pragma passes the
information that this cannot happen to the GC. If the programmer uses the
``acyclic`` pragma for data types that are in reality cyclic, the GC may leak
memory, but nothing worse happens.
Disabling certain messages
@ -244,12 +310,17 @@ efficient than any hand-coded scheme.
The ECMAScript code generator
=============================
Note: As of version 0.7.0 the ECMAScript code generator is not maintained any
longer. Help if you are interested.
Note: I use the term `ECMAScript`:idx: here instead of `JavaScript`:idx:, since
it is the proper term.
The ECMAScript code generator is experimental!
Nimrod targets ECMAScript 1.5 which is supported by any widely used browser.
Since ECMAScript does not have a portable means to include another module,
Nimrod just generate a long ``.js`` file.
Nimrod just generates a long ``.js`` file.
Features or modules that the ECMAScript platform does not support are not
available. This includes:

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff