Initial import

This commit is contained in:
Andreas Rumpf 2008-06-22 16:14:11 +02:00
commit 405b86068e
324 changed files with 163599 additions and 0 deletions

26
doc/docs.txt Executable file
View file

@ -0,0 +1,26 @@
"Incorrect documentation is often worse than no documentation."
-- Bertrand Meyer
The documentation consists of several documents:
- | `Nimrod manual <manual.html>`_
| Read this to get to know the Nimrod programming system.
- | `User guide for the Nimrod Compiler <nimrodc.html>`_
| The user guide lists command line arguments, Nimrodc's special features, etc.
- | `User guide for the Embedded Nimrod Debugger <endb.html>`_
| This document describes how to use the Embedded debugger. The embedded
debugger currently has no GUI. Please help!
- | `Nimrod library documentation <lib.html>`_
| This document describes Nimrod's standard library.
- | `Nimrod internal documentation <intern.html>`_
| The internal documentation describes how the compiler is implemented. Read
this if you want to hack the compiler or develop advanced macros.
- | `Index <theindex.html>`_
| The generated index. Often the quickest way to find the piece of
information you need.

174
doc/endb.txt Executable file
View file

@ -0,0 +1,174 @@
===========================================
Embedded Nimrod Debugger User Guide
===========================================
:Author: Andreas Rumpf
:Version: |nimrodversion|
.. contents::
Nimrod comes with a platform independant debugger -
the `Embedded Nimrod Debugger`:idx: (`ENDB`:idx:). The debugger is
*embedded* into your executable if it has been
compiled with the ``--debugger:on`` command line option.
This also defines the conditional symbol ``ENDB`` for you.
Note: You must not compile your program with the ``--app:gui``
command line option because then there is no console
available for the debugger.
If you start your program the debugger will immediately show
a prompt on the console. You can now enter a command. The next sections
deal with the possible commands. As usual for Nimrod for all commands
underscores and case do not matter. Optional components of a command
are listed in brackets ``[...]`` here.
General Commands
================
``h``, ``help``
Display a quick reference of the possible commands.
``q``, ``quit``
Quit the debugger and the program.
<ENTER>
(Without any typed command) repeat the previous debugger command.
If there is no previous command, ``step_into`` is assumed.
Executing Commands
==================
``s``, ``step_into``
Single step, stepping into routine calls.
``n``, ``step_over``
Single step, without stepping into routine calls.
``f``, ``skip_current``
Continue execution until the current routine finishes.
``c``, ``continue``
Continue execution until the next breakpoint.
``i``, ``ignore``
Continue execution, ignore all breakpoints. This is effectively quitting
the debugger and runs the program until it finishes.
Breakpoint Commands
===================
``b``, ``setbreak`` <identifier> [fromline [toline]] [file]
Set a new breakpoint named 'identifier' for the given file
and line numbers. If no file is given, the current execution point's
filename is used. If the filename has no extension, ``.nim`` is
appended for your convenience.
If no line numbers are given, the current execution point's
line is used. If both ``fromline`` and ``toline`` are given the
breakpoint contains a line number range. Some examples if it is still
unclear:
* ``b br1 12 15 thallo`` creates a breakpoint named ``br1`` that
will be triggered if the instruction pointer reaches one of the
lines 12-15 in the file ``thallo.nim``.
* ``b br1 12 thallo`` creates a breakpoint named ``br1`` that
will be triggered if the instruction pointer reaches the
line 12 in the file ``thallo.nim``.
* ``b br1 12`` creates a breakpoint named ``br1`` that
will be triggered if the instruction pointer reaches the
line 12 in the current file.
* ``b br1`` creates a breakpoint named ``br1`` that
will be triggered if the instruction pointer reaches the
current line in the current file again.
``breakpoints``
Display the entire breakpoint list.
``disable`` <identifier>
Disable a breakpoint. It remains disabled until you turn it on again
with the ``enable`` command.
``enable`` <identifier>
Enable a breakpoint.
Often it happens when debugging that you keep retyping the breakpoints again
and again because they are lost when you restart your program. This is not
necessary: A special pragma has been defined for this:
The ``{.breakpoint.}`` pragma
-----------------------------
The `breakpoint`:idx: pragma is syntactically a statement. It can be used
to mark the *following line* as a breakpoint:
.. code-block:: Nimrod
write("1")
{.breakpoint: "before_write_2".}
write("2")
The name of the breakpoint here is ``before_write_2``. Of course the
breakpoint's name is optional - the compiler will generate one for you
if you leave it out.
Code for the ``breakpoint`` pragma is only generated if the debugger
is turned on, so you don't need to remove it from your source code after
debugging.
Data Display Commands
=====================
``e``, ``eval`` <exp>
Evaluate the expression <exp>. Note that ENDB has no full-blown expression
evaluator built-in. So expressions are limited:
* To display global variables prefix their names with their
owning module: ``nim1.globalVar``
* To display local variables or parameters just type in
their name: ``localVar``. If you want to inspect variables that are not
in the current stack frame, use the ``up`` or ``down`` command.
Unfortunately, only inspecting variables is possible at the moment. Maybe
a future version will implement a full-blown Nimrod expression evaluator,
but this is not easy to do and would bloat the debugger's code.
Since displaying the whole data structures is often not needed and
painfully slow, the debugger uses a *maximal display depth* concept for
displaying.
You can alter the *maximal display depth* with the ``maxdisplay``
command.
``maxdisplay`` <natural>
Sets the maximal display depth to the given integer value. A value of 0
means there is no maximal display depth. Default is 3.
``o``, ``out`` <filename> <exp>
Evaluate the expression <exp> and store its string representation into a
file named <filename>. If the file does not exist, it will be created,
otherwise it will be opened for appending.
``w``, ``where``
Display the current execution point.
``u``, ``up``
Go up in the call stack.
``d``, ``down``
Go down in the call stack.
``stackframe`` [file]
Displays the content of the current stack frame in ``stdout`` or
appends it to the file, depending on whether a file is given.
``callstack``
Display the entire call stack (but not its content).
``l``, ``locals``
Display the available local variables in the current stack frame.
``g``, ``globals``
Display all the global variables that are available for inspection.

44
doc/filelist.txt Executable file
View file

@ -0,0 +1,44 @@
Short description of Nimrod's modules
-------------------------------------
============== ==========================================================
Module Description
============== ==========================================================
lexbase buffer handling of the lexical analyser
scanner lexical analyser
ast type definitions of the abstract syntax tree (AST) and
node constructors
astalgo algorithms for containers of AST nodes; converting the
AST to YAML; the symbol table
trees few algorithms for nodes; this module is less important
types module for traversing type graphs; also contain several
helpers for dealing with types
sigmatch contains the matching algorithm that is used for proc
calls
semexprs contains the semantic checking phase for expressions
semstmts contains the semantic checking phase for statements
semtypes contains the semantic checking phase for types
idents implements a general mapping from identifiers to an internal
representation (``PIdent``) that is used, so that a simple
pointer comparison suffices to say whether two Nimrod
identifiers are equivalent
ropes implements long strings using represented as trees for
lazy evaluation; used mainly by the code generators
ccgobj contains type definitions neeeded for C code generation
and some helpers
ccgmangl contains the name mangler for converting Nimrod
identifiers to their C counterparts
ccgutils contains helpers for the C code generator
ccgtemps contains the handling of temporary variables for the
C code generator
ccgtypes the generator for C types
ccgstmts the generator for statements
ccgexprs the generator for expressions
extccomp this module calls the C compiler and linker; interesting
if you want to add support for a new C compiler
============== ==========================================================

186
doc/grammar.txt Executable file
View file

@ -0,0 +1,186 @@
module ::= ([COMMENT] [SAD] stmt)*
optComma ::= [ ',' ] [COMMENT] [IND]
operator ::= OP0 | OR | XOR | AND | OP3 | OP4 | OP5 | IS | ISNOT | IN | NOTIN
| OP6 | DIV | MOD | SHL | SHR | OP7 | NOT
prefixOperator ::= OP0 | OP3 | OP4 | OP5 | OP6 | OP7 | NOT
optInd ::= [COMMENT] [IND]
lowestExpr ::= orExpr ( OP0 optInd orExpr )*
orExpr ::= andExpr ( OR | XOR optInd andExpr )*
andExpr ::= cmpExpr ( AND optInd cmpExpr )*
cmpExpr ::= ampExpr ( OP3 | IS | ISNOT | IN | NOTIN optInd ampExpr )*
ampExpr ::= plusExpr ( OP4 optInd plusExpr )*
plusExpr ::= mulExpr ( OP5 optInd mulExpr )*
mulExpr ::= dollarExpr ( OP6 | DIV | MOD | SHL | SHR optInd dollarExpr )*
dollarExpr ::= primary ( OP7 optInd primary )*
namedTypeOrExpr ::=
DOTDOT [expr]
| expr [EQUALS (expr [DOTDOT expr] | typeDescK | DOTDOT [expr] )
| DOTDOT [expr]]
| typeDescK
castExpr ::= CAST BRACKET_LE optInd typeDesc BRACKERT_RI
PAR_LE optInd expr PAR_RI
addrExpr ::= ADDR PAR_LE optInd expr PAR_RI
symbol ::= ACC (KEYWORD | IDENT | operator | PAR_LE PAR_RI
| BRACKET_LE BRACKET_RI) ACC | IDENT
accExpr ::= KEYWORD | IDENT | operator [DOT KEYWORD | IDENT | operator]
paramList
primary ::= ( prefixOperator optInd )* ( IDENT | literal | ACC accExpr ACC
| castExpr | addrExpr ) (
DOT optInd symbol
#| CURLY_LE namedTypeDescList CURLY_RI
| PAR_LE optInd
namedExprList
PAR_RI
| BRACKET_LE optInd
(namedTypeOrExpr optComma)*
BRACKET_RI
| CIRCUM
| pragma )*
literal ::= INT_LIT | INT8_LIT | INT16_LIT | INT32_LIT | INT64_LIT
| FLOAT_LIT | FLOAT32_LIT | FLOAT64_LIT
| STR_LIT | RSTR_LIT | TRIPLESTR_LIT
| CHAR_LIT | RCHAR_LIT
| NIL
| BRACKET_LE optInd (expr [COLON expr] optComma )* BRACKET_RI # []-Constructor
| CURLY_LE optInd (expr [DOTDOT expr] optComma )* CURLY_RI # {}-Constructor
| PAR_LE optInd (expr [COLON expr] optComma )* PAR_RI # ()-Constructor
exprList ::= ( expr optComma )*
namedExpr ::= expr [EQUALS expr] # actually this is symbol EQUALS expr|expr
namedExprList ::= ( namedExpr optComma )*
exprOrSlice ::= expr [ DOTDOT expr ]
sliceList ::= ( exprOrSlice optComma )+
anonymousProc ::= LAMBDA paramList [pragma] EQUALS stmt
expr ::= lowestExpr
| anonymousProc
| IF expr COLON expr
(ELIF expr COLON expr)*
ELSE COLON expr
namedTypeDesc ::= typeDescK | expr [EQUALS (typeDescK | expr)]
namedTypeDescList ::= ( namedTypeDesc optComma )*
qualifiedIdent ::= symbol [ DOT symbol ]
typeDescK ::= VAR typeDesc
| REF typeDesc
| PTR typeDesc
| TYPE expr
| PROC paramList [pragma]
typeDesc ::= typeDescK | primary
optSemicolon ::= [SEMICOLON]
macroStmt ::= COLON [stmt] (OF [sliceList] COLON stmt
| ELIF expr COLON stmt
| EXCEPT exceptList COLON stmt )*
[ELSE COLON stmt]
simpleStmt ::= returnStmt
| yieldStmt
| discardStmt
| raiseStmt
| breakStmt
| continueStmt
| pragma
| importStmt
| fromStmt
| includeStmt
| exprStmt
complexStmt ::= ifStmt | whileStmt | caseStmt | tryStmt | forStmt
| blockStmt | asmStmt
| procDecl | iteratorDecl | macroDecl | templateDecl
| constSection | typeSection | whenStmt | varSection
indPush ::= IND # push
stmt ::= simpleStmt [SAD]
| indPush (complexStmt | simpleStmt)
([SAD] (complexStmt | simpleStmt) )*
DED
exprStmt ::= lowestExpr [EQUALS expr | (expr optComma)* [macroStmt]]
returnStmt ::= RETURN [expr]
yieldStmt ::= YIELD expr
discardStmt ::= DISCARD expr
raiseStmt ::= RAISE [expr]
breakStmt ::= BREAK [symbol]
continueStmt ::= CONTINUE
ifStmt ::= IF expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
whenStmt ::= WHEN expr COLON stmt (ELIF expr COLON stmt)* [ELSE COLON stmt]
caseStmt ::= CASE expr (OF sliceList COLON stmt)*
(ELIF expr COLON stmt)*
[ELSE COLON stmt]
whileStmt ::= WHILE expr COLON stmt
forStmt ::= FOR (symbol optComma)+ IN expr [DOTDOT expr] COLON stmt
exceptList ::= (qualifiedIdent optComma)*
tryStmt ::= TRY COLON stmt
(EXCEPT exceptList COLON stmt)*
[FINALLY COLON stmt]
asmStmt ::= ASM [pragma] (STR_LIT | RSTR_LIT | TRIPLESTR_LIT)
blockStmt ::= BLOCK [symbol] COLON stmt
importStmt ::= IMPORT ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) [AS symbol] optComma)+
includeStmt ::= INCLUDE ((symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) optComma)+
fromStmt ::= FROM (symbol | STR_LIT | RSTR_LIT | TRIPLESTR_LIT) IMPORT (symbol optComma)+
pragma ::= CURLYDOT_LE (expr [COLON expr] optComma)+ (CURLYDOT_RI | CURLY_RI)
paramList ::= [PAR_LE ((symbol optComma)+ COLON typeDesc optComma)* PAR_RI] [COLON typeDesc]
genericParams ::= BRACKET_LE (symbol [EQUALS typeDesc] )* BRACKET_RI
procDecl ::= PROC symbol ["*"] [genericParams]
paramList [pragma]
[EQUALS stmt]
macroDecl ::= MACRO symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
iteratorDecl ::= ITERATOR symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
templateDecl ::= TEMPLATE symbol ["*"] [genericParams] paramList [pragma]
[EQUALS stmt]
colonAndEquals ::= [COLON typeDesc] EQUALS expr
constDecl ::= symbol ["*"] [pragma] colonAndEquals [COMMENT | IND COMMENT]
| COMMENT
constSection ::= CONST indPush constDecl (SAD constDecl)* DED
typeDef ::= typeDesc | recordDef | objectDef | enumDef
recordIdentPart ::=
(symbol ["*" | "-"] [pragma] optComma)+ COLON typeDesc [COMMENT | IND COMMENT]
recordWhen ::= WHEN expr COLON [COMMENT] recordPart
(ELIF expr COLON [COMMENT] recordPart)*
[ELSE COLON [COMMENT] recordPart]
recordCase ::= CASE expr COLON typeDesc [COMMENT]
(OF sliceList COLON [COMMENT] recordPart)*
[ELSE COLON [COMMENT] recordPart]
recordPart ::= recordWhen | recordCase | recordIdentPart
| indPush recordPart (SAD recordPart)* DED
recordDef ::= RECORD [pragma] recordPart
objectDef ::= OBJECT [pragma] [OF typeDesc] recordPart
enumDef ::= ENUM [OF typeDesc] (symbol [EQUALS expr] optComma [COMMENT | IND COMMENT])+
typeDecl ::= COMMENT
| symbol ["*"] [genericParams] [EQUALS typeDef] [COMMENT | IND COMMENT]
typeSection ::= TYPE indPush typeDecl (SAD typeDecl)* DED
colonOrEquals ::= COLON typeDesc [EQUALS expr] | EQUALS expr
varPart ::= (symbol ["*" | "-"] [pragma] optComma)+ colonOrEquals [COMMENT | IND COMMENT]
varSection ::= VAR (varPart | indPush (COMMENT|varPart) (SAD (COMMENT|varPart))* DED)

1
doc/html/empty.txt Executable file
View file

@ -0,0 +1 @@
This file keeps several tools from deleting this subdirectory.

575
doc/intern.txt Executable file
View file

@ -0,0 +1,575 @@
=========================================
Internals of the Nimrod Compiler
=========================================
:Author: Andreas Rumpf
:Version: |nimrodversion|
.. contents::
Directory structure
===================
The Nimrod project's directory structure is:
============ ==============================================
Path Purpose
============ ==============================================
``bin`` binary files go into here
``nim`` Pascal sources of the Nimrod compiler; this
should be modified, not the Nimrod version in
``rod``!
``rod`` Nimrod sources of the Nimrod compiler;
automatically generated from the Pascal
version
``data`` data files that are used for generating source
code go into here
``doc`` the documentation lives here; it is a bunch of
reStructuredText files
``dist`` download packages as zip archives go into here
``config`` configuration files for Nimrod go into here
``lib`` the Nimrod library lives here; ``rod`` depends
on it!
``web`` website of Nimrod; generated by ``genweb.py``
from the ``*.txt`` and ``*.tmpl`` files
``koch`` the Koch Build System (written for Nimrod)
``obj`` generated ``*.obj`` files go into here
============ ==============================================
Bootstrapping the compiler
==========================
The compiler is written in a subset of Pascal with special annotations so
that it can be translated to Nimrod code automatically. This conversion is
done by Nimrod itself via the undocumented ``boot`` command. Thus both Nimrod
and Free Pascal can compile the Nimrod compiler.
Requirements for bootstrapping:
- Free Pascal (I used version 2.2); it may not be needed
- Python (should work with 2.4 or higher) and the code generator *cog*
(included in this distribution!)
- C compiler -- one of:
* win32-lcc
* Borland C++ (tested with 5.5)
* Microsoft C++
* Digital Mars C++
* Watcom C++ (currently broken; a fix is welcome!)
* GCC
* Intel C++
* Pelles C
* llvm-gcc
| Compiling the compiler is a simple matter of running:
| ``koch.py boot``
| Or you can compile by hand, this is not difficult.
If you want to debug the compiler, use the command::
koch.py boot --debugger:on
The ``koch.py`` script is Nimrod's maintainance script: Everything that has
been automated is accessible with it. It is a replacement for make and shell
scripting with the advantage that it is more portable and is easier to read.
Coding standards
================
The compiler is written in a subset of Pascal with special annotations so
that it can be translated to Nimrod code automatically. As a generell rule,
Pascal code that does not translate to Nimrod automatically is forbidden.
Porting to new platforms
========================
Porting Nimrod to a new architecture is pretty easy, since C is the most
portable programming language (within certain limits) and Nimrod generates
C code, porting the code generator is not necessary.
POSIX-compliant systems on conventional hardware are usually pretty easy to
port: Add the platform to ``platform`` (if it is not already listed there),
check that the OS, System modules work and recompile Nimrod.
The only case where things aren't as easy is when the garbage
collector needs some assembler tweaking to work. The standard
version of the GC uses C's ``setjmp`` function to store all registers
on the hardware stack. It may be that the new platform needs to
replace this generic code by some assembler code.
Runtime type information
========================
*Runtime type information* (RTTI) is needed for several aspects of the Nimrod
programming language:
Garbage collection
The most important reason for RTTI. Generating
traversal procedures produces bigger code and is likely to be slower on
modern hardware as dynamic procedure binding is hard to predict.
Complex assignments
Sequences and strings are implemented as
pointers to resizeable buffers, but Nimrod requires copying for
assignments. Apart from RTTI the compiler could generate copy procedures
for any type that needs one. However, this would make the code bigger and
the RTTI is likely already there for the GC.
We already knew the type information as a graph in the compiler.
Thus we need to serialize this graph as RTTI for C code generation.
Look at the files ``lib/typeinfo.nim``, ``lib/hti.nim`` for more information.
However, generating type information proved to be difficult and the format
wastes memory. Variant records make problems too. We use a mix of iterator
procedures and constant data structures:
.. code-block:: Nimrod
type
TNimTypeSlot {.export.} = record
offset: int
typ: int
name: CString
TSlotIterator = proc (obj: pointer, field: int): ptr TNimTypeSlot
TNimType {.export.} = record
Kind: TNimTypeKind
baseType, indexType: int
size, len: int
slots: TSlotIterator # instead of: ptr array [0..10_000, TNimTypeSlot]
This is not easy to understand either. Best is to use just the ``rodgen``
module and store type information as string constants.
After thinking I came to the conclusion that this is again premature
optimization. We should just construct the type graph at runtime. In the init
section new types should be constructed and registered:
.. code-block:: Nimrod
type
TSlotTriple = record
offset: int
typ: PRTL_Type
name: Cstring
PSlots = ptr TSlots
TSlots = record
case kind
of linear:
fields: array [TSlotTriple]
of nested:
discriminant: TSlotTriple
otherSlots: array [discriminant, PSlots]
TTypeKind = enum ...
RTL_Type = record
size: int
base: PRTL_Type
case kind
of tyArray, tySequence:
elemSize: int
of tyRecord, tyObject, tyEnum:
slots: PSlots
The Garbage Collector
=====================
Introduction
------------
We use the term *cell* here to refer to everything that is traced
(sequences, refs, strings).
This section describes how the new GC works. The old algorithms
all had the same problem: Too complex to get them right. This one
tries to find the right compromise.
The basic algorithm is *Deferrent reference counting* with cycle detection.
References in the stack are not counted for better performance and easier C
code generation. The GC starts by traversing the hardware stack and increments
the reference count (RC) of every cell that it encounters. After the GC has
done its work the stack is traversed again and the RC of every cell
that it encounters are decremented again. Thus no marking bits in the RC are
needed. Between these stack traversals the GC has a complete accurate view over
the RCs.
Each cell has a header consisting of a RC and a pointer to its type
descriptor. However the program does not know about these, so they are placed at
negative offsets. In the GC code the type ``PCell`` denotes a pointer
decremented by the right offset, so that the header can be accessed easily. It
is extremely important that ``pointer`` is not confused with a ``PCell``
as this would lead to a memory corruption.
When to trigger a collection
----------------------------
Since there are really two different garbage collectors (reference counting
and mark and sweep) we use two different heuristics when to run the passes.
The RC-GC pass is fairly cheap: Thus we use an additive increase (7 pages)
for the RC_Threshold and a multiple increase for the CycleThreshold.
The AT and ZCT sets
-------------------
The GC maintains two sets throughout the lifetime of
the program (plus two temporary ones). The AT (*any table*) simply contains
every cell. The ZCT (*zero count table*) contains every cell whose RC is
zero. This is used to reclaim most cells fast.
The ZCT contains redundant information -- the AT alone would suffice.
However, traversing the AT and look if the RC is zero would touch every living
cell in the heap! That's why the ZCT is updated whenever a RC drops to zero.
The ZCT is not updated when a RC is incremented from zero to one, as
this would be too costly.
The CellSet data structure
--------------------------
The AT and ZCT depend on an extremely efficient datastructure for storing a
set of pointers - this is called a ``PCellSet`` in the source code.
Inserting, deleting and searching are done in constant time. However,
modifying a ``PCellSet`` during traversation leads to undefined behaviour.
.. code-block:: Nimrod
type
PCellSet # hidden
proc allocCellSet: PCellSet # make a new set
proc deallocCellSet(s: PCellSet) # empty the set and free its memory
proc incl(s: PCellSet, elem: PCell) # include an element
proc excl(s: PCellSet, elem: PCell) # exclude an element
proc `in`(elem: PCell, s: PCellSet): bool
iterator elements(s: PCellSet): (elem: PCell)
All the operations have to be performed efficiently. Because a Cellset can
become huge (the AT contains every allocated cell!) a hash table is not
suitable for this.
We use a mixture of bitset and patricia tree for this. One node in the
patricia tree contains a bitset that decribes a page of the operating system
(not always, but that doesn't matter).
So including a cell is done as follows:
- Find the page descriptor for the page the cell belongs to.
- Set the appropriate bit in the page descriptor indicating that the
cell points to the start of a memory block.
Removing a cell is analogous - the bit has to be set to zero.
Single page descriptors are never deleted from the tree. Typically a page
descriptor is only 19 words big, so it does not waste much by not deleting
it. Apart from that the AT and ZCT are rebuilt frequently, so removing a
single page descriptor from the tree is never necessary.
Complete traversal is done like so::
for each page decriptor d:
for each bit in d:
if bit == 1:
traverse the pointer belonging to this bit
Further complications
---------------------
In Nimrod the compiler cannot always know if a reference
is stored on the stack or not. This is caused by var parameters.
Consider this example:
.. code-block:: Nimrod
proc setRef(r: var ref TNode) =
new(r)
proc usage =
var
r: ref TNode
setRef(r) # here we should not update the reference counts, because
# r is on the stack
setRef(r.left) # here we should update the refcounts!
Though it would be possible to produce code updating the refcounts (if
necessary) before and after the call to ``setRef``, it is a complex task to
do so in the code generator. So we don't and instead decide at runtime
whether the reference is on the stack or not. The generated code looks
roughly like this:
.. code-block:: C
void setref(TNode** ref) {
unsureAsgnRef(ref, newObj(TNode_TI, sizeof(TNode)))
}
void usage(void) {
setRef(&r)
setRef(&r->left)
}
Note that for systems with a continous stack (which most systems have)
the check whether the ref is on the stack is very cheap (only two
comparisons). Another advantage of this scheme is that the code produced is
a tiny bit smaller.
The algorithm in pseudo-code
----------------------------
Now we come to the nitty-gritty. The algorithm works in several phases.
Phase 1 - Consider references from stack
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
::
for each pointer p in the stack: incRef(p)
This is necessary because references in the hardware stack are not traced for
better performance. After Phase 1 the RCs are accurate.
Phase 2 - Free the ZCT
~~~~~~~~~~~~~~~~~~~~~~
This is how things used to (not) work::
for p in elements(ZCT):
if RC(p) == 0:
call finalizer of p
for c in children(p): decRef(c) # free its children recursively
# if necessary; the childrens RC >= 1, BUT they may still be in the ZCT!
free(p)
else:
remove p from the ZCT
Instead we do it this way. Note that the recursion is gone too!
::
newZCT = nil
for p in elements(ZCT):
if RC(p) == 0:
call finalizer of p
for c in children(p):
assert(RC(c) > 0)
dec(RC(c))
if RC(c) == 0:
if newZCT == nil: newZCT = allocCellSet()
incl(newZCT, c)
free(p)
else:
# nothing to do! We will use the newZCS
deallocCellSet(ZCT)
ZCT = newZCT
This phase is repeated until enough memory is available or the ZCT is nil.
If still not enough memory is available the cyclic detector gets its chance
to do something.
Phase 3 - Cycle detection
~~~~~~~~~~~~~~~~~~~~~~~~~
Cycle detection works by subtracting internal reference counts::
newAT = allocCellSet()
for y in elements(AT):
# pretend that y is dead:
for c in children(y):
dec(RC(c))
# note that this should not be done recursively as we have all needed
# pointers in the AT! This makes it more efficient too!
proc restore(y: PCell) =
# unfortunately, the recursion here cannot be eliminated easily
if y not_in newAT:
incl(newAT, y)
for c in children(y):
inc(RC(c)) # restore proper reference counts!
restore(c)
for y in elements(AT) with rc > 0:
restore(y)
for y in elements(AT) with rc == 0:
free(y) # if pretending worked, it was part of a cycle
AT = newAT
Phase 4 - Ignore references from stack again
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
::
for each pointer p in the stack:
dec(RC(p))
if RC(p) == 0: incl(ZCT, p)
Now the RCs correctly discard any references from the stack. One can also see
this as a temporary marking operation. Things that are referenced from stack
are marked during the GC's operarion and now have to be unmarked.
The compiler's architecture
===========================
Nimrod uses the classic compiler architecture: A scanner feds tokens to a
parser. The parser builds a syntax tree that is used by the code generator.
This syntax tree is the interface between the parser and the code generator.
It is essential to understand most of the compiler's code.
In order to compile Nimrod correctly, type-checking has to be seperated from
parsing. Otherwise generics would not work. Code generation is done for a
whole module only after it has been checked for semantics.
.. include:: filelist.txt
The first command line argument selects the backend. Thus the backend is
responsible for calling the parser and semantic checker. However, when
compiling ``import`` or ``include`` statements, the semantic checker needs to
call the backend, this is done by embedding a PBackend into a TContext.
The syntax tree
---------------
The synax tree consists of nodes which may have an arbitrary number of
children. Types and symbols are represented by other nodes, because they
may contain cycles. The AST changes its shape after semantic checking. This
is needed to make life easier for the code generators. See the "ast" module
for the type definitions.
We use the notation ``nodeKind(fields, [sons])`` for describing
nodes. ``nodeKind[sons]`` is a short-cut for ``nodeKind([sons])``.
XXX: Description of the language's syntax and the corresponding trees.
How the RTL is compiled
=======================
The system module contains the part of the RTL which needs support by
compiler magic (and the stuff that needs to be in it because the spec
says so). The C code generator generates the C code for it just like any other
module. However, calls to some procedures like ``addInt`` are inserted by
the CCG. Therefore the module ``magicsys`` contains a table
(``compilerprocs``) with all symbols that are marked as ``compilerproc``.
How separate compilation will work
==================================
Soon compiling from scratch every module that's needed will become too slow as
programs grow. For easier cleaning all generated files are generated in the
directory: ``$base/rod_gen``. This cannot be changed. The generated C files
get the names of the modules they result from. A compiled Nimrod module has the
extension ``.rod`` and is a binary file. The format may change from release
to release. The rod-file is mostly a binary representation of the parse trees.
Nimrod currently compiles any module into its own C file. Some things like
type-information, common string literals, common constant sets need to be
shared though. We deal with this problem by writing the shared data
in the main C file. Only "headers" are generated in the other modules. However,
each precompiled Nimrod module lists the shared data it depends on. The same
holds for procedures that have to generated from generics.
A big problem is that the data must get the same name each time it is compiled.
The C compiler is only called for the C files, that changed after the last
compilation (or if their object file does not exist anymore). To work
reliably, in the header comment of the C file these things are listed, so
that the C compiler is called again should they change:
* Nimrod's Version
* the target CC
* the target OS
* the target CPU
The version is questionable: If the resulting C file is the same, it does not
matter that Nimrods's version has increased. We do it anyway to be on the safe
side.
Generation of dynamic link libraries
====================================
Generation of dynamic link libraries or shared libraries is not difficult; the
underlying C compiler already does all the hard work for us. The problem is the
common runtime library, especially the memory manager. Note that Borland's
Delphi had exactly the same problem. The workaround is to not link the GC with
the Dll and provide an extra runtime dll that needs to be initialized.
How to implement closures
=========================
A closure is a record of a proc pointer and a context ref. The context ref
points to a garbage collected record that contains the needed variables.
An example:
.. code-block:: Nimrod
type
TListRec = record
data: string
next: ref TListRec
proc forEach(head: ref TListRec, visitor: proc (s: string) {.closure.}) =
var it = head
while it != nil:
visit(it.data)
it = it.next
proc sayHello() =
var L = new List(["hallo", "Andreas"])
var temp = "jup\xff"
forEach(L, lambda(s: string) =
io.write(temp)
io.write(s)
)
This should become the following in C:
.. code-block:: C
typedef struct ... /* List type */
typedef struct closure {
void (*PrcPart)(string, void*);
void* ClPart;
}
typedef struct Tcl_data {
string temp; // all accessed variables are put in here!
}
void forEach(TListRec* head, const closure visitor) {
TListRec* it = head;
while (it != NIM_NULL) {
visitor.prc(it->data, visitor->cl_data);
it = it->next;
}
}
void printStr(string s, void* cl_data) {
Tcl_data* x = (Tcl_data*) cl_data;
io_write(x->temp);
io_write(s);
}
void sayhello() {
Tcl_data* data = new(...);
asgnRef(&data->temp, "jup\xff");
...
closure cl;
cl.prc = printStr;
cl.cl_data = data;
foreach(L, cl);
}
What about nested closure? - There's not much difference: Just put all used
variables in the data record.

48
doc/lib.txt Executable file
View file

@ -0,0 +1,48 @@
=======================
Nimrod Standard Library
=======================
:Author: Andreas Rumpf
:Version: |nimrodversion|
Though the Nimrod Standard Library is still evolving, it is already quite
usable. It is divided into basic libraries that contains modules that virtually
every program will need and advanced libraries which are more heavy weight.
Advanced libraries are in the ``lib/base`` directory.
Basic libraries
===============
* `System <system.html>`_
Basic procs and operators that every program needs. It also provides IO
facilities for reading and writing text and binary files. It is imported
implicitly by the compiler. Do not import it directly. It relies on compiler
magic to work.
* `Strutils <strutils.html>`_
This module contains common string handling operations like converting a
string into uppercase, splitting a string into substrings, searching for
substrings, replacing substrings.
* `OS <os.html>`_
Basic operating system facilities like retrieving environment variables,
reading command line arguments, working with directories, running shell
commands, etc. This module is -- like any other basic library --
platform independant.
* `Math <math.html>`_
Mathematical operations like cosine, square root.
* `Complex <complex.html>`_
This module implements complex numbers and their mathematical operations.
* `Times <times.html>`_
The ``times`` module contains basic support for working with time.
Advanced libaries
=================
* `Regexprs <regexprs.html>`_
This module contains procedures and operators for handling regular
expressions.

1742
doc/manual.txt Executable file

File diff suppressed because it is too large Load diff

295
doc/nimdoc.css Executable file
View file

@ -0,0 +1,295 @@
/*
:Author: David Goodger
:Contact: goodger@python.org
:Date: $Date: 2006-05-21 22:44:42 +0200 (Sun, 21 May 2006) $
:Revision: $Revision: 4564 $
:Copyright: This stylesheet has been placed in the public domain.
Default cascading style sheet for the HTML output of Docutils.
See http://docutils.sf.net/docs/howto/html-stylesheets.html for how to
customize this style sheet.
*/
/*
Modified for the Nimrod Documenation by
Andreas Rumpf
*/
/* used to remove borders from tables and images */
.borderless, table.borderless td, table.borderless th {
border: 0 }
table.borderless td, table.borderless th {
/* Override padding for "table.docutils td" with "! important".
The right padding separates the table cells. */
padding: 0 0.5em 0 0 ! important }
.first {
/* Override more specific margin styles with "! important". */
margin-top: 0 ! important }
.last, .with-subtitle {
margin-bottom: 0 ! important }
.hidden {
display: none }
a.toc-backref {
text-decoration: none ;
color: black }
blockquote.epigraph {
margin: 2em 5em ; }
dl.docutils dd {
margin-bottom: 0.5em }
/* Uncomment (and remove this text!) to get bold-faced definition list terms
dl.docutils dt {
font-weight: bold }
*/
div.abstract {
margin: 2em 5em }
div.abstract p.topic-title {
font-weight: bold ;
text-align: center }
div.admonition, div.attention, div.caution, div.danger, div.error,
div.hint, div.important, div.note, div.tip, div.warning {
margin: 2em ;
border: medium outset ;
padding: 1em }
div.admonition p.admonition-title, div.hint p.admonition-title,
div.important p.admonition-title, div.note p.admonition-title,
div.tip p.admonition-title {
font-weight: bold ;
font-family: sans-serif }
div.attention p.admonition-title, div.caution p.admonition-title,
div.danger p.admonition-title, div.error p.admonition-title,
div.warning p.admonition-title {
color: red ;
font-weight: bold ;
font-family: sans-serif }
/* Uncomment (and remove this text!) to get reduced vertical space in
compound paragraphs.
div.compound .compound-first, div.compound .compound-middle {
margin-bottom: 0.5em }
div.compound .compound-last, div.compound .compound-middle {
margin-top: 0.5em }
*/
div.dedication {
margin: 2em 5em ;
text-align: center ;
font-style: italic }
div.dedication p.topic-title {
font-weight: bold ;
font-style: normal }
div.figure {
margin-left: 2em ;
margin-right: 2em }
div.footer, div.header {
clear: both;
font-size: smaller }
div.line-block {
display: block ;
margin-top: 1em ;
margin-bottom: 1em }
div.line-block div.line-block {
margin-top: 0 ;
margin-bottom: 0 ;
margin-left: 1.5em }
div.sidebar {
margin-left: 1em ;
border: medium outset ;
padding: 1em ;
background-color: #ffffee ;
width: 40% ;
float: right ;
clear: right }
div.sidebar p.rubric {
font-family: sans-serif ;
font-size: medium }
div.system-messages {
margin: 5em }
div.system-messages h1 {
color: red }
div.system-message {
border: medium outset ;
padding: 1em }
div.system-message p.system-message-title {
color: red ;
font-weight: bold }
div.topic {
margin: 2em;
}
h1.section-subtitle, h2.section-subtitle, h3.section-subtitle,
h4.section-subtitle, h5.section-subtitle, h6.section-subtitle {
margin-top: 0.4em }
h1.title { text-align: center }
h2.subtitle { text-align: center }
hr.docutils { width: 75% }
img.align-left { clear: left }
img.align-right { clear: right }
ol.simple, ul.simple {
margin-bottom: 1em }
ol.arabic {
list-style: decimal }
ol.loweralpha {
list-style: lower-alpha }
ol.upperalpha {
list-style: upper-alpha }
ol.lowerroman {
list-style: lower-roman }
ol.upperroman {
list-style: upper-roman }
p.attribution {
text-align: right ;
margin-left: 50% }
p.caption {
font-style: italic }
p.credits {
font-style: italic ;
font-size: smaller }
p.label {
white-space: nowrap }
p.rubric {
font-weight: bold ;
font-size: larger ;
color: maroon ;
text-align: center }
p.sidebar-title {
font-family: sans-serif ;
font-weight: bold ;
font-size: larger }
p.sidebar-subtitle {
font-family: sans-serif ;
font-weight: bold }
p.topic-title {
font-weight: bold }
pre.address {
margin-bottom: 0 ;
margin-top: 0 ;
font-family: serif ;
font-size: 100% }
pre, span.pre {
background-color:#F9F9F9;
border:1px dotted #2F6FAB;
color:black;
}
pre {padding:1em;}
pre.literal-block, pre.doctest-block {
margin-left: 2em ;
margin-right: 2em }
span.classifier {
font-family: sans-serif ;
font-style: oblique }
span.classifier-delimiter {
font-family: sans-serif ;
font-weight: bold }
span.interpreted {
font-family: sans-serif }
span.option {
white-space: nowrap }
span.pre { white-space: pre }
span.problematic {
color: red }
span.section-subtitle {
/* font-size relative to parent (h1..h6 element) */
font-size: 80% }
table.citation {
border-left: solid 1px gray;
margin-left: 1px }
table.docinfo {
margin: 2em 4em }
table.docutils {
margin-top: 0.5em ;
margin-bottom: 0.5em }
table.footnote {
border-left: solid 1px black;
margin-left: 1px }
table.docutils td, table.docutils th,
table.docinfo td, table.docinfo th {
padding-left: 0.5em ;
padding-right: 0.5em ;
vertical-align: top }
table.docutils th.field-name, table.docinfo th.docinfo-name {
font-weight: bold ;
text-align: left ;
white-space: nowrap ;
padding-left: 0 }
h1 tt.docutils, h2 tt.docutils, h3 tt.docutils,
h4 tt.docutils, h5 tt.docutils, h6 tt.docutils {
font-size: 100% }
ul.auto-toc {
list-style-type: none }
a.reference {
color: #E00000;
font-weight:bold;
}
a.reference:hover {
color: #E00000;
background-color: #ffff00;
display: margin;
font-weight:bold;
}
div.topic ul {
list-style-type: none;
}

241
doc/nimrodc.txt Executable file
View file

@ -0,0 +1,241 @@
===================================
Nimrod Compiler User Guide
===================================
:Author: Andreas Rumpf
:Version: |nimrodversion|
.. contents::
Introduction
============
This document describes the usage of the *Nimrod compiler*
on the different supported platforms. It is not a definition of the Nimrod
programming system (therefore is the Nimrod manual).
Nimrod is free software; it is licensed under the
`GNU General Public License <gpl.html>`_.
Compiler Usage
==============
Command line switches
---------------------
Basis command line switches are:
.. include:: ../data/basicopt.txt
Advanced command line switches are:
.. include:: ../data/advopt.txt
Configuration file
------------------
The ``nimrod`` executable loads the configuration file ``config/nimrod.cfg``
unless this is suppressed by the ``--skip_cfg`` command line option.
Configuration settings can be overwritten in a project specific
configuration file that is read automatically. This specific file has to
be in the same directory as the project and be of the same name, except
that its extension should be ``.cfg``.
Command line settings have priority over configuration file settings.
Nimrod's directory structure
----------------------------
The generated files that Nimrod produces all go into a subdirectory called
``rod_gen``. This makes it easy to write a script that deletes all generated
files. For example the generated C code for the module ``path/modA.nim``
will become ``path/rod_gen/modA.c``.
However, the generated C code is not platform independant! C code generated for
Linux does not compile on Windows, for instance. The comment on top of the
C file lists the OS, CPU and CC the file has been compiled for.
The library lies in ``lib``. Directly in the library directory are essential
Nimrod modules like the ``system`` and ``os`` modules. Under ``lib/base``
are additional specialized libraries or interfaces to foreign libraries which
are included in the standard distribution. The ``lib/extra`` directory is
initially empty. Third party libraries should go there. In the default
configuration the compiler always searches for libraries in ``lib``,
``lib/base`` and ``lib/extra``.
Additional Features
===================
This section describes Nimrod's additional features that are not listed in the
Nimrod manual.
New Pragmas and Options
-----------------------
Because Nimrod generates C code it needs some "red tape" to work properly.
Thus lots of options and pragmas for tweaking the generated C code are
available.
No_decl Pragma
~~~~~~~~~~~~~~
The `no_decl`:idx: pragma can be applied to almost any symbol (variable, proc,
type, etc.) and is one of the most important for interoperability with C:
It tells Nimrod that it should not generate a declaration for the symbol in
the C code. Thus it makes the following possible, for example:
.. code-block:: Nimrod
var
EOF {.import: "EOF", no_decl.}: cint # pretend EOF was a variable, as
# Nimrod does not know its value
Varargs Pragma
~~~~~~~~~~~~~~
The `varargs`:idx: pragma can be applied to procedures only. It tells Nimrod
that the proc can take a variable number of parameters after the last
specified parameter. Nimrod string values will be converted to C
strings automatically:
.. code-block:: Nimrod
proc printf(formatstr: cstring) {.nodecl, varargs.}
printf("hallo %s", "world") # "world" will be passed as C string
Header Pragma
~~~~~~~~~~~~~
The `header`:idx: pragma is very similar to the ``no_decl`` pragma: It can be
applied to almost any symbol and specifies that not only it should not be
declared but also that it leads to the inclusion of a given header file:
.. code-block:: Nimrod
type
PFile {.import: "FILE*", header: "<stdio.h>".} = pointer
# import C's FILE* type; Nimrod will treat it as a new pointer type
The ``header`` pragma expects always a string constant. The string contant
contains the header file: As usual for C, a system header file is enclosed
in angle brackets: ``<>``. If no angle brackets are given, Nimrod
encloses the header file in ``""`` in the generated C code.
No_static Pragma
~~~~~~~~~~~~~~~~
The `no_static`:idx: pragma can be applied to almost any symbol and specifies
that it shall not be declared ``static`` in the generated C code. Note that
symbols in the interface part of a module never get declared ``static``, so
only in special cases is this pragma necessary.
Line_dir Option
~~~~~~~~~~~~~~~
The `line_dir`:idx: option can be turned on or off. If on the generated C code
contains ``#line`` directives.
Stack_trace Option
~~~~~~~~~~~~~~~~~~
If the `stack_trace`:idx: option is turned on, the generated C contains code to
ensure that proper stack traces are given if the program crashes or an
uncaught exception is raised.
Line_trace Option
~~~~~~~~~~~~~~~~~
The `line_trace`:idx: option implies the ``stack_trace`` option. If turned on,
the generated C contains code to ensure that proper stack traces with line
number information are given if the program crashes or an uncaught exception
is raised.
Debugger Option
~~~~~~~~~~~~~~~
The `debugger`:idx: option enables or disables the *Embedded Nimrod Debugger*.
See the documentation of endb_ for further information.
Breakpoint Pragma
~~~~~~~~~~~~~~~~~
The *breakpoint* pragma was specially added for the sake of debugging with
ENDB. See the documentation of `endb <endb.html>`_ for further information.
Volatile Pragma
~~~~~~~~~~~~~~~
The `volatile`:idx: pragma is for variables only. It declares the variable as
``volatile``, whatever that means in C/C++.
Register Pragma
~~~~~~~~~~~~~~~
The `register`:idx: pragma is for variables only. It declares the variable as
``register``, giving the compiler a hint that the variable should be placed
in a hardware register for faster access. C compilers usually ignore this
though and for good reason: Often they do a better job without it anyway.
In highly specific cases (a dispatch loop of interpreters for example) it
may provide benefits, though.
Disabling certain messages
--------------------------
Nimrod generates some warnings and hints ("line too long") that may annoy the
user. Thus a mechanism for disabling certain messages is provided: Each hint
and warning message contains a symbol in brackets. This is the message's
identifier that can be used to enable or disable it:
.. code-block:: Nimrod
{.warning[LineTooLong]: off.} # turn off warning about too long lines
This is often better than disabling all warnings at once.
Debugging with Nimrod
=====================
Nimrod comes with its own *Embedded Nimrod Debugger*. See
the documentation of endb_ for further information.
Optimizing for Nimrod
=====================
Nimrod has no separate optimizer, but the C code that is produced is very
efficient. Most C compilers have excellent optimizers, so usually it is
not needed to optimize one's code. Nimrod has been designed to encourage
efficient code: The most readable code in Nimrod is often the most efficient
too.
However, sometimes one has to optimize. Do it in the following order:
1. switch off the embedded debugger (it is **slow**!)
2. turn on the optimizer and turn off runtime checks
3. profile your code to find where the bottlenecks are
4. try to find a better algorithm
5. do low-level optimizations
This section can only help you with the last item. Note that rewriting parts
of your program in C is *never* necessary to speed up your program, because
everything that can be done in C can be done in Nimrod. Rewriting parts in
assembler *might*.
Optimizing string handling
--------------------------
String assignments are sometimes expensive in Nimrod: They are required to
copy the whole string. However, the compiler is often smart enough to not copy
strings. Due to the argument passing semantics, strings are never copied when
passed to subroutines. The compiler does not copy strings that are returned by
a routine, because a routine returns a new string anyway. Thus it is efficient
to do:
.. code-block:: Nimrod
var s = procA() # assignment will not copy the string; procA allocates a new
# string anyway
However it is not efficient to do:
.. code-block:: Nimrod
var s = varA # assignment has to copy the whole string into a new buffer!
String case statements are optimized too. A hashing scheme is used for them
if several different string constants are used. This is likely to be more
efficient than any hand-coded scheme.

9
doc/overview.txt Executable file
View file

@ -0,0 +1,9 @@
=============================
Nimrod Documentation Overview
=============================
:Author: Andreas Rumpf
:Version: |nimrodversion|
.. include:: ../doc/docs.txt

220
doc/posix.txt Executable file
View file

@ -0,0 +1,220 @@
Function POSIX Description
access Tests for file accessibility
alarm Schedules an alarm
asctime Converts a time structure to a string
cfgetispeed Reads terminal input baud rate
cfgetospeed Reads terminal output baud rate
cfsetispeed Sets terminal input baud rate
cfsetospeed Sets terminal output baud rate
chdir Changes current working directory
chmod Changes file mode
chown Changes owner and/or group of a file
close Closes a file
closedir Ends directory read operation
creat Creates a new file or rewrites an existing one
ctermid Generates terminal pathname
cuserid Gets user name
dup Duplicates an open file descriptor
dup2 Duplicates an open file descriptor
execl Executes a file
execle Executes a file
execlp Executes a file
execv Executes a file
execve Executes a file
execvp Executes a file
_exit Terminates a process
fcntl Manipulates an open file descriptor
fdopen Opens a stream on a file descriptor
fork Creates a process
fpathconf Gets configuration variable for an open file
fstat Gets file status
getcwd Gets current working directory
getegid Gets effective group ID
getenv Gets environment variable
geteuid Gets effective user ID
getgid Gets real group ID
getgrgid Reads groups database based on group ID
getgrnam Reads groups database based on group name
getgroups Gets supplementary group IDs
getlogin Gets user name
getpgrp Gets process group ID
getpid Gets process ID
getppid Gets parent process ID
getpwnam Reads user database based on user name
getpwuid Reads user database based on user ID
getuid Gets real user ID
isatty Determines if a file descriptor is associated with a terminal
kill Sends a kill signal to a process
link Creates a link to a file
longjmp Restores the calling environment
lseek Repositions read/write file offset
mkdir Makes a directory
mkfifo Makes a FIFO special file
open Opens a file
opendir Opens a directory
pathconf Gets configuration variables for a path
pause Suspends a process execution
pipe Creates an interprocess channel
read Reads from a file
readdir Reads a directory
rename Renames a file
rewinddir Resets the readdir() pointer
rmdir Removes a directory
setgid Sets group ID
setjmp Saves the calling environment for use by longjmp()
setlocale Sets or queries a program's locale
setpgid Sets a process group ID for job control
setuid Sets the user ID
sigaction Examines and changes signal action
sigaddset Adds a signal to a signal set
sigdelset Removes a signal to a signal set
sigemptyset Creates an empty signal set
sigfillset Creates a full set of signals
sigismember Tests a signal for a selected member
siglongjmp Goes to and restores signal mask
sigpending Examines pending signals
sigprocmask Examines and changes blocked signals
sigsetjmp Saves state for siglongjmp()
sigsuspend Waits for a signal
sleep Delays process execution
stat Gets information about a file
sysconf Gets system configuration information
tcdrain Waits for all output to be transmitted to the terminal
tcflow Suspends/restarts terminal output
tcflush Discards terminal data
tcgetattr Gets terminal attributes
tcgetpgrp Gets foreground process group ID
tcsendbreak Sends a break to a terminal
tcsetattr Sets terminal attributes
tcsetpgrp Sets foreground process group ID
time Determines the current calendar time
times Gets process times
ttyname Determines a terminal pathname
tzset Sets the timezone from environment variables
umask Sets the file creation mask
uname Gets system name
unlink Removes a directory entry
utime Sets file access and modification times
waitpid Waits for process termination
write Writes to a file
POSIX.1b function calls Function POSIX Description
aio_cancel Tries to cancel an asynchronous operation
aio_error Retrieves the error status for an asynchronous operation
aio_read Asynchronously reads from a file
aio_return Retrieves the return status for an asynchronous operation
aio_suspend Waits for an asynchronous operation to complete
aio_write Asynchronously writes to a file
clock_getres Gets resolution of a POSIX.1b clock
clock_gettime Gets the time according to a particular POSIX.1b clock
clock_settime Sets the time according to a particular POSIX.1b clock
fdatasync Synchronizes at least the data part of a file with the underlying media
fsync Synchronizes a file with the underlying media
kill, sigqueue Sends signals to a process
lio_listio Performs a list of I/O operations, synchronously or asynchronously
mlock Locks a range of memory
mlockall Locks the entire memory space down
mmap Maps a shared memory object (or possibly another file) into process's address space
mprotect Changes memory protection on a mapped area
mq_close Terminates access to a POSIX.1b message queue
mq_getattr Gets POSIX.1b message queue attributes
mq_notify Registers a request to be notified when a message arrives on an empty message queue
mq_open Creates/accesses a POSIX.1b message queue
mq_receive Receives a message from a POSIX.1b message queue
mq_send Sends a message on a POSIX.1b message queue
mq_setattr Sets a subset of POSIX.1b message queue attributes
msync Makes a mapping consistent with the underlying object
munlock Unlocks a range of memory
munlockall Unlocks the entire address space
munmap Undo mapping established by mmap
nanosleep Pauses execution for a number of nanoseconds
sched_get_priority_max Gets maximum priority value for a scheduler
sched_get_priority_min Gets minimum priority value for a scheduler
sched_getparam Retrieves scheduling parameters for a particular process
sched_getscheduler Retrieves scheduling algorithm for a particular purpose
sched_rr_get_interval Gets the SCHED_RR interval for the named process
sched_setparam Sets scheduling parameters for a process
sched_setscheduler Sets scheduling algorithm/parameters for a process
sched_yield Yields the processor
sem_close Terminates access to a POSIX.1b semaphore
sem_destroy De-initializes a POSIX.1b unnamed semaphore
sem_getvalue Gets the value of a POSIX.1b semaphore
sem_open Creates/accesses a POSIX.1b named semaphore
sem_post Posts (signal) a POSIX.1b named or unnamed semaphore
sem_unlink Destroys a POSIX.1b named semaphore
sem_wait, sem_trywait Waits on a POSIX.1b named or unnamed semaphore
shm_open Creates/accesses a POSIX.1b shared memory object
shm_unlink Destroys a POSIX.1b shared memory object
sigwaitinfosigtimedwait Synchronously awaits signal arrival; avoid calling handler
timer_create Creates a POSIX.1b timer based on a particular clock
timer_delete Deletes a POSIX.1b timer
timer_gettime Time remaining on a POSIX.1b timer before expiration
timer_settime Sets expiration time/interval for a POSIX.1b timer
wait, waitpid Retrieves status of a terminated process and clean up corpse
POSIX.1c function calls Function POSIX Description
pthread_atfork Declares procedures to be called before and after a fork
pthread_attr_destroy Destroys a thread attribute object
pthread_attr_getdetachstate Obtains the setting of the detached state of a thread
pthread_attr_getinheritsched Obtains the setting of the scheduling inheritance of a thread
pthread_attr_getschedparam Obtains the parameters associated with the scheduling policy attribute of a thread
pthread_attr_getschedpolicy Obtains the setting of the scheduling policy of a thread
pthread_attr_getscope Obtains the setting of the scheduling scope of a thread
pthread_attr_getstackaddr Obtains the stack address of a thread
pthread_attr_getstacksize Obtains the stack size of a thread
pthread_attr_init Initializes a thread attribute object
pthread_attr_setdetachstate Adjusts the detached state of a thread
pthread_attr_setinheritsched Adjusts the scheduling inheritance of a thread
pthread_attr_setschedparam Adjusts the parameters associated with the scheduling policy of a thread
pthread_attr_setschedpolicy Adjusts the scheduling policy of a thread
pthread_attr_setscope Adjusts the scheduling scope of a thread
pthread_attr_setstackaddr Adjusts the stack address of a thread
pthread_attr_setstacksize Adjusts the stack size of a thread
pthread_cancel Cancels the specific thread
pthread_cleanup_pop Removes the routine from the top of a thread's cleanup stack, and if execute is nonzero, runs it
pthread_cleanup_push Places a routine on the top of a thread's cleanup stack
pthread_condattr_destroy Destroys a condition variable attribute object
pthread_condattr_getpshared Obtains the process-shared setting of a condition variable attribute object
pthread_condattr_init Initializes a condition variable attribute object
pthread_condattr_setpshared Sets the process-shared attribute in a condition variable attribute object to either PTHREAD_PROCESS_SHARED or PTHREAD_PROCESS_PRIVATE
pthread_cond_broadcast Unblocks all threads that are waiting on a condition variable
pthread_cond_destroy Destroys a condition variable
pthread_cond_init Initializes a condition variable with the attributes specified in the specified condition variable attribute object
pthread_cond_signal Unblocks at least one thread waiting on a condition variable
pthread_cond_timedwait Automatically unlocks the specified mutex, and places the calling thread into a wait state
pthread_cond_wait Automatically unlocks the specified mutex, and places the calling thread into a wait state
pthread_create Creates a thread with the attributes specified in attr
pthread_detach Marks a threads internal data structures for deletion
pthread_equal Compares one thread handle to another thread handle
pthread_exit Terminates the calling thread
pthread_getschedparam Obtains both scheduling policy and scheduling parameters of an existing thread
pthread_getspecific Obtains the thread specific data value associated with the specific key in the calling thread
pthread_join Causes the calling thread to wait for the specific thread’s termination
pthread_key_create Generates a unique thread-specific key that's visible to all threads in a process
pthread_key_delete Deletes a thread specific key
pthread_kill Delivers a signal to the specified thread
pthread_mutexattr_destroy Destroys a mutex attribute object
pthread_mutexattr_getprioceiling Obtains the priority ceiling of a mutex attribute object
pthread_mutexattr_getprotocol Obtains protocol of a mutex attribute object
pthread_mutexattr_getpshared Obtains a process-shared setting of a mutex attribute object
pthread_mutexattr_init Initializes a mutex attribute object
pthread_mutexattr_setprioceiling Sets the priority ceiling attribute of a mutex attribute object
pthread_mutexattr_setprotocol Sets the protocol attribute of a mutex attribute object
pthread_mutexattr_setpshared Sets the process-shared attribute of a mutex attribute object to either PTHREAD_PROCESS_SHARED or PTHREAD_PROCESS_PRIVATE
pthread_mutex_destroy Destroys a mutex
pthread_mutex_init Initializes a mutex with the attributes specified in the specified mutex attribute object
pthread_mutex_lock Locks an unlocked mutex
pthread_mutex_trylock Tries to lock a not tested
pthread_mutex_unlock Unlocks a mutex
pthread_once Ensures that init_routine will run just once regardless of how many threads in the process call it
pthread_self Obtains a thread handle of a calling thread
pthread_setcancelstate Sets a thread's cancelability state
pthread_setcanceltype Sets a thread's cancelability type
pthread_setschedparam Adjusts the scheduling policy and scheduling parameters of an existing thread
pthread_setspecific Sets the thread-specific data value associated with the specific key in the calling thread
pthread_sigmask Examines or changes the calling thread's signal mask
pthread_testcancel Requests that any pending cancellation request be delivered to the calling thread

11
doc/readme.txt Executable file
View file

@ -0,0 +1,11 @@
============================
Nimrod's documenation system
============================
This folder contains Nimrod's documentation. The documentation
is written in a format called *reStructuredText*, a markup language that reads
like ASCII and can be converted to HTML, Tex and other formats automatically!
Unfortunately reStructuredText does not allow to colorize source code in the
HTML page. Therefore a postprocessor runs over the generated HTML code, looking
for Nimrod code fragments and colorizing them.

296
doc/regexprs.txt Executable file
View file

@ -0,0 +1,296 @@
Licence of the PCRE library
===========================
PCRE is a library of functions to support regular expressions whose
syntax and semantics are as close as possible to those of the Perl 5
language.
| Written by Philip Hazel
| Copyright (c) 1997-2005 University of Cambridge
----------------------------------------------------------------------
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
* Redistributions of source code must retain the above copyright notice,
this list of conditions and the following disclaimer.
* Redistributions in binary form must reproduce the above copyright
notice, this list of conditions and the following disclaimer in the
documentation and/or other materials provided with the distribution.
* Neither the name of the University of Cambridge nor the names of its
contributors may be used to endorse or promote products derived from
this software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE
LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR
CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF
SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS
INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN
CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
POSSIBILITY OF SUCH DAMAGE.
Regular expression syntax and semantics
=======================================
As the regular expressions supported by this module are enormous,
the reader is referred to http://perldoc.perl.org/perlre.html for the
full documentation of Perl's regular expressions.
Because the backslash ``\`` is a meta character both in the Nimrod
programming language and in regular expressions, it is strongly
recommended that one uses the *raw* strings of Nimrod, so that
backslashes are interpreted by the regular expression engine::
r"\S" # matches any character that is not whitespace
A regular expression is a pattern that is matched against a subject string
from left to right. Most characters stand for themselves in a pattern, and
match the corresponding characters in the subject. As a trivial example,
the pattern::
The quick brown fox
matches a portion of a subject string that is identical to itself.
The power of regular expressions comes from the ability to include
alternatives and repetitions in the pattern. These are encoded in
the pattern by the use of metacharacters, which do not stand for
themselves but instead are interpreted in some special way.
There are two different sets of metacharacters: those that are recognized
anywhere in the pattern except within square brackets, and those that are
recognized in square brackets. Outside square brackets, the metacharacters
are as follows:
============== ============================================================
meta character meaning
============== ============================================================
``\`` general escape character with several uses
``^`` assert start of string (or line, in multiline mode)
``$`` assert end of string (or line, in multiline mode)
``.`` match any character except newline (by default)
``[`` start character class definition
``|`` start of alternative branch
``(`` start subpattern
``)`` end subpattern
``?`` extends the meaning of ``(``
also 0 or 1 quantifier
also quantifier minimizer
``*`` 0 or more quantifier
``+`` 1 or more quantifier
also "possessive quantifier"
``{`` start min/max quantifier
============== ============================================================
Part of a pattern that is in square brackets is called a "character class".
In a character class the only metacharacters are:
============== ============================================================
meta character meaning
============== ============================================================
``\`` general escape character
``^`` negate the class, but only if the first character
``-`` indicates character range
``[`` POSIX character class (only if followed by POSIX syntax)
``]`` terminates the character class
============== ============================================================
The following sections describe the use of each of the metacharacters.
Backslash
---------
The `backslash`:idx: character has several uses. Firstly, if it is followed
by a non-alphanumeric character, it takes away any special meaning that
character may have. This use of backslash as an escape character applies
both inside and outside character classes.
For example, if you want to match a ``*`` character, you write ``\*`` in
the pattern. This escaping action applies whether or not the following
character would otherwise be interpreted as a metacharacter, so it is always
safe to precede a non-alphanumeric with backslash to specify that it stands
for itself. In particular, if you want to match a backslash, you write ``\\``.
Non-printing characters
-----------------------
A second use of backslash provides a way of encoding non-printing characters
in patterns in a visible manner. There is no restriction on the appearance of
non-printing characters, apart from the binary zero that terminates a pattern,
but when a pattern is being prepared by text editing, it is usually easier to
use one of the following escape sequences than the binary character it
represents::
============== ============================================================
character meaning
============== ============================================================
``\a`` alarm, that is, the BEL character (hex 07)
``\e`` escape (hex 1B)
``\f`` formfeed (hex 0C)
``\n`` newline (hex 0A)
``\r`` carriage return (hex 0D)
``\t`` tab (hex 09)
``\ddd`` character with octal code ddd, or backreference
``\xhh`` character with hex code hh
============== ============================================================
After ``\x``, from zero to two hexadecimal digits are read (letters can be in
upper or lower case). In UTF-8 mode, any number of hexadecimal digits may
appear between ``\x{`` and ``}``, but the value of the character code must be
less than 2**31 (that is, the maximum hexadecimal value is 7FFFFFFF). If
characters other than hexadecimal digits appear between ``\x{`` and ``}``, or
if there is no terminating ``}``, this form of escape is not recognized.
Instead, the initial ``\x`` will be interpreted as a basic hexadecimal escape,
with no following digits, giving a character whose value is zero.
After ``\0`` up to two further octal digits are read. In both cases, if there
are fewer than two digits, just those that are present are used. Thus the
sequence ``\0\x\07`` specifies two binary zeros followed by a BEL character
(code value 7). Make sure you supply two digits after the initial zero if
the pattern character that follows is itself an octal digit.
The handling of a backslash followed by a digit other than 0 is complicated.
Outside a character class, PCRE reads it and any following digits as a
decimal number. If the number is less than 10, or if there have been at least
that many previous capturing left parentheses in the expression, the entire
sequence is taken as a back reference. A description of how this works is
given later, following the discussion of parenthesized subpatterns.
Inside a character class, or if the decimal number is greater than 9 and
there have not been that many capturing subpatterns, PCRE re-reads up to
three octal digits following the backslash, and generates a single byte
from the least significant 8 bits of the value. Any subsequent digits stand
for themselves. For example:
============== ============================================================
example meaning
============== ============================================================
``\040`` is another way of writing a space
``\40`` is the same, provided there are fewer than 40 previous
capturing subpatterns
``\7`` is always a back reference
``\11`` might be a back reference, or another way of writing a tab
``\011`` is always a tab
``\0113`` is a tab followed by the character "3"
``\113`` might be a back reference, otherwise the character with
octal code 113
``\377`` might be a back reference, otherwise the byte consisting
entirely of 1 bits
``\81`` is either a back reference, or a binary zero followed by
the two characters "8" and "1"
============== ============================================================
Note that octal values of 100 or greater must not be introduced by a leading
zero, because no more than three octal digits are ever read.
All the sequences that define a single byte value or a single UTF-8 character
(in UTF-8 mode) can be used both inside and outside character classes. In
addition, inside a character class, the sequence ``\b`` is interpreted as the
backspace character (hex 08), and the sequence ``\X`` is interpreted as the
character "X". Outside a character class, these sequences have different
meanings (see below).
Generic character types
-----------------------
The third use of backslash is for specifying `generic character types`:idx:.
The following are always recognized:
============== ============================================================
character type meaning
============== ============================================================
``\d`` any decimal digit
``\D`` any character that is not a decimal digit
``\s`` any whitespace character
``\S`` any character that is not a whitespace character
``\w`` any "word" character
``\W`` any "non-word" character
============== ============================================================
Each pair of escape sequences partitions the complete set of characters into
two disjoint sets. Any given character matches one, and only one, of each pair.
These character type sequences can appear both inside and outside character
classes. They each match one character of the appropriate type. If the
current matching point is at the end of the subject string, all of them fail,
since there is no character to match.
For compatibility with Perl, ``\s`` does not match the VT character (code 11).
This makes it different from the the POSIX "space" class. The ``\s`` characters
are HT (9), LF (10), FF (12), CR (13), and space (32).
A "word" character is an underscore or any character less than 256 that is
a letter or digit. The definition of letters and digits is controlled by
PCRE's low-valued character tables, and may vary if locale-specific matching
is taking place (see "Locale support" in the pcreapi page). For example,
in the "fr_FR" (French) locale, some character codes greater than 128 are
used for accented letters, and these are matched by ``\w``.
In UTF-8 mode, characters with values greater than 128 never match ``\d``,
``\s``, or ``\w``, and always match ``\D``, ``\S``, and ``\W``. This is true
even when Unicode character property support is available.
Simple assertions
-----------------
The fourth use of backslash is for certain `simple assertions`:idx:. An
assertion specifies a condition that has to be met at a particular point in
a match, without consuming any characters from the subject string. The use of
subpatterns for more complicated assertions is described below. The
backslashed assertions are::
============== ============================================================
assertion meaning
============== ============================================================
``\b`` matches at a word boundary
``\B`` matches when not at a word boundary
``\A`` matches at start of subject
``\Z`` matches at end of subject or before newline at end
``\z`` matches at end of subject
``\G`` matches at first matching position in subject
============== ============================================================
These assertions may not appear in character classes (but note that ``\b``
has a different meaning, namely the backspace character, inside a character
class).
A word boundary is a position in the subject string where the current
character and the previous character do not both match ``\w`` or ``\W`` (i.e.
one matches ``\w`` and the other matches ``\W``), or the start or end of the
string if the first or last character matches ``\w``, respectively.
The ``\A``, ``\Z``, and ``\z`` assertions differ from the traditional
circumflex and dollar in that they only ever match at the very start and
end of the subject string, whatever options are set.
The difference between ``\Z`` and ``\z`` is that ``\Z`` matches before
a newline that is the last character of the string as well as at the end
of the string, whereas ``\z`` matches only at the end.
..
Regular expressions in Nimrod itself!
-------------------------------------
'a' -- matches the character a
'a'-'z' -- range operator '-'
'A' | 'B' -- alternative operator |
* 'a' -- prefix * is needed
+ 'a' -- prefix + is needed
? 'a' -- prefix ? is needed
letter -- character classes with real names!
letters
white
whites
any -- any character
() -- are Nimrod syntax
! 'a'-'z'
-- concatentation via proc call:
re('A' 'Z' word * )

111
doc/rst.txt Executable file
View file

@ -0,0 +1,111 @@
===========================================================================
Nimrod's implementation of |rst|
===========================================================================
:Author: Andreas Rumpf
:Version: |nimrodversion|
.. contents::
Introduction
============
This document describes the subset of `Docutils`_' `reStructuredText`_ as it
has been implemented in the Nimrod compiler for generating documentation.
Elements of |rst| that are not listed here have not been implemented.
Unfortunately, the specification of |rst| is quite vague, so Nimrod is not as
compatible to the original implementation as one would like.
Even though Nimrod's |rst| parser does not parse all constructs, it is pretty
usable. The missing features can easily be circumvented. An indication of this
fact is that Nimrod's
*whole* documentation itself (including this document) is
processed by Nimrod's |rst| parser. (Which is an order of magnitude faster than
Docutils' parser.)
Inline elements
===============
Ordinary text may contain *inline elements*.
Bullet lists
============
*Bullet lists* look like this::
* Item 1
* Item 2 that
spans over multiple lines
* Item 3
* Item 4
- bullet lists may nest
- item 3b
- valid bullet characters are ``+``, ``*`` and ``-``
This results in:
* Item 1
* Item 2 that
spans over multiple lines
* Item 3
* Item 4
- bullet lists may nest
- item 3b
- valid bullet characters are ``+``, ``*`` and ``-``
Enumerated lists
================
*Enumerated lists*
Defintion lists
===============
Save this code to the file "greeting.nim". Now compile and run it:
``nimrod run greeting.nim``
As you see, with the ``run`` command Nimrod executes the file automatically
after compilation. You can even give your program command line arguments by
appending them after the filename that is to be compiled and run:
``nimrod run greeting.nim arg1 arg2``
Tables
======
Nimrod only implements simple tables of the form::
================== =============== ===================
header 1 header 2 header n
================== =============== ===================
Cell 1 Cell 2 Cell 3
Cell 4 Cell 5; any Cell 6
cell that is
not in column 1
may span over
multiple lines
Cell 7 Cell 8 Cell 9
================== =============== ===================
This results in:
================== =============== ===================
header 1 header 2 header n
================== =============== ===================
Cell 1 Cell 2 Cell 3
Cell 4 Cell 5; any Cell 6
cell that is
not in column 1
may span over
multiple lines
Cell 7 Cell 8 Cell 9
================== =============== ===================
.. |rst| replace:: reStructuredText
.. _reStructuredText: http://docutils.sourceforge.net/rst.html#reference-documentation
.. _docutils: http://docutils.sourceforge.net/

1297
doc/spec.txt Executable file

File diff suppressed because it is too large Load diff

1436
doc/theindex.txt Executable file

File diff suppressed because it is too large Load diff

215
doc/tutorial.txt Executable file
View file

@ -0,0 +1,215 @@
===========================================
Tutorial of the Nimrod Programming Language
===========================================
:Author: Andreas Rumpf
Motivation
==========
Why yet another programming language?
Look at the trends behind all the new programming languages:
* They try to be dynamic: Dynamic typing, dynamic method binding, etc.
In my opinion the most things the dynamic features buy could be achieved
with static means in a more efficient and *understandable* way.
* They depend on big runtime environments which you need to
ship with your program as each new version of these may break compability
in subtle ways or you use recently added features - thus forcing your
users to update their runtime environment. Compiled programs where the
executable contains all needed code are simply the better solution.
* They are unsuitable for systems programming: Do you really want to
write an operating system, a device driver or an interpreter in a language
that is just-in-time compiled (or interpreted)?
So what lacks are *good* systems programming languages. Nimrod is such a
language. It offers the following features:
* It is readable: It reads from left to right (unlike the C-syntax
languages).
* It is strongly and statically typed: This enables the compiler to find
more errors. Static typing also makes programs more *readable*.
* It is compiled. (Currently this is done via compilation to C.)
* It is garbage collected. Big systems need garbage collection. Manuell
memory management is also supported through *untraced pointers*.
* It scales because high level features are also available: It has built-in
bit sets, strings, enumerations, objects, arrays and dynamically resizeable
arrays (called *sequences*).
* It has high performance: The current implementation compiles to C
and uses a Deutsch-Bobrow garbage collector together with Christoper's
partial mark-sweep garbage collector leading to excellent execution
speed and a small memory footprint.
* It has real modules with proper interfaces and supports separate
compilation.
* It is portable: It compiles to C and platform specific features have
been separated and documented. So even if your platform is not supported
porting should be easy.
* It is flexible: Although primilarily a procedural language, generic,
functional and object-oriented programming is also supported.
* It is easy to learn, easy to use and leads to elegant programs.
* You can link an embedded debugger to your program (ENDB). ENDB is
very easy to use - there is no need to clutter your code with
``echo`` statements for proper debugging.
Introduction
============
This document is a tutorial for the programming language *Nimrod*. It should
be a readable quick tour through the language instead of a dry specification
(which can be found `here <manual.html>`_). This tutorial assumes that
the reader already knows some other programming language such as Pascal. Thus
it is detailed in cases where Nimrod differs from other programming languages
and kept short where Nimrod is more or less the same.
A quick tour through the language
=================================
The first program
-----------------
We start the tour with a modified "hallo world" program:
.. code-block:: Nimrod
# This is a comment
# Standard IO-routines are always accessible
write(stdout, "What's your name? ")
var name: string = readLine(stdin)
write(stdout, "Hi, " & name & "!\n")
Save this code to the file "greeting.nim". Now compile and run it::
nimrod run greeting.nim
As you see, with the ``run`` command Nimrod executes the file automatically
after compilation. You can even give your program command line arguments by
appending them after the filename that is to be compiled and run::
nimrod run greeting.nim arg1 arg2
Though it should be pretty obvious what the program does, I will explain the
syntax: Statements which are not indented are executed when the program
starts. Indentation is Nimrod's way of grouping statements. String literals
are enclosed in double quotes. The ``var`` statement declares a new variable
named ``name`` of type ``string`` with the value that is returned by the
``readline`` procedure. Since the compiler knows that ``readline`` returns
a string, you can leave out the type in the declaration. So this will work too:
.. code-block:: Nimrod
var name = readline(stdin)
Note that this is the only form of type inference that exists in Nimrod:
This is because it yields a good compromise between brevity and readability.
The ``&`` operator concates strings together. ``\n`` stands for the
new line character(s). On several operating systems ``\n`` is represented by
*two* characters: Linefeed and Carriage Return. That is why
*character literals* cannot contain ``\n``. But since Nimrod handles strings
so well, this is a nonissue.
The "hallo world" program contains several identifiers that are already
known to the compiler: ``write``, ``stdout``, ``readLine``, etc. These
built-in items are declared in the system_ module which is implicitly
imported by any other module.
Lexical elements
----------------
Let us look into Nimrod's lexical elements in more detail: Like other
programming languages Nimrod consists of identifiers, keywords, comments,
operators, and other punctation marks. Case is *insignificant* in Nimrod and
even underscores are ignored: ``This_is_an_identifier`` and this is the same
identifier ``ThisIsAnIdentifier``. This feature enables one to use other
peoples code without bothering about a naming convention that one does not
like.
String literals are enclosed in double quotes, character literals in single
quotes. There exist also *raw* string and character literals:
.. code-block:: Nimrod
r"C:\program files\nim"
In raw literals the backslash is not an escape character, so they fit
the principle *what you see is what you get*. *Long string literals*
are also available (``""" ... """``); they can span over multiple lines
and the ``\`` is not an escape character either. They are very useful
for embedding SQL code templates for example.
Comments start with ``#`` and run till the end of the line. (Well this is not
quite true, but you should read the manual for a proper explanation.)
... XXX number literals
The usual statements - if, while, for, case
-------------------------------------------
In Nimrod indentation is used to group statements.
An example showing the most common statement types:
.. code-block:: Nimrod
var name = readLine(stdin)
if name == "Andreas":
echo("What a nice name!")
elif name == "":
echo("Don't you have a name?")
else:
echo("Boring name...")
for i in 0..length(name)-1:
if name[i] == 'm':
echo("hey, there is an *m* in your name!")
echo("Please give your password: \n")
var pw = readLine(stdin)
while pw != "12345":
echo("Wrong password! Next try: \n")
pw = readLine(stdin)
echo("""Login complete!
What do you want to do?
delete-everything
restart-computer
go-for-a-walk
""")
case readline(stdin)
of "delete-everything", "restart-computer":
echo("permission denied")
of "go-for-a-walk": echo("please yourself")
else: echo("unknown command")
..
Types
-----
Nimrod has a rich type system. This tutorial only gives a few examples. Read
the `manual <manual.html>`_ for further information:
.. code-block:: Nimrod
type
TMyRecord = record
x, y: int
Procedures
----------
Procedures are subroutines. They are declared in this way:
.. code-block:: Nimrod
proc findSubStr(sub: string,
.. _strutils: strutils.html
.. _system: system.html