366 lines
12 KiB
Text
366 lines
12 KiB
Text
=========================================
|
|
Internals of the Nimrod Compiler
|
|
=========================================
|
|
|
|
|
|
:Author: Andreas Rumpf
|
|
:Version: |nimrodversion|
|
|
|
|
.. contents::
|
|
|
|
|
|
Directory structure
|
|
===================
|
|
|
|
The Nimrod project's directory structure is:
|
|
|
|
============ ==============================================
|
|
Path Purpose
|
|
============ ==============================================
|
|
``bin`` binary files go into here
|
|
``nim`` Pascal sources of the Nimrod compiler; this
|
|
should be modified, not the Nimrod version in
|
|
``rod``!
|
|
``rod`` Nimrod sources of the Nimrod compiler;
|
|
automatically generated from the Pascal
|
|
version
|
|
``data`` data files that are used for generating source
|
|
code go into here
|
|
``doc`` the documentation lives here; it is a bunch of
|
|
reStructuredText files
|
|
``dist`` download packages as zip archives go into here
|
|
``config`` configuration files for Nimrod go into here
|
|
``lib`` the Nimrod library lives here; ``rod`` depends
|
|
on it!
|
|
``web`` website of Nimrod; generated by ``koch.py``
|
|
from the ``*.txt`` and ``*.tmpl`` files
|
|
``koch`` the Koch Build System (written for Nimrod)
|
|
``obj`` generated ``*.obj`` files go into here
|
|
============ ==============================================
|
|
|
|
|
|
Bootstrapping the compiler
|
|
==========================
|
|
|
|
The compiler is written in a subset of Pascal with special annotations so
|
|
that it can be translated to Nimrod code automatically. This conversion is
|
|
done by Nimrod itself via the undocumented ``boot`` command. Thus both Nimrod
|
|
and Free Pascal can compile the Nimrod compiler.
|
|
|
|
Requirements for bootstrapping:
|
|
|
|
- Free Pascal (I used version 2.2) [optional]
|
|
- Python (should work with version 1.5 or higher)
|
|
|
|
- C compiler -- one of:
|
|
|
|
* win32-lcc (currently broken)
|
|
* Borland C++ (tested with 5.5; currently broken)
|
|
* Microsoft C++
|
|
* Digital Mars C++
|
|
* Watcom C++ (currently broken)
|
|
* GCC
|
|
* Intel C++
|
|
* Pelles C (currently broken)
|
|
* llvm-gcc
|
|
|
|
| Compiling the compiler is a simple matter of running:
|
|
| ``koch.py boot``
|
|
| Or you can compile by hand, this is not difficult.
|
|
|
|
If you want to debug the compiler, use the command::
|
|
|
|
koch.py boot --debugger:on
|
|
|
|
The ``koch.py`` script is Nimrod's maintainance script: Everything that has
|
|
been automated is accessible with it. It is a replacement for make and shell
|
|
scripting with the advantage that it is more portable.
|
|
|
|
|
|
Coding standards
|
|
================
|
|
|
|
The compiler is written in a subset of Pascal with special annotations so
|
|
that it can be translated to Nimrod code automatically. As a general rule,
|
|
Pascal code that does not translate to Nimrod automatically is forbidden.
|
|
|
|
|
|
Porting to new platforms
|
|
========================
|
|
|
|
Porting Nimrod to a new architecture is pretty easy, since C is the most
|
|
portable programming language (within certain limits) and Nimrod generates
|
|
C code, porting the code generator is not necessary.
|
|
|
|
POSIX-compliant systems on conventional hardware are usually pretty easy to
|
|
port: Add the platform to ``platform`` (if it is not already listed there),
|
|
check that the OS, System modules work and recompile Nimrod.
|
|
|
|
The only case where things aren't as easy is when the garbage
|
|
collector needs some assembler tweaking to work. The standard
|
|
version of the GC uses C's ``setjmp`` function to store all registers
|
|
on the hardware stack. It may be that the new platform needs to
|
|
replace this generic code by some assembler code.
|
|
|
|
|
|
Runtime type information
|
|
========================
|
|
|
|
*Runtime type information* (RTTI) is needed for several aspects of the Nimrod
|
|
programming language:
|
|
|
|
Garbage collection
|
|
The most important reason for RTTI. Generating
|
|
traversal procedures produces bigger code and is likely to be slower on
|
|
modern hardware as dynamic procedure binding is hard to predict.
|
|
|
|
Complex assignments
|
|
Sequences and strings are implemented as
|
|
pointers to resizeable buffers, but Nimrod requires copying for
|
|
assignments. Apart from RTTI the compiler could generate copy procedures
|
|
for any type that needs one. However, this would make the code bigger and
|
|
the RTTI is likely already there for the GC.
|
|
|
|
We already knew the type information as a graph in the compiler.
|
|
Thus we need to serialize this graph as RTTI for C code generation.
|
|
Look at the file ``lib/hti.nim`` for more information.
|
|
|
|
|
|
The Garbage Collector
|
|
=====================
|
|
|
|
Introduction
|
|
------------
|
|
|
|
We use the term *cell* here to refer to everything that is traced
|
|
(sequences, refs, strings).
|
|
This section describes how the new GC works.
|
|
|
|
The basic algorithm is *Deferrent reference counting* with cycle detection.
|
|
References in the stack are not counted for better performance and easier C
|
|
code generation.
|
|
|
|
Each cell has a header consisting of a RC and a pointer to its type
|
|
descriptor. However the program does not know about these, so they are placed at
|
|
negative offsets. In the GC code the type ``PCell`` denotes a pointer
|
|
decremented by the right offset, so that the header can be accessed easily. It
|
|
is extremely important that ``pointer`` is not confused with a ``PCell``
|
|
as this would lead to a memory corruption.
|
|
|
|
|
|
The CellSet data structure
|
|
--------------------------
|
|
|
|
The GC depends on an extremely efficient datastructure for storing a
|
|
set of pointers - this is called a ``TCellSet`` in the source code.
|
|
Inserting, deleting and searching are done in constant time. However,
|
|
modifying a ``TCellSet`` during traversation leads to undefined behaviour.
|
|
|
|
.. code-block:: Nimrod
|
|
type
|
|
TCellSet # hidden
|
|
|
|
proc CellSetInit(s: var TCellSet) # initialize a new set
|
|
proc CellSetDeinit(s: var TCellSet) # empty the set and free its memory
|
|
proc incl(s: var TCellSet, elem: PCell) # include an element
|
|
proc excl(s: var TCellSet, elem: PCell) # exclude an element
|
|
|
|
proc `in`(elem: PCell, s: TCellSet): bool # tests membership
|
|
|
|
iterator elements(s: TCellSet): (elem: PCell)
|
|
|
|
|
|
All the operations have to be perform efficiently. Because a Cellset can
|
|
become huge a hash table alone is not suitable for this.
|
|
|
|
We use a mixture of bitset and hash table for this. The hash table maps *pages*
|
|
to a page descriptor. The page descriptor contains a bit for any possible cell
|
|
address within this page. So including a cell is done as follows:
|
|
|
|
- Find the page descriptor for the page the cell belongs to.
|
|
- Set the appropriate bit in the page descriptor indicating that the
|
|
cell points to the start of a memory block.
|
|
|
|
Removing a cell is analogous - the bit has to be set to zero.
|
|
Single page descriptors are never deleted from the hash table. This is not
|
|
needed as the data structures need to be periodically rebuilt anyway.
|
|
|
|
Complete traversal is done in this way::
|
|
|
|
for each page decriptor d:
|
|
for each bit in d:
|
|
if bit == 1:
|
|
traverse the pointer belonging to this bit
|
|
|
|
|
|
Further complications
|
|
---------------------
|
|
|
|
In Nimrod the compiler cannot always know if a reference
|
|
is stored on the stack or not. This is caused by var parameters.
|
|
Consider this example:
|
|
|
|
.. code-block:: Nimrod
|
|
proc setRef(r: var ref TNode) =
|
|
new(r)
|
|
|
|
proc usage =
|
|
var
|
|
r: ref TNode
|
|
setRef(r) # here we should not update the reference counts, because
|
|
# r is on the stack
|
|
setRef(r.left) # here we should update the refcounts!
|
|
|
|
Though it would be possible to produce code updating the refcounts (if
|
|
necessary) before and after the call to ``setRef``, it is a complex task to
|
|
do so in the code generator. So we don't and instead decide at runtime
|
|
whether the reference is on the stack or not. The generated code looks
|
|
roughly like this:
|
|
|
|
.. code-block:: C
|
|
void setref(TNode** ref) {
|
|
unsureAsgnRef(ref, newObj(TNode_TI, sizeof(TNode)))
|
|
}
|
|
void usage(void) {
|
|
setRef(&r)
|
|
setRef(&r->left)
|
|
}
|
|
|
|
Note that for systems with a continous stack (which most systems have)
|
|
the check whether the ref is on the stack is very cheap (only two
|
|
comparisons). Another advantage of this scheme is that the code produced is
|
|
smaller.
|
|
|
|
|
|
The algorithm in pseudo-code
|
|
----------------------------
|
|
To be written.
|
|
|
|
|
|
The compiler's architecture
|
|
===========================
|
|
|
|
Nimrod uses the classic compiler architecture: A scanner feds tokens to a
|
|
parser. The parser builds a syntax tree that is used by the code generator.
|
|
This syntax tree is the interface between the parser and the code generator.
|
|
It is essential to understand most of the compiler's code.
|
|
|
|
In order to compile Nimrod correctly, type-checking has to be seperated from
|
|
parsing. Otherwise generics would not work. Code generation is done for a
|
|
whole module only after it has been checked for semantics.
|
|
|
|
.. include:: filelist.txt
|
|
|
|
The first command line argument selects the backend. Thus the backend is
|
|
responsible for calling the parser and semantic checker. However, when
|
|
compiling ``import`` or ``include`` statements, the semantic checker needs to
|
|
call the backend, this is done by embedding a PBackend into a TContext.
|
|
|
|
|
|
The syntax tree
|
|
---------------
|
|
The synax tree consists of nodes which may have an arbitrary number of
|
|
children. Types and symbols are represented by other nodes, because they
|
|
may contain cycles. The AST changes its shape after semantic checking. This
|
|
is needed to make life easier for the code generators. See the "ast" module
|
|
for the type definitions.
|
|
|
|
We use the notation ``nodeKind(fields, [sons])`` for describing
|
|
nodes. ``nodeKind[sons]`` is a short-cut for ``nodeKind([sons])``.
|
|
XXX: Description of the language's syntax and the corresponding trees.
|
|
|
|
|
|
How the RTL is compiled
|
|
=======================
|
|
|
|
The system module contains the part of the RTL which needs support by
|
|
compiler magic (and the stuff that needs to be in it because the spec
|
|
says so). The C code generator generates the C code for it just like any other
|
|
module. However, calls to some procedures like ``addInt`` are inserted by
|
|
the CCG. Therefore the module ``magicsys`` contains a table
|
|
(``compilerprocs``) with all symbols that are marked as ``compilerproc``.
|
|
|
|
|
|
|
|
Generation of dynamic link libraries
|
|
====================================
|
|
|
|
Generation of dynamic link libraries or shared libraries is not difficult; the
|
|
underlying C compiler already does all the hard work for us. The problem is the
|
|
common runtime library, especially the memory manager. Note that Borland's
|
|
Delphi had exactly the same problem. The workaround is to not link the GC with
|
|
the Dll and provide an extra runtime dll that needs to be initialized.
|
|
|
|
|
|
|
|
How to implement closures
|
|
=========================
|
|
|
|
A closure is a record of a proc pointer and a context ref. The context ref
|
|
points to a garbage collected record that contains the needed variables.
|
|
An example:
|
|
|
|
.. code-block:: Nimrod
|
|
|
|
type
|
|
TListRec = record
|
|
data: string
|
|
next: ref TListRec
|
|
|
|
proc forEach(head: ref TListRec, visitor: proc (s: string) {.closure.}) =
|
|
var it = head
|
|
while it != nil:
|
|
visit(it.data)
|
|
it = it.next
|
|
|
|
proc sayHello() =
|
|
var L = new List(["hallo", "Andreas"])
|
|
var temp = "jup\xff"
|
|
forEach(L, lambda(s: string) =
|
|
io.write(temp)
|
|
io.write(s)
|
|
)
|
|
|
|
|
|
This should become the following in C:
|
|
|
|
.. code-block:: C
|
|
typedef struct ... /* List type */
|
|
|
|
typedef struct closure {
|
|
void (*PrcPart)(string, void*);
|
|
void* ClPart;
|
|
}
|
|
|
|
typedef struct Tcl_data {
|
|
string temp; // all accessed variables are put in here!
|
|
}
|
|
|
|
void forEach(TListRec* head, const closure visitor) {
|
|
TListRec* it = head;
|
|
while (it != NIM_NULL) {
|
|
visitor.prc(it->data, visitor->cl_data);
|
|
it = it->next;
|
|
}
|
|
}
|
|
|
|
void printStr(string s, void* cl_data) {
|
|
Tcl_data* x = (Tcl_data*) cl_data;
|
|
io_write(x->temp);
|
|
io_write(s);
|
|
}
|
|
|
|
void sayhello() {
|
|
Tcl_data* data = new(...);
|
|
asgnRef(&data->temp, "jup\xff");
|
|
...
|
|
|
|
closure cl;
|
|
cl.prc = printStr;
|
|
cl.cl_data = data;
|
|
foreach(L, cl);
|
|
}
|
|
|
|
|
|
What about nested closure? - There's not much difference: Just put all used
|
|
variables in the data record.
|