RST backtick refactor (all *.rst except manual.rst and rst_examples.rst) (#17258)
Co-authored-by: quantimnot <quantimnot@users.noreply.github.com>
This commit is contained in:
parent
15586c7a7a
commit
83ae70cb54
30 changed files with 1402 additions and 1350 deletions
138
doc/intern.rst
138
doc/intern.rst
|
|
@ -1,3 +1,5 @@
|
|||
.. default-role:: code
|
||||
|
||||
=========================================
|
||||
Internals of the Nim Compiler
|
||||
=========================================
|
||||
|
|
@ -19,19 +21,19 @@ The Nim project's directory structure is:
|
|||
============ ===================================================
|
||||
Path Purpose
|
||||
============ ===================================================
|
||||
``bin`` generated binary files
|
||||
``build`` generated C code for the installation
|
||||
``compiler`` the Nim compiler itself; note that this
|
||||
`bin` generated binary files
|
||||
`build` generated C code for the installation
|
||||
`compiler` the Nim compiler itself; note that this
|
||||
code has been translated from a bootstrapping
|
||||
version written in Pascal, so the code is **not**
|
||||
a poster child of good Nim code
|
||||
``config`` configuration files for Nim
|
||||
``dist`` additional packages for the distribution
|
||||
``doc`` the documentation; it is a bunch of
|
||||
`config` configuration files for Nim
|
||||
`dist` additional packages for the distribution
|
||||
`doc` the documentation; it is a bunch of
|
||||
reStructuredText files
|
||||
``lib`` the Nim library
|
||||
``web`` website of Nim; generated by ``nimweb``
|
||||
from the ``*.txt`` and ``*.nimf`` files
|
||||
`lib` the Nim library
|
||||
`web` website of Nim; generated by `nimweb`
|
||||
from the `*.txt` and `*.nimf` files
|
||||
============ ===================================================
|
||||
|
||||
|
||||
|
|
@ -53,7 +55,7 @@ And for a debug version compatible with GDB::
|
|||
nim c koch.nim
|
||||
./koch boot --debuginfo --linedir:on
|
||||
|
||||
The ``koch`` program is Nim's maintenance script. It is a replacement for
|
||||
The `koch` program is Nim's maintenance script. It is a replacement for
|
||||
make and shell scripting with the advantage that it is much more portable.
|
||||
More information about its options can be found in the `koch <koch.html>`_
|
||||
documentation.
|
||||
|
|
@ -67,8 +69,8 @@ Coding Guidelines
|
|||
* Max line length is 80 characters.
|
||||
* Provide spaces around binary operators if that enhances readability.
|
||||
* Use a space after a colon, but not before it.
|
||||
* [deprecated] Start types with a capital ``T``, unless they are
|
||||
pointers/references which start with ``P``.
|
||||
* [deprecated] Start types with a capital `T`, unless they are
|
||||
pointers/references which start with `P`.
|
||||
|
||||
See also the `API naming design <apis.html>`_ document.
|
||||
|
||||
|
|
@ -81,12 +83,12 @@ portable programming language (within certain limits) and Nim generates
|
|||
C code, porting the code generator is not necessary.
|
||||
|
||||
POSIX-compliant systems on conventional hardware are usually pretty easy to
|
||||
port: Add the platform to ``platform`` (if it is not already listed there),
|
||||
port: Add the platform to `platform` (if it is not already listed there),
|
||||
check that the OS, System modules work and recompile Nim.
|
||||
|
||||
The only case where things aren't as easy is when the garbage
|
||||
collector needs some assembler tweaking to work. The standard
|
||||
version of the GC uses C's ``setjmp`` function to store all registers
|
||||
version of the GC uses C's `setjmp` function to store all registers
|
||||
on the hardware stack. It may be necessary that the new platform needs to
|
||||
replace this generic code by some assembler code.
|
||||
|
||||
|
|
@ -111,7 +113,7 @@ Complex assignments
|
|||
|
||||
We already know the type information as a graph in the compiler.
|
||||
Thus we need to serialize this graph as RTTI for C code generation.
|
||||
Look at the file ``lib/system/hti.nim`` for more information.
|
||||
Look at the file `lib/system/hti.nim` for more information.
|
||||
|
||||
Rebuilding the compiler
|
||||
========================
|
||||
|
|
@ -139,7 +141,7 @@ Debugging the compiler
|
|||
======================
|
||||
|
||||
You can of course use GDB or Visual Studio to debug the
|
||||
compiler (via ``--debuginfo --lineDir:on``). However, there
|
||||
compiler (via `--debuginfo --lineDir:on`). However, there
|
||||
are also lots of procs that aid in debugging:
|
||||
|
||||
|
||||
|
|
@ -170,19 +172,19 @@ These procs may not be imported by a module. You can import them directly for de
|
|||
from renderer import renderTree
|
||||
from msgs import `??`
|
||||
|
||||
To create a new compiler for each run, use ``koch temp``::
|
||||
To create a new compiler for each run, use `koch temp`::
|
||||
|
||||
./koch temp c /tmp/test.nim
|
||||
|
||||
``koch temp`` creates a debug build of the compiler, which is useful
|
||||
`koch temp` creates a debug build of the compiler, which is useful
|
||||
to create stacktraces for compiler debugging. See also
|
||||
`Rebuilding the compiler`_ if you need more control.
|
||||
|
||||
Bisecting for regressions
|
||||
=========================
|
||||
|
||||
``koch temp`` returns 125 as the exit code in case the compiler
|
||||
compilation fails. This exit code tells ``git bisect`` to skip the
|
||||
`koch temp` returns 125 as the exit code in case the compiler
|
||||
compilation fails. This exit code tells `git bisect` to skip the
|
||||
current commit.::
|
||||
|
||||
git bisect start bad-commit good-commit
|
||||
|
|
@ -219,15 +221,15 @@ examples how the AST represents each syntactic structure.
|
|||
How the RTL is compiled
|
||||
=======================
|
||||
|
||||
The ``system`` module contains the part of the RTL which needs support by
|
||||
The `system` module contains the part of the RTL which needs support by
|
||||
compiler magic (and the stuff that needs to be in it because the spec
|
||||
says so). The C code generator generates the C code for it, just like any other
|
||||
module. However, calls to some procedures like ``addInt`` are inserted by
|
||||
the CCG. Therefore the module ``magicsys`` contains a table (``compilerprocs``)
|
||||
with all symbols that are marked as ``compilerproc``. ``compilerprocs`` are
|
||||
needed by the code generator. A ``magic`` proc is not the same as a
|
||||
``compilerproc``: A ``magic`` is a proc that needs compiler magic for its
|
||||
semantic checking, a ``compilerproc`` is a proc that is used by the code
|
||||
module. However, calls to some procedures like `addInt` are inserted by
|
||||
the CCG. Therefore the module `magicsys` contains a table (`compilerprocs`)
|
||||
with all symbols that are marked as `compilerproc`. `compilerprocs` are
|
||||
needed by the code generator. A `magic` proc is not the same as a
|
||||
`compilerproc`: A `magic` is a proc that needs compiler magic for its
|
||||
semantic checking, a `compilerproc` is a proc that is used by the code
|
||||
generator.
|
||||
|
||||
|
||||
|
|
@ -254,7 +256,7 @@ This solves the problem without having to special case the logic
|
|||
that fills the internal seqs which are affected by the pragmas.
|
||||
|
||||
In fact, this describes how the AST should be stored in the database,
|
||||
as a "shallow" tree. Let's assume we compile module ``m`` with the
|
||||
as a "shallow" tree. Let's assume we compile module `m` with the
|
||||
following contents:
|
||||
|
||||
.. code-block:: nim
|
||||
|
|
@ -279,21 +281,21 @@ Conceptually this is the AST we store for the module:
|
|||
static:
|
||||
echo "static"
|
||||
|
||||
The symbol's ``ast`` field is loaded lazily, on demand. This is where most
|
||||
The symbol's `ast` field is loaded lazily, on demand. This is where most
|
||||
savings come from, only the shallow outer AST is reconstructed immediately.
|
||||
|
||||
It is also important that the replay involves the ``import`` statement so
|
||||
It is also important that the replay involves the `import` statement so
|
||||
that dependencies are resolved properly.
|
||||
|
||||
|
||||
Shared global compiletime state
|
||||
-------------------------------
|
||||
|
||||
Nim allows ``.global, compiletime`` variables that can be filled by macro
|
||||
Nim allows `.global, compiletime` variables that can be filled by macro
|
||||
invocations across different modules. This feature breaks modularity in a
|
||||
severe way. Plenty of different solutions have been proposed:
|
||||
|
||||
- Restrict the types of global compiletime variables to ``Set[T]`` or
|
||||
- Restrict the types of global compiletime variables to `Set[T]` or
|
||||
similar unordered, only-growable collections so that we can track
|
||||
the module's write effects to these variables and reapply the changes
|
||||
in a different order.
|
||||
|
|
@ -306,7 +308,7 @@ severe way. Plenty of different solutions have been proposed:
|
|||
Since we adopt the "replay the top level statements" idea, the natural
|
||||
solution to this problem is to emit pseudo top level statements that
|
||||
reflect the mutations done to the global variable. However, this is
|
||||
MUCH harder than it sounds, for example ``squeaknim`` uses this
|
||||
MUCH harder than it sounds, for example `squeaknim` uses this
|
||||
snippet:
|
||||
|
||||
.. code-block:: nim
|
||||
|
|
@ -314,12 +316,12 @@ snippet:
|
|||
"\t^self externalCallFailed\C!\C\C")
|
||||
stCode.add(st & "\C\t\"Generated by NimSqueak\"\C\t" & apicall)
|
||||
|
||||
We can "replay" ``stCode.add`` only if the values of ``st``
|
||||
and ``apicall`` are known. And even then a hash table's ``add`` with its
|
||||
We can "replay" `stCode.add` only if the values of `st`
|
||||
and `apicall` are known. And even then a hash table's `add` with its
|
||||
hashing mechanism is too hard to replay.
|
||||
|
||||
In practice, things are worse still, consider ``someGlobal[i][j].add arg``.
|
||||
We only know the root is ``someGlobal`` but the concrete path to the data
|
||||
In practice, things are worse still, consider `someGlobal[i][j].add arg`.
|
||||
We only know the root is `someGlobal` but the concrete path to the data
|
||||
is unknown as is the value that is added. We could compute a "diff" between
|
||||
the global states and use that to compute a symbol patchset, but this is
|
||||
quite some work, expensive to do at runtime (it would need to run after
|
||||
|
|
@ -342,7 +344,7 @@ an alien API and works with some existing Nimble packages, at least.
|
|||
|
||||
On the other hand, in Nim's future I would like to replace the VM
|
||||
by native code. A diff algorithm wouldn't work for that.
|
||||
Instead the native code would work with an API like ``put``, ``get``:
|
||||
Instead the native code would work with an API like `put`, `get`:
|
||||
|
||||
.. code-block:: nim
|
||||
|
||||
|
|
@ -350,7 +352,7 @@ Instead the native code would work with an API like ``put``, ``get``:
|
|||
proc cacheGet*(key: string): NimNode
|
||||
|
||||
The API should embrace the AST diffing notion: See the
|
||||
module ``macrocache`` for the final details.
|
||||
module `macrocache` for the final details.
|
||||
|
||||
|
||||
|
||||
|
|
@ -382,9 +384,9 @@ too. Type converters fall into this category:
|
|||
if 1:
|
||||
echo "ugly, but should work"
|
||||
|
||||
If in the above example module ``B`` is re-compiled, but ``A`` is not then
|
||||
``B`` needs to be aware of ``toBool`` even though ``toBool`` is not referenced
|
||||
in ``B`` *explicitly*.
|
||||
If in the above example module `B` is re-compiled, but `A` is not then
|
||||
`B` needs to be aware of `toBool` even though `toBool` is not referenced
|
||||
in `B` *explicitly*.
|
||||
|
||||
Both the multi method and the type converter problems are solved by the
|
||||
AST replay implementation.
|
||||
|
|
@ -395,7 +397,7 @@ Generics
|
|||
|
||||
We cache generic instantiations and need to ensure this caching works
|
||||
well with the incremental compilation feature. Since the cache is
|
||||
attached to the ``PSym`` datastructure, it should work without any
|
||||
attached to the `PSym` datastructure, it should work without any
|
||||
special logic.
|
||||
|
||||
|
||||
|
|
@ -405,22 +407,22 @@ Backend issues
|
|||
- Init procs must not be "forgotten" to be called.
|
||||
- Files must not be "forgotten" to be linked.
|
||||
- Method dispatchers are global.
|
||||
- DLL loading via ``dlsym`` is global.
|
||||
- DLL loading via `dlsym` is global.
|
||||
- Emulated thread vars are global.
|
||||
|
||||
However the biggest problem is that dead code elimination breaks modularity!
|
||||
To see why, consider this scenario: The module ``G`` (for example the huge
|
||||
To see why, consider this scenario: The module `G` (for example the huge
|
||||
Gtk2 module...) is compiled with dead code elimination turned on. So none
|
||||
of ``G``'s procs is generated at all.
|
||||
of `G`'s procs is generated at all.
|
||||
|
||||
Then module ``B`` is compiled that requires ``G.P1``. Ok, no problem,
|
||||
``G.P1`` is loaded from the symbol file and ``G.c`` now contains ``G.P1``.
|
||||
Then module `B` is compiled that requires `G.P1`. Ok, no problem,
|
||||
`G.P1` is loaded from the symbol file and `G.c` now contains `G.P1`.
|
||||
|
||||
Then module ``A`` (that depends on ``B`` and ``G``) is compiled and ``B``
|
||||
and ``G`` are left unchanged. ``A`` requires ``G.P2``.
|
||||
Then module `A` (that depends on `B` and `G`) is compiled and `B`
|
||||
and `G` are left unchanged. `A` requires `G.P2`.
|
||||
|
||||
So now ``G.c`` MUST contain both ``P1`` and ``P2``, but we haven't even
|
||||
loaded ``P1`` from the symbol file, nor do we want to because we then quickly
|
||||
So now `G.c` MUST contain both `P1` and `P2`, but we haven't even
|
||||
loaded `P1` from the symbol file, nor do we want to because we then quickly
|
||||
would restore large parts of the whole program.
|
||||
|
||||
|
||||
|
|
@ -428,7 +430,7 @@ Solution
|
|||
~~~~~~~~
|
||||
|
||||
The backend must have some logic so that if the currently processed module
|
||||
is from the compilation cache, the ``ast`` field is not accessed. Instead
|
||||
is from the compilation cache, the `ast` field is not accessed. Instead
|
||||
the generated C(++) for the symbol's body needs to be cached too and
|
||||
inserted back into the produced C file. This approach seems to deal with
|
||||
all the outlined problems above.
|
||||
|
|
@ -444,8 +446,8 @@ in mind:
|
|||
keeps allocating memory! Thus a stack overflow may happen, hiding the
|
||||
real issue.
|
||||
* What seem to be C code generation problems is often a bug resulting from
|
||||
not producing prototypes, so that some types default to ``cint``. Testing
|
||||
without the ``-w`` option helps!
|
||||
not producing prototypes, so that some types default to `cint`. Testing
|
||||
without the `-w` option helps!
|
||||
|
||||
|
||||
The Garbage Collector
|
||||
|
|
@ -464,9 +466,9 @@ code generation.
|
|||
|
||||
Each cell has a header consisting of a RC and a pointer to its type
|
||||
descriptor. However the program does not know about these, so they are placed at
|
||||
negative offsets. In the GC code the type ``PCell`` denotes a pointer
|
||||
negative offsets. In the GC code the type `PCell` denotes a pointer
|
||||
decremented by the right offset, so that the header can be accessed easily. It
|
||||
is extremely important that ``pointer`` is not confused with a ``PCell``
|
||||
is extremely important that `pointer` is not confused with a `PCell`
|
||||
as this would lead to a memory corruption.
|
||||
|
||||
|
||||
|
|
@ -474,9 +476,9 @@ The CellSet data structure
|
|||
--------------------------
|
||||
|
||||
The GC depends on an extremely efficient datastructure for storing a
|
||||
set of pointers - this is called a ``TCellSet`` in the source code.
|
||||
set of pointers - this is called a `TCellSet` in the source code.
|
||||
Inserting, deleting and searching are done in constant time. However,
|
||||
modifying a ``TCellSet`` during traversal leads to undefined behaviour.
|
||||
modifying a `TCellSet` during traversal leads to undefined behaviour.
|
||||
|
||||
.. code-block:: Nim
|
||||
type
|
||||
|
|
@ -559,11 +561,11 @@ Code generation for closures is implemented by `lambda lifting`:idx:.
|
|||
Design
|
||||
------
|
||||
|
||||
A ``closure`` proc var can call ordinary procs of the default Nim calling
|
||||
A `closure` proc var can call ordinary procs of the default Nim calling
|
||||
convention. But not the other way round! A closure is implemented as a
|
||||
``tuple[prc, env]``. ``env`` can be nil implying a call without a closure.
|
||||
This means that a call through a closure generates an ``if`` but the
|
||||
interoperability is worth the cost of the ``if``. Thunk generation would be
|
||||
`tuple[prc, env]`. `env` can be nil implying a call without a closure.
|
||||
This means that a call through a closure generates an `if` but the
|
||||
interoperability is worth the cost of the `if`. Thunk generation would be
|
||||
possible too, but it's slightly more effort to implement.
|
||||
|
||||
Tests with GCC on Amd64 showed that it's really beneficial if the
|
||||
|
|
@ -579,7 +581,7 @@ A thunk would need to call 'returnsDefaultCC[i]' somehow and that would require
|
|||
an *additional* closure generation... Ok, not really, but it requires to pass
|
||||
the function to call. So we'd end up with 2 indirect calls instead of one.
|
||||
Another much more severe problem which this solution is that it's not GC-safe
|
||||
to pass a proc pointer around via a generic ``ref`` type.
|
||||
to pass a proc pointer around via a generic `ref` type.
|
||||
|
||||
|
||||
Example code:
|
||||
|
|
@ -695,15 +697,15 @@ Accumulator
|
|||
Internals
|
||||
---------
|
||||
|
||||
Lambda lifting is implemented as part of the ``transf`` pass. The ``transf``
|
||||
Lambda lifting is implemented as part of the `transf` pass. The `transf`
|
||||
pass generates code to setup the environment and to pass it around. However,
|
||||
this pass does not change the types! So we have some kind of mismatch here; on
|
||||
the one hand the proc expression becomes an explicit tuple, on the other hand
|
||||
the tyProc(ccClosure) type is not changed. For C code generation it's also
|
||||
important the hidden formal param is ``void*`` and not something more
|
||||
important the hidden formal param is `void*` and not something more
|
||||
specialized. However the more specialized env type needs to passed to the
|
||||
backend somehow. We deal with this by modifying ``s.ast[paramPos]`` to contain
|
||||
the formal hidden parameter, but not ``s.typ``!
|
||||
backend somehow. We deal with this by modifying `s.ast[paramPos]` to contain
|
||||
the formal hidden parameter, but not `s.typ`!
|
||||
|
||||
|
||||
Integer literals:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue