LLVM tutorial ported (Max Shawabkeh) (Issue #33)
git-svn-id: http://llvm-py.googlecode.com/svn/trunk@94 8d1e9007-1d4e-0410-b67e-1979fd6579aa
This commit is contained in:
parent
8d4addac08
commit
60422c6057
27 changed files with 18381 additions and 102 deletions
|
|
@ -1,4 +1,9 @@
|
|||
|
||||
0.7, in progress:
|
||||
|
||||
* LLVM tutorial ported (Max Shawabkeh) (Issue #33).
|
||||
|
||||
|
||||
0.6, 31-Aug-2010:
|
||||
|
||||
* Add and remove function attributes (Krzysztof Goj) (Issue #21).
|
||||
|
|
|
|||
6
setup.py
6
setup.py
|
|
@ -32,7 +32,7 @@
|
|||
import sys, os
|
||||
from distutils.core import setup, Extension
|
||||
|
||||
LLVM_PY_VERSION = '0.6'
|
||||
LLVM_PY_VERSION = '0.7'
|
||||
|
||||
|
||||
def _run(cmd):
|
||||
|
|
@ -106,8 +106,8 @@ def call_setup(llvm_config):
|
|||
version=LLVM_PY_VERSION,
|
||||
description='Python Bindings for LLVM',
|
||||
author='Mahadevan R',
|
||||
author_email='mdevan.foobar@gmail.com',
|
||||
url='http://mdevan.nfshost.com/llvm-py/',
|
||||
author_email='mdevan@mdevan.org',
|
||||
url='http://www.mdevan.org/llvm-py/',
|
||||
packages=['llvm'],
|
||||
py_modules = [ 'llvm.core' ],
|
||||
ext_modules = [ ext_core ],)
|
||||
|
|
|
|||
|
|
@ -9,27 +9,25 @@ include::example.inc[]
|
|||
LLVM Tutorials
|
||||
--------------
|
||||
|
||||
The http://www.llvm.org/docs/tutorial/[LLVM tutorials] have been
|
||||
ported to llvm-py. Below are the links to the original LLVM tutorial and
|
||||
the corresponding Python code using llvm-py:
|
||||
|
||||
.Simple JIT Tutorials
|
||||
(contributed by Sebastien Binet)
|
||||
|
||||
1. A First Function
|
||||
http://www.llvm.org/docs/tutorial/JITTutorial1.html[LLVM]
|
||||
link:examples/JITTutorial1.html[llvm-py]
|
||||
2. A More Complicated Function
|
||||
http://www.llvm.org/docs/tutorial/JITTutorial2.html[LLVM]
|
||||
link:examples/JITTutorial2.html[llvm-py]
|
||||
The following JIT tutorials were contributed by Sebastien Binet.
|
||||
|
||||
1. link:examples/JITTutorial1.html[A First Function]
|
||||
2. link:examples/JITTutorial2.html[A More Complicated Function]
|
||||
|
||||
|
||||
[[kaleidoscope]]
|
||||
.Kaleidoscope: Implementing a Language with LLVM
|
||||
1. Tutorial Introduction and the Lexer (TODO)
|
||||
2. Implementing a Parser and AST (TODO)
|
||||
3. Implementing Code Generation to LLVM IR (TODO)
|
||||
4. Adding JIT and Optimizer Support (TODO)
|
||||
5. Extending the language: control flow (TODO)
|
||||
6. Extending the language: user-defined operators (TODO)
|
||||
7. Extending the language: mutable variables / SSA construction (TODO)
|
||||
8. Conclusion and other useful LLVM tidbits (TODO)
|
||||
|
||||
The LLVM http://www.llvm.org/docs/tutorial/[Kaleidoscope] tutorial
|
||||
has been ported to llvm-py by Max Shawabkeh.
|
||||
|
||||
1. link:kaleidoscope/PythonLangImpl1.html[Tutorial Introduction and the Lexer]
|
||||
2. link:kaleidoscope/PythonLangImpl2.html[Implementing a Parser and AST]
|
||||
3. link:kaleidoscope/PythonLangImpl3.html[Implementing Code Generation to LLVM IR]
|
||||
4. link:kaleidoscope/PythonLangImpl4.html[Adding JIT and Optimizer Support]
|
||||
5. link:kaleidoscope/PythonLangImpl5.html[Extending the language: control flow]
|
||||
6. link:kaleidoscope/PythonLangImpl6.html[Extending the language: user-defined operators]
|
||||
7. link:kaleidoscope/PythonLangImpl7.html[Extending the language: mutable variables / SSA construction]
|
||||
8. link:kaleidoscope/PythonLangImpl8.html[Conclusion and other useful LLVM tidbits]
|
||||
|
|
|
|||
|
|
@ -18,6 +18,9 @@ a patch.
|
|||
News
|
||||
----
|
||||
|
||||
26-Sep-2010::
|
||||
LLVM tutorial link:examples.html#kaleidoscope[ported] by Max Shawabkeh!
|
||||
|
||||
31-Aug-2010::
|
||||
0.6 released, works with LLVM 2.7.
|
||||
|
||||
|
|
|
|||
389
www/src/kaleidoscope/PythonLangImpl1.html
Normal file
389
www/src/kaleidoscope/PythonLangImpl1.html
Normal file
|
|
@ -0,0 +1,389 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Tutorial Introduction and the Lexer</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<meta name="author" content="Max Shawabkeh">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Tutorial Introduction and the Lexer</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 1
|
||||
<ol>
|
||||
<li><a href="#intro">Tutorial Introduction</a></li>
|
||||
<li><a href="#language">The Basic Language</a></li>
|
||||
<li><a href="#lexer">The Lexer</a></li>
|
||||
</ol>
|
||||
</li>
|
||||
<li><a href="PythonLangImpl2.html">Chapter 2</a>: Implementing a Parser and
|
||||
AST</li>
|
||||
</ul>
|
||||
|
||||
<div class="doc_author">
|
||||
<p>
|
||||
Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
|
||||
and <a href="http://max99x.com">Max Shawabkeh</a>
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="intro">Tutorial Introduction</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Welcome to the "Implementing a language with LLVM" tutorial. This tutorial
|
||||
runs through the implementation of a simple language, showing how fun and
|
||||
easy it can be. This tutorial will get you up and started as well as help to
|
||||
build a framework you can extend to other languages. The code in this tutorial
|
||||
can also be used as a playground to hack on other LLVM specific things.
|
||||
</p>
|
||||
|
||||
<p>The goal of this tutorial is to progressively unveil our language, describing
|
||||
how it is built up over time. This will let us cover a fairly broad range of
|
||||
language design and LLVM-specific usage issues, showing and explaining the code
|
||||
for it all along the way, without overwhelming you with tons of details up
|
||||
front.</p>
|
||||
|
||||
<p>It is useful to point out ahead of time that this tutorial is really about
|
||||
teaching compiler techniques and LLVM specifically, <em>not</em> about teaching
|
||||
modern and sane software engineering principles. In practice, this means that
|
||||
we'll take a number of shortcuts to simplify the exposition. If you dig in and
|
||||
use the code as a basis for future projects, fixing its deficiencies shouldn't
|
||||
be hard.</p>
|
||||
|
||||
<p>We've tried to put this tutorial together in a way that makes chapters easy
|
||||
to skip over if you are already familiar with or are uninterested in the various
|
||||
pieces. The structure of the tutorial is:</p>
|
||||
|
||||
<ul>
|
||||
<li><b><a href="#language">Chapter #1</a>: Introduction to the Kaleidoscope
|
||||
language, and the definition of its Lexer</b> - This shows where we are going
|
||||
and the basic functionality that we want it to do. In order to make this
|
||||
tutorial maximally understandable and hackable, we choose to implement
|
||||
everything in Python instead of using lexer and parser generators. LLVM
|
||||
obviously works just fine with such tools, feel free to use one if you prefer.
|
||||
</li>
|
||||
<li><b><a href="PythonLangImpl2.html">Chapter #2</a>: Implementing a Parser and
|
||||
AST</b> - With the lexer in place, we can talk about parsing techniques and
|
||||
basic AST construction. This tutorial describes recursive descent parsing and
|
||||
operator precedence parsing. Nothing in Chapters 1 or 2 is LLVM-specific,
|
||||
the code doesn't even import the LLVM modules at this point. :)</li>
|
||||
<li><b><a href="PythonLangImpl3.html">Chapter #3</a>: Code generation to LLVM
|
||||
IR</b> - With the AST ready, we can show off how easy generation of LLVM IR
|
||||
really is.</li>
|
||||
<li><b><a href="PythonLangImpl4.html">Chapter #4</a>: Adding JIT and Optimizer
|
||||
Support</b> - Because a lot of people are interested in using LLVM as a JIT,
|
||||
we'll dive right into it and show you the 3 lines it takes to add JIT support.
|
||||
LLVM is also useful in many other ways, but this is one simple and "sexy" way
|
||||
to shows off its power. :)</li>
|
||||
<li><b><a href="PythonLangImpl5.html">Chapter #5</a>: Extending the Language:
|
||||
Control Flow</b> - With the language up and running, we show how to extend it
|
||||
with control flow operations (if/then/else and a 'for' loop). This gives us a
|
||||
chance to talk about simple SSA construction and control flow.</li>
|
||||
<li><b><a href="PythonLangImpl6.html">Chapter #6</a>: Extending the Language:
|
||||
User-defined Operators</b> - This is a silly but fun chapter that talks about
|
||||
extending the language to let the user program define their own arbitrary
|
||||
unary and binary operators (with assignable precedence!). This lets us build a
|
||||
significant piece of the "language" as library routines.</li>
|
||||
<li><b><a href="PythonLangImpl7.html">Chapter #7</a>: Extending the Language:
|
||||
Mutable Variables</b> - This chapter talks about adding user-defined local
|
||||
variables along with an assignment operator. The interesting part about this
|
||||
is how easy and trivial it is to construct SSA form in LLVM: no, LLVM does
|
||||
<em>not</em> require your front-end to construct SSA form!</li>
|
||||
<li><b><a href="PythonLangImpl8.html">Chapter #8</a>: Conclusion and other
|
||||
useful LLVM tidbits</b> - This chapter wraps up the series by talking about
|
||||
potential ways to extend the language, but also includes a bunch of pointers to
|
||||
info about "special topics" like adding garbage collection support, exceptions,
|
||||
debugging, support for "spaghetti stacks", and a bunch of other tips and
|
||||
tricks.</li>
|
||||
|
||||
</ul>
|
||||
|
||||
<p>By the end of the tutorial, we'll have written a bit less than 540 lines of
|
||||
non-comment, non-blank, lines of code. With this small amount of code, we'll
|
||||
have built up a very reasonable compiler for a non-trivial language including
|
||||
a hand-written lexer, parser, AST, as well as code generation support with a JIT
|
||||
compiler. While other systems may have interesting "hello world" tutorials,
|
||||
I think the breadth of this tutorial is a great testament to the strengths of
|
||||
LLVM and why you should consider it if you're interested in language or compiler
|
||||
design.</p>
|
||||
|
||||
<p>A note about this tutorial: we expect you to extend the language and play
|
||||
with it on your own. Take the code and go crazy hacking away at it, compilers
|
||||
don't need to be scary creatures - it can be a lot of fun to play with
|
||||
languages!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="language">The Basic Language</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>This tutorial will be illustrated with a toy language that we'll call
|
||||
"<a href="http://en.wikipedia.org/wiki/Kaleidoscope">Kaleidoscope</a>" (derived
|
||||
from "meaning beautiful, form, and view").
|
||||
Kaleidoscope is a procedural language that allows you to define functions, use
|
||||
conditionals, math, etc. Over the course of the tutorial, we'll extend
|
||||
Kaleidoscope to support the if/then/else construct, a for loop, user defined
|
||||
operators, JIT compilation with a simple command line interface, etc.</p>
|
||||
|
||||
<p>Because we want to keep things simple, the only datatype in Kaleidoscope is a
|
||||
64-bit floating point type. As such, all values are implicitly double precision
|
||||
and the language doesn't require type declarations. This gives the language a
|
||||
very nice and simple syntax. For example, the following simple example computes
|
||||
<a href="http://en.wikipedia.org/wiki/Fibonacci_number">Fibonacci numbers:</a>
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Compute the x'th fibonacci number.
|
||||
def fib(x)
|
||||
if x < 3 then
|
||||
1
|
||||
else
|
||||
fib(x-1)+fib(x-2)
|
||||
|
||||
# This expression will compute the 40th number.
|
||||
fib(40)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>We also allow Kaleidoscope to call into standard library functions (the LLVM
|
||||
JIT makes this completely trivial). This means that you can use the 'extern'
|
||||
keyword to define a function before you use it (this is also useful for mutually
|
||||
recursive functions). For example:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
extern sin(arg);
|
||||
extern cos(arg);
|
||||
extern atan2(arg1 arg2);
|
||||
|
||||
atan2(sin(0.4), cos(42))
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>A more interesting example is included in Chapter 6 where we write a little
|
||||
Kaleidoscope application that <a href="PythonLangImpl6.html#example">displays
|
||||
a Mandelbrot Set</a> at various levels of magnification.</p>
|
||||
|
||||
<p>Lets dive into the implementation of this language!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="lexer">The Lexer</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>When it comes to implementing a language, the first thing needed is
|
||||
the ability to process a text file and recognize what it says. The traditional
|
||||
way to do this is to use a "<a
|
||||
href="http://en.wikipedia.org/wiki/Lexical_analysis">lexer</a>" (aka 'scanner')
|
||||
to break the input up into "tokens". Each token returned by the lexer includes
|
||||
a token type and potentially some metadata (e.g. the numeric value of a number).
|
||||
First, we define the possibilities:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The lexer yields one of these types for each token.
|
||||
class EOFToken(object):
|
||||
pass
|
||||
|
||||
class DefToken(object):
|
||||
pass
|
||||
|
||||
class ExternToken(object):
|
||||
pass
|
||||
|
||||
class IdentifierToken(object):
|
||||
def __init__(self, name): self.name = name
|
||||
|
||||
class NumberToken(object):
|
||||
def __init__(self, value): self.value = value
|
||||
|
||||
class CharacterToken(object):
|
||||
def __init__(self, char): self.char = char
|
||||
def __eq__(self, other):
|
||||
return isinstance(other, CharacterToken) and self.char == other.char
|
||||
def __ne__(self, other): return not self == other
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Each token yielded by our lexer will be of one of the above types. For simple
|
||||
tokens that are always the same, like the "def" keyword, the lexer will yield
|
||||
<tt>DefToken()</tt>. Identifiers, numbers and characters, on the other
|
||||
hand, have extra data, so when the lexer encounteres the number 123.45, it will
|
||||
emit it as <tt>NumberToken(123.45)</tt>. An identifier <tt>foo</tt> will be
|
||||
emitted as <tt>IdentifierToken('foo')</tt>. And finally, an unknown character
|
||||
like '+' will be returned as <tt>CharacterToken('+')</tt>. You may notice that
|
||||
we overload the equality and inequality operators for the characters; this will
|
||||
later simplify character comparisons in the parser code.</p>
|
||||
|
||||
<p>The actual implementation of the lexer is a single function called
|
||||
<tt>Tokenize</tt>, which takes a string and
|
||||
<a href="http://docs.python.org/reference/simple_stmts.html#the-yield-statement">yields</a>
|
||||
tokens. For simplicity, we will use
|
||||
<a href="http://docs.python.org/library/re.html">regular
|
||||
expressions</a> to parse out the tokens. This is terribly inefficient, but
|
||||
perfectly sufficient for our needs.</p>
|
||||
|
||||
<p>First, we define the regular expressions for our tokens. Numbers and strings
|
||||
of digits, optionally followed by a period and another string of digits.
|
||||
Identifiers (and keywords) are alphanumeric string starting with a letter and
|
||||
comments are anything between a hash (<tt>#</tt>) and the end of the line.
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
import re
|
||||
|
||||
...
|
||||
|
||||
# Regular expressions that tokens and comments of our language.
|
||||
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
|
||||
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
|
||||
REGEX_COMMENT = re.compile('#.*')
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>
|
||||
Next, let's start defining the <tt>Tokenize</tt> function itself. The first
|
||||
thing we need to do is set up a loop that scans the string, while ignoring
|
||||
whitespace between tokens:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
def Tokenize(string):
|
||||
while string:
|
||||
# Skip whitespace.
|
||||
if string[0].isspace():
|
||||
string = string[1:]
|
||||
continue
|
||||
|
||||
...
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Next we want to find out what the next token is. For this we run the regexes
|
||||
we defined above on the remainder of the string. To simplify the rest of the
|
||||
code, we run all three regexes each time. As mentioned above, inefficiencies are
|
||||
ignored for the purpose of this tutorial:<p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Run regexes.
|
||||
comment_match = REGEX_COMMENT.match(string)
|
||||
number_match = REGEX_NUMBER.match(string)
|
||||
identifier_match = REGEX_IDENTIFIER.match(string)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Now se check if any of the regexes matched. For comments, we simply
|
||||
ignore the captured match:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Check if any of the regexes matched and yield the appropriate result.
|
||||
if comment_match:
|
||||
comment = comment_match.group(0)
|
||||
string = string[len(comment):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>For numbers, we yield the captured match, converted to a float and tagged
|
||||
with the appropriate token type:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
elif number_match:
|
||||
number = number_match.group(0)
|
||||
yield NumberToken(float(number))
|
||||
string = string[len(number):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>The identifier case is a little more complex. We have to check for keywords
|
||||
to decide whether we have captured an identifier or a keyword:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
elif identifier_match:
|
||||
identifier = identifier_match.group(0)
|
||||
# Check if we matched a keyword.
|
||||
if identifier == 'def':
|
||||
yield DefToken()
|
||||
elif identifier == 'extern':
|
||||
yield ExternToken()
|
||||
else:
|
||||
yield IdentifierToken(identifier)
|
||||
string = string[len(identifier):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Finally, if we haven't recognized a comment, a number of an identifier, we
|
||||
yield the current character as an "unknown character" token. This is used, for
|
||||
example, for operators like <tt>+</tt> or <tt>*</tt>:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
else:
|
||||
# Yield the unknown character.
|
||||
yield CharacterToken(string[0])
|
||||
string = string[1:]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Once we're done with the
|
||||
loop, we return a final end-of-file token:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
yield EOFToken()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>With this, we have the complete lexer for the basic Kaleidoscope language
|
||||
(the <a href="PythonLangImpl2.html#code">full code listing</a> for the Lexer is
|
||||
available in the <a href="PythonLangImpl2.html">next chapter</a> of the
|
||||
tutorial). Next we'll <a href="PythonLangImpl2.html">build a simple parser that
|
||||
uses this to build an Abstract Syntax Tree</a>. When we have that, we'll
|
||||
include a driver so that you can use the lexer and parser together.
|
||||
</p>
|
||||
|
||||
<a href="PythonLangImpl2.html">Next: Implementing a Parser and AST</a>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://max99x.com">Max Shawabkeh</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
1097
www/src/kaleidoscope/PythonLangImpl2.html
Normal file
1097
www/src/kaleidoscope/PythonLangImpl2.html
Normal file
File diff suppressed because it is too large
Load diff
1119
www/src/kaleidoscope/PythonLangImpl3.html
Normal file
1119
www/src/kaleidoscope/PythonLangImpl3.html
Normal file
File diff suppressed because it is too large
Load diff
999
www/src/kaleidoscope/PythonLangImpl4.html
Normal file
999
www/src/kaleidoscope/PythonLangImpl4.html
Normal file
|
|
@ -0,0 +1,999 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Adding JIT and Optimizer Support</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<meta name="author" content="Max Shawabkeh">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Adding JIT and Optimizer Support</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 4
|
||||
<ol>
|
||||
<li><a href="#intro">Chapter 4 Introduction</a></li>
|
||||
<li><a href="#trivialconstfold">Trivial Constant Folding</a></li>
|
||||
<li><a href="#optimizerpasses">LLVM Optimization Passes</a></li>
|
||||
<li><a href="#jit">Adding a JIT Compiler</a></li>
|
||||
<li><a href="#code">Full Code Listing</a></li>
|
||||
</ol>
|
||||
</li>
|
||||
<li><a href="PythonLangImpl5.html">Chapter 5</a>: Extending the Language:
|
||||
Control Flow</li>
|
||||
</ul>
|
||||
|
||||
<div class="doc_author">
|
||||
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
|
||||
and <a href="http://max99x.com">Max Shawabkeh</a>
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="intro">Chapter 4 Introduction</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Welcome to Chapter 4 of the
|
||||
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
|
||||
with LLVM</a>" tutorial. Chapters 1-3 described the implementation of a simple
|
||||
language and added support for generating LLVM IR. This chapter describes
|
||||
two new techniques: adding optimizer support to your language, and adding JIT
|
||||
compiler support. These additions will demonstrate how to get nice, efficient
|
||||
code for the Kaleidoscope language.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="trivialconstfold">Trivial Constant
|
||||
Folding</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Our demonstration for Chapter 3 is elegant and easy to extend. Unfortunately,
|
||||
it does not produce wonderful code. The LLVM Builder, however, does give us
|
||||
obvious optimizations when compiling simple code:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) 1+2+x</b>
|
||||
Read function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 3.000000e+00, %x
|
||||
ret double %addtmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This code is not a literal transcription of the AST built by parsing the
|
||||
input. That would be:
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) 1+2+x</b>
|
||||
Read function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 2.000000e+00, 1.000000e+00
|
||||
%addtmp1 = fadd double %addtmp, %x
|
||||
ret double %addtmp1
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Constant folding, as seen above, in particular, is a very common and very
|
||||
important optimization: so much so that many language implementors implement
|
||||
constant folding support in their AST representation.</p>
|
||||
|
||||
<p>With LLVM, you don't need this support in the AST. Since all calls to build
|
||||
LLVM IR go through the LLVM IR builder, the builder itself checked to see if
|
||||
there was a constant folding opportunity when you call it. If so, it just does
|
||||
the constant fold and return the constant instead of creating an instruction.
|
||||
|
||||
<p>Well, that was easy :). In practice, we recommend always using
|
||||
<tt>llvm.core.Builder</tt> when generating code like this. It has no
|
||||
"syntactic overhead" for its use (you don't have to uglify your compiler with
|
||||
constant checks everywhere) and it can dramatically reduce the amount of
|
||||
LLVM IR that is generated in some cases (particular for languages with a macro
|
||||
preprocessor or that use a lot of constants).</p>
|
||||
|
||||
<p>On the other hand, the <tt>Builder</tt> is limited by the fact that it does
|
||||
all of its analysis inline with the code as it is built. If you take a slightly
|
||||
more complex example:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) (1+2+x)*(x+(1+2))</b>
|
||||
Read a function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 3.000000e+00, %x ; <double> [#uses=1]
|
||||
%addtmp1 = fadd double %x, 3.000000e+00 ; <double> [#uses=1]
|
||||
%multmp = fmul double %addtmp, %addtmp1 ; <double> [#uses=1]
|
||||
ret double %multmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>In this case, the LHS and RHS of the multiplication are the same value. We'd
|
||||
really like to see this generate "<tt>tmp = x+3; result = tmp*tmp;</tt>" instead
|
||||
of computing "<tt>x+3</tt>" twice.</p>
|
||||
|
||||
<p>Unfortunately, no amount of local analysis will be able to detect and correct
|
||||
this. This requires two transformations: reassociation of expressions (to
|
||||
make the add's lexically identical) and Common Subexpression Elimination (CSE)
|
||||
to delete the redundant add instruction. Fortunately, LLVM provides a broad
|
||||
range of optimizations that you can use, in the form of "passes".</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="optimizerpasses">LLVM Optimization
|
||||
Passes</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>LLVM provides many optimization passes, which do many different sorts of
|
||||
things and have different tradeoffs. Unlike other systems, LLVM doesn't hold
|
||||
to the mistaken notion that one set of optimizations is right for all languages
|
||||
and for all situations. LLVM allows a compiler implementor to make complete
|
||||
decisions about what optimizations to use, in which order, and in what
|
||||
situation.</p>
|
||||
|
||||
<p>As a concrete example, LLVM supports both "whole module" passes, which look
|
||||
across as large of body of code as they can (often a whole file, but if run
|
||||
at link time, this can be a substantial portion of the whole program). It also
|
||||
supports and includes "per-function" passes which just operate on a single
|
||||
function at a time, without looking at other functions. For more information
|
||||
on passes and how they are run, see the
|
||||
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html">How to Write a
|
||||
Pass</a> document and the <a href="http://www.llvm.org/docs/Passes.html">List of
|
||||
LLVM Passes</a>.</p>
|
||||
|
||||
<p>For Kaleidoscope, we are currently generating functions on the fly, one at
|
||||
a time, as the user types them in. We aren't shooting for the ultimate
|
||||
optimization experience in this setting, but we also want to catch the easy and
|
||||
quick stuff where possible. As such, we will choose to run a few per-function
|
||||
optimizations as the user types the function in. If we wanted to make a "static
|
||||
Kaleidoscope compiler", we would use exactly the code we have now, except that
|
||||
we would defer running the optimizer until the entire file has been parsed.</p>
|
||||
|
||||
<p>In order to get per-function optimizations going, we need to set up a
|
||||
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html#passmanager">
|
||||
FunctionPassManager</a> to hold and organize the LLVM optimizations that we want
|
||||
to run. Once we have that, we can add a set of optimizations to run. The code
|
||||
looks like this:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The function optimization passes manager.
|
||||
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
|
||||
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
|
||||
...
|
||||
|
||||
def main():
|
||||
# Set up the optimizer pipeline. Start with registering info about how the
|
||||
# target lays out data structures.
|
||||
g_llvm_pass_manager.add(g_llvm_executor.target_data)
|
||||
# Do simple "peephole" optimizations and bit-twiddling optzns.
|
||||
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
|
||||
# Reassociate expressions.
|
||||
g_llvm_pass_manager.add(PASS_REASSOCIATE)
|
||||
# Eliminate Common SubExpressions.
|
||||
g_llvm_pass_manager.add(PASS_GVN)
|
||||
# Simplify the control flow graph (deleting unreachable blocks, etc).
|
||||
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
g_llvm_pass_manager.initialize()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This code defines a <tt>FunctionPassManager</tt>,
|
||||
<tt>g_llvm_pass_manager</tt>. Once it is set up, we use a series of "add" calls
|
||||
to add a bunch of LLVM passes. The first pass is basically boilerplate, it adds
|
||||
a pass so that later optimizations know how the data structures in the program
|
||||
are laid out. (The "<tt>g_llvm_executor</tt>" variable is related to the JIT,
|
||||
which we will get to in the next section.) In this case, we choose to add 4
|
||||
optimization passes. The passes we chose here are a pretty standard set of
|
||||
"cleanup" optimizations that are useful for a wide variety of code. I won't
|
||||
delve into what they do but, believe me, they are a good starting place :).</p>
|
||||
|
||||
<p>Once the pass manager is set up, we need to make use of it. We do this by
|
||||
running it after our newly created function is constructed (in
|
||||
<tt>FunctionNode.CodeGen</tt>), but before it is returned to the client:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
return_value = self.body.CodeGen()
|
||||
g_llvm_builder.ret(return_value)
|
||||
|
||||
# Validate the generated code, checking for consistency.
|
||||
function.verify()
|
||||
|
||||
<b># Optimize the function.
|
||||
g_llvm_pass_manager.run(function)</b>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>As you can see, this is pretty straightforward. The
|
||||
<tt>FunctionPassManager</tt> optimizes and updates the LLVM Function in place,
|
||||
improving (hopefully) its body. With this in place, we can try our test above
|
||||
again:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) (1+2+x)*(x+(1+2))</b>
|
||||
Read a function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double %x, 3.000000e+00 ; <double> [#uses=2]
|
||||
%multmp = fmul double %addtmp, %addtmp ; <double> [#uses=1]
|
||||
ret double %multmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>As expected, we now get our nicely optimized code, saving a floating point
|
||||
add instruction from every execution of this function.</p>
|
||||
|
||||
<p>LLVM provides a wide variety of optimizations that can be used in certain
|
||||
circumstances. Some
|
||||
<a href="http://www.llvm.org/docs/Passes.html">documentation about the various
|
||||
passes</a> is available, but it isn't very complete. Another good source of
|
||||
ideas can come from looking at the passes that <tt>llvm-gcc</tt> or
|
||||
<tt>llvm-ld</tt> run to get started. The "<tt>opt</tt>" tool allows you to
|
||||
experiment with passes from the command line, so you can see if they do
|
||||
anything.</p>
|
||||
|
||||
<p>Now that we have reasonable code coming out of our front-end, lets talk about
|
||||
executing it!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="jit">Adding a JIT Compiler</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Code that is available in LLVM IR can have a wide variety of tools
|
||||
applied to it. For example, you can run optimizations on it (as we did above),
|
||||
you can dump it out in textual or binary forms, you can compile the code to an
|
||||
assembly file (.s) for some target, or you can JIT compile it. The nice thing
|
||||
about the LLVM IR representation is that it is the "common currency" between
|
||||
many different parts of the compiler.
|
||||
</p>
|
||||
|
||||
<p>In this section, we'll add JIT compiler support to our interpreter. The
|
||||
basic idea that we want for Kaleidoscope is to have the user enter function
|
||||
bodies as they do now, but immediately evaluate the top-level expressions they
|
||||
type in. For example, if they type in "1 + 2", we should evaluate and print
|
||||
out 3. If they define a function, they should be able to call it from the
|
||||
command line.</p>
|
||||
|
||||
<p>In order to do this, we first declare and initialize the JIT. This is done
|
||||
by adding and initializing a global variable:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This creates an abstract "Execution Engine" which can be either a JIT
|
||||
compiler or the LLVM interpreter. LLVM will automatically pick a JIT compiler
|
||||
for you if one is available for your platform, otherwise it will fall back to
|
||||
the interpreter.</p>
|
||||
|
||||
<p>Once the <tt>ExecutionEngine</tt> is created, the JIT is ready to be used.
|
||||
We can use the <tt>run_function</tt> method of the execution engine to execute
|
||||
a compiled function and get its return value. In our case, this means that we
|
||||
can change the code that parses a top-level expression to look like this:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
def HandleTopLevelExpression(self):
|
||||
try:
|
||||
function = self.ParseTopLevelExpr().CodeGen()
|
||||
result = g_llvm_executor.run_function(function, [])
|
||||
print 'Evaluated to:', result.as_real(Type.double())
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Recall that we compile top-level expressions into a self-contained LLVM
|
||||
function that takes no arguments and returns the computed double.</p>
|
||||
|
||||
<p>With just these two changes, lets see how Kaleidoscope works now!</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>4+5</b>
|
||||
Read a top level expression:
|
||||
define double @0() {
|
||||
entry:
|
||||
ret double 9.000000e+00
|
||||
}
|
||||
|
||||
Evaluated to: 9.0
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Well this looks like it is basically working. The dump of the function
|
||||
shows the "no argument function that always returns double" that we synthesize
|
||||
for each top-level expression that is typed in. This demonstrates very basic
|
||||
functionality, but can we do more?</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def testfunc(x y) x + y*2</b>
|
||||
Read a function definition:
|
||||
define double @testfunc(double %x, double %y) {
|
||||
entry:
|
||||
%multmp = fmul double %y, 2.000000e+00 ; <double> [#uses=1]
|
||||
%addtmp = fadd double %multmp, %x ; <double> [#uses=1]
|
||||
ret double %addtmp
|
||||
}
|
||||
|
||||
ready> <b>testfunc(4, 10)</b>
|
||||
Read a top level expression:
|
||||
define double @0() {
|
||||
entry:
|
||||
%calltmp = call double @testfunc(double 4.000000e+00, double 1.000000e+01) ; <double> [#uses=1]
|
||||
ret double %calltmp
|
||||
}
|
||||
|
||||
<em>Evaluated to: 24.0</em>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This illustrates that we can now call user code, but there is something a bit
|
||||
subtle going on here. Note that we only invoke the JIT on the anonymous
|
||||
functions that <em>call testfunc</em>, but we never invoked it
|
||||
on <em>testfunc</em> itself. What actually happened here is that the JIT
|
||||
scanned for all non-JIT'd functions transitively called from the anonymous
|
||||
function and compiled all of them before returning from <tt>run_function()</tt>.
|
||||
</p>
|
||||
|
||||
<p>The JIT provides a number of other more advanced interfaces for things like
|
||||
freeing allocated machine code, rejit'ing functions to update them, etc.
|
||||
However, even with this simple code, we get some surprisingly powerful
|
||||
capabilities - check this out (I removed the dump of the anonymous functions,
|
||||
you should get the idea by now :) :</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>extern sin(x)</b>
|
||||
Read an extern:
|
||||
declare double @sin(double)
|
||||
|
||||
ready> <b>extern cos(x)</b>
|
||||
Read an extern:
|
||||
declare double @cos(double)
|
||||
|
||||
ready> <b>sin(1.0)</b>
|
||||
<em>Evaluated to: 0.841470984808</em>
|
||||
|
||||
ready> <b>def foo(x) sin(x)*sin(x) + cos(x)*cos(x)</b>
|
||||
Read a function definition:
|
||||
define double @foo(double %x) {
|
||||
entry:
|
||||
%calltmp = call double @sin(double %x) ; <double> [#uses=1]
|
||||
%calltmp1 = call double @sin(double %x) ; <double> [#uses=1]
|
||||
%multmp = fmul double %calltmp, %calltmp1 ; <double> [#uses=1]
|
||||
%calltmp2 = call double @cos(double %x) ; <double> [#uses=1]
|
||||
%calltmp3 = call double @cos(double %x) ; <double> [#uses=1]
|
||||
%multmp4 = fmul double %calltmp2, %calltmp3 ; <double> [#uses=1]
|
||||
%addtmp = fadd double %multmp, %multmp4 ; <double> [#uses=1]
|
||||
ret double %addtmp
|
||||
}
|
||||
|
||||
ready> <b>foo(4.0)</b>
|
||||
<em>Evaluated to: 1.000000</em>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Whoa, how does the JIT know about sin and cos? The answer is surprisingly
|
||||
simple: in this example, the JIT started execution of a function and got to a
|
||||
function call. It realized that the function was not yet JIT compiled and
|
||||
invoked the standard set of routines to resolve the function. In this case,
|
||||
there is no body defined for the function, so the JIT ended up calling
|
||||
"<tt>dlsym("sin")</tt>" on the Python process that is hosting our Kaleidoscope
|
||||
prompt. Since "<tt>sin</tt>" is defined within the JIT's address space, it
|
||||
simply patches up calls in the module to call the libm version of <tt>sin</tt>
|
||||
directly.</p>
|
||||
|
||||
<p>One interesting application of this is that we can now extend the language
|
||||
by writing arbitrary C++ code to implement operations. For example, we can
|
||||
create a C file with the following simple function:
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#include <stdio.h>
|
||||
|
||||
double putchard(double x) {
|
||||
putchar((char)x);
|
||||
return 0;
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>We can then compile this into a shared library with GCC:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
gcc -shared -fPIC -o putchard.so putchard.c
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Now we can load this library into the Python process using
|
||||
<tt>llvm.core.load_library_permanently</tt> and access it from Kaleidoscope to
|
||||
produce simple output to the console:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
>>> <b>import llvm.core</b>
|
||||
>>> <b>llvm.core.load_library_permanently('/home/max/llvm-py-tutorial/putchard.so')</b>
|
||||
>>> <b>import kaleidoscope</b>
|
||||
>>> <b>kaleidoscope.main()</b>
|
||||
ready> <b>extern putchard(x)</b>
|
||||
Read an extern:
|
||||
declare double @putchard(double)
|
||||
|
||||
ready> <b>putchard(65) + putchard(66) + putchard(67) + putchard(10)</b>
|
||||
<em>ABC</em>
|
||||
Evaluated to: 0.0
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Similar code could be used to implement file I/O, console input, and many
|
||||
other capabilities in Kaleidoscope.</p>
|
||||
|
||||
<p>This completes the JIT and optimizer chapter of the Kaleidoscope tutorial. At
|
||||
this point, we can compile a non-Turing-complete programming language, optimize
|
||||
and JIT compile it in a user-driven way. Next up we'll look into <a
|
||||
href="PythonLangImpl5.html">extending the language with control flow
|
||||
constructs</a>, tackling some interesting LLVM IR issues along the way.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="code">Full Code Listing</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Here is the complete code listing for our running example, enhanced with the
|
||||
LLVM JIT and optimizer:
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#!/usr/bin/env python
|
||||
|
||||
import re
|
||||
from llvm.core import Module, Constant, Type, Function, Builder, FCMP_ULT
|
||||
from llvm.ee import ExecutionEngine, TargetData
|
||||
from llvm.passes import FunctionPassManager
|
||||
from llvm.passes import (PASS_INSTRUCTION_COMBINING,
|
||||
PASS_REASSOCIATE,
|
||||
PASS_GVN,
|
||||
PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
################################################################################
|
||||
## Globals
|
||||
################################################################################
|
||||
|
||||
# The LLVM module, which holds all the IR code.
|
||||
g_llvm_module = Module.new('my cool jit')
|
||||
|
||||
# The LLVM instruction builder. Created whenever a new function is entered.
|
||||
g_llvm_builder = None
|
||||
|
||||
# A dictionary that keeps track of which values are defined in the current scope
|
||||
# and what their LLVM representation is.
|
||||
g_named_values = {}
|
||||
|
||||
# The function optimization passes manager.
|
||||
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
|
||||
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
|
||||
################################################################################
|
||||
## Lexer
|
||||
################################################################################
|
||||
|
||||
# The lexer yields one of these types for each token.
|
||||
class EOFToken(object):
|
||||
pass
|
||||
|
||||
class DefToken(object):
|
||||
pass
|
||||
|
||||
class ExternToken(object):
|
||||
pass
|
||||
|
||||
class IdentifierToken(object):
|
||||
def __init__(self, name): self.name = name
|
||||
|
||||
class NumberToken(object):
|
||||
def __init__(self, value): self.value = value
|
||||
|
||||
class CharacterToken(object):
|
||||
def __init__(self, char): self.char = char
|
||||
def __eq__(self, other):
|
||||
return isinstance(other, CharacterToken) and self.char == other.char
|
||||
def __ne__(self, other): return not self == other
|
||||
|
||||
# Regular expressions that tokens and comments of our language.
|
||||
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
|
||||
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
|
||||
REGEX_COMMENT = re.compile('#.*')
|
||||
|
||||
def Tokenize(string):
|
||||
while string:
|
||||
# Skip whitespace.
|
||||
if string[0].isspace():
|
||||
string = string[1:]
|
||||
continue
|
||||
|
||||
# Run regexes.
|
||||
comment_match = REGEX_COMMENT.match(string)
|
||||
number_match = REGEX_NUMBER.match(string)
|
||||
identifier_match = REGEX_IDENTIFIER.match(string)
|
||||
|
||||
# Check if any of the regexes matched and yield the appropriate result.
|
||||
if comment_match:
|
||||
comment = comment_match.group(0)
|
||||
string = string[len(comment):]
|
||||
elif number_match:
|
||||
number = number_match.group(0)
|
||||
yield NumberToken(float(number))
|
||||
string = string[len(number):]
|
||||
elif identifier_match:
|
||||
identifier = identifier_match.group(0)
|
||||
# Check if we matched a keyword.
|
||||
if identifier == 'def':
|
||||
yield DefToken()
|
||||
elif identifier == 'extern':
|
||||
yield ExternToken()
|
||||
else:
|
||||
yield IdentifierToken(identifier)
|
||||
string = string[len(identifier):]
|
||||
else:
|
||||
# Yield the ASCII value of the unknown character.
|
||||
yield CharacterToken(string[0])
|
||||
string = string[1:]
|
||||
|
||||
yield EOFToken()
|
||||
|
||||
################################################################################
|
||||
## Abstract Syntax Tree (aka Parse Tree)
|
||||
################################################################################
|
||||
|
||||
# Base class for all expression nodes.
|
||||
class ExpressionNode(object):
|
||||
pass
|
||||
|
||||
# Expression class for numeric literals like "1.0".
|
||||
class NumberExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, value):
|
||||
self.value = value
|
||||
|
||||
def CodeGen(self):
|
||||
return Constant.real(Type.double(), self.value)
|
||||
|
||||
# Expression class for referencing a variable, like "a".
|
||||
class VariableExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, name):
|
||||
self.name = name
|
||||
|
||||
def CodeGen(self):
|
||||
if self.name in g_named_values:
|
||||
return g_named_values[self.name]
|
||||
else:
|
||||
raise RuntimeError('Unknown variable name: ' + self.name)
|
||||
|
||||
# Expression class for a binary operator.
|
||||
class BinaryOperatorExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, operator, left, right):
|
||||
self.operator = operator
|
||||
self.left = left
|
||||
self.right = right
|
||||
|
||||
def CodeGen(self):
|
||||
left = self.left.CodeGen()
|
||||
right = self.right.CodeGen()
|
||||
|
||||
if self.operator == '+':
|
||||
return g_llvm_builder.fadd(left, right, 'addtmp')
|
||||
elif self.operator == '-':
|
||||
return g_llvm_builder.fsub(left, right, 'subtmp')
|
||||
elif self.operator == '*':
|
||||
return g_llvm_builder.fmul(left, right, 'multmp')
|
||||
elif self.operator == '<':
|
||||
result = g_llvm_builder.fcmp(FCMP_ULT, left, right, 'cmptmp')
|
||||
# Convert bool 0 or 1 to double 0.0 or 1.0.
|
||||
return g_llvm_builder.uitofp(result, Type.double(), 'booltmp')
|
||||
else:
|
||||
raise RuntimeError('Unknown binary operator.')
|
||||
|
||||
# Expression class for function calls.
|
||||
class CallExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, callee, args):
|
||||
self.callee = callee
|
||||
self.args = args
|
||||
|
||||
def CodeGen(self):
|
||||
# Look up the name in the global module table.
|
||||
callee = g_llvm_module.get_function_named(self.callee)
|
||||
|
||||
# Check for argument mismatch error.
|
||||
if len(callee.args) != len(self.args):
|
||||
raise RuntimeError('Incorrect number of arguments passed.')
|
||||
|
||||
arg_values = [i.CodeGen() for i in self.args]
|
||||
|
||||
return g_llvm_builder.call(callee, arg_values, 'calltmp')
|
||||
|
||||
# This class represents the "prototype" for a function, which captures its name,
|
||||
# and its argument names (thus implicitly the number of arguments the function
|
||||
# takes).
|
||||
class PrototypeNode(object):
|
||||
|
||||
def __init__(self, name, args):
|
||||
self.name = name
|
||||
self.args = args
|
||||
|
||||
def CodeGen(self):
|
||||
# Make the function type, eg. double(double,double).
|
||||
funct_type = Type.function(
|
||||
Type.double(), [Type.double()] * len(self.args), False)
|
||||
|
||||
function = Function.new(g_llvm_module, funct_type, self.name)
|
||||
|
||||
# If the name conflicted, there was already something with the same name.
|
||||
# If it has a body, don't allow redefinition or reextern.
|
||||
if function.name != self.name:
|
||||
function.delete()
|
||||
function = g_llvm_module.get_function_named(self.name)
|
||||
|
||||
# If the function already has a body, reject this.
|
||||
if not function.is_declaration:
|
||||
raise RuntimeError('Redefinition of function.')
|
||||
|
||||
# If F took a different number of args, reject.
|
||||
if len(callee.args) != len(self.args):
|
||||
raise RuntimeError('Redeclaration of a function with different number '
|
||||
'of args.')
|
||||
|
||||
# Set names for all arguments and add them to the variables symbol table.
|
||||
for arg, arg_name in zip(function.args, self.args):
|
||||
arg.name = arg_name
|
||||
# Add arguments to variable symbol table.
|
||||
g_named_values[arg_name] = arg
|
||||
|
||||
return function
|
||||
|
||||
# This class represents a function definition itself.
|
||||
class FunctionNode(object):
|
||||
|
||||
def __init__(self, prototype, body):
|
||||
self.prototype = prototype
|
||||
self.body = body
|
||||
|
||||
def CodeGen(self):
|
||||
# Clear scope.
|
||||
g_named_values.clear()
|
||||
|
||||
# Create a function object.
|
||||
function = self.prototype.CodeGen()
|
||||
|
||||
# Create a new basic block to start insertion into.
|
||||
block = function.append_basic_block('entry')
|
||||
global g_llvm_builder
|
||||
g_llvm_builder = Builder.new(block)
|
||||
|
||||
# Finish off the function.
|
||||
try:
|
||||
return_value = self.body.CodeGen()
|
||||
g_llvm_builder.ret(return_value)
|
||||
|
||||
# Validate the generated code, checking for consistency.
|
||||
function.verify()
|
||||
|
||||
# Optimize the function.
|
||||
g_llvm_pass_manager.run(function)
|
||||
except:
|
||||
function.delete()
|
||||
raise
|
||||
|
||||
return function
|
||||
|
||||
|
||||
################################################################################
|
||||
## Parser
|
||||
################################################################################
|
||||
|
||||
class Parser(object):
|
||||
|
||||
def __init__(self, tokens, binop_precedence):
|
||||
self.tokens = tokens
|
||||
self.binop_precedence = binop_precedence
|
||||
self.Next()
|
||||
|
||||
# Provide a simple token buffer. Parser.current is the current token the
|
||||
# parser is looking at. Parser.Next() reads another token from the lexer and
|
||||
# updates Parser.current with its results.
|
||||
def Next(self):
|
||||
self.current = self.tokens.next()
|
||||
|
||||
# Gets the precedence of the current token, or -1 if the token is not a binary
|
||||
# operator.
|
||||
def GetCurrentTokenPrecedence(self):
|
||||
if isinstance(self.current, CharacterToken):
|
||||
return self.binop_precedence.get(self.current.char, -1)
|
||||
else:
|
||||
return -1
|
||||
|
||||
# identifierexpr ::= identifier | identifier '(' expression* ')'
|
||||
def ParseIdentifierExpr(self):
|
||||
identifier_name = self.current.name
|
||||
self.Next() # eat identifier.
|
||||
|
||||
if self.current != CharacterToken('('): # Simple variable reference.
|
||||
return VariableExpressionNode(identifier_name)
|
||||
|
||||
# Call.
|
||||
self.Next() # eat '('.
|
||||
args = []
|
||||
if self.current != CharacterToken(')'):
|
||||
while True:
|
||||
args.append(self.ParseExpression())
|
||||
if self.current == CharacterToken(')'):
|
||||
break
|
||||
elif self.current != CharacterToken(','):
|
||||
raise RuntimeError('Expected ")" or "," in argument list.')
|
||||
self.Next()
|
||||
|
||||
self.Next() # eat ')'.
|
||||
return CallExpressionNode(identifier_name, args)
|
||||
|
||||
# numberexpr ::= number
|
||||
def ParseNumberExpr(self):
|
||||
result = NumberExpressionNode(self.current.value)
|
||||
self.Next() # consume the number.
|
||||
return result
|
||||
|
||||
# parenexpr ::= '(' expression ')'
|
||||
def ParseParenExpr(self):
|
||||
self.Next() # eat '('.
|
||||
|
||||
contents = self.ParseExpression()
|
||||
|
||||
if self.current != CharacterToken(')'):
|
||||
raise RuntimeError('Expected ")".')
|
||||
self.Next() # eat ')'.
|
||||
|
||||
return contents
|
||||
|
||||
# primary ::= identifierexpr | numberexpr | parenexpr
|
||||
def ParsePrimary(self):
|
||||
if isinstance(self.current, IdentifierToken):
|
||||
return self.ParseIdentifierExpr()
|
||||
elif isinstance(self.current, NumberToken):
|
||||
return self.ParseNumberExpr()
|
||||
elif self.current == CharacterToken('('):
|
||||
return self.ParseParenExpr()
|
||||
else:
|
||||
raise RuntimeError('Unknown token when expecting an expression.')
|
||||
|
||||
# binoprhs ::= (operator primary)*
|
||||
def ParseBinOpRHS(self, left, left_precedence):
|
||||
# If this is a binary operator, find its precedence.
|
||||
while True:
|
||||
precedence = self.GetCurrentTokenPrecedence()
|
||||
|
||||
# If this is a binary operator that binds at least as tightly as the
|
||||
# current one, consume it; otherwise we are done.
|
||||
if precedence < left_precedence:
|
||||
return left
|
||||
|
||||
binary_operator = self.current.char
|
||||
self.Next() # eat the operator.
|
||||
|
||||
# Parse the primary expression after the binary operator.
|
||||
right = self.ParsePrimary()
|
||||
|
||||
# If binary_operator binds less tightly with right than the operator after
|
||||
# right, let the pending operator take right as its left.
|
||||
next_precedence = self.GetCurrentTokenPrecedence()
|
||||
if precedence < next_precedence:
|
||||
right = self.ParseBinOpRHS(right, precedence + 1)
|
||||
|
||||
# Merge left/right.
|
||||
left = BinaryOperatorExpressionNode(binary_operator, left, right)
|
||||
|
||||
# expression ::= primary binoprhs
|
||||
def ParseExpression(self):
|
||||
left = self.ParsePrimary()
|
||||
return self.ParseBinOpRHS(left, 0)
|
||||
|
||||
# prototype ::= id '(' id* ')'
|
||||
def ParsePrototype(self):
|
||||
if not isinstance(self.current, IdentifierToken):
|
||||
raise RuntimeError('Expected function name in prototype.')
|
||||
|
||||
function_name = self.current.name
|
||||
self.Next() # eat function name.
|
||||
|
||||
if self.current != CharacterToken('('):
|
||||
raise RuntimeError('Expected "(" in prototype.')
|
||||
self.Next() # eat '('.
|
||||
|
||||
arg_names = []
|
||||
while isinstance(self.current, IdentifierToken):
|
||||
arg_names.append(self.current.name)
|
||||
self.Next()
|
||||
|
||||
if self.current != CharacterToken(')'):
|
||||
raise RuntimeError('Expected ")" in prototype.')
|
||||
|
||||
# Success.
|
||||
self.Next() # eat ')'.
|
||||
|
||||
return PrototypeNode(function_name, arg_names)
|
||||
|
||||
# definition ::= 'def' prototype expression
|
||||
def ParseDefinition(self):
|
||||
self.Next() # eat def.
|
||||
proto = self.ParsePrototype()
|
||||
body = self.ParseExpression()
|
||||
return FunctionNode(proto, body)
|
||||
|
||||
# toplevelexpr ::= expression
|
||||
def ParseTopLevelExpr(self):
|
||||
proto = PrototypeNode('', [])
|
||||
return FunctionNode(proto, self.ParseExpression())
|
||||
|
||||
# external ::= 'extern' prototype
|
||||
def ParseExtern(self):
|
||||
self.Next() # eat extern.
|
||||
return self.ParsePrototype()
|
||||
|
||||
# Top-Level parsing
|
||||
def HandleDefinition(self):
|
||||
self.Handle(self.ParseDefinition, 'Read a function definition:')
|
||||
|
||||
def HandleExtern(self):
|
||||
self.Handle(self.ParseExtern, 'Read an extern:')
|
||||
|
||||
def HandleTopLevelExpression(self):
|
||||
try:
|
||||
function = self.ParseTopLevelExpr().CodeGen()
|
||||
result = g_llvm_executor.run_function(function, [])
|
||||
print 'Evaluated to:', result.as_real(Type.double())
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
|
||||
def Handle(self, function, message):
|
||||
try:
|
||||
print message, function().CodeGen()
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
|
||||
################################################################################
|
||||
## Main driver code.
|
||||
################################################################################
|
||||
|
||||
def main():
|
||||
# Set up the optimizer pipeline. Start with registering info about how the
|
||||
# target lays out data structures.
|
||||
g_llvm_pass_manager.add(g_llvm_executor.target_data)
|
||||
# Do simple "peephole" optimizations and bit-twiddling optzns.
|
||||
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
|
||||
# Reassociate expressions.
|
||||
g_llvm_pass_manager.add(PASS_REASSOCIATE)
|
||||
# Eliminate Common SubExpressions.
|
||||
g_llvm_pass_manager.add(PASS_GVN)
|
||||
# Simplify the control flow graph (deleting unreachable blocks, etc).
|
||||
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
g_llvm_pass_manager.initialize()
|
||||
|
||||
# Install standard binary operators.
|
||||
# 1 is lowest possible precedence. 40 is the highest.
|
||||
operator_precedence = {
|
||||
'<': 10,
|
||||
'+': 20,
|
||||
'-': 20,
|
||||
'*': 40
|
||||
}
|
||||
|
||||
# Run the main "interpreter loop".
|
||||
while True:
|
||||
print 'ready>',
|
||||
try:
|
||||
raw = raw_input()
|
||||
except KeyboardInterrupt:
|
||||
break
|
||||
|
||||
parser = Parser(Tokenize(raw), operator_precedence)
|
||||
while True:
|
||||
# top ::= definition | external | expression | EOF
|
||||
if isinstance(parser.current, EOFToken):
|
||||
break
|
||||
if isinstance(parser.current, DefToken):
|
||||
parser.HandleDefinition()
|
||||
elif isinstance(parser.current, ExternToken):
|
||||
parser.HandleExtern()
|
||||
else:
|
||||
parser.HandleTopLevelExpression()
|
||||
|
||||
# Print out all of the generated code.
|
||||
print '\n', g_llvm_module
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<a href="PythonLangImpl5.html">Next: Extending the language: control flow</a>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://max99x.com">Max Shawabkeh</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
1607
www/src/kaleidoscope/PythonLangImpl5.html
Normal file
1607
www/src/kaleidoscope/PythonLangImpl5.html
Normal file
File diff suppressed because it is too large
Load diff
1605
www/src/kaleidoscope/PythonLangImpl6.html
Normal file
1605
www/src/kaleidoscope/PythonLangImpl6.html
Normal file
File diff suppressed because it is too large
Load diff
1905
www/src/kaleidoscope/PythonLangImpl7.html
Normal file
1905
www/src/kaleidoscope/PythonLangImpl7.html
Normal file
File diff suppressed because it is too large
Load diff
375
www/src/kaleidoscope/PythonLangImpl8.html
Normal file
375
www/src/kaleidoscope/PythonLangImpl8.html
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Conclusion and other useful LLVM tidbits</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Conclusion and other useful LLVM
|
||||
tidbits</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 8
|
||||
<ol>
|
||||
<li><a href="#conclusion">Tutorial Conclusion</a></li>
|
||||
<li><a href="#llvmirproperties">Properties of LLVM IR</a>
|
||||
<ul>
|
||||
<li><a href="#targetindep">Target Independence</a></li>
|
||||
<li><a href="#safety">Safety Guarantees</a></li>
|
||||
<li><a href="#langspecific">Language-Specific Optimizations</a></li>
|
||||
</ul>
|
||||
</li>
|
||||
<li><a href="#tipsandtricks">Tips and Tricks</a>
|
||||
<ul>
|
||||
<li><a href="#offsetofsizeof">Implementing portable
|
||||
offsetof/sizeof</a></li>
|
||||
<li><a href="#gcstack">Garbage Collected Stack Frames</a></li>
|
||||
</ul>
|
||||
</li>
|
||||
</ol>
|
||||
</li>
|
||||
</ul>
|
||||
|
||||
|
||||
<div class="doc_author">
|
||||
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a></p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="conclusion">Tutorial Conclusion</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Welcome to the the final chapter of the
|
||||
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
|
||||
with LLVM</a>" tutorial. In the course of this tutorial, we have grown
|
||||
our little Kaleidoscope language from being a useless toy, to being a
|
||||
semi-interesting (but probably still useless) toy. :)</p>
|
||||
|
||||
<p>It is interesting to see how far we've come, and how little code it has
|
||||
taken. We built the entire lexer, parser, AST, code generator, and an
|
||||
interactive run-loop (with a JIT!) by-hand in under 540 lines of
|
||||
(non-comment/non-blank) code.</p>
|
||||
|
||||
<p>Our little language supports a couple of interesting features: it supports
|
||||
user defined binary and unary operators, it uses JIT compilation for immediate
|
||||
evaluation, and it supports a few control flow constructs with SSA construction.
|
||||
</p>
|
||||
|
||||
<p>Part of the idea of this tutorial was to show you how easy and fun it can be
|
||||
to define, build, and play with languages. Building a compiler need not be a
|
||||
scary or mystical process! Now that you've seen some of the basics, I strongly
|
||||
encourage you to take the code and hack on it. For example, try adding:</p>
|
||||
|
||||
<ul>
|
||||
<li><b>global variables</b> - While global variables have questional value in
|
||||
modern software engineering, they are often useful when putting together quick
|
||||
little hacks like the Kaleidoscope compiler itself. Fortunately, our current
|
||||
setup makes it very easy to add global variables: just have value lookup check
|
||||
to see if an unresolved variable is in the global variable symbol table before
|
||||
rejecting it. To create a new global variable, make an instance of the LLVM
|
||||
<tt>GlobalVariable</tt> class.</li>
|
||||
|
||||
<li><b>typed variables</b> - Kaleidoscope currently only supports variables of
|
||||
type double. This gives the language a very nice elegance, because only
|
||||
supporting one type means that you never have to specify types. Different
|
||||
languages have different ways of handling this. The easiest way is to require
|
||||
the user to specify types for every variable definition, and record the type
|
||||
of the variable in the symbol table along with its Value*.</li>
|
||||
|
||||
<li><b>arrays, structs, vectors, etc</b> - Once you add types, you can start
|
||||
extending the type system in all sorts of interesting ways. Simple arrays are
|
||||
very easy and are quite useful for many different applications. Adding them is
|
||||
mostly an exercise in learning how the LLVM <a
|
||||
href="http://www.llvm.org/docs/LangRef.html#i_getelementptr">getelementptr</a>
|
||||
instruction works: it is so nifty/unconventional, it <a
|
||||
href="http://www.llvm.org/docs/GetElementPtr.html">has its own FAQ</a>! If you
|
||||
add support for recursive types (e.g. linked lists), make sure to read the <a
|
||||
href="http://www.llvm.org/docs/ProgrammersManual.html#TypeResolve">section in
|
||||
the LLVM Programmer's Manual</a> that describes how to construct them.</li>
|
||||
|
||||
<li><b>standard runtime</b> - Our current language allows the user to access
|
||||
arbitrary external functions, and we use it for things like "putchard". As you
|
||||
extend the language to add higher-level constructs, often these constructs make
|
||||
the most sense if they are lowered to calls into a language-supplied runtime.
|
||||
For example, if you add hash tables to the language, it would probably make
|
||||
sense to add the routines to a runtime, instead of inlining them all the way.
|
||||
</li>
|
||||
|
||||
<li><b>memory management</b> - Currently we can only access the stack in
|
||||
Kaleidoscope. It would also be useful to be able to allocate heap memory,
|
||||
either with calls to the standard libc malloc/free interface or with a garbage
|
||||
collector. If you would like to use garbage collection, note that LLVM fully
|
||||
supports <a href="http://www.llvm.org/docs/GarbageCollection.html">Accurate
|
||||
Garbage Collection</a> including algorithms that move objects and need to
|
||||
scan/update the stack.</li>
|
||||
|
||||
<li><b>debugger support</b> - LLVM supports generation of <a
|
||||
href="http://www.llvm.org/docs/SourceLevelDebugging.html">DWARF Debug info</a>
|
||||
which is understood by common debuggers like GDB. Adding support for debug info
|
||||
is fairly straightforward. The best way to understand it is to compile some
|
||||
C/C++ code with "<tt>llvm-gcc -g -O0</tt>" and taking a look at what it
|
||||
produces.</li>
|
||||
|
||||
<li><b>exception handling support</b> - LLVM supports generation of <a
|
||||
href="http://www.llvm.org/docs/ExceptionHandling.html">zero cost exceptions</a>
|
||||
which interoperate with code compiled in other languages. You could also
|
||||
generate code by implicitly making every function return an error value and
|
||||
checking it. You could also make explicit use of setjmp/longjmp. There are
|
||||
many different ways to go here.</li>
|
||||
|
||||
<li><b>object orientation, generics, database access, complex numbers,
|
||||
geometric programming, ...</b> - Really, there is
|
||||
no end of crazy features that you can add to the language.</li>
|
||||
|
||||
<li><b>unusual domains</b> - We've been talking about applying LLVM to a domain
|
||||
that many people are interested in: building a compiler for a specific language.
|
||||
However, there are many other domains that can use compiler technology that are
|
||||
not typically considered. For example, LLVM has been used to implement OpenGL
|
||||
graphics acceleration, translate C++ code to ActionScript, and many other
|
||||
cute and clever things. Maybe you will be the first to JIT compile a regular
|
||||
expression interpreter into native code with LLVM?</li>
|
||||
|
||||
</ul>
|
||||
|
||||
<p>
|
||||
Have fun - try doing something crazy and unusual. Building a language like
|
||||
everyone else always has, is much less fun than trying something a little crazy
|
||||
or off the wall and seeing how it turns out. If you get stuck or want to talk
|
||||
about it, feel free to email the <a
|
||||
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
|
||||
list</a>: it has lots of people who are interested in languages and are often
|
||||
willing to help out.
|
||||
</p>
|
||||
|
||||
<p>Before we end this tutorial, I want to talk about some "tips and tricks" for
|
||||
generating LLVM IR. These are some of the more subtle things that may not be
|
||||
obvious, but are very useful if you want to take advantage of LLVM's
|
||||
capabilities.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="llvmirproperties">Properties of the LLVM
|
||||
IR</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>We have a couple common questions about code in the LLVM IR form - let's just
|
||||
get these out of the way right now, shall we?</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="targetindep">Target
|
||||
Independence</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Kaleidoscope is an example of a "portable language": any program written in
|
||||
Kaleidoscope will work the same way on any target that it runs on. Many other
|
||||
languages have this property, e.g. LISP, Java, Haskell, Javascript, Python, etc.
|
||||
(note that while these languages are portable, not all their libraries are).</p>
|
||||
|
||||
<p>One nice aspect of LLVM is that it is often capable of preserving target
|
||||
independence in the IR: you can take the LLVM IR for a Kaleidoscope-compiled
|
||||
program and run it on any target that LLVM supports, even emitting C code and
|
||||
compiling that on targets that LLVM doesn't support natively. You can trivially
|
||||
tell that the Kaleidoscope compiler generates target-independent code because it
|
||||
never queries for any target-specific information when generating code.</p>
|
||||
|
||||
<p>The fact that LLVM provides a compact, target-independent, representation for
|
||||
code gets a lot of people excited. Unfortunately, these people are usually
|
||||
thinking about C or a language from the C family when they are asking questions
|
||||
about language portability. I say "unfortunately", because there is really no
|
||||
way to make (fully general) C code portable, other than shipping the source code
|
||||
around (and of course, C source code is not actually portable in general
|
||||
either - ever port a really old application from 32- to 64-bits?).</p>
|
||||
|
||||
<p>The problem with C (again, in its full generality) is that it is heavily
|
||||
laden with target specific assumptions. As one simple example, the preprocessor
|
||||
often destructively removes target-independence from the code when it processes
|
||||
the input text:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#ifdef __i386__
|
||||
int X = 1;
|
||||
#else
|
||||
int X = 42;
|
||||
#endif
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>While it is possible to engineer more and more complex solutions to problems
|
||||
like this, it cannot be solved in full generality in a way that is better than
|
||||
shipping the actual source code.</p>
|
||||
|
||||
<p>That said, there are interesting subsets of C that can be made portable. If
|
||||
you are willing to fix primitive types to a fixed size (say int = 32-bits,
|
||||
and long = 64-bits), don't care about ABI compatibility with existing binaries,
|
||||
and are willing to give up some other minor features, you can have portable
|
||||
code. This can make sense for specialized domains such as an
|
||||
in-kernel language.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="safety">Safety Guarantees</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Many of the languages above are also "safe" languages: it is impossible for
|
||||
a program written in Java to corrupt its address space and crash the process
|
||||
(assuming the JVM has no bugs).
|
||||
Safety is an interesting property that requires a combination of language
|
||||
design, runtime support, and often operating system support.</p>
|
||||
|
||||
<p>It is certainly possible to implement a safe language in LLVM, but LLVM IR
|
||||
does not itself guarantee safety. The LLVM IR allows unsafe pointer casts,
|
||||
use after free bugs, buffer over-runs, and a variety of other problems. Safety
|
||||
needs to be implemented as a layer on top of LLVM and, conveniently, several
|
||||
groups have investigated this. Ask on the <a
|
||||
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
|
||||
list</a> if you are interested in more details.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="langspecific">Language-Specific
|
||||
Optimizations</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>One thing about LLVM that turns off many people is that it does not solve all
|
||||
the world's problems in one system (sorry 'world hunger', someone else will have
|
||||
to solve you some other day). One specific complaint is that people perceive
|
||||
LLVM as being incapable of performing high-level language-specific optimization:
|
||||
LLVM "loses too much information".</p>
|
||||
|
||||
<p>Unfortunately, this is really not the place to give you a full and unified
|
||||
version of "Chris Lattner's theory of compiler design". Instead, I'll make a
|
||||
few observations:</p>
|
||||
|
||||
<p>First, you're right that LLVM does lose information. For example, as of this
|
||||
writing, there is no way to distinguish in the LLVM IR whether an SSA-value came
|
||||
from a C "int" or a C "long" on an ILP32 machine (other than debug info). Both
|
||||
get compiled down to an 'i32' value and the information about what it came from
|
||||
is lost. The more general issue here, is that the LLVM type system uses
|
||||
"structural equivalence" instead of "name equivalence". Another place this
|
||||
surprises people is if you have two types in a high-level language that have the
|
||||
same structure (e.g. two different structs that have a single int field): these
|
||||
types will compile down into a single LLVM type and it will be impossible to
|
||||
tell what it came from.</p>
|
||||
|
||||
<p>Second, while LLVM does lose information, LLVM is not a fixed target: we
|
||||
continue to enhance and improve it in many different ways. In addition to
|
||||
adding new features (LLVM did not always support exceptions or debug info), we
|
||||
also extend the IR to capture important information for optimization (e.g.
|
||||
whether an argument is sign or zero extended, information about pointers
|
||||
aliasing, etc). Many of the enhancements are user-driven: people want LLVM to
|
||||
include some specific feature, so they go ahead and extend it.</p>
|
||||
|
||||
<p>Third, it is <em>possible and easy</em> to add language-specific
|
||||
optimizations, and you have a number of choices in how to do it. As one trivial
|
||||
example, it is easy to add language-specific optimization passes that
|
||||
"know" things about code compiled for a language. In the case of the C family,
|
||||
there is an optimization pass that "knows" about the standard C library
|
||||
functions. If you call "exit(0)" in main(), it knows that it is safe to
|
||||
optimize that into "return 0;" because C specifies what the 'exit'
|
||||
function does.</p>
|
||||
|
||||
<p>In addition to simple library knowledge, it is possible to embed a variety of
|
||||
other language-specific information into the LLVM IR. If you have a specific
|
||||
need and run into a wall, please bring the topic up on the llvmdev list. At the
|
||||
very worst, you can always treat LLVM as if it were a "dumb code generator" and
|
||||
implement the high-level optimizations you desire in your front-end, on the
|
||||
language-specific AST.
|
||||
</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="tipsandtricks">Tips and Tricks</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>There is a variety of useful tips and tricks that you come to know after
|
||||
working on/with LLVM that aren't obvious at first glance. Instead of letting
|
||||
everyone rediscover them, this section talks about some of these issues.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="offsetofsizeof">Implementing portable
|
||||
offsetof/sizeof</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>One interesting thing that comes up, if you are trying to keep the code
|
||||
generated by your compiler "target independent", is that you often need to know
|
||||
the size of some LLVM type or the offset of some field in an llvm structure.
|
||||
For example, you might need to pass the size of a type into a function that
|
||||
allocates memory.</p>
|
||||
|
||||
<p>Unfortunately, this can vary widely across targets: for example the width of
|
||||
a pointer is trivially target-specific. However, there is a <a
|
||||
href="http://nondot.org/sabre/LLVMNotes/SizeOf-OffsetOf-VariableSizedStructs.txt">clever
|
||||
way to use the getelementptr instruction</a> that allows you to compute this
|
||||
in a portable way.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="gcstack">Garbage Collected
|
||||
Stack Frames</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Some languages want to explicitly manage their stack frames, often so that
|
||||
they are garbage collected or to allow easy implementation of closures. There
|
||||
are often better ways to implement these features than explicit stack frames,
|
||||
but <a
|
||||
href="http://nondot.org/sabre/LLVMNotes/ExplicitlyManagedStackFrames.txt">LLVM
|
||||
does support them,</a> if you want. It requires your front-end to convert the
|
||||
code into <a
|
||||
href="http://en.wikipedia.org/wiki/Continuation-passing_style">Continuation
|
||||
Passing Style</a> and the use of tail calls (which LLVM also supports).</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -43,7 +43,7 @@ the llvm-py contributors.</p></div>
|
|||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -91,7 +91,7 @@ update and merge before sending patches etc.</p></div>
|
|||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -48,6 +48,7 @@ below). 0.6 works only with LLVM 2.7.</p></div>
|
|||
package.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="changelog">Changelog</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="listingblock">
|
||||
|
|
@ -129,10 +130,11 @@ package.</p></div>
|
|||
* Initial release.</tt></pre>
|
||||
</div></div>
|
||||
</div>
|
||||
</div>
|
||||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -31,9 +31,11 @@
|
|||
<div id="header">
|
||||
<h1>Examples and LLVM Tutorials</h1>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_examples">Examples</h2>
|
||||
<div class="sectionbody">
|
||||
<h3 id="_a_simple_function">A Simple Function</h3><div style="clear:left"></div>
|
||||
<div class="sect2">
|
||||
<h3 id="_a_simple_function">A Simple Function</h3>
|
||||
<div class="paragraph"><p>Let’s create a (LLVM) module containing a single function, corresponding
|
||||
to the <tt>C</tt> function:</p></div>
|
||||
<div class="listingblock">
|
||||
|
|
@ -106,7 +108,9 @@ entry:
|
|||
ret i32 %tmp
|
||||
}</tt></pre>
|
||||
</div></div>
|
||||
<h3 id="_adding_jit_compilation">Adding JIT Compilation</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_adding_jit_compilation">Adding JIT Compilation</h3>
|
||||
<div class="paragraph"><p>Let’s compile this function in-memory and run it.</p></div>
|
||||
<div class="listingblock">
|
||||
<div class="content"><!-- Generator: GNU source-highlight 3.1.3
|
||||
|
|
@ -151,75 +155,74 @@ retval <span style="color: #990000">=</span> ee<span style="color: #990000">.</s
|
|||
<pre><tt>returned 142</tt></pre>
|
||||
</div></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_llvm_tutorials">LLVM Tutorials</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p>The <a href="http://www.llvm.org/docs/tutorial/">LLVM tutorials</a> have been
|
||||
ported to llvm-py. Below are the links to the original LLVM tutorial and
|
||||
the corresponding Python code using llvm-py:</p></div>
|
||||
<div class="paragraph"><div class="title">Simple JIT Tutorials</div><p>(contributed by Sebastien Binet)</p></div>
|
||||
<div class="paragraph"><div class="title">Simple JIT Tutorials</div><p>The following JIT tutorials were contributed by Sebastien Binet.</p></div>
|
||||
<div class="olist arabic"><ol class="arabic">
|
||||
<li>
|
||||
<p>
|
||||
A First Function
|
||||
<a href="http://www.llvm.org/docs/tutorial/JITTutorial1.html">LLVM</a>
|
||||
<a href="examples/JITTutorial1.html">llvm-py</a>
|
||||
<a href="examples/JITTutorial1.html">A First Function</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
A More Complicated Function
|
||||
<a href="http://www.llvm.org/docs/tutorial/JITTutorial2.html">LLVM</a>
|
||||
<a href="examples/JITTutorial2.html">llvm-py</a>
|
||||
<a href="examples/JITTutorial2.html">A More Complicated Function</a>
|
||||
</p>
|
||||
</li>
|
||||
</ol></div>
|
||||
<div class="olist arabic"><div class="title">Kaleidoscope: Implementing a Language with LLVM</div><ol class="arabic">
|
||||
<div class="paragraph" id="kaleidoscope"><div class="title">Kaleidoscope: Implementing a Language with LLVM</div><p>The LLVM <a href="http://www.llvm.org/docs/tutorial/">Kaleidoscope</a> tutorial
|
||||
has been ported to llvm-py by Max Shawabkeh.</p></div>
|
||||
<div class="olist arabic"><ol class="arabic">
|
||||
<li>
|
||||
<p>
|
||||
Tutorial Introduction and the Lexer (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl1.html">Tutorial Introduction and the Lexer</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Implementing a Parser and AST (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl2.html">Implementing a Parser and AST</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Implementing Code Generation to LLVM IR (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl3.html">Implementing Code Generation to LLVM IR</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Adding JIT and Optimizer Support (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl4.html">Adding JIT and Optimizer Support</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Extending the language: control flow (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl5.html">Extending the language: control flow</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Extending the language: user-defined operators (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl6.html">Extending the language: user-defined operators</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Extending the language: mutable variables / SSA construction (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl7.html">Extending the language: mutable variables / SSA construction</a>
|
||||
</p>
|
||||
</li>
|
||||
<li>
|
||||
<p>
|
||||
Conclusion and other useful LLVM tidbits (TODO)
|
||||
<a href="kaleidoscope/PythonLangImpl8.html">Conclusion and other useful LLVM tidbits</a>
|
||||
</p>
|
||||
</li>
|
||||
</ol></div>
|
||||
</div>
|
||||
</div>
|
||||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -47,10 +47,19 @@ discover that any of these claims are wrong, feel free to send across
|
|||
a patch.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_news">News</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="dlist"><dl>
|
||||
<dt class="hdlist1">
|
||||
26-Sep-2010
|
||||
</dt>
|
||||
<dd>
|
||||
<p>
|
||||
LLVM tutorial <a href="examples.html#kaleidoscope">ported</a> by Max Shawabkeh!
|
||||
</p>
|
||||
</dd>
|
||||
<dt class="hdlist1">
|
||||
31-Aug-2010
|
||||
</dt>
|
||||
<dd>
|
||||
|
|
@ -60,10 +69,11 @@ a patch.</p></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div>
|
||||
</div>
|
||||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
389
www/web/kaleidoscope/PythonLangImpl1.html
Normal file
389
www/web/kaleidoscope/PythonLangImpl1.html
Normal file
|
|
@ -0,0 +1,389 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Tutorial Introduction and the Lexer</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<meta name="author" content="Max Shawabkeh">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Tutorial Introduction and the Lexer</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 1
|
||||
<ol>
|
||||
<li><a href="#intro">Tutorial Introduction</a></li>
|
||||
<li><a href="#language">The Basic Language</a></li>
|
||||
<li><a href="#lexer">The Lexer</a></li>
|
||||
</ol>
|
||||
</li>
|
||||
<li><a href="PythonLangImpl2.html">Chapter 2</a>: Implementing a Parser and
|
||||
AST</li>
|
||||
</ul>
|
||||
|
||||
<div class="doc_author">
|
||||
<p>
|
||||
Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
|
||||
and <a href="http://max99x.com">Max Shawabkeh</a>
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="intro">Tutorial Introduction</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Welcome to the "Implementing a language with LLVM" tutorial. This tutorial
|
||||
runs through the implementation of a simple language, showing how fun and
|
||||
easy it can be. This tutorial will get you up and started as well as help to
|
||||
build a framework you can extend to other languages. The code in this tutorial
|
||||
can also be used as a playground to hack on other LLVM specific things.
|
||||
</p>
|
||||
|
||||
<p>The goal of this tutorial is to progressively unveil our language, describing
|
||||
how it is built up over time. This will let us cover a fairly broad range of
|
||||
language design and LLVM-specific usage issues, showing and explaining the code
|
||||
for it all along the way, without overwhelming you with tons of details up
|
||||
front.</p>
|
||||
|
||||
<p>It is useful to point out ahead of time that this tutorial is really about
|
||||
teaching compiler techniques and LLVM specifically, <em>not</em> about teaching
|
||||
modern and sane software engineering principles. In practice, this means that
|
||||
we'll take a number of shortcuts to simplify the exposition. If you dig in and
|
||||
use the code as a basis for future projects, fixing its deficiencies shouldn't
|
||||
be hard.</p>
|
||||
|
||||
<p>We've tried to put this tutorial together in a way that makes chapters easy
|
||||
to skip over if you are already familiar with or are uninterested in the various
|
||||
pieces. The structure of the tutorial is:</p>
|
||||
|
||||
<ul>
|
||||
<li><b><a href="#language">Chapter #1</a>: Introduction to the Kaleidoscope
|
||||
language, and the definition of its Lexer</b> - This shows where we are going
|
||||
and the basic functionality that we want it to do. In order to make this
|
||||
tutorial maximally understandable and hackable, we choose to implement
|
||||
everything in Python instead of using lexer and parser generators. LLVM
|
||||
obviously works just fine with such tools, feel free to use one if you prefer.
|
||||
</li>
|
||||
<li><b><a href="PythonLangImpl2.html">Chapter #2</a>: Implementing a Parser and
|
||||
AST</b> - With the lexer in place, we can talk about parsing techniques and
|
||||
basic AST construction. This tutorial describes recursive descent parsing and
|
||||
operator precedence parsing. Nothing in Chapters 1 or 2 is LLVM-specific,
|
||||
the code doesn't even import the LLVM modules at this point. :)</li>
|
||||
<li><b><a href="PythonLangImpl3.html">Chapter #3</a>: Code generation to LLVM
|
||||
IR</b> - With the AST ready, we can show off how easy generation of LLVM IR
|
||||
really is.</li>
|
||||
<li><b><a href="PythonLangImpl4.html">Chapter #4</a>: Adding JIT and Optimizer
|
||||
Support</b> - Because a lot of people are interested in using LLVM as a JIT,
|
||||
we'll dive right into it and show you the 3 lines it takes to add JIT support.
|
||||
LLVM is also useful in many other ways, but this is one simple and "sexy" way
|
||||
to shows off its power. :)</li>
|
||||
<li><b><a href="PythonLangImpl5.html">Chapter #5</a>: Extending the Language:
|
||||
Control Flow</b> - With the language up and running, we show how to extend it
|
||||
with control flow operations (if/then/else and a 'for' loop). This gives us a
|
||||
chance to talk about simple SSA construction and control flow.</li>
|
||||
<li><b><a href="PythonLangImpl6.html">Chapter #6</a>: Extending the Language:
|
||||
User-defined Operators</b> - This is a silly but fun chapter that talks about
|
||||
extending the language to let the user program define their own arbitrary
|
||||
unary and binary operators (with assignable precedence!). This lets us build a
|
||||
significant piece of the "language" as library routines.</li>
|
||||
<li><b><a href="PythonLangImpl7.html">Chapter #7</a>: Extending the Language:
|
||||
Mutable Variables</b> - This chapter talks about adding user-defined local
|
||||
variables along with an assignment operator. The interesting part about this
|
||||
is how easy and trivial it is to construct SSA form in LLVM: no, LLVM does
|
||||
<em>not</em> require your front-end to construct SSA form!</li>
|
||||
<li><b><a href="PythonLangImpl8.html">Chapter #8</a>: Conclusion and other
|
||||
useful LLVM tidbits</b> - This chapter wraps up the series by talking about
|
||||
potential ways to extend the language, but also includes a bunch of pointers to
|
||||
info about "special topics" like adding garbage collection support, exceptions,
|
||||
debugging, support for "spaghetti stacks", and a bunch of other tips and
|
||||
tricks.</li>
|
||||
|
||||
</ul>
|
||||
|
||||
<p>By the end of the tutorial, we'll have written a bit less than 540 lines of
|
||||
non-comment, non-blank, lines of code. With this small amount of code, we'll
|
||||
have built up a very reasonable compiler for a non-trivial language including
|
||||
a hand-written lexer, parser, AST, as well as code generation support with a JIT
|
||||
compiler. While other systems may have interesting "hello world" tutorials,
|
||||
I think the breadth of this tutorial is a great testament to the strengths of
|
||||
LLVM and why you should consider it if you're interested in language or compiler
|
||||
design.</p>
|
||||
|
||||
<p>A note about this tutorial: we expect you to extend the language and play
|
||||
with it on your own. Take the code and go crazy hacking away at it, compilers
|
||||
don't need to be scary creatures - it can be a lot of fun to play with
|
||||
languages!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="language">The Basic Language</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>This tutorial will be illustrated with a toy language that we'll call
|
||||
"<a href="http://en.wikipedia.org/wiki/Kaleidoscope">Kaleidoscope</a>" (derived
|
||||
from "meaning beautiful, form, and view").
|
||||
Kaleidoscope is a procedural language that allows you to define functions, use
|
||||
conditionals, math, etc. Over the course of the tutorial, we'll extend
|
||||
Kaleidoscope to support the if/then/else construct, a for loop, user defined
|
||||
operators, JIT compilation with a simple command line interface, etc.</p>
|
||||
|
||||
<p>Because we want to keep things simple, the only datatype in Kaleidoscope is a
|
||||
64-bit floating point type. As such, all values are implicitly double precision
|
||||
and the language doesn't require type declarations. This gives the language a
|
||||
very nice and simple syntax. For example, the following simple example computes
|
||||
<a href="http://en.wikipedia.org/wiki/Fibonacci_number">Fibonacci numbers:</a>
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Compute the x'th fibonacci number.
|
||||
def fib(x)
|
||||
if x < 3 then
|
||||
1
|
||||
else
|
||||
fib(x-1)+fib(x-2)
|
||||
|
||||
# This expression will compute the 40th number.
|
||||
fib(40)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>We also allow Kaleidoscope to call into standard library functions (the LLVM
|
||||
JIT makes this completely trivial). This means that you can use the 'extern'
|
||||
keyword to define a function before you use it (this is also useful for mutually
|
||||
recursive functions). For example:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
extern sin(arg);
|
||||
extern cos(arg);
|
||||
extern atan2(arg1 arg2);
|
||||
|
||||
atan2(sin(0.4), cos(42))
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>A more interesting example is included in Chapter 6 where we write a little
|
||||
Kaleidoscope application that <a href="PythonLangImpl6.html#example">displays
|
||||
a Mandelbrot Set</a> at various levels of magnification.</p>
|
||||
|
||||
<p>Lets dive into the implementation of this language!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="lexer">The Lexer</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>When it comes to implementing a language, the first thing needed is
|
||||
the ability to process a text file and recognize what it says. The traditional
|
||||
way to do this is to use a "<a
|
||||
href="http://en.wikipedia.org/wiki/Lexical_analysis">lexer</a>" (aka 'scanner')
|
||||
to break the input up into "tokens". Each token returned by the lexer includes
|
||||
a token type and potentially some metadata (e.g. the numeric value of a number).
|
||||
First, we define the possibilities:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The lexer yields one of these types for each token.
|
||||
class EOFToken(object):
|
||||
pass
|
||||
|
||||
class DefToken(object):
|
||||
pass
|
||||
|
||||
class ExternToken(object):
|
||||
pass
|
||||
|
||||
class IdentifierToken(object):
|
||||
def __init__(self, name): self.name = name
|
||||
|
||||
class NumberToken(object):
|
||||
def __init__(self, value): self.value = value
|
||||
|
||||
class CharacterToken(object):
|
||||
def __init__(self, char): self.char = char
|
||||
def __eq__(self, other):
|
||||
return isinstance(other, CharacterToken) and self.char == other.char
|
||||
def __ne__(self, other): return not self == other
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Each token yielded by our lexer will be of one of the above types. For simple
|
||||
tokens that are always the same, like the "def" keyword, the lexer will yield
|
||||
<tt>DefToken()</tt>. Identifiers, numbers and characters, on the other
|
||||
hand, have extra data, so when the lexer encounteres the number 123.45, it will
|
||||
emit it as <tt>NumberToken(123.45)</tt>. An identifier <tt>foo</tt> will be
|
||||
emitted as <tt>IdentifierToken('foo')</tt>. And finally, an unknown character
|
||||
like '+' will be returned as <tt>CharacterToken('+')</tt>. You may notice that
|
||||
we overload the equality and inequality operators for the characters; this will
|
||||
later simplify character comparisons in the parser code.</p>
|
||||
|
||||
<p>The actual implementation of the lexer is a single function called
|
||||
<tt>Tokenize</tt>, which takes a string and
|
||||
<a href="http://docs.python.org/reference/simple_stmts.html#the-yield-statement">yields</a>
|
||||
tokens. For simplicity, we will use
|
||||
<a href="http://docs.python.org/library/re.html">regular
|
||||
expressions</a> to parse out the tokens. This is terribly inefficient, but
|
||||
perfectly sufficient for our needs.</p>
|
||||
|
||||
<p>First, we define the regular expressions for our tokens. Numbers and strings
|
||||
of digits, optionally followed by a period and another string of digits.
|
||||
Identifiers (and keywords) are alphanumeric string starting with a letter and
|
||||
comments are anything between a hash (<tt>#</tt>) and the end of the line.
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
import re
|
||||
|
||||
...
|
||||
|
||||
# Regular expressions that tokens and comments of our language.
|
||||
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
|
||||
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
|
||||
REGEX_COMMENT = re.compile('#.*')
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>
|
||||
Next, let's start defining the <tt>Tokenize</tt> function itself. The first
|
||||
thing we need to do is set up a loop that scans the string, while ignoring
|
||||
whitespace between tokens:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
def Tokenize(string):
|
||||
while string:
|
||||
# Skip whitespace.
|
||||
if string[0].isspace():
|
||||
string = string[1:]
|
||||
continue
|
||||
|
||||
...
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Next we want to find out what the next token is. For this we run the regexes
|
||||
we defined above on the remainder of the string. To simplify the rest of the
|
||||
code, we run all three regexes each time. As mentioned above, inefficiencies are
|
||||
ignored for the purpose of this tutorial:<p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Run regexes.
|
||||
comment_match = REGEX_COMMENT.match(string)
|
||||
number_match = REGEX_NUMBER.match(string)
|
||||
identifier_match = REGEX_IDENTIFIER.match(string)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Now se check if any of the regexes matched. For comments, we simply
|
||||
ignore the captured match:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# Check if any of the regexes matched and yield the appropriate result.
|
||||
if comment_match:
|
||||
comment = comment_match.group(0)
|
||||
string = string[len(comment):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>For numbers, we yield the captured match, converted to a float and tagged
|
||||
with the appropriate token type:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
elif number_match:
|
||||
number = number_match.group(0)
|
||||
yield NumberToken(float(number))
|
||||
string = string[len(number):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>The identifier case is a little more complex. We have to check for keywords
|
||||
to decide whether we have captured an identifier or a keyword:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
elif identifier_match:
|
||||
identifier = identifier_match.group(0)
|
||||
# Check if we matched a keyword.
|
||||
if identifier == 'def':
|
||||
yield DefToken()
|
||||
elif identifier == 'extern':
|
||||
yield ExternToken()
|
||||
else:
|
||||
yield IdentifierToken(identifier)
|
||||
string = string[len(identifier):]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Finally, if we haven't recognized a comment, a number of an identifier, we
|
||||
yield the current character as an "unknown character" token. This is used, for
|
||||
example, for operators like <tt>+</tt> or <tt>*</tt>:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
else:
|
||||
# Yield the unknown character.
|
||||
yield CharacterToken(string[0])
|
||||
string = string[1:]
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Once we're done with the
|
||||
loop, we return a final end-of-file token:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
yield EOFToken()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>With this, we have the complete lexer for the basic Kaleidoscope language
|
||||
(the <a href="PythonLangImpl2.html#code">full code listing</a> for the Lexer is
|
||||
available in the <a href="PythonLangImpl2.html">next chapter</a> of the
|
||||
tutorial). Next we'll <a href="PythonLangImpl2.html">build a simple parser that
|
||||
uses this to build an Abstract Syntax Tree</a>. When we have that, we'll
|
||||
include a driver so that you can use the lexer and parser together.
|
||||
</p>
|
||||
|
||||
<a href="PythonLangImpl2.html">Next: Implementing a Parser and AST</a>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://max99x.com">Max Shawabkeh</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
1097
www/web/kaleidoscope/PythonLangImpl2.html
Normal file
1097
www/web/kaleidoscope/PythonLangImpl2.html
Normal file
File diff suppressed because it is too large
Load diff
1119
www/web/kaleidoscope/PythonLangImpl3.html
Normal file
1119
www/web/kaleidoscope/PythonLangImpl3.html
Normal file
File diff suppressed because it is too large
Load diff
999
www/web/kaleidoscope/PythonLangImpl4.html
Normal file
999
www/web/kaleidoscope/PythonLangImpl4.html
Normal file
|
|
@ -0,0 +1,999 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Adding JIT and Optimizer Support</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<meta name="author" content="Max Shawabkeh">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Adding JIT and Optimizer Support</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 4
|
||||
<ol>
|
||||
<li><a href="#intro">Chapter 4 Introduction</a></li>
|
||||
<li><a href="#trivialconstfold">Trivial Constant Folding</a></li>
|
||||
<li><a href="#optimizerpasses">LLVM Optimization Passes</a></li>
|
||||
<li><a href="#jit">Adding a JIT Compiler</a></li>
|
||||
<li><a href="#code">Full Code Listing</a></li>
|
||||
</ol>
|
||||
</li>
|
||||
<li><a href="PythonLangImpl5.html">Chapter 5</a>: Extending the Language:
|
||||
Control Flow</li>
|
||||
</ul>
|
||||
|
||||
<div class="doc_author">
|
||||
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
|
||||
and <a href="http://max99x.com">Max Shawabkeh</a>
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="intro">Chapter 4 Introduction</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Welcome to Chapter 4 of the
|
||||
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
|
||||
with LLVM</a>" tutorial. Chapters 1-3 described the implementation of a simple
|
||||
language and added support for generating LLVM IR. This chapter describes
|
||||
two new techniques: adding optimizer support to your language, and adding JIT
|
||||
compiler support. These additions will demonstrate how to get nice, efficient
|
||||
code for the Kaleidoscope language.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="trivialconstfold">Trivial Constant
|
||||
Folding</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Our demonstration for Chapter 3 is elegant and easy to extend. Unfortunately,
|
||||
it does not produce wonderful code. The LLVM Builder, however, does give us
|
||||
obvious optimizations when compiling simple code:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) 1+2+x</b>
|
||||
Read function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 3.000000e+00, %x
|
||||
ret double %addtmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This code is not a literal transcription of the AST built by parsing the
|
||||
input. That would be:
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) 1+2+x</b>
|
||||
Read function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 2.000000e+00, 1.000000e+00
|
||||
%addtmp1 = fadd double %addtmp, %x
|
||||
ret double %addtmp1
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Constant folding, as seen above, in particular, is a very common and very
|
||||
important optimization: so much so that many language implementors implement
|
||||
constant folding support in their AST representation.</p>
|
||||
|
||||
<p>With LLVM, you don't need this support in the AST. Since all calls to build
|
||||
LLVM IR go through the LLVM IR builder, the builder itself checked to see if
|
||||
there was a constant folding opportunity when you call it. If so, it just does
|
||||
the constant fold and return the constant instead of creating an instruction.
|
||||
|
||||
<p>Well, that was easy :). In practice, we recommend always using
|
||||
<tt>llvm.core.Builder</tt> when generating code like this. It has no
|
||||
"syntactic overhead" for its use (you don't have to uglify your compiler with
|
||||
constant checks everywhere) and it can dramatically reduce the amount of
|
||||
LLVM IR that is generated in some cases (particular for languages with a macro
|
||||
preprocessor or that use a lot of constants).</p>
|
||||
|
||||
<p>On the other hand, the <tt>Builder</tt> is limited by the fact that it does
|
||||
all of its analysis inline with the code as it is built. If you take a slightly
|
||||
more complex example:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) (1+2+x)*(x+(1+2))</b>
|
||||
Read a function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double 3.000000e+00, %x ; <double> [#uses=1]
|
||||
%addtmp1 = fadd double %x, 3.000000e+00 ; <double> [#uses=1]
|
||||
%multmp = fmul double %addtmp, %addtmp1 ; <double> [#uses=1]
|
||||
ret double %multmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>In this case, the LHS and RHS of the multiplication are the same value. We'd
|
||||
really like to see this generate "<tt>tmp = x+3; result = tmp*tmp;</tt>" instead
|
||||
of computing "<tt>x+3</tt>" twice.</p>
|
||||
|
||||
<p>Unfortunately, no amount of local analysis will be able to detect and correct
|
||||
this. This requires two transformations: reassociation of expressions (to
|
||||
make the add's lexically identical) and Common Subexpression Elimination (CSE)
|
||||
to delete the redundant add instruction. Fortunately, LLVM provides a broad
|
||||
range of optimizations that you can use, in the form of "passes".</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="optimizerpasses">LLVM Optimization
|
||||
Passes</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>LLVM provides many optimization passes, which do many different sorts of
|
||||
things and have different tradeoffs. Unlike other systems, LLVM doesn't hold
|
||||
to the mistaken notion that one set of optimizations is right for all languages
|
||||
and for all situations. LLVM allows a compiler implementor to make complete
|
||||
decisions about what optimizations to use, in which order, and in what
|
||||
situation.</p>
|
||||
|
||||
<p>As a concrete example, LLVM supports both "whole module" passes, which look
|
||||
across as large of body of code as they can (often a whole file, but if run
|
||||
at link time, this can be a substantial portion of the whole program). It also
|
||||
supports and includes "per-function" passes which just operate on a single
|
||||
function at a time, without looking at other functions. For more information
|
||||
on passes and how they are run, see the
|
||||
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html">How to Write a
|
||||
Pass</a> document and the <a href="http://www.llvm.org/docs/Passes.html">List of
|
||||
LLVM Passes</a>.</p>
|
||||
|
||||
<p>For Kaleidoscope, we are currently generating functions on the fly, one at
|
||||
a time, as the user types them in. We aren't shooting for the ultimate
|
||||
optimization experience in this setting, but we also want to catch the easy and
|
||||
quick stuff where possible. As such, we will choose to run a few per-function
|
||||
optimizations as the user types the function in. If we wanted to make a "static
|
||||
Kaleidoscope compiler", we would use exactly the code we have now, except that
|
||||
we would defer running the optimizer until the entire file has been parsed.</p>
|
||||
|
||||
<p>In order to get per-function optimizations going, we need to set up a
|
||||
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html#passmanager">
|
||||
FunctionPassManager</a> to hold and organize the LLVM optimizations that we want
|
||||
to run. Once we have that, we can add a set of optimizations to run. The code
|
||||
looks like this:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The function optimization passes manager.
|
||||
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
|
||||
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
|
||||
...
|
||||
|
||||
def main():
|
||||
# Set up the optimizer pipeline. Start with registering info about how the
|
||||
# target lays out data structures.
|
||||
g_llvm_pass_manager.add(g_llvm_executor.target_data)
|
||||
# Do simple "peephole" optimizations and bit-twiddling optzns.
|
||||
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
|
||||
# Reassociate expressions.
|
||||
g_llvm_pass_manager.add(PASS_REASSOCIATE)
|
||||
# Eliminate Common SubExpressions.
|
||||
g_llvm_pass_manager.add(PASS_GVN)
|
||||
# Simplify the control flow graph (deleting unreachable blocks, etc).
|
||||
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
g_llvm_pass_manager.initialize()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This code defines a <tt>FunctionPassManager</tt>,
|
||||
<tt>g_llvm_pass_manager</tt>. Once it is set up, we use a series of "add" calls
|
||||
to add a bunch of LLVM passes. The first pass is basically boilerplate, it adds
|
||||
a pass so that later optimizations know how the data structures in the program
|
||||
are laid out. (The "<tt>g_llvm_executor</tt>" variable is related to the JIT,
|
||||
which we will get to in the next section.) In this case, we choose to add 4
|
||||
optimization passes. The passes we chose here are a pretty standard set of
|
||||
"cleanup" optimizations that are useful for a wide variety of code. I won't
|
||||
delve into what they do but, believe me, they are a good starting place :).</p>
|
||||
|
||||
<p>Once the pass manager is set up, we need to make use of it. We do this by
|
||||
running it after our newly created function is constructed (in
|
||||
<tt>FunctionNode.CodeGen</tt>), but before it is returned to the client:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
return_value = self.body.CodeGen()
|
||||
g_llvm_builder.ret(return_value)
|
||||
|
||||
# Validate the generated code, checking for consistency.
|
||||
function.verify()
|
||||
|
||||
<b># Optimize the function.
|
||||
g_llvm_pass_manager.run(function)</b>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>As you can see, this is pretty straightforward. The
|
||||
<tt>FunctionPassManager</tt> optimizes and updates the LLVM Function in place,
|
||||
improving (hopefully) its body. With this in place, we can try our test above
|
||||
again:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def test(x) (1+2+x)*(x+(1+2))</b>
|
||||
Read a function definition:
|
||||
define double @test(double %x) {
|
||||
entry:
|
||||
%addtmp = fadd double %x, 3.000000e+00 ; <double> [#uses=2]
|
||||
%multmp = fmul double %addtmp, %addtmp ; <double> [#uses=1]
|
||||
ret double %multmp
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>As expected, we now get our nicely optimized code, saving a floating point
|
||||
add instruction from every execution of this function.</p>
|
||||
|
||||
<p>LLVM provides a wide variety of optimizations that can be used in certain
|
||||
circumstances. Some
|
||||
<a href="http://www.llvm.org/docs/Passes.html">documentation about the various
|
||||
passes</a> is available, but it isn't very complete. Another good source of
|
||||
ideas can come from looking at the passes that <tt>llvm-gcc</tt> or
|
||||
<tt>llvm-ld</tt> run to get started. The "<tt>opt</tt>" tool allows you to
|
||||
experiment with passes from the command line, so you can see if they do
|
||||
anything.</p>
|
||||
|
||||
<p>Now that we have reasonable code coming out of our front-end, lets talk about
|
||||
executing it!</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="jit">Adding a JIT Compiler</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Code that is available in LLVM IR can have a wide variety of tools
|
||||
applied to it. For example, you can run optimizations on it (as we did above),
|
||||
you can dump it out in textual or binary forms, you can compile the code to an
|
||||
assembly file (.s) for some target, or you can JIT compile it. The nice thing
|
||||
about the LLVM IR representation is that it is the "common currency" between
|
||||
many different parts of the compiler.
|
||||
</p>
|
||||
|
||||
<p>In this section, we'll add JIT compiler support to our interpreter. The
|
||||
basic idea that we want for Kaleidoscope is to have the user enter function
|
||||
bodies as they do now, but immediately evaluate the top-level expressions they
|
||||
type in. For example, if they type in "1 + 2", we should evaluate and print
|
||||
out 3. If they define a function, they should be able to call it from the
|
||||
command line.</p>
|
||||
|
||||
<p>In order to do this, we first declare and initialize the JIT. This is done
|
||||
by adding and initializing a global variable:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This creates an abstract "Execution Engine" which can be either a JIT
|
||||
compiler or the LLVM interpreter. LLVM will automatically pick a JIT compiler
|
||||
for you if one is available for your platform, otherwise it will fall back to
|
||||
the interpreter.</p>
|
||||
|
||||
<p>Once the <tt>ExecutionEngine</tt> is created, the JIT is ready to be used.
|
||||
We can use the <tt>run_function</tt> method of the execution engine to execute
|
||||
a compiled function and get its return value. In our case, this means that we
|
||||
can change the code that parses a top-level expression to look like this:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
def HandleTopLevelExpression(self):
|
||||
try:
|
||||
function = self.ParseTopLevelExpr().CodeGen()
|
||||
result = g_llvm_executor.run_function(function, [])
|
||||
print 'Evaluated to:', result.as_real(Type.double())
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Recall that we compile top-level expressions into a self-contained LLVM
|
||||
function that takes no arguments and returns the computed double.</p>
|
||||
|
||||
<p>With just these two changes, lets see how Kaleidoscope works now!</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>4+5</b>
|
||||
Read a top level expression:
|
||||
define double @0() {
|
||||
entry:
|
||||
ret double 9.000000e+00
|
||||
}
|
||||
|
||||
Evaluated to: 9.0
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Well this looks like it is basically working. The dump of the function
|
||||
shows the "no argument function that always returns double" that we synthesize
|
||||
for each top-level expression that is typed in. This demonstrates very basic
|
||||
functionality, but can we do more?</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>def testfunc(x y) x + y*2</b>
|
||||
Read a function definition:
|
||||
define double @testfunc(double %x, double %y) {
|
||||
entry:
|
||||
%multmp = fmul double %y, 2.000000e+00 ; <double> [#uses=1]
|
||||
%addtmp = fadd double %multmp, %x ; <double> [#uses=1]
|
||||
ret double %addtmp
|
||||
}
|
||||
|
||||
ready> <b>testfunc(4, 10)</b>
|
||||
Read a top level expression:
|
||||
define double @0() {
|
||||
entry:
|
||||
%calltmp = call double @testfunc(double 4.000000e+00, double 1.000000e+01) ; <double> [#uses=1]
|
||||
ret double %calltmp
|
||||
}
|
||||
|
||||
<em>Evaluated to: 24.0</em>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>This illustrates that we can now call user code, but there is something a bit
|
||||
subtle going on here. Note that we only invoke the JIT on the anonymous
|
||||
functions that <em>call testfunc</em>, but we never invoked it
|
||||
on <em>testfunc</em> itself. What actually happened here is that the JIT
|
||||
scanned for all non-JIT'd functions transitively called from the anonymous
|
||||
function and compiled all of them before returning from <tt>run_function()</tt>.
|
||||
</p>
|
||||
|
||||
<p>The JIT provides a number of other more advanced interfaces for things like
|
||||
freeing allocated machine code, rejit'ing functions to update them, etc.
|
||||
However, even with this simple code, we get some surprisingly powerful
|
||||
capabilities - check this out (I removed the dump of the anonymous functions,
|
||||
you should get the idea by now :) :</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
ready> <b>extern sin(x)</b>
|
||||
Read an extern:
|
||||
declare double @sin(double)
|
||||
|
||||
ready> <b>extern cos(x)</b>
|
||||
Read an extern:
|
||||
declare double @cos(double)
|
||||
|
||||
ready> <b>sin(1.0)</b>
|
||||
<em>Evaluated to: 0.841470984808</em>
|
||||
|
||||
ready> <b>def foo(x) sin(x)*sin(x) + cos(x)*cos(x)</b>
|
||||
Read a function definition:
|
||||
define double @foo(double %x) {
|
||||
entry:
|
||||
%calltmp = call double @sin(double %x) ; <double> [#uses=1]
|
||||
%calltmp1 = call double @sin(double %x) ; <double> [#uses=1]
|
||||
%multmp = fmul double %calltmp, %calltmp1 ; <double> [#uses=1]
|
||||
%calltmp2 = call double @cos(double %x) ; <double> [#uses=1]
|
||||
%calltmp3 = call double @cos(double %x) ; <double> [#uses=1]
|
||||
%multmp4 = fmul double %calltmp2, %calltmp3 ; <double> [#uses=1]
|
||||
%addtmp = fadd double %multmp, %multmp4 ; <double> [#uses=1]
|
||||
ret double %addtmp
|
||||
}
|
||||
|
||||
ready> <b>foo(4.0)</b>
|
||||
<em>Evaluated to: 1.000000</em>
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Whoa, how does the JIT know about sin and cos? The answer is surprisingly
|
||||
simple: in this example, the JIT started execution of a function and got to a
|
||||
function call. It realized that the function was not yet JIT compiled and
|
||||
invoked the standard set of routines to resolve the function. In this case,
|
||||
there is no body defined for the function, so the JIT ended up calling
|
||||
"<tt>dlsym("sin")</tt>" on the Python process that is hosting our Kaleidoscope
|
||||
prompt. Since "<tt>sin</tt>" is defined within the JIT's address space, it
|
||||
simply patches up calls in the module to call the libm version of <tt>sin</tt>
|
||||
directly.</p>
|
||||
|
||||
<p>One interesting application of this is that we can now extend the language
|
||||
by writing arbitrary C++ code to implement operations. For example, we can
|
||||
create a C file with the following simple function:
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#include <stdio.h>
|
||||
|
||||
double putchard(double x) {
|
||||
putchar((char)x);
|
||||
return 0;
|
||||
}
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>We can then compile this into a shared library with GCC:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
gcc -shared -fPIC -o putchard.so putchard.c
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Now we can load this library into the Python process using
|
||||
<tt>llvm.core.load_library_permanently</tt> and access it from Kaleidoscope to
|
||||
produce simple output to the console:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
>>> <b>import llvm.core</b>
|
||||
>>> <b>llvm.core.load_library_permanently('/home/max/llvm-py-tutorial/putchard.so')</b>
|
||||
>>> <b>import kaleidoscope</b>
|
||||
>>> <b>kaleidoscope.main()</b>
|
||||
ready> <b>extern putchard(x)</b>
|
||||
Read an extern:
|
||||
declare double @putchard(double)
|
||||
|
||||
ready> <b>putchard(65) + putchard(66) + putchard(67) + putchard(10)</b>
|
||||
<em>ABC</em>
|
||||
Evaluated to: 0.0
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>Similar code could be used to implement file I/O, console input, and many
|
||||
other capabilities in Kaleidoscope.</p>
|
||||
|
||||
<p>This completes the JIT and optimizer chapter of the Kaleidoscope tutorial. At
|
||||
this point, we can compile a non-Turing-complete programming language, optimize
|
||||
and JIT compile it in a user-driven way. Next up we'll look into <a
|
||||
href="PythonLangImpl5.html">extending the language with control flow
|
||||
constructs</a>, tackling some interesting LLVM IR issues along the way.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="code">Full Code Listing</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>
|
||||
Here is the complete code listing for our running example, enhanced with the
|
||||
LLVM JIT and optimizer:
|
||||
</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#!/usr/bin/env python
|
||||
|
||||
import re
|
||||
from llvm.core import Module, Constant, Type, Function, Builder, FCMP_ULT
|
||||
from llvm.ee import ExecutionEngine, TargetData
|
||||
from llvm.passes import FunctionPassManager
|
||||
from llvm.passes import (PASS_INSTRUCTION_COMBINING,
|
||||
PASS_REASSOCIATE,
|
||||
PASS_GVN,
|
||||
PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
################################################################################
|
||||
## Globals
|
||||
################################################################################
|
||||
|
||||
# The LLVM module, which holds all the IR code.
|
||||
g_llvm_module = Module.new('my cool jit')
|
||||
|
||||
# The LLVM instruction builder. Created whenever a new function is entered.
|
||||
g_llvm_builder = None
|
||||
|
||||
# A dictionary that keeps track of which values are defined in the current scope
|
||||
# and what their LLVM representation is.
|
||||
g_named_values = {}
|
||||
|
||||
# The function optimization passes manager.
|
||||
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
|
||||
|
||||
# The LLVM execution engine.
|
||||
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
|
||||
|
||||
################################################################################
|
||||
## Lexer
|
||||
################################################################################
|
||||
|
||||
# The lexer yields one of these types for each token.
|
||||
class EOFToken(object):
|
||||
pass
|
||||
|
||||
class DefToken(object):
|
||||
pass
|
||||
|
||||
class ExternToken(object):
|
||||
pass
|
||||
|
||||
class IdentifierToken(object):
|
||||
def __init__(self, name): self.name = name
|
||||
|
||||
class NumberToken(object):
|
||||
def __init__(self, value): self.value = value
|
||||
|
||||
class CharacterToken(object):
|
||||
def __init__(self, char): self.char = char
|
||||
def __eq__(self, other):
|
||||
return isinstance(other, CharacterToken) and self.char == other.char
|
||||
def __ne__(self, other): return not self == other
|
||||
|
||||
# Regular expressions that tokens and comments of our language.
|
||||
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
|
||||
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
|
||||
REGEX_COMMENT = re.compile('#.*')
|
||||
|
||||
def Tokenize(string):
|
||||
while string:
|
||||
# Skip whitespace.
|
||||
if string[0].isspace():
|
||||
string = string[1:]
|
||||
continue
|
||||
|
||||
# Run regexes.
|
||||
comment_match = REGEX_COMMENT.match(string)
|
||||
number_match = REGEX_NUMBER.match(string)
|
||||
identifier_match = REGEX_IDENTIFIER.match(string)
|
||||
|
||||
# Check if any of the regexes matched and yield the appropriate result.
|
||||
if comment_match:
|
||||
comment = comment_match.group(0)
|
||||
string = string[len(comment):]
|
||||
elif number_match:
|
||||
number = number_match.group(0)
|
||||
yield NumberToken(float(number))
|
||||
string = string[len(number):]
|
||||
elif identifier_match:
|
||||
identifier = identifier_match.group(0)
|
||||
# Check if we matched a keyword.
|
||||
if identifier == 'def':
|
||||
yield DefToken()
|
||||
elif identifier == 'extern':
|
||||
yield ExternToken()
|
||||
else:
|
||||
yield IdentifierToken(identifier)
|
||||
string = string[len(identifier):]
|
||||
else:
|
||||
# Yield the ASCII value of the unknown character.
|
||||
yield CharacterToken(string[0])
|
||||
string = string[1:]
|
||||
|
||||
yield EOFToken()
|
||||
|
||||
################################################################################
|
||||
## Abstract Syntax Tree (aka Parse Tree)
|
||||
################################################################################
|
||||
|
||||
# Base class for all expression nodes.
|
||||
class ExpressionNode(object):
|
||||
pass
|
||||
|
||||
# Expression class for numeric literals like "1.0".
|
||||
class NumberExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, value):
|
||||
self.value = value
|
||||
|
||||
def CodeGen(self):
|
||||
return Constant.real(Type.double(), self.value)
|
||||
|
||||
# Expression class for referencing a variable, like "a".
|
||||
class VariableExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, name):
|
||||
self.name = name
|
||||
|
||||
def CodeGen(self):
|
||||
if self.name in g_named_values:
|
||||
return g_named_values[self.name]
|
||||
else:
|
||||
raise RuntimeError('Unknown variable name: ' + self.name)
|
||||
|
||||
# Expression class for a binary operator.
|
||||
class BinaryOperatorExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, operator, left, right):
|
||||
self.operator = operator
|
||||
self.left = left
|
||||
self.right = right
|
||||
|
||||
def CodeGen(self):
|
||||
left = self.left.CodeGen()
|
||||
right = self.right.CodeGen()
|
||||
|
||||
if self.operator == '+':
|
||||
return g_llvm_builder.fadd(left, right, 'addtmp')
|
||||
elif self.operator == '-':
|
||||
return g_llvm_builder.fsub(left, right, 'subtmp')
|
||||
elif self.operator == '*':
|
||||
return g_llvm_builder.fmul(left, right, 'multmp')
|
||||
elif self.operator == '<':
|
||||
result = g_llvm_builder.fcmp(FCMP_ULT, left, right, 'cmptmp')
|
||||
# Convert bool 0 or 1 to double 0.0 or 1.0.
|
||||
return g_llvm_builder.uitofp(result, Type.double(), 'booltmp')
|
||||
else:
|
||||
raise RuntimeError('Unknown binary operator.')
|
||||
|
||||
# Expression class for function calls.
|
||||
class CallExpressionNode(ExpressionNode):
|
||||
|
||||
def __init__(self, callee, args):
|
||||
self.callee = callee
|
||||
self.args = args
|
||||
|
||||
def CodeGen(self):
|
||||
# Look up the name in the global module table.
|
||||
callee = g_llvm_module.get_function_named(self.callee)
|
||||
|
||||
# Check for argument mismatch error.
|
||||
if len(callee.args) != len(self.args):
|
||||
raise RuntimeError('Incorrect number of arguments passed.')
|
||||
|
||||
arg_values = [i.CodeGen() for i in self.args]
|
||||
|
||||
return g_llvm_builder.call(callee, arg_values, 'calltmp')
|
||||
|
||||
# This class represents the "prototype" for a function, which captures its name,
|
||||
# and its argument names (thus implicitly the number of arguments the function
|
||||
# takes).
|
||||
class PrototypeNode(object):
|
||||
|
||||
def __init__(self, name, args):
|
||||
self.name = name
|
||||
self.args = args
|
||||
|
||||
def CodeGen(self):
|
||||
# Make the function type, eg. double(double,double).
|
||||
funct_type = Type.function(
|
||||
Type.double(), [Type.double()] * len(self.args), False)
|
||||
|
||||
function = Function.new(g_llvm_module, funct_type, self.name)
|
||||
|
||||
# If the name conflicted, there was already something with the same name.
|
||||
# If it has a body, don't allow redefinition or reextern.
|
||||
if function.name != self.name:
|
||||
function.delete()
|
||||
function = g_llvm_module.get_function_named(self.name)
|
||||
|
||||
# If the function already has a body, reject this.
|
||||
if not function.is_declaration:
|
||||
raise RuntimeError('Redefinition of function.')
|
||||
|
||||
# If F took a different number of args, reject.
|
||||
if len(callee.args) != len(self.args):
|
||||
raise RuntimeError('Redeclaration of a function with different number '
|
||||
'of args.')
|
||||
|
||||
# Set names for all arguments and add them to the variables symbol table.
|
||||
for arg, arg_name in zip(function.args, self.args):
|
||||
arg.name = arg_name
|
||||
# Add arguments to variable symbol table.
|
||||
g_named_values[arg_name] = arg
|
||||
|
||||
return function
|
||||
|
||||
# This class represents a function definition itself.
|
||||
class FunctionNode(object):
|
||||
|
||||
def __init__(self, prototype, body):
|
||||
self.prototype = prototype
|
||||
self.body = body
|
||||
|
||||
def CodeGen(self):
|
||||
# Clear scope.
|
||||
g_named_values.clear()
|
||||
|
||||
# Create a function object.
|
||||
function = self.prototype.CodeGen()
|
||||
|
||||
# Create a new basic block to start insertion into.
|
||||
block = function.append_basic_block('entry')
|
||||
global g_llvm_builder
|
||||
g_llvm_builder = Builder.new(block)
|
||||
|
||||
# Finish off the function.
|
||||
try:
|
||||
return_value = self.body.CodeGen()
|
||||
g_llvm_builder.ret(return_value)
|
||||
|
||||
# Validate the generated code, checking for consistency.
|
||||
function.verify()
|
||||
|
||||
# Optimize the function.
|
||||
g_llvm_pass_manager.run(function)
|
||||
except:
|
||||
function.delete()
|
||||
raise
|
||||
|
||||
return function
|
||||
|
||||
|
||||
################################################################################
|
||||
## Parser
|
||||
################################################################################
|
||||
|
||||
class Parser(object):
|
||||
|
||||
def __init__(self, tokens, binop_precedence):
|
||||
self.tokens = tokens
|
||||
self.binop_precedence = binop_precedence
|
||||
self.Next()
|
||||
|
||||
# Provide a simple token buffer. Parser.current is the current token the
|
||||
# parser is looking at. Parser.Next() reads another token from the lexer and
|
||||
# updates Parser.current with its results.
|
||||
def Next(self):
|
||||
self.current = self.tokens.next()
|
||||
|
||||
# Gets the precedence of the current token, or -1 if the token is not a binary
|
||||
# operator.
|
||||
def GetCurrentTokenPrecedence(self):
|
||||
if isinstance(self.current, CharacterToken):
|
||||
return self.binop_precedence.get(self.current.char, -1)
|
||||
else:
|
||||
return -1
|
||||
|
||||
# identifierexpr ::= identifier | identifier '(' expression* ')'
|
||||
def ParseIdentifierExpr(self):
|
||||
identifier_name = self.current.name
|
||||
self.Next() # eat identifier.
|
||||
|
||||
if self.current != CharacterToken('('): # Simple variable reference.
|
||||
return VariableExpressionNode(identifier_name)
|
||||
|
||||
# Call.
|
||||
self.Next() # eat '('.
|
||||
args = []
|
||||
if self.current != CharacterToken(')'):
|
||||
while True:
|
||||
args.append(self.ParseExpression())
|
||||
if self.current == CharacterToken(')'):
|
||||
break
|
||||
elif self.current != CharacterToken(','):
|
||||
raise RuntimeError('Expected ")" or "," in argument list.')
|
||||
self.Next()
|
||||
|
||||
self.Next() # eat ')'.
|
||||
return CallExpressionNode(identifier_name, args)
|
||||
|
||||
# numberexpr ::= number
|
||||
def ParseNumberExpr(self):
|
||||
result = NumberExpressionNode(self.current.value)
|
||||
self.Next() # consume the number.
|
||||
return result
|
||||
|
||||
# parenexpr ::= '(' expression ')'
|
||||
def ParseParenExpr(self):
|
||||
self.Next() # eat '('.
|
||||
|
||||
contents = self.ParseExpression()
|
||||
|
||||
if self.current != CharacterToken(')'):
|
||||
raise RuntimeError('Expected ")".')
|
||||
self.Next() # eat ')'.
|
||||
|
||||
return contents
|
||||
|
||||
# primary ::= identifierexpr | numberexpr | parenexpr
|
||||
def ParsePrimary(self):
|
||||
if isinstance(self.current, IdentifierToken):
|
||||
return self.ParseIdentifierExpr()
|
||||
elif isinstance(self.current, NumberToken):
|
||||
return self.ParseNumberExpr()
|
||||
elif self.current == CharacterToken('('):
|
||||
return self.ParseParenExpr()
|
||||
else:
|
||||
raise RuntimeError('Unknown token when expecting an expression.')
|
||||
|
||||
# binoprhs ::= (operator primary)*
|
||||
def ParseBinOpRHS(self, left, left_precedence):
|
||||
# If this is a binary operator, find its precedence.
|
||||
while True:
|
||||
precedence = self.GetCurrentTokenPrecedence()
|
||||
|
||||
# If this is a binary operator that binds at least as tightly as the
|
||||
# current one, consume it; otherwise we are done.
|
||||
if precedence < left_precedence:
|
||||
return left
|
||||
|
||||
binary_operator = self.current.char
|
||||
self.Next() # eat the operator.
|
||||
|
||||
# Parse the primary expression after the binary operator.
|
||||
right = self.ParsePrimary()
|
||||
|
||||
# If binary_operator binds less tightly with right than the operator after
|
||||
# right, let the pending operator take right as its left.
|
||||
next_precedence = self.GetCurrentTokenPrecedence()
|
||||
if precedence < next_precedence:
|
||||
right = self.ParseBinOpRHS(right, precedence + 1)
|
||||
|
||||
# Merge left/right.
|
||||
left = BinaryOperatorExpressionNode(binary_operator, left, right)
|
||||
|
||||
# expression ::= primary binoprhs
|
||||
def ParseExpression(self):
|
||||
left = self.ParsePrimary()
|
||||
return self.ParseBinOpRHS(left, 0)
|
||||
|
||||
# prototype ::= id '(' id* ')'
|
||||
def ParsePrototype(self):
|
||||
if not isinstance(self.current, IdentifierToken):
|
||||
raise RuntimeError('Expected function name in prototype.')
|
||||
|
||||
function_name = self.current.name
|
||||
self.Next() # eat function name.
|
||||
|
||||
if self.current != CharacterToken('('):
|
||||
raise RuntimeError('Expected "(" in prototype.')
|
||||
self.Next() # eat '('.
|
||||
|
||||
arg_names = []
|
||||
while isinstance(self.current, IdentifierToken):
|
||||
arg_names.append(self.current.name)
|
||||
self.Next()
|
||||
|
||||
if self.current != CharacterToken(')'):
|
||||
raise RuntimeError('Expected ")" in prototype.')
|
||||
|
||||
# Success.
|
||||
self.Next() # eat ')'.
|
||||
|
||||
return PrototypeNode(function_name, arg_names)
|
||||
|
||||
# definition ::= 'def' prototype expression
|
||||
def ParseDefinition(self):
|
||||
self.Next() # eat def.
|
||||
proto = self.ParsePrototype()
|
||||
body = self.ParseExpression()
|
||||
return FunctionNode(proto, body)
|
||||
|
||||
# toplevelexpr ::= expression
|
||||
def ParseTopLevelExpr(self):
|
||||
proto = PrototypeNode('', [])
|
||||
return FunctionNode(proto, self.ParseExpression())
|
||||
|
||||
# external ::= 'extern' prototype
|
||||
def ParseExtern(self):
|
||||
self.Next() # eat extern.
|
||||
return self.ParsePrototype()
|
||||
|
||||
# Top-Level parsing
|
||||
def HandleDefinition(self):
|
||||
self.Handle(self.ParseDefinition, 'Read a function definition:')
|
||||
|
||||
def HandleExtern(self):
|
||||
self.Handle(self.ParseExtern, 'Read an extern:')
|
||||
|
||||
def HandleTopLevelExpression(self):
|
||||
try:
|
||||
function = self.ParseTopLevelExpr().CodeGen()
|
||||
result = g_llvm_executor.run_function(function, [])
|
||||
print 'Evaluated to:', result.as_real(Type.double())
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
|
||||
def Handle(self, function, message):
|
||||
try:
|
||||
print message, function().CodeGen()
|
||||
except Exception, e:
|
||||
print 'Error:', e
|
||||
try:
|
||||
self.Next() # Skip for error recovery.
|
||||
except:
|
||||
pass
|
||||
|
||||
################################################################################
|
||||
## Main driver code.
|
||||
################################################################################
|
||||
|
||||
def main():
|
||||
# Set up the optimizer pipeline. Start with registering info about how the
|
||||
# target lays out data structures.
|
||||
g_llvm_pass_manager.add(g_llvm_executor.target_data)
|
||||
# Do simple "peephole" optimizations and bit-twiddling optzns.
|
||||
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
|
||||
# Reassociate expressions.
|
||||
g_llvm_pass_manager.add(PASS_REASSOCIATE)
|
||||
# Eliminate Common SubExpressions.
|
||||
g_llvm_pass_manager.add(PASS_GVN)
|
||||
# Simplify the control flow graph (deleting unreachable blocks, etc).
|
||||
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
|
||||
|
||||
g_llvm_pass_manager.initialize()
|
||||
|
||||
# Install standard binary operators.
|
||||
# 1 is lowest possible precedence. 40 is the highest.
|
||||
operator_precedence = {
|
||||
'<': 10,
|
||||
'+': 20,
|
||||
'-': 20,
|
||||
'*': 40
|
||||
}
|
||||
|
||||
# Run the main "interpreter loop".
|
||||
while True:
|
||||
print 'ready>',
|
||||
try:
|
||||
raw = raw_input()
|
||||
except KeyboardInterrupt:
|
||||
break
|
||||
|
||||
parser = Parser(Tokenize(raw), operator_precedence)
|
||||
while True:
|
||||
# top ::= definition | external | expression | EOF
|
||||
if isinstance(parser.current, EOFToken):
|
||||
break
|
||||
if isinstance(parser.current, DefToken):
|
||||
parser.HandleDefinition()
|
||||
elif isinstance(parser.current, ExternToken):
|
||||
parser.HandleExtern()
|
||||
else:
|
||||
parser.HandleTopLevelExpression()
|
||||
|
||||
# Print out all of the generated code.
|
||||
print '\n', g_llvm_module
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<a href="PythonLangImpl5.html">Next: Extending the language: control flow</a>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://max99x.com">Max Shawabkeh</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
1607
www/web/kaleidoscope/PythonLangImpl5.html
Normal file
1607
www/web/kaleidoscope/PythonLangImpl5.html
Normal file
File diff suppressed because it is too large
Load diff
1605
www/web/kaleidoscope/PythonLangImpl6.html
Normal file
1605
www/web/kaleidoscope/PythonLangImpl6.html
Normal file
File diff suppressed because it is too large
Load diff
1905
www/web/kaleidoscope/PythonLangImpl7.html
Normal file
1905
www/web/kaleidoscope/PythonLangImpl7.html
Normal file
File diff suppressed because it is too large
Load diff
375
www/web/kaleidoscope/PythonLangImpl8.html
Normal file
375
www/web/kaleidoscope/PythonLangImpl8.html
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
|
||||
"http://www.w3.org/TR/html4/strict.dtd">
|
||||
|
||||
<html>
|
||||
<head>
|
||||
<title>Kaleidoscope: Conclusion and other useful LLVM tidbits</title>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
|
||||
<meta name="author" content="Chris Lattner">
|
||||
<link rel="stylesheet"
|
||||
href="http://www.llvm.org/docs/llvm.css"
|
||||
type="text/css">
|
||||
</head>
|
||||
|
||||
<body>
|
||||
|
||||
<div class="doc_title">Kaleidoscope: Conclusion and other useful LLVM
|
||||
tidbits</div>
|
||||
|
||||
<ul>
|
||||
<li>
|
||||
<a href="http://www.llvm.org/docs/tutorial/index.html">
|
||||
Up to Tutorial Index
|
||||
</a>
|
||||
</li>
|
||||
<li>Chapter 8
|
||||
<ol>
|
||||
<li><a href="#conclusion">Tutorial Conclusion</a></li>
|
||||
<li><a href="#llvmirproperties">Properties of LLVM IR</a>
|
||||
<ul>
|
||||
<li><a href="#targetindep">Target Independence</a></li>
|
||||
<li><a href="#safety">Safety Guarantees</a></li>
|
||||
<li><a href="#langspecific">Language-Specific Optimizations</a></li>
|
||||
</ul>
|
||||
</li>
|
||||
<li><a href="#tipsandtricks">Tips and Tricks</a>
|
||||
<ul>
|
||||
<li><a href="#offsetofsizeof">Implementing portable
|
||||
offsetof/sizeof</a></li>
|
||||
<li><a href="#gcstack">Garbage Collected Stack Frames</a></li>
|
||||
</ul>
|
||||
</li>
|
||||
</ol>
|
||||
</li>
|
||||
</ul>
|
||||
|
||||
|
||||
<div class="doc_author">
|
||||
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a></p>
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="conclusion">Tutorial Conclusion</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Welcome to the the final chapter of the
|
||||
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
|
||||
with LLVM</a>" tutorial. In the course of this tutorial, we have grown
|
||||
our little Kaleidoscope language from being a useless toy, to being a
|
||||
semi-interesting (but probably still useless) toy. :)</p>
|
||||
|
||||
<p>It is interesting to see how far we've come, and how little code it has
|
||||
taken. We built the entire lexer, parser, AST, code generator, and an
|
||||
interactive run-loop (with a JIT!) by-hand in under 540 lines of
|
||||
(non-comment/non-blank) code.</p>
|
||||
|
||||
<p>Our little language supports a couple of interesting features: it supports
|
||||
user defined binary and unary operators, it uses JIT compilation for immediate
|
||||
evaluation, and it supports a few control flow constructs with SSA construction.
|
||||
</p>
|
||||
|
||||
<p>Part of the idea of this tutorial was to show you how easy and fun it can be
|
||||
to define, build, and play with languages. Building a compiler need not be a
|
||||
scary or mystical process! Now that you've seen some of the basics, I strongly
|
||||
encourage you to take the code and hack on it. For example, try adding:</p>
|
||||
|
||||
<ul>
|
||||
<li><b>global variables</b> - While global variables have questional value in
|
||||
modern software engineering, they are often useful when putting together quick
|
||||
little hacks like the Kaleidoscope compiler itself. Fortunately, our current
|
||||
setup makes it very easy to add global variables: just have value lookup check
|
||||
to see if an unresolved variable is in the global variable symbol table before
|
||||
rejecting it. To create a new global variable, make an instance of the LLVM
|
||||
<tt>GlobalVariable</tt> class.</li>
|
||||
|
||||
<li><b>typed variables</b> - Kaleidoscope currently only supports variables of
|
||||
type double. This gives the language a very nice elegance, because only
|
||||
supporting one type means that you never have to specify types. Different
|
||||
languages have different ways of handling this. The easiest way is to require
|
||||
the user to specify types for every variable definition, and record the type
|
||||
of the variable in the symbol table along with its Value*.</li>
|
||||
|
||||
<li><b>arrays, structs, vectors, etc</b> - Once you add types, you can start
|
||||
extending the type system in all sorts of interesting ways. Simple arrays are
|
||||
very easy and are quite useful for many different applications. Adding them is
|
||||
mostly an exercise in learning how the LLVM <a
|
||||
href="http://www.llvm.org/docs/LangRef.html#i_getelementptr">getelementptr</a>
|
||||
instruction works: it is so nifty/unconventional, it <a
|
||||
href="http://www.llvm.org/docs/GetElementPtr.html">has its own FAQ</a>! If you
|
||||
add support for recursive types (e.g. linked lists), make sure to read the <a
|
||||
href="http://www.llvm.org/docs/ProgrammersManual.html#TypeResolve">section in
|
||||
the LLVM Programmer's Manual</a> that describes how to construct them.</li>
|
||||
|
||||
<li><b>standard runtime</b> - Our current language allows the user to access
|
||||
arbitrary external functions, and we use it for things like "putchard". As you
|
||||
extend the language to add higher-level constructs, often these constructs make
|
||||
the most sense if they are lowered to calls into a language-supplied runtime.
|
||||
For example, if you add hash tables to the language, it would probably make
|
||||
sense to add the routines to a runtime, instead of inlining them all the way.
|
||||
</li>
|
||||
|
||||
<li><b>memory management</b> - Currently we can only access the stack in
|
||||
Kaleidoscope. It would also be useful to be able to allocate heap memory,
|
||||
either with calls to the standard libc malloc/free interface or with a garbage
|
||||
collector. If you would like to use garbage collection, note that LLVM fully
|
||||
supports <a href="http://www.llvm.org/docs/GarbageCollection.html">Accurate
|
||||
Garbage Collection</a> including algorithms that move objects and need to
|
||||
scan/update the stack.</li>
|
||||
|
||||
<li><b>debugger support</b> - LLVM supports generation of <a
|
||||
href="http://www.llvm.org/docs/SourceLevelDebugging.html">DWARF Debug info</a>
|
||||
which is understood by common debuggers like GDB. Adding support for debug info
|
||||
is fairly straightforward. The best way to understand it is to compile some
|
||||
C/C++ code with "<tt>llvm-gcc -g -O0</tt>" and taking a look at what it
|
||||
produces.</li>
|
||||
|
||||
<li><b>exception handling support</b> - LLVM supports generation of <a
|
||||
href="http://www.llvm.org/docs/ExceptionHandling.html">zero cost exceptions</a>
|
||||
which interoperate with code compiled in other languages. You could also
|
||||
generate code by implicitly making every function return an error value and
|
||||
checking it. You could also make explicit use of setjmp/longjmp. There are
|
||||
many different ways to go here.</li>
|
||||
|
||||
<li><b>object orientation, generics, database access, complex numbers,
|
||||
geometric programming, ...</b> - Really, there is
|
||||
no end of crazy features that you can add to the language.</li>
|
||||
|
||||
<li><b>unusual domains</b> - We've been talking about applying LLVM to a domain
|
||||
that many people are interested in: building a compiler for a specific language.
|
||||
However, there are many other domains that can use compiler technology that are
|
||||
not typically considered. For example, LLVM has been used to implement OpenGL
|
||||
graphics acceleration, translate C++ code to ActionScript, and many other
|
||||
cute and clever things. Maybe you will be the first to JIT compile a regular
|
||||
expression interpreter into native code with LLVM?</li>
|
||||
|
||||
</ul>
|
||||
|
||||
<p>
|
||||
Have fun - try doing something crazy and unusual. Building a language like
|
||||
everyone else always has, is much less fun than trying something a little crazy
|
||||
or off the wall and seeing how it turns out. If you get stuck or want to talk
|
||||
about it, feel free to email the <a
|
||||
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
|
||||
list</a>: it has lots of people who are interested in languages and are often
|
||||
willing to help out.
|
||||
</p>
|
||||
|
||||
<p>Before we end this tutorial, I want to talk about some "tips and tricks" for
|
||||
generating LLVM IR. These are some of the more subtle things that may not be
|
||||
obvious, but are very useful if you want to take advantage of LLVM's
|
||||
capabilities.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="llvmirproperties">Properties of the LLVM
|
||||
IR</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>We have a couple common questions about code in the LLVM IR form - let's just
|
||||
get these out of the way right now, shall we?</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="targetindep">Target
|
||||
Independence</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Kaleidoscope is an example of a "portable language": any program written in
|
||||
Kaleidoscope will work the same way on any target that it runs on. Many other
|
||||
languages have this property, e.g. LISP, Java, Haskell, Javascript, Python, etc.
|
||||
(note that while these languages are portable, not all their libraries are).</p>
|
||||
|
||||
<p>One nice aspect of LLVM is that it is often capable of preserving target
|
||||
independence in the IR: you can take the LLVM IR for a Kaleidoscope-compiled
|
||||
program and run it on any target that LLVM supports, even emitting C code and
|
||||
compiling that on targets that LLVM doesn't support natively. You can trivially
|
||||
tell that the Kaleidoscope compiler generates target-independent code because it
|
||||
never queries for any target-specific information when generating code.</p>
|
||||
|
||||
<p>The fact that LLVM provides a compact, target-independent, representation for
|
||||
code gets a lot of people excited. Unfortunately, these people are usually
|
||||
thinking about C or a language from the C family when they are asking questions
|
||||
about language portability. I say "unfortunately", because there is really no
|
||||
way to make (fully general) C code portable, other than shipping the source code
|
||||
around (and of course, C source code is not actually portable in general
|
||||
either - ever port a really old application from 32- to 64-bits?).</p>
|
||||
|
||||
<p>The problem with C (again, in its full generality) is that it is heavily
|
||||
laden with target specific assumptions. As one simple example, the preprocessor
|
||||
often destructively removes target-independence from the code when it processes
|
||||
the input text:</p>
|
||||
|
||||
<div class="doc_code">
|
||||
<pre>
|
||||
#ifdef __i386__
|
||||
int X = 1;
|
||||
#else
|
||||
int X = 42;
|
||||
#endif
|
||||
</pre>
|
||||
</div>
|
||||
|
||||
<p>While it is possible to engineer more and more complex solutions to problems
|
||||
like this, it cannot be solved in full generality in a way that is better than
|
||||
shipping the actual source code.</p>
|
||||
|
||||
<p>That said, there are interesting subsets of C that can be made portable. If
|
||||
you are willing to fix primitive types to a fixed size (say int = 32-bits,
|
||||
and long = 64-bits), don't care about ABI compatibility with existing binaries,
|
||||
and are willing to give up some other minor features, you can have portable
|
||||
code. This can make sense for specialized domains such as an
|
||||
in-kernel language.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="safety">Safety Guarantees</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Many of the languages above are also "safe" languages: it is impossible for
|
||||
a program written in Java to corrupt its address space and crash the process
|
||||
(assuming the JVM has no bugs).
|
||||
Safety is an interesting property that requires a combination of language
|
||||
design, runtime support, and often operating system support.</p>
|
||||
|
||||
<p>It is certainly possible to implement a safe language in LLVM, but LLVM IR
|
||||
does not itself guarantee safety. The LLVM IR allows unsafe pointer casts,
|
||||
use after free bugs, buffer over-runs, and a variety of other problems. Safety
|
||||
needs to be implemented as a layer on top of LLVM and, conveniently, several
|
||||
groups have investigated this. Ask on the <a
|
||||
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
|
||||
list</a> if you are interested in more details.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="langspecific">Language-Specific
|
||||
Optimizations</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>One thing about LLVM that turns off many people is that it does not solve all
|
||||
the world's problems in one system (sorry 'world hunger', someone else will have
|
||||
to solve you some other day). One specific complaint is that people perceive
|
||||
LLVM as being incapable of performing high-level language-specific optimization:
|
||||
LLVM "loses too much information".</p>
|
||||
|
||||
<p>Unfortunately, this is really not the place to give you a full and unified
|
||||
version of "Chris Lattner's theory of compiler design". Instead, I'll make a
|
||||
few observations:</p>
|
||||
|
||||
<p>First, you're right that LLVM does lose information. For example, as of this
|
||||
writing, there is no way to distinguish in the LLVM IR whether an SSA-value came
|
||||
from a C "int" or a C "long" on an ILP32 machine (other than debug info). Both
|
||||
get compiled down to an 'i32' value and the information about what it came from
|
||||
is lost. The more general issue here, is that the LLVM type system uses
|
||||
"structural equivalence" instead of "name equivalence". Another place this
|
||||
surprises people is if you have two types in a high-level language that have the
|
||||
same structure (e.g. two different structs that have a single int field): these
|
||||
types will compile down into a single LLVM type and it will be impossible to
|
||||
tell what it came from.</p>
|
||||
|
||||
<p>Second, while LLVM does lose information, LLVM is not a fixed target: we
|
||||
continue to enhance and improve it in many different ways. In addition to
|
||||
adding new features (LLVM did not always support exceptions or debug info), we
|
||||
also extend the IR to capture important information for optimization (e.g.
|
||||
whether an argument is sign or zero extended, information about pointers
|
||||
aliasing, etc). Many of the enhancements are user-driven: people want LLVM to
|
||||
include some specific feature, so they go ahead and extend it.</p>
|
||||
|
||||
<p>Third, it is <em>possible and easy</em> to add language-specific
|
||||
optimizations, and you have a number of choices in how to do it. As one trivial
|
||||
example, it is easy to add language-specific optimization passes that
|
||||
"know" things about code compiled for a language. In the case of the C family,
|
||||
there is an optimization pass that "knows" about the standard C library
|
||||
functions. If you call "exit(0)" in main(), it knows that it is safe to
|
||||
optimize that into "return 0;" because C specifies what the 'exit'
|
||||
function does.</p>
|
||||
|
||||
<p>In addition to simple library knowledge, it is possible to embed a variety of
|
||||
other language-specific information into the LLVM IR. If you have a specific
|
||||
need and run into a wall, please bring the topic up on the llvmdev list. At the
|
||||
very worst, you can always treat LLVM as if it were a "dumb code generator" and
|
||||
implement the high-level optimizations you desire in your front-end, on the
|
||||
language-specific AST.
|
||||
</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<div class="doc_section"><a name="tipsandtricks">Tips and Tricks</a></div>
|
||||
<!-- *********************************************************************** -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>There is a variety of useful tips and tricks that you come to know after
|
||||
working on/with LLVM that aren't obvious at first glance. Instead of letting
|
||||
everyone rediscover them, this section talks about some of these issues.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="offsetofsizeof">Implementing portable
|
||||
offsetof/sizeof</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>One interesting thing that comes up, if you are trying to keep the code
|
||||
generated by your compiler "target independent", is that you often need to know
|
||||
the size of some LLVM type or the offset of some field in an llvm structure.
|
||||
For example, you might need to pass the size of a type into a function that
|
||||
allocates memory.</p>
|
||||
|
||||
<p>Unfortunately, this can vary widely across targets: for example the width of
|
||||
a pointer is trivially target-specific. However, there is a <a
|
||||
href="http://nondot.org/sabre/LLVMNotes/SizeOf-OffsetOf-VariableSizedStructs.txt">clever
|
||||
way to use the getelementptr instruction</a> that allows you to compute this
|
||||
in a portable way.</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- ======================================================================= -->
|
||||
<div class="doc_subsubsection"><a name="gcstack">Garbage Collected
|
||||
Stack Frames</a></div>
|
||||
<!-- ======================================================================= -->
|
||||
|
||||
<div class="doc_text">
|
||||
|
||||
<p>Some languages want to explicitly manage their stack frames, often so that
|
||||
they are garbage collected or to allow easy implementation of closures. There
|
||||
are often better ways to implement these features than explicit stack frames,
|
||||
but <a
|
||||
href="http://nondot.org/sabre/LLVMNotes/ExplicitlyManagedStackFrames.txt">LLVM
|
||||
does support them,</a> if you want. It requires your front-end to convert the
|
||||
code into <a
|
||||
href="http://en.wikipedia.org/wiki/Continuation-passing_style">Continuation
|
||||
Passing Style</a> and the use of tail calls (which LLVM also supports).</p>
|
||||
|
||||
</div>
|
||||
|
||||
<!-- *********************************************************************** -->
|
||||
<hr>
|
||||
<address>
|
||||
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
|
||||
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
|
||||
<a href="http://validator.w3.org/check/referer"><img
|
||||
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
|
||||
|
||||
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
|
||||
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
|
||||
Last modified: $Date$
|
||||
</address>
|
||||
</body>
|
||||
</html>
|
||||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -74,7 +74,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.</tt></pre>
|
|||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
|
||||
<head>
|
||||
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
|
||||
<meta name="generator" content="AsciiDoc 8.5.3" />
|
||||
<meta name="generator" content="AsciiDoc 8.6.1" />
|
||||
<meta name="description" content="Python bindings for LLVM" />
|
||||
<meta name="keywords" content="llvm python compiler backend bindings" />
|
||||
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
|
||||
|
|
@ -48,6 +48,7 @@ you can setup and use it. A working knowledge of Python and a basic idea
|
|||
of LLVM is assumed.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_introduction">Introduction</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p><a href="http://www.llvm.org/">LLVM</a> (Low-Level Virtual Machine) provides enough
|
||||
|
|
@ -76,6 +77,8 @@ versions.</p></div>
|
|||
<div class="paragraph"><p>llvm-py has been built and tested with Python 2.6. It should work with
|
||||
Python 2.4 and 2.5. It has not been tried with Python 3.x (patches welcome).</p></div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="install">Installation</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p>llvm-py is distributed as a source tarball. You’ll need to build and
|
||||
|
|
@ -109,7 +112,8 @@ distro’s respository has the appropriate version of LLVM!</p></div>
|
|||
<div class="paragraph"><p>It does not matter which compiler LLVM itself was built with (g<tt>,
|
||||
llvm-g</tt> or any other); llvm-py can be built with any compiler. It has
|
||||
been tried only with gcc/g++ though.</p></div>
|
||||
<h3 id="_llvm_and_tt_enable_pic_tt">LLVM and <tt>--enable-pic</tt></h3><div style="clear:left"></div>
|
||||
<div class="sect2">
|
||||
<h3 id="_llvm_and_tt_enable_pic_tt">LLVM and <tt>--enable-pic</tt></h3>
|
||||
<div class="paragraph"><p>The result of an LLVM build is a set of static libraries and object
|
||||
files. The llvm-py contains an extension package that is built into a
|
||||
shared object (_core.so) which links to these static libraries and
|
||||
|
|
@ -121,7 +125,9 @@ configuring LLVM (default is no PIC), like this:</p></div>
|
|||
<div class="content">
|
||||
<pre><tt>~/llvm$ ./configure --enable-pic --enable-optimized</tt></pre>
|
||||
</div></div>
|
||||
<h3 id="_llvm_config">llvm-config</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_llvm_config">llvm-config</h3>
|
||||
<div class="paragraph"><p>Inorder to build llvm-py, it’s build script needs to know from where it
|
||||
can invoke the llvm helper program, <tt>llvm-config</tt>. If you’ve installed
|
||||
LLVM, then this will be available in your <tt>PATH</tt>, and nothing further
|
||||
|
|
@ -131,7 +137,9 @@ of <tt>llvm-config</tt> to the build script.</p></div>
|
|||
<div class="paragraph"><p>You’ll need to be <em>root</em> to install llvm-py. Remember that your <tt>PATH</tt>
|
||||
is different from that of <em>root</em>, so even if <tt>llvm-config</tt> is in your
|
||||
<tt>PATH</tt>, it may not be available when you do <tt>sudo</tt>.</p></div>
|
||||
<h3 id="_steps">Steps</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_steps">Steps</h3>
|
||||
<div class="paragraph"><p>The commands illustrated below assume that the LLVM source is available
|
||||
under <tt>/home/mdevan/llvm</tt>. If you’ve a previous version of llvm-py
|
||||
installed, it is recommended to remove it first, as described
|
||||
|
|
@ -167,7 +175,9 @@ only if you need to debug into LLVM also.</p></div>
|
|||
documentation regarding <a href="http://docs.python.org/inst/inst.html">Installing
|
||||
Python Modules</a> and <a href="http://docs.python.org/dist/dist.html">Distributing
|
||||
Python Modules</a> for more information on such scripts.</p></div>
|
||||
<h3 id="uninstall">Uninstall</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="uninstall">Uninstall</h3>
|
||||
<div class="paragraph"><p>If you’d installed llvm-py with the <tt>--user</tt> option, then llvm-py
|
||||
would be present under <tt>~/.local/lib/python2.6/site-packages</tt>.
|
||||
Otherwise, it might be under <tt>/usr/lib/python2.6/site-packages</tt>
|
||||
|
|
@ -182,11 +192,15 @@ the "egg" can be removed like so:</p></div>
|
|||
<div class="paragraph"><p>See the <a href="http://docs.python.org/install/index.html">Python
|
||||
documentation</a> for more information.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_llvm_concepts">LLVM Concepts</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p>This section explains a few concepts related to LLVM, not specific
|
||||
to llvm-py.</p></div>
|
||||
<h3 id="_intermediate_representation">Intermediate Representation</h3><div style="clear:left"></div>
|
||||
<div class="sect2">
|
||||
<h3 id="_intermediate_representation">Intermediate Representation</h3>
|
||||
<div class="paragraph"><p>The intermediate representation, or IR for short, is an in-memory data
|
||||
structure that represents executable code. The IR data structures allow
|
||||
for creation of types, constants, functions, function arguments,
|
||||
|
|
@ -242,7 +256,9 @@ level than the usual assembly language; for example there are
|
|||
instructions related to variable argument handling, exception handling,
|
||||
and garbage collection. These allow high-level languages to be
|
||||
represented cleanly in the IR.</p></div>
|
||||
<h3 id="_ssa_form_and_phi_nodes">SSA Form and PHI Nodes</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_ssa_form_and_phi_nodes">SSA Form and PHI Nodes</h3>
|
||||
<div class="paragraph"><p>All LLVM instructions are represented in the <em>Static Single Assignment</em>
|
||||
(SSA) form. Essentially, this means that any variable can be assigned to
|
||||
only once. Such a representation facilitates better optimization, among
|
||||
|
|
@ -274,7 +290,9 @@ reached the PHI node. The argument <tt>a1</tt> of the PHI node is associated
|
|||
with the block <tt>"a1 = 1;"</tt> and <tt>a2</tt> with the block <tt>"a2 = 2;"</tt>.</p></div>
|
||||
<div class="paragraph"><p>PHI nodes have to be explicitly created in the LLVM IR. Accordingly the
|
||||
LLVM instruction set has an instruction called <tt>phi</tt>.</p></div>
|
||||
<h3 id="_llvm_assembly_language">LLVM Assembly Language</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_llvm_assembly_language">LLVM Assembly Language</h3>
|
||||
<div class="paragraph"><p>The LLVM IR can be represented offline in two formats
|
||||
- a textual, human-readable form, similar to assembly language, called
|
||||
the LLVM assembly language (files with .ll extension)
|
||||
|
|
@ -337,7 +355,9 @@ specification of the platform ABI (like endianness, sizes of types,
|
|||
alignment etc.).</p></div>
|
||||
<div class="paragraph"><p>The <a href="http://www.llvm.org/docs/LangRef.html">LLVM Language Reference</a>
|
||||
defines the LLVM assembly language including the entire instruction set.</p></div>
|
||||
<h3 id="_modules">Modules</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_modules">Modules</h3>
|
||||
<div class="paragraph"><p>Modules, in the LLVM IR, are similar to a single <tt>C</tt> language source
|
||||
file (.c file). A module contains:</p></div>
|
||||
<div class="ulist"><ul>
|
||||
|
|
@ -361,7 +381,9 @@ global type aliases (typedef-s)
|
|||
contained within modules. Modules may be combined (linked) together to
|
||||
give a bigger resultant module. During this process LLVM attempts to
|
||||
reconcile the references between the combined modules.</p></div>
|
||||
<h3 id="_optimization_and_passes">Optimization and Passes</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_optimization_and_passes">Optimization and Passes</h3>
|
||||
<div class="paragraph"><p>LLVM provides quite a few optimization algorithms that work on the IR.
|
||||
These algorithms are organized as <em>passes</em>. Each pass does something
|
||||
specific, like combining redundant instructions. Passes need not always
|
||||
|
|
@ -384,11 +406,18 @@ any stage, and perform any transforms on it as you like.)</p></div>
|
|||
correct objects to run them on (for example, a pass may work only
|
||||
on functions, individually) and actually runs them. <tt>opt</tt> is a
|
||||
command-line wrapper for the pass manager.</p></div>
|
||||
<h3 id="_bit_code">Bit code</h3><div style="clear:left"></div>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
<h3 id="_execution_engine_jit_and_interpreter">Execution Engine, JIT and Interpreter</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_bit_code">Bit code</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_execution_engine_jit_and_interpreter">Execution Engine, JIT and Interpreter</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_the_llvm_py_package">The llvm-py Package</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p>The llvm-py is a Python package, consisting of 6 modules, that wrap
|
||||
|
|
@ -575,7 +604,8 @@ interpreter or the <tt>object?</tt> of <a href="http://ipython.scipy.org/moin/">
|
|||
to get online help. (Note: not complete yet!)</td>
|
||||
</tr></table>
|
||||
</div>
|
||||
<h3 id="_module_llvm_core">Module (llvm.core)</h3><div style="clear:left"></div>
|
||||
<div class="sect2">
|
||||
<h3 id="_module_llvm_core">Module (llvm.core)</h3>
|
||||
<div class="paragraph"><p>Modules are top-level container objects. You need to create a module
|
||||
object first, before you can add global variables, aliases or functions.
|
||||
Modules are created using the static method <tt>Module.new</tt>:</p></div>
|
||||
|
|
@ -644,7 +674,7 @@ my_module <span style="color: #990000">=</span> Module<span style="color: #99000
|
|||
stringifying them (see below).</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.Module</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="dlist"><div class="title">Static Constructors</div><dl>
|
||||
<dt class="hdlist1">
|
||||
<tt>new(module_id)</tt>
|
||||
|
|
@ -857,7 +887,9 @@ string representations.</p></div>
|
|||
</td>
|
||||
</tr></table>
|
||||
</div>
|
||||
<h3 id="_types_llvm_core">Types (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_types_llvm_core">Types (llvm.core)</h3>
|
||||
<div class="paragraph"><p>Types are what you think they are. A instance of <tt>llvm.core.Type</tt>, or
|
||||
one of its derived classes, represent a type. llvm-py does not use as
|
||||
many classes to represent types as does LLVM itself. Some types are
|
||||
|
|
@ -977,7 +1009,7 @@ cellspacing="0" cellpadding="4">
|
|||
<div class="paragraph"><p>The class-level documentation follows:</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.Type</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="dlist"><div class="title">Static Constructors</div><dl>
|
||||
<dt class="hdlist1">
|
||||
<tt>int(n)</tt>
|
||||
|
|
@ -1176,7 +1208,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.IntegerType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1197,7 +1229,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.FunctionType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1254,7 +1286,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.StructType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1303,7 +1335,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.ArrayType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1332,7 +1364,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.PointerType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1361,7 +1393,7 @@ http://www.gnu.org/software/src-highlite -->
|
|||
</div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.VectorType</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1430,7 +1462,9 @@ f3 <span style="color: #990000">=</span> Type<span style="color: #990000">.</spa
|
|||
fnargs <span style="color: #990000">=</span> <span style="color: #990000">[</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">pointer</span></span><span style="color: #990000">(</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">int</span></span><span style="color: #990000">(</span><span style="color: #993399">8</span><span style="color: #990000">)</span> <span style="color: #990000">)</span> <span style="color: #990000">]</span>
|
||||
printf <span style="color: #990000">=</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">function</span></span><span style="color: #990000">(</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">int</span></span><span style="color: #990000">(),</span> fnargs<span style="color: #990000">,</span> True <span style="color: #990000">)</span>
|
||||
<span style="font-style: italic"><span style="color: #9A1900"># variadic function</span></span></tt></pre></div></div>
|
||||
<h3 id="_typehandle_llvm_core">TypeHandle (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_typehandle_llvm_core">TypeHandle (llvm.core)</h3>
|
||||
<div class="paragraph"><p>TypeHandle objects are used to create recursive types, like this linked
|
||||
list node structure in C:</p></div>
|
||||
<div class="listingblock">
|
||||
|
|
@ -1483,7 +1517,7 @@ in C++. The above example is available as
|
|||
in the source distribution.</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.TypeHandle</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="dlist"><div class="title">Static Constructors</div><dl>
|
||||
<dt class="hdlist1">
|
||||
<tt>new(abstract_ty)</tt>
|
||||
|
|
@ -1508,7 +1542,9 @@ in the source distribution.</p></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_values_llvm_core">Values (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_values_llvm_core">Values (llvm.core)</h3>
|
||||
<div class="paragraph"><p><tt>llvm.core.Value</tt> is the base class of all values computed by a program
|
||||
that may be used as operands to other values. A value has a type
|
||||
associated with it (an object of <tt>llvm.core.Type</tt>).</p></div>
|
||||
|
|
@ -1557,7 +1593,7 @@ a few subclasses that represent interesting instructions.</p></div>
|
|||
<div class="paragraph"><p><tt>Value</tt> objects have a type (read-only), and a name (read-write).</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.Value</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="dlist"><div class="title">Properties</div><dl>
|
||||
<dt class="hdlist1">
|
||||
<tt>name</tt>
|
||||
|
|
@ -1624,14 +1660,16 @@ a few subclasses that represent interesting instructions.</p></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_user_llvm_core">User (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_user_llvm_core">User (llvm.core)</h3>
|
||||
<div class="paragraph"><p><tt>User</tt>-s are values that refer to other values. The values so refered
|
||||
can be retrived by the properties of <tt>User</tt>. This is the reverse of
|
||||
the <tt>Value.uses</tt>. Together these can be used to traverse the use-def
|
||||
chains of the SSA.</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.User</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -1660,7 +1698,9 @@ chains of the SSA.</p></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_constants_llvm_core">Constants (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_constants_llvm_core">Constants (llvm.core)</h3>
|
||||
<div class="paragraph"><p><tt>Constant</tt>-s represents constants that appear within the code. The
|
||||
values of such objects are known at creation time. Constants can be
|
||||
created from Python constants. A constant expression is also a constant — given a <tt>Constant</tt> object, an operation (like addition, subtraction
|
||||
|
|
@ -2074,7 +2114,7 @@ cellspacing="0" cellpadding="4">
|
|||
</div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.Constant</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -2086,7 +2126,9 @@ cellspacing="0" cellpadding="4">
|
|||
<div class="paragraph"><div class="title">Methods</div><p>See table of operations <a href="#constops">above</a> for full list. There are no other
|
||||
methods.</p></div>
|
||||
</div></div>
|
||||
<h3 id="_other_constant_classes_llvm_core">Other Constant* Classes (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_other_constant_classes_llvm_core">Other Constant* Classes (llvm.core)</h3>
|
||||
<div class="paragraph"><p>The following subclasses of <tt>Constant</tt> do not provide additional
|
||||
methods, they serve only to provide richer type information.</p></div>
|
||||
<div class="tableblock">
|
||||
|
|
@ -2165,7 +2207,9 @@ k2 <span style="color: #990000">=</span> Constant<span style="color: #990000">.<
|
|||
|
||||
<span style="font-weight: bold"><span style="color: #0000FF">assert</span></span> <span style="font-weight: bold"><span style="color: #000000">isinstance</span></span><span style="color: #990000">(</span>k1<span style="color: #990000">,</span> ConstantInt<span style="color: #990000">)</span>
|
||||
<span style="font-weight: bold"><span style="color: #0000FF">assert</span></span> <span style="font-weight: bold"><span style="color: #000000">isinstance</span></span><span style="color: #990000">(</span>k2<span style="color: #990000">,</span> ConstantArray<span style="color: #990000">)</span></tt></pre></div></div>
|
||||
<h3 id="_global_value_llvm_core">Global Value (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_global_value_llvm_core">Global Value (llvm.core)</h3>
|
||||
<div class="paragraph"><p>The class <tt>llvm.core.GlobalValue</tt> represents module-scope aliases, variables
|
||||
and functions. Global variables are represented by the sub-class
|
||||
<tt>llvm.core.GlobalVariable</tt> and functions by <tt>llvm.core.Function</tt>.</p></div>
|
||||
|
|
@ -2290,7 +2334,7 @@ global is a declaration or not. The module to which the global belongs
|
|||
to can be retrieved using the <tt>module</tt> property (read-only).</p></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.GlobalValue</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -2352,7 +2396,9 @@ to can be retrieved using the <tt>module</tt> property (read-only).</p></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_global_variable_llvm_core">Global Variable (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_global_variable_llvm_core">Global Variable (llvm.core)</h3>
|
||||
<div class="paragraph"><p>Global variables (<tt>llvm.core.GlobalVariable</tt>) are subclasses of
|
||||
<tt>llvm.core.GlobalValue</tt> and represent module-level variables. These can
|
||||
have optional initializers and can be marked as constants. Global
|
||||
|
|
@ -2407,7 +2453,7 @@ gv<span style="color: #990000">.</span><span style="font-weight: bold"><span sty
|
|||
gv <span style="color: #990000">=</span> None</tt></pre></div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.GlobalVariable</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -2467,7 +2513,9 @@ gv <span style="color: #990000">=</span> None</tt></pre></div></div>
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_function_llvm_core">Function (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_function_llvm_core">Function (llvm.core)</h3>
|
||||
<div class="paragraph"><p>Functions are represented by <tt>llvm.core.Function</tt> objects. They are
|
||||
contained within modules, and can be created either with the method
|
||||
<tt>module_obj.add_function</tt> or the static constructor <tt>Function.new</tt>.
|
||||
|
|
@ -2939,7 +2987,7 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
|
|||
<span style="font-style: italic"><span style="color: #9A1900"># declare i32 @sum(i32, i32) nounwind readonly</span></span></tt></pre></div></div>
|
||||
<div class="exampleblock">
|
||||
<div class="title">llvm.core.Function</div>
|
||||
<div class="exampleblock-content">
|
||||
<div class="content">
|
||||
<div class="ulist"><div class="title">Base Class</div><ul>
|
||||
<li>
|
||||
<p>
|
||||
|
|
@ -2993,7 +3041,7 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
|
|||
<dd>
|
||||
<p>
|
||||
The calling convention for the function, as listed
|
||||
<a href="#callconv">abov</a>.
|
||||
<a href="#callconv">above</a>.
|
||||
</p>
|
||||
</dd>
|
||||
<dt class="hdlist1">
|
||||
|
|
@ -3127,7 +3175,9 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
|
|||
</dd>
|
||||
</dl></div>
|
||||
</div></div>
|
||||
<h3 id="_argument_llvm_core">Argument (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_argument_llvm_core">Argument (llvm.core)</h3>
|
||||
<div class="paragraph"><p>The <tt>args</tt> property of <tt>llvm.core.Function</tt> objects yields
|
||||
<tt>llvm.core.Argument</tt> objects. This allows for setting attributes for
|
||||
functions arguments. <tt>Argument</tt> objects cannot be constructed from user
|
||||
|
|
@ -3223,19 +3273,34 @@ cellspacing="0" cellpadding="4">
|
|||
provide more information.</p></div>
|
||||
<div class="paragraph"><p>The alignment of any parameter can be set via the <tt>alignment</tt>
|
||||
property, to any power of 2.</p></div>
|
||||
<h3 id="_basic_block_llvm_core">Basic Block (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_basic_block_llvm_core">Basic Block (llvm.core)</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
<h3 id="_builder_llvm_core">Builder (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_builder_llvm_core">Builder (llvm.core)</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
<h3 id="_instructions_llvm_core">Instructions (llvm.core)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_instructions_llvm_core">Instructions (llvm.core)</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
<h3 id="_target_data_llvm_ee">Target Data (llvm.ee)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_target_data_llvm_ee">Target Data (llvm.ee)</h3>
|
||||
<div class="paragraph"><p>TODO</p></div>
|
||||
<h3 id="_execution_engine_llvm_ee">Execution Engine (llvm.ee)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_execution_engine_llvm_ee">Execution Engine (llvm.ee)</h3>
|
||||
<div class="paragraph"><p>TODO. For now, see <tt>test/example-jit.py</tt>.</p></div>
|
||||
<h3 id="_pass_manager_and_passes_llvm_passes">Pass Manager and Passes (llvm.passes)</h3><div style="clear:left"></div>
|
||||
</div>
|
||||
<div class="sect2">
|
||||
<h3 id="_pass_manager_and_passes_llvm_passes">Pass Manager and Passes (llvm.passes)</h3>
|
||||
<div class="paragraph"><p>TODO. For now, see <tt>test/passes.py</tt>.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div class="sect1">
|
||||
<h2 id="_about_the_llvm_py_project">About the llvm-py Project</h2>
|
||||
<div class="sectionbody">
|
||||
<div class="paragraph"><p>llvm-py lives at
|
||||
|
|
@ -3259,10 +3324,11 @@ are most welcome. You can checkout the latest SVN HEAD from
|
|||
<div class="paragraph"><p>Mahadevan R wrote llvm-py and works on it in his spare time. He can be
|
||||
reached at <em>mdevan@mdevan.org</em>.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
<div id="footer">
|
||||
<div id="footer-text">
|
||||
Web pages © Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
|
||||
Last updated 2010-08-31.
|
||||
Last updated 2010-09-26.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue