LLVM tutorial ported (Max Shawabkeh) (Issue #33)

git-svn-id: http://llvm-py.googlecode.com/svn/trunk@94 8d1e9007-1d4e-0410-b67e-1979fd6579aa
This commit is contained in:
mdevan.foobar 2010-09-26 08:25:14 +00:00
commit 60422c6057
27 changed files with 18381 additions and 102 deletions

View file

@ -1,4 +1,9 @@
0.7, in progress:
* LLVM tutorial ported (Max Shawabkeh) (Issue #33).
0.6, 31-Aug-2010:
* Add and remove function attributes (Krzysztof Goj) (Issue #21).

View file

@ -32,7 +32,7 @@
import sys, os
from distutils.core import setup, Extension
LLVM_PY_VERSION = '0.6'
LLVM_PY_VERSION = '0.7'
def _run(cmd):
@ -106,8 +106,8 @@ def call_setup(llvm_config):
version=LLVM_PY_VERSION,
description='Python Bindings for LLVM',
author='Mahadevan R',
author_email='mdevan.foobar@gmail.com',
url='http://mdevan.nfshost.com/llvm-py/',
author_email='mdevan@mdevan.org',
url='http://www.mdevan.org/llvm-py/',
packages=['llvm'],
py_modules = [ 'llvm.core' ],
ext_modules = [ ext_core ],)

View file

@ -9,27 +9,25 @@ include::example.inc[]
LLVM Tutorials
--------------
The http://www.llvm.org/docs/tutorial/[LLVM tutorials] have been
ported to llvm-py. Below are the links to the original LLVM tutorial and
the corresponding Python code using llvm-py:
.Simple JIT Tutorials
(contributed by Sebastien Binet)
1. A First Function
http://www.llvm.org/docs/tutorial/JITTutorial1.html[LLVM]
link:examples/JITTutorial1.html[llvm-py]
2. A More Complicated Function
http://www.llvm.org/docs/tutorial/JITTutorial2.html[LLVM]
link:examples/JITTutorial2.html[llvm-py]
The following JIT tutorials were contributed by Sebastien Binet.
1. link:examples/JITTutorial1.html[A First Function]
2. link:examples/JITTutorial2.html[A More Complicated Function]
[[kaleidoscope]]
.Kaleidoscope: Implementing a Language with LLVM
1. Tutorial Introduction and the Lexer (TODO)
2. Implementing a Parser and AST (TODO)
3. Implementing Code Generation to LLVM IR (TODO)
4. Adding JIT and Optimizer Support (TODO)
5. Extending the language: control flow (TODO)
6. Extending the language: user-defined operators (TODO)
7. Extending the language: mutable variables / SSA construction (TODO)
8. Conclusion and other useful LLVM tidbits (TODO)
The LLVM http://www.llvm.org/docs/tutorial/[Kaleidoscope] tutorial
has been ported to llvm-py by Max Shawabkeh.
1. link:kaleidoscope/PythonLangImpl1.html[Tutorial Introduction and the Lexer]
2. link:kaleidoscope/PythonLangImpl2.html[Implementing a Parser and AST]
3. link:kaleidoscope/PythonLangImpl3.html[Implementing Code Generation to LLVM IR]
4. link:kaleidoscope/PythonLangImpl4.html[Adding JIT and Optimizer Support]
5. link:kaleidoscope/PythonLangImpl5.html[Extending the language: control flow]
6. link:kaleidoscope/PythonLangImpl6.html[Extending the language: user-defined operators]
7. link:kaleidoscope/PythonLangImpl7.html[Extending the language: mutable variables / SSA construction]
8. link:kaleidoscope/PythonLangImpl8.html[Conclusion and other useful LLVM tidbits]

View file

@ -18,6 +18,9 @@ a patch.
News
----
26-Sep-2010::
LLVM tutorial link:examples.html#kaleidoscope[ported] by Max Shawabkeh!
31-Aug-2010::
0.6 released, works with LLVM 2.7.

View file

@ -0,0 +1,389 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Tutorial Introduction and the Lexer</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<meta name="author" content="Max Shawabkeh">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Tutorial Introduction and the Lexer</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 1
<ol>
<li><a href="#intro">Tutorial Introduction</a></li>
<li><a href="#language">The Basic Language</a></li>
<li><a href="#lexer">The Lexer</a></li>
</ol>
</li>
<li><a href="PythonLangImpl2.html">Chapter 2</a>: Implementing a Parser and
AST</li>
</ul>
<div class="doc_author">
<p>
Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
and <a href="http://max99x.com">Max Shawabkeh</a>
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="intro">Tutorial Introduction</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Welcome to the "Implementing a language with LLVM" tutorial. This tutorial
runs through the implementation of a simple language, showing how fun and
easy it can be. This tutorial will get you up and started as well as help to
build a framework you can extend to other languages. The code in this tutorial
can also be used as a playground to hack on other LLVM specific things.
</p>
<p>The goal of this tutorial is to progressively unveil our language, describing
how it is built up over time. This will let us cover a fairly broad range of
language design and LLVM-specific usage issues, showing and explaining the code
for it all along the way, without overwhelming you with tons of details up
front.</p>
<p>It is useful to point out ahead of time that this tutorial is really about
teaching compiler techniques and LLVM specifically, <em>not</em> about teaching
modern and sane software engineering principles. In practice, this means that
we'll take a number of shortcuts to simplify the exposition. If you dig in and
use the code as a basis for future projects, fixing its deficiencies shouldn't
be hard.</p>
<p>We've tried to put this tutorial together in a way that makes chapters easy
to skip over if you are already familiar with or are uninterested in the various
pieces. The structure of the tutorial is:</p>
<ul>
<li><b><a href="#language">Chapter #1</a>: Introduction to the Kaleidoscope
language, and the definition of its Lexer</b> - This shows where we are going
and the basic functionality that we want it to do. In order to make this
tutorial maximally understandable and hackable, we choose to implement
everything in Python instead of using lexer and parser generators. LLVM
obviously works just fine with such tools, feel free to use one if you prefer.
</li>
<li><b><a href="PythonLangImpl2.html">Chapter #2</a>: Implementing a Parser and
AST</b> - With the lexer in place, we can talk about parsing techniques and
basic AST construction. This tutorial describes recursive descent parsing and
operator precedence parsing. Nothing in Chapters 1 or 2 is LLVM-specific,
the code doesn't even import the LLVM modules at this point. :)</li>
<li><b><a href="PythonLangImpl3.html">Chapter #3</a>: Code generation to LLVM
IR</b> - With the AST ready, we can show off how easy generation of LLVM IR
really is.</li>
<li><b><a href="PythonLangImpl4.html">Chapter #4</a>: Adding JIT and Optimizer
Support</b> - Because a lot of people are interested in using LLVM as a JIT,
we'll dive right into it and show you the 3 lines it takes to add JIT support.
LLVM is also useful in many other ways, but this is one simple and "sexy" way
to shows off its power. :)</li>
<li><b><a href="PythonLangImpl5.html">Chapter #5</a>: Extending the Language:
Control Flow</b> - With the language up and running, we show how to extend it
with control flow operations (if/then/else and a 'for' loop). This gives us a
chance to talk about simple SSA construction and control flow.</li>
<li><b><a href="PythonLangImpl6.html">Chapter #6</a>: Extending the Language:
User-defined Operators</b> - This is a silly but fun chapter that talks about
extending the language to let the user program define their own arbitrary
unary and binary operators (with assignable precedence!). This lets us build a
significant piece of the "language" as library routines.</li>
<li><b><a href="PythonLangImpl7.html">Chapter #7</a>: Extending the Language:
Mutable Variables</b> - This chapter talks about adding user-defined local
variables along with an assignment operator. The interesting part about this
is how easy and trivial it is to construct SSA form in LLVM: no, LLVM does
<em>not</em> require your front-end to construct SSA form!</li>
<li><b><a href="PythonLangImpl8.html">Chapter #8</a>: Conclusion and other
useful LLVM tidbits</b> - This chapter wraps up the series by talking about
potential ways to extend the language, but also includes a bunch of pointers to
info about "special topics" like adding garbage collection support, exceptions,
debugging, support for "spaghetti stacks", and a bunch of other tips and
tricks.</li>
</ul>
<p>By the end of the tutorial, we'll have written a bit less than 540 lines of
non-comment, non-blank, lines of code. With this small amount of code, we'll
have built up a very reasonable compiler for a non-trivial language including
a hand-written lexer, parser, AST, as well as code generation support with a JIT
compiler. While other systems may have interesting "hello world" tutorials,
I think the breadth of this tutorial is a great testament to the strengths of
LLVM and why you should consider it if you're interested in language or compiler
design.</p>
<p>A note about this tutorial: we expect you to extend the language and play
with it on your own. Take the code and go crazy hacking away at it, compilers
don't need to be scary creatures - it can be a lot of fun to play with
languages!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="language">The Basic Language</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>This tutorial will be illustrated with a toy language that we'll call
"<a href="http://en.wikipedia.org/wiki/Kaleidoscope">Kaleidoscope</a>" (derived
from "meaning beautiful, form, and view").
Kaleidoscope is a procedural language that allows you to define functions, use
conditionals, math, etc. Over the course of the tutorial, we'll extend
Kaleidoscope to support the if/then/else construct, a for loop, user defined
operators, JIT compilation with a simple command line interface, etc.</p>
<p>Because we want to keep things simple, the only datatype in Kaleidoscope is a
64-bit floating point type. As such, all values are implicitly double precision
and the language doesn't require type declarations. This gives the language a
very nice and simple syntax. For example, the following simple example computes
<a href="http://en.wikipedia.org/wiki/Fibonacci_number">Fibonacci numbers:</a>
</p>
<div class="doc_code">
<pre>
# Compute the x'th fibonacci number.
def fib(x)
if x &lt; 3 then
1
else
fib(x-1)+fib(x-2)
# This expression will compute the 40th number.
fib(40)
</pre>
</div>
<p>We also allow Kaleidoscope to call into standard library functions (the LLVM
JIT makes this completely trivial). This means that you can use the 'extern'
keyword to define a function before you use it (this is also useful for mutually
recursive functions). For example:</p>
<div class="doc_code">
<pre>
extern sin(arg);
extern cos(arg);
extern atan2(arg1 arg2);
atan2(sin(0.4), cos(42))
</pre>
</div>
<p>A more interesting example is included in Chapter 6 where we write a little
Kaleidoscope application that <a href="PythonLangImpl6.html#example">displays
a Mandelbrot Set</a> at various levels of magnification.</p>
<p>Lets dive into the implementation of this language!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="lexer">The Lexer</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>When it comes to implementing a language, the first thing needed is
the ability to process a text file and recognize what it says. The traditional
way to do this is to use a "<a
href="http://en.wikipedia.org/wiki/Lexical_analysis">lexer</a>" (aka 'scanner')
to break the input up into "tokens". Each token returned by the lexer includes
a token type and potentially some metadata (e.g. the numeric value of a number).
First, we define the possibilities:</p>
<div class="doc_code">
<pre>
# The lexer yields one of these types for each token.
class EOFToken(object):
pass
class DefToken(object):
pass
class ExternToken(object):
pass
class IdentifierToken(object):
def __init__(self, name): self.name = name
class NumberToken(object):
def __init__(self, value): self.value = value
class CharacterToken(object):
def __init__(self, char): self.char = char
def __eq__(self, other):
return isinstance(other, CharacterToken) and self.char == other.char
def __ne__(self, other): return not self == other
</pre>
</div>
<p>Each token yielded by our lexer will be of one of the above types. For simple
tokens that are always the same, like the "def" keyword, the lexer will yield
<tt>DefToken()</tt>. Identifiers, numbers and characters, on the other
hand, have extra data, so when the lexer encounteres the number 123.45, it will
emit it as <tt>NumberToken(123.45)</tt>. An identifier <tt>foo</tt> will be
emitted as <tt>IdentifierToken('foo')</tt>. And finally, an unknown character
like '+' will be returned as <tt>CharacterToken('+')</tt>. You may notice that
we overload the equality and inequality operators for the characters; this will
later simplify character comparisons in the parser code.</p>
<p>The actual implementation of the lexer is a single function called
<tt>Tokenize</tt>, which takes a string and
<a href="http://docs.python.org/reference/simple_stmts.html#the-yield-statement">yields</a>
tokens. For simplicity, we will use
<a href="http://docs.python.org/library/re.html">regular
expressions</a> to parse out the tokens. This is terribly inefficient, but
perfectly sufficient for our needs.</p>
<p>First, we define the regular expressions for our tokens. Numbers and strings
of digits, optionally followed by a period and another string of digits.
Identifiers (and keywords) are alphanumeric string starting with a letter and
comments are anything between a hash (<tt>#</tt>) and the end of the line.
<div class="doc_code">
<pre>
import re
...
# Regular expressions that tokens and comments of our language.
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
REGEX_COMMENT = re.compile('#.*')
</pre>
</div>
<p>
Next, let's start defining the <tt>Tokenize</tt> function itself. The first
thing we need to do is set up a loop that scans the string, while ignoring
whitespace between tokens:</p>
<div class="doc_code">
<pre>
def Tokenize(string):
while string:
# Skip whitespace.
if string[0].isspace():
string = string[1:]
continue
...
</pre>
</div>
<p>Next we want to find out what the next token is. For this we run the regexes
we defined above on the remainder of the string. To simplify the rest of the
code, we run all three regexes each time. As mentioned above, inefficiencies are
ignored for the purpose of this tutorial:<p>
<div class="doc_code">
<pre>
# Run regexes.
comment_match = REGEX_COMMENT.match(string)
number_match = REGEX_NUMBER.match(string)
identifier_match = REGEX_IDENTIFIER.match(string)
</pre>
</div>
<p>Now se check if any of the regexes matched. For comments, we simply
ignore the captured match:</p>
<div class="doc_code">
<pre>
# Check if any of the regexes matched and yield the appropriate result.
if comment_match:
comment = comment_match.group(0)
string = string[len(comment):]
</pre>
</div>
<p>For numbers, we yield the captured match, converted to a float and tagged
with the appropriate token type:</p>
<div class="doc_code">
<pre>
elif number_match:
number = number_match.group(0)
yield NumberToken(float(number))
string = string[len(number):]
</pre>
</div>
<p>The identifier case is a little more complex. We have to check for keywords
to decide whether we have captured an identifier or a keyword:</p>
<div class="doc_code">
<pre>
elif identifier_match:
identifier = identifier_match.group(0)
# Check if we matched a keyword.
if identifier == 'def':
yield DefToken()
elif identifier == 'extern':
yield ExternToken()
else:
yield IdentifierToken(identifier)
string = string[len(identifier):]
</pre>
</div>
<p>Finally, if we haven't recognized a comment, a number of an identifier, we
yield the current character as an "unknown character" token. This is used, for
example, for operators like <tt>+</tt> or <tt>*</tt>:</p>
<div class="doc_code">
<pre>
else:
# Yield the unknown character.
yield CharacterToken(string[0])
string = string[1:]
</pre>
</div>
<p>Once we're done with the
loop, we return a final end-of-file token:</p>
<div class="doc_code">
<pre>
yield EOFToken()
</pre>
</div>
<p>With this, we have the complete lexer for the basic Kaleidoscope language
(the <a href="PythonLangImpl2.html#code">full code listing</a> for the Lexer is
available in the <a href="PythonLangImpl2.html">next chapter</a> of the
tutorial). Next we'll <a href="PythonLangImpl2.html">build a simple parser that
uses this to build an Abstract Syntax Tree</a>. When we have that, we'll
include a driver so that you can use the lexer and parser together.
</p>
<a href="PythonLangImpl2.html">Next: Implementing a Parser and AST</a>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://max99x.com">Max Shawabkeh</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,999 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Adding JIT and Optimizer Support</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<meta name="author" content="Max Shawabkeh">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Adding JIT and Optimizer Support</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 4
<ol>
<li><a href="#intro">Chapter 4 Introduction</a></li>
<li><a href="#trivialconstfold">Trivial Constant Folding</a></li>
<li><a href="#optimizerpasses">LLVM Optimization Passes</a></li>
<li><a href="#jit">Adding a JIT Compiler</a></li>
<li><a href="#code">Full Code Listing</a></li>
</ol>
</li>
<li><a href="PythonLangImpl5.html">Chapter 5</a>: Extending the Language:
Control Flow</li>
</ul>
<div class="doc_author">
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
and <a href="http://max99x.com">Max Shawabkeh</a>
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="intro">Chapter 4 Introduction</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Welcome to Chapter 4 of the
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
with LLVM</a>" tutorial. Chapters 1-3 described the implementation of a simple
language and added support for generating LLVM IR. This chapter describes
two new techniques: adding optimizer support to your language, and adding JIT
compiler support. These additions will demonstrate how to get nice, efficient
code for the Kaleidoscope language.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="trivialconstfold">Trivial Constant
Folding</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Our demonstration for Chapter 3 is elegant and easy to extend. Unfortunately,
it does not produce wonderful code. The LLVM Builder, however, does give us
obvious optimizations when compiling simple code:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) 1+2+x</b>
Read function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 3.000000e+00, %x
ret double %addtmp
}
</pre>
</div>
<p>This code is not a literal transcription of the AST built by parsing the
input. That would be:
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) 1+2+x</b>
Read function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 2.000000e+00, 1.000000e+00
%addtmp1 = fadd double %addtmp, %x
ret double %addtmp1
}
</pre>
</div>
<p>Constant folding, as seen above, in particular, is a very common and very
important optimization: so much so that many language implementors implement
constant folding support in their AST representation.</p>
<p>With LLVM, you don't need this support in the AST. Since all calls to build
LLVM IR go through the LLVM IR builder, the builder itself checked to see if
there was a constant folding opportunity when you call it. If so, it just does
the constant fold and return the constant instead of creating an instruction.
<p>Well, that was easy :). In practice, we recommend always using
<tt>llvm.core.Builder</tt> when generating code like this. It has no
"syntactic overhead" for its use (you don't have to uglify your compiler with
constant checks everywhere) and it can dramatically reduce the amount of
LLVM IR that is generated in some cases (particular for languages with a macro
preprocessor or that use a lot of constants).</p>
<p>On the other hand, the <tt>Builder</tt> is limited by the fact that it does
all of its analysis inline with the code as it is built. If you take a slightly
more complex example:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) (1+2+x)*(x+(1+2))</b>
Read a function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 3.000000e+00, %x ; &lt;double&gt; [#uses=1]
%addtmp1 = fadd double %x, 3.000000e+00 ; &lt;double&gt; [#uses=1]
%multmp = fmul double %addtmp, %addtmp1 ; &lt;double&gt; [#uses=1]
ret double %multmp
}
</pre>
</div>
<p>In this case, the LHS and RHS of the multiplication are the same value. We'd
really like to see this generate "<tt>tmp = x+3; result = tmp*tmp;</tt>" instead
of computing "<tt>x+3</tt>" twice.</p>
<p>Unfortunately, no amount of local analysis will be able to detect and correct
this. This requires two transformations: reassociation of expressions (to
make the add's lexically identical) and Common Subexpression Elimination (CSE)
to delete the redundant add instruction. Fortunately, LLVM provides a broad
range of optimizations that you can use, in the form of "passes".</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="optimizerpasses">LLVM Optimization
Passes</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>LLVM provides many optimization passes, which do many different sorts of
things and have different tradeoffs. Unlike other systems, LLVM doesn't hold
to the mistaken notion that one set of optimizations is right for all languages
and for all situations. LLVM allows a compiler implementor to make complete
decisions about what optimizations to use, in which order, and in what
situation.</p>
<p>As a concrete example, LLVM supports both "whole module" passes, which look
across as large of body of code as they can (often a whole file, but if run
at link time, this can be a substantial portion of the whole program). It also
supports and includes "per-function" passes which just operate on a single
function at a time, without looking at other functions. For more information
on passes and how they are run, see the
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html">How to Write a
Pass</a> document and the <a href="http://www.llvm.org/docs/Passes.html">List of
LLVM Passes</a>.</p>
<p>For Kaleidoscope, we are currently generating functions on the fly, one at
a time, as the user types them in. We aren't shooting for the ultimate
optimization experience in this setting, but we also want to catch the easy and
quick stuff where possible. As such, we will choose to run a few per-function
optimizations as the user types the function in. If we wanted to make a "static
Kaleidoscope compiler", we would use exactly the code we have now, except that
we would defer running the optimizer until the entire file has been parsed.</p>
<p>In order to get per-function optimizations going, we need to set up a
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html#passmanager">
FunctionPassManager</a> to hold and organize the LLVM optimizations that we want
to run. Once we have that, we can add a set of optimizations to run. The code
looks like this:</p>
<div class="doc_code">
<pre>
# The function optimization passes manager.
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
...
def main():
# Set up the optimizer pipeline. Start with registering info about how the
# target lays out data structures.
g_llvm_pass_manager.add(g_llvm_executor.target_data)
# Do simple "peephole" optimizations and bit-twiddling optzns.
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
# Reassociate expressions.
g_llvm_pass_manager.add(PASS_REASSOCIATE)
# Eliminate Common SubExpressions.
g_llvm_pass_manager.add(PASS_GVN)
# Simplify the control flow graph (deleting unreachable blocks, etc).
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
g_llvm_pass_manager.initialize()
</pre>
</div>
<p>This code defines a <tt>FunctionPassManager</tt>,
<tt>g_llvm_pass_manager</tt>. Once it is set up, we use a series of "add" calls
to add a bunch of LLVM passes. The first pass is basically boilerplate, it adds
a pass so that later optimizations know how the data structures in the program
are laid out. (The "<tt>g_llvm_executor</tt>" variable is related to the JIT,
which we will get to in the next section.) In this case, we choose to add 4
optimization passes. The passes we chose here are a pretty standard set of
"cleanup" optimizations that are useful for a wide variety of code. I won't
delve into what they do but, believe me, they are a good starting place :).</p>
<p>Once the pass manager is set up, we need to make use of it. We do this by
running it after our newly created function is constructed (in
<tt>FunctionNode.CodeGen</tt>), but before it is returned to the client:</p>
<div class="doc_code">
<pre>
return_value = self.body.CodeGen()
g_llvm_builder.ret(return_value)
# Validate the generated code, checking for consistency.
function.verify()
<b># Optimize the function.
g_llvm_pass_manager.run(function)</b>
</pre>
</div>
<p>As you can see, this is pretty straightforward. The
<tt>FunctionPassManager</tt> optimizes and updates the LLVM Function in place,
improving (hopefully) its body. With this in place, we can try our test above
again:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) (1+2+x)*(x+(1+2))</b>
Read a function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double %x, 3.000000e+00 ; &lt;double&gt; [#uses=2]
%multmp = fmul double %addtmp, %addtmp ; &lt;double&gt; [#uses=1]
ret double %multmp
}
</pre>
</div>
<p>As expected, we now get our nicely optimized code, saving a floating point
add instruction from every execution of this function.</p>
<p>LLVM provides a wide variety of optimizations that can be used in certain
circumstances. Some
<a href="http://www.llvm.org/docs/Passes.html">documentation about the various
passes</a> is available, but it isn't very complete. Another good source of
ideas can come from looking at the passes that <tt>llvm-gcc</tt> or
<tt>llvm-ld</tt> run to get started. The "<tt>opt</tt>" tool allows you to
experiment with passes from the command line, so you can see if they do
anything.</p>
<p>Now that we have reasonable code coming out of our front-end, lets talk about
executing it!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="jit">Adding a JIT Compiler</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Code that is available in LLVM IR can have a wide variety of tools
applied to it. For example, you can run optimizations on it (as we did above),
you can dump it out in textual or binary forms, you can compile the code to an
assembly file (.s) for some target, or you can JIT compile it. The nice thing
about the LLVM IR representation is that it is the "common currency" between
many different parts of the compiler.
</p>
<p>In this section, we'll add JIT compiler support to our interpreter. The
basic idea that we want for Kaleidoscope is to have the user enter function
bodies as they do now, but immediately evaluate the top-level expressions they
type in. For example, if they type in "1 + 2", we should evaluate and print
out 3. If they define a function, they should be able to call it from the
command line.</p>
<p>In order to do this, we first declare and initialize the JIT. This is done
by adding and initializing a global variable:</p>
<div class="doc_code">
<pre>
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
</pre>
</div>
<p>This creates an abstract "Execution Engine" which can be either a JIT
compiler or the LLVM interpreter. LLVM will automatically pick a JIT compiler
for you if one is available for your platform, otherwise it will fall back to
the interpreter.</p>
<p>Once the <tt>ExecutionEngine</tt> is created, the JIT is ready to be used.
We can use the <tt>run_function</tt> method of the execution engine to execute
a compiled function and get its return value. In our case, this means that we
can change the code that parses a top-level expression to look like this:</p>
<div class="doc_code">
<pre>
def HandleTopLevelExpression(self):
try:
function = self.ParseTopLevelExpr().CodeGen()
result = g_llvm_executor.run_function(function, [])
print 'Evaluated to:', result.as_real(Type.double())
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
</pre>
</div>
<p>Recall that we compile top-level expressions into a self-contained LLVM
function that takes no arguments and returns the computed double.</p>
<p>With just these two changes, lets see how Kaleidoscope works now!</p>
<div class="doc_code">
<pre>
ready&gt; <b>4+5</b>
Read a top level expression:
define double @0() {
entry:
ret double 9.000000e+00
}
Evaluated to: 9.0
</pre>
</div>
<p>Well this looks like it is basically working. The dump of the function
shows the "no argument function that always returns double" that we synthesize
for each top-level expression that is typed in. This demonstrates very basic
functionality, but can we do more?</p>
<div class="doc_code">
<pre>
ready&gt; <b>def testfunc(x y) x + y*2</b>
Read a function definition:
define double @testfunc(double %x, double %y) {
entry:
%multmp = fmul double %y, 2.000000e+00 ; &lt;double&gt; [#uses=1]
%addtmp = fadd double %multmp, %x ; &lt;double&gt; [#uses=1]
ret double %addtmp
}
ready&gt; <b>testfunc(4, 10)</b>
Read a top level expression:
define double @0() {
entry:
%calltmp = call double @testfunc(double 4.000000e+00, double 1.000000e+01) ; &lt;double&gt; [#uses=1]
ret double %calltmp
}
<em>Evaluated to: 24.0</em>
</pre>
</div>
<p>This illustrates that we can now call user code, but there is something a bit
subtle going on here. Note that we only invoke the JIT on the anonymous
functions that <em>call testfunc</em>, but we never invoked it
on <em>testfunc</em> itself. What actually happened here is that the JIT
scanned for all non-JIT'd functions transitively called from the anonymous
function and compiled all of them before returning from <tt>run_function()</tt>.
</p>
<p>The JIT provides a number of other more advanced interfaces for things like
freeing allocated machine code, rejit'ing functions to update them, etc.
However, even with this simple code, we get some surprisingly powerful
capabilities - check this out (I removed the dump of the anonymous functions,
you should get the idea by now :) :</p>
<div class="doc_code">
<pre>
ready&gt; <b>extern sin(x)</b>
Read an extern:
declare double @sin(double)
ready&gt; <b>extern cos(x)</b>
Read an extern:
declare double @cos(double)
ready&gt; <b>sin(1.0)</b>
<em>Evaluated to: 0.841470984808</em>
ready&gt; <b>def foo(x) sin(x)*sin(x) + cos(x)*cos(x)</b>
Read a function definition:
define double @foo(double %x) {
entry:
%calltmp = call double @sin(double %x) ; &lt;double&gt; [#uses=1]
%calltmp1 = call double @sin(double %x) ; &lt;double&gt; [#uses=1]
%multmp = fmul double %calltmp, %calltmp1 ; &lt;double&gt; [#uses=1]
%calltmp2 = call double @cos(double %x) ; &lt;double&gt; [#uses=1]
%calltmp3 = call double @cos(double %x) ; &lt;double&gt; [#uses=1]
%multmp4 = fmul double %calltmp2, %calltmp3 ; &lt;double&gt; [#uses=1]
%addtmp = fadd double %multmp, %multmp4 ; &lt;double&gt; [#uses=1]
ret double %addtmp
}
ready&gt; <b>foo(4.0)</b>
<em>Evaluated to: 1.000000</em>
</pre>
</div>
<p>Whoa, how does the JIT know about sin and cos? The answer is surprisingly
simple: in this example, the JIT started execution of a function and got to a
function call. It realized that the function was not yet JIT compiled and
invoked the standard set of routines to resolve the function. In this case,
there is no body defined for the function, so the JIT ended up calling
"<tt>dlsym("sin")</tt>" on the Python process that is hosting our Kaleidoscope
prompt. Since "<tt>sin</tt>" is defined within the JIT's address space, it
simply patches up calls in the module to call the libm version of <tt>sin</tt>
directly.</p>
<p>One interesting application of this is that we can now extend the language
by writing arbitrary C++ code to implement operations. For example, we can
create a C file with the following simple function:
</p>
<div class="doc_code">
<pre>
#include &lt;stdio.h&gt;
double putchard(double x) {
putchar((char)x);
return 0;
}
</pre>
</div>
<p>We can then compile this into a shared library with GCC:</p>
<div class="doc_code">
<pre>
gcc -shared -fPIC -o putchard.so putchard.c
</pre>
</div>
<p>Now we can load this library into the Python process using
<tt>llvm.core.load_library_permanently</tt> and access it from Kaleidoscope to
produce simple output to the console:</p>
<div class="doc_code">
<pre>
>>> <b>import llvm.core</b>
>>> <b>llvm.core.load_library_permanently('/home/max/llvm-py-tutorial/putchard.so')</b>
>>> <b>import kaleidoscope</b>
>>> <b>kaleidoscope.main()</b>
ready&gt; <b>extern putchard(x)</b>
Read an extern:
declare double @putchard(double)
ready&gt; <b>putchard(65) + putchard(66) + putchard(67) + putchard(10)</b>
<em>ABC</em>
Evaluated to: 0.0
</pre>
</div>
<p>Similar code could be used to implement file I/O, console input, and many
other capabilities in Kaleidoscope.</p>
<p>This completes the JIT and optimizer chapter of the Kaleidoscope tutorial. At
this point, we can compile a non-Turing-complete programming language, optimize
and JIT compile it in a user-driven way. Next up we'll look into <a
href="PythonLangImpl5.html">extending the language with control flow
constructs</a>, tackling some interesting LLVM IR issues along the way.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="code">Full Code Listing</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Here is the complete code listing for our running example, enhanced with the
LLVM JIT and optimizer:
</p>
<div class="doc_code">
<pre>
#!/usr/bin/env python
import re
from llvm.core import Module, Constant, Type, Function, Builder, FCMP_ULT
from llvm.ee import ExecutionEngine, TargetData
from llvm.passes import FunctionPassManager
from llvm.passes import (PASS_INSTRUCTION_COMBINING,
PASS_REASSOCIATE,
PASS_GVN,
PASS_CFG_SIMPLIFICATION)
################################################################################
## Globals
################################################################################
# The LLVM module, which holds all the IR code.
g_llvm_module = Module.new('my cool jit')
# The LLVM instruction builder. Created whenever a new function is entered.
g_llvm_builder = None
# A dictionary that keeps track of which values are defined in the current scope
# and what their LLVM representation is.
g_named_values = {}
# The function optimization passes manager.
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
################################################################################
## Lexer
################################################################################
# The lexer yields one of these types for each token.
class EOFToken(object):
pass
class DefToken(object):
pass
class ExternToken(object):
pass
class IdentifierToken(object):
def __init__(self, name): self.name = name
class NumberToken(object):
def __init__(self, value): self.value = value
class CharacterToken(object):
def __init__(self, char): self.char = char
def __eq__(self, other):
return isinstance(other, CharacterToken) and self.char == other.char
def __ne__(self, other): return not self == other
# Regular expressions that tokens and comments of our language.
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
REGEX_COMMENT = re.compile('#.*')
def Tokenize(string):
while string:
# Skip whitespace.
if string[0].isspace():
string = string[1:]
continue
# Run regexes.
comment_match = REGEX_COMMENT.match(string)
number_match = REGEX_NUMBER.match(string)
identifier_match = REGEX_IDENTIFIER.match(string)
# Check if any of the regexes matched and yield the appropriate result.
if comment_match:
comment = comment_match.group(0)
string = string[len(comment):]
elif number_match:
number = number_match.group(0)
yield NumberToken(float(number))
string = string[len(number):]
elif identifier_match:
identifier = identifier_match.group(0)
# Check if we matched a keyword.
if identifier == 'def':
yield DefToken()
elif identifier == 'extern':
yield ExternToken()
else:
yield IdentifierToken(identifier)
string = string[len(identifier):]
else:
# Yield the ASCII value of the unknown character.
yield CharacterToken(string[0])
string = string[1:]
yield EOFToken()
################################################################################
## Abstract Syntax Tree (aka Parse Tree)
################################################################################
# Base class for all expression nodes.
class ExpressionNode(object):
pass
# Expression class for numeric literals like "1.0".
class NumberExpressionNode(ExpressionNode):
def __init__(self, value):
self.value = value
def CodeGen(self):
return Constant.real(Type.double(), self.value)
# Expression class for referencing a variable, like "a".
class VariableExpressionNode(ExpressionNode):
def __init__(self, name):
self.name = name
def CodeGen(self):
if self.name in g_named_values:
return g_named_values[self.name]
else:
raise RuntimeError('Unknown variable name: ' + self.name)
# Expression class for a binary operator.
class BinaryOperatorExpressionNode(ExpressionNode):
def __init__(self, operator, left, right):
self.operator = operator
self.left = left
self.right = right
def CodeGen(self):
left = self.left.CodeGen()
right = self.right.CodeGen()
if self.operator == '+':
return g_llvm_builder.fadd(left, right, 'addtmp')
elif self.operator == '-':
return g_llvm_builder.fsub(left, right, 'subtmp')
elif self.operator == '*':
return g_llvm_builder.fmul(left, right, 'multmp')
elif self.operator == '&lt;':
result = g_llvm_builder.fcmp(FCMP_ULT, left, right, 'cmptmp')
# Convert bool 0 or 1 to double 0.0 or 1.0.
return g_llvm_builder.uitofp(result, Type.double(), 'booltmp')
else:
raise RuntimeError('Unknown binary operator.')
# Expression class for function calls.
class CallExpressionNode(ExpressionNode):
def __init__(self, callee, args):
self.callee = callee
self.args = args
def CodeGen(self):
# Look up the name in the global module table.
callee = g_llvm_module.get_function_named(self.callee)
# Check for argument mismatch error.
if len(callee.args) != len(self.args):
raise RuntimeError('Incorrect number of arguments passed.')
arg_values = [i.CodeGen() for i in self.args]
return g_llvm_builder.call(callee, arg_values, 'calltmp')
# This class represents the "prototype" for a function, which captures its name,
# and its argument names (thus implicitly the number of arguments the function
# takes).
class PrototypeNode(object):
def __init__(self, name, args):
self.name = name
self.args = args
def CodeGen(self):
# Make the function type, eg. double(double,double).
funct_type = Type.function(
Type.double(), [Type.double()] * len(self.args), False)
function = Function.new(g_llvm_module, funct_type, self.name)
# If the name conflicted, there was already something with the same name.
# If it has a body, don't allow redefinition or reextern.
if function.name != self.name:
function.delete()
function = g_llvm_module.get_function_named(self.name)
# If the function already has a body, reject this.
if not function.is_declaration:
raise RuntimeError('Redefinition of function.')
# If F took a different number of args, reject.
if len(callee.args) != len(self.args):
raise RuntimeError('Redeclaration of a function with different number '
'of args.')
# Set names for all arguments and add them to the variables symbol table.
for arg, arg_name in zip(function.args, self.args):
arg.name = arg_name
# Add arguments to variable symbol table.
g_named_values[arg_name] = arg
return function
# This class represents a function definition itself.
class FunctionNode(object):
def __init__(self, prototype, body):
self.prototype = prototype
self.body = body
def CodeGen(self):
# Clear scope.
g_named_values.clear()
# Create a function object.
function = self.prototype.CodeGen()
# Create a new basic block to start insertion into.
block = function.append_basic_block('entry')
global g_llvm_builder
g_llvm_builder = Builder.new(block)
# Finish off the function.
try:
return_value = self.body.CodeGen()
g_llvm_builder.ret(return_value)
# Validate the generated code, checking for consistency.
function.verify()
# Optimize the function.
g_llvm_pass_manager.run(function)
except:
function.delete()
raise
return function
################################################################################
## Parser
################################################################################
class Parser(object):
def __init__(self, tokens, binop_precedence):
self.tokens = tokens
self.binop_precedence = binop_precedence
self.Next()
# Provide a simple token buffer. Parser.current is the current token the
# parser is looking at. Parser.Next() reads another token from the lexer and
# updates Parser.current with its results.
def Next(self):
self.current = self.tokens.next()
# Gets the precedence of the current token, or -1 if the token is not a binary
# operator.
def GetCurrentTokenPrecedence(self):
if isinstance(self.current, CharacterToken):
return self.binop_precedence.get(self.current.char, -1)
else:
return -1
# identifierexpr ::= identifier | identifier '(' expression* ')'
def ParseIdentifierExpr(self):
identifier_name = self.current.name
self.Next() # eat identifier.
if self.current != CharacterToken('('): # Simple variable reference.
return VariableExpressionNode(identifier_name)
# Call.
self.Next() # eat '('.
args = []
if self.current != CharacterToken(')'):
while True:
args.append(self.ParseExpression())
if self.current == CharacterToken(')'):
break
elif self.current != CharacterToken(','):
raise RuntimeError('Expected ")" or "," in argument list.')
self.Next()
self.Next() # eat ')'.
return CallExpressionNode(identifier_name, args)
# numberexpr ::= number
def ParseNumberExpr(self):
result = NumberExpressionNode(self.current.value)
self.Next() # consume the number.
return result
# parenexpr ::= '(' expression ')'
def ParseParenExpr(self):
self.Next() # eat '('.
contents = self.ParseExpression()
if self.current != CharacterToken(')'):
raise RuntimeError('Expected ")".')
self.Next() # eat ')'.
return contents
# primary ::= identifierexpr | numberexpr | parenexpr
def ParsePrimary(self):
if isinstance(self.current, IdentifierToken):
return self.ParseIdentifierExpr()
elif isinstance(self.current, NumberToken):
return self.ParseNumberExpr()
elif self.current == CharacterToken('('):
return self.ParseParenExpr()
else:
raise RuntimeError('Unknown token when expecting an expression.')
# binoprhs ::= (operator primary)*
def ParseBinOpRHS(self, left, left_precedence):
# If this is a binary operator, find its precedence.
while True:
precedence = self.GetCurrentTokenPrecedence()
# If this is a binary operator that binds at least as tightly as the
# current one, consume it; otherwise we are done.
if precedence &lt; left_precedence:
return left
binary_operator = self.current.char
self.Next() # eat the operator.
# Parse the primary expression after the binary operator.
right = self.ParsePrimary()
# If binary_operator binds less tightly with right than the operator after
# right, let the pending operator take right as its left.
next_precedence = self.GetCurrentTokenPrecedence()
if precedence &lt; next_precedence:
right = self.ParseBinOpRHS(right, precedence + 1)
# Merge left/right.
left = BinaryOperatorExpressionNode(binary_operator, left, right)
# expression ::= primary binoprhs
def ParseExpression(self):
left = self.ParsePrimary()
return self.ParseBinOpRHS(left, 0)
# prototype ::= id '(' id* ')'
def ParsePrototype(self):
if not isinstance(self.current, IdentifierToken):
raise RuntimeError('Expected function name in prototype.')
function_name = self.current.name
self.Next() # eat function name.
if self.current != CharacterToken('('):
raise RuntimeError('Expected "(" in prototype.')
self.Next() # eat '('.
arg_names = []
while isinstance(self.current, IdentifierToken):
arg_names.append(self.current.name)
self.Next()
if self.current != CharacterToken(')'):
raise RuntimeError('Expected ")" in prototype.')
# Success.
self.Next() # eat ')'.
return PrototypeNode(function_name, arg_names)
# definition ::= 'def' prototype expression
def ParseDefinition(self):
self.Next() # eat def.
proto = self.ParsePrototype()
body = self.ParseExpression()
return FunctionNode(proto, body)
# toplevelexpr ::= expression
def ParseTopLevelExpr(self):
proto = PrototypeNode('', [])
return FunctionNode(proto, self.ParseExpression())
# external ::= 'extern' prototype
def ParseExtern(self):
self.Next() # eat extern.
return self.ParsePrototype()
# Top-Level parsing
def HandleDefinition(self):
self.Handle(self.ParseDefinition, 'Read a function definition:')
def HandleExtern(self):
self.Handle(self.ParseExtern, 'Read an extern:')
def HandleTopLevelExpression(self):
try:
function = self.ParseTopLevelExpr().CodeGen()
result = g_llvm_executor.run_function(function, [])
print 'Evaluated to:', result.as_real(Type.double())
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
def Handle(self, function, message):
try:
print message, function().CodeGen()
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
################################################################################
## Main driver code.
################################################################################
def main():
# Set up the optimizer pipeline. Start with registering info about how the
# target lays out data structures.
g_llvm_pass_manager.add(g_llvm_executor.target_data)
# Do simple "peephole" optimizations and bit-twiddling optzns.
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
# Reassociate expressions.
g_llvm_pass_manager.add(PASS_REASSOCIATE)
# Eliminate Common SubExpressions.
g_llvm_pass_manager.add(PASS_GVN)
# Simplify the control flow graph (deleting unreachable blocks, etc).
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
g_llvm_pass_manager.initialize()
# Install standard binary operators.
# 1 is lowest possible precedence. 40 is the highest.
operator_precedence = {
'&lt;': 10,
'+': 20,
'-': 20,
'*': 40
}
# Run the main "interpreter loop".
while True:
print 'ready&gt;',
try:
raw = raw_input()
except KeyboardInterrupt:
break
parser = Parser(Tokenize(raw), operator_precedence)
while True:
# top ::= definition | external | expression | EOF
if isinstance(parser.current, EOFToken):
break
if isinstance(parser.current, DefToken):
parser.HandleDefinition()
elif isinstance(parser.current, ExternToken):
parser.HandleExtern()
else:
parser.HandleTopLevelExpression()
# Print out all of the generated code.
print '\n', g_llvm_module
if __name__ == '__main__':
main()
</pre>
</div>
<a href="PythonLangImpl5.html">Next: Extending the language: control flow</a>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://max99x.com">Max Shawabkeh</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,375 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Conclusion and other useful LLVM tidbits</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Conclusion and other useful LLVM
tidbits</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 8
<ol>
<li><a href="#conclusion">Tutorial Conclusion</a></li>
<li><a href="#llvmirproperties">Properties of LLVM IR</a>
<ul>
<li><a href="#targetindep">Target Independence</a></li>
<li><a href="#safety">Safety Guarantees</a></li>
<li><a href="#langspecific">Language-Specific Optimizations</a></li>
</ul>
</li>
<li><a href="#tipsandtricks">Tips and Tricks</a>
<ul>
<li><a href="#offsetofsizeof">Implementing portable
offsetof/sizeof</a></li>
<li><a href="#gcstack">Garbage Collected Stack Frames</a></li>
</ul>
</li>
</ol>
</li>
</ul>
<div class="doc_author">
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a></p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="conclusion">Tutorial Conclusion</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Welcome to the the final chapter of the
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
with LLVM</a>" tutorial. In the course of this tutorial, we have grown
our little Kaleidoscope language from being a useless toy, to being a
semi-interesting (but probably still useless) toy. :)</p>
<p>It is interesting to see how far we've come, and how little code it has
taken. We built the entire lexer, parser, AST, code generator, and an
interactive run-loop (with a JIT!) by-hand in under 540 lines of
(non-comment/non-blank) code.</p>
<p>Our little language supports a couple of interesting features: it supports
user defined binary and unary operators, it uses JIT compilation for immediate
evaluation, and it supports a few control flow constructs with SSA construction.
</p>
<p>Part of the idea of this tutorial was to show you how easy and fun it can be
to define, build, and play with languages. Building a compiler need not be a
scary or mystical process! Now that you've seen some of the basics, I strongly
encourage you to take the code and hack on it. For example, try adding:</p>
<ul>
<li><b>global variables</b> - While global variables have questional value in
modern software engineering, they are often useful when putting together quick
little hacks like the Kaleidoscope compiler itself. Fortunately, our current
setup makes it very easy to add global variables: just have value lookup check
to see if an unresolved variable is in the global variable symbol table before
rejecting it. To create a new global variable, make an instance of the LLVM
<tt>GlobalVariable</tt> class.</li>
<li><b>typed variables</b> - Kaleidoscope currently only supports variables of
type double. This gives the language a very nice elegance, because only
supporting one type means that you never have to specify types. Different
languages have different ways of handling this. The easiest way is to require
the user to specify types for every variable definition, and record the type
of the variable in the symbol table along with its Value*.</li>
<li><b>arrays, structs, vectors, etc</b> - Once you add types, you can start
extending the type system in all sorts of interesting ways. Simple arrays are
very easy and are quite useful for many different applications. Adding them is
mostly an exercise in learning how the LLVM <a
href="http://www.llvm.org/docs/LangRef.html#i_getelementptr">getelementptr</a>
instruction works: it is so nifty/unconventional, it <a
href="http://www.llvm.org/docs/GetElementPtr.html">has its own FAQ</a>! If you
add support for recursive types (e.g. linked lists), make sure to read the <a
href="http://www.llvm.org/docs/ProgrammersManual.html#TypeResolve">section in
the LLVM Programmer's Manual</a> that describes how to construct them.</li>
<li><b>standard runtime</b> - Our current language allows the user to access
arbitrary external functions, and we use it for things like "putchard". As you
extend the language to add higher-level constructs, often these constructs make
the most sense if they are lowered to calls into a language-supplied runtime.
For example, if you add hash tables to the language, it would probably make
sense to add the routines to a runtime, instead of inlining them all the way.
</li>
<li><b>memory management</b> - Currently we can only access the stack in
Kaleidoscope. It would also be useful to be able to allocate heap memory,
either with calls to the standard libc malloc/free interface or with a garbage
collector. If you would like to use garbage collection, note that LLVM fully
supports <a href="http://www.llvm.org/docs/GarbageCollection.html">Accurate
Garbage Collection</a> including algorithms that move objects and need to
scan/update the stack.</li>
<li><b>debugger support</b> - LLVM supports generation of <a
href="http://www.llvm.org/docs/SourceLevelDebugging.html">DWARF Debug info</a>
which is understood by common debuggers like GDB. Adding support for debug info
is fairly straightforward. The best way to understand it is to compile some
C/C++ code with "<tt>llvm-gcc -g -O0</tt>" and taking a look at what it
produces.</li>
<li><b>exception handling support</b> - LLVM supports generation of <a
href="http://www.llvm.org/docs/ExceptionHandling.html">zero cost exceptions</a>
which interoperate with code compiled in other languages. You could also
generate code by implicitly making every function return an error value and
checking it. You could also make explicit use of setjmp/longjmp. There are
many different ways to go here.</li>
<li><b>object orientation, generics, database access, complex numbers,
geometric programming, ...</b> - Really, there is
no end of crazy features that you can add to the language.</li>
<li><b>unusual domains</b> - We've been talking about applying LLVM to a domain
that many people are interested in: building a compiler for a specific language.
However, there are many other domains that can use compiler technology that are
not typically considered. For example, LLVM has been used to implement OpenGL
graphics acceleration, translate C++ code to ActionScript, and many other
cute and clever things. Maybe you will be the first to JIT compile a regular
expression interpreter into native code with LLVM?</li>
</ul>
<p>
Have fun - try doing something crazy and unusual. Building a language like
everyone else always has, is much less fun than trying something a little crazy
or off the wall and seeing how it turns out. If you get stuck or want to talk
about it, feel free to email the <a
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
list</a>: it has lots of people who are interested in languages and are often
willing to help out.
</p>
<p>Before we end this tutorial, I want to talk about some "tips and tricks" for
generating LLVM IR. These are some of the more subtle things that may not be
obvious, but are very useful if you want to take advantage of LLVM's
capabilities.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="llvmirproperties">Properties of the LLVM
IR</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>We have a couple common questions about code in the LLVM IR form - let's just
get these out of the way right now, shall we?</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="targetindep">Target
Independence</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Kaleidoscope is an example of a "portable language": any program written in
Kaleidoscope will work the same way on any target that it runs on. Many other
languages have this property, e.g. LISP, Java, Haskell, Javascript, Python, etc.
(note that while these languages are portable, not all their libraries are).</p>
<p>One nice aspect of LLVM is that it is often capable of preserving target
independence in the IR: you can take the LLVM IR for a Kaleidoscope-compiled
program and run it on any target that LLVM supports, even emitting C code and
compiling that on targets that LLVM doesn't support natively. You can trivially
tell that the Kaleidoscope compiler generates target-independent code because it
never queries for any target-specific information when generating code.</p>
<p>The fact that LLVM provides a compact, target-independent, representation for
code gets a lot of people excited. Unfortunately, these people are usually
thinking about C or a language from the C family when they are asking questions
about language portability. I say "unfortunately", because there is really no
way to make (fully general) C code portable, other than shipping the source code
around (and of course, C source code is not actually portable in general
either - ever port a really old application from 32- to 64-bits?).</p>
<p>The problem with C (again, in its full generality) is that it is heavily
laden with target specific assumptions. As one simple example, the preprocessor
often destructively removes target-independence from the code when it processes
the input text:</p>
<div class="doc_code">
<pre>
#ifdef __i386__
int X = 1;
#else
int X = 42;
#endif
</pre>
</div>
<p>While it is possible to engineer more and more complex solutions to problems
like this, it cannot be solved in full generality in a way that is better than
shipping the actual source code.</p>
<p>That said, there are interesting subsets of C that can be made portable. If
you are willing to fix primitive types to a fixed size (say int = 32-bits,
and long = 64-bits), don't care about ABI compatibility with existing binaries,
and are willing to give up some other minor features, you can have portable
code. This can make sense for specialized domains such as an
in-kernel language.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="safety">Safety Guarantees</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Many of the languages above are also "safe" languages: it is impossible for
a program written in Java to corrupt its address space and crash the process
(assuming the JVM has no bugs).
Safety is an interesting property that requires a combination of language
design, runtime support, and often operating system support.</p>
<p>It is certainly possible to implement a safe language in LLVM, but LLVM IR
does not itself guarantee safety. The LLVM IR allows unsafe pointer casts,
use after free bugs, buffer over-runs, and a variety of other problems. Safety
needs to be implemented as a layer on top of LLVM and, conveniently, several
groups have investigated this. Ask on the <a
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
list</a> if you are interested in more details.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="langspecific">Language-Specific
Optimizations</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>One thing about LLVM that turns off many people is that it does not solve all
the world's problems in one system (sorry 'world hunger', someone else will have
to solve you some other day). One specific complaint is that people perceive
LLVM as being incapable of performing high-level language-specific optimization:
LLVM "loses too much information".</p>
<p>Unfortunately, this is really not the place to give you a full and unified
version of "Chris Lattner's theory of compiler design". Instead, I'll make a
few observations:</p>
<p>First, you're right that LLVM does lose information. For example, as of this
writing, there is no way to distinguish in the LLVM IR whether an SSA-value came
from a C "int" or a C "long" on an ILP32 machine (other than debug info). Both
get compiled down to an 'i32' value and the information about what it came from
is lost. The more general issue here, is that the LLVM type system uses
"structural equivalence" instead of "name equivalence". Another place this
surprises people is if you have two types in a high-level language that have the
same structure (e.g. two different structs that have a single int field): these
types will compile down into a single LLVM type and it will be impossible to
tell what it came from.</p>
<p>Second, while LLVM does lose information, LLVM is not a fixed target: we
continue to enhance and improve it in many different ways. In addition to
adding new features (LLVM did not always support exceptions or debug info), we
also extend the IR to capture important information for optimization (e.g.
whether an argument is sign or zero extended, information about pointers
aliasing, etc). Many of the enhancements are user-driven: people want LLVM to
include some specific feature, so they go ahead and extend it.</p>
<p>Third, it is <em>possible and easy</em> to add language-specific
optimizations, and you have a number of choices in how to do it. As one trivial
example, it is easy to add language-specific optimization passes that
"know" things about code compiled for a language. In the case of the C family,
there is an optimization pass that "knows" about the standard C library
functions. If you call "exit(0)" in main(), it knows that it is safe to
optimize that into "return 0;" because C specifies what the 'exit'
function does.</p>
<p>In addition to simple library knowledge, it is possible to embed a variety of
other language-specific information into the LLVM IR. If you have a specific
need and run into a wall, please bring the topic up on the llvmdev list. At the
very worst, you can always treat LLVM as if it were a "dumb code generator" and
implement the high-level optimizations you desire in your front-end, on the
language-specific AST.
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="tipsandtricks">Tips and Tricks</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>There is a variety of useful tips and tricks that you come to know after
working on/with LLVM that aren't obvious at first glance. Instead of letting
everyone rediscover them, this section talks about some of these issues.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="offsetofsizeof">Implementing portable
offsetof/sizeof</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>One interesting thing that comes up, if you are trying to keep the code
generated by your compiler "target independent", is that you often need to know
the size of some LLVM type or the offset of some field in an llvm structure.
For example, you might need to pass the size of a type into a function that
allocates memory.</p>
<p>Unfortunately, this can vary widely across targets: for example the width of
a pointer is trivially target-specific. However, there is a <a
href="http://nondot.org/sabre/LLVMNotes/SizeOf-OffsetOf-VariableSizedStructs.txt">clever
way to use the getelementptr instruction</a> that allows you to compute this
in a portable way.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="gcstack">Garbage Collected
Stack Frames</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Some languages want to explicitly manage their stack frames, often so that
they are garbage collected or to allow easy implementation of closures. There
are often better ways to implement these features than explicit stack frames,
but <a
href="http://nondot.org/sabre/LLVMNotes/ExplicitlyManagedStackFrames.txt">LLVM
does support them,</a> if you want. It requires your front-end to convert the
code into <a
href="http://en.wikipedia.org/wiki/Continuation-passing_style">Continuation
Passing Style</a> and the use of tail calls (which LLVM also supports).</p>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -43,7 +43,7 @@ the llvm-py contributors.</p></div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -91,7 +91,7 @@ update and merge before sending patches etc.</p></div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -48,6 +48,7 @@ below). 0.6 works only with LLVM 2.7.</p></div>
package.</p></div>
</div>
</div>
<div class="sect1">
<h2 id="changelog">Changelog</h2>
<div class="sectionbody">
<div class="listingblock">
@ -129,10 +130,11 @@ package.</p></div>
* Initial release.</tt></pre>
</div></div>
</div>
</div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -31,9 +31,11 @@
<div id="header">
<h1>Examples and LLVM Tutorials</h1>
</div>
<div class="sect1">
<h2 id="_examples">Examples</h2>
<div class="sectionbody">
<h3 id="_a_simple_function">A Simple Function</h3><div style="clear:left"></div>
<div class="sect2">
<h3 id="_a_simple_function">A Simple Function</h3>
<div class="paragraph"><p>Let&#8217;s create a (LLVM) module containing a single function, corresponding
to the <tt>C</tt> function:</p></div>
<div class="listingblock">
@ -106,7 +108,9 @@ entry:
ret i32 %tmp
}</tt></pre>
</div></div>
<h3 id="_adding_jit_compilation">Adding JIT Compilation</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_adding_jit_compilation">Adding JIT Compilation</h3>
<div class="paragraph"><p>Let&#8217;s compile this function in-memory and run it.</p></div>
<div class="listingblock">
<div class="content"><!-- Generator: GNU source-highlight 3.1.3
@ -151,75 +155,74 @@ retval <span style="color: #990000">=</span> ee<span style="color: #990000">.</s
<pre><tt>returned 142</tt></pre>
</div></div>
</div>
</div>
</div>
<div class="sect1">
<h2 id="_llvm_tutorials">LLVM Tutorials</h2>
<div class="sectionbody">
<div class="paragraph"><p>The <a href="http://www.llvm.org/docs/tutorial/">LLVM tutorials</a> have been
ported to llvm-py. Below are the links to the original LLVM tutorial and
the corresponding Python code using llvm-py:</p></div>
<div class="paragraph"><div class="title">Simple JIT Tutorials</div><p>(contributed by Sebastien Binet)</p></div>
<div class="paragraph"><div class="title">Simple JIT Tutorials</div><p>The following JIT tutorials were contributed by Sebastien Binet.</p></div>
<div class="olist arabic"><ol class="arabic">
<li>
<p>
A First Function
<a href="http://www.llvm.org/docs/tutorial/JITTutorial1.html">LLVM</a>
<a href="examples/JITTutorial1.html">llvm-py</a>
<a href="examples/JITTutorial1.html">A First Function</a>
</p>
</li>
<li>
<p>
A More Complicated Function
<a href="http://www.llvm.org/docs/tutorial/JITTutorial2.html">LLVM</a>
<a href="examples/JITTutorial2.html">llvm-py</a>
<a href="examples/JITTutorial2.html">A More Complicated Function</a>
</p>
</li>
</ol></div>
<div class="olist arabic"><div class="title">Kaleidoscope: Implementing a Language with LLVM</div><ol class="arabic">
<div class="paragraph" id="kaleidoscope"><div class="title">Kaleidoscope: Implementing a Language with LLVM</div><p>The LLVM <a href="http://www.llvm.org/docs/tutorial/">Kaleidoscope</a> tutorial
has been ported to llvm-py by Max Shawabkeh.</p></div>
<div class="olist arabic"><ol class="arabic">
<li>
<p>
Tutorial Introduction and the Lexer (TODO)
<a href="kaleidoscope/PythonLangImpl1.html">Tutorial Introduction and the Lexer</a>
</p>
</li>
<li>
<p>
Implementing a Parser and AST (TODO)
<a href="kaleidoscope/PythonLangImpl2.html">Implementing a Parser and AST</a>
</p>
</li>
<li>
<p>
Implementing Code Generation to LLVM IR (TODO)
<a href="kaleidoscope/PythonLangImpl3.html">Implementing Code Generation to LLVM IR</a>
</p>
</li>
<li>
<p>
Adding JIT and Optimizer Support (TODO)
<a href="kaleidoscope/PythonLangImpl4.html">Adding JIT and Optimizer Support</a>
</p>
</li>
<li>
<p>
Extending the language: control flow (TODO)
<a href="kaleidoscope/PythonLangImpl5.html">Extending the language: control flow</a>
</p>
</li>
<li>
<p>
Extending the language: user-defined operators (TODO)
<a href="kaleidoscope/PythonLangImpl6.html">Extending the language: user-defined operators</a>
</p>
</li>
<li>
<p>
Extending the language: mutable variables / SSA construction (TODO)
<a href="kaleidoscope/PythonLangImpl7.html">Extending the language: mutable variables / SSA construction</a>
</p>
</li>
<li>
<p>
Conclusion and other useful LLVM tidbits (TODO)
<a href="kaleidoscope/PythonLangImpl8.html">Conclusion and other useful LLVM tidbits</a>
</p>
</li>
</ol></div>
</div>
</div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -47,10 +47,19 @@ discover that any of these claims are wrong, feel free to send across
a patch.</p></div>
</div>
</div>
<div class="sect1">
<h2 id="_news">News</h2>
<div class="sectionbody">
<div class="dlist"><dl>
<dt class="hdlist1">
26-Sep-2010
</dt>
<dd>
<p>
LLVM tutorial <a href="examples.html#kaleidoscope">ported</a> by Max Shawabkeh!
</p>
</dd>
<dt class="hdlist1">
31-Aug-2010
</dt>
<dd>
@ -60,10 +69,11 @@ a patch.</p></div>
</dd>
</dl></div>
</div>
</div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -0,0 +1,389 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Tutorial Introduction and the Lexer</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<meta name="author" content="Max Shawabkeh">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Tutorial Introduction and the Lexer</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 1
<ol>
<li><a href="#intro">Tutorial Introduction</a></li>
<li><a href="#language">The Basic Language</a></li>
<li><a href="#lexer">The Lexer</a></li>
</ol>
</li>
<li><a href="PythonLangImpl2.html">Chapter 2</a>: Implementing a Parser and
AST</li>
</ul>
<div class="doc_author">
<p>
Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
and <a href="http://max99x.com">Max Shawabkeh</a>
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="intro">Tutorial Introduction</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Welcome to the "Implementing a language with LLVM" tutorial. This tutorial
runs through the implementation of a simple language, showing how fun and
easy it can be. This tutorial will get you up and started as well as help to
build a framework you can extend to other languages. The code in this tutorial
can also be used as a playground to hack on other LLVM specific things.
</p>
<p>The goal of this tutorial is to progressively unveil our language, describing
how it is built up over time. This will let us cover a fairly broad range of
language design and LLVM-specific usage issues, showing and explaining the code
for it all along the way, without overwhelming you with tons of details up
front.</p>
<p>It is useful to point out ahead of time that this tutorial is really about
teaching compiler techniques and LLVM specifically, <em>not</em> about teaching
modern and sane software engineering principles. In practice, this means that
we'll take a number of shortcuts to simplify the exposition. If you dig in and
use the code as a basis for future projects, fixing its deficiencies shouldn't
be hard.</p>
<p>We've tried to put this tutorial together in a way that makes chapters easy
to skip over if you are already familiar with or are uninterested in the various
pieces. The structure of the tutorial is:</p>
<ul>
<li><b><a href="#language">Chapter #1</a>: Introduction to the Kaleidoscope
language, and the definition of its Lexer</b> - This shows where we are going
and the basic functionality that we want it to do. In order to make this
tutorial maximally understandable and hackable, we choose to implement
everything in Python instead of using lexer and parser generators. LLVM
obviously works just fine with such tools, feel free to use one if you prefer.
</li>
<li><b><a href="PythonLangImpl2.html">Chapter #2</a>: Implementing a Parser and
AST</b> - With the lexer in place, we can talk about parsing techniques and
basic AST construction. This tutorial describes recursive descent parsing and
operator precedence parsing. Nothing in Chapters 1 or 2 is LLVM-specific,
the code doesn't even import the LLVM modules at this point. :)</li>
<li><b><a href="PythonLangImpl3.html">Chapter #3</a>: Code generation to LLVM
IR</b> - With the AST ready, we can show off how easy generation of LLVM IR
really is.</li>
<li><b><a href="PythonLangImpl4.html">Chapter #4</a>: Adding JIT and Optimizer
Support</b> - Because a lot of people are interested in using LLVM as a JIT,
we'll dive right into it and show you the 3 lines it takes to add JIT support.
LLVM is also useful in many other ways, but this is one simple and "sexy" way
to shows off its power. :)</li>
<li><b><a href="PythonLangImpl5.html">Chapter #5</a>: Extending the Language:
Control Flow</b> - With the language up and running, we show how to extend it
with control flow operations (if/then/else and a 'for' loop). This gives us a
chance to talk about simple SSA construction and control flow.</li>
<li><b><a href="PythonLangImpl6.html">Chapter #6</a>: Extending the Language:
User-defined Operators</b> - This is a silly but fun chapter that talks about
extending the language to let the user program define their own arbitrary
unary and binary operators (with assignable precedence!). This lets us build a
significant piece of the "language" as library routines.</li>
<li><b><a href="PythonLangImpl7.html">Chapter #7</a>: Extending the Language:
Mutable Variables</b> - This chapter talks about adding user-defined local
variables along with an assignment operator. The interesting part about this
is how easy and trivial it is to construct SSA form in LLVM: no, LLVM does
<em>not</em> require your front-end to construct SSA form!</li>
<li><b><a href="PythonLangImpl8.html">Chapter #8</a>: Conclusion and other
useful LLVM tidbits</b> - This chapter wraps up the series by talking about
potential ways to extend the language, but also includes a bunch of pointers to
info about "special topics" like adding garbage collection support, exceptions,
debugging, support for "spaghetti stacks", and a bunch of other tips and
tricks.</li>
</ul>
<p>By the end of the tutorial, we'll have written a bit less than 540 lines of
non-comment, non-blank, lines of code. With this small amount of code, we'll
have built up a very reasonable compiler for a non-trivial language including
a hand-written lexer, parser, AST, as well as code generation support with a JIT
compiler. While other systems may have interesting "hello world" tutorials,
I think the breadth of this tutorial is a great testament to the strengths of
LLVM and why you should consider it if you're interested in language or compiler
design.</p>
<p>A note about this tutorial: we expect you to extend the language and play
with it on your own. Take the code and go crazy hacking away at it, compilers
don't need to be scary creatures - it can be a lot of fun to play with
languages!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="language">The Basic Language</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>This tutorial will be illustrated with a toy language that we'll call
"<a href="http://en.wikipedia.org/wiki/Kaleidoscope">Kaleidoscope</a>" (derived
from "meaning beautiful, form, and view").
Kaleidoscope is a procedural language that allows you to define functions, use
conditionals, math, etc. Over the course of the tutorial, we'll extend
Kaleidoscope to support the if/then/else construct, a for loop, user defined
operators, JIT compilation with a simple command line interface, etc.</p>
<p>Because we want to keep things simple, the only datatype in Kaleidoscope is a
64-bit floating point type. As such, all values are implicitly double precision
and the language doesn't require type declarations. This gives the language a
very nice and simple syntax. For example, the following simple example computes
<a href="http://en.wikipedia.org/wiki/Fibonacci_number">Fibonacci numbers:</a>
</p>
<div class="doc_code">
<pre>
# Compute the x'th fibonacci number.
def fib(x)
if x &lt; 3 then
1
else
fib(x-1)+fib(x-2)
# This expression will compute the 40th number.
fib(40)
</pre>
</div>
<p>We also allow Kaleidoscope to call into standard library functions (the LLVM
JIT makes this completely trivial). This means that you can use the 'extern'
keyword to define a function before you use it (this is also useful for mutually
recursive functions). For example:</p>
<div class="doc_code">
<pre>
extern sin(arg);
extern cos(arg);
extern atan2(arg1 arg2);
atan2(sin(0.4), cos(42))
</pre>
</div>
<p>A more interesting example is included in Chapter 6 where we write a little
Kaleidoscope application that <a href="PythonLangImpl6.html#example">displays
a Mandelbrot Set</a> at various levels of magnification.</p>
<p>Lets dive into the implementation of this language!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="lexer">The Lexer</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>When it comes to implementing a language, the first thing needed is
the ability to process a text file and recognize what it says. The traditional
way to do this is to use a "<a
href="http://en.wikipedia.org/wiki/Lexical_analysis">lexer</a>" (aka 'scanner')
to break the input up into "tokens". Each token returned by the lexer includes
a token type and potentially some metadata (e.g. the numeric value of a number).
First, we define the possibilities:</p>
<div class="doc_code">
<pre>
# The lexer yields one of these types for each token.
class EOFToken(object):
pass
class DefToken(object):
pass
class ExternToken(object):
pass
class IdentifierToken(object):
def __init__(self, name): self.name = name
class NumberToken(object):
def __init__(self, value): self.value = value
class CharacterToken(object):
def __init__(self, char): self.char = char
def __eq__(self, other):
return isinstance(other, CharacterToken) and self.char == other.char
def __ne__(self, other): return not self == other
</pre>
</div>
<p>Each token yielded by our lexer will be of one of the above types. For simple
tokens that are always the same, like the "def" keyword, the lexer will yield
<tt>DefToken()</tt>. Identifiers, numbers and characters, on the other
hand, have extra data, so when the lexer encounteres the number 123.45, it will
emit it as <tt>NumberToken(123.45)</tt>. An identifier <tt>foo</tt> will be
emitted as <tt>IdentifierToken('foo')</tt>. And finally, an unknown character
like '+' will be returned as <tt>CharacterToken('+')</tt>. You may notice that
we overload the equality and inequality operators for the characters; this will
later simplify character comparisons in the parser code.</p>
<p>The actual implementation of the lexer is a single function called
<tt>Tokenize</tt>, which takes a string and
<a href="http://docs.python.org/reference/simple_stmts.html#the-yield-statement">yields</a>
tokens. For simplicity, we will use
<a href="http://docs.python.org/library/re.html">regular
expressions</a> to parse out the tokens. This is terribly inefficient, but
perfectly sufficient for our needs.</p>
<p>First, we define the regular expressions for our tokens. Numbers and strings
of digits, optionally followed by a period and another string of digits.
Identifiers (and keywords) are alphanumeric string starting with a letter and
comments are anything between a hash (<tt>#</tt>) and the end of the line.
<div class="doc_code">
<pre>
import re
...
# Regular expressions that tokens and comments of our language.
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
REGEX_COMMENT = re.compile('#.*')
</pre>
</div>
<p>
Next, let's start defining the <tt>Tokenize</tt> function itself. The first
thing we need to do is set up a loop that scans the string, while ignoring
whitespace between tokens:</p>
<div class="doc_code">
<pre>
def Tokenize(string):
while string:
# Skip whitespace.
if string[0].isspace():
string = string[1:]
continue
...
</pre>
</div>
<p>Next we want to find out what the next token is. For this we run the regexes
we defined above on the remainder of the string. To simplify the rest of the
code, we run all three regexes each time. As mentioned above, inefficiencies are
ignored for the purpose of this tutorial:<p>
<div class="doc_code">
<pre>
# Run regexes.
comment_match = REGEX_COMMENT.match(string)
number_match = REGEX_NUMBER.match(string)
identifier_match = REGEX_IDENTIFIER.match(string)
</pre>
</div>
<p>Now se check if any of the regexes matched. For comments, we simply
ignore the captured match:</p>
<div class="doc_code">
<pre>
# Check if any of the regexes matched and yield the appropriate result.
if comment_match:
comment = comment_match.group(0)
string = string[len(comment):]
</pre>
</div>
<p>For numbers, we yield the captured match, converted to a float and tagged
with the appropriate token type:</p>
<div class="doc_code">
<pre>
elif number_match:
number = number_match.group(0)
yield NumberToken(float(number))
string = string[len(number):]
</pre>
</div>
<p>The identifier case is a little more complex. We have to check for keywords
to decide whether we have captured an identifier or a keyword:</p>
<div class="doc_code">
<pre>
elif identifier_match:
identifier = identifier_match.group(0)
# Check if we matched a keyword.
if identifier == 'def':
yield DefToken()
elif identifier == 'extern':
yield ExternToken()
else:
yield IdentifierToken(identifier)
string = string[len(identifier):]
</pre>
</div>
<p>Finally, if we haven't recognized a comment, a number of an identifier, we
yield the current character as an "unknown character" token. This is used, for
example, for operators like <tt>+</tt> or <tt>*</tt>:</p>
<div class="doc_code">
<pre>
else:
# Yield the unknown character.
yield CharacterToken(string[0])
string = string[1:]
</pre>
</div>
<p>Once we're done with the
loop, we return a final end-of-file token:</p>
<div class="doc_code">
<pre>
yield EOFToken()
</pre>
</div>
<p>With this, we have the complete lexer for the basic Kaleidoscope language
(the <a href="PythonLangImpl2.html#code">full code listing</a> for the Lexer is
available in the <a href="PythonLangImpl2.html">next chapter</a> of the
tutorial). Next we'll <a href="PythonLangImpl2.html">build a simple parser that
uses this to build an Abstract Syntax Tree</a>. When we have that, we'll
include a driver so that you can use the lexer and parser together.
</p>
<a href="PythonLangImpl2.html">Next: Implementing a Parser and AST</a>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://max99x.com">Max Shawabkeh</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,999 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Adding JIT and Optimizer Support</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<meta name="author" content="Max Shawabkeh">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Adding JIT and Optimizer Support</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 4
<ol>
<li><a href="#intro">Chapter 4 Introduction</a></li>
<li><a href="#trivialconstfold">Trivial Constant Folding</a></li>
<li><a href="#optimizerpasses">LLVM Optimization Passes</a></li>
<li><a href="#jit">Adding a JIT Compiler</a></li>
<li><a href="#code">Full Code Listing</a></li>
</ol>
</li>
<li><a href="PythonLangImpl5.html">Chapter 5</a>: Extending the Language:
Control Flow</li>
</ul>
<div class="doc_author">
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a>
and <a href="http://max99x.com">Max Shawabkeh</a>
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="intro">Chapter 4 Introduction</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Welcome to Chapter 4 of the
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
with LLVM</a>" tutorial. Chapters 1-3 described the implementation of a simple
language and added support for generating LLVM IR. This chapter describes
two new techniques: adding optimizer support to your language, and adding JIT
compiler support. These additions will demonstrate how to get nice, efficient
code for the Kaleidoscope language.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="trivialconstfold">Trivial Constant
Folding</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Our demonstration for Chapter 3 is elegant and easy to extend. Unfortunately,
it does not produce wonderful code. The LLVM Builder, however, does give us
obvious optimizations when compiling simple code:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) 1+2+x</b>
Read function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 3.000000e+00, %x
ret double %addtmp
}
</pre>
</div>
<p>This code is not a literal transcription of the AST built by parsing the
input. That would be:
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) 1+2+x</b>
Read function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 2.000000e+00, 1.000000e+00
%addtmp1 = fadd double %addtmp, %x
ret double %addtmp1
}
</pre>
</div>
<p>Constant folding, as seen above, in particular, is a very common and very
important optimization: so much so that many language implementors implement
constant folding support in their AST representation.</p>
<p>With LLVM, you don't need this support in the AST. Since all calls to build
LLVM IR go through the LLVM IR builder, the builder itself checked to see if
there was a constant folding opportunity when you call it. If so, it just does
the constant fold and return the constant instead of creating an instruction.
<p>Well, that was easy :). In practice, we recommend always using
<tt>llvm.core.Builder</tt> when generating code like this. It has no
"syntactic overhead" for its use (you don't have to uglify your compiler with
constant checks everywhere) and it can dramatically reduce the amount of
LLVM IR that is generated in some cases (particular for languages with a macro
preprocessor or that use a lot of constants).</p>
<p>On the other hand, the <tt>Builder</tt> is limited by the fact that it does
all of its analysis inline with the code as it is built. If you take a slightly
more complex example:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) (1+2+x)*(x+(1+2))</b>
Read a function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double 3.000000e+00, %x ; &lt;double&gt; [#uses=1]
%addtmp1 = fadd double %x, 3.000000e+00 ; &lt;double&gt; [#uses=1]
%multmp = fmul double %addtmp, %addtmp1 ; &lt;double&gt; [#uses=1]
ret double %multmp
}
</pre>
</div>
<p>In this case, the LHS and RHS of the multiplication are the same value. We'd
really like to see this generate "<tt>tmp = x+3; result = tmp*tmp;</tt>" instead
of computing "<tt>x+3</tt>" twice.</p>
<p>Unfortunately, no amount of local analysis will be able to detect and correct
this. This requires two transformations: reassociation of expressions (to
make the add's lexically identical) and Common Subexpression Elimination (CSE)
to delete the redundant add instruction. Fortunately, LLVM provides a broad
range of optimizations that you can use, in the form of "passes".</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="optimizerpasses">LLVM Optimization
Passes</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>LLVM provides many optimization passes, which do many different sorts of
things and have different tradeoffs. Unlike other systems, LLVM doesn't hold
to the mistaken notion that one set of optimizations is right for all languages
and for all situations. LLVM allows a compiler implementor to make complete
decisions about what optimizations to use, in which order, and in what
situation.</p>
<p>As a concrete example, LLVM supports both "whole module" passes, which look
across as large of body of code as they can (often a whole file, but if run
at link time, this can be a substantial portion of the whole program). It also
supports and includes "per-function" passes which just operate on a single
function at a time, without looking at other functions. For more information
on passes and how they are run, see the
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html">How to Write a
Pass</a> document and the <a href="http://www.llvm.org/docs/Passes.html">List of
LLVM Passes</a>.</p>
<p>For Kaleidoscope, we are currently generating functions on the fly, one at
a time, as the user types them in. We aren't shooting for the ultimate
optimization experience in this setting, but we also want to catch the easy and
quick stuff where possible. As such, we will choose to run a few per-function
optimizations as the user types the function in. If we wanted to make a "static
Kaleidoscope compiler", we would use exactly the code we have now, except that
we would defer running the optimizer until the entire file has been parsed.</p>
<p>In order to get per-function optimizations going, we need to set up a
<a href="http://www.llvm.org/docs/WritingAnLLVMPass.html#passmanager">
FunctionPassManager</a> to hold and organize the LLVM optimizations that we want
to run. Once we have that, we can add a set of optimizations to run. The code
looks like this:</p>
<div class="doc_code">
<pre>
# The function optimization passes manager.
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
...
def main():
# Set up the optimizer pipeline. Start with registering info about how the
# target lays out data structures.
g_llvm_pass_manager.add(g_llvm_executor.target_data)
# Do simple "peephole" optimizations and bit-twiddling optzns.
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
# Reassociate expressions.
g_llvm_pass_manager.add(PASS_REASSOCIATE)
# Eliminate Common SubExpressions.
g_llvm_pass_manager.add(PASS_GVN)
# Simplify the control flow graph (deleting unreachable blocks, etc).
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
g_llvm_pass_manager.initialize()
</pre>
</div>
<p>This code defines a <tt>FunctionPassManager</tt>,
<tt>g_llvm_pass_manager</tt>. Once it is set up, we use a series of "add" calls
to add a bunch of LLVM passes. The first pass is basically boilerplate, it adds
a pass so that later optimizations know how the data structures in the program
are laid out. (The "<tt>g_llvm_executor</tt>" variable is related to the JIT,
which we will get to in the next section.) In this case, we choose to add 4
optimization passes. The passes we chose here are a pretty standard set of
"cleanup" optimizations that are useful for a wide variety of code. I won't
delve into what they do but, believe me, they are a good starting place :).</p>
<p>Once the pass manager is set up, we need to make use of it. We do this by
running it after our newly created function is constructed (in
<tt>FunctionNode.CodeGen</tt>), but before it is returned to the client:</p>
<div class="doc_code">
<pre>
return_value = self.body.CodeGen()
g_llvm_builder.ret(return_value)
# Validate the generated code, checking for consistency.
function.verify()
<b># Optimize the function.
g_llvm_pass_manager.run(function)</b>
</pre>
</div>
<p>As you can see, this is pretty straightforward. The
<tt>FunctionPassManager</tt> optimizes and updates the LLVM Function in place,
improving (hopefully) its body. With this in place, we can try our test above
again:</p>
<div class="doc_code">
<pre>
ready&gt; <b>def test(x) (1+2+x)*(x+(1+2))</b>
Read a function definition:
define double @test(double %x) {
entry:
%addtmp = fadd double %x, 3.000000e+00 ; &lt;double&gt; [#uses=2]
%multmp = fmul double %addtmp, %addtmp ; &lt;double&gt; [#uses=1]
ret double %multmp
}
</pre>
</div>
<p>As expected, we now get our nicely optimized code, saving a floating point
add instruction from every execution of this function.</p>
<p>LLVM provides a wide variety of optimizations that can be used in certain
circumstances. Some
<a href="http://www.llvm.org/docs/Passes.html">documentation about the various
passes</a> is available, but it isn't very complete. Another good source of
ideas can come from looking at the passes that <tt>llvm-gcc</tt> or
<tt>llvm-ld</tt> run to get started. The "<tt>opt</tt>" tool allows you to
experiment with passes from the command line, so you can see if they do
anything.</p>
<p>Now that we have reasonable code coming out of our front-end, lets talk about
executing it!</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="jit">Adding a JIT Compiler</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Code that is available in LLVM IR can have a wide variety of tools
applied to it. For example, you can run optimizations on it (as we did above),
you can dump it out in textual or binary forms, you can compile the code to an
assembly file (.s) for some target, or you can JIT compile it. The nice thing
about the LLVM IR representation is that it is the "common currency" between
many different parts of the compiler.
</p>
<p>In this section, we'll add JIT compiler support to our interpreter. The
basic idea that we want for Kaleidoscope is to have the user enter function
bodies as they do now, but immediately evaluate the top-level expressions they
type in. For example, if they type in "1 + 2", we should evaluate and print
out 3. If they define a function, they should be able to call it from the
command line.</p>
<p>In order to do this, we first declare and initialize the JIT. This is done
by adding and initializing a global variable:</p>
<div class="doc_code">
<pre>
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
</pre>
</div>
<p>This creates an abstract "Execution Engine" which can be either a JIT
compiler or the LLVM interpreter. LLVM will automatically pick a JIT compiler
for you if one is available for your platform, otherwise it will fall back to
the interpreter.</p>
<p>Once the <tt>ExecutionEngine</tt> is created, the JIT is ready to be used.
We can use the <tt>run_function</tt> method of the execution engine to execute
a compiled function and get its return value. In our case, this means that we
can change the code that parses a top-level expression to look like this:</p>
<div class="doc_code">
<pre>
def HandleTopLevelExpression(self):
try:
function = self.ParseTopLevelExpr().CodeGen()
result = g_llvm_executor.run_function(function, [])
print 'Evaluated to:', result.as_real(Type.double())
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
</pre>
</div>
<p>Recall that we compile top-level expressions into a self-contained LLVM
function that takes no arguments and returns the computed double.</p>
<p>With just these two changes, lets see how Kaleidoscope works now!</p>
<div class="doc_code">
<pre>
ready&gt; <b>4+5</b>
Read a top level expression:
define double @0() {
entry:
ret double 9.000000e+00
}
Evaluated to: 9.0
</pre>
</div>
<p>Well this looks like it is basically working. The dump of the function
shows the "no argument function that always returns double" that we synthesize
for each top-level expression that is typed in. This demonstrates very basic
functionality, but can we do more?</p>
<div class="doc_code">
<pre>
ready&gt; <b>def testfunc(x y) x + y*2</b>
Read a function definition:
define double @testfunc(double %x, double %y) {
entry:
%multmp = fmul double %y, 2.000000e+00 ; &lt;double&gt; [#uses=1]
%addtmp = fadd double %multmp, %x ; &lt;double&gt; [#uses=1]
ret double %addtmp
}
ready&gt; <b>testfunc(4, 10)</b>
Read a top level expression:
define double @0() {
entry:
%calltmp = call double @testfunc(double 4.000000e+00, double 1.000000e+01) ; &lt;double&gt; [#uses=1]
ret double %calltmp
}
<em>Evaluated to: 24.0</em>
</pre>
</div>
<p>This illustrates that we can now call user code, but there is something a bit
subtle going on here. Note that we only invoke the JIT on the anonymous
functions that <em>call testfunc</em>, but we never invoked it
on <em>testfunc</em> itself. What actually happened here is that the JIT
scanned for all non-JIT'd functions transitively called from the anonymous
function and compiled all of them before returning from <tt>run_function()</tt>.
</p>
<p>The JIT provides a number of other more advanced interfaces for things like
freeing allocated machine code, rejit'ing functions to update them, etc.
However, even with this simple code, we get some surprisingly powerful
capabilities - check this out (I removed the dump of the anonymous functions,
you should get the idea by now :) :</p>
<div class="doc_code">
<pre>
ready&gt; <b>extern sin(x)</b>
Read an extern:
declare double @sin(double)
ready&gt; <b>extern cos(x)</b>
Read an extern:
declare double @cos(double)
ready&gt; <b>sin(1.0)</b>
<em>Evaluated to: 0.841470984808</em>
ready&gt; <b>def foo(x) sin(x)*sin(x) + cos(x)*cos(x)</b>
Read a function definition:
define double @foo(double %x) {
entry:
%calltmp = call double @sin(double %x) ; &lt;double&gt; [#uses=1]
%calltmp1 = call double @sin(double %x) ; &lt;double&gt; [#uses=1]
%multmp = fmul double %calltmp, %calltmp1 ; &lt;double&gt; [#uses=1]
%calltmp2 = call double @cos(double %x) ; &lt;double&gt; [#uses=1]
%calltmp3 = call double @cos(double %x) ; &lt;double&gt; [#uses=1]
%multmp4 = fmul double %calltmp2, %calltmp3 ; &lt;double&gt; [#uses=1]
%addtmp = fadd double %multmp, %multmp4 ; &lt;double&gt; [#uses=1]
ret double %addtmp
}
ready&gt; <b>foo(4.0)</b>
<em>Evaluated to: 1.000000</em>
</pre>
</div>
<p>Whoa, how does the JIT know about sin and cos? The answer is surprisingly
simple: in this example, the JIT started execution of a function and got to a
function call. It realized that the function was not yet JIT compiled and
invoked the standard set of routines to resolve the function. In this case,
there is no body defined for the function, so the JIT ended up calling
"<tt>dlsym("sin")</tt>" on the Python process that is hosting our Kaleidoscope
prompt. Since "<tt>sin</tt>" is defined within the JIT's address space, it
simply patches up calls in the module to call the libm version of <tt>sin</tt>
directly.</p>
<p>One interesting application of this is that we can now extend the language
by writing arbitrary C++ code to implement operations. For example, we can
create a C file with the following simple function:
</p>
<div class="doc_code">
<pre>
#include &lt;stdio.h&gt;
double putchard(double x) {
putchar((char)x);
return 0;
}
</pre>
</div>
<p>We can then compile this into a shared library with GCC:</p>
<div class="doc_code">
<pre>
gcc -shared -fPIC -o putchard.so putchard.c
</pre>
</div>
<p>Now we can load this library into the Python process using
<tt>llvm.core.load_library_permanently</tt> and access it from Kaleidoscope to
produce simple output to the console:</p>
<div class="doc_code">
<pre>
>>> <b>import llvm.core</b>
>>> <b>llvm.core.load_library_permanently('/home/max/llvm-py-tutorial/putchard.so')</b>
>>> <b>import kaleidoscope</b>
>>> <b>kaleidoscope.main()</b>
ready&gt; <b>extern putchard(x)</b>
Read an extern:
declare double @putchard(double)
ready&gt; <b>putchard(65) + putchard(66) + putchard(67) + putchard(10)</b>
<em>ABC</em>
Evaluated to: 0.0
</pre>
</div>
<p>Similar code could be used to implement file I/O, console input, and many
other capabilities in Kaleidoscope.</p>
<p>This completes the JIT and optimizer chapter of the Kaleidoscope tutorial. At
this point, we can compile a non-Turing-complete programming language, optimize
and JIT compile it in a user-driven way. Next up we'll look into <a
href="PythonLangImpl5.html">extending the language with control flow
constructs</a>, tackling some interesting LLVM IR issues along the way.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="code">Full Code Listing</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>
Here is the complete code listing for our running example, enhanced with the
LLVM JIT and optimizer:
</p>
<div class="doc_code">
<pre>
#!/usr/bin/env python
import re
from llvm.core import Module, Constant, Type, Function, Builder, FCMP_ULT
from llvm.ee import ExecutionEngine, TargetData
from llvm.passes import FunctionPassManager
from llvm.passes import (PASS_INSTRUCTION_COMBINING,
PASS_REASSOCIATE,
PASS_GVN,
PASS_CFG_SIMPLIFICATION)
################################################################################
## Globals
################################################################################
# The LLVM module, which holds all the IR code.
g_llvm_module = Module.new('my cool jit')
# The LLVM instruction builder. Created whenever a new function is entered.
g_llvm_builder = None
# A dictionary that keeps track of which values are defined in the current scope
# and what their LLVM representation is.
g_named_values = {}
# The function optimization passes manager.
g_llvm_pass_manager = FunctionPassManager.new(g_llvm_module)
# The LLVM execution engine.
g_llvm_executor = ExecutionEngine.new(g_llvm_module)
################################################################################
## Lexer
################################################################################
# The lexer yields one of these types for each token.
class EOFToken(object):
pass
class DefToken(object):
pass
class ExternToken(object):
pass
class IdentifierToken(object):
def __init__(self, name): self.name = name
class NumberToken(object):
def __init__(self, value): self.value = value
class CharacterToken(object):
def __init__(self, char): self.char = char
def __eq__(self, other):
return isinstance(other, CharacterToken) and self.char == other.char
def __ne__(self, other): return not self == other
# Regular expressions that tokens and comments of our language.
REGEX_NUMBER = re.compile('[0-9]+(?:\.[0-9]+)?')
REGEX_IDENTIFIER = re.compile('[a-zA-Z][a-zA-Z0-9]*')
REGEX_COMMENT = re.compile('#.*')
def Tokenize(string):
while string:
# Skip whitespace.
if string[0].isspace():
string = string[1:]
continue
# Run regexes.
comment_match = REGEX_COMMENT.match(string)
number_match = REGEX_NUMBER.match(string)
identifier_match = REGEX_IDENTIFIER.match(string)
# Check if any of the regexes matched and yield the appropriate result.
if comment_match:
comment = comment_match.group(0)
string = string[len(comment):]
elif number_match:
number = number_match.group(0)
yield NumberToken(float(number))
string = string[len(number):]
elif identifier_match:
identifier = identifier_match.group(0)
# Check if we matched a keyword.
if identifier == 'def':
yield DefToken()
elif identifier == 'extern':
yield ExternToken()
else:
yield IdentifierToken(identifier)
string = string[len(identifier):]
else:
# Yield the ASCII value of the unknown character.
yield CharacterToken(string[0])
string = string[1:]
yield EOFToken()
################################################################################
## Abstract Syntax Tree (aka Parse Tree)
################################################################################
# Base class for all expression nodes.
class ExpressionNode(object):
pass
# Expression class for numeric literals like "1.0".
class NumberExpressionNode(ExpressionNode):
def __init__(self, value):
self.value = value
def CodeGen(self):
return Constant.real(Type.double(), self.value)
# Expression class for referencing a variable, like "a".
class VariableExpressionNode(ExpressionNode):
def __init__(self, name):
self.name = name
def CodeGen(self):
if self.name in g_named_values:
return g_named_values[self.name]
else:
raise RuntimeError('Unknown variable name: ' + self.name)
# Expression class for a binary operator.
class BinaryOperatorExpressionNode(ExpressionNode):
def __init__(self, operator, left, right):
self.operator = operator
self.left = left
self.right = right
def CodeGen(self):
left = self.left.CodeGen()
right = self.right.CodeGen()
if self.operator == '+':
return g_llvm_builder.fadd(left, right, 'addtmp')
elif self.operator == '-':
return g_llvm_builder.fsub(left, right, 'subtmp')
elif self.operator == '*':
return g_llvm_builder.fmul(left, right, 'multmp')
elif self.operator == '&lt;':
result = g_llvm_builder.fcmp(FCMP_ULT, left, right, 'cmptmp')
# Convert bool 0 or 1 to double 0.0 or 1.0.
return g_llvm_builder.uitofp(result, Type.double(), 'booltmp')
else:
raise RuntimeError('Unknown binary operator.')
# Expression class for function calls.
class CallExpressionNode(ExpressionNode):
def __init__(self, callee, args):
self.callee = callee
self.args = args
def CodeGen(self):
# Look up the name in the global module table.
callee = g_llvm_module.get_function_named(self.callee)
# Check for argument mismatch error.
if len(callee.args) != len(self.args):
raise RuntimeError('Incorrect number of arguments passed.')
arg_values = [i.CodeGen() for i in self.args]
return g_llvm_builder.call(callee, arg_values, 'calltmp')
# This class represents the "prototype" for a function, which captures its name,
# and its argument names (thus implicitly the number of arguments the function
# takes).
class PrototypeNode(object):
def __init__(self, name, args):
self.name = name
self.args = args
def CodeGen(self):
# Make the function type, eg. double(double,double).
funct_type = Type.function(
Type.double(), [Type.double()] * len(self.args), False)
function = Function.new(g_llvm_module, funct_type, self.name)
# If the name conflicted, there was already something with the same name.
# If it has a body, don't allow redefinition or reextern.
if function.name != self.name:
function.delete()
function = g_llvm_module.get_function_named(self.name)
# If the function already has a body, reject this.
if not function.is_declaration:
raise RuntimeError('Redefinition of function.')
# If F took a different number of args, reject.
if len(callee.args) != len(self.args):
raise RuntimeError('Redeclaration of a function with different number '
'of args.')
# Set names for all arguments and add them to the variables symbol table.
for arg, arg_name in zip(function.args, self.args):
arg.name = arg_name
# Add arguments to variable symbol table.
g_named_values[arg_name] = arg
return function
# This class represents a function definition itself.
class FunctionNode(object):
def __init__(self, prototype, body):
self.prototype = prototype
self.body = body
def CodeGen(self):
# Clear scope.
g_named_values.clear()
# Create a function object.
function = self.prototype.CodeGen()
# Create a new basic block to start insertion into.
block = function.append_basic_block('entry')
global g_llvm_builder
g_llvm_builder = Builder.new(block)
# Finish off the function.
try:
return_value = self.body.CodeGen()
g_llvm_builder.ret(return_value)
# Validate the generated code, checking for consistency.
function.verify()
# Optimize the function.
g_llvm_pass_manager.run(function)
except:
function.delete()
raise
return function
################################################################################
## Parser
################################################################################
class Parser(object):
def __init__(self, tokens, binop_precedence):
self.tokens = tokens
self.binop_precedence = binop_precedence
self.Next()
# Provide a simple token buffer. Parser.current is the current token the
# parser is looking at. Parser.Next() reads another token from the lexer and
# updates Parser.current with its results.
def Next(self):
self.current = self.tokens.next()
# Gets the precedence of the current token, or -1 if the token is not a binary
# operator.
def GetCurrentTokenPrecedence(self):
if isinstance(self.current, CharacterToken):
return self.binop_precedence.get(self.current.char, -1)
else:
return -1
# identifierexpr ::= identifier | identifier '(' expression* ')'
def ParseIdentifierExpr(self):
identifier_name = self.current.name
self.Next() # eat identifier.
if self.current != CharacterToken('('): # Simple variable reference.
return VariableExpressionNode(identifier_name)
# Call.
self.Next() # eat '('.
args = []
if self.current != CharacterToken(')'):
while True:
args.append(self.ParseExpression())
if self.current == CharacterToken(')'):
break
elif self.current != CharacterToken(','):
raise RuntimeError('Expected ")" or "," in argument list.')
self.Next()
self.Next() # eat ')'.
return CallExpressionNode(identifier_name, args)
# numberexpr ::= number
def ParseNumberExpr(self):
result = NumberExpressionNode(self.current.value)
self.Next() # consume the number.
return result
# parenexpr ::= '(' expression ')'
def ParseParenExpr(self):
self.Next() # eat '('.
contents = self.ParseExpression()
if self.current != CharacterToken(')'):
raise RuntimeError('Expected ")".')
self.Next() # eat ')'.
return contents
# primary ::= identifierexpr | numberexpr | parenexpr
def ParsePrimary(self):
if isinstance(self.current, IdentifierToken):
return self.ParseIdentifierExpr()
elif isinstance(self.current, NumberToken):
return self.ParseNumberExpr()
elif self.current == CharacterToken('('):
return self.ParseParenExpr()
else:
raise RuntimeError('Unknown token when expecting an expression.')
# binoprhs ::= (operator primary)*
def ParseBinOpRHS(self, left, left_precedence):
# If this is a binary operator, find its precedence.
while True:
precedence = self.GetCurrentTokenPrecedence()
# If this is a binary operator that binds at least as tightly as the
# current one, consume it; otherwise we are done.
if precedence &lt; left_precedence:
return left
binary_operator = self.current.char
self.Next() # eat the operator.
# Parse the primary expression after the binary operator.
right = self.ParsePrimary()
# If binary_operator binds less tightly with right than the operator after
# right, let the pending operator take right as its left.
next_precedence = self.GetCurrentTokenPrecedence()
if precedence &lt; next_precedence:
right = self.ParseBinOpRHS(right, precedence + 1)
# Merge left/right.
left = BinaryOperatorExpressionNode(binary_operator, left, right)
# expression ::= primary binoprhs
def ParseExpression(self):
left = self.ParsePrimary()
return self.ParseBinOpRHS(left, 0)
# prototype ::= id '(' id* ')'
def ParsePrototype(self):
if not isinstance(self.current, IdentifierToken):
raise RuntimeError('Expected function name in prototype.')
function_name = self.current.name
self.Next() # eat function name.
if self.current != CharacterToken('('):
raise RuntimeError('Expected "(" in prototype.')
self.Next() # eat '('.
arg_names = []
while isinstance(self.current, IdentifierToken):
arg_names.append(self.current.name)
self.Next()
if self.current != CharacterToken(')'):
raise RuntimeError('Expected ")" in prototype.')
# Success.
self.Next() # eat ')'.
return PrototypeNode(function_name, arg_names)
# definition ::= 'def' prototype expression
def ParseDefinition(self):
self.Next() # eat def.
proto = self.ParsePrototype()
body = self.ParseExpression()
return FunctionNode(proto, body)
# toplevelexpr ::= expression
def ParseTopLevelExpr(self):
proto = PrototypeNode('', [])
return FunctionNode(proto, self.ParseExpression())
# external ::= 'extern' prototype
def ParseExtern(self):
self.Next() # eat extern.
return self.ParsePrototype()
# Top-Level parsing
def HandleDefinition(self):
self.Handle(self.ParseDefinition, 'Read a function definition:')
def HandleExtern(self):
self.Handle(self.ParseExtern, 'Read an extern:')
def HandleTopLevelExpression(self):
try:
function = self.ParseTopLevelExpr().CodeGen()
result = g_llvm_executor.run_function(function, [])
print 'Evaluated to:', result.as_real(Type.double())
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
def Handle(self, function, message):
try:
print message, function().CodeGen()
except Exception, e:
print 'Error:', e
try:
self.Next() # Skip for error recovery.
except:
pass
################################################################################
## Main driver code.
################################################################################
def main():
# Set up the optimizer pipeline. Start with registering info about how the
# target lays out data structures.
g_llvm_pass_manager.add(g_llvm_executor.target_data)
# Do simple "peephole" optimizations and bit-twiddling optzns.
g_llvm_pass_manager.add(PASS_INSTRUCTION_COMBINING)
# Reassociate expressions.
g_llvm_pass_manager.add(PASS_REASSOCIATE)
# Eliminate Common SubExpressions.
g_llvm_pass_manager.add(PASS_GVN)
# Simplify the control flow graph (deleting unreachable blocks, etc).
g_llvm_pass_manager.add(PASS_CFG_SIMPLIFICATION)
g_llvm_pass_manager.initialize()
# Install standard binary operators.
# 1 is lowest possible precedence. 40 is the highest.
operator_precedence = {
'&lt;': 10,
'+': 20,
'-': 20,
'*': 40
}
# Run the main "interpreter loop".
while True:
print 'ready&gt;',
try:
raw = raw_input()
except KeyboardInterrupt:
break
parser = Parser(Tokenize(raw), operator_precedence)
while True:
# top ::= definition | external | expression | EOF
if isinstance(parser.current, EOFToken):
break
if isinstance(parser.current, DefToken):
parser.HandleDefinition()
elif isinstance(parser.current, ExternToken):
parser.HandleExtern()
else:
parser.HandleTopLevelExpression()
# Print out all of the generated code.
print '\n', g_llvm_module
if __name__ == '__main__':
main()
</pre>
</div>
<a href="PythonLangImpl5.html">Next: Extending the language: control flow</a>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://max99x.com">Max Shawabkeh</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,375 @@
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN"
"http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Kaleidoscope: Conclusion and other useful LLVM tidbits</title>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8">
<meta name="author" content="Chris Lattner">
<link rel="stylesheet"
href="http://www.llvm.org/docs/llvm.css"
type="text/css">
</head>
<body>
<div class="doc_title">Kaleidoscope: Conclusion and other useful LLVM
tidbits</div>
<ul>
<li>
<a href="http://www.llvm.org/docs/tutorial/index.html">
Up to Tutorial Index
</a>
</li>
<li>Chapter 8
<ol>
<li><a href="#conclusion">Tutorial Conclusion</a></li>
<li><a href="#llvmirproperties">Properties of LLVM IR</a>
<ul>
<li><a href="#targetindep">Target Independence</a></li>
<li><a href="#safety">Safety Guarantees</a></li>
<li><a href="#langspecific">Language-Specific Optimizations</a></li>
</ul>
</li>
<li><a href="#tipsandtricks">Tips and Tricks</a>
<ul>
<li><a href="#offsetofsizeof">Implementing portable
offsetof/sizeof</a></li>
<li><a href="#gcstack">Garbage Collected Stack Frames</a></li>
</ul>
</li>
</ol>
</li>
</ul>
<div class="doc_author">
<p>Written by <a href="mailto:sabre@nondot.org">Chris Lattner</a></p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="conclusion">Tutorial Conclusion</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>Welcome to the the final chapter of the
"<a href="http://www.llvm.org/docs/tutorial/index.html">Implementing a language
with LLVM</a>" tutorial. In the course of this tutorial, we have grown
our little Kaleidoscope language from being a useless toy, to being a
semi-interesting (but probably still useless) toy. :)</p>
<p>It is interesting to see how far we've come, and how little code it has
taken. We built the entire lexer, parser, AST, code generator, and an
interactive run-loop (with a JIT!) by-hand in under 540 lines of
(non-comment/non-blank) code.</p>
<p>Our little language supports a couple of interesting features: it supports
user defined binary and unary operators, it uses JIT compilation for immediate
evaluation, and it supports a few control flow constructs with SSA construction.
</p>
<p>Part of the idea of this tutorial was to show you how easy and fun it can be
to define, build, and play with languages. Building a compiler need not be a
scary or mystical process! Now that you've seen some of the basics, I strongly
encourage you to take the code and hack on it. For example, try adding:</p>
<ul>
<li><b>global variables</b> - While global variables have questional value in
modern software engineering, they are often useful when putting together quick
little hacks like the Kaleidoscope compiler itself. Fortunately, our current
setup makes it very easy to add global variables: just have value lookup check
to see if an unresolved variable is in the global variable symbol table before
rejecting it. To create a new global variable, make an instance of the LLVM
<tt>GlobalVariable</tt> class.</li>
<li><b>typed variables</b> - Kaleidoscope currently only supports variables of
type double. This gives the language a very nice elegance, because only
supporting one type means that you never have to specify types. Different
languages have different ways of handling this. The easiest way is to require
the user to specify types for every variable definition, and record the type
of the variable in the symbol table along with its Value*.</li>
<li><b>arrays, structs, vectors, etc</b> - Once you add types, you can start
extending the type system in all sorts of interesting ways. Simple arrays are
very easy and are quite useful for many different applications. Adding them is
mostly an exercise in learning how the LLVM <a
href="http://www.llvm.org/docs/LangRef.html#i_getelementptr">getelementptr</a>
instruction works: it is so nifty/unconventional, it <a
href="http://www.llvm.org/docs/GetElementPtr.html">has its own FAQ</a>! If you
add support for recursive types (e.g. linked lists), make sure to read the <a
href="http://www.llvm.org/docs/ProgrammersManual.html#TypeResolve">section in
the LLVM Programmer's Manual</a> that describes how to construct them.</li>
<li><b>standard runtime</b> - Our current language allows the user to access
arbitrary external functions, and we use it for things like "putchard". As you
extend the language to add higher-level constructs, often these constructs make
the most sense if they are lowered to calls into a language-supplied runtime.
For example, if you add hash tables to the language, it would probably make
sense to add the routines to a runtime, instead of inlining them all the way.
</li>
<li><b>memory management</b> - Currently we can only access the stack in
Kaleidoscope. It would also be useful to be able to allocate heap memory,
either with calls to the standard libc malloc/free interface or with a garbage
collector. If you would like to use garbage collection, note that LLVM fully
supports <a href="http://www.llvm.org/docs/GarbageCollection.html">Accurate
Garbage Collection</a> including algorithms that move objects and need to
scan/update the stack.</li>
<li><b>debugger support</b> - LLVM supports generation of <a
href="http://www.llvm.org/docs/SourceLevelDebugging.html">DWARF Debug info</a>
which is understood by common debuggers like GDB. Adding support for debug info
is fairly straightforward. The best way to understand it is to compile some
C/C++ code with "<tt>llvm-gcc -g -O0</tt>" and taking a look at what it
produces.</li>
<li><b>exception handling support</b> - LLVM supports generation of <a
href="http://www.llvm.org/docs/ExceptionHandling.html">zero cost exceptions</a>
which interoperate with code compiled in other languages. You could also
generate code by implicitly making every function return an error value and
checking it. You could also make explicit use of setjmp/longjmp. There are
many different ways to go here.</li>
<li><b>object orientation, generics, database access, complex numbers,
geometric programming, ...</b> - Really, there is
no end of crazy features that you can add to the language.</li>
<li><b>unusual domains</b> - We've been talking about applying LLVM to a domain
that many people are interested in: building a compiler for a specific language.
However, there are many other domains that can use compiler technology that are
not typically considered. For example, LLVM has been used to implement OpenGL
graphics acceleration, translate C++ code to ActionScript, and many other
cute and clever things. Maybe you will be the first to JIT compile a regular
expression interpreter into native code with LLVM?</li>
</ul>
<p>
Have fun - try doing something crazy and unusual. Building a language like
everyone else always has, is much less fun than trying something a little crazy
or off the wall and seeing how it turns out. If you get stuck or want to talk
about it, feel free to email the <a
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
list</a>: it has lots of people who are interested in languages and are often
willing to help out.
</p>
<p>Before we end this tutorial, I want to talk about some "tips and tricks" for
generating LLVM IR. These are some of the more subtle things that may not be
obvious, but are very useful if you want to take advantage of LLVM's
capabilities.</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="llvmirproperties">Properties of the LLVM
IR</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>We have a couple common questions about code in the LLVM IR form - let's just
get these out of the way right now, shall we?</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="targetindep">Target
Independence</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Kaleidoscope is an example of a "portable language": any program written in
Kaleidoscope will work the same way on any target that it runs on. Many other
languages have this property, e.g. LISP, Java, Haskell, Javascript, Python, etc.
(note that while these languages are portable, not all their libraries are).</p>
<p>One nice aspect of LLVM is that it is often capable of preserving target
independence in the IR: you can take the LLVM IR for a Kaleidoscope-compiled
program and run it on any target that LLVM supports, even emitting C code and
compiling that on targets that LLVM doesn't support natively. You can trivially
tell that the Kaleidoscope compiler generates target-independent code because it
never queries for any target-specific information when generating code.</p>
<p>The fact that LLVM provides a compact, target-independent, representation for
code gets a lot of people excited. Unfortunately, these people are usually
thinking about C or a language from the C family when they are asking questions
about language portability. I say "unfortunately", because there is really no
way to make (fully general) C code portable, other than shipping the source code
around (and of course, C source code is not actually portable in general
either - ever port a really old application from 32- to 64-bits?).</p>
<p>The problem with C (again, in its full generality) is that it is heavily
laden with target specific assumptions. As one simple example, the preprocessor
often destructively removes target-independence from the code when it processes
the input text:</p>
<div class="doc_code">
<pre>
#ifdef __i386__
int X = 1;
#else
int X = 42;
#endif
</pre>
</div>
<p>While it is possible to engineer more and more complex solutions to problems
like this, it cannot be solved in full generality in a way that is better than
shipping the actual source code.</p>
<p>That said, there are interesting subsets of C that can be made portable. If
you are willing to fix primitive types to a fixed size (say int = 32-bits,
and long = 64-bits), don't care about ABI compatibility with existing binaries,
and are willing to give up some other minor features, you can have portable
code. This can make sense for specialized domains such as an
in-kernel language.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="safety">Safety Guarantees</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Many of the languages above are also "safe" languages: it is impossible for
a program written in Java to corrupt its address space and crash the process
(assuming the JVM has no bugs).
Safety is an interesting property that requires a combination of language
design, runtime support, and often operating system support.</p>
<p>It is certainly possible to implement a safe language in LLVM, but LLVM IR
does not itself guarantee safety. The LLVM IR allows unsafe pointer casts,
use after free bugs, buffer over-runs, and a variety of other problems. Safety
needs to be implemented as a layer on top of LLVM and, conveniently, several
groups have investigated this. Ask on the <a
href="http://lists.cs.uiuc.edu/mailman/listinfo/llvmdev">llvmdev mailing
list</a> if you are interested in more details.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="langspecific">Language-Specific
Optimizations</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>One thing about LLVM that turns off many people is that it does not solve all
the world's problems in one system (sorry 'world hunger', someone else will have
to solve you some other day). One specific complaint is that people perceive
LLVM as being incapable of performing high-level language-specific optimization:
LLVM "loses too much information".</p>
<p>Unfortunately, this is really not the place to give you a full and unified
version of "Chris Lattner's theory of compiler design". Instead, I'll make a
few observations:</p>
<p>First, you're right that LLVM does lose information. For example, as of this
writing, there is no way to distinguish in the LLVM IR whether an SSA-value came
from a C "int" or a C "long" on an ILP32 machine (other than debug info). Both
get compiled down to an 'i32' value and the information about what it came from
is lost. The more general issue here, is that the LLVM type system uses
"structural equivalence" instead of "name equivalence". Another place this
surprises people is if you have two types in a high-level language that have the
same structure (e.g. two different structs that have a single int field): these
types will compile down into a single LLVM type and it will be impossible to
tell what it came from.</p>
<p>Second, while LLVM does lose information, LLVM is not a fixed target: we
continue to enhance and improve it in many different ways. In addition to
adding new features (LLVM did not always support exceptions or debug info), we
also extend the IR to capture important information for optimization (e.g.
whether an argument is sign or zero extended, information about pointers
aliasing, etc). Many of the enhancements are user-driven: people want LLVM to
include some specific feature, so they go ahead and extend it.</p>
<p>Third, it is <em>possible and easy</em> to add language-specific
optimizations, and you have a number of choices in how to do it. As one trivial
example, it is easy to add language-specific optimization passes that
"know" things about code compiled for a language. In the case of the C family,
there is an optimization pass that "knows" about the standard C library
functions. If you call "exit(0)" in main(), it knows that it is safe to
optimize that into "return 0;" because C specifies what the 'exit'
function does.</p>
<p>In addition to simple library knowledge, it is possible to embed a variety of
other language-specific information into the LLVM IR. If you have a specific
need and run into a wall, please bring the topic up on the llvmdev list. At the
very worst, you can always treat LLVM as if it were a "dumb code generator" and
implement the high-level optimizations you desire in your front-end, on the
language-specific AST.
</p>
</div>
<!-- *********************************************************************** -->
<div class="doc_section"><a name="tipsandtricks">Tips and Tricks</a></div>
<!-- *********************************************************************** -->
<div class="doc_text">
<p>There is a variety of useful tips and tricks that you come to know after
working on/with LLVM that aren't obvious at first glance. Instead of letting
everyone rediscover them, this section talks about some of these issues.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="offsetofsizeof">Implementing portable
offsetof/sizeof</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>One interesting thing that comes up, if you are trying to keep the code
generated by your compiler "target independent", is that you often need to know
the size of some LLVM type or the offset of some field in an llvm structure.
For example, you might need to pass the size of a type into a function that
allocates memory.</p>
<p>Unfortunately, this can vary widely across targets: for example the width of
a pointer is trivially target-specific. However, there is a <a
href="http://nondot.org/sabre/LLVMNotes/SizeOf-OffsetOf-VariableSizedStructs.txt">clever
way to use the getelementptr instruction</a> that allows you to compute this
in a portable way.</p>
</div>
<!-- ======================================================================= -->
<div class="doc_subsubsection"><a name="gcstack">Garbage Collected
Stack Frames</a></div>
<!-- ======================================================================= -->
<div class="doc_text">
<p>Some languages want to explicitly manage their stack frames, often so that
they are garbage collected or to allow easy implementation of closures. There
are often better ways to implement these features than explicit stack frames,
but <a
href="http://nondot.org/sabre/LLVMNotes/ExplicitlyManagedStackFrames.txt">LLVM
does support them,</a> if you want. It requires your front-end to convert the
code into <a
href="http://en.wikipedia.org/wiki/Continuation-passing_style">Continuation
Passing Style</a> and the use of tail calls (which LLVM also supports).</p>
</div>
<!-- *********************************************************************** -->
<hr>
<address>
<a href="http://jigsaw.w3.org/css-validator/check/referer"><img
src="http://jigsaw.w3.org/css-validator/images/vcss" alt="Valid CSS!"></a>
<a href="http://validator.w3.org/check/referer"><img
src="http://www.w3.org/Icons/valid-html401" alt="Valid HTML 4.01!"></a>
<a href="mailto:sabre@nondot.org">Chris Lattner</a><br>
<a href="http://llvm.org">The LLVM Compiler Infrastructure</a><br>
Last modified: $Date$
</address>
</body>
</html>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -74,7 +74,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.</tt></pre>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>

View file

@ -2,7 +2,7 @@
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8" />
<meta name="generator" content="AsciiDoc 8.5.3" />
<meta name="generator" content="AsciiDoc 8.6.1" />
<meta name="description" content="Python bindings for LLVM" />
<meta name="keywords" content="llvm python compiler backend bindings" />
<link rel="stylesheet" href="style/xhtml11.css" type="text/css" />
@ -48,6 +48,7 @@ you can setup and use it. A working knowledge of Python and a basic idea
of LLVM is assumed.</p></div>
</div>
</div>
<div class="sect1">
<h2 id="_introduction">Introduction</h2>
<div class="sectionbody">
<div class="paragraph"><p><a href="http://www.llvm.org/">LLVM</a> (Low-Level Virtual Machine) provides enough
@ -76,6 +77,8 @@ versions.</p></div>
<div class="paragraph"><p>llvm-py has been built and tested with Python 2.6. It should work with
Python 2.4 and 2.5. It has not been tried with Python 3.x (patches welcome).</p></div>
</div>
</div>
<div class="sect1">
<h2 id="install">Installation</h2>
<div class="sectionbody">
<div class="paragraph"><p>llvm-py is distributed as a source tarball. You&#8217;ll need to build and
@ -109,7 +112,8 @@ distro&#8217;s respository has the appropriate version of LLVM!</p></div>
<div class="paragraph"><p>It does not matter which compiler LLVM itself was built with (g<tt>,
llvm-g</tt> or any other); llvm-py can be built with any compiler. It has
been tried only with gcc/g++ though.</p></div>
<h3 id="_llvm_and_tt_enable_pic_tt">LLVM and <tt>--enable-pic</tt></h3><div style="clear:left"></div>
<div class="sect2">
<h3 id="_llvm_and_tt_enable_pic_tt">LLVM and <tt>--enable-pic</tt></h3>
<div class="paragraph"><p>The result of an LLVM build is a set of static libraries and object
files. The llvm-py contains an extension package that is built into a
shared object (_core.so) which links to these static libraries and
@ -121,7 +125,9 @@ configuring LLVM (default is no PIC), like this:</p></div>
<div class="content">
<pre><tt>~/llvm$ ./configure --enable-pic --enable-optimized</tt></pre>
</div></div>
<h3 id="_llvm_config">llvm-config</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_llvm_config">llvm-config</h3>
<div class="paragraph"><p>Inorder to build llvm-py, it&#8217;s build script needs to know from where it
can invoke the llvm helper program, <tt>llvm-config</tt>. If you&#8217;ve installed
LLVM, then this will be available in your <tt>PATH</tt>, and nothing further
@ -131,7 +137,9 @@ of <tt>llvm-config</tt> to the build script.</p></div>
<div class="paragraph"><p>You&#8217;ll need to be <em>root</em> to install llvm-py. Remember that your <tt>PATH</tt>
is different from that of <em>root</em>, so even if <tt>llvm-config</tt> is in your
<tt>PATH</tt>, it may not be available when you do <tt>sudo</tt>.</p></div>
<h3 id="_steps">Steps</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_steps">Steps</h3>
<div class="paragraph"><p>The commands illustrated below assume that the LLVM source is available
under <tt>/home/mdevan/llvm</tt>. If you&#8217;ve a previous version of llvm-py
installed, it is recommended to remove it first, as described
@ -167,7 +175,9 @@ only if you need to debug into LLVM also.</p></div>
documentation regarding <a href="http://docs.python.org/inst/inst.html">Installing
Python Modules</a> and <a href="http://docs.python.org/dist/dist.html">Distributing
Python Modules</a> for more information on such scripts.</p></div>
<h3 id="uninstall">Uninstall</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="uninstall">Uninstall</h3>
<div class="paragraph"><p>If you&#8217;d installed llvm-py with the <tt>--user</tt> option, then llvm-py
would be present under <tt>~/.local/lib/python2.6/site-packages</tt>.
Otherwise, it might be under <tt>/usr/lib/python2.6/site-packages</tt>
@ -182,11 +192,15 @@ the "egg" can be removed like so:</p></div>
<div class="paragraph"><p>See the <a href="http://docs.python.org/install/index.html">Python
documentation</a> for more information.</p></div>
</div>
</div>
</div>
<div class="sect1">
<h2 id="_llvm_concepts">LLVM Concepts</h2>
<div class="sectionbody">
<div class="paragraph"><p>This section explains a few concepts related to LLVM, not specific
to llvm-py.</p></div>
<h3 id="_intermediate_representation">Intermediate Representation</h3><div style="clear:left"></div>
<div class="sect2">
<h3 id="_intermediate_representation">Intermediate Representation</h3>
<div class="paragraph"><p>The intermediate representation, or IR for short, is an in-memory data
structure that represents executable code. The IR data structures allow
for creation of types, constants, functions, function arguments,
@ -242,7 +256,9 @@ level than the usual assembly language; for example there are
instructions related to variable argument handling, exception handling,
and garbage collection. These allow high-level languages to be
represented cleanly in the IR.</p></div>
<h3 id="_ssa_form_and_phi_nodes">SSA Form and PHI Nodes</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_ssa_form_and_phi_nodes">SSA Form and PHI Nodes</h3>
<div class="paragraph"><p>All LLVM instructions are represented in the <em>Static Single Assignment</em>
(SSA) form. Essentially, this means that any variable can be assigned to
only once. Such a representation facilitates better optimization, among
@ -274,7 +290,9 @@ reached the PHI node. The argument <tt>a1</tt> of the PHI node is associated
with the block <tt>"a1 = 1;"</tt> and <tt>a2</tt> with the block <tt>"a2 = 2;"</tt>.</p></div>
<div class="paragraph"><p>PHI nodes have to be explicitly created in the LLVM IR. Accordingly the
LLVM instruction set has an instruction called <tt>phi</tt>.</p></div>
<h3 id="_llvm_assembly_language">LLVM Assembly Language</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_llvm_assembly_language">LLVM Assembly Language</h3>
<div class="paragraph"><p>The LLVM IR can be represented offline in two formats
- a textual, human-readable form, similar to assembly language, called
the LLVM assembly language (files with .ll extension)
@ -337,7 +355,9 @@ specification of the platform ABI (like endianness, sizes of types,
alignment etc.).</p></div>
<div class="paragraph"><p>The <a href="http://www.llvm.org/docs/LangRef.html">LLVM Language Reference</a>
defines the LLVM assembly language including the entire instruction set.</p></div>
<h3 id="_modules">Modules</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_modules">Modules</h3>
<div class="paragraph"><p>Modules, in the LLVM IR, are similar to a single <tt>C</tt> language source
file (.c file). A module contains:</p></div>
<div class="ulist"><ul>
@ -361,7 +381,9 @@ global type aliases (typedef-s)
contained within modules. Modules may be combined (linked) together to
give a bigger resultant module. During this process LLVM attempts to
reconcile the references between the combined modules.</p></div>
<h3 id="_optimization_and_passes">Optimization and Passes</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_optimization_and_passes">Optimization and Passes</h3>
<div class="paragraph"><p>LLVM provides quite a few optimization algorithms that work on the IR.
These algorithms are organized as <em>passes</em>. Each pass does something
specific, like combining redundant instructions. Passes need not always
@ -384,11 +406,18 @@ any stage, and perform any transforms on it as you like.)</p></div>
correct objects to run them on (for example, a pass may work only
on functions, individually) and actually runs them. <tt>opt</tt> is a
command-line wrapper for the pass manager.</p></div>
<h3 id="_bit_code">Bit code</h3><div style="clear:left"></div>
<div class="paragraph"><p>TODO</p></div>
<h3 id="_execution_engine_jit_and_interpreter">Execution Engine, JIT and Interpreter</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_bit_code">Bit code</h3>
<div class="paragraph"><p>TODO</p></div>
</div>
<div class="sect2">
<h3 id="_execution_engine_jit_and_interpreter">Execution Engine, JIT and Interpreter</h3>
<div class="paragraph"><p>TODO</p></div>
</div>
</div>
</div>
<div class="sect1">
<h2 id="_the_llvm_py_package">The llvm-py Package</h2>
<div class="sectionbody">
<div class="paragraph"><p>The llvm-py is a Python package, consisting of 6 modules, that wrap
@ -575,7 +604,8 @@ interpreter or the <tt>object?</tt> of <a href="http://ipython.scipy.org/moin/">
to get online help. (Note: not complete yet!)</td>
</tr></table>
</div>
<h3 id="_module_llvm_core">Module (llvm.core)</h3><div style="clear:left"></div>
<div class="sect2">
<h3 id="_module_llvm_core">Module (llvm.core)</h3>
<div class="paragraph"><p>Modules are top-level container objects. You need to create a module
object first, before you can add global variables, aliases or functions.
Modules are created using the static method <tt>Module.new</tt>:</p></div>
@ -644,7 +674,7 @@ my_module <span style="color: #990000">=</span> Module<span style="color: #99000
stringifying them (see below).</p></div>
<div class="exampleblock">
<div class="title">llvm.core.Module</div>
<div class="exampleblock-content">
<div class="content">
<div class="dlist"><div class="title">Static Constructors</div><dl>
<dt class="hdlist1">
<tt>new(module_id)</tt>
@ -857,7 +887,9 @@ string representations.</p></div>
</td>
</tr></table>
</div>
<h3 id="_types_llvm_core">Types (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_types_llvm_core">Types (llvm.core)</h3>
<div class="paragraph"><p>Types are what you think they are. A instance of <tt>llvm.core.Type</tt>, or
one of its derived classes, represent a type. llvm-py does not use as
many classes to represent types as does LLVM itself. Some types are
@ -977,7 +1009,7 @@ cellspacing="0" cellpadding="4">
<div class="paragraph"><p>The class-level documentation follows:</p></div>
<div class="exampleblock">
<div class="title">llvm.core.Type</div>
<div class="exampleblock-content">
<div class="content">
<div class="dlist"><div class="title">Static Constructors</div><dl>
<dt class="hdlist1">
<tt>int(n)</tt>
@ -1176,7 +1208,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.IntegerType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1197,7 +1229,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.FunctionType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1254,7 +1286,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.StructType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1303,7 +1335,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.ArrayType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1332,7 +1364,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.PointerType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1361,7 +1393,7 @@ http://www.gnu.org/software/src-highlite -->
</div></div>
<div class="exampleblock">
<div class="title">llvm.core.VectorType</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1430,7 +1462,9 @@ f3 <span style="color: #990000">=</span> Type<span style="color: #990000">.</spa
fnargs <span style="color: #990000">=</span> <span style="color: #990000">[</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">pointer</span></span><span style="color: #990000">(</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">int</span></span><span style="color: #990000">(</span><span style="color: #993399">8</span><span style="color: #990000">)</span> <span style="color: #990000">)</span> <span style="color: #990000">]</span>
printf <span style="color: #990000">=</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">function</span></span><span style="color: #990000">(</span> Type<span style="color: #990000">.</span><span style="font-weight: bold"><span style="color: #000000">int</span></span><span style="color: #990000">(),</span> fnargs<span style="color: #990000">,</span> True <span style="color: #990000">)</span>
<span style="font-style: italic"><span style="color: #9A1900"># variadic function</span></span></tt></pre></div></div>
<h3 id="_typehandle_llvm_core">TypeHandle (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_typehandle_llvm_core">TypeHandle (llvm.core)</h3>
<div class="paragraph"><p>TypeHandle objects are used to create recursive types, like this linked
list node structure in C:</p></div>
<div class="listingblock">
@ -1483,7 +1517,7 @@ in C++. The above example is available as
in the source distribution.</p></div>
<div class="exampleblock">
<div class="title">llvm.core.TypeHandle</div>
<div class="exampleblock-content">
<div class="content">
<div class="dlist"><div class="title">Static Constructors</div><dl>
<dt class="hdlist1">
<tt>new(abstract_ty)</tt>
@ -1508,7 +1542,9 @@ in the source distribution.</p></div>
</dd>
</dl></div>
</div></div>
<h3 id="_values_llvm_core">Values (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_values_llvm_core">Values (llvm.core)</h3>
<div class="paragraph"><p><tt>llvm.core.Value</tt> is the base class of all values computed by a program
that may be used as operands to other values. A value has a type
associated with it (an object of <tt>llvm.core.Type</tt>).</p></div>
@ -1557,7 +1593,7 @@ a few subclasses that represent interesting instructions.</p></div>
<div class="paragraph"><p><tt>Value</tt> objects have a type (read-only), and a name (read-write).</p></div>
<div class="exampleblock">
<div class="title">llvm.core.Value</div>
<div class="exampleblock-content">
<div class="content">
<div class="dlist"><div class="title">Properties</div><dl>
<dt class="hdlist1">
<tt>name</tt>
@ -1624,14 +1660,16 @@ a few subclasses that represent interesting instructions.</p></div>
</dd>
</dl></div>
</div></div>
<h3 id="_user_llvm_core">User (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_user_llvm_core">User (llvm.core)</h3>
<div class="paragraph"><p><tt>User</tt>-s are values that refer to other values. The values so refered
can be retrived by the properties of <tt>User</tt>. This is the reverse of
the <tt>Value.uses</tt>. Together these can be used to traverse the use-def
chains of the SSA.</p></div>
<div class="exampleblock">
<div class="title">llvm.core.User</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -1660,7 +1698,9 @@ chains of the SSA.</p></div>
</dd>
</dl></div>
</div></div>
<h3 id="_constants_llvm_core">Constants (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_constants_llvm_core">Constants (llvm.core)</h3>
<div class="paragraph"><p><tt>Constant</tt>-s represents constants that appear within the code. The
values of such objects are known at creation time. Constants can be
created from Python constants. A constant expression is also a constant&#8201;&#8212;&#8201;given a <tt>Constant</tt> object, an operation (like addition, subtraction
@ -2074,7 +2114,7 @@ cellspacing="0" cellpadding="4">
</div>
<div class="exampleblock">
<div class="title">llvm.core.Constant</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -2086,7 +2126,9 @@ cellspacing="0" cellpadding="4">
<div class="paragraph"><div class="title">Methods</div><p>See table of operations <a href="#constops">above</a> for full list. There are no other
methods.</p></div>
</div></div>
<h3 id="_other_constant_classes_llvm_core">Other Constant* Classes (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_other_constant_classes_llvm_core">Other Constant* Classes (llvm.core)</h3>
<div class="paragraph"><p>The following subclasses of <tt>Constant</tt> do not provide additional
methods, they serve only to provide richer type information.</p></div>
<div class="tableblock">
@ -2165,7 +2207,9 @@ k2 <span style="color: #990000">=</span> Constant<span style="color: #990000">.<
<span style="font-weight: bold"><span style="color: #0000FF">assert</span></span> <span style="font-weight: bold"><span style="color: #000000">isinstance</span></span><span style="color: #990000">(</span>k1<span style="color: #990000">,</span> ConstantInt<span style="color: #990000">)</span>
<span style="font-weight: bold"><span style="color: #0000FF">assert</span></span> <span style="font-weight: bold"><span style="color: #000000">isinstance</span></span><span style="color: #990000">(</span>k2<span style="color: #990000">,</span> ConstantArray<span style="color: #990000">)</span></tt></pre></div></div>
<h3 id="_global_value_llvm_core">Global Value (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_global_value_llvm_core">Global Value (llvm.core)</h3>
<div class="paragraph"><p>The class <tt>llvm.core.GlobalValue</tt> represents module-scope aliases, variables
and functions. Global variables are represented by the sub-class
<tt>llvm.core.GlobalVariable</tt> and functions by <tt>llvm.core.Function</tt>.</p></div>
@ -2290,7 +2334,7 @@ global is a declaration or not. The module to which the global belongs
to can be retrieved using the <tt>module</tt> property (read-only).</p></div>
<div class="exampleblock">
<div class="title">llvm.core.GlobalValue</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -2352,7 +2396,9 @@ to can be retrieved using the <tt>module</tt> property (read-only).</p></div>
</dd>
</dl></div>
</div></div>
<h3 id="_global_variable_llvm_core">Global Variable (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_global_variable_llvm_core">Global Variable (llvm.core)</h3>
<div class="paragraph"><p>Global variables (<tt>llvm.core.GlobalVariable</tt>) are subclasses of
<tt>llvm.core.GlobalValue</tt> and represent module-level variables. These can
have optional initializers and can be marked as constants. Global
@ -2407,7 +2453,7 @@ gv<span style="color: #990000">.</span><span style="font-weight: bold"><span sty
gv <span style="color: #990000">=</span> None</tt></pre></div></div>
<div class="exampleblock">
<div class="title">llvm.core.GlobalVariable</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -2467,7 +2513,9 @@ gv <span style="color: #990000">=</span> None</tt></pre></div></div>
</dd>
</dl></div>
</div></div>
<h3 id="_function_llvm_core">Function (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_function_llvm_core">Function (llvm.core)</h3>
<div class="paragraph"><p>Functions are represented by <tt>llvm.core.Function</tt> objects. They are
contained within modules, and can be created either with the method
<tt>module_obj.add_function</tt> or the static constructor <tt>Function.new</tt>.
@ -2939,7 +2987,7 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
<span style="font-style: italic"><span style="color: #9A1900"># declare i32 @sum(i32, i32) nounwind readonly</span></span></tt></pre></div></div>
<div class="exampleblock">
<div class="title">llvm.core.Function</div>
<div class="exampleblock-content">
<div class="content">
<div class="ulist"><div class="title">Base Class</div><ul>
<li>
<p>
@ -2993,7 +3041,7 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
<dd>
<p>
The calling convention for the function, as listed
<a href="#callconv">abov</a>.
<a href="#callconv">above</a>.
</p>
</dd>
<dt class="hdlist1">
@ -3127,7 +3175,9 @@ f<span style="color: #990000">.</span><span style="font-weight: bold"><span styl
</dd>
</dl></div>
</div></div>
<h3 id="_argument_llvm_core">Argument (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_argument_llvm_core">Argument (llvm.core)</h3>
<div class="paragraph"><p>The <tt>args</tt> property of <tt>llvm.core.Function</tt> objects yields
<tt>llvm.core.Argument</tt> objects. This allows for setting attributes for
functions arguments. <tt>Argument</tt> objects cannot be constructed from user
@ -3223,19 +3273,34 @@ cellspacing="0" cellpadding="4">
provide more information.</p></div>
<div class="paragraph"><p>The alignment of any parameter can be set via the <tt>alignment</tt>
property, to any power of 2.</p></div>
<h3 id="_basic_block_llvm_core">Basic Block (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_basic_block_llvm_core">Basic Block (llvm.core)</h3>
<div class="paragraph"><p>TODO</p></div>
<h3 id="_builder_llvm_core">Builder (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_builder_llvm_core">Builder (llvm.core)</h3>
<div class="paragraph"><p>TODO</p></div>
<h3 id="_instructions_llvm_core">Instructions (llvm.core)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_instructions_llvm_core">Instructions (llvm.core)</h3>
<div class="paragraph"><p>TODO</p></div>
<h3 id="_target_data_llvm_ee">Target Data (llvm.ee)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_target_data_llvm_ee">Target Data (llvm.ee)</h3>
<div class="paragraph"><p>TODO</p></div>
<h3 id="_execution_engine_llvm_ee">Execution Engine (llvm.ee)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_execution_engine_llvm_ee">Execution Engine (llvm.ee)</h3>
<div class="paragraph"><p>TODO. For now, see <tt>test/example-jit.py</tt>.</p></div>
<h3 id="_pass_manager_and_passes_llvm_passes">Pass Manager and Passes (llvm.passes)</h3><div style="clear:left"></div>
</div>
<div class="sect2">
<h3 id="_pass_manager_and_passes_llvm_passes">Pass Manager and Passes (llvm.passes)</h3>
<div class="paragraph"><p>TODO. For now, see <tt>test/passes.py</tt>.</p></div>
</div>
</div>
</div>
<div class="sect1">
<h2 id="_about_the_llvm_py_project">About the llvm-py Project</h2>
<div class="sectionbody">
<div class="paragraph"><p>llvm-py lives at
@ -3259,10 +3324,11 @@ are most welcome. You can checkout the latest SVN HEAD from
<div class="paragraph"><p>Mahadevan R wrote llvm-py and works on it in his spare time. He can be
reached at <em>mdevan@mdevan.org</em>.</p></div>
</div>
</div>
<div id="footer">
<div id="footer-text">
Web pages &copy; Mahadevan R. Generated with <a href="http://www.methods.co.nz/asciidoc/">asciidoc</a>.
Last updated 2010-08-31.
Last updated 2010-09-26.
</div>
</div>
</div>