Initial revision

git-svn-id: https://swig.svn.sourceforge.net/svnroot/swig/trunk@444 626c5289-ae23-0410-ae9c-e8d60b6d4f22
This commit is contained in:
Dave Beazley 2000-05-17 02:21:07 +00:00
commit 45b51a07f8
176 changed files with 95839 additions and 0 deletions

24
swigweb/papers/Perl98/readme Executable file
View file

@ -0,0 +1,24 @@
Perl Extension Building with SWIG
Authors : David Beazley
David Fletcher
Dominique Dumont
Contact : Dave Beazley
beazley@cs.utah.edu
(801) 359-5705
Session : Perl Internals and Linking to C
Thursday, August 20, 1:30 pm.
The presentation consists of a paper about SWIG
and its use with Perl. Three files are available :
swigperl.html - HTML Version of the paper
swigperl.ps - Postscript version (2 column)
swigperl.pdf - PDF Version of the postscript paper
The HTML version contains links to the postscript and PDF
versions.

2198
swigweb/papers/Perl98/swigperl.htm Executable file

File diff suppressed because it is too large Load diff

2120
swigweb/papers/Perl98/swigperl.pdf Executable file

File diff suppressed because one or more lines are too long

572
swigweb/papers/Py96/python96.html Executable file
View file

@ -0,0 +1,572 @@
<html>
<title>
Using SWIG to Control, Prototype, and Debug C Programs with Python
</title>
<body BGCOLOR="ffffff">
<h1> Using SWIG to Control, Prototype, and Debug C Programs with Python </h1>
<b>
(Submitted to the 4th International Python Conference) <br> <br>
David M. Beazley <br>
Department of Computer Science <br>
University of Utah <br>
Salt Lake City, Utah 84112 <br>
<a href="http://www.cs.utah.edu/~beazley"> beazley@cs.utah.edu </a> <br>
</b>
<h2> Abstract </h2>
<em>
I discuss early results in using SWIG to interact with C and C++ programs
using Python. Examples of using SWIG are provided along with a
discussion of the benefits and limitations of this approach. This
paper describes work in progress with the hope of getting feedback and
ideas from the Python community.
</em> <br> <br>
Note : SWIG is freely available at <a href="http://www.cs.utah.edu/~beazley/SWIG"> http://www.cs.utah.edu/~beazley/SWIG </a>
<h2> Introduction </h2>
SWIG (Simplified Wrapper and Interface Generator) is a code
development
tool designed to make it easy for scientists and engineers to add scripting
language interfaces to programs and libraries written in C and C++.
The first version of SWIG was developed in July 1995
for use with
very large scale scientific applications running on the Connection
Machine 5 and Cray T3D at Los Alamos National Laboratory.
Since that time, the system has been extended to support several
interface languages including Python, Tcl, Perl4, Perl5, and Guile [1].
In this paper, I will describe how SWIG can be used to interact with
C and C++ code from Python. I'll also discuss some
applications and limitations of this approach. It is important to
emphasize that SWIG was primarily designed for scientists and engineers
who would like to have a nice interface, but who would rather
work on more interesting problems than figuring out how to build
a user interface or using a complicated interface generation tool.
<h2> Introducing SWIG </h2>
The idea behind SWIG is really quite simple---most interface
languages such as Python provide some mechanism for making
calls to functions written in C or C++. However, this almost
always requires the user to write special "wrapper" functions
that provide the glue between the scripting language and the
C functions. As an example, if you wanted to add the
<tt> getenv() </tt> function to Python, you would need to write
the following wrapper code [4] :
<blockquote>
<tt>
<pre>
static PyObject *wrap_getenv(PyObject *self, PyObject *args) {
char * result;
char * arg0;
if(!PyArg_ParseTuple(args, "s",&arg0))
return NULL;
result = getenv(arg0);
return Py_BuildValue("s", result);
}
</pre> </tt>
</blockquote>
While writing a single wrapper function isn't too difficult,
it quickly becomes tedious and error prone as the number
of functions increases.
SWIG automates this process by generating
wrapper code from a list ANSI C function and variable declarations.
As a result, adding scripting languages to C applications can become
an almost trivial exercise. <br> <br>
The core of the SWIG consists of a YACC parser and a collection
of general purpose utility functions. The output of the parser
is sent to two different modules--a language module for writing
wrapper functions and a documentation module for producing a simple
reference guide. Each of the target languages and documentation methods
is implemented as C++ class that can be plugged in to the system.
Currently, Python, Tcl, Perl4, Perl5, and Guile are supported as
target languages while documentation can be produced in ASCII, HTML,
and LaTeX. Different target languages can be implemented as a new
C++ class that is linked with a general purpose SWIG library file.
<h2> Previous work </h2>
Most Python users are probably aware that automatic wrapper generation is
not a new idea. In particular, packages such as ILU are capable
of providing bindings between C, C++, and Python using specifications
given in IDL (Interface Definition Language) [3]. SWIG is not really meant to compete with this
approach, but is designed to be a no-nonsense, easy to use tool
that scientists and engineers can use to build interesting interfaces
without having to worry about grungy programming details or learning
a complicated interface specification language.
With that said, SWIG is certainly not designed to turn Python into
a C or C++ clone (I'm not sure that turning Python into interpreted
C/C++ compiler would be a good idea anyways). However, SWIG can allow
Python to interface with a surprisingly wide variety of C functions.
</body>
</html>
<h2> A SWIG example </h2>
While it is not my intent to provide a tutorial, I hope to
illustrate how SWIG works by building a module from part of the
UNIX socket library.
While there is probably no need to do this in practice (since Python
already has a socket module), it illustrates most of SWIG's capabilities
on a moderately complex example and provides a basis for
comparison between an existing module and one created by SWIG. <br> <br>
As input, SWIG takes an input file referred to as an "interface file."
The file starts with a preamble containing the name of the
module and a code block where special header files, additional C
code, and other things can be added. After that, ANSI C function
and variable declarations are listed in any order. Since SWIG
uses C syntax, it's usually fairly easy to build an interface
from scratch or simply by copying an existing header file.
The following SWIG interface file will be used for our socket example :
<blockquote>
<tt>
<pre>
// socket.i
// SWIG Interface file to play with some sockets
%init sock // Name of our module
%{
#include &lt;sys/types.h&gt
#include &lt;sys/socket.h&gt
#include &lt;netinet/in.h&gt
#include &lt;arpa/inet.h&gt
#include &lt;netdb.h&gt
/* Set some values in the sockaddr_in structure */
struct sockaddr *new_sockaddr_in(short family, unsigned long hostid, int port) {
struct sockaddr_in *addr;
addr = (struct sockaddr_in *) malloc(sizeof(struct sockaddr_in));
bzero((char *) addr, sizeof(struct sockaddr_in));
addr-&gt;sin_family = family;
addr-&gt;sin_addr.s_addr = hostid;
addr-&gt;sin_port = htons(port);
return (struct sockaddr *) addr;
}
/* Get host address, but return as a string instead of hostent */
char *my_gethostbyname(char *hostname) {
struct hostent *h;
h = gethostbyname(hostname);
if (h) return h-&gt;h_name;
else return "";
}
%}
// Add these constants
enum {AF_UNIX, AF_INET, SOCK_STREAM, SOCK_DGRAM, SOCK_RAW,
IPPROTO_UDP, IPPROTO_TCP, INADDR_ANY};
#define SIZEOF_SOCKADDR sizeof(struct sockaddr)
// Wrap these functions
int socket(int family, int type, int protocol);
int bind(int sockfd, struct sockaddr *myaddr, int addrlen);
int connect(int sockfd, struct sockaddr *servaddr, int addrlen);
int listen(int sockfd, int backlog);
int accept(int sockfd, struct sockaddr *peer, %val int *addrlen);
int close(int fd);
struct sockaddr *new_sockaddr_in(short family, unsigned long, int port);
%name gethostbyname { char *my_gethostbyname(char *); }
unsigned long inet_addr(const char *ip);
%include unixio.i
</pre>
</tt>
</blockquote>
There are several things to notice about this file
<ul>
<li> The module name is specified with the <tt> %init </tt> directive.
<li> Header files and supporting C code is placed in <tt> %{,%} </tt>.
<li> Special support functions may be necessary. For example, <tt>
new_sockaddr_in() </tt> creates a new <tt> struct sockaddr_in </tt> type
and sets some of its values.
<li> Constants are created with <tt> enum </tt> or <tt> #define. </tt>
<li> Functions can take most C basic datatypes as well as pointers
to complex types.
<li> Functions can be renamed with the <tt> %name </tt> directive.
<li> The <tt> %val </tt> directive forces a function argument to be called by
value.
<li> The <tt> %include </tt> directive can be used to include other
SWIG interface files. In this case, the file <tt> unixio.i </tt>
might look like the following :
<blockquote>
<tt>
<pre>
// File : unixio.i
// Some file I/O and memory functions
%{
%}
int read(int fd, void *buffer, int n);
int write(int fd, void *buffer, int n);
typedef unsigned int size_t;
void *malloc(size_t nbytes);
</pre>
</tt>
</blockquote>
<li> Since interface files can include other interface files, it can be very
easy to build up libraries and collections of useful functions.
<li> <tt> typedef </tt> can be used to provide mapping between datatypes.
<li> C and C++ comments are allowed and like C, SWIG ignores whitespace.
</ul>
The key thing to notice about SWIG interface files is that they support
real C functions, constants, and datatypes including C pointers. As
a general rule these files are also independent of the target scripting
language--the primary reason why it's easy to support different
languages.
<h2> Building a Python module </h2>
Building a Python module with SWIG is a simple process that proceeds as
follows (assuming you're running Solaris 2.x) :
<blockquote>
<tt>
<pre>
unix > wrap -python socket.i
unix > gcc -c socket_wrap.c -I/usr/local/include/Py
unix > ld -G socket_wrap.o -lsocket -lnsl -o sockmodule.so
</pre>
</tt>
</blockquote>
The <tt> wrap </tt> command takes the interface file and produces a C file
called <tt> socket_wrap.c </tt>. This file is then compiled and built into
a shared object file.
<h2> A sample Python script </h2>
Our new socket module can be used normally within a Python script.
For example, we could write a simple server process to echo all
data received back to a client :
<blockquote>
<tt>
<pre>
# Echo server program using SWIG module
from sock import *
PORT = 5000
sockfd = socket(AF_INET, SOCK_STREAM, 0)
if sockfd < 0:
raise IOError, 'server : unable to open stream socket'
addr = new_sockaddr_in(AF_INET,INADDR_ANY,PORT)
if (bind(sockfd, addr, SIZEOF_SOCKADDR)) < 0 :
raise IOError, 'server : unable to bind local address'
listen(sockfd,5)
client_addr = new_sockaddr_in(AF_INET, 0, 0)
newsockfd = accept(sockfd, client_addr, SIZEOF_SOCKADDR)
buffer = malloc(1024)
while 1:
nbytes = read(newsockfd, buffer, 1024)
if nbytes > 0 :
write(newsockfd,buffer,nbytes)
else : break
close(newsockfd)
close(sockfd)
</pre>
</tt>
</blockquote>
In our Python module, we are manipulating a socket, in almost
exactly the same way as would be done in a C program. We'll take
a look at a client written using the Python socket library shortly.
<h2> Pointers and run-time type checking </h2>
SWIG imports C pointers as ASCII strings containing both the
address and pointer type. Thus, a typical SWIG pointer might look
like the following : <br> <br>
<center>
<tt> _fd2a0_struct_sockaddr_p </tt> <br> <br>
</center>
Some may view the idea of importing C pointers into Python as unsafe or even
truly insane. However, handling pointers seems to be needed in order
for SWIG to handle most interesting types of C code. To provide some safety,
the wrapper functions generated by SWIG perform pointer-type checking
(since importing C pointers into
Python has in effect, bypassed type-checking in the C compiler).
When incompatible pointer types are used, this is what happens :
<blockquote>
<tt>
<pre>
unix > python
Python 1.3 (Apr 12 1996) [GCC 2.5.8]
Copyright 1991-1995 Stichting Mathematisch Centrum, Amsterdam
>>> from sock import *
>>> sockfd = socket(AF_INET, SOCK_STREAM,0);
>>> buffer = malloc(8192);
>>> bind(sockfd,buffer,SIZEOF_SOCKADDR);
Traceback (innermost last):
File "<stdin>", line 1, in ?
TypeError: Type error in argument 2 of bind. Expected _struct_sockaddr_p.
>>>
</pre>
</tt>
</blockquote>
SWIG's run-time type checking provides a certain degree of safety from
using invalid parameters and making simple mistakes. Of course, any system
supporting pointers can be abused, but SWIG's implementation has proven
to be quite reliable under normal use.
<h2> Comparison with a real Python module </h2>
As a point of comparison, we can now write a client script using the
Python socket module (this example courtesy of the Python library
reference manual) [5].
<blockquote>
<tt>
<pre>
# Echo client program
from socket import *
HOST = 'tjaze.lanl.gov'
PORT = 5000
s = socket(AF_INET, SOCK_STREAM)
s.connect(HOST,PORT)
s.send('Hello world')
data = s.recv(1024)
s.close()
print 'Received', `data`
</pre>
</tt>
</blockquote>
As one might expect, the two scripts look fairly similar. However
there are certain obvious differences. First, the SWIG generated
module is clearly a quick and dirty approach. We must explicitly
check for errors and create a few C data structures. Furthermore,
while Python functions such as <tt> socket() </tt> can take optional
arguments, SWIG generated functions can not (since the underlying
C function can't). Of course, the apparent ugliness
of the SWIG module compared to the Python equivalent is not the
point here.
The real idea to emphasize is that SWIG
can take a set of real C functions with moderate complexity and
produce a fully functional Python module in a very short period of
time. In this case, about 10 minutes worth of effort.
(I'm
guessing that the Python socket module required more time than that).
With a bit of work, the SWIG generated module could probably be
used to build a more user-friendly version that hid many of the details.
<h2> Controlling C programs </h2>
Of course, the real goal of SWIG is not to produce quirky
replacements for modules in the Python library. Instead, it
better suited as a mechanism for
controlling a variety of C programs. <br> <br>
In the SPaSM molecular dynamics code used at Los Alamos National
Laboratory, SWIG
is used to build an interface out of about 200 C functions [2].
This happens at compile time so most users don't
notice that it has occurred. However, as a result, it is now
extremely easy for the physicists using the code to
extend it with new functions.
Whenever new functionality is added, the new
functions are put into a SWIG interface file and the functions
become available when the code is recompiled (a process
which usually only involves the new C functions and SWIG generated
wrapper file). Of course, when running under Python, it is possible
to build extensions and dynamically load them as needed. <br> <br>
Another point, not to be overlooked, is the fact that SWIG makes it
very easy for a user to combine bits and pieces of completely different
software packages without waiting for someone to write a special
purpose module. For example, SWIG can be used to import the C API
of Matlab 4.2 into Python. This combined with a simulation module
can allow a scientist to perform both simulation and data analysis in
interactively or from a script. One could even build a Tkinter interface to control
the entire system. The ease of using SWIG makes it possible for
scientists to build applications out of components that
might not have been considered earlier.
<h2> Prototyping and Debugging </h2>
When debugging new modules or libraries, it would be nice to be able
to test out the functions without having to repeatedly recompile the
module with the rest of a large application. With SWIG, it is often
possible to prototype and debug modules independently. SWIG can be
used to build Python wrappers around the functions and scripts
can be used for testing. In many cases, it is possible to create
the C data structures and other information that would have been generated
by the larger application all from within a Python script. Since
SWIG requires no changes to the C code, integrating the new C code into
the large package is usually pretty easy. <br> <br>
A related benefit of the SWIG approach is that it is now possible
to easily experiment with large software libraries and packages.
Peter-Pike Sloan at the
University of Utah, used SWIG to wrap the entire contents of the Open-GL
library (including more than 560 constants and 340 C functions).
This ``interpreted'' Open-GL could then be used interactively to
determine various image parameters and to draw simple figures.
Is this case, it proved to be a remarkably effective way of
figuring out the effects of different parameters and functions without
having to recompile after every modification.
<h2> Living dangerously with datatypes </h2>
One of the reasons SWIG is easy to use for prototyping and control
applications is its extremely forgiving treatment of datatypes.
<tt> typedef </tt> can be used to map datatypes, but whenever
SWIG encounters an unknown datatype, it simply assumes that it's
some sort of complex datatype. As a result, it's possible to
wrap some fairly complex C code with almost no effort. For example,
one could create a module allowing Python to create a Tcl interpreter
and execute Tcl commands as follows :
<blockquote>
<tt>
<pre>
// tcl.i
// Wrap Tcl's C interface
%init tcl
%{
#include &lt;tcl.h&gt;
%}
Tcl_Interp *Tcl_CreateInterp(void);
void Tcl_DeleteInterp(Tcl_Interp *interp);
int Tcl_Eval(Tcl_Interp *interp, char *script);
int Tcl_EvalFile(Tcl_Interp *interp, char *file);
</pre>
</tt>
</blockquote>
SWIG has no idea what <tt> Tcl_Interp </tt> is, but it doesn't
really care as long as it's used as a pointer. When used in a
Python script, this module would work as follows :
<blockquote>
<tt>
<pre>
unix > python
Python 1.3 (Mar 26 1996) [GCC 2.7.0]
Copyright 1991-1995 Stichting Mathematisch Centrum, Amsterdam
>>> import tcl
>>> interp = tcl.Tcl_CreateInterp();
>>> tcl.Tcl_Eval(interp,"for {set i 0} {$i < 5} {incr i 1} {puts $i}");
0
1
2
3
4
0
>>>
</pre>
</tt>
</blockquote>
<h2> Limitations </h2>
This approach represents a balance of flexibility and ease of use.
I have always felt that the tool shouldn't be more complicated than
the original problem. Therefore, SWIG certainly won't do everything.
Some of SWIG's limitations include it's limited understanding of
complex datatypes. Often the user must write special functions
to extract information from structures even though passing complex
datatypes around by reference works remarkably well in practice.
SWIG's support for C++ is also quite limited. At the moment SWIG
can be used to pass pointers to C++ objects around, but actually
transforming a C++ object into some sort of Python object is far
beyond the capabilities of the current implementation. Finally, SWIG
lacks an exception model which may be of concern to some users.
<h2> Limitations in the Python implementation </h2>
SWIG's support for Python is in many ways, limited by the fact that
SWIG was originally developed for use with other languages. For example,
representing C pointers as a character string can be seen as a carry-over
from an earlier Tcl implementation. With a little work, I believe
that one could create a Python "pointer" datatype while still retaining
type-checking and other features. <br> <br>
Another limitation in the current Python implementation is the apparent lack
of variable linking support. Many applications, particularly scientific ones,
have global variables that a user may want to access. SWIG currently
creates Python functions for accessing these variables, but other
scripting languages provide a more natural mechanism.
For example, in Perl, it is
possible to attach "set" and "get" functions to a Perl object. These
are evaluated whenever that variable is accessed in a script.
For all practical purposes this special variable looks like any other Perl
variable, but is really linked to an C global variable (of course,
maybe there is a reason why Perl calls these "magic" variables). <br> <br>
If an equivalent variable linking mechanism is available in Python, I
couldn't find it [4].
If not, it might be possible to create a new kind of Python datatype
that supports variable linking, but which interacts seamlessly
with existing Python integer, floating point, and character string types.
<h2> Summary and future work </h2>
This paper has described work in progress with the SWIG system and
its support for Python. I believe that this approach shows great
promise for future work--especially within the scientific and
engineering community. Most of the people introduced to SWIG and
Python have
found both systems extremely easy and straightforward to use.
I am currently
quite interested in improving SWIG's interface to Python and exploring
its use with numerical Python in order to build interesting
and flexible scientific applications [6].
<h2> Acknowledgments </h2>
This project would not have been possible without the support of a number
of people. Peter Lomdahl, Shujia Zhou, Brad Holian, Tim Germann, and Niels
Jensen at Los Alamos National Laboratory were the first users and were
instrumental in testing out the first designs. Patrick Tullmann,
John Schmidt,Peter-Pike Sloan, and Kurtis Bleeker at the University of Utah have also
been instrumental in testing out some of the newer versions and
providing feedback.
John Buckman has suggested many interesting
improvements and is currently working on porting SWIG to non-Unix
platforms including NT, OS/2 and Macintosh.
Finally
I'd like to thank Chris Johnson and the members of the Scientific
Computing and Imaging group at the University of Utah for their continued
support. Some of this work was performed under the auspices of the
Department of Energy.
<h2> References </h2>
[1] D.M. Beazley. <em> SWIG : An Easy to Use Tool for Integrating
Scripting Languages with C and C++ </em>. Proceedings of the 4th
USENIX Tcl/Tk Workshop (1996) (to appear). <br> <br>
[2] D.M. Beazley and P.S. Lomdahl <em> Message-Passing Multi-Cell
Molecular Dynamics on the Connection Machine 5 </em>, Parallel
Computing 20 (1994), p. 173-195. <br> <br>
[3] Bill Janssen and Mike Spreitzer, <em> ILU : Inter-Language
Unification via Object Modules</em>, OOPSLA 94 Workshop on
Multi-Language Object Models.
<br> <br>
[4] Guido van Rossum, <em> Extending and Embedding the Python Interpreter,
</em> October 1995. <br> <br>
[5] Guido van Rossum, <em> Python Library Reference </em>, October
1995. <br> <br>
[6] Paul F. Dubois, Konrad Hinsen, and James Hugunin, <em> Numerical
Python, </em> Computers in Physics, (1996) (to appear)
</body>
</html>

View file

@ -0,0 +1,762 @@
<html>
<title> Feeding a Large-scale Physics Application to Python </title>
<body bgcolor="#ffffff">
<h1> Feeding a Large-scale Physics Application to Python </h1>
David M. Beazley <br>
Department of Computer Science <br>
University of Utah <br>
Salt Lake City, Utah 84112 <br>
<tt> beazley@cs.utah.edu </tt> <br>
<p>
Peter S. Lomdahl <br>
Theoretical Division <br>
Los Alamos National Laboratory <br>
Los Alamos, New Mexico 87545 <br>
<tt> pxl@lanl.gov </tt> <br>
<h2> Abstract </h2>
<em>
We describe our experiences using Python with the SPaSM molecular
dynamics code at Los Alamos National Laboratory. Originally developed
as a large monolithic application for massively parallel processing
systems, we have used Python to transform our application into a
flexible, highly modular, and extremely powerful system for performing
simulation, data analysis, and visualization. In addition, we
describe how Python has solved a number of important problems related
to the development, debugging, deployment, and maintenance of
scientific software. </em>
<h2> Background </h2>
For the past 5 years, we have been developing a large-scale physics
code for performing molecular-dynamics simulations of
materials. This code, SPaSM (Scalable Parallel Short-range
Molecular-dynamics), was originally developed for the Connection
Machine 5 massively parallel supercomputing system and later moved to
a number of other machines including the Cray T3D, multiprocessor Sun
and SGI systems, and Unix workstations [1,2]. Our goal has
been to investigate material properties such as fracture, crack
propagation, dislocation generation, friction, and ductile-brittle transitions [3].
The SPaSM code
helps us investigate these problems by performing atomistic
simulations--that is, we simulate the dynamics of every atom in a
material and hope to make sense of what happens. Of course, the
underlying physics is not so important for this discussion.
<p>
While the development of SPaSM is an ongoing effort, we
have been hampered by a number of serious problems.
First, typical
simulations generate tens to hundreds of gigabytes of data that must
be analyzed. This task is not easily performed on a user's
workstation, nor is it economically feasible to buy everyone their own
personal desktop supercomputer. A second problem is that of
interactivity and control. We are constantly making changes to
investigate new physical models, different materials, and so forth.
This would usually require changes to the underlying C
code--a process that was tedious and not very user friendly. We
wanted a more flexible mechanism.
Finally, there were many difficulties associated with the
development and maintenance of our software. While we are only a small
group, it was not uncommon for different users to have their
own private copies of the software that had been modified in some manner.
This, in turn, led to a maintenance nightmare that made it almost impossible
to update the software or apply bug-fixes in a consistent manner.
<p>
To address these problems, we started investigating the use of scripting
languages. In 1995, we wrote a special purpose parallel-scripting language,
but replaced it with Python a year later (although the interface
generation tool for that scripting language lives on as SWIG). In this
paper, we hope to describe some of our experiences with Python and how
it has helped us solve practical scientific computing problems. In particular,
we describe the organization of our system, module building process,
interesting tools that Python has helped us develop, and why we think this
approach is particularly well-suited for scientific computing research.
<h2> Why Python? </h2>
Although our code originally used a custom scripting language, we decided
to switch to Python for a number of reasons :
<p>
<ul>
<li> Features. Python had a rich set of datatypes, support for
object oriented programming, namespaces, exceptions, dynamic loading,
and a large number of useful modules.
<li> Syntax. Our system is controlled through
text-based commands and scripts. We felt that Python had a nice
syntax that does not require users to
type weird symbols ($,%,@, etc...) or use syntax that is radically
different than C.
<li> A small core. The modular structure of Python makes it easy for us
to remove or add modules as needed. Given the difficult task of running
on parallel machines and supercomputers, we felt that this structure
would make it easier for us to port Python to a variety of special-purpose
machines (since we could just remove problematic modules).
<li> Availability of documentation. Users could go to the bookstore
to find out more about Python.
<li> Support and stability. Python appeared to be highly stable and
well supported through newsgroups, special interest groups, and the PSA.
<li> Freely available. We have been able to
use and modify Python as needed for our application. This would have
been impossible without access to the Python source.
</ul>
<p>
Another factor, not to be overlooked, is the increasing acceptance of
Python elsewhere in the computational science community. Efforts at
Lawrence Livermore National Laboratory and elsewhere were
attractive--in particular, we saw Python as being a potential vehicle
for sharing modules and utilizing third-party tools [4,9,10].
<h2> System Organization </h2>
Our goal was to build a highly modular system that could perform
simulation, data analysis, and visualization--often performing
all of these tasks simultaneously. Unfortunately, there is a tendency
to do this by building a tightly integrated
monolithic package (perhaps using a well-structured C++ class
hierarchy for example). In our view, this is too
formal and restrictive. We wanted to support modules that might only
be loosely related to each other. For example, there is no need for a
graphics library to depend on the same structures as a simulation code
or to even be written in the same language for that matter. Likewise,
we wanted to exploit third-party modules where no assumptions
could be made about their internal structure.
<p>
The major components of our system take the form of C libraries. When
Python is used, the libraries are compiled into shared libraries and
dynamically loaded into Python as extension modules. The
functionality of each library is exposed as a collection of Python
"commands." Because of this, the C and Python programming
environments are closely related. In many cases, it is possible to
implement the same code in both C or Python. For example :
<blockquote>
<tt><pre>
/* A simple function written in C */
#include "SPaSM.h"
void run(int nsteps, double Dt, int freq) {
int i;
char filename[64];
for (i = 0; i < nsteps; i++) {
integrate_adv_coord(Dt);
boundary_periodic();
redistribute();
force_eam();
integrate_adv_velocity(Dt);
if ((i % freq) == 0) {
sprintf(filename,"Dat%d",i);
output_particles(filename);
}
}
}</pre>
</tt>
</blockquote>
Now, in Python :
<blockquote><tt><pre>
# A function written in Python
from SPaSM import *
def run(nsteps, Dt, freq):
for i in xrange(0, nsteps):
integrate_adv_coord(Dt)
boundary_periodic()
redistribute()
force_eam()
integrate_adv_velocity(Dt)
if (i % freq) == 0) :
output_particles('Dat'+str(i))
</pre>
</tt>
</blockquote>
While C libraries provide much of the underlying functionality, the real power
of the system comes in the form of modules and scripts written entirely in Python.
Users write scripts to set up and control simulations. Several major
components such as the visualization and data-analysis system make heavy
use of Python. We also utilize a variety of modules in the Python
library and have dynamically loadable versions of Tkinter and the
Python Imaging Library [11].
<h2> Embedding Python and Hiding System Dependencies </h2>
One of the biggest implementation problems we have encountered is the
fact that the SPaSM code is a parallel application that relies heavily
upon the proper implementation of low-level system services such as
I/O and process management. Currently, it is possible to run the code
in two different configurations--one that uses message passing via the
MPI library and another using Solaris threads. Using Python with
both versions of code requires a degree of care--in particular, we
have found it to be necessary to provide Python with enhanced I/O
support to run properly in parallel. This work has been described
elsewhere [5,6].
<p>
Handling two operational modes introduces a number of problems related
to code maintenance and installation. Traditionally, we would
recompile the entire system and all of its modules for each
configuration (placing the files in an architecture dependent
subdirectory). Users would have to decide which system they wanted to
use and compile all of their modules by linking against the
appropriate libraries and setting the right compile-time options.
Changing configurations would typically require a complete recompile.
<p>
With dynamic loading and shared libraries however, we have been able
to devise a different approach to this problem. Rather
than recompiling everything for each configuration, we use an
implementation independent layer of system wrappers. These wrappers
provide a generic implementation of message passing, parallel I/O, and
thread management. All of the core modules are then compiled
using these generic wrappers. This makes the modules independent
of the underlying operational mode---to use MPI or threads, we simply
need to supply a different implementation of the system wrapper libraries.
This is easily accomplished by building two different versions of
Python that are linked against the appropriate system libraries.
To run in a particular mode, we now just run the appropriate
version of Python (i.e. 'python' or 'pythonmpi'). The neat
part about this approach is that all of the modules work with both
operational modes without any recompilation or reconfiguration. If a
user is using threads, but wants to switch to MPI, they simply run
a different version of Python--no recompilation of modules is necessary.
<p>
A full discussion of writing system-wrappers can be found elsewhere.
In particular, a discussion of writing parallel I/O wrappers for
Python can be found in [5]. An earlier discussion of the technique we
have used for writing message passing and I/O wrappers can also be
found in [7].
<h2> Module Building with SWIG </h2>
To build modules, we have been using SWIG [8].
Each module is described by a SWIG interface file containing the ANSI
C declarations of functions, structures, and variables in that module.
For example :
<blockquote><pre>
// SWIG interface file
%module SPaSM
%{
#include "SPaSM.h"
%}
void integrate_adv_coord(double Dt);
void boundary_periodic();
void redistribute();
void force_eam();
void integrate_adv_velocity(double Dt);
int output_particles(char *filename);
</pre></blockquote>
SWIG provides a logical mapping of the underlying C implementation into
Python. During compilation, interface files are automatically
converted into wrapper code and compiled into Python modules. This
process is entirely transparent--changes made to the interface
are automatically propagated to Python whenever a module is recompiled.
Given the constantly evolving nature of research applications, this makes
it easy to extend and maintain the system.
<h3> Separation of Implementation and Interface </h3>
An important aspect of SWIG is that it is requires no modifications to
existing C code which allows us to maintain a strict
separation between the implementation of C modules and their
Python interface. We believe that this results in code that
is more organized and generally reusable. There is no particular
reason why a C module should depend on Python (one might
want to use it as a stand-alone package or in a different application).
Despite using Python extensively, the SPaSM code can still be
compiled with no Python interface (of course, you lose all of
the benefits gained by having Python).
We feel that most physics codes tend
to have a rather long life-span. Maintaining a separation of
implementation and interface helps insure that our physics code will
be usable in the future--even if there are drastic changes in the
interface along the way.
<h3> Providing Access to Data Structures </h3>
For the purposes of debugging and data exploration, we have
used SWIG to provide wrappers around C data structures. For example,
a collection of C structures such as the following
<blockquote><pre>
typedef struct {
double x,y,z;
} Vector;
typedef struct {
int type;
Vector r;
Vector v;
Vector f;
} Particle;
</pre></blockquote>
can be turned into Python wrapper classes. In addition, SWIG can extend structures with member functions as follows :
<blockquote> <pre>
// SWIG interface for vector and particle structures
%include datatypes.h
// Extend data structures with a few useful methods
%addmethods Vector {
char *__str__() {
static char a[1024];
sprintf(a,"[ %0.10f, %0.10f, %0.10f ]", self->x, self->y, self->z);
return a;
}
}
%addmethods Particle {
Particle *__getitem__(int index) {
return self+index;
}
char *__str__() {
// print out a particle
...
}
}
</pre>
</blockquote>
When the Python interface is built, C structures now appear like
Python objects. By providing Python-specific methods (such as
<tt>__getitem__</tt>) we can even provide array access.
For example, the following Python code would print out
all of the coordinates of stored particles :
<blockquote>
<pre>
# Print out all coordinates to a file
f = open("part.data","w")
p = SPaSM_first_particle() # Get first particle pointer
for i in xrange(0,SPaSM_count_particles()):
f.write("%f, %f, %f\n" % (p[i].r.x, p[i].r.y, p[i].r.z))
f.close()
</pre>
</blockquote>
While only a simple example, having direct access to underlying data has
proven to be quite valuable since we view the internal representation
of data, check values, and perform diagnostics.
<h3> Improving the Reliability of Modules </h3>
When instrumenting our original physics application to use scripting,
we found that many parts of the code were not written in an entirely
"reliable" manner. In particular, the code had never been operated in
an event-driven manner. Functions often made assumptions about
initializations and rarely checked the validity of input parameters.
To address these issues, most sections of code were gradually changed
to provide some kind of validation of input values and error recovery.
The SWIG compiler has also been extended to provide some of these
capabilities as well.
<p>
One of Python's most powerful features is its exception handling mechanism.
Exceptions are easily raised and handled in Python scripts as follows :
<blockquote><pre>
# A Python function that throws an exception
def allocate(nbytes):
ptr = SPaSM_malloc(nbytes)
if ptr == "NULL": raise MemoryError,"Out of memory!"
# A Python function that catches an exception
def foo():
try:
allocate(NBYTES)
except:
return # Bailing out
</blockquote></pre>
We have borrowed this idea and implemented a similar exception handling
mechanism for our C code. This is accomplished using functions in the
<tt>&lt;setjmp.h&gt;</tt> library and defining a few C macros for "<tt>Try</tt>", "<tt>Except</tt>",
"<tt>Throw</tt>", etc... Using these macros, many of our library functions now look like the following :
<blockquote><pre>
/* A C function that throws an exception */
void *SPaSM_malloc(size_t nbytes) {
void *ptr = (void *) malloc(nbytes);
if (!ptr) Throw("SPaSM_malloc : Out of memory!");
return ptr;
}
</pre></blockquote>
Like Python, we allow C functions to
catch exceptions and provide their own recovery as follows :
<blockquote> <pre>
/* A C function catching an exception */
int foo() {
void *p;
Try {
p = SPaSM_malloc(NBYTES);
} Except {
printf("Unable to allocate memory. Returning!\n");
return -1;
}
}
</pre></blockquote>
In the case of a stand-alone C application, exceptions can
be caught and handled internally. Should an uncaught exception
occur, the code prints an error message and terminates. However,
when Python is used, we can
generate Python exceptions using a SWIG user-defined
exception handler such as the following :
<blockquote> <pre>
%module SPaSM
// A SWIG user defined exception handler
%except(python) {
Try {
$function
} Except {
PyErr_SetString(PyExc_RuntimeError,SPaSM_error_msg());
}
}
// C declarations
...
</pre></blockquote>
The handler code gets placed into all of the Python
"wrapper" functions and is responsible for translating C exceptions
into Python exceptions. Doing this makes our physics code operate
in a more seamless manner and gives it a precisely defined error recovery
procedure (i.e. internal errors always result Python exceptions). For example :
<blockquote><pre>
>>> SPaSM_malloc(1000000000)
RuntimeError: SPaSM_malloc(1000000000). Out of memory!
(Line 52 in memory.c)
>>>
</pre></blockquote>
We have found error recovery to be critical. Without it, simulations may
continue to run, only to generate wrong answers or a mysterious
system crash.
<h2> More Than Scripting </h2>
When we originally started using scripting languages, we thought
they would mainly be a convenient mechanism for gluing C libraries
together and controlling them in an interactive manner. However,
we have come to realize that scripting languages are
much more powerful than this. Now, we find ourselves
implementing significant functionality entirely in Python--bypassing
C altogether. The most surprising (well, not really that surprising)
fact is that Python makes it possible to build very powerful tools with
only a small amount of extra programming. In this section, we present
a brief overview of some of the tools we have developed--most of
which have been implemented largely in Python.
<h3> Object-Oriented Visualization and Data Analysis </h3>
Our original goal was to add a powerful data analysis and
visualization component to our application. To this end, we developed
a lightweight high-performance graphics library implemented in C.
This library supports both 2D and 3D plotting and produces output in
the form of GIF images. To make plots, a user writes simple C
functions such as the following :
<blockquote><pre>
void plot_spheres(Plot3D *p3, DataFunction func, double min, double max,
double radius) {
Particle *p = Particles;
int i, npart, color;
npart = SPaSM_count_particles();
for (i = 0; i < npart; i++, p++) {
/* Compute color value */
value = (*func)(p);
color = (value-min)/(max-min)*255;
/* Plot it */
Plot3D_sphere(p3,p->r.x,p->r.y,p->r.z,radius,color);
}
}
</pre></blockquote>
When executed, this function produces a raw image such as the following
showing stacking-faults generated by a passing shock wave in an fcc crystal
of 10 million atoms :
<p>
<center>
<img src="bigshockraw.gif">
</center>
<p>
While simple, we often want our images to contain more information
including titles, axis labels, colorbars, time-stamps, and bounding
boxes. To do this, we have built an object-oriented visualization
framework in Python. Python classes provide methods for graph annotation
and common graph operations. If a user wants to make a new kind of plot,
they simply inherit from an appropriate base class and provide a function
to plot the desired data. For example, a "Sphere" plot using the above C
function would look like this :
<blockquote><pre>
class Spheres(Image3D):
def __init__(self, func, min, max, radius=0.5):
Image3D.__init__(self, ...
self.func = func
self.min = min
self.max = max
self.radius = radius
def draw(self):
self.newplot()
plot_spheres(self.p3,self.func,self.min,self.max,self.radius)
</pre></blockquote>
To use the new image, one simply creates an object of that type and
manipulates it. For example :
<blockquote><pre>
>>> s = Spheres(PE,-8,-3.5) # Create a new image
>>> s.rotd(45) # Rotate it down
>>> s.zoom(200) # Zoom in
>>> s.show() # generate Image and display it
>>> ...
</pre></blockquote>
<center>
<img src="bigshock.gif">
</center>
<p>
Thus, with a simple C function and a simple Python class, we get a lot
of functionality for free, including methods for image manipulation,
graph annotation, and display. In the current system, there are about
a dozen different types of images for 2D and 3D plotting. It is also
possible to perform filtering, apply clipping planes, and other
advanced operations. Most of this is supported by about 2000 lines of
Python code and a relatively small C library containing performance
critical operations.
<h3> Web-Based Simulation Monitoring </h3>
One feature of many physics codes is that they tend to run for a long
time. In our case, a simulation may run for tens to hundreds of
hours. During the course of a simulation, it is desirable to check on
its status and see how it is progressing. Given the strong Internet
support already bundled with Python, we decided to write a simple
physics web-server that could be used for this purpose. Unlike a
more traditional server, our server is used by "registering" various
objects that one wants to look at during the course of a simulation.
The simulation code then periodically polls a network socket to see if
anyone has requested anything. If so, the simulation will stop for a
moment, generate the requested information, send it to the user, and
then continue on with the calculation. A simplified example of using
the server is as follows :
<blockquote> <pre>
# Simplified script using a web-server
from vis import *
from web import *
web = SPaSMWeb()
web.add(WebMain(""))
web.add(WebFileText("Msg"+`run_no`))
# Create an image object
set_graphics_mode(HTTP)
ke = Spheres(KE,0,20)
ke.title = "Kinetic Energy"
web.add(WebImage("ke.gif",ke))
def run(nsteps):
for i in xrange(0,nsteps):
integrate(1) # Integrate 1 timestep
web.poll() # Anyone looking for me?
# Run it
run(10000)
</pre></blockquote>
While simple, this example gives the basic idea. We create a
server object and register a few "links." When a user connects with
the server, they will be presented with some status information and a
list of available links. When the links are accessed, the server
feeds real-time data back to the user. In the case of
images, they are generated immediately and sent back in GIF format. During
this process, no temporary files are created, nor is the server
transmitting previously stored information (i.e. everything is
created on-the-fly).
<p>
While we are still refining the implementation, this approach has turned
out to be useful--not only can we periodically check up
on a running simulation, this can can be done from any machine that
has a Web-browser, including PCs and Macs running over a modem line
(allowing a bored physicist to check on long-running jobs from home
or while on travel). Of course, the most amazing fact of all is that
this was implemented by one person in an afternoon and only involved
about 150 lines of Python code (with generous help from the Python library).
<h3> Development Support </h3>
Finally, we have found Python to be quite useful for supporting
the future development of our code. Some of this comes in the form of
sophisticated debugging--Python can be used to analyze internal data
structures and track down problems. Since Python can read
and modify most of the core C data structures, it is possible to
prototype new functions or to perform one-time operations without ever
loading up the C compiler.
<p>
We have also used Python in conjunction with code
management. One of the newer problems we have faced is the task of
finding the definitions of C functions (for example,
we might want to know how a particular command has been
implemented in C). To support this, we have written tools for
browsing source directories, displaying definitions, and spawning
editors. All of this is implemented in Python and can be performed
directly from the Physics application. Python has made these kinds
of tools very easy to write--often only requiring a few hours of
effort.
<h2> Python and the Development of Scientific Software </h2>
The adoption of Python has had a profound effect on the overall
structure of our application. With time, the code became more
modular, more reliable, and better organized.
Furthermore, these gains have been achieved
without significant losses in performance, increased coding complexity,
or substantially increased development cost. It should also be stressed
that automated tools such as SWIG have played a critical role in this
effort by hiding the Python-C interface and allowing us to concentrate
on the real problems at hand (i.e. physics).
<p>
One point, that we would like to strongly emphasize, is the dynamic
nature of many scientific applications. Our application is not a huge
monolithic package that never changes and which no one is supposed to
modify. Rather, the code is <em>constantly</em> changing to explore
new problems, try new numerical methods, or to provide better
performance. The highly modular nature of Python works great in
this environment. We provide a common repository of modules that are shared by all
of the users of the system. Different users are typically responsible
for different modules and all users may be working on different modules
at the same time.
Whenever changes to a module are made, they are immediately propagated
to everyone else. In the case of bug-fixes, this means that patched
versions of code are available immediately. Likewise, new
features show up automatically and can be easily tested by everyone.
And of course, if something breaks--others tend to notice almost immediately
(which we feel is a good thing).
<h2> Limitations of Python </h2>
While we consider our Python "experiment" to be highly successful, we
have experienced a number of problems along the way. The first is
merely a technical issue--it would be nice if the mechanism for
building static extensions to Python were made easier. While we are
currently using dynamic loading, many machines, especially
supercomputing systems, do not support it. On these machines it is
necessary to staticly link everything. Libraries help simplify the
task, but combining different statically linked modules together into
a single Python interpreter still remains magical and somewhat
problematic (especially since we are creating an application that
changes a lot--not merely trying to patch the Python core with a new
extension module).
<p>
A second problem is that of education--we have found that some users
have a very difficult time making sense of the system (even we're
somewhat amazed that it really works).
Interestingly enough, this does not seem to be directly related to the
use of C code or Python for that matter, but to issues related to the
configuration, compilation, debugging, and installation of modules. While most
users have written programs before, they have never dealt with shared
libraries, automated code generators, high-level languages, and third-party packages
such as Python. In other cases, there seems to be a
"fear of autoconf"--if a program requires a configure script, it must
be too complicated to understand (which is not the case
here). The combination of these
factors has led some users to believe that the system is far more
complicated and fragile than it really is. We're not sure how to combat
this problem other than to try and educate the users
(i.e. it's essentially the same code as before, but with a really cool
hack).
<h2> Conclusions and Future Work </h2>
Python is cool and we plan to keep using it. However much work
remains. One of the more exciting aspects of Python is its solid
support for modular programming and the culture of cooperation in the
Python community. There are already many Python-related scientific
computing projects underway. If these efforts were coordinated
in some manner, we feel that the results could be spectacular.
<h2> Acknowledgments </h2>
We would like acknowledge our collaborators Brad Holian, Tim Germann, Shujia Zhou, and Wanshu Huang at Los Alamos National Laboratory. We would also like
to acknowledge Paul Dubois and Brian Yang at Lawrence Livermore
National Laboratory for many interesting conversations concerning Python
and its use in physics applications. We would also like to
acknowledge the Scientific Computing and Imaging Group at the
University of Utah for their continued support. Finally, we'd like to
thank the entire Python development community for making a truly
awesome tool for solving real problems. Development of the SPaSM code
has been performed under the auspices of the United States Department
of Energy.
<h2> References </h2>
<p>
[1] D.M.Beazley and P.S. Lomdahl, "Message
Passing Multi-Cell Molecular Dynamics on the Connection Machine 5,"
Parallel Computing 20 (1994), p. 173-195. <p>
<p>
[2] P.S.Lomdahl, P.Tamayo, N.Gronbech-Jensen,
and D.M.Beazley, "50 Gflops Molecular Dynamics on the CM-5," Proceedings
of Supercomputing 93, IEEE Computer Society (1993), p.520-527. <p>
<p>
[3] S.J.Zhou, D.M. Beazley, P.S. Lomdahl, B.L. Holian, Physical Review
Letters. 78, 479 (1997).
<p>
[4] P. Dubois, K. Hinsen, and J. Hugunin, Computers in Physics (10), (1996)
p. 262-267.
<p>
[5] D. M. Beazley and P.S. Lomdahl, "Extensible Message Passing Application
Development and Debugging with Python", Proceedings of IPPS'97, Geneva
Switzerland, IEEE Computer Society (1997).
<p>
[6] T.-Y.B. Yang, D.M. Beazley, P. F. Dubois, G. Furnish, "Steering Object-oriented Computations with
Python", Python Workshop 5, Washington D.C., Nov. 4-5, 1996.
<p>
[7] D.M. Beazley and P.S. Lomdahl, "High
Performance Molecular Dynamics Modeling with SPaSM : Performance and
Portability Issues," Proceedings of the Workshop on Debugging and
Tuning for Parallel Computer Systems, Chatham, MA, 1994 (IEEE Computer
Society Press, Los Alamitos, CA, 1996), pp. 337-351.
<p>
[8] D.M. Beazley, "Using SWIG to Control, Prototype, and Debug C Programs with Python",
Proceedings of the 4th International Python Conference, Lawrence Livermore
National Laboratory, June 3-6, 1996.
<p>
[9] K. Hinsen, The Molecular Modeling Toolkit, <tt>http://starship.skyport.net/crew/hinsen/mmtk.html</tt>
<p>
[10] Jim Hugunin, Numeric Python, <tt>http://www.sls.lcs.mit.edu/jjh/numpy</tt>
<p>
[11] Fredrik Lundh, Python Imaging Library, <tt>http://www.python.org/sigs/image-sig/Imaging.html</tt>
</body>
</html>

763
swigweb/papers/Py97/beazley.html Executable file
View file

@ -0,0 +1,763 @@
<html>
<title> Feeding a Large-scale Physics Application to Python </title>
<body bgcolor="#ffffff">
<h1> Feeding a Large-scale Physics Application to Python </h1>
David M. Beazley <br>
Department of Computer Science <br>
University of Utah <br>
Salt Lake City, Utah 84112 <br>
<tt> beazley@cs.utah.edu </tt> <br>
<p>
Peter S. Lomdahl <br>
Theoretical Division <br>
Los Alamos National Laboratory <br>
Los Alamos, New Mexico 87545 <br>
<tt> pxl@lanl.gov </tt> <br>
<p>
(Presented at the 6th International Python Conference, San Jose, California. October 14-17, 1997).
<h2> Abstract </h2>
<em>
We describe our experiences using Python with the SPaSM molecular
dynamics code at Los Alamos National Laboratory. Originally developed
as a large monolithic application for massively parallel processing
systems, we have used Python to transform our application into a
flexible, highly modular, and extremely powerful system for performing
simulation, data analysis, and visualization. In addition, we
describe how Python has solved a number of important problems related
to the development, debugging, deployment, and maintenance of
scientific software. </em>
<h2> Background </h2>
For the past 5 years, we have been developing a large-scale physics
code for performing molecular-dynamics simulations of
materials. This code, SPaSM (Scalable Parallel Short-range
Molecular-dynamics), was originally developed for the Connection
Machine 5 massively parallel supercomputing system and later moved to
a number of other machines including the Cray T3D, multiprocessor Sun
and SGI systems, and Unix workstations [1,2]. Our goal has
been to investigate material properties such as fracture, crack
propagation, dislocation generation, friction, and ductile-brittle transitions [3].
The SPaSM code
helps us investigate these problems by performing atomistic
simulations--that is, we simulate the dynamics of every atom in a
material and hope to make sense of what happens. Of course, the
underlying physics is not so important for this discussion.
<p>
While the development of SPaSM is an ongoing effort, we
have been hampered by a number of serious problems.
First, typical
simulations generate tens to hundreds of gigabytes of data that must
be analyzed. This task is not easily performed on a user's
workstation, nor is it economically feasible to buy everyone their own
personal desktop supercomputer. A second problem is that of
interactivity and control. We are constantly making changes to
investigate new physical models, different materials, and so forth.
This would usually require changes to the underlying C
code--a process that was tedious and not very user friendly. We
wanted a more flexible mechanism.
Finally, there were many difficulties associated with the
development and maintenance of our software. While we are only a small
group, it was not uncommon for different users to have their
own private copies of the software that had been modified in some manner.
This, in turn, led to a maintenance nightmare that made it almost impossible
to update the software or apply bug-fixes in a consistent manner.
<p>
To address these problems, we started investigating the use of scripting
languages. In 1995, we wrote a special purpose parallel-scripting language,
but replaced it with Python a year later (although the interface
generation tool for that scripting language lives on as SWIG). In this
paper, we hope to describe some of our experiences with Python and how
it has helped us solve practical scientific computing problems. In particular,
we describe the organization of our system, module building process,
interesting tools that Python has helped us develop, and why we think this
approach is particularly well-suited for scientific computing research.
<h2> Why Python? </h2>
Although our code originally used a custom scripting language, we decided
to switch to Python for a number of reasons :
<p>
<ul>
<li> Features. Python had a rich set of datatypes, support for
object oriented programming, namespaces, exceptions, dynamic loading,
and a large number of useful modules.
<li> Syntax. Our system is controlled through
text-based commands and scripts. We felt that Python had a nice
syntax that does not require users to
type weird symbols ($,%,@, etc...) or use syntax that is radically
different than C.
<li> A small core. The modular structure of Python makes it easy for us
to remove or add modules as needed. Given the difficult task of running
on parallel machines and supercomputers, we felt that this structure
would make it easier for us to port Python to a variety of special-purpose
machines (since we could just remove problematic modules).
<li> Availability of documentation. Users could go to the bookstore
to find out more about Python.
<li> Support and stability. Python appeared to be highly stable and
well supported through newsgroups, special interest groups, and the PSA.
<li> Freely available. We have been able to
use and modify Python as needed for our application. This would have
been impossible without access to the Python source.
</ul>
<p>
Another factor, not to be overlooked, is the increasing acceptance of
Python elsewhere in the computational science community. Efforts at
Lawrence Livermore National Laboratory and elsewhere were
attractive--in particular, we saw Python as being a potential vehicle
for sharing modules and utilizing third-party tools [4,9,10].
<h2> System Organization </h2>
Our goal was to build a highly modular system that could perform
simulation, data analysis, and visualization--often performing
all of these tasks simultaneously. Unfortunately, there is a tendency
to do this by building a tightly integrated
monolithic package (perhaps using a well-structured C++ class
hierarchy for example). In our view, this is too
formal and restrictive. We wanted to support modules that might only
be loosely related to each other. For example, there is no need for a
graphics library to depend on the same structures as a simulation code
or to even be written in the same language for that matter. Likewise,
we wanted to exploit third-party modules where no assumptions
could be made about their internal structure.
<p>
The major components of our system take the form of C libraries. When
Python is used, the libraries are compiled into shared libraries and
dynamically loaded into Python as extension modules. The
functionality of each library is exposed as a collection of Python
"commands." Because of this, the C and Python programming
environments are closely related. In many cases, it is possible to
implement the same code in both C or Python. For example :
<blockquote>
<tt><pre>
/* A simple function written in C */
#include "SPaSM.h"
void run(int nsteps, double Dt, int freq) {
int i;
char filename[64];
for (i = 0; i < nsteps; i++) {
integrate_adv_coord(Dt);
boundary_periodic();
redistribute();
force_eam();
integrate_adv_velocity(Dt);
if ((i % freq) == 0) {
sprintf(filename,"Dat%d",i);
output_particles(filename);
}
}
}</pre>
</tt>
</blockquote>
Now, in Python :
<blockquote><tt><pre>
# A function written in Python
from SPaSM import *
def run(nsteps, Dt, freq):
for i in xrange(0, nsteps):
integrate_adv_coord(Dt)
boundary_periodic()
redistribute()
force_eam()
integrate_adv_velocity(Dt)
if (i % freq) == 0) :
output_particles('Dat'+str(i))
</pre>
</tt>
</blockquote>
While C libraries provide much of the underlying functionality, the real power
of the system comes in the form of modules and scripts written entirely in Python.
Users write scripts to set up and control simulations. Several major
components such as the visualization and data-analysis system make heavy
use of Python. We also utilize a variety of modules in the Python
library and have dynamically loadable versions of Tkinter and the
Python Imaging Library [11].
<h2> Embedding Python and Hiding System Dependencies </h2>
One of the biggest implementation problems we have encountered is the
fact that the SPaSM code is a parallel application that relies heavily
upon the proper implementation of low-level system services such as
I/O and process management. Currently, it is possible to run the code
in two different configurations--one that uses message passing via the
MPI library and another using Solaris threads. Using Python with
both versions of code requires a degree of care--in particular, we
have found it to be necessary to provide Python with enhanced I/O
support to run properly in parallel. This work has been described
elsewhere [5,6].
<p>
Handling two operational modes introduces a number of problems related
to code maintenance and installation. Traditionally, we would
recompile the entire system and all of its modules for each
configuration (placing the files in an architecture dependent
subdirectory). Users would have to decide which system they wanted to
use and compile all of their modules by linking against the
appropriate libraries and setting the right compile-time options.
Changing configurations would typically require a complete recompile.
<p>
With dynamic loading and shared libraries however, we have been able
to devise a different approach to this problem. Rather
than recompiling everything for each configuration, we use an
implementation independent layer of system wrappers. These wrappers
provide a generic implementation of message passing, parallel I/O, and
thread management. All of the core modules are then compiled
using these generic wrappers. This makes the modules independent
of the underlying operational mode---to use MPI or threads, we simply
need to supply a different implementation of the system wrapper libraries.
This is easily accomplished by building two different versions of
Python that are linked against the appropriate system libraries.
To run in a particular mode, we now just run the appropriate
version of Python (i.e. 'python' or 'pythonmpi'). The neat
part about this approach is that all of the modules work with both
operational modes without any recompilation or reconfiguration. If a
user is using threads, but wants to switch to MPI, they simply run
a different version of Python--no recompilation of modules is necessary.
<p>
A full discussion of writing system-wrappers can be found elsewhere.
In particular, a discussion of writing parallel I/O wrappers for
Python can be found in [5]. An earlier discussion of the technique we
have used for writing message passing and I/O wrappers can also be
found in [7].
<h2> Module Building with SWIG </h2>
To build modules, we have been using SWIG [8].
Each module is described by a SWIG interface file containing the ANSI
C declarations of functions, structures, and variables in that module.
For example :
<blockquote><pre>
// SWIG interface file
%module SPaSM
%{
#include "SPaSM.h"
%}
void integrate_adv_coord(double Dt);
void boundary_periodic();
void redistribute();
void force_eam();
void integrate_adv_velocity(double Dt);
int output_particles(char *filename);
</pre></blockquote>
SWIG provides a logical mapping of the underlying C implementation into
Python. During compilation, interface files are automatically
converted into wrapper code and compiled into Python modules. This
process is entirely transparent--changes made to the interface
are automatically propagated to Python whenever a module is recompiled.
Given the constantly evolving nature of research applications, this makes
it easy to extend and maintain the system.
<h3> Separation of Implementation and Interface </h3>
An important aspect of SWIG is that it is requires no modifications to
existing C code which allows us to maintain a strict
separation between the implementation of C modules and their
Python interface. We believe that this results in code that
is more organized and generally reusable. There is no particular
reason why a C module should depend on Python (one might
want to use it as a stand-alone package or in a different application).
Despite using Python extensively, the SPaSM code can still be
compiled with no Python interface (of course, you lose all of
the benefits gained by having Python).
We feel that most physics codes tend
to have a rather long life-span. Maintaining a separation of
implementation and interface helps insure that our physics code will
be usable in the future--even if there are drastic changes in the
interface along the way.
<h3> Providing Access to Data Structures </h3>
For the purposes of debugging and data exploration, we have
used SWIG to provide wrappers around C data structures. For example,
a collection of C structures such as the following
<blockquote><pre>
typedef struct {
double x,y,z;
} Vector;
typedef struct {
int type;
Vector r;
Vector v;
Vector f;
} Particle;
</pre></blockquote>
can be turned into Python wrapper classes. In addition, SWIG can extend structures with member functions as follows :
<blockquote> <pre>
// SWIG interface for vector and particle structures
%include datatypes.h
// Extend data structures with a few useful methods
%addmethods Vector {
char *__str__() {
static char a[1024];
sprintf(a,"[ %0.10f, %0.10f, %0.10f ]", self->x, self->y, self->z);
return a;
}
}
%addmethods Particle {
Particle *__getitem__(int index) {
return self+index;
}
char *__str__() {
// print out a particle
...
}
}
</pre>
</blockquote>
When the Python interface is built, C structures now appear like
Python objects. By providing Python-specific methods (such as
<tt>__getitem__</tt>) we can even provide array access.
For example, the following Python code would print out
all of the coordinates of stored particles :
<blockquote>
<pre>
# Print out all coordinates to a file
f = open("part.data","w")
p = SPaSM_first_particle() # Get first particle pointer
for i in xrange(0,SPaSM_count_particles()):
f.write("%f, %f, %f\n" % (p[i].r.x, p[i].r.y, p[i].r.z))
f.close()
</pre>
</blockquote>
While only a simple example, having direct access to underlying data has
proven to be quite valuable since we view the internal representation
of data, check values, and perform diagnostics.
<h3> Improving the Reliability of Modules </h3>
When instrumenting our original physics application to use scripting,
we found that many parts of the code were not written in an entirely
"reliable" manner. In particular, the code had never been operated in
an event-driven manner. Functions often made assumptions about
initializations and rarely checked the validity of input parameters.
To address these issues, most sections of code were gradually changed
to provide some kind of validation of input values and error recovery.
The SWIG compiler has also been extended to provide some of these
capabilities as well.
<p>
One of Python's most powerful features is its exception handling mechanism.
Exceptions are easily raised and handled in Python scripts as follows :
<blockquote><pre>
# A Python function that throws an exception
def allocate(nbytes):
ptr = SPaSM_malloc(nbytes)
if ptr == "NULL": raise MemoryError,"Out of memory!"
# A Python function that catches an exception
def foo():
try:
allocate(NBYTES)
except:
return # Bailing out
</blockquote></pre>
We have borrowed this idea and implemented a similar exception handling
mechanism for our C code. This is accomplished using functions in the
<tt>&lt;setjmp.h&gt;</tt> library and defining a few C macros for "<tt>Try</tt>", "<tt>Except</tt>",
"<tt>Throw</tt>", etc... Using these macros, many of our library functions now look like the following :
<blockquote><pre>
/* A C function that throws an exception */
void *SPaSM_malloc(size_t nbytes) {
void *ptr = (void *) malloc(nbytes);
if (!ptr) Throw("SPaSM_malloc : Out of memory!");
return ptr;
}
</pre></blockquote>
Like Python, we allow C functions to
catch exceptions and provide their own recovery as follows :
<blockquote> <pre>
/* A C function catching an exception */
int foo() {
void *p;
Try {
p = SPaSM_malloc(NBYTES);
} Except {
printf("Unable to allocate memory. Returning!\n");
return -1;
}
}
</pre></blockquote>
In the case of a stand-alone C application, exceptions can
be caught and handled internally. Should an uncaught exception
occur, the code prints an error message and terminates. However,
when Python is used, we can
generate Python exceptions using a SWIG user-defined
exception handler such as the following :
<blockquote> <pre>
%module SPaSM
// A SWIG user defined exception handler
%except(python) {
Try {
$function
} Except {
PyErr_SetString(PyExc_RuntimeError,SPaSM_error_msg());
}
}
// C declarations
...
</pre></blockquote>
The handler code gets placed into all of the Python
"wrapper" functions and is responsible for translating C exceptions
into Python exceptions. Doing this makes our physics code operate
in a more seamless manner and gives it a precisely defined error recovery
procedure (i.e. internal errors always result Python exceptions). For example :
<blockquote><pre>
>>> SPaSM_malloc(1000000000)
RuntimeError: SPaSM_malloc(1000000000). Out of memory!
(Line 52 in memory.c)
>>>
</pre></blockquote>
We have found error recovery to be critical. Without it, simulations may
continue to run, only to generate wrong answers or a mysterious
system crash.
<h2> More Than Scripting </h2>
When we originally started using scripting languages, we thought
they would mainly be a convenient mechanism for gluing C libraries
together and controlling them in an interactive manner. However,
we have come to realize that scripting languages are
much more powerful than this. Now, we find ourselves
implementing significant functionality entirely in Python--bypassing
C altogether. The most surprising (well, not really that surprising)
fact is that Python makes it possible to build very powerful tools with
only a small amount of extra programming. In this section, we present
a brief overview of some of the tools we have developed--most of
which have been implemented largely in Python.
<h3> Object-Oriented Visualization and Data Analysis </h3>
Our original goal was to add a powerful data analysis and
visualization component to our application. To this end, we developed
a lightweight high-performance graphics library implemented in C.
This library supports both 2D and 3D plotting and produces output in
the form of GIF images. To make plots, a user writes simple C
functions such as the following :
<blockquote><pre>
void plot_spheres(Plot3D *p3, DataFunction func, double min, double max,
double radius) {
Particle *p = Particles;
int i, npart, color;
npart = SPaSM_count_particles();
for (i = 0; i < npart; i++, p++) {
/* Compute color value */
value = (*func)(p);
color = (value-min)/(max-min)*255;
/* Plot it */
Plot3D_sphere(p3,p->r.x,p->r.y,p->r.z,radius,color);
}
}
</pre></blockquote>
When executed, this function produces a raw image such as the following
showing stacking-faults generated by a passing shock wave in an fcc crystal
of 10 million atoms :
<p>
<center>
<img src="bigshockraw.gif">
</center>
<p>
While simple, we often want our images to contain more information
including titles, axis labels, colorbars, time-stamps, and bounding
boxes. To do this, we have built an object-oriented visualization
framework in Python. Python classes provide methods for graph annotation
and common graph operations. If a user wants to make a new kind of plot,
they simply inherit from an appropriate base class and provide a function
to plot the desired data. For example, a "Sphere" plot using the above C
function would look like this :
<blockquote><pre>
class Spheres(Image3D):
def __init__(self, func, min, max, radius=0.5):
Image3D.__init__(self, ...
self.func = func
self.min = min
self.max = max
self.radius = radius
def draw(self):
self.newplot()
plot_spheres(self.p3,self.func,self.min,self.max,self.radius)
</pre></blockquote>
To use the new image, one simply creates an object of that type and
manipulates it. For example :
<blockquote><pre>
>>> s = Spheres(PE,-8,-3.5) # Create a new image
>>> s.rotd(45) # Rotate it down
>>> s.zoom(200) # Zoom in
>>> s.show() # generate Image and display it
>>> ...
</pre></blockquote>
<center>
<img src="bigshock.gif">
</center>
<p>
Thus, with a simple C function and a simple Python class, we get a lot
of functionality for free, including methods for image manipulation,
graph annotation, and display. In the current system, there are about
a dozen different types of images for 2D and 3D plotting. It is also
possible to perform filtering, apply clipping planes, and other
advanced operations. Most of this is supported by about 2000 lines of
Python code and a relatively small C library containing performance
critical operations.
<h3> Web-Based Simulation Monitoring </h3>
One feature of many physics codes is that they tend to run for a long
time. In our case, a simulation may run for tens to hundreds of
hours. During the course of a simulation, it is desirable to check on
its status and see how it is progressing. Given the strong Internet
support already bundled with Python, we decided to write a simple
physics web-server that could be used for this purpose. Unlike a
more traditional server, our server is used by "registering" various
objects that one wants to look at during the course of a simulation.
The simulation code then periodically polls a network socket to see if
anyone has requested anything. If so, the simulation will stop for a
moment, generate the requested information, send it to the user, and
then continue on with the calculation. A simplified example of using
the server is as follows :
<blockquote> <pre>
# Simplified script using a web-server
from vis import *
from web import *
web = SPaSMWeb()
web.add(WebMain(""))
web.add(WebFileText("Msg"+`run_no`))
# Create an image object
set_graphics_mode(HTTP)
ke = Spheres(KE,0,20)
ke.title = "Kinetic Energy"
web.add(WebImage("ke.gif",ke))
def run(nsteps):
for i in xrange(0,nsteps):
integrate(1) # Integrate 1 timestep
web.poll() # Anyone looking for me?
# Run it
run(10000)
</pre></blockquote>
While simple, this example gives the basic idea. We create a
server object and register a few "links." When a user connects with
the server, they will be presented with some status information and a
list of available links. When the links are accessed, the server
feeds real-time data back to the user. In the case of
images, they are generated immediately and sent back in GIF format. During
this process, no temporary files are created, nor is the server
transmitting previously stored information (i.e. everything is
created on-the-fly).
<p>
While we are still refining the implementation, this approach has turned
out to be useful--not only can we periodically check up
on a running simulation, this can can be done from any machine that
has a Web-browser, including PCs and Macs running over a modem line
(allowing a bored physicist to check on long-running jobs from home
or while on travel). Of course, the most amazing fact of all is that
this was implemented by one person in an afternoon and only involved
about 150 lines of Python code (with generous help from the Python library).
<h3> Development Support </h3>
Finally, we have found Python to be quite useful for supporting
the future development of our code. Some of this comes in the form of
sophisticated debugging--Python can be used to analyze internal data
structures and track down problems. Since Python can read
and modify most of the core C data structures, it is possible to
prototype new functions or to perform one-time operations without ever
loading up the C compiler.
<p>
We have also used Python in conjunction with code
management. One of the newer problems we have faced is the task of
finding the definitions of C functions (for example,
we might want to know how a particular command has been
implemented in C). To support this, we have written tools for
browsing source directories, displaying definitions, and spawning
editors. All of this is implemented in Python and can be performed
directly from the Physics application. Python has made these kinds
of tools very easy to write--often only requiring a few hours of
effort.
<h2> Python and the Development of Scientific Software </h2>
The adoption of Python has had a profound effect on the overall
structure of our application. With time, the code became more
modular, more reliable, and better organized.
Furthermore, these gains have been achieved
without significant losses in performance, increased coding complexity,
or substantially increased development cost. It should also be stressed
that automated tools such as SWIG have played a critical role in this
effort by hiding the Python-C interface and allowing us to concentrate
on the real problems at hand (i.e. physics).
<p>
One point, that we would like to strongly emphasize, is the dynamic
nature of many scientific applications. Our application is not a huge
monolithic package that never changes and which no one is supposed to
modify. Rather, the code is <em>constantly</em> changing to explore
new problems, try new numerical methods, or to provide better
performance. The highly modular nature of Python works great in
this environment. We provide a common repository of modules that are shared by all
of the users of the system. Different users are typically responsible
for different modules and all users may be working on different modules
at the same time.
Whenever changes to a module are made, they are immediately propagated
to everyone else. In the case of bug-fixes, this means that patched
versions of code are available immediately. Likewise, new
features show up automatically and can be easily tested by everyone.
And of course, if something breaks--others tend to notice almost immediately
(which we feel is a good thing).
<h2> Limitations of Python </h2>
While we consider our Python "experiment" to be highly successful, we
have experienced a number of problems along the way. The first is
merely a technical issue--it would be nice if the mechanism for
building static extensions to Python were made easier. While we are
currently using dynamic loading, many machines, especially
supercomputing systems, do not support it. On these machines it is
necessary to staticly link everything. Libraries help simplify the
task, but combining different statically linked modules together into
a single Python interpreter still remains magical and somewhat
problematic (especially since we are creating an application that
changes a lot--not merely trying to patch the Python core with a new
extension module).
<p>
A second problem is that of education--we have found that some users
have a very difficult time making sense of the system (even we're
somewhat amazed that it really works).
Interestingly enough, this does not seem to be directly related to the
use of C code or Python for that matter, but to issues related to the
configuration, compilation, debugging, and installation of modules. While most
users have written programs before, they have never dealt with shared
libraries, automated code generators, high-level languages, and third-party packages
such as Python. In other cases, there seems to be a
"fear of autoconf"--if a program requires a configure script, it must
be too complicated to understand (which is not the case
here). The combination of these
factors has led some users to believe that the system is far more
complicated and fragile than it really is. We're not sure how to combat
this problem other than to try and educate the users
(i.e. it's essentially the same code as before, but with a really cool
hack).
<h2> Conclusions and Future Work </h2>
Python is cool and we plan to keep using it. However much work
remains. One of the more exciting aspects of Python is its solid
support for modular programming and the culture of cooperation in the
Python community. There are already many Python-related scientific
computing projects underway. If these efforts were coordinated
in some manner, we feel that the results could be spectacular.
<h2> Acknowledgments </h2>
We would like acknowledge our collaborators Brad Holian, Tim Germann, Shujia Zhou, and Wanshu Huang at Los Alamos National Laboratory. We would also like
to acknowledge Paul Dubois and Brian Yang at Lawrence Livermore
National Laboratory for many interesting conversations concerning Python
and its use in physics applications. We would also like to
acknowledge the Scientific Computing and Imaging Group at the
University of Utah for their continued support. Finally, we'd like to
thank the entire Python development community for making a truly
awesome tool for solving real problems. Development of the SPaSM code
has been performed under the auspices of the United States Department
of Energy.
<h2> References </h2>
<p>
[1] D.M.Beazley and P.S. Lomdahl, "Message
Passing Multi-Cell Molecular Dynamics on the Connection Machine 5,"
Parallel Computing 20 (1994), p. 173-195. <p>
<p>
[2] P.S.Lomdahl, P.Tamayo, N.Gronbech-Jensen,
and D.M.Beazley, "50 Gflops Molecular Dynamics on the CM-5," Proceedings
of Supercomputing 93, IEEE Computer Society (1993), p.520-527. <p>
<p>
[3] S.J.Zhou, D.M. Beazley, P.S. Lomdahl, B.L. Holian, Physical Review
Letters. 78, 479 (1997).
<p>
[4] P. Dubois, K. Hinsen, and J. Hugunin, Computers in Physics (10), (1996)
p. 262-267.
<p>
[5] D. M. Beazley and P.S. Lomdahl, "Extensible Message Passing Application
Development and Debugging with Python", Proceedings of IPPS'97, Geneva
Switzerland, IEEE Computer Society (1997).
<p>
[6] T.-Y.B. Yang, D.M. Beazley, P. F. Dubois, G. Furnish, "Steering Object-oriented Computations with
Python", Python Workshop 5, Washington D.C., Nov. 4-5, 1996.
<p>
[7] D.M. Beazley and P.S. Lomdahl, "High
Performance Molecular Dynamics Modeling with SPaSM : Performance and
Portability Issues," Proceedings of the Workshop on Debugging and
Tuning for Parallel Computer Systems, Chatham, MA, 1994 (IEEE Computer
Society Press, Los Alamitos, CA, 1996), pp. 337-351.
<p>
[8] D.M. Beazley, "Using SWIG to Control, Prototype, and Debug C Programs with Python",
Proceedings of the 4th International Python Conference, Lawrence Livermore
National Laboratory, June 3-6, 1996.
<p>
[9] K. Hinsen, The Molecular Modeling Toolkit, <tt>http://starship.skyport.net/crew/hinsen/mmtk.html</tt>
<p>
[10] Jim Hugunin, Numeric Python, <tt>http://www.sls.lcs.mit.edu/jjh/numpy</tt>
<p>
[11] Fredrik Lundh, Python Imaging Library, <tt>http://www.python.org/sigs/image-sig/Imaging.html</tt>
</body>
</html>

BIN
swigweb/papers/Py97/bigshock.gif Executable file

Binary file not shown.

After

Width:  |  Height:  |  Size: 80 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 101 KiB

File diff suppressed because one or more lines are too long

1084
swigweb/papers/Tcl96/.~newtcl.html Executable file

File diff suppressed because it is too large Load diff

BIN
swigweb/papers/Tcl96/Tcl_Talk.pdf Executable file

Binary file not shown.

BIN
swigweb/papers/Tcl96/fig1.gif Executable file

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.6 KiB

BIN
swigweb/papers/Tcl96/hits.gif Executable file

Binary file not shown.

After

Width:  |  Height:  |  Size: 21 KiB

BIN
swigweb/papers/Tcl96/steer2.gif Executable file

Binary file not shown.

After

Width:  |  Height:  |  Size: 9.7 KiB

1084
swigweb/papers/Tcl96/tcl96.html Executable file

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

File diff suppressed because one or more lines are too long