version 0.8.2
This commit is contained in:
parent
581572b28c
commit
053309e60a
125 changed files with 6564 additions and 1308 deletions
180
doc/pegdocs.txt
Executable file
180
doc/pegdocs.txt
Executable file
|
|
@ -0,0 +1,180 @@
|
|||
PEG syntax and semantics
|
||||
========================
|
||||
|
||||
A PEG (Parsing expression grammar) is a simple deterministic grammar, that can
|
||||
be directly used for parsing. The current implementation has been designed as
|
||||
a more powerful replacement for regular expressions. UTF-8 is supported.
|
||||
|
||||
The notation used for a PEG is similar to that of EBNF:
|
||||
|
||||
=============== ============================================================
|
||||
notation meaning
|
||||
=============== ============================================================
|
||||
``A / ... / Z`` Ordered choice: Apply expressions `A`, ..., `Z`, in this
|
||||
order, to the text ahead, until one of them succeeds and
|
||||
possibly consumes some text. Indicate success if one of
|
||||
expressions succeeded. Otherwise do not consume any text
|
||||
and indicate failure.
|
||||
``A ... Z`` Sequence: Apply expressions `A`, ..., `Z`, in this order,
|
||||
to consume consecutive portions of the text ahead, as long
|
||||
as they succeed. Indicate success if all succeeded.
|
||||
Otherwise do not consume any text and indicate failure.
|
||||
The sequence's precedence is higher than that of ordered
|
||||
choice: ``A B / C`` means ``(A B) / Z`` and
|
||||
not ``A (B / Z)``.
|
||||
``(E)`` Grouping: Parenthesis can be used to change
|
||||
operator priority.
|
||||
``{E}`` Capture: Apply expression `E` and store the substring
|
||||
that matched `E` into a *capture* that can be accessed
|
||||
after the matching process.
|
||||
``&E`` And predicate: Indicate success if expression `E` matches
|
||||
the text ahead; otherwise indicate failure. Do not consume
|
||||
any text.
|
||||
``!E`` Not predicate: Indicate failure if expression E matches the
|
||||
text ahead; otherwise indicate success. Do not consume any
|
||||
text.
|
||||
``E+`` One or more: Apply expression `E` repeatedly to match
|
||||
the text ahead, as long as it succeeds. Consume the matched
|
||||
text (if any) and indicate success if there was at least
|
||||
one match. Otherwise indicate failure.
|
||||
``E*`` Zero or more: Apply expression `E` repeatedly to match
|
||||
the text ahead, as long as it succeeds. Consume the matched
|
||||
text (if any). Always indicate success.
|
||||
``E?`` Zero or one: If expression `E` matches the text ahead,
|
||||
consume it. Always indicate success.
|
||||
``[s]`` Character class: If the character ahead appears in the
|
||||
string `s`, consume it and indicate success. Otherwise
|
||||
indicate failure.
|
||||
``[a-b]`` Character range: If the character ahead is one from the
|
||||
range `a` through `b`, consume it and indicate success.
|
||||
Otherwise indicate failure.
|
||||
``'s'`` String: If the text ahead is the string `s`, consume it
|
||||
and indicate success. Otherwise indicate failure.
|
||||
``i's'`` String match ignoring case.
|
||||
``y's'`` String match ignoring style.
|
||||
``v's'`` Verbatim string match: Use this to override a global
|
||||
``\i`` or ``\y`` modifier.
|
||||
``.`` Any character: If there is a character ahead, consume it
|
||||
and indicate success. Otherwise (that is, at the end of
|
||||
input) indicate failure.
|
||||
``_`` Any Unicode character: If there is an UTF-8 character
|
||||
ahead, consume it and indicate success. Otherwise indicate
|
||||
failure.
|
||||
``A <- E`` Rule: Bind the expression `E` to the *nonterminal symbol*
|
||||
`A`. **Left recursive rules are not possible and crash the
|
||||
matching engine.**
|
||||
``\identifier`` Built-in macro for a longer expression.
|
||||
``\ddd`` Character with decimal code *ddd*.
|
||||
``\"``, etc Literal ``"``, etc.
|
||||
=============== ============================================================
|
||||
|
||||
|
||||
Built-in macros
|
||||
---------------
|
||||
|
||||
============== ============================================================
|
||||
macro meaning
|
||||
============== ============================================================
|
||||
``\d`` any decimal digit: ``[0-9]``
|
||||
``\D`` any character that is not a decimal digit: ``[^0-9]``
|
||||
``\s`` any whitespace character: ``[ \9-\13]``
|
||||
``\S`` any character that is not a whitespace character:
|
||||
``[^ \9-\13]``
|
||||
``\w`` any "word" character: ``[a-zA-Z_]``
|
||||
``\W`` any "non-word" character: ``[^a-zA-Z_]``
|
||||
``\n`` any newline combination: ``\10 / \13\10 / \13``
|
||||
``\i`` ignore case for matching; use this at the start of the PEG
|
||||
``\y`` ignore style for matching; use this at the start of the PEG
|
||||
``\ident`` a standard ASCII identifier: ``[a-zA-Z_][a-zA-Z_0-9]*``
|
||||
============== ============================================================
|
||||
|
||||
A backslash followed by a letter is a built-in macro, otherwise it
|
||||
is used for ordinary escaping:
|
||||
|
||||
============== ============================================================
|
||||
notation meaning
|
||||
============== ============================================================
|
||||
``\\`` a single backslash
|
||||
``\*`` same as ``'*'``
|
||||
``\t`` not a tabulator, but an (unknown) built-in
|
||||
============== ============================================================
|
||||
|
||||
|
||||
Supported PEG grammar
|
||||
---------------------
|
||||
|
||||
The PEG parser implements this grammar (written in PEG syntax)::
|
||||
|
||||
# Example grammar of PEG in PEG syntax.
|
||||
# Comments start with '#'.
|
||||
# First symbol is the start symbol.
|
||||
|
||||
grammar <- rule* / expr
|
||||
|
||||
identifier <- [A-Za-z][A-Za-z0-9_]*
|
||||
charsetchar <- "\\" . / [^\]]
|
||||
charset <- "[" "^"? (charsetchar ("-" charsetchar)?)+ "]"
|
||||
stringlit <- identifier? ("\"" ("\\" . / [^"])* "\"" /
|
||||
"'" ("\\" . / [^'])* "'")
|
||||
builtin <- "\\" identifier / [^\13\10]
|
||||
|
||||
comment <- '#' !\n* \n
|
||||
ig <- (\s / comment)* # things to ignore
|
||||
|
||||
rule <- identifier \s* "<-" expr ig
|
||||
identNoArrow <- identifier !(\s* "<-")
|
||||
primary <- (ig '&' / ig '!')* ((ig identNoArrow / ig charset / ig stringlit
|
||||
/ ig builtin / ig '.' / ig '_'
|
||||
/ (ig "(" expr ig ")"))
|
||||
(ig '?' / ig '*' / ig '+')*)
|
||||
|
||||
# Concatenation has higher priority than choice:
|
||||
# ``a b / c`` means ``(a b) / c``
|
||||
|
||||
seqExpr <- primary+
|
||||
expr <- seqExpr (ig "/" expr)*
|
||||
|
||||
|
||||
Examples
|
||||
--------
|
||||
|
||||
Check if `s` matches Nimrod's "while" keyword:
|
||||
|
||||
.. code-block:: nimrod
|
||||
s =~ peg" y'while'"
|
||||
|
||||
Exchange (key, val)-pairs:
|
||||
|
||||
.. code-block:: nimrod
|
||||
"key: val; key2: val2".replace(peg"{\ident} \s* ':' \s* {\ident}", "$2: $1")
|
||||
|
||||
Determine the ``#include``'ed files of a C file:
|
||||
|
||||
.. code-block:: nimrod
|
||||
for line in lines("myfile.c"):
|
||||
if line =~ peg"""s <- ws '#include' ws '"' {[^"]+} '"' ws
|
||||
comment <- '/*' (!'*/' . )* '*/' / '//' .*
|
||||
ws <- (comment / \s+)* """:
|
||||
echo matches[0]
|
||||
|
||||
PEG vs regular expression
|
||||
-------------------------
|
||||
As a regular expression ``\[.*\]`` maches longest possible text between ``'['``
|
||||
and ``']'``. As a PEG it never matches anything, because a PEG is
|
||||
deterministic: ``.*`` consumes the rest of the input, so ``\]`` never matches.
|
||||
As a PEG this needs to be written as: ``\[ ( !\] . )* \]``
|
||||
|
||||
Note that the regular expression does not behave as intended either:
|
||||
``*`` should not be greedy, so ``\[.*?\]`` should be used.
|
||||
|
||||
|
||||
PEG construction
|
||||
----------------
|
||||
There are two ways to construct a PEG in Nimrod code:
|
||||
(1) Parsing a string into an AST which consists of `TPeg` nodes with the
|
||||
`peg` proc.
|
||||
(2) Constructing the AST directly with proc calls. This method does not
|
||||
support constructing rules, only simple expressions and is not as
|
||||
convenient. It's only advantage is that it does not pull in the whole PEG
|
||||
parser into your executable.
|
||||
|
||||
Loading…
Add table
Add a link
Reference in a new issue