Reweave readme

Also fix some syntax errors in the RST
This commit is contained in:
Flaviu Tamas 2015-05-11 15:45:57 -04:00
commit a75cfc9887
2 changed files with 57 additions and 44 deletions

View file

@ -30,21 +30,14 @@ By default, NRE compiles it’s own PCRE. If this is undesirable, pass
``-d:pcreDynlib`` to use whatever dynamic library is available on the ``-d:pcreDynlib`` to use whatever dynamic library is available on the
system. This may have unexpected consequences if the dynamic library system. This may have unexpected consequences if the dynamic library
doesn’t have certain features enabled. doesn’t have certain features enabled.
Types Types
----- -----
``type Regex* = ref object`` ``type Regex* = ref object``
~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Represents the pattern that things are matched against, constructed with Represents the pattern that things are matched against, constructed with
``re(string, string)``. Examples: ``re"foo"``, ``re(r"foo # comment", ``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo #
"x<anycrlf>")``, ``re"(?x)(*ANYCRLF)foo # comment"``. For more details comment".``
on the leading option groups, see the `Option
Setting <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#OPTION_SETTING>`__
and the `Newline
Convention <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#NEWLINE_CONVENTION>`__
sections of the `PCRE syntax
manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
``pattern: string`` ``pattern: string``
the string that was used to create the pattern. the string that was used to create the pattern.
@ -56,34 +49,36 @@ manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
a table from the capture names to their numeric id. a table from the capture names to their numeric id.
Flags Options
..... .......
- ``8`` - treat both the pattern and subject as UTF8 The following options may appear anywhere in the pattern, and they affect
- ``9`` - prevents the pattern from being interpreted as UTF, no matter the rest of it.
what
- ``A`` - as if the pattern had a ``^`` at the beginning - ``(?i)`` - case insensitive
- ``E`` - DOLLAR\_ENDONLY - ``(?m)`` - multi-line: ``^`` and ``$`` match the beginning and end of
- ``f`` - fails if there is not a match on the first line
- ``i`` - case insensitive
- ``m`` - multi-line, ``^`` and ``$`` match the beginning and end of
lines, not of the subject string lines, not of the subject string
- ``N`` - turn off auto-capture, ``(?foo)`` is necessary to capture. - ``(?s)`` - ``.`` also matches newline (*dotall*)
- ``s`` - ``.`` matches newline - ``(?U)`` - expressions are not greedy by default. ``?`` can be added
- ``U`` - expressions are not greedy by default. ``?`` can be added to to a qualifier to make it greedy
a qualifier to make it greedy. - ``(?x)`` - whitespace and comments (``#``) are ignored (*extended*)
- ``u`` - same as ``8`` - ``(?X)`` - character escapes without special meaning (``\w`` vs.
- ``W`` - Unicode character properties; ``\w`` matches ``к``. ``\a``) are errors (*extra*)
- ``X`` - "Extra", character escapes without special meaning (``\w``
vs. ``\a``) are errors One or a combination of these options may appear only at the beginning
- ``x`` - extended, comments (``#``) and newlines are ignored of the pattern:
(extended)
- ``Y`` - pcre.NO\_START\_OPTIMIZE, - ``(*UTF8)`` - treat both the pattern and subject as UTF-8
- ``<cr>`` - newlines are separated by ``\r`` - ``(*UCP)`` - Unicode character properties; ``\w`` matches ``я``
- ``<crlf>`` - newlines are separated by ``\r\n`` (Windows default) - ``(*U)`` - a combination of the two options above
- ``<lf>`` - newlines are separated by ``\n`` (UNIX default) - ``(*FIRSTLINE*)`` - fails if there is not a match on the first line
- ``<anycrlf>`` - newlines are separated by any of the above - ``(*NO_AUTO_CAPTURE)`` - turn off auto-capture for groups;
- ``<any>`` - newlines are separated by any of the above and Unicode ``(?<name>...)`` can be used to capture
- ``(*CR)`` - newlines are separated by ``\r``
- ``(*LF)`` - newlines are separated by ``\n`` (UNIX default)
- ``(*CRLF)`` - newlines are separated by ``\r\n`` (Windows default)
- ``(*ANYCRLF)`` - newlines are separated by any of the above
- ``(*ANY)`` - newlines are separated by any of the above and Unicode
newlines: newlines:
single characters VT (vertical tab, U+000B), FF (form feed, U+000C), single characters VT (vertical tab, U+000B), FF (form feed, U+000C),
@ -92,10 +87,15 @@ Flags
are recognized only in UTF-8 mode. are recognized only in UTF-8 mode.
— man pcre — man pcre
- ``<bsr_anycrlf>`` - ``\R`` matches CR, LF, or CRLF - ``(*JAVASCRIPT_COMPAT)`` - JavaScript compatibility
- ``<bsr_unicode>`` - ``\R`` matches any unicode newline - ``(*NO_STUDY)`` - turn off studying; study is enabled by default
- ``<js>`` - Javascript compatibility
- ``<no_study>`` - turn off studying; study is enabled by deafault For more details on the leading option groups, see the `Option
Setting <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#OPTION_SETTING>`__
and the `Newline
Convention <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#NEWLINE_CONVENTION>`__
sections of the `PCRE syntax
manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
``type RegexMatch* = object`` ``type RegexMatch* = object``
@ -146,14 +146,24 @@ fields are as follows:
same as ``match`` same as ``match``
``type SyntaxError* = ref object of Exception`` ``type RegexInternalError* = ref object of RegexException``
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Internal error in the module, this probably means that there is a bug
``type InvalidUnicodeError* = ref object of RegexException``
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Thrown when matching fails due to invalid unicode in strings
``type SyntaxError* = ref object of RegexException``
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Thrown when there is a syntax error in the Thrown when there is a syntax error in the
regular expression string passed in regular expression string passed in
``type StudyError* = ref object of Exception`` ``type StudyError* = ref object of RegexException``
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Thrown when studying the regular expression failes Thrown when studying the regular expression failes
for whatever reason. The message contains the error for whatever reason. The message contains the error
code. code.
@ -244,3 +254,5 @@ If a given capture is missing, a ``ValueError`` exception is thrown.
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Escapes the string so it doesn’t match any special characters. Escapes the string so it doesn’t match any special characters.
Incompatible with the Extra flag (``X``). Incompatible with the Extra flag (``X``).

View file

@ -47,7 +47,8 @@ from unicode import runeLenAt
type type
Regex* = ref object Regex* = ref object
## Represents the pattern that things are matched against, constructed with ## Represents the pattern that things are matched against, constructed with
## ``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo # comment". ## ``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo #
## comment".``
## ##
## ``pattern: string`` ## ``pattern: string``
## the string that was used to create the pattern. ## the string that was used to create the pattern.