Reweave readme
Also fix some syntax errors in the RST
This commit is contained in:
parent
0056ebdd15
commit
a75cfc9887
2 changed files with 57 additions and 44 deletions
98
README.rst
98
README.rst
|
|
@ -30,21 +30,14 @@ By default, NRE compiles it’s own PCRE. If this is undesirable, pass
|
||||||
``-d:pcreDynlib`` to use whatever dynamic library is available on the
|
``-d:pcreDynlib`` to use whatever dynamic library is available on the
|
||||||
system. This may have unexpected consequences if the dynamic library
|
system. This may have unexpected consequences if the dynamic library
|
||||||
doesn’t have certain features enabled.
|
doesn’t have certain features enabled.
|
||||||
|
|
||||||
Types
|
Types
|
||||||
-----
|
-----
|
||||||
|
|
||||||
``type Regex* = ref object``
|
``type Regex* = ref object``
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
Represents the pattern that things are matched against, constructed with
|
Represents the pattern that things are matched against, constructed with
|
||||||
``re(string, string)``. Examples: ``re"foo"``, ``re(r"foo # comment",
|
``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo #
|
||||||
"x<anycrlf>")``, ``re"(?x)(*ANYCRLF)foo # comment"``. For more details
|
comment".``
|
||||||
on the leading option groups, see the `Option
|
|
||||||
Setting <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#OPTION_SETTING>`__
|
|
||||||
and the `Newline
|
|
||||||
Convention <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#NEWLINE_CONVENTION>`__
|
|
||||||
sections of the `PCRE syntax
|
|
||||||
manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
|
|
||||||
|
|
||||||
``pattern: string``
|
``pattern: string``
|
||||||
the string that was used to create the pattern.
|
the string that was used to create the pattern.
|
||||||
|
|
@ -56,34 +49,36 @@ manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
|
||||||
a table from the capture names to their numeric id.
|
a table from the capture names to their numeric id.
|
||||||
|
|
||||||
|
|
||||||
Flags
|
Options
|
||||||
.....
|
.......
|
||||||
|
|
||||||
- ``8`` - treat both the pattern and subject as UTF8
|
The following options may appear anywhere in the pattern, and they affect
|
||||||
- ``9`` - prevents the pattern from being interpreted as UTF, no matter
|
the rest of it.
|
||||||
what
|
|
||||||
- ``A`` - as if the pattern had a ``^`` at the beginning
|
- ``(?i)`` - case insensitive
|
||||||
- ``E`` - DOLLAR\_ENDONLY
|
- ``(?m)`` - multi-line: ``^`` and ``$`` match the beginning and end of
|
||||||
- ``f`` - fails if there is not a match on the first line
|
|
||||||
- ``i`` - case insensitive
|
|
||||||
- ``m`` - multi-line, ``^`` and ``$`` match the beginning and end of
|
|
||||||
lines, not of the subject string
|
lines, not of the subject string
|
||||||
- ``N`` - turn off auto-capture, ``(?foo)`` is necessary to capture.
|
- ``(?s)`` - ``.`` also matches newline (*dotall*)
|
||||||
- ``s`` - ``.`` matches newline
|
- ``(?U)`` - expressions are not greedy by default. ``?`` can be added
|
||||||
- ``U`` - expressions are not greedy by default. ``?`` can be added to
|
to a qualifier to make it greedy
|
||||||
a qualifier to make it greedy.
|
- ``(?x)`` - whitespace and comments (``#``) are ignored (*extended*)
|
||||||
- ``u`` - same as ``8``
|
- ``(?X)`` - character escapes without special meaning (``\w`` vs.
|
||||||
- ``W`` - Unicode character properties; ``\w`` matches ``к``.
|
``\a``) are errors (*extra*)
|
||||||
- ``X`` - "Extra", character escapes without special meaning (``\w``
|
|
||||||
vs. ``\a``) are errors
|
One or a combination of these options may appear only at the beginning
|
||||||
- ``x`` - extended, comments (``#``) and newlines are ignored
|
of the pattern:
|
||||||
(extended)
|
|
||||||
- ``Y`` - pcre.NO\_START\_OPTIMIZE,
|
- ``(*UTF8)`` - treat both the pattern and subject as UTF-8
|
||||||
- ``<cr>`` - newlines are separated by ``\r``
|
- ``(*UCP)`` - Unicode character properties; ``\w`` matches ``я``
|
||||||
- ``<crlf>`` - newlines are separated by ``\r\n`` (Windows default)
|
- ``(*U)`` - a combination of the two options above
|
||||||
- ``<lf>`` - newlines are separated by ``\n`` (UNIX default)
|
- ``(*FIRSTLINE*)`` - fails if there is not a match on the first line
|
||||||
- ``<anycrlf>`` - newlines are separated by any of the above
|
- ``(*NO_AUTO_CAPTURE)`` - turn off auto-capture for groups;
|
||||||
- ``<any>`` - newlines are separated by any of the above and Unicode
|
``(?<name>...)`` can be used to capture
|
||||||
|
- ``(*CR)`` - newlines are separated by ``\r``
|
||||||
|
- ``(*LF)`` - newlines are separated by ``\n`` (UNIX default)
|
||||||
|
- ``(*CRLF)`` - newlines are separated by ``\r\n`` (Windows default)
|
||||||
|
- ``(*ANYCRLF)`` - newlines are separated by any of the above
|
||||||
|
- ``(*ANY)`` - newlines are separated by any of the above and Unicode
|
||||||
newlines:
|
newlines:
|
||||||
|
|
||||||
single characters VT (vertical tab, U+000B), FF (form feed, U+000C),
|
single characters VT (vertical tab, U+000B), FF (form feed, U+000C),
|
||||||
|
|
@ -92,10 +87,15 @@ Flags
|
||||||
are recognized only in UTF-8 mode.
|
are recognized only in UTF-8 mode.
|
||||||
— man pcre
|
— man pcre
|
||||||
|
|
||||||
- ``<bsr_anycrlf>`` - ``\R`` matches CR, LF, or CRLF
|
- ``(*JAVASCRIPT_COMPAT)`` - JavaScript compatibility
|
||||||
- ``<bsr_unicode>`` - ``\R`` matches any unicode newline
|
- ``(*NO_STUDY)`` - turn off studying; study is enabled by default
|
||||||
- ``<js>`` - Javascript compatibility
|
|
||||||
- ``<no_study>`` - turn off studying; study is enabled by deafault
|
For more details on the leading option groups, see the `Option
|
||||||
|
Setting <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#OPTION_SETTING>`__
|
||||||
|
and the `Newline
|
||||||
|
Convention <http://man7.org/linux/man-pages/man3/pcresyntax.3.html#NEWLINE_CONVENTION>`__
|
||||||
|
sections of the `PCRE syntax
|
||||||
|
manual <http://man7.org/linux/man-pages/man3/pcresyntax.3.html>`__.
|
||||||
|
|
||||||
|
|
||||||
``type RegexMatch* = object``
|
``type RegexMatch* = object``
|
||||||
|
|
@ -146,14 +146,24 @@ fields are as follows:
|
||||||
same as ``match``
|
same as ``match``
|
||||||
|
|
||||||
|
|
||||||
``type SyntaxError* = ref object of Exception``
|
``type RegexInternalError* = ref object of RegexException``
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
Internal error in the module, this probably means that there is a bug
|
||||||
|
|
||||||
|
|
||||||
|
``type InvalidUnicodeError* = ref object of RegexException``
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
|
Thrown when matching fails due to invalid unicode in strings
|
||||||
|
|
||||||
|
|
||||||
|
``type SyntaxError* = ref object of RegexException``
|
||||||
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
Thrown when there is a syntax error in the
|
Thrown when there is a syntax error in the
|
||||||
regular expression string passed in
|
regular expression string passed in
|
||||||
|
|
||||||
|
|
||||||
``type StudyError* = ref object of Exception``
|
``type StudyError* = ref object of RegexException``
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
Thrown when studying the regular expression failes
|
Thrown when studying the regular expression failes
|
||||||
for whatever reason. The message contains the error
|
for whatever reason. The message contains the error
|
||||||
code.
|
code.
|
||||||
|
|
@ -244,3 +254,5 @@ If a given capture is missing, a ``ValueError`` exception is thrown.
|
||||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||||
Escapes the string so it doesn’t match any special characters.
|
Escapes the string so it doesn’t match any special characters.
|
||||||
Incompatible with the Extra flag (``X``).
|
Incompatible with the Extra flag (``X``).
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -47,7 +47,8 @@ from unicode import runeLenAt
|
||||||
type
|
type
|
||||||
Regex* = ref object
|
Regex* = ref object
|
||||||
## Represents the pattern that things are matched against, constructed with
|
## Represents the pattern that things are matched against, constructed with
|
||||||
## ``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo # comment".
|
## ``re(string)``. Examples: ``re"foo"``, ``re(r"(*ANYCRLF)(?x)foo #
|
||||||
|
## comment".``
|
||||||
##
|
##
|
||||||
## ``pattern: string``
|
## ``pattern: string``
|
||||||
## the string that was used to create the pattern.
|
## the string that was used to create the pattern.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue