This wiki is in the process of being archived due to lack of usage and the resources necessary to serve it — predominately to bots, crawlers, and LLM companies. Edits are discouraged.
Pages are preserved as they were at the time of archival. For current information, please visit python.org.
If a change to this archive is absolutely needed, requests can be made via the infrastructure@python.org mailing list.

Here's a way to make wiki-like syntaxes in Python.

Lines as Tokens

Usually, you tokenize input word-by-word.

For a wiki-like syntax, it can be easier to go line by line.

Data Flow

We start with:

= Hello, world! =
This is an example.

We want to demonstrate how you take text with
multiple lines, and turn it into one paragraph.

...turn it into...

   1 [("HEADER", ("=", " Hello, world! ")),
   2  ("TEXT", ("This is an example.",)),
   3  ("BLANK", ()),
   4  ("TEXT", ("We want to demonstrate how you take text with",)),
   5  ("TEXT", ("multiple lines, and turn it into one paragraph.",)),
   6 ]

...and from there into...

   1 [("HEADER", ("=", " Hello, world! ")),
   2  ("PARAGRAPH", ["This is an example."]),
   3  ("BLANK", ()),
   4  ("PARAGRAPH", ["We want to demonstrate how you take text with",
   5                 "multiple lines, and turn it into one paragraph."]),
   6 ]

...and finally into:

<h1> Hello, world! </h1>
<p>This is an example.</p>
<p>We want to demonstrate how you take text with multple lines, and turn it into one paragraph.</p>

Spec out Types of Lines

Our first task is to spec out the types of lines that exist.

Use RegularExpressions!

Let's start with just three types of lines. (It'll be clear how to add more.)

We'll have header lines, (like in wiki,) blank lines (seperating paragraphs,) and text lines, which will be anything that doesn't match the others.

Examples:

== Level 2 Heading ==

Here's paragraph 1.
Ladee dah dee dah.

Here's paragraph 2.

=== Level 3 Header ===
Look, a paragraph without a blank line above it!
No problem, we can parse it!

Here are our RegularExpressions:

   1 header = re.compile(r"^(=+)(.+?)\1$")
   2 blank = re.compile(r"^(\s*)$")
   3 text = re.compile(r"^(.+)$")  # if nothing else matched, use this

(Note that we store interesting information in regex groups. This is so we can get to it later on.)

This is good to start with. You can add more types now, if you want, though!

Tokenize a Line

Now, we'll teach Python how to tokenize a line.

It goes like so:

   1 def tokenize_line(line):
   2     """Tokenize a line, returning the token type and recognized groups."""
   3     mo = header.match(line)
   4     if mo is not None:
   5         return ("HEADER", mo.groups())
   6     mo = blank.match(line)
   7     if mo is not None:
   8         return ("BLANK", mo.groups())
   9     mo = text.match(line)
  10     if mo is not None:
  11         return ("TEXT", mo.groups())
  12     return ("ERROR", "this should never happen")

There's a pattern in the expression above; If you make a lot of line types, exploit it. But it's easier for me explaining this to show it all to you unrolled.

Okay! We can tokenize a line.

Next up, read all the lines in a string.

Tokenize All Lines

   1 tokens = [tokenize_line(line) for line in text.split_lines()]

Uuuuuuunh-hunh!

Well, that about wraps that one up.

Of course, you're going to have to get some text on your own. Not my concern.

Isolating Paragraphs

Lines and Paragraphs

We've done all the tokenizing. Now comes the trickier part.

We don't want text to just be a series of lines. If we do that, then we can't have anything like the following:

This is some text with ''some parts
italic'' and with '''some parts that
are bold.'''

You see the problem, right? The problem here is that "some parts italic" spans across lines. If you want your nifty bold-izing, italic-izing, link-izing, whatever-itizing magic substitutions to work on paragraphs, you're going to have to transcend the boundary of the line.

That, and: We want to put paragraph markers around paragraphs, not every single line they type.

But we can't just be naive, and just work on blank lines; Because, there are headers nestled right against paragraphs, and we want to recognize, "No, this is a header, and this is a paragraph, and even thought they are sitting right next to one another, they are actually very different things."

What we're going to do is: Whever text lines follow one another in series- roll that whole thing up into one big paragraph.

Turn Text into Paragraphs

There are probably better ways to do this. Please adjust this text here, if you know one. But, this is how I did it.

   1 def text_to_paragraph(tokens):
   2 
   3     """Group TEXTs into PARAGRAPHs."""
   4