Here's a way to make wiki-like syntaxes in Python.
Contents
Lines as Tokens
Usually, you tokenize input word-by-word.
For a wiki-like syntax, it can be easier to go line by line.
Data Flow
We start with:
= Hello, world! = This is an example. We want to demonstrate how you take text with multiple lines, and turn it into one paragraph.
...turn it into...
...and from there into...
...and finally into:
<h1> Hello, world! </h1> <p>This is an example.</p> <p>We want to demonstrate how you take text with multple lines, and turn it into one paragraph.</p>
Spec out Types of Lines
Our first task is to spec out the types of lines that exist.
Use RegularExpressions!
Let's start with just three types of lines. (It'll be clear how to add more.)
We'll have header lines, (like in wiki,) blank lines (seperating paragraphs,) and text lines, which will be anything that doesn't match the others.
Examples:
== Level 2 Heading == Here's paragraph 1. Ladee dah dee dah. Here's paragraph 2. === Level 3 Header === Look, a paragraph without a blank line above it! No problem, we can parse it!
Here are our RegularExpressions:
(Note that we store interesting information in regex groups. This is so we can get to it later on.)
This is good to start with. You can add more types now, if you want, though!
Tokenize a Line
Now, we'll teach Python how to tokenize a line.
It goes like so:
1 def tokenize_line(line):
2 """Tokenize a line, returning the token type and recognized groups."""
3 mo = header.match(line)
4 if mo is not None:
5 return ("HEADER", mo.groups())
6 mo = blank.match(line)
7 if mo is not None:
8 return ("BLANK", mo.groups())
9 mo = text.match(line)
10 if mo is not None:
11 