Python and TOML: New Best Friends

Python and TOML: Read, Write, and Configure with tomllib

by Geir Arne Hjelle Updated Reading time estimate 1h 6m intermediate data-structures

TOML stands for Tom’s Obvious Minimal Language. Its human-readable syntax makes TOML convenient to parse into data structures across various programming languages. In Python, you can use the built-in tomllib module to work with TOML files. TOML plays an essential role in the Python ecosystem. Many of your favorite tools rely on TOML for configuration, and you’ll use pyproject.toml when you build and distribute your own packages.

By the end of this tutorial, you’ll understand that:

  • TOML in Python refers to a minimal configuration file format that’s convenient to read and parse.
  • TOML files are primarily used for configuration, separating code from settings for flexibility.
  • pyproject.toml is crucial for package configuration and specifies the build system and dependencies.
  • Loading a TOML file in Python involves using tomli or tomllib to parse it into a dictionary.
  • tomli and tomllib differ mainly in origin, with tomllib being a standard library module in modern Python.

If you want to know more about why tomllib was added to Python, then have a look at the companion tutorial, Python 3.11 Preview: TOML and tomllib.

Use TOML as a Configuration Format

TOML is short for Tom’s Obvious Minimal Language and is humbly named after its creator, Tom Preston-Werner. It was designed expressly to be a configuration file format that should be “easy to parse into data structures in a wide variety of languages” (Source).

In this section, you’ll start thinking about configuration files and look at what TOML brings to the table.

Configurations and Configuration Files

A configuration is an important part of almost any application or system. It’ll allow you to change settings or behavior without changing the source code. Sometimes you’ll use a configuration to specify information needed to connect to another service like a database or cloud storage. Other times you’ll use configuration settings to allow your users to customize their experience with your project.

Using a configuration file for your project is a good way to separate your code from its settings. It also encourages you to be conscious about which parts of your system are genuinely configurable, giving you a tool to name magic values in your source code. For now, consider this configuration file for a hypothetical tic-tac-toe game:

Language: Config File
player_x_color = blue
player_o_color = green
board_size     = 3
server_url     = https://tictactoe.example.com/

You could potentially code this directly in your source code. However, by moving the settings into a separate file, you achieve a few things:

  • You give explicit names to values.
  • You provide these values more visibility.
  • You make it simpler to change the values.

Look more closely at your hypothetical configuration file. Those values are conceptually different. The colors are values that your framework probably supports changing. In other words, if you replaced blue with red, that would be honored without any special handling in your code. You could even consider if it’s worth exposing this configuration to your end users through your front end.

However, the board size may or may not be configurable. A tic-tac-toe game is played on a three-by-three grid. It’s not certain that your logic would still work for other board sizes. It may still make sense to keep the value in your configuration file, both to give a name to the value and to make it visible.

Finally, the project URL is usually essential when deploying your application. It’s not something that a typical user will change, but a power user may want to redeploy your game to a different server.

To be more explicit about these different use cases, you may want to add some organization to your configuration. One popular option is to separate your configuration into additional files, each dealing with a different concern. Another option is to group your configuration values somehow. For example, you can organize your hypothetical configuration file as follows:

Language: Config File
[user]
player_x_color = blue
player_o_color = green

[constant]
board_size = 3

[server]
url = https://tictactoe.example.com

The organization of the file makes the role of each configuration item clearer. You can also add comments to the configuration file with instructions to anyone thinking about making changes to it.

There are many ways for you to specify a configuration. Windows has traditionally used INI files, which resemble your configuration file from above. Unix systems have also relied on plain-text, human-readable configuration files, although the actual format varies between different services.

Over time, more and more applications have come to use well-defined formats like XML, JSON, or YAML for their configuration needs. These formats were designed as data interchange or serialization formats, usually meant for computer communication.

On the other hand, configuration files are often written or edited by humans. Many developers have gotten frustrated with JSON’s strict comma rules when updating their Visual Studio Code settings or with YAML’s nested indentations when setting up a cloud service. Despite their ubiquity, these file formats aren’t the easiest to write by hand.

TOML: Tom’s Obvious Minimal Language

TOML is a fairly new format. The first format specification, version 0.1.0, was released in 2013. From the beginning, it focused on being a minimal configuration file format that’s human-readable. According to the TOML web page, TOML’s goals are the following:

TOML aims to be a minimal configuration file format that’s easy to read due to obvious semantics. TOML is designed to map unambiguously to a hash table. TOML should be easy to parse into data structures in a wide variety of languages. (Source, emphasis added)

As you work through this tutorial, you’ll see how well TOML hits these targets. It’s clear, though, that TOML has gotten quite popular over its short life span. More and more Python tools, including Black, pytest, mypy, and isort, use TOML for their configuration. TOML parsers are available for most popular programming languages.

Recall your configuration from the previous subsection. One way to express it in TOML is the following:

Language: TOML
[user]
player_x.color = "blue"
player_o.color = "green"

[constant]
board_size = 3

[server]
url = "https://tictactoe.example.com"

You’ll learn more about the details of the TOML format in the next section. For now, just try to read and parse the information yourself. Note that it’s not much different from earlier. The biggest change is the addition of quotation marks (") in some of the values.

TOML’s syntax is inspired by traditional configuration files. Its one major advantage over Windows INI files and Unix configuration files is that TOML has a specification that spells out precisely what’s allowed in a TOML document and how different values should be interpreted. The specification is stable and mature after reaching version 1.0.0 in early 2021.

In contrast, the INI format doesn’t have a formal specification. Instead, there are many variants and dialects, most of them defined by an implementation. Python comes bundled with support for reading INI files in the standard library. While ConfigParser is quite lenient, it doesn’t support all kinds of INI files.

Another difference between TOML and many traditional formats is that TOML values have types. In the example above, "blue" is interpreted as a string, while 3 is a number. One potential criticism of TOML is that humans writing TOML need to be aware of types. In simpler formats, that responsibility lies with the programmer parsing the configuration.

TOML is not meant to be a data serialization format like JSON or YAML. In other words, you shouldn’t try to store general data in TOML to recover it later. TOML is restrictive in a few aspects:

  • All keys are interpreted as strings. You can’t easily use, say, a number as a key.
  • TOML has no null type.
  • Some whitespace is important, which makes it less efficient to compress the size of TOML documents.

Even though TOML is a good hammer, not all data files are nails. You should primarily use TOML for configurations.

TOML Schema Validation

You’ll dive deeper into TOML syntax in the next section. There you’ll learn about some of the syntax requirements of TOML files. However, in practice, a given TOML file may also come with some non-syntactical requirements.

These are schema requirements. For example, your tic-tac-toe application may require that the configuration file contain the server URL. On the other hand, the player colors may be optional because the application defines a default color.

Currently, TOML doesn’t include a schema language that can specify required and optional fields in a TOML document. Several proposals exist, although it’s not clear if any of them will be accepted anytime soon.

In simple applications, you can validate your TOML configuration manually. For example, you can use structural pattern matching, which was introduced in Python 3.10. Assume that you’ve parsed the configuration into Python and named it config. You can then check its structure as follows:

Language: Python
match config:
    case {
        "user": {"player_x": {"color": str()}, "player_o": {"color": str()}},
        "constant": {"board_size": int()},
        "server": {"url": str()},
    }:
        pass
    case _:
        raise ValueError(f"invalid configuration: {config}")

The first case statement spells out the structure that you expect. If config matches, then you use pass to continue your code. Otherwise, you raise an error.

This approach may not scale well if your TOML document is more complicated. You also need to do more work if you want to provide good error messages. A better alternative is to use pydantic which utilizes type annotations to do data validation at runtime. One advantage of pydantic is that it has precise and helpful error messages built in.

There are also tools that take advantage of the existing schema validations that exist for formats like JSON. For example, Taplo is a TOML tool kit that can validate TOML documents against JSON schemas. Taplo is also available for Visual Studio Code, bundled into the Even Better TOML extension.

You won’t worry about schema validation in the rest of this tutorial. Instead, you’ll get more familiar with the TOML syntax and see all the different data types that are available to you. Later on, you’ll see examples of how you can interact with TOML in Python, and you’ll explore some of the use cases where TOML is a good fit.

Get to Know TOML: Key-Value Pairs

TOML is built around key-value pairs that map nicely to hash table data structures. TOML values have different types. Each value must have one of the following types:

Additionally, you can use tables and arrays of tables as collections that organize several key-value pairs. You’ll learn more about all of these—and how you can specify them in TOML—in the rest of this section.

You’ll see all the different elements of TOML in this tutorial. However, some details and edge cases will be glossed over. Check out the documentation if you’re interested in the fine print.

As noted, key-value pairs are your basic building blocks in a TOML document. You specify them with a <key> = <value> syntax, where the key is separated from the value with an equal sign. The following is a valid TOML document with one key-value pair:

Language: TOML
greeting = "Hello, TOML!"

In this example, greeting is the key, while "Hello, TOML!" is the value. Values have types. In this example, the value is a text string. You’ll learn about the different value types in the following subsections.

Keys are always interpreted as strings, even if quotation marks don’t surround them. Consider the following example:

Language: TOML
greeting = "Hello, TOML!"
42 = "Life, the universe, and everything"

Here, 42 is a valid key, but it’s interpreted as a string, not a number. Usually, you want to use bare keys. These are keys that consist only of ASCII letters and numbers as well as underscores and dashes. All such keys can be written without quotation marks, as in the examples above.

TOML documents must be encoded in UTF-8 Unicode. This gives you great flexibility when representing your values. Despite the restrictions on bare keys, you can also use Unicode when spelling out your keys. This comes at a cost, though. To use Unicode keys, you must add quotation marks around them:

Language: TOML
"realpython.com" = "Real Python"
"blåbærsyltetøy" = "blueberry jam"
"Tom Preston-Werner" = "creator"

All these keys contain characters that aren’t allowed in bare keys: the dot (.), Norwegian characters (å, æ, and ø), and a space. You’re allowed to use quotation marks around any key, but in general, you want to stick to bare keys that don’t use or require quotation marks.

Dots (.) play a special role in TOML keys. You can use dots in unquoted keys, but in that case, they’ll trigger grouping by splitting the dotted key at each dot. Consider the following example:

Language: TOML
player_x.symbol = "X"
player_x.color = "purple"

Here, you specify two dotted keys. Since they both start with player_x, the keys symbol and color will be grouped together inside a section named player_x. You’ll learn more about dotted keys when you start exploring tables.

Next, turn your attention to the values. In the next section, you’ll learn about the most basic data types in TOML.

Strings, Numbers, and Booleans

TOML uses familiar syntax for the basic data types. Coming from Python, you’ll recognize strings, integers, floats, and Booleans:

Language: TOML
string = "Text with quotes"
integer = 42
float = 3.11
boolean = true

The immediate difference between TOML and Python is that TOML’s Boolean values are lowercase: true and false.

A TOML string should typically use double quotation marks ("). Inside strings, you can escape special characters with the help of backslashes: "\u03c0 is less than four". Here, \u03c0 denotes the Unicode character with codepoint U+03c0, which happens to be the Greek letter π. This string will be interpreted as "π is less than four".

You can also specify TOML strings using single quotation marks ('). Single-quoted strings are called literal strings and behave similarly to raw strings in Python. Nothing is escaped and interpreted in a literal string, so '\u03c0 is the Unicode codepoint of π' starts with the literal \u03c0 characters.

Finally, TOML strings can also be specified using triple quotation marks (""" or '''). Triple-quoted strings allow you to write a string over multiple lines, similar to Python multiline strings:

Language: TOML
partly_zen = """
Flat is better than nested.
Sparse is better than dense.
"""

Control characters, including literal newlines, aren’t allowed in basic strings. You can use \n to represent a newline inside a basic string, though. You must use a multiline string if you want to format your strings over several lines. You can also use triple-quoted literal strings. In addition to being multiline, these are the only way to include a single quotation mark inside a literal string: '''Use '\u03c0' to represent π'''.

Numbers in TOML are either integers or floating-point numbers. Integers represent whole numbers and are specified as plain, numeric characters. As in Python, you can use underscores to enhance readability:

Language: TOML
number = 42
negative = -8
large = 60_481_729

Floating-point numbers represent decimal numbers and include an integer part, a dot representing the decimal point, and a fractional part. Floats can use scientific notation to represent very small or very large numbers. TOML also supports special float values like infinity and not a number (NaN):

Language: TOML
number = 3.11
googol = 1e100
mole = 6.22e23
negative_infinity = -inf
not_a_number = nan

Note that the TOML specification requires that integers at least are represented as 64-bit signed integers. Python handles arbitrarily large integers, but only integers with up to about 19 digits are guaranteed to work on all TOML implementations.

Non-negative integer values may also be represented as hexadecimal, octal, or binary values by using a 0x, 0o, or 0b prefix, respectively. For example, 0xffff00 is a hexadecimal representation, and 0b00101010 is a binary representation.

Boolean values are represented as true and false. These must be lowercase.

TOML also includes several time and date types. Before you explore those, though, you’ll see how you can use tables to organize and structure your key-value pairs.

Tables

You’ve learned that a TOML document consists of one or more key-value pairs. When represented in a programming language, these should be stored in a hash table data structure. In Python, that would be a dictionary or another dictionary-like data structure. To organize your key-value pairs, you can use tables.

TOML supports three different ways of specifying tables. You’ll see examples of each of these shortly. The end results will be the same, independently of how you represent your tables. Still, the different tables do have slightly different use cases:

  • Use regular tables with headers in most cases.
  • Use dotted key tables when you need to specify a few key-value pairs that are closely tied to their parent table.
  • Use inline tables only for very small tables with up to three key-value pairs, where the data makes up a clearly defined entity.

The different table representations are mostly interchangeable. You should default to regular tables and only switch to dotted key tables or inline tables if you think it improves your configuration’s readability or clarifies your intent.

How do these different table types look in practice? Start with the regular tables. They’re defined by adding a table header above your key-value pairs. A header is a key without a value, wrapped inside square brackets ([]). The following example, which you encountered earlier, defines three tables:

Language: TOML
[user]
player_x.color = "blue"
player_o.color = "green"

[constant]
board_size = 3

[server]
url = "https://tictactoe.example.com"

The three highlighted lines are the table headers. They specify three tables, named user, constant, and server, respectively. The contents or value of a table are all the key-value pairs listed below the header and above the next header. For example, constant and server contain one nested key-value pair each.

You can also find dotted key tables in the configuration above. Inside user, you have the following:

Language: TOML
[user]
player_x.color = "blue"
player_o.color = "green"

The period, or dot (.), in the keys creates a table named by the part of the key before the dot. You can also represent the same part of the configuration by nesting regular tables:

Language: TOML
[user]

    [user.player_x]
    color = "blue"

    [user.player_o]
    color = "green"

Indentation isn’t important in TOML. You use it here to represent the nesting of the tables. You can see that the user table contains two sub-tables, player_x and player_o. Each of those sub-tables contains one key-value pair.

Note that you need to use a dotted key in the header of nested tables and name all the intermediate tables. This makes TOML header specifications quite verbose. In a similar specification in, for example, JSON or YAML, you’d only specify the sub-table name without repeating the names of the outer tables. At the same time, this makes TOML very explicit, and it’s harder to lose track inside deeply nested structures.

Now, you’ll expand a bit on the user table by including a label or symbol for each player as well. You’ll represent this table in three different forms, first using only regular tables, then using dotted key tables, and finally using inline tables. You haven’t seen the latter yet, so this will be an introduction to inline tables and how those are represented.

Start with nested, regular tables:

Language: TOML
[user]

    [user.player_x]
    symbol = "X"
    color = "blue"

    [user.player_o]
    symbol = "O"
    color = "green"

This representation makes it very clear that you have two different player tables. You don’t need to explicitly define tables that only contain sub-tables and not any regular keys. In the previous example, you could remove the line [user].

Compare the nested tables with the dotted key configuration:

Language: TOML
[user]
player_x.symbol = "X"
player_x.color = "blue"
player_o.symbol = "O"
player_o.color = "green"

This is shorter and more concise than the nested tables above. However, the structure is now less clear, and you need to expend some effort on parsing the keys before you realize that there are two player tables nested within user. The dotted key tables are more beneficial when you have a few nested tables with one key each, like in the earlier example with only color sub-keys.

Next, you’ll represent user with inline tables:

Language: TOML
[user]
player_x = { symbol = "X", color = "blue" }
player_o = { symbol = "O", color = "green" }

An inline table is defined with curly braces ({}) wrapped around comma-separated key-value pairs. In this example, the inline table brings a nice balance of readability and compactness, as the grouping of the player tables becomes clear.

Still, you should use inline tables sparingly, and mostly in cases like this where a table represents a small and well-defined entity, like a player. Inline tables are intentionally limited compared to regular tables. In particular, an inline table must be written on one line in the TOML file, and you can’t use conveniences like trailing commas.

To wrap up your tour of tables in TOML, you’ll have a brief look at a few minor points. In general, you can define your tables in any order, and you should strive to order your configuration in a way that makes sense to your users.

A TOML document is represented by a nameless root table that contains all other tables and key-value pairs. Key-value pairs that you write at the top of your TOML configuration, before any table header, are stored directly in the root table:

Language: TOML
title = "Tic-Tac-Toe"

[constant]
board_size = 3

In this example, title is a key in the root table, constant is a table that’s nested within the root table, and board_size is a key in the constant table.

Be aware that a table includes all key-value pairs written between its header and the next table header. In practice, this means that you must define nested sub-tables below the key-value pairs belonging to the table. Consider this document:

Language: TOML
[user]

    [user.player_x]
    color = "blue"

    [user.player_o]
    color = "green"

background_color = "white"

The indentation suggests that background_color is supposed to be a key in the user table. However, TOML ignores the indentation and checks only the table headers. In this example, background_color is part of the user.player_o table. To correct this, background_color should be defined before the nested tables:

Language: TOML
[user]
background_color = "white"

    [user.player_x]
    color = "blue"

    [user.player_o]
    color = "green"

In this case, background_color is a key in user, as intended. If you use a dotted key table, then you can more freely use any order of the keys:

Language: TOML
[user]
player_x.color = "blue"
player_o.color = "green"
background_color = "white"

Now there are no explicit table headers except for [user], so background_color will be a key in the user table.

You’ve learned about the basic data types in TOML and how you can use tables to organize your data. In the next subsections, you’ll see the final data types that you can play with in your TOML documents.

Times and Dates

TOML supports defining times and dates directly in your documents. You can choose between four different representations, each with its own specific use case:

  • An offset date-time is a timestamp with time zone information, representing a specific instant in time.
  • A local date-time is a timestamp without time zone information.
  • A local date is a date without any time zone information. You typically use this to represent a full day.
  • A local time is a time with any date or time zone information. You use a local time to represent a time of day.

TOML bases its representation of times and dates on RFC 3339. This document defines a time and date format that is commonly used to represent timestamps on the Internet. A fully defined timestamp would look something like this: 2021-01-12T01:23:45.654321+01:00. The timestamp is composed of several fields, split by different separators:

Field Example Details
Year 2021
Month 01 Two digits from 01 (January) to 12 (December)
Day 12 Two digits, zero padded when below ten
Hour 01 Two digits from 00 to 23
Minute 23 Two digits from 00 to 59
Second 45 Two digits from 00 to 59
Microsecond 654321 Six digits from 000000 to 999999
Offset +01:00 Time zone as offset from UTC, with Z representing UTC

An offset date-time is a timestamp that includes the offset information. A local date-time is a timestamp that doesn’t include this. Local timestamps are also called naive timestamps.

In TOML, the microsecond field is optional for all date-time and time types. You’re also allowed to replace the T that separates the date and time with a space. Here, you can see examples of each of the timestamp-related types:

Language: TOML
offset_date-time     = 2021-01-12 01:23:45+01:00
offset_date-time_utc = 2021-01-12 00:23:45Z
local_date-time      = 2021-01-12 01:23:45
local_date           = 2021-01-12
local_time           = 01:23:45
local_time_with_us   = 01:23:45.654321

Note that you can’t wrap the timestamp values within quotation marks, since that would turn them into text strings.

These different time and date types give you a fair amount of flexibility. If you have use cases that aren’t covered by these—for example, if you want to specify a time interval like 1 day—then you can use strings and use your application to properly process those.

The final data type that TOML supports is arrays. These allow you to combine several other values in a list. Read on to learn more.

Arrays

TOML arrays represent an ordered list of values. You specify them using square brackets ([]), so that they resemble Python’s lists:

Language: TOML
packages = ["tomllib", "tomli", "tomli_w", "tomlkit"]

In this example, the value of packages is an array containing four string elements: "tomllib", "tomli", "tomli_w", and "tomlkit".

You can use any TOML data type, including other arrays, inside arrays, and one array can contain different data types. You’re allowed to specify an array over several lines, and you can use a trailing comma after the last element in the array. All the following examples are valid TOML arrays:

Language: TOML
potpourri = ["flower", 1749, { symbol = "X", color = "blue" }, 1994-02-14]
skiers = ["Thomas", "Bjørn", "Mika"]
players = [
    { symbol = "X", color = "blue", ai = true },
    { symbol = "O", color = "green", ai = false },
]

This defines three arrays. potpourri is an array with four elements with different data types, while skiers is an array containing three strings. The final array, players, adapts your earlier example to represent two inline tables as elements in an array. Note that players is defined over four lines and that there’s an optional comma after the last inline table.

The last example shows one way that you can create arrays of tables. You can put inline tables inside square brackets. However, as you saw previously, inline tables don’t scale well. If you want to represent an array of tables where the tables are bigger, you should use a different syntax.

In general, you should express an array of tables by writing table headers inside double square brackets ([[]]). The syntax isn’t necessarily pretty, but it’s quite effective. You can represent players from the example above as follows:

Language: TOML
[[players]]
symbol = "X"
color = "blue"
ai = true

[[players]]
symbol = "O"
color = "green"
ai = false

This array of tables is equivalent to the array of inline tables that you wrote above. The double square brackets define an array of tables instead of a regular table. You need to repeat the array name for each nested table inside the array.

For a more extensive example, consider the following excerpt from a TOML document listing questions for a quiz application:

Language: TOML
[python]
label = "Python"

[[python.questions]]
question = "Which built-in function can get information from the user"
answers = ["input"]
alternatives = ["get", "print", "write"]

[[python.questions]]
question = "What's the purpose of the built-in zip() function"
answers = ["To iterate over two or more sequences at the same time"]
alternatives = [
    "To combine several strings into one",
    "To compress several files into one archive",
    "To get information from the user",
]

In this example, the python table has two keys, label and questions. The value of questions is an array of tables with two elements. Each element is a table with three keys: question, answers, and alternatives.

You’ve now seen all the data types that TOML has to offer. In addition to simple data types like strings, numbers, Booleans, times, and dates, you can also combine and organize your keys and values with tables and arrays. There are some details and edge cases that you’ve glossed over in this overview. You can learn all the details in the TOML specification.

In the upcoming sections, you’ll get more practical as you learn how you can use TOML in Python. You’ll learn about how you can read and write TOML documents and explore how you can organize your applications to use configuration files effectively.

Load TOML With Python

It’s time to get your hands dirty. In this section, you’ll fire up your Python interpreter and load TOML documents into Python. You’ve seen how the main use case for the TOML format is configuration files. These are often written by hand, so in this section, you’ll look at how you can read such configuration files with Python and work with them inside your project.

Read TOML Documents With tomli and tomllib

Since the TOML specification first arrived in 2013, there have been several packages available for working with the format. Over time, some of these packages have become unmaintained. Some libraries that used to be popular are no longer compliant with the latest version of TOML.

In this section, you’ll work with a package called tomli and its sibling tomllib. These are great libraries when you only want to load a TOML document into Python. You’ll also explore tomlkit in a future section. That package brings more advanced functionality to the table, and opens up some new use cases for you.

It’s time to explore how you can read TOML files. Start by creating the following TOML file and saving it as tic_tac_toe.toml:

Language: TOML Filename: tic_tac_toe.toml
[user]
player_x.color = "blue"
player_o.color = "green"

[constant]
board_size = 3

[server]
url = "https://tictactoe.example.com"

This is the same configuration that you worked with in the previous section. Next, use pip to install tomli into your virtual environment