This wiki is in the process of being archived due to lack of usage and the resources necessary to serve it — predominately to bots, crawlers, and LLM companies. Edits are discouraged.
Pages are preserved as they were at the time of archival. For current information, please visit python.org.
If a change to this archive is absolutely needed, requests can be made via the infrastructure@python.org mailing list.

Resources to help you learn how to handle Unicode in your Python programs:

General Unicode Resources

Python-Specific Resources

Standard Reference

Search the Python reference for:

Tutorials

Sample Code

Pitfalls

Supported Encodings

There is a list of standard encodings in the Python documentation.

Encodings can be registered at runtime, as well, with the codecs module.

Python2.4 supports many codecs that 2.2 and 2.3 do not, including Chinese bg2312.

Encodings are specified in files found in a directory called "encodings"; one way to find the encodings with your Python distribution is to check the contents of this directory:

>>> import encodings, os
>>> [n for n in os.listdir(os.path.dirname(encodings.__file__))
...     if n[0] != '_' and n.endswith('.py')]
['aliases.py', 'ascii.py', 'base64_codec.py', 'charmap.py', 'cp037.py', ...]

Another is to list aliases from the encodings module.

>>> import encodings
>>> from encodings import aliases
>>> aliases.aliases
{'base64': 'base64_codec', 'us_ascii': 'ascii', ...}

Contributors: LionKimbro, FredDrake, JürgenHermann.

"The Truth about Unicode in Python"

The Truth about Unicode in Python

Discussion

Here's a conversation that I had on CommunityWiki; I'd like to bring the main ideas into here.