Resources to help you learn how to handle Unicode in your Python programs:
Contents
General Unicode Resources
The Absolute Minimum Every Software Developer Must Know about Unicode - short intro to Unicode
Unicode -- Wikipedia unicode entry
Python-Specific Resources
Standard Reference
Search the Python reference for:
- unichr builtin
string handling - example: u'Hello\u0020World !'
unicodedata module
regular expressions - see the (?u) flag, and the re.UNICODE constant
exceptions - UnicodeEncodeError
Tutorials
Python and Unicode (pdf talk) He also has a brief tutorial.
Unicode for Programmers - Java and Python info
Python Unicode Objects - brief notes
Sample Code
HTMLifying and UnHTMLifying - see atomef.py
Pitfalls
The standard encodings list is for the current version of python. GB2312 (PRC Chinese,) for example, is in Python2.4, but not in Python2.2, nor Python2.3.
Supported Encodings
There is a list of standard encodings in the Python documentation.
Encodings can be registered at runtime, as well, with the codecs module.
Python2.4 supports many codecs that 2.2 and 2.3 do not, including Chinese bg2312.
Encodings are specified in files found in a directory called "encodings"; one way to find the encodings with your Python distribution is to check the contents of this directory:
>>> import encodings, os
>>> [n for n in os.listdir(os.path.dirname(encodings.__file__))
... if n[0] != '_' and n.endswith('.py')]
['aliases.py', 'ascii.py', 'base64_codec.py', 'charmap.py', 'cp037.py', ...]Another is to list aliases from the encodings module.
>>> import encodings
>>> from encodings import aliases
>>> aliases.aliases
{'base64': 'base64_codec', 'us_ascii': 'ascii', ...}Contributors: LionKimbro, FredDrake, JürgenHermann.
"The Truth about Unicode in Python"
The Truth about Unicode in Python
Discussion
Here's a conversation that I had on CommunityWiki; I'd like to bring the main ideas into here.
