From: Austin Ziegler Date: 2007-09-01T01:24:56+09:00 Subject: Re: Encodings of string literals; explicit codepoint escapes? On 8/31/07, Robin Stocker wrote: > Note that in Python 3, the u"" notation will be gone and all strings > will be unicode strings by default [1]. If what you want is a string of > bytes, this is what the new bytes type is for. I like this approach and > think this will help get rid of the confusion of having the same type > for character strings and byte strings. > > [1] http://www.artima.com/weblogs/viewpost.jsp?thread=157004 I've argued in the past that this is an unnecessary distinction. I still hold to this. All strings are bytes. Period. Characters are what you get out of those bytes when you view them through a particular "lens" (e.g., an encoding). It should be easy to specify strings in various encodings, including the empty (e.g., "binary") encoding. It is possible to view a standard US-ASCII (7-bit) string through the UTF-8 lens and it's valid in either one. -austin -- Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ * austin@halostatue.ca * http://www.halostatue.ca/feed/ * austin@zieglers.ca