From: Anthony DeRobertis Date: 2006-03-11T05:57:15+09:00 Subject: Re: unicode in ruby Austin Ziegler wrote: > Unix support for > Unicode is still in the stone ages because of the nonsense that POSIX > put on Unix ages ago. (When Unix filesystems can write UTF-16 as their > native filename format, then we're going to be much better. That will, > however, break some assumptions by really stupid programs.) Ummm, no. UTF-16 filenames would break *every* correctly-implemented UNIX program: UTF-16 allows the octect 0x00, which has always been the end-of-string marker. Personally, my file names have been in UTF-8 for quite some time now, and it works well: What exactly is this 'stone age' you refer to? UTF-8 can take multiple octets to represent a character. So can UTF-16, UTF-32, and every other variation of Unicode. Depending on content, a string in UTF-8 can consume more octects than the same string in UTF-16, or vice versa. Ah! But wait. I can see an advantage to UTF-16. With UTF-8, you don't get to have the fun of picking between big- and little-endian!