From: Eric Hodel Date: 2011-10-22T10:45:02+09:00 Subject: [ruby-core:40268] Re: Counseling On Oct 21, 2011, at 9:43 AM, Perry Smith wrote: > Rails, thin, rvm, almost nothing is really and truly ruby 1.9 compliant -- not really. Not when you include all the encoding problems that are still very common in my life and I assume in the life of anyone trying to use Ruby in any serious fashion. In order to generate correct documentation RDoc may need to transcode your source files into the output encoding you desire. > Is there a compile time option (or can one be added) that says "I don't care!!! -- just cram the two strings together and eat your spinach!" Mashing strings of different encodings together destroys data, but Ruby will automatically handle compatible encodings. You can concatenate a UTF-8 string with a US-ASCII string, for example. > Because, ultimately, I've yet to find anything except ruby that actually cares. I care, because I hate to see text like 'This �������encoding������� stuff is ���' on a website. (Yeah, I know fancy-quotes are an equal abomination, but it's not that hard to be aware of encodings, is it?) On Oct 21, 2011, at 5:52 PM, Perry Smith wrote: > Just as good of an alternative would be to change my default to UTF-8 instead of US-ASCII. This will not fix your problem, nor will -KU fix your problem. They'll only mask your problem. The correct solution is to add the encoding magic comment to files that matches the expected encoding of the strings they create. Blindly forcing all strings to UTF-8 will break libraries that depend on their strings being in US-ASCII encoding. See: https://github.com/rdoc/rdoc/commit/ca7651a8b9e6ef32dfa56f4ca618d9cff6ba8b74 https://github.com/rdoc/rdoc/issues/63 You will need to send patches to the library maintainers to mark their required encodings correctly, or file tickets. > My first attempt to solve this was to put a UTF-8 coding into all my ruby files. This appeared to help but upon reflection, I don't think it really did. Adding the -KU in scripts like thin's startup script helps more. In fact, I think it solves 99.99% of my problems. But when I update thin (for example) I forget to add the -KU to the script and hit errors until I add the -KU back. Let's get concrete. Show us an error you get when running thin without any modification and I can help you and the maintainer of thin (or whatever other library) find the appropriate changes to make for it to work correctly. Through our combined efforts at a concrete task we may even be able to make it easier for authors to avoid such a pitfall. > One recent saga involved memcache-client (which I've mentioned). memcache-client tries to concatenate a command, key, and Marshall'ed data. If the Marshalled data really is ASCII-8BIT, then the concatenation dies. Marshal data is always ASCII-8BIT. If memcache-client doesn't set the encoding to US-ASCII (compatible with ASCII-8BIT) then -KU will break it: $ ruby19 -e 'a = "text"; b = "\xFF"; a.force_encoding Encoding::US_ASCII; b.force_encoding Encoding::BINARY; p (a + b).encoding' #