From: 7stud -- Date: 2007-09-20T14:19:25+09:00 Subject: Re: How to parse a unicode url? Dan The man wrote: > I would really like to be able to do the following. Is this even > possible? > > Thanks, > nerdytenor > > uri = URI.parse('http://www.h旦ren.com') # not a real url (that I know > of) > URI::InvalidURIError: bad URI(is not URI?): http://www.h旦ren.com > from /usr/lib/ruby/1.8/uri/common.rb:432:in `split' > from /usr/lib/ruby/1.8/uri/common.rb:481:in `parse' > from (irb):26 You can do this: require "uri" url = "http://www.h旦ren.co"旦 enc_url = URI.encode(url) puts enc_url to get this: http://www.h%C3%B6ren.co which according to wikipedia here: http://en.wikipedia.org/wiki/Percent-encoding is a legal uri. But when I do this: require "uri" url = "http://www.h旦ren.co" enc_url = URI.encode(url) puts enc_url uri = URI.parse(enc_url) I get this: http://www.h%C3%B6ren.co /usr/lib/ruby/1.8/uri/generic.rb:194:in `initialize': the scheme http does not accept registry part: www.h%C3%B6ren.co (or bad hostname?) (URI::InvalidURIError) from /usr/lib/ruby/1.8/uri/http.rb:46:in `initialize' from /usr/lib/ruby/1.8/uri/common.rb:484:in `new' from /usr/lib/ruby/1.8/uri/common.rb:484:in `parse' from r3test.rb:7 which as far as I can tell means that URI.parse() is broken. -- Posted via http://www.ruby-forum.com/.