From: Jacob Fugal Date: 2005-12-15T07:28:51+09:00 Subject: Re: English Ruby Home as a second class citizen On 12/14/05, Eero Saynatkari wrote:> Jacob Fugal wrote:> > I tried integrating that, but the 150 char limit meant that most -- if> > not all -- of the bodies couldn't be translated. Having translated> > titles would be a slight improvement, but not enough. Thanks for the> > direction though!>> You could split the strings into 150-char sections and loop over> them sending a new request for each. A Google translation might> not have such limits (quality of translation notwithstanding). Well, truth be told, I couldn't even get the snippet working for me,regardless of the 150ish-char limit[1]. The oddity is this: I can echoout the URI I'm going to request, past that URI into my browser andget a page with a decent translation. However, if I take that same URIand run it through wget (or Ruby's open-uri) I get a different result.And the difference isn't just in encoding, there are differences inthe returned HTML! For instance, the URIhttp://babelfish.altavista.com/tr?doit=done&intl=1&tt=urltext&trtext=Webrick+DoS%E8%84%86%E5%BC%B1%E6%80%A7%E3%81%AB%E3%81%A4%E3%81%84%E3%81%A6&lp=ja_en&btnTrTxt=Translate,relating to the Japanese string: "Webrick DoS脆弱性について". If I put thatURI in my browser, the resulting page says "About Webrick DoSvulnerability". Perfect. If I fetch that same URI via wget, the HTML returned has "Webrick DoS^@ <86> ^@ ± ^@ SECT. ^@ ^@ ^@ ^@ ^@ <84> ^@ ^@" instead (looking atit in a non-unicode vim over putty). Even stranger, the page in mybrowser has a hidden input with name="q" and value="About Webrick DoSvulnerability". The HTML returned by wget has a hidden input in thesame position (surrounding HTML identical) but with name="kls" andvalue="0". Any one have any ideas what's going on? Jacob Fugal [1] The limit imposed by bablefish is actually 150 *words*. However,as a GET request, URI raises an exception about the URL being too longfar before 150 words are actually reached.