From: duerst via ruby-core Date: 2026-04-01T06:36:19+00:00 Subject: [ruby-core:125173] [Ruby Feature#21975] Add "UTF-八" as an alias for UTF-8 encoding Issue #21975 has been updated by duerst (Martin D��rst). Thanks to @ko1 for this timely news. It looks like the current Japanese government is recently taking some steps that in some ways have felt long overdue. On December 22, 2025, they changed the Romanization used by the Government from 'Kunrei' to 'Hepburn' (see e.g. https://en.wikipedia.org/wiki/Hepburn_romanization). Kunrei reflects the structure of the Japanese syllabaries (Hiragana, Katakana), but Hepburn makes it easier for foreigners to pronounce Japanese words more or less correctly. Anyway, with respect to @ko1's proposal, I think it's a good idea to allow "UTF-���" (and probably also full-width "���������-���") as an alternative to "UTF-8" for internal Ruby use. However, it shouldn't be used on the Internet unless it is formally registered (see https://www.iana.org/assignments/character-sets/character-sets.xhtml). As the expert reviewer for that registry (rather than as a Rubyist) I would have to reject such a registration because currently, "charset"s have to be US-ASCII. Rewriting the relevant RFCs (not to speak about all the software that uses them) would be a lot of work :-). ---------------------------------------- Feature #21975: Add "UTF-���" as an alias for UTF-8 encoding https://bugs.ruby-lang.org/issues/21975#change-116902 * Author: ko1 (Koichi Sasada) * Status: Open ---------------------------------------- In Japan, legal texts must write all characters - including digits - using full-width or kanji forms. As a result, the encoding name "UTF-8" appears as "UTF-���" (��� = eight in kanji) in official government notices. Specifically, it appears in a notice issued by the Digital Agency and the Ministry of Internal Affairs and Communications (������8���������������������������������������12���), which defines character sets and encoding for local government information systems: > ������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������ Reference: https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/d12bde7e-a950-493b-987c-0f8d4bbd1b6b/66117898/20260324_laws_notice_text_02.pdf This patch adds "UTF-���" as an encoding alias for UTF-8, so that Ruby is compliant with Japanese law. ```ruby # encoding: UTF-��� p __ENCODING__ #=> # p Encoding.find("UTF-���") #=> # p "hello".encode("UTF-���") #=> "hello" p "���������������".force_encoding("UTF-���") #=> "���������������" p "���������������".encode("��� ��� ��� ��� ���") #=> "���������������" ``` -- https://bugs.ruby-lang.org/ ______________________________________________ ruby-core mailing list -- ruby-core@ml.ruby-lang.org To unsubscribe send an email to ruby-core-leave@ml.ruby-lang.org ruby-core info -- https://ml.ruby-lang.org/mailman3/lists/ruby-core.ml.ruby-lang.org/