From: "Dan0042 (Daniel DeLorme) via ruby-core" Date: 2026-04-01T14:09:36+00:00 Subject: [ruby-core:125175] [Ruby Feature#21975] Add "UTF-八" as an alias for UTF-8 encoding Issue #21975 has been updated by Dan0042 (Daniel DeLorme). duerst (Martin D��rst) wrote in #note-2: > I think it's a good idea to allow "UTF-���" (and probably also full-width "���������-���") as an alternative to "UTF-8" for internal Ruby use. Indeed, but I believe "���������������" would be a better alias here, since hyphen does indeed have a fullwidth version (U+FF0D) distinct from the prolonged sound mark ��� (U+30FC) ---------------------------------------- Feature #21975: Add "UTF-���" as an alias for UTF-8 encoding https://bugs.ruby-lang.org/issues/21975#change-116906 * Author: ko1 (Koichi Sasada) * Status: Open ---------------------------------------- In Japan, legal texts must write all characters - including digits - using full-width or kanji forms. As a result, the encoding name "UTF-8" appears as "UTF-���" (��� = eight in kanji) in official government notices. Specifically, it appears in a notice issued by the Digital Agency and the Ministry of Internal Affairs and Communications (������8���������������������������������������12���), which defines character sets and encoding for local government information systems: > ������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������������ Reference: https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/d12bde7e-a950-493b-987c-0f8d4bbd1b6b/66117898/20260324_laws_notice_text_02.pdf This patch adds "UTF-���" as an encoding alias for UTF-8, so that Ruby is compliant with Japanese law. ```ruby # encoding: UTF-��� p __ENCODING__ #=> # p Encoding.find("UTF-���") #=> # p "hello".encode("UTF-���") #=> "hello" p "���������������".force_encoding("UTF-���") #=> "���������������" p "���������������".encode("��� ��� ��� ��� ���") #=> "���������������" ``` -- https://bugs.ruby-lang.org/ ______________________________________________ ruby-core mailing list -- ruby-core@ml.ruby-lang.org To unsubscribe send an email to ruby-core-leave@ml.ruby-lang.org ruby-core info -- https://ml.ruby-lang.org/mailman3/lists/ruby-core.ml.ruby-lang.org/