HZ+ Specification, Version 0.77 Copyright 1994 by Stephen G. Simpson 0. Availability. The latest version of this document should always be available for anonymous FTP at ftp.math.psu.edu, /pub/simpson/chinese/hzp/hzp.doc. Related documents and software are maintained at the same FTP site. 1. Introduction. By Big5 I mean the ETen implementation of the popular Big5 Chinese character coding scheme. I choose the ETen implementation because it is by far the most popular Big5 implementation. It is described precisely in my document "ETen Bitmaps and Character Codes," available for anonymous FTP at ftp.math.psu.edu, /pub/simpson/chinese/hzp/etfonts.doc. The purpose of HZ+ is to provide a standardized 7-bit representation of mixed Big5, GB, and ASCII text for convenient e-mail transmission, news posting, etc. HZ+ is compatible with HZ, the popular 7-bit representation of mixed GB and ASCII. The name HZ+ is intended to suggest an enhancement of HZ. 2. Details of the HZ+ specification. The Big5 code range is naturally partitioned into two blocks. Blocks 1 and 2 consist of the ranges 0xA140 to 0xC8FE and 0xC940 to 0xFEFE respectively, where the second byte is restricted to the ranges 0x40 to 0x7E and 0xA1 to 0xFE. Here 0x denotes hexadecimal notation. (Note that block 1 contains the more frequently used Chinese characters plus numerals, Roman and Greek and Russian and Bopomofo alphabets, symbols, numerals, Japanese kana, etc. Block 2 contains the less frequently used Chinese characters plus some additional symbols. The less frequently used Chinese characters start at 0xC940, which is also the beginning of the block 2 range.) In order to give Big5 a convenient 7-bit representation, HZ+ packs each of the block 1 and 2 ranges in order onto an initial segment of the 2-byte graphical ASCII range range, 0x2121 to 0x7D7E. Here the first byte is restricted to the range 0x21 to 0x7D and the second byte is restricted to the range 0x21 to 0x7E. Thus each character in each Big5 block is mapped to a code consisting of two graphical ASCII bytes. In addition to the the above 7-bit representation of Big5, HZ+ needs a way to distinguish between Big5 block 1, Big5 block 2, GB, and ASCII. This is accomplished by using four tilde escape sequences: ~{ to switch into GB (just as in the existing HZ standard) ~} to switch into ASCII (just as in the existing HZ standard) ~> to switch into Big5 block 1 (this is new in HZ+) ~< to switch into Big5 block 2 (this is new in HZ+) Here > and < are intended to suggest "more" and "less", referring to the fact that Big5 blocks 1 and 2 contain the more and less frequently used Chinese characters, respectively. HZ+ also includes a few additional tilde escape sequences as defined in the HZ specification, notably ~~ and ~\n. All HZ+ implementations are required to support the above mentioned escape sequences for ASCII, GB, and Big5 blocks 1 and 2. (In homogenous Big5 environments, GB support may be omitted.) In addition, HZ+ implementations may optionally support the ETen Big5 user-defined range from 0x8140 to 0xA0FE. This range is treated as another block, Big5 block 0. Just as in the case of blocks 1 and 2, block 0 is packed in order onto an initial segment of the range 0x2121 to 0x7D7E, where the first byte is restricted to the range 0x21 to 0x7D and the second byte is restricted to the range 0x21 to 0x7E. The escape sequence ~= is used to switch into block 0. 3. Advantages of HZ+. I now present a list of the advantages of HZ+. (1). The HZ+ standard is based firmly on two existing standards which are already very popular: HZ and Big5. This may enable HZ+ to win a degree of immediate, popular acceptance. (2). HZ+ is very easy to implement. I have already written some Big5<->HZ+ conversion filters. This experimental software is based on Fung F. Lee's GB<->HZ conversion filters and is available for anonymous FTP at ftp.math.psu.edu, /pub/simpson/chinese/hzp/. (3). HZ+ maps Big5 into the graphical ASCII range in an extremely simple manner, with no arbitrary features whatsoever. Blocks 1 and 2 are packed in order, with no gaps. (4). HZ+ fully respects the natural partition of the Big5 code range into blocks containing the more and less frequently used Chinese characters. This leads to further advantages, listed in (5)-(9) below. (5). HZ+ can use existing Big5 fonts essentially unchanged. (6). HZ+ uses escape sequences sparingly. Most of the time, we use only ~> to switch into Big5 block 1, plus ~{ and ~} as in HZ. (7). HZ+ is compatible not only with Big5 but also with corrupt variants of Big5 (e.g. so-called "HKU-Big5") in which the less frequently used Chinese characters begin somewhere prior to 0xC940. Thus HZ+ may help to overcome the difficulties associated with these corrupt variants. (8). HZ+ works not only with ETen Big5 but also with other vendor-specific extensions of Big5. No matter which vendor's Big5 is being used, round-trip Big5 to HZ+ to Big5 conversion is always 100 percent accurate. (9). HZ+ is compatible with an unofficial extension of ISO-2022 to include Big5, similar to the way CNS-11643 is officially included. (So far as I know, this unofficial extension of ISO-2022 to include Big5 was first implemented in version 1.0 of Mule, the multilingual Emacs.) 4. Acknowledgements. The originator of the HZ standard is Fung F. Lee, who mentions significant contributions from Edmund Lai, Tin-Fook Ngai, Ya-Gui Wei, and Ricky Yeung. HZ+ has benefitted greatly from my discussions with Yidao Cai, Jing-Shin Chang, Nelson Chin, John Delacour, Jianqing Hu, Edmund Lai, Fung F. Lee, Carlos McEvilly, Ross Paterson, Wei-Chang Shann, Ya-Gui Wei, Ricky Yeung, and other members of soft-authors@ifcss.org. Please let me know your comments on HZ+. Stephen G. Simpson / simpson@math.psu.edu / September 25, 1994