X J D I C V1.1 INTRODUCTION XJDIC is an electronic Japanese-English dictionary program designed to operate in the X11 window environment. In particular, it must run in an "xterm" environment which has Japanese language support such as provided by "kterm" or internationalized xterm, aixterm, etc. It is based on JDIC and JREADER which were developed to run under MS-DOS on IBM PCs or clones. XJDIC functions as: (a) an English to Japanese dictionary (eiwa jiten), searching for and displaying entries for key-words entered in English; (b) a Japanese to English dictionary (waei jiten), searching for and displaying entries for keywords or phrases entered in Japanese (kanji, hiragana or katakana); (c) a Japanese-English Character dictionary (kanei jiten), capable of selecting kanji characters by JIS code, radical, stroke count, Nelson Index number or reading, and displaying compounds containing that kanji. XJDIC is typically run in a window of its own. The user can then use it as a free-standing on-line dictionary. It can also be used as an accessory when reading or writing text in another window (e.g. reading the "fj" Japanese news groups.) Strings of text, either English or Japanese, can be moved to and from XJDIC using X11's mouse "cut-and-paste" operations. The source code, documentation and dictionaries of XJDIC are hereby released to the "public domain". All usage of this program is at the user's risk, and there is no warranty on its performance. Copies may be distributed by any means, but at no charge apart from that of the distribution medium. No modifications may be made to the software without the author's permission, except where minor changes are required to install the software on a particular platform. The software is not to be incorporated into any other software product without the permission of the author. All the Japanese displayed by XJDIC is in kana and kanji, so if you cannot read at least hiragana and katakana, this is not the program for you. The author has no intention whatsover of producing a version using romanized Japanese. OPERATION The invocation of XJDIC is: xjdic The command line options are: -l nn the number of lines of dictionary display produced at a time. The default is set by the MAXLINES #define (currently 20). After this number of lines are displayed, there is a prompt giving the user the choice of having more lines displayed, or abandoning the search. -d dictionary-path_and_filename the path and file-name of the Japanese-English dictionary to use. If not present, the dictionary file "EDICT" will be used, along with the index file "EDICT.XJDX". These must be either in the current directory, or the directory specified in the XJDIC environment variable. The dictionary can also be specified in the .xjdicrc file (see below). -k kanji_dictionary-path_and_filename the path and file-name of the Kanji dictionary to use. If not present, the dictionary file "KANJIDIC" will be used, along with the index file "KANJIDIC.XJDX". These must be either in the current directory, or the directory specified in the XJDIC environment variable. The dictionary can also be specified in the .xjdicrc file (see below). -j japanese_output_code_type (j, e or s) XJDIC uses "New-JIS" codes as its default output method. This is quite acceptable if you are running under kterm. Some other environments which are internationalized (e.g. aixterm) can only handle EUC or Shift-JIS codes. XJDIC can be made to output in these codes by the "-j e" or "-j s" command-line options. "-j j" sets it to New-JIS (the default). -v To disable the verb deinflection function. -h this option results in the display of a simple summary of command-line options. ENTERING SEARCH KEYS (a) Japanese-English Dictionary XJDIC operates in two modes: Japanese-English Dictionary, and Kanji Dictionary. In the case of the Japanese-English Dictionary, search keys are entered in response to the "XJDIC SEARCH KEY:" prompt. Search keys can be either in English (typically entered from the keyboard) or kana and/or kanji (entered via a "front-end" program such as kinput, entered using XJDIC's internal romaji/kana converter (see below), or cut from another window using X11 mouse operations.) To invoke the romaji/kana converter, begin a search key with either "@" for hiragana or "#" for katakana. Then as you type the key, it will be converted to the selected kana. (See below for details of the romaji-to-kana conversion.) A multi-line display will be produced of all the dictionary entries which contain matches with the search key. The display format is: matched_word [kanji] (yomikata) english_1, english_2, etc where "matched_word" is either English or Japanese, depending on the search key. If the search key was kana (i.e. the match was on the yomikata), the separate yomikata display is omitted. A line is only displayed once per search, regardless of the number of matches which occur within it. If the search resulted in more entries than will fit on a screen, a further prompt occurs at the bottom of the screen giving you the option of requesting the next screen-full. Once all the matches on a key are exhausted, the keyword is shortened by one character, and the display is continued. The matching of kana keys is insensitive to whether they are in katakana or hiragana, however note that the convention for long vowels differs between Japanese words and gairaigo. Matching of English keywords is insensitive to case. The display is in "dictionary" order for the words matched, i.e. alphabetical for the English search, and JIS code order for the Japanese search. JIS order is very close to the "gojuuon" kana order used in Japanese dictionaries except that it separates the syllables with the nigori and maru diacritic marks. If the word being used to search the dictionary consists of a kanji followed by two or more hiragana, the kana is matched against common verb and adjective inflections. If a match is found, the search is initially made for the plain or "dictionary form" of the word. The possible combinations of inflections or conjugations is taken from the VCONJ file. It is possible to set up a number of "filters" which either restrict the display to dictionary entries which contain certain strings of characters, or suppress the display of entries with certain strings. This feature is useful if the user wants to avoid the large number of proper nouns in the dictionary. See the section FILTERS below for details of how to set up such filter strings. Individual filters can be activated or deactivated using the ";" command. (b) Kanji Dictionary XJDIC has the capability to select individual kanji characters by a variety of techniques, and to display information about that character. The character can then be "cut" into the main dictionary search to display all dictionary entries starting with or containing that particular character. The Kanji Dictionary used by XJDIC is the KANJIDIC file, some details of which are included below. The search of the Kanji Dictionary is triggered by entering "\", which causes the "KANJI LOOKUP TYPE:" prompt to appear. The kanji lookup types are specified by entering a further single character: J - by its "JIS" code. This is the standard 4-digit hexadecimal code used to identify each Japanese character. Alternatively the 4-digit Kuten code may be entered preceded by a "-", and the 4-digit Shift-JIS code may be entered preceded by a "s". C - by one of the identifying codes within the Kanji Dictionary. The codes presently in KANJIDIC are: Nnnnn - the "Nelson" index number. This refers the kanji index numbers used by the late Professor Andrew N. Nelson in his famous "Japanese-English Character Dictionary", published by Tuttle. Over 5000 kanji in XJDIC's files have Nelson numbers. Snn - the stroke count. A display of all the characters with that stroke count is produced. As above, the desired character can be selected for the display of its compounds. Bnn - the primary radical (Bushu). The Bushu numbers used by XJDIC are those from Nelson (as depicted inside the front cover of Nelson's "Japanese English Character Dictionary"). To use this method, you will either need to have a copy of the Nelson radical table with you, or be prepared to use the "R" command to display the radical numbers. (For the Bushu search, you will be asked for a stroke count, and only the kanji with that bushu/stroke combination will be displayed. If you want a display of all the kanji for that bushu number, enter a stroke count of 0.) Cnn - the "classical" radical, where this is listed. Hnnnn - the Halpern number, which is the index in Jack Halpern's Character Dictionary. About 3000 characters in KANJIDIC have Halpern numbers assigned. Pn-n-n - the SK*P code used by Halpern and others to find kanji. Unnnn - the Unicode code for the kanji. (Other codes may be added, so check the kanjidic.doc file.) K - by the reading (or yomikata) of a character. Both on and kun readings are used for this search. A display of all kanji with that particular yomikata is produced, and the desired character can be selected using the mouse. A kanji can also be entered if its characteristics are to be examined. As with the other mode of usage, an automatic romaji/kana conversion can be invoked by beginning the key with either "@" or "#". M - by its English "meaning". R - initiates a display of all the Bushu along with their numbers. Once a kanji has been identified by any of the techniques described above, XJDIC displays the following information about it: KANJI character KUTEN code in decimal ([nnnn]) JIS code in hexadecimal [Kuten-code:Shift-JIS-code] Unicode code in hexadecimal (Uxxxx) Nelson No (Nnnnn) Bushu No (Bnn, and Cnn if the "classical radical differs) Stroke Count (Snn) Halpern No (Hnnnn) SK*P code (Pn-n-n) Grade (G1 to G6 if taught in those grades, G8 for Joyo kanji, and G9 for the supplementary Jinmeiyou kanji.) "on" readings of the character in katakana "kun" readings in hiragana meanings ascribed to the kanji (taken from the popular kanji dictionaries, such as Nelson, Heisig, Halpern and Spahn & Hadamitsky.) Some other indices are included here, such as S&H. At this stage, the user can request a display of all the compounds containing that character by using the mouse to select the kanji and entering it as the search key for a main dictionary search. XJDIC has two modes for displaying compounds containing a particular sequence of one or more kanji. Either the display is restricted to only those compounds which begin with the sequence, or all compounds containing the sequence can be displayed. When XJDIC loads it is in the more limited mode, however the mode can be toggled using the "/" key. The verb deinflection function can be toggled on and off with the ":" key. EXITING To exit XJDIC, type "bye". Ctrl-C will work, but may leave echo turned off. ON-LINE HELP Basic operating information can be obtained by typing "?". A summary of the command-line options can be obtained by invoking XJDIC with the "-h" option. ROMAJI-TO-KANA CONVERSION To enter a search key in kana, initiate it with either "@" (hiragana) or "#" (katakana), then type it in romaji and it will be converted to kana as you type. The romaji->kana translation is almost identical to that used in "front-end-processors" such as kinput, and MOKE and other Japanese word processors, i.e. for a small "tsu" you can type either a double consonant, e.g. "shippai", or "t-", e.g. shit-pai, and for "n" you can type n' if necessary (e.g. as in "hon'ya"). Most of the time just typing ordinary Hepburn or kunrei romaji works. Note that the romaji must follow the kana style for long vowels. Tokyo must be toukyou, NOT tookyoo. The actual romaji to kana conversions are specified in the file "romkana.cnv". This file provides the capability for inputing all the kana characters. It may, however, be edited if you want to add extra mappings, e.g. some of the modern katakana mora constructions. JAPANESE CODES Kterm can operate with the JIS, EUC or Shift-JIS code sets (as specified by the command-line, or by Ctrl-middle_mouse_button). XJDIC uses EUC internally and displays in (new) JIS, EUC or Shift-JIS. New-JIS is the default, and the others can be specified by command-line option or in the .xjdicrc file. It will accept input in any code type. In fact, XJDIC's operation is smoothest in JIS mode. This is because it detects the closing "shift-out" sequence which is present in this code, and immediately invokes the dictionary search. Thus it is possible to cut a string from a document being read, and initiate a dictionary scan, solely by using the mouse. (Entering a kana/kanji string in response to almost all of XJDIC's prompts will result in a dictionary search on that string.) DICTIONARIES XJDIC depends for its performance on a pair of dictionary files; a Japanese <-> English dictionary and a Kanji dictionary, and a file of radicals. It has been designed to work with the EDICT dictionary, which is the author's extension of MOKE's EDICT, and the KANJIDIC character dictionary file, compiled by the author from various sources. EDICT has now over 45,000 entries, while KANJIDIC has an entry for each of the JIS1 and JIS2 kanji. (The file of radicals is RADICALS.TM, compiled by Theresa Martin for the earlier JDIC program; the ROMKANA.CNV file and VCONJ file were compiled by the author, the former partly from one of the .hlp files in MOKE.) The format each entry of EDICT is: Kanji [kana] /english_1/english_2/..../ or kana /english_1/english_2/..../ For full information about EDICT, see the edict.doc file. KANJIDIC is a Public Domain compilation of information about each of the JIS1 and JIS2 kanji. It has the format: Kanji hex_JIS_code Unnnn Bnnn Snn on_reading(s) kun_reading(s) {meaning(s)} where N, H, B, S and G flag the Nelson number, Halpern number, Bushu number, stroke count and (school) grade respectively. The Pn-n-n codes are Halpern's SK*P codes for finding kanji. On readings are in katakana and kun readings in hiragana. For full information about this file, see the kanjidic.doc file. FILTERS Up to 10 sets of filters can be specified using "filt" lines in .xjdicrc. These allow the option of only displaying dictionary entries which contain, or do not contain certain text strings. There are three types of filters: (a) inclusion filters (Type 0). If one of these is active, only those entries which contain one of the specified text strings will be displayed. (b) exclusion filters (Type 1 & 2). If one or more of these is active, lines which contain the specified text strings will not be displayed. In the case of Type 2 filters, they only function if the dictionary entry has just ONE English entry. The format of the filter lines in xjdicrc is; filt f t on|off "filter name" string_1 string_2 .... where: f - the filter number (0 to 9) t - the filter type (0, 1 or 2) on|off - sets the initial state of the filter "filter name" - the " " delimited name of the filter, up to 50 characters long string_n - the space-separated strings which are to be matched as part of the filter operation. Up to 10 strings per filter, each up to 10 characters. Here are some sample filter entries: filt 0 2 "Suppress proper name entries" (pl, (pn pn) pl) [This filter, if activated, would prevent the display of entries which only relate to proper names.] filt 1 0 "Show only place names" (pl, pl) [This filter would enable xjdic to be used as a place-name dictionary.] filt 2 1 "Suppress colloquialisms" (col) (col.) The ";" command initiates a dialogue in which individual filters to be activated or deactivated. Use caution when setting up filters, as their operation may make xjdic examine many dictionary entries, resulting in a slow display of information. Note that once a filter condition has been met for a dictionary entry, no further testing is carried out for that entry. LOGGING Users of the author's JREADER program will notice that XJDIC has no logging facilities. This is because the X11 environment makes logging possible via another window running an editor such as jstevie or nemacs against a logfile. JREADER has a facility to look up kanji compounds which are not in EDICT in MOKE's Kanji->Kana file (WSKTOK.DAT), and to log the match for the later addition of a translation. If you wish to have this capability in XJDIC, obtain the file WSKTOK.DAT and either incorporate it into EDICT, or run a separate window with XJDIC running with WSKTOK as the dictionary. WSKTOK.DAT contains many entries (kanji plus kana) not currently in EDICT as well. (See XJDIC.INSTALL.) If a match is found against one of these entries, you can use the mouse to cut it into the logfile. CONTROL FILE XJDIC uses a control file called ".xjdicrc". XJDIC will look for this file in the directory identified by the XJDIC environment variable, in the HOME directory,and finally in the current (PWD) directory. XJDIC will function quite well without a .xjdicrc file, but it is a useful way of setting various options, and it is the only way to set up search filters. .xjdicrc contains lines of text which consist of: line_type The line_types are: filt set up filter details (see the FILTERS section) omode e|j|s set the screen output codes to EUC, JIS or Shift-JIS dicfile path_name alternative dictionary name kdicfile path_name alternative kanji dictionary name romfile path_name alternative romaji conversion file verbfile path_name alternative conjugation file radfile path_name alternative radical/bushu no. file lines nn set the number of lines per display radlines nn set the number of radicals/per line in the \R display jverb on|off enable or disable the verb deinflection function Note that some of these are also command-line options. If both are used, the command-line request takes precedence. OTHER FILES Apart from the .xjdicrc control file, XJDIC requires three other file: radicals.tm - the list of bushu numbers and descriptive kanji, originally prepared by Theresa Martin for JDIC. romkana.cnv - the list of romaji to kana mappings used in the input conversion routines. vconj - the verb/adjective inflections used to identify the dictionary forms of words prior to lookup. All of these files are in the public domain, and can be modified by the user. Exercise extreme caution if you do change these files, particularly if you change the order of entries. INSTALLATION See the document XJDIC.INSTALL for information on compiling the XJDIC program and setting up the dictionary files and index files. Make sure you have the XJDIC executable in your path, and that the dictionary, index and radical files are in your current directory, in the directory specified by the XJDIC environment variable, or in the places specified by the .xjdicrc file. AUTHOR'S COMMENT XJDIC is port/rework of my earlier JDIC/JREADER programs which were written for PCs or clones. Most of the code came from JREADER. In producing XJDIC I have relied heavily on the Japanese environment such as is provided by kterm, with the result that XJDIC is smaller than either JDIC or JREADER. Also I took a different approach with the kanji dictionary. Whereas in JDIC/JREADER I use a compressed kanji dictionary file with separate index files for Nelson number, stroke count, yomikata, etc. (originally devised by Stephen Chung for his JWP Word Processor package), in XJDIC I have used the same indexing and lookup approach as with the main dictionary. XJDIC's output format is perhaps not quite as elegant as that in JDIC and JREADER, largely because it does not have as much control over essential aspects such as window and font size. This is more than compensated for by the inherent advantages of the windowing environment. My thanks to the many people who helped and gave advice to me, and particularly to Lars Huttar, Scott Trent, Philip Moore, Ken Lunde and the other XJDIC beta-testers, whose many suggestions and critical comment have played a considerable part in the program's development.. Cameron Blackwood helped me with the cbreak code, Paul Buchard provided the pure BSD versions of this, and Hitoshi Doi (who ran it on the 64-bit DEC Alpha) pointed out my invalid assumption that long integers were invariably 4 bytes long. A special mention to Andrew Moore, my Department's sysadmin, who laboured long and hard to install wnn/kterm/kinput on our DEC5000/3000/2000 (Ultrix) network without knowing a word of Japanese. With the release of XJDIC, the source is now available to the world. It has successfully been installed on many Unix platforms. A highly successful Macintosh KanjiTalk port has been undertaken by Dan Crevier to produce MacJDic. Will an Amiga port be next? As ever comments and constructive criticism are welcome. VERSION 1.1 The additions in Version 1.1 include: o the built-in romaji to kana code, which was in JDIC, but not included in the original XJDIC. o backspacing on the english, kana and kanji input lines. o the filter system* o the verb deinflection function* o a more flexible kanji index selection o the ability to specify a stroke count/bush combination o the .xjdicrc control file * these features also became available in V2.3 of JDIC & JREADER Jim Breen Department of Robotics & Digital Technology Monash University Melbourne, Australia (jwb@capek.rdt.monash.edu.au) July 1992 - July 1993