
    [42m                                          [40m
    [42m   [33;41mO[31;42mptical [33;41mC[31;42mharacter [33;41mR[31;42mecognition          [40m
    [42m                                          [40m    [33;3mby John Collett[0m
    [42m   Step by Step through an OCR session    [40m
    [42m                                          [40m


[32m   <> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <>
[0m

   [42m Background 1 [40m

    Every week I read nearly all of the [4mGuardian Weekly[0m, including its
    selections from [4mThe Washington Post[0m and [4mLe Monde[0m.  I usually finish
    the crossword, too!  I do not buy but have access to the weekly
    international edition of [4mLe Monde[0m, and I read about a third of it (not
    without difficulty, and I haven't even tried its crossword.)

    Quite often the same article can be found in English in the [4mGuardian[0m
    [4mWeekly[0m and in French in [4mLe Monde[0m.  Reading the two versions side
    by side can be interesting and revealing.

    [42m  Background  2  [40m

    At the moment I am responsible for a 4th Year University French course
    called 'Professional and Technical Translation', and am always on the
    lookout for varied, up-to-date, interesting, and challenging texts to
    use in this course.  The newspapers referred to above are just the job.
    The original version of [4mLe Monde[0m usually contains substantial articles
    on a wide range of topics including science, the Arts, history, etc. as
    well as current political, economic, and financial matters.  We have
    permission to make such copies as we require for teaching purposes.
    When a translation appears in the [4mGuardian Weekly[0m, the two versions
    provide samples of actual work currently being done by professional
    translators in the real world, and students can study, evaluate, and
    emulate them.

    [42m  Duplication  [40m

    It is possible to Xerox the copies we need for student use, but though
    quick, cheap, and legal, it is not always convenient. The size and
    shape of articles are rarely such that they can be transferred to A4
    paper without cutting and pasting, and I sometimes wish either to omit
    parts of an article or to insert explanatory comments.  This is where,
    at last, my AlfaData scanner and Migraph OCR software enter the story.

    [42m  A sample OCR run  [40m

    I'll use for this discussion the example of a recently scanned
    newspaper article.  It is in [4mLe Monde Sélection Hebdomadaire[0m, issue
    2322, Thursday 6 May 1993, p.10.  Its title is "Les Merveilles
    démystifiés" and it is about the recent decipherment of thousands of
    prehistoric engravings in two valleys in the French Alps.

    The article is arranged in six columns, each column being 44 mm wide
    (1.75 inches) by 133 mm high (5.25 inches), with a small print size of
    about 20 (twenty!) chars per inch, and 9 (nine!) lines per inch.  The
    output was printed on two and a half sides of A4 paper, at 12 chars and
    6 lines per inch.

    I don't intend here to describe all the options and refinements which
    are available in OCR, but simply to trace one standard task, step by
    step.

    [42m Setting the scanner and preparing the document [40m

    Apart from one complication explained later, I scanned each of the six
    columns in turn, with the scanner set to its finest setting of 400 dots
    per inch to cope with the small print.  And set to 'Text' of course.
    You feel such a fool when it doesn't work because you've overlooked the
    obvious, says the voice of recent experience.

    The column lengths were within the maximum possible in one scan, and
    for each scan I just had to move the original across the Migraph 'Tray'
    and refold it as necessary so that each column was within the scanner's
    vision, checking vertical alignment carefully each time.

    [42m Migraph Control Panel [40m

    The Migraph control panel is a requester of which the major sections
    are arranged in terms of Input, Document, and Output, and allow the
    selection of Language, Dictionary, Text type, Pitch, etc. as required.
    Note that 'dictionary' here means 'stored data about letter shapes',
    not meanings of words.  More on that later.

    This is also the time to select 'Automatic' or 'Interactive'.

    [4mAutomatic[0m  Given the quality of the source document, you think the
    software will be able to identify correctly all or nearly all the
    characters. You are prepared to accept the raw results, and to polish
    the text later on, in a text editor, if needed.

    [4mInteractive[0m  There may be problems, and you would rather clear them up
    as they arise during the processing. You intend to help the program
    with the characters it finds difficult, and may choose to train it
    towards dealing with them on its own in the future.

    [42m Scanner Settings [40m

    Another requester, to confirm the DPI setting at 400 dpi, the required
    length of scan (the only one I actually needed to set), and the
    direction : portrait ( = vertical, the default) or landscape.

    [42m Scan [40m

    When you click to do the scan, you have to confirm 'OK' in response to
    a warning 'This will destroy the current image'.

    Then you do the scan itself, moving the scanner steadily, and keeping
    an eye on the little red monitor light.  If it flickers, you're going
    too fast. If it goes out, either you've finished, or you might as well
    start again.  If you can also glance at the screen, you should see a
    portion of what you are scanning, large and clear.

    [42m Zoom [40m

    Of the three zoom choices, F ( = Full) is needed at this stage. This
    option squeezes the whole image on to the screen, making the printing
    itself more or less illegible.  Everything that passed under the
    scanner window is there, but you will normally wish to clip one or more
    particular areas out of the full image.  You are able to distinguish
    paragraph shapes etc. as needed for the next step.

    [42m Set clip area [40m

    For a plain rectangle of text, click on the Text and Clip Box icons,
    and then a couple of clicks is all it takes to define the area of the
    screen to be processed.

    Then it's time to click on the OCR gadget, and believe me, the
    process leading up to this point is much quicker than might appear to
    be the case from the above description.

    [42m OCR 'phases' [40m

    A mighty amount of complex and clever work goes on behind the scenes.
    You are kept informed of progress with slider bars making their
    way, in turn, from 0% to 100%.  Sometimes they appear to be stuck.
    Don't worry - the program is working very hard.  Some phases are
    faster than other.  It presumably depends on the nature of the task,
    text size, clarity, etc. You can watch the progress of at least some
    of :
          [32mAuto learning phase            Page reconstruction phase
          [32mCharacter recognition phase    Linguistical (sic!) phase

    [31;42m Interaction [40m

    Assuming 'Interactive' and not 'Automatic' was selected, you now have
    the chance to help the program to identify problem characters, and
    the option of training it for the future.  It will already have
    consulted its dictionaries and lexicons to help it make the best
    possible guess at the identity of each character.  If it feels unsure
    about a character, it presents its 'best guess' for you to confirm or
    correct.  Completely unrecognized symbols are represented by '@'.

    In this interactive phase of correction, confirmation, and learning,
    three boxes are used to display progress.

     [42m In this box the text is built up word by word as it is       [40m
     [42m confirmed.  The current word is [43mhilighted[42m                    [40m
     [42m                                                              [40m
     [42m                                                              [40m

                          [42m  hilighte[43md[42m  [40m

                             [42m    d  [42m

    The program is not sure about the 'd'.  There may have been a flaw on
    the paper.  You are asked to confirm the 'd' in the third box by
    pressing Return or clicking an 'Accept' gadget. If a letter is wrong
    you change it by typing over it.

    Then you have the following choice of actions.  Their general meaning
    is fairly obvious, though there are some subtleties which I'll not
    go into here.

      [42m  Stop   [40m  [42m  Undo   [40m  [42m Confirm [40m              [42m  Train  [40m   [42m  [40m

      [42m  Auto   [40m  [42m  Delete [40m                         [42m  Accept [40m   [41m  [40m

    Either 'Train' or 'Accept' is on, as indicated by the filled-in square.
    'Train' is used for a character which you wish to be added to a
    dictionary, 'Accept' for one which you wish to be included in the
    current file, but not in a dictionary.

    And so you proceed to the end of the file.  With a good clear source
    text you may have no work to do.  With very small and closely packed
    print, as in [4mLe Monde[0m, there will be a few queries and corrections,
    but it is not a long or tedious task.

    When all is done, you are told how many words and characters have been
    processed, and asked to confirm that it is is OK to output the results
    to a text file.  You can preset its mode to 'new' or 'append'.  Another
    slider shows the progress of the output stage.  The target text now
    exists as an ASCII file, and you can do with it whatever you do with
    ASCII files.

    [42m The kerning problem [40m

    From Collins English Dictionary : "kern : the part of the character on
    a piece of printer's type that projects beyond the body."
    More widely used now to refer to pairs of letters which are printed so
    closely that they actually touch, sometimes deliberately, but sometimes
    just the result of cramming too many letters in a given space. Such
    shapes cannot be recognised by an OCR program unless it has been
    specially trained to do so.  With printing at 20 cpi as in [4mLe Monde[0m,
    problems do arise.  In interactive mode, they can be dealt with as they
    arise.

    To give just a couple of examples, I met a few cases of a joined 's'
    and 't', and of an 'f' followed by 'i' in which the end of the upper
    hook on the 'f' is vertically above the space occupied by the 'i'.

    A trickier problem is vertical overlap of letters as in jag  pay
                                                            hub  but.

    Look in vain for a space between 'jag' and 'hub' or between 'pay' and
    'but'. The program will read those two lines as one, because there is
    no uninterrupted space between them, and the entire pair of lines
    becomes problematical.  Nine lines per inch, as in [4mLe Monde[0m, is tight,
    and means that this problem of vertical touching is going to occur from
    time to time. It is more of a problem than horizontal kerning, since,
    as far as I know, there is no way of handling it within OCR.

    [42m More on the dictionaries [40m

    I still have some learning to do here, but you can use the supplied
    MIOCR.ALD dictionary in 'read' mode, append additional material to an
    existing dictionary in 'learn' mode, or create a new dictionary for
    special purposes in 'new' mode.  I have tried all three.

    Although the printing in [4mLe Monde[0m is cramped, the actual type face
    used looks standard.  Apart from kerning problems as described above, I
    expected the default dictionary to be able to read all the characters
    correctly, and initially opted for 'read' mode. There were a few
    characters which kept on being mis-read.  For example, the lower case
    'l' has a rather clunky shape at its top, and it kept being read as 'ì'
    ('i' with a grave accent!).  So I used the 'learn' mode and selected
    the 'train' option to teach my dictionary to recognise that shape
    correctly.  That seemed to work.  I got it to the stage where it was
    only confused by unusual kerning, badly formed letters, specks of dust
    on the paper, and other such hazards. (The scanning process and the
    zoom options give a high degree of magnification. It is astonishing to
    see just how deformed some of the letter shapes really are, which makes
    the performance of the software all the more impressive.)

    I referred earlier to a complication in one of the columns of the
    article, about half of which was printed in italics.  The standard
    dictionary would keep failing to identify characters correctly if it
    was used on that patch. This is the kind of circumstance in which
    special purpose dictionaries are needed, and I have created one called
    ITALICS.ALD (ALD is the standard file extension for dictionaries).
    With use, it has been getting progressively better at recognising
    italic letters.  I did the italic patch separately, using that
    dictionary, and it worked reasonably well.


    [33;3mHamilton, New Zealand  May, 1993[0m

                    ---------------------------

[32m   <> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <%> 34 <>
[0m


