A computer-readable medium stores a first lexicon data structure for lexicon words. The first data structure includes a host form variant field containing a host form variant such as a clitic host form variant, a host form field containing the host form of the host form variant (only present if the forms differ) such as a clitic host verbal form, and a verification field indicative of whether the host form variant is a valid word. The first data structure also includes a segment association field containing data or segmentation bits associating the host form variant with certain types of attachment entries in the lexicon, which also contain data or segmentation bits, to define valid combinations between the host form variant and at least one of the attachment entries in the lexicon. A second lexicon data structure for each of the attachment entries in the lexicon is also stored.
The present application is based on and claims the benefit of U.S. provisional patent application Ser. No. 60/513,921, filed Oct. 23, 2003, the content of which is hereby incorporated by reference in its entirety.
CROSS-REFERENCE TO RELATED APPLICATIONS
Reference is hereby made to the following co-pending and commonly assigned patent applications. U.S. application Ser. No. 10/804,930, filed on Mar. 19, 2004, entitled "COMPOUND WORD BREAKER AND SPELL CHECKER" and U.S. application Ser. No. 10/804,998, filed Mar. 19, 2004, entitled "FULL-FORM LEXICON WITH TAGGED DATA AND METHODS OF CONSTRUCTING AND USING THE SAME", both of which are incorporated by reference in their entirety.