( 1 of 1 ) |
United States Patent | 5,937,422 |
Nelson , et al. | August 10, 1999 |
A method of automatically generating a topical description of text by receiving the text containing input words; stemming each input word to its root form; assigning a user-definable part-of-speech score to each input word; assigning a language salience score to each input word; assigning an input-word score to each input word; creating a tree structure under each input word, where each tree structure contains the definition of the corresponding input word; assigning a definition-word score to each definition word; collapsing each tree structure to a corresponding tree-word list; assigning a tree-word-list score to each entry in each tree-word list; combining the tree-word lists into a final word list; assigning each word in the final word list a final-word-list score; and choosing the top N scoring words in the final word list as the topic description of the input text. Document searching and sorting may be accomplished by performing the method described above on each document in a database and then comparing the similarity of the resulting topical descriptions.
Inventors: | Nelson; Douglas J. (Columbia, MD); Schone; Patrick John (Elkridge, MD); Bates; Richard Michael (Greenbelt, MD) |
Assignee: | The United States of America as represented by the National Security (Washington, DC) |
Appl. No.: | 834263 |
Filed: | April 15, 1997 |
U.S. Class: | 707/531; 707/4; 707/532; 707/535; 707/512 |
Intern'l Class: | G06F 017/30 |
Field of Search: | 704/10 707/512,532,535,531,3-5,7 |
4965763 | Oct., 1990 | Zamora | 704/1. |
5371673 | Dec., 1994 | Fan | 704/1. |
5384703 | Jan., 1995 | Withgott et al. | 707/531. |
5434962 | Jul., 1995 | Kyojima et al. | 707/531. |
5619410 | Apr., 1997 | Emori et al. | 704/7. |
5845278 | Dec., 1998 | Kirsch et al. | 707/3. |
5873660 | Feb., 1999 | Walsh et al. | 400/63. |