Buch, Englisch, Band 278, 305 Seiten, Format (B × H): 160 mm x 241 mm, Gewicht: 1390 g
Reihe: The Springer International Series in Engineering and Computer Science
Buch, Englisch, Band 278, 305 Seiten, Format (B × H): 160 mm x 241 mm, Gewicht: 1390 g
Reihe: The Springer International Series in Engineering and Computer Science
ISBN: 978-0-7923-9468-6
Verlag: Springer US
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.
Zielgruppe
Research
Autoren/Hrsg.
Fachgebiete
- Mathematik | Informatik EDV | Informatik Informatik Künstliche Intelligenz Wissensbasierte Systeme, Expertensysteme
- Mathematik | Informatik EDV | Informatik Informatik Natürliche Sprachen & Maschinelle Übersetzung
- Technische Wissenschaften Elektronik | Nachrichtentechnik Elektronik Robotik
- Mathematik | Informatik EDV | Informatik Informatik Künstliche Intelligenz Spracherkennung, Sprachverarbeitung
- Geisteswissenschaften Sprachwissenschaft Computerlinguistik, Korpuslinguistik
Weitere Infos & Material
1 INTRODUCTION.- 2 SEMANTIC EXTRACTION.- 2.1 Historical Overview.- 2.2 Cognitive Science Approaches.- 2.3 Recycling Approaches.- 2.4 Knowledge-Poor Approaches.- 3 SEXTANT.- 3.1 Philosophy.- 3.2 Methodology.- 3.3 Other examples.- 3.4 Discussion.- 4 EVALUATION.- 4.1 Deese Antonyms Discovery.- 4.2 Artificial Synonyms.- 4.3 Gold Standards Evaluations.- 4.4 Webster’s 7th.- 4.5 Syntactic vs. Document Co-occurrence.- 4.6 Summary.- 5 APPLICATIONS.- 5.1 Query Expansion.- 5.2 Thesaurus enrichment.- 5.3 Word Meaning Clustering.- 5.4 Automatic Thesaurus Construction.- 5.5 Discussion and Summary.- 6 CONCLUSION.- 6.1 Summary.- 6.2 Criticisms.- 6.3 Future Directions.- 6.4 Vision.- 1 PREPROCESSORS.- 2 WEBSTER STOPWORD LIST.- 3 SIMILARITY LIST.- 4 SEMANTIC CLUSTERING.- 5 AUTOMATIC THESAURUS GENERATION.- 6 CORPORA TREATED.- 6.1 ADI.- 6.2 AI.- 6.3 AIDS.- 6.4 ANIMALS.- 6.5 BASEBALL.- 6.6 BROWN.- 6.7 CACM.- 6.8 CISI.- 6.9 CRAN.- 6.10 HARVARD.- 6.11 JFK.- 6.12 MED.- 6.13 MERGERS.- 6.14 MOBYDICK.- 6.15 NEJM.- 6.16 NPL.- 6.17 SPORTS.- 6.18 TIME.- 6.19 XRAY.- 6.20 THESIS.