תקציר
Parsing texts into universal dependencies (UD) in realistic scenarios requires infrastructure for morphological analysis and disambiguation (MA&D) of typologically different languages as a first tier. MA&D is particularly challenging in morphologically rich languages (MRLs), where the ambiguous space-delimited tokens ought to be disambiguated with respect to their constituent morphemes. Here we present a novel, language-agnostic, framework for MA&D, based on a transition system with two variants, word-based and morpheme-based, and a dedicated transition to mitigate the biases of variable-length morpheme sequences. Our experiments on a Modern Hebrew case study outperform the state of the art, and we show that the morpheme-based MD consistently outperforms our word-based variant. We further illustrate the utility and multilingual coverage of our framework by morphologically analyzing and disambiguating the large set of languages in the UD treebanks.
שפה מקורית | אנגלית |
---|---|
כותר פרסום המארח | COLING 2016 - 26th International Conference on Computational Linguistics, Proceedings of COLING 2016 |
כותר משנה של פרסום המארח | Technical Papers |
מוציא לאור | Association for Computational Linguistics, ACL Anthology |
עמודים | 337-348 |
מספר עמודים | 12 |
מסת"ב (מודפס) | 9784879747020 |
סטטוס פרסום | פורסם - 2016 |
אירוע | 26th International Conference on Computational Linguistics, COLING 2016 - Osaka, יפן משך הזמן: 11 דצמ׳ 2016 → 16 דצמ׳ 2016 |
סדרות פרסומים
שם | COLING 2016 - 26th International Conference on Computational Linguistics, Proceedings of COLING 2016: Technical Papers |
---|
כנס
כנס | 26th International Conference on Computational Linguistics, COLING 2016 |
---|---|
מדינה/אזור | יפן |
עיר | Osaka |
תקופה | 11/12/16 → 16/12/16 |
הערה ביבליוגרפית
Publisher Copyright:© 1963-2018 ACL.