تخطي إلى التنقل الرئيسي تخطي إلى البحث تخطي إلى المحتوى الرئيسي

MRL Parsing Without Tears: The Case of Hebrew

  • Shaltiel Shmidman
  • , Avi Shmidman
  • , Moshe Koppel
  • , Reut Tsarfaty

نتاج البحث: فصل من :كتاب / تقرير / مؤتمرمنشور من مؤتمرمراجعة النظراء

ملخص

Syntactic parsing remains a critical tool for relation extraction and information extraction, especially in resource-scarce languages where LLMs are lacking. Yet in morphologically rich languages (MRLs), where parsers need to identify multiple lexical units in each token, existing systems suffer in latency and setup complexity. Some use a pipeline to peel away the layers: first segmentation, then morphology tagging, and then syntax parsing; however, errors in earlier layers are then propagated forward. Others use a joint architecture to evaluate all permutations at once; while this improves accuracy, it is notoriously slow. In contrast, and taking Hebrew as a test case, we present a new "flipped pipeline": decisions are made directly on the whole-token units by expert classifiers, each one dedicated to one specific task. The classifier predictions are independent of one another, and only at the end do we synthesize their predictions. This blazingly fast approach requires only a single huggingface call, without the need for recourse to lexicons or linguistic resources. When trained on the same training set used in previous studies, our model achieves near-SOTA performance on a wide array of Hebrew NLP tasks. Furthermore, when trained on a newly enlarged training corpus, our model achieves a new SOTA for Hebrew POS tagging and dependency parsing. We release this new SOTA model to the community. Because our architecture does not rely on any language-specific resources, it can serve as a model to develop similar parsers for other MRLs.

اللغة الأصليةالإنجليزيّة
عنوان منشور المضيفThe 62nd Annual Meeting of the Association for Computational Linguistics
العنوان الفرعي لمنشور المضيفFindings of the Association for Computational Linguistics, ACL 2024
المحررونLun-Wei Ku, Andre Martins, Vivek Srikumar
ناشرAssociation for Computational Linguistics (ACL)
الصفحات4537-4550
عدد الصفحات14
رقم المعيار الدولي للكتب (الإلكتروني)9798891760998
المعرِّفات الرقمية للأشياء
حالة النشرنُشِر - 2024
منشور خارجيًانعم
الحدثFindings of the 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024 - Hybrid, Bangkok, تايلند
المدة: ١١ أغسطس ٢٠٢٤١٦ أغسطس ٢٠٢٤

سلسلة المنشورات

الاسمProceedings of the Annual Meeting of the Association for Computational Linguistics
رقم المعيار الدولي للدوريات (المطبوع)0736-587X

!!Conference

!!ConferenceFindings of the 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024
الدولة/الإقليمتايلند
المدينةHybrid, Bangkok
المدة١١/٠٨/٢٤١٦/٠٨/٢٤

ملاحظة ببليوغرافية

Publisher Copyright:
© 2024 Association for Computational Linguistics.

بصمة

أدرس بدقة موضوعات البحث “MRL Parsing Without Tears: The Case of Hebrew'. فهما يشكلان معًا بصمة فريدة.

قم بذكر هذا