Clarissa Forbes scite author profile

Clarissa Forbes

2Publications

0Citation Statements Received

39Citation Statements Given

How they've been cited

How they cite others

Affiliations

Publications

Order By: Most citations

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages

Forbes¹,

Samir²,

Oliver³

et al. 2022

View full text Add to dashboard Cite

Recent progress in NLP is driven by pretrained models leveraging massive datasets and has predominantly benefited the world's political and economic superpowers. Technologically underserved languages are left behind because they lack such resources. Hundreds of underserved languages, nevertheless, have available data sources in the form of interlinear glossed text (IGT) from language documentation efforts. IGT remains underutilized in NLP work, perhaps because its annotations are only semistructured and often language-specific. With this paper, we make the case that IGT data can be leveraged successfully provided that target language expertise is available. We specifically advocate for collaboration with documentary linguists. Our paper provides a roadmap for successful projects utilizing IGT data: (1) It is essential to define which NLP tasks can be accomplished with the given IGT data and how these will benefit the speech community.(2) Great care and target language expertise is required when converting the data into structured formats commonly employed in NLP. (3) Task-specific and user-specific evaluation can help to ascertain that the tools which are created benefit the target language speech community. We illustrate each step through a case study on developing a morphological reinflection system for the Tsimchianic language Gitksan.

show abstract

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Underdocumented Languages

Forbes¹,

Samir²,

Oliver³

et al. 2022

Preprint

View full text Add to dashboard Cite

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Clarissa Forbes

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Underdocumented Languages

Contact Info

Product

Resources

About