Changbing Yang scite author profile

Changbing Yang

5Publications

29Citation Statements Received

69Citation Statements Given

How they've been cited

How they cite others

Affiliations

Publications

Order By: Most citations

CLiMP: A Benchmark for Chinese Language Model Evaluation

Xiang¹,

Yang²,

Liu³

et al. 2021

View full text Add to dashboard Cite

Linguistically informed analyses of language models (LMs) contribute to the understanding and improvement of these models. Here, we introduce the corpus of Chinese linguistic minimal pairs (CLiMP), which can be used to investigate what knowledge Chinese LMs acquire. CLiMP consists of sets of 1,000 minimal pairs (MPs) for 16 syntactic contrasts in Mandarin, covering 9 major Mandarin linguistic phenomena. The MPs are semiautomatically generated, and human agreement with the labels in CLiMP is 95.8%. We evaluate 11 different LMs on CLiMP, covering n-grams, LSTMs, and Chinese BERT. We find that classifier-noun agreement and verb complement selection are the phenomena that models generally perform best at. However, models struggle the most with the bǎ construction, binding, and filler-gap dependencies. Overall, Chinese BERT achieves an 81.8% average accuracy, while the performances of LSTMs and 5-grams are only moderately above chance level.

show abstract

IGT2P: From Interlinear Glossed Texts to Paradigms

Moeller

Liu

Yang³

et al. 2020

View full text Add to dashboard Cite

An intermediate step in the linguistic analysis of an under-documented language is to find and organize inflected forms that are attested in natural speech. From this data, linguists generate unseen inflected word forms in order to test hypotheses about the language's inflectional patterns and to complete inflectional paradigm tables. To get the data linguists spend many hours manually creating interlinear glossed texts (IGTs). We introduce a new task that speeds this process and automatically generates new morphological resources for natural language processing systems: IGTto-paradigms (IGT2P). IGT2P generates entire morphological paradigms from IGT input. We show that existing morphological reinflection models can solve the task with 21% to 64% accuracy, depending on the language. We further find that (i) having a language expert spend only a few hours cleaning the noisy IGT data improves performance by as much as 21 percentage points, and (ii) POS tags, which are generally considered a necessary part of NLP morphological reinflection input, have no effect on the accuracy of the models considered here.

show abstract

Linguist vs. Machine: Rapid Development of Finite-State Morphological Grammars

Beemer¹,

Boston²,

Bukoski³

et al. 2020

View full text Add to dashboard Cite

CLiMP: A Benchmark for Chinese Language Model Evaluation

Xiang¹,

Yang²,

Liu³

et al. 2021

Preprint

View full text Add to dashboard Cite

show abstract

Unsupervised Paradigm Clustering Using Transformation Rules

Yang¹,

Nicolai²,

Silfverberg³

2021

View full text Add to dashboard Cite

This paper describes the submission of the CU-UBC team for the SIGMORPHON 2021 Shared Task 2: Unsupervised morphological paradigm clustering. Our system generates paradigms using morphological transformation rules which are discovered from raw data. We experiment with two methods for discovering rules. Our first approach generates prefix and suffix transformations between similar strings. Secondly, we experiment with more general rules which can apply transformations inside the input strings in addition to prefix and suffix transformations. We find that the best overall performance is delivered by prefix and suffix rules but more general transformation rules perform better for languages with templatic morphology and very high morpheme-to-word ratios.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Changbing Yang

CLiMP: A Benchmark for Chinese Language Model Evaluation

IGT2P: From Interlinear Glossed Texts to Paradigms

Linguist vs. Machine: Rapid Development of Finite-State Morphological Grammars

CLiMP: A Benchmark for Chinese Language Model Evaluation

Unsupervised Paradigm Clustering Using Transformation Rules

Contact Info

Product

Resources

About