Proceedings of the 4th Workshop on Computational Approaches to Historical Language Change 2023
DOI: 10.18653/v1/2023.lchange-1.8
|View full text |Cite
|
Sign up to set email alerts
|

Multi-lect automatic detection of Swadesh list items from raw corpus data in East Slavic languages

Ilia Afanasev

Abstract: The article introduces a novel task of multilect automatic detection of Swadesh list items from raw corpora. The task aids the early stage of historical linguistics study by helping the researcher compile word lists for further analysis.In this paper, I test multi-lect automatic detection on the East Slavic lects' data. The training data consists of Ukrainian, Belarusian, and Russian material. I introduce a new dataset for the Ukrainian language. I implement data augmentation techniques to give automatic tools… Show more

Help me understand this report

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

0
0
0

Publication Types

Select...

Relationship

0
0

Authors

Journals

citations
Cited by 0 publications
references
References 23 publications
0
0
0
Order By: Relevance

No citations

Set email alert for when this publication receives citations?