The quality of speech transcripts is of great significance to power system management and is an important basis for supporting subsequent data analysis. In this paper, based on the characteristics of speech transcription texts of customer service, this paper proposes an analysis method combining manual processing and latent Dirichlet allocation topic model, analyzing the transcribed texts. First, data preprocessing is performed on the State Grid’s work order data, and then the text topic distribution calculation is performed by the LDA topic model, and the topic parameter is set to a total of 100 topics. Next, the unsupervised clustering of the documents is performed by the k-means method, and the similarity between the files is obtained. Finally, the quality of the data is analyzed by combining manual labeling and manual evaluation. For the first time, this paper marks and identifies the State Grid’s work order analysis data, which is a pioneering work for natural language processing technology in the field of power grid.
scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.