Songtao Shang scite author profile

Text preprocessing is one of the key problems in pattern recognition and plays an important role in the process of text classification. Text preprocessing has two pivotal steps: feature selection and feature weighting. The preprocessing results can directly affect the classifiers’ accuracy and performance. Therefore, choosing the appropriate algorithm for feature selection and feature weighting to preprocess the document can greatly improve the performance of classifiers. According to the Gini Index theory, this paper proposes an Improved Gini Index algorithm. This algorithm constructs a new feature selection and feature weighting function. The experimental results show that this algorithm can improve the classifiers’ performance effectively. At the same time, this algorithm is applied to a sensitive information identification system and has achieved a good result. The algorithm’s precision and recall are higher than those of traditional ones. It can identify sensitive information on the Internet effectively.

show abstract

An Improved Video Recommendations Based on the Hyperlink-Graph Model

Shang

Shang²,

Feng

et al. 2016

View full text Add to dashboard Cite

An Improved Hybrid Feature Selection Algorithm for Electric Charge Recovery Risk

Suxiang

Shi

et al. 2020

Mathematical Problems in Engineering

View full text Add to dashboard Cite

In order to extract more information that affects customer arrears behavior, the feature extraction method is used to extend the low-dimensional features to the high-dimensional features for the warning problem of user arrears risk model of electric charge recovery (ECR). However, there are many irrelevant or redundant features in data, which affect prediction accuracy. In order to reduce the dimension of the feature and improve the prediction result, an improved hybrid feature selection algorithm is proposed, integrating nonlinear inertia weight binary particle swarm optimization with shrinking encircling and exploration mechanism (NBPSOSEE) with sequential backward selection (SBS), namely, NBPSOSEE-SBS, for selecting the optimal feature subset. NBPSOSEE-SBS can not only effectively reduce the redundant or irrelevant features from the feature subset selected by NBPSOSEE but also improve the accuracy of classification. The experimental results show that the proposed NBPSOSEE-SBS can effectively reduce a large number of redundant features and stably improve the prediction results in the case of low execution time, compared with one state-of-the-art optimization algorithm, and seven well-known wrapper-based feature selection approaches for the risk prediction of ECR for power customers.

show abstract

An Improved Distributed File System Based on GPU Acceleration

Shang

Gan

2018

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Songtao Shang

A micro-video recommendation system based on big data

Improved Feature Weight Algorithm and Its Application to Text Classification

An Improved Video Recommendations Based on the Hyperlink-Graph Model

An Improved Hybrid Feature Selection Algorithm for Electric Charge Recovery Risk

An Improved Distributed File System Based on GPU Acceleration

Contact Info

Product

Resources

About