Jointly Learning Heterogeneous Features for RGB-D Activity Recognition

Hu, Jiefeng; Zheng, Wei‐Shi; Lai, Jianhuang; Zhang, Jianguo

doi:10.1109/tpami.2016.2640292

Cited by 265 publications

(260 citation statements)

References 42 publications

Supporting

Mentioning

258

Contrasting

Order By: Relevance

“…SYSU dataset: For the empirical evaluations, we compare our DACNN algorithm to other baselines including CNN+DPRL [28], ST-LSTM+Trust Gate [13], Dynamic Skeletons [35], LAFF(SKL) [45], SR-TSL [46], VA-LSTM [47], and GCA-LSTM [48], which includes the most recent deep learning applications (CNN, LSTM, etc.) on this dataset.…”

Section: Action Recognition Resultsmentioning

confidence: 99%

See 1 more Smart Citation

Deep-Aligned Convolutional Neural Network for Skeleton-Based Action Recognition and Segmentation

Hosseini

Montagne

Hammer

2019

2019 IEEE International Conference on Data Mining (ICDM)

View full text Add to dashboard Cite

Convolutional neural networks (CNNs) are deep learning frameworks which are well-known for their notable performance in classification tasks. Hence, many skeleton-based action recognition and segmentation (SBARS) algorithms benefit from them in their designs. However, a shortcoming of such applications is the general lack of spatial relationships between the input features in such data types. Besides, non-uniform temporal scalings is a common issue in skeletonbased data streams which leads to having different input sizes even within one specific action category. In this work, we propose a novel deep-aligned convolutional neural network (DACNN) to tackle the above challenges for the particular problem of SBARS. Our network is designed by introducing a new type of filters in the context of CNNs which are trained based on their alignments to the local subsequences in the inputs. These filters result in efficient predictions as well as learning interpretable patterns in the data. We empirically evaluate our framework on real-world benchmarks showing that the proposed DACNN algorithm obtains a competitive performance compared to the state-of-the-art while benefiting from a less complicated yet more interpretable model. * Preprint of the publication [1] including extended experiments, as provided by the authors. The final publication is available at https://ieeexplore.ieee.org/ networks (RNN) [8,12,13]. RNN methods can learn the temporal dynamics of the sequential data; nevertheless, they have practical shortcomings in the training of their stacked structures [14,15].Compared to RNN architectures, CNN-based methods provide more effective solutions by extracting local features from their input and finding discriminative patterns in the data [16,10]. However, regardless of the promising feature extraction capability of CNN, its specific convolutional structure is designed originally for image-based input data and primarily relies on spatial dependencies between the neighboring points. In contrast, such a direct relationship does not generally exist in skeleton-based action datasets. Although some works tried to solve this problem by using 1-dimensional filters (only for the temporal dimension), it is still not an efficient solution to this specific shortcoming of CNN-based frameworks [17].A crucial step before analyzing any motion data is the temporal segmentation phase, in which we predict the action to which each time-frame belongs. Although there exist unsupervised algorithms for temporal segmentation of motion data [18,19], they usually oversegment actions into smaller sub-actions or segment also the blank parts of the stream.Motivations: An essential group of techniques for classification of the sequential data is time-series alignment methods [20]. It is shown that via comparing each data sequence to some predefined or learned sequences, we can discriminate or segment the data samples with high accuracy [21,22]. Also, in algorithms similar to [23], finding a small distinct subsequence in the input data can reveal i...

show abstract

Section: Action Recognition Resultsmentioning

confidence: 99%

“…SYSU-3D Human-Object Interaction dataset (SYSU) [35]: This dataset contains 12 different action classes recorded in 480 video sequences. SYSU dataset is captured from 40 human subjects, and in each time-frame, it represents the 3D coordinates of 20 body joints.…”

Section: Montalbano V2 Datasetmentioning

confidence: 99%

Deep-Aligned Convolutional Neural Network for Skeleton-Based Action Recognition and Segmentation

Hosseini

Montagne

Hammer

2019

2019 IEEE International Conference on Data Mining (ICDM)

View full text Add to dashboard Cite

show abstract

“…An intuitive way to combine multimodal features is to directly concatenate them together . To mine more useful information among multimodal features for better performance, researchers propose to explicitly learn shared‐specific structures among features …”

Section: Related Workmentioning

confidence: 99%

“…52 To mine more useful information among multimodal features for better performance, researchers propose to explicitly learn shared-specific structures among features. 11,53 Early on, Liu and Shao 54 utilized a genetic programming framework to improve not only RGB and depth descriptors but also their fusion simultaneously through an iterative evolution. Ni et al 55 concatenated depth descriptor-and RGB-based representations of spatiotemporal interest points for better RGB+D information fusion.…”

Section: Multimodal 3d Action Recognitionmentioning

confidence: 99%

“…Generally speaking, there are two kinds of methods for validation in action recognition, ie, leave-one-out cross validation and 2-fold cross validation. We follow the idea in the works of Liu et al, 13 Xia et al, 33 and Hu et al 53 to utilize leave-one-out cross validation for the experiments. Florence 3D data set collects data through a stationary Kinect as well, collecting nine common indoor action categories, such as "watching," "drinking water," "calling," and so on.…”

Section: Data Setsmentioning

confidence: 99%

See 1 more Smart Citation

Deep spatiotemporal LSTM network with temporal pattern feature for 3D human action recognition

Wei²,

Duan

2019

Computational Intelligence

View full text Add to dashboard Cite

With the rapid development of RGB‐D cameras and pose estimation techniques, action recognition based on three‐dimensional skeleton data has gained significant attention in the artificial intelligence community. In this paper, we incorporate temporal pattern descriptors of joint positions with the currently popular long short‐term memory (LSTM)–based learning scheme to obtain accurate and robust action recognition. Considering that actions are essentially formed by small subactions, we first utilize a two‐dimensional wavelet transform to extract temporal pattern descriptors in the frequency domain for each subaction. Afterward, we design a novel LSTM structure to extract deep features, which model a long‐term spatiotemporal correlation between body parts. Since temporal pattern descriptors and LSTM deep features can be regarded as multimodal representations for actions, we fuse them with an autoencoder network to achieve a more effective feature descriptor for action recognition. Experimental results on three challenging data sets with several comparative methods demonstrate the effectiveness of the proposed method for three‐dimensional action recognition.

show abstract

Research on multi‐center assisted diagnosis of ASD based on multimodal feature fusion

Zhang,

Sheng,

Wang

et al. 2024

Int J Imaging Syst Tech

View full text Add to dashboard Cite

Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder, and structural magnetic resonance imaging (sMRI) and functional magnetic resonance imaging (fMRI) provide information about brain structure and function, aiding in objective ASD diagnosis. However, existing ASD classification methods face challenges such as sample scarcity, inter‐imaging center variations, insufficient single‐modality information, and inconsistent feature dimensions. This study introduced a method based on the Local Global Multimodal Domain Adaptation (LGMDA)‐Sparse Adaptive Prior Coupled Dictionary Learning (SACDL) framework. Initially, the LGMDA method was introduced to achieve multi‐source DA. By minimizing differences between different data domains while maximizing inter‐class differences within the same domain, it expands the sample size of multi‐modal data, addressing the issues of sample scarcity and heterogeneity in ASD data. Subsequently, the SACDL method was employed for multimodal fusion. It initialized dictionaries using the ATGP algorithm, combined sMRI and fMRI data for dictionary learning, adaptively adjusted sparsity parameters, and integrated ASD phenotype data for constrained optimization. It enables joint learning of shared and modality‐specific features, balancing differences in feature dimensions. Experimental results show that this model effectively utilizes multi‐center, multi‐modal information to achieve better auxiliary diagnosis than single‐modal small samples. This method has the potential to provide effective solutions for ASD multi‐source and multi‐modal classification problems, which are significant for ASD research and clinical diagnosis.

show abstract

Jointly Learning Heterogeneous Features for RGB-D Activity Recognition

Cited by 265 publications

References 42 publications

Deep-Aligned Convolutional Neural Network for Skeleton-Based Action Recognition and Segmentation

Deep-Aligned Convolutional Neural Network for Skeleton-Based Action Recognition and Segmentation

Deep spatiotemporal LSTM network with temporal pattern feature for 3D human action recognition

Research on multi‐center assisted diagnosis of ASD based on multimodal feature fusion

Contact Info

Product

Resources

About