Abstract
This paper addresses the problem of learning meaningful human action attributes from high-dimensional video sequences based on union-of-subspaces (UoS) model. The model hypothesizes that each action attribute is represented by a subspace. It puts forth an extension of existing low-rank representation (LRR), termed the clustering-aware structure-constrained low-rank representation (CS-LRR) model, for unsupervised learning of human action attributes. The proposed CS-LRR model overcomes the shortcomings of existing techniques by its ability to handle disjoint subspaces, and by performing optimal spectral clustering of the subspaces. An efficient linear alternating direction method (LADM) is developed to solve the CS-LRR optimization problem. A human action or activity is represented as a sequence of transitions from one action attribute to another and can be uniquely represented by a subspace transition vector. These subspace transition vectors are used for human action recognition. The effectiveness of the proposed model is demonstrated through experiments on two real-world datasets for action recognition.