Hand pose recognition has been one of the most fundamental tasks in computer vision and pattern recognition, and substantial effort has been devoted to this field. However, owing to lack of public large‐scale benchmark dataset, there is little literature to specially study pen‐holding hand pose (PHHP) recognition. As an attempt to fill this gap, in this paper, a PHHP image dataset, consisting of 18,000 PHHP samples is established. To the best of the authors’ knowledge, this is the largest vision‐based PHHP dataset ever collected. Furthermore, the authors design a coarse‐to‐fine PHHP recognition network consisting of a coarse multi‐feature learning network and a fine pen‐grasping‐specific feature learning network, where the coarse learning network aims to extensively exploit the multiple discriminative features by sharing a hand‐shape‐based spatial attention information, and the fine learning network further learns the pen‐grasping‐specific features by embedding a couple of convolutional block attention modules into three convolution blocks models. Experimental results show that the authors’ proposed method can achieve a very competitive PHHP recognition performance when compared with the baseline recognition models.