ViXNet: Vision Transformer with Xception Network for Facial Expression Recognition
DOI:
https://doi.org/10.15849/ijasca.147Keywords:
Facial Expression Recognition (FER), Vision Transformer (ViT), Xception-ViTAbstract
Facial expression recognition (FER) is paramount to many computer vision tasks. It has garnered a lot of interest. However, the FER performance is limited by either expression ambiguity (high inter-class and low intra-class similarity) or noisy annotations in existing FER techniques. The structural relationships between various local regions in a facial expression are difficult for Xception-based models to represent. Techniques built on the principles of Vision Transformer (ViT) have been presented to capture long-range relationships between nearby regions. Nevertheless, because of its self-attention mechanism, ViT-based methods can develop redundant correlation representations and are susceptible to facing regions unrelated to expressions. We suggest an Xception-ViT to solve these problems by combining Xception’s extensive feature extraction capabilities with the Transformer’s capacity to comprehend intricate correlations between these features. An Xception-based pre-trained encoder is used on a lab-controlled and controlled FER dataset to extract feature parameters for facial expression recognition. Next, the features extracted by Xception are processed by a Vision Transformer (ViT) model to learn global and contextual relationships in the image. The prototypes are developed in order to verify the classification algorithm for specific facial expression recognition. The support and query sets are segregated into a FER dataset obtained from diverse environments and the classification episodes are constructed. The carefully configured Xception encoder with the help of the ViT model serves as a feature extractor that generates the prototype of each category of the support set. Xception-ViT can achieve recognition rates of 96.11%, 95.94%, 92.80%, and 100% on the FER24-CK+, JAFFE, FER+, and CK+ datasets respectively according to the extensive experimental results. Therefore, the use of Xception-ViT can improve the performance of specific facial expressions in these datasets.
Downloads
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Mohamed Ouhammou, Nabil Ababou, Mohamed BaslamCopyright © The Author(s).
Articles published in the International Journal of Advances in Soft Computing and its Applications (IJASCA) are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
This license permits anyone to copy, redistribute, remix, transform, and build upon the material for any purpose, including commercial use, provided appropriate credit is given to the original author(s), a link to the license is provided, and any modifications are indicated.
Authors retain the copyright of their published work and grant the journal right of first publication, with the work simultaneously licensed under the terms above.
Link
