ViXNet: Vision Transformer with Xception Network for Facial Expression Recognition

Authors

  • Mohamed Ouhammou Sultan Moulay Slimane University, Faculty of Science & Technology, Beni Mellal, Morocco.
  • Nabil Ababou Sultan Moulay Slimane University, Faculty of Science & Technology, Beni Mellal, Morocco.
  • Mohamed Baslam Sultan Moulay Slimane University, Faculty of Science & Technology, Beni Mellal, Morocco.

DOI:

https://doi.org/10.15849/ijasca.147

Keywords:

Facial Expression Recognition (FER), Vision Transformer (ViT), Xception-ViT

Abstract

Facial expression recognition (FER) is paramount to many computer vision tasks. It has garnered a lot of interest. However, the FER performance is limited by either expression ambiguity (high inter-class and low intra-class similarity) or noisy annotations in existing FER techniques. The structural relationships between various local regions in a facial expression are difficult for Xception-based models to represent. Techniques built on the principles of Vision Transformer (ViT) have been presented to capture long-range relationships between nearby regions. Nevertheless, because of its self-attention mechanism, ViT-based methods can develop redundant correlation representations and are susceptible to facing regions unrelated to expressions. We suggest an Xception-ViT to solve these problems by combining Xception’s extensive feature extraction capabilities with the Transformer’s capacity to comprehend intricate correlations between these features. An Xception-based pre-trained encoder is used on a lab-controlled and controlled FER dataset to extract feature parameters for facial expression recognition. Next, the features extracted by Xception are processed by a Vision Transformer (ViT) model to learn global and contextual relationships in the image. The prototypes are developed in order to verify the classification algorithm for specific facial expression recognition. The support and query sets are segregated into a FER dataset obtained from diverse environments and the classification episodes are constructed. The carefully configured Xception encoder with the help of the ViT model serves as a feature extractor that generates the prototype of each category of the support set. Xception-ViT can achieve recognition rates of 96.11%, 95.94%, 92.80%, and 100% on the FER24-CK+, JAFFE, FER+, and CK+ datasets respectively according to the extensive experimental results. Therefore, the use of Xception-ViT can improve the performance of specific facial expressions in these datasets.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-11

How to Cite

ViXNet: Vision Transformer with Xception Network for Facial Expression Recognition (M. Ouhammou, N. Ababou, & M. . Baslam, Trans.). (2026). International Journal of Advances in Soft Computing and Its Applications , 18(3), 411–432. https://doi.org/10.15849/ijasca.147
Total Downloads: 0

Google Scholar Link

Similar Articles

1-10 of 14

You may also start an advanced similarity search for this article.