Detecting Plagiarized Text in Images using OCR and NLP-Based Deep Learning Approaches

Authors

  • Belal Zaqaibeh Department of Data Science and Artificial Intelligence, Jadara University, Jordan
  • Ahmad Alhami Department of Data Science and Artificial Intelligence, Jadara University, Jordan

DOI:

https://doi.org/10.15849/ijasca.v18i2.108

Keywords:

DEPTI, Text Image Plagiarism, Deep Learning, Natural Language Processing

Abstract

This paper evaluates plagiarism detection using deep learning and natural language processing (NLP) techniques. A novel model, Detecting Embedded Plagiarized Text in Images (DEPTI), is introduced to identify plagiarized text embedded within images, demonstrating high accuracy and robust performance. DEPTI effectively recognizes paraphrased, translated, and artificial intelligence generated content, achieving strong detection capabilities across diverse scenarios. The model integrates PAN-PC-11, TF-IDF, Tesseract OCR, DistilBERT, and LSTM to extract and analyze text from images, enabling advanced plagiarism detection beyond conventional approaches. Experimental results confirm DEPTI’s effectiveness, highlighting its potential as a reliable tool for safeguarding academic integrity in the digital era.

Downloads

All Downloads: 2

Download data is not yet available.

Downloads

Published

2026-07-03

How to Cite

Zaqaibeh, B., & Alhami , A. . (2026). Detecting Plagiarized Text in Images using OCR and NLP-Based Deep Learning Approaches. International Journal of Advances in Soft Computing and Its Applications, 18(2), 376–384. https://doi.org/10.15849/ijasca.v18i2.108

Google Scholar Link