Skip to main navigation Skip to search Skip to main content

Deep learning–driven image captioning: Progress through transformers and large language models

  • Priyanka Panchal
  • , Vishal Polara
  • , U. Siddaraj*
  • , Abdullah Baz
  • , Shobhit K. Patel
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

This paper provides a novel deep learning model for captioning of images by using an advanced vision transformer architecture with a powerful LLM. Proposed models show a significant improvement over traditional CNN-RNN hybrids and existing transformer-based approaches by integrating a unique cross-attention mechanism that enables deep alignment between linguistic context and visual features. We show the superiority of our proposed architecture through extensive evaluation on different datasets like MSCOCO, Flickr30K, and NoCaps. The proposed model consistently shows good performance for leading methods such as GIT, BLIP-2, and CoCa across a comprehensive suite of metrics. On the MS COCO dataset, the BLEU-4, METEOR, and CIDEr scores of proposed models are equal to 0.495, 0.390, and 1.32, respectively. In this paper, we have critically analyzed the key challenges of this field, like enhancing caption diversity, ensuring robust multimodal alignment, and mitigating inherent biases. By providing a new performance level, the proposed model provides a source of reference for the next generation of image captioning systems. The results show the efficiency of our fusion strategy and facilitate the development of techniques that use models that can produce more precise, contextually rich, and human-like image depictions. This work supports SDG 9 (Industry, Innovation, and Infrastructure) by advancing multimodal AI systems, and SDG 4 (Quality Education) by enabling intelligent and accessible image understanding technologies.

Original languageEnglish
Article numbere0345012
JournalPLoS One
Volume21
Issue number3 March
DOIs
Publication statusPublished - 03-2026

All Science Journal Classification (ASJC) codes

  • General

Fingerprint

Dive into the research topics of 'Deep learning–driven image captioning: Progress through transformers and large language models'. Together they form a unique fingerprint.

Cite this