"VocPix: Image Captioning using CNN and LSTM"
Abstract & Details
Research Area
Computer Engineering
Keywords
Deep learning
CNN
LSTM
Machine learning
Neural Networks
Text-To-Speech
Abstract
With billions of users on sites like Facebook, Twitter, Instagram, and YouTube, social media has become an essential component of modern life. Social media has fundamentally altered how we communicate, share information, and engage with one another. It has also changed how we both consume and produce material. Therefore, an image caption generator has become essential in today's culture because it is necessary for social media addicts or people who are blind. It is a kind of algorithm or software program that makes use of Deep Learning techniques and analyses an image's visual information before converting it into plain language.
It can be used as a plugin on the popular social networking sites of today to suggest appropriate captions for users to include with their postings. The goal of the suggested research is to create an image caption, also known as a description of an image, and to translate it into different languages, using CNN-LSTM architecture. To use the current word as input for the prediction of the following word, CNN layers will help in retrieving input data, and LSTM will extract essential information as it processes input. Python 3 and machine learning will be the programming languages used. This study will go into great detail about the many Neural networks involved, including their structures and functions. The program that converts created captions into spoken words is called a text-to-speech synthesizer, and it uses Natural Language Processing (NLP) to analyze and process the text. The text is subsequently converted into a synthesized speech representation using digital signal processing (DSP) technology. Here, we've created a practical text-to-speech synthesizer in the form of an easy-to-use application that reads aloud generated captions as synthesized speech.
The proposed deep learning approach aims at generating the best caption for a particular image by analyzing and extracting various features from images and converting that textual caption into speech using Text-To-Speech (TTS).
License
This work is licensed under a Creative
Commons
Attribution-ShareAlike 4.0 International License.
Author Information
| # | Name | Institute / Affiliation |
|---|---|---|
| 1 | Bhushan Arun Ambhore | Sinhgad College Of Engineering, Pune |
| 2 | Linal Anil Patil | Sinhgad College Of Engineering, Pune |
| 3 | Manasi Shekhar Patil | Sinhgad College Of Engineering, Pune |
| 4 | Niket Sitaram Sharma | Sinhgad College Of Engineering, Pune |
| 5 | Hitesh E. Chaudhari | Sinhgad College Of Engineering, Pune |
How to Cite
Use the following formats to cite this article in your research.
APA Style
Ambhore, Bhushan Arun, Patil, Linal Anil, Patil, Manasi Shekhar, Sharma, Niket Sitaram, & Chaudhari, Hitesh E. (2023). "VocPix: Image Captioning using CNN and LSTM". International Journal of Advance Research and Innovative Ideas In Education, 9(2), 2873-2880.
MLA Style
Ambhore, Bhushan Arun, et al. ""VocPix: Image Captioning using CNN and LSTM"." International Journal of Advance Research and Innovative Ideas In Education, vol. 9, no. 2, 2023, pp. 2873-2880.
IEEE Style
Bhushan Arun Ambhore, Linal Anil Patil, Manasi Shekhar Patil, Niket Sitaram Sharma, and Hitesh E. Chaudhari, ""VocPix: Image Captioning using CNN and LSTM"," International Journal of Advance Research and Innovative Ideas In Education, vol. 9, no. 2, pp. 2873-2880, 2023.
Vancouver Style
Ambhore Bhushan Arun, Patil Linal Anil, Patil Manasi Shekhar, Sharma Niket Sitaram, Chaudhari Hitesh E.. "VocPix: Image Captioning using CNN and LSTM". International Journal of Advance Research and Innovative Ideas In Education. 2023;9(2):2873-2880.
Harvard Style
Ambhore, Bhushan Arun, Patil, Linal Anil, Patil, Manasi Shekhar, Sharma, Niket Sitaram, & Chaudhari, Hitesh E. (2023) '"VocPix: Image Captioning using CNN and LSTM"', International Journal of Advance Research and Innovative Ideas In Education, 9(2), pp. 2873-2880.
Chicago Style
Ambhore, Bhushan Arun, et al. ""VocPix: Image Captioning using CNN and LSTM"." International Journal of Advance Research and Innovative Ideas In Education 9, no. 2 (2023): 2873-2880.
Turabian Style
Ambhore, Bhushan Arun, et al. ""VocPix: Image Captioning using CNN and LSTM"." International Journal of Advance Research and Innovative Ideas In Education 9, no. 2 (2023): 2873-2880.
Related Research
Comprehensive Review of Existing Chatbot Systems for Career Assistance, Resume Support, and ATS-Aware Guidance
PDF Unavailable
Development of an AI-Powered Multimodal Web Assistant with Intelligent Resume Building and ATS Enhancement
PDF Unavailable
A Deep Learning-Based Framework for Mood-Oriented Music Recommendation Using Facial Expression Analysis
PDF Unavailable
Survey On : Intelligent Payroll and Human Resource Management Systems: A Systematic Review of Automation, Security, and Analytics
PDF Unavailable
Civic Engagement & Empowerment Platform
PDF Unavailable
Recent Developments in Microneedle Technology and Its Diverse Biomedical Applications
PDF Unavailable
RAG System Development with Pydantic AI ChromaDB & Groq
PDF Unavailable
Machine Learning Based Early Stage Diabetes Detection System
PDF Unavailable
A Survey on Skillsense:AI Career Analyzer App
PDF Unavailable
Employee Performance Portal
PDF Unavailable