ASL Chatbot: Sign to Speech
Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted3 hours ago
I want to create an AI-powered web chatbot that recognises American Sign Language in real time, turns each recognised sign into clear written text, and then voices that text through natural-sounding speech. The end goal is an easy-to-use educational tool that helps people who are deaf and non-verbal practise everyday communication with hearing users directly from their browser.
Scope of work
• Build or fine-tune a computer-vision model (e.g., MediaPipe, TensorFlow, PyTorch, or similar) to detect and classify ASL signs from a webcam stream.
• Pipe the recognised signs to a text layer, then feed that text into a speech-synthesis engine so the conversation flows naturally.
• Develop a responsive web interface where users can sign into the camera, read the live transcript, and hear the spoken output instantly.
• Keep latency low enough that the interaction feels conversational (sub-second end-to-end is my target).
• Provide clean, well-commented source code, a brief deployment guide, and a short demo video showing the system working in a standard desktop browser.
Acceptance criteria
1. At least 90 % accuracy on the most common ASL alphabet and core vocabulary in normal indoor lighting.
2. Text and speech output appear within one second of a completed sign on a 10 Mbps connection.
3. Runs inside Chrome, Edge, and Firefox without extra plugins.
4. Setup instructions allow me to deploy the service on my own VPS (Ubuntu 22.04, Docker available).
Future phases may expand to mobile apps and additional sign languages, so modular, well-documented code is essential. If you have previous work in gesture recognition, ASL datasets, or browser-based inference, please highlight it when you respond; it will make collaboration much smoother.
Scope of work
• Build or fine-tune a computer-vision model (e.g., MediaPipe, TensorFlow, PyTorch, or similar) to detect and classify ASL signs from a webcam stream.
• Pipe the recognised signs to a text layer, then feed that text into a speech-synthesis engine so the conversation flows naturally.
• Develop a responsive web interface where users can sign into the camera, read the live transcript, and hear the spoken output instantly.
• Keep latency low enough that the interaction feels conversational (sub-second end-to-end is my target).
• Provide clean, well-commented source code, a brief deployment guide, and a short demo video showing the system working in a standard desktop browser.
Acceptance criteria
1. At least 90 % accuracy on the most common ASL alphabet and core vocabulary in normal indoor lighting.
2. Text and speech output appear within one second of a completed sign on a 10 Mbps connection.
3. Runs inside Chrome, Edge, and Firefox without extra plugins.
4. Setup instructions allow me to deploy the service on my own VPS (Ubuntu 22.04, Docker available).
Future phases may expand to mobile apps and additional sign languages, so modular, well-documented code is essential. If you have previous work in gesture recognition, ASL datasets, or browser-based inference, please highlight it when you respond; it will make collaboration much smoother.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.