Applied AI Case Study
Gesto — Sign to Speech
A privacy-first sign-to-speech web application that combines MediaPipe landmarks, ONNX Runtime Web, browser-local TensorFlow.js training, dataset portability, and accessible speech output for a focused ASL gesture demo.
Project Snapshot
Gesto needed a stronger digital surface that could explain the offer clearly, carry more visual weight, and stay structured as the product or content model evolved.
Visibility
Public case study
Year
2026
Primary Track
Interactive Experiences
Scope
Privacy-first browser ASL recognition
Challenge
The hackathon prototype needed to turn live signing into useful feedback and speech without uploading camera frames, while keeping recognition stable enough for a narrow demo vocabulary and still giving collaborators a path to collect data and train custom gestures.
Approach
Gesto keeps webcam processing and inference in the browser. MediaPipe extracts hand and face landmarks, a sequence model runs through ONNX Runtime Web for the default gesture set, and an optional TensorFlow.js workflow lets users build browser-local datasets and custom models with ONNX fallback when confidence is insufficient.
Results
What changed after the rebuild.
This project uses qualitative results and operational outcomes rather than invented vanity metrics.
Privacy model
Client-side
Webcam frames and landmarks stay in the browser for MediaPipe processing, ONNX inference, local sample recording, and TensorFlow.js training.
Recognition
Hybrid
The public demo combines a focused ONNX vocabulary with optional browser-trained TensorFlow.js labels and confidence-based fallback.
Data workflow
Portable JSON
Datasets can be recorded locally, merged or replaced through import, exported for collaboration, and paired with custom translations.
Demo
See the project in action.
A concise walkthrough of the shipped experience and its core interaction flow.
Deliverables
Real-time browser ASL gesture recognition with MediaPipe and ONNX Runtime Web
Browser-local dataset recording, JSON import/export, and custom translations
TensorFlow.js custom-model training, evaluation, and hybrid ONNX fallback
Sentence output, speech controls, accessibility preferences, and privacy-first camera handling
Engagement Tracks
This case study is intentionally light on client-identifying details. The service tracks below show the capability mix behind the final build.
Screens
Selected visuals from the case study.
Public work uses screenshots. Anonymized work uses abstracted layout boards instead of identifiable client surfaces.



Need the same level of structure, motion discipline, and delivery quality on your next project?