Start a Project

Applied AI Case Study

Gesto — Sign to Speech

A privacy-first sign-to-speech web application that combines MediaPipe landmarks, ONNX Runtime Web, browser-local TensorFlow.js training, dataset portability, and accessible speech output for a focused ASL gesture demo.

Project Snapshot

Gesto needed a stronger digital surface that could explain the offer clearly, carry more visual weight, and stay structured as the product or content model evolved.

Interactive Experiences

Visibility

Public case study

Year

2026

Primary Track

Interactive Experiences

Scope

Privacy-first browser ASL recognition

Challenge

The hackathon prototype needed to turn live signing into useful feedback and speech without uploading camera frames, while keeping recognition stable enough for a narrow demo vocabulary and still giving collaborators a path to collect data and train custom gestures.

Approach

Gesto keeps webcam processing and inference in the browser. MediaPipe extracts hand and face landmarks, a sequence model runs through ONNX Runtime Web for the default gesture set, and an optional TensorFlow.js workflow lets users build browser-local datasets and custom models with ONNX fallback when confidence is insufficient.

Results

What changed after the rebuild.

This project uses qualitative results and operational outcomes rather than invented vanity metrics.

Privacy model

Client-side

Webcam frames and landmarks stay in the browser for MediaPipe processing, ONNX inference, local sample recording, and TensorFlow.js training.

Recognition

Hybrid

The public demo combines a focused ONNX vocabulary with optional browser-trained TensorFlow.js labels and confidence-based fallback.

Data workflow

Portable JSON

Datasets can be recorded locally, merged or replaced through import, exported for collaboration, and paired with custom translations.

Demo

See the project in action.

A concise walkthrough of the shipped experience and its core interaction flow.

Gesto — Sign to Speech demo

Deliverables

Real-time browser ASL gesture recognition with MediaPipe and ONNX Runtime Web

Browser-local dataset recording, JSON import/export, and custom translations

TensorFlow.js custom-model training, evaluation, and hybrid ONNX fallback

Sentence output, speech controls, accessibility preferences, and privacy-first camera handling

Engagement Tracks

This case study is intentionally light on client-identifying details. The service tracks below show the capability mix behind the final build.

Interactive Experiences

Screens

Selected visuals from the case study.

Public work uses screenshots. Anonymized work uses abstracted layout boards instead of identifiable client surfaces.

Gesto application overview from the project README.
Application overview
Gesto starter dataset workflow from the project README.
Starter dataset workflow
Gesto browser-local training view from the project README.
Browser-local model training

Need the same level of structure, motion discipline, and delivery quality on your next project?