Sign language → text · in your browser
Watch your hands become words.
SL2T tracks your hands, face, and body through your webcam and turns six ASL signs into live English text. The whole pipeline — tracking, recognition, translation — runs on your device. No uploads, no account, no server.
Needs a webcam · ~18 MB of models on first load, cached after · best in Chrome
Translation
HELLOTHANK YOU
How it works
Camera in. Words out. Nothing in between.
Camera on
Grant camera access and face the lens. Video stays inside the tab — it is never recorded, stored, or sent anywhere.
553 points, tracked live
MediaPipe maps your hands, body, and face as a moving constellation of landmarks — on your device, at webcam frame rate.
Motion becomes text
A recognizer reads each sign's shape and movement — a bobbing fist, a palm circling the chest — and streams the matching word onto the page.
Vocabulary
Six signs, recognized live.
Small on purpose. Each sign is detected from real motion — position, shape, and movement over time — not a single frozen pose.
Privacy
Your camera feed never leaves your device.
The models download once (about 18 MB, then cached) and every frame is processed and discarded in memory. Nothing from your camera is uploaded, recorded, or analysed elsewhere — close the tab and it's gone. DeepMind's SL2T demo streams video to a server for inference; this independent demo runs the entire pipeline locally, so it never has to.
The dictionary is a separate thing and does use a server. Entries and their demonstration recordings are stored, every recording names its signer and the consent it was given under, and consent can be withdrawn — which deletes the recording, not just the link to it.