Sanket: Indian Sign Language recognition that never uploads your video
A 4.6 MB model that reads signs from a webcam at 12.5 ms per window, fully offline in the browser.
- What
- My own product · M.Tech AI case study, SRM
- Role
- Researcher and engineer
- When
- Jul 2026 – present
- per window, offline in the browser, with a 4.6 MB model
- 12.5 ms
- words in an open vocabulary
- 8,165
- accuracy on signers the model had never seen
- 33% → 50%
The situation
Sign-language recognition usually means uploading video to a server. Sanket runs entirely on the device, so nothing leaves the browser, and it has to work for people who were not in any training set.
What I built
- MediaPipe landmarks feed a 4.6 MB int8 ONNX model in a Web Worker: 12.5 ms per window, offline.
- An open vocabulary of 8,165 words through an embedding bank instead of a fixed classifier.
- Audited the deployed model on signers it had never seen: 82% on its own corpus fell to 33%. Rebuilt it around a self-supervised (JEPA) encoder and synthetic signers, which lifted unseen-signer accuracy to 50%.
- Caught one benchmark leaking into an older model's training data, and moved all model ranking to a split no model had seen.
- GPU pipelines on Modal: a corpus re-extracted in about 15 minutes for about $1.50, against an estimated 3.5–4 hours elsewhere.
Stack
- PyTorch
- ONNX Runtime Web
- MediaPipe
- Self-supervised learning
- TypeScript
- Web Workers
- Modal