
Head of ML
- calibration
- Temporal
- Language Models
- Python
- PyTorch
- FPGA
- Machine Learning
- CUDA
About The Subvocal Company
Talk to your computer without talking
Tech
At a high level, we are trying to detect the incredibly subtle physiological changes that happen when someone forms words internally, and turn those signals into continuous language. The broader sensing design space includes RF sensing, EMG, EEG, mmWave, and other non-invasive methods. For competitive and IP reasons, we are not publicly disclosing the exact architecture of our current system yet, although we can share much more during the interview process. The hard part is not just getting a model to work on one person in one recording session. These signals can change when the device moves slightly, when the same person comes back the next day, or when you move to someone with completely different anatomy. We need the system to work across people, devices, and environments, then adapt to a new user from only a few minutes of calibration. Solving that involves a mix of signal processing, self-supervised learning, large temporal models, personalization, and language decoding. We currently experiment with architectures including Conformers, Mamba-style sequence models, and pretrained speech and language models. Most of our ML stack is built in Python and PyTorch, with custom infrastructure for data collection, signal processing, distributed training, and evaluation. At the same time, we are taking a research system built from laboratory equipment and turning it into a wearable. That means building custom sensing and mixed-signal electronics, embedded and FPGA systems, custom ASICs, high-speed data acquisition, firmware, power systems, and eventually fitting everything into a small wearable. Sensor placement, electrical design, mechanical design, and the model are all tightly connected. Moving something by a few millimeters or changing part of the electronics can alter the data distribution enough to affect the model. A large part of our next phase is also a data problem. We are building the infrastructure to collect thousands of hours of high-quality subvocal speech data across thousands of people, train models continuously as that dataset grows, and understand exactly how performance scales with more users and more variation. This is what makes the work unusually interesting. There is no established playbook, and the important breakthroughs can come from a better model, a better sensing configuration, a clever piece of hardware, or simply discovering that we have been framing the problem incorrectly. The people joining now will have room to work across those boundaries and make decisions that directly determine whether the system works.
The role
We are looking for the person who will own the central machine learning problem at Subvocal: turning weak, noisy, highly variable physiological signals into continuous language. We have already built prototypes that decode subvocal speech at more than 200 words per minute. The much harder problem now is generalization. A model that works on one person, in one session, with one device placement is not a product. It needs to work when the same person returns the next day, when the hardware moves slightly, and eventually when a completely new person puts it on and gives us only a few minutes of calibration data. You will lead that effort end to end. You will work directly with the founders and our sensing and hardware leads to decide what data we collect, how we represent the signal, which model families we pursue, and how we evaluate whether we are actually making progress. Some of the problems you will work on include: * Learning general representations from large amounts of unlabeled and weakly labeled physiological time-series data. * Building continuous sequence models using approaches such as Conformers, Mamba-style architectures, CTC, transducers, and pretrained speech or language models. * Separating speech-related information from anatomy, placement, session, device, and environmental variation. * Adapting a large cross-user model to a new person from roughly 15 minutes of calibration data. * Designing honest user-held-out, session-held-out, and device-held-out evaluations. * Scaling training across thousands of hours and thousands of people. * Getting the complete system to run continuously with low enough latency for real-time use. Our current ML stack is primarily Python, PyTorch, CUDA, and distributed GPU training, with custom infrastructure for signal processing, data collection, experiment tracking, and evaluation. You might be a great fit if you have unusually strong experience in deep learning for speech (ASR), time series, biosignals, neuroscience, BCIs, radar, or another domain where signals are noisy and data distributions shift constantly. We are especially interested in people with PhD-level research ability, whether or not that came through a formal PhD, who are also comfortable writing production-quality code and moving quickly when the research direction changes. This is not a role where you will be handed a model architecture and asked to improve it incrementally. You will help decide how the problem should be framed in the first place, build the initial ML organization around you, and directly determine whether this technology becomes a real product. This is a full-time, in-person role in San Francisco.
Skills
- Machine Learning
- Speech Recognition
Head of ML Β· The Subvocal Company