
Company Overview
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are 'Powered by Deepgram', including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box.
Company Operating Rhythm
At Deepgram, we expect an AI-first mindset—AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance. Every team member is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work.
The Opportunity
Voice is the most natural modality for human interaction with machines. However, current sequence modeling paradigms based on jointly scaling model and data cannot deliver voice AI capable of universal human interaction due to the fundamental data problems posed by audio.
The Role
As a Member of the Research Staff, you will pioneer the development of Latent Space Models (LSMs), a new approach that aims to solve the fundamental data, scale, and cost challenges associated with building robust, contextualized voice AI. Your research will focus on solving one or more of the following problems:
- Build next-generation neural audio codecs that achieve extreme, low bit-rate compression and high fidelity reconstruction across a world-scale corpus of general audio.
- Pioneer steerable generative models that can synthesize the full diversity of human speech from the codec latent representation.
- Develop embedding systems that cleanly factorize the codec latent space into interpretable dimensions of speaker, content, style, environment, and channel effects.
- Leverage latent recombination to generate synthetic audio data at previously impossible scales.
- Design model architectures, training schemes, and inference algorithms that are adapted for hardware at the bare metal.
Requirements
- Strong mathematical foundation in statistical learning theory, particularly in areas relevant to self-supervised and multimodal learning.
- Deep expertise in foundation model architectures, with an understanding of how to scale training across multiple modalities.
- Proven ability to bridge theory and practice—someone who can both derive novel mathematical formulations and implement them efficiently.
- Demonstrated ability to build data pipelines that can process and curate massive datasets while maintaining quality and diversity.
- Track record of designing controlled experiments that isolate the impact of architectural innovations and validate theoretical insights.
- Experience optimizing models for real-world deployment, including knowledge of hardware constraints and efficiency techniques.
- History of open-source contributions or research publications that have advanced the state of the art in speech/language AI.
Timezone overlap
UTC-8–-4
Benefits
Open to
US · San Francisco · United States · Ann Arbor
Sign in to track applications and earn points.