
About Deepgram
Deepgram is the leading platform underpinning the emerging Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and production-grade voice agents. More than 200,000 developers and 1,300+ organizations build voice offerings powered by Deepgram, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box.
Company Operating Rhythm
At Deepgram, we expect an AI-first mindset—AI use and comfort are core to how we operate, innovate, and measure performance. Team members actively experiment with advanced AI tools and integrate them into everyday workflows. We move at the fast pace of AI, so adaptability and continuous learning are essential.
About the Role
Deepgram's speech models run at scale on NVIDIA GPUs, but customers increasingly need those same models on non-NVIDIA accelerators, edge servers, and embedded platforms. Getting Deepgram models onto target hardware with minimal model changes and zero hardware paradigm shifts is the core objective of this role.
As an Applied ML Engineer on the Partner Platform Engineering team, you will adapt existing Deepgram models to run correctly and efficiently within a target platform's existing kernel and runtime paradigm (swapping operators, rewriting graphs, adjusting precision, and validating performance on physical devices).
What You'll Do
- Port Deepgram speech models to non-NVIDIA and edge platforms, adapting model structure and parameters to fit target runtimes and operator sets.
- Own serving-side model decisions for edge targets, including quantization, precision choices, operator substitution, and architecture tweaks while holding accuracy and latency.
- Validate ports on real hardware by building benchmark suites for accuracy, latency, throughput, and memory consumption.
- Build deployment pipelines for model packaging, conversion, versioning, and automated delivery to edge targets.
- Collaborate with Embedded AI Engineers when custom low-level kernels are required.
- Partner with silicon and platform vendors on runtimes, toolchains, and model format conventions.
- Feed edge hardware constraints back to Research and Engineering teams so future models are built to port easily.
Requirements
- Hands-on experience deploying ML models to edge or non-NVIDIA hardware in production (Cloud-only or GPU-only serving experience does not qualify).
- Working knowledge of quantization and precision tradeoffs (INT8, FP16, mixed precision, calibration) and their performance impacts on physical hardware.
- Experience with edge/vendor inference runtimes and conversion toolchains (e.g., ONNX Runtime, TFLite, ExecuTorch, OpenVINO, Qualcomm AI Engine, NPU SDKs).
- Proficiency in modifying model architectures: graph rewrites, operator swapping, and structural adjustments without breaking accuracy.
- Strong Python and PyTorch skills with solid production engineering practices (tests, reproducibility, clean architecture).
- Comfort building automated pipelines around model conversion and deployment.
Nice to Have
- Experience with speech, audio, or streaming/real-time AI models.
- Exposure to low-level kernel code (CUDA, Metal, NEON, DSP assembly).
- Knowledge of edge model security, signature verification, and encrypted storage.
- Familiarity with multiple accelerator families (Qualcomm, Apple, ARM, Intel, AMD, custom NPUs).
Timezone overlap
UTC-8–-4
Open to
US
Sign in to track applications and earn points.