About PhonicPhonic is a product and research lab focused on powering the most realistic, human-like voice AI conversations. We've re-thought the entire stack in pursuit of this goal, from models to product, to create voice agents that feel like they truly understand you, respond emotionally and perform agentic tasks with frontier intelligence.Our team includes top-tier AI researchers, international olympiad medalists, and former founders.Our customers include companies that are building voice-native AI products in industries such as customer support, healthcare, and logistics. We have raised over $30M from tier 1 VCs.About The RoleAs a Research Engineer at Phonic, you'll sit at the intersection of cutting-edge ML research and production engineering, working directly on the core systems that make Phonic's voice AI feel genuinely human.
You'll design, train, and iterate on models across the voice stack (speech, audio, language, and beyond), while also building the infrastructure and tooling needed to move fast from research idea to deployed product.This is a role for someone who thrives in ambiguity, moves with urgency, and takes full ownership of problems. You'll work closely with researchers and product engineers in a high trust, in-person environment in our SF office to push the frontier of what voice AI can do.What You'll DoDesign, implement, and iterate on models across the voice stack from audio and speech to language and beyondBuild the training pipelines, evaluation frameworks, and tooling that let us experiment and iterate quicklyTranslate research results into production-grade systems alongside our engineering teamForm hypotheses, chase down interesting results, and take ownership over the problems you work onWhat You'll BringHands-on experience in ML research or research engineering, industry or PhD-level academiaProficiency in Python and PyTorch (or JAX), and the ability to implement models cleanly from papersExperience running ML experiments end-to-end: data processing, training, evaluation, and iterationComfort in fast-moving, ambiguous environments - defining the problem is part of the jobThe ability to clearly explain what you tried, what worked, and whyNice To HaveResearch experience in speech, audio, or language modeling (ASR, TTS, LLMs, codec models)Familiarity with diffusion, flow matching, or autoregressive generative modelsExperience with distributed training, quantization, or inference optimizationCompetitive programming or olympiad backgroundBenefitsTop-tier compensation: in order to get the best talent, we provide salary and equity that recognize your skillsetMeals: free breakfast, lunch, and dinner provided in the officeHealthcare: Comprehensive health, dental, and visionWe have regular off-sites and team celebrations401(k) – Let us help you plan for the future. We've got you covered.