Applied ML Engineer, Audio
We’re a team of audio experts solving one of the hardest problems in speech technology. Real conversations are messy. People interrupt each other, move around, sit at different distances from the microphone, and speak in rooms filled with reverb and background noise.Most models are trained on synthetic mixtures because high-quality paired recordings of real conversations are extremely difficult to collect. Models that perform well in simulated conditions often struggle when they meet the real world. The solution:We’ve built technology that captures both the room and the individual speakers at the same time. This has given us something rare: a proprietary dataset of real conversations with the reference signals needed to train and evaluate a new generation of speech enhancement, separation, dereverberation, and spatial audio models. The RoleWe’re looking for an exceptional Applied ML Engineer to turn this advantage into audio improvements people can immediately hear. You’ll own the journey from data to production building datasets and listening tests, training models, investigating failures, and shipping improvements to real users. A few things we’re looking for:Hands-on experience training speech or audio modelsStrong PyTorch skillsExperience with real-world, imperfect audio not only benchmark datasetsA solid understanding of evaluation, including perceptual listening testsThe ability to take models from experimentation into productionA first-principles mindset and genuine curiosity about sound Experience with speech enhancement, source separation, dereverberation, room acoustics, microphone arrays, ambisonics, or spatial audio would be particularly valuable. If this problem fascinates you as much as we are obsessed, please apply and we will get back to you right away.
