DeepL AI Labs
Bringing Mixhalo to the DeepL platform hasn’t just added new capabilities around real-time speech-to-speech translation at scale. It’s transforming the scope of what we can build in voice. When you’re able to control the speed of audio delivery, and remove it as a constraint, the previously impossible can quickly become a product roadmap.
Together, our teams are building a spoken language layer that’s available for every type of experience and interaction. Through DeepL Voice API, they’re empowering innovators everywhere to build live, real-time, multilingual voice into every kind of customer and audience journey.
That includes capturing all the emotion of a football commentator describing a goal or a touchdown, instantly accessible voice experiences that remove the feeling of isolation for non-native speakers in healthcare settings, careful design of customer support call experiences to eliminate stress and frustration, and a multilingual experience of live events that delivers translated speech far faster than human interpreters can.
Mixhalo Co-Founder, Vik Singh, and DeepL Head of Voice, Leo Doin, have been coming at the challenge of reducing latency for multilingual audio from two complementary directions. Mixhalo was born in the music industry, streaming the sound of a drummer hitting a snare to a stadium full of thousands of fans, faster than the soundwaves themselves could reach them. DeepL’s language models are able to handle the translation of real-time audio, end-to-end, with no handoffs to different APIs, bringing the latency of translation to its lowest possible level.
Bespoke audio capture daemons, custom network protocols, real-time error detection, audio engines optimized for speed in playback: all of these technical innovations create time and space to design a more optimized experience of live, audio translation. Add in language intelligence and that optimization can happen instantly, through models that know the best time to translate, and how to synchronize languages with different grammatical structures.
The ideal live experience of translation doesn’t mean playing translated audio instantly. It’s about adjusting playback intelligently, through AI that deeply understands how each language works.
This all comes together in live speech-to-speech translation for call centers. It’s one of the most powerful applications of DeepL Voice API. Vik and the Mixhalo team are engineering it ensure customers always feel a sense of connection to support agents, and that translation never compromises the feeling that they’re being heard.
Perhaps the greatest impact of bringing DeepL and Mixhalo together will be felt across live events: conferences with thousands of attendees and hundreds of speakers, all able to share ideas and follow sessions in their preferred language, through audio that’s streamed to their devices, across the venue.
It’s in scenarios such as this that the value of a real-time, multilingual voice layer becomes truly apparent. Ultra-low latency audio means that attendees can listen to the speaker of their choice as they move around. They can follow sessions even when they’re not in the room. Live, instantaneous speech translation means that every attendee hears the language they understand best, with all the emotion and implicit expression captured. It also means that every speaker can share their ideas in the language they formed them
Perhaps most striking, though, will be the speed at which all of this happens. We’re developing innovative new ways for DeepL’s AI models to capture and ingest context around keynotes, panel discussions and other speaker sessions, in much the same way that live event interpreters prepare in advance. When this is combined with AI’s greater processing speed, we’re able to generate live, speech-to-speech translations far faster than human translators are capable of. This will reset expectations about what the experience of multilingual content shall be. It’s going to feel futuristic, magical even. And it will start to transform expectations for how spoken communication works.