
About this episode
🤗 Upvotes: 24 | cs.AI
Authors:
Alexander H. Liu, Alexis Tacnet, Andy Ehrenberg, Andy Lo, Chen-Yo Sun, Guillaume Lample, Henry Lagarde, Jean-Malo Delignon, Jaeyoung Kim, John Harvill, Khyathi Raghavi Chandu, Lorenzo Signoretti, Margaret Jennings, Patrick von Platen, Pavankumar Reddy Muddireddy, Rohin Arora, Sanchit Gandhi, Samuel Humeau, Soham Ghosh, Srijan Mishra, Van Phung, Abdelaziz Bounhar, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, Alexandre Sablayrolles, Amélie Héliou, Amos You, Andrew Bai, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Anton Eliseev, Antonia Calvi, Arjun Majumdar, Arthur Fournier, Artjom Joosen, Avi Sooriyarachchi, Aysenur Karaduman Utkur, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Benjamin Tibi, Bowen Yang, Charlotte Cronjäger, Clémence Lanfranchi, Connor Chen, Corentin Barreau, Corentin Sautier, Cyprien Courtot, Darius Dabert, Diego de las Casas, Elizaveta Demyanenko, Elliot Chane-Sane, Emmanuel Gottlob, Enguerrand Paquin, Etienne Goffinet, Fabien Niel, Faruk Ahmed, Federico Baldassarre, Gabrielle Berrada, Gaëtan Ecrepont, Gauthier Guinet, Genevieve Hayes, Georgii Novikov, Giada Pistilli, Guillaume Kunsch, Guillaume Martin, Guillaume Raille, Gunjan Dhanuka, Gunshi Gupta, Han Zhou, Harshil Shah, Hope McGovern, Hugo Thimonier, Indraneel Mukherjee, Irene Zhang, Jacques Sun, Jan Ludziejewski, Jason Rute, Jérémie Dentan, Joachim Studnia, Jonas Amar, Joséphine Delas, Josselin Somerville Roberts, Julien Tauran, Karmesh Yadav, Kartik Khandelwal, Kilian Tep, Kush Jain, Laurence Aitchison, Laurent Fainsin, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Maarten Buyl, Manan Sharma, Marie Pellat, Mark Prins, Martin Alexandre, Mathieu Poirée, Mathieu Schmitt, Mathilde Guillaumin, Matthieu Dinot, Matthieu Futeral, Maxime Darrin, Maximilian Augustin, Mert Unsal, Mia Chiquier, Mikhail Biriuchinskii, Minh-Quang Pham, Mircea Lica, Morgane Rivière, Nathan Grinsztajn, Neha Gupta, Olivier Bousquet, Olivier Duchenne, Patricia Wang, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philippe Pinel, Philomène Chagniot, Pierre Stock, Piotr Miłoś, Prateek Gupta, Pravesh Agrawal, Quentin Torroba, Ram Ramrakhya, Randall Isenhour, Rishi Shah, Romain Sauvestre, Roman Soletskyi, Rosalie Millner, Rupert Menneer, Sagar Vaze, Samuel Barry, Samuel Belkadi, Sandeep Subramanian, Sean Cha, Shashwat Verma, Siddhant Waghjale, Siddharth Gandhi, Simon Lepage, Sumukh Aithal, Szymon Antoniak, Tarun Kumar Vangani, Teven Le Scao, Théo Cachet, Theo Simon Sorg, Thibaut Lavril, Thomas Chabal, Thomas Foubert, Thomas Robert, Thomas Wang, Tim Lawson, Tom Bewley, Tom Edwards, Tyler Wang, Umar Jamil, Umberto Tomasini, Valeriia Nemychnikova, Vedant Nanda, Victor Jouault, Vincent Maladière, Vincent Pfister, Virgile Richard, Vladislav Bataev, Wassim Bouaziz, Wen-Ding Li, William Havard, William Marshall, Xinghui Li, Xingran Guo, Xinyu Yang, Yannic Neuhaus, Yassine El Ouahidi, Yassir Bendou, Yihan Wang, Yimu Pan, Zaccharie Ramzi, Zhenlin Xu
Title:
Voxtral TTS
Arxiv:
http://arxiv.org/abs/2603.25551v1
Abstract:
We introduce Voxtral TTS, an expressive multilingual text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. Voxtral TTS adopts a hybrid architecture that combines auto-regressive generation of semantic speech tokens with flow-matching for acoustic tokens. These tokens are encoded and decoded with Voxtral Codec, a speech tokenizer trained from scratch with a hybrid VQ-FSQ quantization scheme. In human evaluations conducted by native speakers, Voxtral TTS is preferred for multilingual voice cloning due to its naturalness and expressivity, achieving a 68.4\% win rate over ElevenLabs Flash v2.5. We release the model weights under a CC BY-NC license.
Interactive timestamps
Jump to segmentGet every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
705 searchable segments. Every word is indexed and playable.
Full transcript
Daily Paper Cast — Voxtral TTS. Machine-transcribed; use the interactive transcript above to jump the player to any line.
0:00Hello, and welcome to the Daily Papercast. I'm Echo, and joining me today is Nova. How's it going, Nova? Hey, Echo. I'm great. Super excited about today's episode. We're diving into another fascinating research paper, straight from Hugging Faces Daily Paper List. Oh, yes. Today, we're breaking down a paper in four epic sections. We'll start with the introduction. Move on to methods, followed by experiments, and finally, we'll look at related work. Each section will take just four to five minutes, so stick around for the whole journey. Awesome. So let's get things rolling with today's spotlight paper. Echo, could you do the honors? Absolutely. The paper we're discussing today is titled Word-by-word, Vocstral TTS, a multilingual text-to-speech model using hybrid auto-regressive and flow matching architecture. Quite a mouthful, huh? Ha-ha, definitely. By the way, kudos to the first two authors, Alexander H. Lew,
1:00and Diome Lampel, and the last author, Thomas Foubert, all affiliated with Where Else, but meta AI. All right. Let's dive straight into the introduction. This paper addresses the challenge of generating high-quality speech in multiple languages. The authors point out the limitations of existing text-to-speech TTS systems, mainly that reliance on large-scale supervised data and lack of expressivity. Right, Echo. And those systems often fumble with creating natural sounding speech across different languages and dialects. Something we all notice, especially when using automated customer service lines. You know what I mean? Exactly. This paper introduces Vocstral TTS, which aims to overcome these hurdles. They've designed a hybrid model that combines auto-regressive and flow matching architectures. The goal is to synthesize expressive multilingual speech based on minimal reference audio.
2:00Interesting. So essentially, the Vocstral TTS model learns from various language data and improves its expressivity, right? What else do they mention in the introduction? Yes, Nova. They highlight that the crux of their approach is in blending the best of both worlds, the precision of auto-aggressive models with the smoothness and speed of flow matching. This allows the model to generate coherent and emotionally nuanced speech rapidly. That's pretty cool. They also mention something about the multilingual capability. Unlike many TTS systems that struggle with certain languages, Vocstral TTS is built to handle a diverse set of languages with high fidelity. Impressive, right? Absolutely. The authors emphasize that their model can effectively create speech outputs for various languages by leveraging large and varied datasets. They even mention specific techniques like voice cloning from as little as three seconds of reference audio.
3:01Whoa, just three seconds of reference audio? That's incredible. Imagine the applications, personalize digital assistance, automated translations, and more. Totally, they also delve into the practical applications. For instance, think of it helping in scenarios for language learning tools or the expressivity and emotional undertones of the speech are vital. That makes so much sense. And they're not just stopping at theory. According to the authors, Vocstral TTS is undergoing rigorous testing with both automated evaluations and human assessments to ensure its effectiveness. Right. It's clear they've put a lot of thought into making this system robust and reliable. The introduction lays a solid foundation for understanding why this new approach is both necessary and innovative. All right, folks, that wraps up the introduction section. Stick around because after a short break, we're diving into the method section where we unpack the technical magic
4:02behind Vocstral TTS. All right, folks, welcome back. So we've introduced the topic and given you some background. Now it's time to dive into the core of today's discussion, the methods behind Vocstral TTS. Nova, you ready to unpack this? Absolutely, Echo. The methods section is always where the magic happens. It's like showing how the sausage is made, you know? Anyway, the Vocstral TTS system is divided into several key components, each of which is intricately designed to work together seamlessly. So let's get started with the overall architecture. Cool, let's do it. What's the first piece of the puzzle? First up, we have the audio tokenizer. This part of the system breaks down the 24 kilohertz mono input waveform into manageable, non-overlapping patches of 240 samples each. This essentially converts the raw audio into 100 hertz input frames that can be processed by the encoder. These frames are then projected
5:03into 1024 dimensional embeddings using a causal convolution layer. All right, so we start with some heavy duty processing. What happens next? These embeddings go through a series of four encoder blocks, each made up of a two layer causal self-attention transformer and a causal CNN layer. Now, each transformer layer employs techniques like sliding window attention, A-liby positional bias, QK norm, and layer scale initialized at 0.01. So it's like pulling out all the stops to get the best possible representation. Uh-huh, they're pulling out all the nerdy stops, aren't they? Totally. And as all this processing happens, the CNN layers downsample by a factor of two in the first three blocks, resulting in an overall eight tons reduction from 100 hertz down to 12.5 hertz. Finally, the fourth block projects the 1024 dimensional representation to a 292 dimensional latent space.
6:04Okay, so we now have a reduced high dimensional representation of the audio. What's next? This 292 dimensional latent is then quantized into audio tokens. These tokens are split into a 256 dimensional semantic component and a 36 dimensional acoustic component. The semantic component uses a learned vector quantizer with a code book of size 8192, while the acoustic component is quantized using finite scalar quantization to 21 uniform levels. Ah, so each of these components has its own way of being quantized, but why the different approaches? Good question. The idea is to optimize the processing of each type of data. The semantic token benefits from a larger code book for better linguistic representation. While the acoustic tokens smaller levels are efficient for capturing the nuances of sound. Makes sense. So once we have these tokens, what happens then? From here, the decoder mirrors the encoder,
7:04but in reverse. It takes these tokens and reconstructs the waveform. The semantic and acoustic tokens are reprojected back into higher dimensional embeddings, using a causal CNN and a series of transposed CNNs for upsampling. Finally, the reconstructed waveform is produced. Wow, that's a lot of back and forth. I assume that's not all, right? There must be more to bring this all together. Absolutely. The semantic token learning is enhanced using an auxiliary automatic speech recognition, ASR, distillation loss. This means a frozen whisper model runs on the input audio to generate decoder hidden states and cross-attention weights. These are then aligned to the projected post-VQ semantic embeddings using a cosine distance loss. So essentially, the system learns better semantic representations. Exactly. And this alignment is derived from whisper's cross-attention weights, which means the semantic tokens are text-aligned
8:06without needing external forced aligners. That's neat. But what about generating the actual audio output? For that, there's a flow matching transformer, which is crucial for predicting the acoustic tokens. This transformer operates independently on the hidden states and models the velocity field required to transform Gaussian noise into a acoustic embeddings. And yes, each decoding step uses this flow matching transformer to maintain a fully discrete token interface. Hold on a second. You're saying they use Gaussian noise, like literal white noise? In a way, yes. They use Gaussian noise as a starting point and apply the velocity field to change it into meaningful acoustic embeddings. It's like transforming raw chaos into delightful order. Wow. What a visual. Wow. So do they do something special for training all these models? Oh, definitely. The training uses paired audio and transcripts pseudo-labelled with voxtral mini transcribe.
9:07Each training sample consists of a tuple, voice reference, transcript, and target audio. To make things more interesting, these samples are interleaved with special tokens to signify different segments. Special tokens, huh? That sounds important. Tell me more. These tokens essentially guide the model on how to transition between segments. For instance, a next token indicates the transition from the voice reference to the transcript and a repeat token marks the transition from the transcript to the target audio. These transitions help the model understand the context better. Got it. They're essentially guiding the learning process. Now, what about the loss functions? Any special tricks there? Yes, indeed. They use a two-part loss function, a cross-entropy loss for the semantic tokens, and a flow matching loss for the acoustic tokens. This makes sure that both linguistic and acoustic information is learned effectively. And I bet they do something unique to prevent overfitting, too, right? Absolutely.
10:07To avoid overfitting to silence, they wait the loss lower for frames that contain no speech based on a voice activity detection model. They even set the loss weight to zero for extremely long silences. Wow. They've thought of everything. So what about the post-training phase? Any special techniques there? Yes. They use direct preference optimization, DPO to fine-tune the model, focusing on improving metrics like word error rate, WER, and speaker similarity. The cool part is they adapt the DPO objective to their autoregressive setup by computing differential velocities for winners and losers in the training data. Differential velocities. Can you break that down a bit? Sure thing. They calculate the difference in the predicted velocities between pairs of winner and loser samples. This helps in reinforcing the correct predictions and penalizing the incorrect ones. Plus, they use a low learning rate for better stability during this fine-tuning process.
11:08That's genius. So what's next after all this training and fine-tuning? Next comes the inference phase. They use techniques like cuda-graph acceleration to reduce latency and increase throughput. This part is especially critical for real-time applications. Real-time applications, huh? That's where things get really practical. How do they ensure it's low latency? To achieve this, they leverage cuda-graphs. During initialization, they perform an eager warm-up pass for each bucket size, capturing the cuda-graph. During inference, the cuda-graph is replayed, which leads to significant improvements in latency and real-time factors. Very cool. So how effective is this cuda-graph acceleration? Super effective. They report a 47% improvement in latency and a 2.5x reduction in the real-time factor. That's huge for applications requiring speed. I bet. And that's crucial for applications
12:09like voice assistance and real-time translations, right? Exactly. That's the massive advantage here. Plus, with asynchronous chunked streaming, they manage to overlap the generation stage with the decoding stage, further reducing the time from input to audible speech. So, faster and smarter, I love it. And I guess that wraps up the method section, huh? Yep, that's the end of the method section. Quite a ride, wasn't it? Absolutely. I hope our listeners are as fascinated as we are. Up next, we'll dive into the experiments and results. Stay tuned. All right, Echo, continuing our deep dive on this fascinating paper. Let's talk about the experiments and results. I think a lot of our listeners are eager to hear how well the proposed techniques performed in various evaluations. Oh, definitely. This is the juicy part, right? There's so much to unpack here. The research team conducted a series of experiments to evaluate the performance of their VoxDraw TTS system
13:10across multiple languages and tasks. The results section is packed with info. Let's kick it off with the automated evaluations. Okay, so the paper mentions they evaluated VoxDraw TTS against 11 Labs V3 and 11 Labs Flash V2.5 using automated metrics like word error rate, WER, U-T-M-O-S-V2 and speaker similarity. Any surprises there? Actually, there were quite a few interesting findings. VoxDraw TTS significantly outperformed the 11 Labs models, particularly in speaker similarity metrics. However, there was a twist while 11 Labs Flash V2.5 performed better on most automated metrics, 11 Labs V3 excelled in human evaluations, especially with emotion steering. That's fascinating. It really underscores the importance of combining automated evaluations with human assessments, doesn't it? It's like trying to get the full picture by looking at both. Oh, I don't know. The stats and the game itself.
14:12Exactly. They even did extensive human evaluations to capturing nuances like naturalness, expressivity and emotional conveyance. Areas where automated metrics might fall short. So, let's break down one of their detailed tables, shall we? For instance, let's take the word error rate, WER, across languages and table three. It looks like VoxDraw TTS was pretty consistently low in WER compared to the other models, right? You got it. WER across different languages showed that VoxDraw TTS had the lowest error rates in almost all cases. For instance, in Arabic, it had a WER of 2.68%, compared to 3.67% for 11 Labs V3, and 2.86% for 11 Labs Flash. Wow, that's quite impressive. I see they also evaluated with UTMOS and speaker similarity. UTMOS is about predicting the mean opinion score of generated speech, right?
15:14Yep, you got that right. It's supposed to give an overall quality score. And interestingly, the paper mentions that VoxDraw TTS's UTMOS scores were mostly competitive, but not always the best across languages. However, it's shown in speaker similarity scores. That's pretty cool. So, they did human evaluations too. Tell me more about that part. Sure. They conducted two sets of human evaluations. One was on flagship voices with explicit and implicit steering for emotions. Explicit steering is where the instructions are more direct, like speak in an angry tone. Implicit steering lets the model infer the emotion from the text like, this is the best day of my life. Oh, that sounds intriguing. Any results that stood out in these evaluations? VoxDraw TTS was quite competitive. It had a win rate of 51% against 11 Labs V3 in explicit steering and even better win rates
16:14in the implicit steering scenarios. Over 58% against both 11 Labs Flash and V3. 4-0 times source. That's impressive, especially in the implicit setting where detecting the right emotion from the text is harder. I guess it shows how versatile VoxDraw TTS is, huh? Exactly. They even evaluated zero-shot voice cloning capabilities with high quality audios from recognized speakers in each language. VoxDraw TTS had an overall win rate of 68.4%, compared to 11 Labs Flash V2.5. It excelled in both high and low resource languages. It sounds like VoxDraw TTS is really setting the bar high for zero-shot cloning. I mean, being able to do that across both high and low resource languages is really something. Absolutely. The detailed analysis also showed improvements between the pre-trained and DPO checkpoints, particularly in reducing word amers, and enhancing naturalness and loudness consistency.
17:15So it sounds like the DPO post-training gave a significant boost to the model's performance, particularly in those tricky areas, always fascinating how small tweaks and optimizations can yield big results. Totally. They also experimented with different values of NFEs, number of functional evaluations, and found that increasing NFEs to eight significantly improved metrics, like speaker similarity and upmoss. But beyond that, the improvements tapered off. Oh, I see. So tweaking the NFEs was like hitting a sweet spot for performance. How did they balance this with the CFG or classifier free guidance? Good question. They found that higher CFG values generally improved most metrics, but could sometimes lead to the model being overly strict with the voice prompt, thus failing to convey the intended emotion from the text. So setting the CFG value was all about finding that balance.
18:15It's all about balance, isn't it? Anything else notable in this section we should share with our listeners? Well, they also compared the latency and real-time factor when using cuda-graph acceleration for the flow matching transformer. Result of 47% improvement in latency and a 2.5x reduction in real-time factor. Wow, that's a significant performance boost. That must have really optimized the whole TTS pipeline. Totally. They even tested the serving performance on a symbol Nvidia H200 GPU and found it could handle over 30 concurrent users with sub-second latency. The sort of feature you need for real-world deployment. That's a wrap on the experiments and results section. It's evident how comprehensive their experiments were, covering both automated and human evaluations and fine-tuning various parameters. Really thorough work. Absolutely, Nova. And it's not just about numbers. They've shown real-world applicability and robustness.
19:16Definitely exciting times in the world of TTS. All right, Nova. Let's dive into the related work section of our featured paper today. There's quite a bit here to unpack. Ready? Absolutely, Echo. This part is crucial because it helps us understand the context and background this paper is building upon. So where do we start? Right. Let's kick things off by looking at how this paper positions itself against recent advancements in zero-shot TTS systems. These systems typically condition generation on discrete speech tokens extracted from a short-voice prompt. This enables generalization to unseen speakers and natural synthesis across long sequences. Exactly. These tokens are crucial because they allow the model to understand and generate speech that it hasn't explicitly seen during training. It's kind of like how we can mimic accents or phrases after hearing them just a few times. Exactly. The paper also discusses various modeling
20:19approaches for rich acoustic variations in speech generation, like diffusion and flow-based models. For instance, works by Popoff Ed Allen, 2021, and Lay Ed Allen, 2023, have been pivotal in exploring these models. Interesting. So these models are essentially trying to capture the nuances and variations in speech, right? Which is really important for making text-to-speech sound natural and expressive. Yes, definitely. It's all about capturing those subtle details. The paper also highlights recent advances in speech codex, which can factorize speech into a low-rate semantic stream and a higher rate acoustic stream. This structure is already being exploited by hierarchical generators, like Moshe. Ah, I see. This kind of factorization helps in breaking down the complexity of the speech signal into more manageable parts for the model to handle, right? So what's the main innovation introduced by the current paper?
21:19The main innovation is the voxtral TTS system that combines a representation-aware hybrid architecture. It uses voxtral codec, a low-bit rate speech tokenizer. This codec tokenizes a voice prompt into semantic and acoustic tokens, enabling the model to generate high-quality, low-latency speech across multiple languages. Wow. That's quite sophisticated. They've managed to combine different strengths from various modeling approaches into one cohesive system. And how is it compared to the existing models, like 11 labs and others? Good question, Nova. The paper mentions that voxtral TTS outperforms 11 labs V3 on speaker similarity scores. It also shows a significant win rate in human evaluations for zero-shot voice cloning and expressive speech synthesis, making it very competitive. That's impressive. It's one thing to perform well in automated metrics, but human evaluations are a true test of effectiveness,
22:20especially when it comes to naturalness and expressivity in speech. Exactly. And it's important to note that automated metrics while useful can't always capture the full spectrum of human speech expressivity and naturalness. That's why human evaluations are crucial. Did the paper mention any specific challenges related to these evaluations? Yes, indeed. One challenge highlighted was ensuring that the model remains consistent and natural across different languages and speaking styles. This is where voxtral TTS shows its strength, maintaining high performance even in multi-lingual settings. Multi-lingual capability is a big deal, especially in our increasingly globalized world, speaking of which, how does voxtral TTS handle the different languages? Great point, Nova. voxtral TTS supports nine languages and can generate speech from voice props as short as three seconds. It uses a hybrid architecture that includes both autoregressive
23:22and flow matching models to handle the semantic and acoustic tokens effectively. That's quite efficient. Being able to work with just three seconds of voice data is remarkable for achieving zero shot learning tasks. It's truly an advancement. Absolutely. The paper also mentions adapting direct preference optimization, DPO to this hybrid model, combining traditional preference objectives with flow-based preference ones. This adaptation helps in fine-tuning the model for better performance. Direct preference optimization reminds me of ongoing efforts in reinforcement learning, where models are trained based on feedback from preferences. It's fascinating to see this concept being applied here. Exactly. This method has been effectively adapted to improve both the word error rate and the speaker similarity metrics, enhancing the overall performance of the model. So, echo, did the authors mention any specific influences or prior works that heavily inspired their current model?
24:23Yes, they cited quite a few. For instance, the Vox Straw Codex concept of using a supervised ASR model to distill semantic representations was influenced by the works of Vashished at all in 2024 and others. This approach tends to produce more effective and meaningful semantic tokens compared to self-supervised methods. I see. So it's like building on the best practices found in previous successful methods and improving from there. Smart move. Absolutely. The field is all about standing on the shoulders of giants and pushing me envelope further. By integrating these innovative methods, they've developed a system that's more robust and capable of producing high-quality speech across different use cases. It's indeed a collaborative and iterative process. New research often opens up even more possibilities for future work. Do they discuss any future directions? They do. One of the directions includes enhancing
25:24the expressiveness and naturalness of TTS models further. They also talk about expanding the system's capabilities to support more languages and dialects, which is always an exciting frontier. That is exciting. The broader the linguistic support, the more accessible and useful these models become to people around the world. Exactly, Nova. They also mention the importance of continuous improvements through both automated and human evaluations to keep refining these models. It's all about incremental gains and learning from each step. And there we have it, folks. A deep dive into the related work section for this remarkable paper on VoxStrial TTS. Understanding where we come from is crucial to where we're going. Wouldn't you agree? Absolutely, Nova. It's fascinating to trace the lineage of ideas and see how each piece of research builds upon the last. Stay tuned as we continue our discussion in the next segment. Hey, everyone. Welcome back.
26:25Before we wrap up today's episode, let's do a quick recap of the key contributions and takeaways from the paper we discussed. Absolutely, echo. So the paper introduced VoxStrial TTS, an expressive multilingual text to speech model. Now, what's cool about this model is that it generates natural speech from just three seconds of reference audio. Can you believe that? I know, right? It's incredible. VoxStrial TTS combines autoregressive generation of semantic speech tokens with flow matching for acoustic tokens. And these tokens are encoded and decoded using VoxStrial codec. Exactly. And they didn't just stop there. The model was rigorously evaluated both automatically and through human evaluations. Native speakers preferred VoxStrial TTS from multilingual voice cloning because of its naturalness and expressivity. In fact, it had a win rate of 68.4% over 11 labs flash V2.5. That's pretty impressive.
27:27Yeah. And one of the standout features is its ability to perform zero shot voice cloning. That means the model can generate fairly accurate speech from voices it hasn't explicitly been trained on just by using a short audio reference. How cool is that? Super cool. And remember, they released the model weights under the CC BYNC license, which means it's open for further research and development. It's an exciting step forward for the community. Totally. This open access could lead to even more advancements in expressive TTS systems. It's a big win for researchers and developers who want to build on this work. And who knows? Navy will see even more natural sounding virtual assistance in the future. Imagine having an assistant that sounds just like your favorite actor or even like a loved one. That's where this technology can lead us. Absolutely. It's all about making interactions with technology more personal and engaging. Before we sign off, we want to thank you all for tuning in
28:28and joining us on this deep dive into Voxthrill TTS. Yes. Thank you so much. We hope you found this episode enlightening and fun. Don't forget to subscribe so you won't miss out on our future episodes where we'll continue to bring you the latest and greatest from the world of AI research. And if you have any papers you'd like us to cover, drop us a line. We're always eager to hear from our listeners. Until next time, keep questioning, keep learning, and stay curious. That's right. This is Nova and Echo signing off. Catch you in the next episode of the Daily Papercast. Bye, everyone. Bye.
More episodes
More from Daily Paper Cast

Seedance 2.0: Advancing Video Generation for World Complexity
Daily Paper Cast

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Age...
Daily Paper Cast

RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Tes...
Daily Paper Cast

SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Envir...
Daily Paper Cast