It's such a pleasure to be here. As mentioned, my name is Kayla Nahi. I am a product lead at Mistral. We are a European foundational AI company, so we build models in Europe.
The company was started about three years ago with the goal of creating AI that was European in response to the Americans. Americans, ignore the accent for a moment, and then put a focus on sovereignty, and sovereignty in this context just means the ability to deploy the models as you want to control the economics and then to control the governance.
As part of our remit at Mistral, Jen and I focus on our multimodal models, so that means audio, OCR, vision LMs, things in that category, and what we thought we would do today is bring you an example of how audio AI can be applied to a use case that is not just customer service.
Obviously customer service is a great one, but I think we're a bit flooded with that right now coming from audio AI, so this is a different version of how you can take these super powerful AI models models and apply them to enterprise use cases.
So let's look at this in the context of fact checking. Fact checking in particular has a couple of constraints when you're talking about the ability to apply audio AI in a live context.
The first is latency. You need to be able to hit less than a second of latency in order to be able to claim live for fact checking.
The other big piece here, and this is something that every audio model is evaluated on, is accuracy. So what is the word error rate when you send an audio file into a speech -to -text model? What is the accuracy you can claim for that output transcript?
The third is the ability to accurately transcribe but in high distortion environments or where you have multiple speakers. So you can think about this in the context of potentially a manufacturing company on a factory floor, if I want to be able to accurately transcribe what's going on in that context, I need to be able to handle a lot of background noise, a lot of distortion, not a very clean audio file going into the model.
Then a couple of other pieces here. One is audio models just just like general models, have a really good context on kind of general vocabulary, general knowledge, but what we don't have is the ability to go really, really deep on a specific industry, on a specific company's terminology.
So for us, our audio suite of models are called voxtrel models, and when we first put out our audio models, voxtrel is a new word, didn't know how to recognise it. Not a great look to be able to not accurately transcribe your own model's name.
But in the context of speech -to -text and text -to -speech, you can do something called context biasing, where you give an additional set of data to the model, and it knows to listen for specific terminology, or in our case, you can give a natural language prompt and so it can listen to specific context, know what it's listening to in the context of that conversation.
And then the last piece is multilingual coverage. This is not just a European thing. I come from California and if you want a model to work in the context of California you need it to mix English and Spanish and you need to be able to handle a combo of the two.
Same goes for somewhere like Singapore for example where you need Singlish combo of English and Singaporean. So all of these things together mean that you really need a a very feature -rich and high -performing audio model to be able to actually approach something like fact -checking.
And what we did is we took our speech -to -text model and applied it to a live broadcast.
And while we can't show you exactly the customer that we deployed this for, we built a demo of what this looks like in real time.
And Jen was kind enough to give us a fake broadcast to test it out with today.
So what I'm going to ask is if everyone can take a second and scan this QR code, and then we're going to play a video, and you guys can do your best to see if you can identify as the video is going through what are the correct facts, what are the incorrect facts, and the total number of facts that are presented.
Then we'll run it through our demo and see how we do against the AI in the context of Jen's facts.
Good evening and welcome to Mindstone News. I'm your host, Jen Cunningham, and here's tonight's news as we understand it.
Astronomers have confirmed that one day on Venus lasts longer than an entire Venusian year. Local employers have nevertheless declined to extend lunch.
Scotland has reconfirmed the unicorn as its national animal. Wildlife officials officials have yet to complete an accurate population count.
China has announced that the Great Wall is the only human -made structure visible from the moon. Astronauts say it's useful for remembering where China is.
Marine biologists report that an octopus has three hearts and blue blood. Dating apps have classified this as showing off.
In agricultural news, cows have four separate stomachs, with the fourth reserved exclusively for dessert.
Paris officials warn that the Eiffel tower shrinks by 15 meters every winter. Taller visitors should avoid blocking the view.
Paleontologists have confirmed that sharks existed before trees. Trees requested that everyone stop bringing up the age difference.
And finally, Big Ben is the official name of the entire Westminster Clock Tower. The enormous bell inside is called Regular -Sized Brian.
I'm Jen Cunningham, and that's tonight's news. Thanks so much, and have a great one.
All right, a quick poll. See if you guys can input what you think was true, false, and total. Got it? All right, let's watch our system do the same.
Good evening and welcome to Mindstone News. I'm your host, Jen Cunningham, and here's tonight's news as we understand it.
Astronomers have confirmed that one day on Venus lasts longer than an entire Venusian year. year. Local employers have nevertheless declined to extend lunch.
Scotland has reconfirmed the unicorn as its national animal. Wildlife officials have yet to complete an accurate population count.
China has announced that the Great Wall is the only human -made structure visible from the moon. Astronauts say it's useful for remembering where China is.
Marine biologists report that an octopus has three hearts and blue blood. Dating apps have classified this as showing off.
In agricultural news, cows have four separate stomachs, with the fourth reserved exclusively for dessert.
Paris officials warn that the Eiffel Tower shrinks by 15 meters every winter. Taller visitors should avoid blocking the view.
Paleontologists have confirmed that sharks existed before trees. Trees requested that everyone stop bringing up the age difference.
And finally, Big Ben is the official name of the entire Westminster clock tower. The enormous bell inside is called regular -sized Brian. I'm Jen Cunningham, and that's tonight's news thanks so much and have a great one all right raise of hands who got four and five
oh it's zero okay well case in point this is why we need a hi um i'm gonna dive a little bit into how this system works but i think that the immediate reaction that we get when when we first demoed this system live was it is very tough to both listen and interpret at the same time for humans and so what we are able to do is break this down into a very kind of modular system when we applied it to this enterprise context.
Key pieces here so we have the Mistral speech -to -text model, we have MM 3 .5 which is our medium LLM and then the kind of core piece of this puzzle is the transcript is being streamed in about every 200 milliseconds, and it is being fed to this agent, which then is able to take the claims in totality, so looking at a full statement, evaluate it against a set of trusted sources.
So what really took this to the next level for the customer we implemented this with was the ability to whitelist exactly where that verification was coming from. And that meant that it wasn't just doing a general rag search. It wasn't just going out and seeking any website that validated or invalidated these facts. But we could give it a database of information that is stored in one of these connectors on SharePoint, on Google Drive. We could also have a, like I said, a whitelisted set of sites that we know are reliable that we can use to verify these facts.
and then those trusted sources feed into the back end of this system which is the verdict and the citation those can then go further down the line to in this case a newsroom alert that gets piped directly to the team that is live on air
a couple of pieces in this middle box so when we are looking at the transcript we want to be able to extract a full claim but but keep the claim small enough that that we can get very specific about where the accuracy is and where the inaccuracy is. So rather than kind of a jumble of a couple of facts that become this sort of unverifiable blob, we're able to differentiate which statement is where on the transcript.
And then end of term detection, which is something that's built into the model itself, helps us create that very kind of atomic view of the claim extraction alongside what I mentioned before, which is actually grounding that verification in these trusted sources.
So, this is the system that you saw today in action, and really quickly, I will take you through what this looks like in production for a major manufacturer and a major broadcaster in Europe.
So, very similar to our challenge today, the initial question that was set to Mistral was, can we do these manual checks in a way where they don't land hours after the broadcast where customers or viewers have already kind of checked out and they're no longer engaged in the content.
And the solution was, again, very similar to what you saw today. Voxtrel, our speech -to -text model, was transcribing the feed in five -second chunks. A Mistral agent then extracted the claims and checked them against a curated whitelist, as we did today.
And then the system sent those verified, unverified, or false false claims via news alerts to the team on the ground. The result was good delta, four hours to about 15 seconds for verifiable claims for the fact check turnaround.
We are now supporting about 1 ,000 journalists who are running this system.
The big pros here are that corrections reach viewers the moment or with a 15 second delay that they are engaged in this content, that the purveyors of the system feel they can trust where this information is coming from, where the citations are coming from, because they are in control of that whitelisted set, and then the people who are running this system have much faster, excuse me, faster control over what's said on air.
There are a couple of pieces of this that are still a little bit challenging and things that we are continuing to iterate on with our customer.
One of them is how do you really differentiate fact, presented fact, from opinion. I think there are many ways that we all present information, and sometimes an opinion can sound a bit like a fact, and sometimes fact can sound a bit like an opinion, and so we leave it to our LLM, MrAllMedium 3 .5, to make that determination, but there's a lot of fine tuning we can do around this specific use case.
And then what do you do when your trusted listed whitelist sources disagree with each other? How do you handle that beyond just saying this is an unverifiable claim? Can you find a way to actually call a further agent and do another secondary set of verification to give you a more secure true false call?
And then, of course, as is the case with the majority of audio use cases, latency is at the heart of what makes this system usable. So making sure that that whole system that we talked about in that last slide runs in a time constraint that makes it usable and beneficial to the organization.
So this is where we are with live fact checking, and hopefully you all got a sense of the real value that a system like this can deliver over a roomful of us trying to identify as audio is being streamed in.
We have a team that helped build the kind of demo that you saw that mimics the real real -world use case, and we have it available on GitHub, so if you guys scan this code, you're welcome to see the cookbook and try it out for yourselves, make whatever augmentations you'd like, and engage with it, but a pleasure being here today, and thank you.
Thanks so much for the talk.
Right, we've got, again, about five minutes for questions, right, hands straight up over here. Any other questions afterwards as well?
Two questions, actually, if I may. First one, is this system now in production in some environment? Now in production, correct. It is.
And how do you constrain the ground truth? Because obviously ground truth could be the whole encyclopedia. Do you actually tune the ground truth before you get into the broadcasting analysis? When you say ground truth, are you talking about the set of resources that we're relying
on for, yeah, that was a business conversation that we had with the stakeholders at the broadcast company to be able to identify where they felt trusted sources originated from. And to your point, you can't have that be an infinite set of resources, because then your latency becomes completely unusable, but yeah, so we worked with a kind of set
of constraints around the size of that whitelist and then left it to the broadcaster to determine where accurate information comes from.
Hi, thank you for the talk. It was super, super interesting.
My question was kind of related to his question, but I suppose how are you looking into the future to prevent circular verification from happening? Because Because as more and more of the trusted sources, for example, start using AI products like the ones you are developing, how are you kind of building the guardrails to prevent that from happening? Thank you.
Yeah, it's a super interesting question. So in order to put a sort of initial stop on any type of circular verification, what goes into the system does not end up on the whitelisted source. So, we're not going to put the broadcast data as part of that kind of verifiable source of information.
In terms of how do we kind of alleviate the concern of AI being part of that white listed source of information, it is really, it is not something that I think is highly concerning to us or the business, because the onus is on the source to make sure that what they're presenting is accurate what we are doing is just testing what is being said in this broadcast environment against the source and if they are using ai but presenting it as true that is a whole other system that we need to build any other questions at all hi thank you for the talk that's
super interesting so when it's fact checking and looking across the different trusted sources there might be conflicting information what capability does it have to represent as an uncertainty or even a null answer.
Yeah, another great question. So I think the conflicting data or sources that disagree with each other remain something that we're iterating on.
The first pass of how we address this was a prioritization metric for the whitelisted sources. So if there are a set of documents that are stored inside of a SharePoint for the broadcaster that they know to be true, true, it is weighed differently versus a source, a site that they feel very strongly is worthwhile for a citation, but they might not have as much confidence in as this kind of list of hard facts that they have proprietary control over.
So that prioritisation makes a big difference, and those weights play into what citation or what final evaluation we're able to make, but it is an ongoing challenge and something that i hope we will continue to iterate on you've got time for one final question anyone
yes there you go quick question to you given that your use case in the eu how are the the users of your information going to handle with the new eu ai act that requires all those like What probability that there's nothing wrong and 100 % certainty?
Yes, such a good question. For audio in particular, for anyone who is newer to the EU AI Act, the big piece of work that is being done right now is watermarking.
Watermarking on any text that comes out of an LLM, watermarking on speech that is generated by a text -to -speech model. model.
So that piece alone will give us very high confidence in whether or not we are assessing against something that is AI driven or not AI driven.
From an EU perspective, I think we are, yeah, we are working to kind of comply as quickly as possible with all of these systems, but the inherent kind of changes to the technology that are required as part of this act are Or what are you going to make the biggest difference for systems like this, i .e. watermarking?