Hi, I'm Peter. I'm a product engineer at MindStone. Normally you'd recognise someone working at MindStone because we wear these lovely purple t -shirts. But mine was in the wash, so I've had to opt for this tortoise purple t -shirt instead.
What is MindStone? It's an AI transformation company. Drummond kind of introduced us already.
We have a number of different services that we offer. offer. We teach people how to use AI, both from a non -technical and technical perspective in their jobs.
We offer a product called Rebel, which is a bit like Claude Co -Work, which allows you to do agentic AI. And this talk is really the story of how we took our learning platform, one of our core products, and we rebuilt it from scratch with AI.
And the question is really if we build something with ai from scratch how can we trust it how do we know it's
right can we just have a hands up who's an engineer in the room or from a technical background okay about a half a third it's normally the case uh okay and put your hand up if you're not from a technical background but you've built something with ai oh wow that's amazing okay watch out for these people because they're going to take your jobs.
Okay, so, as I mentioned, our core product is our learning platform. It was also our main revenue stream in MindStone. And
MindStone has been around for like six years now. So we've pivoted around quite a lot of times to different products.
And our learning platform was built or architected by an ex Monzo engineer who was our old CTO so it's it's kind of designed on Monzo's architecture and for those who don't know Monzo is a bank and they have
around 400 engineers and the way it works is that like each service is divvied up into different teams so it's quite easy to work on per team on this architecture but Mindstone only has four engineers so it's kind of quite complex for a team of four to work on this architecture.
Just for the technical people, it's 86 microservices running in Kubernetes with Go, NoSQL, Kafka, Redis, and all communicating together with RPC.
For non -technical people, just think of a big machine with 86 moving parts all trying to talk to each other to try and get one job done, but it's quite complicated to do that one thing.
And the thing is, because it was complicated, no one really wanted to touch it. And we can't break it because it was our main revenue stream.
And just for a comparison, I mentioned Rebel, our Gentic product. This was built in around three months and has a similar code base, so around 500 ,000 lines of code.
And we want to be able to build at the same speed with Rebel as we come with our learning. Well, the other way around, sorry. want to be able to build at the same speed with our learning platform as we can with rebel
so how do we go about doing that i've got to mention one thing is it's quite a technical subject this but if you bear with me it's got stuff for technical non -technical people and
technical people in the talk so we want to take these 86 different services and condense them down into one service in technical speak we call this a monolith um we also wanted to go from the
the Go programming language. So one called TypeScript. Our whole other engineering stack is already in TypeScript.
And AI just love, AI agents just love writing in TypeScript. It's like what they're good at, what they're best trained on. And we don't really need the performance of Go.
All it is, is like a learning platform. It doesn't need to be complicated. Originally, it was built to scale, but it's not really required to do that anymore.
And also we have this philosophy at MindStone of we don't want just the engineers to contribute, everyone who wasn't from a tech background who put their hands up. We want those people to be able to contribute to the codebase too.
So we want it to be really easy for if there's an issue in a particular application for someone non -technical to jump on and start contributing to that application.
Now before pre -AI we estimated this to take around six to seven months to do this rewrite so this would just like never really happen like in traditional terms and if you think about banks they have all
these legacy systems that they will wants to touch and it's running and we which is there and it's written in studying from the 60s but with AI we we can change that so we set us off a target of just doing it in one week
which is kind of ridiculous target where we just said it would take six or seven other months normally, but now we've got AI on our side.
And
why set a ridiculous target? Well, it kind of drives creativity and innovation.
We don't want to just sit there and try to hand crank this thing out. But we want to build a system or a factory to do this work for us.
So how do we go back to this?
Well, we could just boot up our favorite coding agent, Claude code, or whatever it is, and just say, rewrite this into something else.
But then how would you know that it's right? How do you know it's working?
Well, we have these things in developer speak called tests.
And for non -technical people, a test is like a contract where given a certain input, I get an output back and I can store this output and I can run that test again. And if it comes back with the same output, then I know the test is passing.
So given given this input, do I expect this output? That's basically what a test is.
And the idea is we could write these tests against the old system and just run them against the new system.
And as long as the tests passed in both, we knew that the two systems were the same.
And one of the other key things... Oh, yeah, sorry. And also, you can see at the bottom here, we've got email and then there's a database.
So, we're also looking at any calls that our systems make, so say we go to send an email, we also need to verify the email is sent and with the right information.
And the same with writing and reading from the database.
And the other key piece was creating it bug for bug. So it's important that we first made it exactly the same. And then once it's the same, it's easier to develop going forward.
So there's issues in our code base. There's always issues in code bases. Every light bit of code has problems.
And we could fix those problems as we migrate to the new system, but then you're kind of missing the single source of truth.
So then we can just fix all the bugs.
So you might recognize this guy. This is Ralph Wiggum from The Simpsons. And there's this thing in agentic AI called a Ralph Wiggum loop.
And the idea is that you can can give it a goal.
So ours was coverage at the top. And coverage is, again, it's a technical thing, but it's very simple. All it means is, is this line of code tested?
And if you have all of your application and every single line of code is tested, that's 100 % coverage. So we gave it a goal to get 100 % coverage on all of our tests.
And the way the Ralph Wiggum loop works is you go, go, okay, we want 100 % coverage. Let's run the tests. Let's check for gaps.
Let's implement any extra tests we need to cover those gaps. Then we run the tests again, and we keep going around until we've reached the maximum coverage that we can.
And this was great. We were writing loads of tests, and I was kind of like a few days in, almost a week in, and I was like, this is amazing.
Like, we've got all these tests against our code. It's going to be really easy now.
And then I realised the AI was cheating me. Here he is holding up his lovely green tick of truth in developer speak.
We're like, oh, all our tests are passing. But then behind his back, he's holding all the dodgy wires.
And the thing is, we were aiming for great coverage, but the quality of the tests were awful. So it basically ran the code. So the code was going through and we're saying, okay,
okay, this code is tested. But then within the actual test itself, we weren't asserting that thing was right.
For example, if we were sending an email in a particular test case, we weren't checking the contents of the email were what we're expecting. So it was kind of missing a bunch of stuff in the tests out.
And even though I instructed the agents to do all these different things, the sub -agents just want to finish as fast as possible. So they will just
cheat to try and get to the fastest possible route of completing their task.
So we cut this got to a point where we can't assume things are right. We can't ask AI if it's right,
we need to put a system in place to make the agent prove it's right. Yeah.
And then, in October, I had an accident and I fell off a cliff. And I got a call from the hospital.
This is real, by the way, last time I did this talk, someone didn't think it was real. This is real.
So I'm a week in, I got a call from hospital saying my surgery was in a week's time. So the one week deadline actually became like a two week real deadline now.
So yeah, I had to get it done. And at this point, I've got a bunch of rubbish tests that don't do anything.
So you may think I'm failing at this point, which I probably was. But
the job now became building a system that checks the AI. So rather than just asking the AI if it was right, actually building a system around it.
So
So what we did was deterministic checks so that if something happened, the test would fail if we don't assert that thing. For example, with the email case, if it tries to send an email and we don't test that the inputs and the outputs of the email are right, then it will just automatically fail. fail.
What this did was create an instant feedback loop for the AI to know if it had done its job correctly. That was the trick, creating this instant feedback loop.
Once we built this into the AI agents and this feedback loop for them to understand what was going on, I tried to fix all the tests because I was like, I've got these 7 ,500 tests and there's just a little bit not working in them, but they're all there.
And they were really struggling. They were going around burning loads of tokens.
And it was really hard for the agents to match up the code to the tests and try and work out what was missing. And in the end, it was way easier for the agents just to build new tests from scratch.
So a bit of a mindset shift that was basically just to throw all these tests out, which we're not really used to doing as engineers, just throwing all this code away, which we spent a week working on.
But code's quite cheap now, and actually burning more tokens was worse. So it was easier just to throw everything out and start from scratch.
So while the tests were being created, I put my engineering brain down into doing a playbook. So for each service, I kind of said how it should be created using all the engineering best practices. And then again, I went through like a loop of building it, running the tests that I've created against it, fixing it, reviewing and implementing it.
And the anti -climax of this talk is that it worked. So having the good tests, creating the truth for truth thing, implementing this, it worked really well. well.
And we were able to ship this new system out two days before I went into surgery. And that's kind of crazy as well, because normally you have to give a big handover in engineering, like normally you kind of have meetings and all this stuff and document all this thing.
Handover has now become, made sure the agents .md or claude .md, which is kind of the instructions for the agent, make sure that's well documented and signposted so that if someone comes along and needs to fix a bug or they need to push to production or anything they can just ask the agent to do it and it will do it so that's become our handover now just making sure that is really
well defined and signposted so can we trust what ai builds i would say yes but not blindly trust it and the trick here is to create the unbluffable test and this applies again going back to knowledge -based work and not just development but it applies to both code and knowledge -based work
How do we check if the output is good in a knowledge -based thing?
Imagine you were building a report and you had a bunch of numbers in there. You could add up the totals. You could manually check it against what you have in your systems. You could check the formulas are correct that it's using.
You could create a script which deterministically builds the report. and then you can test that script so that you know that that script is always going to build the report correctly rather than relying on the LLM to sort of non -determinously build this report each time.
If you're speaking to chat GPT, you could demand like citations or sources and then you could go and fact check those and make sure they're right. You could do it or you get another agent to do it.
So the thing is, is like, yeah, we don't want to be asking AI to be right. right, we want to build a way to prove it's right, which is, I think, what I've written there.
Finally, I just want to touch on the human cost side of things. Again, both knowledge -based work and development, we're kind of running this swarm of agents now, and it can be incredibly taxing cognitively.
There's a risk of burnout. You're trying to manage all these things. there's like loads of context switching going on and it's kind of another talk for another time but we're kind of struggling with it at mindstone as an engineering team and as a non -engineering
team who are heavily reliant on like agentic ai so we're just sort of saying you're not alone out there if you're struggling about this already um and yeah it's another talk for another time but I'm happy to discuss that after as well.
That's it.
There is a blog post of loads of more technical detail on there.
And I just want to mention a couple of more things. I said earlier we're an AI transformation company. We offer training mainly for non -technical people. So if a company company wants to get all of their employees up to speed with AI from like zero to 100. We work with them to do that.
And we're currently hiring two product engineers. It is London based or two days a week in London.
And I know not everyone here is an engineer. But if you know, like the best engineer you can think of, I'd love to hear about them.
And And yeah, that's it. Thank you very much.