From the event: Mindstone Newcastle September AI MeetupCan You Trust What AI Builds?
View event

Can You Trust What AI Builds?

Rebuilding a Learning Platform with AI

So, I've said who I am already. I'm Peter. I'm a product engineer at MindStone. I've said a little bit about us already.

So, we're an AI transformation company. And one of our core products is our learning platform. And this is the story about how we rebuilt our learning platform from scratch with AI. AI. And if we built the whole thing from scratch with AI, how can we trust that it's correct?

So I just want a quick show of hands. Who's from a technical background in the room? Yeah, quite a few. Okay. 75%.

Put your hands up if you're not from a technical background, but you've built something with AI, which kind of Chris asked already. Okay. These are people you need to watch out for because they're going to steal your jobs so like i said our core

The Legacy Platform Problem

product um was our learning platform and it was our it's our main revenue stream really um mindset's been around for like six years now and we've pivoted around like lots of startups doing lots of different things um it actually started off they were trying to catalog the whole internet and then allow people to learn through this system i don't really understand because i was i've only been here six months um but our platform was actually built by an ex monzo engineer who was

our cto and it was um developed it was like built on monzo's architecture if you don't know monzo is a bank they've got around 400 engineers and um they work with micro micro yeah microservices and the idea is they have lots of different services and they have a team that works in different services and it works really nicely because they can break up their infrastructure to have lots of teams working on it but i'm not saying we're just a team of four engineers

so it massively over complicates things so for the technical people out there for this just this simple learning platform we had 86 go microservices running in kubernetes with kafka redis and or communicating via rpc for everyone else imagine like 86 moving parts in one big machine that could kind of go wrong somewhere and trying to do one job but no one really wants to touch it and we can't just we can't break it because like i said it's our main revenue stream

so we can't just mess it up um for contrast and i know people have built stuff available i already know this but we have rebel which is our take on co -work it actually came out before co -work and it's a multi -model so you can use any any model like anthropic we use open router so you can kind of pick whatever you want agent and you can yeah sorry we built that in just three months and it's around five million so half a million lines of code in those three months and this was

a similar kind of size and we want to be able to move like as fast as we could on that building with AI as we can with the learning platform. But with this complexity, we couldn't, basically. So, we wanted to simplify...

I forgot to say. It's quite a technical talk, but I'm trying to make it valuable for everyone. And if you bear with me, there are points for non -technical people as well. So, hopefully, it will make sense for everyone.

A Simpler Architecture

So, we want to take these 86 different different services and make them into one. And in technical speak, we call this a monolith.

It currently was in Go, which was one of those language of choice. And AIs love TypeScript or JavaScript, they write everything in TypeScript.

So we kind of wanted to move to TypeScript, we wanted to move it all into one service that was easy for the AI to kind of look at. And we didn't need the performance of Go. Go is very good at like performance and we didn't need that for just this training platform

and we also wanted to give the ability for anyone to contribute to the learning platform not just engineers but kind of as chris mentioned where have you gone kind of as chris mentioned already you have these like transcripts or you have like people with product ideas and we want anyone to be able to come along and implement them without having to like tell an engineer to do it so that's kind of the goal we want to get to.

And we kind of estimated originally to do this kind of big re -architecture before AI, it would be like six to seven months for humans. So like pre AI, this would just like never happen, basically.

An Ambitious AI-Powered Rewrite

And you can think about banks that run on like these old legacy systems that like for years, like from, I was working for a visa once and they had stuff running from like the 1980s or something and no one wanted to touch it basically um but we kind of set this ridiculous deadline of trying to do it in a week um which

kind of seems ridiculous um i guess but we want to do this to like drive creativity and innovation around the process and we didn't want to like hand crank the whole thing uh we wanted to like build a factory or a system around to rewrite it for us so we could just crack up our favorite agent like claude code or codex and be like rewrite everything and hopefully it will do

Using Tests as the Contract

it but then how do we know if it's correct so we have these things in technical speak called tests in non -technical talk they're like a contract so this is a little diagram explaining it here so So we've got the old system, and we pass in some input, and we get some output back, and we check that output's correct.

We can also verify that when we send something in, and if it sends an email or it writes the database, we can check that it's doing those things and it's writing and reading the correct things and sending the correct emails.

So we can basically write all these tests around the old system, and then when we create the new system, we just run the same tests against it and the idea was to write the new system bug for bug

we all know that like all systems have so many well lots of issues in them normally they will have bugs in them and it's kind of a bit of a bizarre thing to be like why don't we fix it as we go forward but then it gets incredibly difficult to kind of have a single source of truth if we're

trying to fix it and recreate it at the same time so the idea was go for bug for bug recreate the issues and then once we're in a better state with the whole system we can then fix those issues

um i think that's everything on there

Why Test Coverage Wasn't Enough

so how do we go about doing this um well first of all uh there's this chap called ralph wiggum from the simpsons you might know him and uh there's a thing in ai called the ralph wiggum loop which basically means given a goal get the AI to go around in a loop and continue iterating until you've reached that goal so we originally went for like test coverage and that mean basically

in simple talk that just means it has every line of code being tested because if it has then if we write enough tests to cover all these issues then we can build up a big collection of these tests then we can run it against the new one, and if it passes, we know we're good. So that's what we did.

I spun up my coding agents, I got some really good agent instructions going on, I set off all these sub -agents, and it was amazing. I was like, this is the best thing I've done, because this was a few months ago, and I was like, oh my god, this is amazing.

As a developer, everyone hates writing tests, so this was incredible at the time. Also, we started off with zero tests, I forgot to say that. So, yeah, I thought we were winning at this point. We got loads of tests.

And then I found the AI was cheating me. So here he is with his, or she, or they, with their green tick of truth, but behind their back, they were holding all the cables.

And we had great coverage, average, but the quality of the tests were horrendous.

Even though I'd asked the agents to review them and do all these checks, because they were sub -agents and other stuff, they just wanted to complete the tests and do as good a job as possible. It was like, yes, they're good. Yes, great. They pass.

It was really hard for the tests to marry up against the code and check that we were testing all the stuff that we needed to. It was really hard for the agents to do that.

So we kind of needed to put a system in place where the agents could understand if they were doing a poor job. And then it sort of became, rather than chasing coverage, we chased correctness of the tests.

Building Checks AI Couldn't Bluff

So I was a week in, I had 7 ,500 tests. And then I had a little climbing accident in October and I got a call from the hospital saying that I had surgery in a week's time so I wasn't doing very well I basically we'd gone from that deadline of like a week as a kind of a joke and then I actually did have a week left now to do it and I had 7 ,500 rubbish tests so as I said the job became building a system that checks the AI

and the way we did that so for example if a test fired off a piece of code that sent an email we'd make the test fail we'd put a deterministic check in place if an email is sent and that email isn't tested to like check the body so say it said like hello peter and we're not actually testing that then it would just fail the whole test and this would give like a really quick and

instant feedback loops the ai to be like oh i've actually missed something whereas before it was trying to do this like in the the llm and it was really struggling to work out if it had missed things so just putting this like simple feedback loop in mean it couldn't cheat anymore basically yeah so yeah the deterministic check was very reliable even though ai is not always reliable and it can hallucinate.

So then I put these checks in place, and I was like, okay, go and fix all these broken tests, because that's what you do as a developer. You don't throw 7 ,500 tests in the bin. You go and fix the stuff you've already got.

Actually, it was just going around in circles trying to fix these tests, and because it had all these difficult checks in place now, again, it was really struggling to fix them because they were so badly written.

And actually, it was way easier just to bin them off and cheaper deeper to bin them off and just get them to write them from scratch, which kind of is a bit of a mindset shift because we're not used to doing that, I guess, just bidding things and starting from scratch again.

So, yeah, the anti -climax of the talk is after we did these, put this deterministic thing in place, it kind of just all worked nicely.

From Playbooks to Production

While these tests were being written, I worked on playbooks for different services, so I looked at the different patterns of the different services and how they worked from a developer point of view and used my developer brain to actually work out how I wanted to architect these applications.

Then I'd write a factory to build these things, run the tests, fix it, review it, roll back if there's any issues.

use.

In the end, it took around two weeks, not six months, so kind of a whim. I shipped it two days before I went to surgery.

The handover has now become... Well, before, handover for developers or moving off a project was quite a big job. You'd have loads of documentation, but now it's just making sure the agents .md file is up to date so that when people come in to work on it they can just ask the agent and it knows how to deploy it knows how to fix things it knows it signposted all the relevant documents so it knows what to do so that's what i did the

Trust AI Through Proof, Not Assumption

day before i left so can we trust what ai builds yes but not blindly and i think the trick is to make an unbluffable check or this fast feedback loop that i talked about this can be deterministic but maybe not um so i've spoken a lot about code but for knowledge workers um how can you check it's a good output um well if you did if you bought a report you could check the totals and add them up against the data where you got them from to make sure it's correct um you can rather

than trying to pull a report together via an mcp every time you could build a script with the with AI to pull the data out and then you could put tests around that script so you know that that script's right and then you can just run it every time deterministically and if you asking for like uh if you're asking chat to be to a question and you want to know you can basically demand citations to check the sources so there's like a bunch of things you can do rather than just assuming what it's telling you is correct and you probably know all this already but um

So, yeah, don't ask AI to be right. Build a way to prove it.

The Human Cost of Working with AI

And I just want to touch on the human cost of things because, like, it's a massive thing we're dealing with at MindStone, both for developers and knowledge workers

because, obviously, we're running, like, multiple agents in parallel and you're constantly contact switching between all these different things you're trying to do because if you're not doing that then you're not being efficient.

So there's a lot of cognitive load going on and there's a lot of burnout risk and it's just incredibly taxing.

So it's probably another talk for another time really because there's like so much I could say about this,

but yeah I just wanted to say if you're experiencing this in the AI world then and you're not alone. And that's it.

Finished reading?