From the event: Mindstone Newcastle September AI MeetupWhat actually breaks: a year of running AI agents in production
View event

What actually breaks: a year of running AI agents in production

Building AI Agents for Production

Hello everyone, Chris from RightBrain. Just before I kind of get into it, I just want to find out kind of, you know, where everyone's at in terms of their AI experience.

So who's actually built an agent and is using it regularly for their own purposes as an individual, like quick news fans? And then what about deploying an agent in a business? So letting it out in the wild. You guys don't count, the RightBrain people.

but yeah so I think if you've ever built an agent into production especially if it's gone outside of just yourself as the user I think you'll understand is only really one profound truth when it comes to agents and that is that shit happens and it's really it's really a case of how you deal with that and you

RightBrain’s Approach to Deployment

So RightBrain is a company. We build AI agents for businesses, and then we run them on behalf of the businesses as well.

And my team, Forward Deployment Engineering, we're really interested in getting embedded in the client's use case and the problems that they have and that they're trying to solve with AI.

We built a platform at RightBrain that is purpose -built for that. And we provide the platform, but we provide the people as well. to actually run it.

Because a lot of these clients, they know that they want to use AI for something, for doing something and automating a process within the business, but they're not really quite sure how to do it. So we help them do that.

Incidentally, we are hiring for a forward deployed engineer in our team. So I guess if you're at an event like this and you put your hand up before, then definitely drop me a direct message because we're actively recruiting for that role at the moment and you get to work with me and these guys as well so lucky lucky you um so

What Makes a Production Agent Work

really when it um you know when it comes to ai and i know this gets repeated a lot and and i think that it's one of those things that um is it's almost a bit bit too obvious now but the demo really is the the easy bit and when you get these agents into production and touching real world data.

It's messy, it's unpredictable, and it just takes a bit of time to iterate once it's in production. And you've got to use all the lessons learned from previous productions and bring that forward into what you're deploying, you know, currently.

So I think I want to really focus on three things within a production agent that we've built and that is running for a client at at the moment.

Useful, Trusted, and Observable

The first thing is really that it does useful work. And again, sounds obvious, but a lot of the use cases that we get presented to from clients that they want to use AI for aren't really optimal candidates for AI because yes, it might look good on a demo, but is it gonna save you a lot of time? Is it gonna be a repetitive piece of work that an agent is really suitable to do? So you really have to do useful work with the agent, so it has to be high value, and it's got to be viable in terms of the cost.

So we do a lot of production agents where we'll put in a really high -powered frontier model to begin with, and then over time, once we know what good looks like, then we optimize, and we might use a lighter weight, faster model. we might then try and optimize the architecture of the agent and and that's how you kind of iterate on it over time moving from that gold standard something that's more optimal and more useful and then the second bit is that the team

trusts it and this is all about the internal team within the client whoever's interacting with the agent it might not just be the person that's responsible for AI within the business it could be just the users and the guys that are actually providing feedback to the agent outputs. So this feedback is human in the loop. It's got to work during the initial deployment phase, but also as you scale up. And we've built in a couple of features, and I'll talk through that as well.

You've also got to phase the deployment in the right way. And then it's got to be auditable as well for it to work for any sort of serious business. business.

And then the third thing is you've got to trust it yourself. So if you're responsible as an FTE or whatever you want to call that kind of role, whether you're an AI consultant, whether you're working within a business and you're responsible for the AI, you've got to trust it. So you've got to have the observability tools. You've got to be able to debug and fix things quickly before the client or the users notice.

And then we built recently, this is something we've just introduced really but it's working really well which are watches which is a type of monitoring system for agents and i'll go into a little bit more detail with that as well

Case Study: A Meeting Concierge Agent

um so the the production agent i want to talk about is um for a client of ours that's a marketing agency they've got about 100 clients under management and they've got about a team of 70. so they are constantly having these client meetings and I wanted to talk through this

agent in particular because on the surface it feels like something that is quite simple to put together but then when you actually build this and get it working at scale you realize the kind of the edge cases and the messiness of the real world that actually you know that you encounter counter as an FTE or a consultant for AI.

What I'll do is I'm going to take you through a little, which I'm hoping you can see that. It doesn't matter if you can't see everything.

Turning Meeting Transcripts into Action

This particular agent, we call it a meeting concierge agent, and it's a foundational agent because what it does is it takes a transcript from a meeting. In this case it's a client meeting and it can expand beyond that.

But then with that type of meeting where you've got humans in a Google Meet or it can be a Teams or whatever and you've got a transcription service, our client's got Fathom set up but you could have Granola or whatever, you've got all this rich data from the human interaction, the creative ideation, whatever is going on within that meeting.

meeting and with that data and after that meeting there's a whole host of tasks and actions and follow -on bits and bobs that everyone's got to do and it might just take you kind of an hour after that meeting which you think well you know is an agent really going to add too much value doing something you know where actions are on our transcript after a meeting but when you scale scale it up to teams of 70 where they're having 20, 30, 40 meetings a day, it actually really adds up and you get a lot of value from that.

And when you actually see it working in real world and it actually works well, you realise how much time it's having. It's kind of a psychological thing as well. So things don't get missed as much and it really adds a lot of value to a business.

So what we're finding is that this type of agent is actually really really popular with clients that are kind of considering AI but not exactly sure what use case to handle. So the meeting transcript, you know, getting a transcript from the meeting is kind of pretty common now and everyone's got their own transcriber.

Classifying Meetings Safely

And the way that we set up agents is that they work within existing workflows. flows so um you know in our case the fathom transcriber uh that triggers our agent to kick into action and you might think well okay so just classifying the meeting which is a first step in the process really simple but actually there's there's always a bit of ambiguity there so you

know i'll put a few examples there but um it could be an internal meeting so an internal meeting you You might not want to transcribe that, so it could be a one -to -one and they've got Fathom set up and they didn't really want to transcribe the meeting, so it's really a classified meeting. So the agent has got to be intelligent enough and have the guardrails in place that it won't post a summary of the meeting to Slack or Teams or send an email out. So it's got to recognise that, not process the transcript and not reveal any sensitive data.

so that's something you've got to handle a prospect meeting so it could be a client that is pre contract or post contract so the agents got to work that out it could be a kickoff meeting so if you've just you know if you're having a kickoff meeting with a new client it might be like you know an hour and a half two hours long loads of information in there and there's loads of like follow on either reports or updates or setting CRM's up or whatever there's a a lot of tasks that needs need to be done with that but the agent needs to pick up that it is a kickoff meeting so not to set off a chain of chain of events afterwards that are not relevant

to that meeting it could be client delivery so a client workshop it could be a supplier meeting so it kind of falls outside of those categories and and you might be thinking well look like from a transcript an agent it's definitely going to be intelligent enough to work all that out and you'd be right if you just passed an agent the full transcript it could work it out but then that doesn't really work at scale because that's extremely expensive so um we we have typically

Efficient Architecture Through Specialized Tasks

an orchestrator agent and if we passed the whole transcript into that agent so you know say a a a a long meeting could be 100 000 tokens and that agent's got a lot of follow -on actions as you can kind of see in the in the flow chart so um it's got access to a lot of tools as well that it uses so the whole of that transcript is in that one agent and that is being used as context in every turn tool call so it can be really inefficient and you notice it at scale especially

if you've got quite a few meetings going through so what we did to kind of tackle that particular issue is we connected to fathom just via a webhook and we grabbed a summary of the meeting which fathom generates after every meeting anyway the meeting title which again can trip you up because sometimes you know big company they don't put the meeting titles in the naming convention that an agent would expect so you've got to handle that as well so but what it does is takes in a

a tiny bit of context, classifies it, and then it passes down to a very deterministic sub -agent, we call them tasks, so structured input in very isolated specific purpose in terms of AI, and then structured output out. Now we give to that task the meeting ID, a recording ID for the meeting along with the summary in the title of the meeting and a an estimate with confidence of the classification of that meeting that

specific task then calls for the transcript based on that ID already knowing with some confidence what classification of meeting it is so that means that we can build a very specific task for that type of meeting and only only then taken to that task, the transcript, which is at that point, it's not being used with a bunch of tools, it's not being passed around on every turn the agent makes because there's not much reasoning steps within that specific task. And then that means it's super efficient. So we can have a highly configured task for that particular meeting.

And we've got these sub -agents, these tasks set up and configured for each type of meeting that the client has and just to give you an idea of kind of numbers i mean um you're talking uh an order of magnitude in terms of cost um savings and um and then um on from that as well um because you've got this uh transcript in a very specific uh functional part of the process this task um you can then play around with the the model and um we we tested some open weight models at the moment we had we had gemini flash uh working on this particular uh task and then we

we tested it and evaluated it with a number of open weight models and deep seek we found was 97 percent cheaper like for that particular task and we ran it on the gold standard kind of test runs runs, so you could see going from an assumption that, oh we'll just bung in the transcript, let it work it out and it will figure it all out, to actually something that's 100 times cheaper or whatever and it works just as well. So that's the sort of the, I guess just a couple of examples of what we found in production. I could go on about this all day, by the way, in terms of what we found, but I'll just head to one more example on this.

Handling Real-World Details

Attendee list at a meeting, you think, well, it'll just be there in the calendar invite and you can just check it, call a calendar tool and you can check who was invited to the meeting, simple.

Well, actually, in reality, people don't turn up to meetings and also it's an additional tool that you're adding into the orchestration agent to figure out you know from the meeting invite who was at the meeting so what you really want to do is figure out and infer from the transcript who was actually at the meeting so we bring that little step of reasoning into a

into the very specific task so again it's just it's just something that you think is going to to be easy but it's not but then once you've got it like going and you're thinking about it logically then uh you can you can unravel it unpack it and and deal with it um so in terms of you know

Human Oversight at Scale

something like this you know an obvious kind of use case is a configured summary that's posted to slack or a follow -up email to the client that's drafted that picks out their actions or it could be an RFI to the client or something like that, or it could be a report, whatever it is, you'll need that human in the loop step.

So you'll need to obviously assign a person who is responsible for approving, rejecting, or adding a comment and getting that meeting summary regenerated.

And the other thing that we kind of had to deal with quite quickly was what happens at scale with this so you know it's fine at low volumes to get asked you know by the agent to approve everything but what happens when you've got 100 runs a day 200 runs a day on the same kind of tool call it becomes more of a problem than a solution then for an

unlucky individual so and and then the other kind of problem that we encountered was well what if someone doesn't have access to the right brain platform or whatever platform you're using to manage your kind of human in the loop steps and and it you know these platforms you know you know what it's like you go in there and it's and it's and it's easy for uh regular users of the platform to work things out and understand what's going on and navigate around it but

um you know in the marketing agency for example they've got about 50 account managers they just work in slack all day and they don't want another platform to learn and figure out and go in and try and approve something review it regenerate it and everything else so you need to set up human in the loop for people who just work in slack so they get a draft of the output they can comment on it they can either approve it reject it add comments regenerate it and then only when they say within within Slack that they're happy with it, then it gets published to the client Slack channel or whatever they've got set up, or if it's a draft email, something like that.

So we've had to build features around that and within the platform, and that's something we're constantly improving anyway.

Tapering Human-in-the-Loop Review

Anyway, another good one is applying a sample rate in terms of human in the loop. So when you initially deploy an agent with human in the loop at any step, any reasoning step or any tool call, you don't have that much confidence in it, really. Even if you know you've done as much testing as possible, once you get it in the real world, you're not exactly sure exactly how it's going to react to that messy real world data.

Now, what you need to do is you need to taper the human in the loop down over time. And you can do that within our platform, and I'm sure there's other sort of platforms you use as well. but yeah it can give you like a probabilistic chance of human in the loop being activated for that agent run so it might be like you know 20 % chance it might be 10 % chance and then in the end once you've got so much confidence that that agent can do that step in the chain really well and reliably you can just remove it and then you're saving more time for people that are involved in it

So I've talked about a couple of kind of follow -on actions from a meeting transcript and talked about kind of human in the loop and how you sort of redraft anything that comes out of it.

CRM Updates and Agent Collaboration

I'll just pick up on kind of one example, which is probably quite relevant, which is a CRM update. update.

So I talked about agent tools and how you can have a very deterministic task like with the transcript processing. But something like a CRM update, which again you think that's quite easy, we use ATEO for example.

We have built in sub -agents, so these kind of reasoning orchestration agents as tools so the meeting concierge orchestrator agent can when they when they know they've got to update a CRM they will then use an agent that is specifically designed for updating the CRM and they've got to

check whether that client prospect is already in the CRM they've then got to confirm what you know who the account manager is and make sure that's attributed properly they've then got to check for any duplication of notes to make sure that they're not just you know two people have recorded the transcript and it's gone through and it just it just becomes a mess so there's lots of little steps that that agent makes

but if you design that well and you've designed the meeting concierge agent well well, then they actually work beautifully together, seamlessly together.

Once you get into production, you realise more and more that you need to refine that over time. That's another one we've encountered recently.

Audience Questions: Feedback, Adoption, and Reliability

I'm going to take a little breath there. Has anyone got any questions at this stage?

Managing Duplicate Information and Emerging Requirements

If you have a situation where you've got a very big organisation and they're having a lot of meetings and putting the input into that, there may be some area of duplication. So you might have some people who attempt two or three stand -ups hypothetically saying similar things. So is there intelligence built into that to say, hang on, in meeting one Fred said this and then he said that. so you can trail an escalation and I say produce a RAG report for an end user so that type of thing

you're better off not presuming that's exactly what you need the agent to do if it becomes something that you find during the initial deployment phase then you deal with that I mean, what we do with clients, so for things like that, we give them the heads up that, oh, we think, you know, this is going to be a problem that we're going to try and solve. But anything that requires reasoning and a high level intelligence from the agent is costly at scale. So you have to be really kind of considered with that. But what you're describing is something that, yeah, we've had to deal with for real.

Using Feedback to Improve Agents

actually in the loop did you start putting predefined prompts in about right make sure you got more context of what was going on do you mean do you mean what the feedback going back in from you yeah so um so we because i'm saying users are well if you're not used to yeah ai they will just put the least mild comments context possible of what's yeah so we um

we're always well I'm always surprised about the feedback that we get from the agent output and originally we actually just had this kind of right approve or reject and human in the loop is really about sort of critical decisions that AI has to make that might impact the business but then we realized super quickly that all this like subjective feedback that won't kind of break the agent it'll still run but it could affect like the agent's output and it's actually pretty detrimental to the business but no one be alerted to it um so we um so i mentioned in the earlier

slide but we've now got watchers so um i built agents that um look at um a an agent sort of the initial deployment of an agent in the pilot phase we call it so the first two weeks after we deploy an agent and it's it keeps an eye on all the agent runs and periodically checks them but it doesn't really check for critical errors I mean it does but it's not responsible for that like the platform is what it does is it checks kind of any feedback coming in on slack about the agent anything that is kind of subjective in terms of is that a good output or not it records all of that and And then anything kind of breaking critical, I'll get an alert that's like, right, go and fix that now. But at the end of the day, I'll get a bit of a digest and a summary of those things that are recommended to be updated within the agent based on the feedback that's come back from the team. And then as the FTE, I can then make a judgment call on which ones of those are actually valid feedback and we need to fix it or whether it's just someone who's had a bad day in the office. yeah yeah um and and you've usually got one person within the company that you're working with who kind of helps filter that out as well yeah um but it definitely makes it definitely makes life a lot easier and it's it's a really good point and it's definitely something that's important in production

Organizational Adoption and Trust

yeah i'm interested to understand um your sense of how the organization how the client as a kind of organism responds to this kind of standardization because us as humans and me in particular We're kind of a bit sketchy at times and inconsistent and I suppose we're familiar with that. With this I'm thinking that there's more standardisation, actually more stuff gets done. How does the organisation kind of respond to that as a whole? How do the people in the organisation respond to that? Which is like a different, it's a kind of cyborg organisation almost.

Yeah, I think you obviously get different reactions, but we try and manage the client's expectations through it and try and educate them and make them feel comfortable with it. That's why when I said about the three key things, trust is super important and you've got to earn that.

That's why we really focus on working so closely in that initial period and we set the expectation that look this is not going to be perfect from from day one and um and i think that um you know companies are definitely getting more and more comfortable with it especially because the models are more capable and they're more predictable um we try and build into the agent architecture always uh to be as deterministic as robust as possible from the the outset um and again using

more powerful models i'm guessing that you'll have a you'll have some early adopters who get yeah this is great yeah yeah you have some people who might be called cultural custodians or laggards who basically hate it yeah then you'll have the kind of the bulk of the people in the middle yeah is that fair to say yeah totally so i mean we um we work in uh sort of further

education and you get um you know in colleges you get a mix of personalities that are in that type of uh environment i imagine it's the same for like a lot of public sector kind of kind of organizations as well and you'll get a couple of people that are just all in it you know they are like right i believe the hype this is where everything's going to go and then you'll get

another person that's just like this is just like you know too far -fetched it's not going to work it's just going to cause an issue you get the other ones that are so far down the hype rabbit hole that they're like ah is this it it just takes transcription and does that um it's going to save me half an hour that's you know i thought it was going to just do my whole job for me and you get

that it's such a wide spectrum and you i think you know we've got to manage expectations without dampening the excitement too much and an optimism for it um but yeah it's just part of the role

Classification and Model Selection

does this does this agent takes full transcript to decide the classification So what happens is we make a first guess at the classification based on the summary and then when it goes to the task that is responsible for processing the whole transcript, it actually verifies that and, you know, if the confidence category is high for that, it's usually right and we've proven that in production.

but if it's not then what it will do is it will change the classification but then what happens is it's a slightly more expensive agent run because it will return that result back to the agent orchestrator it will reclassify it then send it to the appropriate task so there's kind of like two layers of uh you know resilience there um and and we just felt that was uh that you know i just felt that was the best architecture and that's worked out yeah it's worked out pretty

Some people would put a couple of words, meeting, meeting title. Yeah, exactly, yeah, it's just... No description. Yeah, exactly, exactly.

So the transcript gets feeded to the highly capable model, and then classic... Well, actually, not particularly.

So, you know, if you've got something that's very deterministic, and there's, like, one function, and you've got structured data in, structured data out, you can use really reliably a lighter weight model so you know I really like sort of Gemini family of models for that type of task there's not much reasoning it's not much kind of like back and forth with tool calls so you know I mean I love I love like open AI models like Astra is just amazing but you don't you don't need that you don't need a sledgehammer for everything so

So this is why breaking it up into, you know, this way really works very well. And then you can bring the model, sort of the size of the model or the power, you know, you can bring it down. And like I said before, you can use open weights. So if you haven't used like Kimi or DeepSeek, they're like amazing.

Have you seen the gel? Yeah, so yeah, we spoke about this actually. And we're already going to try and implement it for classification. we tested it on some classification test site today actually and it was ridiculous like ridiculous like it basically cost nothing and it's like instant but i mean i haven't i haven't tested it enough to kind of have a proper informed view on it really probably got time for

Verification, Oversight, and Retrieval

one more question yeah obviously the agentic system has lots of stages to it how do you

yeah yeah and it's it's one of the the biggest challenges and um you know you have to have that human oversight for those critical things you have to and it might be that yes you can i know that um if you're um if critical decisions are based on source data and the model needs to reference that source data then that's kind of pretty standard now you know you can do that and reference source data but you still need that human oversight and human in the loop i mean i am

so back the previous life i used to work for um an oxford university spin out and they created models for highly regulated industries and this was before chat gpt came about and transformers were were the rage and um you know we we would build models over you know periods of months for these clients to make these critical uh decisions in terms of risk and the impact on on human life in some cases in the defense sector um but now we've kind of sort of skipped that process a lot. And it is a challenge.

I mean, we don't deal with that type of, you know, life or death decision. So we really kind of rely on the robustness of the agent architecture, having that human in the loop, building that trust. Building more. Yeah.

I mean, you know, by nature, you know, they hallucinate, right? That's the kind of, you know that's that's that's how it works and yes they can act as orchestrators so you've got tools that are much more tightly defined and and the llms will use the output of that tool but that's that's that's the way you're doing at each consensus to make the decision on the next stage and then it passes forward exactly exactly exactly but there's no verification there

When you say verification, do you mean coming from some source material that is the source of truth, like the ground? Yeah, like a corpus, whether it's a company corpus, whether it's Wikipedia for one, whether it's a word or a scientific literature or an ego corpus. Yeah, so we do have kind of research agents that will cite... More like a rag system. We use a web search for that, but we have a separate kind of primitive, which is collections, which uses RAG, so that's a corpus of documents. You can adjust the settings, so you can adjust the p -value and the chunks that it uses.

I don't use much RAG because the models are so good at handling context, like all of the context, it works. It's retrieval augmented generation. So blimey.

But if you've got a large body of documents, you basically embed that data to turn the data to numbers, and then you put it into chunks, which you can set. And then when you ask a question to the, the AI, the LLM, then it looks for relevant terms uh or embeddings number representations of the data and then it picks out those relevant chunks within that big corpus of documents and then it formulates an answer from those so um it's it's a way to very and it's super quick and it's a way to really kind of um make useful and cost efficient large volumes of data as context i think i might have explained find that okay but yeah there you go thank you so much okay cool

Finished reading?