So today I want to talk about the caveman skill, but this is really an excuse to talk about token efficiency or token economics. But we'll go back to the caveman skill. Big brain, small mouth.
You've all heard about AI tokens, I presume. There are many different definitions as to what tokens are. I looked it up in the dictionary. None of these apply to really AI. AI. But there were a few that I liked.
Obviously, it's the token idea that you would have something, a memento, a relic that you could take with you almost. But what I particularly liked is the IOU. Because in AI terms, the AIU has actually become a UOAI.
And there are a number of different cases that I'm sure you've become very familiar with, where the AI bill has come as a shock. We call that bill shock.
Many organisations and individuals have found themselves having to increase their subscriptions or pay more money to the big AI vendors than they perhaps had ever realised they were going to have to spend.
The return on that investment Investment is another conversation. I won't cover that tonight. I think that could do with a couple of more talks, to be fair.
But what I do want to talk about is the efficiency of the token usage. And this is a relatively new term, but this is actually thinking about how we can be more efficient with our tokens, but still maintain the same quality, behavior, accuracy, and value that we get from using AI. So I'm going to blow your minds a little bit if you haven't done this before, so apologies.
It's pretty straightforward, but let's go through this over here, because this is how we normally are introduced to AI. There is a system prompt behind the scenes that we rarely see, and then as users, we ourselves provide a message, a user prompt, if you like. that goes through the AI system it thinks about it and it responds with the output
that's a rather simplistic view of what takes place a little bit more complicated view is that actually there are all kinds of other information that we need to consider as well this provides the overall context that the AI needs in order to understand what the hell we're talking about about.
That could be the documents and files that you've attached. It could be the different servers that you've connected to it, the MCP servers if you use those.
A number of different tools with long lists of all the different functions and features that they can all provide. And there's probably a memory file as well, a set of instructions if you use a project for instance.
And of course, some specific domain knowledge, especially if you're using a rag or another system to provide it even more context and useful information. And these can be repeated a few times. It then compresses that a little bit, and you have a context window that is then passed to the AI to make a decision.
It thinks about it, it spits out an answer, but in order to understand the conversation that's taking place it needs to have all of this reinjected in the next round all of that information and the message history are essential for it to be able to answer the next question that you give it within that same chat or or conversation window.
So far, so good?
Now, the simplistic, oh, that's a nice wobble there. The simplistic issue is that when we actually provide our message and we get an output, the output that we get is three to five times the size of the message that we've given it, okay? These consume tokens. tokens.
When we turn round and ask it another question in the second turn, that message gets added to everything that's gone before it and it spits out another response and so on and so on. Every turn, every time we ask it another question or give it another piece piece of context it grows but it's got a limit and that's the context window limit and everybody here I'm pretty sure knows what happens when we hit that context limit
what happens you've got to subscribe sometimes yes but no very good point but no actually something happens in the chat as you ask it a question and you're right in the middle of a really dense conversation about something and it goes over that limit what do you think happens you can create a new chat but then you've got to rebuild that context what actually happens within that chat compaction correct sorry i didn't hear you the first time it compresses it so that it can actually carry carry on working.
And what happens when it compresses all of that information? It forgets half of the stuff, because it can't hold it all in memory. So actually, anything that we could do to reduce that growth would help us have more meaningful conversations with the AI systems, and that's really the purpose of this evening's conversation. How do we do that, and how do we reduce the overall token burn?
Because in real life, we actually have a lot more complexity to this than meets the eye. As well as the messages going back and forwards, we have the tools, everything that the tools can do, the thinking process is also added to the sequence, and anything that we use in terms of a tool, like for example we examine a file system and it does a directory listing, all of that gets added at every single turn from that moment on. So you can see how quickly we run out of the 200 ,000 or now 1 million tokens that you get on some of the frontier models.
So there are four different ways in which we can reduce the token burn. The first is how we actually ask the questions. questions now these could be the user messages it could be the system prompt obviously we can't usually modify that but it could be the developer prompt if you're using api calls
you can compact it in terms of the way that you're asking the question you can compact the data that you provide to it you can also create macros so if you're going to be repeating the same terms all all of the time, don't write it out longhand, just say variable or macro definition equals and then the long -winded explanation and that saves a lot of tokens.
You can in a similar way use other references, IDs and similar. Or you can tell it in JSON format how to output the information so that it reduces the amount of tokens that it uses in the output.
In terms of the context, we can remove some of the redundant information. We can actually add a rag, which is a more efficient way of going through huge amounts of information, and, and, and, and.
But the one I'm going to be talking about is the one at the far right, which is the output. I mentioned before, the output tends to be three to five times larger than the input. so if we can reduce that we can make a sizeable difference through the number of tokens that we're burning through hence the caveman skill
so the idea is very very straightforward if we can reduce the complexity of the answers in terms of the verbosity how many words it uses how many tokens it's consuming but still keep it valid and relevant and intelligent we can actually get a work done without consuming so many tokens obviously if you're having long deep meaningful
conversations with your ai system this really does start to add up and it can lead to up to 75 increase or decrease in the number of tokens that you're using i should say
so i've got the caveman skill to actually describe what the caveman skill did and this is what it says it basically cuts all the filler all the pleasantries that's a really good point that's a really great question or you are so right none of that it just gets to the key matter it doesn't have any of the adverbs adjectives and so on no hedging it could be this it could be that or or it's this but not that. It just cuts all of that to get to the key core information.
And if you are working in something that's really sensitive, like code or a press release or whatever, you can tell it specifically which bits of output never to change in this way. So that bit comes out properly, but everything else, all your interactions with it, for instance, are greatly simplified.
This was one of the most downloaded repositories on GitHub a couple of months ago, and it remains one of the top three.
Would you like to see it in action? Cool.
What shall we ask it? Okay, I've got an idea. That seems like a sensible one, doesn't it?
We'll work that one out, but I am a little bit anal in that way, so I won't correct that.
I've got to say it's going to take a couple more seconds than usual, because the way I have it set up with Notion, It's just going to get the instructions for that skill from there and load it that way. If you're using this as a normal skill or GPT in your normal AI system, it'll be a lot faster. Let's see how accurate this is.
I think you can tell very quickly why it's called the caveman skill. Year 2029, machines rule. Skynet AI nearly wipe out humans. Human resistance led by John Connor about to win. Skynet sends Cyborg assassin, etc., etc. You understand that, don't you? You get the gist straight away. You don't need the filler. That just adds tokens. So if we remove that, we can actually get on with the job at hand, whatever that may be, without having to spend too much.
So if you're interested in downloading the caveman skill and having played with it yourselves, that's the repository for it. As I said, it's one of the most downloaded repositories at the moment on GitHub, and for good reason. It allows people to have a lower -level subscription than they would otherwise have to have with an AI platform. But it's also good fun. You can turn it on and off simply by asking it to do it in caveman mode or not. Simples.
But this is, of course, the one for the output. What about the import?
port and that's where there's yet another one this one known as the headroom skill which I've stolen an image of Mac's headroom I don't think anyone else in this room is old enough to remember Mac's headroom but there you go back from his MTV days and the concept here is that it
actually goes through all of the context that you're supplying to it and it compresses it it before passing it across. It's a lossless compression, so you don't lose any information, but it's compressed, and any redundancy in it is cut out, which makes it hugely efficient. You can use both of these skills together to great effect.
And last but by no means least, developed by the very same organization that's putting together this fantastic meetup, Mindstone themselves have also developed the Super MCP Router. You may also see this in Rebel, their AI front -end tool platform.
The concept here is that the MCP servers are very very chatty. When you connect one they give you a long list of all the different things that they can do.
So by the time you had the three or four different agentic MCP servers, you end up burning through loads of tokens without ever using them, just by giving them all these lists of capabilities that it never needs.
Super and CP actually work slightly differently. It has one command available to the AI, which is show tools.
So if the AI says, oh, I would love to be able to read their Outlook email. Have we a tool for that? It can actually use show tools to see what tools are available and only at that moment get the listing for the relevant information and that makes it again very token efficient.
Use all three and you're really on to a winner.
But of course what happens if you actually work for a company? What I've just been showing you now is great for the consumer and you could use it in a smaller organization but when you get to a larger organization it does become a little bit more complex.
Token efficiency and the bill shock that I mentioned earlier are very very important to large organizations for self -evident reasons especially as they're trying to work out the ROI of these platforms and tools but we now have scalable solutions for actually being able to track token usage across not only the AI models that they're using, but across individual users, departments, and actually starting to track the return on investment in that way.
You also get further insights into auditing, so you can actually see the logs, the conversation logs at enterprise level, which helps tremendously in identifying opportunities for further enhancing enhancing efficiencies around token usage and of course this is actually part of a bigger conversation around AI governance and strategy and frankly that's what we do at QL security we help organizations mid to large size organizations mainly deal with AI security and governance if that's of interest please see me afterwards or connect to me on LinkedIn thank you