It’s been a wild two years building agents, from request response word salad demos to full, looping, long running agents. What we had then to now, are worlds apart but as I speak with friends, read comments on LinkedIn from experts, it’s increasingly apparent that AI and Agents are the new Big Data meme:
Agents really aren’t all that complicated, in fact, as technologies go, Agents are probably the simplest hype-train technology to really understand and I hope with this series, I can help you peer into the token generating black box to show you how agents are built, how simple they are, and arm you with some learning along the way.
Understanding the basics:
For the sake of simplicity, i’m going to focus on the foundational models standard APIs, they all have their quirks, but I’ll focus on the commonalities.
Interacting with an LLM, like a Claude or OpenAI model is for all intents and purposes just a standard REST call:
The caller (the user, the code, whomever) sends a structured input.
That input can be:
text
image(s)
audio
even pdf with some vendors
Their api takes this structured content and returns some words. Thats it! Of course I am simplifying a little, the input payload does contain some specifics, the model you wish to use (in most cases), the limitations around tokens, etc. But at face value for typical models, that’s it.
Users send messages with a role of User, Assistants respond with the role of Assistant (this varies by provider, Gemini used to prefer model) that’s it.
There is one other special message Role, System (or more recently Developer, for OpenAI) this message is the models instructions on how to behave and is weighted higher for the models behavior. You’ll commonly hear this referred to as the ‘system prompt’.
The input can even be numerous messages combined in a list:

but wait a second, if it’s just REST, why do i get hit by countless words from a chain gun when i talk to these models?
Streaming:
As i described above, A simple request response is great, but model generation can take a while, how can we keep the user engaged? With streaming aka Server Sent Events (SSE):
When you request streaming (usually part of that message contract i mentioned earlier, with stream = true, gemini uses a path instead), the provider returns a 200 immediately with a very specific header:
Content-type: text/event-stream
this tells the listener to buckle up, we’re going to start spitting bars, and chunks of inference start flying off the model in realtime down the http response, for you to emit to your users or a stream, concatenate, whatever you like. with a natural conclusion when the stream complete.
fun fact: if you want the fastest possible response? use request/response, the nature of streaming, parsing, etc will add overhead. If you want speed to first token, or to update the UI in realtime, use streaming. Streaming and combining chunks, particularly when you’re dealing with JSON is a bit “ick” so only use it, if you need it.
API Specifications:
In the above example, i mentioned Claude and OpenAI, but there are a few more you should be aware of, from the foundational models:
Gemini Generate Content
Ollama Generate and Ollama Chat (forget this exists).
Chat completions is the near term OG spec, the one that most people utilized. If you watched the market grow, you would have seen many vendors like Ollama, Deepseek, Moonshot AI, Grok, etc all adopt Chat Completions as their standard.
So much was built on chat completions that Gemini readily supported it, Claude too, reluctantly, I suspect. OpenAI are aggressively abandoning this specification, but it’s still the most likely API specification in use today.
OpenAI’s Responses (And Open Responses) specifications followed shortly after, when they realised they had painted themselves into a corner. To socialize and encourage others to move over, OpenAI opened the specification up for others to play too.
Gemini’s GenerateContent (I’m not even sure this has a name, but whatever) Specification remained proprietary to them. Parts all the way down.
Anthropic’s messages API has been stable for two years now, and with the meteoric rise of Claude code, we’re seeing vendors (Kimi, Ollama, etc) now implement this standard to allow Claude Code to operate through their API’s.
Ollama has it’s own spec too, and probably the only one who truly saw that SSE was JSONL and treated it as such, but it’s largely dead.
All of this to say, there are a few protocols and their support varies from provider to provider. Some vendors do bridging for these things, like OpenRouter, but lets cover that later in AI Gateways.
Maintaining state:
So we send an input, we get a response, but how do we build on that? Easy, just replay the entire conversation each time!
Hey wait, that seems a bit simplified and redundant? You may say to yourself, but honestly, it’s not, it’s how it works 99% of the time.
While some later API’s like OpenAI’s Responses allow you to chain messages, this is exactly how Anthropic, Gemini, Ollama and OpenAI’s work! you can even use this approach with the Responses API if you choose.
For many years, until the advent of Functions and Tool calling, this is all we had. We would prompt the model to behave, or respond in a certain way (Personality or JSON formatting mostly) and Prompt Engineers earned themselves great outputs with a lot of tuning.
All was well, simply capture the response, append it, and send it all back to the model until? The context window.
The context window:
When you send words to an LLM, the llm converts them to numbers. I’ll spare you the gory details on the why and the how, but just bear in mind, that on average 4 characters are considered a single “number”, aka a token, and that sentence is now chopped up into a list of tokens.
When you take that large body of text, and chop it up, those tokens add up fast. in the early days of gpt 3.5 turbo, a considerable model in it’s time, could support roughly 16k tokens on input before it would start to get upset. it sounds like a lot right? 65 thousand characters? It really wasn’t, and isn’t.
The context window management was and remains painful even today.
Awareness of how many tokens you’re currently playing with is key. In some cases you can accurately count, in others they would “guestimate” by counting characters, or rely on the response count, which is great if you haven’t ALREADY sent too much. But awareness is necessary to avoid spurious errors.
As you begin to send messages and encroach the context window, there were two primary solutions you could rely on:
roll off / discard old messages
compaction
Rolling off old messages is risky and error prone, you couldn’t roll off the system prompt and the detail in the first messages were nearly always the needed breadcrumbs to the conversation.
Compaction on the other hand was tricky too, you prompted the model to compact the conversation, but inevitably precious context was lost when compacting.
Today, in tools like Claude code, you’ll see compaction, you can even request it on demand. In older versions of ChatGPT you would see compaction from time to time.
Lately though, Anthropic and OpenAI have baked compaction into their API’s and will do it on demand for you, just last night I found that the OpenAI responses specification has dedicated Compaction endpoints, too for explicit control. I fully expect this to become table-stakes, but not quite yet. On the topic of Anthropic, they also have some really nice tooling around reasoning and tool tokens, but more on that later.
Ok so we have text in, text out. Prompting. How do we actually do things? Functions.
Functions / Tooling:
In the early, dark days. Prompt engineers got pretty good at getting the LLM’s to respond in a reproducible way, they could prompt the llm to respond a certain way if it wanted to check the weather for example.
The llm would respond with a specific payload that their code would recognise, and call a function if the llm wanted to. This was noticed quickly by the industry and a pattern was formed.
By mid 2023 OpenAI was the first to introduce official function calling. For the first time, instead of dark arts, hopes, prayers and prompts, you could provide a structured model in JSON Schema to the LLM, along with a name and description and the model could “Elect” to use the tool.

Shortly after OpenAI’s announcement, with a new name, Anthropic announced “Tool Use”.
I am of course simplifying here, OpenAI and Anthropic’s approach were rooted in differences, but the result to consumers was the same. I can now write “Tooling” that the LLM can elect to use and that tooling can be code, hurray!
The llm would request the tool
the code would do the thing
The code would append the request and response to the messages
The cycle would continue.
LLM’s could now perform actions, request data and get data back to perform further functions.
And just like that, the agentic loop was born.
It was a beautiful, little re-entrancy loop, you would track the LLM’s response, was there a tool call, and you would handle the tools. The workflow was simply weave the LLM’s response, and tool result back into the original payload you had and it would continue! It was black magic and it was impressive.
Other vendors quickly reacted too, Gemini, etc added function / tooling support and we all went on our merry way, writing all fashions of advanced tooling, independently of each other until MCP arrived. (More on MCP later).
The astute reader will notice, that with two response types, streaming and request response, it’s not quite as simple as loop loop loop, streams present a challenge re-composing the deltas back to a full message, but i’ll cover some strategies for this in a further post. For the sake of simplicity, loop until finished is all you need to take away.
Quick detour, Server Side Tools:
A note to say, while many of us clambered to write tooling for common features, the providers saw this and features almost everyone was bolting on, websearch as an example became critically important:
The emergence of server side tools started to negate the need for web search tools and the provider simply charged you for the tooling they provided. You’ll see this heavily in Anthropics sandboxing, Websearch, etc. When a server side tool is declared as available in the message list, the server handles this tool internally, tells you they did it and handles the loop internally.
This is infinitely faster than handling it yourself, so where needed, this model is becoming more and more prevalent to tackle common problems and my advice is LEAN IN, deprecate your tooling early and often for the sake of speed and less to maintain.
Reasoning / Thinking:
So at this point, we now had LLM’s that could take increasingly larger context windows, return structured payloads or call tools, locally or over MCP. But, at the time, multi-step processes, including tooling, were tricky. We asked for something, the LLM understood the tooling it had available, but after a few cycles it would get confused, exit early or simply hallucinate a result. Often you had to prompt the LLM with the workflow steps:
“If asked to do W, do X, then Y and return Z.”
Which really negated the benefit of having a deterministic super power at the wheel if you had to steer it for every use case.
Thankfully, in and around the same time, concepts like chain-of-thought reasoning already existed and it was starting to gain traction with the frontier labs. It was around this time too, that DeepSeek came bursting out of nowhere, like the cool-aid man with their incredible model which showed the model reasoning in realtime.
“The model is thinking, look!” was awash on the usual social media circles and sure enough, for the first time, for many, internal prompted reasoning was seen by users.
With models “thinking” (and i use thinking loosely, there’s far more here, but lets stay simple) and the advent of “Interleaved Thinking” (thinking between tool turns) suddenly LLM’s were spending a LOT longer in their agentic loop, navigating challenging workflows, evaluating the responses and navigating towards their end goal.
The agentic loop became far less prompted, far less brittle. The loop was now starting to behave truly reliably with simple to large complexities with ease.
Model Context Protocol:
Meanwhile, in tooling land, with Functions / Tools in hand, LLM’s could now do the things, but making your LLM interactions useful, suddenly required many tools, of which everyone was writing independently.
The problem was clear and Anthropic answered, if everyone is out there writing tools, surely there’s a standard here we could apply to help share tools? MCP was born.
With MCP, you had a client and a server architecture. The client, ran close to your tooling loop and requested the tools available on initialization. With the tools in hand, you continued on your merry way just like with the old tool approaches.
MCP does a lot more than just tooling, but for the sake of this conversation, I’ll leave it here.
If the LLM requested a tool call, you just offloaded it to the MCP server and waited for the response. Easy stuff.
MCP took a distributed problem and created a solution to allow people to aggressively write, host or distribute Tooling bundles to be made available to LLM’s. The spec was simple and people went batshit writing their own tool stacks and servers.
People wrote ports of OpenAPI specifications, and wrapped them in public MCP servers hosted on Github, vendors added support (some badly, looking at you Atlassian), for a time you couldn’t avoid it. While MCP was born in 2024, MCP was the Agent hype-cycle of 2025 and it was impossible to ignore it.
Hype was high, but reality was close behind it, we’d opened the doors to many, MANY tools, the trustworthiness of these MCP servers were questionable and our context window could only handle so many without negative impacts. More on this later, but just know, the early MCP specs were wild, awkward and downright silly at times.
This massive hype cycle was driven predominately by the first emergence of the Agentic Loop users could touch and feel: Claude Code.
Claude Code - The birth of the “In-Process” agent.
in late 2024, Claude did something revolutionary. It launched it’s research preview of Claude Code. To this point, interactions with LLM’s were programmatic or chat orientated. But Claude Code, followed by Codex, Gemini, etc was the start of the Agentic Revolution in my eyes.
Claude Code’s premise was simple:
run the process.
tell it what to do.
plug in your things.
It does its thing.
Really, in the early days, it was largely: Watch it defecate in it’s hands and clap excitedly.
Claude took everything we have discussed in this article:
Messages
Reasoning
Function / Tooling
Looping
Conversation compaction
And showed the power of the harness that it could yield.
Claude presented pre-created tools to allow the LLM explore the local file system, read and write files, and run processes via Bash. For the first time, it was not request, response, instead you could steer your long running process, interleaving messages into the loop as it ran and interrupt the agent to course correct.
Taking all that we had up to this point, we now had an agentic loop that users could actually play with, which they did and it was a joyous mixture of diamonds in the rough for a few quarters.
We’ll cover this message interleaving in a follow up article, but to point out, if you’ve ever been frustrated by having to stop Claude or ChatGPT mid turn, because you forgot a fact or said something stupid, it’s quite easily fixed and is present in pretty much all agentic loops and co-work, but not in the chat interfaces.
Skills:
For a while, this is where we were, and we were relatively stable. If you wrote thoughtful MCP tooling, guarded that context window and kept the loop stable with resilience, the harness was powerful. Context windows continued to grow, LLM’s got better at long tasks, but some very interesting breadcrumbs dropped by Cloudflare late 2025 about what they called “Code Mode”.
The article itself is well worth the read, but if I was to summarize, “MCP is fine, but have you considered just letting the LLM write some code?”.
As we were seeing in the wild, Claude Code was doing a much better job at this point, and the LLM was more than capable of taking on meaty assignments. Cloudflare called out the need for a Sandbox, obviously, because writing and executing it’s own code was enough to give most people the terrified feeling, but it certainly was a pre-cursor to something greater.
A month later, Anthropic announces skills and the rest was history. Skills were a specification of YAML files, that contained instructions on how to do things, how to approach problems, how to think differently.
Skills offered:
Modularity: Developers could package domain knowledge, scripts and context retrieval logic into a directory. Skills enabled teams to share reusable capabilities across projects.
Progressive disclosure: Only the skill name and description are initially loaded; the full instructions are read when relevant, keeping the prompt compact .
Bash integration: Claude uses its built‑in bash tool to read
SKILL.mdand to execute any scripts included with the skill . This confirmed the community’s belief that giving agents a shell enables them to manage their own context and code.
Skills combined prompts, scripts, and leveraged in place tooling for the LLM to start working with, and the best part was? Anyone could write them, no plumbing or protocol, no hosting:
Just tell the LLM what you wanted, have it write a skill, iterate until you were happy.
Nobody needed to write functions, nobody needed to write MCP tooling, just an idea or process, some documentation, bash commands. You could scan for skills and provide them as breadcrumbs in the prompt and be on your merry way, and even that was optional, telling the model that the directory existed was often enough.
Skills exploded in popularity, were adopted by the now rich ecosystem of claude code style harnesses, like Codex, Pi, Gemini Cli, etc. and really upset the sycophantic MCP hype trajectory.
They addressed so many use cases, were great if your agent was running in your file system, but it wasn’t all rosey in the garden.
Skills particularly excelled at things that MCP failed at, context heavy tasks you did not want to route via the LLM, such as downloading, uploading etc, but skills suffered a problem MCP tackled quite succinctly, Cloud Agents.
How do we get these little nuggets of functionality into a cloud based agent? And should they require credentials, how does one handle that? More on that later.
Note: While i was browsing around, i spotted a really impressive, forward thinking blog by Adam Lucek here. Well worth the read if you want the technical deepdive.
And that’s it for part 1! I hope you found this interesting or it answered some fundamental questions you weren’t sure about!
In the follow up episodes I’ll be covering:
Part two:
Agentic loop patterns, making it easier on yourself.
Guardrails, how trivial guardrails work.
Advanced Context Window management.
Part Three:
Swarms and Orchestration, handoff vs collaboration.
Moving the Agents from the operating system to the cloud.
Sandboxing, a pragmatic approach.
Part Four:
Building gateway services.
Telemetry and Governance.
If you enjoyed this, or would like a topic covered drop me a comment!
A


















Nice article, well done! part 2?