From Chatbot to Agent: How I Designed an AI Assistant That Executes Actions with Human Confirmation

For a while now I've been noticing the same pattern in the product teams I work with: they all have a chatbot. It answers frequently asked questions, searches the documentation, sometimes even sounds intelligent. But at some point in the conversation with the client, the uncomfortable question comes up: "and can this actually do something, or does it just talk?"
That's where the difference between a chatbot and an agent comes in. A chatbot responds. An agent operates your platform: it creates a resource, updates a status, triggers a process. And that leap — from responding to executing — is the one that adds the most value, but also the one that demands the most responsibility. In this post I'll tell you how we approached it, what pieces we used, and why human confirmation isn't a detail, it's the heart of the design.
The FAQ chatbot is no longer enough
An assistant that only answers questions about your product has a low ceiling. It's useful for reducing support tickets, sure, but it doesn't change the way a user interacts with your platform. The question I ask every product team that consults me is simple: what would happen if this assistant, besides explaining how to do something, could do it directly?
That's where function-calling comes in: instead of the model only generating text, we give it the ability to invoke concrete functions that operate on the real entities of the business. It's not a generic "connect your API" integration; it's modeling which actions make sense for the agent to be able to execute on your entities —an order, a user, a subscription, a ticket— and exposing that as tools that the model can decide to use depending on the context of the conversation.
RAG on the organization's state: so the agent knows where it stands
For the agent to make good decisions it needs something more than the user's instructions: it needs real context about the current state of the organization. That's where RAG (Retrieval-Augmented Generation) comes in, but not only applied to static documentation, but to the live state of the platform: which resources exist, what situation they're in, which business rules apply.
We built our own RAG engine that wasn't designed exclusively for the agent. We reuse it on three fronts: the landing page (where it answers visitors' questions about the product), the documentation (where it helps users understand features), and the agent itself (where it feeds the decisions before proposing an action). One single engine, three surfaces. This isn't just development efficiency: it means there's a single source of truth, and when you update the knowledge base, it updates for all three use cases at the same time.
For embeddings we use pgvector directly on Postgres. We didn't add a separate vector database for this: if you already have Postgres running your business, it makes sense for the vectorized knowledge to live there too, with the same backups, the same infrastructure, the same query language you already know.
Function-calling with a handbrake: human confirmation
Here we arrive at the point I'm most interested in conveying in this post, because it's the one that makes the difference between a useful agent and a risky agent.
When the agent decides that an action needs to be executed —creating something, modifying something, triggering a process— that action is not executed directly. It's recorded as a pending action, which in our implementation we call AiPendingAction, and it remains in that state until a human reviews it and confirms it explicitly.
This completely changes the conversation with the client about "AI that acts on its own." We're not delegating blind control to a language model; we're giving the agent the ability to propose with context and precision, but leaving the final decision where it needs to be: with a person. The agent does the heavy lifting of understanding what needs to be done and preparing it; the human does the light work of saying "yes, go ahead" or "no, not like that."
For product teams this tends to be a relief rather than a limitation. The first question I get asked when I present this architecture is "what if it makes a mistake?" With mandatory human confirmation before any execution, that question loses urgency: the worst case isn't the agent executing something wrong, it's proposing something wrong and someone rejecting it. That's a huge difference in the level of risk you're willing to take on when launching this kind of feature.
Controlling cost: measure before spending, charge upon confirmation
An agent that does function-calling and RAG can quickly become expensive if you don't carefully design the lifecycle of each interaction. The logic we apply is fairly straightforward: first we measure, then we spend.
Before executing the heaviest part of the flow —the one that actually consumes model resources significantly— we evaluate whether the request warrants that expense. And in flows where the agent proposes a concrete action, the charge (or the quota consumption, depending on the platform's business model) happens at the moment of confirmation, not at the moment the agent simply "thought about" the action. This aligns the actual cost with the value delivered: you don't charge anyone for a proposal that later gets discarded.
As the main provider for language capabilities we use Gemini, with a fallback to OpenRouter when additional resilience or access to other models is needed. This combination gives us room to operate without depending on a single point of failure, without that implying reinventing the architecture every time a new model appears on the market.
From responding to operating, without losing control
If your product already has an FAQ chatbot, you're probably at the right point to take the next step. The key isn't to add more "intelligence" to the assistant just for the sake of adding it, but to clearly define which business actions make sense for an agent to be able to propose, with what context (RAG on the real state of your organization) and under what control mechanism (human confirmation before executing).
That combination —function-calling on real entities, RAG reusable across your whole platform, pending actions until human confirmation, and cost control that measures before spending— is what separates a chatbot from an agent that truly operates your product without you losing control over what happens.
If you're thinking about taking that leap in your platform, I invite you to visit estebanburgos.com.ar/servicios/inteligencia-artificial and let's talk about how to apply this to your specific case.

About the author
ESTEBAN BURGOS · Forward Deployed Engineer
I build software end to end for startups and companies: websites, platforms and AI solutions. I write about what I learn on real projects.
See my backgroundWant something like this for your company?
Tell Tuki about your idea and get a price range in your currency within minutes.
Keep reading
- How I built Tuki: the RAG-powered quoting tool that's already in productionA first-hand account of Tuki, the RAG-powered quoting assistant running in production at estebanburgos.com.ar/quote. I go over the architecture decisions, why I chose pgvector without an HNSW index, how I keep costs under control with rate limiting, and what broke along the way.
- LangChain in 2026: the key to intelligent chatbots with RAG, tools, and unbeatable UXDiscover what LangChain is and why it has become one of the central pieces for building modern intelligent chatbots. We go over how to combine RAG, chat-tools, and a smooth user experience to create assistants that truly solve problems, not just answer questions.
- DisplayAds: the real architecture behind a digital signage SaaS that reaches all the way to the hardwareAn end-to-end walkthrough of DisplayAds, a production product that spans a real-time NestJS API all the way to Raspberry Pi and kiosk-mode Android apps. I share the 8 pieces of the system, how remote updates work with autonomous rollback, the tech stack, and how I kept the project organized with Spec-Driven Development and per-agent cost telemetry.
Comments
Be the first to comment.