How I built Tuki: the RAG-powered quoting tool that's already in production
I'm Esteban Burgos, Forward Deployed Engineer, and this is the post I wish I had read before putting an AI-powered assistant into production. It's not a "how to do RAG in 10 minutes" tutorial. It's what I decided, what went well, and what broke while building Tuki, the quoting tool that now lives at estebanburgos.com.ar/quote and provides price ranges based on the Argentine freelance market.
If you're evaluating adding a conversational assistant to your product, this post is for you: not so you copy code, but so you anticipate the decisions that really matter when the assistant stops being a demo and starts receiving real traffic.
What Tuki is and why it exists
Tuki answers a very concrete question: "how much does a project like the one I'm thinking of cost?". Instead of a static form or a generic price table, it's a conversation that uses a proprietary knowledge base on how freelance work is quoted in Argentina, and it returns a reasonable range adjusted to what the visitor describes.
The most important design constraint was this: it had to be useful without becoming a source of unpredictable expenses or an abuse vector. An assistant that "works" in a demo and an assistant that survives real traffic, bots, and creative users testing the limits of the prompt are two different products. Everything that follows is built from that tension.
Architecture: Next.js, AI SDK, and tool calling
Tuki runs on Next.js, which I was already using for the rest of the site, so I didn't add a new system: the chat lives as just another route in the same app. For the AI layer I used Vercel's AI SDK, which today is basically the standard for not being locked into a single model provider. That gives me two things I value as a Forward Deployed Engineer: the ability to switch models without rewriting business logic, and a streaming and tool calling API that I don't have to reinvent.
The model doesn't quote "from memory". It quotes using tools: when it detects it has enough information about the project (type of work, scope, urgency), it calls a function that searches the knowledge base and another that assembles the response with the corresponding range. This is key to the next point: the model doesn't make up numbers, it looks them up.
Separating "reasoning about the conversation" from "getting the correct data" is the architecture decision I'd repeat the most if I had to do it again. The LLM is good at understanding what the user is asking for and carrying the conversation; it's much less reliable at making up figures. Tool calling solves that division of responsibilities without needing fine-tuning or fragile prompt engineering tricks.
RAG over a proprietary knowledge base
Tuki's knowledge base isn't a giant PDF or internet scraping: it's proprietary content about how a freelance quote is structured in Argentina, with its nuances (project type, complexity, timelines). That content is split into chunks, embedded, and stored in Postgres with pgvector.
When a query comes in, it's also embedded and matched for similarity against those vectors to bring in only the relevant context before the model generates the final response. It's RAG in its most boring form and, for this case, the most correct one: no multi-stage pipelines, no sophisticated re-ranking. The knowledge base is limited and curated by me, so retrieval doesn't need to be heroic to be accurate.
pgvector without an HNSW index: a conscious decision, not an oversight
This is the part that generates the most questions when I mention it: I use pgvector without an HNSW index. It's not that I don't know it exists, it's that I decided not to use it, for now.
HNSW (and its alternative, IVFFlat) exist to speed up similarity searches when you have millions of vectors and need to avoid a full sequential scan. The problem is that this index isn't free: it consumes memory, adds maintenance complexity, and, at small volumes, the latency improvement is marginal compared to a sequential scan on a table that fits comfortably in RAM.
Tuki's knowledge base is the size of "everything I need to quote well", not "the entire internet". At that volume, an exact cosine similarity search without an approximate index responds in milliseconds that aren't even noticeable next to the latency of the generative model, which will always be the real bottleneck. Adding HNSW today would mean optimizing the part of the system that isn't the problem, at the cost of adding one more piece to maintain and introducing approximation where I currently have exactness.
The lesson for a CTO evaluating this: don't adopt "LLM startup scale with millions of documents" infrastructure if your knowledge base is curated and small. Measure the actual volume before copying the architecture from a blog post that solves a problem ten times bigger than yours.
Limits per IP and per conversation to control costs and abuse
An assistant connected to a paid model is, as soon as you publish it, a variable cost surface that you don't fully control unless you set explicit limits. In Tuki the limits are simple and deliberately set:
- 30 messages every 10 minutes per IP.
- 40 messages per conversation.
The first cuts off abuse and bots trying to flood requests. The second prevents a single conversation from stretching on indefinitely, whether from a legitimate user going back and forth without reaching a conclusion or from someone trying to force the model off-topic through sheer message volume.
Neither limit replaces cost monitoring, but they act as a floor: they're the difference between "a weird traffic spike costs me a bit more that day" and "a weird traffic spike generates a bill I have to explain." For anyone evaluating adding an AI chat to their product, this isn't optional and it isn't an implementation detail: it's the first thing you need to design, before writing the first prompt.
Price in the visitor's currency: MEP dollar for Argentina
Tuki detects the visitor's country and adjusts the quote's currency accordingly. For Argentina, that means showing the range using the MEP dollar as a reference, instead of forcing each user to mentally convert a number in another currency.
This sounds like a cosmetic detail but it isn't: a quoting tool that shows a price in the "wrong" currency for the visitor's context generates friction and distrust before the conversation even begins. Detecting the country and choosing the correct conversion reference is part of what makes the range feel like something designed for that person, rather than a generic table forcibly translated.
What failed in production and how it was resolved
No system with an LLM in the middle behaves in production exactly as it does in the testing environment, and Tuki was no exception. The rate limiting thresholds per IP and per conversation that I mentioned earlier weren't born in the initial design with those exact numbers: they were adjusted after seeing real traffic behavior, which is exactly the kind of signal that can't be fully simulated before going live.
The other constant source of adjustment was the separation between the model's reasoning and obtaining the price data through tools. Any conversational quoting tool that lets the model "calculate" or "remember" a number instead of fetching it from a controlled source will sooner or later show inconsistencies between two similar responses. Reinforcing that the price always comes from the tool and never from the model's free generation was, more than a one-off fix, the criterion I've used to review every change I make to Tuki ever since.
If you're about to launch something similar, the concrete advice is: instrument from day one so you can see message volume per IP, per conversation, and the proportion of responses that come from tool calls versus free generation. Without that visibility, any post-launch adjustment is done blind.
Resuming conversations: browser or link
A quoting conversation rarely closes in a single exchange. People come back, compare, check with someone else. That's why Tuki allows resuming the conversation in two ways: automatically from the browser, if you return on the same device, or by sharing the link in the format ?tuki=<id>, which takes you straight back to the exact state where you left it.
It's a decision that seems small, but it changes the perception of the assistant: it stops being an ephemeral chat that resets every time and starts behaving like part of a real quoting process, something you pick up over time, not something you use once and discard.
What I'd tell a tech lead before starting
If you're evaluating adding an AI assistant to your product, the easy part is connecting a model and making it respond with something coherent. The hard part, and the one that decides whether the project survives in production, is everything around that: separating what the model reasons from what the model looks up in a reliable source, setting usage limits from day one even if they seem conservative, and choosing infrastructure suited to the actual volume you'll have, not the volume a use case ten times bigger would have.
Tuki is small on purpose. That is, to a large extent, the reason it works.
If you want to see all of this working in a real case, go to https://estebanburgos.com.ar/quote and try quoting a project.

About the author
ESTEBAN BURGOS · Forward Deployed Engineer
I build software end to end for startups and companies: websites, platforms and AI solutions. I write about what I learn on real projects.
See my backgroundWant something like this for your company?
Tell Tuki about your idea and get a price range in your currency within minutes.
Keep reading
- LangChain in 2026: the key to intelligent chatbots with RAG, tools, and unbeatable UXDiscover what LangChain is and why it has become one of the central pieces for building modern intelligent chatbots. We go over how to combine RAG, chat-tools, and a smooth user experience to create assistants that truly solve problems, not just answer questions.
- Are devs trembling? The illusion of the technical user and the true ceiling of AIThe debate on whether artificial intelligence will replace developers is often framed as a false dilemma: **the traditional technical expert versus the average user empowered by an intelligent agent**. However, the real friction is neither syntactic nor superficial. It's not about who writes a Python function faster or who sets up an automation flow in half an hour. The real gap is **architectural and logical**.
- Podcast: No more spaghetti code, how SDD-Harness brings order to development with AIThe speed at which artificial intelligences can write code today is astonishing, but it comes with a very high hidden cost: **technical anarchy**. When multiple developers integrate AI agents into their daily workflows in an improvised manner, the result is often inconsistent code, silent regressions, undocumented patches, and an alarming loss of traceability. Ultimately, a new type of **"agentic spaghetti code"** that stifles long-term software maintainability. To solve this problem at its root, I designed **SDD-Harness**, a portable, language-agnostic command-line interface (CLI) framework that implements the **Spec-Driven Development (SDD)** methodology. The premise is as simple as it is inflexible: **no line of code is written without a formal, structured, and pre-validated technical specification**. Below, I share a deep dive into how the tool works and why agent governance is the indispensable step to scale AI-assisted software engineering.
Comments
Be the first to comment.