Skip to content
Profile
E. Burgos
October 5, 2026By Esteban Burgos8 min read

From loose AI to agentic development: how we adopted SDD in a team at YPF

#SDD#Spec-Driven Development#desarrollo agéntico#IA en equipos de ingeniería#YPF
From loose AI to agentic development: how we adopted SDD in a team at YPF

When I joined a development team at YPF, AI was already part of the daily work: all of us used it, but each person with their own judgment, their own prompts, and results that were hard to compare. I proposed adopting SDD (Spec-Driven Development) as a shared framework, and we implemented it step by step together with the team. In this article I describe how that process went: where we started, what steps we took, and what changed in the way we work. It includes no code or internal YPF information, but it does include the process detail needed so that it can be useful to other teams in a similar situation.

The problem: loose AI, no process

Before SDD, the situation was the typical one for a team that adopted generative AI organically: each developer had their own workflow with Copilot, ChatGPT, Claude, or whichever assistant was on hand. Some pasted giant prompts into the chat, others let the agent touch half the repo without review, others didn't use it at all for anything critical out of distrust.

The symptoms were clear:

  • No documentation of the "why." The code changed, but there was no record of what was decided and why. When someone came back to touch that part three weeks later, they had to rebuild the context from scratch.
  • No cost measurement. Nobody knew how much was being spent on tokens, or whether that spend translated into real speed or into more rework.
  • Uneven quality. A PR generated with AI could be flawless or could bring a silent regression, depending on how good a prompt that person had written that day.
  • Little traceability between teams. Operations backoffice, funds disbursement, and platform catalog —the three fronts where we later introduced SDD— each lived with their own informal conventions, with no shared standard.

In short: AI was generating value, but it was also generating entropy. And entropy in an engineering team, sooner or later, gets paid for with incidents or invisible technical debt.

The proposal: Spec-Driven Development as a process

SDD is not "write a document before coding." It's a process where the specification is the source of truth, the work cycle is defined, and AI agents operate within specific roles, with gates that prevent moving forward if something isn't ready.

In practice, this translates into three pieces:

Specs. Every relevant change starts with a spec: what problem it solves, what its scope is, what acceptance criteria apply. It's not bureaucracy because the spec is short and lives in the repo, versioned alongside the code. It's the contract between what was requested and what gets built.

Cycles. Work is organized into cycles with clear stages —design, implementation, validation— and each stage has a concrete output that the next one can consume. This gives the AI a bounded context instead of "the whole repo at once."

Agents by role. Instead of a generic agent that does everything, we defined agents with specific responsibilities: one that helps draft and refine the spec, another that implements within the defined scope, another that reviews against the acceptance criteria, another that documents. Each role has its own context and its own constraints.

And here comes the part that changed the team's behavior the most: the gates. A gate is a control point that stops the flow if something doesn't meet a minimum condition —an incomplete spec, ambiguous acceptance criteria, lack of review— and doesn't let you move to the next stage until it's resolved. It's not a "please review this," it's an actual block in the workflow. That helped us reduce variability: the outcome no longer depended on each person's individual discipline, but on the process.

Full, reduced, and lite flows: not every change carries the same weight

One of the most common mistakes when adopting SDD is applying the same rigor to a one-line change as to a brand-new end-to-end feature. That complicates adoption quickly, because the team starts feeling like the process is heavier than the problem it solves.

That's why we work with three flows depending on the size of the change:

  • Full: for new features or changes with cross-cutting impact. Full cycle, detailed spec, all role-based agents, all gates.
  • Reduced: for medium-scope changes, where risk exists but is bounded. Lighter spec, fewer stages, critical gates still active.
  • Lite: for small fixes, minor tweaks, low-risk tasks. Minimal spec, short cycle, just enough gates to not lose traceability without slowing down velocity.

This segmentation was key so that SDD wouldn't feel like a bureaucratic layer imposed from above, but rather a process that adapts to the actual size of the problem.

Multiple devs, same repo: per-author IDs and additive context

Another concrete challenge: when several developers work in parallel on the same repository using agents, context can overlap. If two people are in different cycles and one agent's context bleeds into the other's, you get contaminated responses and decisions that don't belong to that branch of work.

The solution we implemented was simple: per-author IDs. Each developer has their own context identifier, so the agent knows which cycle and which spec it's dealing with at any given moment, without mixing different work sessions.

On top of that, we use additive context fragments: instead of rewriting the entire context every time something changes, incremental fragments are added that document decisions and progress. This avoids two typical problems: losing history (if everything gets overwritten) and overloading the agent (if it's sent the full history on every interaction). The agent consumes only the fragments relevant to the stage it's in.

Where we applied it at YPF

We implemented SDD on three fronts with different levels of rigor, depending on the risk and maturity of each team:

  • Operations backoffice: there, SDD is mandatory. It's a domain where errors carry a high cost and decision traceability is fundamental, so we left no room to work outside the process.
  • Funds disbursement: we applied role-based agents strictly, given that it's a sensitive domain where each stage —design, implementation, review— needs a specialized agent bounded to its function, with no overlap of responsibilities.
  • Platform catalog: a domain with relatively lower risk, where we were able to iterate faster and use the reduced and lite flows more freely, validating the process before taking it to the more critical fronts.

This combination —mandatory where risk demands it, flexible in flow where it doesn't— helped adoption not feel like a uniform imposition, but rather a process calibrated to the reality of each team.

SDD-Harness: the tool behind the process

To keep all of this from remaining abstract methodology, we use SDD-Harness, an open source project I published on npm (current version 0.15.1). It includes 8 role-based agents and 19 skills that cover the different stages of the cycle: from spec drafting to validation against acceptance criteria. It's the layer that operationalizes specs, cycles, gates, and flows so that each team doesn't have to reinvent the wheel.

Being open source, any team can install it, adapt it to their own domain, and start measuring results without depending on a closed platform.

How to get started: diagnosis, pilot, and rollout

For teams in the position we were in before this —loose AI, no process, no cost visibility— the sequence that worked best for us was:

  1. Diagnosis. Before touching anything, understand how each person uses AI today: what tools, what type of tasks, what level of review exists. Without this step, any proposed process will clash with invisible habits.
  2. Pilot in one repo. Choose a bounded domain, neither the most critical nor the most trivial, and run SDD there first. Define specs, cycles, role-based agents, and the corresponding flow (full, reduced, or lite depending on the size of the changes made in that repo).
  3. Rollout. With results from the pilot —speed, quality, traceability— extend to other repos, calibrating mandatoriness and rigor according to the risk of each domain, just as we did across operations backoffice, funds disbursement, and platform catalog.

There's no need to convert the whole team overnight. What matters is that the process exists, that it's measurable, and that the gates genuinely stop what isn't ready.

What I took away

  • Gates reduce variability better than asking each person for discipline: the process sustains what individual judgment can't.
  • Rigor has to be proportional to the change; a single level of strictness for everything makes the team abandon it.
  • Starting with a pilot in a bounded domain lets you validate the process before taking it where the risk is higher.
  • With several developers in the same repo, separating context by author and recording it additively keeps agents from mixing decisions.
  • Adopting a shared framework is, above all, teamwork: it gets adjusted along the way rather than imposed overnight.

Closing

Moving from "everyone uses AI their own way" to a process with specs, cycles, role-based agents, and gates changed how we document, how we measure, and how much we trust what AI produces. If you're working on something similar or your team is evaluating how to incorporate AI, I'd be glad to exchange experiences: feel free to reach out on LinkedIn. If it helps as a reference, my website describes how I support teams through this kind of adoption: Agentic SDD adoption.

Written by Esteban Burgos
ESTEBAN BURGOS

About the author

ESTEBAN BURGOS · Forward Deployed Engineer

I build software end to end for startups and companies: websites, platforms and AI solutions. I write about what I learn on real projects.

See my background

Want something like this for your company?

Tell Tuki about your idea and get a price range in your currency within minutes.

Keep reading

Comments

Be the first to comment.