Skip to main content

Search...

Agentic Engineering: AI Accelerates Existing Principles

Agentic engineering is less about prompts than about the agent environment: your prompt is about 20 of 20,000 tokens. Cleanup agents keep quality up.

• • Updated: • 12 min read
Cover of the expert talk on 'Agentic Engineering: AI Accelerates Existing Principles' with Benedikt Stemmildt and Richard Seidl.

Agentic engineering is a systematic approach to building software with AI agents, in which architecture, quality principles and the design of the agent environment matter more than writing individual prompts. Good principles speed up progress, bad principles speed up decline. Specialized cleanup agents take care of ongoing tasks such as deduplication, refactoring and test coverage.

Key Takeaways

  • With bad principles, AI makes you worse faster. With good principles, it makes you better faster, because AI simply accelerates whatever course you are already on.
  • Agentic engineering replaces prompt engineering as the core skill: your own prompt accounts for about 20 of the 20,000 tokens in the context, so what really counts is how the agent environment is configured.
  • A cleanup agent that runs continuously and checks every change for duplication, missing tests and quality attributes takes over the manual refactoring that used to follow the first draft.
  • Five developers all working with agents at the same time step on each other’s toes, because agents touch the entire stack right away. The consequence is smaller cross-functional teams with a clearly partitioned architecture.
  • When you switch to a new model, it pays to throw away all your existing skills and configurations, because outdated settings make the model’s new default behavior worse instead of better.

Vibe Coding or Agentic Engineering: Why the Term Matters

Agentic engineering is a more precise name for AI-assisted software development than the popular term vibe coding. Andrej Karpathy, who coined “vibe coding” on X a little over a year ago, later corrected himself and now prefers agentic engineering himself.

The difference lies in the word. Engineering brings architecture, quality assurance and sound engineering practice back into the conversation. The point isn’t just to generate code. It is engineering work, with principles and a method.

That doesn’t make existing knowledge about architecture, paradigms and quality any less valuable. It becomes more valuable. The job shifts toward systematizing that knowledge and putting it into agents and their configuration, instead of keeping it only in developers’ heads.

Why the Prompt Matters Less than You Think

In agentic setups, how you phrase the prompt hardly makes a big difference anymore. Classic prompt engineering, which was the big topic a year ago, still matters when you interact with a model directly, but in agentic setups it carries less weight.

The reason is context. Through its system prompt, an agent already sends around 20,000 tokens in which it introduces itself to the model and describes its tools, such as which files it can read and edit. A 20-token prompt of your own barely registers next to that.

The prompt’s real job is to let the agent find the information it needs by itself. The better the environment is configured, the less the prompt has to say. Known difficulties go into the prompt. Difficulties you have already solved belong in the configuration.

Reverse Engineering Beats Better Prompts

The bigger lever is watching the model’s default behavior, not polishing the prompt. You deliberately start with a weak prompt and no configuration and see how the model behaves.

The fix follows from what you observe. Instead of asking what a better prompt would look like, you ask what was missing from the agent’s context file or tools that would have let it solve the task with the weak prompt alone.

This works through a guided conversation. You steer the agent by hand and correct its course, and once the result is right, you have it review the conversation: where was it right, where did it have to redo work, where did you step in? Those findings go into the configuration file, so the same corrections aren’t needed next time.

It is a lot like a retrospective in agile work. You ask what went well and what didn’t, and you capture the essence. Tool vendors have since recognized this pattern as a best practice and built it into their products, for example as a command in Claude Code that reviews the conversation and writes down the findings.

A Cleanup Crew Tidies Up After the First Draft

Instead of making the code perfect on the first pass, separate agents handle the cleanup in the background. One such agent loops continuously over all changes and looks for ways to consolidate and simplify.

This division of labor can be specialized. One agent checks security, another checks quality attributes, a third does nothing but deduplication. If it notices that someone has built a new modal window even though a component for it already exists, it refactors the code to use that component. If a new function is missing a test, it adds one.

That changes the bar. Building cleanly from the start is still better, but a feature doesn’t have to be perfect on the first pass if another agent tidies up afterward. Many developers already work this way: ship something first, refactor later. The only difference is that they no longer do the refactoring themselves.

Why Choosing a Model Becomes an Architecture Decision

Switching models devalues the configuration you worked so hard on, because every model comes with different defaults and a different system prompt. The intuition you build through reverse engineering is always tied to a specific model and a specific harness.

The recommendation is to work with the current state-of-the-art models and to commit to one model as a rule. Splitting by task makes sense: create the plan with an expensive model and the implementation with a cheaper, faster one, because the planning step has already gathered the context that’s needed.

When a new model comes out, it pays to throw away your skills and configurations and start over. Otherwise the old skills push the new model in the wrong direction and hold it back. That is the worst case: a new, more capable model dragged down to its predecessor’s level by an old configuration.

One thing to keep in mind: new models usually don’t get smarter in the sense of knowing more. They often have the same training cutoff as their predecessor. What changes is that they act closer to what a person would expect for the task.

The Production Line Replaces the Single Feature

The next stage stops building the feature and builds the line that produces features. Instead of implementing a search function yourself, you build a production line that turns incoming feature requests into an implementation, in the context of your specific system and company.

This approach also handles the model question more cleanly. In a production line, you can hardwire the right model into each step and test that choice deliberately. Manual work doesn’t follow a fixed process that closely, while a production line lets you assign a model to each step explicitly.

A Different Abstraction, Not a New Compiler

AI isn’t simply the next level of abstraction, the way Java sits on top of assembly. The comparison falls short because there is no compiler that guarantees a deterministic result.

With assembly, nobody cares about the generated code, because the compiler translates reliably. With AI, that guarantee is gone. If you no longer want to look after the code yourself, you need other safeguards, such as the cleanup agent, because you can’t count on a deterministic translation.

That shifts the skills that matter. The code moves into the background, and the principles and approaches that produce good code move to the front. That is architecture work, and there is more of it than before.

Agentic Engineering Accelerates the Principles You Already Have

Bad principles plus AI lead to worse results faster, good principles plus AI lead to better results faster. Current studies from the DORA research community already factor in the effects of AI and show exactly this mechanism.

The consequence has two sides. Teams that bet on AI blindly now may simply hit the wall faster. And the gap between high-performing and struggling organizations is getting wider, not smaller.

The remedies aren’t new. Unit testing, integration testing, shift left and a working pipeline are homework you have to do anyway. If you produce pull requests 500 times faster without them and then put a central QA department in front of them, you lose the productivity gain right away.

“If you have bad principles and apply AI, you get worse faster. And if you have good principles, you get better faster.”

(Benedikt Stemmildt)

Why Teams Step on Each Other’s Toes in Agentic Software Development

Most experience reports come from individuals. Hardly anyone has figured out how to work agentically as a team, and that is exactly where the biggest problems show up.

Five developers who all work with agents get badly in each other’s way. In classic planning, you split up the work and everyone takes a slice. With agents, everyone develops across the whole stack from the start. Merge conflicts are the smaller problem, and the AI resolves them itself. The harder part is keeping any overview of how the product is evolving.

The answer is smaller units. Product manager, product owner and developer are increasingly merging into one role: the product engineer. Three of these people plus an agent make a team that works well as an ensemble.

Team Structure and Architecture Have to Fit Together

Cross-functional teams are the prerequisite. Separate backend and frontend teams can’t be sensibly cut into smaller units for agentic work, and five frontend teams plus five backend teams don’t add up to a workable structure.

Cutting up the organization alone isn’t enough. Conway’s Law and sociotechnical architecture come together here. The architecture has to be split so that the agents don’t step on each other’s toes in the code either.

In practice, that points to self-contained systems or a modulith rather than microservices, where teams risk blocking each other again. Units that are self-contained end to end give agents and teams working in parallel the room they need.

Systems with AI Inside Are Built Differently from Deterministic Software

As soon as a system contains AI itself, for example an agent that reads through a regulation, the development process changes fundamentally. Such systems aren’t deterministic like the software we have built so far.

That calls for smaller iterations and a lot more trial and error. With traditional software, a developer has an idea of the solution and implements an algorithm. With systems that contain AI, you have to experiment much more, because you can’t pin down the behavior in advance.

This is where old management models hit their limits. Detailed requirements and specification documents are out of the question anyway, but even agile principles are no longer enough for this kind of development.

Frequently Asked Questions

Is prompt engineering still worthwhile when working with coding agents?

It remains relevant for direct interaction with a model, but its importance diminishes in agent-based setups. An agent already sends about 20,000 tokens via its system prompt, in which it describes itself and its tools. In comparison, a custom prompt of 20 tokens is hardly significant. Known challenges belong in the prompt; resolved challenges belong in the configuration.

Does traditional architectural and quality knowledge lose value due to AI agents?

No, it becomes more valuable. The term “Agentic Engineering” emphasizes exactly that: architecture, quality assurance, and an engineering-based approach. What is shifting is the location of that knowledge. Instead of remaining solely in the minds of developers, it must be systematized and transferred into agents and their configurations.

How do you figure out what’s missing from an agent’s configuration?

By observing default behavior rather than relying on better prompts. You deliberately start with a poor prompt without any configuration, manually guide the agent through the conversation, and then have it evaluate at the end where it got it right, where it needed to correct itself, and where you had to intervene. These insights are incorporated into the configuration, similar to a retrospective.

Does the code have to be clean on the first pass when agents are involved in development?

Not necessarily, if a separate agent cleans it up afterward. A cleanup agent running continuously goes through all changes and looks for ways to consolidate and simplify them. The work can be specialized: one agent checks security, another checks quality characteristics, and another focuses solely on deduplication. If a modal component already exists and someone builds a second one, the agent merges the two. It also adds any missing tests.

Can agent configurations that have been developed be transferred to a new model?

It’s better not to. Each model comes with its own defaults and system prompts; the intuition gained is tied to the specific model and harness. Old skills steer a new model in the wrong direction and, in the worst case, drag it down to the level of its predecessor. New models are rarely more knowledgeable anyway; they often have the same training cutoff and simply behave more closely to expectations.

Is AI-assisted development simply the next level of abstraction, like a high-level language over assembly language?

No, the comparison doesn’t hold up because there’s no compiler to guarantee a deterministic result. With assembly language, the generated code doesn’t matter because the translation is reliable. That guarantee doesn’t apply here. Anyone who no longer wants to worry about the code needs other safeguards, such as cleanup agents and robust principles.

Why do some organizations benefit more from AI in development than others?

Because AI accelerates the existing trajectory. Poor principles combined with AI lead more quickly to worse results, while good principles lead more quickly to better ones. Studies from the DORA community demonstrate this mechanism, factoring in the effects of AI. This widens the gap between strong and weak organizations. Unit testing, integration testing, “shift left,” and a functioning pipeline remain prerequisites.

What architecture works best when multiple agents are working on the same product in parallel?

End-to-end self-contained units (in practice, self-contained systems or a “modulith”) rather than microservices, which risk creating new bottlenecks. A cross-functional team structure is essential: separate front-end and back-end teams cannot be meaningfully subdivided further. Simply breaking up the organization isn’t enough; the code must also give the agents room to operate.

Share this page

Related Posts