Understanding AI agents · 9 min read

Why your AI agent is slow, and what to fix before changing models

When an agent takes nine seconds to respond, the first impulse is to change models. That is rarely where time is lost. The wait is divided across three layers: how long the model takes to write, how long the tools it queries take to respond, and the time added by the sequence of steps. Three API calls launched one after another take the sum of their durations; launched together, they take the duration of the slowest. And before all that, a business owner's question: is anyone actually waiting in front of the screen?

A small robot sweats while struggling forward with a stack of packages in its arms.

Why does an AI agent take several seconds to respond?

Because what we call “the agent” is not one program, but a sequence of steps. The n8n team, whose automation platform is used to build the systems we install, divides this wait into three layers (September 19, 2026 article):

  • Model time. It reads what it is given, then writes its answer word by word. Reading is fast; writing is sequential, so it is proportional to the length of the answer.
  • Tool time. Every time the agent retrieves information elsewhere (your catalog, CRM, document database, or a supplier API), it waits for the response. These calls often take several seconds.
  • Sequencing time. Moving from one step to the next costs tens or hundreds of milliseconds. Invisible in isolation. Visible when repeated across ten steps.

Each layer is fixed differently, and that is the whole issue. According to the same article, changing models does nothing for a workflow that spends four seconds in three successive API calls, and parallelizing those calls fixes nothing if the problem comes from the infrastructure running the workflow. Optimizing the wrong layer costs time and produces no visible result.

Even before latency, there is a more common cause: the agent does not know where to find the information. We explain what to organize before changing models.

Where does the time go in practice?

The clearest example in the n8n article fits in one line: three calls of 800 milliseconds each, executed one after another, take 2.4 seconds; executed in parallel, about 800 milliseconds, excluding sequencing time. Same work, same model, same API. Only the order changes.

This is not a figure measured by us or one of our customers: it is an order of magnitude provided by the publisher to illustrate the mechanism. It is worth what it is worth, but the mechanism applies everywhere. As soon as two pieces of information do not depend on each other (recognizing the customer and looking up a catalog reference, for example), there is no reason to wait for the first before launching the second.

What causes delayCommon reactionWhat works better
Three searches launched one after anotherChoose a faster modelLaunch them together when they are independent
A supplier API that takes eight secondsWait, or retry everythingSet a maximum time and define what happens when it is exceeded
Long answers that no one reads in fullLet the model writeSet a maximum length and a short format
Simple sorting or classification assigned to the large modelUse one model for everythingSend simple steps to a smaller model
Forty executions arriving at the same timeBlame model slownessCap the number of simultaneous executions and queue them

The final row deserves attention: when everything arrives at once, a throughput problem disguises itself as a slowness problem. The system is not slow, it is saturated. The two are not fixed in the same way.

How long is an agent allowed to take?

It depends entirely on what it does, and this is the first question we ask. n8n publishes reference points by workflow type: under 500 milliseconds for real-time processing, 5 to 20 seconds for batch processing, 30 seconds or more for background processing. For an interaction where the user sees the answer being written, the reference point is the delay before the first displayed word: between 300 and 500 milliseconds, or less.

Now look at the tasks in a small business. A price request received by email, a quote follow-up, inbox sorting, a prospect record to prepare: no one watches the screen during that time. The message arrives, the proposal is prepared in the background, and the only thing that matters is that it is ready when the person opens the inbox again. Nine seconds on those tasks is not a problem to solve. An agent responding in a chat while a customer waits, or carrying on a phone conversation, is in another category.

That is why we start with scoping rather than a technical choice: until we have said who waits and for how long, we do not know whether speed is an issue.

Is your agent taking too long, or do you not yet have one and want to know what it would look like in your company?

Book a free scoping call

What does this change for you?

Three things, from the perspective of someone who runs a company rather than a technical team.

The question to ask your provider changes. It is not “which model do you use?” but “where does time go in my workflow, step by step?” If no one can show you that breakdown, no one knows where to optimize, and the answer that follows will be an intuition.

A slower answer is not necessarily bad news. An agent that looks up a reference with three suppliers before answering takes longer than an agent that answers from memory, and serves you better. In our systems, the agent that prepares inquiries and quotes still waits for a person to review and approve before anything is sent: the second saved during preparation is invisible beside that approval time, and that approval is not negotiable.

The maximum delay is a management decision, not a technical setting. When a supplier API does not respond after ten seconds, someone must have decided in advance what happens: abandon that line and flag it, fall back to the internal catalog, or place the inquiry in a human queue. This choice affects your customer relationship. It belongs to you, and it is written before launch.

Maximum delay is not the only management decision to write before launch: access, approvals, the log, and the stop procedure are part of it.

What are the limits of all this?

Several, and it is better to know them before starting an optimization project.

  • Parallel execution applies only to independent tasks. If step 2 needs the result of step 1, there is nothing to gain. Much of a real workflow is sequential by nature.
  • Retries are expensive in time. n8n gives the calculation: a call that normally takes one second, retried three times with a three-second delay at each attempt, consumes a much larger share of the total time than a single call. Retrying improves reliability, not speed: you must choose, and set a cap.
  • Round trips also have a cost. They cost time and money: what an AI agent really costs depends in particular on the number of turns required to finish the task.
  • A slow provider remains slow. If its API takes eight seconds, no orchestration trick will make it faster. What can be done is to isolate that step so it does not block the rest, give it its own rules, and decide whether the main workflow must wait for it.
  • Shortening answers has a price. The n8n article notes, based on OpenAI recommendations, that cutting the generated text in half reduces latency by roughly the same amount because writing is sequential. This is the most direct setting. It is also the one that can remove useful information: a commercial proposal without its sources or confidence level is faster, and less usable.
  • Sometimes the wait protects you. An input check (n8n's guardrail node, for example, detects secret keys or personal data and replaces them) adds a few milliseconds and prevents sensitive data from going to the wrong place, or a malformed request from sending the agent into a dead end lasting several seconds.
  • Finally, slowness is not always the real problem. n8n says it in one sentence we are happy to repeat: first check that latency is actually the problem, rather than the accuracy of the answers. You can spend weeks saving two seconds on an agent whose real defect was being wrong.

Where should you start in your company?

In this order, without buying anything for the first two steps.

  1. Name the task and say who waits. An email processed in the background, a chat where the customer waits, a phone call: three different requirements. This sets the acceptable time budget.
  2. Look at the time for each step on real cases. Not a demonstration case: the inquiries you actually receive, including ambiguous ones.
  3. Fix the layer that weighs the most. Parallel calls if time is spent in tools, a shorter answer and smaller model for simple steps if time is spent in the model, a queue if time is spent in saturation.
  4. Write what happens when the limit is exceeded. Maximum delay, number of retries, fallback behavior, and what is escalated to a human. This is the part most often neglected and the one that costs the most.
  5. Measure again. Otherwise you will not know whether you saved time or merely moved the problem.

If you are starting from zero, the same logic applies to design: this is what we examine when scoping an automation or a prospecting agent, before writing a single line.

FAQ

Does changing models make my agent faster?

Only if time is being lost in the model. If your workflow spends most of its time querying tools or sequencing steps, a faster model will not be visible. Assigning simple steps (classification, short extraction) to a smaller model, however, is a real improvement because those models write faster.

What response time should an AI agent target?

According to n8n's published reference points: under 500 milliseconds for real-time processing, 5 to 20 seconds for batch processing, 30 seconds or more for background processing. For a response streamed on screen, the target is 300 to 500 milliseconds before the first word. Most small-business tasks (emails, quotes, follow-ups) belong to the last two categories.

My agent takes 9 seconds: is that serious?

It depends on who is waiting. For a request received by email and processed in the background, no. For a live conversation with a customer, yes. Ask the question before starting work: the answer determines whether there is a problem.

Can an agent be made fast without losing reliability?

In part. Parallelizing independent calls and avoiding unnecessary steps costs nothing in reliability. Reducing retries, shortening answers, or removing checks does: these are tradeoffs, to be made with an understanding of what is exchanged. When an action commits the company, we choose verification.

Do you need a particular tool to measure all this?

Above all, you need a system that keeps a step-by-step record of every execution, with its duration. Automation platforms do this natively, including n8n. If your system does not, that is the first point to fix: without this record, every optimization is an assumption.

We review one task, your tools, and what must remain decided by a human. The written conclusion is a simple automation, an AI agent, or nothing useful to build.

Book a free scoping call

Source of the cited figures and patterns: “Reducing AI Workflow Latency: Patterns That Actually Work,” Yulia Dmitrievna and the n8n team, September 19, 2026, blog.n8n.io. The timing references and the three-call example are the publisher's orders of magnitude; they do not come from measurements made by Équipage IA.

Read next

Understanding AI agents7 min read

AI knowledge base: organize what your agent needs to know before changing models

When an AI agent gives poor answers in your company, the first impulse is to change models. That is almost always the wrong place to look. The model knows how to reason; what it does not know is which quote is current, which of your two price lists applies to this customer, and that “Dupont SARL” in email, “DUPONT S.A.R.L.” in management software, and “Ets Dupont” in the pricing spreadsheet refer to the same company. That information exists in your company, but it is written nowhere.

Understanding AI agents7 min read

AI agent cost: pay for the task, not the token

On September 22, 2026, OpenAI released two models at half the price of the previous generation, and Anthropic reduced its prices by one fifth on the same day. If you were waiting for AI to become affordable enough to try in your company, it has. But the token rate will not tell you what your project will cost: the bill is determined by the work you assign to the agent and how many times it rereads its context to complete it.

Understanding AI agents6 min read

Four questions to ask before connecting an AI agent to your tools

On September 18, 2026, security researchers described how they got into OpenAI. It started with the company's community forum: a flaw in an image-processing library, then a weakness in single sign-on, and they found themselves in an employee account connected to the internal development environment. Remember the mechanism rather than the feat: one compromised account carried everything connected to it.

Free discovery audit

Start with one task.

Prospecting that stalls. A request answered too late. A quote that never gets a follow-up. Choose one task and use a 30-minute conversation to decide whether it deserves a solution.

Free and with no commitment. You leave with a written opinion.