Why does an AI agent take several seconds to respond?
Because what we call “the agent” is not one program, but a sequence of steps. The n8n team, whose automation platform is used to build the systems we install, divides this wait into three layers (September 19, 2026 article):
- Model time. It reads what it is given, then writes its answer word by word. Reading is fast; writing is sequential, so it is proportional to the length of the answer.
- Tool time. Every time the agent retrieves information elsewhere (your catalog, CRM, document database, or a supplier API), it waits for the response. These calls often take several seconds.
- Sequencing time. Moving from one step to the next costs tens or hundreds of milliseconds. Invisible in isolation. Visible when repeated across ten steps.
Each layer is fixed differently, and that is the whole issue. According to the same article, changing models does nothing for a workflow that spends four seconds in three successive API calls, and parallelizing those calls fixes nothing if the problem comes from the infrastructure running the workflow. Optimizing the wrong layer costs time and produces no visible result.
Even before latency, there is a more common cause: the agent does not know where to find the information. We explain what to organize before changing models.
Where does the time go in practice?
The clearest example in the n8n article fits in one line: three calls of 800 milliseconds each, executed one after another, take 2.4 seconds; executed in parallel, about 800 milliseconds, excluding sequencing time. Same work, same model, same API. Only the order changes.
This is not a figure measured by us or one of our customers: it is an order of magnitude provided by the publisher to illustrate the mechanism. It is worth what it is worth, but the mechanism applies everywhere. As soon as two pieces of information do not depend on each other (recognizing the customer and looking up a catalog reference, for example), there is no reason to wait for the first before launching the second.
| What causes delay | Common reaction | What works better |
|---|---|---|
| Three searches launched one after another | Choose a faster model | Launch them together when they are independent |
| A supplier API that takes eight seconds | Wait, or retry everything | Set a maximum time and define what happens when it is exceeded |
| Long answers that no one reads in full | Let the model write | Set a maximum length and a short format |
| Simple sorting or classification assigned to the large model | Use one model for everything | Send simple steps to a smaller model |
| Forty executions arriving at the same time | Blame model slowness | Cap the number of simultaneous executions and queue them |
The final row deserves attention: when everything arrives at once, a throughput problem disguises itself as a slowness problem. The system is not slow, it is saturated. The two are not fixed in the same way.
How long is an agent allowed to take?
It depends entirely on what it does, and this is the first question we ask. n8n publishes reference points by workflow type: under 500 milliseconds for real-time processing, 5 to 20 seconds for batch processing, 30 seconds or more for background processing. For an interaction where the user sees the answer being written, the reference point is the delay before the first displayed word: between 300 and 500 milliseconds, or less.
Now look at the tasks in a small business. A price request received by email, a quote follow-up, inbox sorting, a prospect record to prepare: no one watches the screen during that time. The message arrives, the proposal is prepared in the background, and the only thing that matters is that it is ready when the person opens the inbox again. Nine seconds on those tasks is not a problem to solve. An agent responding in a chat while a customer waits, or carrying on a phone conversation, is in another category.
That is why we start with scoping rather than a technical choice: until we have said who waits and for how long, we do not know whether speed is an issue.
Is your agent taking too long, or do you not yet have one and want to know what it would look like in your company?
What does this change for you?
Three things, from the perspective of someone who runs a company rather than a technical team.
The question to ask your provider changes. It is not “which model do you use?” but “where does time go in my workflow, step by step?” If no one can show you that breakdown, no one knows where to optimize, and the answer that follows will be an intuition.
A slower answer is not necessarily bad news. An agent that looks up a reference with three suppliers before answering takes longer than an agent that answers from memory, and serves you better. In our systems, the agent that prepares inquiries and quotes still waits for a person to review and approve before anything is sent: the second saved during preparation is invisible beside that approval time, and that approval is not negotiable.
The maximum delay is a management decision, not a technical setting. When a supplier API does not respond after ten seconds, someone must have decided in advance what happens: abandon that line and flag it, fall back to the internal catalog, or place the inquiry in a human queue. This choice affects your customer relationship. It belongs to you, and it is written before launch.
Maximum delay is not the only management decision to write before launch: access, approvals, the log, and the stop procedure are part of it.
What are the limits of all this?
Several, and it is better to know them before starting an optimization project.
- Parallel execution applies only to independent tasks. If step 2 needs the result of step 1, there is nothing to gain. Much of a real workflow is sequential by nature.
- Retries are expensive in time. n8n gives the calculation: a call that normally takes one second, retried three times with a three-second delay at each attempt, consumes a much larger share of the total time than a single call. Retrying improves reliability, not speed: you must choose, and set a cap.
- Round trips also have a cost. They cost time and money: what an AI agent really costs depends in particular on the number of turns required to finish the task.
- A slow provider remains slow. If its API takes eight seconds, no orchestration trick will make it faster. What can be done is to isolate that step so it does not block the rest, give it its own rules, and decide whether the main workflow must wait for it.
- Shortening answers has a price. The n8n article notes, based on OpenAI recommendations, that cutting the generated text in half reduces latency by roughly the same amount because writing is sequential. This is the most direct setting. It is also the one that can remove useful information: a commercial proposal without its sources or confidence level is faster, and less usable.
- Sometimes the wait protects you. An input check (n8n's guardrail node, for example, detects secret keys or personal data and replaces them) adds a few milliseconds and prevents sensitive data from going to the wrong place, or a malformed request from sending the agent into a dead end lasting several seconds.
- Finally, slowness is not always the real problem. n8n says it in one sentence we are happy to repeat: first check that latency is actually the problem, rather than the accuracy of the answers. You can spend weeks saving two seconds on an agent whose real defect was being wrong.
Where should you start in your company?
In this order, without buying anything for the first two steps.
- Name the task and say who waits. An email processed in the background, a chat where the customer waits, a phone call: three different requirements. This sets the acceptable time budget.
- Look at the time for each step on real cases. Not a demonstration case: the inquiries you actually receive, including ambiguous ones.
- Fix the layer that weighs the most. Parallel calls if time is spent in tools, a shorter answer and smaller model for simple steps if time is spent in the model, a queue if time is spent in saturation.
- Write what happens when the limit is exceeded. Maximum delay, number of retries, fallback behavior, and what is escalated to a human. This is the part most often neglected and the one that costs the most.
- Measure again. Otherwise you will not know whether you saved time or merely moved the problem.
If you are starting from zero, the same logic applies to design: this is what we examine when scoping an automation or a prospecting agent, before writing a single line.
FAQ
Does changing models make my agent faster?
Only if time is being lost in the model. If your workflow spends most of its time querying tools or sequencing steps, a faster model will not be visible. Assigning simple steps (classification, short extraction) to a smaller model, however, is a real improvement because those models write faster.
What response time should an AI agent target?
According to n8n's published reference points: under 500 milliseconds for real-time processing, 5 to 20 seconds for batch processing, 30 seconds or more for background processing. For a response streamed on screen, the target is 300 to 500 milliseconds before the first word. Most small-business tasks (emails, quotes, follow-ups) belong to the last two categories.
My agent takes 9 seconds: is that serious?
It depends on who is waiting. For a request received by email and processed in the background, no. For a live conversation with a customer, yes. Ask the question before starting work: the answer determines whether there is a problem.
Can an agent be made fast without losing reliability?
In part. Parallelizing independent calls and avoiding unnecessary steps costs nothing in reliability. Reducing retries, shortening answers, or removing checks does: these are tradeoffs, to be made with an understanding of what is exchanged. When an action commits the company, we choose verification.
Do you need a particular tool to measure all this?
Above all, you need a system that keeps a step-by-step record of every execution, with its duration. Automation platforms do this natively, including n8n. If your system does not, that is the first point to fix: without this record, every optimization is an assumption.
We review one task, your tools, and what must remain decided by a human. The written conclusion is a simple automation, an AI agent, or nothing useful to build.
Source of the cited figures and patterns: “Reducing AI Workflow Latency: Patterns That Actually Work,” Yulia Dmitrievna and the n8n team, September 19, 2026, blog.n8n.io. The timing references and the three-call example are the publisher's orders of magnitude; they do not come from measurements made by Équipage IA.
