Rebuttal Latency in Real-Time Voice Agents
Why debate voice agents feel slow even when components are fast.

Rebuttal latency in a live debate voice agent is the accumulated cost of three separate pipeline stages, speech-to-text, language model reasoning, and text-to-speech, each with its own minimum processing time, and in a debate exchange that sum routinely exceeds the window in which a reply still feels like a rebuttal. A rebuttal only works as a rebuttal if it lands while the adversarial pressure of the prior argument is still active. Once the pipeline tax pushes the reply past that window, the exchange stops resembling a debate and starts resembling two monologues separated by a pause. Because the three stages run in series, their latencies add rather than average, so the realistic end-to-end figure for most deployed agents runs up to 2 seconds, even though individual components can be made fast in isolation. Deepgram's Nova-3 reaches sub-300ms time-to-first-token, and ElevenLabs' Flash v2.5 model reaches approximately 75ms of model inference latency for short inputs, but neither figure describes what happens once the two are chained together with a reasoning step in between and a live human waiting on the other end. A debate agent carries one more burden before any of this even starts: it has to decide, from an audio stream alone, whether the opponent has actually finished the argument or just paused mid-thought, and that decision is already spending part of the latency budget before a single token of rebuttal gets generated.
Each pipeline stage's latency floor in a debate context
Each stage in the pipeline has a floor set by what the underlying models and physics allow, not by sloppy engineering, and a debate use case tends to push every stage toward the slow end of its range. Speech-to-text adds a meaningful amount of latency even in its streaming best case, and that figure can stretch considerably further depending on the model and on real-world audio conditions, which matters because debate speech is fast, dense with subordinate clauses, and full of the kind of rhetorical hedging that trips up automatic speech recognition trained on cleaner, more conversational audio. Streaming ASR, which processes audio in chunks as it arrives rather than waiting for a complete utterance, can shave meaningful time off the transcription step, but feeding partial transcripts downstream before the opponent has finished speaking invites reasoning errors that cost more time to fix than the streaming approach saved.
The language model is where the largest single chunk of latency gets spent, and debate reasoning is a harder task than the query answering that most voice agents are built around. A well-optimized stack gives the model something like 200 to 700 milliseconds between the end of transcription and the start of synthesis, and a model that scores well on chat benchmarks but has a slow time-to-first-token is not usable inside that window no matter how good its arguments are. Text-to-speech introduces its own tension between naturalness and speed: a rebuttal that arrives quickly but sounds flat or robotic undercuts the illusion of adversarial exchange that the latency budget was built to protect, because a debate opponent who sounds like a kiosk reading a receipt does not feel like an opponent. Of the three stages, only the language model puts capability and speed in direct opposition to each other. Faster models tend to reason less carefully, and in a debate context a fast but wrong rebuttal does more damage to the exchange than a slower but accurate one, because it misrepresents the opponent's point in front of the very audience the debate is meant to persuade.
Turn Detection and Debate-Specific Latency
Before any of the three pipeline stages can even begin, the system has to solve a problem that acoustic voice-activity detection alone cannot cleanly solve: telling a finished argument apart from a thinking pause. In debate, getting this wrong in either direction carries a real cost. A single silence threshold is being asked to do two incompatible jobs at once. Setting the threshold short causes the agent to cut off the opponent mid-sentence, rebutting an argument that was not yet complete; this reads as both bad reasoning and plain rudeness inside a format built on taking the other side seriously. Setting the threshold long lets dead air stack up after every completed turn, breaking the rhythm of the exchange and signaling that the agent is stalling rather than thinking, and moving the threshold from 400 milliseconds to 800 milliseconds cuts down on premature interruptions while adding that same 400 milliseconds of dead air to every single completed turn's response time, which is a trade a debate agent cannot afford to make carelessly given how much of the format's credibility rests on pace.
Semantic turn detection offers a structural way out of this bind, combining acoustic signal with an understanding of whether the sentence itself is grammatically complete. By pairing acoustic voice-activity detection with a lightweight classifier running on the partial transcript, the system can treat an incomplete phrase as a signal to keep waiting and a grammatically complete phrase as a cue to shorten the silence window before responding. This separates two questions that pure acoustic detection collapses into one: whether the sentence is actually finished, and whether the speaker has simply stopped making noise. Barge-in handling presents the sharpest debate-specific version of this problem, since debaters interrupt each other as a matter of strategy, and the system has to distinguish a genuine counter-argument cutting in from a backchannel acknowledgment like "right" or "exactly" that should not stop the agent from finishing its point. A minimum-duration guard on barge-in detection adds a small fixed cost of around 200 milliseconds but meaningfully cuts down on false positives, which is a worthwhile trade given how disruptive it is for an agent to yield the floor to a filler word. Retell AI points to a proprietary turn-taking model as a distinguishing part of its platform, citing evaluation results in which the system handled interruptions and barge-in without breaking the flow of conversation. The implication carries past this one example: an agent can have fast transcription and a fast language model and still feel fundamentally broken as a debate partner if its sense of when to speak is naive, because turn detection is where conversational intelligence either gets built in or gets left out.
Context Accumulation and Late-Round Latency
Latency in a live debate is not a flat tax applied evenly to every turn. It grows as the round progresses, and it grows for a specific structural reason: most language models process the entire accumulated session context on every single turn, so a rebuttal late in an exchange carries the weight of everything said before it, while an early turn carries almost nothing. This means the pipeline gets slower exactly as the stakes of the exchange rise, since the final rebuttals in a debate round tend to matter more to the outcome than the opening exchanges, and that is precisely when the system is carrying its heaviest prompt. A customer service call typically resolves in a handful of turns and rarely accumulates enough context to strain this mechanism. A debate round is a long-horizon exchange where the context keeps compounding turn over turn rather than staying flat, which makes it a genuinely different engineering problem from the voice agent use cases that most latency benchmarks are built around.
The gpt-realtime-2 line addresses part of this with a 128,000-token context window and an adjustable reasoning-effort control, giving a debate agent more room before older turns need to be dropped or summarized. A larger context window changes where the boundary sits, not how much it costs to process everything inside that boundary, so the latency cost of reasoning over a long transcript does not go away just because the window got bigger. OpenAI's gpt-realtime-2.1, released in July 2026, added improved silence and noise handling and reduced p95 latency substantially across the line, which helps late-round performance without resolving the underlying accumulation problem on its own. The practical implication for anyone building a debate agent is that raw model speed cannot substitute for an actual context management strategy: selective summarization of prior turns, or retrieval limited to the claims most relevant to the current exchange, rather than feeding the model the full, ever-growing transcript on every turn. The round does not just get longer as it goes. It gets structurally heavier, and a system that does not plan for that will feel sharp in round one and sluggish by the final rebuttal, right when the sluggishness is most visible and most costly.
The two architectural paths for debate: cascaded pipeline versus speech-to-speech
Everything established so far, the pipeline tax, the per-stage floors, how turn detection gets handled, and the context accumulation effect, feeds into a single architectural choice that has no clean winner for debate. Choosing between a cascaded pipeline and a speech-to-speech model is a question of which failure mode is easier to live with: a system that is slower but fully debuggable, or one that is faster but harder to see inside.
The cascaded pipeline, running speech-to-text into a language model into text-to-speech, remains the dominant approach in production because every stage is observable and can be evaluated and swapped independently. For a debate agent, this means each component can be tuned to the task: a faster speech-to-text model, a language model trained or prompted specifically for debate reasoning, a text-to-speech voice tuned for expressiveness, and when a rebuttal goes wrong, the failure can be traced to a specific stage. Word error rates from the transcription step and reasoning errors from the language model step can be caught and addressed separately.
Speech-to-speech models take a different approach entirely, with a single model taking audio in and producing audio out without an intermediate text representation. This preserves the paralinguistic signal that cascaded pipelines discard, and in production, when the model is well matched to the task, it can achieve lower end-to-end latency than a chained system. The cost is that the reasoning happening inside the model becomes opaque: a transcript of what was said exists, but the model's internal deliberation about how it arrived at a given rebuttal does not, which makes it genuinely difficult to diagnose why a rebuttal misread or misrepresented the opponent's argument. One leading speech-to-speech model represents the frontier of this approach as of July 2026, carrying forward the advanced reasoning of its predecessor into a live audio loop, with the audio path now eligible for sensitive regulated use under its provider's data-handling agreements, though the cost structure at scale runs materially higher than a chained pipeline. Neither architecture resolves that tension on its own.
Latency targets a debate voice agent must hit to stay inside the human conversational window
The human conversational window is a hard ceiling, and a debate agent running above 800 milliseconds of median turn latency is a different kind of adversary, one whose pacing signals disengagement rather than active listening, regardless of how sound its arguments are. The targets that follow describe what closing that gap actually requires.
Sub-300 millisecond end-to-end latency is the threshold enterprise teams are currently pushing toward, and it is the point at which the adversarial illusion of a live debate is most sustainable, since a reply at that speed reads as genuine engagement. Retell AI's production data puts roughly 600 milliseconds as the threshold at which callers stop noticing they are talking to an AI system at all, which offers a reasonable floor for a debate agent: not necessarily invisible, but natural enough that the exchange keeps its rhythm. Above roughly 1,500 milliseconds, conversational experience degrades consistently, with callers interrupting, talking over the system, or disengaging. In a debate context, that threshold marks the point where the reciprocal pressure that makes the exchange worth having breaks down entirely, since a debate depends on each side responding to the other while the argument is still live, and a system that misses that window has stopped debating and started merely replying.
