Skip to content

fix: preserve adaptive interruption across tool calls - #6564

Open
chenghao-mou wants to merge 1 commit into
mainfrom
chenghao/fix/AGT-3180-adaptive-interruption-tool-calls
Open

fix: preserve adaptive interruption across tool calls#6564
chenghao-mou wants to merge 1 commit into
mainfrom
chenghao/fix/AGT-3180-adaptive-interruption-tool-calls

Conversation

@chenghao-mou

@chenghao-mou chenghao-mou commented Jul 27, 2026

Copy link
Copy Markdown
Member

Bug: Adaptive interruption could silently drop a user overlap when agent playout resumed after a tool call because the interruption stream received mismatched speech boundaries.

Treat thinking gaps as real speech boundaries and restart overlap inference when playout resumes mid-utterance. Agent-ended overlaps remain inconclusive and no longer count as backchannels; remote forwarding is noted until the protocol carries that marker.

Fixes #6548. Fixes #6548.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Open in Devin Review

@biztex

biztex commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

This looks like the right fix — treating the thinking gap as a real speech boundary and restarting overlap inference when playout resumes addresses the root cause at the signal source, where my #6550 only made the stream tolerate the redundant start sentinel. The metrics and remote-session coverage go beyond what mine attempted, too.

Closing #6550 in favor of this. If the hermetic stream-level regression tests from it would be useful here (a fake bargein gateway driving InterruptionWebSocketStream through a mid-overlap second segment), happy to adapt them to the new sentinel semantics — just say the word.


def _on_pipeline_reply_done(self, _: asyncio.Task[None]) -> None:
if not self._speech_q and (not self._current_speech or self._current_speech.done()):
was_speaking = self._session.agent_state == "speaking"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question: what is the difference between session.agent_state and self._agent_speaking?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(Chiming in from #6550, since this distinction was the heart of #6548 — hope it helps.)

They live at different layers and deliberately diverge mid-turn:

  • session.agent_state is the session-level lifecycle (listening / thinking / speaking), updated via AgentActivity._update_agent_state(). A reply that produces tool calls flips it to "thinking" between segments; it is what applications observe through agent_state_changed.

  • AudioRecognition._agent_speaking is a private recognition-side flag meaning "an agent playout window is open for interruption/endpointing purposes". It is set by _on_start_of_agent_speech() (first audio frame of playout) and cleared only inside _on_end_of_agent_speech() — which is intentionally not called when the reply produced tool calls, so it stays True straight through the thinking gap. It gates overlap tracking in _on_start_of_speech(), the transcript-ignore window, and _should_ignore_input_audio().

So during a tool-call gap the two disagree by design: agent_state == "thinking" while _agent_speaking == True. was_speaking as written captures "playout was audibly active at reply-done" (session view) rather than "the recognition-side speech window is still open" — which one is wanted here is of course the author's call.

@chenghao-mou
chenghao-mou requested a review from longcw July 31, 2026 11:54
@longcw

longcw commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

from claude code:

#6662 section (b) is the same bug with a different trigger: a second agent-speech boundary inside one turn resets the detector stream while an overlap is open. It corrected the false-interruption pause. It did not touch the tool branch:

if not speech_handle.interrupted and len(tool_output.output) > 0:
self._session._update_agent_state("thinking")
elif self._session.agent_state == "speaking":
self._session._update_agent_state("listening")
if self._audio_recognition:
self._audio_recognition._on_end_of_agent_speech(
ignore_user_transcript_until=time.time()
)
if self.interruption_enabled:
self._restore_interruption_by_audio_activity()

The tool-response reply is a new _pipeline_reply_task on the same speech_handle, so its playout start reaches the same call site as a first segment, where resumed defaults to False. Second start sentinel, _reset_state(), verdict gone. _restore_interruption_by_audio_activity() is in the skipped elif, so there is no VAD fallback either.

Two edits instead:

  1. Call _on_end_of_agent_speech(..., paused=True) and _restore_interruption_by_audio_activity() in the tool branch. A pause keeps _overlap_open and sends no reset, so user audio keeps reaching the gateway while the tools run.
  2. Add resumed = resumed or self._overlap_open in _on_start_of_agent_speech. _overlap_open is set only while an overlap is unresolved, so a new turn still resets. This covers the queued say() case too.

Then the synthesized _on_start_of_speech, the _agent_speaking guard, both was_speaking gates and the self._speaking and condition can all go. Three of them have side effects: the synthesized onset overwrites _utterance_started_at with the resume time, the guard lets _turn_tracker reset mid-utterance, and speech_duration=0.0 with _overlap_count = 0 shifts the buffer down to _prefix_size.

Keep the sentinel reorder. It fixes a separate bug, and it is real: the reset currently goes out before the agent_ended=True sentinel, so the stream drops that event and audio_recognition.py:502 never emits. The num_backchannels change needs the reorder to reach the metrics at all.

@chenghao-mou
chenghao-mou force-pushed the chenghao/fix/AGT-3180-adaptive-interruption-tool-calls branch from c124787 to 4ad72da Compare August 4, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Adaptive interruption: a second agent speech segment (e.g. after a tool call) disarms an open overlap and silently swallows the interruption

3 participants