feat(realtime): speaker-aware conversations - surface identity to client and LLM - #10424
Merged
Conversation
Add Enforce *bool and Identity *VoiceIdentityConfig to PipelineVoiceRecognition, plus EnforceGate/IdentityEnabled/ AnnounceEnabled/PersonalizeEnabled helpers. Enforce nil defaults to gating (backward compatible); identity surfacing is independent of the gate. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Split the speaker authorization into a Resolve step (embed once, produce a types.Speaker identity) and a pure authorize policy step, with a 0..100 confidence score mirroring /v1/voice/identify. The legacy Authorize wrapper is kept so existing specs stay green. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…peaker Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Set the per-message name field on each recognized user turn and append a current-speaker note to the system message, both gated by the voice recognition identity config. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Document the new voice_recognition keys (enforce, identity.*) and the LocalAI-extension conversation.item.speaker server event in the realtime feature docs. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
…aker history Add two integration specs to harden the speaker-aware realtime path: - when:first with an Identity block re-resolves the speaker every turn even though re-authorization is skipped after the first match: a later resolve error now fails closed, while a clean later resolve still surfaces and names the speaker. - multi-speaker history attribution: each user turn carries its own per-message name and the injected system note reflects the latest speaker. Test-only change; no production behavior was modified. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Carry the registered speaker's labels (identify mode) on types.Speaker so they flow into the conversation.item.speaker event and the stored item. Verify mode has no labels, so the field is omitted there. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add a realtime-pipeline-identity config (verify mode, enforce:false, identity announce+announce_unknown+personalize) and two e2e specs driving the real server over a real WebSocket with the mock VoiceEmbed backend: an authorized speaker yields a conversation.item.speaker event naming e2e-speaker (matched true) and reaches response.done; an unauthorized speaker yields an unknown (matched false, no name) event and still responds, proving enforce:false never drops a turn. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The meta registry coverage test (TestAllFieldsHaveRegistryEntries) requires every config field to have an entry in core/config/meta/registry.go. The new voice_recognition.enforce and voice_recognition.identity.* fields were missing, failing tests-linux and tests-apple. Add registry entries (toggles) so the fields are surfaced in the model-config editor and the coverage test passes. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Speaker-aware realtime conversations
Extends the realtime voice-recognition gate so the recognized speaker is surfaced to the client and fed to the LLM for personalized replies (e.g. "Hey Jeremy, ..."), decoupled from the authorization gate and fully backward compatible.
What it does
conversation.item.speaker(a LocalAI extension; OpenAI clients ignore unknown event types, so it is non-breaking):{ "type": "conversation.item.speaker", "item_id": "...", "speaker": { "name": "Jeremy", "id": "spk_1", "confidence": 92.0, "distance": 0.1, "matched": true } }namefield on each user turn and/or appends aThe current speaker is <Name>.note to the system message. Each user item carries its own speaker, so multi-speaker histories stay correctly attributed.pipeline.voice_recognition:enforce(defaulttrue): the authorization gate.falseresolves/surfaces the speaker without ever dropping a turn.identity.{announce, announce_unknown, personalize, inject_name, inject_system_note, note_unknown}: surfacing + personalization, independent ofenforce. When set, identity is resolved on every turn even underwhen: first.Design
voiceGateis split intoResolve(embed once →types.Speakeridentity) and a pureauthorizepolicy step; the legacyAuthorizewrapper is preserved so existing behavior and tests are unchanged.enforce: falsenever drops a turn.Backward compatibility
A
voice_recognitionconfig with neitherenforcenoridentitybehaves exactly as before. The new event and item field are additive.Tests
TDD throughout (Ginkgo/Gomega). New unit + integration specs cover: the
Resolve/authorizesplit and confidence math, event marshaling, per-turn surfacing,enforce:falsenever dropping,when:first×identityper-turn re-resolution, and multi-speaker history attribution.core/configandcore/http/endpoints/openaiboth green.Docs
docs/content/features/openai-realtime.mdextended with the new config keys and theconversation.item.speakerevent.Assisted-by: Claude:claude-opus-4-8
🤖 Generated with Claude Code