feat: add support for audio-text-to-text with m serve CLI and include examples#1443
Open
markstur wants to merge 6 commits into
Open
feat: add support for audio-text-to-text with m serve CLI and include examples#1443markstur wants to merge 6 commits into
markstur wants to merge 6 commits into
Conversation
Extend the serve-layer ChatMessage model to accept OpenAI's
`input_audio` content-part schema and expose audio blocks to
serve functions.
- Add InputAudioData and InputAudioContent Pydantic models matching
the OpenAI `{"type": "input_audio", "input_audio": {"data": ...,
"format": "wav|mp3"}}` wire format
- Extend MessageContent union to include InputAudioContent
- Add ChatMessage.get_audio_blocks() returning list[AudioBlock],
mirroring the existing get_image_blocks() pattern
- Re-export new types from mellea.serve and cli.serve.models
to be consistent with current practice
- Add unit tests covering: string/None content → [], valid WAV →
AudioBlock, invalid base64 → ValueError, text extraction ignores
audio parts, mixed text+audio deserialisation
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
* Client uses OpenAI API * llama-server with gemma is simple audio-text-to-text support * ollama with granite shows same results using dual models Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
…ite.py Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
…, and grounding requirement Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com> Assisted-by: IBM Bob
…xample * sync with granite/ollama example improvements except this version is one-step with no separate transcribe to process Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Issue
Fixes #1395
Description
audio-text-to-text examples using m serve.
This shows that our OpenAI API support can handle multimodal.
The backends are limited by Mellea and by the servers and models.
For now we show 2 examples:
1 - use llama-server/gemma to demonstrate a clean openai API audio-text-to-text implementation
2 - use ollama with 2 step granite (speech for transcribe, then LLM to query w/ transcription).
Also trying to show more Mellea ability in 2 (e.g. a requirement).
Testing
Attribution
Adding a new component, requirement, sampling strategy, or tool?
If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.
NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.