Skip to content

feat: add support for audio-text-to-text with m serve CLI and include examples#1443

Open
markstur wants to merge 6 commits into
generative-computing:mainfrom
markstur:issue_1395
Open

feat: add support for audio-text-to-text with m serve CLI and include examples#1443
markstur wants to merge 6 commits into
generative-computing:mainfrom
markstur:issue_1395

Conversation

@markstur

@markstur markstur commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

Pull Request

Issue

Fixes #1395

Description

audio-text-to-text examples using m serve.
This shows that our OpenAI API support can handle multimodal.
The backends are limited by Mellea and by the servers and models.
For now we show 2 examples:
1 - use llama-server/gemma to demonstrate a clean openai API audio-text-to-text implementation
2 - use ollama with 2 step granite (speech for transcribe, then LLM to query w/ transcription).
Also trying to show more Mellea ability in 2 (e.g. a requirement).

Testing

  • Tests added to the respective file if code was changed
  • New code has 100% coverage if code was added
  • Ensure existing tests and github automation passes (a maintainer will kick off the github automation when the rest of the PR is populated)

Attribution

  • AI coding assistants used

Adding a new component, requirement, sampling strategy, or tool?

If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.

  • Component
  • Requirement
  • Sampling Strategy
  • Tool

NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.

markstur added 6 commits July 23, 2026 13:00
Extend the serve-layer ChatMessage model to accept OpenAI's
`input_audio` content-part schema and expose audio blocks to
serve functions.

- Add InputAudioData and InputAudioContent Pydantic models matching
  the OpenAI `{"type": "input_audio", "input_audio": {"data": ...,
  "format": "wav|mp3"}}` wire format
- Extend MessageContent union to include InputAudioContent
- Add ChatMessage.get_audio_blocks() returning list[AudioBlock],
  mirroring the existing get_image_blocks() pattern
- Re-export new types from mellea.serve and cli.serve.models
  to be consistent with current practice
- Add unit tests covering: string/None content → [], valid WAV →
  AudioBlock, invalid base64 → ValueError, text extraction ignores
  audio parts, mixed text+audio deserialisation

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
* Client uses OpenAI API
* llama-server with gemma is simple audio-text-to-text support
* ollama with granite shows same results using dual models

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
…ite.py

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
…, and grounding requirement

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
…xample

* sync with granite/ollama example improvements except this version is one-step
  with no separate transcribe to process

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
@markstur
markstur requested a review from a team as a code owner July 24, 2026 23:56
@markstur markstur changed the title Issue 1395 feat: add support for audio-text-to-text with m serve CLI and include examples Jul 24, 2026
@github-actions github-actions Bot added the enhancement New feature or request label Jul 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: audio-text-to-text support in CLI with example

1 participant