Skip to content

bug: generate_from_raw hanging when ollama under load #650

Description

@psschwei

In the process of testing #567 I encountered an issue where test/core/test_component_typing.py::test_generating was timing out after 15 minutes and failing. Seems like I had overloaded ollama by running for i in {1..10}; do uv run pytest test/core -v; done and the test failed on the fourth or fifth run.

Doing a little bit of digging, it seems like the issue happened here:

responses = await asyncio.gather(*coroutines, return_exceptions=True)

Because there's no timeout, if Ollama stalls on any of the concurrent requests (like it seems happened when I ran this test over and over and over again) the entire call hangs indefinitely.

There's a couple of different ways we could fix this, but probably the easiest would be to set a timeout on the underlying ollama.AsyncClient:

_async_client = ollama.AsyncClient(self._base_url)

Ideally we'd make the timeout configurable too (so users could set it when creating the backend?)

Metadata

Metadata

Assignees

Labels

area/backendsProvider-specific work: Ollama, HF, LiteLLM, OpenAI, Bedrock, vLLMbugSomething isn't workinggood first issueGood for newcomershelp wantedExtra attention is needed

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions