From 1f2146e151a1dd62f3095e9b73de1e1025ed7daa Mon Sep 17 00:00:00 2001 From: Vinod Muthusamy Date: Mon, 6 Apr 2026 12:27:48 -0500 Subject: [PATCH] docs: restructure index page and add results page Reorganize the landing page to lead with value proposition and streamlined getting started links. Add results page with AppWorld benchmark data and paper reference. --- docs/index.md | 28 +++++++++++----------------- docs/results/index.md | 26 ++++++++++++++++++++++++++ mkdocs.yaml | 2 ++ 3 files changed, 39 insertions(+), 17 deletions(-) create mode 100644 docs/results/index.md diff --git a/docs/index.md b/docs/index.md index 69b0eff4..f74688e2 100644 --- a/docs/index.md +++ b/docs/index.md @@ -4,36 +4,30 @@ template: main.html # Evolve ***_Self-improving agents through iterations._*** -Evolve is a system designed to help agents improve over time by learning from their trajectories. It uses a combination of an MCP server for tool integration, vector storage for memory, and LLM-based conflict resolution to refine its knowledge base. +Coding agents repeat the same mistakes because they start fresh every session. Evolve gives agents memory — they learn from what worked and what didn't, so each session is better than the last. -## Features +On the AppWorld benchmark, Evolve improved agent reliability by **+8.9 points** overall, with a **74% relative increase** on hard multi-step tasks. See the [full results](results/index.md) and the [paper (arXiv:2603.10600)](https://arxiv.org/abs/2603.10600). -- **MCP Server**: Exposes tools to get guidelines and save trajectories. -- **Conflict Resolution**: Intelligently merges new insights with existing guidelines using LLMs. -- **Trajectory Analysis**: Automatically analyzes agent trajectories to generate guidelines and best practices. -- **Milvus Integration**: Uses Milvus (or Milvus Lite) for efficient vector storage and retrieval. - -## Start Here +## Get Started - [Installation](installation/index.md): Set up Evolve on Bob or Claude Code. -- [Configuration](guides/configuration.md): Configure models, backends, and environment variables. -- [CLI Reference](reference/cli.md): Manage namespaces, entities, and sync jobs from the command line. +- [Hello World with IBM Bob](examples/hello_world/bob.md): A simple walkthrough that shows how memory gets learned. +- [Hello World with Claude Code](examples/hello_world/claude.md): Get started with Evolve Lite in Claude Code. ## Guides +- [Configuration](guides/configuration.md): Configure models, backends, and environment variables. - [Low-Code Tracing](guides/low-code-tracing.md): Instrument agents with Phoenix and verify end-to-end tracing. - [Phoenix Sync](guides/phoenix-sync.md): Pull trajectories from Phoenix and generate stored guidelines. - [Extract Trajectories](guides/extract-trajectories.md): Export Phoenix traces into an OpenAI-style message format. -## Integrations and Tutorials - -- [Evolve Lite (Claude Code)](integrations/claude/evolve-lite.md): Lightweight Claude Code integration with local entity storage. -- [Claude Code Demo](tutorials/claude-code-demo.md): Run the filesystem demo with Claude Code and the Evolve MCP server. -- [Hello World with IBM Bob](examples/hello_world/bob.md): A simple Bob walkthrough that shows how memory gets learned. - ## Reference +- [CLI Reference](reference/cli.md): Manage namespaces, entities, and sync jobs from the command line. - [Policies](reference/policies.md): Structured policy entities and how to retrieve them with MCP tools. -## Architecture +## How It Works + +Evolve analyzes agent trajectories to extract guidelines and best practices, then recalls them in future sessions. It supports both a lightweight file-based mode (Evolve Lite) and a full mode backed by an MCP server with vector storage and LLM-based conflict resolution. + ![Architecture](assets/architecture.png) diff --git a/docs/results/index.md b/docs/results/index.md new file mode 100644 index 00000000..178aca7d --- /dev/null +++ b/docs/results/index.md @@ -0,0 +1,26 @@ +# Results + +## AppWorld Benchmark + +We evaluated Evolve on [AppWorld](https://appworld.dev), where agents complete realistic multi-step tasks via APIs, averaging 9.5 APIs across 1.8 apps. Hard tasks require more complex control flow across multiple services. + +A ReAct agent received the task instruction plus the top 5 retrieved guidelines generated from one prior run on train/dev and was tested on an unseen partition (test-normal). We report Scenario Goal Completion (SGC), a strict consistency metric requiring success across scenario variants. + +| Difficulty | Baseline SGC | + Evolve | Gain | +|---|---:|---:|---:| +| Easy | 79.0% | 84.2% | +5.2 | +| Medium | 56.2% | 62.5% | +6.3 | +| **Hard** | **19.1%** | **33.3%** | **+14.2** | +| **Aggregate** | **50.0%** | **58.9%** | **+8.9** | + +### Key findings + +- **Generalization:** The agent improves on unseen test tasks, showing it learns transferable principles rather than memorizing solutions. +- **Complexity scaling:** The harder the task, the more the agent benefits from learned guidelines. Hard tasks saw a 74% relative increase in success rate. +- **Consistency:** SGC gains exceeded raw pass-rate improvements, reducing "flaky" behavior across scenario variants. Guidelines help the agent solve tasks reliably, not just occasionally. + +## Paper + +For full details on the architecture, experiments, and analysis, see: + +> [Trajectory-Informed Memory Generation for Self-Improving Agent Systems](https://arxiv.org/abs/2603.10600) (arXiv:2603.10600) diff --git a/mkdocs.yaml b/mkdocs.yaml index d673503d..a3b8582c 100644 --- a/mkdocs.yaml +++ b/mkdocs.yaml @@ -88,6 +88,8 @@ nav: - Hello World (IBM Bob IDE): examples/hello_world/bob.md - Claude Code Demo: tutorials/claude-code-demo.md - Starter Example 2 (Claude Code): examples/hello_world/claude.md + - Results: + - Benchmarks: results/index.md - Operations: - Docker Testing: operations/docker-testing.md