Skip to content

docs(bench): OpenClaw + Dream 500d Executive Assistant results - #4

Open
spqian wants to merge 1 commit into
Stevenic:mainfrom
spqian:dream-500d-results
Open

docs(bench): OpenClaw + Dream 500d Executive Assistant results#4
spqian wants to merge 1 commit into
Stevenic:mainfrom
spqian:dream-500d-results

Conversation

@spqian

@spqian spqian commented Jul 9, 2026

Copy link
Copy Markdown

OpenClaw + Dream — 500-day Executive Assistant results

Per the /dream skill discussion — @Stevenic offered to publish this run in the recall bench. This PR adds the writeup + evidence following the existing docs/bench/results/ convention (mirrors openclaw-ea.md).

What this is: the same openclaw[vector:text-embedding-3-small+agent] answer loop and judges, but with the memory backend pointed at the Dreamweave nightly consolidation engine (report→judge→apply merges, supersede/sequence edges, graph recall). It's the natural A/B against the published plain-vector 500d run.

Headline (vs plain-vector 500d baseline)

metric Dream Vector baseline
Overall composite (Q1→Q4) 5.60 → 5.15 5.24 → 4.98
Hallucination (Q1→Q4) 4.5% → 10.2% 13.6% → 18.2%
Hallucination peak 16.2% 27.1%
temporal-reasoning 4.60 3.89
recency-bias-resistance 5.32 (only rising cat) 4.74

Consolidation roughly halves the hallucination rate and lifts every category; temporal reasoning stays the weakest and is called out as the frontier.

Contents

  • docs/bench/results/openclaw-dream-ea.md — full writeup (TL;DR, per-category quartiles, hallucination trajectory, findings, reproduce)
  • docs/bench/results/openclaw-dream-ea-500d-heatmap.png
  • docs/bench/results/index.md — new row
  • bench-results/openclaw/ea-500d-dream/ — raw result.json, progress.jsonl, failures.jsonl, heatmap.png (per-question log omitted for size, ~23 MB)

Includes an honest fairness note (engine-level comparison, not a single isolated knob) and an operational note on the tier2=50000 ingest-cost artifact.

Add a Recall Bench results writeup for the OpenClaw agent loop backed by
the Dreamweave nightly consolidation engine, on the Executive Assistant
persona over a 500-day corpus (50 checkpoints, 3,164 evals).

Headline vs the published plain-vector 500d run: overall composite 5.32
(last-quartile 5.15 vs 4.98), hallucination roughly halved (7.7% avg,
16.2% peak vs 13.6% floor / 27.1% peak), and every category improved.
Temporal reasoning remains the weakest category (4.60).

Includes the heatmap, per-category quartile tables, hallucination
trajectory, an operational note on tier2=50000 ingest cost, and the raw
run artifacts under bench-results/openclaw/ea-500d-dream/.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant