Skip to content

TCP-MT-7: Run scale stress test for TCP per-task filtering at 500+ tool corpus #21

Description

@samscarrow

Notion Work Item

TCP-MT-7 (Measurement Track) — TCP Harness Prototype project

Objective

Run a scale stress test for TCP per-task filtering by expanding the current 90-tool corpus to 500+ tools with realistic synthetic supplements, then re-running the benchmark to test whether candidate-set size and correctness still hold at production-like scale.

Benchmark Design

Corpus Expansion

  • Start from the current 90 real MCP-tool descriptors.
  • Generate realistic synthetic tools to reach 500+ total tools.
  • Synthetic tools must have plausible names, descriptions, and capability flags.
  • The synthetic corpus must include near-collisions, adjacent capabilities, tempting-but-wrong candidates, mixed permission profiles, and realistic clusters so the filter is genuinely stressed.

Run Conditions

  • Model: claude-sonnet-4-6 only
  • Environment: offline only
  • Repetitions: 3 reps
  • Compare ungated vs per-task filtered behavior at 500+ scale

Primary Question

Does per-task filtering still produce roughly 1-3 tools per task rather than degrading to dozens, and does correctness still hold at this larger corpus size?

Secondary Questions

  • Does preflight still pass cleanly at 500+ scale?
  • Do token savings continue to scale roughly linearly with corpus growth?

Kill Condition

Stop if the synthetic corpus cannot be made realistic enough to stress the filter meaningfully, or if per-task filtering drops below 80% correctness at scale.

Required Evidence

  • Corpus construction method and realism rationale
  • Total corpus size and category distribution
  • Candidate-set size per task
  • Task correctness at 500+ scale
  • Preflight outcome
  • Input token deltas vs ungated
  • Clear conclusion about whether per-task filtering remains production-credible at scale

Completion Contract

  • Epistemic. Take measurements as specified. Report results with: (1) methodology, (2) raw data or summary statistics, (3) any measurement caveats or confounds encountered.
  • No PR expected unless the objective requires instrumentation code.
  • Handoff: Post measurements + methodology to this issue → Close issue → Stop.

Repo: git-scarrow/tool-capability-protocol | Branch: main | Environment: dev

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions