Community · E2 · artifact verified

Benchmark Jev experiments against baselines

A reproducible experiments repo where each folder tests one Jev idea against a baseline: first up, Jev as an MCP tool router - picking the 5-10 tools an agent needs instead of ~250 - measured on MCP-Atlas, MCP-Bench and MCP-Universe vs OpenAI's decision model.

01 · Role in the system

What Jev does here

The headline experiment targets the tool-context bloat in MCP agents: instead of loading ~250 tool definitions, Jev picks the 5-10 an agent actually needs for the task. Measured across MCP-Atlas, MCP-Bench and MCP-Universe against OpenAI's decision model, Jev's top-10 contained the right tool 80% of the time on MCP-Atlas versus 64%, with far better calibration (ECE 0.11 vs 0.33) at roughly half the cost and ~96% less tool context. The repo's format is the discipline: one folder per idea, every experiment run against a baseline, benchmarks published for reproduction, all judged through the Decisions API (OpenRouter or Vercel AI Gateway).

02 · Control boundary

Where Jev sits

Jev as the tool pre-selector for MCP agents, studied rigorously: typed decisions shrink ~250 tool definitions to the relevant 5-10, with per-benchmark scores, calibration metrics (ECE) and cost comparisons published against a vendor baseline.

Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.

03 · Known limits

What this evidence does not prove

  • Created 2026-10-09, 0 stars; single-author benchmarks.
  • Headline numbers are author-run; the repo invites reproduction but none is published yet.
  • First experiment only - the folder format promises more ideas.

04 · Attribution

Public sources

This is a Community record: the project was published by a third-party community author.