Community · E2 · artifact verified
Benchmark Jev experiments against baselines
A reproducible experiments repo where each folder tests one Jev idea against a baseline: first up, Jev as an MCP tool router - picking the 5-10 tools an agent needs instead of ~250 - measured on MCP-Atlas, MCP-Bench and MCP-Universe vs OpenAI's decision model.
01 · Role in the system
What Jev does here
The headline experiment targets the tool-context bloat in MCP agents: instead of loading ~250 tool definitions, Jev picks the 5-10 an agent actually needs for the task. Measured across MCP-Atlas, MCP-Bench and MCP-Universe against OpenAI's decision model, Jev's top-10 contained the right tool 80% of the time on MCP-Atlas versus 64%, with far better calibration (ECE 0.11 vs 0.33) at roughly half the cost and ~96% less tool context. The repo's format is the discipline: one folder per idea, every experiment run against a baseline, benchmarks published for reproduction, all judged through the Decisions API (OpenRouter or Vercel AI Gateway).
02 · Control boundary
Where Jev sits
Jev as the tool pre-selector for MCP agents, studied rigorously: typed decisions shrink ~250 tool definitions to the relevant 5-10, with per-benchmark scores, calibration metrics (ECE) and cost comparisons published against a vendor baseline.
Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.
03 · Known limits
What this evidence does not prove
- Created 2026-10-09, 0 stars; single-author benchmarks.
- Headline numbers are author-run; the repo invites reproduction but none is published yet.
- First experiment only - the folder format promises more ideas.
04 · Attribution
Public sources
This is a Community record: the project was published by a third-party community author.
- NeelakandanNC ↗Community · github · public · checked 2026-10-09