Community · E2 · artifact verified

Answer finance questions from two judged passages

A finance RAG benchmark where Jev picks the passages: the agentic baseline reads whole SEC filings - 82,908 tokens for 46 of 50 right - while Jev-judged retrieval reads about 840 tokens and answers all 50; benchmark code and write-up are public.

01 · Role in the system

What Jev does here

The pipeline fetches the latest 10-K for a ticker, chunks the filing text, and lets BM25 rank chunks before Jev sees the shortlist. One call then asks two typed questions: a Choice over the candidate passages - which one states the information needed to answer - and a Noul asking whether the passages contain the answer at all. The writer is restricted to the picked passages and answers in one or two sentences. In the published run this reads about 840 tokens per question where the agentic baseline - the Vals AI finance agent reading whole filings - consumes 82,908, and scores 50 of 50 against its 46. The harness is public: multi-question benchmark runs, a retrieval benchmark, retries with backoff, and a write-up on the author's site; judge calls route through OpenRouter.

02 · Control boundary

Where Jev sits

Jev as the passage judge inside RAG: BM25 narrows, a Choice picks the passage that bears on the question and a Noul gates whether it exists, and the writer sees only what was picked - two orders of magnitude fewer tokens than whole-document agentic reading.

Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.

03 · Known limits

What this evidence does not prove

  • Benchmark is author-run against one named baseline; no independent replication yet.
  • Judge calls route through OpenRouter, not the native api.typesafe.ai endpoint.
  • Single-issuer 10-K pipeline; other filings and languages are untested.

04 · Attribution

Public sources

This is a Community record: the project was published by a third-party community author.