Community · E2 · artifact verified
Judge Azure Foundry agents with typed evaluators
Four drop-in evaluators for Azure AI Foundry, rebuilt on Jev: Intent Resolution, Task Adherence, Tool Call Accuracy, Groundedness - each metric becomes small typed questions answered in one call, combined into a 1-5 score listing its checks.
01 · Role in the system
What Jev does here
Every Foundry evaluator is decomposed: instead of asking an LLM judge for a holistic grade, the metric becomes several narrow typed questions - did the agent resolve the intent, adhere to the task, call tools accurately, stay grounded - all sent in a single systemone call per conversation, with plain code combining the calibrated answers into the 1-5 score Foundry expects. Each verdict enumerates the checks that drove it, so a low score comes with its reasons attached. A demo web app runs Jev next to Foundry's own LLM-judge evaluators on your own dataset, measuring latency, cost and agreement between the two judges.
02 · Control boundary
Where Jev sits
Jev as the decomposed judge inside a vendor evaluation harness: each official metric becomes narrow typed questions answered in one call, code combines them into the expected score, and agreement against the vendor's LLM judge is itself measured.
Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.
03 · Known limits
What this evidence does not prove
- Azure AI Foundry-specific; other evaluation harnesses need their own adapters.
- Agreement numbers are author-measured on the demo dataset.
- 1 star; single-maintainer project.
04 · Attribution
Public sources
This is a Community record: the project was published by a third-party community author.
- nguyennhianhtri ↗Community · github · public · checked 2026-10-03