Community · E2 · artifact verified
Evaluate Jev on the SNIPS NLU benchmark
An evaluation of Jev on the SNIPS natural-language-understanding benchmark - intent detection and slot filling - asking how far a model that never generates text gets on a task normally solved by a trained tagger, using label names alone.
01 · Role in the system
What Jev does here
SNIPS is the standard NLU benchmark: pick the intent behind an utterance and tag the slots inside it. The harness asks Jev typed questions over each example - one choice for the intent, per-slot questions for the fill - using nothing but the label names as criteria, with no training, fine-tuning or few-shot examples. A Python package wraps the run (src/jevsnips with a run script and a Jev client), tests cover the judge integration, and the results quantify where calibrated zero-shot judgments land on a tagging task that classically belongs to trained sequence models.
02 · Control boundary
Where Jev sits
Jev as the zero-shot judge on a standard NLU benchmark: one typed choice per intent and per-slot questions using label names as the only criteria, wrapped in a tested Python package that reports the scores.
Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.
03 · Known limits
What this evidence does not prove
- SNIPS only; other NLU benchmarks are out of scope.
- Zero-shot by design - no comparison against fine-tuned tagger accuracy is claimed.
- Created 2026-10-04, 2 stars; single-author evaluation.
04 · Attribution
Public sources
This is a Community record: the project was published by a third-party community author.
- will-rice ↗Community · github · public · checked 2026-10-05