Community · E2 · artifact verified

Evaluate Jev on the SNIPS NLU benchmark

An evaluation of Jev on the SNIPS natural-language-understanding benchmark - intent detection and slot filling - asking how far a model that never generates text gets on a task normally solved by a trained tagger, using label names alone.

01 · Role in the system

What Jev does here

SNIPS is the standard NLU benchmark: pick the intent behind an utterance and tag the slots inside it. The harness asks Jev typed questions over each example - one choice for the intent, per-slot questions for the fill - using nothing but the label names as criteria, with no training, fine-tuning or few-shot examples. A Python package wraps the run (src/jevsnips with a run script and a Jev client), tests cover the judge integration, and the results quantify where calibrated zero-shot judgments land on a tagging task that classically belongs to trained sequence models.

02 · Control boundary

Where Jev sits

Jev as the zero-shot judge on a standard NLU benchmark: one typed choice per intent and per-slot questions using label names as the only criteria, wrapped in a tested Python package that reports the scores.

Code owns the loop, permissions, thresholds, validation, and side effects. Jev owns only the bounded judgments described above.

03 · Known limits

What this evidence does not prove

  • SNIPS only; other NLU benchmarks are out of scope.
  • Zero-shot by design - no comparison against fine-tuned tagger accuracy is claimed.
  • Created 2026-10-04, 2 stars; single-author evaluation.

04 · Attribution

Public sources

This is a Community record: the project was published by a third-party community author.