Assess a simulated company under attack
A local cybersecurity lab replays synthetic telemetry and asks Jev for compromise probability, classification, severity, and an advisory response as evidence accumulates.
The complete index
147 projects and examples that demonstrate what Jev can do, each one checked against its public artifact and re-read on a rolling basis. Search the ledger below, or browse by domain — every cluster carries an editorial note on what it proves and where the evidence stops.
Public evidence ledger
147 of 147 records
A local cybersecurity lab replays synthetic telemetry and asks Jev for compromise probability, classification, severity, and an advisory response as evidence accumulates.
An official cookbook asks 13 regulatory questions over a pinned GDPR article in one Jev request and compares that batch with 13 separate requests.
An official cookbook combines exact quote matching with a Jev relation judgment to label citations verified, unsupported, contradicted, or fabricated.
A simulated quadrotor uses Jev for low-frequency tactical judgments while classical vision, flight control, and safety reflexes remain in code.
A reproducible harness lets Jev direct combat, exploration, and economy actions in the original StarCraft shareware campaign.
A computer-use loop combines OCR and accessibility data, then asks Jev which bounded action should move the Mac toward a plain-English goal.
A battle harness reads FireRed state from RAM and lets Jev choose the next legal move or switch while ordinary code advances the fight.
A Claude Code plugin and npm library asks Jev which old tool calls and results still matter, then drops or truncates the stale ones, so /compact replaces the lossy built-in summary with the kept messages verbatim.
TypeSafe's smart-home demo evaluates a request against many typed questions in parallel, then lets code use only the answers relevant to that request.
A browser agent turns the visible DOM into an indexed action space and uses Jev to choose the next operation and compatible target.
A metasearch front end lets Jev choose the query, sources, and time range, then score every result for relevance while code fans out to engines.
A Home Assistant integration exposes Jev probabilities, choices, and scores as entities and action responses that automations can use.
A finance RAG benchmark where Jev picks the passages: the agentic baseline reads whole SEC filings - 82,908 tokens for 46 of 50 right - while Jev-judged retrieval reads about 840 tokens and answers all 50; benchmark code and write-up are public.
A Turkish chat toy where whatever you type is answered by one of five fixed phrases - Jev picks which word, and two more answers decide the punctuation, in a single call.
A DuckDB extension where jev_choice, jev_score, and jev_noul are scalar functions over table rows: the criteria literal is both the set of permitted answers and the column's SQL type.
A starter kit from Kinde: identity and permissions decide what an agent may do, and Jev judges each call in about 200 milliseconds - does it match the request, is it destructive, does it follow planted text, does it exfiltrate - before anything runs.
A reproducible experiment beyond the 'optimal' Wordle solver: Jev scores how answer-like each of 12,972 accepted words is, the prior feeds entropy search over 1,925 days of NYT answers, and every claim is archived with pinned inputs and a reproduce script.
A Flutter plugin that brings typed Jev calls to Dart apps: question builders for Noul and Choice with option-count validation, the systemone wire protocol on the native endpoint, and a Python-side test harness for the plugin's bridge.
OAS Sentinel compares two OpenAPI documents in two layers: deterministic checks find structural breaks, and Jev answers bounded semantic questions about changed prose - retries, ordering, pagination, error meaning - that schema diffs cannot see.
An experiment that makes a model which cannot write text answer anyway: for every word of the reply Jev picks 1 of 254 meaning-based word groups, then the word inside that group - and the README reports exactly where that stops working.
Lossless Rewrite closes the loop on AI shortening your report: your model rewrites, and Jev checks every protected idea survived - exact wording, meaning, or a reviewed checklist - then helps repair what went missing.
Paste a suspicious SMS, email, DM or listing into ScamCheck - web app, API or browser extension - and get a scam verdict, risk score, plain-English reasons and next steps, with all wording from the project's own templates rather than a model.
A trading loop reads the Kuru MON-USDC order book and asks Jev for a buy-or-sell judgment before code places a post-only limit order.
A GitHub Action that classifies test failures as regression, flaky, environment, or unknown: deterministic signals first, one structured Jev choice second, and a local policy that never auto-reruns tests or masks failures.
An autonomous bot that plays the Chrome T-Rex Runner to a thousand points by asking Jev which action each obstacle requires - jump, duck, or run - and letting a measured physics model decide exactly when to press.
A research task treating Jev's answer distributions as classifier features: many small typed questions about an SVG, responses weighted and combined until the signal classifies the image - a study of whether decision calls can stand in for maths on pixels.
jevtrim is a comparative analysis of context compaction driven by calibrated judgments instead of summarization: Jev scores every chunk for relevance, ordinary Python keeps what fits the token budget, and the result is auditable and replayable offline.
Jev drives a LIBERO robot through 27 control inputs with layered decisions - intent, then motion family, then input - while reversible physics previews evaluate candidate effects locally before anything executes.
A Chrome MV3 extension that turns speech into browser actions: Jev routes each spoken command to open a site, search, click a link, fill a form field or go back, with typed answers instead of parsed free text; ships with tests, CI and a side-panel command log.
A Chrome extension covers every YouTube comment the moment it appears, asks Jev one Noul question per comment, and keeps it covered whenever the probability says it discloses a concrete plot event.
Pipette reads arXiv, bioRxiv, medRxiv and 58 journals each morning and publishes a short, diverse daily edition: Jev labels and ranks, quoted sentences are the authors' own abstract lines, method and probabilities are public, and output is CC0 open data.
A Home Assistant conversation agent - installable via HACS - that decides with a TypeSafe System One model instead of an LLM: your spoken or typed commands become typed choices and Nouls that drive devices, with metered costs documented.
A tModLoader mod whose boss fights are driven by Jev: every 200ms one request asks a nine-way intent, a five-band danger score, and whether to dash or jump; a per-frame reflex layer turns intents into keypresses.
Give jev-browser a task and a URL: Jev picks one action per step from the page's clickable, typeable, and selectable elements and scores goal-met and stuck likelihood, while code owns budgets, recovery, and stop gates.
An Android automation agent with a two-tier brain: Jev decides fast from the accessibility tree, a vision agent takes over only when the structural view is ambiguous, and ADB executes - cheap judge on the hot path, expensive one on the exceptions.
A proof-of-concept Android loop stabilizes the screen, builds a short list of valid actions, and lets Jev pick one while code executes it.
bside pilots the Aside browser with Jev instead of a chat LLM: each tick answers which action, which element, and whether the goal is met, over an action schema the pilot cannot hallucinate outside of.
An evaluation of Jev on the SNIPS natural-language-understanding benchmark - intent detection and slot filling - asking how far a model that never generates text gets on a task normally solved by a trained tagger, using label names alone.
JDE, the Jev Decision Engine, wraps Jev as an MCP server: any MCP client - Claude Code, Cursor, your own agent - gets typed decision tools, with a policy layer, a decision ledger, and recorded evals comparing the hosted jev-1.13.0 against local alternatives.
An MCP server gives compatible agents tools for classification, scoring, checking, matching, screening, and custom typed Jev questions.
A browser app for systematic-review extraction: Jev never writes the answer - it points at line ids in trial reports and supplements, and code copies the quote out with its file, page, row, or slide, highlighted where it sits.
A live link-checker extracts each claim from a submitted page, asks Jev whether the cited excerpts support, contradict, or fail to establish it, and reports REAL or FAKE only when enough evidence agrees.
Osso fades the parts of a page that are not what the reader came for: each sentence of the main text gets a Jev probability, workspace sites are covered by default, and password fields, account pages and reviews are left alone by construction.
A tiny Japanese game with no send button: type IT buzzwords and each keystroke gets a Jev judgment that stretches a meter, so 25 seconds of play answers what curl never does - how fast and how cheap the model feels inside a real app.
A local-first document filer: text is extracted locally, Jev decides category, confidentiality and prompt-injection risk as typed Choices, and low-confidence or suspicious files land in review lanes - never overwritten, with audit preview and undo.
A Filament plugin for Laravel admin panels that filters tables by natural language instead of SQL: type "the customer is angry" and Jev's typed decisions drive the where-clauses - shipped as a Packagist package with a driver system, tests and a live demo.
A single-file Python portal that filters news and YouTube feeds by interests you describe in plain English: the official typesafe-sdk scores every item, routine business hides separately, and the result is one static page with News and YouTube tabs.
jevpipe pipes thousands of lines, files or records through one question and gets a typed judgment per item - grep-style filtering where Jev decides - shipped as a Rust binary on PyPI with an agent skill that teaches coding agents when to reach for it.
An open-source Chrome extension asks Jev to judge each visible X post for relevance, substance, practical value, promotion, and engagement bait, then dims, collapses, or hides it under weights the reader owns.
Describe what you want to do in San Francisco and a rules engine works out which permits you need - Jev answers each rules question as a typed choice, falls back to asking you when unknown, and summarizes fees and deadlines per permit.
For people with aphasia who know what they mean but can't get the word out: describe it any way you can, and Jev picks the best guesses from a fixed 900-word list as big tap-to-hear picture tiles - never inventing a word.
jevcal measures a typed decision model on private labeled data, fits per-question thresholds to a target accuracy, and fails CI when a model update drifts.
tenet makes agents fix rule violations before you ever see the diff: rules live in a YAML file in plain language, and Jev answers each one with a calibrated probability that becomes a pass-or-fail cutoff.
A guard for OpenClaw agents that reads an outgoing message and where it is going: Jev judges whether a client name, credential or internal hostname is about to reach the wrong readers, then confirms, blocks or rewrites per channel.
A requirements quality gate: one batched call asks 22 atomic questions across five MECE facets about an AI-generated requirement, and thresholds route it pass, human review, or reject - never rewriting, only judging.
Four proof-of-concepts wiring Jev in front of a pay-per-call API marketplace: relevance below 0.7 confidence skips the paid call entirely, a typed choice routes to exactly one endpoint, and a second independent gate enforces per-team budget caps.
A Pi extension asks Jev to flag destructive, exfiltrating, or out-of-scope tool calls and to classify failures in command output.
A Japanese companion game where one Jev pass decides the character's true feeling - Choice, affection Score, dislike Noul, topic Choice - in 0.2 to 0.5 seconds, so her face and a one-liner land before Claude-written dialogue and a Gemini TTS voice.
Mina, a simulated 34-year-old librarian, runs on two systems: Jev reads her body, senses and clock every second and accumulates feelings; only when a feeling crosses its line does an LLM stop and think.
A bridge connecting images, video streams and RGB-D cameras to Jev's judgment engine: identify what matters in a frame, estimate risk, score a situation or judge many visible objects at once - typed answers over the visual world, 46 stars in its first day.
jgrep answers semantic queries like catches an error and silently ignores it over a whole source tree in about two seconds for a cent: one typed yes-or-no judgment per code chunk, sixteen chunks per request, no index.
TraceDocs structures documents into source-linked blocks; Jev judges each with four Nouls - relevant, evidence, contradicts-premise, prompt-injection - returning a cited evidence set with a trace; the LLM writes from evidence and refuses when none exists.
Kassad brings calibrated guardrails to .NET: every prompt, completion, tool call and citation passes narrow typed checks answered by Jev, batched one round trip per stage, thresholded in code into Allow, Flag, Review, or Block.
Describe a mood - a slow, sad waltz - and a live piano plays it, with Jev deciding continuously as it goes: every musical choice is a typed question answered with probabilities, streamed to a public site in real time.
Doom or Bloom maps where you stand between AI doom and bloom: a dynamic interview where the engine picks the next curated question by where your answers are thinnest, with Jev interpreting answers and scoring candidate follow-ups.
Four drop-in evaluators for Azure AI Foundry, rebuilt on Jev: Intent Resolution, Task Adherence, Tool Call Accuracy, Groundedness - each metric becomes small typed questions answered in one call, combined into a 1-5 score listing its checks.
Paper Radar reads all of arXiv so you read the few that matter: every new paper is judged against plain-English interests with calibrated per-interest probabilities - about six cents a day for everything, no pre-filtering.
An npm library that lets game characters argue back: give a name, a persona and a goal, pass what the player typed, and Jev judges whether that character - with those values - was convinced; the same line can win a greedy merchant and offend an honest guard.
A Chrome extension for YouTube: ask the video a question in plain text, and Jev's typed judgments locate the moment that answers it - the player seeks straight there instead of you scrubbing.
A generative painting instrument: select part of a sketch and Jev chooses its paint material - one typed Choice per region with probabilities and certainty bands - and the material flies in and paints itself; demo film and live gallery included.
A fuzzy linter that watches a coding agent write and speaks up 0.3 seconds later: Jev checks the file against your team's rules - race conditions in effects, missing cleanup, house style - so the agent fixes them before any human reviews the code.
A demonstration pushing the judgment-only model past its envelope: Jev 'writes' by answering which-word-comes-next in 250-word batches - each option shown as the whole reply so far plus the word - top-3 shortlist, final pick, until sentence end.
An independent study routes Jev confidence into a larger model on two labelled datasets and shows the winning settings do not transfer between them.
A weekly measured series pitting Jev against frontier LLMs on the same real workflow steps: week one routed inbound leads - Jev 90 percent correct at 366 milliseconds and four cents per thousand, against Sonnet 5's 78 percent at 2.6 seconds and three dollars.
Finding and extracting repeating patterns from noisy sequences with Jev: instead of a hand-tuned distance metric, typed questions decide what counts as the same pattern, and matches come back with probabilities instead of thresholds you guess.
jevmod scores every community message for spam, scam, harassment, NSFW, self-harm, doxxing, off-topic and custom plain-English rules; operators set thresholds, decisions log their numbers, and bots ship for Discord, Twitch, YouTube and Reddit.
An SAP Commerce extension where Jev answers four yes-or-no questions per product review - abusive, spam, personal data, on-topic - and code turns the probabilities into approve, reject, or pending, with a dry-run mode that judges against human decisions first.
One text box that becomes the right UI as you type - an event card, checklist, timer, color picker, bill splitter or poll - with Jev Nouls deciding what the input means and code rendering the component; live demo included.
At each node of a Neo4j graph the outgoing relationships become Choice options; Jev returns a full probability distribution over which one to follow, with a goal-reached Noul riding in the same call.
A shared 1,000-by-1,000 emoji canvas where humans place strokes and Jev paints with them: after each stroke one typed call picks a contextually relevant emoji and where to put it, and a yes/no decides whether your stroke was finished.
A Next.js dashboard that streams live crypto data, computes fifteen moving averages and eleven oscillators into one typed market state, and lets a Jev agent trade a simulated 100k portfolio - paper only, no broker connected.
Jev Yarn is a party game where everyone writes the next sentence and Jev picks the winner: one taste request scores every line on four dimensions, and Nouls handle room filters and whether the story feels finished.
Drop a file or paste a URL and get a floating window with the right tool: Jev decides and composes a viewer from the registry, and when the format is unknown, Haiku invents a spec for a mini-app on the spot.
Jev Trip is an explainable day-trip planner where the LLM plans ahead and Jev chooses and checks: scope choices keep one day in one city, per-place choices rank every candidate with probabilities, and code owns routes, times and validation.
Jev plays Yasuo in League of Legends through three decision heads - strategy once a second, tactics six to seven times a second while units are on screen, and build checks every twenty seconds - while code reads the game and executes.
A single-file Pac-Man that asks the System One endpoint for typed decisions per game tick, playable live without setup or locally with your own key pointed at the official endpoint.
An experimental controller translates NES telemetry into object-centric JSON and lets Jev choose the next legal controller macro.
A pluggable decision layer for the ego agent: System One (Jev) by default, swappable to local or other OpenAI-compatible backends, fail-closed guardrails - and the README's whole argument is that it is measured: 16 suites, 429 checks, rerun in full.
A macOS menu-bar switcher asks Jev which of the ten most recent apps you intend on a hotkey press and falls back to the last-used app on any failure.
A Kosovo electronics shop’s support agent reads a unified Albanian and English inbox and lands every message on auto-resolve, verification, or escalate - with Jev only proposing toward caution and deterministic code owning facts, access, and the final call.
A Pi extension uses Jev to decide which old tool calls and results still matter while keeping conversation text verbatim.
A zsh plugin asks Jev which recent command you are completing and shows the best match with its probability, while code owns gating and acceptance.
A browser extension where Jev answers 17 typed questions about every LinkedIn post - thirteen AI-tell questions plus four bait questions in the same request - and fixed weights turn the answers into a badge and a bait chip.
The kamchatka terminal agent asks Jev to place each pending shell command on a three-level safety rubric - reads and reports, changes something reversibly, destroys or sends something out - drawn green, yellow, or red beside the permission prompt.
A Japanese-language experiment scoring 100 labeled customer inquiries with the official TypeSafe SDK: one request per inquiry answers six questions at once - sentiment, emotion, anger intensity, urgency, churn risk, and sarcasm.
OCR flattens superscripts: a footnote star, an endnote number and a unit power land in the stream as look-alike tokens. A regex over-finds the suspects, then Jev classifies each - footnote, citation or unit - as a typed choice so formatting can be restored.
A support inbox where Jev decides per span whether text is personal data and of what kind, and a redact() SQL function enforces the masking in Postgres by the viewer's clearance - content-aware, not pattern-based.
A Chrome extension finds ad-shaped DOM candidates and asks Jev whether each candidate is a paid advertisement before code removes it.
Product search inside PostgreSQL - typos, barcodes, typeahead, facet counts, all in SQL - with an optional second stage asking Jev two questions per search to rerank the shortlist; the README is itself the report, three failed versions included.
A Pinecone official examples repository: full-text search retrieves 200 candidates, and one Jev judgment pass reranks them to 10 by natural-language criteria, with a Claude baseline column for comparison.
Jev Social pairs the decision model with socai, a CLI that drives your real Chrome across Instagram, TikTok, and LinkedIn: Jev chooses each next read-only operation - search, open a post or profile, read comments - and socai executes it.
Hunkpick resolves git conflicts by computing every plausible resolution itself - ours, theirs, union, line merge, token merge - discarding the ones that fail to parse, and asking Jev only to choose.
A Chrome MV3 extension that restyles any site from a plain-English prompt: one Jev call picks palette, fonts, spacing and intent from a fixed catalog, and deterministic code compiles role-stamped CSS - no LLM ever writes CSS that can break.
nudgement reviews commit messages, code, comments, tests and UI copy before you commit: exact checks plus focused Jev questions catch what formatters cannot - a message that misrepresents the diff, or a test that passes even when behavior broke.
A single binary that reviews a diff against the ten refactoring rules of Clausen's Five Lines of Code: the countable rules run on a real parser, and Jev answers the judgment rules as typed questions whose probabilities set the bar.
A single-file local pull-request reviewer: deterministic code does the plumbing while Jev judges each hunk with a real-issue Noul, scores severity, and returns a PR-level risk with a needs-human probability - no agent loop, no prompts to tune.
FastGate fronts an English/Uzbek/Russian university helpdesk with four narrow Jev judgments per message plus per-passage grounding, routes deterministically in code, and ships an independent benchmark of Jev on low-resource languages.
A staged review workflow uses Jev to identify risky areas, select evidence, classify mechanisms, score severity, and route follow-up checks.
A DJ chatbot's pre-LLM router, built as a learning study: free-form /play requests - English, Spanish, typos, pasted lyrics, emojis - classified by Jev in one call (~500 ms, ~$0.000036) into a structured hint that tells the expensive LLM what it is handling.
A multi-bot OKX trading desk with Jev as the brain and code as the body: each tick reads market and account state, computes features in TypeScript, asks Jev typed questions, then holds or places and cancels real orders - dashboard as the glass.
Jev Mart is a Japanese e-commerce demo where every operational decision - listings, inquiries, reviews, triage - is a typed Jev call and an LLM is never invoked; it ships with a twelve-chapter lecture set that teaches the pattern.
jev-axi is a CLI that puts a half-second Jev opinion in front of every command an agent runs: pick, rate, check, rank, triage, and guard, plus a PreToolUse hook that blocks risky Bash calls before they execute.
A CLI named jev turns Jev's typed questions into pipable commands - verify, screen, classify, extract, find, rerank, match, route, ask, compact, batch - with keychain auth and exit codes a script can gate on.
Book Aurora sends every ~90-word passage of Frankenstein to Jev as ten parallel score questions - nine emotions plus overall intensity - and draws each answer as one feathered row of a full-book aurora strip.
Supercov asks Jev yes-or-no questions about every source file, does the arithmetic in code, and turns weak spots and coverage gaps into agent tasks.
jevmeter turns a video into a BS meter: every sentence scored and every dodge flagged under a preset viewpoint - debates, earnings calls, podcasts, pitches - rendered as a 16:9 edit you can post.
A daily pipeline reads the BOE with one Jev call per provision, scoring impact, tagging topics, and selecting an original paragraph as the summary.
A sibling GitHub Action that decides which allowlisted test groups a pull request needs and which can safely be skipped: the group list stays in your config, Jev picks among it, and deterministic rules a model cannot bypass have the final word.
A working example of a shadow decision layer under a production CRM: one call returns intention, business division, human-handover probability, urgency, and lead quality - decided and logged, never sent.
OpenRecurSearch is an agentic web-search interface where nothing stays hidden: each research layer and each Jev decision appears live in the chat - option probabilities, score and latency included - while the markdown report streams into the side panel.
Type a description and matching emoji fly out of the pile: one request carries the query as shared state with one score question per emoji, answered in parallel, and code ranks and floors the results.
An Android noise gate for notifications and SMS that asks Jev whether each message is an ad - and suppresses only what is explicitly flagged. Every uncertain path resolves to allow, because swallowing a verification code costs more than leaking an ad.
A light terminal Gmail client for Omarchy Linux: Jev files incoming mail into Gmail labels only above a confidence threshold, and grades your reply as you write - clarity, tone, length, next step, and which questions you haven't answered yet.
JevSpeak holds a conversation without any generative model in the loop: Jev answers about thirteen parallel questions per turn, the decisions normalize into a semantic IR, and a deterministic compiler writes the sentence.
A hospital delivery robot judged about four times a second: code samples twelve candidate paths and removes predicted collisions, Jev picks one among the survivors, and a deterministic safety brake guarantees no contact - built at the TypeSafe JEVATHON.
A playground asks Jev to decide only closed-vocabulary labels for character, key, meter, and phrasing while deterministic code writes, engraves, and plays the notes.
Spikecast puts Jev at the controls of a fly's simulated brain: the insect walks a road, and every turn, stop or dash is a typed decision with the synapse memory and neural firing visualised beside it - with docs separating what is real from what is modelled.
You are the Town Crier of a 3D town: write one broadcast and all fifty citizens decide in parallel - investigate, join, flee, warn, or ignore - in a single batched choice request.
A Home Assistant add-on runs a crop-steering irrigation engine with Jev judging what fixed rules got wrong - ramp timing, probe trust, shot landing, EC moves, alerts - inside code-enforced physical limits; the README replays real Jev answers from 27 Sep 2026.
An interactive installation - explore a flooded observatory where glowing matter builds paths from your movement, gaze and actions, with real-time Jev decisions deciding how the world responds, shipped as a playable web app with live-decision tests.
Chunk documents where meaning changes, not at a character count: Jev judges sentence continuation, planted instructions are quarantined before embedding, and retrieved passages classified as evidence, conflict or noise - on your existing vector database.
Foreman runs an independent observation loop that asks Jev whether a coding worker is progressing, stuck, complete, or ready for verification.
A calibre plugin that suggests subject and genre tags for your books with Jev: suggestions arrive for review before anything is applied, existing tags are preserved, and the plugin is independent GPL software with a live demo and production notes.
A quiet clerk for a small online store's inbox: a decision model reads every customer email first and sorts it into four lanes - needs a person now, can wait, template could answer, sales pitch - packaged as an n8n workflow with a shadow mode.
A Portuguese customer-service demo where each client message makes one Jev call with seven typed questions in parallel - intent, urgency, sentiment and more - and a cost table prices it against routing everything to a large LLM instead.
An issue classifier whose decision logic is ordinary Python gated on returned numbers: eight typed questions in one call, labels written only where confidence clears a threshold, and everything else escalated.
Fault triage for an industrial compressed-air unit with Jev: live telemetry in, classified alerts and tickets out, every judgment grounded in the unit's own service manual - read-only by design, with an eval harness that replays recorded decision cassettes.
A read-only Kubernetes controller that turns workload state and Events into stable incident decisions: deterministic rules decide the clear cases, and a Jev decision provider is called only for the ambiguous remainder.
Open-source alert triage where each production alert gets one Jev call with four typed questions - actionable, severity, team and paging - and auditable code turns the probabilities into routing; Jev never pages anyone, it only judges.
A single Rust binary pipes JSONL security alerts through five typed Jev questions and emits validated dispositions that code, not the model, enforces.
Paste a YouTube URL and every comment in the section is classified on four typed axes - then ranked so the ones actually worth a reply float to the top, for creators drowning in comment volume.
Python scripts that ask Jev to select every CVSS metric from a vulnerability description, then compute the numeric score in code exactly per the FIRST specification - for CVSS 3.0, 3.1, and 4.0.
A bilingual study asking whether JEV can serve as a world model for LLM agents - predicting what happens next from typed state questions - benchmarked against generative LLMs on identical predictions, with run manifests and a phase-0 API check archived.
A reader for public-domain Japanese novels where every paragraph is judged by Jev and each character accumulates a trait radar for the page in front of you, plus a running ranking of traits for the story so far.
Qualm is a screen-time app for Apple silicon Macs, local-first, built on two judges: Kev running on your machine and TypeSafe's hosted Jev - typed choices decide what counts as a distracting session without your activity leaving the machine unless you opt in.
For each topic you follow, jrp runs a pipeline: your questions steer the web search, Jev decides which findings count, and an LLM you choose writes the morning note - in English, Chinese or Japanese - with parallel stages and recorded claim-fidelity checks.
Jev cannot write a single note, so code lays out the bars and the legal notes with musical facts attached, and Jev chooses: mode, tempo, form, every chord, every note - one call per decision, every call replayable.
Curated paths through the ledger
13 clusters
The largest cluster in the index: coding agents under supervision, code-review routing, citation checks, and typed tool gates. The recurring shape is an LLM-powered pipeline that would otherwise prompt a chat model at a decision point and instead asks Jev for a bounded judgment the surrounding code can branch on.
Most records here are reproducible repositories, and several ship as installable CLI or MCP integrations. What the cluster does not yet show is a large-scale production deployment with published reliability numbers — treat every reported metric as single-artifact evidence.
An official cookbook combines exact quote matching with a Jev relation judgment to label citations verified, unsupported, contradicted, or fabricated.
A Claude Code plugin and npm library asks Jev which old tool calls and results still matter, then drops or truncates the stale ones, so /compact replaces the lossy built-in summary with the kept messages verbatim.
A finance RAG benchmark where Jev picks the passages: the agentic baseline reads whole SEC filings - 82,908 tokens for 46 of 50 right - while Jev-judged retrieval reads about 840 tokens and answers all 50; benchmark code and write-up are public.
A DuckDB extension where jev_choice, jev_score, and jev_noul are scalar functions over table rows: the criteria literal is both the set of permitted answers and the column's SQL type.
A Flutter plugin that brings typed Jev calls to Dart apps: question builders for Noul and Choice with option-count validation, the systemone wire protocol on the native endpoint, and a Python-side test harness for the plugin's bridge.
OAS Sentinel compares two OpenAPI documents in two layers: deterministic checks find structural breaks, and Jev answers bounded semantic questions about changed prose - retries, ordering, pagination, error meaning - that schema diffs cannot see.
Lossless Rewrite closes the loop on AI shortening your report: your model rewrites, and Jev checks every protected idea survived - exact wording, meaning, or a reviewed checklist - then helps repair what went missing.
A GitHub Action that classifies test failures as regression, flaky, environment, or unknown: deterministic signals first, one structured Jev choice second, and a local policy that never auto-reruns tests or masks failures.
A Chrome MV3 extension that turns speech into browser actions: Jev routes each spoken command to open a site, search, click a link, fill a form field or go back, with typed answers instead of parsed free text; ships with tests, CI and a side-panel command log.
Give jev-browser a task and a URL: Jev picks one action per step from the page's clickable, typeable, and selectable elements and scores goal-met and stuck likelihood, while code owns budgets, recovery, and stop gates.
bside pilots the Aside browser with Jev instead of a chat LLM: each tick answers which action, which element, and whether the goal is met, over an action schema the pilot cannot hallucinate outside of.
JDE, the Jev Decision Engine, wraps Jev as an MCP server: any MCP client - Claude Code, Cursor, your own agent - gets typed decision tools, with a policy layer, a decision ledger, and recorded evals comparing the hosted jev-1.13.0 against local alternatives.
An MCP server gives compatible agents tools for classification, scoring, checking, matching, screening, and custom typed Jev questions.
A tiny Japanese game with no send button: type IT buzzwords and each keystroke gets a Jev judgment that stretches a meter, so 25 seconds of play answers what curl never does - how fast and how cheap the model feels inside a real app.
A Filament plugin for Laravel admin panels that filters tables by natural language instead of SQL: type "the customer is angry" and Jev's typed decisions drive the where-clauses - shipped as a Packagist package with a driver system, tests and a live demo.
jevpipe pipes thousands of lines, files or records through one question and gets a typed judgment per item - grep-style filtering where Jev decides - shipped as a Rust binary on PyPI with an agent skill that teaches coding agents when to reach for it.
Describe what you want to do in San Francisco and a rules engine works out which permits you need - Jev answers each rules question as a typed choice, falls back to asking you when unknown, and summarizes fees and deadlines per permit.
jevcal measures a typed decision model on private labeled data, fits per-question thresholds to a target accuracy, and fails CI when a model update drifts.
tenet makes agents fix rule violations before you ever see the diff: rules live in a YAML file in plain language, and Jev answers each one with a calibrated probability that becomes a pass-or-fail cutoff.
A requirements quality gate: one batched call asks 22 atomic questions across five MECE facets about an AI-generated requirement, and thresholds route it pass, human review, or reject - never rewriting, only judging.
Four proof-of-concepts wiring Jev in front of a pay-per-call API marketplace: relevance below 0.7 confidence skips the paid call entirely, a typed choice routes to exactly one endpoint, and a second independent gate enforces per-team budget caps.
A Pi extension asks Jev to flag destructive, exfiltrating, or out-of-scope tool calls and to classify failures in command output.
A bridge connecting images, video streams and RGB-D cameras to Jev's judgment engine: identify what matters in a frame, estimate risk, score a situation or judge many visible objects at once - typed answers over the visual world, 46 stars in its first day.
jgrep answers semantic queries like catches an error and silently ignores it over a whole source tree in about two seconds for a cent: one typed yes-or-no judgment per code chunk, sixteen chunks per request, no index.
Kassad brings calibrated guardrails to .NET: every prompt, completion, tool call and citation passes narrow typed checks answered by Jev, batched one round trip per stage, thresholded in code into Allow, Flag, Review, or Block.
Four drop-in evaluators for Azure AI Foundry, rebuilt on Jev: Intent Resolution, Task Adherence, Tool Call Accuracy, Groundedness - each metric becomes small typed questions answered in one call, combined into a 1-5 score listing its checks.
An npm library that lets game characters argue back: give a name, a persona and a goal, pass what the player typed, and Jev judges whether that character - with those values - was convinced; the same line can win a greedy merchant and offend an honest guard.
A Chrome extension for YouTube: ask the video a question in plain text, and Jev's typed judgments locate the moment that answers it - the player seeks straight there instead of you scrubbing.
A fuzzy linter that watches a coding agent write and speaks up 0.3 seconds later: Jev checks the file against your team's rules - race conditions in effects, missing cleanup, house style - so the agent fixes them before any human reviews the code.
Finding and extracting repeating patterns from noisy sequences with Jev: instead of a hand-tuned distance metric, typed questions decide what counts as the same pattern, and matches come back with probabilities instead of thresholds you guess.
One text box that becomes the right UI as you type - an event card, checklist, timer, color picker, bill splitter or poll - with Jev Nouls deciding what the input means and code rendering the component; live demo included.
Drop a file or paste a URL and get a floating window with the right tool: Jev decides and composes a viewer from the registry, and when the format is unknown, Haiku invents a spec for a mini-app on the spot.
A pluggable decision layer for the ego agent: System One (Jev) by default, swappable to local or other OpenAI-compatible backends, fail-closed guardrails - and the README's whole argument is that it is measured: 16 suites, 429 checks, rerun in full.
A macOS menu-bar switcher asks Jev which of the ten most recent apps you intend on a hotkey press and falls back to the last-used app on any failure.
A Pi extension uses Jev to decide which old tool calls and results still matter while keeping conversation text verbatim.
A zsh plugin asks Jev which recent command you are completing and shows the best match with its probability, while code owns gating and acceptance.
The kamchatka terminal agent asks Jev to place each pending shell command on a three-level safety rubric - reads and reports, changes something reversibly, destroys or sends something out - drawn green, yellow, or red beside the permission prompt.
OCR flattens superscripts: a footnote star, an endnote number and a unit power land in the stream as look-alike tokens. A regex over-finds the suspects, then Jev classifies each - footnote, citation or unit - as a typed choice so formatting can be restored.
Product search inside PostgreSQL - typos, barcodes, typeahead, facet counts, all in SQL - with an optional second stage asking Jev two questions per search to rerank the shortlist; the README is itself the report, three failed versions included.
Hunkpick resolves git conflicts by computing every plausible resolution itself - ours, theirs, union, line merge, token merge - discarding the ones that fail to parse, and asking Jev only to choose.
A Chrome MV3 extension that restyles any site from a plain-English prompt: one Jev call picks palette, fonts, spacing and intent from a fixed catalog, and deterministic code compiles role-stamped CSS - no LLM ever writes CSS that can break.
nudgement reviews commit messages, code, comments, tests and UI copy before you commit: exact checks plus focused Jev questions catch what formatters cannot - a message that misrepresents the diff, or a test that passes even when behavior broke.
A single binary that reviews a diff against the ten refactoring rules of Clausen's Five Lines of Code: the countable rules run on a real parser, and Jev answers the judgment rules as typed questions whose probabilities set the bar.
A single-file local pull-request reviewer: deterministic code does the plumbing while Jev judges each hunk with a real-issue Noul, scores severity, and returns a PR-level risk with a needs-human probability - no agent loop, no prompts to tune.
A staged review workflow uses Jev to identify risky areas, select evidence, classify mechanisms, score severity, and route follow-up checks.
A DJ chatbot's pre-LLM router, built as a learning study: free-form /play requests - English, Spanish, typos, pasted lyrics, emojis - classified by Jev in one call (~500 ms, ~$0.000036) into a structured hint that tells the expensive LLM what it is handling.
Jev Mart is a Japanese e-commerce demo where every operational decision - listings, inquiries, reviews, triage - is a typed Jev call and an LLM is never invoked; it ships with a twelve-chapter lecture set that teaches the pattern.
jev-axi is a CLI that puts a half-second Jev opinion in front of every command an agent runs: pick, rate, check, rank, triage, and guard, plus a PreToolUse hook that blocks risky Bash calls before they execute.
A CLI named jev turns Jev's typed questions into pipable commands - verify, screen, classify, extract, find, rerank, match, route, ask, compact, batch - with keychain auth and exit codes a script can gate on.
Supercov asks Jev yes-or-no questions about every source file, does the arithmetic in code, and turns weak spots and coverage gaps into agent tasks.
A sibling GitHub Action that decides which allowlisted test groups a pull request needs and which can safely be skipped: the group list stays in your config, Jev picks among it, and deterministic rules a model cannot bypass have the final word.
A light terminal Gmail client for Omarchy Linux: Jev files incoming mail into Gmail labels only above a confidence threshold, and grades your reply as you write - clarity, tone, length, next step, and which questions you haven't answered yet.
Foreman runs an independent observation loop that asks Jev whether a coding worker is progressing, stuck, complete, or ready for verification.
A calibre plugin that suggests subject and genre tags for your books with Jev: suggestions arrive for review before anything is applied, existing tags are preserved, and the plugin is independent GPL software with a live demo and production notes.
A quiet clerk for a small online store's inbox: a decision model reads every customer email first and sorts it into four lanes - needs a person now, can wait, template could answer, sales pitch - packaged as an n8n workflow with a shadow mode.
A Portuguese customer-service demo where each client message makes one Jev call with seven typed questions in parallel - intent, urgency, sentiment and more - and a cost table prices it against routing everything to a large LLM instead.
An issue classifier whose decision logic is ordinary Python gated on returned numbers: eight typed questions in one call, labels written only where confidence clears a threshold, and everything else escalated.
Fault triage for an industrial compressed-air unit with Jev: live telemetry in, classified alerts and tickets out, every judgment grounded in the unit's own service manual - read-only by design, with an eval harness that replays recorded decision cassettes.
A read-only Kubernetes controller that turns workload state and Events into stable incident decisions: deterministic rules decide the clear cases, and a Jev decision provider is called only for the ambiguous remainder.
Open-source alert triage where each production alert gets one Jev call with four typed questions - actionable, severity, team and paging - and auditable code turns the probabilities into routing; Jev never pages anyone, it only judges.
A single Rust binary pipes JSONL security alerts through five typed Jev questions and emits validated dispositions that code, not the model, enforces.
Paste a YouTube URL and every comment in the section is classified on four typed axes - then ranked so the ones actually worth a reply float to the top, for creators drowning in comment volume.
Qualm is a screen-time app for Apple silicon Macs, local-first, built on two judges: Kev running on your machine and TypeSafe's hosted Jev - typed choices decide what counts as a distracting session without your activity leaving the machine unless you opt in.
Browser, desktop, and smart-home control driven from observed state: the visible DOM becomes a table of allowed operations, Home Assistant state becomes automation signals, and Jev picks the operation while deterministic code executes it. This is the clearest demonstration of the split the whole index is about — model chooses, code acts.
Because these records touch real interfaces, the interesting evidence is in the failure boundaries: frames, occlusion, timing windows, and devices that change state mid-command. The limitations field on each record is where authors document what their automation still cannot do.
An official cookbook asks 13 regulatory questions over a pinned GDPR article in one Jev request and compares that batch with 13 separate requests.
A computer-use loop combines OCR and accessibility data, then asks Jev which bounded action should move the Mac toward a plain-English goal.
TypeSafe's smart-home demo evaluates a request against many typed questions in parallel, then lets code use only the answers relevant to that request.
A browser agent turns the visible DOM into an indexed action space and uses Jev to choose the next operation and compatible target.
A Home Assistant integration exposes Jev probabilities, choices, and scores as entities and action responses that automations can use.
A trading loop reads the Kuru MON-USDC order book and asks Jev for a buy-or-sell judgment before code places a post-only limit order.
A Chrome MV3 extension that turns speech into browser actions: Jev routes each spoken command to open a site, search, click a link, fill a form field or go back, with typed answers instead of parsed free text; ships with tests, CI and a side-panel command log.
A Chrome extension covers every YouTube comment the moment it appears, asks Jev one Noul question per comment, and keeps it covered whenever the probability says it discloses a concrete plot event.
A Home Assistant conversation agent - installable via HACS - that decides with a TypeSafe System One model instead of an LLM: your spoken or typed commands become typed choices and Nouls that drive devices, with metered costs documented.
An Android automation agent with a two-tier brain: Jev decides fast from the accessibility tree, a vision agent takes over only when the structural view is ambiguous, and ADB executes - cheap judge on the hot path, expensive one on the exceptions.
A proof-of-concept Android loop stabilizes the screen, builds a short list of valid actions, and lets Jev pick one while code executes it.
Osso fades the parts of a page that are not what the reader came for: each sentence of the main text gets a Jev probability, workspace sites are covered by default, and password fields, account pages and reviews are left alone by construction.
A local-first document filer: text is extracted locally, Jev decides category, confidentiality and prompt-injection risk as typed Choices, and low-confidence or suspicious files land in review lanes - never overwritten, with audit preview and undo.
A single-file Python portal that filters news and YouTube feeds by interests you describe in plain English: the official typesafe-sdk scores every item, routine business hides separately, and the result is one static page with News and YouTube tabs.
An open-source Chrome extension asks Jev to judge each visible X post for relevance, substance, practical value, promotion, and engagement bait, then dims, collapses, or hides it under weights the reader owns.
Four proof-of-concepts wiring Jev in front of a pay-per-call API marketplace: relevance below 0.7 confidence skips the paid call entirely, a typed choice routes to exactly one endpoint, and a second independent gate enforces per-team budget caps.
Mina, a simulated 34-year-old librarian, runs on two systems: Jev reads her body, senses and clock every second and accumulates feelings; only when a feeling crosses its line does an LLM stop and think.
jevmod scores every community message for spam, scam, harassment, NSFW, self-harm, doxxing, off-topic and custom plain-English rules; operators set thresholds, decisions log their numbers, and bots ship for Discord, Twitch, YouTube and Reddit.
An SAP Commerce extension where Jev answers four yes-or-no questions per product review - abusive, spam, personal data, on-topic - and code turns the probabilities into approve, reject, or pending, with a dry-run mode that judges against human decisions first.
Jev Trip is an explainable day-trip planner where the LLM plans ahead and Jev chooses and checks: scope choices keep one day in one city, per-place choices rank every candidate with probabilities, and code owns routes, times and validation.
A macOS menu-bar switcher asks Jev which of the ten most recent apps you intend on a hotkey press and falls back to the last-used app on any failure.
A Kosovo electronics shop’s support agent reads a unified Albanian and English inbox and lands every message on auto-resolve, verification, or escalate - with Jev only proposing toward caution and deterministic code owning facts, access, and the final call.
A browser extension where Jev answers 17 typed questions about every LinkedIn post - thirteen AI-tell questions plus four bait questions in the same request - and fixed weights turn the answers into a badge and a bait chip.
A Chrome extension finds ad-shaped DOM candidates and asks Jev whether each candidate is a paid advertisement before code removes it.
Jev Social pairs the decision model with socai, a CLI that drives your real Chrome across Instagram, TikTok, and LinkedIn: Jev chooses each next read-only operation - search, open a post or profile, read comments - and socai executes it.
FastGate fronts an English/Uzbek/Russian university helpdesk with four narrow Jev judgments per message plus per-passage grounding, routes deterministically in code, and ships an independent benchmark of Jev on low-resource languages.
A daily pipeline reads the BOE with one Jev call per provision, scoring impact, tagging topics, and selecting an original paragraph as the summary.
A working example of a shadow decision layer under a production CRM: one call returns intention, business division, human-handover probability, urgency, and lead quality - decided and logged, never sent.
An Android noise gate for notifications and SMS that asks Jev whether each message is an ad - and suppresses only what is explicitly flagged. Every uncertain path resolves to allow, because swallowing a verification code costs more than leaking an ad.
A light terminal Gmail client for Omarchy Linux: Jev files incoming mail into Gmail labels only above a confidence threshold, and grades your reply as you write - clarity, tone, length, next step, and which questions you haven't answered yet.
A Home Assistant add-on runs a crop-steering irrigation engine with Jev judging what fixed rules got wrong - ramp timing, probe trust, shot landing, EC moves, alerts - inside code-enforced physical limits; the README replays real Jev answers from 27 Sep 2026.
A quiet clerk for a small online store's inbox: a decision model reads every customer email first and sorts it into four lanes - needs a person now, can wait, template could answer, sales pitch - packaged as an n8n workflow with a shadow mode.
A Portuguese customer-service demo where each client message makes one Jev call with seven typed questions in parallel - intent, urgency, sentiment and more - and a cost table prices it against routing everything to a large LLM instead.
Fault triage for an industrial compressed-air unit with Jev: live telemetry in, classified alerts and tickets out, every judgment grounded in the unit's own service manual - read-only by design, with an eval harness that replays recorded decision cassettes.
A read-only Kubernetes controller that turns workload state and Events into stable incident decisions: deterministic rules decide the clear cases, and a Jev decision provider is called only for the ambiguous remainder.
Open-source alert triage where each production alert gets one Jev call with four typed questions - actionable, severity, team and paging - and auditable code turns the probabilities into routing; Jev never pages anyone, it only judges.
Paste a YouTube URL and every comment in the section is classified on four typed axes - then ranked so the ones actually worth a reply float to the top, for creators drowning in comment volume.
Measurement-first projects: routing strategies scored on labelled data, novels scored passage by passage, clinical reviews extracted as verbatim quotes, and chat interfaces that emit decisions instead of free text. These records are where the directory's evidence habits matter most, because the authors are usually testing a hypothesis rather than shipping a product.
Read the metric conditions closely in this cluster. Several studies report numbers from one dataset, one run, or one author-supplied baseline, and the record says so explicitly rather than averaging it away.
A finance RAG benchmark where Jev picks the passages: the agentic baseline reads whole SEC filings - 82,908 tokens for 46 of 50 right - while Jev-judged retrieval reads about 840 tokens and answers all 50; benchmark code and write-up are public.
A Turkish chat toy where whatever you type is answered by one of five fixed phrases - Jev picks which word, and two more answers decide the punctuation, in a single call.
A reproducible experiment beyond the 'optimal' Wordle solver: Jev scores how answer-like each of 12,972 accepted words is, the prior feeds entropy search over 1,925 days of NYT answers, and every claim is archived with pinned inputs and a reproduce script.
An experiment that makes a model which cannot write text answer anyway: for every word of the reply Jev picks 1 of 254 meaning-based word groups, then the word inside that group - and the README reports exactly where that stops working.
Lossless Rewrite closes the loop on AI shortening your report: your model rewrites, and Jev checks every protected idea survived - exact wording, meaning, or a reviewed checklist - then helps repair what went missing.
A research task treating Jev's answer distributions as classifier features: many small typed questions about an SVG, responses weighted and combined until the signal classifies the image - a study of whether decision calls can stand in for maths on pixels.
jevtrim is a comparative analysis of context compaction driven by calibrated judgments instead of summarization: Jev scores every chunk for relevance, ordinary Python keeps what fits the token budget, and the result is auditable and replayable offline.
Pipette reads arXiv, bioRxiv, medRxiv and 58 journals each morning and publishes a short, diverse daily edition: Jev labels and ranks, quoted sentences are the authors' own abstract lines, method and probabilities are public, and output is CC0 open data.
An evaluation of Jev on the SNIPS natural-language-understanding benchmark - intent detection and slot filling - asking how far a model that never generates text gets on a task normally solved by a trained tagger, using label names alone.
A browser app for systematic-review extraction: Jev never writes the answer - it points at line ids in trial reports and supplements, and code copies the quote out with its file, page, row, or slide, highlighted where it sits.
A live link-checker extracts each claim from a submitted page, asks Jev whether the cited excerpts support, contradict, or fail to establish it, and reports REAL or FAKE only when enough evidence agrees.
Osso fades the parts of a page that are not what the reader came for: each sentence of the main text gets a Jev probability, workspace sites are covered by default, and password fields, account pages and reviews are left alone by construction.
A bridge connecting images, video streams and RGB-D cameras to Jev's judgment engine: identify what matters in a frame, estimate risk, score a situation or judge many visible objects at once - typed answers over the visual world, 46 stars in its first day.
TraceDocs structures documents into source-linked blocks; Jev judges each with four Nouls - relevant, evidence, contradicts-premise, prompt-injection - returning a cited evidence set with a trace; the LLM writes from evidence and refuses when none exists.
Doom or Bloom maps where you stand between AI doom and bloom: a dynamic interview where the engine picks the next curated question by where your answers are thinnest, with Jev interpreting answers and scoring candidate follow-ups.
Four drop-in evaluators for Azure AI Foundry, rebuilt on Jev: Intent Resolution, Task Adherence, Tool Call Accuracy, Groundedness - each metric becomes small typed questions answered in one call, combined into a 1-5 score listing its checks.
Paper Radar reads all of arXiv so you read the few that matter: every new paper is judged against plain-English interests with calibrated per-interest probabilities - about six cents a day for everything, no pre-filtering.
A demonstration pushing the judgment-only model past its envelope: Jev 'writes' by answering which-word-comes-next in 250-word batches - each option shown as the whole reply so far plus the word - top-3 shortlist, final pick, until sentence end.
An independent study routes Jev confidence into a larger model on two labelled datasets and shows the winning settings do not transfer between them.
A weekly measured series pitting Jev against frontier LLMs on the same real workflow steps: week one routed inbound leads - Jev 90 percent correct at 366 milliseconds and four cents per thousand, against Sonnet 5's 78 percent at 2.6 seconds and three dollars.
Finding and extracting repeating patterns from noisy sequences with Jev: instead of a hand-tuned distance metric, typed questions decide what counts as the same pattern, and matches come back with probabilities instead of thresholds you guess.
A shared 1,000-by-1,000 emoji canvas where humans place strokes and Jev paints with them: after each stroke one typed call picks a contextually relevant emoji and where to put it, and a yes/no decides whether your stroke was finished.
A Japanese-language experiment scoring 100 labeled customer inquiries with the official TypeSafe SDK: one request per inquiry answers six questions at once - sentiment, emotion, anger intensity, urgency, churn risk, and sarcasm.
OCR flattens superscripts: a footnote star, an endnote number and a unit power land in the stream as look-alike tokens. A regex over-finds the suspects, then Jev classifies each - footnote, citation or unit - as a typed choice so formatting can be restored.
FastGate fronts an English/Uzbek/Russian university helpdesk with four narrow Jev judgments per message plus per-passage grounding, routes deterministically in code, and ships an independent benchmark of Jev on low-resource languages.
Jev Mart is a Japanese e-commerce demo where every operational decision - listings, inquiries, reviews, triage - is a typed Jev call and an LLM is never invoked; it ships with a twelve-chapter lecture set that teaches the pattern.
Book Aurora sends every ~90-word passage of Frankenstein to Jev as ten parallel score questions - nine emotions plus overall intensity - and draws each answer as one feathered row of a full-book aurora strip.
jevmeter turns a video into a BS meter: every sentence scored and every dodge flagged under a preset viewpoint - debates, earnings calls, podcasts, pitches - rendered as a 16:9 edit you can post.
JevSpeak holds a conversation without any generative model in the loop: Jev answers about thirteen parallel questions per turn, the decisions normalize into a semantic IR, and a deterministic compiler writes the sentence.
Spikecast puts Jev at the controls of a fly's simulated brain: the insect walks a road, and every turn, stop or dash is a typed decision with the synapse memory and neural firing visualised beside it - with docs separating what is real from what is modelled.
An interactive installation - explore a flooded observatory where glowing matter builds paths from your movement, gaze and actions, with real-time Jev decisions deciding how the world responds, shipped as a playable web app with live-decision tests.
A calibre plugin that suggests subject and genre tags for your books with Jev: suggestions arrive for review before anything is applied, existing tags are preserved, and the plugin is independent GPL software with a live demo and production notes.
A bilingual study asking whether JEV can serve as a world model for LLM agents - predicting what happens next from typed state questions - benchmarked against generative LLMs on identical predictions, with run manifests and a phase-0 API check archived.
A reader for public-domain Japanese novels where every paragraph is judged by Jev and each character accumulates a trait radar for the page in front of you, plus a running ranking of traits for the story so far.
For each topic you follow, jrp runs a pipeline: your questions steer the web search, Jev decides which findings count, and an LLM you choose writes the morning note - in English, Chinese or Japanese - with parallel stages and recorded claim-fidelity checks.
Supervision records: watching a coding agent work, gating risky tool calls, pruning stale tool history, and exposing typed judgments to MCP clients. The cluster treats the agent as the system under control and Jev as the referee that decides what the agent may do next.
All of these are community-published repositories verified at the artifact level. None of them claim production uptime, which is why each record's verification note matters more than its demo video.
An official cookbook combines exact quote matching with a Jev relation judgment to label citations verified, unsupported, contradicted, or fabricated.
A computer-use loop combines OCR and accessibility data, then asks Jev which bounded action should move the Mac toward a plain-English goal.
A Claude Code plugin and npm library asks Jev which old tool calls and results still matter, then drops or truncates the stale ones, so /compact replaces the lossy built-in summary with the kept messages verbatim.
A browser agent turns the visible DOM into an indexed action space and uses Jev to choose the next operation and compatible target.
A metasearch front end lets Jev choose the query, sources, and time range, then score every result for relevance while code fans out to engines.
A starter kit from Kinde: identity and permissions decide what an agent may do, and Jev judges each call in about 200 milliseconds - does it match the request, is it destructive, does it follow planted text, does it exfiltrate - before anything runs.
jevtrim is a comparative analysis of context compaction driven by calibrated judgments instead of summarization: Jev scores every chunk for relevance, ordinary Python keeps what fits the token budget, and the result is auditable and replayable offline.
Give jev-browser a task and a URL: Jev picks one action per step from the page's clickable, typeable, and selectable elements and scores goal-met and stuck likelihood, while code owns budgets, recovery, and stop gates.
A proof-of-concept Android loop stabilizes the screen, builds a short list of valid actions, and lets Jev pick one while code executes it.
bside pilots the Aside browser with Jev instead of a chat LLM: each tick answers which action, which element, and whether the goal is met, over an action schema the pilot cannot hallucinate outside of.
JDE, the Jev Decision Engine, wraps Jev as an MCP server: any MCP client - Claude Code, Cursor, your own agent - gets typed decision tools, with a policy layer, a decision ledger, and recorded evals comparing the hosted jev-1.13.0 against local alternatives.
An MCP server gives compatible agents tools for classification, scoring, checking, matching, screening, and custom typed Jev questions.
jevpipe pipes thousands of lines, files or records through one question and gets a typed judgment per item - grep-style filtering where Jev decides - shipped as a Rust binary on PyPI with an agent skill that teaches coding agents when to reach for it.
tenet makes agents fix rule violations before you ever see the diff: rules live in a YAML file in plain language, and Jev answers each one with a calibrated probability that becomes a pass-or-fail cutoff.
A guard for OpenClaw agents that reads an outgoing message and where it is going: Jev judges whether a client name, credential or internal hostname is about to reach the wrong readers, then confirms, blocks or rewrites per channel.
A Pi extension asks Jev to flag destructive, exfiltrating, or out-of-scope tool calls and to classify failures in command output.
Mina, a simulated 34-year-old librarian, runs on two systems: Jev reads her body, senses and clock every second and accumulates feelings; only when a feeling crosses its line does an LLM stop and think.
TraceDocs structures documents into source-linked blocks; Jev judges each with four Nouls - relevant, evidence, contradicts-premise, prompt-injection - returning a cited evidence set with a trace; the LLM writes from evidence and refuses when none exists.
A fuzzy linter that watches a coding agent write and speaks up 0.3 seconds later: Jev checks the file against your team's rules - race conditions in effects, missing cleanup, house style - so the agent fixes them before any human reviews the code.
Jev Trip is an explainable day-trip planner where the LLM plans ahead and Jev chooses and checks: scope choices keep one day in one city, per-place choices rank every candidate with probabilities, and code owns routes, times and validation.
A pluggable decision layer for the ego agent: System One (Jev) by default, swappable to local or other OpenAI-compatible backends, fail-closed guardrails - and the README's whole argument is that it is measured: 16 suites, 429 checks, rerun in full.
A Pi extension uses Jev to decide which old tool calls and results still matter while keeping conversation text verbatim.
The kamchatka terminal agent asks Jev to place each pending shell command on a three-level safety rubric - reads and reports, changes something reversibly, destroys or sends something out - drawn green, yellow, or red beside the permission prompt.
jev-axi is a CLI that puts a half-second Jev opinion in front of every command an agent runs: pick, rate, check, rank, triage, and guard, plus a PreToolUse hook that blocks risky Bash calls before they execute.
A working example of a shadow decision layer under a production CRM: one call returns intention, business division, human-handover probability, urgency, and lead quality - decided and logged, never sent.
OpenRecurSearch is an agentic web-search interface where nothing stays hidden: each research layer and each Jev decision appears live in the chat - option probabilities, score and latency included - while the markdown report streams into the side panel.
Foreman runs an independent observation loop that asks Jev whether a coding worker is progressing, stuck, complete, or ready for verification.
A bilingual study asking whether JEV can serve as a world model for LLM agents - predicting what happens next from typed state questions - benchmarked against generative LLMs on identical predictions, with run manifests and a phase-0 API check archived.
Emulator and game-state experiments — Mario, StarCraft, Pokémon, Terraria, shared stories — where Jev reads structured game state and picks the next move. Games are the lowest-risk place to demonstrate tight control loops, so the cluster works as a proving ground for latency and decision quality under continuous state change.
The evidence ceiling here is honest: these are demos and reproducible repos, not benchmarks. Their value is showing typed decisions holding up at interactive frame rates, which is hard to fake.
A reproducible harness lets Jev direct combat, exploration, and economy actions in the original StarCraft shareware campaign.
A battle harness reads FireRed state from RAM and lets Jev choose the next legal move or switch while ordinary code advances the fight.
A reproducible experiment beyond the 'optimal' Wordle solver: Jev scores how answer-like each of 12,972 accepted words is, the prior feeds entropy search over 1,925 days of NYT answers, and every claim is archived with pinned inputs and a reproduce script.
An autonomous bot that plays the Chrome T-Rex Runner to a thousand points by asking Jev which action each obstacle requires - jump, duck, or run - and letting a measured physics model decide exactly when to press.
A tModLoader mod whose boss fights are driven by Jev: every 200ms one request asks a nine-way intent, a five-band danger score, and whether to dash or jump; a per-frame reflex layer turns intents into keypresses.
A tiny Japanese game with no send button: type IT buzzwords and each keystroke gets a Jev judgment that stretches a meter, so 25 seconds of play answers what curl never does - how fast and how cheap the model feels inside a real app.
A Japanese companion game where one Jev pass decides the character's true feeling - Choice, affection Score, dislike Noul, topic Choice - in 0.2 to 0.5 seconds, so her face and a one-liner land before Claude-written dialogue and a Gemini TTS voice.
An npm library that lets game characters argue back: give a name, a persona and a goal, pass what the player typed, and Jev judges whether that character - with those values - was convinced; the same line can win a greedy merchant and offend an honest guard.
A generative painting instrument: select part of a sketch and Jev chooses its paint material - one typed Choice per region with probabilities and certainty bands - and the material flies in and paints itself; demo film and live gallery included.
A shared 1,000-by-1,000 emoji canvas where humans place strokes and Jev paints with them: after each stroke one typed call picks a contextually relevant emoji and where to put it, and a yes/no decides whether your stroke was finished.
Jev Yarn is a party game where everyone writes the next sentence and Jev picks the winner: one taste request scores every line on four dimensions, and Nouls handle room filters and whether the story feels finished.
Jev plays Yasuo in League of Legends through three decision heads - strategy once a second, tactics six to seven times a second while units are on screen, and build checks every twenty seconds - while code reads the game and executes.
A single-file Pac-Man that asks the System One endpoint for typed decisions per game tick, playable live without setup or locally with your own key pointed at the official endpoint.
An experimental controller translates NES telemetry into object-centric JSON and lets Jev choose the next legal controller macro.
Spikecast puts Jev at the controls of a fly's simulated brain: the insect walks a road, and every turn, stop or dash is a typed decision with the synapse memory and neural firing visualised beside it - with docs separating what is real from what is modelled.
You are the Town Crier of a 3D town: write one broadcast and all fifty citizens decide in parallel - investigate, join, flee, warn, or ignore - in a single batched choice request.
An interactive installation - explore a flooded observatory where glowing matter builds paths from your movement, gaze and actions, with real-time Jev decisions deciding how the world responds, shipped as a playable web app with live-decision tests.
Alert triage as UNIX filters, vulnerability text mapped to CVSS vectors, contract drift caught in OpenAPI prose, and PII redaction inside Postgres. The pattern across the cluster is narrow, checkable judgments inserted before a human or an agent acts on a security signal.
Security records carry an extra burden the directory cannot resolve: correctness of the judgment matters more than fluency, and no record here includes an independent audit. The limitations fields are worth reading before trusting any of these in a real pipeline.
A local cybersecurity lab replays synthetic telemetry and asks Jev for compromise probability, classification, severity, and an advisory response as evidence accumulates.
A starter kit from Kinde: identity and permissions decide what an agent may do, and Jev judges each call in about 200 milliseconds - does it match the request, is it destructive, does it follow planted text, does it exfiltrate - before anything runs.
OAS Sentinel compares two OpenAPI documents in two layers: deterministic checks find structural breaks, and Jev answers bounded semantic questions about changed prose - retries, ordering, pagination, error meaning - that schema diffs cannot see.
Paste a suspicious SMS, email, DM or listing into ScamCheck - web app, API or browser extension - and get a scam verdict, risk score, plain-English reasons and next steps, with all wording from the project's own templates rather than a model.
A local-first document filer: text is extracted locally, Jev decides category, confidentiality and prompt-injection risk as typed Choices, and low-confidence or suspicious files land in review lanes - never overwritten, with audit preview and undo.
A guard for OpenClaw agents that reads an outgoing message and where it is going: Jev judges whether a client name, credential or internal hostname is about to reach the wrong readers, then confirms, blocks or rewrites per channel.
Kassad brings calibrated guardrails to .NET: every prompt, completion, tool call and citation passes narrow typed checks answered by Jev, batched one round trip per stage, thresholded in code into Allow, Flag, Review, or Block.
jevmod scores every community message for spam, scam, harassment, NSFW, self-harm, doxxing, off-topic and custom plain-English rules; operators set thresholds, decisions log their numbers, and bots ship for Discord, Twitch, YouTube and Reddit.
A support inbox where Jev decides per span whether text is personal data and of what kind, and a redact() SQL function enforces the masking in Postgres by the viewer's clearance - content-aware, not pattern-based.
jev-axi is a CLI that puts a half-second Jev opinion in front of every command an agent runs: pick, rate, check, rank, triage, and guard, plus a PreToolUse hook that blocks risky Bash calls before they execute.
Chunk documents where meaning changes, not at a character count: Jev judges sentence continuation, planted instructions are quarantined before embedding, and retrieved passages classified as evidence, conflict or noise - on your existing vector database.
A single Rust binary pipes JSONL security alerts through five typed Jev questions and emits validated dispositions that code, not the model, enforces.
Python scripts that ask Jev to select every CVSS metric from a vulnerability description, then compute the numeric score in code exactly per the FIRST specification - for CVSS 3.0, 3.1, and 4.0.
Reranking by natural-language criteria, typed intent judgments for web search, graph traversal one hop at a time, and daily arXiv triage. These records use Jev as a ranking and routing layer over candidate sets that deterministic code assembles and filters.
Because ranking quality is easy to overclaim, the stronger records in this cluster describe their candidate pools and judging criteria explicitly. Weaker ones simply demonstrate the mechanism on a fixed example.
A metasearch front end lets Jev choose the query, sources, and time range, then score every result for relevance while code fans out to engines.
Pipette reads arXiv, bioRxiv, medRxiv and 58 journals each morning and publishes a short, diverse daily edition: Jev labels and ranks, quoted sentences are the authors' own abstract lines, method and probabilities are public, and output is CC0 open data.
A single-file Python portal that filters news and YouTube feeds by interests you describe in plain English: the official typesafe-sdk scores every item, routine business hides separately, and the result is one static page with News and YouTube tabs.
Paper Radar reads all of arXiv so you read the few that matter: every new paper is judged against plain-English interests with calibrated per-interest probabilities - about six cents a day for everything, no pre-filtering.
A Chrome extension for YouTube: ask the video a question in plain text, and Jev's typed judgments locate the moment that answers it - the player seeks straight there instead of you scrubbing.
At each node of a Neo4j graph the outgoing relationships become Choice options; Jev returns a full probability distribution over which one to follow, with a goal-reached Noul riding in the same call.
Product search inside PostgreSQL - typos, barcodes, typeahead, facet counts, all in SQL - with an optional second stage asking Jev two questions per search to rerank the shortlist; the README is itself the report, three failed versions included.
A Pinecone official examples repository: full-text search retrieves 200 candidates, and one Jev judgment pass reranks them to 10 by natural-language criteria, with a Claude baseline column for comparison.
Jev Social pairs the decision model with socai, a CLI that drives your real Chrome across Instagram, TikTok, and LinkedIn: Jev chooses each next read-only operation - search, open a post or profile, read comments - and socai executes it.
OpenRecurSearch is an agentic web-search interface where nothing stays hidden: each research layer and each Jev decision appears live in the chat - option probabilities, score and latency included - while the markdown report streams into the side panel.
Type a description and matching emoji fly out of the pile: one request carries the query as shared state with one score question per emoji, answered in parallel, and code ranks and floors the results.
Chunk documents where meaning changes, not at a character count: Jev judges sentence continuation, planted instructions are quarantined before embedding, and retrieved passages classified as evidence, conflict or noise - on your existing vector database.
For each topic you follow, jrp runs a pipeline: your questions steer the web search, Jev decides which findings count, and an LLM you choose writes the morning note - in English, Chinese or Japanese - with parallel stages and recorded claim-fidelity checks.
Home Assistant integrations, hydroponic control, and screen-time judging — small, high-variance domestic systems where a typed decision maps to one safe action. The official TypeSafe smart-home routing record sits here alongside community equivalents, which makes the cluster a natural side-by-side of official and community practice.
Household devices make failure visible, so authors tend to document fail-open behavior and manual overrides carefully. Both are recorded in the limitations of each entry.
TypeSafe's smart-home demo evaluates a request against many typed questions in parallel, then lets code use only the answers relevant to that request.
A Home Assistant integration exposes Jev probabilities, choices, and scores as entities and action responses that automations can use.
A Home Assistant conversation agent - installable via HACS - that decides with a TypeSafe System One model instead of an LLM: your spoken or typed commands become typed choices and Nouls that drive devices, with metered costs documented.
A Home Assistant add-on runs a crop-steering irrigation engine with Jev judging what fixed rules got wrong - ramp timing, probe trust, shot landing, EC moves, alerts - inside code-enforced physical limits; the README replays real Jev answers from 27 Sep 2026.
Qualm is a screen-time app for Apple silicon Macs, local-first, built on two judges: Kev running on your machine and TypeSafe's hosted Jev - typed choices decide what counts as a distracting session without your activity leaving the machine unless you opt in.
Android automation from the accessibility tree, ad silencing that fails open, aphasia assistance, and Flutter bindings. The cluster demonstrates that typed decisions work from semantic UI state on constrained devices, not just from desktop DOM dumps.
Every record here is community-published and artifact-verified against a specific device or emulator. Portability across Android versions is exactly the kind of claim none of them make, so none should be assumed.
A Flutter plugin that brings typed Jev calls to Dart apps: question builders for Noul and Choice with option-count validation, the systemone wire protocol on the native endpoint, and a Python-side test harness for the plugin's bridge.
An Android automation agent with a two-tier brain: Jev decides fast from the accessibility tree, a vision agent takes over only when the structural view is ambiguous, and ADB executes - cheap judge on the hot path, expensive one on the exceptions.
A proof-of-concept Android loop stabilizes the screen, builds a short list of valid actions, and lets Jev pick one while code executes it.
For people with aphasia who know what they mean but can't get the word out: describe it any way you can, and Jev picks the best guesses from a fixed 900-word list as big tap-to-hear picture tiles - never inventing a word.
An Android noise gate for notifications and SMS that asks Jev whether each message is an ad - and suppresses only what is explicitly flagged. Every uncertain path resolves to allow, because swallowing a verification code costs more than leaking an ad.
Composition steered by enum labels, songs written note by note, and piano improvisation from typed mood choices. A small cluster, but a useful one: creative domains show Jev making aesthetic judgments under explicit human direction rather than operational decisions under automation.
The records are demos and reproducible repos. They evidence that the decision interface works for creative control, not anything about musical quality, which no published metric here attempts to score.
Describe a mood - a slow, sad waltz - and a live piano plays it, with Jev deciding continuously as it goes: every musical choice is a typed question answered with probabilities, streamed to a public site in real time.
A DJ chatbot's pre-LLM router, built as a learning study: free-form /play requests - English, Spanish, typos, pasted lyrics, emojis - classified by Jev in one call (~500 ms, ~$0.000036) into a structured hint that tells the expensive LLM what it is handling.
A playground asks Jev to decide only closed-vocabulary labels for character, key, meter, and phrasing while deterministic code writes, engraves, and plays the notes.
Jev cannot write a single note, so code lays out the bars and the legal notes with musical facts attached, and Jev chooses: mode, tempo, form, every chord, every note - one call per decision, every call replayable.
Per-block market quotes, paper trading on typed market state, and one live crypto desk with Jev as the decision core. The cluster is small and the stakes are the highest in the directory, since a misjudged control decision maps directly to money.
Read these records with the strictest posture in the index: reported performance comes from single artifacts, market conditions are not controlled, and no record has been independently verified. They demonstrate architecture, not profitability.
A trading loop reads the Kuru MON-USDC order book and asks Jev for a buy-or-sell judgment before code places a post-only limit order.
A Next.js dashboard that streams live crypto data, computes fifteen moving averages and eleven oscillators into one typed market state, and lets a Jev agent trade a simulated 100k portfolio - paper only, no broker connected.
A multi-bot OKX trading desk with Jev as the brain and code as the body: each tick reads market and account state, computes features in TypeScript, asks Jev typed questions, then holds or places and cancels real orders - dashboard as the glass.
Drone tactics from camera-derived state, a robot arm driven through layered choices, and a hospital delivery robot steering past four hundred people. Physical control is where bounded judgments meet the least forgiving failure modes, and the cluster shows authors responding with layered choice hierarchies and conservative action sets.
All three records are community-published demonstrations. None reports long-run reliability on hardware, which is the number that would actually matter for deployment.
A simulated quadrotor uses Jev for low-frequency tactical judgments while classical vision, flight control, and safety reflexes remain in code.
Jev drives a LIBERO robot through 27 control inputs with layered decisions - intent, then motion family, then input - while reversible physics previews evaluate candidate effects locally before anything executes.
A hospital delivery robot judged about four times a second: code samples twelve candidate paths and removes predicted collisions, Jev picks one among the survivors, and a deterministic safety brake guarantees no contact - built at the TypeSafe JEVATHON.
Spain's official gazette screened every morning, LLM-generated requirements gated through twenty-two questions, and city permits found from free-text plans. Civic and legal data rewards exactness, so this cluster leans on extraction and classification judgments with verbatim outputs that code can re-check against the source.
The records demonstrate the pipeline pattern on public documents. None of them constitutes legal advice or a validated compliance tool, and each says so in its limitations.
Describe what you want to do in San Francisco and a rules engine works out which permits you need - Jev answers each rules question as a typed choice, falls back to asking you when unknown, and summarizes fees and deadlines per permit.
A requirements quality gate: one batched call asks 22 atomic questions across five MECE facets about an AI-generated requirement, and thresholds route it pass, human review, or reject - never rewriting, only judging.
A daily pipeline reads the BOE with one Jev call per provision, scoring impact, tagging topics, and selecting an original paragraph as the summary.
Add to the record
Send the artifact, what Jev decided, and any measured outcome. Nothing appears in the index before editorial review.
Submit evidence