Ecosystem report · computed from the public index
147 Jev records, one pattern
Six weeks after Jev's announcement, the public record of what it can do has grown into something you can count. This report counts it: all 147 artifacts in our index as of October 8, 2026 — where they cluster, which judgments people actually delegate to a System One model, the architecture almost all of them share, and the holes in the evidence that nobody has filled yet. Every figure below is computed from the public dataset this site publishes.
What the corpus is
The index holds 147 records captured between August 23 and October 7, 2026 — one August outlier, 134 in September as the ecosystem detonated, and 12 in the first week of October. Three are official TypeSafe examples; 144 are community projects. 144 of the 147 rest on public GitHub repositories, with the remainder demonstrated through hosted demos and cookbook examples. Every record in the index sits at evidence level E2: an editor inspected the artifact and it says what the record claims. None has yet been independently reproduced.
One property matters before any other statistic: 126 of the 147 are reproducible repositories — not screenshots, not vaporware, but code you can clone and run. Whatever else the Jev ecosystem is, six weeks in it is disproportionately a working-code ecosystem, which is unusual for a model launch and is the reason an index like this one can exist at all.
Where the projects cluster
Developer tools dominate: 63 records, 43% of the index, spanning coding-agent supervision, code-review routing, semantic code search, and citation checking. The closest reading of this cluster is that the people building with Jev first are the people who build tools for a living — and the judgment they reach for first is “which of these candidate items deserves attention.”
Automation (37) and agents (28) form the control half of the index: browsers driven from DOM tables, macOS driven from screen inventories, Home Assistant state turned into typed signals, coding agents gated by referee processes. Research and measurement (35) is the third pillar — scored novels, screened gazettes, judged paper feeds, labelled-data routing studies — and it is where the most honest metrics live, because measurement-minded authors write down their conditions. The long tail is genuinely long: games as proving grounds, security triage, mobile accessibility loops, a small but serious trading cluster, and single records in music, robotics, and legal tech that show the pattern reaching far outside software.
What people delegate: classification, overwhelmingly
The job tags tell a blunt story. Classification appears in 104 of 147 records — 71% — and the next jobs, routing (44) and scoring (44), are classifications wearing different hats. Extraction, the least common job at 9 records, is the one that most demands the model not improvise, which is exactly why practitioners keep it rare and gate it with verbatim-copying patterns.
Nobody in this corpus uses Jev to write. The primitives in play are Choice (117 records), Noul (107), and Score (55) — closed lists, calibrated yes-or-no positions, and small ordinal scales. Open-ended generation appears nowhere in the index, and when a project needs prose it summons a text model explicitly and keeps Jev on the decision boundary. Six weeks of public experimentation have converged on a division of labor: the language model drafts, Jev decides what the draft may claim, code does everything else.
The one architecture, and its four variations
Strip the domains away and nearly every record is the same machine: code observes a state and reduces it to a typed description; the description and a set of narrow questions go to Jev in one call; calibrated answers come back; code gates them against thresholds and executes. The four recurring variations are worth naming, because choosing among them is most of the integration work:
Speculative fan-out asks all plausible questions at once and lets code discard the irrelevant answers — the official smart-home router is the reference. Legal-action snapshots enumerate only what may be done before the model is consulted, so an illegal choice is unrepresentable — the Pokémon battle harness is the purest form. Verbatim extraction lets the model point and forces code to copy, so nothing can be paraphrased — the clinical-review and gazette tools. And layered authority wraps the judgment in faster deterministic loops that can veto it — the drone and hospital-robot records, where a reflex layer outranks the model.
What the numbers say — and what they are worth
Only 27 of 147 records publish any quantitative metric at all. The ones that do are striking: a full Elite Four Pokémon run for about three cents of model time; all of arXiv judged daily for roughly six cents; a 13-question regulatory briefing batched 12.2× cheaper than asking one question at a time; a Google Flights search completed in 7.1 seconds including page loads. Every one of these numbers is author-reported or vendor-reported, most are single runs, and not one has been independently reproduced. They are best read as cost demonstrations — evidence that the per-judgment economics of a System One model permit workloads that per-token generation pricing would forbid — rather than as performance benchmarks.
The holes nobody has filled
The corpus's gaps are as informative as its clusters. There are no published reliability numbers: no win rates over many seeds, no false-positive rates on labelled sets, no uptime telemetry from long-running deployments. Zero records have been editor-reproduced. Coverage skews Mac and Linux; Windows, the largest desktop installed base, is nearly absent. Production deployments with real users appear only as claims, because closed systems cannot be artifact-verified under this index's rules. And the security cluster, 13 records, is entirely demonstration-grade — nobody has published an adversarial evaluation.
These are the records we would most like to add. If you have run a multi-seed evaluation, measured a false-positive rate, or kept a Jev loop alive in production long enough to publish its numbers, the submission channel is open — and the reproduction program exists precisely to turn a single independent re-run into an upgrade: a record moved from artifact-verified to editor-reproduced, with your conditions and credit attached. A reproduction that upgrades an existing record to E3 is the single most valuable artifact the ecosystem is currently missing.