Editorial methodology · written by the editors
How records earn their place
Anyone can claim a Jev demo. Screenshots are cheap, Threads are fluent, and a forty-second video proves nothing about the hundredth run. This index exists to separate what has been publicly demonstrated from what has merely been said, and this page documents the exact system we use to make that separation — including the parts that are still weak.
Gate one: it starts from a public artifact
A record begins with something a reader can open: a repository, a deployed demo, a cookbook example, a recorded benchmark, or a public post that links to one of these. We store the source URL, who published it, and the date we captured it — because artifacts move, and a record that cannot trace its source is gossip.
The corollary is what gets rejected. A private screenshot is not an artifact. A login-only Discord message can lead us to a project, but it cannot carry a record on its own, and a claim whose only evidence is the author asking you to trust them stays out of the index. Of the sources behind the current 147 records, 144 are public GitHub repositories — not because we prefer repositories aesthetically, but because they are the artifact type a stranger can actually inspect.
Gate two: classify the source, and cap the official share
Every record is labeled Official — published by TypeSafe on a TypeSafe-owned site or in its GitHub organization — or Community — published by anyone else, even if TypeSafe later links to it. The distinction sounds bureaucratic until you see what it protects: an official demo is marketing-adjacent evidence, chosen by the vendor, while a community project is evidence the vendor did not select. A directory of only official examples would tell you what TypeSafe wants shown, not what the model can do.
We cap official records at 30% of the directory — currently three of 147 — and the cap is enforced in the validation schema, not in policy prose. If the collection drifts above it, the data itself fails validation and the build stops.
Gate three: set the evidence level, honestly
Three levels describe how far checking went. E1 means the author reported it and the artifact has not been inspected in depth. E2 means an editor opened the public artifact — repository, demo, or benchmark — and it matches what the record claims. E3, the strongest, means an editor reproduced the result independently and recorded the conditions.
Today, all 147 records sit at E2. Zero are E3. We publish that fact rather than dressing E2 up as something stronger, because the gap matters: E2 confirms the artifact exists and says what its author says it says. It does not confirm the metric inside the README. E3 reproduction is the system's main growth path — there is a standing reproduction program for it — and records that ship reproducible harnesses, which is most of this directory, are the candidates.
Gate four: limitations are mandatory, and metrics keep their conditions
Two rules govern what a record may say. First, every record must list at least one limitation — what the evidence does not prove. A record without a limitation field fails validation. This is the rule reviewers thank us for and submitters resent: “single run, one browser profile, not a benchmark” travels with the record everywhere it appears.
Second, a reported number never appears without the conditions its author attached. The 7.073-second Google Flights search stays labeled as one recorded run on one browser profile. The 12.2× batching cost saving stays labeled as a five-run vendor comparison over one long shared document. Model confidence is never treated as correctness. A number quoted without its conditions is exactly the kind of claim this index was built to prevent.
Gate five: the loop, not the line
Verification here is a maintenance schedule, not a checkpoint. Records carry their last-verified date and return to review as repositories move, demos expire, or authors publish better results — the currently verified set was re-checked across 21 distinct days between late August and early October 2026. Corrections and takedown requests have a standing public policy with a named channel; when a repo's claims change, the record changes with it or comes down.
Where the system falls short
A methodology page that only lists virtues is a brochure. The honest list: we have not yet reproduced anything ourselves (no E3), so every metric in the index is author-reported or vendor-reported by definition. Verification depth necessarily varies — an afternoon with a repository is not a security audit, and we do not audit security claims. Records carry our reading of a project at a moment in time; fast-moving repos can outrun their record between re-checks. And the system structurally favors projects that publish in public repositories, which under-represents closed-source production deployments — the most interesting evidence class, and the one we currently cannot verify at all.
If you spot a record that outruns its evidence, the corrections channel is the fastest way to get it fixed, and submissions of public artifacts are always open.