Hands-on · ~15 minutes · Python

Build your first Jev decision loop

Every project in this index — 147 of them — is the same machine at different scales: code observes a state, asks Jev a narrow typed question, and acts on the calibrated answer. This tutorial builds the smallest useful version of that machine: a support-inbox triage loop that decides what to do with each incoming message, using the real System One API. No chat, no prompt engineering — three questions, two thresholds, one rule you can read.

Published 2026-10-08 · API details checked against TypeSafe's official docs

Diagram of the decision loop: code observes state and asks typed questions, Jev answers in parallel, code gates on confidence — high confidence executes, low confidence escalates to a human or text model — then the loop repeats
The loop you are about to build. Everything on the left is your code; everything on the right is one API call. The escalation path is not optional — it is what makes the loop safe to leave running.

Setup

You need a TypeSafe API key (the console playground at console.typesafe.ai/playground lets you try questions without code first) and the official SDK:

pip install typesafe-sdk
export TYPE_SAFE_API_KEY="your-key-here"

The client reads the key from the environment. All calls in this tutorial go to api.typesafe.ai/v1/systemone under the hood, with jev-latest as the model.

The three question types, in one breath

Jev answers three kinds of typed questions, and they can all ride in a single call: a Choice picks one option from a closed list your code provides; a Score places something on a small ordinal scale you define; a Noul answers one yes-or-no criterion with a calibrated probability. Each answer carries a confidence — the docs describe it as calibration about the answer's correctness, not enthusiasm — and Choice and Score answers carry the full probability distribution across options. TypeSafe's own guidance is to keep questions atomic: one gut-check judgment each, with multi-factor decisions composed in your code.

One state, three questions, one call

Our loop triages a support inbox. The state is the message plus a little account context; the questions are: what kind of request is this (Choice), how urgent is it (Score), and can an auto-reply handle it (Noul). One API call answers all three in parallel — adding questions to a call barely changes response time, which is the economic fact most of this index is built on.

from typesafe_sdk import TypeSafeClient, Choice, Score, Noul

client = TypeSafeClient()

def triage(message: str, plan: str, account_age_days: int) -> dict:
    state = {
        "message": message,
        "customer_plan": plan,
        "account_age_days": account_age_days,
    }
    result = client.system_one(
        state=state,
        model="jev-latest",
        questions={
            "kind": Choice(
                instructions="What kind of request is this message?",
                criteria={
                    "bug": 1,
                    "billing": 2,
                    "feature_request": 3,
                    "meeting_request": 4,
                    "other": 5,
                },
            ),
            "urgency": Score(
                instructions="How urgent is a correct response?",
                criteria={
                    "0": "no time pressure",
                    "1": "this week is fine",
                    "2": "today",
                    "3": "customer is blocked right now",
                },
            ),
            "auto_reply_ok": Noul(
                instructions="Could a static help-page reply handle this adequately?",
            ),
        },
    )
    return result["answers"]

The answer comes back as typed structures, not prose to parse — roughly {kind: {choice, confidence, probabilities}, urgency: {score, ...}, auto_reply_ok: {noul, ...}} with a usage block alongside. If you prefer raw HTTP, the same call is a POST to https://api.typesafe.ai/v1/systemone with a bearer key and a JSON body of state, model, and questions.

Gate on confidence, not on vibes

The step that separates a decision loop from a demo is the gate. Confidence answers are calibrated estimates, not guarantees — and the only defensible policy is: act on high confidence, escalate the rest. Two gates work well in practice, and both appear all over this index: a floor on the winning answer's confidence, and a margin between the top two options.

CONFIDENT = 0.75
MARGIN = 0.20

def route(answers: dict) -> str:
    kind = answers["kind"]
    ranked = sorted(kind["probabilities"].items(),
                    key=lambda kv: kv[1], reverse=True)
    (top, top_p), (second, second_p) = ranked[0], ranked[1]

    if kind["confidence"] < CONFIDENT or top_p - second_p < MARGIN:
        return "human"          # escalate, never guess
    if answers["urgency"]["score"] in ("2", "3") and not answers["auto_reply_ok"]["noul"]:
        return f"page_oncall:{top}"
    if answers["auto_reply_ok"]["noul"]:
        return f"auto_reply:{top}"
    return f"queue:{top}"

Read that routing function carefully: it is ordinary code, it is testable with plain unit tests, and nobody can change the policy by editing a prompt. That is the whole philosophy of this index in one function — the model supplies calibrated judgments, and your code owns the policy for what those judgments may cause.

Three habits from 147 shipped loops

First, enumerate only legal options. The strongest records here — the Pokémon battle harness, the merge-conflict picker — hand the model a list where every option is already valid, so an illegal action is not merely discouraged but unrepresentable.

Second, keep questions atomic and combine in code. The official docs say it and the corpus obeys it: one gut-check per question, multi-factor decisions composed by your own formula. If a question contains the word “and”, it is probably two questions.

Third, log every judgment. Records that survived contact with real usage — the trading desk, the security filter — all persist the full request, answer, probabilities, and confidence per decision. It is your audit trail, your calibration dataset, and your regression suite, and it costs one JSON line per decision.

Where to steal patterns next

Once the basic loop runs, the index has production-grade variations to borrow from: a Rust filter that runs five typed questions per alert, a daily arXiv judge that costs about six cents a day, and the official speculative fan-out router that answers every plausible question in one call and discards what it does not need. TypeSafe's patterns documentation covers the same moves from the vendor side. And when your loop produces something public, submit it — the index grows by artifact, not by claim.