Reflex

Resources · the education page

One family. Three names.
The same idea, three rungs deep.

Reflex is the product you install. Instinct is the idea that upgrades it. Rethink is that idea, served from our GPU hosts.

Numbers law: every measured claim on this page is a link into the benchmark — none is typed.

New here? Start with this

How Jev works — and how Reflex goes further

A two-minute primer on decision models, the wire they speak, the numbers that judge them — and where a free local engine changes the picture.

What is a decision model?

Think of a triage nurse with a clipboard. You describe the situation; the clipboard has a few fixed questions — how urgent?, which ward?, needs a specialist, yes or no? The nurse does not write you an essay. They tick one box per question and tell you how sure they are.

A decision model is that nurse, as software: it takes a state (the situation, as text or structured data) plus a set of typed questions, and returns one typed answer per question with a probability for every option — never free text. Jev, by TypeSafe AI, is the best-known example: a hosted model you call over the internet. The request shape it uses — the Jev wire — has become a common language: Cloudflare's open Clef models speak it too, and Reflex speaks the same vocabulary.

decision model
Software that answers typed questions about a situation with probabilities, instead of writing free text.
Jev
TypeSafe AI's hosted decision model, and the request shape (the Jev wire) many decision models now share.
state
The situation under decision — a paragraph of text, or a JSON object.
questions
What you want decided, each with an id and a type. The answer space is part of the request, not baked into a model.
noul
A typed yes/no question; the answer is the probability of “yes”.
choice
Pick one option from a list you supply.
score
Pick a level on an ordered scale you supply (lowest first), e.g. routine → critical.
criteria
The options themselves: a map of option → description for choice, a list of levels for score. noul needs none.
prefill
A neural model reading the whole input in one forward pass, before it produces anything.
autoregressive
Generating an answer one token at a time, each step feeding the next — how chatbots write. Slow for decisions, and the text still has to be parsed.
non-autoregressive
Scoring every option in parallel straight from the prefill — no text is generated. This is what decision models like Clef do, per Cloudflare.
calibration
Whether stated confidence matches reality: of all answers given “eight in ten” confidence, about eight in ten should be right.
ECE
Expected calibration error — the average gap between confidence and accuracy, measured in confidence buckets. Lower is better; zero is perfectly calibrated.
Brier score
The mean squared distance between the predicted probabilities and the true answer. Lower is better; it punishes being confidently wrong.
abstain
Saying I'm not sure instead of guessing. In Reflex it is a designed output: the full probability distribution still comes back with it.
coverage
The share of questions a system actually answered (did not abstain on, did not error on).
chance-corrected skill
(score − chance) / (1 − chance): zero means no better than random guessing on that task, one means perfect. It lets a two-option task and a many-option task share one scale.
macro-F1
Accuracy-like score computed per label and then averaged, so rare labels count as much as common ones.
JDI
The Jev Decision Index — a community leaderboard that scores Jev-wire models on one frozen suite of typed-decision benchmarks: chance-corrected, with an unanswered question counted as wrong, and calibration published per entrant.
More words used further down this page
modelless
No neural network and no training run — answers come from documents you author plus a few small numbers fitted from your labels.
lane
One way of answering the same questions — an engine plus its setup. The board compares lanes side by side.
suite
One benchmark task: a fixed set of questions with known answers.
selective accuracy
How often answers are right, counting only the questions the system chose to answer (abstentions left out). Read it beside coverage.
cc
Chance-corrected accuracy: zero is random guessing on that suite's option count, one is perfect, below zero is worse than guessing.
conformal floor
A simple, well-understood baseline for honest confidence (split conformal prediction). A calibration claim has to beat it.
record-only
Measured and published on the board, but not yet served to anyone.
tier-fallback
Where a product lane has no answer of its own on a suite, the cell shows what is actually served there — the cheaper tier's answer — marked ↩ in tables and ▲ on the radar.
encoder head
A small trained layer on top of a text encoder that turns the encoder's reading of the question into option probabilities.
vessel
The package a hosted head is built to ship in: hashed, signed and encrypted, so only the hosted reader can open it. The format is built but not yet the serving path — today's specialists are files checked against a pinned BLAKE3 digest, run on our servers.
rung
One step on the ladder: free local Reflex first; a hosted head only when the step below abstains.

The workflow

Inputs (a state + typed questions) → the decision model (reads it all, scores every option) → typed outputs (one answer per question, each with probabilities). Your code reads the answer and, if it is careful, the confidence.

One request, one response

A security-incident triage, in the Jev wire shape Cloudflare documents for Clef. Three questions, one of each type.

request
{
  "state": "Login failures spiked overnight; an admin account then created a new API key.",
  "questions": {
    "severity": {
      "type": "score",
      "instructions": "How severe is this incident?",
      "criteria": ["routine", "needs attention", "critical"]
    },
    "action": {
      "type": "choice",
      "instructions": "What should happen next?",
      "criteria": {
        "page_oncall": "Wake the on-call responder now",
        "open_ticket": "File a ticket for business hours",
        "ignore": "No action needed"
      }
    },
    "rotate_keys": {
      "type": "noul",
      "instructions": "Should every API key be rotated?"
    }
  }
}
response
{
  "answers": {
    "severity": {
      "type": "score",
      "probabilities": { "0": p₀, "1": p₁, "2": p₂ },
      "score": Σ i·pᵢ,
      "confidence": max(p₀, p₁, p₂)
    },
    "action": {
      "type": "choice",
      "choice": "page_oncall",
      "probabilities": {
        "page_oncall": q₁, "open_ticket": q₂, "ignore": q₃
      },
      "confidence": q₁
    },
    "rotate_keys": {
      "type": "noul",
      "noul": P(true)
    }
  }
}

The p and q values are placeholders, not results: each is a probability between zero and one, and one question's probabilities add up to one. A score answer keys its levels by position (lowest first) and reports the expected level; envelopes and extra fields (model, usage) are trimmed here, and field spellings vary slightly by server; Reflex's own spelling (a list of questions with kind, prompt and options) is in the integration skill.

How Reflex differs

Same questions, same typed answers — different place and different honesty. Jev is a hosted model: every decision is a network round-trip to someone else's GPU, and it always returns an answer. Reflex is a modelless engine (no neural network, no training run) on your own machine: it answers from a corpus you author, in microseconds, and when the evidence is thin it abstains instead of guessing. Only those abstained questions need to go anywhere else — to a hosted Rethink head, paid only where Reflex was not sure.

Two ways to answer the same typed question. Jev: your app sends the state and typed questions over a network round-trip to a hosted model, which always returns a typed answer with probabilities, even when unsure. Reflex: your app asks the same questions over loopback on your own machine; the modelless Reflex floor answers from your corpus in microseconds; if it is sure it returns a typed answer with calibrated confidence, if not it abstains with the full distribution attached and only then escalates to a hosted Rethink head, paid only for those questions
Swipe the figure sideways. Jev (top) always answers, from a hosted model. Reflex (bottom) answers locally when it is sure and abstains honestly when it is not — the abstain is what makes paying for a hosted head optional rather than constant. Not affiliated with TypeSafe AI or Cloudflare; their products are named only to compare.

Where it runs

Jev: their cloud. Reflex: localhost — nothing leaves your machine unless you escalate.

When it is unsure

Jev: answers anyway; thresholding is your job. Reflex: abstains, distribution attached, so your code can route the hard case.

How it is judged

Accuracy, chance-corrected skill and calibration — every lane, losses included, on the benchmark.

coming Clef is coming to the board. Cloudflare's Clef and Clef-flash are open decision models that speak the Jev wire. We are adding a Clef comparison lane to the benchmark, measured by the same harness as every other lane. No numbers here until that lane publishes — the Jev Decision Index is the community board in the meantime.

Try it

What is this family?

Two products and one idea. The products ship binaries; the idea is how they compose.

Reflex — the product

The modelless [no neural network — no training run] decision engine. Open source, MIT, installs on your machine, answers from documents you author. The free floor of the family.

Read the Reflex section

Instinct — the idea

When the free engine abstains, a specialist trained for exactly that domain scores the survivors — composed on top, never instead. Embodied by the open teaching lane, and one rung deeper in Rethink.

See it on top of Reflex

Rethink — the product

The same composition idea at the encoder tier, served HOSTED-ONLY [weights that never leave our servers — you call, we think] from our GPU hosts. Private by design.

Read the Rethink section
The Reflex decision flow: state plus a typed question (choice, score, or yes/no) is embedded, routed to a domain, options scored corpus-is-the-model, calibrated by a sigmoid gate, then answered in microseconds with confidence or an honest abstain — loopback only, thresholds fitted offline from your own labeled data
The free lane: Reflex's decision flow — the same figure as the home page (one asset, two readers). Write-up: decision_flow.md.
The Instinct composition flow in two bands: the top band is the question in — the same typed question runs Reflex's modelless engine first, and a confident fused-gate answer returns in microseconds with the specialist never paid; on abstain the bottom band takes over — top-k prune keeps the candidates, the locked per-domain specialist scores them, an H1 cascade or H2 prior fusion calibrates the pick, and the answer ships with a decision receipt
The composition: Instinct's flow — the same figure as the benchmark's Instinct section, in learning vocabulary instead of measurement vocabulary. Write-up: instinct_flow.md.

Where each lane runs.

Structural facts only — a shipped artifact, a deploy shape, a design law. The measured side of every claim is one click away.

Runs onReflex (product)Instinct (idea · open lane)Rethink (product)
Browser (wasm)✅ arena head + playground— (a server-side lane)— (hosted-only by law)
Mac / PC / Linux✅ release binaries✅ serve bin (CPU hosts)— (no public binary)
Mobilevia the browser (wasm)via a hosted APIvia a hosted API
Edge device (ESP32-class)not shipped — possible in principle, no tested surface today ³not shipped (a server-side lane)not shipped (GPU hosts only)
Self-hosted server✅ one static binary✅ container image (CPU tier)private deploy — our GPU hosts only
Hosted APIloopback-only by design ¹✅ demo vessels after the open✅ the product surface (HOSTED-ONLY)

¹ The engine listens on your machine and answers browser calls only from origins you allow — see how it works. ³ An edge lane exists in the wider workspace but serves a different product — the honest cell today is “not shipped” for every tier.

Reflex — the free floor.

Reflex is the family's free floor — a modelless decision engine that runs on your machine, answers from a corpus you author, and abstains by design when the evidence is thin. [Modelless: no neural network — no training run. Corpus: your documents — the corpus is the model. Abstain: the engine says “I don't know” instead of guessing.]

The Reflex decision flow: state plus a typed question (choice, score, or yes/no) is embedded, routed to a domain, options scored corpus-is-the-model, calibrated by a sigmoid gate, then answered in microseconds with confidence or an honest abstain — loopback only, thresholds fitted offline from your own labeled data
The Reflex decision flow — the same figure as the home page (one asset, two readers, no copy).
BEFORE · THE WORLD YOU KNOW

Knowledge lives in someone else's training run.

  1. Build: a training run produces gigabyte-scale weights you download and trust.
  2. Serve: a model forward pass — it always answers, right or wrong.
  3. Guard: you threshold its confidence yourself, and hope.
AFTER · CORPUS-IS-THE-MODEL

Knowledge lives in documents you wrote.

  1. Build: write docs per domain, label a small slice, refit — no GPU.
  2. Serve: microsecond-class answers, or an abstain when the question is off-corpus.
  3. Guard: the abstain is a designed output — the full distribution still rides with it.

Instinct — the idea on top.

Instinct is the idea that upgrades Reflex without replacing it: when the free engine abstains, a specialist trained for exactly that domain scores the survivors on top, never instead.

The fused gate [the confidence check that decides whether to consult a specialist] answers with confidence and the specialist is never paid. On an abstain, the survivors are pruned to the top candidates and a locked per-domain specialist scores them. The answer carries a receipt — an audit trail of what answered, from which lane, checksum-committed in both directions.

The idea is taught by the open lane gist-rs/riir-instinct — and a specialist serves only where it strictly beat the free engine on a frozen test read. A tie or a loss sells nothing; the free engine keeps answering.

What does an abstain give my code?

A typed answer — the full distribution rides with it, so your code decides what a no-answer is worth instead of catching an exception. That is exactly the signal Instinct composes on: the abstain is the moment a specialist earns its keep.

Why does a free engine need a specialist at all?

The corpus bounds it. Modelless answers are excellent inside the domain you authored and honest about the edge of it. A specialist trained for exactly that hard domain covers the cases the corpus cannot — paid only where it was needed, per the winner law.

Deep write-up

The lane's own explainer, mirrored on this site: resources.md (local mirror) · on GitHub.

Rethink — the same idea, one rung deeper.

Rethink is the same Instinct idea one rung deeper: where a bag specialist is too coarse, a trained encoder head thinks — served HOSTED-ONLY from our GPU servers, so the weights never leave controlled hardware. [HOSTED-ONLY: weights that never leave our servers — you call, we think.]

One rung deeper than the bag specialist: when the modelless engine abstains, hashed-bag specialists answer fast on CPU, and the harder cases rise to a trained encoder head served HOSTED-ONLY from our GPU hosts, so the weights never leave our servers
The Rethink flow — from question to answer, with the rung model drawn in. Write-up: resources.md (mirrored — the source repo is private).
RUNG BELOW · BAG SPECIALISTS (INSTINCT)

Hashed word-count features, scored on CPU.

  1. Features: the question becomes a bag of words — cheap and fast, but blind to word order and nuance.
  2. Think: microsecond-class, per domain.
  3. Where: your machine or any CPU host.
THIS RUNG · THE ENCODER (RETHINK)

A trained encoder head reads the question whole.

  1. Features: the full encoder substrate — context, order, nuance.
  2. Think: millisecond-class thoughts, on our GPU hosts.
  3. Guard: a HOSTED-ONLY vessel — hashed, signed and encrypted, unreadable without the hosted reader. Planned: the format is built; until the hosted lane opens, heads are checked against a pinned BLAKE3 digest.

The free floor always answers first. Rethink is paid only where the free engine's confidence check abstains — the same fused gate as Instinct, one tier deeper. Every answer carries a receipt, and the seated cells are published either way: the benchmark's Rethink rows are the measured side, live today.

Why hosted-only? Why not ship the weights?

Because the class is structural, not a policy. The heads ship inside a HOSTED-ONLY vessel, and a build without the hosted reader cannot open one at all — so the weights cannot land on hardware nobody controls, not by accident and not by option.

Can I run Rethink myself?

No — and that is the honest answer by design. Rethink ships its own serve binary to our own GPU hosts (private distribution, never a public download). The open lanes — Reflex and the Instinct teaching lane — are the ones you can run.

Where is the product surface?

The storefront is live at rethink.gist.rs — what the hosted tier does, how your data is handled, and the buy card, which carries its honest state (a waitlist until purchases open). The bench Rethink rows are the measured side of everything this section claims.

Build your own.

Three development flows, one per rung. The compact write-ups live in each lane's docs — mirrored on this site, with the source of truth on GitHub for the public lanes.

Reflex — author a lane.

The Reflex development loop: author rules and grammar (state and options as sentences), label a small calibration slice, refit thresholds, serve on loopback, decide with an answer or an honest abstain, ship the binary
Author your domain as closed-grammar sentences, label a small slice, refit, serve on loopback, ship the binary. Write-up: dev_flow.md (local mirror) · on GitHub.

Instinct — train a specialist.

The Instinct development build flow: a corpus is made from public data, a per-domain specialist is trained on it, a stratified holdout read applies the winner law so only a strict win over the free engine survives, the passing artifact is BLAKE3-locked, registered as a manifest row, and served with decision receipts
Make a corpus from public data, train one specialist per domain, apply the winner law on a stratified holdout, lock the winner, register it, serve with receipts. Write-up: dev_flow.md (local mirror) · on GitHub.

Rethink — mint a head.

The Rethink development flow: train the head on the trainer side, lock it, mint a HOSTED-ONLY vessel, deploy to our GPU hosts only, and serve decisions through the storefront
Train the head on the trainer side, lock it, mint a HOSTED-ONLY vessel, deploy to our GPU hosts only, serve through the storefront. No step ever places the weights on hardware nobody controls. Write-up: dev_flow.md (local mirror — the source repo is private).