Knowledge lives in someone else's training run.
- Build: a training run produces gigabyte-scale weights you download and trust.
- Serve: a model forward pass — it always answers, right or wrong.
- Guard: you threshold its confidence yourself, and hope.
Resources · the education page
Reflex is the product you install. Instinct is the idea that upgrades it. Rethink is that idea, served from our GPU hosts.
New here? Start with this
A two-minute primer on decision models, the wire they speak, the numbers that judge them — and where a free local engine changes the picture.
Think of a triage nurse with a clipboard. You describe the situation; the clipboard has a few fixed questions — how urgent?, which ward?, needs a specialist, yes or no? The nurse does not write you an essay. They tick one box per question and tell you how sure they are.
A decision model is that nurse, as software: it takes a state (the situation, as text or structured data) plus a set of typed questions, and returns one typed answer per question with a probability for every option — never free text. Jev, by TypeSafe AI, is the best-known example: a hosted model you call over the internet. The request shape it uses — the Jev wire — has become a common language: Cloudflare's open Clef models speak it too, and Reflex speaks the same vocabulary.
choice, a list
of levels for score. noul needs none.(score − chance) / (1 − chance): zero means no better than
random guessing on that task, one means perfect. It lets a two-option task and a many-option task
share one scale.Inputs (a state + typed questions) → the decision model (reads it all, scores every option) → typed outputs (one answer per question, each with probabilities). Your code reads the answer and, if it is careful, the confidence.
A security-incident triage, in the Jev wire shape Cloudflare documents for Clef. Three questions, one of each type.
{
"state": "Login failures spiked overnight; an admin account then created a new API key.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is this incident?",
"criteria": ["routine", "needs attention", "critical"]
},
"action": {
"type": "choice",
"instructions": "What should happen next?",
"criteria": {
"page_oncall": "Wake the on-call responder now",
"open_ticket": "File a ticket for business hours",
"ignore": "No action needed"
}
},
"rotate_keys": {
"type": "noul",
"instructions": "Should every API key be rotated?"
}
}
}
{
"answers": {
"severity": {
"type": "score",
"probabilities": { "0": p₀, "1": p₁, "2": p₂ },
"score": Σ i·pᵢ,
"confidence": max(p₀, p₁, p₂)
},
"action": {
"type": "choice",
"choice": "page_oncall",
"probabilities": {
"page_oncall": q₁, "open_ticket": q₂, "ignore": q₃
},
"confidence": q₁
},
"rotate_keys": {
"type": "noul",
"noul": P(true)
}
}
}
The p and q values are placeholders, not results: each is a probability
between zero and one, and one question's probabilities add up to one. A score answer keys its
levels by position (lowest first) and reports the expected level; envelopes and extra fields
(model, usage) are trimmed here, and field spellings vary slightly by server; Reflex's own spelling (a list of questions with kind, prompt and
options) is in the integration skill.
Same questions, same typed answers — different place and different honesty. Jev is a hosted model: every decision is a network round-trip to someone else's GPU, and it always returns an answer. Reflex is a modelless engine (no neural network, no training run) on your own machine: it answers from a corpus you author, in microseconds, and when the evidence is thin it abstains instead of guessing. Only those abstained questions need to go anywhere else — to a hosted Rethink head, paid only where Reflex was not sure.
Jev: their cloud. Reflex: localhost — nothing
leaves your machine unless you escalate.
Jev: answers anyway; thresholding is your job. Reflex: abstains, distribution attached, so your code can route the hard case.
Accuracy, chance-corrected skill and calibration — every lane, losses included, on the benchmark.
coming Clef is coming to the board. Cloudflare's Clef and Clef-flash are open decision models that speak the Jev wire. We are adding a Clef comparison lane to the benchmark, measured by the same harness as every other lane. No numbers here until that lane publishes — the Jev Decision Index is the community board in the meantime.
Two products and one idea. The products ship binaries; the idea is how they compose.
The modelless [no neural network — no training run] decision engine. Open source, MIT, installs on your machine, answers from documents you author. The free floor of the family.
Read the Reflex sectionWhen the free engine abstains, a specialist trained for exactly that domain scores the survivors — composed on top, never instead. Embodied by the open teaching lane, and one rung deeper in Rethink.
See it on top of ReflexThe same composition idea at the encoder tier, served HOSTED-ONLY [weights that never leave our servers — you call, we think] from our GPU hosts. Private by design.
Read the Rethink sectionStructural facts only — a shipped artifact, a deploy shape, a design law. The measured side of every claim is one click away.
| Runs on | Reflex (product) | Instinct (idea · open lane) | Rethink (product) |
|---|---|---|---|
| Browser (wasm) | ✅ arena head + playground | — (a server-side lane) | — (hosted-only by law) |
| Mac / PC / Linux | ✅ release binaries | ✅ serve bin (CPU hosts) | — (no public binary) |
| Mobile | via the browser (wasm) | via a hosted API | via a hosted API |
| Edge device (ESP32-class) | not shipped — possible in principle, no tested surface today ³ | not shipped (a server-side lane) | not shipped (GPU hosts only) |
| Self-hosted server | ✅ one static binary | ✅ container image (CPU tier) | private deploy — our GPU hosts only |
| Hosted API | loopback-only by design ¹ | ✅ demo vessels after the open | ✅ the product surface (HOSTED-ONLY) |
¹ The engine listens on your machine and answers browser calls only from origins you allow — see how it works. ³ An edge lane exists in the wider workspace but serves a different product — the honest cell today is “not shipped” for every tier.
Reflex is the family's free floor — a modelless decision engine that runs on your machine, answers from a corpus you author, and abstains by design when the evidence is thin. [Modelless: no neural network — no training run. Corpus: your documents — the corpus is the model. Abstain: the engine says “I don't know” instead of guessing.]
Instinct is the idea that upgrades Reflex without replacing it: when the free engine abstains, a specialist trained for exactly that domain scores the survivors on top, never instead.
The fused gate [the confidence check that decides whether to consult a specialist] answers with confidence and the specialist is never paid. On an abstain, the survivors are pruned to the top candidates and a locked per-domain specialist scores them. The answer carries a receipt — an audit trail of what answered, from which lane, checksum-committed in both directions.
The idea is taught by the open lane gist-rs/riir-instinct — and a specialist serves only where it strictly beat the free engine on a frozen test read. A tie or a loss sells nothing; the free engine keeps answering.
A typed answer — the full distribution rides with it, so your code decides what a no-answer is worth instead of catching an exception. That is exactly the signal Instinct composes on: the abstain is the moment a specialist earns its keep.
The corpus bounds it. Modelless answers are excellent inside the domain you authored and honest about the edge of it. A specialist trained for exactly that hard domain covers the cases the corpus cannot — paid only where it was needed, per the winner law.
The lane's own explainer, mirrored on this site: resources.md (local mirror) · on GitHub.
Rethink is the same Instinct idea one rung deeper: where a bag specialist is too coarse, a trained encoder head thinks — served HOSTED-ONLY from our GPU servers, so the weights never leave controlled hardware. [HOSTED-ONLY: weights that never leave our servers — you call, we think.]
The free floor always answers first. Rethink is paid only where the free engine's confidence check abstains — the same fused gate as Instinct, one tier deeper. Every answer carries a receipt, and the seated cells are published either way: the benchmark's Rethink rows are the measured side, live today.
Because the class is structural, not a policy. The heads ship inside a HOSTED-ONLY vessel, and a build without the hosted reader cannot open one at all — so the weights cannot land on hardware nobody controls, not by accident and not by option.
No — and that is the honest answer by design. Rethink ships its own serve binary to our own GPU hosts (private distribution, never a public download). The open lanes — Reflex and the Instinct teaching lane — are the ones you can run.
The storefront is live at rethink.gist.rs — what the hosted tier does, how your data is handled, and the buy card, which carries its honest state (a waitlist until purchases open). The bench Rethink rows are the measured side of everything this section claims.
Three development flows, one per rung. The compact write-ups live in each lane's docs — mirrored on this site, with the source of truth on GitHub for the public lanes.