I'm new to this
Read Welcome for the five-minute mental model, then run the quickstart: a calibrated answer from a model on your own machine.
A System One model answers typed questions about your program’s state: yes or no, pick one, rate on a scale. It returns a probability for every possible answer, in one pass, without generating text. You get the fast, cheap, testable reflex that sits in front of slower reasoning.
This site is where you learn to use one well. Every code sample runs against open models you can host yourself, such as Kenning, and the same code works with hosted engines that speak the same format.
from systemone import Client, Noul
client = Client("http://localhost:8093") # Kenning, running locally
r = client.system_one( state={"request": {"path": request.path, "query": request.query, "ip": request.ip}}, questions={"sqli": Noul("Is this request a SQL injection attempt?")},)
p = r.nouls["sqli"].noul # probability of "yes"if p >= 0.95: firewall.block(request.ip) # sure: actelif p > 0.05: route_to_analyst(request, p) # unsure: a person decidesThere’s no parser and no retry loop, and the answer can’t come back malformed. The thresholds are yours, and you set them from measured data.
res = llm.chat(messages=[ {"role": "system", "content": "You are a security classifier. ONLY RETURN VALID JSON..."}, {"role": "user", "content": raw_request},])try: decision = json.loads(strip_code_fences(res.text))except json.JSONDecodeError: decision = retry_with_backoff(raw_request) # ...and hope round two parses
# decision["confidence"] is a number the model wrote down, not a measured probability.It takes seconds, needs a parser and a retry loop, and returns a confidence that was never calibrated.
I'm new to this
Read Welcome for the five-minute mental model, then run the quickstart: a calibrated answer from a model on your own machine.
I have a latency or cost problem
An LLM in your hot path costs seconds and tokens. Read the fast loop and the fuzzy if-statement.
I have a reliability problem
Parsing model output is your biggest source of incidents. See what “no hallucination” means and calibrated confidence.
I want to build something
Ship an end-to-end system: phishing and alert triage, a data pipeline, or a moderation queue.
Kenning
An open System One model: Apache-2.0 weights, a ~2 GB GPU footprint, deterministic, with its training data and licences documented and its benchmarks published. On Hugging Face →
SystemOne Builder
Serve, train, distil and benchmark your own System One models on one GPU, from a dashboard or the command line. On GitHub →
The systemone client
One Python client for Kenning, Cloudflare’s Clef and TypeSafe’s Jev: same questions, same answers, switch with a URL. Engines →
SystemOne.dev is community-run and engine-neutral. It is built by the maintainers of the open SystemOne Builder and Kenning, and covers every engine that speaks the System One format. It is not affiliated with TypeSafe AI or Cloudflare.