Small, open models that live on your machine.
Resident Labs builds compact models you can download, run and own. Our first one makes decisions: thousands of them, calibrated, on one GPU.
tılde
The decision model.
Give it a state (a document, a chat, a web page, a tool result, an agent's trajectory) and any number of typed questions about it. It reads the state once and answers every question in a single forward pass, each with a calibrated probability distribution. No text is generated, so there is nothing to parse.
illustrative values · click an example
Not a chatbot. A decision.
Ask a chat model to triage a ticket and you get a paragraph to parse. Ask Tilde and you get numbers you can put a threshold on.
“Based on the message, this looks like it's probably a billing issue, although support could also handle it. The customer seems fairly urgent, maybe a 4 out of 5? It doesn't appear to be spam, though I can't be completely certain without more context.”
Now write a parser. And a retry. And hope “fairly” means the same thing tomorrow.
// POST /v1/systemone → answers (illustrative values) { "spam": { "noul": 0.02, "confidence": 0.96 }, "team": { "choice": "billing", "probabilities": { "billing": 0.91, "support": 0.08, "sales": 0.01 } }, "urgency": { "score": 3.06, "probabilities": { "0": .02, "1": .04, "2": .07, "3": .60, "4": .27 } } }
if answers.team.probabilities.billing > 0.85: route("billing")
Three kinds of question
Classify, gate, triage, rank and route, on one 24 GB GPU, with up to 32k tokens per state + question and 64k per request.
choiceA probability for each of up to 255 named options.
noulThe probability that a yes/no statement about the state is true.
scoreA distribution over 2 to 10 described levels.
What people build with it
Agent guardrails
Gate every tool call: is it safe, is it what the user asked, does a human need to look?
Triage & routing
Tickets, emails and requests sorted by team, intent and urgency before anyone opens them.
Grounding checks
Is this RAG answer supported by the retrieved passages? A probability, not a vibe.
Evals at scale
Judge thousands of model outputs per minute on your own card, with no per-token bill.
Moderation
Spam, phishing, prompt injection and policy checks with thresholds you control.
Ranking
Score candidates against a rubric and sort by the full distribution, not a single guess.
Tilde vs Jev, head to head.
Same questions, same scorer. Ten public benchmarks, sorted by margin.
Show as a table
Jev is jev-1.13.0, queried through its API on identical questions and scored by the same code. Decision Index is an aggregate, shown separately.
Quickstart
Tilde speaks the /v1/systemone protocol. Send a state and your questions; get a distribution back for each.
curl -s http://localhost:18796/v1/systemone -H 'Content-Type: application/json' -d '{ "state": "Charged twice for my March invoice. Need the refund before Friday.", "questions": { "spam": {"type": "noul", "instructions": "Is this spam?"}, "team": {"type": "choice", "instructions": "Which team should handle it?", "criteria": {"billing": null, "support": null, "sales": null}}, "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["not urgent", "low", "medium", "high", "critical"]} }}'