Resident

Small, open models that live on your machine.

Resident Labs builds compact models you can download, run and own. Our first one makes decisions: thousands of them, calibrated, on one GPU.

Our first model

tılde

tilde-8b-a1bopen weightsMIT

The decision model.

Give it a state (a document, a chat, a web page, a tool result, an agent's trajectory) and any number of typed questions about it. It reads the state once and answers every question in a single forward pass, each with a calibrated probability distribution. No text is generated, so there is nothing to parse.

8B MoE1B active64K per request1 GPU
state
One forward pass · three answers

illustrative values · click an example

Not a chatbot. A decision.

Ask a chat model to triage a ticket and you get a paragraph to parse. Ask Tilde and you get numbers you can put a threshold on.

A chat modelgenerates text

“Based on the message, this looks like it's probably a billing issue, although support could also handle it. The customer seems fairly urgent, maybe a 4 out of 5? It doesn't appear to be spam, though I can't be completely certain without more context.”

Now write a parser. And a retry. And hope “fairly” means the same thing tomorrow.

Tildeone forward pass
// POST /v1/systemone → answers (illustrative values)
{
  "spam":    { "noul": 0.02, "confidence": 0.96 },
  "team":    { "choice": "billing",
               "probabilities": {
                 "billing": 0.91, "support": 0.08, "sales": 0.01 } },
  "urgency": { "score": 3.06,
               "probabilities": {
                 "0": .02, "1": .04, "2": .07, "3": .60, "4": .27 } }
}

if answers.team.probabilities.billing > 0.85: route("billing")

Three kinds of question

Classify, gate, triage, rank and route, on one 24 GB GPU, with up to 32k tokens per state + question and 64k per request.

choice

A probability for each of up to 255 named options.

noul

The probability that a yes/no statement about the state is true.

score

A distribution over 2 to 10 described levels.

What people build with it

01

Agent guardrails

Gate every tool call: is it safe, is it what the user asked, does a human need to look?

02

Triage & routing

Tickets, emails and requests sorted by team, intent and urgency before anyone opens them.

03

Grounding checks

Is this RAG answer supported by the retrieved passages? A probability, not a vibe.

04

Evals at scale

Judge thousands of model outputs per minute on your own card, with no per-token bill.

05

Moderation

Spam, phishing, prompt injection and policy checks with thresholds you control.

06

Ranking

Score candidates against a rubric and sort by the full distribution, not a single guess.

Tilde vs Jev, head to head.

Same questions, same scorer. Ten public benchmarks, sorted by margin.

TildeJev
accuracy, %
7of 10
benchmarks ahead of Jev
+16.4pts
CLadder, the widest margin
1GPU
runs on your own 24 GB card
Show as a table

Jev is jev-1.13.0, queried through its API on identical questions and scored by the same code. Decision Index is an aggregate, shown separately.

Quickstart

Tilde speaks the /v1/systemone protocol. Send a state and your questions; get a distribution back for each.

curl -s http://localhost:18796/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "Charged twice for my March invoice. Need the refund before Friday.",
  "questions": {
    "spam":    {"type": "noul", "instructions": "Is this spam?"},
    "team":    {"type": "choice", "instructions": "Which team should handle it?",
                "criteria": {"billing": null, "support": null, "sales": null}},
    "urgency": {"type": "score", "instructions": "How urgent is it?",
                "criteria": ["not urgent", "low", "medium", "high", "critical"]}
  }}'

Decisions that live on your machine.