Kodiak · open decision model · v0.2

Decisions,
not tokens.

Most of what software asks an LLM to do isn’t writing. It’s deciding: which intent, how urgent, which tool, safe or not. Kodiak answers those questions directly. You send it typed questions and get back calibrated answers in one forward pass, with nothing to parse.

View on GitHub Join the API waitlist

kodiak.decide()
state

Customer: My card was charged twice for order #4411 and I need it fixed before Friday.

  1. intent · What does the customer want?choice · p_null 0.02
    refund or billing fix
    94%
    order status
    3%
    cancel order
    2%
    technical support
    1%
  2. urgency · How urgent is this?score · p_null 0.03
    no time pressure0.78 [0.57, 0.93]immediate
  3. card_brand · Which card brand?choice · p_null 0.91
    ∅ can’t tell from this stateabstain · 91%
1 forward pass · 0 tokens generated Illustrative output, not a benchmark
Ticket routingTriageTool selectionGuardrails Policy checksPrompt-injection screeningLead scoringEscalation Ticket routingTriageTool selectionGuardrails Policy checksPrompt-injection screeningLead scoringEscalation

01 — The model

A small encoder that does one job very deliberately.

Kodiak is built on Ettin-encoder-1B, an open encoder (ModernBERT architecture) that reads text but can’t write it. On top of that we add a decision layer: a packed request format, a structured attention mask, and choice, score and abstain heads. Every answer stays inside the answer space you defined.

requestPOST /v1/decide
{
  "state": ["Customer: My card was charged twice
    for order #4411 and I need it fixed
    before Friday."],
  "questions": [
    { "id": "intent", "type": "choice",
      "labels": ["refund or billing fix",
                 "order status", "cancel order"] },
    { "id": "urgency", "type": "score",
      "min": 0, "max": 1 },
    { "id": "card_brand", "type": "choice",
      "labels": ["Visa", "Mastercard", "Amex"] }
  ]
}
responseillustrative values
{
  "intent": {
    "answer": "refund or billing fix",
    "confidence": 0.94, "p_null": 0.02 },
  "urgency": {
    "answer": 0.78,
    "interval": [0.57, 0.93] },
  "card_brand": {
    "answer": null,
    "abstain_reason": "unanswerable",
    "p_null": 0.91 }
}
  1. i

    Nothing to parse

    Kodiak never generates text. A choice is always one of your labels, and a score is always inside your range. No retries for broken JSON.

  2. ii

    Labels you invent today

    Choice labels are defined per request, so a new label set works without retraining. Short, distinct labels work best.

  3. iii

    An honest “can’t tell”

    Any question can come back as unanswerable from this state, with its own probability, so the model doesn’t have to guess.

  4. iv

    Confidence that means something

    Trained with proper scoring rules and evaluated on calibration (ECE, Brier), not only accuracy. When it says 90%, that’s meant to be true.

  5. v

    Small, fast, cheap

    About 38 ms per decision on a GPU, and it also runs on a CPU (a few hundred milliseconds), with no per-token bill.

  6. vi

    Open, Apache-2.0

    Built in the open with every decision and dead end documented. Download the weights on Hugging Face and run it yourself. The hosted API is coming.

02 — Results

Kodiak-v0.2-1B, measured.

An open 1B decision model that matches an 8B LLM on decisions it has never been trained for, about 40× faster.

Metric Kodiak-v0.2-1B v0.2 accuracy mode Qwen3-8B LLM, our own baseline run
Never-seen tasks, accuracy0.689 ± 0.0080.7060.688
Calibration error, never-seen lower is better0.0850.0620.293
When it says “can’t tell”, it’s right0.880.905–
Speed GPU, one request~38 ms~3× the single model~1,500 ms

Frozen eval set v0.2, choice questions. ‘Never-seen’ = tasks and label sets the model was never trained on. Kodiak-v0.2-1B figures with ± are the mean of three training runs. Accuracy mode averages three models.

New in v0.2

  • Checks whether an answer is grounded in its source.
  • Picks an assistant’s next step from API specs (call a tool, ask for missing info, or answer directly).
  • Verifies claims against a text.
  • Judges product relevance.

Known limits

Kodiak has limits worth knowing before you rely on it, including sensitivity to how labels are worded. Read the known limits on the model card ↗

Weights and code are Apache-2.0; the base model is MIT.

03 — How teams use it

Automate the sure ones. Hand off the rest.

Calibrated confidence is what makes this work. Choose a threshold. Decisions above it run automatically, and the ones Kodiak isn’t sure about go to a person or a larger LLM. Drag the gate to see the trade-off.

Simulated queue of 60 decisions, for illustration.

Automated
0
Escalated
0
Likely errors
0

04 — Services

We also build the systems around the model.

Cortex Agent LLC is a small, senior shop. We ship automation that plugs into what you already run, and we tell you when you don’t need AI at all.

Automation

AI workflow automation

Find the repetitive read-and-decide steps in your operation and automate them end to end, with human review where it counts.

Integration

Systems & API integration

Connect models and agents to your CRM, ticketing, data stores and internal APIs. We also build MCP servers so your agents can use real tools.

Kodiak

Custom decision models

Teach Kodiak your team’s labels and policies, then run it on your own infrastructure or ours. Your data stays yours.

Advisory

Architecture & evaluation

Model selection, eval design, cost modelling and the hard question of what should be automated in the first place.

05 — Get in touch

The hosted Kodiak API is coming.

Kodiak-v0.2-1B is open and free to download. The hosted API is coming. If you have a decision-heavy workflow and want early access, or you want help automating one now, send a note.

What’s this about?

Prefer email? [email protected] · Kodiak on GitHub