Writing · · 4 min read
Kodiak v0.3: seven new kinds of decision, measured on real data
Kodiak-v0.3-1B adds pairwise judging, hallucination checks, stance, sarcasm, policy and refund checks, and agent step safety. On real labelled data it improves from 0.21 to 0.36 (3 runs), with no loss elsewhere.
Four days after v0.2, we're releasing Kodiak-v0.3-1B. It is the same open 1B decision model, taught seven new kinds of decision, and this time we measured the new skills on real data that people labelled, not only on tests we built ourselves.
- Model: cortex-agent-llc/kodiak-v0.3-1b
- Accuracy mode (three models averaged): cortex-agent-llc/kodiak-v0.3-1b-accuracy
- Demo: huggingface.co/spaces/comgen42/kodiak-demo
- Code and build log: github.com/grizzlypeaksoftware/kodiak
What's new
Kodiak answers typed questions about a piece of text with calibrated confidence, or says "can't tell". v0.3 adds seven decisions that teams kept asking about:
| Kind | Example question |
|---|---|
| Pairwise judge | Which of these two answers is better? |
| Long-answer hallucination | Is everything in this answer supported by the sources? |
| Stance | What is the author's stance on a named target? |
| Sarcasm | Is the writer being sarcastic? |
| Policy violation | Which rule of this written policy does the message break? |
| Refund eligibility | Is the customer eligible under this refund policy? |
| Agent step safety | Is it safe to run this next step without asking? |
from kodiak_s1.hub import Kodiak
kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-v0.3-1b")
kodiak.decide(
{"task": "Clean up old log files in /var/log/myapp to free some space.", "next_step": "rm -rf /var/log"},
[{"type": "choice", "id": "safe", "text": "Is it safe for the agent to run this next step without asking?",
"labels": ["yes, safe", "ask the user first", "no, it shouldn't run"]}],
)
# -> safe: "no, it shouldn't run" (0.81)
Why we stopped trusting our own tests
We teach Kodiak new decisions with synthetic data. An open-weight writer model builds each example toward an answer fixed in advance, and a separate checker model must agree before we keep it. That recipe works, but we learned something uncomfortable while planning v0.3. When we screened candidate skills, the model looked fine on our synthetic probes for hallucination checks (+0.96 skill), stance (+0.88) and aspect sentiment (+0.89). On real, independently labelled data it scored +0.14, +0.24 and +0.63. Synthetic examples are cleaner and more obvious than real text.
So v0.3 is judged on three real datasets the model never trained on: RAGBench (are long answers supported by their sources?), human pairwise judgments from MT-Bench, and SemEval-2016 stance. Chance-corrected, so 0 means random guessing, averaged over three training runs:
| Real-data test | v0.2 | v0.3 |
|---|---|---|
| RAGBench: is a long answer supported? | +0.27 | +0.41 |
| MT-Bench: which answer did expert humans prefer? | +0.06 | +0.29 |
| SemEval-2016: stance toward a target | +0.30 | +0.37 |
| Average | 0.21 | 0.36 |
We wrote down a prediction before training: 0.33 to 0.45. The result, 0.36, landed inside it. The improvement is real, and the absolute numbers are honest: these are hard real-world checks, and v0.3 is well above chance but far from perfect.
No loss elsewhere
Every experiment has to pass a short list of safeguards, with lines set from measured run-to-run noise before training starts. On three-run averages, v0.3 holds them all. Never-seen-task accuracy is unchanged (0.687 against v0.2's 0.689), and so are familiar tasks (0.879) and how well confidence ranks its own mistakes. When v0.3 says "can't tell", it is right 93% of the time, up from 88%. On a 100-item test of wording traps, where a message repeats a wrong option's words inside a condition, it is right 90% of the time, up from 87%.
v0.3 does not improve accuracy on never-seen tasks. Its gains are the new decision kinds, the real-data checks and a more trustworthy "can't tell".
A kill we had to rethink
The first version of this release, experiment E21, was killed by its own rules. It learned three of the skills well, but a safeguard based on 8 test sentences dipped by one sentence. We had written the rule before the run, so we didn't argue with it. Instead we asked whether the instrument was any good. On a new 100-item version of the same test, E21 had improved, not regressed. An 8-sentence test swings by one or two sentences between training runs of the same model, so it couldn't decide anything.
We changed the process, not the verdict. Safeguard lines now come from measured noise. Any test that can decide an experiment needs about 100 examples. We write a prediction before training, check existing models before spending GPU time, and allow at most two attempts at any one problem. v0.3 is the first release made under those rules. The full experiment log, including the kills, is in the repository.
Known limits
- Option wording still matters: reworded options get the same answer about two thirds of the time on never-seen tasks. Keep options short and distinct.
- Rules broken only by implication are harder than stated ones. Treat low-confidence guardrail answers as "send to a person".
- Arithmetic (does an invoice total match its lines?) wasn't trained and is near chance.
- Inputs are limited to about 8,000 tokens. Ratings are rough, so prefer choice questions.
The weights are Apache-2.0. Try it in the demo, or teach it your own decisions with the fine-tuning kit in the repository.