 

AI・27 min read・Posted on October 7, 2026

# Jev and Laya, tested on our own data: what held up

 

 

 

 

 



 

   Table of contents  - [Key takeaways](#key-takeaways)
- [Why I ran this on our own data](#why-i-ran-this-on-our-own-data)
- [The dataset: a Slack channel full of tiny decisions](#the-dataset-a-slack-channel-full-of-tiny-decisions)
- [Five decisions per message](#five-decisions-per-message)
- [Where the correct answers came from](#where-the-correct-answers-came-from)
- [What are Jev and Laya?](#what-are-jev-and-laya)
- [What is Jev?](#what-is-jev)
- [What is Laya?](#what-is-laya)
- [What independent testers found](#what-independent-testers-found)
- [The claims &amp; what held up](#the-claims-what-held-up)
- [Keeping colleagues' data private, &amp; where that falls short](#keeping-colleagues-data-private-where-that-falls-short)
- [What "de-identified" means here](#what-de-identified-means-here)
- [How I set up a fair benchmark](#how-i-set-up-a-fair-benchmark)
- [Split by time, not at random](#split-by-time-not-at-random)
- [A baseline that knows the posting time](#a-baseline-that-knows-the-posting-time)
- [Same examples for every model](#same-examples-for-every-model)
- [Running the models](#running-the-models)
- [Fine-tuning Laya](#fine-tuning-laya)
- [Hardware](#hardware)
- [Terms used in this post](#terms-used-in-this-post)
- [Result 1: with no labels, Jev is ready &amp; base Laya isn't](#result-1-with-no-labels-jev-is-ready-base-laya-isn-t)
- [Calibration fixed Laya's confidence, but not its 32.6% accuracy](#calibration-fixed-laya-s-confidence-but-not-its-32-6-accuracy)
- [Jev beat the regex with zero examples](#jev-beat-the-regex-with-zero-examples)
- [Why Jev loses to the regex on everyday messages](#why-jev-loses-to-the-regex-on-everyday-messages)
- [Result 2: fine-tuned, Laya needs about half the labels](#result-2-fine-tuned-laya-needs-about-half-the-labels)
- [Where Laya's lead actually comes from](#where-laya-s-lead-actually-comes-from)
- [Result 3: only TF-IDF knows when it is wrong](#result-3-only-tf-idf-knows-when-it-is-wrong)
- [Good average calibration and a good gate are different tests](#good-average-calibration-and-a-good-gate-are-different-tests)
- [Nine of Laya's confident mistakes are debatable labels](#nine-of-laya-s-confident-mistakes-are-debatable-labels)
- [What broke along the way](#what-broke-along-the-way)
- [The fp16 GradScaler bug](#the-fp16-gradscaler-bug)
- [Judge epochs by accuracy, not loss](#judge-epochs-by-accuracy-not-loss)
- [Where Jev fits &amp; where Laya fits](#where-jev-fits-where-laya-fits)
- [A checklist for reading any decision-model benchmark](#a-checklist-for-reading-any-decision-model-benchmark)
- [Caveats: what could make these numbers wrong](#caveats-what-could-make-these-numbers-wrong)
- [What we'd pick for this channel: TF-IDF &amp; human review](#what-we-d-pick-for-this-channel-tf-idf-human-review)
- [Frequently asked questions](#frequently-asked-questions)
- [What is the difference between Jev and Laya?](#what-is-the-difference-between-jev-and-laya)
- [Does Laya work without fine-tuning?](#does-laya-work-without-fine-tuning)
- [How much does Jev cost and how fast is it?](#how-much-does-jev-cost-and-how-fast-is-it)
- [How fast is Laya on a CPU?](#how-fast-is-laya-on-a-cpu)
- [How many labels should I collect before fine-tuning Laya?](#how-many-labels-should-i-collect-before-fine-tuning-laya)
- [How do I check if a model's confidence is good enough to send unsure cases to a person?](#how-do-i-check-if-a-model-s-confidence-is-good-enough-to-send-unsure-cases-to-a-person)
- [Would a better question wording fix Jev's score on everyday messages?](#would-a-better-question-wording-fix-jev-s-score-on-everyday-messages)
- [What GradScaler starting scale should I use for a small fp16 fine-tune?](#what-gradscaler-starting-scale-should-i-use-for-a-small-fp16-fine-tune)
- [Is it safe to send employee messages to Jev?](#is-it-safe-to-send-employee-messages-to-jev)
- [Accuracy is the easy number](#accuracy-is-the-easy-number)
- [Method details](#method-details)
- [Calling Laya](#calling-laya)
- [Calling Jev](#calling-jev)
- [The TF-IDF baseline](#the-tf-idf-baseline)
- [Raw zero-shot scores](#raw-zero-shot-scores)
- [Fine-tuned results, in numbers](#fine-tuned-results-in-numbers)
- [Sources](#sources)
 
  

 

Jev &amp; Laya landed in my feed in the same week, and the claims were big. Decisions in a few milliseconds. Honest, calibrated probabilities. No prompt engineering, no LLM bill. One of them open, one of them closed, and each model card quietly comparing itself to the other.

Launch-week charts weren't going to settle it, so I spent a weekend running **an experiment on a dataset we own**: a year of QED42's `#availability` Slack channel, where people post when they start, step out or take leave. I turned each message into a few structured decisions &amp; checked the launch claims against them. I ran both models against a classic baseline, on a free Kaggle GPU account and about four cents of API credit, and wrote down everything, including what broke.

**No labels: start with Jev. A few hundred labels: fine-tune Laya. Need to send unsure cases to a person: test the confidence first.**

## Key takeaways

- **The model that scored best didn't know when it was wrong.** Fine-tuned Laya made 19 of its 25 status mistakes at 90% or higher confidence, while plain TF-IDF made 1 of 28.
- **With no training labels, Jev is the better start.** It got all five answers right on 64.8% of hard messages, against 47.0% for our regex rules and 10.2% for Laya out of the box (zero-shot).
- **With a few hundred labels, fine-tuned Laya pulls ahead.** On hard messages, it beat a classic TF-IDF baseline (a word-counting text classifier) at every training size and passed Jev at roughly 100 to 250 labels. With all 2,262 labels it scored 83.7% against Jev's 64.8%.
- **Jev is cheap:** $0.044 per 1,000 messages, at a median of 395 ms per message through OpenRouter.

## Why I ran this on our own data

At least one headline number came from a model fine-tuned on its benchmark's own training split. So I wanted private, messy data that neither model had seen.

Our `#availability` channel fits that. It has about 27,000 posts a year, and most are stock phrases. The long tail of sick days, half days, late starts &amp; outages is where simple rules slip. Support tickets that need a priority, invoice exceptions that need a reason code and alerts that need a severity look much the same. So I think the results carry over to that kind of work.

The usual options are hand-written rules, a classifier trained on labels, or an LLM paid per call. **Jev &amp; Laya claim a fourth**: LLM-level reading at classifier speed &amp; cost, with structured output. That's the claim I wanted to test.

## The dataset: a Slack channel full of tiny decisions

Every working morning, the channel fills up. "Starting now." "Lunch break." "Not well, taking the day off." People post there to tell their team about their own availability. A handful of regex rules can already tag these posts, &amp; they're our first baseline. The rules handle the stock phrases fine and **fall apart on everything else**, which is exactly where the sick days, half days &amp; power cuts live. This experiment only measured how well each system reads those posts. It produced no per-person statistics.

I had a year of the channel to work with: **27,326 messages** from September 2025 to September 2026.

About 89% of them are one of roughly 25 stock phrases. Any method gets those right. The other 11%, the **long tail**, is where people announce sick days, half days, late starts &amp; outages. There, **the regex gets the status wrong about 30% of the time**. That's why the long tail gets its own test set below. Some typical misses:

- "Logging off now, on leave next Monday" is filed as leave *today*.
- "Half day today" posted at 2:15 PM is filed as leave in the *first* half.
- "Leaving now" at 5 PM is an early exit, but the rules ignore the time and call it a normal day end.

### Five decisions per message

Each message comes down to **five small decisions**:

QuestionTypeAnswers`status`: what is the person doing right now?choice12 options: `day_start`, `day_end`, `meal_break`, `late_start`, `early_leave`, `leave_full_day`, …`availability_impact`: how much of today is lost?score0 none, 1 brief, 2 partial day, 3 full day`reason`: why?choice8 options: health, family, errand, travel, power or internet, …`future_leave_notice`: leave announced for a later date?yes/no–`makeup_commitment`: promises to make up the time?yes/no–

I defined each question once, as data, &amp; gave both models the same definitions. Simplified, the `status` question looks like this:

 ```
{
  "id": "status",
  "type": "choice",
  "question": "What is the person doing right now?",
  "options": ["day_start", "day_end", "meal_break", "late_start", "early_leave", "leave_full_day", "..."]
}
```

Each model sees the message plus a little context. Here's one input. The text is made up, in the style of the channel:

 ```
{
  "date": "2026-07-15",
  "posted_at": "14:15",
  "weekday": "Wed",
  "public_holiday": false,
  "text": "Taking the second half off today, will cover the pending review tonight"
}
```

The right answers: `leave_second_half`, impact `2`, reason `none_given`, no future leave, and yes to making up the time. **Posting time matters here.** The same words at 9 AM would mean something else.

### Where the correct answers came from

Two Claude models (Sonnet &amp; Opus) labelled every message independently, following a written labelling guide. The final label is the majority vote of those two plus the regex. I dropped questions with no majority. The two annotators agreed on 97.6% of status labels, so **roughly 2% of the "correct" answers are debatable**. The guide isn't public, but two of its rules matter later. Status means what the person is doing at the moment of posting, &amp; an ordinary sign-off or lunch break loses no working time.

## What are Jev and Laya?

Jev &amp; Laya are **typed decision models**: models that make decisions instead of writing text. You give them an input and a few questions with fixed answer options, and they return **a probability for every option in one pass**. No tokens are generated, so they are fast and the output always matches the schema. You can use them *zero-shot*, meaning with no examples from your own data, just the question wording, or fine-tune them on labelled examples where the model allows it. These models were everywhere in AI feeds in mid-September 2026.

Both support the same three question types:

- **Choice:** pick one option from a list.
- **Score:** pick a level on a scale.
- **Noul:** a yes/no answer returned as a probability.

### What is Jev?

Jev started the trend. [TypeSafe AI](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/), founded by a former OpenAI researcher, launched it on [Hacker News](https://news.ycombinator.com/item?id=49717558) on 15 September 2026, where it got nearly 2,000 points. Jev is a **closed, hosted API**. TypeSafe keeps the architecture private. [One write-up](https://www.digitalocean.com/resources/articles/what-is-jev) notes that outsiders suspect an open-weight LLM underneath. TypeSafe hasn't confirmed it.

### What is Laya?

Laya appeared three days later, on 18 September 2026, from Convai Innovations. It uses the same three question types, but **it is open**: Apache-2.0 weights on [Hugging Face](https://huggingface.co/convaiinnovations/laya) and a `pip install laya` package. It's a ModernBERT-large encoder with a small decision head (421M parameters), trained with RLCD, a scoring rule that rewards honest probabilities. The [model card](https://huggingface.co/convaiinnovations/laya) has the rest.

The Laya card also compares itself to Jev on the [typed-decisions benchmark](https://huggingface.co/datasets/LocalLLaMA/typed-decisions), and here it pays to read closely. The Jev numbers in that table were measured by third parties on different samples, and **Laya's headline 0.766 comes from a checkpoint fine-tuned on that benchmark's own training split**. To its credit, the card has an "Honest Limits" section that says the base model is near chance zero-shot and calls Laya a fast base to specialise.

  ![Screenshot of the Laya model card on Hugging Face. Its comparison table lists Jev 1.13.0 (published) at 0.727. The Honest Limits section says the base checkpoints are near chance on typed decisions zero-shot, that the 0.766 belongs to the checkpoint fine-tuned on that benchmark's own training split, and that Laya is a fast base to specialise, not a zero-shot decision engine.](/sites/default/files/styles/uncropped_600w_webp/public/2026-10/laya-hugging-face-model-card-honest-limits.png.webp?itok=E5VtchXy) *Laya's model card on Hugging Face: the end of its comparison table, with the published Jev 1.13.0 row, and the "Honest Limits" section.*

### What independent testers found

The few independent tests show a mixed picture:

- [AbdelStark's benchmark](https://github.com/AbdelStark/jev-benchmarks) measured Jev at 0.910 on AG News and 0.870 on Banking77. It also found Jev badly calibrated on the DAIR Emotion dataset, with zero probability on the true label 16% of the time.
- [A stress test of both models](https://github.com/gazelle93/decision-models-under-pressure) found that, with 64 options, just reordering them **changed 49.4% of Laya's answers** and 14.6% of Jev's. With 128 answer options, Jev scored 60% and Laya 39%.
- I couldn't find an independent test of either model on a real team's data, or any test of a fine-tuned Laya. **That's the gap this post fills.**

The launch weeks were noisy: copycat projects, blog posts that compared models they had not actually run, and a prior-art argument between the two camps. The safest reading of every number above is **"vendor claim"** unless an independent source re-ran it. That's what this weekend was for.

## The claims &amp; what held up

The claimVerdictWhat I foundJev works out of the box, no training needed**Held up**64.8% of hard messages fully right with zero labels, against 47.0% for the regexLaya is "a fast base to specialise, not a zero-shot decision engine" (its own model card)**Held up, both halves**10.2% out of the box. Fine-tuned, it beat a classic TF-IDF model on hard messages at every training size and passed Jev at roughly 100 to 250 labelsLaya beats Jev (Laya's comparison table)**Only after training**Zero-shot Laya is far behind. With all 2,262 labels: Laya 83.7%, Jev 64.8%Both give honest, calibrated probabilities**Partly**Average calibration looked fine for fine-tuned Laya and Jev, but 19 of fine-tuned Laya's 25 mistakes came with 90%+ confidence. Jev: 26 of 90. Plain TF-IDF: 1 of 28Laya answers in about 33 to 40 ms per question on a T4**Roughly**About 110 ms per message for all five questions together on a T4, roughly 22 ms per question. The card's own 5-question batch figure is 40.1 ms, faster than I sawJev is fast and cheap**Cheap, yes**$0.044 per 1,000 messages. A median of 395 ms per message through OpenRouter, network included

**The short version:** Jev is the better start when you have no labels. Fine-tuned Laya pulls ahead once you have a few hundred. And on knowing when it's wrong, **the oldest model in the test beat both**.

One more thing shaped every step. These messages describe colleagues' health &amp; family lives, so I de-identified them before they left my machine. **De-identified isn't the same as anonymous**, and I explain the gap below. The correct answers were labelled by Claude models, and an AI coding assistant (Claude) wrote most of the scripts. Both matter for some comparisons (see [Caveats](#caveats-what-could-make-these-numbers-wrong)).

## Keeping colleagues' data private, &amp; where that falls short

This data is a year of named colleagues' sick days, family emergencies &amp; leave. **I treated it like HR data.** Every stage got the least it needed, and nothing in this post quotes a real message.

StageWhat it sawHow it was protectedLabelling (before this experiment)Every message, with names replaced by placeholdersWe replaced names with placeholders before two Claude models labelled itMy laptopThe full, named dataIt stayed there. Scripts printed counts and scores. I printed a few sample rows while checking the data, and those passed through my AI assistant too.Kaggle, for trainingA de-identified copy: training subsets, dev &amp; test setsPrivate dataset. Staged in temporary storage, so it's not saved with the notebook. I kept only ids and scores as output.Jev, via OpenRouterThe de-identified test sets only, 881 messagesI sent them after explicit sign-off. The model page lists a no-training data policy, and prompt logging was off. The script refuses to read the named files.This postAggregate numbersExample messages are invented. No per-person statistics.

### What "de-identified" means here

A script removed every author name and Slack ID, and replaced every @-mention with `@colleague`. That covered the full names of the 95 people in the channel, first names of people outside the channel, and lower-case Slack handles. `@here` was kept. Then it checked its own work:

 ```
"slack_ids_left": 0,
"author_fields_left": 0,
"full_names_left": 0,
"texts_with_an_author_surname": 0,
"at_mentions_not_colleague": 0
```

**That check earned its place.** The first version missed two things: mentions of people outside the 95-person list, and lower-case Slack handles (one of them appeared 276 times). The check caught both, and I fixed them before uploading anything.

**De-identified still isn't anonymous.** The text itself can point to a person: a sick day on a known date, a family event, or a writing style colleagues would recognise. That's why even the de-identified copy went only to a private dataset &amp; one approved API, and why this post shows numbers, not messages.

## How I set up a fair benchmark

I compared five systems on the same test messages:

SystemTraining labelsWhat it isRegex rules0Hand-written rules for the stock phrases, the zero-label baselineJev 1.130TypeSafe's closed decision model, called through OpenRouter, no fine-tuning offeredLaya zero-shot0The public Laya checkpoint, given only the question wordingTF-IDF + logistic regression25 to 2,262A classic text classifier, one per questionLaya fine-tuned25 to 2,262Laya trained further on labelled messages

### Split by time, not at random

Training uses messages from before July 2026. Testing uses July to September 2026, so **no test message can leak into training**.

**The training pool isn't a random sample.** A fixed rule picked 2,262 messages from before July: every long-tail message (1,780, apart from 300 held back as the dev set) and up to 3 examples of each stock phrase per time-of-day band (482). A message counts as long tail if its text appears 3 times or fewer in the whole year. The cap is how "leaving now" at 13:30 &amp; the same words at 19:45 both made it into training.

There are two test sets, and the rest of this post calls them hard and everyday:

- **Hard messages** (381) are every long-tail message from the test period. **This is where the methods differ.**
- **Everyday messages** (500) are a random sample at the channel's real mix, where 88% are stock phrases.

### A baseline that knows the posting time

**TF-IDF sees the posting hour &amp; weekday as tokens.** So it can learn that "leaving now" at 17:00 is different from the same words at 19:45. It uses character- &amp; word-level TF-IDF features with a logistic regression on top, one model per question (code in [Method details](#method-details)).

### Same examples for every model

For each size (25, 50, 100, 250, 500 and 1,000 labels), I drew five random subsets from that pool with fixed seeds. I trained one model per subset and averaged the five scores. The full 2,262-label set ran once. The seeds are shared, so **a difference between models comes from the model, not from luckier examples**:

 ```
for seed in range(5):
    subset = random.Random(seed).sample(train, n)   # identical rows for TF-IDF and Laya
```

### Running the models

Both models got exactly the same input for every message: the state shown above and the same five typed questions. Laya ran on a GPU through its `laya` Python package, &amp; I called Jev through OpenRouter's Decisions API. Both return an answer and a probability for each question. The calling code for both is in [Method details](#method-details).

### Fine-tuning Laya

I reused the training loop from Laya's own Kaggle notebook, ported to run on one GPU. Each training example is one message paired with one question, so 25 labelled messages make 125 examples:

- 4 passes over the data (epochs), an effective batch of 64, and Laya's RLCD loss, which rewards honest probabilities rather than just right answers.
- After training, I kept the epoch with the best accuracy on a separate 300-message dev set.
- Then I fitted one confidence "temperature" per question type on that same dev set.

### Hardware

My MacBook (Apple M4) could train, but slowly: about 0.6 message-question pairs per second. From a short timing test, I estimated about 3 hours for the 250-label run, dev checks included. On Kaggle's free tier (two NVIDIA T4 GPUs), the same run took **about 8 minutes**. All 31 fine-tunes, with scoring, took about 4 hours. I ran two jobs at once, one per GPU, and started every run from the terminal with the Kaggle command-line tool.

## Terms used in this post

A few terms come up again &amp; again. Here's what each one means in this experiment.

- **Dev set:** 300 long-tail messages from before July, held back from training &amp; used to pick the best epoch &amp; tune confidence. The dev loss &amp; accuracy charts come from it.
- **Hard messages:** the 381 long-tail test messages, meaning texts that appear 3 times or fewer in the year: the sick days, half days, late starts &amp; outages. `test_longtail` in the code &amp; screenshots.
- **Everyday messages:** 500 random test messages at the channel's real mix, mostly stock phrases like "Starting now". `test_natural` in the code.
- **All five correct:** the main score. A message counts only if status, impact, reason, future leave &amp; make-up time are all right.
- **Calibration &amp; ECE:** calibration is whether a model's confidence, the probability it gives its own answer, is honest. ECE (expected calibration error) is the average gap between how sure a model says it is &amp; how often it's right. Lower is better.
- **Confidence gate:** send a message to a person when the model isn't sure enough, keep its answer otherwise. It only works if wrong answers come with low confidence.

## Result 1: with no labels, Jev is ready &amp; base Laya isn't

With no training labels, Jev got all five answers right on **64.8%** of the hard messages, the regex rules on **47.0%**, and Laya on **10.2%**. **On everyday messages the order flips:** the regex rules lead, because they were written for exactly those stock phrases.

  ![Bar chart of messages with all five answers correct and zero training labels. Hard messages: Jev 1.13 64.8%, regex rules 47.0%, Laya out of the box 10.2%. Everyday messages: regex rules 94.2%, Jev 59.6%, Laya 38.0%.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_01-zero-shot-jev-vs-laya-vs-regex.png.webp?itok=eAXmwn5c) *With zero training labels, Jev wins on the 381 hard messages, while the regex rules still win on the 500 everyday ones. Jev and Laya saw de-identified text.*

Here are the full numbers, including status accuracy on its own:

MetricRegex rulesJev 1.13Laya zero-shotHard messages, all five correct47.0%**64.8%**10.2%Hard messages, status correct70.3%**76.3%**32.6%Everyday messages, all five correct**94.2%**59.6%38.0%Everyday messages, status correct96.4%**97.0%**84.0%

The per-question scores (full output in [Method details](#method-details)) show where zero-shot Laya went wrong.

Two things stand out. **Laya called 69 ordinary messages a full day off.** And on future leave it flags far too much: **only 6% of its "yes" answers were right**.

This matches what Laya's own [model card](https://huggingface.co/convaiinnovations/laya) says under its limits: the base checkpoints are close to chance on typed decisions, and the accuracy comes from fine-tuning. I expected zero-shot Laya to at least beat the regex on the hard messages. **It didn't come close.**

### Calibration fixed Laya's confidence, but not its 32.6% accuracy

I fitted temperatures on the dev set, a cheap step that rescales probabilities without changing any answer. Status calibration error (ECE) fell from 0.445 to 0.127. Accuracy stayed at 32.6%, as expected.

The published checkpoint also ships with a broken temperature for questions with 11 or more options. The `laya` package (0.3.22) clamps it &amp; prints a warning, so **zero-shot status confidence should not be trusted as is**.

### Jev beat the regex with zero examples

With the same question wording and no examples, **Jev was already better than the regex on the hard set**. It got the reason right 96.6% of the time and caught every advance leave notice, at a median of 395 ms and $0.044 per 1,000 messages. Its full per-question scores are in [Method details](#method-details).

### Why Jev loses to the regex on everyday messages

Jev scores 59.6% against the regex's 94.2% on everyday messages **almost entirely because of one question**. On the everyday set it rated 141 ordinary "logging off" messages as losing part or all of the day, and 36 lunch breaks as a brief loss. The labelling guide says both lose nothing.

The question asks *"How much of today's working time does this message say the person will lose?"*, and reading a sign-off as lost time is a fair reading of those words, just not the one the labels use. **A zero-shot model only knows what the question says.** A better question wording might fix this, but I kept the wording identical for every model, so I didn't try.

## Result 2: fine-tuned, Laya needs about half the labels

Trained on the same examples, fine-tuned Laya beat TF-IDF on hard messages at every size, and **matched TF-IDF's score with about half as many labels**.

  ![Line chart of all-five accuracy on 381 hard messages by number of training labels. Fine-tuned Laya rises from 37% at 25 labels to 84% at 2,262, and TF-IDF from 28% to 80%. Laya passes zero-shot Jev (65%) between 100 and 250 labels. Regex rules score 47% and untrained Laya 10%.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_02-learning-curve-hard-messages.png.webp?itok=0ZQ3mp-2) *All-five accuracy on hard messages by training size (log scale): 31 fine-tuned Laya runs on Kaggle T4 x2 (4 Oct 2026) against TF-IDF on the same seeded subsets. Bars show the worst to best of five runs. Dashed lines are the zero-label systems.*

**The dashed Jev line is the price of labelling.** Jev needs no labels and scores 64.8%. Fine-tuned Laya passes it somewhere between 100 and 250 labels, and TF-IDF between 250 and 500. If you can label a few hundred messages &amp; keep your data in-house, I think you can beat a hosted model. If you can't, Jev is the strongest option here.

The lead is about 10 points up to 250 labels and **shrinks to 4 points with all 2,262**. That's what I expected for a channel this repetitive: with enough examples, a word-counting model learns the stock phrases too.

### Where Laya's lead actually comes from

The overall number hides where the gap is. On status alone, Laya's lead disappears from 500 labels: 88.9% vs 88.8% there, and 93.4% vs 92.6% with all labels. With full data, **the lead comes from the harder questions**: reason (97.4% vs 93.4%) &amp; availability impact (91.3% vs 90.0%).

On the everyday test set, **both models are erratic with small training sets**. Below 1,000 labels, the gap between the worst and best of five runs is wide for both, and at 500 labels TF-IDF is actually the steadier of the two. With all labels, Laya scores 98.4% and TF-IDF 97.8%. The exact numbers for both test sets are in [Method details](#method-details).

  ![Chart of all-five accuracy on everyday messages for Laya and TF-IDF at each training size, showing the mean and the worst-to-best range of five runs. Both models vary widely below 1,000 labels, and both reach about 97% at 1,000 labels.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_03-everyday-messages-spread-by-size.png.webp?itok=wNB-4ma7) *On everyday messages, which subset you happen to label matters as much as which model you pick, until about 1,000 labels. Dot: mean of five runs. Bar: worst to best run.*

## Result 3: only TF-IDF knows when it is wrong

**The best-scoring model is not automatically the best system.** I planned to use confidence as a gate: keep the model's answer when it is sure, send the message to a person when it is not. For that, the model's mistakes must come with low confidence. Fine-tuned Laya's mostly don't.

On the hard set, the full-data Laya and TF-IDF made a similar number of status mistakes: 25 and 28. Jev, with no training, made 90. **What matters for a gate is how sure each model was when it was wrong**:

- **Fine-tuned Laya:** 19 of its 25 mistakes with 90% or higher confidence.
- **Jev:** 26 of 90.
- **TF-IDF:** 1 of 28.

  ![Bar chart of status mistakes on 381 hard messages, split by confidence. TF-IDF made 1 of its 28 mistakes at 90% or higher confidence, fine-tuned Laya 19 of 25, and Jev 1.13 26 of 90.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_04-confident-mistakes.png.webp?itok=GVB4TjEY) *Status mistakes on the 381 hard messages, split by how sure each model was. The solid part of each bar is mistakes made with 90% or higher confidence, which a confidence gate can never catch.*

Here is how many messages a person would need to review:

Gate (send to a person if confidence is below)Messages reviewedStatus accuracy after reviewTF-IDF, 0.7032.5%**98.7%**Laya fine-tuned, 0.909.4%95.0%Laya fine-tuned, 0.9570.9%96.6%Jev, 0.7040.2%87.9%Jev, 0.9069.3%93.2%

Sweeping the threshold across its whole range shows the same picture:

  ![Line chart of status accuracy after human review against the share of hard messages sent to a person. TF-IDF reaches 99% with about a third reviewed. Fine-tuned Laya stays near 96 to 97% until almost every message is reviewed. Jev climbs slowly from 76%.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_05-human-review-cascade.png.webp?itok=bWamB_ut) *Status accuracy after human review on the hard messages. Each line sweeps the confidence threshold. A message goes to a person if any of its five answers is unsure, and the person is assumed to be always right.*

**TF-IDF reaches about 99% if a person checks a third of the hard messages. Neither decision model gets there.** Laya's confidence piles up near 1.0, so raising the threshold barely helps until suddenly most messages fall below it. That's the long, flat blue line. Jev's confidence does sort its answers better, but it starts from a lower accuracy, so a person would have to check most of the hard messages.

### Good average calibration and a good gate are different tests

Laya's status ECE was 0.031, which looks excellent. But ECE and gating measure different things:

- **ECE** measures whether confidence is right *on average*: if the model says 95% on 100 answers, about 95 should be right.
- **A gate** needs the wrong answers to have *lower* confidence than the right ones.

**A model can pass the first test and fail the second, and this one did.**

### Nine of Laya's confident mistakes are debatable labels

9 of the 19 match what the regex rules or one of the two annotators said, so the "correct" label is debatable there. The most common one is a message like "starting now, taking the second half off": the labelling guide calls that `day_start` (what the person is doing at posting time), while Laya answers `leave_second_half`. **A person might well agree with Laya.**

## What broke along the way

The biggest single improvement in this project came from **fixing a bug in my own training script**. Here's everything that went wrong, in the order I hit it:

ProblemWhat I sawFixMy first successful sweep ran in an interactive editor sessionIts results zip vanished when the session endedOnly a saved version keeps its output files**My training script skipped the first updates of small runs**The 25-label runs' dev loss didn't move from 0.894 for the whole first epoch, and for some runs the second tooStart the fp16 `GradScaler` at 210 instead of 216

### The fp16 GradScaler bug

On a T4, I trained in 16-bit floats (fp16) to save memory and time. Small fp16 numbers can underflow to zero, so PyTorch's `GradScaler` multiplies the loss by a large number before backpropagation. If the gradients then overflow, **it skips that update and halves the scale**.

The default starting scale is 65,536 (216), and the first few updates usually overflow. On a big dataset, losing a handful of updates doesn't matter. **A 25-label run has 125 examples, so it only gets 8 updates in total, &amp; it lost most of its training**:

 ```
# before: the first ~5 optimizer steps are skipped as overflow
scaler = torch.amp.GradScaler("cuda", enabled=use_amp)

# after
scaler = torch.amp.GradScaler("cuda", init_scale=2.0 ** 10, enabled=use_amp)
```

You can see the skipped updates in the dev loss. Before the fix, not one of the five 25-label runs moved in epoch 1. After it, every run improved straight away:

  ![Two line charts of dev loss for the five 25-label runs. Before the fix, with GradScaler starting at 2^16, no run's loss moved in epoch 1. After the fix, starting at 2^10, every run's loss dropped in epoch 1. All-five accuracy at 25 labels rose from 22.9% to 37.1%.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_06-gradscaler-bug-before-after.png.webp?itok=_j_QbvFs) *Dev loss of the five 25-label runs, before (*`<em>GradScaler</em>` *starting at 216) and after (210) the fix.*

**After the fix, the 25-label score went from 22.9% to 37.1%.** Before it, Laya lost to TF-IDF at 25 labels. After it, Laya won on hard messages at every size.

### Judge epochs by accuracy, not loss

I also misread a chart. In the first sweep, dev *loss* bottomed out after epoch 1 &amp; then climbed. You can see the same rise in the right-hand panel above. That looked like overfitting, so I added "keep the best epoch" to the script. It made almost no difference, **because dev** ***accuracy*** **kept rising through epoch 4**:

  ![Line chart of Laya's dev accuracy per answer by epoch for each training size from 25 to 2,262 labels. Accuracy rises at every size through epoch 4, finishing between about 80% and 97%.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669307_08-training-curves-by-size.png.webp?itok=hV1GDGlh) *Dev accuracy per answer, averaged over the runs at each training size. Loss rose after epoch 1 for many runs, but accuracy never fell.*

The model was getting over-confident. Loss punishes that and accuracy doesn't. The temperature fit corrects it afterwards. **Lesson: for this model, judge epochs by accuracy, not loss.**

Here are all 31 runs after the GradScaler fix, one cell per run. The 20 runs at 25 to 250 labels ran in an earlier session. The final sweep covered the other 11, at 500, 1,000 and 2,262 labels, and took 8,686 seconds on Kaggle's GPU T4 x2, including 60 minutes for the full-data run:

  ![Heatmap of all-five accuracy for all 31 fine-tuned Laya runs by training size and random seed. On hard messages, scores rise steadily from 31 to 46% at 25 labels to 84% at 2,262. On everyday messages, small runs swing widely, from 6% to 64% at 25 labels, and settle at 95 to 98% from 1,000 labels.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/1791356669308_07-every-fine-tuning-run.png.webp?itok=8djcS4JL) *Every fine-tuning run: all five answers correct, per run. Each row is a training size and each column a different random subset. Kaggle GPU T4 x2, October 2026.*

Both numbers come straight from Kaggle. Here is the run log for the final sweep:

  ![Screenshot of the laya_availability_kaggle notebook logs on Kaggle: successfully ran in 8,685.8 seconds on a GPU T4 x2 accelerator, with two Tesla T4 GPUs of 15,360 MiB each listed in the log.](/sites/default/files/styles/uncropped_960w_webp/public/2026-10/kaggle-run-log-laya-sweep-gpu-t4x2.png.webp?itok=g7YQuogV) *Kaggle run log for the final sweep: it ran successfully in 8,685.8 seconds on a GPU T4 x2 accelerator. The notebook itself is private.*

And the notebook's own output, one line per run as each one finished on either GPU:

  ![Screenshot of the Kaggle notebook output for the final sweep, one line per fine-tuning run across two GPUs. The full-data run laya_ft_n2262_s0 scored all-five 0.837 and status 0.934 on test_longtail in 60 minutes; the 1,000-label runs scored 0.769 to 0.801 and the 500-label runs 0.714 to 0.740.](/sites/default/files/styles/uncropped_720w_webp/public/2026-10/kaggle-notebook-output-laya-fine-tuning-runs.png.webp?itok=m7G8bjNP) *Notebook output for the final sweep, which covered the 500-, 1,000- and 2,262-label runs. The full-data run scored 83.7% in 60 minutes.*

## Where Jev fits &amp; where Laya fits

The two models suit different situations. **IMO, Jev is the easier start.** It needs no labels &amp; no GPU, cost four cents for 881 messages, &amp; beat the regex on day one. **Laya fits** when you can label a few hundred examples and want the model, and the data, under your own control.

**Data control was a real constraint for me.** My first draft of this post had no Jev numbers at all, because sending colleagues' leave messages to an outside API needed sign-off. When I got it, I sent only a de-identified copy of the test set. Laya never needed that step: its training ran on a private Kaggle dataset and could have run on my laptop.

These are the questions that decide it:

- **Can you label a few hundred examples?** No: use Jev, since Laya out of the box was useless here. Yes: fine-tuned Laya passed Jev at roughly 100 to 250 labels.
- **Can your data leave the building?** No: Laya, or any model you can run yourself.
- **Do you need to route unsure cases to a person?** Neither decision model was good enough at that here. Test the confidence, not just the accuracy ([Result 3](#result-3-only-tf-idf-knows-when-it-is-wrong)).
- **Does the question wording match what you mean?** Jev only knows what the question says. The "impact" question meant one thing to the labels and another to Jev. Fine-tuning learns the intent from examples.
- **How many options does a question have?** An independent test found both models getting worse with many options, and reordering 64 options changed half of Laya's answers. My biggest question had 12 options, which worked.

### A checklist for reading any decision-model benchmark

This applies to mine too:

1. **Was the model fine-tuned on that benchmark's own training data?** A fine-tuned score is not a zero-shot score.
2. **Were all models run on the same examples, by the same people?** Numbers borrowed from different samples are not a comparison.
3. **Who wrote the correct answers?** If an LLM labelled the data, scores partly measure agreement with that LLM.
4. **Is there a simple baseline (majority class, regex, TF-IDF)?** Without one, 0.77 means nothing.
5. **Is latency per question or per input, and on what hardware?** I measured about 110 ms per message for all five questions together on a T4, roughly 22 ms per question, while the Laya card quotes 33 to 40 ms for one question asked on its own. Always check which of the two a number means. Jev took a median of 395 ms per message, but that includes the network trip through OpenRouter.

## Caveats: what could make these numbers wrong

These results hold for one channel, one team and one labelling process. Keep these limits in mind:

- **Claude wrote the correct answers.** Two Claude models labelled the data, with the regex as a third vote. That third vote slightly favours the regex baseline wherever the two Claude models disagreed. Any Claude-based system would be flattered on this test, and the fine-tuned models are partly learning to agree with Claude. If Jev is built on an LLM, as some outsiders suspect, it may share some of the annotators' reading of the questions too. I can't tell.
- **About 2% of the labels are debatable.** The two annotators agreed on 97.6% of status labels, so no system can score much above ~98%. Some "confident mistakes" are cases where a model sided with one annotator.
- **I tested Jev zero-shot only, once, as version 1.13.** It has no fine-tuning option, and a hosted model can change under you. The same questions might score differently next month.
- **I wrote the questions for the annotators, not for Jev.** Jev's biggest loss came from one question it read differently from the labelling guide. A prompt tuned for Jev would likely score higher. I didn't tune prompts for any model.
- **Same people in training and test.** The split is by date, but the same colleagues write on both sides. This measures how well a model learns *this team's* style, not how it handles strangers.
- **The full-data point is one run.** Every size up to 1,000 labels is the average of five random subsets. The 2,262-label result is a single run. Its spread is unknown.
- **The training pool is mostly long tail.** 79% of the 2,262 training messages are long-tail, while the channel is about 89% stock phrases. Every model trained on the same pool, so the comparison stays fair. But the label counts on the learning curve are draws from this pool. Labelling 250 random channel messages would give you mostly stock phrases, &amp; probably lower scores on the hard set.
- **Small subsets are fragile.** One 25-label subset had only 4 routine messages, and that model labelled 278 ordinary "starting now" and "logging off" posts as full-day leave. That's the 6% cell in the heatmap above. TF-IDF shares this problem, because it trains on the same subsets.
- **De-identified text.** Jev and every cloud run saw `@colleague` in place of names. A zero-shot Laya check scored slightly higher without names (10.2% vs 8.9%), so this probably didn't hurt, but the two runs also used different hardware.
- **Latency isn't like for like.** I timed Laya on the GPU itself. Jev's 395 ms median includes the network trip through OpenRouter.
- **Fixed recipe.** I didn't tune learning rates or epochs for Laya, or TF-IDF beyond its defaults. A tuned version of either could do better.

## What we'd pick for this channel: TF-IDF &amp; human review

If we had to automate this channel, IMO the pick is **TF-IDF with a person reviewing the low-confidence messages**, &amp; I'd keep working on Laya. That's the decision this experiment was set up to make.

The labels already exist, so Jev's no-labels advantage doesn't help here, and the data would have to leave the building. With all labels, Laya is only 4 points ahead of TF-IDF overall and level on status. TF-IDF runs in about 6 ms on any CPU, against about 110 ms on a T4 GPU for Laya. Most importantly, I think **its confidence can pick out the messages a person should check**.

A different team would choose differently:

- **No labels, and data that may leave the building:** start with Jev. It beat the regex with zero work, for $0.044 per 1,000 messages.
- **Fewer than about 500 labels, and data that must stay in-house:** fine-tuned Laya, which clearly leads TF-IDF there.
- **You need the finer answers:** with full data, Laya leads on reason (97.4% vs TF-IDF's 93.4%) &amp; impact (91.3% vs 90.0%).

## Frequently asked questions

### What is the difference between Jev and Laya?

Both are typed decision models that return a probability for every answer option instead of generating text. Jev is a **closed, hosted API** from TypeSafe AI that you use zero-shot. Laya is an **open, Apache-2.0 model** from Convai Innovations (421M parameters, built on ModernBERT-large) that is designed to be fine-tuned on your own labels.

### Does Laya work without fine-tuning?

**Not well.** On the hard messages in this benchmark, zero-shot Laya got all five answers right on 10.2% of messages, against 47.0% for simple regex rules. Laya's own model card says the base checkpoints are near chance zero-shot and that accuracy comes from fine-tuning.

### How much does Jev cost and how fast is it?

Called through OpenRouter's Decisions API, Jev 1.13 cost **$0.044 per 1,000 messages** with five questions each. Median latency was 395 ms per message (p95 761 ms), including the network round trip.

### How fast is Laya on a CPU?

I didn't measure it. Laya's [model card](https://huggingface.co/convaiinnovations/laya) reports 193 to 464 ms per request on CPU, with the model preloaded. On a T4 GPU I saw about 110 ms per message, and TF-IDF took about 6 ms on a CPU.

### How many labels should I collect before fine-tuning Laya?

In this experiment, **about 250 labels** put fine-tuned Laya ahead of Jev on hard messages. Those labels came from a pool weighted towards unusual messages, though. Random messages from a channel like ours would be mostly stock phrases, so you'd likely need more. On everyday messages, results changed a lot depending on which examples were picked, until about 1,000 labels.

### How do I check if a model's confidence is good enough to send unsure cases to a person?

On your own labelled test set, look at the model's wrong answers &amp; count how many came with 90% or higher confidence. Then try different thresholds &amp; plot accuracy after review against the share of messages a person checks. **The average calibration score alone won't tell you.**

### Would a better question wording fix Jev's score on everyday messages?

I think most of it. Jev lost nearly all its everyday points on the impact question, because it read normal sign-offs as lost time. A wording that says "an ordinary sign-off or lunch break loses no working time" targets exactly that. I kept the wording the same for every model, so I didn't test it.

### What GradScaler starting scale should I use for a small fp16 fine-tune?

**210 worked for me.** Check that the dev loss moves in the first epoch. If it doesn't, the scaler is probably skipping your first updates.

### Is it safe to send employee messages to Jev?

**Only with care &amp; approval.** I removed names first, sent only the de-identified test set, turned prompt logging off &amp; relied on the no-training policy on the model page, all after internal sign-off. Text without names can still point to a person, so treat it as sensitive.

## Accuracy is the easy number

A weekend and four cents of API credit settled most of the launch claims. Laya's own "Honest Limits" section was right: it's a strong base to fine-tune &amp; weak out of the box. Jev was right about working with zero labels. The calibration claims held on average and slipped exactly where a confidence gate needs them, and the model that knew its own mistakes best was the oldest one in the test.

**I think accuracy is the easy number to compare.** Knowing when a model is wrong, and where your data has to go to use it, decided more for us than any leaderboard. **Run the confidence check on your own data before you trust a leaderboard, including this one.**

## Method details

The code and raw outputs behind the results above, for readers who want to check the numbers or try a similar setup. The snippets are simplified excerpts, not a runnable package. I used `laya` 0.3.22. Pin the version if you try this: it had about 35 releases in its first 17 days.

### Calling Laya

Asking Laya is one call. You pass the state and the typed questions, and get back an answer and a probability for each:

 ```
from laya import Router

router = Router(device="cuda", preload=True)
res = router.predict(row["state"], row["questions"])
res["answers"]["status"]
# {"choice": "leave_second_half", "answer_confidence": 0.91, ...}
```

### Calling Jev

Asking Jev looks almost the same. OpenRouter serves it through a separate Decisions API, and it takes the exact same question definitions, so **both models saw identical inputs**:

 ```
req = urllib.request.Request(
    "https://openrouter.ai/api/alpha/decisions",
    data=json.dumps({
        "model": "typesafe/jev-1.13",
        "state": row["state"],          # same state object Laya saw
        "questions": row["questions"],
    }).encode(),
    headers={"Authorization": f"Bearer {key}", "Content-Type": "application/json"},
)
answers = json.loads(urllib.request.urlopen(req).read())["answers"]
```

### The TF-IDF baseline

The posting hour &amp; weekday are added to the text as tokens, then character- &amp; word-level TF-IDF features feed a logistic regression:

 ```
def text_of(r):
    s = r["state"]
    return f"[t{s['posted_at'][:2]}] [{s['weekday']}] {s['text']}"

model = make_pipeline(
    make_union(
        TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5), sublinear_tf=True),
        TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True),
    ),
    LogisticRegression(max_iter=3000, C=8.0, class_weight="balanced"),
)
```

### Raw zero-shot scores

Scorer output for the two zero-shot runs on the 381 hard messages. In the run names, *zs* means zero-shot and *deid* means the model saw de-identified text. Macro-F1 averages the score of each answer option equally, so rare options count as much as common ones. ECE is explained in [Terms used in this post](#terms-used-in-this-post).

 ```
### laya_zs_deid  (n=381, missing=0, all-5-correct=0.102)

| question            | acc   | macro-F1 | ECE   | extra                                 |
| status              | 0.326 | 0.429    | 0.445 |                                       |
| availability_impact | 0.583 | 0.541    | 0.093 | full-day missed 10, false full-day 69 |
| reason              | 0.646 | 0.583    | 0.132 |                                       |
| future_leave_notice | 0.792 | 0.497    | 0.061 | recall 0.625, precision 0.062         |
| makeup_commitment   | 0.908 | 0.632    | 0.085 | recall 0.615, precision 0.211         |
```

 ```
### jev_deid  (n=381, missing=0, all-5-correct=0.648)

| question            | acc   | macro-F1 | ECE   | extra                         |
| status              | 0.763 | 0.721    | 0.111 |                               |
| availability_impact | 0.791 | 0.774    | 0.101 | MAE 0.27, false full-day 10   |
| reason              | 0.966 | 0.888    | 0.020 |                               |
| future_leave_notice | 0.989 | 0.897    | 0.068 | recall 1.000, precision 0.667 |
| makeup_commitment   | 0.987 | 0.888    | 0.101 | recall 0.692, precision 0.900 |

Latency p50 395 ms, p95 761 ms (through OpenRouter, network included)
Cost $0.044 per 1,000 messages
```

### Fine-tuned results, in numbers

Fine-tuned Laya and TF-IDF with all 2,262 labels, and the worst-to-best spread of the five runs on everyday messages at each training size. In the output, `test_longtail` is the hard set and `test_natural` the everyday set.

 ```
test_longtail model            status  impact  reason  all-5
              Laya fine-tuned  0.934   0.913   0.974   0.837
              TF-IDF           0.926   0.900   0.934   0.798
test_natural  model            status  impact  reason  all-5
              Laya fine-tuned  0.994   0.988   0.998   0.984
              TF-IDF           0.994   0.982   1.000   0.978

test_natural, all-5, worst to best of 5 runs
  N=25    Laya 0.06-0.64   TF-IDF 0.24-0.49
  N=50    Laya 0.42-0.69   TF-IDF 0.25-0.60
  N=100   Laya 0.45-0.78   TF-IDF 0.36-0.79
  N=250   Laya 0.73-0.97   TF-IDF 0.55-0.91
  N=500   Laya 0.61-0.98   TF-IDF 0.72-0.97
  N=1000  Laya 0.95-0.98   TF-IDF 0.94-0.98
```

Status ECE for the full-data Laya run on the hard set was 0.031, the figure used in Result 3.

## Sources

All pages accessed 5 October 2026.

- [Laya model card](https://huggingface.co/convaiinnovations/laya), [Laya GitHub](https://github.com/NandhaKishorM/laya) and [Laya releases on PyPI](https://pypi.org/project/laya/)
- [TechCrunch on Jev and TypeSafe AI](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/) and the [Jev launch on Hacker News](https://news.ycombinator.com/item?id=49717558)
- [Jev on OpenRouter](https://openrouter.ai/docs/guides/community/jev.md) and the [Jev tutorial](https://openrouter.ai/docs/guides/community/jev-tutorial.md) (the Decisions API I called)
- [TypeSafe privacy policy](https://typesafe.ai/legal/privacy-policy) and [TypeSafe docs](https://docs.typesafe.ai/introduction)
- [DigitalOcean: what is Jev](https://www.digitalocean.com/resources/articles/what-is-jev)
- [AbdelStark jev-benchmarks](https://github.com/AbdelStark/jev-benchmarks)
- [Decision models under pressure](https://github.com/gazelle93/decision-models-under-pressure)
- [typed-decisions dataset card](https://huggingface.co/datasets/LocalLLaMA/typed-decisions)

 



Written by

Piyuesh Kumar

Director of Technology

 

 

 

 

 

 

 ## We'd love to talk about your business objectives

 [ Contact Us  ](/contact)