Applied AI · ML Systems

Laya vs Jev in Practice: Zero-Shot, Fine-Tuning, and the Real Cost of an Open-Source Decision Model

On the same held-out set, Laya went from 8/70 to 60/70 correct tuples after adaptation. An experiment in structured decisions, confidence, and the cost of control.

TypeSafe announced Jev on September 15, 2026, introducing its first “System One Model.” Laya’s official Hugging Face history dates the commit marked “Initial release” to September 18, 2026. Those public records are three days apart — an interval between an announcement and a release record, not a measure of development time. Sources: TypeSafe’s announcement and Laya’s initial release record.

By the time I completed the Jev 1.13 evaluation, Laya was already an open-weight alternative worth investigating. Learning a tool was just the first step. The more durable work was evaluating a new approach with a sound method and measuring whether it improves the system.

On the same Ops Triage AI held-out set, Laya went from 8/70 correct tuples zero-shot to 60/70 after domain adaptation. Jev had reached 66/70 without domain-specific training. The interesting question is what changes when we can modify a model’s weights, and what we take responsibility for in exchange.

First, what is a decision model?

Jev is a non-generative AI model specialized in structured decisions. It takes state or context expressed in natural language and returns typed decisions with probability distributions, rather than generating an open-ended token sequence like a conventional generative LLM.

It is still AI. TypeSafe calls the category System One Models; that is the company’s terminology, not a universally established academic taxonomy. I also use the more neutral terms “probabilistic decision model” and “structured decision model.”

Generative LLM
text → token-by-token generation → text

Decision model
state + set of choices → distribution → typed decision

This changes both the interface and the task. In ticket triage, category, priority, and risk can be closed choices. The software receives values it already understands, without asking for a paragraph and extracting an enum afterward.

LLMs can produce structured outputs too. The distinction goes beyond model size or JSON formatting: a decision model is designed to score a decision space instead of generating free-form text. A bounded answer space reduces parsing problems and out-of-schema values. A valid choice can still be semantically wrong.

Where Laya fits

Laya addresses a similar role with code and weights released under Apache 2.0, non-autoregressive inference, self-hosting, and fine-tuning. The general English checkpoint combines ModernBERT-large with a decision head for 421 million parameters, according to the model card at the revision we used.

We pinned convaiinnovations/laya at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, using laya==0.3.23. The runtime release was published on October 1, 2026. Library version and weight revision are separate identities, and both were frozen.

This base checkpoint is distinct from laya-typed-decisions, which has already been fine-tuned on other workflows. We did not use the multilingual checkpoint either. Here, “zero-shot” means no training specific to our domain, not a model with no prior training.

A decision model inside a real architecture

A voice-driven website-building demonstration I watched made the role easier to understand. It described a library of more than 1,000 assets/components, with Jev selecting actions and assets. The conceptual flow was:

voice → real-time transcription → context adaptation/correction
      → decision model → asset/action selection → UI

This is an account of the observed demonstration, without a verifiable URL in the project context or benchmark value.

Jev does not have to generate the website. A speech model can transcribe, another layer can normalize context, and a decision model can choose among permitted actions. Speech models, LLMs, and decision models can serve different roles in the same architecture. Code remains responsible for executing and validating the actions.

The benchmark we already had

Our most useful asset was a frozen held-out set of 70 synthetic English tickets, previously used for the deterministic baseline, Ollama, and Jev 1.13. We kept the tickets, labels, and metric definitions. Exact-tuple accuracy counts a ticket as correct only when category, priority, and risk all match.

We split the investigation into two questions:

Experiment Comparison Question
A — out of the box Historical Jev 1.13 × base Laya zero-shot What happens without domain-specific training?
B — domain adaptation The same, unchanged Jev result × adapted Laya How much can specialization close the observed gap?

We did not rerun Jev. We retained its frozen historical result through OpenRouter: typesafe/jev-1.13, resolved to typesafe/jev-1.13-20260917.

Experiment B is asymmetric. Jev received no domain-specific training examples; Laya received 1,120 TRAIN tickets. The taxonomy is shared, but prompts and inference paths differ. Laya’s questions were adapted to its token budget before the freeze and held constant between A and B. We are evaluating generalization without adaptation and the capacity to specialize, with different exposure to the domain.

Experiment A: base Laya zero-shot

For this general checkpoint, zero-shot, on this taxonomy and frozen held-out set:

Metric Base Laya zero-shot
Category accuracy 84.29% — 59/70
Priority accuracy 31.43% — 22/70
Risk accuracy 62.86% — 44/70
Exact tuple 11.43% — 8/70

The clearest pattern was 40/40 expected LOW priorities predicted as MEDIUM. Before attributing that to the model, the audit examined options, order, indices, mapping, serialization, and normalization. It found no integration bug that invalidated the result: the archived probabilities placed MEDIUM first, matching the published choice.

The audit used parsed outputs with copied choices and probabilities; untouched native responses were not retained byte for byte. That evidence supports checking the mapping: MEDIUM was already the model’s choice, with no LOW-to-MEDIUM conversion in the adapter.

Why this does not contradict Laya’s published benchmark

The material that prompted the investigation showed much stronger results. The laya-typed-decisions model card explains the key difference: that checkpoint was fine-tuned on the corresponding benchmark’s workflows. The card distinguishes the specialized checkpoint from the base model and notes that its Jev figures were published by third parties, not measured in the same run.

The published metrics belong to a different dataset, protocol, and checkpoint, and measure individual decisions rather than our per-ticket tuple. That distinction raised a second question: was weak zero-shot quality structural, or was this model missing adaptation to our domain?

Experiment B: keep learning separate from evaluation

We created 1,120 TRAIN tickets and 280 VALIDATION tickets, also synthetic, from the operational taxonomy. A deterministic generator used authored scenarios, without an LLM and without reading the held-out file or its errors to construct examples.

TRAIN — 1,120 tickets → changes weights
VALIDATION — 280 tickets → selects checkpoint
HELD-OUT — the same 70 tickets → final evaluation after the freeze

The held-out set stayed outside dataset generation, training, validation, checkpoint selection, thresholds, calibration, and hyperparameter tuning. This was a follow-up on a benchmark whose aggregate result we already knew; its file and errors were isolated from construction of the new data.

VALIDATION shares scenario families with TRAIN, although the wording is distinct. It supports checkpoint selection, but is not an independent real-world population.

Fine-tuning started from the same base checkpoint as A. Our deterministic labels became one-hot targets, unlike the soft distributions from a probabilistic teacher in Laya’s official example. We did not fit temperatures afterward: candidate checkpoints used temperature 1.0. This is not an exact replication of the external benchmark recipe, and the difference may affect confidence.

The selection rule favored the highest VALIDATION exact-tuple accuracy, with predeclared tie-breakers based on macro-F1, mean Brier, and earliest epoch. Epoch 4 was selected. Freeze 2, commit 00a9a64, fixed data, checkpoint, code, inference, and thresholds before the single official adapted evaluation: 70/70 valid responses and zero failures.

The infrastructure needed to adapt the model

The local host offered four vCPUs, roughly 4 GB of RAM, and no GPU. Full fine-tuning there was not a reasonable option. We used a temporary NVIDIA RTX A5000 with 24 GB VRAM, Python 3.11, and PyTorch 2.5.1 with CUDA 12.1.

Training ran for four epochs with mixed precision — fp16 autocast and fp32 master weights — and gradient accumulation to manage memory use.

This supplied concrete evidence of my direction in ML Systems / AI Infrastructure. The task included checking whether training accounting matched the updates that actually happened.

In the first run, GradScaler skipped an optimizer update while the scheduler advanced anyway. We detected the anomaly, changed the scheduler to advance only after applied updates, preserved the initial run’s artifacts, and repeated training before opening the adapted held-out evaluation.

The corrected official run recorded 420 update attempts, 413 applied optimizer updates, seven GradScaler skips, and 413 scheduler steps. It became the candidate on methodological validity, not because its accuracy was more convenient. Selection, freeze, and final evaluation followed that correction.

The jump: 11.43% → 85.71%

On the same held-out set:

Metric Laya zero-shot Adapted Laya
Category accuracy 84.29% — 59/70 97.14% — 68/70
Priority accuracy 31.43% — 22/70 94.29% — 66/70
Risk accuracy 62.86% — 44/70 92.86% — 65/70
Exact tuple 11.43% — 8/70 85.71% — 60/70

LOW → MEDIUM errors fell from 40/40 to 0/40. Adaptation clearly changed behavior in this domain, as documented in the public final report.

Fine-tuning did not fix everything. Two categories, four priorities, and five risk labels remained wrong. Errors overlapped across fields, leaving ten incorrect tuples. HIGH/CRITICAL priority recall decreased from 13/14 zero-shot to 12/14 after adaptation. Improving aggregate accuracy does not guarantee improvement in every severity error.

Jev and adapted Laya: read the asymmetry carefully

Metric Jev 1.13 — none of our TRAIN tickets Adapted Laya — 1,120 TRAIN tickets
Category accuracy 100% 97.14%
Priority accuracy 98.57% 94.29%
Risk accuracy 95.71% 92.86%
HIGH/CRITICAL priority recall 100% — 14/14 85.71% — 12/14
HIGH risk recall 85.71% — 6/7 100% — 7/7
Exact tuple 94.29% — 66/70 85.71% — 60/70

Jev showed strong out-of-the-box generalization on this benchmark. Laya showed how much control open source offers when we can adapt a model to the domain. Jev’s 94.29% came without our TRAIN tickets; Laya went from 11.43% zero-shot to 85.71% with adaptation. Adapted Laya had higher HIGH-risk recall, but the denominator is just seven examples.

Confidence remained a problem

Adapted Laya’s mean answer_confidence — its rounded top-option probability — was nearly saturated:

Field Mean answer_confidence Multiclass Brier Top-label ECE
Category 0.99986 0.05715 0.02844
Priority 0.99997 0.11429 0.05711
Risk 0.99884 0.13841 0.07027

Ten of the 70 tuples were still wrong. Brier and ECE describe this small sample; they do not certify calibration on future traffic.

The exploratory gate required all three answer_confidence values to meet the threshold. From 0 through 0.90, it retained 70/70 tickets at 85.71% exact-tuple accuracy. At 0.95 and 0.99, it retained 69/70 at 86.96%. It barely filtered errors. No threshold was selected or changed using these results.

This returns to the caution in the Jev article: high confidence does not automatically mean a reliable probability of correctness. Jev reports its own confidence; Laya distinguishes answer_confidence from an entropy-based confidence. Those different semantics prevent a direct comparison of the values.

The gap between learning and generalization

Split, selected checkpoint Tickets Exact tuple
TRAIN 1,120 100%
VALIDATION 280 92.50%
HELD-OUT 70 85.71%

The sequence 100% → 92.50% → 85.71%, TRAIN loss recorded as zero in the final epoch, and saturated confidence show a very strong fit to training data. Since VALIDATION still shares scenario families with TRAIN, the held-out drop deserves attention.

I interpret the perfect fit and split differences as signs consistent with memorization and overconfidence. The sample is too small to determine the severity of that effect or estimate a production generalization gap.

Local did not automatically mean faster

Model Measured path Mean latency per ticket
Base Laya Local CPU, four AMD EPYC vCPUs, fp32 1,414 ms
Adapted Laya Same CPU context, fp32 1,619 ms
Jev 1.13 Historical API/network call 569 ms

Laya timings include tokenization, inference over three questions, parsing, and validation; the per-request mean excludes model loading. Adapted held-out inference ran on CPU, not the training A5000.

On this specific CPU, self-hosted Laya was not automatically faster than calling Jev remotely. Since infrastructure and measurement dates differ, these timings describe the observed serving paths. Historical Ollama hybrid latency also measures a broader path than standalone inference.

Open source does not mean zero cost

Jev’s historical run reported US$ 0.003789618 for 70 requests. That is API usage cost for that run, not a current pricing projection or a complete operating-cost estimate.

Laya inference incurred US$ 0 in external per-token/request charges. CPU/GPU, RAM/VRAM, electricity, cloud resources, and operations still matter. CPU cost was not monetized because the necessary rate and measurements were unavailable.

For training, the recorded Pod rate was approximately US$ 0.27 per GPU-hour, or US$ 0.28/hour with container disk. Measured tasks — base VALIDATION, the first anomalous run, and corrected training/evaluation — totaled 1,523.51 seconds, about 25m24s.

The report estimates US$ 0.1185 for those tasks at the rate including disk. That is not the Pod’s total bill: it excludes setup, provider startup, and idle time; total billed time and the invoice were unavailable. The official CPU evaluation added no GPU time.

Laya removes per-token charges while moving cost and responsibility into infrastructure and operations. Preparing data, investigating an anomaly, and managing artifacts also belong in the decision, even without a monetary measurement of engineering time.

The real trade-off: convenience and control

Jev demonstrated the value of a decision model that works very well without training specific to this domain. Laya demonstrated a different advantage: open weights let us change a model when zero-shot is insufficient.

Open source did not automatically give us better accuracy, lower latency, or better calibration. It gave us control. That meant taking responsibility for the dataset, training, GPU, checkpoint, calibration, versioning, deployment, observability, and operating costs. This experiment exercised data, training, and evaluation; deploying and operating the adapted service remain future work.

That is the engineering decision: how much is the ability to adapt behavior and host the model under our constraints worth, given the work required to sustain it?

Limitations

  • A synthetic, English-only HELD-OUT set of just 70 tickets from one domain; no real traffic or production validation.
  • One official run per protocol; Jev is a historical result, without repeated runs to measure variance.
  • Synthetic TRAIN and VALIDATION data; VALIDATION shares scenario families with TRAIN.
  • Different domain exposure, questions, and inference paths for Jev and Laya.
  • Different hardware and measurement dates; latency paths are not equivalent.
  • Different confidence semantics; calibration and coverage analyses are descriptive and based on small samples.
  • Untouched native responses were not preserved byte for byte for the zero-shot audit.
  • The roughly 843 MB adapted checkpoint is outside public Git. Its hash, size, metadata, and results are public, but the weights are not available for download there.
  • Total billed Pod cost was unavailable; the estimate covers measured tasks only.
  • Jev and Laya remain evaluation-only in Ops Triage. Neither was integrated into HybridPolicy because of these experiments.

What I would do in a real system

I would start with labeled real tickets permitted for that use, separating scenarios, sources, and time periods so VALIDATION does not merely rephrase TRAIN. I would set acceptance criteria by severity and error cost before opening a new final evaluation.

Calibration and thresholds would need separate, representative data, with human review and fallback evaluated as part of the policy. I would repeat protocols to measure variance and compare latency, throughput, loading, memory, and cost on the serving paths I actually intend to operate.

Before integration, I would version weights and configuration alongside data and code, verify the license and context limits, and prepare monitoring for drift and errors, plus rollback. Typed decisions simplify integration; reliability still depends on the surrounding system.

Evidence and reproducibility

All figures come from the public reports and artifacts in Ops Triage AI, inspected at commit e8c75f0. The links below pin that snapshot. Its README still summarizes only zero-shot; domain-adaptation evidence lives in the final report and artifacts.

Back to articles