JEV Concept Overview and Comparison of Two Open-Source Implementations

JEV is the name TypeSafe AI gives to its “System One” model concept. The company describes it as a function-call-like model: unstructured state goes in, and typed probabilistic decisions come out. This post reviews the theory behind that idea and compares two open-source implementations. TypeSafe has not disclosed many design details, so the descriptions here are based on the company’s public blog post and the two repositories’ READMEs.

Category : Deep dive

I wanted to know whether this approach could be used for personal information detection, so I compared two open-source implementations of the JEV concept. I also kept in mind privyscope, a personal information detection engine I built.

Origin of the name and the concept

The company says the name comes from Daniel Kahneman’s System 1, the fast and intuitive mode of thinking, as opposed to the slow, deliberate System 2. JEV is meant for repeated, structured judgments inside software systems, such as which team should receive a request or whether a customer qualifies for a refund.

The company also cites response times between 70 ms and 500 ms. These are the company’s own claims, and we have not found independent verification. It also says the model does not hallucinate. Whether that means the output format is guaranteed or the content is guaranteed to be accurate is a separate question.

Theory 1: judging instead of generating

Most language models work auto-regressively, predicting one next token at a time. That suits long text, but for problems where the answer is one of a few options or a score on a few levels, it costs more computation than necessary. When the output format is fixed in advance, the model can score the options in one pass instead of writing sentences.

Theory 2: typed outputs

JEV fixes the shape of each output as a type. There are three basic forms.

  • Yes/no (Noul): returns the probability p(true) that a statement is true. Several statements can be asked independently.
  • Choice: returns a probability distribution over only the options supplied at call time. An answer outside those options cannot be produced.
  • Ordinal (Score): returns a distribution over fixed levels and its expected value.

Type constraints prevent malformed answers. They do not guarantee that the answer is correct. The kyegomez/open-jev README makes this point directly.

Theory 3: calibrated probabilities

A probability is useful only if its values match observed frequencies. If a model says 0.8 across many cases, about 80% of them should turn out to be correct. Calibration is measured with the Brier score or the expected calibration error (ECE). Raw logits are often miscalibrated, so a temperature is usually fitted on separate calibration data. Open-Jev also fits its temperature after training, using only calibration data.

Calibration can break again when the input distribution shifts. The kyegomez/open-jev README says calibration should be measured again under distribution shift before deployment.

Theory 4: training objectives

Training against a target probability distribution, instead of a single correct label, keeps ambiguous cases from being forced into false certainty. Both repositories follow this direction. Open-Jev uses soft cross-entropy together with the Brier loss. kyegomez/open-jev’s RLCDLoss adds terms for consistency (equivalent inputs should get the same answers), confidence, and ECE. Those terms are the author’s reconstruction from public material, not a disclosed TypeSafe design.

Three ways to build a JEV-style system

There are three broad ways to turn the JEV concept into a working system. The choice determines the data needed, the compute cost, and how the model is served.

  1. Head fine-tuning: attach a LoRA adapter and a decision head to a pretrained model and train them. It works with little data and a small GPU, but the model was originally trained for next-token prediction, so that trace remains in its representations.
  2. Training from scratch: design the state encoder and decision heads from the beginning and train them. This can remove the generation loop entirely, but it needs large data and compute, and it is hard to use existing serving tools.
  3. Knowledge distillation: use the probabilities from a large teacher model as targets to train a smaller student model dedicated to decisions. Because the teacher’s distribution is the target, the student can learn without ground-truth labels. The student is cheaper to run, but the teacher is still needed, and you must check whether the student also inherits the teacher’s calibration.

Mapped onto these categories, zefan-cai/Open-Jev is head fine-tuning. kyegomez/open-jev is a from-scratch design, but it currently implements only the structure, without trained weights. Distillation is not implemented in either repository. kyegomez/open-jev lists a pretraining and distillation pipeline as a todo item.

The same recipe appears in other domains. For example, privyscope attaches a token-level BIOES classification head to a pretrained encoder to locate personal information in text. The head is replaced and fine-tuned in the same way, but its output is a tag on each token. That is a different problem and output shape from JEV, which gives a probability for each question.

Two implementations compared

Itemkyegomez/open-jevzefan-cai/Open-Jev
PurposeA research reconstruction of the JEV concept from first principlesA family of decision models meant for practical use
BaseA bidirectional transformer encoder designed in-house, with random weightsPretrained Qwen3.5 2B and 9B and 27B checkpoints, with LoRA and a scalar decision head
OutputTyped heads: Noul (truth probability), Choice (option distribution), Score (ordinal distribution)Probabilities from Yes/No logits for each candidate
TrainingRLCDLoss (soft-target NLL, Brier, consistency, confidence, ECE)Soft cross-entropy and Brier, with a temperature fitted on calibration data
Public statusNo weights, no performance claimsCheckpoints released, with self-reported evaluations and limitations

kyegomez/open-jev: a structure-first reconstruction

This implementation encodes the state once, and question slots read that representation through cross-attention. Questions do not attend to each other, so several questions can be answered in one parallel forward pass. Answers are restricted to the declared types. For example, a Choice answer is always one of the options passed in the call.

As the repository says, this design is a hypothesis inferred from public material. TypeSafe’s actual architecture is not public. The weights are random, so the output values carry no meaning. What this repository lets you check is the structure and the interface.

zefan-cai/Open-Jev: adapting an existing LLM

This repository attaches a LoRA adapter and a scalar decision head to pretrained Qwen checkpoints. For each candidate it produces a probability from Yes/No logits. Candidates are processed as independent sequences during both training and inference. Shared prefix caching is still experimental, so the compute grows with the number of candidates.

The repository reports the following. In isolated operator benchmarks on an H200, the speedup was 2.63 to 2.75 times. For the whole model, the P50 ratio was only 1.006 to 1.020, and the authors state that statistical significance was not established. In a natural-language support pilot, the 2B model answered 166 of 256 correctly (64.8%), below BM25 at 211 of 256 (82.4%). There are no external user validations yet.

Limitations and what to verify

  • Format versus correctness: type constraints prevent malformed answers, but they do not guarantee that the answers are correct.
  • Information in outputs: because the output is a probability rather than free text, verbatim leakage is less of a concern. However, the probabilities themselves can reveal information about inputs and training data, so fine-grained outputs and repeated queries need controls.
  • Calibration over time: when the input distribution shifts, calibration can break, so it must be measured again on real data.
  • Speed claims: the company’s response-time figures have not been independently verified. The whole-model speed advantage of zefan-cai/Open-Jev was also not statistically significant.
  • Reproducibility: kyegomez/open-jev has no trained weights, so its decision quality cannot yet be compared.

Viewing it as a trade-off

Some people see JEV as a major technical advance, but I think it is better understood as a trade-off. Whether a machine or a person makes the judgment begins with a difference in output format, but it also changes how decisions are explained, how errors show up, and who is accountable. The gain in speed and format guarantees comes with a heavier burden of verification and calibration management.

Conclusion

What is new about JEV lies less in any single technique than in the combination of design goals. Outputs are typed decisions rather than generated text, calibration is part of the training objective, and the whole thing is designed to be called like a function inside a system. Each piece, such as multiple-choice logit scoring, task-specific heads, temperature calibration, and soft targets, was already known. The claims about speed and the absence of hallucination mostly follow from the single choice not to generate text, and whether that choice actually pays off has to be measured.

Even so, the core idea of JEV is to make probabilistic judgments without generating text. The two repositories start from the same concept but answer different questions. kyegomez/open-jev helps you understand and experiment with the structure. zefan-cai/Open-Jev provides checkpoints and measurements that you can refer to when adding a judgment capability on top of a local LLM. If you adopt this approach, measure accuracy, calibration, and latency on your own data first.

In the end, this is less a leap forward than a trade: faster, format-guaranteed judgments in exchange for more verification and calibration work.

Glossary

The terms used in this post are briefly explained below.

  • System 1 and System 2: Two modes of thinking distinguished by the psychologist Daniel Kahneman. System 1 is fast and intuitive. System 2 is slow and step-by-step reasoning.
  • Auto-regressive generation: A method that builds output one token at a time, each conditioned on the tokens already produced. Most LLM text generation works this way.
  • Token: The smallest unit a model processes. It can be a whole word, part of a word, or a single character.
  • Encoder: The part of a neural network that turns input text into a numeric vector.
  • Logits: The raw output scores of a model, before they are converted into probabilities.
  • Fine-tuning: Further training of an already trained model on a new task.
  • LoRA: A fine-tuning method that trains small added matrices instead of changing the whole model.
  • Head: A small layer attached at the end of a model that produces the final answer.
  • Calibration: How well predicted probabilities match observed frequencies. If the model says 0.8 across many cases and 80% are correct, it is well calibrated.
  • Temperature: A value that controls how sharp a probability distribution is. It is usually fitted on calibration data.
  • Soft target: A training target expressed as a probability distribution rather than a single correct label.
  • Knowledge distillation: Training a small student model to imitate the outputs of a large teacher model.
  • Distribution shift: When the data seen in deployment differs in its properties from the training data.
  • Prefix caching: Computing the shared beginning of several requests once and reusing it.
  • Regular expression (regex): A pattern rule that finds strings with a fixed format, such as phone numbers or resident registration numbers.
  • BIOES tagging: A scheme that tags each token as beginning, inside, end, single, or outside of an entity, so spans can be located.

This post is based on TypeSafe AI’s public blog post and the READMEs and published figures of the two repositories. We have not independently reproduced the results.

By Mark

-_-