Jevstral: a decision model on Ministral 3 8B
I wanted to understand how decision models work. For me, the best way to understand a model is to build one. So I built Jevstral, a decision model on Ministral 3 8B Base, an open-weight model from Mistral AI.[1]
A decision model reads a document and a set of questions. For each option of each question, it gives a probability. It does not write text. It does all of its work in one forward pass.
This post explains each part of Jevstral, in the order that the data goes through the model. Most parts have a small playground: change the inputs and see what the part does. The code, the weights and the benchmark results are public.[2]
Jevstral is a personal learning project. It is not affiliated with Mistral AI, or with TypeSafe, the company that makes Jev.
What's the hype?
Much software makes small decisions on text. Which team must handle this support ticket? Is this tool call safe to run without a person?
Classifiers are not new. A logistic regression or a fine-tuned BERT model can route support tickets well, and it is fast and cheap. Traditional ML works.
But a classifier knows only the labels that it saw in training. If the support department adds a new team, you must label new tickets and train the classifier again. If you also want to ask "Is this ticket urgent?", you need a second classifier. Large language models removed this limit. You can ask any question in plain words, with any options that the model has never seen before (zero-shot).
The problem is that a chat model answers with text. The software must parse the text, and the text can name an option that does not exist. When a chat model writes "I am 90 % sure", that number is also text. It is not a calibrated probability.
Language models have become much better at reasoning tasks. But when a model explains why it chose option B and not option A, the explanation is often not the real cause of the choice. In one study, a hidden hint changed the answers of reasoning models, but their explanations mentioned the hint in only about 25 % (Claude 3.7 Sonnet) and 39 % (DeepSeek R1) of these cases.[3] Other studies found explanations that justify an answer that the model had already chosen.[4] This is not true for all tasks: on math problems, the written reasoning often does change the answer.[5]
A decision model tries to keep the good part of each. Like a chat model, it gets the question and the options with the request. Like a classifier, it gives one probability for each option. Then software can use a threshold: if p ≥ 0.9, do the action automatically; if not, send the case to a person.
Why not just use a chat model?
Speed is the other reason. A chat model works in two steps. First, it reads the full prompt in one parallel pass. This is the prefill. Then it writes the answer one token at a time, and each token is one more pass. This is the decode. A decision model does only the prefill. A small head then reads the hidden states and gives one score for each option.
Thus a decision takes approximately the time to the first token of the backbone. And the answer is always one of the given options.
I wanted to know how such a model works inside. The rest of this post follows the data through Jevstral, one part at a time.
What's inside
Jevstral has the Ministral 3 8B text decoder, LoRA adapters in its 34 layers, and a pointer head. The decoder is frozen. Training changes only 46.7 million weights: 44.6 million in LoRA, 2.1 million in the head and 20,480 in five embedding rows. That is approximately 0.6 % of the decoder.[6]
The figure below follows one forward pass. The next sections explain each step.
The tokenizer cuts the text into tokens and gives each token an id, a whole number. Five reserved ids mark the parts of the row: <state>, <q>, <opt>, </opt> and <decide>. L is the number of tokens.
The input
The input is a state and one or more questions. The state is the document. Each question has a list of options. Each question becomes one row of tokens:
<s> <state> document <q> question <opt> option 1 </opt> <opt> option 2 </opt> ... <decide>
Five reserved tokens of the Mistral tokenizer mark the parts. I call them delimiters.[7] The code inserts them as token ids. A request with three questions gives three rows, and each row contains the full state. Thus one question cannot see another question.
Ministral never trained these five tokens: their embeddings were all zeros. So Jevstral learns their embeddings during training. They are the only part of the embedding table that changes.[8]
User text cannot make a delimiter. By default, the tokenizer changes the text <SPECIAL_22> into token 22, the <opt> token. Then a support ticket could contain a fake option border. Jevstral tokenizes all user text with split_special_tokens=True, so the text stays plain text.
Attention and the KV cache
The row goes through 34 decoder layers. In each layer, attention lets a token read other tokens. The attention is causal: a token reads only itself and the tokens before it.
The order of the row uses this rule. Each </opt> comes after the text of its option, so its hidden state holds the meaning of that option. <decide> comes last, so its hidden state can hold the state, the question and all the options. Click a row in the grid to see what each token reads.
| <s> | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| <state> | |||||||||||||||
| charged | |||||||||||||||
| twice | |||||||||||||||
| refund | |||||||||||||||
| <q> | |||||||||||||||
| Which | |||||||||||||||
| team? | |||||||||||||||
| <opt> | |||||||||||||||
| billing | |||||||||||||||
| ▸ </opt> | |||||||||||||||
| <opt> | |||||||||||||||
| shipping | |||||||||||||||
| ▸ </opt> | |||||||||||||||
| ▸ <decide> |
<decide> can read 15 of 15 tokens. It comes last, so it reads the state, the question and every option. That is why the head uses it as the query.
The rule has a cost. An early option cannot read the options after it, so the result can change a little when the order of the options changes. Training shuffles the options to make this effect smaller.
The rule also gives a free speed-up. No state token can read a question token, so the state part is the same in every row of a request. For a request with more than one question, Jevstral computes the state once and keeps its keys and values in a cache (the KV cache).[9]
LoRA
The decoder has 8 billion weights. To train all of them, the GPU must keep each weight, its gradient and two optimizer values. That needs too much memory. Also, the model can forget what it learned before.
LoRA (low-rank adaptation) freezes each weight matrix W and adds a small trained path next to it. The path has two thin matrices, A and B. For an input x, the layer gives:
A has rows and B has columns. Jevstral uses and . B starts at zero, so at the first step each LoRA layer gives exactly the output of the base layer.[10] After training, is added to W. The result has the same shape as W, so inference has no extra cost.
Jevstral puts LoRA on all 7 projections of each layer: q, k, v and o in attention, and gate, up and down in the MLP.
The pointer head
After the last layer, each position has a vector of 4,096 numbers. For K options, we need K scores.
A simple head gives the <decide> vector to a fixed layer with one output for each option slot. But then the number of options has a maximum, and the last slots get almost no training. Jevstral uses a pointer instead. Two linear layers, q and k, change each vector into 256 numbers. The score of each option is a dot product:
This is the calculation of one attention head: <decide> is the query, and each option is a key. The same k scores every option. So the number of options has no maximum, and the score comes from the text of the option, not from its position.
A softmax changes the scores into probabilities. In training, the loss is −log of the probability of the correct option. In use, software compares the top probability with a threshold. Try both modes:
Move T. The bars become flatter or sharper, but the blue option never changes: division by a positive number keeps the order of the scores.
Calibration
A model is calibrated if, of all the cases where it gives 0.8, it is correct in approximately 80 %. After training, Jevstral was overconfident on data that it did not train on: it was wrong more often, but it stayed just as sure.
The correction is one number, the temperature T. The model divides all scores by T before the softmax. A T above 1 makes the probabilities less sure. For example, with the scores 4, 1 and 0:
| billing | shipping | returns | |
|---|---|---|---|
| T = 1 | 0.94 | 0.05 | 0.02 |
| T = 2.04 | 0.73 | 0.17 | 0.10 |
Division by a positive number keeps the order of the scores. So the top option never changes, and the accuracy cannot change. You can see this with the T slider in the playground above.
How to measure the error: ECE
The expected calibration error (ECE) measures how far the confidence of the model is from its accuracy. A reminder of how it works:
- Put the predictions in 10 groups by their top probability: 0 to 0.1, 0.1 to 0.2, and so on.
- In each group, compare the mean top probability with the fraction of correct answers.
- Calculate the mean of these differences, weighted by the size of each group.
0 is perfect. For example, a group where the model gives approximately 0.9 but is correct in 70 % of the cases has a difference of 0.2. An ECE of 0.14 means that the confidence is, on average, 14 points away from the accuracy.
The result
I ran the trained model on 648 questions from datasets that no training stage uses. Then I selected the T with the lowest log loss.[11] The result is T = 2.04.
On these questions, the ECE went from 0.141 to 0.043, and the accuracy stayed at 0.679. Two limits apply. First, I measured the ECE on the same 648 questions that I used to select T, so 0.043 is optimistic. A fair test needs a second held-out set. Second, the accuracy of 0.679 is lower than on the development sets (0.71 to 0.89), because these questions come from other datasets.
Training
Training has four stages.[12] Stage 1 teaches the format and the basic decision skill. Stage 2 teaches date arithmetic, and what to do when the document cannot decide. Stage 3 teaches long, real documents. Stage 4 teaches long policies, trade-offs, multi-hop reasoning, judging and developer tools. Stages 2, 3 and 4 also replay records from stage 1, so the model keeps its earlier skills.
The model must learn the meaning of each option, not its position or its words. So training changes the records. It shuffles the options. It sometimes adds a "none of the above" option, with or without the correct option. It sometimes adds an unrelated option, for example "A recipe for pancakes".[13]
The four stages took approximately 3 hours and 25 minutes on one H100.[14]
Stage 4 trains the hardest skills. Between stage 2 and stage 4 (these sets were not measured after stage 3), the accuracy on hard-v1 goes from 0.517 to 0.810. On devtools-v1, it goes from 0.563 to 0.710. With replay, decision-v7 stays at 0.88 through all four stages. I did not test training without replay.
Results
The Decision Index 0.2.1 is a public benchmark for decision models. It has 38 benchmarks in five areas. Each score is chance-corrected: 0 is random guessing and 100 is perfect. I ran the full suite, 150,317 requests, on one NVIDIA RTX PRO 6000. All requests were answered, with no errors.[15]
Jevstral scored 34.3. If my run were on the leaderboard of 1 October, it would rank 36 of 74 (the submission is in review). Ahem: this model was trained on the $30 of free Modal credits. My aim was not the top of the benchmark. I wanted to see what this architecture can do. Jevstral is inspired by Kev, and Kev 4B scored 34.6. The median latency was 28.3 ms for one request.[16]
Jevstral is in the middle of the table. It is better than every model smaller than 4 B. But models of 8–10 B score from 7.4 (CLM 8B) to 46.9 (JPT-9B), so Jevstral is in the lower half of its size class.
It is fast, but at least three 4 B models are both faster and better: Decider 4B (40.7 at 12.6 ms), Hopper 1.2 (40.8 at 23.3 ms) and JevK5 (38.8 at 22.0 ms). The table below compares Jevstral with Kev and Jev in each of the five areas.
| Area (weight) | Jevstral 8B | Kev 4B | Kev 9B | Jev |
|---|---|---|---|---|
| Knowledge and reasoning (25.8 %) | 24.5 | 22.9 | 26.2 | 51.4 |
| Language understanding (25.8 %) | 34.6 | 35.3 | 41.7 | 62.0 |
| Retrieval and classification (20 %) | 39.3 | 41.0 | 43.7 | 55.4 |
| Tools and automation (18.3 %) | 52.9 | 52.6 | 54.5 | 75.1 |
| Arts and human taste (10 %) | 14.8 | 17.9 | 22.4 | 37.7 |
| Decision Index | 34.3 | 34.6 | 38.5 | 57.9 |
The gap to Jev (57.9) is large in every area. Other open models show that most of this gap comes from training data, not from model size. I think that this is also true for Jevstral: it has an 8 B backbone, but with the same data it has the quality of a 4 B model.
Acknowledgements
The method, the training recipe and the training data come from Kev. The design follows the analysis in Jev's Architecture Unmasked. The benchmark and its harness come from the Decision Index. The base model comes from Mistral AI, under Apache 2.0.
Jevstral is a personal learning project. It is not affiliated with Mistral AI, with TypeSafe (the company that makes Jev) or with the author of Kev.
Notes
- Mistral AI publishes the model as mistralai/Ministral-3-8B-Base-2512, under Apache 2.0. Jevstral uses only its text decoder. ↩
- Code: github.com/AmirBraham/jevstral. Weights: AmirBraham/jevstral-8b. Full benchmark run: AmirBraham/jevstral-decision-index. All under Apache 2.0. ↩
- Chen et al., Reasoning Models Don't Always Say What They Think (2025). The hints were placed in MMLU and GPQA questions. The percentages are averages over six types of hint. ↩
- Turpin et al., Language Models Don't Always Say What They Think (NeurIPS 2023): when the example prompts always had the answer "(A)", the models chose (A) more often, and their explanations never mentioned this pattern. Accuracy fell by up to 36 points. Arcuschin et al., Chain-of-Thought Reasoning In The Wild Is Not Always Faithful (2025): some models answered "yes" to both "Is X bigger than Y?" and "Is Y bigger than X?", with a plausible explanation for each answer. The rate goes from about 13 % for GPT-4o-mini to below 1 % for recent reasoning models. ↩
- Lanham et al., Measuring Faithfulness in Chain-of-Thought Reasoning (2023): when the written reasoning is cut or changed, the answer changes most often on math tasks. Sprague et al., To CoT or not to CoT? (ICLR 2025): written reasoning improves accuracy mainly on math and symbolic tasks. ↩
- LoRA: 44,564,480 weights. Pointer head: 2,097,664. Delimiter rows: 5 × 4,096 = 20,480. Total: 46,682,624. The vision encoder of Ministral is not used. ↩
- The ids are 20, 21, 22, 23 and 26, for
<state>,<q>,<opt>,</opt>and<decide>. Mistral gives these tokens no role. ↩ - Five zero rows are all the same, so the model cannot tell the five delimiters apart. So stage 1 starts each row as a random sample from the distribution of the real embeddings, with their mean and covariance. The other 131,067 rows stay frozen. Norms at the start of stage 1: <state> 0.3633, <q> 0.3526, <opt> 0.3643, </opt> 0.3705, <decide> 0.3801. After stage 1: <state> 0.3774, <q> 0.3641, <opt> 0.3772, </opt> 0.3753, <decide> 0.3814. Ordinary tokens: mean 0.3646, standard deviation 0.0442. ↩
- On a sample of 100 benchmark requests, the cache halved the p95 latency. It made requests with one question slower, so Jevstral uses it only when a request has more than one question. ↩
- The code computes , not . has only numbers, so the extra cost is approximately 0.8 % for q_proj. The product B A is a full matrix and costs much more. ↩
- I tried T from 0.50 to 5.00 in steps of 0.01. Log loss is a proper scoring rule: only honest probabilities give the lowest value. The fit uses held-out data because, on training data, the model is correct most of the time, so high confidence costs nothing. A fit there gives a T that is too small. ↩
- All training data comes from the public Kev decision suites, at pinned revisions. Before training, the code checks each file against its SHA-256. ↩
- "None" with the correct option removed: 10 % of questions. "None" with the correct option kept: 12 %. Unrelated option: 15 %. Stage 1 also adds minimal pairs: two copies of one question that differ in one option. One copy has the correct option, and the other does not. ↩
- Stage 1: 70 minutes. Stage 2: 10. Stage 3: 45. Stage 4: 80. Calibration: 5. Settings: AdamW with weight decay 0.01, a one-cycle learning rate with 10 % warm-up, gradient clipping at 1.0, fp32 weights with bf16 computation. All GPU work ran on Modal. Training and checks cost approximately $29, and the benchmark run approximately $12. ↩
- The run took 3 hours 21 minutes with the public harness. Two overlaps: BANKING77 is in the benchmark and in the training data, and 200 MMLU-Pro questions were in the calibration set. The temperature does not change any answer. ↩
- p95 latency: 185.1 ms. Mean: 79.3 ms. One request at a time. Jev is a hosted API, so its latency includes the network. ↩