I needed a model for one narrow decision. When a new fact about someone comes in, does it replace something already known? "I moved to Berlin" should retire "I live in Pune". "Watched two episodes of The Bear" should leave "works at Infosys" alone.
It sounds trivial, but any assistant with long-term memory makes this call all the time, and both kinds of mistake hurt. Miss an update and it keeps repeating stale facts. Replace too eagerly and it throws away things that are still true.
This post covers the models I compared, how I fine-tuned the best starting point, how I tested it, and what it took to train. The short version: a 0.8B open model, fine-tuned on 6,000 examples on a single T4, beat every hosted LLM I tested at this one job.

Why a System-One model
The obvious approach is to prompt an LLM and parse what comes back. That works, but it's slow, every decision is an API call, and the answer is free text that you then have to interpret.
TypeSafe's Jev takes a different approach. It's a "System One" model: instead of writing text, it answers typed questions with calibrated probabilities, like a choice from a list or a yes/no. For this task I send it the new fact and up to eight existing ones, and ask two questions about each existing fact:
How does the new fact relate to it: duplicate, updates, elaborates, related or unrelated?
Is the existing fact now outdated?
The score I use is the average of the two: p_update = (P(updates) + P(outdated)) / 2. At 0.70 or above, the old fact is replaced.
The candidates
Jev (hosted, by TypeSafe). Solid, but tuned for general decisions rather than my definition of an update.
Laya, an open 421M ModernBERT model that answers in Jev's format. Fast and light, but off the shelf it replaced too many everyday events.
Kev, Jared Palmer's open-weight model, compatible with Jev. I tried two sizes. Kev-0.8B was the best of everything as released on stated updates. Kev-4B scored lower on that, and on the T4 it ran out of memory on longer requests and took 5.5 seconds per decision.
For reference I also ran a zero-shot NLI model (DeBERTa-v3-large) and several hosted LLMs, each prompted with the same definition of an update. Kev-0.8B was the clear starting point: the best as released, and the only size that both trains and runs on a 16 GB T4.
The data
Each training example is one request: a new fact, up to eight existing facts, and the typed questions, with a label for every question. I built them from two pools:
6,818 logged Jev decisions. Training on these is distillation: Kev learns to make Jev's calls on real requests.
10,219 decisions with ground truth. I simulated 48 lives of a year each. People move city, change jobs, partners and managers, switch gyms, cars, courses, diets and laptops, reschedule appointments, mention crucial facts and make small talk. Because the simulator knows what actually changed, every pair has the right answer: same attribute with a new value is an update, the same value is a duplicate, anything else is related or unrelated. That's 53,870 fact pairs, 2,298 of them real updates.
The 6,000 training examples took every decision that contains a real update first, then filled up from the shuffled mix. Nothing from the test data went in.
Stage 1: supervised fine-tuning
LoRA on Qwen3.5-0.8B-Base, starting from Kev's released weights: rank 16, alpha 32, dropout 0.05, on every linear and linear-attention projection, plus a small decision head. One epoch over the 6,000 examples at a learning rate of 5e-5, with label smoothing of 0.05.
The GPU was a single NVIDIA T4 with 16 GB. It has no bf16, so everything ran in fp32, with batch 1, gradient accumulation of 8 and gradient checkpointing. The requests are short (87 tokens on average, 138 at the 95th percentile), so nothing was truncated. That came to 750 optimizer steps and 5.52M tokens in 6.55 hours, peaking at 5.09 GB of GPU memory.
Stage 2: reward-ranked RL (ReST)
Supervised training teaches the format and the definition, but it doesn't know which mistakes matter more. So I let the stage-1 model make the real decision, replace or keep at the 0.70 threshold, on 2,000 training decisions (10,574 pairs), and scored every pair:
+1 for catching a real update
−1 for missing one
−2 for replacing a fact that was still true
+0.25 for correctly keeping a fact
The mean reward was 0.265. There were 15 wrong replacements and 65 missed updates, spread over 52 decisions with a negative reward. Those 52 were relabelled with the truth (repeated twice when they included a −2) and mixed with replayed stage-1 examples, for 366 records in total. A second, gentler pass at a learning rate of 1e-5 took 46 steps, 0.34M tokens and 23 minutes, peaking at 4.98 GB.

I picked the final checkpoint on six separate validation lives (400 decisions), where it scored 0.9991 AUC against 0.9983 for stage 1. The test sets were never used to choose anything.
How I tested it
Two held-out tests, both kept out of training:
A two-year simulated life with template text: 122 decisions, 428 pairs, 186 updates, 45 of them stated outright.
A live run where an LLM wrote the facts from a three-month simulated life, so they read more like real conversation: 127 decisions, 498 pairs, 104 updates, 36 stated. This is the harder and more realistic one.
A stated update says it directly ("I moved to Berlin"). An implied one doesn't, like a new address in another city with no mention of moving. I report both.
The main metric is ROC AUC. Take one real update and one non-update at random: AUC is the chance the model scores the real update higher. 0.5 is a coin flip and 1.0 is perfect. It isn't accuracy, so 0.954 doesn't mean "right 95% of the time". I also check precision and recall at the 0.70 threshold, because that's where the model actually acts.
Results

On the live test, stated updates went from 0.752 as released to 0.929 after stage 1 and 0.945 after stage 2. For any update, including implied ones, it went from 0.708 to 0.873 and 0.872. Jev scores 0.781 and 0.570 on the same test. On the two-year life the stated score went from 0.952 to 0.962, and for any update from 0.583 to 0.840.
At the 0.70 threshold on the live test, precision was 1.00: it never replaced a fact that was still true. The cost is a recall of 0.14, so it only replaces when it's very sure. Kev and Laya as released never crossed the threshold at all. For this job, a cautious model is the right default, and a lower threshold chosen on validation data recovers more of the misses.
Averaged over both tests on stated updates, with every hosted LLM given the same definition:
my fine-tuned Kev (kev-mem): 0.954
Gemini 3.7 Flash: 0.892
Gemma 4 31B: 0.887
GPT-5.4 mini: 0.873
Kev as released: 0.852
Jev: 0.834
DeBERTa-v3-large NLI, zero-shot: 0.815
Laya: 0.711
Size and cost
The model has about 885M parameters: 873.4M in the base, 10.8M in the LoRA adapter and 0.5M in the decision head, so only 11.3M were trained. The adapter is 43 MB. It runs in about 4 GB of GPU memory and takes 0.81 seconds per decision on the T4, without any of the optimised kernels. Counting the baseline tests and a second stage-1 run, the whole project used about 12 GPU-hours.
Caveats
This is my use case and my definition of an update. Jev and Kev are general-purpose models and the hosted LLMs were only prompted, so a lot of the gain comes from specialising. That's the point, but it isn't a general ranking.
These are single training runs, with no confidence intervals yet.
Jev answered 234 of the 249 test decisions; the rest failed after retries.
The test lives are simulated, apart from the LLM-written text of the live run.
Try it
kev-mem-0.8b on Hugging Face, served through Kev's own server, so any Jev client works with it
kev-mem on Kaggle and the benchmark notebook
The memory-update detection benchmark for hosted LLMs
Thanks to TypeSafe for Jev, Jared Palmer for Kev, the Laya authors for open-sourcing their work (Apache-2.0), and the Qwen team for Qwen3.5.
