Jeff: 0.8B Decision Models Trained at Home
A developer using the handle firelex released "Jeff", a set of very small decision models that do zero-shot classification in about 22 milliseconds. The project appeared on Hacker News and gathered 367 points and 145 comments in its first few hours. It is built on Qwen3.5 (0.8B and 2B) and Gemma 4 (E2B), fine-tuned to return calibrated probabilities for options you describe in plain words.
The pitch is simple: you describe a situation and list the options in text. Jeff returns a probability for each option from a single forward pass. No generated text, no parsing. Support queues, user intents, moderation labels, voice commands, game moves. Your categories do not need to appear in the training data. You describe them, and Jeff picks.
The benchmark results are strong for the size. On an overall score across five public benchmarks, Jeff-Qwen3.5-2B hits 83.1%, matching Jev (the larger model it is based on) at 83.0%. On Financial PhraseBank, Jeff scores 96.3% versus Jev's 77.0%. On RAGTruth, Jeff-Qwen3.5-2B ties with AutoJev-27B at 88.9%. The model does not win on reasoning-heavy benchmarks like BBH and JudgeBench, which is expected at 0.8B to 2B parameters.
The fun part is the game tests. They had Jeff play Doom, Frogger, and Pac-Man zero-shot. Each turn, the code describes the situation and legal moves in words, and the model picks one. Jeff-Qwen3.5-0.8B matched a hand-coded rule bot on Doom kills (6.55) and Frogger crossings (10.3), and scored 57 pellets in Pac-Man. Not bad for a model that decides in 29 to 49 ms per move on an M4 Max.
Everything was trained on local hardware. One RTX PRO 6000 workstation GPU for training (the 0.8B trains in about 2 hours, the 2B in about 3.5). Synthetic training data written by Qwen3.8-Flash-Next on two DGX Sparks. No cloud GPUs, no closed-model output in the training data. A closed model was used only to spot-check the quality of a sample of the synthetic data.
The practical takeaway: if you need fast, calibrated decisions between described options and do not need multi-step reasoning, a 0.8B model fine-tuned for the task is enough. And if zero-shot is not good enough, a short fine-tune on your own examples takes you much further. Their voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.
Source: GitHub (firelex/jeff) [1]
Hacker News discussion (367 points, 145 comments) [2]
Model on Hugging Face [3]