Open-source decision models Jeff-Qwen3.5-0.8B and Jeff-Gemma4-E2B run in ~30ms on local hardware
Jeff is a set of small, fast decision models fine-tuned from Qwen3.5 and Gemma 4, designed for zero-shot classification tasks. The 0.8B-parameter model runs in about 22ms on an RTX PRO 6000 and 28ms on an Apple M4 Max, returning calibrated probabilities for predefined options. It uses the same request format as Jev but is not affiliated with it. The project emphasizes speed, calibration, and…
Key points
- Jeff-Qwen3.5-0.8B and Jeff-Gemma4-E2B run in ~22ms (RTX 6000) and ~28ms (M4 Max) for zero-shot decisions
- Fine-tuning on ~11k examples boosted accuracy from 31.7% to 95.8% in 30 minutes on one GPU
- Models use local hardware only, with synthetic data generated by Qwen3.8-Flash-Next on DGX Sparks
The models excel at fast, well-calibrated choices between options—such as routing customer service calls or moderating content—without multi-step reasoning. Benchmark tests show they approach or surpass Jev’s performance on classification tasks but lag in reasoning-heavy scenarios like BBH or JevBench. Fine-tuning on domain-specific data (e.g., voice navigation) can drastically improve accuracy: a 31.7% to 95.8% jump in held-out accuracy was achieved in under 30 minutes on one GPU. The project releases weights under Apache 2.0 and code under MIT, with training data sources listed separately.
Model pages: Jev → · Jeff-Qwen3.5-0.8B →
The story so far
2 episodes →- Open-source decision models Jeff-Qwen3.5-0.8B and Jeff-Gemma4-E2B run in ~30ms on local hardwarethis story
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
github.com · 28 September 2026Loading the full article…
This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Hacker News discussion · 59 pointsnews.ycombinator.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Google DeepMind’s AI model adds five days to hurricane warnings · 1 src
- OpenAI claims new model solved Navier-Stokes and 100+ math problems · 3 src
- ConTP uses contrastive learning to predict transporter substrate specificity · 1 src
- German researchers introduce ReImaGin for image-based chain-of-thought reasoning · 1 src
- OpenAI contractors fired for using banned AI tools on ChatGPT review tasks · 1 src
Comments
via GitHub Discussions