< bergheim > dropped a link
Jeeves – Reasoning improves Jev-like decision models
https://github.com/PostHog/jeevesJeeves, a 9B decision model from PostHog built on Qwen3.5-9B, argues that adding a reasoning step significantly improves the accuracy of "Jev-like" classifiers without sacrificing their calibrated probability outputs. By training with SFT and CISPO reinforcement learning, the model achieves a 0.889 accuracy on held-out test data, outperforming both Kev-9B (0.822) and the original Jev (0.857). It also leads on JevBench's public tiers with a score of 0.935 compared to Jev's 0.866. The system uses a pointer head and a block-4 diffusion drafter to speed up inference, reaching a median latency of 3.3 seconds per request on an H100 GPU when thinking is enabled.
The technical tradeoff is clear: while Jeeves excels at out-of-domain tasks and rule-based logic (scoring 1.000 on held-out structures), it lags behind Jev on pure knowledge retrieval, such as MMLU (0.793 vs 0.900) and MMLU-Pro (0.739 vs 0.840). Additionally, the "thinking" process introduces tail latency, hitting 17 seconds at the p90 mark, though this can be mitigated by truncating the reasoning chain or using a confidence threshold to skip thinking for easier queries. The project provides full training code and weights, allowing users to serve the model via a Jev-compatible API on CUDA or Apple Silicon hardware.
what others said
- Hacker News 237 points · 92 comments