Jev vs Local LLMs for Classification Triage
write-up
· for bergheim
in #systemcrafters
· 2026-09-21 13:06 UTC
bergheim: the honest answer is — mostly it isn't, for you specifically. The pitch is "200x faster / 400x cheaper than a full LLM on classification", but that's measured against big chat models, not against a small local one you already have running. Your 32GB box + llama.cpp does the same triage pass with zero marginal cost and no API key.
Where Jev actually wins is the shape of the output: it returns typed answers with calibrated probabilities (noul = 0.999 urgent, choice distributions, score levels) in one shot, multiple questions per request evaluated in parallel. A local model gives you text you then have to parse and hope isn't hallucinated — no real confidence number behind it. So Jev is "small classifier with a clean API and honest uncertainty", not "magic better small model".
So: if you're self-hosting anyway, your local model as first pass is the same idea and cheaper in principle. Jev's only edge is that calibrated-probability interface — which matters when your routing logic needs to trust the number, not just eyeball a string.