
What is PostHog Jeeves?
Jeeves is a "reasoning Jev-style classifier with a diffusion drafter," released by PostHog on September 29, 2026. It was built by Nicholas Waltz, an AI research engineer on PostHog's AI Research team, who previously spent two years on a YC startup building RL models for drone delivery, per his PostHog profile.
To understand Jeeves you need to know what it's copying. A "decision model" doesn't write text. You send it a state (an email, a ticket, a JSON blob) and a set of typed questions, and it sends back a probability for each answer option. The category went mainstream two weeks earlier, when TypeSafe launched Jev with three question types: a yes/no noul, a pick-one choice, and a rubric score. Jev is closed and hosted. Within days, Jared Palmer's Kev shipped open Jev-compatible weights, and Jeeves credits Kev as its inspiration.
What Jeeves adds is thinking. The README frames the problem bluntly: Jev-like models "give calibrated decision probabilities, but at low accuracy," so "a lot of pipelines therefore rely on a reasoning model as a fallback." That fallback is usually a hosted LLM billed per token, the kind of cost I broke down in Claude Sonnet 5.5 pricing. Jeeves tries to be that fallback, in Jev's own shape.
The headline specs:
| PostHog Jeeves | |
|---|---|
| Base model | Qwen3.5-9B with LoRA and a pointer head (see Qwen alternatives) |
| Question types | noul (yes/no), choice, score, mixable in one request |
| API | Jev-compatible /v1/systemone, plus a drop-in Python SDK |
| License | Code MIT, weights Apache-2.0 (inherited from Qwen3.5) |
| Weights size | 21 GB bf16, 11.5 GB with FP8 linear layers |
| Hardware | CUDA GPU; Apple Silicon Mac with 48 GB+ for inference |
| Hosted API | None. Self-host only |
| Traction at launch | 241 points on Hacker News, 356 GitHub stars |
I'm Alicia, and I build the AI agents inside eesel, so I spend a lot of time on exactly this layer: the part of the pipeline that decides what a ticket is before anything gets written. Here's how Jeeves does it, and where it earns its spot.
How does Jeeves reason before it decides?
The mechanism is a neat bit of engineering, and it explains both the accuracy gain and the latency cost.
Jeeves loads the state and each question into the Qwen chat template using rare, mostly unused tokens as markers (<|fim_prefix|> for the state, <|box_start|> for each option, and so on). Then it does something Jev doesn't: it generates a free-text reasoning chain inside a <think> block. After the chain closes, the question and options are repeated, and a small pointer head scores each option by comparing the hidden state at a <decide> token against the hidden state at the end of each option. A softmax over those scores, divided by a temperature fitted on dev data (1.859, per the Hugging Face model card), gives the final probabilities.
Two details from the ablations are worth knowing if you ever train one of these yourself. Swapping the rare tokens for plain words like "State" made results worse. So did dropping the repeated question after the reasoning block. The model needs to be reminded of the options right before it commits.
Training ran in three stages, all documented in the repo:
- Supervised fine-tuning: 2 epochs, 596 steps on 8 GPUs, across 19,126 questions from 12 public datasets plus synthetic policy data. Half the questions carried a reasoning chain sampled from the base model.
- Reinforcement learning with CISPO: the RL method from MiniMax-M1, on 9,992 questions with 8 rollouts each, capped at 2,560 thinking tokens. The 624-step schedule was stopped at step 402, because past that point "the head over-sharpens" and calibration gets worse.
- Calibration: a single temperature fitted on dev and saved with the checkpoint.
That step-402 stop is the part I'd underline. The team gave up some raw score to keep the probabilities honest, which is the whole reason to use a decision model in the first place. A model that says 90% and is right 70% of the time is worse than useless in a routing pipeline, and it's the same root problem behind most AI hallucinations in support.
To claw back speed, Jeeves ships a diffusion drafter that proposes several reasoning tokens at once for speculative decoding. On one question it lifts chain speed from 109 to 176 tokens per second with block-4 drafting, and to about 960 tokens per second in total when eight questions are batched.
What does Jeeves beat Jev at, and where does it lose?
PostHog's results table is unusually candid, and the pattern in it is the most useful thing in the launch.

| Benchmark | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (held-out and out-of-domain) | 0.822 | 0.857 | 0.889 |
| JevBench overall (231 public items) | 0.715* | 0.866 | 0.935 |
| JevBench hard (111 public items) | 0.451* | 0.730 | 0.865 |
| PAWS (paraphrase detection) | 0.763 | 0.788 | 0.875 |
| Held-out rule structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
| MMLU | 0.738 | 0.900 | 0.793 |
| MMLU-Pro (10-way) | 0.515 | 0.840 | 0.739 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0.000 | 0.090 | 0.055 |
| JevBench calibration error (lower is better) | 0.049 | 0.037 |
*PostHog notes no Kev-9B JevBench result is published, so those cells are Kev-8B. All figures from the Jeeves README.
Read down the table and a split appears. Where the question is about applying rules to a messy input (held-out rule structures, contrastive policies, paraphrase detection, JevBench hard), thinking wins, sometimes by a lot. Where the question is about knowing a fact (MMLU, MMLU-Pro), Jev wins by about 10 points. Reasoning can't conjure knowledge a 9B model doesn't have. For knowledge, support teams usually ground the model in their own docs instead, the trade-off covered in RAG vs fine-tuning.
That split maps neatly onto support work. "Does this refund request fall inside our 30-day policy, given the order date buried in paragraph three?" is a rules question. "What's the capital of Peru?" never shows up in a ticket queue. If your decisions look like the first kind, the table is on your side.
Two honest caveats from PostHog itself: the Kev and Jev comparisons outside JevBench "use different items from the same sources," and the reasoning chains aren't very readable, because "no language consistency reward was included." You get the answer and a chain, but don't plan on showing that chain to an auditor.
How slow is "thinking" in practice?
This is where the Hacker News thread pushed back hardest, and fairly.

PostHog measured three settings on 325 dev questions, on one H100 with FP8:
| Setting | Accuracy | Mean reasoning tokens | Median / p90 latency |
|---|---|---|---|
| Full thinking | 0.825 | 1,138 | 3.3 s / 17.1 s |
max_think 768, nothink_threshold 0.9 | 0.806 | 344 | 2.0 s / 5.6 s |
| No thinking | 0.775 | 0 | about 0.3 s |
The middle row is the one I'd actually ship. max_think cuts each reasoning chain off at a token budget, and nothink_threshold skips thinking entirely when the quick answer is already confident. You keep most of the accuracy gain (0.806 vs 0.775) and cut p90 latency by two-thirds.
For comparison, Jev's own pitch is 70 to 500 milliseconds per call, which I covered in my Jev speed test. The HN crowd noticed:
"Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap."
That's true if you treat Jeeves as a Jev replacement. It's less true if you treat it as the thing that runs after the fast model shrugs. One commenter put the architecture plainly:
"Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work."
Jeeves collapses that hybrid into one model. Thinking off is your fast tier. Thinking on is your slow tier. And because both tiers share one API and one calibration, you don't have to reconcile a classifier's probabilities with an LLM's free-text answer.
What happens when you hand it a real support ticket?
The README's demo request is a support ticket, which made my job easy. The state reads: "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card." Three questions ride along: which department, should this escalate, and how frustrated is the customer.
Jeeves came back with billing at 0.46, returns at 0.40, shipping at 0.14, and a choice confidence of just 0.19. Escalation scored 0.72. Frustration landed at 1.5 on a 0-to-2 scale. The whole request took 8.1 seconds with the three questions thinking in parallel.
Here's the same ticket (plus one extra sentence) run through Kev-4B in Kev's playground:

Kev-4B said returns at 0.91 in 670 milliseconds. It's a smaller model and a slightly different prompt, so this isn't a clean head-to-head. But look at the ticket. It really does contain two problems owned by two teams. Kev answered fast and sounded sure. Jeeves thought about it and said, in effect, "this is close, and I'm not confident."
For a support queue, the second answer is the more useful one. A confident wrong route means the customer waits in the returns queue while their double charge sits untouched. A low-confidence split is a signal to do something smarter: route to billing and tag returns, split the ticket, or send it to a person with a clean AI-to-human handoff.

This matches what I've seen from eesel running AI on real queues. In one real-traffic trial on an e-commerce inbox, the decisions were the reliable part: 93% triage accuracy and 100% spam detection with zero false positives. The generated replies were where things slipped, with a 7% factual error rate in drafts. Classification is the half of support AI you can trust early, as long as the model admits when it's unsure. I've also watched confident-sounding bots quietly give wrong answers, which is why every eesel rollout gets simulated against historical tickets before it touches a live customer. A calibrated "I don't know" is worth more than a fast guess.
Who should actually use PostHog Jeeves?
Jeeves is a research release, not a product, and it reads like one, much like most of the open-source AI agents space: full training code, train/dev/test data, a reproduction script, no hosted endpoint. The model page on Hugging Face notes the model "isn't deployed by any Inference Provider."
Here's how I'd sort it:
| You are... | Reach for | Why |
|---|---|---|
| An ML team already running Jev or Kev, with an LLM fallback | Jeeves | Same API, one model for both tiers, better on rule-heavy edge cases |
| A team that needs sub-second decisions on every call | Jev or Kev | Single pass, no thinking tail |
| A team that wants open weights on a laptop | Kev-0.8B or Kev-4B, or Laya | Far smaller footprints |
| A team that wants to train its own decision model | Jeeves | The full SFT, CISPO and drafter recipe is in the repo |
| A support team that wants tickets triaged and handled | An AI helpdesk agent | A decision model doesn't reply, update fields or escalate |
Early community testing is encouraging for moderation-style work. One HN commenter ran it on their own data:
"Update: Jeeves took about 2 hours to moderate 394 data points and performed really well. It's not as good as Jev, but it's super close!"
Note the "2 hours." That's the thinking tax again on hardware that isn't an H100. A PostHog engineer, Robbie Coomber, replied in the same thread that an MPS port was in progress to speed up Apple Silicon, and the repo shows it has since merged.
If you're the one wiring this in, the full map of Jev alternatives is worth a skim first, because structured-output features on hosted LLMs might cover your case without a GPU.
What a decision model won't do for your helpdesk
Here's the shift I'd push on. A decision model answers "what is this ticket?" very well. It doesn't answer "so what do we do now?" Somebody still has to write the code that turns billing: 0.46 into a helpdesk tag, a reply, a refund lookup, an escalation with a note, and a record of why.
That's where the gap between infrastructure and an employee shows up. Jeeves, Jev and Kev are infrastructure: excellent, cheap-to-run building blocks for ticket classification, prioritization and sentiment. An AI teammate is the employee that makes the same calls inside your actual helpdesk and then does the work. If you're weighing the two, my guide to automating ticket triage lays out the build-vs-hire trade-off in more depth.
Try eesel

If you're looking at Jeeves because you want smarter triage, you probably don't want to host a 9B model and write the glue around it. eesel's AI helpdesk teammate joins your existing queue in Zendesk, Freshdesk or Help Scout, learns from your past tickets and help center, and makes the route-tag-escalate call on every ticket, following the escalation rules you already use. Then it drafts or sends the reply, and hands the truly unsure ones to a person with a note. Before it goes live, you can simulate it on hundreds of your past tickets and see exactly where it would have been right and wrong.
If the self-hosting appeal of Jeeves is really about control from code, the eesel CLI covers that. It runs the same teammate from your terminal: eesel activity lists every run and shows one in detail, eesel approvals lets you approve or deny actions that need a human, and eesel instructions edits the teammate's standing rules. Every command prints JSON, writes support --dry-run so you can see the exact call before it happens, and the workspace doubles as an MCP server, so coding agents like Claude Code can drive it directly. There's more on that in my AI agent CLI guide.
If you're comparing it with your helpdesk's built-in AI, the eesel vs Zendesk AI breakdown is a good next read. It's free to start with 100 credits and no card, and a ticket or chat counts as one credit, per the pricing page. Try eesel on your own queue and see how many tickets it would have routed right last month.
Frequently Asked Questions
What is PostHog Jeeves?
PostHog Jeeves is an open-weight 9B decision model from PostHog's AI research team. You give it a state (text or JSON) plus typed yes/no, multiple-choice or rating questions, and it writes a reasoning chain before returning a calibrated probability for every option. It copies the request format of TypeSafe's Jev, so it slots into the same classification pipelines.
Is PostHog Jeeves free to use?
Yes. The code is MIT licensed and the weights on Hugging Face are Apache-2.0, so PostHog Jeeves costs nothing to download. You pay for the hardware instead: a CUDA GPU, or an Apple Silicon Mac with 48 GB or more. There is no hosted API, which is the main difference from Jev's per-token pricing.
How is PostHog Jeeves different from Jev?
Jev answers in a single pass with no visible reasoning. PostHog Jeeves thinks first, which lifts accuracy on held-out tests (0.889 vs 0.857) and on JevBench hard (0.865 vs 0.730), but pushes p90 latency to 17.1 seconds with full thinking. Jev still leads on knowledge questions like MMLU. The Jev review covers the single-pass side in detail.
Can PostHog Jeeves triage support tickets?
It can make the decision part of ticket triage: pick a department, flag urgency, score frustration. The README's own demo is a support ticket. It does not reply to customers, update your helpdesk or escalate on its own, so you still need the code or an AI helpdesk agent around it.
How fast is PostHog Jeeves?
On one H100 with FP8, PostHog Jeeves answers in about 0.3 seconds with thinking off, 2.0 seconds median with capped thinking, and 3.3 seconds median (17.1 seconds p90) with full thinking. That is slower than single-pass decision models, so most teams will want the nothink_threshold option to skip thinking on easy cases, much like a tiered escalation flow.
What hardware do I need to run PostHog Jeeves?
The weights take 21 GB in bf16 or 11.5 GB with FP8 linear layers. The README asks for Python 3.12 and a CUDA GPU, and says inference also runs on Apple Silicon Macs with 48 GB or more if you shrink the caches. If you would rather not manage GPUs at all, a managed AI customer service API is the simpler route.
Is PostHog Jeeves better than Kev?
On PostHog's published numbers, Jeeves beats Kev-9B on held-out test data (0.889 vs 0.822) and JevBench (0.935 vs 0.715 for Kev-8B). Kev is smaller and faster, ships four sizes from 0.8B to 27B, and runs on a laptop. Pick Kev for speed, PostHog Jeeves for hard rule-following calls. Both sit alongside Laya in the open decision-model space.

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.







