[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fqo5i0pn53rld":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":57,"categories":58,"source":59,"lang":62,"author":63,"audioState":66,"stats":67,"publishedAt":70,"renderer":71},"6abbaf48ca21c797c7e9d399","github---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9","GitHub - PostHog\u002Fjeeves: Jeeves – Reasoning improves Jev-like decision models","A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.","news",[10,12,17,22,27,32,37,42,47,52],{"headline":6,"body":7,"imageUrl":11,"sourceImageUrl":11},"https:\u002F\u002Fopengraph.githubassets.com\u002Fa5c318f88ac0d21dcb69f526d11682f2f3a71faafb8c24f9fcae1d23b4261316\u002FPostHog\u002Fjeeves",{"headline":13,"body":14,"imageUrl":15,"images":16},"A 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head)","A 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, with a block-4 diffusion drafter and the full training code and train\u002Fdev\u002Ftest data.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F1.webp",{"local":15},{"headline":18,"body":19,"imageUrl":20,"images":21},"Beats Kev-9B and Jev on test data it","Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench's public tiers (0.935 vs 0.866 for Jev).","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F2.webp",{"local":20},{"headline":23,"body":24,"imageUrl":25,"images":26},"Supports yes\u002Fno (noul), multiple-choice (choice), and rating (score)","Supports yes\u002Fno (noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F3.webp",{"local":25},{"headline":28,"body":29,"imageUrl":30,"images":31},"About 0.3 s per request without thinking and","About 0.3 s per request without thinking and a 3.3 s median with it on one H100. Can be sped up by truncating chain length. Runs on CUDA (Hopper for the FP8 kernel).","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F4.webp",{"local":30},{"headline":33,"body":34,"imageUrl":35,"images":36},"Jev-like models give calibrated decision probabilities, but at","Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides. This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public). Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes. * No Kev-9B JevBench result is published. These are Kev-8B (Qwen3).","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F5.webp",{"local":35},{"headline":38,"body":39,"imageUrl":40,"images":41},"All JevBench numbers are on the public easy","All JevBench numbers are on the public easy, standard and hard tiers (231 items). The sealed judge tier is not included, and the Jev and Kev numbers are restricted to the same public items. Without thinking the same checkpoint scores 0.804 on our test split (2,962 items), against 0.840 with it. Requirements: Python 3.12 and a CUDA GPU. Download the released weights and serve them: Or fuse your own trained checkpoint into a standalone model and serve it with a drafter: Then send a request in Jev's format: Response on one H100 (FP8), with the three questions thinking in parallel: sdk\u002F is a drop-in replacement for Jev's Python SDK (typesafe-sdk): The client connects to http:\u002F\u002F127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F6.webp",{"local":40},{"headline":43,"body":44,"imageUrl":45,"images":46},"options is optional and ignored by Jev clients","options is optional and ignored by Jev clients that don't send it. Server-wide defaults are set with the matching serve flags. Questions, states and answers are loaded into the Qwen chat template like The model then rolls out its reasoning chain, and after the \u003C\u002Fthink> token we append","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F7.webp",{"local":45},{"headline":48,"body":49,"imageUrl":50,"images":51},"A pointer head scores each option with a","A pointer head scores each option with a scaled dot product between a query projection of the hidden state at \u003Cdecide> and a key projection of the hidden state at that option's \u003C\u002Fopt>, where","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F8.webp",{"local":50},{"headline":53,"body":54,"imageUrl":55,"images":56},"These are rare, largely unused tokens in the","These are rare, largely unused tokens in the Qwen tokenizer. Ablations found that using plain text like \"State\" in the prompt instead worsened performance. Likewise, not repeating the questions after the reasoning block also decreases performance. The final probabilities are a softmax over the option scores, divided by a temperature fitted on the dev set.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---posthogjeeves-jeeves-reasoning-improves-jev-like-de-ebdae3b9\u002F9.webp",{"local":55},[],[],{"name":60,"url":61},"Hacker News","https:\u002F\u002Fgithub.com\u002FPostHog\u002Fjeeves","en",{"handle":64,"displayName":65},"spots","Spots","queued",{"views":68,"likes":69,"saves":69,"shares":69,"completions":69,"opens":69,"skips":69,"depthSum":69},3,0,"2026-09-29T12:30:00.555Z","local"]