
Testing the 9x Smaller Local LLM Claim on a GPU-less VPS
Last month I did something no sane person should try on a Tuesday night: I downloaded a custom fork of llama.cpp from a company I'd never heard of, just to run a 27 billion parameter model that supposedly fits in under 6 gigabytes. I'd read about it in Simon Willison's comment on the release, a ternary-quantized model called Bonsai 2 27B, pitched as "near-lossless compression in a 9x smaller footprint." Near-lossless is a big claim for something throwing away most of the bits in every weight. I wanted to know whether that phrase meant what I hoped, or whether it was the kind of number that only survives contact with a benchmark table. So instead of running it on a beefy Mac like Willison did, I pointed it at the cheapest box I own: the small Hetzner VPS that runs this very blog's publishing pipeline. Here's what actually happened. What "ternary" buys you and what it costs

