[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fk8b7cw1vkqkj":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":58,"categories":60,"source":62,"lang":65,"author":66,"audioState":69,"stats":70,"publishedAt":73,"renderer":74},"6abc7b93ca21c797c7e9f936","github---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4","GitHub - lateos-ai\u002Freflex: A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start…","A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time \"System 1\" agent decision loops — process launch to first token, not sustained server throughput.","news",[10,13,18,23,28,33,38,43,48,53],{"headline":11,"body":7,"imageUrl":12,"sourceImageUrl":12},"GitHub - lateos-ai\u002Freflex: A high-performance, GGUF-native Rust & CUDA…","https:\u002F\u002Fopengraph.githubassets.com\u002F52f2b9fce2945f5d79a2ffd1d7cd923d47978754e09b04e99d443020cc579816\u002Flateos-ai\u002Freflex",{"headline":14,"body":15,"imageUrl":16,"images":17},"Every CUDA kernel is compiled ahead-of-time by nvcc","Every CUDA kernel is compiled ahead-of-time by nvcc at build time and shipped inside the binary — never compiled at runtime via NVRTC — so there's no multi-second JIT tax on first use, the way there is with a runtime-compilation design. That's the whole bet: be the fastest way to turn a cold process into one output token, then get out of the way.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"Requires the CUDA toolkit (nvcc on PATH, or","Requires the CUDA toolkit (nvcc on PATH, or CUDA_PATH\u002FCUDA_HOME set) and an NVIDIA GPU. REFLEX_SKIP_CUDA=1 cargo build skips kernel compilation for editing\u002Ftype-checking on a machine without CUDA (no subcommand will actually run kernels in that mode).","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"reflex is a single binary with subcommands: generate","reflex is a single binary with subcommands: generate (load a GGUF and generate tokens), system1 (single-pass, non-autoregressive candidate scoring — the \"System 1\" decision-loop path), smoke (the AOT-pipeline check above), bench (warm-latency microbenchmark), check (byte-exact-vs-reference correctness check, CI-scriptable), and stdio\u002Fuds (local JSON-line IPC, both need --features ipc). Run reflex \u003Csubcommand> with no further arguments to see that subcommand's own usage. Example: download a model from Hugging Face, then run a System1 test","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"system1 takes a local GGUF path, so download","system1 takes a local GGUF path, so download the file first with the hf CLI (pip install -U huggingface_hub), then point system1 at it — this scores each --candidate against the prompt in a single pass, with no autoregressive decode loop: Real output from this exact command (RTX A6000):","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F4.webp",{"local":31},{"headline":34,"body":35,"imageUrl":36,"images":37},"probability is relative to this candidate set only","probability is relative to this candidate set only, not a vocab-wide probability — see Model::system1_evaluate's doc comment in src\u002Fmodel.rs. system1 currently supports dense\u002FMoE Qwen3 only; the Qwen3.5 hybrid mixer and DeepSeek-V2\u002FV3 (MLA) are rejected with a clear error (generate supports all four architectures — swap in a DeepSeek GGUF the same way for a generate run instead).","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F5.webp",{"local":36},{"headline":39,"body":40,"imageUrl":41,"images":42},"generate can also pull a GGUF straight from","generate can also pull a GGUF straight from the Hub itself, via this project's own Rust hf-hub integration — --model \u003Corg\u002Frepo:file.gguf> or --quickstart, both requiring cargo build --features download:","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F6.webp",{"local":41},{"headline":44,"body":45,"imageUrl":46,"images":47},"Closing the steady-state-throughput gap with llama.cpp\u002FvLLM is a","Closing the steady-state-throughput gap with llama.cpp\u002FvLLM is a kernel-optimization race against projects with a multi-year head start — Rust as a language doesn't change who wins it. What none of llama.cpp, vLLM, or candle are built for or measured against is a cold invocation — serverless\u002FFaaS, single-shot CLI\u002Fdev-tool calls, batch\u002Fcron jobs, edge devices that wake on demand. A naive runtime-JIT design pays a real, measured multi-second tax on first kernel use; vLLM took ~235s to become ready (CUDA graph capture) before serving one request; llama.cpp avoids both because its kernels are compiled by nvcc at build time, not at process start.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F7.webp",{"local":46},{"headline":49,"body":50,"imageUrl":51,"images":52},"Target metric: energy-to-first-token from cold start (joules, process","Target metric: energy-to-first-token from cold start (joules, process launch to first generated token) — a real, underexplored gap. Existing energy benchmarks measure warm\u002Fsteady-state joules-per-token, not full-lifecycle cold-start cost. Target models: Qwen and DeepSeek families.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F8.webp",{"local":51},{"headline":54,"body":55,"imageUrl":56,"images":57},"All comparisons are cold-start (process launch to first","All comparisons are cold-start (process launch to first token\u002Fresult), same ThunderCompute A6000, n=3, external wall-clock (\u002Fusr\u002Fbin\u002Ftime -v — process launch to exit, not just Reflex's own internal timer). The llama.cpp\u002FOllama\u002FJev \"loses\" results above are reported as-is, not smoothed over.","\u002Fapi\u002Fmedia\u002Fposts\u002Fgithub---lateos-aireflex-a-high-performance-gguf-native-rust-a88b20f4\u002F9.webp",{"local":56},[59],"dev",[61],"Technology",{"name":63,"url":64},"Show HN","https:\u002F\u002Fgithub.com\u002Flateos-ai\u002Freflex","en",{"handle":67,"displayName":68},"spots","Spots","queued",{"views":71,"likes":72,"saves":72,"shares":72,"completions":72,"opens":72,"skips":72,"depthSum":72},1,0,"2026-09-30T03:01:39.351Z","local"]