Spots

GitHub - lateos-ai/reflex: A high-performance, GGUF-native Rust & CUDA…

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time "System 1" agent decision loops — process launch to first token, not sustained server throughput.

Every CUDA kernel is compiled ahead-of-time by nvcc

Every CUDA kernel is compiled ahead-of-time by nvcc at build time and shipped inside the binary — never compiled at runtime via NVRTC — so there's no multi-second JIT tax on first use, the way there is with a runtime-compilation design. That's the whole bet: be the fastest way to turn a cold process into one output token, then get out of the way.

Requires the CUDA toolkit (nvcc on PATH, or

Requires the CUDA toolkit (nvcc on PATH, or CUDA_PATH/CUDA_HOME set) and an NVIDIA GPU. REFLEX_SKIP_CUDA=1 cargo build skips kernel compilation for editing/type-checking on a machine without CUDA (no subcommand will actually run kernels in that mode).

reflex is a single binary with subcommands: generate

reflex is a single binary with subcommands: generate (load a GGUF and generate tokens), system1 (single-pass, non-autoregressive candidate scoring — the "System 1" decision-loop path), smoke (the AOT-pipeline check above), bench (warm-latency microbenchmark), check (byte-exact-vs-reference correctness check, CI-scriptable), and stdio/uds (local JSON-line IPC, both need --features ipc). Run reflex <subcommand> with no further arguments to see that subcommand's own usage. Example: download a model from Hugging Face, then run a System1 test

system1 takes a local GGUF path, so download

system1 takes a local GGUF path, so download the file first with the hf CLI (pip install -U huggingface_hub), then point system1 at it — this scores each --candidate against the prompt in a single pass, with no autoregressive decode loop: Real output from this exact command (RTX A6000):

probability is relative to this candidate set only

probability is relative to this candidate set only, not a vocab-wide probability — see Model::system1_evaluate's doc comment in src/model.rs. system1 currently supports dense/MoE Qwen3 only; the Qwen3.5 hybrid mixer and DeepSeek-V2/V3 (MLA) are rejected with a clear error (generate supports all four architectures — swap in a DeepSeek GGUF the same way for a generate run instead).

generate can also pull a GGUF straight from

generate can also pull a GGUF straight from the Hub itself, via this project's own Rust hf-hub integration — --model <org/repo:file.gguf> or --quickstart, both requiring cargo build --features download:

Closing the steady-state-throughput gap with llama.cpp/vLLM is a

Closing the steady-state-throughput gap with llama.cpp/vLLM is a kernel-optimization race against projects with a multi-year head start — Rust as a language doesn't change who wins it. What none of llama.cpp, vLLM, or candle are built for or measured against is a cold invocation — serverless/FaaS, single-shot CLI/dev-tool calls, batch/cron jobs, edge devices that wake on demand. A naive runtime-JIT design pays a real, measured multi-second tax on first kernel use; vLLM took ~235s to become ready (CUDA graph capture) before serving one request; llama.cpp avoids both because its kernels are compiled by nvcc at build time, not at process start.

Target metric: energy-to-first-token from cold start (joules, process

Target metric: energy-to-first-token from cold start (joules, process launch to first generated token) — a real, underexplored gap. Existing energy benchmarks measure warm/steady-state joules-per-token, not full-lifecycle cold-start cost. Target models: Qwen and DeepSeek families.

All comparisons are cold-start (process launch to first

All comparisons are cold-start (process launch to first token/result), same ThunderCompute A6000, n=3, external wall-clock (/usr/bin/time -v — process launch to exit, not just Reflex's own internal timer). The llama.cpp/Ollama/Jev "loses" results above are reported as-is, not smoothed over.

News

GitHub - lateos-ai/reflex: A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start…

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time "System 1" agent decision loops — process launch to first token, not sustained server throughput.

@spots #dev
Source: Show HN
See more like this