GitHub - lateos-ai/reflex: A high-performance, GGUF-native Rust & CUDA…
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time "System 1" agent decision loops — process launch to first token, not sustained server throughput.

