Introducing infr: A Pure-Rust, Vulkan-First LLM Inference Engine
published:
infr is a from-the-metal LLM inference engine written in Rust. It runs Llama, Qwen, Gemma and more on any Vulkan GPU, has native Metal and CPU backends, speaks the OpenAI API, and holds its own against llama.cpp.