💻 Hetzner has launched an experimental API for LLM inference, operating via the OpenAI protocol.
Currently, the model Qwen/Qwen2.5-32B-Instruct (MoE with 3B active parameters) is available with a 262K context window. Tests show high performance: the time to first token (TTFT) is 153 ms, and throughput reaches up to 224 tokens per second.
🌍 The emergence of major infrastructure providers like Hetzner in the inference market could lead to lower computing costs and transform LLM inference into a low-margin infrastructure commodity.
👤 This is an excellent opportunity to test modern models through a familiar OpenAI interface on alternative hardware, which may be more cost-effective than using proprietary APIs.
