💻 Hetzner has launched an experimental API for LLM inference, operating via the OpenAI protocol.

Currently, the model Qwen/Qwen2.5-32B-Instruct (MoE with 3B active parameters) is available with a 262K context window. Tests show high performance: the time to first token (TTFT) is 153 ms, and throughput reaches up to 224 tokens per second.

🌍 The emergence of major infrastructure providers like Hetzner in the inference market could lead to lower computing costs and transform LLM inference into a low-margin infrastructure commodity.

👤 This is an excellent opportunity to test modern models through a familiar OpenAI interface on alternative hardware, which may be more cost-effective than using proprietary APIs.

Source 1: https://sliplane.io/blog/hetzner-inference