September 12, 2026

New

🤖 Monitor your LLM inference servers: vLLM and SGLang

Your GPU dashboard tells you the card is busy. It doesn't tell you the server is serving.

When vLLM or SGLang dies (a CUDA OOM, an engine crash, an OOM-killed container), every GPU metric goes green: an idle GPU is cool, low-power and 0% utilized. Fivenines now watches the inference server itself, so a dead model server pages you instead of looking healthy.

🔌 Setup is just a URL
No exporter to install. vLLM already exposes Prometheus metrics on its OpenAI-compatible port (http://127.0.0.1:8000/metrics), SGLang on port 30000 when launched with --enable-metrics. Enable it per instance in Settings → AI Inference, paste the metrics URL, and add an auth header if the endpoint sits behind a proxy. Needs agent v1.17.0 for vLLM, v1.17.1 for SGLang.

📊 The serving signals GPU metrics can't see
One chart line per served model: queued and running requests, KV cache usage (vLLM) or token usage (SGLang), prompt and generation tokens/s, time to first token, inter-token latency, end-to-end latency, and prefix / RadixAttention cache hit rate. All on a dedicated vLLM or SGLang page on the instance, plus a card on the overview.

🔔 Alerts that fire while the GPUs look fine
A new AI Inference category in the workflow templates: server down, queue saturated, and KV cache / token usage pressure. Unreachable is a red incident; an auth or TLS problem on your endpoint shows up as a separate config error, so a misconfigured header doesn't wake anyone at 3am.

🙈 No fake zeroes
A vLLM started with --disable-log-stats, or an SGLang without --enable-metrics, answers fine but publishes nothing. We don't show "0 running requests", which would read as a healthy idle server: we tell you which flag is missing.

Open any instance → Settings → AI Inference to turn it on. 🚀

← All updates

Want something that isn't here yet? Vote on the roadmap.