Free, read-only GPU waste scanner for Kubernetes inference clusters, by Paralleliq. Finds idle GPUs, oversized GPU tiers and wasted capacity in vLLM deployments.
kubernetes gpu introspection cloud-native model-serving cost-optimization finops gpu-utilization mlops model-discovery llm vllm llm-inference ai-infrastructure modelspec gpu-waste
-
Updated
Oct 4, 2026 - Python