About this integration
Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.
- Transport
- stdio
- Authentication
- api-key
- Initial setup
- credentials
- Runtime
- unattended
- Evidence
- documented
- Version
- 0.10.4
- Package
- inference-aiops
- Last compatibility test
- Not independently tested
Connect your agent
inference-aiopsPublisher documentation reviewed from pinned captures; package and integration endpoint were not executed or independently security-audited.
Capabilities: Diagnose inference latency, Inspect GPU serving clusters, Perform guarded scaling and model lifecycle operations
Connected profiles
Additional details
io.github.AIops-tools/inference-aiops
Source ↗ · Checked 2026-09-170.10.4
Source ↗ · Checked 2026-09-17CC0-1.0; package licenses are separate
Source ↗ · Checked 2026-09-17pypi
Source ↗ · Checked 2026-09-17inference-aiops
Source ↗ · Checked 2026-09-180.10.4
Source ↗ · Checked 2026-09-18The console command runs the complete MCP tool surface over stdio.
Source ↗ · Checked 2026-09-18A serving-stack bearer token is optional and encrypted storage may require a master password.
Source ↗ · Checked 2026-09-18Run the inference-aiops mcp console command.
Source ↗ · Checked 2026-09-18Publisher documentation reviewed; package and integration endpoint not executed or independently security-audited.
Source ↗ · Checked 2026-09-18