Chapter 9. Supported deployment environments


The following deployment environments for Red Hat AI Inference are supported.

Important

Red Hat AI Inference is available only as a container image. You can download Red Hat AI Inference container images from registry.redhat.io or browse available images in the Red Hat Ecosystem Catalog. To find Red Hat AI Inference container images in the catalog, search for "AI Inference".

The host operating system and kernel must support the required accelerator drivers. For more information, see Supported AI accelerators.

Expand
Table 9.1. Red Hat AI Inference supported deployment environments
EnvironmentSupported versionsDeployment notes

OpenShift Container Platform (self‑managed)

4.14+

Deploy on bare‑metal hosts or virtual machines.

Red Hat OpenShift Service on AWS (ROSA)

4.14+

Requires a ROSA cluster with STS and GPU‑enabled P5 or G5 node types. See Prepare your environment for more information.

Red Hat Enterprise Linux AI

3.0+

Deploy on bare‑metal hosts or virtual machines.

Red Hat Enterprise Linux (RHEL)

9.2+

Deploy on bare‑metal hosts or virtual machines.

Linux (not RHEL)

-

Supported under third‑party policy deployed on bare‑metal hosts or virtual machines. OpenShift Container Platform Operators are not required.

Kubernetes (not OpenShift Container Platform)

-

Supported under third‑party policy deployed on bare‑metal hosts or virtual machines.

Important
  • Single-host deployments for IBM Spyre AI accelerators on IBM Z and IBM Power are supported for RHEL AI 9.6+.
  • Cluster deployments for IBM Spyre AI accelerators on IBM Z are supported as part of Red Hat OpenShift AI version 3.0+ only.
Expand
Table 9.2. Distributed Inference with llm-d supported deployment environments (Technology Preview)
EnvironmentSupported versionsDeployment notes

OpenShift Container Platform

4.19+

Deployed by using the Red Hat AI Helm chart. Requires the Red Hat AI Operator, KServe, Istio, cert-manager, and Gateway API.

Azure Kubernetes Service (AKS)

Kubernetes 1.33+

Deployed by using the Red Hat AI Helm chart. Requires Helm 3.17+, kubectl 1.33+, and a GPU-enabled node pool.

CoreWeave Kubernetes Service (CKS)

Kubernetes 1.33+

Deployed by using the Red Hat AI Helm chart. Requires Helm 3.17+, kubectl 1.33+, and GPU instances (A100, H100, H200, or B200).

Important

Distributed Inference with llm-d is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.

For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.

Red Hat logoGithubredditYoutubeTwitter

Learn

Try, buy, & sell

Communities

About Red Hat

We deliver hardened solutions that make it easier for enterprises to work across platforms and environments, from the core datacenter to the network edge.

Making open source more inclusive

Red Hat is committed to replacing problematic language in our code, documentation, and web properties. For more details, see the Red Hat Blog.

About Red Hat Documentation

Legal Notice

Theme

© 2026 Red Hat
Back to top