Release notes

Red Hat OpenShift AI Self-Managed 3.0

Features, enhancements, resolved issues, and known issues associated with this release

Abstract

These release notes provide an overview of new features, enhancements, resolved issues, and known issues in version 3.0 of Red Hat OpenShift AI.

Chapter 1. Upgrade to OpenShift AI 3.0 not supported
Copy link

You cannot upgrade from OpenShift AI 2.25 or any earlier version to 3.0. OpenShift AI 3.0 introduces significant technology and component changes and is intended for new installations only. To use OpenShift AI 3.0, install the Red Hat OpenShift AI Operator on a cluster running OpenShift Container Platform 4.19 or later and select the fast-3.x channel.

Support for upgrades will be available in a later release, including upgrades from OpenShift AI 2.25 to a stable 3.x version.

For more information, see the Why upgrades to OpenShift AI 3.0 are not supported Knowledgebase article.

Chapter 2. Overview of OpenShift AI
Copy link

Red Hat OpenShift AI is a platform for data scientists and developers of artificial intelligence and machine learning (AI/ML) applications.

OpenShift AI provides an environment to develop, train, serve, test, and monitor AI/ML models and applications on-premise or in the cloud.

For data scientists, OpenShift AI includes Jupyter and a collection of default workbench images optimized with the tools and libraries required for model development, and the TensorFlow and PyTorch frameworks. Deploy and host your models, integrate models into external applications, and export models to host them in any hybrid cloud environment. You can enhance your projects on OpenShift AI by building portable machine learning (ML) workflows with AI pipelines by using Docker containers. You can also accelerate your data science experiments through the use of graphics processing units (GPUs) and Intel Gaudi AI accelerators.

For administrators, OpenShift AI enables data science workloads in an existing Red Hat OpenShift or ROSA environment. Manage users with your existing OpenShift identity provider, and manage the resources available to workbenches to ensure data scientists have what they require to create, train, and host models. Use accelerators to reduce costs and allow your data scientists to enhance the performance of their end-to-end data science workflows using graphics processing units (GPUs) and Intel Gaudi AI accelerators.

OpenShift AI has two deployment options:

Self-managed software that you can install on-premise or in the cloud. You can install OpenShift AI Self-Managed in a self-managed environment such as OpenShift Container Platform, or in Red Hat-managed cloud environments such as Red Hat OpenShift Dedicated (with a Customer Cloud Subscription for AWS or GCP), Red Hat OpenShift Service on Amazon Web Services (ROSA classic or ROSA HCP), or Microsoft Azure Red Hat OpenShift.
A managed cloud service, installed as an add-on in Red Hat OpenShift Dedicated (with a Customer Cloud Subscription for AWS or GCP) or in Red Hat OpenShift Service on Amazon Web Services (ROSA classic).
For information about OpenShift AI Cloud Service, see Product Documentation for Red Hat OpenShift AI.

For information about OpenShift AI supported software platforms, components, and dependencies, see the Supported Configurations for 3.x Knowledgebase article.

For a detailed view of the 3.0 release lifecycle, including the full support phase window, see the Red Hat OpenShift AI Self-Managed Life Cycle Knowledgebase article.

Chapter 3. New features and enhancements
Copy link

This section describes new features and enhancements in Red Hat OpenShift AI 3.0.

3.1. New features
Copy link

PyTorch v2.8.0 KFTO training images now generally available

You can now use PyTorch v2.8.0 training images for distributed workloads in OpenShift AI.

The following new images are available:

ROCm-compatible KFTO training image: quay.io/modh/training:py312-rocm63-torch280 Compatible with AMD accelerators supported by ROCm 6.3.
CUDA-compatible KFTO training image: quay.io/modh/training:py312-cuda128-torch280 Compatible with NVIDIA GPUs supported by CUDA 12.8.

Hardware profiles

A new Hardware Profiles feature replaces the previous Accelerator Profiles and legacy Container Size selector for workbenches.

Hardware profiles provide a more flexible and consistent way to define compute configurations for AI workloads, simplifying resource management across different hardware types.

Important

The Accelerator Profiles and legacy Container Size selector are now deprecated and will be removed in a future release.

Connections API now generally available

The Connections API is now available as a general availability (GA) feature in OpenShift AI.

This API enables you to create and manage connections to external data sources and services directly within OpenShift AI. Connections are stored as Kubernetes Secrets with standardized annotations, allowing protocol-based validation and routing across integrated components.

IBM Power and IBM Z architecture support

OpenShift AI now supports the IBM Power (ppc64le) and IBM Z (s390x) architectures.

This expanded platform support enables deployment of AI and machine learning workloads on IBM enterprise hardware, providing greater flexibility and scalability across heterogeneous environments.

For more information about supported software platforms, components, and dependencies, see the Knowledgebase article: Supported Configurations for 3.x.

IBM Power and IBM Z architecture support for TrustyAI

TrustyAI is now available as a general availability (GA) feature for the IBM Power (ppc64le) and IBM Z (s390x) architectures.

TrustyAI is an open-source Responsible AI toolkit that provides a suite of tools to support responsible and transparent AI workflows. It offers capabilities such as fairness and data drift metrics, local and global model explanations, text detoxification, language model benchmarking, and language model guardrails.

These capabilities help ensure transparency, accountability, and the ethical use of AI systems within OpenShift AI environments on IBM Power and IBM Z systems.

IBM Power and IBM Z architecture support for Model Registry

Model Registry is now available for the IBM Power (ppc64le) and IBM Z (s390x) architectures.

Model Registry is an open-source component that simplifies and standardizes the management of AI and machine learning (AI/ML) model lifecycles. It provides a centralized platform for storing, versioning, and governing models, enabling seamless collaboration across data science and MLOps teams.

Model Registry supports capabilities such as model versioning and lineage tracking, metadata management and model discovery, model approval and promotion workflows, integration with CI/CD and deployment pipelines, and governance, auditability, and compliance features on IBM Power and IBM Z systems.

IBM Power and IBM Z architecture support for Notebooks

Notebooks are now available for the IBM Power (ppc64le) and IBM Z (s390x) architectures.

Notebooks provide containerized, browser-based development environments for data science, machine learning, research, and coding within the OpenShift AI ecosystem. These environments can be launched through Workbenches and include the following options:

Jupyter Minimal notebook: A lightweight JupyterLab IDE for basic Python development and model prototyping.
Jupyter Data Science notebook: Preconfigured with popular data science libraries and tools for end-to-end workflows.
Jupyter TrustyAI notebook: An environment for Responsible AI tasks, including model explainability, fairness, data drift detection, and text detoxification.
Code Server: A browser-based VS Code environment for collaborative development with familiar IDE features.
Runtime Minimal and Runtime Data Science: Headless environments for automated workflows and consistent pipeline execution.

IBM Power architecture support for Feature Store

Feature Store is now supported on the IBM Power (ppc64le) architecture.

This support enables users to build, register, and manage features for machine learning models directly on IBM Power-based environments, with full integration with OpenShift AI.

IBM Power architecture support for AI Pipelines

AI Pipelines are now supported on the IBM Power (ppc64le) architecture.

This capability enables users to define, run, and monitor AI pipelines natively within OpenShift AI, leveraging the performance and scalability of IBM Power systems for AI workloads.

AI Pipelines executed on IBM Power systems maintain functional parity with x86 deployments.

Support for IBM Power accelerated Triton Inference Server

You can now enable Power architecture support for Triton inference server (CPU only) with FIL, PyTorch, Python and ONNX backend. You can deploy Triton inference server as a custom model serving runtime on IBM Power architecture in Red Hat OpenShift AI.

For details, see Triton Inference Server image.

Support for IBM Z accelerated Triton Inference Server

You can now enable Z architecture support for the Triton Inference Server (Telum I/Telum II) with multiple backend options, including ONNX-MLIR, Snap ML (C++), and PyTorch. The Triton Inference Server can be deployed as a custom model serving runtime on IBM Z architecture as a Technology Preview feature in Red Hat OpenShift AI.

For details, see IBM Z accelerated Triton Inference Server.

IBM Spyre AI Accelerator model serving support on IBM Z platforms

Model serving with the IBM Spyre AI Accelerator is now available as a general availability (GA) feature for IBM Z platforms.

The IBM Spyre Operator automates installation and integrates key components such as the device plugin, secondary scheduler, and monitoring.

For more information, see the IBM Spyre Operator catalog entry: IBM Spyre Operator — Red Hat Ecosystem Catalog.

Note

On IBM Z and IBM LinuxONE, Red Hat OpenShift AI supports deploying large language models (LLMs) with vLLM on IBM Spyre. The Triton Inference Server is supported on Telum (CPU) only.

For more information, see the following documentation:

Model customization components

OpenShift AI 3.0 introduces a suite of model customization components that streamline and enhance the process of preparing, fine-tuning, and deploying AI models.

The following components are now available:

Red Hat AI Python Index: A Red Hat-maintained Python package index that hosts supported builds of packages useful for AI and machine learning notebooks. Using the Red Hat AI Python Index ensures reliable and secure access to these packages in both connected and disconnected environments.
docling: A powerful Python library for advanced data processing that converts unstructured documents, such as PDFs or images, into clean, machine-readable formats for AI and ML workloads.
Synthetic Data Generation Hub (sdg-hub): A toolkit for generating high-quality synthetic data to augment datasets, improve model robustness, and address edge cases.
Training Hub: A framework that simplifies and accelerates the fine-tuning and customization of foundation models by using your own data.
Kubeflow Trainer: A Kubernetes-native capability that enables distributed training and fine-tuning of models while abstracting the underlying infrastructure complexity.
AI Pipelines: A Kubeflow-native capability for building configurable workflows across AI components, including all other model customization modules in this suite.

3.2. Enhancements
Copy link

Hybrid search support for remote vector databases in Llama Stack

You can now enable hybrid search on remote vector databases in Llama Stack in OpenShift AI.

This enhancement allows enterprises to use their existing managed vector database infrastructure while maintaining high retrieval performance and flexibility across different database types.

IBM Spyre support for IBM Z with Caikit-TGIS adapter

You can now serve models with IBM Spyre AI accelerators on IBM Z (s390x architecture) by using the vLLM Spyre s390x ServingRuntime for KServe with the Caikit-TGIS gRPC adapter.

This integration enables high-performance model serving and inference for generative AI workloads on IBM Z systems within OpenShift AI.

Data Science Pipelines renamed to AI Pipelines

OpenShift AI now uses the term "AI Pipelines" instead of "Data Science Pipelines" to better reflect the broader range of AI and generative AI use cases supported by the platform.

In the default DataScienceCluster (default-dsc), the datasciencepipelines component has been renamed to aipipelines to align with this terminology update.

This is a naming change only. The AI pipelines functionality remains the same.

Model Catalog enhancements with model validation data

The Model Details page in the OpenShift AI Model Catalog now includes comprehensive model validation data, such as performance benchmarks, hardware compatibility, and other key metrics.

This enhancement provides a unified and detailed view consistent with the Jounce UI model details layout, enabling users to evaluate models more effectively from a single interface.

Model Catalog performance data with search and filtering

The Model Catalog now includes detailed performance and validation data for Red Hat-validated third-party models, such as benchmarks and hardware compatibility metrics.

Enhanced search and filtering capabilities, such as filtering by latency or hardware profile, help users quickly identify models optimized for their specific use cases and available resources, providing a unified discovery experience within the Red Hat AI Hub.

Distributed Inference with llm-d is now generally available (GA)

Distributed Inference with llm-d supports multi-model serving, intelligent inference scheduling, and disaggregated serving for improved GPU utilization on generative AI models.

Note

The following capabilities are not fully supported:

Wide Expert-Parallelism multi-node: Developer Preview.
Wide Expert-Parallelism on Blackwell B200: Not available but can be provided as a Technology Preview.
Multi-node on GB200: Not supported.
Gateway discovery and association are not supported in the UI during model deployment in this release. Users must associate models with Gateways by applying the resource manifests directly through the API or CLI.

User interface for Distributed Inference with llm-d deployment configuration

OpenShift AI now includes a user interface (UI) for configuring large language model (LLM) deployments that run on the llm-d Serving Runtime.

This streamlined interface simplifies common deployment scenarios by providing essential configuration options with sensible defaults while still allowing explicit selection of the llm-d runtime for your deployment.

The new UI reduces setup complexity and helps users deploy distributed inference workloads more efficiently.

New navigation system

OpenShift AI 3.0 introduces a redesigned, streamlined navigation system that improves usability and workflow efficiency.

The new layout enables users to move seamlessly between features, simplifying access to key capabilities and supporting a smoother end-to-end experience.

Enhanced authentication for AI Pipelines

OpenShift AI 3.0 replaces oauth-proxy with kube-rbac-proxy for AI Pipelines as part of the platform-wide authentication transition.

This update improves security and compatibility, particularly for environments without an internal OAuth server, such as Red Hat OpenShift Service on AWS.

When migrating to kube-rbac-proxy, SubjectAccessReview (SAR) requirements and RBAC permissions change accordingly. Users who rely on the built-in ds-pipeline-user-access-<dspa-name> role are updated automatically, while others must ensure their roles include access to the datasciencepipelinesapplications/api subresource with the following verbs: create, update, patch, delete, get, list, and watch.

Observability and Grafana integration for Distributed Inference with llm-d

In OpenShift AI 3.0, platform administrators can connect observability components to Distributed Inference with llm-d deployments and integrate with self-hosted Grafana instances to monitor inference workloads.

This capability allows teams to collect and visualize Prometheus metrics from Distributed Inference with llm-d for performance analysis and custom dashboard creation.

Chapter 4. Technology Preview features
Copy link

Important

This section describes Technology Preview features in Red Hat OpenShift AI 3.0. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.

For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.

TrustyAI–Llama Stack integration for safety, guardrails, and evaluation

You can now use the Guardrails Orchestrator from TrustyAI with Llama Stack as a Technology Preview feature.

This integration enables built-in detection and evaluation workflows to support AI safety and content moderation. When TrustyAI is enabled and the FMS Orchestrator and detectors are configured, no manual setup is required.

To activate this feature, set the following field in the DataScienceCluster custom resource for the OpenShift AI Operator: spec.llamastackoperator.managementState: Managed

For more information, see the TrustyAI FMS Provider on GitHub: TrustyAI FMS Provider.

AI Available Assets page for deployed models and MCP servers

A new AI Available Assets page enables AI engineers and application developers to view and consume deployed AI resources within their projects.

This enhancement introduces a filterable UI that lists available models and Model Context Protocol (MCP) servers in the selected project, allowing users with appropriate permissions to identify accessible endpoints and integrate them directly into the AI Playground or other applications.

Generative AI Playground for model testing and evaluation

The Generative AI (GenAI) Playground introduces a unified, interactive experience within the OpenShift AI dashboard for experimenting with foundation and custom models.

Users can test prompts, compare models, and evaluate Retrieval-Augmented Generation (RAG) workflows by uploading documents and chatting with their content. The GenAI Playground also supports integration with approved Model Context Protocol (MCP) servers and enables export of prompts and agent configurations as runnable code for continued iteration in local IDEs.

Chat context is preserved within each session, providing a suitable environment for prompt engineering and model experimentation.

Support for air-gapped Llama Stack deployments

You can now install and operate Llama Stack and RAG/Agentic components in fully disconnected (air-gapped) OpenShift AI environments.

This enhancement enables secure deployment of Llama Stack features without internet access, allowing organizations to use AI capabilities while maintaining compliance with strict network security policies.

Feature Store integration with Workbenches and new user access capabilities

This feature is available as a Technology Preview.

The Feature Store is now integrated with OpenShift AI, data science projects, and workbenches. This integration also introduces centrally managed, role-based access control (RBAC) capabilities for improved governance.

These enhancements provide two key capabilities:

Feature development within the workbench environment.
Administrator-controlled user access.
This update simplifies and accelerates feature discovery and consumption for data scientists while allowing platform teams to maintain full control over infrastructure and feature access.

Feature Store user interface

The Feature Store component now includes a web-based user interface (UI).

You can use the UI to view registered Feature Store objects and their relationships, such as features, data sources, entities, and feature services.

To enable the UI, edit your FeatureStore custom resource (CR) instance. When you save the change, the Feature Store Operator starts the UI container and creates an OpenShift route for access.

For more information, see Setting up the Feature Store user interface for initial use.

IBM Spyre AI Accelerator model serving support on x86 platforms: Model serving with the IBM Spyre AI Accelerator is now available as a Technology Preview feature for x86 platforms. The IBM Spyre Operator automates installation and integrates the device plugin, secondary scheduler, and monitoring. For more information, see the IBM Spyre Operator catalog entry.

Build Generative AI Apps with Llama Stack on OpenShift AI

With this release, the Llama Stack Technology Preview feature enables Retrieval-Augmented Generation (RAG) and agentic workflows for building next-generation generative AI applications. It supports remote inference, built-in embeddings, and vector database operations. It also integrates with providers like TrustyAI’s provider for safety and Trusty AI’s LM-Eval provider for evaluation.

This preview includes tools, components, and guidance for enabling the Llama Stack Operator, interacting with the RAG Tool, and automating PDF ingestion and keyword search capabilities to enhance document discovery.

Centralized platform observability

Centralized platform observability, including metrics, traces, and built-in alerts, is available as a Technology Preview feature. This solution introduces a dedicated, pre-configured observability stack for OpenShift AI that allows cluster administrators to perform the following actions:

View platform metrics (Prometheus) and distributed traces (Tempo) for OpenShift AI components and workloads.
Manage a set of built-in alerts (alertmanager) that cover critical component health and performance issues.
Export platform and workload metrics to external 3rd party observability tools by editing the DataScienceClusterInitialization (DSCI) custom resource.
You can enable this feature by integrating with the Cluster Observability Operator, Red Hat build of OpenTelemetry, and Tempo Operator. For more information, see Monitoring and observability. For more information, see Managing observability.

Support for Llama Stack Distribution version 0.3.0

The Llama Stack Distribution now includes version 0.3.0 as a Technology Preview feature.

This update introduces several enhancements, including expanded support for retrieval-augmented generation (RAG) pipelines, improved evaluation provider integration, and updated APIs for agent and vector store management. It also provides compatibility updates aligned with recent OpenAI API extensions and infrastructure optimizations for distributed inference.

The previously supported version was 0.2.22.

Support for Kubernetes Event-driven Autoscaling (KEDA)

OpenShift AI now supports Kubernetes Event-driven Autoscaling (KEDA) in its KServe RawDeployment mode. This Technology Preview feature enables metrics-based autoscaling for inference services, allowing for more efficient management of accelerator resources, reduced operational costs, and improved performance for your inference services.

To set up autoscaling for your inference service in KServe RawDeployment mode, you need to install and configure the OpenShift Custom Metrics Autoscaler (CMA), which is based on KEDA.

For more information about this feature, see: Configuring metrics-based autoscaling.

LM-Eval model evaluation UI feature: TrustyAI now offers a user-friendly UI for LM-Eval model evaluations as Technology Preview. This feature allows you to input evaluation parameters for a given model and returns an evaluation-results page, all from the UI.

Use Guardrails Orchestrator with LlamaStack

You can now run detections using the Guardrails Orchestrator tool from TrustyAI with Llama Stack as a Technology Preview feature, using the built-in detection component. To use this feature, ensure TrustyAI is enabled, the FMS Orchestrator and detectors are set up, and KServe RawDeployment mode is in use for full compatibility if needed. There is no manual set up required. Then, in the DataScienceCluster custom resource for the Red Hat OpenShift AI Operator, set the spec.llamastackoperator.managementState field to Managed.

For more information, see Trusty AI FMS Provider on GitHub.

Support for creating and managing Ray Jobs with the CodeFlare SDK

You can now create and manage Ray Jobs on Ray Clusters directly through the CodeFlare SDK.

This enhancement aligns the CodeFlare SDK workflow with the KubernetesFlow Training Operator (KFTO) model, where a job is created, run, and completed automatically. This enhancement simplifies manual cluster management by preventing Ray Clusters from remaining active after job completion.

Support for direct authentication with an OIDC identity provider

Direct authentication with an OpenID Connect (OIDC) identity provider is now available as a Technology Preview feature.

This enhancement centralizes OpenShift AI service authentication through the Gateway API, providing a secure, scalable, and manageable authentication model. You can configure the Gateway API with your external OIDC provider by using the GatewayConfig custom resource.

Custom flow estimator for Synthetic Data Generation pipelines

You can now use a custom flow estimator for synthetic data generation (SDG) pipelines.

For supported and compatible tagged SDG teacher models, the estimator helps you evaluate a chosen teacher model, custom flow, and supported hardware on a sample dataset before running full workloads.

Llama Stack support and optimization for single node OpenShift (SNO)

Llama Stack core can now deploy and run efficiently on single node OpenShift (SNO).

This enhancement optimizes component startup and resource usage so that Llama Stack can operate reliably in single-node cluster environments.

FAISS vector storage integration

You can now use the FAISS (Facebook AI Similarity Search) library as an inline vector store in OpenShift AI.

FAISS is an open-source framework for high-performance vector search and clustering, optimized for dense numerical embeddings with both CPU and GPU support. When enabled with an embedded SQLite backend in the Llama Stack Distribution, FAISS stores embeddings locally within the container, removing the need for an external vector database service.

New Feature Store component

You can now install and manage Feature Store as a configurable component in OpenShift AI. Based on the open-source Feast project, Feature Store acts as a bridge between ML models and data, enabling consistent and scalable feature management across the ML lifecycle.

This Technology Preview release introduces the following capabilities:

Centralized feature repository for consistent feature reuse
Python SDK and CLI for programmatic and command-line interactions to define, manage, and retrieve features for ML models
Feature definition and management
Support for a wide range of data sources
Data ingestion via feature materialization
Feature retrieval for both online model inference and offline model training
Role-Based Access Control (RBAC) to protect sensitive features
Extensibility and integration with third-party data and compute providers
Scalability to meet enterprise ML needs
Searchable feature catalog
Data lineage tracking for enhanced observability
For configuration details, see Configuring Feature Store.

FIPS support for Llama Stack and RAG deployments

You can now deploy Llama Stack and RAG or agentic solutions in regulated environments that require FIPS compliance.

This enhancement provides FIPS-certified and compatible deployment patterns to help organizations meet strict regulatory and certification requirements for AI workloads.

Validated sdg-hub notebooks for Red Hat AI Platform

Validated sdg_hub example notebooks are now available to provide a notebook-driven user experience in OpenShift AI 3.0.

These notebooks support multiple Red Hat platforms and enable customization through SDG pipelines. They include examples for the following use cases:

Knowledge and skills tuning, including annotated examples for fine-tuning models.
Synthetic data generation with reasoning traces to customize reasoning models.
Custom SDG pipelines that demonstrate using default blocks and creating new blocks for specialized workflows.

RAGAS evaluation provider for Llama Stack (inline and remote)

You can now use the Retrieval-Augmented Generation Assessment (RAGAS) evaluation provider to measure the quality and reliability of RAG systems in OpenShift AI.

RAGAS provides metrics for retrieval quality, answer relevance, and factual consistency, helping you identify issues and optimize RAG pipeline configurations.

The integration with the Llama Stack evaluation API supports two deployment modes:

Inline provider: Runs RAGAS evaluation directly within the Llama Stack server process.
Remote provider: Runs RAGAS evaluation as distributed jobs using OpenShift AI pipelines.
The RAGAS evaluation provider is now included in the Llama Stack distribution.

Enable targeted deployment of workbenches to specific worker nodes in Red Hat OpenShift AI Dashboard using node selectors

Hardware profiles are now available as a Technology Preview. The hardware profiles feature enables users to target specific worker nodes for workbenches or model-serving workloads. It allows users to target specific accelerator types or CPU-only nodes.

This feature replaces the current accelerator profiles feature and container size selector field, offering a broader set of capabilities for targeting different hardware configurations. While accelerator profiles, taints, and tolerations provide some capabilities for matching workloads to hardware, they do not ensure that workloads land on specific nodes, especially if some nodes lack the appropriate taints.

The hardware profiles feature supports both accelerator and CPU-only configurations, along with node selectors, to enhance targeting capabilities for specific worker nodes. Administrators can configure hardware profiles in the settings menu. Users can select the enabled profiles using the UI for workbenches, model serving, and AI pipelines where applicable.

RStudio Server workbench image

With the RStudio Server workbench image, you can access the RStudio IDE, an integrated development environment for R. The R programming language is used for statistical computing and graphics to support data analysis and predictions.

To use the RStudio Server workbench image, you must first build it by creating a secret and triggering the BuildConfig, and then enable it in the OpenShift AI UI by editing the rstudio-rhel9 image stream. For more information, see Building the RStudio Server workbench images.

Important

Disclaimer: Red Hat supports managing workbenches in OpenShift AI. However, Red Hat does not provide support for the RStudio software. RStudio Server is available through rstudio.org and is subject to their licensing terms. You should review their licensing terms before you use this sample workbench.

CUDA - RStudio Server workbench image

With the CUDA - RStudio Server workbench image, you can access the RStudio IDE and NVIDIA CUDA Toolkit. The RStudio IDE is an integrated development environment for the R programming language for statistical computing and graphics. With the NVIDIA CUDA toolkit, you can enhance your work by using GPU-accelerated libraries and optimization tools.

To use the CUDA - RStudio Server workbench image, you must first build it by creating a secret and triggering the BuildConfig, and then enable it in the OpenShift AI UI by editing the rstudio-rhel9 image stream. For more information, see Building the RStudio Server workbench images.

Important

The CUDA - RStudio Server workbench image contains NVIDIA CUDA technology. CUDA licensing information is available in the CUDA Toolkit documentation. You should review their licensing terms before you use this sample workbench.

Support for multinode deployment of very large models

Serving models over multiple graphical processing unit (GPU) nodes when using a single-model serving runtime is now available as a Technology Preview feature. Deploy your models across multiple GPU nodes to improve efficiency when deploying large models such as large language models (LLMs). For more information, see Deploying models by using multiple GPU nodes.

Chapter 5. Developer Preview features
Copy link

Important

This section describes Developer Preview features in Red Hat OpenShift AI 3.0. Developer Preview features are not supported by Red Hat in any way and are not functionally complete or production-ready. Do not use Developer Preview features for production or business-critical workloads. Developer Preview features provide early access to functionality in advance of possible inclusion in a Red Hat product offering. Customers can use these features to test functionality and provide feedback during the development process. Developer Preview features might not have any documentation, are subject to change or removal at any time, and have received limited testing. Red Hat might provide ways to submit feedback on Developer Preview features without an associated SLA.

For more information about the support scope of Red Hat Developer Preview features, see Developer Preview Support Scope.

Model-as-a-Service (MaaS) integration