Red Hat AI Inference 3.5

What's New

リリースノート

今回の Red Hat AI Inference リリースにおける新機能と変更点のハイライト

Get started

スタートガイド

Red Hat AI Inference のスタートガイド

Distributed Inference with llm-d

Kubernetes でのスケーラブルな LLM サービス向けの llm-d を使用した分散推論のアーキテクチャー、コンポーネント、およびデプロイメント

Plan

Validated models

Red Hat AI validated models

Supported product and hardware configurations

Supported product and hardware configurations for deploying Red Hat AI software

Inference Operations

OpenShift Container Platform でのスタンドアロンの Red Hat AI 推論コンテナーのデプロイ

サポートされている AI アクセラレーターがインストールされている OpenShift Container Platform クラスターにスタンドアロンの Red Hat AI Inference コンテナーをデプロイします。

非接続環境でのスタンドアロンの Red Hat AI Inference コンテナーのデプロイ

OpenShift Container Platform と切断されたミラーイメージレジストリーを使用した非接続環境での Red Hat AI 推論のデプロイ

OCI 準拠のモデルコンテナーの推論サービング言語モデル

Red Hat AI Inference での OCI 準拠のモデルの推論

投機的デコーディング

Red Hat AI Inference での投機的なデコード

推論提供の Mistral 3 モデル

Red Hat AI 推論による推論サービス Mistral 3 モデル

地理空間基盤モデルに関する推論

Red Hat AI 推論による地理空間基盤モデルの推論

Red Hat AI Model Optimization Toolkit

LLM Compressor ライブラリーを使用した大規模言語モデルの圧縮

vLLM のサーバー引数

Red Hat AI Inference を実行するためのサーバー引数

機能呼び出しツールによる Red Hat AI 推論の拡張

AI Inference 用のツール呼び出しおよびチャットテンプレートの設定

Distributed Inference Operations

Openshift Container Platform での llm-d を使用した分散推論のデプロイ

Openshift Container Platform 上の大規模な言語モデルをデプロイし、これを提供します。

Azure または CoreWeave Kubernetes Service の llm-d を使用した分散推論のデプロイ

Azure または CoreWeave Kubernetes Service の llm-d を使用した分散推論のデプロイ

llm-d デプロイメントによる分散推論の監視およびトラブルシューティング

llm-d デプロイメントによる分散推論の監視およびトラブルシューティング

Batch process inference requests with Distributed Inference with llm-d

Batch process inference requests with Distributed Inference with llm-d

Deploy large MoE models across multiple GPUs with WideEP

Deploy Distributed Inference with llm-d on Azure or CoreWeave Kubernetes Service

Deploy Models as a Service on non-OpenShift Kubernetes

Deploy and configure Models as a Service on non-OpenShift Kubernetes environments

Manage mixed workloads by using priority queuing

Manage mixed workloads by using priority queuing

Tool calling for Distributed Inference with llm-d deployments

Tool calling for Distributed Inference with llm-d deployments

Upgrade Distributed Inference with llm-d on managed Kubernetes

Upgrade the Distributed Inference with llm-d infrastructure stack on managed Kubernetes clusters

Additional Resources

Red Hat AI Inference Server 3.3

Switch to the Red Hat AI Inference Server 3.3 documentation

Red Hat AI Foundations

Explore no-cost courses to boost your AI knowledge and get hands-on experience with Red Hat AI products while earning a certificate

Red Hat AI learning hub

Explore the curated set of third‑party models validated for Red Hat AI products, ready for fast, reliable deployment

Red Hat logoGithubredditYoutubeTwitter

詳細情報

試用、購入および販売

コミュニティー

会社概要

Red Hat は、企業がコアとなるデータセンターからネットワークエッジに至るまで、各種プラットフォームや環境全体で作業を簡素化できるように、強化されたソリューションを提供しています。

多様性を受け入れるオープンソースの強化

Red Hat では、コード、ドキュメント、Web プロパティーにおける配慮に欠ける用語の置き換えに取り組んでいます。このような変更は、段階的に実施される予定です。詳細情報: Red Hat ブログ.

Red Hat ドキュメントについて

Legal Notice

Theme

© 2026 Red Hat
トップに戻る