Skip to content

[WIP] feat: Update precise profile - #176

Open
albertoperdomo2 wants to merge 24 commits into
openshift-psap:mainfrom
albertoperdomo2:feat/precise-prefix-cache-v2
Open

[WIP] feat: Update precise profile #176
albertoperdomo2 wants to merge 24 commits into
openshift-psap:mainfrom
albertoperdomo2:feat/precise-prefix-cache-v2

Conversation

@albertoperdomo2

Copy link
Copy Markdown
Collaborator

No description provided.

Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
@openshift-ci openshift-ci Bot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 18, 2026
@openshift-ci

openshift-ci Bot commented Aug 18, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign ashishkamra for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@coderabbitai

coderabbitai Bot commented Aug 18, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: c36841b1-8504-4bcc-a9b2-6097d92ea8db


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@psap-forge-bot

Copy link
Copy Markdown

🟢 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🟢

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 2 seconds

01 Test 16 minutes, 56 seconds

Test Description

This FORGE test executes release validation for the llm_d project, testing the gpt-oss-120b model on RHOAI 3.5 EA2 with H200 GPUs using the release-precise-prefix-cache deployment profile under multi-turn workloads to evaluate prefix caching performance and throughput with high-prefix-length requests.

🔄 02 Export-Artifacts

Post-processing Status

Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
…omo2/forge into feat/precise-prefix-cache-v2
@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 3 seconds

01 Test 6 minutes, 2 seconds

Test Description

This FORGE test benchmarks the gpt-oss-120b model within the llm_d project on RHOAI 3.5 EA2 using H200 GPUs. It validates performance across three deployment configurations (distributed default, precise prefix cache, and approximate prefix cache) and three concurrent workload patterns (1k-1k, heavy heterogeneous, and multi-turn text completions).

🔄 02 Export-Artifacts

Post-processing Status

@psap-forge-bot

Copy link
Copy Markdown

🟢 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🟢

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 3 seconds

01 Test 17 minutes, 12 seconds

Test Description

This test benchmarks the llm_d project using the openai/gpt-oss-120b model to evaluate multi-turn text completion performance on RHOAI 3.5 EA2 with H200 GPUs. It applies the release-precise-prefix-cache deployment profile and configures the workload for text completions.

🔄 02 Export-Artifacts

Post-processing Status

@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 2 seconds

01 Test 3 minutes, 16 seconds

Test Description

This test performs release validation for the openai/gpt-oss-120b model on RHOAI 3.5-EA2 using H200 GPUs within the llm_d project, benchmarking release-distributed-default, release-precise-prefix-cache, and release-approximate-prefix-cache deployment profiles across concurrent-1k-1k, heavy-heterogeneous, and multi-turn workload scenarios.

🔄 02 Export-Artifacts

Post-processing Status

Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 3 seconds

01 Test 11 minutes, 56 seconds

Test Description

This test validates the llm_d project release on RHOAI 3.5 EA2 using the openai/gpt-oss-120b model on H200 GPUs. It benchmarks performance across a matrix of deployment profiles (distributed default, precise and approximate prefix caching) and workload scenarios (concurrent 1k-1k, heavy heterogeneous, and multi-turn).

Failure Review 002 Deploy Llmisvc

001__llmd__multi-turn__release-precise-prefix-c/002__deploy_llmisvc

The wait_service_ready task for the LLMInferenceService failed because the release-precise-prefix-cache deployment profile configured a Gateway API route referencing an unsupported InferencePool backend, causing the cluster controller to reject the configuration with an InvalidKind error. This incompatibility blocked network routing, led to workload pod readiness probe failures and a subsequent restart, which triggered the test harness's fail-fast mechanism to abort the deployment.

🔄 02 Export-Artifacts

Post-processing Status

@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 2 seconds

01 Test 10 minutes, 48 seconds

Test Description

This test evaluates the llm_d project on RHOAI-3.5-EA2 using an NVIDIA H200 GPU, specifically benchmarking the openai/gpt-oss-120b model across three deployment profiles (release-distributed-default, release-precise-prefix-cache, release-approximate-prefix-cache) and three concurrent workload scenarios (concurrent-1k-1k, heavy-heterogeneous, multi-turn). It validates release readiness, scalability, and performance under varied request loads and prefix-caching configurations.

Failure Review 002 Deploy Llmisvc

001__llmd__multi-turn__release-precise-prefix-c/002__deploy_llmisvc

The KServe LLMInferenceService readiness validation failed when the test harness aborted the wait task after detecting a pod restart, caused by the inference deployment failing to stabilize due to routing errors. The definitive root cause is a Gateway API controller incompatibility where the cluster does not recognize the InferencePool custom resource (inference.networking.x-k8s.io), resulting in HTTPRoute validation failures that block service initialization and trigger container crashes.

🔄 02 Export-Artifacts

Post-processing Status

@albertoperdomo2

Copy link
Copy Markdown
Collaborator Author

/test fournos llm_d cpt-release-testing-llama-33-70b
/cluster athena-fire
/pipeline forge-prepare-only
/var model_cache.pvc.access_mode: ReadWriteMany
/var model_cache.pvc.storage_class_name: nfs-rwx
/rhoai.rc-image quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be stable-3.5
/var platform.cluster.skip_gpu_readiness: true

@psap-forge-bot

Copy link
Copy Markdown

🟢 Execution of llm_d cpt-release-testing-llama-33-70b 🟢

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-llama-33-70b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
  project: llm_d

Artifact Links

Test Logs

00 Prepare 55 seconds

01 Preflight 2 seconds

🔄 02 Export-Artifacts

@psap-forge-bot

Copy link
Copy Markdown
🟢 Submission of llm_d cpt-release-testing-llama-33-70b succeeded after 2 minutes, 56 seconds 🟢
/test fournos llm_d cpt-release-testing-llama-33-70b
/var model_cache.pvc.access_mode: ReadWriteMany
/var model_cache.pvc.storage_class_name: nfs-rwx
/var platform.cluster.skip_gpu_readiness: true
/cluster athena-fire
/pipeline forge-prepare-only
/rhoai.rc-image quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be stable-3.5

@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 3 seconds

01 Test 5 minutes, 40 seconds

Test Description

This llm_d test benchmarks the gpt-oss-120b model on H200 GPUs for the RHOAI 3.5 EA2 release, evaluating performance across distributed-default, precise-prefix-cache, and approximate-prefix-cache deployment profiles using a multi-turn text completion workload with prefix caching variations.

Failure Review 002 Deploy Llmisvc

001__llmd__multi-turn__release-precise-prefix-c/002__deploy_llmisvc

The wait_service_ready task aborted the KServe deployment after detecting a pod restart, preventing the LLMInferenceService from reaching a ready state. The failure stems from a dual-layer misconfiguration where the cluster lacks the required InferencePool CRD, causing HTTPRoute routing rejections, and the inference containers are crashing due to OOMKilled events when loading the 120B parameter model, which exceeds the provisioned resource limits.

🔄 02 Export-Artifacts

Post-processing Status

@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of llm_d cpt-gpt-oss-120b precise-v2 multi-turn 🔴

Execution Engine Configuration

forge:
  args:
  - cpt-release-testing-gpt-oss-120b
  configOverrides:
    model_cache.pvc.access_mode: ReadWriteMany
    model_cache.pvc.storage_class_name: nfs-rwx
    platform.cluster.skip_gpu_readiness: true
    platform.operators.rhods-operator.channel: stable-3.5
    platform.rhoai.custom_catalog.enabled: true
    platform.rhoai.custom_catalog.image: quay.io/rhoai/rhoai-fbc-fragment@sha256:fd718eb8e74e22a5f3fb8bcd4899229ac36045f55a790243adc355a218c877be
    runtime.benchmark_key: multi-turn
    runtime.deployment_profile: release-precise-prefix-cache
    workloads.pvc_storage_class: nfs-rwx
  project: llm_d

Artifact Links

Test Logs

00 Preflight 2 seconds

01 Test 6 minutes, 2 seconds

Test Description

This test benchmarks the llm_d project by evaluating the gpt-oss-120b model on H200 GPUs within the RHOAI 3.5 EA2 release environment. It specifically tests concurrent request scaling, heavy heterogeneous workloads, and multi-turn inference across distributed serving and precise/approximate prefix caching deployment profiles.

Failure Review 002 Deploy Llmisvc

001__llmd__multi-turn__release-precise-prefix-c/002__deploy_llmisvc

The wait_service_ready task in the deploy_llmisvc component failed with a TaskExecutionError after the inference pod llm-d-release-prec-439d77-kserve-86b8c47597-2n5bf crashed and restarted during the readiness check. This restart triggered a strict safety guard in the task logic that immediately aborts the wait to prevent masking unstable deployments, causing the LLMInferenceService to remain unavailable and the pipeline to fail.

🔄 02 Export-Artifacts

Post-processing Status

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant