Skip to content

[rhaiis] Add custom model and workload stubs for dashboard overrides - #179

Open
ssaketh-ch wants to merge 2 commits into
openshift-psap:mainfrom
ssaketh-ch:feat/custom-model-workload-stubs
Open

[rhaiis] Add custom model and workload stubs for dashboard overrides#179
ssaketh-ch wants to merge 2 commits into
openshift-psap:mainfrom
ssaketh-ch:feat/custom-model-workload-stubs

Conversation

@ssaketh-ch

Copy link
Copy Markdown
Contributor

Summary

  • Add custom model and workload entries so the dashboard can set
    hf_model_id, vllm_args, and GuideLLM profile fields at job time
    through configOverrides.
  • Remove rampup from profile1 through profile4 in workloads.yaml.

Motivation

Quick fix based on feed back to override the yaml directly from the UI.

Adds custom entries in models.yaml and workloads.yaml so the dashboard UI can override hf_model_id and GuideLLM profile fields at runtime. Also removes rampup from profile1 through profile4 in workloads.yaml.
@openshift-ci

openshift-ci Bot commented Aug 19, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign tosokin for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@coderabbitai

coderabbitai Bot commented Aug 19, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: aa6473fa-27d2-457d-b6ea-6c972f9b3b2c


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ssaketh-ch

Copy link
Copy Markdown
Contributor Author

/test fournos rhaiis nvidia benchmark hera ci-quick
/pipeline forge-full
/exclusive false
/cluster hera

@psap-forge-bot

Copy link
Copy Markdown

🔴 Execution of rhaiis nvidia benchmark hera ci-quick 🔴

Execution Engine Configuration

forge:
  args:
  - nvidia
  - benchmark
  - hera
  - ci-quick
  configOverrides: {}
  project: rhaiis

Artifact Links

Test Logs

00 Pre-Cleanup 2 seconds

01 Prepare 3 seconds

02 Preflight 2 seconds

03 Test 4 hours, 2 minutes, 37 seconds

04 Post-Cleanup 5 seconds

🔄 05 Export-Artifacts

Post-processing Status

  • parse: success
  • artifacts_to_kpis: success
  • kpis_to_mlflow: success
  • kpis_to_csv: success
  • ⏭️ artifacts_to_ai_data: disabled

    kpi.artifacts_to_ai_data disabled

  • ⏭️ s3_import: disabled

    s3_import disabled

  • ⏭️ analyse_kpis: disabled

    analyze disabled

  • ⏭️ s3_export: disabled

    s3_export disabled

@psap-forge-bot

Copy link
Copy Markdown
🔴 Submission of rhaiis nvidia benchmark hera ci-quick failed after 4 hours, 5 minutes, 22 seconds 🔴

Error: FournosJobFailureError: FOURNOS Job 'forge-rhaiis-20260819-200435' failed: Tasks Completed: 6 (Failed: 1, Cancelled 0), Skipped: 0

/test fournos rhaiis nvidia benchmark hera ci-quick
/pipeline forge-full
/exclusive false
/cluster hera

profile7:
tests.rhaiis.workload_key: profile7

custom:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these presets are not required.

@Harshith-umesh

Copy link
Copy Markdown
Member

@ssaketh-ch could you re trigger the ci test, with recommendation from Kevin. there was a transient pvc error on Hera which should be fixed now.

Adds agentic failure review capability based on feedback. Also drops custom presets since they're not required.
@ssaketh-ch

Copy link
Copy Markdown
Contributor Author

/test fournos rhaiis nvidia benchmark hera custom
/pipeline forge-full
/exclusive false
/cluster hera
/var tests.rhaiis.model_key: custom
/var models.custom.name: Qwen3-0.6B
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var tests.rhaiis.workload_key: custom
/var workloads.custom.rates: [1, 5]

@psap-forge-bot

Copy link
Copy Markdown
🔴 Submission of rhaiis nvidia benchmark hera custom failed after 27 seconds 🔴

Error: FournosJobFailureError: FOURNOS Job 'forge-rhaiis-20260820-171343' failed: Job failed in its early stages: Resolution failed: Job has reached the specified backoff limit

/test fournos rhaiis nvidia benchmark hera custom
/var tests.rhaiis.model_key: custom
/var models.custom.name: Qwen3-0.6B
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var tests.rhaiis.workload_key: custom
/var workloads.custom.rates: [1, 5]
/pipeline forge-full
/exclusive false
/cluster hera

metrics: {}

agentic:
enabled: true

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need this?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

that's the agentic configure & failure review I presented yesterday

@ssaketh-ch

Copy link
Copy Markdown
Contributor Author

/test fournos rhaiis nvidia benchmark hera ci-quick
/var tests.rhaiis.model_key: custom
/var tests.rhaiis.workload_key: custom
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var models.custom.name: Qwen3-0.6B
/var workloads.custom.rates: [1, 5]
/pipeline forge-full
/exclusive false
/cluster hera

@ssaketh-ch

Copy link
Copy Markdown
Contributor Author

/test fournos rhaiis nvidia benchmark hera ci-quick
/var tests.rhaiis.model_key: custom
/var tests.rhaiis.workload_key: custom
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var models.custom.name: Qwen3-0.6B
/var workloads.custom.rates: [1, 5]
/pipeline forge-full
/exclusive false
/cluster hera

@psap-forge-bot

Copy link
Copy Markdown

🟢 Execution of rhaiis nvidia benchmark hera ci-quick 🟢

Execution Engine Configuration

forge:
  args:
  - nvidia
  - benchmark
  - hera
  - ci-quick
  configOverrides:
    models.custom.hf_model_id: Qwen/Qwen3-0.6B
    models.custom.name: Qwen3-0.6B
    tests.rhaiis.model_key: custom
    tests.rhaiis.workload_key: custom
    workloads.custom.rates:
    - 1
    - 5
  project: rhaiis

Artifact Links

Test Logs

00 Pre-Cleanup 2 seconds

01 Prepare 3 seconds

02 Preflight 2 seconds

03 Test 18 minutes, 2 seconds

Test Description

This test performs a rapid CI validation of the rhaiis project by benchmarking vLLM on NVIDIA GPUs using the ci-quick preset. It specifically evaluates inference performance at request rates of 1 and 5 with a fixed 1000-token workload, while disabling warmup, profiling, and dashboard exports to ensure fast feedback.

04 Post-Cleanup 5 seconds

🔄 05 Export-Artifacts

Post-processing Status

  • parse: success
  • artifacts_to_kpis: success
  • kpis_to_mlflow: success
  • kpis_to_csv: success
  • ⏭️ artifacts_to_ai_data: disabled

    kpi.artifacts_to_ai_data disabled

  • ⏭️ s3_import: disabled

    s3_import disabled

  • ⏭️ analyse_kpis: disabled

    analyze disabled

  • ⏭️ s3_export: disabled

    s3_export disabled

@psap-forge-bot

Copy link
Copy Markdown
🟢 Submission of rhaiis nvidia benchmark hera ci-quick succeeded after 22 minutes, 32 seconds 🟢
/test fournos rhaiis nvidia benchmark hera ci-quick
/var tests.rhaiis.model_key: custom
/var tests.rhaiis.workload_key: custom
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var models.custom.name: Qwen3-0.6B
/var workloads.custom.rates: [1, 5]
/pipeline forge-full
/exclusive false
/cluster hera

@ssaketh-ch

Copy link
Copy Markdown
Contributor Author
/test fournos rhaiis nvidia benchmark hera ci-quick 
/var tests.rhaiis.model_key: custom 
/var tests.rhaiis.workload_key: custom 
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B 
/var models.custom.name: Qwen3-0.6B 
/var workloads.custom.rates: [1, 5] 
/pipeline forge-full 
/exclusive false
/cluster hera

The first trigger failed due to a parsing issue. The second time succeeded so it doesn't seems to be a command error. Need to be careful when triggering via comments.

@psap-forge-bot

Copy link
Copy Markdown
🔴 Submission of rhaiis nvidia benchmark hera ci-quick failed after 5 minutes, 11 seconds 🔴

Error: RetryFailure: All 30 attempts failed for task wait_for_job_to_resolve : Wait for FOURNOS job to successfully resolve (reach Pending state) (last reason: Unknown status)

/test fournos rhaiis nvidia benchmark hera ci-quick
/var tests.rhaiis.model_key: custom
/var tests.rhaiis.workload_key: custom
/var models.custom.hf_model_id: Qwen/Qwen3-0.6B
/var models.custom.name: Qwen3-0.6B
/var workloads.custom.rates: [1, 5]
/pipeline forge-full
/exclusive false
/cluster hera

@openshift-ci

openshift-ci Bot commented Aug 20, 2026

Copy link
Copy Markdown

@ssaketh-ch: The following test failed, say /retest to rerun all failed tests or /retest-required to rerun all mandatory failed tests:

Test name Commit Details Required Rerun command
ci/prow/fournos 566c536 link true /test fournos

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@psap-forge-bot

Copy link
Copy Markdown

🟢 Execution of rhaiis nvidia benchmark hera ci-quick 🟢

Execution Engine Configuration

forge:
  args:
  - nvidia
  - benchmark
  - hera
  - ci-quick
  configOverrides:
    models.custom.hf_model_id: Qwen/Qwen3-0.6B
    models.custom.name: Qwen3-0.6B
    tests.rhaiis.model_key: custom
    tests.rhaiis.workload_key: custom
    workloads.custom.rates:
    - 1
    - 5
  project: rhaiis

Artifact Links

Test Logs

00 Pre-Cleanup 2 seconds

01 Prepare 3 seconds

02 Preflight 2 seconds

03 Test 18 minutes, 1 second

Test Description

This test executes a rapid CI benchmark for the rhaiis project using vLLM on NVIDIA GPUs hosted on the Hera cluster. It specifically targets the ci-quick workload with low concurrency rates (1 and 5) and no warmup to quickly assess engine stability and performance under minimal load.

04 Post-Cleanup 4 seconds

🔄 05 Export-Artifacts

Post-processing Status

  • parse: success
  • artifacts_to_kpis: success
  • kpis_to_mlflow: success
  • kpis_to_csv: success
  • ⏭️ artifacts_to_ai_data: disabled

    kpi.artifacts_to_ai_data disabled

  • ⏭️ s3_import: disabled

    s3_import disabled

  • ⏭️ analyse_kpis: disabled

    analyze disabled

  • ⏭️ s3_export: disabled

    s3_export disabled

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants