Skip to content

[data, torch_profile] test: add st of torch and mstx - #23

Merged
tardis-key merged 16 commits into
verl-project:mainfrom
756017542:main
Mar 25, 2026
Merged

[data, torch_profile] test: add st of torch and mstx#23
tardis-key merged 16 commits into
verl-project:mainfrom
756017542:main

Conversation

@756017542

@756017542 756017542 commented Mar 18, 2026

Copy link
Copy Markdown
Contributor

What does this PR do?

add st of torch and mstx

Add concise overview of what this PR aims to achieve or accomplish. Reference related GitHub issues and PRs that help with the review.

Checklist Before Starting

  • Search for similar PRs. Paste at least one query link here: ...
  • Format the PR title as [{modules}] {type}: {description} (This will be checked by the CI)
    • {modules} include mstx, mvtx, torch_profile, deployment, perf, algo, env, doc, data, cfg, ci, misc,
    • If this PR involves multiple modules, separate them with , like [mstx, ci]
    • {type} is in feat, fix, refactor, chore, test
    • If this PR breaks any API (CLI arguments, config, function signature, etc.), add [BREAKING] to the beginning of the title.
    • Example: [BREAKING][mstx, torch_profile] feat: support timeline parsing

Test

For changes that can not be tested by CI (e.g., algorithm implementation, new model support), validate by experiment(s) and show results like training curve plots, evaluation results, etc.

API and Usage Example

Demonstrate how the API changes if any, and provide usage example(s) if possible.

# Add code snippet or script demonstrating how to use this

Design & Code Changes

Demonstrate the high-level design if this PR is complex, and list the specific changes.

Checklist Before Submitting

Important

Please check all the following items before requesting a review, otherwise the reviewer might deprioritize this PR for review.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the testing infrastructure by introducing end-to-end tests for MSTX and PyTorch profilers. It provides concrete examples of profiler outputs and a utility for data compression, ensuring the robustness and maintainability of the profiling analysis tools.

Highlights

  • New End-to-End Tests: Added dedicated end-to-end tests for both MSTX and PyTorch profilers to validate their output generation and functionality.
  • Sample Profiling Data: Included new sample profiling data files for both MSTX and PyTorch, which are used by the newly added E2E tests.
  • JSON to JSON.GZ Conversion Utility: Introduced a Python script to convert standard JSON files into a gzipped JSON format, potentially for storage or transfer efficiency.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for torch and mstx profilers, including new test data and end-to-end tests. My review focuses on improving the quality and robustness of the new tests and utility scripts. I've suggested using pytest fixtures to make the tests cleaner and more reliable. For the new data conversion script, I've recommended changes to make it more reusable and to improve its error handling. I also pointed out some inconsistencies in the test data and a best practice of not committing generated files to the repository.

Comment thread tests/special_e2e/test_mstx_e2e.py Outdated
Comment thread tests/special_e2e/test_torch_e2e.py Outdated
Comment thread torch_data/jsontojsongz.py Outdated
Comment thread torch_data/jsontojsongz.py Outdated
Comment thread torch_data/rl_timeline.html Outdated
@756017542 756017542 changed the title add st of torch and mstx add st of torch Mar 18, 2026
Comment thread torch_data/jsontojsongz.py Outdated
@756017542 756017542 changed the title add st of torch [data, torch_profile] test: add st of torch Mar 19, 2026
Comment thread tests/special_e2e/test_torch_e2e.py
Comment thread tests/special_e2e/test_torch_e2e.py Outdated
Comment thread tests/special_e2e/test_torch_e2e.py Outdated
Comment thread tests/special_e2e/test_torch_e2e.py Outdated
Comment thread docs/data/data_directory.md
@tardis-key

Copy link
Copy Markdown
Collaborator

ci name is not correct.

@756017542 756017542 changed the title [data, torch_profile] test: add st of torch [data, torch_profile] test: add st of torch and mstx Mar 19, 2026

- name: Install dependencies
run: |
pip install -r requirements.txt

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pip install 没有使用 --no-cache-dir,建议添加 --no-cache-dir 避免缓存问题


- name: Run profiling_data_analysis_st tests
run: |
pytest -s -x tests/special_e2e

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pytest 没有超时限制,建议添加 --timeout=300 或其他合适的超时设置


- name: Run profiling_data_analysis_st tests
run: |
pytest -s -x tests/special_e2e

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

文件末尾缺少空行,添加空行以符合POSIX标准

Comment thread docs/data/data_directory.md Outdated
```
### 数据解析文件 prof_*.json.gz,解析文件缺少字段见解析日志warning,解析文件内容示例:

![img_1.png](img_1.png)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

use text instead of png

Comment thread docs/data/data_directory.md Outdated
└── <role>/
└── prof_*.json.gz
```
### 数据解析文件 prof_*.json.gz,解析文件内容包含distrubutedInfo、traceEvent等字段,数据内容一般包含ts、dur等字段,解析文件内容示例:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

标题太长,建议拆解成标题和正文

Comment thread docs/data/data_directory.md Outdated
└── ASCEND_PROFILER_OUTPUT/
└── trace_view.json
```
### 数据解析文件 trace_view.json,解析文件内容必须包含"ph": "M",且"name": "Overlap Analysis"对应"pid"的数据,该数据一般包含ts、dur等字段,解析文件内容示例:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同上

@@ -0,0 +1,101 @@
[

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

修改文件路径,文件夹名字非公共部分可以用xxx表述

@@ -0,0 +1,50 @@
{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同上

@tardis-key
tardis-key merged commit 19f838b into verl-project:main Mar 25, 2026
4 checks passed
@tardis-key tardis-key mentioned this pull request Mar 25, 2026
19 tasks
@Rhetee

Rhetee commented Mar 26, 2026

Copy link
Copy Markdown
Collaborator

/lgtm

tardis-key added a commit that referenced this pull request Jun 8, 2026
* [misc] feat: optimize console output (#29)

Optimize console output

* [data, torch_profile] test: add st of torch and mstx (#23)

* add st of torch

* delate json

* add special_e2e

* code check

* code check

* add special_e2e.yml

* mv data and fix bug

* bugfix

* add st of mstx

* Revise inspection comments

* change st name

* Add data description

* add data file

* delate st --timeout=300

* delate img.png

* revise data_directory

---------

Co-authored-by: gcw_fbonFwWl <gcw_fbonFwWl@noreply.gitcode.com>

* [data] feat: add DataChecker (#28)

* [data] feat: add base data class framework with validation support (#26)

* [data] feat: add base data class framework with validation support

Signed-off-by: Debonex <debonexx@gmail.com>

* [refactor] move data module to rl_insight package and fix imports

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] refactor: simplify validation system with class method approach

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] feat: update data related interfaces and tests

Signed-off-by: Debonex <debonexx@gmail.com>

---------

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] feat: add datachecker (#27)

* add datachecker

* Delete rl_timeline.html

* Unify logging and address other code review comments

* update  requirement

* address code review comments from gemini

* adjust ut

* pre-commit

* pre-commit

* format uworkflow yml

* Address review comments

---------

Signed-off-by: Debonex <debonexx@gmail.com>
Co-authored-by: Debonet <37174444+Debonex@users.noreply.github.com>

* [mstx] fix: do not perform filtering when total_rank is too small (#34)

* do not perform filtering when total_rank is too small

* Update rl_insight/parser/parser.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [mstx] fix: skip mstx_preprocessing if necessary (#32)

* skip mstx_preprocessing if necessary

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update mstx_preprocessing.py

* Update mstx_preprocessing.md

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [visualizer] feat: optimize the encapsulation and registration of the visualizer (#30)

* Optimize the encapsulation and registration of the visualizer

* update pr-title check

* [data] feat: add data check rule func ParserOutputValidatorRule (#36)

* add rule func ParserOutputValidatorRule

* fix check error

* fix data checker test error

* fix start_time_ms error

* Fixed review comments and added test cases.

* rename docs/data/data_directory.md to  docs/data/data_specification.md
add empty summary event data test case

* fix data directory reference

---------

Co-authored-by: zhangning <zhangning42@huawei.com>

* [ci] test: add doc url validation check (#43)

test: add doc url check in CI

Co-authored-by: zhangning <zhangning42@huawei.com>

* [doc] chore: add guideline for offlinepipeline (#40)

* doc optimization

* add guideline

* Update main.py

* Update architecture_and_guideline.md

* [visualizer] fix: delete the remaining visualizer type. (#39)

* remove vis_type from visualizer

* update ut

* [data] feat: add VerlLogData validation rules (#35)

* [data] feat: add VERL_LOG validation rules and local check script

* fix(data): refine VERL_LOG validation per review

- Move VeRL rules to verl_log_rules.py; rename PresentRule to ExistRule
- Require single .log file path; share path validation helper
- CLI errors to stderr; update tests and data_directory.md

* docs(data): use data/verl_data for VeRL log sample paths

Made-with: Cursor

* chore(data): track VeRL sample logs under data/verl_data (override *.log)

Made-with: Cursor

* feat(data): require Training Progress in VeRL logs; trim full sample

- Add Training Progress: to DEFAULT_REQUIRED_KEYWORDS and docs
- Replace good_full_verl.log with good_minimal_verl.log; add bad_* fixtures
- Update tests; data spec lists only minimal log for check example

Made-with: Cursor

* style: ruff-format data rules; Apache header on check_verl_log

- Keep shebang first in check_verl_log.py for direct execution

Made-with: Cursor

* [visualizer] feat: add RL timeline PNG generator (#38)

* [visualizer]feat:update generate timeline type of PNG

* [visualizer]fix:add ut and code optimization

* [visualizer]feat:add ut

* [visualizer] fix: clean code and fix rank sorting

* fix:update gemini assist

* fix:update gemini assist again

* fix:update gemini assist

* fix:fix gemini assist

* fix:update pre-commit

* fix: update ut

* fix: update ci test again

* fix: update ci test again

* feat: add kaleido depend

* fix: test ci again

* fix: test ci again

* fix: test ci again

* fix: test ci again

---------

Co-authored-by: ChenGary13 <chenjiayu31@huawei.com>

* [data] feat: add verify rules for MSTX profiling files (#41)

* Add validation rules for Mstx JSON files

1、add MstxJsonFileExistsRule
2、add MstxJsonFieldValidRule

* Added tests for MstxJsonFileExistsRule and MstxJsonFieldValidRule

Added tests for MstxJsonFileExistsRule and MstxJsonFieldValidRule.

* Update data_checker.py

* split MULTI_JSON to MULTI_JSON_MSTX and MULTI_JSON_TORCH.

* Fix bug

* modify 'multi_json' to 'multi_json_mstx'

* fix 'tmp_path' bug in test case

* use pre-commit code consistent

---------

Co-authored-by: pqhgitee <pqhgitee@noreply.gitcode.com>

* [doc] fix: Update roadmap links in README.md

* [data] feat: add verify rules for Torch profiling files (#52)

* Add validation rules for Torch Profile files

* Update rl_insight/data/rules.py

changing this to > 0

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [doc] refactor: reorganize documentation structure (#51)

* Update data_specification.md

* Update data_specification.md

* update pipeline docs&mstx_exec

* update

* update docs

* update docs

* clean DS_Store

* clean

* Update index.rst

* update

* Enhance input/output section with directory structure

Add directory structure example for input data in documentation.

* Update input/output section in documentation

Removed example directory structure for input data.

* Update baseclusterparser_interface.md

* [parser, data, doc, ci] feat: add nvtx parser function (#50)

* add nvtx parser function

* Update examples/nvtx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* fix error

* resume pre-commit yaml

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update docs/cluster_analysis.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* add nvtx data and e2e test script

* add verify rules for nvtx data

* refactor nvtx parser

* Update rl_insight/parser/nvtx_parser.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [visualizer, pipeline] feat: add MoE expert load heatmap visualization (#47)

* support moe load visualization with heatmap

* standardize gmm_exec.sh

* [parser,data] fix: harden gmm parsing, filtering, and output schema

* [pipeline,cfg] refactor: normalize gmm cli parameters and config passing

* [visualizer] refactor: improve gmm heatmap layout and segment rendering

* [doc] docs: update gmm guide, example script and sample heatmap asset

* [test,doc] test: add st for gmm_heatmap and document dump data layout

* refactor: move GMM CLI groups to parser/visualizer packages for centralized registration

* add minimal data example & format code (pre-commit compliant)

* fix(types): resolve parser mypy errors for nvtx and shared DataMap keys

---------

Co-authored-by: chenjiao.angel <chenjiao.angel@bytedance.com>

* [misc] fix: harden parser validation and stabilize cross-platform test behavior (#56)

* fix: improve parser robustness, cross-platform path handling, and test stability

* docs: refine RL timeline quickstart profiling link text

* fix: improve GMM parsing and restore scalable heatmap metadata

* fix: restore matplotlib gmm visualizer

* fix: restore original gmm visualizer

* fix: improve MSTX ordering and harden docs URL validation

* fix: accept Path inputs in validators and normalize GMM paths

* fix: apply pre-commit cleanup for validator tests

* update requirements.txt

* fix ut error & update requirements

* [pipeline] feat: rl insight support online monitor (#53)

rl insight support online monitor

* [cfg] refactor: Use omegaconf structured configuration system (#58)

* Use omegaconf structured configuration system and migrate CLI

* Modular configuration of parameters and reduction of redundancy.

* [parser] feat: add memory parser (#54)

* [memory] feat: add memory parser

* [memory]fix: fix reviews

* [memory]fix: fix pre check problems

---------

Co-authored-by: mookies1 <zhanghaoyong1@huawei.com>

* [pipeline] feat: rl monitor add grafana config file (#60)

rl monitor add grafana config file

* [pipeline] feat: rl online monitor statetimeline optime (#61)

rl online monitor statetimeline optime

* [doc] fix: fix pypi display issue and expand readme with more details (#65)

* Update README.md and provide more details

* remove dead code

* Update schema.py

* [doc] fix: align quickstart installation instructions with README (#67)

Unify installation sections in timeline and GMM heatmap quickstarts to match README: pip install first, then optional source install.

* [misc] fix: remove dead code (#71)

* remove dead code

* Update schema.py

---------

Signed-off-by: Debonex <debonexx@gmail.com>
Co-authored-by: JIANG-PENGJUN <52533600+756017542@users.noreply.github.com>
Co-authored-by: gcw_fbonFwWl <gcw_fbonFwWl@noreply.gitcode.com>
Co-authored-by: Debonet <37174444+Debonex@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: tifning <44485561+tifning@users.noreply.github.com>
Co-authored-by: zhangning <zhangning42@huawei.com>
Co-authored-by: duesdues <55529526+duesdues@users.noreply.github.com>
Co-authored-by: Gary-cjy <71553064+Gary-cjy@users.noreply.github.com>
Co-authored-by: ChenGary13 <chenjiayu31@huawei.com>
Co-authored-by: panqihan <39796558+pqhgit@users.noreply.github.com>
Co-authored-by: pqhgitee <pqhgitee@noreply.gitcode.com>
Co-authored-by: a550580874 <82751568+a550580874@users.noreply.github.com>
Co-authored-by: zhengxiaojun <alwayszxj@gmail.com>
Co-authored-by: hswei88 <129183149+hswei88@users.noreply.github.com>
Co-authored-by: chenjiao.angel <chenjiao.angel@bytedance.com>
Co-authored-by: Zhen <295632982@qq.com>
Co-authored-by: TMC <87188729+mengchengTang@users.noreply.github.com>
Co-authored-by: Ruowei Zheng <892882856@qq.com>
Co-authored-by: Moocharr <1123277477@qq.com>
Co-authored-by: mookies1 <zhanghaoyong1@huawei.com>
Co-authored-by: zyang6 <zhouyang271@huawei.com>
tardis-key added a commit that referenced this pull request Jun 29, 2026
* add st of torch

* delate json

* add special_e2e

* code check

* code check

* add special_e2e.yml

* mv data and fix bug

* bugfix

* add st of mstx

* Revise inspection comments

* change st name

* Add data description

* add data file

* delate st --timeout=300

* delate img.png

* revise data_directory

---------

Co-authored-by: gcw_fbonFwWl <gcw_fbonFwWl@noreply.gitcode.com>
tardis-key added a commit that referenced this pull request Jul 17, 2026
* [misc] feat: optimize console output (#29)

Optimize console output

* [data, torch_profile] test: add st of torch and mstx (#23)

* add st of torch

* delate json

* add special_e2e

* code check

* code check

* add special_e2e.yml

* mv data and fix bug

* bugfix

* add st of mstx

* Revise inspection comments

* change st name

* Add data description

* add data file

* delate st --timeout=300

* delate img.png

* revise data_directory

---------

Co-authored-by: gcw_fbonFwWl <gcw_fbonFwWl@noreply.gitcode.com>

* [data] feat: add DataChecker (#28)

* [data] feat: add base data class framework with validation support (#26)

* [data] feat: add base data class framework with validation support

Signed-off-by: Debonex <debonexx@gmail.com>

* [refactor] move data module to rl_insight package and fix imports

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] refactor: simplify validation system with class method approach

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] feat: update data related interfaces and tests

Signed-off-by: Debonex <debonexx@gmail.com>

---------

Signed-off-by: Debonex <debonexx@gmail.com>

* [data] feat: add datachecker (#27)

* add datachecker

* Delete rl_timeline.html

* Unify logging and address other code review comments

* update  requirement

* address code review comments from gemini

* adjust ut

* pre-commit

* pre-commit

* format uworkflow yml

* Address review comments

---------

Signed-off-by: Debonex <debonexx@gmail.com>
Co-authored-by: Debonet <37174444+Debonex@users.noreply.github.com>

* [mstx] fix: do not perform filtering when total_rank is too small (#34)

* do not perform filtering when total_rank is too small

* Update rl_insight/parser/parser.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [mstx] fix: skip mstx_preprocessing if necessary (#32)

* skip mstx_preprocessing if necessary

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update examples/mstx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update mstx_preprocessing.py

* Update mstx_preprocessing.md

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [visualizer] feat: optimize the encapsulation and registration of the visualizer (#30)

* Optimize the encapsulation and registration of the visualizer

* update pr-title check

* [data] feat: add data check rule func ParserOutputValidatorRule (#36)

* add rule func ParserOutputValidatorRule

* fix check error

* fix data checker test error

* fix start_time_ms error

* Fixed review comments and added test cases.

* rename docs/data/data_directory.md to  docs/data/data_specification.md
add empty summary event data test case

* fix data directory reference

---------

Co-authored-by: zhangning <zhangning42@huawei.com>

* [ci] test: add doc url validation check (#43)

test: add doc url check in CI

Co-authored-by: zhangning <zhangning42@huawei.com>

* [doc] chore: add guideline for offlinepipeline (#40)

* doc optimization

* add guideline

* Update main.py

* Update architecture_and_guideline.md

* [visualizer] fix: delete the remaining visualizer type. (#39)

* remove vis_type from visualizer

* update ut

* [data] feat: add VerlLogData validation rules (#35)

* [data] feat: add VERL_LOG validation rules and local check script

* fix(data): refine VERL_LOG validation per review

- Move VeRL rules to verl_log_rules.py; rename PresentRule to ExistRule
- Require single .log file path; share path validation helper
- CLI errors to stderr; update tests and data_directory.md

* docs(data): use data/verl_data for VeRL log sample paths

Made-with: Cursor

* chore(data): track VeRL sample logs under data/verl_data (override *.log)

Made-with: Cursor

* feat(data): require Training Progress in VeRL logs; trim full sample

- Add Training Progress: to DEFAULT_REQUIRED_KEYWORDS and docs
- Replace good_full_verl.log with good_minimal_verl.log; add bad_* fixtures
- Update tests; data spec lists only minimal log for check example

Made-with: Cursor

* style: ruff-format data rules; Apache header on check_verl_log

- Keep shebang first in check_verl_log.py for direct execution

Made-with: Cursor

* [visualizer] feat: add RL timeline PNG generator (#38)

* [visualizer]feat:update generate timeline type of PNG

* [visualizer]fix:add ut and code optimization

* [visualizer]feat:add ut

* [visualizer] fix: clean code and fix rank sorting

* fix:update gemini assist

* fix:update gemini assist again

* fix:update gemini assist

* fix:fix gemini assist

* fix:update pre-commit

* fix: update ut

* fix: update ci test again

* fix: update ci test again

* feat: add kaleido depend

* fix: test ci again

* fix: test ci again

* fix: test ci again

* fix: test ci again

---------

Co-authored-by: ChenGary13 <chenjiayu31@huawei.com>

* [data] feat: add verify rules for MSTX profiling files (#41)

* Add validation rules for Mstx JSON files

1、add MstxJsonFileExistsRule
2、add MstxJsonFieldValidRule

* Added tests for MstxJsonFileExistsRule and MstxJsonFieldValidRule

Added tests for MstxJsonFileExistsRule and MstxJsonFieldValidRule.

* Update data_checker.py

* split MULTI_JSON to MULTI_JSON_MSTX and MULTI_JSON_TORCH.

* Fix bug

* modify 'multi_json' to 'multi_json_mstx'

* fix 'tmp_path' bug in test case

* use pre-commit code consistent

---------

Co-authored-by: pqhgitee <pqhgitee@noreply.gitcode.com>

* [doc] fix: Update roadmap links in README.md

* [data] feat: add verify rules for Torch profiling files (#52)

* Add validation rules for Torch Profile files

* Update rl_insight/data/rules.py

changing this to > 0

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [doc] refactor: reorganize documentation structure (#51)

* Update data_specification.md

* Update data_specification.md

* update pipeline docs&mstx_exec

* update

* update docs

* update docs

* clean DS_Store

* clean

* Update index.rst

* update

* Enhance input/output section with directory structure

Add directory structure example for input data in documentation.

* Update input/output section in documentation

Removed example directory structure for input data.

* Update baseclusterparser_interface.md

* [parser, data, doc, ci] feat: add nvtx parser function (#50)

* add nvtx parser function

* Update examples/nvtx_exec.sh

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* fix error

* resume pre-commit yaml

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update docs/cluster_analysis.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* add nvtx data and e2e test script

* add verify rules for nvtx data

* refactor nvtx parser

* Update rl_insight/parser/nvtx_parser.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [visualizer, pipeline] feat: add MoE expert load heatmap visualization (#47)

* support moe load visualization with heatmap

* standardize gmm_exec.sh

* [parser,data] fix: harden gmm parsing, filtering, and output schema

* [pipeline,cfg] refactor: normalize gmm cli parameters and config passing

* [visualizer] refactor: improve gmm heatmap layout and segment rendering

* [doc] docs: update gmm guide, example script and sample heatmap asset

* [test,doc] test: add st for gmm_heatmap and document dump data layout

* refactor: move GMM CLI groups to parser/visualizer packages for centralized registration

* add minimal data example & format code (pre-commit compliant)

* fix(types): resolve parser mypy errors for nvtx and shared DataMap keys

---------

Co-authored-by: chenjiao.angel <chenjiao.angel@bytedance.com>

* [misc] fix: harden parser validation and stabilize cross-platform test behavior (#56)

* fix: improve parser robustness, cross-platform path handling, and test stability

* docs: refine RL timeline quickstart profiling link text

* fix: improve GMM parsing and restore scalable heatmap metadata

* fix: restore matplotlib gmm visualizer

* fix: restore original gmm visualizer

* fix: improve MSTX ordering and harden docs URL validation

* fix: accept Path inputs in validators and normalize GMM paths

* fix: apply pre-commit cleanup for validator tests

* update requirements.txt

* fix ut error & update requirements

* [pipeline] feat: rl insight support online monitor (#53)

rl insight support online monitor

* [cfg] refactor: Use omegaconf structured configuration system (#58)

* Use omegaconf structured configuration system and migrate CLI

* Modular configuration of parameters and reduction of redundancy.

* [parser] feat: add memory parser (#54)

* [memory] feat: add memory parser

* [memory]fix: fix reviews

* [memory]fix: fix pre check problems

---------

Co-authored-by: mookies1 <zhanghaoyong1@huawei.com>

* [pipeline] feat: rl monitor add grafana config file (#60)

rl monitor add grafana config file

* [pipeline] feat: rl online monitor statetimeline optime (#61)

rl online monitor statetimeline optime

* [doc] fix: fix pypi display issue and expand readme with more details (#65)

* Update README.md and provide more details

* [doc] fix: align quickstart installation instructions with README (#67)

Unify installation sections in timeline and GMM heatmap quickstarts to match README: pip install first, then optional source install.

* [misc] fix: remove dead code (#71)

* remove dead code

* Update schema.py

* [pipeline] refactor: support server backend (#73)

* [pipeline] fix: online monitor bug fix and grafana json update (#80)

* optim rl insight

* optim env config

* update grafana json

* remove docker dev guide

* refactor1 close rename to finish

* refactor2 add client base class

* refactor 3 add collector base

* refactor 4 move constant to constant.py

* refactor 5 move load config to utils

* refactor 6 config file path refactor

* refactor 7 readme refactor

* fix

* [BREAKING][pipeline] refactor: rl-insight refactor directory structure (#82)

rl-insight refactor directory structure

* [env, ci] fix: remove specific versions in requirements (#81)

* remove torch 2.7.1 requirement

* Update requirements.txt

* Update pyproject.toml

* [doc] refactor: rl-insight docs fix (#85)

rl-insight docs fix

* [monitor-server, monitor-config] feat: prometheus config file support add labels (#87)

* prometheus config file support add labels

* grafana dashborad support seprate replica

* gemini review fix

* fix dashboard

* fix title ci

* [doc, deployment] feat: flexible server install (#84)

* Keep the manifest binary field name consistent with config.

* Add binary_path field in config.yaml, with documented binary discovery fallback logic

* Make download URLs configurable

* Convert instance methods to static methods

* Print the planned download list before downloading, and support offline installation from a local archive directory.

* Update server_installation.md

* Add guard for --local-archive when the path is not a directory

* Remove duplicate _resolve_release calls across plan_install and install

* test bugfix

* pre-commit

* Fix url-check ci

* Gemini review

* remove approach 2

* Update server_installation.md

* Remove the URL template exposure from the config

* [doc] feat: readme optim and add demo grafana json (#90)

* readme optim

* Unify between README and actual files

---------

Co-authored-by: Xiaobo Hu <huxiaobo@zju.edu.cn>

* [monitor-config] fix: config optim (#92)

config optim

* [recipe] feat: add memory timeline HTML visualizer with DataChecker and e2e test (#93)

* [visualizer] feat: add memory timeline HTML visualizer

- Plotly-based interactive Gantt chart with operator grouping
- Time-window segmentation for large datasets (max 20 segments)
- Chart1 memory trend line + Chart2 operator Gantt synchronized zoom
- Segment navigation with smart hint (suggest best segment for range)
- Call stack display in detail panel on bar click
- Overlap detection: hover and click show all stacked bars per operator
- Absolute time display on x-axis tick labels and hover tooltip

* [visualizer] refactor: rename MEMORY_DATA to MEMORY_SUMMARY and optimize performance

* [memory] feat: add DataChecker rules, e2e test, and update docs for memory visualizer

- Add MEMORYKEYS schema, MemoryContentRule, and wire MEMORY_SUMMARY rules
- Fix memory_visualizer.py imports (rl_insight.* -> recipe.*)
- Add e2e test with sample data (data/recipe/memory_data/)
- Rename docs: memory_parser_* -> memory_guide/quickstart
- Update all docs for visualizer output, DataChecker, CLI format

* [memory] fix: add OmegaConf DictConfig type annotation to MemoryClusterParser init

* [recipe] feat: integrate memory analysis pipeline (#97)

* [pipeline] fix: add yaml and examples

* [pipeline] fix: Generate one HTML per rank

* [pipeline] feat: add Ascend Memory data checker

* [pipeline] feat: optim memory html

* [pipeline] feat: fix review

* [pipeline] feat: fix review

* Update recipe/visualizer/memory_template.html

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [pipeline] feat: fix review

* Update recipe/visualizer/memory_template.html

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update recipe/visualizer/memory_template.html

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update recipe/visualizer/memory_visualizer.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update recipe/visualizer/memory_visualizer.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [pipeline] feat: fix review

* Update recipe/visualizer/memory_visualizer.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [pipeline] feat: fix review

* [pipeline] feat: fix review

* [pipeline] feat: fix review

* Update recipe/data/rules.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update recipe/data/rules.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [pipeline] feat: fix review

---------

Co-authored-by: mookies1 <zhanghaoyong1@huawei.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [BREAKING][monitor-server, monitor-config] feat: rl-insight support server api and supprot ipv6 (#96)

* [monitor-config] feat: Json update to support transfer queue monitor (#99)

grafana json update support transfer_queu

* [monitor-api] fix: trace api optim (#101)

trace api optim

* [monitor-collector] fix: opentelemetry log optim (#102)

opentelemetry log optim

* [monitor-config] feat: json update to support verl v1 trainer better (#103)

json update

* [ci] feat: add st (#100)

* add st

* pre-commit

* refactor test

* [BREAKING][monitor-api] refactor: api rename to promethues style (#104)

api rename to promethues style

* [ci] test: rl-insight monitor add ci (#86)

monitor e2e ci

* [doc] fix: readme update (#109)

* readme update

* Update README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* [ci] test: monitor add ut (#107)

monitor add ut

* [monitor-config] feat: add verl_tainer_v1_with_sglang_engine (#108)

add verl_tainer_v1_with_sglang_engine

* [deployment] feat: rl-insight log optim (#110)

rl-insight log optim

* [monitor-api, doc] feat: rl-insight support hardware (#111)

rl-insight support hardware

* [monitor-server] fix: ipv6 support and actor isolation (#112)

* bugfix for ipv6 support

* pre-commit

* Update test_prometheus_utils.py

* Update test_prometheus_utils.py

* job-level isolation, remove the detached lifecycle

* pre-commit & ci fix

* keep the original try-catch logic, ci update

* Update ray_monitor_client.py

* fix: catch RayActorError when health check detects dead actor

* Update ray_monitor_client.py

* test: remove health-check assertions and dead-actor test, no longer applicable

* [doc] feat: docs optim (#113)

docs optim

* [misc] chore: update version to 0.2 (#114)

Update pyproject.toml

* [doc] fix: docs update (#115)

docs update

---------

Signed-off-by: Debonex <debonexx@gmail.com>
Co-authored-by: JIANG-PENGJUN <52533600+756017542@users.noreply.github.com>
Co-authored-by: gcw_fbonFwWl <gcw_fbonFwWl@noreply.gitcode.com>
Co-authored-by: Debonet <37174444+Debonex@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: tifning <44485561+tifning@users.noreply.github.com>
Co-authored-by: zhangning <zhangning42@huawei.com>
Co-authored-by: duesdues <55529526+duesdues@users.noreply.github.com>
Co-authored-by: Gary-cjy <71553064+Gary-cjy@users.noreply.github.com>
Co-authored-by: ChenGary13 <chenjiayu31@huawei.com>
Co-authored-by: panqihan <39796558+pqhgit@users.noreply.github.com>
Co-authored-by: pqhgitee <pqhgitee@noreply.gitcode.com>
Co-authored-by: a550580874 <82751568+a550580874@users.noreply.github.com>
Co-authored-by: zhengxiaojun <alwayszxj@gmail.com>
Co-authored-by: hswei88 <129183149+hswei88@users.noreply.github.com>
Co-authored-by: chenjiao.angel <chenjiao.angel@bytedance.com>
Co-authored-by: Zhen <295632982@qq.com>
Co-authored-by: TMC <87188729+mengchengTang@users.noreply.github.com>
Co-authored-by: Ruowei Zheng <892882856@qq.com>
Co-authored-by: Moocharr <1123277477@qq.com>
Co-authored-by: mookies1 <zhanghaoyong1@huawei.com>
Co-authored-by: zyang6 <zhouyang271@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants