Commit 7af08d2
fix(kubeflow): stream logs once, not per replica
torchx calls scheduler.log_iter(app_id, role_name, k=...) once per replica
(k = 0..num_nodes-1). The Kubeflow log_iter ignored k and re-ran
fetch_logs — which tails the entire jobset via the jobset-name selector — for
every replica, producing N independent tail streams (each with its own dedup
state) and N-fold-duplicating every console line (prefixed <role>/<k>). At 16
nodes that's 16x the log volume, which also overruns the CI job-log limit on
long runs. Stream only for k == 0; that single tail already covers all ranks
(and writes log-allranks_0.out once).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: oliver könig <okoenig@nvidia.com>1 parent 636ec99 commit 7af08d2
1 file changed
Lines changed: 8 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
188 | 188 | | |
189 | 189 | | |
190 | 190 | | |
| 191 | + | |
| 192 | + | |
| 193 | + | |
| 194 | + | |
| 195 | + | |
| 196 | + | |
| 197 | + | |
| 198 | + | |
191 | 199 | | |
192 | 200 | | |
193 | 201 | | |
| |||
0 commit comments