| #799 — fix: preserve optional fields in structured_outputs strictify_schema_json |
@jkyi-nvidia |
review requested |
2026-03-02 20:59 UTC |
2026-03-03 20:59 UTC |
123d 2h |
| #879 — Yev/scicode env v1 |
@dhruvnathawani, @jubick1337 |
review requested |
2026-03-20 18:09 UTC |
2026-03-23 18:09 UTC |
109d 5h |
| #998 — Add ruler answer prefix |
@bxyu-nvidia, @cmunley1 |
review requested |
2026-04-02 23:04 UTC |
2026-04-03 23:04 UTC |
100d 0h |
| #1005 — fix: bound retries in server_utils.request() |
@bxyu-nvidia |
review requested |
2026-04-15 08:24 UTC |
2026-04-16 08:24 UTC |
91d 15h |
| #946 — Turing/ Playwright-based CUA environment with multi-provider adapter support |
@bxyu-nvidia |
review requested |
2026-04-29 19:09 UTC |
2026-04-30 19:09 UTC |
81d 4h |
| #1193 — docs: address repo-health audit nits in README and CONTRIBUTING |
@cmunley1 |
review requested |
2026-04-30 13:30 UTC |
2026-05-01 13:30 UTC |
80d 9h |
| #1264 — feat: artifact-root resolution for external plugins |
@ananthsub, @bxyu-nvidia |
review requested |
2026-05-07 20:20 UTC |
2026-05-08 20:20 UTC |
75d 3h |
| #1334 — structured_outputs: add LLM-as-a-judge semantic verification |
@jkyi-nvidia |
review requested |
2026-05-18 20:37 UTC |
2026-05-19 20:37 UTC |
68d 2h |
| #1491 — Add compute_eval benchmark and resource server |
@ananthsub, @bxyu-nvidia |
review requested |
2026-06-02 00:07 UTC |
2026-06-03 00:07 UTC |
57d 23h |
| #1560 — docs: clarify resources server example-data contract and license validation |
@cwing-nvidia, @lbliii |
review requested |
2026-06-10 12:04 UTC |
2026-06-11 12:04 UTC |
51d 11h |
| #1541 — Add SC_Bench benchmark |
@adil-a, @ananthsub, @bxyu-nvidia, @cwing-nvidia |
review requested |
2026-06-13 00:57 UTC |
2026-06-15 00:57 UTC |
49d 22h |
| #1598 — Eshachar/openclaw swe |
@sdevare-nv |
review requested |
2026-06-16 18:19 UTC |
2026-06-17 18:19 UTC |
47d 5h |
| #1653 — docs: remove model recipe column from RL framework compatibility table |
@cwing-nvidia |
review requested |
2026-06-23 03:23 UTC |
2026-06-24 03:23 UTC |
42d 20h |
| #1752 — fix: handle vllm context length errors in nano v3 recipe |
@yfw |
review requested |
2026-06-26 05:53 UTC |
2026-06-29 05:53 UTC |
39d 17h |
| #2026 — feat: Switchyard token-capture integration for OpenHands SWE runs (zero-fork) |
@bxyu-nvidia |
review requested |
2026-07-14 22:36 UTC |
2026-07-15 22:36 UTC |
27d 0h |
| #1964 — feat(model-server): add OpenAI-compatible /v1/embeddings endpoint |
@ffrujeri |
review requested |
2026-07-15 16:18 UTC |
2026-07-16 16:18 UTC |
26d 7h |
| #1727 — fix(vllm_model): handle explicit null metadata in preprocessing |
@ananthsub |
review requested |
2026-07-15 17:11 UTC |
2026-07-16 17:11 UTC |
26d 6h |
| #2048 — feat: sandbox_agent swe resources server |
@adil-a, @ananthsub, @bxyu-nvidia, @hemildesai |
review requested |
2026-07-15 21:13 UTC |
2026-07-16 21:13 UTC |
26d 2h |
| #2052 — hermes_agent: repin to upstream hermes and fix the wrapper for its API |
@Glorf |
review requested |
2026-07-17 22:38 UTC |
2026-07-20 22:38 UTC |
24d 0h |
| #2141 — feat: add switchyard_model server for router-backed evals |
@ananthsub, @bxyu-nvidia |
review requested |
2026-07-27 19:55 UTC |
2026-07-28 19:55 UTC |
18d 3h |
| #2180 — feat(token-id-capture): resolve each call's parent at request time |
@pthombre |
review requested |
2026-07-29 13:59 UTC |
2026-07-30 13:59 UTC |
16d 9h |
| #2181 — feat(vllm-model): supply the previous call's exact training tokens |
@pthombre |
review requested |
2026-07-29 14:00 UTC |
2026-07-30 14:00 UTC |
16d 9h |
| #2155 — [GDPVal] Fix Stirrup reference paths and make head-server wait configurable |
@agronskiy |
review requested |
2026-07-29 16:40 UTC |
2026-07-30 16:40 UTC |
16d 6h |
| #2160 — [GDPVal] Fix what reaches the rubric judge, and stop silent zeros |
@agronskiy |
review requested |
2026-07-29 16:40 UTC |
2026-07-30 16:40 UTC |
16d 6h |
| #2162 — [GDPVal] Add model-agnostic Slurm orchestration for multi-allocation rollouts |
@agronskiy |
review requested |
2026-07-29 16:40 UTC |
2026-07-30 16:40 UTC |
16d 6h |
| #2301 — feat(sandbox): add E2B sandbox provider |
@hemildesai |
review requested |
2026-08-04 13:34 UTC |
2026-08-05 13:34 UTC |
12d 9h |
| #2304 — default port range to the 5000-5999 band |
@adil-a |
review requested |
2026-08-04 13:48 UTC |
2026-08-05 13:48 UTC |
12d 9h |
| #2311 — docs: rewrite Configure Agents index and document external agent harnesses |
@cwing-nvidia, @lbliii |
review requested |
2026-08-04 14:30 UTC |
2026-08-05 14:30 UTC |
12d 8h |
| #2386 — feat: harden SWE against filesystem overflow |
@sdevare-nv |
review requested |
2026-08-06 17:31 UTC |
2026-08-07 17:31 UTC |
10d 5h |
| #2142 — Add EnterpriseOps-Gym benchmark: resources server, benchmark registration, and per-turn telemetry agent |
@bxyu-nvidia |
review requested |
2026-08-08 08:52 UTC |
2026-08-10 08:52 UTC |
9d 14h |
| #2427 — feat: automation bench, general-agent, tau2-synth environments and verifiers 0.3 bump |
@ananthsub, @d-molinari |
review requested |
2026-08-07 22:39 UTC |
2026-08-10 22:39 UTC |
9d 0h |
| #2172 — feat: make Legal Agent Bench agent harness and sandbox-backends configurable |
@cmunley1 |
review requested |
2026-08-10 16:56 UTC |
2026-08-11 16:56 UTC |
8d 6h |
| #2481 — feat: swe_trace_converter change for the opencode backend |
@bxyu-nvidia, @sdevare-nv |
review requested |
2026-08-11 18:44 UTC |
2026-08-12 18:44 UTC |
7d 4h |
| #2458 — feat: deepswe environment |
@Glorf, @hemildesai |
review requested |
2026-08-12 04:41 UTC |
2026-08-13 04:41 UTC |
6d 18h |
| #2507 — fix(anthropic): make Messages conversion loss-aware |
@bxyu-nvidia, @cmunley1, @ffrujeri, @linj-glitch |
review requested |
2026-08-12 06:45 UTC |
2026-08-13 06:45 UTC |
6d 16h |
| #2383 — feat: judge endpoint resiliency |
@e-dobrowolska |
review requested |
2026-08-12 16:43 UTC |
2026-08-13 16:43 UTC |
6d 6h |
| #2346 — feat(sandbox): add Tenki provider |
@hemildesai |
review requested |
2026-08-12 16:45 UTC |
2026-08-13 16:45 UTC |
6d 6h |
| #2420 — feat: strands agent |
@Glorf, @ffrujeri |
review requested |
2026-08-13 17:47 UTC |
2026-08-14 17:47 UTC |
5d 5h |
| #2457 — feat: terminus 2 agent integration |
@bxyu-nvidia, @hemildesai |
review requested |
2026-08-13 17:51 UTC |
2026-08-14 17:51 UTC |
5d 5h |
| #2151 — feat: agentic math environments |
@ananthsub, @ffrujeri |
review requested |
2026-08-13 18:23 UTC |
2026-08-14 18:23 UTC |
5d 5h |
| #2120 — [rollout-observability][7/7] Add SWE agent rollout observations |
@mlazuka |
review requested |
2026-08-17 16:09 UTC |
2026-08-18 16:09 UTC |
3d 7h |
| #2592 — Fix compute_pass_majority_metrics: score unanswered rollouts as failures in pass@k |
@gwarmstrong, @jkyi-nvidia |
review requested |
2026-08-18 16:55 UTC |
2026-08-19 16:55 UTC |
2d 6h |
| #2600 — feat: record the routing condition an eval runs under |
@ananthsub, @bxyu-nvidia |
review requested |
2026-08-18 19:04 UTC |
2026-08-19 19:04 UTC |
2d 4h |
| #2474 — Add repeat-level variability statistics to reward profiling |
@mlazuka |
review requested |
2026-08-18 20:46 UTC |
2026-08-19 20:46 UTC |
2d 2h |
Pull request handoff SLA tracker
Review during the daily stand-up alongside the architecture/RFC decision tracker.
The SLA is one business day per handoff; weekends are excluded and public holidays are not modeled.
Drafts, bots, and PRs labeled
stale,paused, ordo-not-trackare excluded.Last refreshed: 2026-08-21T23:27+00:00
Newly breached (2)
Waiting on reviewers (2)
Older breach backlog (87)
Waiting on reviewers (44)
Waiting on authors (5)
Needs review assignment (38)