Summary
Full convergence is demonstrated only for the 0.6B simple_gsm8k recipe. The
8B long_context_qa (LoongRL) and txt2sql (BIRD) recipes have so far only
been run for short windows (≤10 steps) showing healthy, monotonic learning,
not full convergence curves.
Validated so far
- simple_gsm8k (Qwen3-0.6B): eval
pass@1 0.29 → 0.64, reward 0.28 → 0.80
over 70 steps. ✅ Full convergence.
- long_context_qa 8B (LoongRL): reward climbing over 9 steps. ✅ Learning,
⏳ not a full curve.
- txt2sql 8B (BIRD): reward climbing over 10 steps. ✅ Learning, ⏳ not a
full curve.
Acceptance
Summary
Full convergence is demonstrated only for the 0.6B
simple_gsm8krecipe. The8B
long_context_qa(LoongRL) andtxt2sql(BIRD) recipes have so far onlybeen run for short windows (≤10 steps) showing healthy, monotonic learning,
not full convergence curves.
Validated so far
pass@10.29 → 0.64, reward 0.28 → 0.80over 70 steps. ✅ Full convergence.
⏳ not a full curve.
full curve.
Acceptance