Living research notebook · started September 2026

Read papers to build a model you can redraw.

A public survey and pre-read guide. Each entry points to a verified source, a stored PDF, the figures that matter, and questions to answer from memory. The diagrams remain yours to reshape in draw.io.

1 paper reviewedAgentic reinforcement learningPDFs stored separately in Drive

Start with MiMo-V2.6

LLM-Core Xiaomi · technical report · 44 pages · reviewed 27 Sep 2026

Take-home message

Large agentic RL in this report is a systems recipe: more rollout throughput, more varied tasks and harnesses, and better grading signals must work together with stability controls.

The stored Drive copy allows anyone with its link to view this report. The reviewed PDF's SHA-256 begins fb81e6e08380.

Scale the rollouts

Each reported RL step uses 1,568 prompts with 16 rollouts each: 25,088 trajectories and 2.7–3.7B training tokens. Pro and Flash RL costs are reported as about $2.6M and $0.9M. §4.1, p. 8.

Diversify the experience

The task mix covers code, general tools, visual design, context following, and cybersecurity. Multi-harness code training is tested on held-out harnesses. §5.1, p. 20; Figure 10, p. 23.

Improve the feedback

GRS makes task-specific rubrics offline; GAR compares passing and failing trajectories online, then shifts advantage toward better passing solutions. §4.3; Figure 7, p. 17.

Interpretation boundary. The full Pro/Flash gains are measured for an integrated recipe. They do not isolate the contribution of each scaling axis. Several evaluations are internal, and the report describes a finite 30-step RL run rather than open-ended recursive self-improvement.

Read in four passes

Use the guide as a prediction to test, then correct it from the paper.

1 · ClaimAbstract and introduction, pp. 1–3. What exactly is being scaled?
2 · Mechanism§4.1 and §4.3, pp. 8–9 and 16–19. Trace the rollout and reward paths.
3 · Evidence§5, pp. 20–24. Compare metrics, held-out harnesses, and stability.
4 · System & release§6–7, pp. 26–35. Map infrastructure and released experiments.

Figures worth studying

Open the PDF at the printed page; redraw Figure 7 without looking back.

Figure 3 · p. 8

Score versus RL cost

DeepSWE average@3 rises Pro 58.4→72.6 and Flash 48.7→65.7 over the reported run. Inspect the cost split too.

Increasing cost accompanies several simultaneous changes.

Figure 7 · p. 17

Two grading paths

GRS: offline rollouts → solution/behavior rubrics → per-rollout scores. GAR: online group comparison → hack check → advantage redistribution.

This is a mechanism diagram, not an efficacy experiment.

Figure 8 · p. 19

GAR comparison

In code-only RL, online grading accompanies steadier turns/token lengths and sustained pass-rate gains.

Separate setting with batch size 128, not the full mixed-task run.

Figures 9–11 · pp. 22–24

Progress, transfer, stability

Follow benchmark scores and token counts; then held-out harness Pass@1; then expert load with and without router freeze.

Average@3, Pass@1, and avg@n are different metrics.

MiMo-V2.6 Figure 7: GRS creates reusable rubrics from offline rollouts; GAR grades passing patches online and redistributes advantage while zeroing confirmed hacks
Figure 7, p. 17, cropped from the MiMo-V2.6 technical report by LLM-Core Xiaomi. This diagram explains the grading routes; its bars are schematic, not measured gains. Redraw the two paths from memory after reading §4.3.

The diagram to rebuild

A compressed map for draw.io. The arrows describe the training flow; the reported figures test selected parts of it.

Task mix
+ harnesses
Grouped
rollouts
Tests +
groupwise grader
Sequence
advantages
Policy
update
Throughput: async large batch + Sample MixerFeedback: GRS / GAR + hack checksStability: frozen MoE router + train/inference consistency

In your own map, add the environment, grader, control/data plane, and evaluation as separate branches. Label observed benchmark changes beside the diagram, rather than turning them into mechanism arrows.

Close the paper and recall

  1. Why can two passing patches receive different learning signals under GRS or GAR?
  2. Which result supports transfer to a held-out harness, and what remains untested?
  3. Draw the path from a mixed task sample to a policy update. Where do the grader, Sample Mixer, and router freeze matter?

Answer these before checking §4–6. Revisit after one day and one week; send corrections to Codex for the survey.

Survey · agentic RL at scale

Question: Which system components enable useful large mixed-task training?

PaperContributionBest evidenceReservation
MiMo-V2.6 (2026)Scales rollout volume, environment/harness diversity, and grader compute with supporting infrastructure.Figures 3, 7–11; Table 6.One model family; internal benchmarks; limited component isolation.

One reviewed paper is a starting point, not a cross-paper consensus. Next comparison: controlled ablations at a fixed rollout budget with public held-out evaluations.

How a paper enters the loop

1 · Codex prepares

Verify the PDF and version, store a copy, update the survey, and write a source-grounded brief with figures and a diagram seed.

2 · You reconstruct

Read with the brief in mind. Draw the causal or procedural map in draw.io from memory. Keep personal derivations in OneNote when helpful.

3 · Revise together

Send your corrected diagram, unanswered questions, and disagreements. Codex updates the brief and survey; revisit the recall prompts later.