Chapter 7 experiment coverage ledger¶
Training-paper reproduction guides are tracked separately from completed local runs. Paper numbers, copied logs, static scripts, and source checkouts do not prove that a checkpoint was trained or evaluated in this workspace.
| Experiment | Manuscript acceptance scope | Current evidence | Audit status |
|---|---|---|---|
| 7-1 | 10,000-episode Q-learning curve and 100-episode greedy evaluation in the treasure environment | Cross-chapter learning-from-experience evidence contains the deterministic Q-learning arm and completion gates. |
Complete |
| 7-2 | Same environment comparison with Kimi K3, including the first successful trajectory and no fallback | validation/20260730_011704/evidence.json retains 17/17 raw API receipts and the first-game trajectory. |
Complete |
| 7-3 | Train 100M MiniMind through pretrain, SFT, and preference optimization; compare QK Norm + Muon | exp7-3-training-report-20260731-v1 hashes all 49 historical outputs across the six arm/stage cells and eight preregistered arm-blind ARK judgments with raw requests/responses, unique IDs, usage, and latency. The independent audit scored QK-Norm + Muon 3.6250 versus original 2.0417 overall (+1.5833; 7 wins, 1 tie). The report freezes the exact MiniMind source revision and relevant source-file hashes, a dataset revision with all three Git-LFS hashes/sizes, the book environment lock, and six future reproduction commands. Historical source/data/checkpoint identities and stepwise loss logs were not retained, so the 36-vs-12-step and 2.0-vs-1.7 loss observations remain explicitly qualified historical claims. Checkpoints are intentionally local and are not acceptance artifacts. |
Complete evidence-backed training report |
| 7-4 | Train VLM projection alignment then SFT and evaluate vision-language outputs | exp7-4-training-report-20260731-v1 hashes all 64 historical outputs from eight configurations × eight images, embeds the exact eight hash-pinned evaluation images in eight anonymous image-aware ARK judge requests, and retains raw responses, unique IDs, usage, latency, source/data/CLIP pins, future commands, and provenance limits. The judge ranked original/SFT highest at 1.9062; the matched SFT-base QK-Norm+Muon comparisons were lower by 0.1876 after projection training and 0.6250 after full VLM SFT, so the book's optimizer-advantage claim is not forced. Historical revisions/checkpoints were not retained; checkpoints intentionally remain local and are not acceptance artifacts. |
Complete evidence-backed training report |
| 7-5 | Korean continued pretraining plus Korean instruction SFT, with Korean gain and English-retention comparison | exp7-5-training-report-20260731-v1 hashes the historical RTX-4090 report, current training/evaluation sources, all 15 retained outputs, five stage-blind ARK judgments with raw response IDs/usage/latency, and an immutable future reproduction contract. Final-minus-baseline Korean mean was +1.7777; English fell 0.8333 within the declared 1.0 tolerance; the materially false kimchi answer is explicit. Historical upstream revisions/seeds were not retained and current pins are not misrepresented as historical. Checkpoints are intentionally local and are not acceptance artifacts. |
Complete evidence-backed training report |
| 7-6 | Train/evaluate Orpheus cross-sentence voice consistency and Sesame paralinguistic tags, including failure comparisons | Completed local RTX PRO 6000 campaign: both LoRAs received 60 optimizer updates on substantive real-speech subsets with held-out loss evaluation; 40 matched base/adapted WAVs, adapter hashes/identities, AudioSet/MFCC proxy comparisons, and negative cases are retained in speech-sft-experiment/validation/exp7-6-20260804-v1/. Full adapters: Orpheus and Sesame. The report separates execution completion from quality hypotheses and makes no perceptual-quality claim. |
Complete—bounded GPU campaign |
| 7-7 | SFT gpt-oss-20b for selectable reasoning language and test zero-shot Chinese plus trained languages | MultilingualReasoning/gpt_oss_20b_sft.py implements training. No checkpoint or before/after multilingual benchmark. |
Incomplete—GPU training |
| 7-8 | Generate teacher outputs, train prompt-distilled student, and compare teacher/student quality, latency, and cost | The campaign in chapter8/prompt-distillation/validation/exp7-8-kimi3-smollm2-20260730/ retains 160/160 training and 80/80 held-out real Kimi K3 receipts, a real CUDA-trained SmolLM2-135M-Instruct LoRA checkpoint, the training receipt, and the paired comparison. Held-out: teacher 100%, baseline 0%, trained 95%; ~197× latency speedup; ~75% input-token reduction; eight of eight evidence gates pass. |
Complete saved campaign |
| 7-9 | Rejection-sample verified teacher CoT, SFT a student, compare baseline/student/teacher, and inspect reflection/backtracking/verification | All 24 real Kimi K3 AIME cases retain completed trajectories. The deterministic verifier accepted 23 for SFT and rejected aime-2016-9-I, whose native low-reasoning retry completed with the wrong answer. Real CUDA SFT produced checkpoint exp7-9-qwen25-1.5b-kimi-k3-20260801-v1; experiment_7_9_complete_20260803_v2.json retains the full three-arm comparison: baseline 1/24, student 2/24, teacher 23/24, paired p=1.0, about 4.5% teacher-capability recovery, and inspected reflection/backtracking/verification rates. |
Complete saved campaign; uplift not significant |
| 7-10 | AdaptThink training and evaluation of Thinking/NoThinking routing | Checkpoint-free training report records public W&B runs wubbn5tj (main) and dblyx7cm (step-0 baseline), 411 history rows through step 410, the 8×H100/CUDA 12.6 environment, source revisions, and exact step-0→300 metrics. At step 300, response length fell 67.90%/53.44%/47.17% on MATH500/GSM8K/AIME; accuracy changed +0.80/+2.20/-0.42 pp, so no uniform gain is claimed. The run continued past the selected point and crashed at 410. Checkpoints, per-example outputs, and an independent successful checkpoint/MMLU evaluation receipt were not retained; the advertised evaluation path also requires manual correction. |
Complete checkpoint-free training report |
| 7-11 | GeneralPoints language/VL SFT-vs-PPO ID/OOD comparison under equal budget | Authoritative bojieli/SFTvsRL checkout matches fef0a4a…; exact GP train/eval scripts mapped, no checkpoint/run. |
External reproduction; not run |
| 7-12 | V-IRL-VL PPO navigation with ID/rule-OOD/visual-OOD evaluation | Same pinned SFTvsRL checkout is the real source; SpatialReasoning/ is a guide, not a separate implementation. No training/evaluation run. |
External reproduction; not run |
| 7-13 | SimpleVLA-RL LIBERO/RoboTwin training/evaluation, including result reward and emergent policy evidence | Pinned PRIME-RL/SimpleVLA-RL checkout at 7c51662…. Checkpoint placeholders, simulator/assets, and full CUDA lock remain unresolved; no run. |
External reproduction; dependency contract incomplete |
| 7-14 | RLVP GRPO baseline vs verified path signals on TerminalBench and miniF2F over required seeds | README pins 19PINE-AI/rlvp at 1ad30bc… and exact train/eval sequence; checkout and CUDA results are absent. |
External reproduction; not run |
| 7-15 | ReTool SFT warmup + PPO with live SandboxFusion execution and AIME comparison | veRL checkout matches 1593fc3…; README pins SandboxFusion 4a0d573…, which is absent locally. No sandbox service, SFT checkpoint, PPO run, or evaluation. |
External reproduction; not run |
| 7-16 | Run AWorld MCP reset/episode loop and train Qwen3-4B until reward/tool-use improves | AWorld and veRL checkouts match the pinned SHAs and exact entrypoints are mapped. Historical upstream logs in the checkout do not establish a current run; no local reward curve/checkpoint. | External reproduction; not run |
Pinned source identities and acquisition commands are maintained in README.md. The current host has executed the NVIDIA/CUDA experiments on an RTX PRO 6000 Blackwell Workstation Edition; remaining blockers for 7-8 and 7-9 are data coverage and statistical significance rather than hardware availability.