Shockingly Simple Self-retrospection Improves Agentic Models Without RL
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior.
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior.
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point.