This hands-on tutorial explores response engineering techniques that improve language-model outputs at inference time, including Best-of-N selection, self-consistency, Reflexion, and Mixture-of-Agents.
This notebook was developed and curated by Son The Nguyen for the NAIRR 2026 Training-Free Alignment of LLMs tutorial.