Research Study
Topic archive • 1 matches
2026-09-25
Technology
Study shows self-distillation can degrade multi-turn LLM agent performance: Researchers found that on-policy self-distillation teaches multi-turn LLM agents to act with confidence without the underlying information, sometimes performing worse than untrained base models. To address this, they proposed Privileged Self-Practice, which moves privileged information from the loss to the sampler.
LLM • arXiv
Permalink