Skip to content

Research Study

Topic archive • 1 matches

Back to home • GEO summary endpoint

2026-09-25

Technology

  • Study shows self-distillation can degrade multi-turn LLM agent performance: Researchers found that on-policy self-distillation teaches multi-turn LLM agents to act with confidence without the underlying information, sometimes performing worse than untrained base models. To address this, they proposed Privileged Self-Practice, which moves privileged information from the loss to the sampler.

    LLM • arXiv

    Permalink