Supervised Fine Tuning
Topic archive • 1 matches
2026-09-12
Technology
Researchers train Nemotron 3 Ultra checkpoints to generate olympiad math proofs: Researchers trained two specialist checkpoints from Nemotron 3 Ultra using supervised fine-tuning and reinforcement learning to generate natural-language proofs for olympiad mathematics. The resulting test-time-compute pipeline operates entirely in natural language without formal provers or external tools.
Nemotron • arXiv
Permalink