MikhbarMIKHBAR
Artificial Intelligence

NVIDIA Fine-Tunes Nemotron for Gold-Level IMO and IOI Results

By combining supervised fine-tuning, reinforcement learning, and advanced test-time inference loops, NVIDIA adapted the Nemotron foundation model family to achieve gold-medal level performance at both the International Mathematical Olympiad and the International Informatics Olympiad.

NVIDIA Fine-Tunes Nemotron for Gold-Level IMO and IOI Results

Proving Adaptability Across Demanding Domains

The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test distinct skill sets under stringent conditions. While the IOI requires algorithms and code to pass hidden tests within strict time and submission limits, the IMO demands rigorous natural-language proofs. Achieving success at either competition is exceptionally difficult, but recent findings detailed on the Hugging Face Blog indicate that a single model family can master both challenges through targeted specialization.

NVIDIA teams demonstrated that starting from Nemotron 3, supervised fine-tuning, reinforcement learning, and feedback-driven inference can transform foundation models into world-class specialists. Rather than building a brand-new foundation model for every distinct challenge, the approach focused on adapting the existing architecture with a clear, reusable training recipe.

Specializing Nemotron for Competitive Programming

For competitive programming tasks, teams curated 22,000 problems and generated synthetic reasoning traces to train specialized architectures. Two primary variants were developed: Nemotron-3-Nano-CC, featuring 30 billion total parameters and 3 billion active parameters, which received both supervised fine-tuning and reinforcement learning; and Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, which underwent supervised fine-tuning.

Progression data from IOI 2025 highlighted the direct value of this specialization strategy. Nano improved from 130 points before post-training to 280 after supervised fine-tuning, and further to 291 after reinforcement learning. When paired with GenCorrect—an iterative generate-evaluate-refine test-time strategy—Nano reached 468 points, surpassing the gold-medal threshold of 438.3. The larger Ultra-CC configuration achieved 502 points using the same evaluation framework, while the competition-specific Ultra-CC system deployed for IOI 2026 scored 535.4 out of 600.

Teaching Nemotron to Prove, Check, and Revise

The parallel project focusing on olympiad mathematics adapted the Nemotron 3 Ultra architecture using both supervised fine-tuning and reinforcement learning checkpoints. The supervised fine-tuning corpus incorporated 414,890 quality-filtered examples spanning 15,818 unique proof problems. This dataset went beyond final answers, explicitly covering proof generation, refinement, verification, and meta-verification so the model could construct logical arguments, identify flaws, respond to critiques, and evaluate completeness.

Additionally, the reinforcement learning model was trained on 9,597 proof problems selected near its capability frontier. Because the SFT and RL checkpoints exhibited complementary strengths during development, the final IMO 2026 system integrated both specialist models alongside the general model. Operating entirely in natural language without formal provers or external tools, the system scored 30 out of 42 points—securing full credit on four of the six problems and exceeding the official gold threshold.

Combining Fine-Tuning and Test-Time Compute

The successful outcomes across both competitions emphasized that performance was not driven by fine-tuning alone or brute-force sampling alone. Instead, the results emerged from co-designing the underlying model, the training datasets, and the test-time inference loop. Better specialization provided the inference system with higher-quality candidates, more accurate critics, and more effective refinement cycles.

Further technical documentation regarding the mathematics framework is detailed in the academic literature associated with IMO 2026, providing deeper insight into the underlying methodologies and evaluation metrics.

Open Availability on Hugging Face

To encourage broader community research and application, NVIDIA released the Nemotron Labs IMO 2026 collection on Hugging Face. This release includes the supervised fine-tuning and reinforcement learning checkpoints, both training datasets, and Nemotron-IMO-Bench—a new benchmark comprising 200 olympiad-level math problems. Furthermore, the Nemotron-3-Ultra-CC competitive programming model and associated IOI inference pipelines have been made available for developers.

Sources

  • Hugging Face BlogOne Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Continue chronologically

You are readingNVIDIA Fine-Tunes Nemotron for Gold-Level IMO and IOI Results
Google Introduces Playground for Custom Games
Older storyGoogle Introduces Playground for Custom GamesOctober 7, 2026 · 3 min

Related entity coverage