ATRD Pipeline Workspace

Adaptive Test-Time Reasoning Distillation workspace. Fine-tuning Nemotron-3-Nano-30B on failure-grounded datasets with PRM-guided GRPO and adaptive budget-forcing.

PHASE 01

Curation

Identify failure modes and generate synthetic corrections.

PHASE 02

SFT Training

Instruction tune on Nemotron formatting tags.

PHASE 03

GRPO RL

Align steps via Process Reward Model policy updates.

PHASE 04

Budget Scaling

Extend search tokens dynamically for hard problems.

View Git Source
BASE: Nemotron-3-Nano-30BACCELERATION: TF32 matmulMAX RANK: LoRA Rank 32HARDWARE: RTX PRO 6000 Blackwell