The 500M instruction model improved against a learned reward model (best-of-N). Ask a question or give an instruction.
The 500M instruction model aligned with RLAIF: a Bradley-Terry reward model trained on the pairs, then best-of-N reward-weighted fine-tuning, on 4,000 AI-feedback preference pairs spanning closed-book QA and instruction-following failure modes (wrong figures, invented citations, broken format constraints, ignored instructions), prompts held out of every SFT set. Lineage: base → QA SFT → instruction SFT → RLAIF v2.
Served scale-to-zero on Modal, so the first request may take ~20–60s while the model wakes.