A frontier model for long-running agents and ambitious interactive and visual work.
Our most capable model — 1M context, frontier vision and tool use.
Sunlight 2 Pro is our frontier — built for agents that run for hours and ship production code. With 1M tokens, native tool use, and state-of-the-art vision (86.2% on MMMU), it leads on SWE-bench Verified (71.4%) and AIME 2024 (52.1%). It reasons with verifiable traces and can self-correct over long horizons. If Sunlight 2 is the sprinter, Pro is the marathon runner.
400B dense transformer with MoE inference, 32k native generation, and 16B active parameters. Trained on 18T tokens including 4T code + vision + trajectory data. RL with verifiable rewards on SWE-bench, terminal tasks and visual diffs. Architecture shares weights with Sunlight 2 but with 4× depth and extended reasoning. Safety-tuned with constitutional AI and red-team coverage.
Frontier results — Sunlight 2 Pro vs strongest closed models. Higher is better.
| Benchmark | Sunlight 2 Pro | Claude 4 Opus | GPT-4o | Gemini 2.0 |
|---|---|---|---|---|
| SWE-bench Verified | 71.4% | 62.1% | 48.1% | 55.3% |
| AIME 2024 (pass@1) | 52.1% | 38.9% | 30.1% | 34.5% |
| GPQA Diamond | 61.8% | 53.4% | 49.9% | 51.2% |
| MMMU (vision) | 68.9% | 62.0% | 58.4% | 65.1% |
| LiveCodeBench | 68.9% | 60.2% | 53.2% | 57.8% |
| MMLU | 88.9% | 87.2% | 86.1% | 87.5% |
Evaluations run with identical prompts, temperature 0, 1× attempt. SWE-bench Verified = 500 human-verified tasks. See model card for details.