Agentic benchmark
Grade the answer.
Keep the evidence.
2,700 trials across 225 hidden aerospace tasks. The primary measure is complete numerical agreement on valid calculation cases.
- Trials
- 2,700
- Tasks
- 225
- Goldens
- Runner-derived
Complete numerical agreement by arm
Each point requires every numerical value in the runner result, within tolerance. Object shape and unit-field formatting do not gate this measure.
Inspect every numeric result.
See the model’s final numeric values beside the runner values for all 2,256 valid calculation trials. A missing final result is shown explicitly. This excludes prompts, inputs, reasoning traces, and response prose.
Loading value records…
| Trial | Arm | Model output | Runner output | Matched | Outcome |
|---|
Match rule: max(1e-6, 0.5% relative error), one returned number per runner number. “All outputs” is the headline metric; “at least one” is a directional diagnostic, not a substitute for a complete answer.
What each arm cost
Provider-recorded when available. Direct OpenAI costs use standard published token rates and are clearly marked as estimates.
The signal was there.
Terra returned every runner number on 15.6% of valid calculation cases, and at least one runner number on 51.2%. Its earlier 0.2% rate was the much stricter runner-equivalent measure.
It returned valid JSON for 100.0% of trials and correctly rejected 98.2% of invalid-input cases. Hidden operation selection was 1.6% correct; it remains diagnostic rather than a gate on numerical accuracy.
Where the models landed
Each operation contributes 15 total trials per arm: three repeats of five cases. In the frozen run, 188 fixtures produced a successful runner golden and 37 were invalid-input goldens; skill rates aggregate the valid calculation trials.
- Nemotron 3.5 Lightning
- Nemotron 3.5 Lightning + skills
- GPT-5.6 Sol
- GPT-5.6 Terra
aerodynamics
15.7% skills
aircraft-performance
21.4% skills
atmosphere
30.0% skills
coordinate-frames
82.1% skills
flight-planning
24.5% skills
guidance-navigation-control
0.0% skills
orbital-mechanics
10.7% skills
propeller-systems
7.7% skills
rocket-analysis
13.9% skills
trajectory-optimization
0.0% skills
Read the study correctly.
This is runner-consistency and agent-orchestration evidence. It does not independently validate aerospace physics or operational suitability.
Controlled interface. Skills can read one relevant guide and use at most two local runners. Baselines receive the same prompt but no aerospace tools or files.
Bounded scope. Goldens derive from the bundled runtime. The fixed prompt set excludes weapon, targeting, operational control, dispatch, and safety-of-life scenarios.
Provider provenance. Nemotron 3.5 Lightning trials use OpenRouter’s provider-specific serving path. GPT-5.6 Sol is deterministically split across direct OpenAI and OpenRouter.
Read the complete methodology