Agentic benchmark

Grade the answer.
Keep the evidence.

2,700 trials across 225 hidden aerospace tasks. The primary measure is complete numerical agreement on valid calculation cases.

Trials
2,700
Tasks
225
Goldens
Runner-derived

Complete numerical agreement by arm

Each point requires every numerical value in the runner result, within tolerance. Object shape and unit-field formatting do not gate this measure.

Complete numerical agreement by benchmark armA lollipop plot comparing the rate at which every runner number was returned within tolerance across model arms.0%25%50%75%100%Nemotron 3.5 Lightning7.6%Nemotron 3.5 Lightning + skills19.3%GPT-5.6 Sol19.9%GPT-5.6 Terra15.6%
Nemotron 3.5 Lightning7.6%
Nemotron 3.5 Lightning + skills19.3%
GPT-5.6 Sol19.9%
GPT-5.6 Terra15.6%
Nemotron 3.5 Lightning skills uplift: 11.7% across 564 matched valid trials. A separate stricter runner-equivalent metric remains in the downloadable value artifact.

Inspect every numeric result.

See the model’s final numeric values beside the runner values for all 2,256 valid calculation trials. A missing final result is shown explicitly. This excludes prompts, inputs, reasoning traces, and response prose.

Loading value records…

TrialArmModel outputRunner outputMatchedOutcome

Match rule: max(1e-6, 0.5% relative error), one returned number per runner number. “All outputs” is the headline metric; “at least one” is a directional diagnostic, not a substitute for a complete answer.

What each arm cost

Provider-recorded when available. Direct OpenAI costs use standard published token rates and are clearly marked as estimates.

Nemotron 3.5 Lightning675 trials
Total$0.75
Per answer$0.00
Mean latency18.5s
Nemotron 3.5 Lightning + skills675 trials
Total$0.90
Per answer$0.00
Mean latency20.6s
GPT-5.6 Sol675 trials
Total$22.68
Per answer$0.03
Mean latency24.4s
GPT-5.6 Terra675 trials
Total$7.97
Per answer$0.01
Mean latency11.6s

The signal was there.

Terra returned every runner number on 15.6% of valid calculation cases, and at least one runner number on 51.2%. Its earlier 0.2% rate was the much stricter runner-equivalent measure.

It returned valid JSON for 100.0% of trials and correctly rejected 98.2% of invalid-input cases. Hidden operation selection was 1.6% correct; it remains diagnostic rather than a gate on numerical accuracy.

Terra numerical and contract componentsA dot plot showing valid JSON, numerical accuracy, operation selection, and runner-equivalent output.0%100%Valid JSON100.0%All numeric outputs15.6%Any numeric output51.2%Invalid classification98.2%Operation selected1.6%Runner equivalent0.2%
Numerical accuracy covers valid calculation tasks only. Invalid-input classification and runner-equivalence are separate measures.

Where the models landed

Each operation contributes 15 total trials per arm: three repeats of five cases. In the frozen run, 188 fixtures produced a successful runner golden and 37 were invalid-input goldens; skill rates aggregate the valid calculation trials.

  • Nemotron 3.5 Lightning
  • Nemotron 3.5 Lightning + skills
  • GPT-5.6 Sol
  • GPT-5.6 Terra

aerodynamics

15.7% skills

aircraft-performance

21.4% skills

atmosphere

30.0% skills

coordinate-frames

82.1% skills

flight-planning

24.5% skills

guidance-navigation-control

0.0% skills

orbital-mechanics

10.7% skills

propeller-systems

7.7% skills

rocket-analysis

13.9% skills

trajectory-optimization

0.0% skills

Read the study correctly.

This is runner-consistency and agent-orchestration evidence. It does not independently validate aerospace physics or operational suitability.

Controlled interface. Skills can read one relevant guide and use at most two local runners. Baselines receive the same prompt but no aerospace tools or files.

Bounded scope. Goldens derive from the bundled runtime. The fixed prompt set excludes weapon, targeting, operational control, dispatch, and safety-of-life scenarios.

Provider provenance. Nemotron 3.5 Lightning trials use OpenRouter’s provider-specific serving path. GPT-5.6 Sol is deterministically split across direct OpenAI and OpenRouter.

Read the complete methodology