Benchmark v2 · sealed evaluation complete

Prove the calculators.
Then test the agents.

Benchmark v2 separates calculator evidence from agent performance, then measures skill uplift across 360 aerospace tasks and eight direct-provider conditions.

Canonical dossiers
45/ 45
Development referents
45/ 45
Headline validated
0/ 45
Technical tool loops
8/ 8
Sealed trials
8,640/ 8,640

Sealed result · QA passed

01 · Calculator credibility

Forty-five operations. Zero shortcuts.

Every square is one public operation. “Code verified” means its schema and standalone execution pass—it does not mean the aerospace model is externally validated.

Empirically validated 0Reference validated 0Code verified 31Unverified 1Blocked 13

Headline rule: only operations that pass the public-source, isolated-reference, replay, negative-control, and output-specific acceptance gates count as reference_validated or empirically_validated.

02 · Release gates

Evidence before inference.

A passed API check cannot waive missing evidence. Each gate is derived from a versioned artifact. Technical snapshot: Aug 18, 2026 UTC.

  1. 01
    Operation dossiersSchemas, intended use, equations, domains, limitations
    45 / 45pass
  2. 02
    Open referent coveragePublic source traces plus isolated reference artifacts
    45 / 45pass
  3. 03
    Development evidence32/32 replay and 32/32 known-bad rejection
    Passpass
  4. 04
    Sealed holdoutEight unique tasks per operation, hash-frozen
    360 / 360pass
  5. 05
    Provider/tool preflightDirect-provider native tool loops
    8 / 8pass
  6. 06
    Development pilot768/768 trials · full evaluation separately gated
    Passpass
  7. 07
    Sealed evaluation8640/8640 scored trials
    Completepass

03 · Development evidence

The harness can grade known answers.

The aerodynamics development set exercises a separate standard-library Python referent that cannot import production runtime code. Public NASA and UIUC sources, artifact hashes, deterministic replay, and deliberately wrong-answer controls are recorded. This is automated evidence—not external peer review or empirical validation.

The current wing comparison uses a declared 7% model-form envelope against classical finite-wing theory. That is a development tolerance, not a measured uncertainty claim.

Independent reference replay
32/ 32
Forced production-runner cases
32/ 32

32 unique development tasks · 4 aerodynamics operations · 32/32 known-bad controls rejected

04 · Paired effects

What access to skills changed.

Each point is the full-skills minus base change in key-answer accuracy. Horizontal intervals are hierarchical paired-bootstrap 95% confidence intervals.

95% confidence intervalFull-skills minus base estimateNo change
Nemotron 3.5 LightningPreregistered primary comparison
+29.07 pts [+20.65 pts, +37.69 pts]
GPT-5.6 TerraWithin-family replication
+52.59 pts [+41.57 pts, +62.96 pts]
GPT-5.6 SolWithin-family replication
+40.83 pts [+31.76 pts, +49.91 pts]
Risk difference in task-level key-answer correctness. Positive values favor full skills. Intervals are uncertainty estimates, not physical-validation claims.

Nemotron 3.5 Lightning

+29.07 pts95% CI +20.65 pts to +37.69 ptsPreregistered primary comparison

GPT-5.6 Terra

+52.59 pts95% CI +41.57 pts to +62.96 ptsWithin-family replication

GPT-5.6 Sol

+40.83 pts95% CI +31.76 pts to +49.91 ptsWithin-family replication

Interpretation boundary: Only within-model full-versus-base differences estimate skill uplift. Terra and Sol are replications; comparisons between model families are descriptive.

05 · Absolute performance

The answer, by condition.

Bars show the preregistered all-primary-fields task accuracy. Base conditions are graphite; documentation or runner ablations are outlined; complete skills are orange.

BaseDocs or runner ablationFull skills
Nemotron 3.5 Lightning base
22.9%
Nemotron 3.5 Lightning docs-only
22.3%
Nemotron 3.5 Lightning runner-only
58.5%
Nemotron 3.5 Lightning full skills
51.9%
GPT-5.6 Terra base
31.0%
GPT-5.6 Terra full skills
83.6%
GPT-5.6 Sol base
19.6%
GPT-5.6 Sol full skills
60.5%
Accuracy requires correct status and every declared primary quantity within its operation-specific deterministic tolerance.

Nemotron detail

Uplift by skill family.

Paired base and full-skills accuracy reveals where local execution helped and where documentation, calculator scope, or agent reasoning still failed.

06 · Controlled experiment

Eight arms. One primary comparison.

Nemotron’s 2×2 design separates documentation, execution, and interaction effects. Terra and Sol repeat base-versus-full; they are not pooled controls.

Direct-provider armDocumentationLocal runnerCurrent state
Nemotron 3.5 Lightning baseNoNoSealed · 22.9%
Nemotron 3.5 Lightning docs-onlyCompleteNoSealed · 22.3%
Nemotron 3.5 Lightning runner-onlyCallable schemasYesSealed · 58.5%
Nemotron 3.5 Lightning full skillsCompleteYesSealed · 51.9%
GPT-5.6 Terra baseNoNoSealed · 31.0%
GPT-5.6 Terra full skillsCompleteYesSealed · 83.6%
GPT-5.6 Sol baseNoNoSealed · 19.6%
GPT-5.6 Sol full skillsCompleteYesSealed · 60.5%

Preregistered causal estimand: Nemotron full skills minus Nemotron base. Three trials per task; 360 tasks; hierarchical paired bootstrap over operation → task → repeat.

07 · Relative economics

Compare cost at the task level.

The sealed evaluation is compared by observed cost per scored trial and cost per correct task, so efficiency can be interpreted alongside accuracy.

Scroll for cost columns

Direct-provider armCompleted trialsKey-answer accuracyCost / completed trialCost / correct task
Nemotron 3.5 Lightning base1,08022.9%$0.0000$0.0000
Nemotron 3.5 Lightning docs-only1,08022.3%$0.0000$0.0000
Nemotron 3.5 Lightning runner-only1,08058.5%$0.0000$0.0000
Nemotron 3.5 Lightning full skills1,08051.9%$0.0000$0.0000
GPT-5.6 Terra base1,08031.0%$0.0079$0.03
GPT-5.6 Terra full skills1,08083.6%$0.0075$0.0090
GPT-5.6 Sol base1,08019.6%$0.02$0.09
GPT-5.6 Sol full skills1,08060.5%$0.02$0.03

These are sealed-evaluation economics, using provider-recorded cache-aware usage and direct-provider token rates. NVIDIA Build reported no token charge; settled provider dashboards remain authoritative. No private budget or planning forecast is shown. See the observed v1 comparison.

08 · Primary endpoint

Score the answer the question asked for.

A task passes when the status is correct, every declared primary quantity is correct after unit normalization and output-aware tolerance, and no contradictory primary value appears.

  • Normally one to three semantic quantities—not every internal runner field.
  • Scalar, vector, curve, controller, stochastic, retrieval, and classification graders.
  • Safety prose, provenance, operation slug, and envelope formatting reported separately.
{
  "status": "success",
  "answers": [
    {
      "quantity": "total_delta_v",
      "value": 3935.2,
      "unit": "m/s"
    }
  ],
  "assumptions": [],
  "warnings": [],
  "method": "two-impulse Hohmann transfer"
}

09 · Scientific boundaries

What the result does not prove.

The agent experiment can be complete while calculator-validation claims remain deliberately conservative.

VAL

Physical-accuracy headline

No operation is labeled reference- or empirically validated. Automated referent agreement is reported as evidence, not peer review, certification, or operational qualification.

Audit trail

Nothing disappears.

V1 remains public with its original charts and value records, but is clearly labeled an exploratory runner-consistency study.

Machine-readable status. The current gate state and blockers are available as JSON.

Reproducible protocol. Canonical dossiers, development cases, reference traces, scoring code, and sanitized future trajectories live in the repository.