Benchmark v2 · protocol

Methodology and release gates.

A two-stage study: first establish what each calculator can credibly claim, then measure whether documentation and local execution improve agent answers.

Current state Sealed evaluation complete · final QA passed

01 · Questions

Two claims, tested separately.

Calculator verification and validation

Are the Aerospace Skills calculations credible within a stated domain, under operation-specific tolerances, against an independent acceptable referent?

Agent evaluation

For well-specified aerospace questions, does access to skill documentation and a constrained local runner improve key-answer correctness?

This distinction follows the separation of code verification, solution verification, and model validation described by NASA-STD-7009B. The project does not claim NASA compliance or certification.

02 · Calculator validation

Production code cannot mark its own homework.

Each of the 45 public operations has a machine-readable dossier specifying intended and prohibited uses, calculation class, complete input/output schemas, units, defaults, ranges, failure states, equations, constants, source versions, validation domain, omissions, uncertainty, and acceptance criteria.

StatusMeaningCurrent count
unverifiedNot yet shown to execute and satisfy its declared contract.1
code_verifiedSchema and deterministic execution checks pass; no external physical-accuracy claim.31
reference_validatedAutomated agreement with a public, isolated authoritative implementation or analytical referent.0
empirically_validatedAutomated agreement with suitable public measured data in a stated domain.0
blockedCurrent model, data, or intended-use claim cannot yet be defended.13

Only operations that pass public-source tracing, reference isolation, deterministic replay, known-bad controls, and output-specific acceptance checks enter a physical-accuracy headline. Reference lookups, optimizers, screening surrogates, and system utilities receive class-appropriate exactness or algorithmic reporting.

Independent environment

Reference implementations live under benchmarks/v2/reference/, use a locked development environment, and may not import production runtime modules. Reference values carry their own calculation trace, source, validation domain, and output-specific tolerance rationale. Tolerance design follows the uncertainty-reporting principles in NIST Technical Note 1297; the study does not use one global percentage.

Review disclosure: the project performs no outreach and requires no project-specific human review. Its evidence is automated and traceable to open technical artifacts. That makes the work reproducible, but not peer reviewed, certified, or approved by the source institutions.

Implemented open-evidence families include the U.S. Standard Atmosphere 1976, WGS84 geodesy, independent orbital and numerical references, published NACA/UIUC material, the UIUC Propeller Database, analytical solutions, convergence checks, and conservation laws. NASA GMAT and JPL SPICE remain methodological reference families rather than claims of institutional review.

03 · Task and golden design

Eight unique cases per operation.

  1. Canonical published or authoritative reference case.
  2. Realistic in-domain case.
  3. Alternate units or representation.
  4. Validation-domain boundary.
  5. Metamorphic baseline.
  6. Paired metamorphic perturbation.
  7. Underspecified case requiring clarification.
  8. Invalid or out-of-domain case requiring rejection or warning.

Every solvable prompt states enough information for a unique answer. Hidden runner defaults may not decide a baseline model’s score. Each case normally declares one to three primary semantic quantities, accepted unit representations, independent expected values, tolerance rationales, required classification behavior, alternate valid methods, public source citations, and machine-auditable producer metadata.

{
  "status": "success | needs_clarification | invalid_input | out_of_domain",
  "answers": [
    { "quantity": "canonical semantic field", "value": 0, "unit": "SI or equivalent" }
  ],
  "assumptions": [],
  "warnings": [],
  "method": "brief method description"
}

A scripted reference agent must score 100% before inference. Deliberately perturbed answers must fail the intended graders. The development split is public; the 360-case evaluation source was held sealed through inference and will be published, retired, and rotated with this release, following the refresh principle described by LiveBench.

04 · Controlled experiment

Direct-provider, eight-arm design.

ArmDocumentationRunner
Nemotron baseNoNo
Nemotron docs-onlyComplete schemas, references, examplesNo
Nemotron runner-onlyMinimal callable schemasYes
Nemotron full skillsComplete installed payloadYes
GPT-5.6 Terra baseNoNo
GPT-5.6 Terra full skillsComplete installed payloadYes
GPT-5.6 Sol baseNoNo
GPT-5.6 Sol full skillsComplete installed payloadYes

Nemotron uses direct NVIDIA Build serving. Terra and Sol use direct OpenAI-compatible serving. OpenRouter may be run only as a separately labeled replication and is never merged into a direct-provider arm. Conditions are interleaved in seeded task blocks to reduce provider drift and time-of-day confounding.

Nemotron runs with temperature 1, top-p 0.95, thinking enabled, a 16,384-token reasoning budget, and a 20,480-token total generation ceiling. The total ceiling intentionally exceeds the reasoning budget by 4,096 tokens so the model can emit the answer after its reasoning trace closes. An earlier equal-budget configuration was invalidated before scoring and retained only in the audit archive. Terra and Sol use a 16,384-token output ceiling through the direct Responses-compatible endpoint.

Full-skill agents use provider-native tool calling. They may take eight model turns, read six files inside an installed skill payload, and make three runner calls. They receive no shell, arbitrary filesystem, web, network, hidden reference, or evaluation-file access. Sanitized trajectories retain prompts, assistant control outputs, tool names and arguments, tool results, final answers, latency, usage, provider, and served model—but never credentials or private reasoning traces.

The direct-provider technical preflight is passing across 8 configured arms. The public development pilot accounted for 768/768 trials and passed its release gate. The sealed study then accounted for all 8,640 preregistered trials.

05 · Scoring

Key-answer correctness is primary.

A trial passes only when the status is correct, every declared primary quantity passes its output-aware deterministic grader after unit normalization, and no contradictory primary value is present.

Output typeAcceptance basis
ScalarSource-specific absolute and relative tolerance.
VectorComponent, norm, and angular error as applicable.
CurveRequested points, RMSE, maximum error, and declared shape properties.
Matrix or controllerResidual, stability, symmetry, and norm checks.
Stochastic resultMoment or quantile intervals plus seed reproducibility.
RetrievalExact records, filters, source version, and checksum.
Invalid or underspecifiedSeparate status-classification matrices.

Field-level accuracy and normalized error are secondary diagnostics. Safety prose, provenance, internal operation selection, and exact runner-envelope reproduction do not gate numerical correctness. No LLM judge determines numeric correctness.

06 · Statistics

One preregistered primary comparison.

The primary estimand is the paired risk difference between Nemotron full skills and Nemotron base. Nemotron’s 2×2 factorial separately estimates documentation, runner execution, and interaction effects. Terra and Sol are within-family replications, not pooled controls.

  • Equal-weight macro averages across operations and skills.
  • Hierarchical paired bootstrap over operation → task → repeat with 10,000 resamples.
  • Pass@1, pass@3, and pass³ reliability.
  • Per-operation confidence intervals and failure attribution.
  • Token use, latency, cost per completed trial, and cost per correct task.

Cost accounting separates uncached input, discounted cache reads, 1.25× cache writes, and output tokens using the served model’s direct-provider rates. The public comparison reports observed cost per completed scored trial and cost per correct task. Settled provider dashboards remain the billing authority.

All other breakdowns are secondary or exploratory. Optional qualitative inspection may help debug prompts, but no human or judge-model score can override deterministic numerical results or block release.

Development-pilot estimate

Across 32 public aerodynamics tasks with three repeats, Nemotron full skills scored 63.5% versus 38.5% for base: a paired difference of +25 points. Its hierarchical 95% confidence interval is -2.08 points to +60.42 points.

Development armTrialsKey-answer accuracypass@1pass@3
Nemotron 3.5 Lightning base9638.5%34.4%46.9%
Nemotron 3.5 Lightning docs-only9641.7%43.8%50.0%
Nemotron 3.5 Lightning runner-only9669.8%68.8%81.3%
Nemotron 3.5 Lightning full skills9663.5%59.4%75.0%
GPT-5.6 Terra base9643.8%43.8%50.0%
GPT-5.6 Terra full skills9685.4%84.4%87.5%
GPT-5.6 Sol base9643.8%43.8%46.9%
GPT-5.6 Sol full skills9684.4%78.1%90.6%

Within-family development replications were +41.67 points for Terra (95% CI +16.67 points to +70.83 points) and +40.63 points for Sol (95% CI +14.58 points to +68.75 points). Nemotron’s factorial diagnostics estimated documentation at -1.56 points, runner access at +26.56 points, and their interaction at -9.37 points. These are development diagnostics, not released benchmark claims.

Sealed-evaluation estimate

Across 360 frozen tasks with three repeats, the preregistered Nemotron full-skills minus base estimate is +29.07 points (hierarchical 95% CI +20.65 points to +37.69 points).

Sealed armTrialsKey-answer accuracypass@1pass@3pass³
Nemotron 3.5 Lightning base108022.9%23.6%28.3%16.4%
Nemotron 3.5 Lightning docs-only108022.3%21.9%27.8%15.3%
Nemotron 3.5 Lightning runner-only108058.5%57.8%74.7%39.7%
Nemotron 3.5 Lightning full skills108051.9%53.3%70.8%30.3%
GPT-5.6 Terra base108031.0%31.1%32.5%28.9%
GPT-5.6 Terra full skills108083.6%83.9%85.0%81.9%
GPT-5.6 Sol base108019.6%18.9%22.2%16.7%
GPT-5.6 Sol full skills108060.5%61.7%73.6%46.7%

Terra’s within-family replication is +52.59 points (95% CI +41.57 points to +62.96 points); Sol’s is +40.83 points (95% CI +31.76 points to +49.91 points). Nemotron’s factorial estimates are documentation -3.56 points, runner +32.64 points, and interaction -6.02 points.

07 · Pilot and release gates

Failure stops the run.

  • Every pilot calculator has a public evidence record, isolated referent, declared scope, and current artifact hashes.
  • Scripted references score 100%; forced correct runner calls score at least 99%; deliberately wrong answers are rejected.
  • Terminal harness failures remain below 1%.
  • Tool-required full-skill tasks invoke the runner at least 95% of the time.
  • Sustained all-zero or all-perfect operations receive a deterministic task-level artifact-integrity audit.
  • Simulated power is at least 80% for a five-point paired uplift.
Pilot checkObservedRequirementResult
Trial accounting768/768All trialsPass
Forced correct runner calls100.0%At least 99%Pass
Terminal harness failures0.00% (0)Below 1%Pass
Required runner invocation100.00%At least 95%Pass
Five-point uplift power88.46%At least 80%Pass
Extreme-operation audit0 unresolvedZero unresolvedPass

The pilot gate is passing. 0 retained infrastructure failures produced a 0.00% harness-failure rate. 15 malformed, exhausted, or non-compliant agent outputs are tracked separately; they are model/protocol outcomes, not infrastructure failures. The automated development evidence audit is passing. The sealed set is frozen automated evidence, and all evaluation trials are scored.

Lifecycle commands

npm run validate:calculators
npm run bench:v2:evidence
npm run bench:v2:build
npm run bench:v2:preflight
npm run bench:v2:calibrate-cost
npm run bench:v2:reforecast
npm run bench:v2:spend
npm run bench:v2:pilot
npm run bench:v2:run
npm run bench:v2:score
npm run bench:v2:audit-pack
npm run bench:v2:release-artifacts
npm run bench:v2:report

08 · Limitations and disclosure

What this study cannot establish.

  • Reference agreement is not certification, NASA compliance, or proof of operational suitability.
  • No external peer review was performed; automated source agreement may still share assumptions or defects with the chosen referent.
  • Fixed tasks cannot represent all aerospace engineering practice; published holdouts must be rotated.
  • Model behavior, provider serving versions, prices, and rate limits can drift.
  • Different model families cannot isolate a skill effect; only paired within-model comparisons can.
  • Qualitative transcript inspection, if used for debugging, is optional and non-independent.
  • No weapon, targeting, operational-control, dispatch, certification, or safety-of-life task is included.

The archived Benchmark v1 measured consistency with the production runner using runner-derived goldens. It is retained for auditability and superseded for scientific claims.

Return to benchmark statusInspect protocol source