01 · Questions
Two claims, tested separately.
Calculator verification and validation
Are the Aerospace Skills calculations credible within a stated domain, under operation-specific tolerances, against an independent acceptable referent?
Agent evaluation
For well-specified aerospace questions, does access to skill documentation and a constrained local runner improve key-answer correctness?
This distinction follows the separation of code verification, solution verification, and model validation described by NASA-STD-7009B. The project does not claim NASA compliance or certification.
02 · Calculator validation
Production code cannot mark its own homework.
Each of the 45 public operations has a machine-readable dossier specifying intended and prohibited uses, calculation class, complete input/output schemas, units, defaults, ranges, failure states, equations, constants, source versions, validation domain, omissions, uncertainty, and acceptance criteria.
| Status | Meaning | Current count |
|---|---|---|
unverified | Not yet shown to execute and satisfy its declared contract. | 1 |
code_verified | Schema and deterministic execution checks pass; no external physical-accuracy claim. | 31 |
reference_validated | Automated agreement with a public, isolated authoritative implementation or analytical referent. | 0 |
empirically_validated | Automated agreement with suitable public measured data in a stated domain. | 0 |
blocked | Current model, data, or intended-use claim cannot yet be defended. | 13 |
Only operations that pass public-source tracing, reference isolation, deterministic replay, known-bad controls, and output-specific acceptance checks enter a physical-accuracy headline. Reference lookups, optimizers, screening surrogates, and system utilities receive class-appropriate exactness or algorithmic reporting.
Independent environment
Reference implementations live under benchmarks/v2/reference/, use a locked development environment, and may not import production runtime modules. Reference values carry their own calculation trace, source, validation domain, and output-specific tolerance rationale. Tolerance design follows the uncertainty-reporting principles in NIST Technical Note 1297; the study does not use one global percentage.
Review disclosure: the project performs no outreach and requires no project-specific human review. Its evidence is automated and traceable to open technical artifacts. That makes the work reproducible, but not peer reviewed, certified, or approved by the source institutions.
Implemented open-evidence families include the U.S. Standard Atmosphere 1976, WGS84 geodesy, independent orbital and numerical references, published NACA/UIUC material, the UIUC Propeller Database, analytical solutions, convergence checks, and conservation laws. NASA GMAT and JPL SPICE remain methodological reference families rather than claims of institutional review.
03 · Task and golden design
Eight unique cases per operation.
- Canonical published or authoritative reference case.
- Realistic in-domain case.
- Alternate units or representation.
- Validation-domain boundary.
- Metamorphic baseline.
- Paired metamorphic perturbation.
- Underspecified case requiring clarification.
- Invalid or out-of-domain case requiring rejection or warning.
Every solvable prompt states enough information for a unique answer. Hidden runner defaults may not decide a baseline model’s score. Each case normally declares one to three primary semantic quantities, accepted unit representations, independent expected values, tolerance rationales, required classification behavior, alternate valid methods, public source citations, and machine-auditable producer metadata.
{
"status": "success | needs_clarification | invalid_input | out_of_domain",
"answers": [
{ "quantity": "canonical semantic field", "value": 0, "unit": "SI or equivalent" }
],
"assumptions": [],
"warnings": [],
"method": "brief method description"
}A scripted reference agent must score 100% before inference. Deliberately perturbed answers must fail the intended graders. The development split is public; the 360-case evaluation source was held sealed through inference and will be published, retired, and rotated with this release, following the refresh principle described by LiveBench.
04 · Controlled experiment
Direct-provider, eight-arm design.
| Arm | Documentation | Runner |
|---|---|---|
| Nemotron base | No | No |
| Nemotron docs-only | Complete schemas, references, examples | No |
| Nemotron runner-only | Minimal callable schemas | Yes |
| Nemotron full skills | Complete installed payload | Yes |
| GPT-5.6 Terra base | No | No |
| GPT-5.6 Terra full skills | Complete installed payload | Yes |
| GPT-5.6 Sol base | No | No |
| GPT-5.6 Sol full skills | Complete installed payload | Yes |
Nemotron uses direct NVIDIA Build serving. Terra and Sol use direct OpenAI-compatible serving. OpenRouter may be run only as a separately labeled replication and is never merged into a direct-provider arm. Conditions are interleaved in seeded task blocks to reduce provider drift and time-of-day confounding.
Nemotron runs with temperature 1, top-p 0.95, thinking enabled, a 16,384-token reasoning budget, and a 20,480-token total generation ceiling. The total ceiling intentionally exceeds the reasoning budget by 4,096 tokens so the model can emit the answer after its reasoning trace closes. An earlier equal-budget configuration was invalidated before scoring and retained only in the audit archive. Terra and Sol use a 16,384-token output ceiling through the direct Responses-compatible endpoint.
Full-skill agents use provider-native tool calling. They may take eight model turns, read six files inside an installed skill payload, and make three runner calls. They receive no shell, arbitrary filesystem, web, network, hidden reference, or evaluation-file access. Sanitized trajectories retain prompts, assistant control outputs, tool names and arguments, tool results, final answers, latency, usage, provider, and served model—but never credentials or private reasoning traces.
The direct-provider technical preflight is passing across 8 configured arms. The public development pilot accounted for 768/768 trials and passed its release gate. The sealed study then accounted for all 8,640 preregistered trials.
05 · Scoring
Key-answer correctness is primary.
A trial passes only when the status is correct, every declared primary quantity passes its output-aware deterministic grader after unit normalization, and no contradictory primary value is present.
| Output type | Acceptance basis |
|---|---|
| Scalar | Source-specific absolute and relative tolerance. |
| Vector | Component, norm, and angular error as applicable. |
| Curve | Requested points, RMSE, maximum error, and declared shape properties. |
| Matrix or controller | Residual, stability, symmetry, and norm checks. |
| Stochastic result | Moment or quantile intervals plus seed reproducibility. |
| Retrieval | Exact records, filters, source version, and checksum. |
| Invalid or underspecified | Separate status-classification matrices. |
Field-level accuracy and normalized error are secondary diagnostics. Safety prose, provenance, internal operation selection, and exact runner-envelope reproduction do not gate numerical correctness. No LLM judge determines numeric correctness.
06 · Statistics
One preregistered primary comparison.
The primary estimand is the paired risk difference between Nemotron full skills and Nemotron base. Nemotron’s 2×2 factorial separately estimates documentation, runner execution, and interaction effects. Terra and Sol are within-family replications, not pooled controls.
- Equal-weight macro averages across operations and skills.
- Hierarchical paired bootstrap over operation → task → repeat with 10,000 resamples.
- Pass@1, pass@3, and pass³ reliability.
- Per-operation confidence intervals and failure attribution.
- Token use, latency, cost per completed trial, and cost per correct task.
Cost accounting separates uncached input, discounted cache reads, 1.25× cache writes, and output tokens using the served model’s direct-provider rates. The public comparison reports observed cost per completed scored trial and cost per correct task. Settled provider dashboards remain the billing authority.
All other breakdowns are secondary or exploratory. Optional qualitative inspection may help debug prompts, but no human or judge-model score can override deterministic numerical results or block release.
Development-pilot estimate
Across 32 public aerodynamics tasks with three repeats, Nemotron full skills scored 63.5% versus 38.5% for base: a paired difference of +25 points. Its hierarchical 95% confidence interval is -2.08 points to +60.42 points.
| Development arm | Trials | Key-answer accuracy | pass@1 | pass@3 |
|---|---|---|---|---|
| Nemotron 3.5 Lightning base | 96 | 38.5% | 34.4% | 46.9% |
| Nemotron 3.5 Lightning docs-only | 96 | 41.7% | 43.8% | 50.0% |
| Nemotron 3.5 Lightning runner-only | 96 | 69.8% | 68.8% | 81.3% |
| Nemotron 3.5 Lightning full skills | 96 | 63.5% | 59.4% | 75.0% |
| GPT-5.6 Terra base | 96 | 43.8% | 43.8% | 50.0% |
| GPT-5.6 Terra full skills | 96 | 85.4% | 84.4% | 87.5% |
| GPT-5.6 Sol base | 96 | 43.8% | 43.8% | 46.9% |
| GPT-5.6 Sol full skills | 96 | 84.4% | 78.1% | 90.6% |
Within-family development replications were +41.67 points for Terra (95% CI +16.67 points to +70.83 points) and +40.63 points for Sol (95% CI +14.58 points to +68.75 points). Nemotron’s factorial diagnostics estimated documentation at -1.56 points, runner access at +26.56 points, and their interaction at -9.37 points. These are development diagnostics, not released benchmark claims.
Sealed-evaluation estimate
Across 360 frozen tasks with three repeats, the preregistered Nemotron full-skills minus base estimate is +29.07 points (hierarchical 95% CI +20.65 points to +37.69 points).
| Sealed arm | Trials | Key-answer accuracy | pass@1 | pass@3 | pass³ |
|---|---|---|---|---|---|
| Nemotron 3.5 Lightning base | 1080 | 22.9% | 23.6% | 28.3% | 16.4% |
| Nemotron 3.5 Lightning docs-only | 1080 | 22.3% | 21.9% | 27.8% | 15.3% |
| Nemotron 3.5 Lightning runner-only | 1080 | 58.5% | 57.8% | 74.7% | 39.7% |
| Nemotron 3.5 Lightning full skills | 1080 | 51.9% | 53.3% | 70.8% | 30.3% |
| GPT-5.6 Terra base | 1080 | 31.0% | 31.1% | 32.5% | 28.9% |
| GPT-5.6 Terra full skills | 1080 | 83.6% | 83.9% | 85.0% | 81.9% |
| GPT-5.6 Sol base | 1080 | 19.6% | 18.9% | 22.2% | 16.7% |
| GPT-5.6 Sol full skills | 1080 | 60.5% | 61.7% | 73.6% | 46.7% |
Terra’s within-family replication is +52.59 points (95% CI +41.57 points to +62.96 points); Sol’s is +40.83 points (95% CI +31.76 points to +49.91 points). Nemotron’s factorial estimates are documentation -3.56 points, runner +32.64 points, and interaction -6.02 points.
07 · Pilot and release gates
Failure stops the run.
- Every pilot calculator has a public evidence record, isolated referent, declared scope, and current artifact hashes.
- Scripted references score 100%; forced correct runner calls score at least 99%; deliberately wrong answers are rejected.
- Terminal harness failures remain below 1%.
- Tool-required full-skill tasks invoke the runner at least 95% of the time.
- Sustained all-zero or all-perfect operations receive a deterministic task-level artifact-integrity audit.
- Simulated power is at least 80% for a five-point paired uplift.
| Pilot check | Observed | Requirement | Result |
|---|---|---|---|
| Trial accounting | 768/768 | All trials | Pass |
| Forced correct runner calls | 100.0% | At least 99% | Pass |
| Terminal harness failures | 0.00% (0) | Below 1% | Pass |
| Required runner invocation | 100.00% | At least 95% | Pass |
| Five-point uplift power | 88.46% | At least 80% | Pass |
| Extreme-operation audit | 0 unresolved | Zero unresolved | Pass |
The pilot gate is passing. 0 retained infrastructure failures produced a 0.00% harness-failure rate. 15 malformed, exhausted, or non-compliant agent outputs are tracked separately; they are model/protocol outcomes, not infrastructure failures. The automated development evidence audit is passing. The sealed set is frozen automated evidence, and all evaluation trials are scored.
Lifecycle commands
npm run validate:calculators
npm run bench:v2:evidence
npm run bench:v2:build
npm run bench:v2:preflight
npm run bench:v2:calibrate-cost
npm run bench:v2:reforecast
npm run bench:v2:spend
npm run bench:v2:pilot
npm run bench:v2:run
npm run bench:v2:score
npm run bench:v2:audit-pack
npm run bench:v2:release-artifacts
npm run bench:v2:report08 · Limitations and disclosure
What this study cannot establish.
- Reference agreement is not certification, NASA compliance, or proof of operational suitability.
- No external peer review was performed; automated source agreement may still share assumptions or defects with the chosen referent.
- Fixed tasks cannot represent all aerospace engineering practice; published holdouts must be rotated.
- Model behavior, provider serving versions, prices, and rate limits can drift.
- Different model families cannot isolate a skill effect; only paired within-model comparisons can.
- Qualitative transcript inspection, if used for debugging, is optional and non-independent.
- No weapon, targeting, operational-control, dispatch, certification, or safety-of-life task is included.
The archived Benchmark v1 measured consistency with the production runner using runner-derived goldens. It is retained for auditability and superseded for scientific claims.