Nemotron 3.5 Lightning
+29.07 pts95% CI +20.65 pts to +37.69 ptsPreregistered primary comparisonBenchmark v2 · sealed evaluation complete
Prove the calculators.
Then test the agents.
Benchmark v2 separates calculator evidence from agent performance, then measures skill uplift across 360 aerospace tasks and eight direct-provider conditions.
- Canonical dossiers
- 45/ 45
- Development referents
- 45/ 45
- Headline validated
- 0/ 45
- Technical tool loops
- 8/ 8
- Sealed trials
- 8,640/ 8,640
Sealed result · QA passed
01 · Calculator credibility
Forty-five operations. Zero shortcuts.
Every square is one public operation. “Code verified” means its schema and standalone execution pass—it does not mean the aerospace model is externally validated.
Headline rule: only operations that pass the public-source, isolated-reference, replay, negative-control, and output-specific acceptance gates count as reference_validated or empirically_validated.
02 · Release gates
Evidence before inference.
A passed API check cannot waive missing evidence. Each gate is derived from a versioned artifact. Technical snapshot: Aug 18, 2026 UTC.
- 01Operation dossiersSchemas, intended use, equations, domains, limitationspass
- 02Open referent coveragePublic source traces plus isolated reference artifactspass
- 03Development evidence32/32 replay and 32/32 known-bad rejectionpass
- 04Sealed holdoutEight unique tasks per operation, hash-frozenpass
- 05Provider/tool preflightDirect-provider native tool loopspass
- 06Development pilot768/768 trials · full evaluation separately gatedpass
- 07Sealed evaluation8640/8640 scored trialspass
03 · Development evidence
The harness can grade known answers.
The aerodynamics development set exercises a separate standard-library Python referent that cannot import production runtime code. Public NASA and UIUC sources, artifact hashes, deterministic replay, and deliberately wrong-answer controls are recorded. This is automated evidence—not external peer review or empirical validation.
The current wing comparison uses a declared 7% model-form envelope against classical finite-wing theory. That is a development tolerance, not a measured uncertainty claim.
- Independent reference replay
- 32/ 32
- Forced production-runner cases
- 32/ 32
32 unique development tasks · 4 aerodynamics operations · 32/32 known-bad controls rejected
04 · Paired effects
What access to skills changed.
Each point is the full-skills minus base change in key-answer accuracy. Horizontal intervals are hierarchical paired-bootstrap 95% confidence intervals.
GPT-5.6 Terra
+52.59 pts95% CI +41.57 pts to +62.96 ptsWithin-family replicationGPT-5.6 Sol
+40.83 pts95% CI +31.76 pts to +49.91 ptsWithin-family replicationInterpretation boundary: Only within-model full-versus-base differences estimate skill uplift. Terra and Sol are replications; comparisons between model families are descriptive.
05 · Absolute performance
The answer, by condition.
Bars show the preregistered all-primary-fields task accuracy. Base conditions are graphite; documentation or runner ablations are outlined; complete skills are orange.
Nemotron detail
Uplift by skill family.
Paired base and full-skills accuracy reveals where local execution helped and where documentation, calculator scope, or agent reasoning still failed.
06 · Controlled experiment
Eight arms. One primary comparison.
Nemotron’s 2×2 design separates documentation, execution, and interaction effects. Terra and Sol repeat base-versus-full; they are not pooled controls.
| Direct-provider arm | Documentation | Local runner | Current state |
|---|---|---|---|
| Nemotron 3.5 Lightning base | No | No | Sealed · 22.9% |
| Nemotron 3.5 Lightning docs-only | Complete | No | Sealed · 22.3% |
| Nemotron 3.5 Lightning runner-only | Callable schemas | Yes | Sealed · 58.5% |
| Nemotron 3.5 Lightning full skills | Complete | Yes | Sealed · 51.9% |
| GPT-5.6 Terra base | No | No | Sealed · 31.0% |
| GPT-5.6 Terra full skills | Complete | Yes | Sealed · 83.6% |
| GPT-5.6 Sol base | No | No | Sealed · 19.6% |
| GPT-5.6 Sol full skills | Complete | Yes | Sealed · 60.5% |
Preregistered causal estimand: Nemotron full skills minus Nemotron base. Three trials per task; 360 tasks; hierarchical paired bootstrap over operation → task → repeat.
07 · Relative economics
Compare cost at the task level.
The sealed evaluation is compared by observed cost per scored trial and cost per correct task, so efficiency can be interpreted alongside accuracy.
Scroll for cost columns
| Direct-provider arm | Completed trials | Key-answer accuracy | Cost / completed trial | Cost / correct task |
|---|---|---|---|---|
| Nemotron 3.5 Lightning base | 1,080 | 22.9% | $0.0000 | $0.0000 |
| Nemotron 3.5 Lightning docs-only | 1,080 | 22.3% | $0.0000 | $0.0000 |
| Nemotron 3.5 Lightning runner-only | 1,080 | 58.5% | $0.0000 | $0.0000 |
| Nemotron 3.5 Lightning full skills | 1,080 | 51.9% | $0.0000 | $0.0000 |
| GPT-5.6 Terra base | 1,080 | 31.0% | $0.0079 | $0.03 |
| GPT-5.6 Terra full skills | 1,080 | 83.6% | $0.0075 | $0.0090 |
| GPT-5.6 Sol base | 1,080 | 19.6% | $0.02 | $0.09 |
| GPT-5.6 Sol full skills | 1,080 | 60.5% | $0.02 | $0.03 |
These are sealed-evaluation economics, using provider-recorded cache-aware usage and direct-provider token rates. NVIDIA Build reported no token charge; settled provider dashboards remain authoritative. No private budget or planning forecast is shown. See the observed v1 comparison.
08 · Primary endpoint
Score the answer the question asked for.
A task passes when the status is correct, every declared primary quantity is correct after unit normalization and output-aware tolerance, and no contradictory primary value appears.
- Normally one to three semantic quantities—not every internal runner field.
- Scalar, vector, curve, controller, stochastic, retrieval, and classification graders.
- Safety prose, provenance, operation slug, and envelope formatting reported separately.
{
"status": "success",
"answers": [
{
"quantity": "total_delta_v",
"value": 3935.2,
"unit": "m/s"
}
],
"assumptions": [],
"warnings": [],
"method": "two-impulse Hohmann transfer"
}09 · Scientific boundaries
What the result does not prove.
The agent experiment can be complete while calculator-validation claims remain deliberately conservative.
Physical-accuracy headline
No operation is labeled reference- or empirically validated. Automated referent agreement is reported as evidence, not peer review, certification, or operational qualification.
Audit trail
Nothing disappears.
V1 remains public with its original charts and value records, but is clearly labeled an exploratory runner-consistency study.
Machine-readable status. The current gate state and blockers are available as JSON.
Reproducible protocol. Canonical dossiers, development cases, reference traces, scoring code, and sanitized future trajectories live in the repository.