Know when to proceed.
Check whether the required research conditions are in place. Proceed with a supported analysis, ask for missing facts, or stop when the request cannot proceed.
nomue gives AI-driven statistical analysis a clear path to proceed, verifiable results, and next actions grounded in evidence.
Life sciences. Finance. Chemistry. Wherever claims must hold.
Why
Before a calculation, an agent needs to know whether the required research conditions are met. Afterward, it needs to know what the result establishes and what to do next.
nomue turns declared conditions and evidence into machine-readable decisions.
What nomue does
Check whether the required research conditions are in place. Proceed with a supported analysis, ask for missing facts, or stop when the request cannot proceed.
Check supported calculations and claims against the supplied data and evidence. See what passed, what failed, and what each check establishes.
Give the agent a structured outcome, reasons, and next actions, so it can continue the workflow or bring a question back to the researcher.
Evidence
216 / 216
Compared with 132 / 216 using general Python and 131 / 216 using a stable Welch wrapper.
72 constructed Welch-report questions · three repetitions · 648 sessions
78.1%lower API cost
62.9%less task time
Measured across 103 matched pairs where both configurations reached the correct answer.
48 synthetic Welch tasks · three repetitions · 288 scheduled sessions
These results apply to the tested Welch workflows. Both studies were led by nomue's developer and published as preprints, not peer-reviewed findings.
Scope
Available today
The public nomue package checks supported two-sided Welch analyses locally on Linux, macOS and Windows.
Hosted Welch checks are available to approved recipients in limited Release 1.
Long-term direction
The long-term scope is any field where AI-generated claims can be checked against declared data, calculations or rules.
Each new capability will state what it checks, the evidence behind it and where it stops.
Open work
We test the software scientific claims depend on, report the failures we can reproduce, and publish the research behind nomue.
14 reports submitted 5 fixes merged upstream
Each report records the finding and the upstream decision. How to read the outcomes
SciPy
Closed · NumPy fix not planned
Returned two-sided p → independent reference
0.0≈ 3.505e-316
jStat
Reported · confirmation pending
Returned CDF → independent reference
0.0≈ 0.4410
statsmodels
Fix merged upstream
Before patch → after patch
0.0≈ 7.076e-18
statsmodels
Fix merged upstream
Before patch → after patch
NaN≈ 0.26657
SciPy
Repair proposed · checks passing
Returned degrees of freedom → exact reference
14
SciPy
Reported · confirmation pending
Reported SF → exact special-case reference
0.0≈ 2.0e-8
SciPy
Repair proposed · checks passing
Original scale → scaled by 2^-511
0.026500.05611
R / agricolae
Emailed · confirmation pending
Same observations: original labels → renamed labels
0.03630.0791
SciPy
Under discussion · remedy undecided
Same pair: alone → batched with a tied pair
0.047980.05132
SciPy
Fix proposed · high-precision results posted
Returned after rescaling → scale-invariant reference
0.0 / 1.00.2048 / 0.03510
Julia / HypothesisTests.jl
Matching fix released
Registered 0.11.8 → merge commit
1.251.0
SciPy
Fix merged upstream
Before patch → after patch
0.08.12511917099255e-17
Boost.Math / SciPy
Fix merged upstream
Returned → expected
+∞finite negative quantile
R
Reported · unconfirmed
R result → exact reference
-7.55e-152.59e-18
Research papers
Adding nomue reduced API cost by 78.1% and task time by 62.9% on historical-analysis reuse decisions the ordinary-tool comparator also answered correctly.
Preprint v1 — not peer reviewed; synthetic Welch tasks; aggregate reproduction supplement available
A 648-session fixed-panel study compares general Python, a stable Welch wrapper, and a bound verifier for complete five-field report correctness.
Preprint v1.0 — not peer reviewed; fixed constructed panel; reproducibility package available
An audit of R, SciPy, and Julia shows why matching statistical decisions can still hide differences in definitions, numerical precision, and valid probability values.
Preprint v1.0 — not peer reviewed; reproducibility data and code available
A preprint on checking the numerical accuracy of paired Student-t test results before software returns them.
Preprint v0.2 — not peer reviewed
Follow the evidence
New results are published with their tested scope, evidence and limits.
Company
nomue is developed by Licklider, Inc. We publish what it checks, the evidence behind it and where it stops.