28 paired LowToHi wins · p = 7.45e-09
Measured, versioned, reproducible
Evidence strong enough to inspect — and narrow enough to trust.
LowToHi publishes the task, model, runtime, evaluator and statistical boundary behind each result. The goal is not marketing theatre; it is falsifiable progress.
10 paired LowToHi wins · p = 0.001953125
16 paired LowToHi wins · p < 0.001
How the comparisons are governed
The same provider model is tested in RAW and LowToHi conditions against frozen tasks and evaluators.
Frozen task sets
Inputs, fixtures, hidden tests and evaluator rules are fixed before the campaign.
Paired conditions
RAW and LowToHi runs share the same underlying model and declared generation settings.
Saved evidence
Responses, outcomes, token usage, costs and manifests are retained for offline verification.
Fail-closed claims
Missing or unverifiable measurements do not become positive marketing statements.
What the evidence does and does not establish
- It supports a material runtime effect on named agentic programming tasks.
- It does not establish universal intelligence or superiority across every domain.
- It does not imply that every product surface is production-ready.
- It does not isolate each internal mechanism unless an ablation is run.
- It does not grant operational authority to the model or runtime.
A category should be built in public evidence, not adjectives.
LowToHi treats benchmark methodology, negative results and claim boundaries as part of the product.
Low cost in. High trust out.
Build the intelligence your discipline actually needs.
Explore the product, inspect the evidence and see how LowToHi turns replaceable models into durable professional capability.