Skip to content

Measured, versioned, reproducible

Evidence strong enough to inspect — and narrow enough to trust.

LowToHi publishes the task, model, runtime, evaluator and statistical boundary behind each result. The goal is not marketing theatre; it is falsifiable progress.

Agentic holdout v1RAW and LowToHi · same model · frozen holdout
RAW1 / 30
LOWTOHI29 / 30

28 paired LowToHi wins · p = 7.45e-09

Agentic holdout v2Independent 20-task holdout · no retries
RAW0 / 20
LOWTOHI10 / 20

10 paired LowToHi wins · p = 0.001953125

N20 agentic programmingFrozen campaign · temp 0.2 · max 4096
RAW3 / 20
LOWTOHI19 / 20

16 paired LowToHi wins · p < 0.001

How the comparisons are governed

The same provider model is tested in RAW and LowToHi conditions against frozen tasks and evaluators.

01

Frozen task sets

Inputs, fixtures, hidden tests and evaluator rules are fixed before the campaign.

02

Paired conditions

RAW and LowToHi runs share the same underlying model and declared generation settings.

03

Saved evidence

Responses, outcomes, token usage, costs and manifests are retained for offline verification.

04

Fail-closed claims

Missing or unverifiable measurements do not become positive marketing statements.

What the evidence does and does not establish

  • It supports a material runtime effect on named agentic programming tasks.
  • It does not establish universal intelligence or superiority across every domain.
  • It does not imply that every product surface is production-ready.
  • It does not isolate each internal mechanism unless an ablation is run.
  • It does not grant operational authority to the model or runtime.
CLAIM BOUNDARY

A category should be built in public evidence, not adjectives.

LowToHi treats benchmark methodology, negative results and claim boundaries as part of the product.

Low cost in. High trust out.

Build the intelligence your discipline actually needs.

Explore the product, inspect the evidence and see how LowToHi turns replaceable models into durable professional capability.