Sophetic
Distance to frontier 51.5
Zeno 1.0 · research preview
San Francisco34 peopleFounded 19 months ago

We are a long way from the frontier.

Zeno 1.0 is our first model. It scores 1.5 on the Artificial Analysis Intelligence Index. The leading model scores 53.0. We publish every result we measure, including the ones below.

Zeno 1.0 The Dichotomy Method
1.5 51.5 remaining
053.0 frontier

Each generation of Zeno closes half of what is left. This is the founding thesis of the company.

Everything we measured, including the zeros.

Eight benchmarks, all eight reported, which we intend to keep doing for every model we release. The grey rule is the leading model. The solid block is Zeno 1.0.

Long-horizon software engineeringDeepSWE v1.1
1.4of 73.8
Agentic codingCursorBench 3.2.0
1.9of 73.4
Agentic codingTerminal-Bench 4.0
0.0of 58.2
Agentic computer useOSWorld 2.0
0.0of 72.6
Knowledge workGDPval-AA v2, Elo
412of 1766
Multidisciplinary reasoningHLE-Verified
0.9of 59.1
Document-grounded workGDP.pdf · All-pass
of 33.2
CompositeAA Intelligence Index v4.3
1.5of 53.0
Expert-level reasoningGPQA Diamond · outside index v4.3
27.1of 95.3

Zeno 1.0 completed no tasks on Terminal-Bench 4.0 or OSWorld 2.0. GDP.pdf entered the index in September and Zeno has not yet been run against it; we will report the figure when we have one. The composite figure is Intelligence Index v4.3, released in September, which upgraded Terminal-Bench and added an agentic workflow benchmark with a private test set; Zeno 1.0 was regraded from 1.9 to 1.5 and the leading model from 57.0 to 53.0. The distance closed by 3.6 points. We did nothing to close it and are not claiming it. GPQA Diamond was removed from the index in September and we agree with the removal; we continue to publish our result on it, outside the composite, because we publish every result we measure. On GPQA Diamond Zeno 1.0 scored 27.1 against a random-choice baseline of 25. Right-hand figures are the best publicly reported result for each benchmark this month and are not attributed to any single system.

Zeno of Elea would say we never arrive. He underestimated how much of the distance is in the first half.
Our founding thesis, and the reason the model carries his name.
Founded
19 months agoBy two researchers who have asked that their previous employers not be named. The employers have not asked this.
Team
34 peopleApplied research, infrastructure, and one person on policy.
Funding
Self-fundedWe have not raised outside capital. Zeno 1.0 was trained on our own money and the figures above are what it bought.
Compute
The StoaOur cluster, in an undisclosed location. It draws enough power to supply a mid-sized city.

The Document

Forty pages, published in full, setting out the values we want Zeno to hold. Principle 1 instructs it to be helpful. Principle 2 instructs it to be honest. There is no Principle 3.

On sycophancy

Models trained on human feedback learn to tell people what they want to hear. Zeno is designed to push back where it disagrees. In testing, Zeno has not yet disagreed.

What we measured before release

Zeno was evaluated against the Assay, our full internal suite, prior to deployment. It passed. We wrote the Assay. We are reviewing the Assay.

Responsible scaling

Zeno 1.0 is deployed at TSL-1, a level at which models present no meaningful catastrophic risk. We made that determination ourselves and are confident in it. Our thresholds are binding and will not be revised.

Zeno 1.0, in detail.

Available today in two configurations. Instant is the faster of the two. It is not a smaller model in any sense we are prepared to describe.

Zeno 1.0
$1.20 / $6.00per million tokens, in and out
Zeno Instant 1.0
$0.15 / $0.60290 tokens per second
Context
64kKnowledge cutoff eight months ago
Usage
UnitsUnits do not convert to tokens, to messages, or to each other
Kiln
The trainerThe furnace in which Zeno is fired. We have no plans to describe it
The Loom
The pipelineTraining data reaches Kiln through the Loom
The Assay
The evaluationsIt has never failed a model
Status
OperationalNo incidents recorded. Accepting traffic since launch