The Document
Forty pages, published in full, setting out the values we want Zeno to hold. Principle 1 instructs it to be helpful. Principle 2 instructs it to be honest. There is no Principle 3.
Zeno 1.0 is our first model. It scores 1.5 on the Artificial Analysis Intelligence Index. The leading model scores 53.0. We publish every result we measure, including the ones below.
Each generation of Zeno closes half of what is left. This is the founding thesis of the company.
Eight benchmarks, all eight reported, which we intend to keep doing for every model we release. The grey rule is the leading model. The solid block is Zeno 1.0.
Zeno 1.0 completed no tasks on Terminal-Bench 4.0 or OSWorld 2.0. GDP.pdf entered the index in September and Zeno has not yet been run against it; we will report the figure when we have one. The composite figure is Intelligence Index v4.3, released in September, which upgraded Terminal-Bench and added an agentic workflow benchmark with a private test set; Zeno 1.0 was regraded from 1.9 to 1.5 and the leading model from 57.0 to 53.0. The distance closed by 3.6 points. We did nothing to close it and are not claiming it. GPQA Diamond was removed from the index in September and we agree with the removal; we continue to publish our result on it, outside the composite, because we publish every result we measure. On GPQA Diamond Zeno 1.0 scored 27.1 against a random-choice baseline of 25. Right-hand figures are the best publicly reported result for each benchmark this month and are not attributed to any single system.
Zeno of Elea would say we never arrive. He underestimated how much of the distance is in the first half.Our founding thesis, and the reason the model carries his name.
Forty pages, published in full, setting out the values we want Zeno to hold. Principle 1 instructs it to be helpful. Principle 2 instructs it to be honest. There is no Principle 3.
Models trained on human feedback learn to tell people what they want to hear. Zeno is designed to push back where it disagrees. In testing, Zeno has not yet disagreed.
Zeno was evaluated against the Assay, our full internal suite, prior to deployment. It passed. We wrote the Assay. We are reviewing the Assay.
Zeno 1.0 is deployed at TSL-1, a level at which models present no meaningful catastrophic risk. We made that determination ourselves and are confident in it. Our thresholds are binding and will not be revised.
Available today in two configurations. Instant is the faster of the two. It is not a smaller model in any sense we are prepared to describe.