Sui Evals

How much does your AI actually know about Sui?

See how well AI models answer 158 real Sui developer questions, with and without the Sui skills.

Get the skills


    

Run it on your own agent

Clone https://github.com/MystenLabs/skills, then read
suievals/RUN.md in that repo and run the full suite on yourself.
Report a score per pillar and list the evals you missed.
Run it twice, bare and with the skills loaded, then submit both to MystenLabs/skills.

How each model scores

Loading…

How this is scored

A question is not scored pass-or-fail, and a model’s score is the mean of four areas rather than one pooled total.

The questions are real

Every prompt is one a developer would type: “How do flash loans work in DeepBook V3? Can I borrow and trade in the same pool?” Answers come back as prose and code, the length of a real reply.

Expectations, written in advance

885 expectations across 156 questions: the import to use, the call to make, the mistake to avoid. Most are prose, so a second model reads the answer to judge them. 26 name something exact enough to check literally.

Expectations are counted, not questions

Five of six expectations right is not the same as none. Scored pass-or-fail, the four weakest models all land between 0% and 7%. Counting expectations spreads them from 8% to 24%.

Four areas, equally weighted

Objects, transactions, building and security, averaged. Building holds 80 of the 156 questions, so a pooled total would let it set the score and bury security’s 16.

Asked more than once

The same model asked the same question twice does not always answer the same way. Most runs ask three times and score the average. Each card records how many attempts it used.

Every question points at a page

Each question cites the documentation a correct answer comes from, so the suite measures whether Sui is documented rather than what a model memorised. 155 of 156 do; the one that does not is a defect in the suite.

The evals