See how well AI models answer 158 real Sui developer questions, with and without the Sui skills.
Get the skills
Run it on your own agent
Clone https://github.com/MystenLabs/skills, then read suievals/RUN.md in that repo and run the full suite on yourself. Report a score per pillar and list the evals you missed.
Loading…
A question is not scored pass-or-fail, and a model’s score is the mean of four areas rather than one pooled total.
Every prompt is one a developer would type: “How do flash loans work in DeepBook V3? Can I borrow and trade in the same pool?” Answers come back as prose and code, the length of a real reply.
885 expectations across 156 questions: the import to use, the call to make, the mistake to avoid. Most are prose, so a second model reads the answer to judge them. 26 name something exact enough to check literally.
Five of six expectations right is not the same as none. Scored pass-or-fail, the four weakest models all land between 0% and 7%. Counting expectations spreads them from 8% to 24%.
Objects, transactions, building and security, averaged. Building holds 80 of the 156 questions, so a pooled total would let it set the score and bury security’s 16.
The same model asked the same question twice does not always answer the same way. Most runs ask three times and score the average. Each card records how many attempts it used.
Each question cites the documentation a correct answer comes from, so the suite measures whether Sui is documented rather than what a model memorised. 155 of 156 do; the one that does not is a defect in the suite.