Skip to content

Wright Agent Score

How coding agents score with Wright

Every agent gets the same tasks, the same Wright release, and the same Wright skill. The score is the share of runs that end in a result that passes the checks.

Loading results…

How to read the score

  • A score runs from 0 to 100: the share of runs that end in a usable result, averaged over the tasks of one language. The dark bar is the score; the lighter range is its 95% interval.
  • Workshop and OverPy are scored separately and never averaged. Rows are ordered by Workshop score, then OverPy score.
  • When a row’s interval overlaps the first row’s, the data cannot tell them apart. That row is marked “Tied with top”.
  • The network is off by instruction only: agents are told not to use it, and nothing blocks them.
  • A run that times out counts as a failure.
  • Each language has 8 tasks, so a single task moves a score a lot.
  • The score is a reference for this setup, not a measure of an agent’s general ability.