SWE-Bench
A benchmark that tests how well AI models solve real software bugs from open-source projects, widely used to rank coding ability.
Analogy
Like a report card or coding exam for AI programmers.
Why it matters
It's the headline number for comparing coding models, helping you pick the best one for building automations.
In practice
A model's SWE-Bench score is cited as evidence it's the strongest for coding.
