A benchmark for testing whether coding agents can resolve engineering tasks in scientific software
What is the OpenMOSS/SWE-bench-Science GitHub project? Description: "A benchmark for testing whether coding agents can resolve engineering tasks in scientific software". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.
Question is copied to clipboard — paste it after the AI opens.
Clone via HTTPS
Clone via SSH
Download ZIP
Download main.zipReport bugs or request features on the SWE-bench-Science issue tracker:
Open GitHub Issues