| HumanEval+ HumanEval with 80x more tests to catch incorrect-but-plausible solutions. | Code | Liu et al., 2023, arXiv:2305.01210 | Apache-2.0 | Live |
| LiveCodeBench Continuously updated contest problems to avoid training-set contamination. | Code | Jain et al., 2024, arXiv:2403.07974 | CC-BY-4.0 | Live |
| BigCodeBench Practical programming tasks chaining many real library function calls. | Code | Zhuo et al., 2024, arXiv:2406.15877 | Apache-2.0 | Live |
| RepoBench Repository-level completion requiring cross-file retrieval and context. | Code | Liu et al., 2023, arXiv:2306.03091 | MIT | Live |
| SWE-Lancer Real freelance software tasks priced by their actual payout. | Code | Miserendino et al., 2025, arXiv:2502.12115 | Custom (OpenAI) | Live |