Snapshot 2026-09-21 · rebuilt daily
RL environments and AI benchmarks, cataloged daily.
105 RL environments. 3,144 AI benchmarks. 24,924 model scores. One page per entry: builder, domains, source, related work.
Domain map
Environments and benchmarks by domain
Click a hub to list its entries. Click a circle (RL environment) or diamond (AI benchmark) to open its page.
Select a node to trace its connections. Zoom in for more space; drag to pan.
Biggest domains
Ranked by total entries. 22 domains in all.
Show 16 smaller domains ↓ Hide smaller domains ↑
- 07 Multi-Domain 8 env · 4 bench 12
- 08 Math 4 env · 7 bench 11
- 09 ML 6 env · 1 bench 7
- 10 Private Codebases 7 env · 0 bench 7
- 11 Reasoning 0 env · 4 bench 4
- 12 Creative 3 env · 0 bench 3
- 13 Science 3 env · 0 bench 3
- 14 Security 3 env · 0 bench 3
- 15 Chip Design 2 env · 0 bench 2
- 16 Finance 2 env · 0 bench 2
- 17 Healthcare 1 env · 1 bench 2
- 18 Tool Use 2 env · 0 bench 2
- 19 Alignment 1 env · 0 bench 1
- 20 Multilingual 0 env · 1 bench 1
- 21 Multimodal 0 env · 1 bench 1
- 22 Search 0 env · 1 bench 1
Today's signals · 2026-09-21
What changed in the catalog
September 21, 2026 was a near-static day. The sole change across all seven tracked datasets was a single leaderboard row added to Design Arena — React Native Apps, bringing its total to 46 rows. All other counts held flat: 3,144 benchmarks, 24,924 models, 1,702 companies, 105 environments, 10 tools, 66 occupations, and 1 research article.
Design Arena — React Native Apps Reaches 46 Rows
The multimodal mobile-app generation leaderboard gained one submission today, ticking from 45 to 46 rows. This marks continued incremental community engagement with Design Arena's mobile category, consistent with the steady row growth seen across Design Arena leaderboards throughout September.
The catalog
Environments, benchmarks, domains, and models
Each entry has its own page, and you can sort or filter any of the tables. The snapshot rebuilds daily.
RL environments
Training and eval environments for agents, with builder and domains.
Browse → 80AI benchmarks
Public eval suites and leaderboards referenced across the ecosystem.
Browse → 22Domains
Coding, computer use, enterprise workflows, and more, each with its entries.
Browse → 50Model leaderboard
Ranked by BenchLM overall score, with creator and rank.
Browse →