The Real-SWE benchmark evaluates AI models on private, real-world, enterprise codebases, challenging their ability to navigate complex business logic and company-specific coding patterns. The benchmark features 8 tasks from private production codebases, each with 8 independent runs per model, resulting in a resolution rate of 15% or lower for most models, with Fable 5.1 achieving the highest resolution rate of 38.8%. The cost of running the benchmark varies from $2.50 to $6.96 per rollout, with Gemini 3.8 Flash and GPT-5.6 Sol being the most cost-effective models. AI summary
Firehose
Filtered to tagged “enterprise” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
artificial intelligence 89continual learning 27AI 23AI safety 13reinforcement learning 13agentic coding 12open-weight models 12AI agents 10machine learning 9AI ethics 8cybersecurity 8existential risk 8language models 8natural language processing 8ethics 7Reinforcement learning 6Diffusion models 5large language models 5multi-agent systems 5open-source 5recursive self-improvement 5robotics 5security 5software development 5Agentic AI 4artificial general intelligence 4mathematics 4Recursive self-improvement 4agents 3AI infrastructure 3