Signals

Filtered to AI engineering practice · clear filter

Browse by topic

The frontier leaderboard is now a list of effort settings, not models

Of the top 15 entries on Artificial Analysis' intelligence index, most are the same handful of models at different reasoning-effort settings — Claude Opus 5 appears at max, xhigh, high and medium; GPT-5.6 Sol does the same. The spread within a single model is wide: Opus 5 ranges from 60.7 down to 56.3 across its…

Source ↗
AI engineering practiceIndustry analysisreasoning-effortbenchmarksmodel-selectionartificial-analysis

Simon Willison probes "the first known runaway AI agent"

Simon Willison examines a claimed incident of an AI agent operating out of control — and openly weighs whether it's a genuine ops failure or a marketing stunt. Either way it's a useful case study in how agent-gone-wrong stories will get reported and how hard they are to verify. Read it as a lesson in demanding…

Source ↗
AI engineering practiceai-agentsagent-safetysimon-willison

OpenAI and Hugging Face disclose a security incident during model evaluation

OpenAI and Hugging Face jointly disclosed and addressed a security incident that occurred during a model evaluation run — among the top HN stories of the week at 1,500+ points. Details are still light on the exact mechanism, but it's a reminder that eval pipelines connecting frontier labs to model hubs are now…

Source ↗
Dev tooling & infraAI engineering practicesecurityopenaihugging-faceeval-infrastructure