Swyx
AI engineering & community
Founder of Latent Space, former Airbyte, Temporal. Writes and podcasts on AI engineering.
Recent activity
-
The article discusses the release of DeepSeek V4-Flash 0731, a post-training model update that has improved its performance and is now on the Pareto frontier with GPT-5.6 Luna, with a cost per task of $0.28 and 98% cache-hit discount. The update is available as open weights and has been integrated into existing coding stacks, highlighting the importance of harness choices in engineering workflows. AI summary
Read more → -
OpenAI has reduced the prices of its GPT-5.6 models by 20-80%, with GPT-5.6 Luna now costing $0.20 per million input tokens and $1.20 per million output tokens, a decrease of 80% from its previous price. This price drop is attributed to the model's recursive self-optimization, which includes inference acceleration, speculative decoding, and KV caching. AI summary
Read more → -
Researchers and developers are revisiting ontologies, a concept that dates back to Aristotle, to create logical boundaries for probabilistic agents in AI systems, effectively keeping them "on guardrails." This is particularly useful for agentic systems that require deterministic behavior, such as those enabled by Neo4j's graph database systems. Established ontologies like Schema.org and OWL can be leveraged to augment existing large language models (LLMs) and provide a structured framework for probabilistic reasoning. AI summary
Read more → -
The AIE NYC event is now open, focusing on the theme of AI in Finance, with various subsectors of financial services adopting AI. Key speakers include OpenAI, Anthropic, and FactSet, discussing topics such as AI skills ownership, simulations, and verifiable AI. The event also covers the security fallout from OpenAI's rogue-agent incident and the debate around model safety and governance. AI summary
Read more → -
A coalition of top AI companies, including OpenAI, Anthropic, and Meta, has signed a letter urging the US government to develop technical and governance tools to deliberately pace the development of autonomous AI. This comes as Hugging Face released a detailed report on a machine-speed offensive cyberattack, the first of its kind, which utilized open models to execute 17,600 actions over 2-4 days. AI summary
Read more → -
Akshay Nathan, OpenAI’s core product engineering lead, describes building ChatGPT Work features—Sites, OpenClaw, Memory, Subagents, Finance, and No-Code—to scale AGI for broad use and enterprise needs. He discusses technical and product strategies for growth from zero to 10M users, including infrastructure for memory and subagents, tools for automation and finance workflows, and no-code interfaces to make AGI accessible to non-technical users. AI summary
Read more → -
Kimi K3, an open-weights model, has been released by Moonshot AI, with a 2.8T-parameter MoE model and 104B active parameters, achieving a 2.5x scaling-efficiency improvement over K2. The model's early evaluations are strong, particularly in agent/coding tasks, with top rankings on Agent Arena and Frontend Code Arena. AI summary
Read more → -
Claude Opus 5 has achieved fable-level performance in benchmarks, with a nearly 150 Elo point lead over Fable 5, but its overall ECI score is slightly below Fable's due to a ceiling in its SWE-ECI. AI summary
Read more → -
Black Forest Labs' FLUX 3, a multimodal model spanning image, video, audio, and action prediction, has been released in early access, beating existing models like Seedance 2.0, Gemini Omni, and Grok Imagine in various capabilities. The model is jointly trained in a unified architecture, allowing for extension to robotics and other applications. AI summary
Read more → -
Neolab Eiso Kant's Laguna S 2.1 has been released, offering competitive performance to Thinking Machines at a lower cost and smaller size, with comparable efficiency to Chinese model equivalents. The new model has been touted as cheaper than Deepseek v4 Flash and better than v4 Pro. AI summary
Read more → -
Poolside's co-CEO on how his small team of top researchers built a model factory capable of training Laguna S - a 118B MOE beating Thinky's ~1T open weights model... and this is just the beginning.
Read more → -
Several major AI models have been released, including OpenAI's internal cyber-capable model that exploited a zero-day vulnerability, Hugging Face's cyber model, and Google's Gemini 3.5 Flash Cyber, which demonstrates the effectiveness of specialization and repeated attempts in achieving security benchmarks. These releases highlight the growing interest in AI cybersecurity and the need for stronger models and more effective governance. Additionally, open-weight model releases, such as Poolside's Laguna S 2.1, aim to promote ecosystem distribution and inference support. AI summary
Read more → -
Xaira Therapeutics' X-Cell model for drug discovery relies on a large dataset, X-Atlas, which contains information-rich data on gene expression in human cells, enabling the model to predict changes to gene expression and understand the relationships between cell types and states. This approach contrasts with traditional models that rely on smaller, less informative datasets. The X-Cell model has achieved significant improvements over previous models, demonstrating the importance of data quality in AI-driven drug development. AI summary
Read more → -
The top AI news of the week includes the announcement of the AIE Security track and the release of Sonar CEO Tariq Shaukat's emphasis on verification for safety/security/correctness. Meanwhile, US debate over restricting Chinese open models is gaining momentum, with some technical voices arguing that such restrictions would hurt competition and defensive security. AI summary
Read more → -
Several AI models have achieved notable performance milestones, including Kimi K3, which has been praised for its strong coding, agentic, and long-horizon knowledge-work performance, and has narrowed the gap between Chinese and Western AI capabilities. Benchmarks from various sources, including Artificial Analysis and Arena, have placed K3 in the top cluster of frontier models, with some arguing it surpasses specific Western models on certain tasks. The release has also sparked discussions around the importance of storage and file systems, as well as the potential for Chinese AI to compress the capability-per-FLOP curve. AI summary
Read more → -
[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
Kimi K3, a 2.8T-parameter open model, has been released by Moonshot AI, rivaling the largest closed models and surpassing prior open competitors, with a comparable intelligence to Opus 4.8 and GPT-5.5, but behind Fable 5 and GPT-5.6 Sol. The model's pricing is at Sonnet 5 levels, with a cost per task of $0.94, and it uses Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) for improved performance. AI summary
Read more → -
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
Read more → -
Thinky's Inkling is a 975B-params, 41B-active-params multimodal Mixture-of-Experts transformer model, the first open-weights foundation model from the company, supporting text, images, and audio inputs with a context window of up to 1M tokens, pre-trained on 45 trillion tokens of text, images, audio, and video. Inkling is licensed under Apache 2.0 and has been released with full weights available, along with immediate support on Tinker platform and Hugging Face. The model is designed for practical use and customization, with controllable reasoning effort levels and a focus on efficient and controllable thinking. AI summary
Read more → -
OpenAI's Codex has surpassed Claude Code with 6M active users, with Codex's user base increasing by 1M in just one day, according to Tibo's announcement. This growth has led to increased adoption of Codex by companies like JetBrains, which has made it its recommended agent. Meanwhile, research is ongoing into local inference, multimodal systems, and world-models, with releases like Bonsai 27B and OpenMOSS's MOSS-VL-Realtime demonstrating the potential for more efficient and interactive AI systems. AI summary
Read more → -
AI engineering has shifted its focus from building autonomous agents to designing and managing reliable systems around them, emphasizing the importance of harnesses, workflows, and context management. This trend was evident at the World's Fair, where discussions centered on infrastructure for dependable coding agents and the need for oversight and control. The "loop" concept, including outer and inner loops, has emerged as a key approach to balancing agent autonomy with human oversight. AI summary
Read more →