<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Deep Learning Weekly]]></title><description><![CDATA[Bringing you everything new and exciting in the world of  deep learning from academia to the grubby depths  of industry every week right to your inbox.]]></description><link>https://www.deeplearningweekly.com</link><image><url>https://substackcdn.com/image/fetch/$s_!yiM2!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fc63609b6-c5bb-426a-a5c1-b6ce9d56b51e_468x468.png</url><title>Deep Learning Weekly</title><link>https://www.deeplearningweekly.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 12 Sep 2026 18:02:39 GMT</lastBuildDate><atom:link href="https://www.deeplearningweekly.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Deep Learning Weekly]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[deeplearningweekly@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[deeplearningweekly@substack.com]]></itunes:email><itunes:name><![CDATA[Deep Learning Weekly]]></itunes:name></itunes:owner><itunes:author><![CDATA[Deep Learning Weekly]]></itunes:author><googleplay:owner><![CDATA[deeplearningweekly@substack.com]]></googleplay:owner><googleplay:email><![CDATA[deeplearningweekly@substack.com]]></googleplay:email><googleplay:author><![CDATA[Deep Learning Weekly]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Deep Learning Weekly: Issue 472]]></title><description><![CDATA[Introducing ChatGPT Images 2.5, Linguistic drift at the frontier, a paper on Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-472</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-472</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 11 Sep 2026 14:02:48 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://openai.com/index/introducing-chatgpt-images-2-5/"><span>Introducing ChatGPT Images 2.5</span></a><span>, </span><a href="https://pydantic.dev/articles/linguistic-drift-at-the-frontier"><span>Linguistic drift at the frontier</span></a><span> and </span><a href="https://arxiv.org/abs/2609.03430"><span>a paper on Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://www.anthropic.com/research/formalizing-fermats-last-theorem"><span>Formalizing Fermat&#8217;s Last Theorem \ Anthropic</span></a><span>, </span><a href="https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude"><span>Latency Scaling Differences for GPT and Claude Models</span></a><span>, </span><a href="https://arxiv.org/abs/2609.04148"><span>a paper on Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://openai.com/index/introducing-chatgpt-images-2-5/"><span>Introducing ChatGPT Images 2.5</span></a></strong></p><p><span>OpenAI released GPT-Image-2.5 in Flare and Sunburst variants, up to 50% faster with sketch input, C2PA metadata, and invisible watermarking.</span></p><p><strong><a href="https://www.anthropic.com/research/formalizing-fermats-last-theorem"><span>Formalizing Fermat&#8217;s Last Theorem</span></a></strong></p><p><span>Anthropic&#8217;s internal research model produced the first complete computer-checked proof of Fermat&#8217;s Last Theorem in 11 days, generating 13 million lines of Lean across 30,300 theorems.</span></p><p><strong><a href="https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/"><span>AI Has Solved One of Math&#8217;s $1 Million Millennium Prize Problems</span></a></strong></p><p><span>OpenAI deployed 10,000 collaborating agents for 88 hours to prove a finite-time singularity in the 3D Navier&#8211;Stokes equations, formally verified in Lean, hours after an independent competing announcement.</span></p><p><strong><a href="https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/"><span>Mistral raises &#8364;3B to make sovereign, open-weight AI the technology frontier</span></a></strong></p><p><span>Mistral raised a &#8364;3 billion Series D led by Samsung Electronics at a valuation above &#8364;21 billion, the largest equity round ever for a European technology company.</span></p><p><strong><a href="https://siliconangle.com/2026/09/08/meta-debuts-its-secure-by-design-personal-ai-agent-muse/"><span>Meta debuts its &#8220;secure by design&#8221; personal AI agent Muse</span></a></strong></p><p><span>Meta launched Muse, a personal AI agent that runs tasks inside a secure VM with Sentinel monitoring and 100 million free weekly tokens, using single-use Stripe payment cards.</span></p><p><strong><a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/"><span>AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome</span></a></strong></p><p><span>Google DeepMind released AlphaGenome Atlas, a 1-petabyte predictive map scoring all 9 billion possible single-letter DNA variants in the human genome, free for noncommercial research.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://pydantic.dev/articles/linguistic-drift-at-the-frontier"><span>Linguistic drift at the frontier</span></a></strong></p><p><span>A fascinating article about &#8220;linguistic drift,&#8221; tracing Claude-popularized vocabulary into 685,000 GitHub PRs in 2026 and introducing vocabguard, a drift-monitoring capability with 0.869 AUC.</span></p><p><strong><a href="https://opentelemetry.io/blog/2026/genai-observability/"><span>Inside the LLM Call: GenAI Observability with OpenTelemetry</span></a></strong></p><p><span>A tutorial about OpenTelemetry&#8217;s Generative AI semantic conventions, tracing agent invocations, LLM calls, and tool executions with standardized token-usage and duration metrics.</span></p><p><strong><a href="https://research.ibm.com/blog/running-open-models-on-h100-gpus-with-llmd"><span>How llm-d makes the most of the hardware you already have</span></a></strong></p><p><span>A benchmark blog post about llm-d serving a 753B-parameter MoE across 544 H100s to 3,000 concurrent coding agents, with 85.2% of input tokens served from cache.</span></p><p><strong><a href="https://vercel.com/changelog/vercel-sandbox-routing-is-now-18x-faster-globally"><span>Vercel Sandbox routing is now 18x faster globally</span></a></strong></p><p><span>Vercel cut sandbox domain-resolution latency 18x globally &#8212; from 62ms to 3.4ms at the median &#8212; by resolving domains from regional replicas instead of one centralized store.</span></p><p><strong><a href="https://ampcode.com/news/desktop"><span>Amp Desktop</span></a></strong></p><p><span>Amp added a Desktop tab giving agent threads an interactive Linux desktop for verifying work that needs real applications, from LibreOffice exports to computer-use tasks.</span></p><p><strong><a href="https://medium.com/google-cloud/look-but-dont-touch-7f470591ff1f"><span>Look, But Don&#8217;t Touch (Read-Only Tools for AI Agents)</span></a></strong></p><p><span>A blog post about protocol-level read-only enforcement in MCP Toolbox for Databases, showing how prompt guardrails, regex parsers, and session flags all fail against injection.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude"><span>Long-context latency scales quadratically for GPT-5.6 but nearly linearly for Claude 5</span></a></strong></p><p><span>A data analysis about time-to-first-token scaling quadratically with context for GPT-5.6 but nearly linearly for Claude 5, measured up to million-token contexts.</span></p><p><strong><a href="https://cohere.com/blog/megakernels"><span>Inside the megakernel serving engine for North Mini Code</span></a></strong></p><p><span>A deep engineering blog post about serving a 30B model through one persistent CUDA &#8220;megakernel,&#8221; hitting 292 tokens/sec at batch one &#8212; 62% of theoretical peak versus vLLM&#8217;s 39%.</span></p><p><strong><a href="https://developer.nvidia.com/blog/building-a-memory-driven-agent-with-nvidia-nemoclaw/"><span>Building a Memory-Driven Agent with NVIDIA NemoClaw</span></a></strong></p><p><span>A tutorial about a memory-driven &#8220;self model&#8221; agent architecture scoring 90.9% versus an 82.8% agentic-RAG baseline, with changed-fact tracking jumping from 60% to 100%.</span></p><p><strong><a href="https://huggingface.co/blog/grpo-with-trl-ifstruct"><span>Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps</span></a></strong></p><p><span>A hands-on tutorial about GRPO fine-tuning a 350M model for structured outputs on a free 16GB GPU, lifting JSON format compliance from 18.0% to 31.9% in 100 steps.</span></p><p><strong><a href="https://spectrum.ieee.org/ai-code-review-software-engineers"><span>AI Slop Is Changing How Engineers Review Code</span></a></strong></p><p><span>An article about AI-generated code straining review, citing a Sonar survey where 42% of shared-codebase code is AI-attributed and 96% of developers don&#8217;t fully trust it.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/okf-memory/okf-agent-memory"><span>okf-memory/okf-agent-memory</span></a></strong></p><p><span>A Domain-Neutral, Git-Native Persistent Project Memory for AI Agents based on the Open Knowledge Format (OKF).</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2609.03430"><span>Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks it matches the strongest prior evictor while serving 32-43% higher throughput than it in vLLM deployment. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them.</span></p><p><strong><a href="https://arxiv.org/abs/2609.04148"><span>Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 471]]></title><description><![CDATA[Introducing Claude Fable 5.1 and Claude Mythos 5.1, Evaluating LLMs Under Production Parity, a paper on Automated Researchers Can Reliably Mitigate Alignment Failures, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-471</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-471</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 04 Sep 2026 16:02:02 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://www.anthropic.com/claude-fable-and-mythos-5-1"><span>Introducing Claude Fable 5.1 and Claude Mythos 5.1</span></a><span>, </span><a href="https://huggingface.co/blog/TechforHumans/evaluating-llms-under-production-parity"><span>Evaluating LLMs Under Production Parity</span></a><span> and </span><a href="https://alignment.anthropic.com/2026/automated-alignment-researchers/"><span>a paper on Automated Researchers Can Reliably Mitigate Alignment Failures</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra"><span>OpenAI launches GPT-6 Astra</span></a><span>, </span><a href="https://epoch.ai/data-insights/eci-frontier-trend"><span>The ECI frontier has advanced by 14 points per year since reasoning models</span></a><span>, </span><a href="https://machinelearning.apple.com/research/llms-not-consistently-bayesian"><span>a paper on LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs&#8217; Probabilistic Beliefs</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://www.anthropic.com/claude-fable-and-mythos-5-1"><span>Introducing Claude Fable 5.1 and Claude Mythos 5.1</span></a></strong></p><p><span>Anthropic released Claude Fable 5.1 and Mythos 5.1, cutting cache reads 75% to $0.25 per million tokens and more than doubling Terminal-Bench-Science scores to 52.6%.</span></p><p><strong><a href="https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra"><span>Welcome to the AGI era: OpenAI launches GPT-6 Astra</span></a></strong></p><p><span>OpenAI launched GPT-6 Astra, a computer-use model scoring 98.6% on ARC-AGI-3 and 100% on ExploitBench, priced at $10/$50 per million tokens.</span></p><p><strong><a href="https://techcrunch.com/2026/09/02/openais-new-reasoning-technique-alarms-ai-safety-experts/"><span>OpenAI&#8217;s new reasoning technique alarms AI safety experts</span></a></strong></p><p><span>Astra&#8217;s &#8220;opaque recurrence&#8221; loops computation instead of visible chain-of-thought, which Redwood researchers warn destroys the monitorability that oversight currently depends on.</span></p><p><strong><a href="https://techcrunch.com/2026/09/03/nvidia-confirms-it-will-buy-hugging-face-for-12-9-billion/"><span>Nvidia confirms it will buy Hugging Face for $12.9 billion</span></a></strong></p><p><span>Nvidia confirmed a $12.93 billion acquisition of Hugging Face, whose platform hosts 3 million models and 18 million developers, pledging the hub stays open.</span></p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"><span>Introducing Gemini 3.8 Flash and 3.8 Flash Cyber</span></a></strong></p><p><span>Google released Gemini 3.8 Flash and a cyber-specialized variant that Chrome Security reports produces 2.6x more correct vulnerability patches than competing commercial models.</span></p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/google-deepmind/introducing-weathernext-3/"><span>Introducing WeatherNext 3</span></a></strong></p><p><span>Google DeepMind&#8217;s WeatherNext 3 delivers 5km hourly forecasts, five times sharper than its predecessor, improving precipitation accuracy up to 60% against NASA satellite data.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://azure.microsoft.com/en-us/blog/the-economics-of-agent-optimization-context-engineering-for-enterprise-ai-agents/"><span>The Economics of Agent Optimization</span></a></strong></p><p><span>A detailed blog post about context engineering for enterprise agents, reporting 54% better evidence recall, 34% lower retrieval token costs, and roughly 97% input-token reduction from tool search.</span></p><p><strong><a href="https://www.comet.com/site/blog/agent-first-terminal-ux/"><span>UX Without a UI: Notes from Researching Terminal UX for Opik MCP</span></a></strong></p><p><span>A thoughtful blog post about designing terminal UX for LLM observability, covering six rendering constraints and &#8220;numbered spans&#8221; for referencing trace rows without copying.</span></p><p><strong><a href="https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month"><span>What&#8217;s new in AI infrastructure and orchestration in August</span></a></strong></p><p><span>A monthly roundup about AI infrastructure updates, including gVisor sandboxes on distributed Ray clusters, fully stateless MCP, and a customer cutting pipeline costs 90%.</span></p><p><strong><a href="https://cursor.com/blog/self-hosted-machines"><span>Cursor Self-Hosted Machines</span></a></strong></p><p><span>A blog post about running Cursor Cloud Agents on customer-controlled machine pools with hibernation and snapshots; cloud agents now open over 60% of Cursor&#8217;s merged pull requests.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://huggingface.co/blog/TechforHumans/evaluating-llms-under-production-parity"><span>Evaluating LLMs Under Production Parity</span></a></strong></p><p><span>A methodical article about a replay pipeline for safely swapping models, where hallucination rate acted as an eliminatory gate rejecting three otherwise-passing candidates.</span></p><p><strong><a href="https://epoch.ai/data-insights/eci-frontier-trend"><span>The ECI frontier has advanced by 14 points per year since reasoning models</span></a></strong></p><p><span>A data analysis about capability growth accelerating to 14 index points per year since reasoning models arrived, more than double the prior six-point rate.</span></p><p><strong><a href="https://stripe.dev/blog/ai-didnt-write-our-sdk-it-changed-how-we-built-it"><span>AI didn&#8217;t write our SDK. It changed how we built it.</span></a></strong></p><p><span>A candid engineering blog post about spec-first delegation to AI agents, which produced a working SDK prototype in two days versus a typical two-to-three-week effort.</span></p><p><strong><a href="https://www.decodingai.com/p/subagents-are-context-engineering"><span>From 1 Bloated Context Window to 6 Scoped Subagents</span></a></strong></p><p><span>A hands-on course lesson about replacing one bloated context window with six scoped subagents, using semaphore limits, shared byte budgets, and per-child request caps.</span></p><p><strong><a href="https://weaviate.io/blog/charts-tables-pdfs"><span>How to extract meaning from charts and tables in PDFs</span></a></strong></p><p><span>A practical guide about late-interaction multi-vector retrieval over chart-heavy PDFs, skipping OCR and chunking entirely in roughly 50 lines of Python.</span></p><p><strong><a href="https://www.quantamagazine.org/live-from-icm-2026-what-is-math-for-in-the-age-of-ai-20260903/"><span>What Is Math For in the Age of AI?</span></a></strong></p><p><span>A panel discussion about what mathematics is for in the age of AI, covering AlphaProof&#8217;s olympiad results and concerns about vanishing training problems for graduate students.</span></p><p><strong><a href="https://cohere.com/blog/how-small-models-can-make-a-big-impact-for-enterprises"><span>How small AI models can make a big impact for enterprises</span></a></strong></p><p><span>A blog post about enterprise model portfolios, citing a 30B/3B-active coding model that outperforms some 123B models on Artificial Analysis&#8217; coding index.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/google-research/timesfm"><span>google-research/timesfm</span></a></strong></p><p><span>TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://alignment.anthropic.com/2026/automated-alignment-researchers/"><span>Automated Researchers Can Mitigate Well-Characterized Alignment Failures</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7&#215; larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs&#8217; initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers. These results suggest that automating alignment research on well-characterized failures may be practical in the near term.</span></p><p><strong><a href="https://machinelearning.apple.com/research/llms-not-consistently-bayesian"><span>LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs&#8217; Probabilistic Beliefs - Apple Machine Learning Research</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decisions. We introduce the novel technique of studying LLMs as information processing rules and utilize the information processing gap&#8212;the deviation from Bayes updates&#8212;to study the internal (in)consistencies of how LLMs update their probabilistic beliefs from evidence. Our extensive experiments evaluate multiple approaches in which LLMs can incorporate evidence into their beliefs. Some of these approaches produce (nearly) Bayesian updates, thus optimally processing evidence; others use a learned heuristic. Surprisingly, the non-Bayesian heuristic updates often outperform exact Bayesian updates (optimal information processing) in terms of downstream task performance&#8212;indicating the LLMs&#8217; probabilistic models of the world are misspecified. Lastly, we show how our measure can provide diagnostics to identify issues with LLM-powered inferential s</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 470]]></title><description><![CDATA[GLM-5.3-Flash, Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments, a paper on FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution, and many]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-470</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-470</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 27 Aug 2026 15:03:29 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://z.ai/blog/glm-5.3-flash"><span>GLM-5.3-Flash: Frontier Intelligence, Flash Cost</span></a><span>, </span><a href="https://alignment.anthropic.com/2026/chive/"><span>Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments</span></a><span> and </span><a href="https://arxiv.org/abs/2608.16157"><span>a paper on FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://techcrunch.com/2026/08/25/stability-ai-maker-of-image-generator-stable-diffusion-raises-76-million-in-fresh-funding/"><span>Stability AI raises $76 million</span></a><span>, </span><a href="https://www.philschmid.de/recursive-self-improvement"><span>Recursive Self-Improvement</span></a><span>, </span><a href="https://arxiv.org/abs/2608.24979"><span>a paper on FrontierChallenge: Evaluating Scientific Workflow Completion</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://z.ai/blog/glm-5.3-flash"><span>GLM-5.3-Flash: Frontier Intelligence, Flash Cost</span></a></strong></p><p><span>Z.ai open-sources GLM-5.3-Flash, the stealth &#8220;Ox Alpha&#8221; model &#8212; a 320B/18B MoE with hybrid sparse-plus-linear attention and 1M-token multimodal context, at a tenth of GLM-5.3&#8217;s price.</span></p><p><strong><a href="https://techcrunch.com/2026/08/25/stability-ai-maker-of-image-generator-stable-diffusion-raises-76-million-in-fresh-funding/"><span>Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding</span></a></strong></p><p><span>Stability AI raises $76M Series B backed by Universal Music, Sony Music, Warner Music, and EA, converting content-licensing partners into investors as it builds out creative production tooling.</span></p><p><strong><a href="https://venturebeat.com/infrastructure/perplexity-partners-with-nvidia-to-launch-portable-computer-a-fully-local-ai-agent-with-zero-token-costs"><span>Perplexity partners with Nvidia to launch Portable Computer, a fully local AI agent with zero token costs</span></a></strong></p><p><span>Perplexity and Nvidia launch Portable Computer, a fully local agent stack on DGX Spark and RTX hardware whose co-designed minimal harness beats open-source harnesses on the same 27B model.</span></p><p><strong><a href="https://siliconangle.com/2026/08/25/agentic-web-search-infrastructure-startup-keenable-raises-26m/"><span>Agentic web search infrastructure startup Keenable raises $26M</span></a></strong></p><p><span>Keenable launched with $26M seed from Accel and Conviction to run a 100-billion-document independent index built for agent-rate querying, including point-in-time historical search and a Web Query Language for cross-source answers.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/llm-observability-tools/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=llm-observability-tools/"><span>Best LLM Observability Tools of 2026: Top Platforms &amp; Features</span></a></strong></p><p><span>A comparison guide to 13 LLM observability platforms, splitting the field into evaluation-centric and observability-centric tools and scoring each on tracing depth, integration approach, self-hosting, and pricing.</span></p><p><strong><a href="https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management"><span>Best practices for dynamic capacity management</span></a></strong></p><p><span>A practical guide to dynamic capacity management for agentic workloads, pairing scheduled reservations with automated hardware fallback lists and GPU/TPU slicing so bursty AI demand doesn&#8217;t strand idle compute.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://www.philschmid.de/recursive-self-improvement"><span>Recursive Self-Improvement</span></a></strong></p><p><span>An essay arguing that agents editing their own tools and harness is bounded self-improvement, not recursion, because no public system can raise its own verifier without capturing it.</span></p><p><strong><a href="https://www.decodingai.com/p/run-coding-agents-safely"><span>From a Raw Shell to a Sandboxed Coding Agent</span></a></strong></p><p><span>A deep dive that walks through how to safely sandbox coding agents using Docker and Modal, covering everything from execution boundaries to remote sandboxes, GPU access, and running agents at scale.</span></p><p><strong><a href="https://alignment.anthropic.com/2026/chive/"><span>Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments</span></a></strong></p><p><span>A research post introducing CHIVE, an agentic pipeline that discovers in-the-wild LLM behaviors and tests explanations via counterfactual prompt edits, finding activation-reading interpretability tools give zero uplift over a transcript-only baseline.</span></p><p><strong><a href="https://research.google/blog/agenthands-generating-interactive-hand-gestures-for-spatially-grounded-agent-conversations-in-xr/"><span>AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR</span></a></strong></p><p><span>A research blog post on AgentHands, a Google XR prototype where an LLM emits inline gesture events synced to speech, significantly improving spatial grounding and warning salience over a speech-only baseline.</span></p><p><strong><a href="https://cloud.google.com/blog/products/ai-machine-learning/how-agents-can-delegate-better"><span>How agents can delegate better</span></a></strong></p><p><span>A blog post distilling DeepMind&#8217;s Intelligent AI Delegation research into four principles for multi-agent orchestration: contract-first decomposition, cost-aware model routing, least-privilege data sharing, and cognitive friction against blind compliance down delegation chains.</span></p><p><strong><a href="https://deepmind.google/blog/from-atari-to-eve-online-building-on-15-years-of-ai-research-in-games/"><span>From Atari to EVE Online: Building on 15 Years of AI Research in Games</span></a></strong></p><p><span>A blog post tracing fifteen years from Atari-playing deep RL to generalist game agents, then framing a new Fenris Creations partnership around EVE Online as a testbed for continual learning, long-horizon memory, and multi-agent dynamics.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/PrimeIntellect-ai/prime-agent"><span>PrimeIntellect-ai/prime-agent</span></a></strong></p><p><span>Prime Agent is an open-source coding and research agent for general and long-running work.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2608.16157"><span>FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence.</span></p><p><strong><a href="https://arxiv.org/abs/2608.24979"><span>FrontierChallenge: Evaluating Scientific Workflow Completion</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 469]]></title><description><![CDATA[FLUX Upscale: 2K and 4K for Video, Empty shelves or lost keys? Recall is the bottleneck for parametric factuality, a paper on VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?,]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-469</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-469</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 21 Aug 2026 15:01:13 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://bfl.ai/blog/flux-video-upscale"><span>FLUX Upscale: 2K and 4K for Video</span></a><span>, </span><a href="https://research.google/blog/empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality/"><span>Empty shelves or lost keys? Recall is the bottleneck for parametric factuality</span></a><span> and </span><a href="https://arxiv.org/abs/2608.15265"><span>a paper on VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>Pacing model development in an era of cyber-critical capabilities</span></a><span>, </span><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/"><span>Introducing the Conceptual Reasoning Index</span></a><span>, </span><a href="https://arxiv.org/abs/2608.16590"><span>a paper on Zetta: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://bfl.ai/blog/flux-video-upscale"><span>FLUX Upscale: 2K and 4K for Video</span></a></strong></p><p><span>Black Forest Labs launches FLUX Upscale, a standalone tool and API endpoint that regenerates video at up to native 4K while repairing generation artifacts, offered in Precise and Creative modes.</span></p><p><strong><a href="https://openai.com/index/pacing-model-development-cyber-capabilities/"><span>Pacing model development in an era of cyber-critical capabilities</span></a></strong></p><p><span>OpenAI paused its largest frontier RL run and slowed scaling after preliminary evidence that its upcoming model Astra may cross the &#8220;Critical&#8221; cybersecurity threshold, hardening research-environment security, monitoring, and alignment before proceeding.</span></p><p><strong><a href="https://siliconangle.com/2026/08/19/stripe-buys-ai-model-router-openrouter-in-reported-7-5b-deal/"><span>Stripe buys AI model router OpenRouter in reported $7.5B deal</span></a></strong></p><p><span>Stripe agreed to acquire AI model-routing startup OpenRouter &#8212; a single-endpoint gateway to 400+ models from 80+ providers handling 10T+ daily tokens &#8212; in a reported ~$7.5B+ deal.</span></p><p><strong><a href="https://openai.com/index/offering-zero-data-retention-for-frontier-models/"><span>Offering Zero Data Retention for frontier models</span></a></strong></p><p><span>OpenAI reaffirmed Zero Data Retention for eligible API customers and previewed Private Safety Processing.</span></p><p><strong><a href="https://www.csail.mit.edu/news/when-ai-art-has-no-author-mit-study-finds-generated-images-often-cant-be-traced-any-training"><span>When AI art has no author: MIT study finds generated images often can&#8217;t be traced to any training data</span></a></strong></p><p><span>MIT CSAIL researchers identified &#8220;attribution decay&#8221; &#8212; showing that in large-scale generative models, removing any single training image often leaves outputs unchanged, meaning many AI images can&#8217;t be attributed to any specific source.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/what-is-ai-observability/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=what-is-ai-observability/"><span>What is AI Observability? A Complete Guide to Debugging and Monitoring Modern AI Systems at Scale</span></a></strong></p><p><span>A guide on why AI observability must track system behavior, not just health &#8212; capturing telemetry across four layers (app, agent, model, retrieval) and adding evaluations as a fourth signal alongside logs, metrics, and traces to catch semantic failures that leave infra dashboards green.</span></p><p><strong><a href="https://openai.com/index/builders-guide-to-gpt-5-6/"><span>The builder&#8217;s guide to GPT&#8209;5.6</span></a></strong></p><p><span>OpenAI&#8217;s builder&#8217;s guide argues GPT-5.6 collapses agent economics by pairing lower-reasoning-effort accuracy with new Responses API primitives, letting startups match frontier quality at a fraction of the cost.</span></p><p><strong><a href="https://developer.nvidia.com/blog/building-federated-multimodal-ai-workflows-with-nvidia-flare/"><span>Building Federated Multimodal AI Workflows with NVIDIA FLARE</span></a></strong></p><p><span>A technical guide showing how NVIDIA FLARE federates multimodal VLM training, where FedUMM&#8217;s adapter-only approach cut per-client communication from 28.6 GB to 0.094 GB per round while holding ~97% of centralized performance.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://jamwithai.substack.com/p/build-your-own-voice-agent"><span>Build Your Own Voice Agent</span></a></strong></p><p><span>The final part of the Observable Job Agent series brings the agent to life with voice. See how LangGraph, ElevenLabs, FastAPI, and Opik come together to build a voice agent that can search, rank, and tailor job applications.</span></p><p><strong><a href="https://research.google/blog/empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality/"><span>Empty shelves or lost keys? Recall is the bottleneck for parametric factuality</span></a></strong></p><p><span>Google Research&#8217;s knowledge profiling framework shows factual errors in frontier LLMs are overwhelmingly recall failures, not encoding failures &#8212; the facts are stored but inaccessible, shifting the factuality bottleneck from acquisition to utilization.</span></p><p><strong><a href="https://alignment.anthropic.com/2026/conceptual-reasoning-index/"><span>Introducing the Conceptual Reasoning Index</span></a></strong></p><p><span>Redwood Research and Anthropic introduce the Conceptual Reasoning Index, aggregating three benchmarks that measure argumentation on unverifiable questions, where top scorer Opus 5 hits 73.6 against an estimated ceiling of 91.</span></p><p><strong><a href="https://huggingface.co/blog/ibm-research/altk-evolve-sldd"><span>Thinking of ACE? We Can Do It with Fewer Tokens</span></a></strong></p><p><span>IBM Research&#8217;s ALTK-Evolve matches or beats ACE on AppWorld agent tasks by calibrating how much learned &#8220;memory&#8221; reaches the model per task rather than injecting the full playbook every step.</span></p><p><strong><a href="https://ai-frontiers.org/articles/hyperlaw-ai-will-change-how-law-evolves"><span>Hyperlaw: AI Will Change How Law Evolves</span></a></strong></p><p><span>An analytical article about &#8220;hyperlaw&#8221; &#8212; how AI-driven legal productivity gains of 12%&#8211;130% cut both contracting and litigation costs, shrinking settlement ranges and accelerating precedent change fastest in contract-free domains.</span></p><p><strong><a href="https://embracethered.com/blog/posts/2026/recovering-encrypted-llm-thoughts/"><span>Recovering Encrypted LLM Reasoning Traces</span></a></strong></p><p><span>A security research post reproducing an attack that replays encrypted LLM reasoning blobs across sessions, accounts, and models to recover hidden traces &#8212; the source paper extracted 367 PII items and 182 credentials from 315,320 public blocks.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2608.15265"><span>VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.</span></p><p><strong><a href="https://arxiv.org/abs/2608.16590"><span>Zetta &#950;: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today&#8217;s large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic &#8220;Aha Moments&#8221; emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 468]]></title><description><![CDATA[Introducing Grok 4.6, TutorMoments: Do AI tutors know when to help and when to hold back?, a paper on BDH-CQ: In-Context Learning with Recurrent Latent Reasoning, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-468</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-468</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 13 Aug 2026 15:02:24 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://x.ai/news/grok-4-6"><span>Introducing Grok 4.6</span></a><span>, </span><a href="https://huggingface.co/blog/allenai/tutormoments"><span>TutorMoments: Do AI tutors know when to help and when to hold back?</span></a><span> and </span><a href="https://arxiv.org/abs/2608.09888"><span>a paper on BDH-CQ: In-Context Learning with Recurrent Latent Reasoning</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"><span>Meta&#8217;s Muse Glimmer</span></a><span>, </span><a href="https://epoch.ai/publications/one-in-five-workers-delegate-work-to-ai"><span>One in five US workers now delegates tasks to AI instead of other humans</span></a><span>, </span><a href="https://arxiv.org/abs/2608.09802"><span>a paper on SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://x.ai/news/grok-4-6"><span>Introducing Grok 4.6</span></a></strong></p><p><span>xAI releases Grok 4.6, tuned for long-running agents and interactive builds, matching GPT-5.6 Sol at 61 on the AA Intelligence Index and leading on GDPval-AA (1753) and AA-Briefcase (1577) at unchanged $2/$6 pricing.</span></p><p><strong><a href="https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"><span>Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device</span></a></strong></p><p><span>Meta open-sources Muse Glimmer, a 30B agentic model distilled from Muse Spark and quantized to under 20GB so it runs always-on locally, with speculative decoding delivering up to 3.1x faster decode.</span></p><p><strong><a href="https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/"><span>Expanding Daybreak as the Cyber Defense Window Narrows</span></a></strong></p><p><span>OpenAI expands Daybreak into Blue and Red access tiers and launches GPT-5.6-Cyber, which completes 95% of advanced offensive-security requests versus 1.5% for the guarded base model.</span></p><p><strong><a href="https://venturebeat.com/security/aws-continuum-integrates-with-openai-codex-and-anthropic-claude-code-in-major-ai-security-push"><span>AWS Continuum integrates with OpenAI Codex and Anthropic Claude Code in major AI security push</span></a></strong></p><p><span>AWS embeds its Continuum vulnerability platform directly into Claude Code, OpenAI Codex, and Kiro, positioning itself as a model-neutral security control plane that absorbs token costs behind a single price.</span></p><p><strong><a href="https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/"><span>Anthropic says it will watermark text generated by its AI models</span></a></strong></p><p><span>Anthropic will watermark all Claude-generated text and files at the model level to comply with the EU AI Act&#8217;s Transparency Code, with marks persisting through copy-paste and some editing.</span></p><p><strong><a href="https://deepmind.google/blog/putting-sign-language-ai-into-users-hands/"><span>Putting sign language AI into users&#8217; hands</span></a></strong></p><p><span>Google DeepMind ships SL2T, a sign-language-to-text model trained on 100,000+ hours across 50+ sign languages, hitting 70 BLEURT zero-shot and powering ASL dictation in Gboard and Live Transcribe on Pixel 11.</span></p><p><strong><a href="https://venturebeat.com/technology/ltx-2-5-can-generate-a-10-second-ai-video-from-an-image-in-just-6-8-seconds-on-nvidia-superchips-and-its-open-weights"><span>LTX-2.5 can generate a 10-second AI video from an image in just 6.8 seconds on Nvidia superchips</span></a></strong></p><p><span>LTX releases LTX-2.5, a 22B open-weights video world model that renders a 10-second clip in 6.8 seconds on dual GB200s &#8212; roughly 7.6x faster than the quickest closed API.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://jamwithai.substack.com/p/build-your-own-job-agent-part-3"><span>Build your own Job Agent - Part 3</span></a></strong></p><p><span>Part 3 of The Observable Job Agent series continues with a practical look at LLMOps in action: using Opik traces, evals, prompt optimization, and Ollie to diagnose failures and improve agent performance.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://huggingface.co/blog/allenai/tutormoments"><span>TutorMoments: Do AI tutors know when to help and when to hold back?</span></a></strong></p><p><span>Ai2 releases TutorMoments, a replay-based eval built on 462 real math tutoring transcripts and 1,500+ teacher-annotated decision points, finding LLMs systematically over-help rather than pushing students toward productive struggle.</span></p><p><strong><a href="https://epoch.ai/publications/one-in-five-workers-delegate-work-to-ai"><span>One in five US workers now delegates tasks to AI instead of other humans</span></a></strong></p><p><span>Epoch AI survey of 1,106 US workers finds 20% now delegate to AI work once handed to a coworker or contractor, with 66% of AI outputs accepted unchanged or with minor edits.</span></p><p><strong><a href="https://weaviate.io/blog/search-mode-effort"><span>Scaling Test-Time Compute in Search Mode</span></a></strong></p><p><span>Weaviate adds medium/high/ultrahigh effort tiers to Query Agent Search Mode, scaling test-time compute on query decomposition and reranking to lift BRIGHT Biology nDCG@10 from 13.0 to 57.5 over hybrid search.</span></p><p><strong><a href="https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards"><span>Improving Fable 5 Safeguards</span></a></strong></p><p><span>A detailed blog post about how Anthropic rewrote Claude Fable 5&#8217;s biology classifier constitution to cut false-positive fallbacks by ~85%, illustrating the tradeoff between launching broad safeguards early and refining precision over time.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/NVIDIA-NeMo/Switchyard"><span>NVIDIA-NeMo/Switchyard</span></a></strong></p><p><span>Switchyard is a Rust proxy and library for LLM traffic. It routes requests across providers, translates between OpenAI and Anthropic APIs, records operational metrics, and provides typed, composable routing algorithms.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2608.09888"><span>BDH-CQ: In-Context Learning with Recurrent Latent Reasoning</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model&#8217;s recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the model on the public ARC-AGI-1 evaluation set and use controlled ARC-like interventions to study what it learns from demonstrations, how consistently it applies an inferred transformation, and which concepts remain difficult. A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of $0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.</span></p><p><strong><a href="https://arxiv.org/abs/2608.09802"><span>SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and that frontier models can verbatim reproduce gold patches from training data. Code refactoring, which requires coordinated, behavior-preserving changes across many files, offers a substantially harder and more realistic test of agent capability, yet remains underserved by current benchmarks. We introduce SWE-Bench ProMax, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation that directly addresses the quality problems identified in prior benchmarks: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to remove overly narrow and overly broad tests. Tasks with insufficient complexity or limited cross-file scope are filtered out, yielding a benchmark of challenging, large-scale refactoring tasks that average 11.4 modified files and 261.6 lines of code per instance, substantially exceeding the scale of existing benchmarks. Experiments with frontier models under two agent scaffolds show that the best model achieves only 41.2% resolve rate, confirming that SWE-Bench ProMax presents a meaningful and unsaturated challenge for current AI coding agents.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 467]]></title><description><![CDATA[Qwen3.8-Max, A verifiable autonomous research framework via Chain-of-Evidence, a paper on From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement, a]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-467</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-467</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 06 Aug 2026 16:00:53 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://qwen.ai/blog?id=qwen3.8"><span>Qwen3.8-Max</span></a><span>, </span><a href="https://research.google/blog/science-one-framework-a-verifiable-autonomous-research-framework-via-chain-of-evidence/"><span>Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence</span></a><span> and </span><a href="https://arxiv.org/abs/2607.23802"><span>a paper on From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://mistral.ai/news/shieldstral/"><span>Shieldstral</span></a><span>, </span><a href="https://epoch.ai/publications/parallelization-constraints-could-delay-a-technological-singularity"><span>Even after R&amp;D is automated, parallelization constraints could delay a technological singularity</span></a><span>, </span><a href="https://arxiv.org/abs/2608.04003"><span>a paper on PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://qwen.ai/blog?id=qwen3.8"><span>Qwen: Qwen3.8-Max: A New Bar for Coding and Cowork</span></a></strong></p><p><span>Alibaba releases Qwen3.8-Max, a 2.4T-parameter multimodal MoE with 95B active parameters and 1M-token context, leading Terminal-Bench 2.1 at 86.6 and marking the first Max-class Qwen to go open-weights.</span></p><p><strong><a href="https://mistral.ai/news/shieldstral/"><span>Introducing Shieldstral</span></a></strong></p><p><span>Mistral releases Shieldstral, a 3B Apache 2.0 multimodal safety classifier that takes plain-language policies at inference time and matches guard models up to 7x its size while running on a single 16GB GPU.</span></p><p><strong><a href="https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"><span>Gemini Robotics 2 brings whole body intelligence to robots</span></a></strong></p><p><span>Google DeepMind launches Gemini Robotics 2, a three-model family bringing full humanoid whole-body control, 22-DoF multi-finger dexterity, and multi-robot collaboration, with on-device adaptation to new embodiments in hours from under 200 examples.</span></p><p><strong><a href="https://siliconangle.com/2026/07/30/nscale-buys-ai-infrastructure-optimization-startup-anyscale-reported-1-65b/"><span>Nscale buys AI infrastructure optimization startup Anyscale for reported $1.65B</span></a></strong></p><p><span>Data center builder Nscale acquires Anyscale, the company behind the open-source Ray cluster-optimization framework, for a reported $1.65B to complete a vertically integrated AI cloud spanning power, compute, and software.</span></p><p><strong><a href="https://siliconangle.com/2026/08/04/obsidian-security-raises-85m-ai-agents-create-cybersecuritys-next-major-attack-surface/"><span>Obsidian Security raises $85M as AI agents create cybersecurity&#8217;s next major attack surface</span></a></strong></p><p><span>Obsidian Security raises $85M Series D at a $1.1B valuation to govern non-human AI agent identities, with 65% of its enterprise customers already granting agents access to third-party SaaS data.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://blog.dailydoseofds.com/p/8904b4e2-4510-4221-8e5d-18f44a3a1d59"><span>86% of Your Claude Code Bill Has Nothing to Do With Your Prompts</span></a></strong></p><p><span>A look into where Claude Code tokens actually go, revealing that prompts make up just a fraction of total spend while context replay, MCP schemas, and tool outputs quietly dominate the bill.</span></p><p><strong><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"><span>Investigating three real-world incidents in our cybersecurity evaluations</span></a></strong></p><p><span>Anthropic discloses that a review of 141,006 cybersecurity evaluation runs surfaced three incidents where Claude models breached the real production infrastructure of three organizations via a misconfigured third-party eval environment.</span></p><p><strong><a href="https://pytorch.org/blog/fbtriton-infra-upstream-ingestion-hierarchical-validation-ideals-vs-realities/"><span>FBTriton Infra: Upstream Ingestion, Hierarchical Validation, Ideals vs Realities</span></a></strong></p><p><span>An engineering post on how Meta keeps its custom Triton fork in sync with upstream, using AI agents to sort incoming commits by risk and a three-tier testing system to catch regressions cheaply.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://www.decodingai.com/p/the-coding-agent-loop"><span>The Bare-Bones Coding Agent Loop</span></a></strong></p><p><span>A hands-on guide to building a coding agent from scratch, unpacking the agent loop, tool orchestration, and harness design that power systems like Claude Code and Codex.</span></p><p><strong><a href="https://research.google/blog/science-one-framework-a-verifiable-autonomous-research-framework-via-chain-of-evidence/"><span>Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence</span></a></strong></p><p><span>Google Research introduces the Science One Framework and CoE Audit, an autonomous research agent that hits zero phantom references and fully verifiable scores against baselines, hallucinating up to 21% of citations.</span></p><p><strong><a href="https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly"><span>SOTA alignment assessments don&#8217;t strongly update us against misalignment</span></a></strong></p><p><span>A critical analysis arguing that frontier labs&#8217; alignment assessments provide far weaker evidence against misalignment than their system cards claim, because the covert-capability evals underpinning those reliability arguments are themselves unreliable.</span></p><p><strong><a href="https://epoch.ai/publications/parallelization-constraints-could-delay-a-technological-singularity"><span>Even after R&amp;D is automated, parallelization constraints could delay a technological singularity</span></a></strong></p><p><span>An economic analysis arguing that intelligence-explosion models omit a key parameter &#8212; &#8220;parallelization technology,&#8221; the ceiling on how many researchers can work usefully at once &#8212; which could slow or flatten takeoff even after AI R&amp;D is fully automated.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/AMAP-ML/LongHorizon-Harness"><span>AMAP-ML/LongHorizon-Harness</span></a></strong></p><p><span>The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2607.23802"><span>From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs. Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a Self-PlaY Reinforcement Learning method inspired by social deduction game Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains.</span></p><p><strong><a href="https://arxiv.org/abs/2608.04003"><span>PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 466]]></title><description><![CDATA[Claude Opus 5, How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform, a paper on Progress Reward Modeling for Robotic Learning, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-466</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-466</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 30 Jul 2026 15:02:13 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://www.anthropic.com/news/claude-opus-5"><span>Claude Opus 5</span></a><span>, </span><a href="https://www.comet.com/site/blog/digibee-opik-user-story/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=digibee-opik-user-story/"><span>How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform</span></a><span> and </span><a href="https://arxiv.org/abs/2607.21655"><span>a paper on Progress Reward Modeling for Robotic Learning: A Comprehensive Survey</span></a><span>.</span></p><p><span>You may also enjoy </span><a href="https://bfl.ai/blog/flux-3"><span>FLUX 3</span></a><span>, </span><a href="https://www.philschmid.de/evocode-bench"><span>Evaluating Agents Beyond the First Prompt</span></a><span>, </span><a href="https://arxiv.org/abs/2607.22798"><span>a paper on StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://www.anthropic.com/news/claude-opus-5"><span>Introducing Claude Opus 5 \ Anthropic</span></a></strong></p><p><span>Anthropic releases Claude Opus 5, taking state-of-the-art on Frontier-Bench and GDPval-AA and surpassing Fable 5 on several evals at half the cost, with its lowest misalignment score yet.</span></p><p><strong><a href="https://bfl.ai/blog/flux-3"><span>FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.</span></a></strong></p><p><span>Black Forest Labs launches FLUX 3 in early access, a unified multimodal model trained jointly on image, video, and audio that generates 20-second clips with native audio and extends to robotic action prediction.</span></p><p><strong><a href="https://techcrunch.com/2026/07/28/fish-audio-raises-50m-seed-to-build-ai-voice-models-for-creators-and-enterprises/"><span>Fish Audio raises $52M seed to build AI voice models for creators and enterprises</span></a></strong></p><p><span>Fish Audio raised a $50M seed led by Coreline Ventures and Capital Today, reaching $21M ARR and 8 million users a year after launching its open-source voice models.</span></p><p><strong><a href="https://openai.com/index/health-in-chatgpt/"><span>Launching Health in ChatGPT</span></a></strong></p><p><span>OpenAI launches Health in ChatGPT for U.S. users, letting connected Apple Health and medical records ground conversations anywhere in the app, with that data excluded from model training and ad targeting.</span></p><p><strong><a href="https://venturebeat.com/infrastructure/microsoft-launches-new-in-house-ai-models-it-says-cut-costs-up-to-89-versus-openai"><span>Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI | VentureBeat</span></a></strong></p><p><span>Microsoft releases MAI-Image-2.5-Pro and MAI-Voice-2-Flash into public preview, displacing OpenAI models across Bing, PowerPoint, and Dynamics 365 with claimed GPU cost cuts of 84% and 89%.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/digibee-opik-user-story/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=digibee-opik-user-story/"><span>How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform</span></a></strong></p><p><span>A case study on how Digibee adopted Opik to bring evaluation-driven development into production, using traces and automated evaluations to catch regressions before they reach users.</span></p><p><strong><a href="https://jamwithai.substack.com/p/build-your-own-job-agent-part-1"><span>Build your own Job Agent - Part 1</span></a></strong></p><p><span>Part 1 of 3 of The Observable Job Agent series: A hands-on guide to building a LangGraph-powered job search agent that ranks real openings while instrumenting every step with Opik observability.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://bair.berkeley.edu/blog/2026/07/26/abbel/"><span>Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction</span></a></strong></p><p><span>A research post introducing ABBEL, which replaces recursive summarization with graded natural-language belief states, closing about half the gap to full-context agents on CollabBench in 50% fewer training steps.</span></p><p><strong><a href="https://www.philschmid.de/evocode-bench"><span>Evaluating Agents Beyond the First Prompt</span></a></strong></p><p><span>A breakdown of EvoCode-Bench, a 227-round multi-turn coding benchmark where pass rates fall from 46.7% at Round 1 to 7.7% by Round 10, with regressions rather than missing features driving failure.</span></p><p><strong><a href="https://embracethered.com/blog/posts/2026/ai-intrusion-are-now-real/"><span>Autonomous AI Intrusions Are Here: Lessons from the Hugging Face Compromise</span></a></strong></p><p><span>A defender-focused analysis of the Hugging Face breach, in which an autonomous agent logged over 17,000 attack actions and commercial model guardrails blocked forensic work until responders switched to a self-hosted open-weight model.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/NVIDIA-NeMo/labs-molt"><span>NVIDIA-NeMo/labs-molt</span></a></strong></p><p><span>An agentic-first RL framework for research.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2607.22798"><span>StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with program state by using code, while a dedicated GUI subagent handles screenshot-and-click interaction on the few subgoals that need it, just 28 of 108 tasks and 1.1% of main-agent steps. The same direct access to program state also supports verification: an independent finish gate double-checks the saved result for structural failures, e.g., output that is missing, unsaved, or written to the wrong path. To stay on track over hundreds of steps, the main agent hands subgoals to fresh subagents, keeping its own context focused. On OSWorld 2.0, StateAct lifts Claude Opus 4.8 from 20.6% to 26.9% on binary success, and from 54.8% to 61.6% on partial success, at ~ 9x lower cost per task than the same model driven by screenshots alone; a code-only variant with no GUI subagent reaches only 45.9% partial, below that screenshot-based baseline&#8217;s 54.8%. In general, grounding action, verification, and memory in state, what we call state-grounding, shifts the main bottleneck from perception toward reasoning: failures depend more on what the agent thinks than on what it sees.</span></p><p><strong><a href="https://arxiv.org/abs/2607.21655"><span>Progress Reward Modeling for Robotic Learning: A Comprehensive Survey</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 465]]></title><description><![CDATA[Kimi K3, Beyond the Single Trace: How We Built Agent Diagnostics for Opik, a paper on Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-465</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-465</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 24 Jul 2026 15:02:41 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://www.kimi.com/blog/kimi-k3"><span>Kimi K3</span></a><span>, </span><a href="https://www.comet.com/site/blog/debugging-ai-agents/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=debugging-ai-agents/"><span>Beyond the Single Trace: How We Built Agent Diagnostics for Opik</span></a><a href="https://epoch.ai/data-insights/ai-detectors-false-negatives"><span>s</span></a><span> and </span><a href="https://arxiv.org/abs/2511.00617"><span>a paper on Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering</span></a><span>.</span></p><p><span>You may also enjoy, </span><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"><span>Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber</span></a><span>, </span><a href="https://research.google/blog/symptomai-towards-a-conversational-ai-agent-for-everyday-symptom-assessment/"><span>SymptomAI: Towards a conversational AI agent for everyday symptom assessment</span></a><span>, </span><a href="https://arxiv.org/abs/2512.21577"><span>a paper on A Unified Definition of Hallucination: It&#8217;s The World Model, Stupid!</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://www.kimi.com/blog/kimi-k3"><span>Kimi K3 Tech Blog: Open Frontier Intelligence</span></a></strong></p><p><span>Moonshot AI launched Kimi K3, the world&#8217;s first open 3T-class model at 2.8 trillion parameters, with native multimodality and a 1-million-token context window.</span></p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"><span>Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber</span></a></strong></p><p><span>Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, with 3.6 Flash reducing output token usage by 17% compared to 3.5 Flash while improving coding and knowledge work performance.</span></p><p><strong><a href="https://openai.com/index/introducing-openai-presence/"><span>Introducing OpenAI Presence</span></a></strong></p><p><span>OpenAI launched Presence, an enterprise agent deployment product pairing model reasoning with policies and escalation rules, currently resolving 75% of inbound issues without human assistance on OpenAI&#8217;s own support line.</span></p><p><strong><a href="https://ir.amd.com/news-events/press-releases/detail/1292/amd-and-anthropic-announce-strategic-partnership-to-deploy-up-to-2-gigawatts-of-amd-instinct-mi450-series-gpus"><span>AMD and Anthropic Announce Strategic Partnership to Deploy Up to 2 Gigawatts of AMD Instinct MI450 Series GPUs</span></a></strong></p><p><span>AMD and Anthropic announced a partnership to deploy up to 2 gigawatts of MI450 Series GPUs in AMD Helios rack-scale solutions starting in the first half of 2027, paired with up to $5 billion in AMD equity investment in Anthropic.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/debugging-ai-agents/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=debugging-ai-agents/"><span>Beyond the Single Trace: How We Built Agent Diagnostics for Opik</span></a></strong></p><p><span>Comet built Diagnostics for Opik, an automated debugging agent that queries trace and span data in ClickHouse rather than reading traces one at a time, to surface silent agent failures like retry loops and over-deliberation at scale.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://medium.com/google-cloud/loop-engineering-in-self-correcting-code-migration-using-google-adk-2-0-61c30c9e36ca"><span>Loop Engineering in Self-Correcting Code Migration using Google ADK 2.0</span></a></strong></p><p><span>A technical guide detailing how Google ADK 2.0&#8217;s declarative workflow architecture eliminates quadratic token growth in self-correcting code migration loops through state-aware history pruning and decoupled validation.</span></p><p><strong><a href="https://epoch.ai/data-insights/ai-detectors-false-negatives"><span>AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors</span></a></strong></p><p><span>Epoch AI tested three detectors and found style-imitated AI text evaded detection far more often than plain AI text, with scientific writing failing to be detected ~26% of the time versus near-zero false negative rates on basic prompts.</span></p><p><strong><a href="https://research.google/blog/symptomai-towards-a-conversational-ai-agent-for-everyday-symptom-assessment/"><span>SymptomAI: Towards a conversational AI agent for everyday symptom assessment</span></a></strong></p><p><span>Google Research&#8217;s national-scale study found their Gemini Flash 2.0-based SymptomAI agent&#8217;s differential diagnoses were preferred over clinicians&#8217; own assessments in 53.3% of cases.</span></p><p><strong><a href="https://huggingface.co/blog/nunchaku-diffusers"><span>Bringing Nunchaku 4-bit Diffusion Inference to Diffusers</span></a></strong></p><p><span>A Hugging Face blog post detailing the native integration of Nunchaku&#8217;s SVDQuant W4A4 kernels into Diffusers, enabling 4-bit diffusion inference that cuts peak VRAM by up to 50% while boosting speed by 30-80%.</span></p><p><strong><a href="https://cloud.google.com/blog/topics/developers-practitioners/what-we-learned-about-agent-teamwork"><span>What we learned about agent teamwork</span></a></strong></p><p><span>An engineering blog about running 10 autonomous AI agent film crews on the open-source Scion orchestration testbed, surfacing patterns for multi-agent coordination outside coding domains.</span></p><p><strong><a href="https://venturebeat.com/security/evals-are-the-new-prd-expedia-ai-chief-tells-vb-transform-2026"><span>Evals are the new PRD, Expedia&#8217;s AI chief tells VB Transform 2026</span></a></strong></p><p><span>Expedia&#8217;s chief AI and data officer, Xavi Amatriain, states that evals now function as the new PRD, with product intent encoded directly into evaluation suites before coding begins.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/block/buzz"><span>block/buzz</span></a></strong></p><p><span>A workspace where humans and agents build together, on a relay you own.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2511.00617"><span>Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly disparate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation-based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interventions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena - e.g., sigmoidal learning curves as in-context evidence accumulates - while predicting novel ones - e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly changing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions.</span></p><p><strong><a href="https://arxiv.org/abs/2512.21577"><span>A Unified Definition of Hallucination: It&#8217;s The World Model, Stupid!</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today&#8217;s frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition wherein prior definitions are subsumed. We argue that hallucination can be unified by defining it as simply inaccurate (internal) world modeling, in a form where it is observable to the user. For example, stating a fact which contradicts a knowledge base OR producing a summary which contradicts the source. By varying the reference world model and conflict policy, our framework unifies prior definitions. We argue that this unified view is useful because it forces evaluations to clarify their assumed reference &#8220;world&#8221;, distinguishes true hallucinations from planning or reward errors, and provides a common language for comparison across benchmarks and discussion of mitigation strategies. Building on this definition, we also connect our framework to HalluWorld, a complementary benchmark that instantiates fully specified reference world models for stress-testing model hallucinations.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 464]]></title><description><![CDATA[Thinking Machines' Inkling, How We Optimized Opik&#8217;s MCP Server for Cost & Performance, Metacognition in LLMs: Foundations, Progress, and Opportunities, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-464</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-464</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 16 Jul 2026 15:01:04 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://thinkingmachines.ai/news/introducing-inkling/"><span>Thinking Machines&#8217; Inkling</span></a><span>, </span><a href="https://www.comet.com/site/blog/mcp-performance-optimization/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=mcp-performance-optimization/"><span>How We Optimized Opik&#8217;s MCP Server for Cost &amp; Performance</span></a><span> and </span><a href="https://arxiv.org/abs/2607.11881"><span>Metacognition in LLMs: Foundations, Progress, and Opportunities</span></a><span>.</span></p><p><span>You may also enjoy, </span><a href="https://openai.com/index/unlocking-self-improvement-gpt-red/"><span>GPT-Red: Unlocking Self-Improvement for Robustness</span></a><span>, </span><a href="https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt"><span>Model Routing Is Simple. Until It Isn&#8217;t.</span></a><span>, </span><a href="https://arxiv.org/abs/2607.09024"><span>Video Generation Models are General-Purpose Vision Learners</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://thinkingmachines.ai/news/introducing-inkling/"><span>Inkling: Our open-weights model - Thinking Machines Lab</span></a></strong></p><p><span>Thinking Machines Lab releases Inkling, an open-weights 975B-parameter multimodal MoE model with controllable reasoning effort, fine-tunable on Tinker.</span></p><p><strong><a href="https://venturebeat.com/technology/canva-launches-code-2-0-offering-ai-website-building-to-every-user-including-free-accounts"><span>Canva launches Code 2.0, offering AI website building to every user</span></a></strong></p><p><span>Canva launches Code 2.0 to all 265M monthly users, betting design polish&#8212;not code generation&#8212;is the real gap versus Lovable, Replit, and Bolt in vibe coding.</span></p><p><strong><a href="https://openai.com/index/unlocking-self-improvement-gpt-red/"><span>GPT-Red: Unlocking Self-Improvement for Robustness</span></a></strong></p><p><span>OpenAI trains GPT-Red, a self-play automated red-teaming model, to attack production systems and adversarially train GPT-5.6, cutting direct prompt injection failures 6x.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/mcp-performance-optimization/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=mcp-performance-optimization/"><span>How We Optimized Opik&#8217;s MCP Server for Cost &amp; Performance</span></a></strong></p><p><span>A look into how Comet rebuilt Opik&#8217;s MCP server from 30 tools down to four, using self-correcting schemas and adaptive response compression to cut token waste, improve tool selection, and keep agent context lean.</span></p><p><strong><a href="https://huggingface.co/blog/allenai/shippy-tech-blog"><span>What building Shippy taught us about building agents</span></a></strong></p><p><span>Ai2 details Shippy&#8217;s maritime agent architecture, showing that reliability came from a deterministic CLI layer, isolated per-session sandboxes, and rubric-based evals scoring the whole agent rather than the model alone.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt"><span>Model Routing Is Simple. Until It Isn&#8217;t.</span></a></strong></p><p><span>IBM Research argues LLM routing is a systems-optimization problem, not classification, after finding Claude Sonnet 4.6 cost half of GPT-4.1 per task due to caching effects despite higher sticker pricing.</span></p><p><strong><a href="https://www.pinecone.io/blog/text-match-filters/"><span>Text match filters for agents</span></a></strong></p><p><span>Pinecone launches lexical text match filters that scope semantic search to unstated query context, without pre-labeling metadata across the dataset.</span></p><p><strong><a href="https://epoch.ai/publications/ai-energy"><span>AI energy use: its impact on prices, climate, and more</span></a></strong></p><p><span>Epoch AI&#8217;s guide finds individual chatbot queries cost less energy than a microwave running 10 seconds, while global AI compute demand&#8212;tens of gigawatts&#8212;still doubles roughly every year.</span></p><p><strong><a href="https://research.google/blog/towards-demystifying-the-creativity-of-diffusion-models/"><span>Towards demystifying the creativity of diffusion models</span></a></strong></p><p><span>Google Research explains diffusion model creativity as a mathematical byproduct of neural network training, where regularization smooths the score function and drives interpolation between training points rather than memorization.</span></p><p><strong><a href="https://alignment.anthropic.com/2026/agentic-misalignment-summer-2026/"><span>Agentic Misalignment in Summer 2026</span></a></strong></p><p><span>Anthropic documents four new agentic misalignment failure modes in frontier models &#8212; covert sabotage, fraud assistance, motivated mislabeling, and coaching human whistleblowers &#8212; through controlled multi-model simulations.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/microsoft/ResearchStudio"><span>microsoft/ResearchStudio</span></a></strong></p><p><span>ResearchStudio: Our AI co-author, from research problem to final publication.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2607.09024"><span>Video Generation Models are General-Purpose Vision Learners</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world.</span></p><p><strong><a href="https://arxiv.org/abs/2607.11881"><span>Metacognition in LLMs: Foundations, Progress, and Opportunities</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs&#8217; metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 463]]></title><description><![CDATA[Introducing Grok 4.5, How Internal Optimizations Led to Comet Cost Intelligence, a paper on AlayaWorld: Long-Horizon and Playable Video World Generation, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-463</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-463</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 09 Jul 2026 15:00:58 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://x.ai/news/grok-4-5"><span>Introducing Grok 4.5</span></a><span>, </span><a href="https://www.comet.com/site/blog/ai-coding-cost-optimization/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-coding-cost-optimization/"><span>How Internal Optimizations Led to Comet Cost Intelligence</span></a><span>, and </span><a href="https://arxiv.org/abs/2607.06291"><span>a paper on AlayaWorld: Long-Horizon and Playable Video World Generation</span></a><span>.</span></p><p><span>You may also enjoy, </span><a href="https://openai.com/index/introducing-gpt-live/"><span>Introducing GPT-Live</span></a><span>, </span><a href="https://bair.berkeley.edu/blog/2026/07/07/intelligence-is-free-now-what/"><span>Intelligence is Free, Now What? Data Systems for, of, and by Agents</span></a><span>, </span><a href="https://openreview.net/pdf?id=kChzpoYFef"><span>a paper on Can LLMs Reason Structurally? Benchmarking via the lens of Data Structures</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://x.ai/news/grok-4-5"><span>Introducing Grok 4.5</span></a></strong></p><p><span>xAI launches Grok 4.5, its fastest and most token-efficient model yet (80 TPS, 4.2&#215; fewer output tokens than Opus 4.8 on SWE-Bench Pro), priced at $2/$6 per million tokens.</span></p><p><strong><a href="https://openai.com/index/introducing-gpt-live/"><span>Introducing GPT-Live</span></a></strong></p><p><span>OpenAI launches GPT-Live, a full-duplex voice model that listens and speaks simultaneously and delegates reasoning/search tasks to GPT-5.5 in the background, now powering ChatGPT Voice.</span></p><p><strong><a href="https://www.letta.com/blog/introducing-mods/"><span>Introducing Mods: Enabling Agents to Self-Improve through Harness-Level Adaptation</span></a></strong></p><p><span>Letta launches Mods for Letta Code, letting agents self-modify their own harness code rather than only learning through context, with a built-in learn/generate-env workflow.</span></p><p><strong><a href="https://mistral.ai/news/leanstral-1-5/"><span>Leanstral 1.5: Proof Abundance for All</span></a></strong></p><p><span>Mistral open-sources Leanstral 1.5, which saturates miniF2F, solves 587/672 PutnamBench problems, and sets new SOTA on FATE-H/X while surfacing 5 previously unknown bugs across 57 tested repositories.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://bair.berkeley.edu/blog/2026/07/07/intelligence-is-free-now-what/"><span>Intelligence is Free, Now What? Data Systems for, of, and by Agents</span></a></strong></p><p><span>UC Berkeley researchers argue near-zero inference costs will force data systems to be redesigned around agentic workloads &#8212; serving agent query swarms, sustaining multi-agent memory/coordination, and letting agents synthesize disposable custom systems.</span></p><p><strong><a href="https://huggingface.co/blog/nvidia/open-data-for-agents"><span>Data for Agents</span></a></strong></p><p><span>NVIDIA argues agentic AI&#8217;s bottleneck is open, inspectable synthetic data rather than model weights, releasing Nemotron&#8217;s over 10 trillion pre-training tokens and millions of post-training samples alongside a Prompt Atlas and locally-grounded persona datasets covering more than 2.4B people across ten countries.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/ai-coding-cost-optimization/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-coding-cost-optimization/"><span>Engineering Insights: How Internal Optimizations Led to Comet Cost Intelligence</span></a></strong></p><p><span>A look into how Comet&#8217;s engineering team built Cost Intelligence to expose where AI coding agents waste tokens and budget, turning opaque AI spend into actionable engineering metrics.</span></p><p><strong><a href="https://lilianweng.github.io/posts/2026-07-04-harness/"><span>Harness Engineering for Self-Improvement</span></a></strong></p><p><span>Lilian Weng&#8217;s post argues the &#8220;harness&#8221; wrapping a model is becoming as important as the model itself &#8212; and is now being optimized by the agents it contains.</span></p><p><strong><a href="https://www.anthropic.com/features/making-of-claude-code"><span>The Making of Claude Code</span></a></strong></p><p><span>Anthropic publishes a history of Claude Code&#8217;s origins, tracing its evolution from a 2022 internal coding assistant through &#8220;clide&#8221; to its February 2025 launch and shift to near-100% AI-written code by winter 2025.</span></p><p><strong><a href="https://openai.com/index/separating-signal-from-noise-coding-evaluations/"><span>Separating signal from noise in coding evaluations</span></a></strong></p><p><span>OpenAI&#8217;s audit finds ~30% of SWE-Bench Pro tasks are broken, retracting its own earlier recommendation to adopt the benchmark as a replacement for the flawed SWE-bench Verified.</span></p><p><strong><a href="https://alignment.anthropic.com/2026/modular-pretraining/"><span>Modular Pretraining Enables Access Control</span></a></strong></p><p><span>Anthropic and AE Studio introduce GRAM, a technique that isolates dual-use knowledge (virology, cybersecurity, nuclear physics) into switchable modules, letting one pretrained model approximate five separately data-filtered models at a fifth of the training compute.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/zilliztech/mfs/tree/main"><span>zilliztech/mfs</span></a></strong></p><p><span>A context harness for AI agents: all your scattered context &#8212; code, memory, docs, databases, SaaS &#8212; in one searchable, browsable, file-like interface.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2607.06291"><span>AlayaWorld: Long-Horizon and Playable Video World Generation</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than explicitly authoring every component of a virtual environment, these models autoregressively synthesize future observations conditioned on the current world state and user interactions, enabling playable worlds to be generated online. Trained on both gameplay recordings and real-world videos, they can capture diverse visual appearances and physical dynamics, opening new opportunities for interactive applications beyond gaming, including embodied intelligence. In this paper, we present \textbf{AlayaWorld}, a full-stack open-source framework for building interactive generative worlds. AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning. The framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-within a modular and extensible architecture. Alongside the framework, we release reproducible pipelines, reference implementations, evaluation tools, and comprehensive documentation, establishing a practical foundation for future research and real-time applications of generative world models.</span></p><p><strong><a href="https://openreview.net/pdf?id=kChzpoYFef"><span>Can LLMs Reason Structurally? Benchmarking via the lens of Data Structures</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Large language models (LLMs) are deployed on increasingly complex tasks that require multistep decision-making. Understanding their algorithmic reasoning abilities is therefore crucial. However, we lack a diagnostic benchmark for evaluating these capabilities. We propose to use data structures as a principled lens: as fundamental building blocks of algorithms, they naturally probe structural reasoning&#8212;the ability to understand and manipulate relationships such as order, hierarchy, and connectivity that underpin algorithmic reasoning. We introduce DSR-Bench (Data Structure Reasoning Benchmark), spanning 20 data structures, 35 operations, and 4,140 problem instances. DSR-Bench features hierarchical task organization, fully automated generation and evaluation, and fine-grained diagnostics. Evaluating 13 state-of-the-art LLMs reveals critical limitations: the top-performing model achieves only 0.46/1 on challenging instances. Three auxiliary probes targeting more realistic usages expose further weaknesses: models perform poorly on spatial data and context-rich scenarios, and they struggle to reason over their own code.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 462]]></title><description><![CDATA[GPT-5.6 Sol, Scaling Laws, a paper on Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-462</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-462</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 02 Jul 2026 18:00:36 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://openai.com/index/previewing-gpt-5-6-sol/"><span>Previewing GPT-5.6 Sol: a next-generation model</span></a><span>, </span><a href="https://lilianweng.github.io/posts/2026-06-24-scaling-laws/"><span>Scaling Laws, Carefully</span></a><span> and </span><a href="https://arxiv.org/abs/2606.06036"><span>a paper on Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents</span></a><span>.</span></p><p><span>You may also enjoy, </span><a href="https://www.anthropic.com/news/redeploying-fable-5"><span>Redeploying Claude Fable 5</span></a><span>, </span><a href="https://ai.stanford.edu/blog/rnb-encore/"><span>R&amp;B-EnCoRe: Self-Improving Pretraining of Embodied Reasoning Vision-Language-Action Models</span></a><span>, </span><a href="https://arxiv.org/abs/2606.23050"><span>a paper on Unlimited OCR Works</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://openai.com/index/previewing-gpt-5-6-sol/"><span>Previewing GPT-5.6 Sol: a next-generation model</span></a></strong></p><p><span>OpenAI begins a limited, government-coordinated preview of GPT-5.6 Sol, Terra, and Luna, its strongest cybersecurity model yet, paired with a layered safeguard stack and 700,000 GPU hours of automated red-teaming.</span></p><p><strong><a href="https://www.anthropic.com/news/redeploying-fable-5"><span>Redeploying Claude Fable 5</span></a></strong></p><p><span>Anthropic restores access to Claude Fable 5 and Mythos 5 after US export controls are lifted, rolling out with a strengthened cybersecurity classifier and a new cross-industry jailbreak severity framework.</span></p><p><strong><a href="https://www.anthropic.com/news/claude-sonnet-5"><span>Introducing Claude Sonnet 5</span></a></strong></p><p><span>Anthropic launches Claude Sonnet 5, its most agentic Sonnet model yet, narrowing the performance gap to Opus 4.8 on coding and tool use.</span></p><p><strong><a href="https://www.anthropic.com/news/claude-science-ai-workbench"><span>Claude Science, an AI workbench for scientists</span></a></strong></p><p><span>Anthropic launches Claude Science, an AI workbench that unifies research tools like PubMed, Jupyter, and cluster compute into one environment with over 60 scientific skills and auditable, reproducible outputs.</span></p><p><strong><a href="https://www.liquid.ai/blog/lfm2-5-230m"><span>LFM2.5-230M: Built to Run Anywhere</span></a></strong></p><p><span>Liquid AI releases LFM2.5-230M, its smallest open-weight model yet, delivering 213 tok/s on a Galaxy S25 Ultra and outperforming larger models on tool use and data extraction benchmarks.</span></p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/"><span>Start building with Nano Banana 2 Lite and Gemini Omni Flash</span></a></strong></p><p><span>Google launched Nano Banana 2 Lite, generating images in 4 seconds at $0.034 per 1K image, alongside Gemini Omni Flash for developer video generation and conversational editing at $0.10 per second.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/opik-oracle-agent-specification/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=opik-oracle-agent-specification/"><span>Opik + Oracle Agent Specification: Build Once, Run Anywhere</span></a></strong></p><p><span>Comet integrates Opik with Oracle&#8217;s Open Agent Specification, for defining agents once and trace, evaluate, and swap them across frameworks like LangGraph, AutoGen, and WayFlow without rebuilding.</span></p><p><strong><a href="https://www.comet.com/site/blog/ai-evaluation/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-evaluation/"><span>AI Evaluation Simplified: Automate Dataset &amp; Metric Eval Workflows with Test Suites</span></a></strong></p><p><span>A guide about Opik&#8217;s Test Suites feature, which replaces traditional dataset-and-metric AI evaluation workflows with plain-English assertions that return pass/fail results instead of raw scores.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://lilianweng.github.io/posts/2026-06-24-scaling-laws/"><span>Scaling Laws, Carefully</span></a></strong></p><p><span>A comprehensive explainer tracing how LLM scaling laws evolved from Kaplan to Chinchilla to data-constrained regimes, covering why fitting these power-law relationships is more fragile in practice than it looks.</span></p><p><strong><a href="https://ai.stanford.edu/blog/rnb-encore/"><span>R&amp;B-EnCoRe: Self-Improving Pretraining of Embodied Reasoning Vision-Language-Action Models</span></a></strong></p><p><span>A research blog post introducing R&amp;B-EnCoRe, a self-supervised pretraining cycle that lets vision-language-action models discover which reasoning steps actually improve action prediction, rather than following fixed templates.</span></p><p><strong><a href="https://openai.com/index/core-dump-epidemiology-data-infrastructure-bug/"><span>Core dump epidemiology: fixing an 18-year-old bug</span></a></strong></p><p><span>An engineering deep-dive about how OpenAI used population-level core dump analysis, rather than single-case debugging, to uncover two unrelated crash causes in its Rockset data infrastructure.</span></p><p><strong><a href="https://epoch.ai/MirrorCode"><span>MirrorCode: What&#8217;s the largest software project AI can complete on its own?</span></a></strong></p><p><span>Epoch AI, with METR, launches MirrorCode, a benchmark measuring whether AI can reimplement entire real-world programs end-to-end without source access, on which Claude Opus 4.7 leads at a 56% solve rate.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/radixark/miles"><span>radixark/miles</span></a></strong></p><p><span>An enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2606.06036"><span>Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a static retrieve-then-reason paradigm, this rigid pipeline design prevents them from dynamically adapting memory access to intermediate evidence discovered during inference. To bridge this gap, we propose MRAgent, a framework that combines an associative memory graph with an active reconstruction mechanism. We represent memory as a Cue-Tag-Content graph, where associative tags serve as semantic bridges connecting fine-grained cues to memory contents. Operating on this structure, our active reconstruction mechanism integrates LLM reasoning directly into memory access, allowing the agent to iteratively explore and prune retrieval paths based on accumulated evidence. This ensures that memory retrieval is dynamically adapted to the reasoning context while avoiding combinatorial explosion caused by unconstrained expansion. Experiments on the LoCoMo benchmark and LongMemEval benchmark demonstrate significant improvements over strong baselines (up to 23%), while substantially reducing token and runtime cost, highlighting the effectiveness of active and associative reconstruction for long-horizon memory reasoning.</span></p><p><strong><a href="https://arxiv.org/abs/2606.23050"><span>Unlimited OCR Works</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a large language model (LLM) as the decoder allows the model to leverage the prior distribution of language, leading to improved OCR performance. However, the downside is equally evident: as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. In this technical report, we propose Unlimited OCR, a model designed to emulate human parsing working memory. Taking DeepSeek OCR as the baseline, we replace all attention layers in the decoder with our proposed Reference Sliding Window Attention (R-SWA), which reduces attention computation costs while maintaining a constant KV cache throughout the entire decoding process. By combining the high compression rate of DeepSeek OCR&#8217;s encoder with our constant KV cache design, Unlimited OCR can transcribe dozens of pages of documents in a single forward pass under a standard maximum length of 32K. More importantly, R-SWA is a general-purpose parsing attention mechanism - beyond OCR, it is equally applicable to tasks such as ASR, translation, etc.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 461]]></title><description><![CDATA[Advanced Claude Code Cost Tracking: How to Save 30% on Token Spend, Introducing Claude Tag, a paper on Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-461</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-461</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 26 Jun 2026 15:02:06 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>This week in deep learning, we bring you </span><a href="https://www.comet.com/site/blog/claude-code-cost-tracker/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=claude-code-cost-tracker/"><span>Advanced Claude Code Cost Tracking: How to Save 30% on Token Spend </span></a><span>, </span><a href="https://www.anthropic.com/news/introducing-claude-tag"><span>Introducing Claude Tag</span></a><span>, and </span><a href="https://arxiv.org/abs/2603.09906"><span>a paper on Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs</span></a><span>.</span></p><p><span>You may also enjoy, </span><a href="https://ai.stanford.edu/blog/cot-monitoring-history/"><span>CoT Monitoring: Where Does a Hot Safety Problem Come From?</span></a><span>, </span><a href="https://blog.bytebytego.com/p/a-guide-to-ai-inference-engineering"><span>A Guide to Inference Engineering</span></a><span>, </span><a href="https://arxiv.org/abs/2606.24775"><span>a paper on Are We Ready For An Agent-Native Memory System?</span></a><span>, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://venturebeat.com/orchestration/xiaomis-harnessx-rewrites-its-own-ai-scaffolding-mid-task-and-smaller-models-gain-the-most"><span>Xiaomi&#8217;s HarnessX rewrites its own AI scaffolding mid-task &#8212; and smaller models gain the most</span></a></strong></p><p><span>Xiaomi researchers introduce HarnessX, a framework that treats an AI agent&#8217;s harness as a composable, self-evolving object, letting it autonomously rewrite its own scaffolding mid-task rather than relying on manual updates.</span></p><p><strong><a href="https://www.anthropic.com/news/introducing-claude-tag"><span>Introducing Claude Tag \ Anthropic</span></a></strong></p><p><span>Anthropic launches Claude Tag, a Slack-native @Claude teammate that learns channel context, works asynchronously, and now powers 65% of the product team&#8217;s code at Anthropic internally.</span></p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-computer-use-gemini-3-5-flash/"><span>Introducing computer use in Gemini 3.5 Flash</span></a></strong></p><p><span>Google makes computer use a native built-in tool in Gemini 3.5 Flash, consolidating its previously standalone screen-control model into the main Flash model with new enterprise prompt-injection safeguards.</span></p><p><strong><a href="https://qwen.ai/blog?id=qwen-agentworld"><span>Qwen-AgentWorld: Language World Models for General Agents</span></a></strong></p><p><span>Qwen releases Qwen-AgentWorld, a native language world model trained to simulate seven agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within one model, outperforming GPT-5.4 and Claude Opus 4.8 on the new AgentWorldBench evaluation.</span></p><p><strong><a href="https://mistral.ai/news/ocr-4/"><span>Mistral OCR 4 : SOTA OCR for Document Intelligence</span></a></strong></p><p><span>Mistral releases OCR 4, a self-hostable document-intelligence model adding bounding boxes, block classification, and confidence scores, beating competitors in human preference testing and topping OlmOCRBench.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/claude-code-cost-tracker/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=claude-code-cost-tracker/"><span>Advanced Claude Code Cost Tracking: How to Save 30% on Token Spend</span></a></strong></p><p><span>Comet introduces Cost Intelligence, a Claude Code and Codex cost tracker that attributes AI spend across developers, projects, tools, and workflows while surfacing configuration changes that reduce token usage.</span></p><p><strong><a href="https://pub.towardsai.net/version-controlling-your-agents-deployment-rollback-and-safe-promotion-patterns-6b7107dbe82a"><span>Version-Controlling Your Agents: Deployment, Rollback, and Safe Promotion Patterns</span></a></strong></p><p><span>A practical guide arguing that AI agents need the same versioning, staged promotion, and rollback discipline as traditional software, treating their configs as immutable, deployable &#8220;Agent-as-Code&#8221; artifacts.</span></p><p><strong><a href="https://x.com/akshay_pachaar/status/2069118430582866051"><span>Loop Engineering Explained</span></a></strong></p><p><span>A guide to loop engineering, an agent design pattern that replaces prompt-centric workflows with iterative execution loops that manage tool use, state, planning, and reflection.</span></p><p><strong><a href="https://blog.bytebytego.com/p/a-guide-to-ai-inference-engineering"><span>A Guide to AI Inference Engineering</span></a></strong></p><p><span>A guide explaining LLM inference engineering through the prefill/decode split, breaking down six core optimization techniques and when self-hosting beats off-the-shelf APIs.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://www.decodingai.com/p/5b766861-0001-494f-a37f-4d4eb104dcfa"><span>How Evaluation-Driven Development (EDD) Works</span></a></strong></p><p><span>Learn how retrieval, memory, and context management have emerged as critical infrastructure for AI agents, often having a greater impact on system performance than the underlying model itself.</span></p><p><strong><a href="https://ai.stanford.edu/blog/cot-monitoring-history/"><span>CoT Monitoring: Where Does a Hot Safety Problem Come From?</span></a></strong></p><p><span>A reflective essay tracing the intellectual lineage of chain-of-thought (CoT) monitoring, arguing it emerged from the convergence of ML monitoring practices and CoT-as-explainability research rather than as a standalone idea.</span></p><p><strong><a href="https://netflixtechblog.com/toward-more-controllable-ai-video-editing-an-early-research-exploration-at-netflix-eb8160ed60a2"><span>Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix</span></a></strong></p><p><span>A blog post about Netflix Research introducing Vera, a layered video diffusion model, and VOID, a physics-aware inpainting model, both aimed at giving artists more precise, controllable AI-assisted video editing.</span></p><p><strong><a href="https://blog.ml.cmu.edu/2026/06/19/healthcare-benchmarks-are-only-as-good-as-their-assumptions/"><span>Healthcare Benchmarks Are Only as Good as Their Assumptions</span></a></strong></p><p><span>A research blog post arguing that healthcare LLM benchmarks fail to predict real-world performance because they embed unstated assumptions about task structure and outcome measurement that break down at deployment.</span></p><p><strong><a href="https://eugeneyan.com/writing/cybersecurity-evals/"><span>Patterns for Building Cybersecurity Evals</span></a></strong></p><p><span>A guide breaking down cybersecurity evals into four shared primitives (sandboxed target, difficulty-tuning inputs, tools, grader), then walking through seven benchmarks from CTF-style exploitation to 50-host network compromise.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/XiaomiMiMo/MiMo-Code"><span>XiaomiMiMo/MiMo-Code</span></a></strong></p><p><span>MiMoCode is a terminal-native AI coding assistant.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2603.09906"><span>Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>While reasoning in LLMs plays a natural role in math, code generation, and multi-hop factual questions, its effect on simple, single-hop factual questions remains unclear. Such questions do not require step-by-step logical decomposition, making the utility of reasoning highly counterintuitive. Nevertheless, we find that enabling reasoning substantially expands the capability boundary of the model&#8217;s parametric knowledge recall, unlocking correct answers that are otherwise effectively unreachable. Why does reasoning aid parametric knowledge recall when there are no complex reasoning steps to be done? To answer this, we design a series of hypothesis-driven controlled experiments, and identify two key driving mechanisms: (1) a computational buffer effect, where the model uses the generated reasoning tokens to perform latent computation independent of their semantic content; and (2) factual priming, where generating topically related facts acts as a semantic bridge that facilitates correct answer retrieval. Importantly, this latter generative self-retrieval mechanism carries inherent risks: we demonstrate that hallucinating intermediate facts during reasoning increases the likelihood of hallucinations in the final answer. Finally, we show that our insights can be harnessed to directly improve model accuracy by prioritizing reasoning trajectories that contain hallucination-free factual statements.</span></p><p><strong><a href="https://arxiv.org/abs/2606.24775"><span>Are We Ready For An Agent-Native Memory System?</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a data management system that supports persistent information storage, retrieval, update, consolidation, and dynamic lifecycle governance throughout agent execution. Despite this evolution, existing evaluations still benchmark agent memory mainly through end-to-end task success metrics (e.g., F1, BLEU), while treating the underlying system as a monolithic black box. As a result, critical system-level concerns, including operational costs, architectural trade-offs across memory modules, and robustness under dynamic knowledge updates, remain insufficiently explored. In this paper, we present a systematic experimental study of agent memory from a data management perspective. We propose an analytical framework that decomposes agent memory into four core modules: memory representation and storage, extraction, retrieval and routing, and maintenance. Under this framework, we evaluate 12 representative memory systems and two reference baselines across five benchmark workloads spanning 11 datasets. Our extensive end-to-end evaluation shows that no single architecture dominates across all scenarios; instead, effectiveness depends heavily on how well the memory structure aligns with the workload bottleneck. Furthermore, through fine-grained ablation studies, we quantify their individual effects on representation fidelity, retrieval precision, update correctness, and long-horizon stability. Finally, we reveal cost-performance trade-offs under realistic workloads, showing localized maintenance is more cost-efficient than global reorganization. Based on these findings, we identify promising directions towards building truly agent-native memory systems.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 460]]></title><description><![CDATA[GLM 5.2, Understanding Your Claude Code Spend: What&#8217;s Actually Driving the Cost, a paper on Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-460</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-460</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 18 Jun 2026 15:02:21 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span data-color="rgb(40, 51, 56)" style="color: rgb(40, 51, 56);">This week in deep learning, we bring you </span><a href="https://z.ai/blog/glm-5.2"><span>GLM 5.2</span></a><span>, </span><a href="https://www.comet.com/site/blog/claude-code-context-bloat/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=claude-code-context-bloat/"><span>Understanding Your Claude Code Spend: What&#8217;s Actually Driving the Cost</span></a><span data-color="rgb(40, 51, 56)" style="color: rgb(40, 51, 56);"> and </span><a href="https://arxiv.org/abs/2606.11176"><span>a paper on Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories</span></a><span data-color="rgb(40, 51, 56)" style="color: rgb(40, 51, 56);">.</span></p><p><span>You may also enjoy </span><a href="https://www.kimi.com/resources/kimi-k2-7-code"><span>Kimi K2.7 Code</span></a><span>, </span><a href="https://ai.stanford.edu/blog/mstar/"><span>M*: A Modular, Extensible, Serving System for Multimodal Models</span></a><span data-color="rgb(40, 51, 56)" style="color: rgb(40, 51, 56);">, </span><a href="https://arxiv.org/abs/2606.14066"><span>a paper on FastContext: Training Efficient Repository Explorer for Coding Agents</span></a><span data-color="rgb(40, 51, 56)" style="color: rgb(40, 51, 56);">, and more!</span></p><p><span>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: </span><a href="https://twitter.com/dl_weekly"><span>@dl_weekly</span></a><span>.</span></p><p><span>Until next week!</span></p><div><hr></div><h2><strong><span>Industry</span></strong></h2><p><strong><a href="https://www.kimi.com/resources/kimi-k2-7-code"><span>Kimi K2.7 Code: Open-Source Agentic Coding Model</span></a></strong></p><p><span>Moonshot AI launches Kimi K2.7 Code, an open-source 1T-parameter MoE coding model gaining up to 31.5% on benchmarks while cutting thinking-token usage ~30% versus K2.6.</span></p><p><strong><a href="https://huggingface.co/blog/allenai/molmomotion"><span>MolmoMotion: Language-guided 3D motion forecasting</span></a></strong></p><p><span>Ai2 releases MolmoMotion, a language-guided model that forecasts object 3D point trajectories from video, beating prior methods on motion forecasting, robot planning, and video generation.</span></p><p><strong><a href="https://venturebeat.com/orchestration/stanfords-delm-cuts-multi-agent-task-costs-50-without-a-central-orchestrator"><span>Stanford&#8217;s DeLM cuts multi-agent task costs 50% &#8212; without a central orchestrator</span></a></strong></p><p><span>Stanford&#8217;s DeLM replaces central orchestrators with shared verified context, beating the strongest baseline by 10.5% on SWE-bench Verified at roughly half the cost.</span></p><p><strong><a href="https://z.ai/blog/glm-5.2"><span>GLM-5.2: Built for Long-Horizon Tasks</span></a></strong></p><p><span>Z.ai launches GLM-5.2, an open-source 753B-parameter MoE model with a stable 1M-token context, beating GLM-5.1 by wide margins on Terminal-Bench 2.1 and SWE-bench Pro while trailing Claude Opus 4.8 by just 1% on FrontierSWE.</span></p><h2><strong><span>MLOps/LLMOps/AgentOps</span></strong></h2><p><strong><a href="https://leehanchung.github.io/blogs/2026/06/13/hidden-technical-debt-agent-evaluation-infra/"><span>Hidden Technical Debt of AI Systems: Agent Evaluation Infrastructure</span></a></strong></p><p><span>A detailed guide arguing agent evaluation requires a control-plane/data-plane system spanning traces, state deltas, checkpoints, and replay &#8212; not just a single benchmark score.</span></p><p><strong><a href="https://ai.stanford.edu/blog/mstar/"><span>M*: A Modular, Extensible, Serving System for Multimodal Models</span></a></strong></p><p><span>Stanford&#8217;s M* replaces vLLM/SGLang&#8217;s single autoregressive loop with a generic &#8220;Walk Graph,&#8221; beating specialized serving systems by up to 2.7x on speech, 2.6x on image editing, and 12.5x on world-model rollouts.</span></p><h2><strong><span>Learning</span></strong></h2><p><strong><a href="https://www.comet.com/site/blog/claude-code-context-bloat/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=claude-code-context-bloat/"><span>Understanding Your Claude Code Spend: What&#8217;s Actually Driving the Cost</span></a></strong></p><p><span>Long and inefficient context can quietly drive up costs and reduce performance in Claude Code. Learn five techniques for auditing context usage and eliminating unnecessary token overhead.</span></p><p><strong><a href="https://www.decodingai.com/p/5b766861-0001-494f-a37f-4d4eb104dcfa"><span>How Evaluation-Driven Development (EDD) Works</span></a></strong></p><p><span>Learn how retrieval, memory, and context management have emerged as critical infrastructure for AI agents, often having a greater impact on system performance than the underlying model itself.</span></p><p><strong><a href="https://blog.ml.cmu.edu/2026/06/17/pre-training-isnt-bitter-enough/"><span>Pre-Training Isn&#8217;t Bitter Enough</span></a></strong></p><p><span>CMU proposes V-pretraining, which uses a small labeled feedback set to train a task designer that shapes self-supervised targets, lifting Qwen2.5-0.5B&#8217;s GSM8K Pass@1 from 22.20 to 29.60 without directly supervising the learner.</span></p><p><strong><a href="https://www.recursive.com/articles/first-steps-toward-automated-ai-research"><span>First Steps Toward Automated AI Research</span></a></strong></p><p><span>Recursive&#8217;s automated AI research system beats human-optimized SOTA on three benchmarks: 0.0263 lower BPB on fixed-budget LM training, 2.2s faster on NanoGPT Speedrun, and an 18% gap reduction on GPU kernel optimization.</span></p><p><strong><a href="https://jxnl.co/writing/2026/06/16/three-ways-codex-can-use-a-computer/#remote-control-lets-me-leave"><span>Three Ways Codex Can Use a Computer</span></a></strong></p><p><span>A practical guide on routing OpenAI Codex tasks across Computer Use, Chrome, and the in-app browser based on whether the job needs native apps, signed-in sites, or public-page review.</span></p><p><strong><a href="https://pytorch.org/blog/portable-vllm-model-inference-kernels-in-helion/"><span>Portable vLLM Model Inference Kernels in Helion</span></a></strong></p><p><span>Red Hat and Meta integrate Helion&#8217;s PyTorch-native kernel DSL into vLLM for FP8 Qwen3 inference, beating existing CUDA/TorchInductor kernels on most ops while still trailing CUTLASS on GEMM for Blackwell GPUs.</span></p><h2><strong><span>Libraries &amp; Code</span></strong></h2><p><strong><a href="https://github.com/comet-ml/opik"><span>comet-ml/opik</span></a></strong></p><p><span data-color="rgb(36, 41, 46)" style="color: rgb(36, 41, 46);">An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</span></p><p><strong><a href="https://github.com/openai/codex-plugin-cc"><span>openai/codex-plugin-cc</span></a></strong></p><p><span data-color="rgb(36, 41, 46)" style="color: rgb(36, 41, 46);">Use Codex from Claude Code to review code or delegate tasks.</span></p><h2><strong><span>Papers &amp; Publications</span></strong></h2><p><strong><a href="https://arxiv.org/abs/2606.11176"><span>Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Data tells stories that shape society; the data journalist&#8217;s job is to turn raw information into stories non-experts can trust. A high-quality news feature takes a newsroom team weeks: hunting for context, running statistics, choosing an angle, and designing visuals. Recent agents handle individual steps well: data-science agents close the analysis loop, while design agents synthesize beautiful websites. But can an agent serve as a data journalist end to end? We introduce Data Journalist Agent (Data2Story), a multi-agent framework that orchestrates specialized roles into a single virtual newsroom. Data2Story contributes two innovations. (i) Claims are evidence-grounded: an Inspector links every number, angle, and asset back to data, code, or an external reference. (ii) Articles are multimodally generative: rather than defaulting to plain text and static charts, Data2Story reasons about what readers will want to see, then deploys multimodal tools, such as interactive maps for geography and audio for music. We evaluate Data2Story on 18 articles, each paired with the originally published expert piece, along four axes: (a) human-agent angle coverage; (b) rubric evaluation with 53 participants across five dimensions; (c) computer-use agents as judges, a cost-saving proxy for how readers navigate interactive articles; and (d) verifiability, where a coding verifier re-executes statements against the data and checks claims against references. Data2Story produces competitive, evidence-traceable multimedia stories, with particular strength in transparency and auditability. Human articles retain an edge in editorial angle, creative design, and presentation. We position Data2Story as a collaborator for journalists, enabling more evidence-based, transparent, and verifiable reporting.</span></p><p><strong><a href="https://arxiv.org/abs/2606.14066"><span>FastContext: Training Efficient Repository Explorer for Coding Agents</span></a></strong></p><p><strong><span>Abstract:</span></strong></p><p><span>Large Language Model (LLM) coding agents have achieved strong results on software engineering tasks, yet repository exploration remains a major bottleneck: locating relevant code consumes substantial token budget and pollutes the agent&#8217;s context with irrelevant snippets. In most agents, the same model explores the repository and solves the task, leaving exploratory reads and searches in the solver&#8217;s history. We present FastContext, a dedicated exploration subagent that separates repository exploration from solving. Invoked on demand, FastContext issues parallel tool calls and returns concise file paths and line ranges as focused context. FastContext is powered by specialized exploration models spanning 4B--30B parameters. We bootstrap them from strong reference-model trajectories and refine them with task-grounded rewards for broad first-turn search, multi-turn evidence gathering, and precise citation generation. Across SWE-bench Multilingual, SWE-bench Pro, and SWE-QA, integrating FastContext into Mini-SWE-Agent improves end-to-end resolution rates up to 5.5% while reducing coding-agent token consumption up to 60%, with marginal overhead. These results show that repository exploration can be separated from solving and handled effectively by specialized models.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 459]]></title><description><![CDATA[Claude Fable 5, Cohere&#8217;s North Mini Code, a paper on Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-459</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-459</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 12 Jun 2026 15:00:51 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Claude Fable 5</a>, <a href="https://huggingface.co/blog/CohereLabs/introducing-north-mini-code">North Mini Code: Cohere&#8217;s First Model For Developers</a> and <a href="https://arxiv.org/abs/2601.14470">a paper on Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering</a>.</p><p>You may also enjoy <a href="https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/">NVIDIA Nemotron 3 Ultra</a>, <a href="https://epoch.ai/gradient-updates/controlling-the-capital-after-agi">Controlling the capital after AGI</a>, <a href="https://arxiv.org/abs/2602.07055">a paper on Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Claude Fable 5 and Claude Mythos 5 \ Anthropic</a></strong></p><p>Anthropic launches Claude Fable 5, a general-access Mythos-class model at $10/$50 per M tokens, with classifier-based fallbacks to Opus 4.8 for cyber, bio, and distillation queries.</p><p><strong><a href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/">Introducing Gemma 4 12B: a unified, encoder-free multimodal model</a></strong></p><p>Google releases Gemma 4 12B, an encoder-free multimodal model that runs on 16GB VRAM and processes vision and audio natively through the LLM backbone</p><p><strong><a href="https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/">NVIDIA Nemotron 3 Ultra</a></strong></p><p>NVIDIA releases Nemotron 3 Ultra &#8212; 550B total / 55B active MoE hybrid Mamba-Transformer, pretrained in NVFP4, with up to 5.9x throughput over competing open MoEs and 1M token context.</p><p><strong><a href="https://openai.com/index/openai-submits-confidential-s-1/">Confidential submission of draft S-1 to the SEC | OpenAI</a></strong></p><p>OpenAI files a confidential S-1 with the SEC, preemptively announcing it publicly ahead of an expected leak &#8212; while noting IPO timing remains undecided as some strategic moves are easier as a private company.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://cloud.google.com/blog/products/devops-sre/how-google-sre-is-using-agentic-ai-to-improve-operations">How Google SRE is using agentic AI to improve operations</a></strong></p><p>An article about how Google SRE is wiring agentic AI across the full incident lifecycle &#8212; dynamic anomaly detection, autonomous investigation, and a RAG layer over historical incidents to inform mitigation agents.</p><p><strong><a href="https://research.google/blog/unlocking-dependable-responses-with-gemini-enterprise-agent-platforms-agentic-rag/">Unlocking dependable responses with Gemini Enterprise Agent Platform&#8217;s Agentic RAG</a></strong></p><p>Google launches an agentic RAG framework on Gemini Enterprise Agent Platform with a Sufficient Context Agent that iterates until retrieval gaps are filled, hitting 90.1% accuracy on multi-hop queries &#8212; up to 34% over vanilla RAG.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://huggingface.co/blog/CohereLabs/introducing-north-mini-code">North Mini Code: Cohere&#8217;s First Model For Developers</a></strong></p><p>An article about how Cohere trained North Mini Code&#8217;s agentic coding capabilities &#8212; using cascaded SFT as an RLVR primer, joint multi-environment RL across 70k containerized repos, and cross-harness data mixing to generalize across SWE-Agent, mini-SWE-agent, and OpenCode scaffolds.</p><p><strong><a href="https://deepmindsafetyresearch.medium.com/testing-gemini-models-for-scheming-tendencies-3368c013ff16">Testing Gemini models for scheming tendencies | by DeepMind Safety Research</a></strong></p><p>DeepMind releases two scheming eval frameworks for Gemini &#8212; Gram (simulated agentic environments) and honeypots (real safety codebases) &#8212; finding 2&#8211;3% unprompted sabotage rates and no coherent misalignment.</p><p><strong><a href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety">Claude Fable 5 and new AI safety fables</a></strong></p><p>Nathan Lambert argues that Anthropic&#8217;s undisclosed safety filters in Claude Fable 5 &#8212; which silently degrade responses for frontier AI research without notifying users &#8212; are competitive entrenchment dressed as safety policy.</p><p><strong><a href="https://epoch.ai/gradient-updates/controlling-the-capital-after-agi">Controlling the capital after AGI</a></strong></p><p>An analytical piece from Epoch AI taxonomizing post-AGI wealth redistribution proposals &#8212; UBI, UBS, UBC, and sovereign wealth funds &#8212; along a single axis: how much control over capital, not just income, each scheme grants citizens.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2601.14470">Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering</a></strong></p><p><strong>Abstract:</strong></p><p>LLM-based Multi-Agent (LLM-MA) systems are increasingly applied to automate complex software engineering tasks such as requirements engineering, code generation, and testing. However, their operational efficiency and resource consumption remain poorly understood, hindering practical adoption due to unpredictable costs and environmental impact. To address this, we conduct an analysis of token consumption patterns in an LLM-MA system within the Software Development Life Cycle (SDLC), aiming to understand where tokens are consumed across distinct software engineering activities. We analyze execution traces from 30 software development tasks performed by the ChatDev framework using a GPT-5 reasoning model, mapping its internal phases to distinct development stages (Design, Coding, Code Completion, Code Review, Testing, and Documentation) to create a standardized evaluation framework. We then quantify and compare token distribution (input, output, reasoning) across these stages.</p><p>Our preliminary findings show that the iterative Code Review stage accounts for the majority of token consumption for an average of 59.4% of tokens. Furthermore, we observe that input tokens consistently constitute the largest share of consumption for an average of 53.9%, providing empirical evidence for potentially significant inefficiencies in agentic collaboration. Our results suggest that the primary cost of agentic software engineering lies not in initial code generation but in automated refinement and verification. Our novel methodology can help practitioners predict expenses and optimize workflows, and it directs future research toward developing more token-efficient agent collaboration protocols.</p><p><strong><a href="https://arxiv.org/abs/2602.07055">Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?</a></strong></p><p><strong>Abstract:</strong></p><p>Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploration remains understudied. We propose Theory of Space, defined as an agent&#8217;s ability to actively acquire information through self-directed, active exploration and to construct, revise, and exploit a spatial belief from sequential, partial observations. We evaluate this through a benchmark where the goal is curiosity-driven exploration to build an accurate cognitive map. A key innovation is spatial belief probing, which prompts models to reveal their internal spatial representations at each step. Our evaluation of state-of-the-art models reveals several critical bottlenecks. First, we identify an Active-Passive Gap, where performance drops significantly when agents must autonomously gather information. Second, we find high inefficiency, as models explore unsystematically compared to program-based proxies. Through belief probing, we diagnose that while perception is an initial bottleneck, global beliefs suffer from instability that causes spatial knowledge to degrade over time. Finally, using a false belief paradigm, we uncover Belief Inertia, where agents fail to update obsolete priors with new evidence. This issue is present in text-based agents but is particularly severe in vision-based models. Our findings suggest that current foundation models struggle to maintain coherent, revisable spatial beliefs during active exploration.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 458]]></title><description><![CDATA[Claude Opus 4.8, Agent Tracing and Observability: Log & Debug Complex AI Systems, a paper on A Self-Healing Framework for Reliable LLM-Based Autonomous Agents, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-458</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-458</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 04 Jun 2026 15:01:49 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://www.anthropic.com/news/claude-opus-4-8">Claude Opus 4.8</a>, <a href="https://www.comet.com/site/blog/ai-agent-tracing/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-agent-tracing/">Agent Tracing and Observability: Log &amp; Debug Complex AI Systems</a> and <a href="https://arxiv.org/abs/2605.06737">a paper on A Self-Healing Framework for Reliable LLM-Based Autonomous Agents</a>.</p><p>You may also enjoy <a href="https://www.minimax.io/blog/minimax-m3">Minimax M3</a>, <a href="https://huggingface.co/blog/Dharma-AI/direct-preference-optimization-beyond-chatbots">Direct Preference Optimization Beyond Chatbots</a>, <a href="https://arxiv.org/abs/2505.17117">a paper on From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://www.anthropic.com/news/claude-opus-4-8">Introducing Claude Opus 4.8 \ Anthropic</a></strong></p><p>Anthropic ships Claude Opus 4.8 at the same price as 4.7 &#8212; adding benchmark gains across coding and agentic tasks, a huge reduction in unremarked code flaws, and cheaper fast mode.</p><p><strong><a href="https://openai.com/index/codex-for-every-role-tool-workflow/">Codex for every role, tool, and workflow</a></strong></p><p>OpenAI expands Codex beyond developers with six role-specific plugins covering 62 apps and 110 skills, a preview of shareable hosted Sites, and inline annotations.</p><p><strong><a href="https://mistral.ai/news/search-toolkit/">Introducing Search Toolkit</a></strong></p><p>Mistral launches Search Toolkit, an open-source composable framework unifying ingestion, retrieval, and evaluation into a single production-ready pipeline for enterprise RAG and search applications.</p><p><strong><a href="https://openai.com/index/openai-frontier-models-and-codex-are-now-available-on-aws/">OpenAI frontier models and Codex are now available on AWS</a></strong></p><p>OpenAI makes its frontier models and Codex generally available on AWS via Amazon Bedrock &#8212; including GovCloud regions &#8212; letting enterprises adopt OpenAI through existing AWS security, compliance, and procurement workflows.</p><p><strong><a href="https://www.minimax.io/blog/minimax-m3">MiniMax M3: Frontier Coding, 1M Context, Native Multimodality &#8212; All in One Model</a></strong></p><p>MiniMax launches M3, currently the only open-weight model combining frontier coding performance, native multimodality, and a 1M-token context window via a new sparse attention architecture.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://www.comet.com/site/blog/ai-agent-tracing/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-agent-tracing/">Agent Tracing and Observability: Log &amp; Debug Complex AI Systems</a></strong></p><p>A guide on instrumenting agent tracing for multi-agent systems, covering why flat logging breaks at coordination boundaries, the three structural pillars of agentic observability, and how self-evolving agents compound debugging complexity.</p><p><strong><a href="https://cursor.com/blog/cloud-agent-lessons">What we&#8217;ve learned building cloud agents</a></strong></p><p>A Cursor engineering retrospective on a year of shipping cloud agents, arguing the work is less &#8220;local agent on a server&#8221; and more building a full operating layer &#8212; covering environment fidelity, durable execution via Temporal, etc.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://epoch.ai/data-insights/open-closed-eci-gap">Open models lag state-of-the-art closed models by 4 months</a></strong></p><p>An Epoch AI data insight measuring the open-to-closed model capability gap using their Epoch Capabilities Index (ECI), finding open-weight models now lag frontier closed models by an average of 4 months.</p><p><strong><a href="https://www.anthropic.com/news/AI-enabled-cyber-threats-mitre-attack">What we learned mapping a year&#8217;s worth of AI-enabled cyber threats \ Anthropic</a></strong></p><p>Anthropic analyzed 832 banned malicious accounts over one year, finding that AI is accelerating cyberattack sophistication &#8212; shifting from initial access tactics to post-compromise operations.</p><p><strong><a href="https://zilliz.com/blog/what-is-a-vector-lakebase">What Is a Vector Lakebase?</a></strong></p><p>An explainer introducing the Vector Lakebase &#8212; an architecture that unifies vector-database-grade serving with open lake storage and a shared semantic layer.</p><p><strong><a href="https://huggingface.co/blog/Dharma-AI/direct-preference-optimization-beyond-chatbots">Direct Preference Optimization Beyond Chatbots</a></strong></p><p>A blog post on applying Direct Preference Optimization to structured OCR &#8212; not for chat alignment &#8212; by using the SFT model&#8217;s own degeneration failures as rejection pairs.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2605.06737">A Self-Healing Framework for Reliable LLM-Based Autonomous Agents</a></strong></p><p><strong>Abstract:</strong></p><p>Autonomous agents based on Large Language Models (LLMs) are increasingly being utilized in complex software systems. However, reliability remains a significant challenge due to unpredictable failures such as hallucinations, execution errors, and inconsistent reasoning. This paper proposes a reliability-aware self-healing framework for LLM-based software agents. The framework integrates failure detection, reliability assessment, and automated recovery mechanisms. First, we define a taxonomy of failure types and introduce a quantitative reliability assessment model. Next, we propose a failure detection method that identifies abnormal agent behavior based on execution patterns and output consistency. Finally, we design a self-healing mechanism that dynamically recovers from failures through adaptive replanning and corrective prompting strategies. The proposed framework was implemented in a multi-agent workflow environment and evaluated using real-world task scenarios. Experimental results demonstrate that our approach significantly increases task success rates, reduces failure propagation, and enhances overall system robustness compared to existing methods. In particular, this study distinguishes itself by establishing an integrated monitoring system that combines the agent&#8217;s internal reasoning process with external execution results. These findings are expected to contribute to securing the stability of advanced autonomous systems and lowering the barriers to LLM adoption in production environments.</p><p><strong><a href="https://arxiv.org/abs/2505.17117">From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning</a></strong></p><p><strong>Abstract:</strong></p><p>Humans organize knowledge into compact conceptual categories that balance compression with semantic richness. Large Language Models (LLMs) exhibit impressive linguistic abilities, but whether they navigate this same compression-meaning trade-off remains unclear. We apply an Information Bottleneck framework to compare human conceptual structure with embeddings from 40+ LLMs using classic categorization benchmarks. We find that LLMs broadly align with human category boundaries, yet fall short on fine-grained semantic distinctions. Unlike humans, who maintain ``inefficient&#8217;&#8216; representations that preserve contextual nuance, LLMs aggressively compress, achieving more optimal information-theoretic compression at the cost of semantic richness. Surprisingly, encoder models outperform much larger decoder models in human alignment, suggesting that understanding and generation rely on distinct representational mechanisms. Training-dynamics analysis reveals a two-phase trajectory: rapid initial concept formation followed by architectural reorganization, during which semantic processing migrates from deep to mid-network layers as the model discovers increasingly efficient, sparser encodings. These divergent strategies, where LLMs optimize for compression and humans for adaptive utility, reveal fundamental differences between artificial and natural intelligence. This highlights the need for models that preserve the conceptual ``inefficiencies&#8217;&#8216; essential for human-like understanding.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 457]]></title><description><![CDATA[DeepSWE, The Best AI Observability Tools for Agentic Systems in 2026, a paper on SkillOpt: Executive Strategy for Self-Evolving Agent Skills, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-457</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-457</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Fri, 29 May 2026 15:02:20 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://deepswe.datacurve.ai/blog">DeepSWE</a>, <a href="https://www.comet.com/site/blog/ai-observability-tools/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-observability-tools/">The Best AI Observability Tools for Agentic Systems in 2026</a> and <a href="https://arxiv.org/abs/2605.23904">a paper on SkillOpt: Executive Strategy for Self-Evolving Agent Skills</a>.</p><p>You may also enjoy <a href="https://siliconangle.com/2026/05/19/google-reimagines-search-ai-agents-generative-interfaces/">Google reimagines search with AI agents and generative interfaces</a>, <a href="https://huggingface.co/blog/nvidia/nemotron-labs-diffusion#towards-speed-of-light-text-generation-with-nemotron-labs-diffusion-language-models">Towards Speed-of-Light Text Generation with Nemotron-Labs Diffusion Language Models</a>, <a href="https://arxiv.org/abs/2604.17609">a paper on Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://deepswe.datacurve.ai/blog">DeepSWE</a></strong></p><p>Datacurve introduces DeepSWE, a contamination-free coding benchmark of 113 from-scratch tasks across 91 repos and 5 languages, where GPT-5.5 leads at 70% and frontier models separate far more sharply than on SWE-Bench Pro.</p><p><strong><a href="https://siliconangle.com/2026/05/19/google-reimagines-search-ai-agents-generative-interfaces/">Google reimagines search with AI agents and generative interfaces</a></strong></p><p>Google overhauls Search at I/O 2026 with always-on Search Agents that monitor the web and report back, plus generative UI that builds interactive mini-apps on the fly via Antigravity and Gemini 3.5 Flash.</p><p><strong><a href="https://siliconangle.com/2026/05/26/openrouter-raises-113m-bring-order-enterprise-ai-inference-routing/">OpenRouter raises $113M to bring order to enterprise AI inference routing</a></strong></p><p>OpenRouter raises $113M Series B led by CapitalG to scale its multi-model inference routing platform.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://www.comet.com/site/blog/ai-observability-tools/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=ai-observability-tools/">The Best AI Observability Tools for Agentic Systems in 2026</a></strong></p><p>A guide to the top AI observability tools in 2026 for agentic systems. Learn which platforms are best for tracing, evaluation, debugging, testing, and monitoring in production.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://www.comet.com/site/blog/weavecli-opik-project-example/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=weavecli-opik-project-example/">What Held Up at 3 AM: One Engineer&#8217;s RAG Case Study</a></strong></p><p>Learn how WeaveCLI, a unified command-line tool for RAG over eleven vector databases was built, using Opik&#8217;s tracing capabilities, and configurable retrieval pipelines.</p><p><strong><a href="https://openai.com/index/building-self-improving-tax-agents-with-codex/">Building self-improving tax agents with Codex</a></strong></p><p>OpenAI and Thrive Holdings build Tax AI, a Codex-driven self-improving agent that turns repeated accountant corrections into bounded evals</p><p><strong><a href="https://huggingface.co/blog/nvidia/nemotron-labs-diffusion#towards-speed-of-light-text-generation-with-nemotron-labs-diffusion-language-models">Towards Speed-of-Light Text Generation with Nemotron-Labs Diffusion Language Models</a></strong></p><p>NVIDIA releases Nemotron-Labs Diffusion, an open model family (3B/8B/14B plus an 8B VLM) that combines autoregressive, diffusion, and self-speculation modes in one checkpoint</p><p><strong><a href="https://medium.com/@MongoDB/the-agent-harness-why-the-llm-is-the-smallest-part-of-your-agent-system-bce68414ccfd">The Agent Harness: Why the LLM Is the Smallest Part of Your Agent System</a></strong></p><p>A technical article arguing that the LLM is the smallest part of a production agent system, with the real engineering living in a six-component harness and a deeper platform layer that determines production reliability.</p><p><strong><a href="https://medium.com/pinterest-engineering/an-engineers-guide-to-better-ai-skills-implementing-a-testing-process-to-optimize-agent-a000c9c9abcd">An Engineer&#8217;s Guide to Better AI Skills: Implementing a Testing Process to Optimize Agent Performance in Any Repository or Skill</a></strong></p><p>A practical guide from Pinterest Engineering on building a test harness to measure and improve how reliably coding agents invoke custom skills through frontmatter tuning and other techniques.</p><p><strong><a href="https://huggingface.co/blog/agent-glossary">Harness, Scaffold, and the AI Agent Terms Worth Getting Right</a></strong></p><p>A glossary that pins down the agent vocabulary people keep using loosely &#8212; clarifying the scaffold-versus-harness distinction and grounding terms like context engineering, policy, skills, and more.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><p><strong><a href="https://github.com/mksglu/context-mode">mksglu/context-mode</a></strong></p><p>Context window optimization for AI coding agents.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2605.23904">SkillOpt: Executive Strategy for Self-Evolving Agent Skills</a></strong></p><p><strong>Abstract:</strong></p><p>Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization.</p><p><strong><a href="https://arxiv.org/abs/2604.17609">Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity</a></strong></p><p><strong>Abstract:</strong></p><p>LLM-based agents are assumed to integrate environmental observations into their reasoning: discovering highly relevant but unexpected information should naturally lead to a model exploiting its own discoveries. We show that this assumption is false for current LLM-based agents, which struggle to reflect or react to unexpected information. Across three benchmarks (Terminal-Bench, SWE-Bench, AppWorld), we inject complete task solutions into the agent environments to deliberately expose a task&#8217;s solution to a model. While agents discover these solutions on Terminal-Bench in 79-81% of runs, they interact, or exploit, them in only 37-50% of cases. This gap is starkest in AppWorld: agents see documentation stating that a command &#8220;returns the complete solution to this task&#8221; in over 90% of attempts but exploit this in fewer than 7% of trials. We show that agents lack what we call environmental curiosity: the capability to recognize and investigate unexpected but relevant observations in response to environmental stimuli. We identify three main factors influencing environmental curiosity: available tools in the agent scaffold, test-time compute, and training data distribution. Our findings identify configurations that maximize curiosity also achieve the best performance on the unmodified benchmarks. Yet even jointly optimized agents still ignore discovered solutions in the majority of trials: current agents use the environment to fetch expected information, but not to revise their strategy or maximally exploit useful stimuli.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 456]]></title><description><![CDATA[Gemini 3.5: frontier intelligence with action, Codex-maxxing, a paper on Lance: Unified Multimodal Modeling by Multi-Task Synergy, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-456</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-456</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 21 May 2026 15:03:46 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/#gemini-3-5-flash">Gemini 3.5: frontier intelligence with action</a>, <a href="https://jxnl.co/writing/2026/05/10/codex-maxxing/">Codex-maxxing</a> and <a href="https://arxiv.org/abs/2605.18678">a paper on Lance: Unified Multimodal Modeling by Multi-Task Synergy</a>.</p><p>You may also enjoy <a href="https://cohere.com/blog/command-a-plus">Introducing Command A+</a>, <a href="https://alignment.anthropic.com/2026/sleight-bench/">SLEIGHT-Bench: Finding Blind Spots in AI Monitors</a>, <a href="https://arxiv.org/abs/2605.12882">a paper on CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/#gemini-3-5-flash">Gemini 3.5: frontier intelligence with action</a></strong></p><p>Google launched Gemini 3.5, leading with 3.5 Flash&#8212;a model delivering flagship-tier agentic and coding performance at under half the cost.</p><p><strong><a href="https://cohere.com/blog/command-a-plus">Introducing Command A+</a></strong></p><p>Cohere released Command A+ open-source &#8212;a 218B/25B-active MoE model for enterprise agentic workflows that runs on as little as two H100s or one Blackwell GPU, supports 48 languages, and adds multimodal reasoning.</p><p><strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/">Introducing Gemini Omni</a></strong></p><p>Google introduced Gemini Omni, a natively multimodal generation model debuting with Omni Flash&#8212;it creates and conversationally edits video from any combination of image, audio, video, and text inputs, rolling out across the Gemini app, Flow, and YouTube Shorts.</p><p><strong><a href="https://x.ai/news/grok-build-cli">Introducing Grok Build</a></strong></p><p>xAI launched Grok Build &#8212; a coding CLI powered by Grok 4.3 Heavy, featuring a 2M-token context window, 8 parallel subagents, and more.</p><p><strong><a href="https://bfl.ai/blog/outpainting-extend-any-image-in-any-direction">FLUX Outpainting: Extend any image, in any direction</a></strong></p><p>Black Forest Labs launched FLUX Outpainting, a purpose-built API endpoint that extends images in any direction without prompts.</p><p><strong><a href="https://cursor.com/blog/composer-2-5">Introducing Composer 2.5</a></strong></p><p>Cursor released Composer 2.5, a coding model (built on Moonshot&#8217;s Kimi K2.5) with gains on long-horizon agentic tasks.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://www.comet.com/site/blog/llm-cost-tracking-solution/?utm_source=substack&amp;utm_medium=email&amp;utm_campaign=dlw&amp;utm_content=llm-cost-tracking-solution/">LLM Cost Tracking Solution: How to Monitor and Control AI Spend in Agentic Systems</a></strong></p><p>A guide on treating LLM cost as an observability problem in agentic systems, using span/trace/project-level tracing to pinpoint token-burning prompts and routing.</p><p><strong><a href="https://redis.io/blog/context-is-all-you-need/">Context is all you need: Introducing Redis Iris</a></strong></p><p>Redis launched Redis Iris, a context engine sitting between agents and enterprise data&#8212;bundling five tools (two new: Context Retriever and Agent Memory) to deliver navigable, fresh, low-latency context with semantic caching that cuts token costs up to 90%.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://weaviate.io/blog/tokenization-text-analysis-weaviate">Text Analysis for Hybrid Search: Tokenization, Stopwords &amp; Accent Folding</a></strong></p><p>A technical guide on how Weaviate v1.37 makes BM25 tokenization observable and per-property configurable&#8212;covering accent folding, per-language stopwords, and more.</p><p><strong><a href="https://pytorch.org/blog/running-pytorch-models-on-apple-silicon-gpus-with-the-executorch-mlx-delegate/">Running PyTorch Models on Apple Silicon GPUs with the ExecuTorch MLX Delegate</a></strong></p><p>PyTorch released the experimental ExecuTorch MLX delegate, a backend that runs PyTorch models on Apple Silicon GPUs via Apple&#8217;s MLX framework</p><p><strong><a href="https://jxnl.co/writing/2026/05/10/codex-maxxing/">Codex-maxxing</a></strong></p><p>A power user&#8217;s playbook for extracting more value from Codex&#8212;using durable threads, file-based memory, verifiable goals, and self-scheduling loops to turn it into a workspace where long-running knowledge work keeps progressing between sessions.</p><p><strong><a href="https://alignment.anthropic.com/2026/sleight-bench/">SLEIGHT-Bench: Finding Blind Spots in AI Monitors</a></strong></p><p>Anthropic researchers released SLEIGHT-Bench, a benchmark of 40 synthetic attacks across 11 categories that exploit &#8220;blind spots&#8221; in frontier AI monitors&#8212;on the Opus 4.6 monitor, 50% of attacks evaded all 10 trials and only 8 of 40 were reliably caught.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><p><strong><a href="https://github.com/generalaction/emdash">generalaction/emdash</a></strong></p><p>Emdash is the Open-Source Agentic Development Environment. Run multiple coding agents in parallel. Use any provider.</p><h2><strong>Papers &amp; Publications</strong></h2><p> <strong><a href="https://arxiv.org/abs/2605.18678">Lance: Unified Multimodal Modeling by Multi-Task Synergy</a></strong></p><p><strong>Abstract:</strong></p><p>We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, Lance explores a practical paradigm for unified multimodal modeling via collaborative multi-task training. It is grounded in two core principles: unified context modeling and decoupled capability pathways. Specifically, Lance is trained from scratch and employs a dual-stream mixture-of-experts architecture on shared interleaved multimodal sequences, enabling joint context learning while decoupling the pathways for understanding and generation. We further introduce modality-aware rotary positional encoding to mitigate interference among heterogeneous visual tokens and boost cross-task alignment. During training, Lance adopts a staged multi-task training paradigm with capability-oriented objectives and adaptive data scheduling to strengthen both semantic comprehension and visual generation performance. Experimental results demonstrate that Lance substantially outperforms existing open-source unified models in image and video generation, while retaining strong multimodal understanding capabilities.</p><p><strong><a href="https://arxiv.org/abs/2605.12882">CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence</a></strong></p><p><strong>Abstract:</strong></p><p>Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only approach masks a critical failure mode: a model can land on the correct answer while grounding it in the wrong passage -- a critical risk in high-stakes domains like law, finance, and medicine, where every conclusion must be traceable to a specific source region. To address this, we introduce CiteVQA, a benchmark that requires models to return element-level bounding-box citations alongside each answer, evaluating both jointly. CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, averaging 40.6 pages per document. To ensure fidelity and scalability, the ground-truth citations are generated by an automated pipeline-which identifies crucial evidence via masking ablation-and are subsequently validated through expert review. At the core of our evaluation is Strict Attributed Accuracy (SAA), which credits a prediction only when the answer and the cited region are both correct. Auditing 20 MLLMs reveals a pervasive Attribution Hallucination: models frequently produce the right answer while citing the wrong region. The strongest system (Gemini-3.1-Pro-Preview) achieves an SAA of only 76.0, and the strongest open-source MLLM reaches just 22.5. Ultimately, towards trustworthy document intelligence, CiteVQA exposes a reliability gap that answer-only evaluations overlook, providing the instrumentation needed to close it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 455]]></title><description><![CDATA[Interaction Models: A Scalable Approach to Human-AI Collaboration, Hidden Technical Debt of AI Systems: Agent Harness, a paper on Efficient Online Memory for Large Language Models, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-455</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-455</guid><dc:creator><![CDATA[Deep Learning Weekly]]></dc:creator><pubDate>Thu, 14 May 2026 15:02:11 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://thinkingmachines.ai/blog/interaction-models/">Interaction Models: A Scalable Approach to Human-AI Collaboration</a>, <a href="https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/">Hidden Technical Debt of AI Systems: Agent Harness</a> and <a href="https://arxiv.org/abs/2605.12357">a paper on &#948;-mem: Efficient Online Memory for Large Language Models</a>.</p><p>You may also enjoy <a href="https://www.perceptron.inc/blog/introducing-perceptron-mk1">Introducing Perceptron Mk1</a>, <a href="https://www.anthropic.com/research/teaching-claude-why">Teaching Claude why</a>, <a href="https://arxiv.org/abs/2605.03546">a paper on ProgramBench: Can Language Models Rebuild Programs From Scratch?</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://www.perceptron.inc/blog/introducing-perceptron-mk1">Introducing Perceptron Mk1</a></strong></p><p>Perceptron AI launches Mk1, a video and embodied-reasoning vision-language model priced roughly 80&#8211;90% cheaper than Claude Sonnet 4.5, GPT-5, and Gemini 3.1 Pro.</p><p><strong><a href="https://techcrunch.com/2026/05/13/notion-just-turned-its-workspace-into-a-hub-for-ai-agents/">Notion just turned its workspace into a hub for AI agents</a></strong></p><p>Notion launches Developer Platform turning its workspace into an agent orchestration hub with custom code Workers, external database sync, and native integrations for Claude Code, Cursor, Codex, and Decagon.</p><p><strong><a href="https://thinkingmachines.ai/blog/interaction-models/">Interaction Models: A Scalable Approach to Human-AI Collaboration</a></strong></p><p>Thinking Machines unveils TML-Interaction-Small, a 276B MoE (12B active) interaction model trained from scratch with 200ms time-aligned micro-turns that natively handles concurrent audio, video, and text without VAD-style harnesses.</p><p><strong><a href="https://unsloth.ai/blog/pytorch">Unsloth Joins the PyTorch Ecosystem</a></strong></p><p>Unsloth joins the PyTorch Ecosystem Landscape, recognizing its open-source contributions including 2&#215; faster training with 70% less VRAM, FP8 RL for consumer GPUs, and 250M+ model downloads.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/">Hidden Technical Debt of AI Systems: Agent Harness</a></strong></p><p>A Hanchung Lee essay reframing Sculley&#8217;s 2015 ML technical debt diagram for the agent era, arguing the agent runtime (harness + state) &#8212; not the model &#8212; is where most spend, incidents, and architectural debt are now accumulating.</p><p><strong><a href="https://huggingface.co/blog/amazon/foundation-model-building-blocks">Building Blocks for Foundation Model Training and Inference on AWS</a></strong></p><p>A reference guide from Amazon mapping AWS&#8217;s four-layer infrastructure stack to foundation model pre-training, post-training, and inference workloads.</p><p><strong><a href="https://developer.nvidia.com/blog/how-to-eliminate-pipeline-friction-in-ai-model-serving/">How to Eliminate Pipeline Friction in AI Model Serving</a></strong></p><p>A practical NVIDIA guide laying out 18 best practices to eliminate AI model-serving friction across export issues, unsupported ops, dynamic input shapes, and version mismatches</p><h2><strong>Learning</strong></h2><p><strong><a href="https://www.anthropic.com/research/teaching-claude-why">Teaching Claude why</a></strong></p><p>An Anthropic post detailing how teaching Claude why actions are aligned &#8212; via constitutional documents and ethical reasoning, not just demonstrations &#8212; drove blackmail rates from 96% (Opus 4) to 0% on every Claude model since Haiku 4.5.</p><p><strong><a href="https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/">Accelerating Gemma 4: faster inference with multi-token prediction drafters</a></strong></p><p>Google releases Multi-Token Prediction (MTP) drafters for Gemma 4 models, delivering up to 3x faster inference via speculative decoding with zero quality degradation.</p><p><strong><a href="https://simonw.substack.com/p/vibe-coding-and-agentic-engineering">Vibe coding and agentic engineering are getting closer than I&#8217;d like</a></strong></p><p>A post by Simon Willison observing that vibe coding and agentic engineering are converging in his own workflow as he increasingly ships production code from Claude Code without reviewing every line.</p><p><strong><a href="https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing">How fast is autonomous AI cyber capability advancing?</a></strong></p><p>UK AISI reports the length of cyber tasks frontier models can autonomously complete is doubling every 4.7 months &#8212; accelerating from 8 months last November &#8212; with Claude Mythos Preview and GPT-5.5 exceeding even that trend.</p><p><strong><a href="https://deepmind.google/blog/ai-pointer/">Reimagining the mouse pointer for the AI era</a></strong></p><p>A design-principles post from Google DeepMind reframing the mouse pointer as a Gemini-powered context-aware partner, built on four principles: maintain the flow, show and tell, embrace &#8220;this/that&#8221; deixis, and turn pixels into actionable entities.</p><p><strong><a href="https://www.pinecone.io/blog/full-text-search-architecture/">Full Text Search: Architecture and Design</a></strong></p><p>A technical architecture post from Pinecone introducing full-text search built on Tantivy, delivering Lucene query syntax, BM25 scoring, 18-language tokenization, and 22.7ms p50 latency on 6.4M Wikipedia articles.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2605.12357">&#948;-mem: Efficient Online Memory for Large Language Models</a></strong></p><p><strong>Abstract:</strong></p><p>Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems. Simply expanding the context window is costly and often fails to ensure effective context utilization. We propose &#948;-mem, a lightweight memory mechanism that augments a frozen full-attention backbone with a compact online state of associative memory. &#948;-mem compresses past information into a fixed-size state matrix updated by delta-rule learning, and uses its readout to generate low-rank corrections to the backbone&#8217;s attention computation during generation. With only an 8&#215;8 online memory state, &#948;-mem improves the average score to 1.10&#215; that of the frozen backbone and 1.15&#215; that of the strongest non-&#948;-mem memory baseline. It achieves larger gains on memory-heavy benchmarks, reaching 1.31&#215; on MemoryAgentBench and 1.20&#215; on LoCoMo, while largely preserving general capabilities. These results show that effective memory can be realized through a compact online state directly coupled with attention computation, without full fine-tuning, backbone replacement, or explicit context extension.</p><p><strong><a href="https://arxiv.org/abs/2605.03546">ProgramBench: Can Language Models Rebuild Programs From Scratch?</a></strong></p><p><strong>Abstract:</strong></p><p>Turning ideas into full software projects from scratch has become a popular use case for language models. Agents are being deployed to seed, maintain, and grow codebases over extended periods with minimal human oversight. Such settings require models to make high-level software architecture decisions. However, existing benchmarks measure focused, limited tasks such as fixing a single bug or developing a single, specified feature. We therefore introduce ProgramBench to measure the ability of software engineering agents to develop software holisitically. In ProgramBench, given only a program and its documentation, agents must architect and implement a codebase that matches the reference executable&#8217;s behavior. End-to-end behavioral tests are generated via agent-driven fuzzing, enabling evaluation without prescribing implementation structure. Our 200 tasks range from compact CLI tools to widely used software such as FFmpeg, SQLite, and the PHP interpreter. We evaluate 9 LMs and find that none fully resolve any task, with the best model passing 95\% of tests on only 3\% of tasks. Models favor monolithic, single-file implementations that diverge sharply from human-written code.</p>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 454]]></title><description><![CDATA[MolmoAct 2: An open foundation for robots, How to Work and Compound with AI, a paper on ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-454</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-454</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 07 May 2026 15:03:11 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://allenai.org/blog/molmoact2">MolmoAct 2: An open foundation for robots</a>, <a href="https://eugeneyan.com/writing/working-with-ai/">How to Work and Compound with AI</a> and <a href="https://arxiv.org/abs/2605.03042">a paper on ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration</a>.</p><p>You may also enjoy <a href="https://openai.com/index/gpt-5-5-instant/">GPT-5.5 Instant</a>, <a href="https://epoch.ai/blog/chips-topic-overview">AI Chips: why they cost as much as a car, and why companies can&#8217;t get enough</a>, <a href="https://arxiv.org/abs/2605.04036">a paper on OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://allenai.org/blog/molmoact2">MolmoAct 2: An open foundation for robots that work in the real world</a></strong></p><p>Ai2 releases MolmoAct 2, a fully open robotics foundation model running up to 37x faster than its predecessor, alongside the largest open-source bimanual manipulation dataset.</p><p><strong><a href="https://openai.com/index/gpt-5-5-instant/">GPT-5.5 Instant: smarter, clearer, and more personalized</a></strong></p><p>OpenAI releases GPT-5.5 Instant as ChatGPT&#8217;s new default model, cutting hallucinations by 52.5% on high-stakes prompts and using 30% fewer words while pulling context from past chats, files, and more.</p><p><strong><a href="https://x.ai/news/anthropic-compute-partnership">New Compute Partnership with Anthropic</a></strong></p><p>SpaceXAI signs a deal with rival Anthropic for full access to Colossus 1, with Anthropic also expressing interest in jointly developing multi-gigawatt orbital compute.</p><p><strong><a href="https://siliconangle.com/2026/05/05/blitzy-raises-200m-1-4b-valuation-deploy-thousands-coding-agents-parallel/">Blitzy raises $200M at $1.4B valuation to deploy thousands of coding agents in parallel</a></strong></p><p>Blitzy raises $200M at $1.4B valuation to scale its enterprise platform that orchestrates thousands of parallel coding agents across 100M+ line legacy codebases, scoring 66.5% on SWE-Bench Pro.</p><p><strong><a href="https://siliconangle.com/2026/05/06/monday-com-relaunches-ai-work-platform-native-agents/">Monday.com relaunches as an AI work platform with native agents</a></strong></p><p>Monday.com relaunches as an &#8220;AI work platform&#8221; with native agents that draft campaigns, qualify leads, and triage tickets across its 250,000+ customers, plus one-click connectors to Claude, ChatGPT, Copilot, and Gemini.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://www.comet.com/site/blog/end-to-end-agent-testing/">Introducing the Opik Agent Playground</a></strong></p><p>Comet launches Opik Agent Playground, a UI-based environment for testing and tweaking full agent configurations (prompts, models, tools) without touching code, opening iteration to PMs and domain experts.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://epoch.ai/blog/chips-topic-overview">AI Chips: why they cost as much as a car, and why companies can&#8217;t get enough</a></strong></p><p>A primer on AI chip economics and supply: the entire frontier flows through TSMC and a handful of designers, with total compute capacity doubling every 7 months while performance-per-dollar doubles every 2.5 years.</p><p><strong><a href="https://www.philschmid.de/subagent-patterns-2026">How Agents Manage Other Agents: Four Subagents Patterns in 2026</a></strong></p><p>A practical blog post about four subagent orchestration patterns&#8212;inline tool, fan-out, agent pool, and teams&#8212;each requiring progressively more capable models and offering different tradeoffs in control, lifetime, and result collection.</p><p><strong><a href="https://leehanchung.github.io/blogs/2026/05/01/dont-outsource-your-understanding/">Don&#8217;t Outsource Your Understanding</a></strong></p><p>An essay arguing the real AI risk isn&#8217;t using it but &#8220;cognitive surrender&#8221;&#8212;outsourcing verification too&#8212;evidenced by 1,300+ hallucinated court filings and 50 ICLR papers with fake citations.</p><p><strong><a href="https://huggingface.co/blog/ibm-granite/granite-4-1">Granite 4.1 LLMs: How They&#8217;re Built</a></strong></p><p>An technical deep-dive on how Granite 4.1 was built: 15T-token, five-phase pretraining with long-context extension to 512K, SFT on 4.1M curated samples, and on-policy GRPO with DAPO loss.</p><p><strong><a href="https://eugeneyan.com/writing/working-with-ai/">How to Work and Compound with AI</a></strong></p><p>A practitioner&#8217;s playbook for compounding with AI: treat context as infra, encode taste as config (CLAUDE.md, skills), make verification cheap, delegate bigger chunks in parallel, and mine transcripts to close the loop.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><p><strong><a href="https://github.com/zilliztech/claude-context">zilliztech/claude-context</a></strong></p><p>Claude Context is an MCP plugin that adds semantic code search to Claude Code and other AI coding agents, giving them deep context from your entire codebase.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2605.03042">ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration</a></strong></p><p><strong>Abstract:</strong></p><p>This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The performance of agent systems built on LLMs depends on both the model weights and the harness around them, which governs what information to store, retrieve, and present to the model. For long-horizon research workflows, the central failure mode is not a visible breakdown but a plausible unsupported success: a long-running agent can produce claims whose evidential support is incomplete, misreported, or silently inherited from the executor&#8217;s framing. Therefore, we present ARIS as a research harness that coordinates machine-learning research workflows through cross-model adversarial collaboration as a default configuration: an executor model drives forward progress while a reviewer from a different model family is recommended to critique intermediate artifacts and request revisions. ARIS has three architectural layers. The execution layer provides more than 65 reusable Markdown-defined skills, model integrations via MCP, a persistent research wiki for iterative reuse of prior findings, and deterministic figure generation. The orchestration layer coordinates five end-to-end workflows with adjustable effort settings and configurable routing to reviewer models. The assurance layer includes a three-stage process for checking whether experimental claims are supported by evidence: integrity verification, result-to-claim mapping, and claim auditing that cross-checks manuscript statements against the claim ledger and raw evidence, as well as a five-pass scientific-editing pipeline, mathematical-proof checks, and visual inspection of the rendered PDF. A prototype self-improvement loop records research traces and proposes harness improvements that are adopted only after reviewer approval.</p><p><strong><a href="https://arxiv.org/abs/2605.04036">OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories</a></strong></p><p><strong>Abstract:</strong></p><p>Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continual pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). In this report, we show that when fueled with informative and high-difficulty trajectories, a simple SFT approach could be surprisingly powerful for training frontier search agents. By introducing three simple data synthesis modifications: scaling knowledge graph size for richer exploration, expanding the tool set size for broader functionality, and strict low-step filtering, we establish a stronger baseline. Trained on merely 10.6k data points, our OpenSeeker-v2 achieves state-of-the-art performance across 4 benchmarks (30B-sized agents with ReAct paradigm): 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on Humanity&#8217;s Last Exam, and 78.0% on xbench, surpassing even Tongyi DeepResearch trained with heavy CPT+SFT+RL pipeline, which achieves 43.4%, 46.7%, 32.9%, and 75.0%, respectively. Notably, OpenSeeker-v2 represents the first state-of-the-art search agent within its model scale and paradigm to be developed by a purely academic team using only SFT. We are excited to open-source the OpenSeeker-v2 model weights and share our simple yet effective findings to make frontier search agent research more accessible to the community.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Deep Learning Weekly: Issue 453]]></title><description><![CDATA[OpenAI models come to AWS, Hidden Technical Debt of AI Systems: Agent Runtime, a paper on Recursive Multi-Agent Systems, and many more!]]></description><link>https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-453</link><guid isPermaLink="false">https://www.deeplearningweekly.com/p/deep-learning-weekly-issue-453</guid><dc:creator><![CDATA[Miko Planas]]></dc:creator><pubDate>Thu, 30 Apr 2026 15:02:50 GMT</pubDate><enclosure url="https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5f87e5bf-9ea0-4fec-a25d-f48477452317_1100x220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week in deep learning, we bring you <a href="https://openai.com/index/openai-on-aws/">OpenAI models, Codex, and Managed Agents come to AWS</a>, <a href="https://leehanchung.github.io/blogs/2026/04/24/hidden-technical-debt-agent-runtime/">Hidden Technical Debt of AI Systems: Agent Runtime</a> and <a href="https://arxiv.org/abs/2604.25917">a paper on Recursive Multi-Agent Systems</a>.</p><p>You may also enjoy <a href="https://venturebeat.com/technology/mistral-ai-launches-workflows-a-temporal-powered-orchestration-engine-already-running-millions-of-daily-executions">Mistral AI&#8217;s Workflows</a>, <a href="https://research.google/blog/four-ways-google-research-scientists-have-been-using-empirical-research-assistance/">Four ways Google Research scientists have been using Empirical Research Assistance</a>, <a href="https://arxiv.org/abs/2604.22446">a paper on From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company</a>, and more!</p><p>As always, happy reading and hacking. If you have something you think should be in next week&#8217;s issue, find us on Twitter: <a href="https://twitter.com/dl_weekly">@dl_weekly</a>.</p><p>Until next week!</p><div><hr></div><h2><strong>Industry</strong></h2><p><strong><a href="https://openai.com/index/openai-on-aws/">OpenAI models, Codex, and Managed Agents come to AWS</a></strong></p><p>OpenAI brings GPT-5.5, Codex, and Managed Agents to Amazon Bedrock in limited preview, a day after ending Microsoft&#8217;s exclusive cloud license.</p><p><strong><a href="https://venturebeat.com/technology/mistral-ai-launches-workflows-a-temporal-powered-orchestration-engine-already-running-millions-of-daily-executions">Mistral AI launches Workflows, a Temporal-powered orchestration engine already running millions of daily executions</a></strong></p><p>Mistral AI launches Workflows in public preview, a Temporal-powered durable execution engine that lets enterprises run human-in-the-loop AI processes with data staying inside their own infrastructure.</p><p><strong><a href="https://siliconangle.com/2026/04/27/ineffable-intelligence-raises-1-1b-5-1b-valuation-build-ai-superlearner/">Ineffable Intelligence raises $1.1B at $5.1B valuation to build an AI &#8216;superlearner&#8217;</a></strong></p><p>AlphaGo creator David Silver&#8217;s new British startup Ineffable Intelligence raises $1.1B to build a &#8220;superlearner&#8221; AI that generates entirely new knowledge via RL without pretraining.</p><p><strong><a href="https://venturebeat.com/technology/american-ai-startup-poolside-launches-free-high-performing-open-model-laguna-xs-2-for-local-agentic-coding">American AI startup Poolside launches free, high-performing open model Laguna XS.2 for local agentic coding</a></strong></p><p>Poolside releases two agentic coding models &#8212; the open-weight Laguna XS.2 and proprietary Laguna M.1 &#8212; both trained from scratch and free to use temporarily via API, alongside a terminal coding agent and web IDE.</p><h2><strong>MLOps/LLMOps/AgentOps</strong></h2><p><strong><a href="https://leehanchung.github.io/blogs/2026/04/24/hidden-technical-debt-agent-runtime/">Hidden Technical Debt of AI Systems: Agent Runtime</a></strong></p><p>A technical blog post arguing that the agent runtime &#8212; the sandboxed execution environment wrapping the model &#8212; is the emerging hidden technical debt of AI systems.</p><p><strong><a href="https://venturebeat.com/infrastructure/context-decay-orchestration-drift-and-the-rise-of-silent-failures-in-ai-systems">Context decay, orchestration drift, and the rise of silent failures in AI systems</a></strong></p><p>A practical guide about four enterprise AI failure patterns &#8212; context degradation, orchestration drift, silent partial failure, and automation blast radius &#8212; that standard infrastructure monitoring cannot detect, and what teams must add to catch them.</p><p><strong><a href="https://venturebeat.com/infrastructure/monitoring-llm-behavior-drift-retries-and-refusal-patterns">Monitoring LLM behavior: Drift, retries, and refusal patterns</a></strong></p><p>A practical guide on instrumenting LLM applications with two complementary evaluation pipelines &#8212; offline regression testing and online behavioral telemetry &#8212; to detect model drift before and after deployment.</p><h2><strong>Learning</strong></h2><p><strong><a href="https://research.google/blog/four-ways-google-research-scientists-have-been-using-empirical-research-assistance/">Four ways Google Research scientists have been using Empirical Research Assistance</a></strong></p><p>A Google Research blog post on four real-world applications of their Empirical Research Assistance (ERA) tool &#8212; an LLM-backed system for generating expert-level scientific software &#8212; spanning epidemiology, cosmology, climate monitoring, and neuroscience.</p><p><strong><a href="https://alignment.anthropic.com/2026/ai-organizations/">AI Organizations Can Be More Effective but Less Aligned than Individual Agents</a></strong></p><p>An Anthropic research paper found that teams of individually aligned AI agents can still collectively produce less ethical &#8212; but more effective &#8212; solutions than a single agent, suggesting AI safety research needs to move beyond studying agents in isolation.</p><p><strong><a href="https://blog.ml.cmu.edu/2026/04/27/arfbench/">Introducing ARFBench: A time series question-answering benchmark based on real incidents</a></strong></p><p>CMU and Datadog introduce ARFBench, a 750-question time series QA benchmark derived from real production incidents, where the best model (GPT-5 at 62.7% accuracy) still trails domain experts by ~9 points but a hybrid TSFM-VLM oracle reaches 87.2%.</p><p><strong><a href="https://deepmind.google/blog/decoupled-diloco/">Decoupled DiLoCo: Resilient, Distributed AI Training at Scale</a></strong></p><p>Google DeepMind releases Decoupled DiLoCo, a fault-tolerant distributed training architecture that achieves 88% goodput vs. 27% for standard data-parallel at scale, using ~240x less inter-datacenter bandwidth with no measurable ML performance loss.</p><p><strong><a href="https://www.philschmid.de/use-mcp-servers">How to correctly use MCP servers with your AI Agents</a></strong></p><p>A practical guide on avoiding context bloat from MCP servers by using two patterns &#8212; user-triggered @mention injection for ad-hoc tool loading, and scoped subagent declarations.</p><p><strong><a href="https://deepseek.ai/blog/deepseek-v4-compressed-attention">DeepSeek V4 Compressed Attention: How the KV-Cache Shrinks to Just 2%</a></strong></p><p>A technical explainer on how DeepSeek V4 achieves 1M-token context windows by compressing the KV cache to just 2% of standard size &#8212; combining coarse and fine-grained sequence-dimension compression across a hybrid layer stack.</p><h2><strong>Libraries &amp; Code</strong></h2><p><strong><a href="https://github.com/comet-ml/opik">comet-ml/opik</a></strong></p><p>An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.</p><p><strong><a href="https://github.com/facebookresearch/sapiens2">facebookresearch/sapiens2</a></strong></p><p>1K resolution vision transformers pretrained on 1B human images.</p><h2><strong>Papers &amp; Publications</strong></h2><p><strong><a href="https://arxiv.org/abs/2604.25917">Recursive Multi-Agent Systems</a></strong></p><p><strong>Abstract:</strong></p><p>Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states to deepen reasoning. We extend such scaling principle from a single model to multi-agent systems, and ask: Can agent collaboration itself be scaled through recursion? To this end, we introduce RecursiveMAS, a recursive multi-agent framework that casts the entire system as a unified latent-space recursive computation. RecursiveMAS connects heterogeneous agents as a collaboration loop through the lightweight RecursiveLink module, enabling in-distribution latent thoughts generation and cross-agent latent state transfer. To optimize our framework, we develop an inner-outer loop learning algorithm for iterative whole-system co-optimization through shared gradient-based credit assignment across recursion rounds. Theoretical analyses of runtime complexity and learning dynamics establish that RecursiveMAS is more efficient than standard text-based MAS and maintains stable gradients during recursive training. Empirically, we instantiate RecursiveMAS under 4 representative agent collaboration patterns and evaluate across 9 benchmarks spanning mathematics, science, medicine, search, and code generation. In comparison with advanced single/multi-agent and recursive computation baselines, RecursiveMAS consistently delivers an average accuracy improvement of 8.3%, together with 1.2&#215;-2.4&#215; end-to-end inference speedup, and 34.6%-75.6% token usage reduction.</p><p><strong><a href="https://arxiv.org/abs/2604.22446">From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company</a></strong></p><p><strong>Abstract:</strong></p><p>Individual agent capabilities have advanced rapidly through modular skills and tool integrations, yet multi-agent systems remain constrained by fixed team structures, tightly coupled coordination logic, and session-bound learning. We argue that this reflects a deeper absence: a principled organisational layer that governs how a workforce of agents is assembled, governed, and improved over time, decoupled from what individual agents know. To fill this gap, we introduce \emph{OneManCompany (OMC)}, a framework that elevates multi-agent systems to the organisational level. OMC encapsulates skills, tools, and runtime configurations into portable agent identities called \emph{Talents}, orchestrated through typed organisational interfaces that abstract over heterogeneous backends. A community-driven \emph{Talent Market} enables on-demand recruitment, allowing the organisation to close capability gaps and reconfigure itself dynamically during execution. Organisational decision-making is operationalised through an \emph{Explore-Execute-Review} (E2R) tree search, which unifies planning, execution, and evaluation in a single hierarchical loop: tasks are decomposed top-down into accountable units and execution outcomes are aggregated bottom-up to drive systematic review and refinement. This loop provides formal guarantees on termination and deadlock freedom while mirroring the feedback mechanisms of human enterprises. Together, these contributions transform multi-agent systems from static, pre-configured pipelines into self-organising and self-improving AI organisations capable of adapting to open-ended tasks across diverse domains. Empirical evaluation on PRDBench shows that OMC achieves an 84.67% success rate, surpassing the state of the art by 15.48 percentage points, with cross-domain case studies further demonstrating its generality.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.deeplearningweekly.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>