This week in deep learning, we bring you Gemini 4 Argon, Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle and a paper on Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability.
You may also enjoy Cloudflare’s Clef, a paper on Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks, and more!
As always, happy reading and hacking. If you have something you think should be in next week’s issue, find us on Twitter: @dl_weekly.
Until next week!
Industry
Gemini 4 Argon: our next era of frontier intelligence
Gemini 4 Argon scores 77.9% on DeepSWE v1.1 against Claude Opus 5.5’s 74.2% and ships first to 650-plus Fairwind Program defenders at 2/10 per million tokens.
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
A blog post about Cloudflare’s Apache-2.0 Clef decision models, which add a vision encoder and a 64k context window against Jev’s 32k and classified a domain in 2.2s versus 4.7s for gpt-oss-120b.
Mistral opens a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters, with open weights promised by month’s end.
Google expands EmbeddingGemma beyond text to images, audio and video
EmbeddingGemma 2 reaches 740 million parameters and 78.68 on MTEB’s code section, nearly ten points above v1, with a 270M text core using about 191MB quantized.
North 2: Enterprise AI Without Compromises
Cohere rebuilds North around a new agent harness, adds memory, skills and artifacts, and gives admins granular cost controls, rate limits and org-wide token caps.
AWS debuts Strands Decider 2B, a first lightweight decision model for accelerate agentic workflows
AWS open-sources Strands Decider 2B, a laptop-runnable decision model that returns confidence-scored choices instead of generated text, aimed at the structural limits its team sees in Jev.
Ai2 releases Olmo-core 3 to make developing large mixture-of-experts LLMs more efficient
The Allen Institute for AI ships Olmo-core 3, a training framework that carries mixture-of-experts models to trillion-parameter scale while holding down expert-routing and networking overhead.
MLOps / LLMOps / AgentOps
UX Without a UI: Notes from Researching Terminal UX for Opik MCP
A look at what changes when designing developer tools for AI agents instead of traditional interfaces, including the challenges of displaying complex data in terminals, handling unpredictable LLM behavior, and designing for both humans and agents.
How to implement long-term AI agent memory in AlloyDB and Memorystore for Valkey
A guide about a two-tier agent memory architecture — Valkey for short-term buffer, AlloyDB AI for long-term persistence — that Google says can cut token spend by up to 70%.
Build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances
A hands-on blog post about AgentCore’s new Runtime Instances, which add multi-day sessions, GPUs, persistent volumes and colocated agents to the serverless MicroVM option.
Learning
Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle
A research blog post about framing agent privacy and security through contextual integrity, setting out the open problems that arise once agents act on a user’s behalf.
A critical blog post about trying Anthropic’s build_eval and hill-climb commands on real leasing-assistant traces, arguing they push you to write an eval before you have looked at data.
Research: Qwen3.8 27B addition in words
A benchmark note about a local Qwen3.8-27B quant scoring 23.57% numeric accuracy across 5,070 reasoning-disabled addition cases, falling from 97.04% on short operands to 6.44% on long ones.
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
A blog post about an objective-metric TTS leaderboard built because arenas cannot keep pace with over 8K Hub TTS models, only 16 of 92 Artificial Analysis entries being open-weights.
A guest research post about a physicist dropping the human-scientist workflow and building BootLoops, an open-source harness that lets Claude work on problems shaped to its strengths.
How many AI agents could we run?
An analysis about chips shipped through 2027 supporting tens to hundreds of millions of concurrent frontier agents — as many weekly hours as 140-720 million full-time employees.
Principles for Embedded Evaluations
A position blog post about designing embedded third-party evaluations around verifying or falsifying a developer’s own safety claims, with ongoing access to training data, rollouts and checkpoints.
Libraries & Code
An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Skills for real engineers. Straight from Matt Pocock’s .agents directory.
Papers & Publications
On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.


