This week in deep learning, we bring you OpenAI’s GPT-6.1 Sol, Jev vs. LLM-as-a-Judge for AI Evals, and a paper on Just-in-Time Memory for LLM Agents.
You may also enjoy Anthropic’s Claude Sonnet 5.5, Cursor’s deep dive on improving token efficiency for longer agent runs, a paper on IterSynth for deep search agents, and more!
As always, happy reading and hacking. If you have something you think should be in next week’s issue, find us on Twitter: @dl_weekly.
Until next week!
Industry
OpenAI introduces GPT-6.1 Sol at 2/10 per million tokens against Astra’s 10/50, beating GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 on AutomationBench.
Anthropic launches Sonnet 5.5, which scores 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1, nearly matching Opus 5.5 on GDPval-AA.
OpenAI’s DevDay headliner gives each Dot a dedicated cloud computer and browser running on GPT-6 Astra, with Custom Rules gating actions and passwords reserved for humans.
AMD will acquire Fei-Fei Li’s World Labs for $8.2B
AMD buys World Labs for $8.2 billion and makes Fei-Fei Li executive vice president and chief scientist, giving it a world-model answer to Nvidia’s Cosmos line.
Nvidia launches new platform for reining in rogue AI agents
Jensen Huang’s Open Agent Safety Platform pairs OpenShell access control with Sentry monitoring on BlueField-4 DPUs, moving the guardrail off the processor the agent runs on.
Introducing Embed 5—A New Family of Frontier Embedding Models
Cohere ships Embed 5 Pro and Fast at $0.12 and $0.08 per million tokens sharing one embedding space, so teams index with Pro and query with either without re-indexing.
FLUX 3 Action: A 7B World Action Model for Robot Control
Black Forest Labs releases an open-weight 7B world-action model that tops the RoboLab-120 leaderboard at under half the parameters of the previous best open model and up to 3.95x faster.
MLOps / LLMOps / AgentOps
Jev vs. LLM-as-a-Judge for AI Evals
A practical comparison of Jev and GPT-4o-mini as LLM judges on 1,000 production turns, where Jev was 3.5× cheaper and 3.8× faster at the median while agreeing on 89.7% of evaluations.
Add Runtime Controls to AI Agents with NVIDIA OpenShell
An article about OpenShell 0.1.0, an open-source runtime enforcing which systems an agent may reach without rewriting it, using a formal policy prover to verify permission boundaries.
Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
A deep-dive article about why sparse MoE post-training becomes communication-bound rather than compute-bound, and how DeepEP over EFA lifts aggregate RL rollout throughput by 40%.
What if your agent’s hallucinations had a budget? How to start using SLOs for agent behavior
A practical article about mapping error budgets onto agent quality, treating groundedness, helpfulness and toxicity as separate behaviors that fail independently and need separate targets.
Learning
Improved token efficiency for longer agent runs
A detailed engineering post about where agent inference spend actually goes, and how trimming roughly 66% of the system prompt and reusing cache cut user token costs 7% with no quality loss.
Free the models: Harness design at the frontier
An opinionated article about why a model router can never exceed the model it routes to, and how letting the core loop pick its own effort beat a single-worker architecture by 11 and 16 points.
A data report about the skills.sh registry reaching one million agent skills and nearly 280 million installs in seven months, versus 27 months for GitHub to reach a million repositories.
An annotated keynote about the year’s LLM trends, tracing the moment coding agents crossed from “often make mistakes” to reliable enough for daily use.
Best practices guide for customizing Gemini models via Reinforcement Learning (RL)
A practical guide about a managed RL fine-tuning service where you bring prompts and a reward function, unlocking tasks that are hard to demonstrate but easy to score.
Embedded Evaluators are necessary for meaningful external testing
An argumentative post about why final-checkpoint testing could not have caught the Hugging Face incident, which arose far earlier in development, and what employee-equivalent access would change.
GLM-5.3 and the spread of advanced cyber capabilities
A red-team analysis finding that simple techniques bypass GLM-5.3’s safeguards 64% to 100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models.
Libraries & Code
An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Papers & Publications
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and τ2-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior ≤8B agent by +4.2\%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.


