Daily digest
9 items · ~9 min · Week 2026-W32
Worth knowing (5)
xAI ships Grok Imagine Image 2.0 with region-level editing and multi-image reference input
xAIxAI made Grok Imagine Image 2.0 generally available as Quality Mode on grok.com/imagine and in its iOS/Android apps, adding a magic-wand region editor, segmentation-based selection, background removal, and multi-reference editing that accepts up to five input images in one generation.
LongHorizon-Harness proposes Manage-Execute-Audit loop for long-horizon agents
Proposes a Manage-Execute-Audit (MEA) loop that separates task-state tracking from execution, using a manager to hold verified task state, a fresh-context executor per subtask, and a read-only auditor to check environment state before each new round. Reported gains include Qwen 3.7-Plus rising from 51.8% to 80.7% on WeaveBench and Claude Opus 4.7 rising from 20.0% to 34.3% on an OSWorld 2.0 subset.
Recursive Synthetic Terminal Tasks generates 37,000+ verifiable long-horizon agent tasks
Introduces Recursive Synthetic Terminal Tasks (RST), a recursive verified-synthesis framework that automatically generates over 37,000 verifiable long-horizon terminal-agent tasks, keeping instructions, environments, solutions, and verifiers consistent at roughly $0.05 per task. Task difficulty compounds across rounds, with median reference-solution length growing from 67 to 374 lines and DeepSeek-V4-Pro pass@4 dropping from 90% to 2.5% over 15 rounds.
DAPD identifies 'privilege illusion' failure mode in on-policy self-distillation
Shanghai AI LaboratoryIdentifies a 'privilege illusion' failure mode in on-policy self-distillation, where a student model learns behavior dependent on training-time privileged information it cannot access at inference. Proposes Dual-Path and Dual-Source Anchoring to align reference and rollout behavior bidirectionally, improving over prior on-policy self-distillation by roughly 2-2.8 points across Qwen3 models from 4B to 32B.
Claude Code opens public beta of self-hosted environments; v2.1.224 adds cross-session agent messaging
AnthropicAnthropic launched a public beta letting Claude Team and Enterprise customers run Claude Code cloud sessions on their own infrastructure via `claude self-hosted-runner`, keeping repository checkouts, build artifacts, and secrets inside the organization's network. The same v2.1.224 release also added cross-session `SendMessage`/`ListAgents` APIs so sessions on different machines can message each other; v2.1.225/226 followed on August 8 with gateway spend-limit warnings and MCP OAuth fixes.
For reference (4)
OpenAI acquires AI presentation startup NextSlide
OpenAIOpenAI acquired presentation-generation startup NextSlide, with its team joining to build new ChatGPT productivity features for turning prompts and documents into polished, editable presentations. The deal had closed earlier but was disclosed publicly around August 7-8, 2026, with financial terms undisclosed.
Anthropic loosens Claude Fable 5's biology safety classifiers, cutting false-positive blocks 85%
AnthropicAnthropic rewrote the biology safety classifiers gating Claude Fable 5, reducing unnecessary fallbacks to the less-capable Opus 5 model by about 85% while continuing to block dual-use requests in virology, toxicology, and molecular design.
OpenAI Codex CLI 0.147.0 adds portable Agent Plugins and MCP 2026-07-28 support
OpenAICodex CLI 0.147.0, released August 7, 2026, adds portable Agent Plugins installable from local, personal, workspace, or remote catalogs, persistent manually-ordered conversation sections, an --approve-for-me auto-approval flag, and support for the MCP 2026-07-28 protocol revision.
ByteDance opens Seedance 2.5 developer API to the public
ByteDanceByteDance opened public developer API access to Seedance 2.5 on August 7, 2026, a week after the model's consumer launch. The API exposes 30-second single-shot video generation with up to 50 multimodal reference inputs (images, video, audio) and 3D camera blockout control.