Daily digest
20 items · ~20 min · Week 2026-W38
Worth knowing (11)
Trump rejects frontier labs' call to slow AI development, attacks Amodei
After Amodei's 'pace the frontier' essay and public agreement from Altman and Musk, President Trump rejected any slowdown and new AI guardrails, dismissing the safety warnings and personally attacking Anthropic's CEO. His administration framed pacing as the labs' own responsibility and reiterated that whoever wins AI wins.
DeepSeek names first-ever CFO, investment banker Yan Wentao, ahead of possible STAR Market IPO
DeepSeekDeepSeek has hired dealmaker Yan Wentao, a partner at GL Ventures, as its first chief financial officer, Reuters and Chinese media reported on September 14. The appointment comes as the startup works with CITIC Securities on a potential listing on Shanghai's STAR Market.
Z.AI closes around $5 billion Hong Kong financing for next-generation GLM models
Z.AI (Zhipu)Z.AI has completed roughly $5 billion in financing through a $2 billion share placement and $3 billion in convertible bonds, filings and media reported on September 13-14, with about 60 percent of proceeds earmarked for next-generation GLM model R&D and infrastructure. The stock traded below the placement price the following day.
OpenAI acquires camera startup Glass Imaging for over $300 million
OpenAIOpenAI has bought Glass Imaging, an Israeli-founded startup built by former Apple employees that develops AI-enhanced smartphone camera technology, for more than $300 million, per a WSJ report. The deal is read as accelerating OpenAI's consumer hardware plans.
GPT-6 Astra arrives on Amazon Bedrock for enterprise customers
OpenAIAWS's weekly roundup confirms OpenAI's GPT-6 Astra model is now available on Amazon Bedrock, giving AWS enterprise customers native access to OpenAI's flagship work-focused model. Third-party API platforms added the model the same day.
DataFlex-RL: evaluation platform finds RLVR data policies don't beat uniform sampling
Peking UniversityA controlled evaluation platform comparing 13 data-policy configurations for RLVR under a common GRPO recipe (Qwen2.5-7B-Base and Llama-3.1-8B-Base, 12 matched seeds, 12 benchmarks). No rollout-selection, reweighting, or adaptive domain-mixing method achieved a reproducible improvement over uniform GRPO, and benchmark-subset rescaling flipped rankings with a correlation of -0.33.
Feyospace-v1: a seven-person team trains open-weight frontier cyber agents
feyospaceA data-centric post-training framework for cybersecurity agents built from five training systems plus a resettable-environment data engine, yielding 164,269 verified trajectories for long-context SFT. The released checkpoints improve over their base models by 23.76% on CyberGym and 10.49% on pooled CTF suites; Feyospace-s1 reached 63.24% verified success on the CyberGym leaderboard as of Sep 1, 2026.
Tencent Hunyuan open-sources SAS sparse-attention routers trained end-to-end on Qwen3
Tencent HunyuanTencent's Hunyuan team released Simple Attention Sparsification (SAS), a gated sparse-attention method that learns to rank and select KV blocks via continuous gates optimized end-to-end with the language-modeling loss, instead of distilling dense attention scores. Router-only checkpoints (33-42M gate parameters) for Qwen3-4B/8B/14B went up on Hugging Face on September 14, alongside an SGLang-compatible seer_attn backend and paper arXiv 2609.13141.
Andon Labs launches Pion, an agent platform that runs real companies autonomously
Andon LabsAndon Labs released Pion (Sep 14, research preview), a platform handing real businesses to persistent AI agents with email, phone, banking, and browser tools. It grew out of Vending-Bench and real deployments including a San Francisco store and a Stockholm cafe run by agents directing human staff over Slack.
Grok models land in Microsoft Copilot across Word, Excel and PowerPoint as a subprocessor
xAIMicrosoft is rolling xAI's Grok models into Copilot inside Word, Excel and PowerPoint, with xAI acting as a subprocessor handling prompts under Microsoft's data-processing terms. Reports note system prompts keep Grok outputs strictly work-appropriate.
ByteDance launches Doubao Phone Assistant consumer version with beta screen-automation controls
ByteDanceByteDance officially launched the consumer version of its system-level Doubao Phone Assistant on September 14, adding screen-based Q&A, local search across photos and messages, voice activation, and a beta 'operate phone' screen-automation mode built on the SAEP protocol, which lets third-party apps declare boundaries and opt out of AI-driven operations. It ships first on the Nubia NaviX Ultra agent phone.
For reference (9)
xAI and X Corp drop Apple from antitrust lawsuit over App Store ranking
xAIMusk's xAI, X Corp and SpaceX moved to dismiss Apple from their antitrust lawsuit, which had accused Apple of suppressing apps like Grok in App Store rankings. No settlement terms were disclosed, and the parallel claims against OpenAI remain live.
Benchmark Radar: a living database and search engine for AI benchmarks
Carnegie Mellon UniversityA continuously updated catalog and search engine for AI benchmarks spanning LLM, agentic, coding, reasoning, and safety evaluation, aggregating daily from 37 sources into 1,283 records and 12,916 score observations with preserved citation trails. It also ships analyses of benchmark saturation, adoption trends, and a prior-art search workflow for designing new evaluations.
COBRA-Skills: contextual bandit-guided evolution for agent skills
Chinese University of Hong Kong, ShenzhenTreats LLM agent skill-library optimization as budgeted sequential optimization over an evolving candidate space, pairing contextual-bandit-guided prioritization with evidence-grounded skill evolution from execution feedback. Across six agent benchmarks and three models it matches or beats prior skill-optimization methods while cutting optimization cost by 55-58% relative to SkillOpt, using only 50 examples per benchmark.
Claude Code v2.1.271 adds fast mode in remote sessions and per-command sandbox domains
AnthropicClaude Code v2.1.271 (Sep 14) brings fast mode to Remote sessions, per-command allowed_domains for Bash/PowerShell/Monitor sandboxing, omitClaudeMd for subagents, and --accept-command (sha256) for plugin installs, plus MCP OAuth and --resume fixes. A follow-up v2.1.272 (Sep 15) ships only bug fixes.
llama.cpp v0.4.1 adds Maple 20B, Tencent Hy 4, and Spark2.5 support
ggmlllama.cpp v0.4.1 (Sep 14) adds support for Maple 20B-A1B, Tencent Hy 4, and Spark2.5 models, refactors JSON schema handling, and adds JSONL logging via --log-jsonl. The mmap/mlock/direct-io flags are replaced by a unified --load-mode option.
GitHub Copilot lets users tune cost and quality in auto model selection
GitHubGitHub announced (Sep 14) that Copilot's automatic model selection can now be configured for cost versus quality, letting users bias routing toward cheaper or more capable models per task.
qwen-code v0.23.4 adds native subagent builtins and cross-session agent board
Alibaba (Qwen)Alibaba's open-source Qwen Code CLI shipped v0.23.4 on September 14, bringing native agent builtins for subagents, a Codex executor, an 'agent board' for sharing work across independently started agents, a peer endpoint for external programs to join cross-session messaging, configurable Goal limits, and about 60 bug fixes.
Anthropic launches Claude for financial advisers with Schwab partnership
AnthropicAnthropic pitched a Claude product aimed at financial advisers, automating prep work such as meeting preparation and compliance-heavy workflows. Coverage reports a distribution partnership with Charles Schwab and an emphasis on regulatory compliance for wealth management.
OpenCode v1.18.31 restores ACP session state on resume and fork
SSTOpenCode v1.18.31 (Sep 14) fixes restoration of ACP session model, effort, mode, and reasoning chunk boundaries when loading, resuming, or forking sessions, and surfaces remote config auth errors at startup. Extensions now summarize adaptive thinking for GitHub Copilot models.