-
OpenAI Previews GPT-5.6 Family: Sol, Terra, and Luna in Government-Gated Limited Release
OpenAI
models-llm
-
Tencent Officially Releases Hunyuan Hy3: 295B MoE Model with Agent and Reasoning Capabilities
Tencent
models-llm
-
OpenAI GPT-5.6 Sol, Terra, and Luna Launch Publicly After Government Review
OpenAI
models-llm
-
OpenAI Launches GPT-5.6 Family: Sol, Terra, and Luna
OpenAI
models-llm
-
VibeThinker-3B Reaches Frontier-Level Reasoning Benchmarks via Curriculum RL
WeiboAI
research
-
Orca: BAAI's General World Foundation Model Trained on 125K Hours of Video
BAAI
research
-
Tencent Releases Hunyuan Hy3: 295B Open-Weight MoE Model Under Apache 2.0
Tencent
models-llm
-
Tencent Releases Hunyuan Hy3: 295B Open-Weight MoE Under Apache 2.0
Tencent
models-llm
-
Thinking Machines Lab Releases Inkling: 975B Open-Weight Multimodal MoE
Thinking Machines Lab
models-llm
-
Recursive Multi-Agent Systems: agent communication in latent space
Stanford University
research
-
Zyphra Releases ZAYA1-8B: Open Reasoning MoE Model Trained on AMD Hardware
Zyphra
models-llm
-
Google DeepMind's AI Co-Mathematician Reaches 48% on FrontierMath Tier 4
Google DeepMind
research
-
RubricEM: Meta-RL with Rubric-Guided Policy Decomposition Beyond Verifiable Rewards
Google
research
-
SU-01: Gold-Medal-Level Olympiad Reasoning via Curriculum SFT and Two-Stage RL
SU-01 Team
research
-
SOOHAK: Frontier LLMs Solve Hard Math But Fail to Recognize Unsolvable Problems
research
-
Code as Agent Harness: Survey Positions Code as the Substrate for Executable Agent Systems (159 HF upvotes)
Multi-institution (42 authors)
research
-
SkillsVote: Lifecycle Governance of Agent Skills — Collection, Recommendation, Evolution (219 HF upvotes)
Memtensor Research Group / IAAR-Shanghai
research
-
Grok 4.3 Now Available on Amazon Bedrock with 1M-Token Context
xAI
models-llm
-
Qwen-AgentWorld: Language World Models for General Agents at 35B and 397B Scale
Qwen Team, Alibaba
research
-
ByteDance Releases Seed2.0 Model Card for Frontier Model
ByteDance
research
-
DeepSeek Confirms V4 Official Launch for Mid-July with Peak-Time API Pricing
DeepSeek
models-llm
-
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
Tianjin University / Alibaba
research
-
AREX: Towards a Recursively Self-Improving Agent for Deep Research
research
-
OpenAI's Next Model Astra Solves Ten Open Problems in Math and Theoretical CS
OpenAI
research
-
RoPE Provably Fails at Long Contexts: Locality Bias and Token Consistency Both Break
research
-
DRPO: Rethinking Divergence Regularization in LLM Reinforcement Learning
Tencent Hunyuan
research
-
MaxProof: MiniMax Model Exceeds IMO and USAMO Gold-Medal Thresholds on Formal Math
MiniMax
research
-
Mistral Releases Leanstral 1.5: Open Formal-Verification Model for Lean 4
Mistral
research
-
LLM-as-a-Verifier: verification as an independent scaling axis for LLMs
Stanford University / UC Berkeley / NVIDIA
research
-
Ring-Zero: Scaling Zero RL to 1 Trillion Parameters with Emergent Reasoning Behaviors
Ant Group
research
-
Kimi K3 Technical Report: Kimi Delta Attention and Stable LatentMoE Architecture Detailed
Moonshot AI
research
-
Ctx2Skill: Self-Improving Framework for Autonomous Context-Skill Discovery in LLMs
research
-
AI Co-Mathematician: Google DeepMind Achieves 48% on FrontierMath Tier 4
Google DeepMind
research
-
SDAR: Self-Distilled Agentic Reinforcement Learning for Multi-Turn Agents
Zhejiang University / Meituan
research
-
MMSkills: Reusable Multimodal Skills for General Visual Agents (105 HF upvotes)
Shanghai Jiao Tong University
research
-
GrepSeek: Training Search Agents for Direct Corpus Interaction via Shell Commands (93 HF Upvotes)
University of Massachusetts Amherst
research
-
ThoughtFold: Introspective Preference Learning Cuts Reasoning Tokens by 56% Without Accuracy Loss
research
-
The Deterministic Horizon: Information-Theoretic Proof That Extended CoT Fails and Tool Use Is Necessary
research
-
The Self-Correction Illusion: LLMs Fix Others' Errors but Not Their Own — Role Labels Are the Cause
research
-
GitHub Copilot Gets 1M Token Context Window and Configurable Reasoning Levels
GitHub / Microsoft
tools
-
Agentic Transformers Provably Learn Depth-First Search via Reinforcement Learning
Carnegie Mellon University / Ohio State University
research
-
Arbor: Generalist Autonomous ML Research via Hypothesis-Tree Refinement
NLPIR Lab
research
-
DeNovoSWE: Full Repository Generation Jumps from 5.8% to 47.2% with Synthetic Training Data
AweAI Team
research
-
Z-Reward: Score Distributions Instead of Scalar Rewards for Image Generation RLHF
Alibaba
research
-
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
Meta
research
-
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
research
-
OPID: On-Policy Skill Distillation Improves Long-Horizon Agent RL
Institute of Automation, Chinese Academy of Sciences
research
-
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
Zhipu AI / Tsinghua University
research
-
LLM-Driven Formal Mathematics Review: Where Current Systems Fall Short
UCLA
research
-
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Ant Group / Renmin University
research
-
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
research
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
MCG-NJU (Nanjing University)
research
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
SLAI
research
-
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
research
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Horizon Research, Frontis.AI, Tsinghua University
research
-
ESamp: LLMs explore by latent distilling for semantic-novelty sampling
ShanghaiTech University
research
-
Odysseus: Training VLMs for 100+ Turn Interactive Decision-Making via RL
Princeton University
research
-
Soohak: 64 Mathematicians Build Research-Level Benchmark That Stumps Frontier LLMs
Seoul National University
research
-
AutoTTS: LLM Agents Automatically Discover Test-Time Scaling Strategies for $40
research
-
TrOPD: Trust-Region On-Policy Distillation Stabilizes LLM Training When Teacher-Student Gap Is Large
Samsung Research
research
-
Do Language Models Need Sleep? Offline Recurrence as Memory Consolidation for Improved Inference
Google / CMU
research
-
InterleaveThinker: RL Framework for Agentic Text-and-Image Interleaved Generation
research
-
Astra: RL-Trained VLM Queries World Simulator for Spatial Reasoning
research
-
Ideas Have Genomes: Frontier LLMs Score Only 27% on Scientific Lineage Reasoning
Shanghai Jiao Tong University
research
-
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Zhejiang University
research
-
HeavySkill: Internalizing Heavy Thinking as a Trainable Agentic Skill via RL
research
-
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents
Shanghai Jiao Tong University
research
-
Executable World Models for ARC-AGI-3: Coding-Agent Approach Without Game-Specific Logic
research
-
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized AI Research Automation
Shanghai AI Lab
research
-
TMAS: Scaling Test-Time Compute via Multi-Agent Synergy with Hierarchical Memory
research
-
Learning, Fast and Slow: Dual-Weight Architecture for Continual LLM Adaptation
research
-
BetaPRM: Uncertainty-Aware Process Rewards Cut Reasoning Token Use by 33%
research
-
NudgeRL: Strategy-Level Context Nudges for Efficient RLVR Exploration
KAIST AI
research
-
QUBRIC: Co-Designing Queries and Rubrics Extends RLVR to Open-Ended Reasoning Domains
research
-
Quantifying Faithful Confidence Expression in Large Reasoning Models
Yale NLP
research
-
SubtleMemory: Benchmark Reveals Agents Systematically Fail Fine-Grained Relational Memory
research
-
VideoKR: 315K-Example Training Corpus for Knowledge- and Reasoning-Intensive Video Understanding
Yale University
research
-
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
Rutgers University
research
-
SearchSwarm: Delegation Intelligence for LLM Agents in Long-Horizon Deep Research
research
-
Memory is Reconstructed, Not Retrieved: Graph Memory Improves LLM Agent Recall by 23%
National University of Singapore
research
-
ZPPO: Teacher-in-Prompts Knowledge Distillation Outperforms Gradient Methods for Small Reasoners
NVIDIA
research
-
Diffusion-Proof: Formal Theorem Proving via Diffusion Language Models
research
-
DreamReasoner-8B: Block-Size Curriculum for Diffusion Reasoning Models
research
-
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence in VLMs
Nanyang Technological University
research
-
Agentic Transformers Provably Learn to Search via Reinforcement Learning
research
-
Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models
research
-
OPRD: On-Policy Representation Distillation for Post-Training LLMs
research
-
Formalizing Latent Thoughts: Axiomatic Framework for Evaluating LLM Reasoning Representations
University of British Columbia
research
-
OpenCode v1.17.13: Reasoning Mode Fixes and Searchable Model Picker
SST
tools
-
Will Scaling Improve Social Simulation with LLMs? A Study of 85 Models
Stanford / Columbia / Tsinghua
research
-
VRRL: Visually Grounded Self-Reflection for Vision-Language Models via RL
UT Austin / Cornell
research
-
Weak-to-Strong Generalization via Direct On-Policy Distillation
ByteDance / Tsinghua University
research
-
RL Post-Training Actively Builds Compositional Reasoning Strategies, Not Just Amplifies Base Skills
research
-
OpenCode 1.17.19–1.17.20: OpenAI Pro Reasoning Mode and Per-Prompt Model Selection
SST
tools
-
AdvancedMathBench: Benchmark Suite for Advanced Mathematical Proof Generation and Verification
InternLM
research
-
Metacognition in LLMs: Foundations, Progress, and Opportunities — Yale Survey
Yale NLP Lab
research
-
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Mind Lab
research
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
inclusionAI
research
-
Understanding Reasoning from Pretraining to Post-Training
research
-
Recursive Harness Self-Improvement (RHI): Refining Agent Harnesses from Execution Feedback
Sakana AI
research
-
LLMs Get Lost in Evolving User Intent
Microsoft Research
research
-
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
Alibaba
research
-
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
research