1DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionThe widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains coDeepSeek09-18933
2Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative ModelRecent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimHaoyu Zhao09-1877422
3When EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationWe study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between baMicrosoft09-187556
4SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedbNVIDIA09-186832344
5An Empirical Study of Harness Design for Coding AgentsCoding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systeZoom Communications09-18553
6JEPA-Anything: Learning Predictive Models across Different WorldsWorld modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support wtaesiri09-1838448
7Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns Google09-18333
8RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web InvestigationPlatform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with leonliuzx09-182831
9RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement LearningMulti-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision fromZhengxi Lu09-18263
10Self-Evolving Search IndexInformation retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through indexSangam Lee09-18252
11Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI AgentsGUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before execution. Recent agent-skill frZhengxi Lu09-1825213
12Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video GenerationVideo diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and taesiri09-18232
13WeVisDoc: From Coverage to Capability for Robust End-to-End Document ParsingDocument parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common docTencent09-1822249
14VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric ControlSpatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act onDaLian University of Technology09-1820212
15When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning ModelsLarge Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based onMicrosoft09-18182
16FAMOS: Feed-Forward 3D Articulation Modeling from Sparse ObservationsModeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a Stanford University09-18172
17What Does Privileged Information Add to On-Policy Self-Distillation?On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the studeNational University of Singapore09-181724
18Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time ScalingTest-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number iowa state university09-18164
19Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RLAgent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as conteAmazon09-18162
20UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image GenerationMulti-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains laUniversity of Science and Technology of China09-18162