Submitted by Kaixiang Zhao 75 When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Microsoft 6 5
Submitted by Jaejun Shim 18 When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Microsoft 2
Submitted by Sungho Park 64 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Microsoft 224 2
Submitted by Niels Rogge 12 One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Microsoft 81 2
Submitted by Sukmin Cho 25 OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching Microsoft 2
Submitted by Dong Yan 19 AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Microsoft 15 2
Submitted by Senqiao Yang 36 Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Microsoft 1.65k 2