DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 4 days ago • 129
view post Post 4890 📣 HF Viewer now has a HF space! 🤗 embedl/hfviewerVisualize any model directly on Hugging Face - now 4,727 graphs!If you like it, feel free to give the space a heart to help it grow! ❤️And you can reply with any feedback or feature requests here! See translation 🔥 16 16 🧠 1 1 👍 1 1 + Reply
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? Paper • 2609.10226 • Published 12 days ago • 19
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 13 days ago • 171
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 18 days ago • 111
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 17 days ago • 24
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated 19 days ago • 770k • • 923
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 18 days ago • 241
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 20 days ago • 89
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 21 days ago • 62