Collections
Discover the best community collections!
Collections including paper arxiv:2604.13074
-
PersonaVLM: Long-Term Personalized Multimodal LLMs
Paper • 2604.13074 • Published • 46 -
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
Paper • 2604.16044 • Published • 73 -
Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems
Paper • 2604.04936 • Published • 26 -
Qwen3.5-Omni Technical Report
Paper • 2604.15804 • Published • 60
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 10.8k • 1.49k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 73 • • 18 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Paper • 2604.05015 • Published • 236 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 373 -
LPM 1.0: Video-based Character Performance Model
Paper • 2604.07823 • Published • 82 -
A Simple Baseline for Streaming Video Understanding
Paper • 2604.02317 • Published • 74
-
The Art of Scaling Reinforcement Learning Compute for LLMs
Paper • 2510.13786 • Published • 34 -
Attention Is All You Need for KV Cache in Diffusion LLMs
Paper • 2510.14973 • Published • 43 -
BitNet Distillation
Paper • 2510.13998 • Published • 63 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 54
-
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
Paper • 2503.20756 • Published • 7 -
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Paper • 2505.09568 • Published • 100 -
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Paper • 2508.18265 • Published • 224 -
Qwen3-Omni Technical Report
Paper • 2509.17765 • Published • 154
-
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Paper • 2604.05015 • Published • 236 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 373 -
LPM 1.0: Video-based Character Performance Model
Paper • 2604.07823 • Published • 82 -
A Simple Baseline for Streaming Video Understanding
Paper • 2604.02317 • Published • 74
-
PersonaVLM: Long-Term Personalized Multimodal LLMs
Paper • 2604.13074 • Published • 46 -
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
Paper • 2604.16044 • Published • 73 -
Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems
Paper • 2604.04936 • Published • 26 -
Qwen3.5-Omni Technical Report
Paper • 2604.15804 • Published • 60
-
The Art of Scaling Reinforcement Learning Compute for LLMs
Paper • 2510.13786 • Published • 34 -
Attention Is All You Need for KV Cache in Diffusion LLMs
Paper • 2510.14973 • Published • 43 -
BitNet Distillation
Paper • 2510.13998 • Published • 63 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 54
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 10.8k • 1.49k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 73 • • 18 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
Paper • 2503.20756 • Published • 7 -
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Paper • 2505.09568 • Published • 100 -
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Paper • 2508.18265 • Published • 224 -
Qwen3-Omni Technical Report
Paper • 2509.17765 • Published • 154