| 扩散语言模型 |
dQwen3.5: Adapting Hybrid Language Models into Bidirectional Diffusion Language Models |
University of Texas at Austin,2026-09-17 |
未发现官方代码 |
dqwen35 |
| 注意力与长上下文 |
On-Demand Attention: Efficient Long-Context Decoding with Learned Recall |
Southern University of Science and Technology,2026-09-17 |
未发现官方代码 |
oda |
| 推测解码 |
ASPIRE: Asynchronous Batch Self-Speculative Decoding |
University of Southern California,2026-09-16 |
未发现官方代码 |
aspire |
| 推测解码 |
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding |
Wuhan University,2026-09-15 |
已开源 |
echo |
| 推测解码 |
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers |
Seoul National University / KAIST,2026-09-15 |
未发现官方代码 |
loopspec |
| 可复现训练 |
OPEN-1B: A Fully Auditable Training Run |
Gensyn,2026-09-15 |
已开源 |
open-1b-audit |
| 网络架构 |
Persistent Recurrent Memory Between Transformer Layers Improves Language Model Generalization |
FITec Labs / Ericsson São Paulo,2026-09-15 |
未发现官方代码 |
persistent-recurrent-memory |
| Agent 推理 |
AgentKV: Phase-Aware KV Eviction for Agentic LLMs |
University of Cambridge,2026-09-14 |
已开源 |
agentkv |
| 扩散语言模型 |
Register Tokens for Bounded-State Reasoning in Diffusion Language Models |
University of Wisconsin–Madison,2026-09-14 |
未发现官方代码 |
register-tokens-dllm |
| 注意力与长上下文 |
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking |
Tencent HY LLM Frontier / HKUST (Guangzhou) / HKUST,2026-09-11 |
已开源 |
sas-attention |
| 长视频理解 |
Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding |
Queen Mary University of London,2026-09-10 |
未发现官方代码 |
frames-on-demand |
| MoE |
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data |
Stanford University,2026-09-10 |
未发现官方代码 |
repeat-aware-moe |
| 优化器 |
Musec: MomentUm SpEctral Clipping for Stable Muon-type Training |
National University of Singapore,2026-09-10 |
未发现官方代码 |
musec |
| 多模态推理 |
OmniKVQuant: KV Cache Quantization for Omni-LLMs |
KAIST,2026-09-10 |
已开源 |
omnikvquant |
| 统一多模态 |
SenseNova-U1.5: Towards Native Unified Visual Intelligence |
SenseTime Research,2026-09-10 |
已开源 |
sensenova-u1-5 |
| 模型路由 |
SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations |
Shanghai Jiao Tong University,2026-09-10 |
未发现官方代码 |
swrouter |
| 推理与系统效率 |
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference |
Hanyang University,2026-09-04 |
未发现官方代码 |
beaconkv |
| 推理与系统效率 |
KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU |
Shanghai University of Finance and Economics,2026-09-04 |
已开源 |
kvmem |
| 网络架构 |
Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations |
Beijing University of Posts and Telecommunications / Kuaishou Technology,2026-09-03 |
未发现官方代码 |
lngram-v2 |
| 推理与系统效率 |
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning |
Salesforce AI Research / University of Illinois Urbana-Champaign,2026-09-03 |
已开源 |
random-attention |
| 弱模型失败模式 ICL |
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes |
The Ohio State University,2026-08-27 |
未发现官方代码 |
criticl |
| 多模态基础模型 |
PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference |
Sun Yat-sen University,2026-08-27 |
已开源 |
pace-vlm |
| 注意力与长上下文 |
TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy |
The Hong Kong University of Science and Technology (Guangzhou),2026-08-27 |
未发现官方代码 |
twinkv |
| 多模态基础模型 |
MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations |
KAIST / Sony AI,2026-08-26 |
未发现官方代码 |
mllmclip |
| 多模态基础模型 |
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning |
Nanyang Technological University / VBVR Community,2026-08-26 |
已开源 |
vbvr-pro |
| 多模态基础模型 |
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report |
WeChat Vision, Tencent,2026-08-25 |
已开源 |
wemm-embedding |
| 网络架构 |
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models |
Huazhong University of Science and Technology,2026-08-21 |
未发现官方代码 |
rare |
| 预训练与数据 |
Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing |
New York University,2026-08-13 |
未发现官方代码 |
tcab |
| 多模态基础模型 |
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction |
ByteDance,2026-08-12 |
未发现官方代码 |
gas |
| 注意力与长上下文 |
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension |
Ai2,2026-08-10 |
已开源 |
olmpool-long-context |
| 推理与系统效率 |
DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference |
Oklahoma State University,2026-08-09 |
未发现官方代码 |
distillcache |
| 注意力与长上下文 |
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry |
Institute of Computing Technology, Chinese Academy of Sciences,2026-08-07 |
未发现官方代码 |
autonomy-heads |
| 推理与系统效率 |
BaKron: Efficient Quantization with Kronecker-Factored Hessians |
University of California, San Diego,2026-08-06 |
未发现官方代码 |
bakron |
| 预训练与数据 |
Hierarchical Latent Prediction for Language Models |
University of Texas at Austin,2026-08-06 |
未发现官方代码 |
hilp |
| 网络架构 |
MACRO: Markov Chain Routing of Transformer Layers |
Heinrich Heine University Düsseldorf,2026-08-06 |
已开源 |
macro |
| 推理与系统效率 |
DBLast: Dependent Block Drafting for Stochastic Speculative Decoding |
Huawei Technologies Canada,2026-08-05 |
未发现官方代码 |
dblast |
| 推理与系统效率 |
QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding |
Indian Institute of Technology Roorkee,2026-08-05 |
未发现官方代码 |
qevict |
| 多模态基础模型 |
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes |
Meta,2026-08-05 |
未发现官方代码 |
physics-mm-pretraining |
| 网络架构 |
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling |
Zhejiang University,2026-08-03 |
未发现官方代码 |
dart |
| 注意力与长上下文 |
Learning What to Remember: Test-Time Training via Context Distillation |
Princeton University,2026-08-03 |
已开源 |
ttcd |
| 网络架构 |
Role-Decoupled Attention Residuals |
Kehan Wang(论文未列机构),2026-08-03 |
未发现官方代码 |
rd-attnres |
| 网络架构 |
TransMem: Transforming Hidden States into Memory for Large Language Models |
Authors did not disclose affiliation,2026-07-31 |
未发现官方代码 |
transmem |
| 多模态基础模型 |
ReToken: One Token to Improve Vision–Language Models for Visual Retrieval |
UIUC / Microsoft Research / Google DeepMind,2026-07-30 |
已开源 |
retoken |
| 推理与系统效率 |
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning |
EIT-NLP / LMU Munich,2026-07-30 |
已开源 |
wide |
| 网络架构 |
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning |
Academic author team,2026-07-28 |
未发现官方代码 |
penelope |
| 预训练与数据 |
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data |
Fudan University / Shanghai Jiao Tong University / SII-GAIR,2026-07-27 |
已开源 |
data-orchestra |
| 推理与系统效率 |
Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs |
Huawei ACS Lab / Southern University of Science and Technology,2026-07-23 |
未发现官方代码 |
adadsf |
| 注意力与长上下文 |
Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable |
Independent researcher,2026-07-23 |
未发现官方代码 |
mobius-rope |
| 网络架构 |
Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory |
Independent researchers,2026-07-23 |
未发现官方代码 |
naju |
| 注意力与长上下文 |
Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection |
Pennsylvania State University,2026-07-23 |
未发现官方代码 |
gzip-sparse-attention |
| 推理与系统效率 |
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context |
NVIDIA,2026-07-23 |
已开源 |
windowed-mtp |
| 推理与系统效率 |
GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries |
University of Cambridge,2026-07-22 |
已开源 |
gaugequant |
| 网络架构 |
Convolution for Large Language Models |
Huawei / Peking University / Tsinghua University,2026-07-20 |
未发现官方代码 |
conv-llm |
| 推理与系统效率 |
C²KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference |
Shanghai Jiao Tong University,2026-07-20 |
已开源 |
c2kv |
| 预训练与数据 |
PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning |
McGill University,2026-07-20 |
未发现官方代码 |
ppl-factory |
| 预训练与数据 |
OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research |
Indian Institute of Technology Madras,2026-07-18 |
已开源 |
open-language-model |
| 注意力与长上下文 |
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers |
University of Maryland / Meta AI,2026-07-16 |
未发现官方代码 |
looped-latent-attention |
| 注意力与长上下文 |
MiniMax Sparse Attention |
MiniMax,2026-06-11 |
已开源 |
minimax-sparse-attention |
| 网络架构 |
Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory |
Tsinghua University / Microsoft Research Asia,2026-05-20 |
未发现官方代码 |
memory-grafting |
| 注意力与长上下文 |
Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers |
Peking University / Huawei Technologies,2026-03-27 |
未发现官方代码 |
switch-attention |
| 生成式检索 |
Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion |
University of Illinois Urbana-Champaign / Google Research,2026-03-06 |
未发现官方代码 |
r4t |
| 网络架构 |
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models |
DeepSeek,2026-01-12 |
已开源 |
engram |
| 网络架构 |
mHC: Manifold-Constrained Hyper-Connections |
DeepSeek-AI,2025-12-31 |
未发现官方代码 |
mhc |
| 动态候选分类 |
GLiClass: Generalist Lightweight Model for Sequence Classification Tasks |
Knowledgator,2025-08-11 |
已开源 |
gliclass |
| 注意力与长上下文 |
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free |
Qwen / Alibaba,2025-05-10 |
已开源 |
gated-attention |
| 多模态基础模型 |
SmolVLM: Redefining small and efficient multimodal models |
Hugging Face,2025-04-07 |
已开源 |
smolvlm |
| 预训练与数据 |
Muon is Scalable for LLM Training |
Moonshot AI / UCLA,2025-02-24 |
已开源 |
muon |
| 多模态基础模型 |
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features |
Google DeepMind,2025-02-20 |
已开源 |
siglip2 |
| 注意力与长上下文 |
MoBA: Mixture of Block Attention for Long-Context LLMs |
Moonshot AI,2025-02-18 |
已开源 |
moba |
| 注意力与长上下文 |
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention |
DeepSeek,2025-02-16 |
未发现官方代码 |
native-sparse-attention |
| 网络架构 |
Byte Latent Transformer: Patches Scale Better Than Tokens |
Meta FAIR,2024-12-13 |
已开源 |
blt |
| 网络架构 |
Hymba: A Hybrid-head Architecture for Small Language Models |
NVIDIA,2024-11-20 |
已开源 |
hymba |
| 预训练与数据 |
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance |
University of Cambridge / Shanghai AI Laboratory,2024-03-25 |
已开源 |
data-mixing-laws |
| 推理与系统效率 |
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads |
Together AI / Princeton University / University of Illinois Urbana-Champaign,2024-01-19 |
已开源 |
medusa |
| 网络架构 |
Mamba: Linear-Time Sequence Modeling with Selective State Spaces |
Carnegie Mellon University / Princeton University,2023-12-01 |
已开源 |
mamba |
| 推理与系统效率 |
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration |
MIT / NVIDIA / Harvard / SJTU,2023-06-01 |
已开源 |
awq |
| 注意力与长上下文 |
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints |
Google Research,2023-05-22 |
未发现官方代码 |
gqa |
| 预训练与数据 |
DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining |
Stanford University / Google Research,2023-05-17 |
已开源 |
doremi |
| 多模态基础模型 |
Visual Instruction Tuning |
University of Wisconsin-Madison / Microsoft Research / Columbia University,2023-04-17 |
已开源 |
llava |
| 多模态基础模型 |
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
Salesforce Research,2023-01-30 |
已开源 |
blip2 |
| 推理与系统效率 |
Fast Inference from Transformers via Speculative Decoding |
Google Research,2022-11-30 |
已开源 |
speculative-decoding |
| 注意力与长上下文 |
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation |
University of Washington / Meta AI,2021-08-27 |
已开源 |
alibi |
| 注意力与长上下文 |
RoFormer: Enhanced Transformer with Rotary Position Embedding |
Zhuiyi Technology,2021-04-20 |
已开源 |
rope |
| 多模态基础模型 |
Learning Transferable Visual Models From Natural Language Supervision |
OpenAI,2021-02-26 |
已开源 |
clip |
| 网络架构 |
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity |
Google Brain,2021-01-11 |
已开源 |
switch-transformer |