跳转至

基础模型论文与资料索引

本页由 docs/research-manifest.json 自动生成;论文元数据只在统一 manifest 维护。背景、架构、公式、原文效果和本地实验请进入独立详情页。

已实现论文与资料

方向 方法 机构与日期 原作者代码 本地入口
扩散语言模型 dQwen3.5: Adapting Hybrid Language Models into Bidirectional Diffusion Language Models University of Texas at Austin,2026-09-17 未发现官方代码 dqwen35
注意力与长上下文 On-Demand Attention: Efficient Long-Context Decoding with Learned Recall Southern University of Science and Technology,2026-09-17 未发现官方代码 oda
推测解码 ASPIRE: Asynchronous Batch Self-Speculative Decoding University of Southern California,2026-09-16 未发现官方代码 aspire
推测解码 ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding Wuhan University,2026-09-15 已开源 echo
推测解码 LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers Seoul National University / KAIST,2026-09-15 未发现官方代码 loopspec
可复现训练 OPEN-1B: A Fully Auditable Training Run Gensyn,2026-09-15 已开源 open-1b-audit
网络架构 Persistent Recurrent Memory Between Transformer Layers Improves Language Model Generalization FITec Labs / Ericsson São Paulo,2026-09-15 未发现官方代码 persistent-recurrent-memory
Agent 推理 AgentKV: Phase-Aware KV Eviction for Agentic LLMs University of Cambridge,2026-09-14 已开源 agentkv
扩散语言模型 Register Tokens for Bounded-State Reasoning in Diffusion Language Models University of Wisconsin–Madison,2026-09-14 未发现官方代码 register-tokens-dllm
注意力与长上下文 SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Tencent HY LLM Frontier / HKUST (Guangzhou) / HKUST,2026-09-11 已开源 sas-attention
长视频理解 Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding Queen Mary University of London,2026-09-10 未发现官方代码 frames-on-demand
MoE Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data Stanford University,2026-09-10 未发现官方代码 repeat-aware-moe
优化器 Musec: MomentUm SpEctral Clipping for Stable Muon-type Training National University of Singapore,2026-09-10 未发现官方代码 musec
多模态推理 OmniKVQuant: KV Cache Quantization for Omni-LLMs KAIST,2026-09-10 已开源 omnikvquant
统一多模态 SenseNova-U1.5: Towards Native Unified Visual Intelligence SenseTime Research,2026-09-10 已开源 sensenova-u1-5
模型路由 SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations Shanghai Jiao Tong University,2026-09-10 未发现官方代码 swrouter
推理与系统效率 BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference Hanyang University,2026-09-04 未发现官方代码 beaconkv
推理与系统效率 KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU Shanghai University of Finance and Economics,2026-09-04 已开源 kvmem
网络架构 Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations Beijing University of Posts and Telecommunications / Kuaishou Technology,2026-09-03 未发现官方代码 lngram-v2
推理与系统效率 Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Salesforce AI Research / University of Illinois Urbana-Champaign,2026-09-03 已开源 random-attention
弱模型失败模式 ICL CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes The Ohio State University,2026-08-27 未发现官方代码 criticl
多模态基础模型 PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference Sun Yat-sen University,2026-08-27 已开源 pace-vlm
注意力与长上下文 TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy The Hong Kong University of Science and Technology (Guangzhou),2026-08-27 未发现官方代码 twinkv
多模态基础模型 MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations KAIST / Sony AI,2026-08-26 未发现官方代码 mllmclip
多模态基础模型 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Nanyang Technological University / VBVR Community,2026-08-26 已开源 vbvr-pro
多模态基础模型 WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report WeChat Vision, Tencent,2026-08-25 已开源 wemm-embedding
网络架构 RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models Huazhong University of Science and Technology,2026-08-21 未发现官方代码 rare
预训练与数据 Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing New York University,2026-08-13 未发现官方代码 tcab
多模态基础模型 Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction ByteDance,2026-08-12 未发现官方代码 gas
注意力与长上下文 Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension Ai2,2026-08-10 已开源 olmpool-long-context
推理与系统效率 DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference Oklahoma State University,2026-08-09 未发现官方代码 distillcache
注意力与长上下文 Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry Institute of Computing Technology, Chinese Academy of Sciences,2026-08-07 未发现官方代码 autonomy-heads
推理与系统效率 BaKron: Efficient Quantization with Kronecker-Factored Hessians University of California, San Diego,2026-08-06 未发现官方代码 bakron
预训练与数据 Hierarchical Latent Prediction for Language Models University of Texas at Austin,2026-08-06 未发现官方代码 hilp
网络架构 MACRO: Markov Chain Routing of Transformer Layers Heinrich Heine University Düsseldorf,2026-08-06 已开源 macro
推理与系统效率 DBLast: Dependent Block Drafting for Stochastic Speculative Decoding Huawei Technologies Canada,2026-08-05 未发现官方代码 dblast
推理与系统效率 QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding Indian Institute of Technology Roorkee,2026-08-05 未发现官方代码 qevict
多模态基础模型 Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Meta,2026-08-05 未发现官方代码 physics-mm-pretraining
网络架构 DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Zhejiang University,2026-08-03 未发现官方代码 dart
注意力与长上下文 Learning What to Remember: Test-Time Training via Context Distillation Princeton University,2026-08-03 已开源 ttcd
网络架构 Role-Decoupled Attention Residuals Kehan Wang(论文未列机构),2026-08-03 未发现官方代码 rd-attnres
网络架构 TransMem: Transforming Hidden States into Memory for Large Language Models Authors did not disclose affiliation,2026-07-31 未发现官方代码 transmem
多模态基础模型 ReToken: One Token to Improve Vision–Language Models for Visual Retrieval UIUC / Microsoft Research / Google DeepMind,2026-07-30 已开源 retoken
推理与系统效率 WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning EIT-NLP / LMU Munich,2026-07-30 已开源 wide
网络架构 Penelope: Localized Latent Recurrence for Efficient Structured Reasoning Academic author team,2026-07-28 未发现官方代码 penelope
预训练与数据 DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data Fudan University / Shanghai Jiao Tong University / SII-GAIR,2026-07-27 已开源 data-orchestra
推理与系统效率 Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs Huawei ACS Lab / Southern University of Science and Technology,2026-07-23 未发现官方代码 adadsf
注意力与长上下文 Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Independent researcher,2026-07-23 未发现官方代码 mobius-rope
网络架构 Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory Independent researchers,2026-07-23 未发现官方代码 naju
注意力与长上下文 Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection Pennsylvania State University,2026-07-23 未发现官方代码 gzip-sparse-attention
推理与系统效率 Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context NVIDIA,2026-07-23 已开源 windowed-mtp
推理与系统效率 GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries University of Cambridge,2026-07-22 已开源 gaugequant
网络架构 Convolution for Large Language Models Huawei / Peking University / Tsinghua University,2026-07-20 未发现官方代码 conv-llm
推理与系统效率 C²KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference Shanghai Jiao Tong University,2026-07-20 已开源 c2kv
预训练与数据 PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning McGill University,2026-07-20 未发现官方代码 ppl-factory
预训练与数据 OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research Indian Institute of Technology Madras,2026-07-18 已开源 open-language-model
注意力与长上下文 Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers University of Maryland / Meta AI,2026-07-16 未发现官方代码 looped-latent-attention
注意力与长上下文 MiniMax Sparse Attention MiniMax,2026-06-11 已开源 minimax-sparse-attention
网络架构 Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory Tsinghua University / Microsoft Research Asia,2026-05-20 未发现官方代码 memory-grafting
注意力与长上下文 Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers Peking University / Huawei Technologies,2026-03-27 未发现官方代码 switch-attention
生成式检索 Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion University of Illinois Urbana-Champaign / Google Research,2026-03-06 未发现官方代码 r4t
网络架构 Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models DeepSeek,2026-01-12 已开源 engram
网络架构 mHC: Manifold-Constrained Hyper-Connections DeepSeek-AI,2025-12-31 未发现官方代码 mhc
动态候选分类 GLiClass: Generalist Lightweight Model for Sequence Classification Tasks Knowledgator,2025-08-11 已开源 gliclass
注意力与长上下文 Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Qwen / Alibaba,2025-05-10 已开源 gated-attention
多模态基础模型 SmolVLM: Redefining small and efficient multimodal models Hugging Face,2025-04-07 已开源 smolvlm
预训练与数据 Muon is Scalable for LLM Training Moonshot AI / UCLA,2025-02-24 已开源 muon
多模态基础模型 SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Google DeepMind,2025-02-20 已开源 siglip2
注意力与长上下文 MoBA: Mixture of Block Attention for Long-Context LLMs Moonshot AI,2025-02-18 已开源 moba
注意力与长上下文 Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek,2025-02-16 未发现官方代码 native-sparse-attention
网络架构 Byte Latent Transformer: Patches Scale Better Than Tokens Meta FAIR,2024-12-13 已开源 blt
网络架构 Hymba: A Hybrid-head Architecture for Small Language Models NVIDIA,2024-11-20 已开源 hymba
预训练与数据 Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance University of Cambridge / Shanghai AI Laboratory,2024-03-25 已开源 data-mixing-laws
推理与系统效率 Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Together AI / Princeton University / University of Illinois Urbana-Champaign,2024-01-19 已开源 medusa
网络架构 Mamba: Linear-Time Sequence Modeling with Selective State Spaces Carnegie Mellon University / Princeton University,2023-12-01 已开源 mamba
推理与系统效率 AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration MIT / NVIDIA / Harvard / SJTU,2023-06-01 已开源 awq
注意力与长上下文 GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints Google Research,2023-05-22 未发现官方代码 gqa
预训练与数据 DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining Stanford University / Google Research,2023-05-17 已开源 doremi
多模态基础模型 Visual Instruction Tuning University of Wisconsin-Madison / Microsoft Research / Columbia University,2023-04-17 已开源 llava
多模态基础模型 BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Salesforce Research,2023-01-30 已开源 blip2
推理与系统效率 Fast Inference from Transformers via Speculative Decoding Google Research,2022-11-30 已开源 speculative-decoding
注意力与长上下文 Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation University of Washington / Meta AI,2021-08-27 已开源 alibi
注意力与长上下文 RoFormer: Enhanced Transformer with Rotary Position Embedding Zhuiyi Technology,2021-04-20 已开源 rope
多模态基础模型 Learning Transferable Visual Models From Natural Language Supervision OpenAI,2021-02-26 已开源 clip
网络架构 Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity Google Brain,2021-01-11 已开源 switch-transformer

分类浏览: