ArchSpace-Collection/NCP_ArchPreview_dolma3_8.9B_Stage1 Text Generation • 9B • Updated 6 days ago • 1.15k • 8
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 8 days ago • 318
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 8 days ago • 318
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 62
MemSFT: Mitigating Alignment Tax with an External Parametric Memory Paper • 2607.25614 • Published Jul 28 • 23
Depth-Attention: Cross-Layer Value Mixing for Language Models Paper • 2606.05014 • Published Jun 3 • 1
K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization Paper • 2306.05064 • Published Sep 13, 2023
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models Paper • 2602.08984 • Published Feb 9
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space Paper • 2509.23184 • Published Sep 27, 2025 • 2
RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL Paper • 2205.06983 • Published May 14, 2022
Training LLMs to be Better Text Embedders through Bidirectional Reconstruction Paper • 2509.03020 • Published Sep 3, 2025
Theano: A Python framework for fast computation of mathematical expressions Paper • 1605.02688 • Published May 9, 2016 • 2
Critical Data Size of Language Models from a Grokking Perspective Paper • 2401.10463 • Published Jan 19, 2024 • 1
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence Paper • 2502.13943 • Published Feb 19, 2025 • 8
FreqKV: Frequency Domain Key-Value Compression for Efficient Context Window Extension Paper • 2505.00570 • Published May 1, 2025 • 1