-
Attention Is All You Need
Paper • 1706.03762 • Published • 138 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 116M • • 2.73k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 13.7M • 3.41k
Collections
Discover the best community collections!
Collections including paper arxiv:1707.06347
-
High-Resolution Image Synthesis with Latent Diffusion Models
Paper • 2112.10752 • Published • 17 -
Adding Conditional Control to Text-to-Image Diffusion Models
Paper • 2302.05543 • Published • 60 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149
-
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149 -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Paper • 2501.12948 • Published • 463
-
Attention Is All You Need
Paper • 1706.03762 • Published • 138 -
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 65 -
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Paper • 2101.03961 • Published • 13 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12
-
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Paper • 2304.09842 • Published • 2 -
ReAct: Synergizing Reasoning and Acting in Language Models
Paper • 2210.03629 • Published • 38 -
Gorilla: Large Language Model Connected with Massive APIs
Paper • 2305.15334 • Published • 7 -
Reflexion: Language Agents with Verbal Reinforcement Learning
Paper • 2303.11366 • Published • 9
-
Attention Is All You Need
Paper • 1706.03762 • Published • 138 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 116M • • 2.73k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 13.7M • 3.41k
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149
-
High-Resolution Image Synthesis with Latent Diffusion Models
Paper • 2112.10752 • Published • 17 -
Adding Conditional Control to Text-to-Image Diffusion Models
Paper • 2302.05543 • Published • 60 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149 -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Paper • 2501.12948 • Published • 463
-
Attention Is All You Need
Paper • 1706.03762 • Published • 138 -
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 65 -
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Paper • 2101.03961 • Published • 13 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12
-
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Paper • 2304.09842 • Published • 2 -
ReAct: Synergizing Reasoning and Acting in Language Models
Paper • 2210.03629 • Published • 38 -
Gorilla: Large Language Model Connected with Massive APIs
Paper • 2305.15334 • Published • 7 -
Reflexion: Language Agents with Verbal Reinforcement Learning
Paper • 2303.11366 • Published • 9