-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 88 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 235 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 161 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
Collections
Discover the best community collections!
Collections including paper arxiv:2608.15062
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 73 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 10 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 9
-
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
Paper • 2608.15062 • Published • 10 -
Amr-Hegazy/grt-large-isoflop
Text Generation • 0.3B • Updated • 61 -
Amr-Hegazy/grt-large-isoparam
Text Generation • 0.8B • Updated • 61 -
Amr-Hegazy/grt-medium-isoflop
Text Generation • 0.4B • Updated • 65
-
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
Paper • 2401.02994 • Published • 52 -
MambaByte: Token-free Selective State Space Model
Paper • 2401.13660 • Published • 59 -
Repeat After Me: Transformers are Better than State Space Models at Copying
Paper • 2402.01032 • Published • 25 -
BlackMamba: Mixture of Experts for State-Space Models
Paper • 2402.01771 • Published • 25
-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 88 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 235 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 161 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
-
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
Paper • 2608.15062 • Published • 10 -
Amr-Hegazy/grt-large-isoflop
Text Generation • 0.3B • Updated • 61 -
Amr-Hegazy/grt-large-isoparam
Text Generation • 0.8B • Updated • 61 -
Amr-Hegazy/grt-medium-isoflop
Text Generation • 0.4B • Updated • 65
-
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
Paper • 2401.02994 • Published • 52 -
MambaByte: Token-free Selective State Space Model
Paper • 2401.13660 • Published • 59 -
Repeat After Me: Transformers are Better than State Space Models at Copying
Paper • 2402.01032 • Published • 25 -
BlackMamba: Mixture of Experts for State-Space Models
Paper • 2402.01771 • Published • 25
-
AI for Auto-Research: Roadmap & User Guide
Paper • 2605.18661 • Published • 73 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper • 2605.16865 • Published • 10 -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper • 2603.28069 • Published • 9