Volver a Radar
Impacto Alto 3 min

Las Novedades: TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI

Fuente: Hugging Face 31 Aug 2026
Las Novedades: TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
Resumen

We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive

expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regular and efficient expert execution, while retaining dropless routing during pretraining.

Turing-20B-A2B also employs a hybrid attention architecture that combines Lightning Attention with a small number of full-attention layers for efficient long-context modeling. The model is pretrained with a progre...

Leer noticia completa en Hugging Face →

Serás redirigido al portal original.

#Inteligencia Artificial #Tendencias #Radar