DeepSeek Lanza DeepSeek-V4.1-Flash: MoE Multimodal de 552B con Arquitectura Causal Encoder-Decoder y 1M de Contexto
DeepSeek presenta un modelo MoE de 552B que activa solo 8B de parámetros en prefill y 16B en decode gracias a su arquitectura Causal Encoder-Decoder (CED), reduciendo drásticamente la huella de memoria con licencia MIT.
DeepSeek has published **deepseek-harness**, the inference and plugin orchestration harness powering its web and API infrastructure.
DeepSelect & DeepGEMM deepseek-harness open-sources DeepSelect—a GPU-accelerated routing kernel that evaluates attention sparsity in real-time, executing Dynamic Sparse Attention (DSA) with near-zero latency overhead. Paired with DeepGEMM FP8 matrix operations, self-hosted clusters achieve up to a 3.4× boost in tokens-per-second-per-GPU.
Publicidad
Infraestructura de APIs e Inteligencia Artificial (728x90)
Fuente y Verificación
Este informe técnico fue contrastado contra la documentación primaria publicada por DeepSeek GitHub.
Leer Anuncio Original en DeepSeek GitHub →