Publicidad
Zona de Patrocinadores de Herramientas de Desarrollo y Nube

StepFun Releases Step-3.7-Flash in NVFP4: Sub-10ms Time-to-First-Token for Real-Time Voice and Agents

Por Marcus Sterling Publicado el 2026-09-07 4 min read Fuente: StepFun AI
StepFun collaborates with NVIDIA to release Step-3.7-Flash in native NVFP4 quantization, achieving sub-10ms TTFT for conversational agent workflows.

StepFun has unveiled **Step-3.7-Flash**, bringing sub-10-millisecond time-to-first-token (TTFT) performance to production deployments.

Native NVFP4 Optimization Optimized directly for Blackwell B200 and Hopper Tensor Cores with NVIDIA NVFP4 quantization, Step-3.7-Flash operates with minimal latency jitter, making it the preferred backend for synchronous voice-to-voice agents and high-frequency code completion engines.

Publicidad
Infraestructura de APIs e Inteligencia Artificial (728x90)

Fuente y Verificación

Este informe técnico fue contrastado contra la documentación primaria publicada por StepFun AI.

Leer Anuncio Original en StepFun AI →