StepFun Releases Step-3.7-Flash in NVFP4: Sub-10ms Time-to-First-Token for Real-Time Voice and Agents
StepFun collaborates with NVIDIA to release Step-3.7-Flash in native NVFP4 quantization, achieving sub-10ms TTFT for conversational agent workflows.
StepFun has unveiled **Step-3.7-Flash**, bringing sub-10-millisecond time-to-first-token (TTFT) performance to production deployments.
Native NVFP4 Optimization Optimized directly for Blackwell B200 and Hopper Tensor Cores with NVIDIA NVFP4 quantization, Step-3.7-Flash operates with minimal latency jitter, making it the preferred backend for synchronous voice-to-voice agents and high-frequency code completion engines.
Publicidad
Infraestructura de APIs e Inteligencia Artificial (728x90)
Fuente y Verificación
Este informe técnico fue contrastado contra la documentación primaria publicada por StepFun AI.
Leer Anuncio Original en StepFun AI →