OpenAI Announces GPT-5.6 Sol: Dominating Terminal Bench 3.0 (34.6%) and DeepSWE (72.7%) with Sol Architecture
OpenAI unveils GPT-5.6 Sol, featuring autonomous agent test-time compute, setting all-time highs across SWE-Bench Verified, DeepSWE, and Terminal Bench 3.0.
OpenAI has officially launched **GPT-5.6 Sol**, the commercial flagship designed specifically for end-to-end software engineering agents and autonomous systems.
Sol Test-Time Compute Architecture GPT-5.6 Sol integrates deep search during code generation. When tasked with fixing an open-source issue, Sol simulates compiler outputs, generates test fixtures in ephemeral containers, and verifies solutions before returning final code diffs.
Record Scores across Benchmark Suites - **DeepSWE (v1.1)**: **72.7%** (World Record). - **Terminal Bench 3.0**: **34.6%** (World Record). - **ExploitGym (6h)**: 293 verified exploit chains, establishing unmatched proficiency in automated security patching.
Publicidad
Infraestructura de APIs e Inteligencia Artificial (728x90)
Fuente y Verificación
Este informe técnico fue contrastado contra la documentación primaria publicada por OpenAI Research.
Leer Anuncio Original en OpenAI Research →