GPT-5.5 is OpenAI's most capable general-purpose model, designed for tasks requiring sustained multi-step reasoning, complex tool orchestration, and deep coding. It handles a 1M+ token context window, making it suitable for large-scale document analysis and code review. The model excels at agentic workflows where it needs to plan, execute, and iterate across multiple tool calls. Cost and latency are at the high end of OpenAI's lineup — reach for GPT-5.5 Mini or Nano for lighter tasks.
| Benchmark | Score | Type | Recorded |
|---|---|---|---|
| LiveCodeBench | 85.3 | accuracy | 29d ago |
| IFBench | 75.9 | accuracy | 29d ago |
| SciCode | 55.8 | accuracy | 29d ago |
| LCR | 84.3 | accuracy | 29d ago |
| MMLU-Pro | 88.1 | accuracy | 29d ago |
| GPQA Diamond | 93.6 | accuracy | 29d ago |
| Humanity's Last Exam | 41.4 | accuracy | 29d ago |
| TAU2 | 98.0 | accuracy | 29d ago |