GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. $2.50 per million input tokens, $15 per million output tokens. 1,050,000 token context window, maximum output of 128,000 tokens. Higher uptime with 3 providers. Includes independent benchmarks from Artificial Analysis.
Tried 5.6. Didn’t like it. Specifically, it plunged right into the codebase and started making changes without asking clarifying questions or offering analysis of the problem or any visible explanations, unlike Opus or other frontier models. Disconcerting.
It’s excellent if you can keep it in hand, but holy hell is it tough to get it to stay between the lines. I’ve gone back to 5.5, it’s just way less handholding. And 5.5 gets things done well enough, in about a tenth of the time.
Sol is good on Ultra for orchestrating if you specify the models you want it to use for subagents. Other than that, it’s too much work.
It is miles ahead of Opus when designing scientific experiments. If you want to discover novel techniques or combine known techniques in new ways and you want to be sure that what you are measuring is real and not dependent on your specific dataset, Opus or even Fable just does not compete. Opus is fine for running experiments, but when it comes to interpreting results or recommending next steps, it makes mistakes.
Tried 5.6. Didn’t like it. Specifically, it plunged right into the codebase and started making changes without asking clarifying questions or offering analysis of the problem or any visible explanations, unlike Opus or other frontier models. Disconcerting.
It’s excellent if you can keep it in hand, but holy hell is it tough to get it to stay between the lines. I’ve gone back to 5.5, it’s just way less handholding. And 5.5 gets things done well enough, in about a tenth of the time.
Sol is good on Ultra for orchestrating if you specify the models you want it to use for subagents. Other than that, it’s too much work.
It is miles ahead of Opus when designing scientific experiments. If you want to discover novel techniques or combine known techniques in new ways and you want to be sure that what you are measuring is real and not dependent on your specific dataset, Opus or even Fable just does not compete. Opus is fine for running experiments, but when it comes to interpreting results or recommending next steps, it makes mistakes.