Fewer tokens, so why is GPT-5.5 in Codex actually more expensive?
Stunned. / Dumbfounded.
The official ChatGPT side doesn’t make it easy to track tokens and costs directly, so I found a third-party platform and ran a round of similar tasks using GPT-5.4 and GPT-5.5 in Codex, setting the thinking mode to high. The result was very straightforward: simple questions were relatively mild (in terms of cost), with GPT-5.5 being about 30% more expensive than GPT-5.4; however, once complex tasks were involved, the costs shot up to 2.6 times, and both the request count and token consumption increased simultaneously.
My current assessment is very straightforward: this isn’t something that can be decided just because of the statement “5.5 has a higher unit price.” In simple tasks, the cost mainly comes from the unit price; but in complex tasks, what is actually expensive is the entire calling chain (or execution flow). However, looking at it another way, 5.5 does genuinely feel like it’s absorbing your rework costs for you. The model is more willing to think through multiple steps, perform more actions, and check things more thoroughly. Ultimately, the billing isn’t based on a single answer; it’s based on the complete set of actions, which also minimizes the number of back-and-forth cycles required from the human user.