Tech AI News - 테카이
생각하는 토큰들이 겹겹이 쌓여 올라가는 모습의 AI 실루엣

Reasoning models — why thinking AI costs more

· 4 min read

This post was translated from the Korean original by AI.한국어 원문 읽기 →

“It’s only half-right to say ‘reasoning models are pricier because they’re smarter.’ A big chunk of the price gap exists because invisible ‘thinking tokens’ are fully billed at output rates. If you don’t understand this, even a simple question can rack up a bill dozens of times higher.”


Background / why this term now

Starting with OpenAI’s o1 (launched in September 2024), then o3, Anthropic Claude’s extended thinking and adaptive thinking, Google Gemini’s thinking models, and DeepSeek R1 released in January 2025, major AI companies have all rolled out models that “think before they answer.” These models clearly improved performance over previous models on math, coding, and complex reasoning tasks, but they also triggered more user complaints like “why did the same question get so expensive?”

The reason isn’t that the models got bigger. It’s that the amount of “reasoning tokens” the model generates internally before producing an answer has increased, and even though these tokens are invisible to users, they’re still included in billing.

Key data & current landscape

항목OpenAI 추론 모델Anthropic ClaudeDeepSeek R1
추론 토큰 처리응답에 노출되지 않지만 output_tokens 안의 reasoning_tokens로 집계·과금thinking 블록이 output_tokens_details.thinking_tokens로 집계·과금사고 과정이 출력에 포함되는 구조, 출력 단가 자체가 낮게 책정
조절 파라미터reasoning effort(저/중/고 등 단계)budget_tokens(수동) 또는 effort(적응형)별도 공개 파라미터 문서 없음
근거OpenAI 공식 Reasoning 가이드Anthropic 공식 Extended thinking 문서Epoch AI 분석(2025년 1월)

Interpretation: All three companies follow the same principle—“thinking tokens are billed at the output token rate.” The differences lie in how finely developers can control that thinking budget and what margin each company adds for the same amount of thinking. Epoch AI analyzed DeepSeek R1’s output prices as about 1/27 of OpenAI o1’s, arguing that much of the gap comes from pricing policy (margin), not just technical efficiency.

Deep dive

1) What exactly is a reasoning model?

A standard LLM takes a question and immediately predicts the next tokens to generate an answer. A reasoning model inserts a “thinking phase” first. It breaks the problem into sub-steps, checks intermediate results, and sometimes tries multiple solution paths before giving the final answer. This thinking process itself is a sequence of tokens, and OpenAI’s official docs describe them as “tokens the model uses to decompose the prompt and consider multiple approaches.” In training, reinforcement learning is also used to encourage “longer and more accurate chains of thought,” which is different from merely scaling the model up.

2) A common misconception—why do people read it as “it’s pricier because it’s smarter”?

Users only see the final answer, so when prices rise, it’s easy to assume “the model just got better.” But your invoice includes invisible thinking tokens. According to OpenAI’s official docs, reasoning tokens don’t appear in the API response content, yet they occupy context and are billed as output tokens. Anthropic applies the same rule and explicitly states that tokens used in the thinking block are billed at the output token rate. So part of the price gap is not “because it’s smarter,” but because “more invisible tokens were generated.”

3) What actually drives the decision

Both companies provide knobs to control the amount of thinking. OpenAI lets you set reasoning effort (e.g., low/medium/high) to cap the generation of thinking tokens. Anthropic lets you directly target a thinking-token budget with budget_tokens (manual mode) or, in adaptive thinking, specify an effort level so the model decides whether to think at all. In other words, the right cost metric is not “how smart is the model,” but “how much thinking budget does this task need?”

4) When it’s useful and when it isn’t

For tasks with genuine “thinking load”—multi-step math proofs, complex code debugging, or design problems with many simultaneous constraints—longer reasoning markedly boosts accuracy. Stanford’s “s1” paper (January 2025) showed that simply increasing the thinking-token budget improves reasoning performance. But more thinking doesn’t help indefinitely. A 2026 study analyzing models with elongated reasoning found that beyond a point, gains drop sharply and models even flip correct answers into wrong ones—an “overthinking” effect. For simple fact checks or short summaries with little to ponder, leaving reasoning models on by default often raises costs without improving quality.

In practice — 4 rules for reading reasoning-model costs

  1. When reading price sheets, interpret them as “how many extra thinking tokens are being minted as output” rather than “the model got smarter.” Both OpenAI and Anthropic state this billing rule in their docs.
  2. If your API exposes knobs like reasoning effort or budget_tokens, don’t accept the defaults. Lower them to match task difficulty; for simple tasks, low settings are often enough.
  3. Track costs by inspecting reasoning_tokens/thinking_tokens in the usage response, not by eyeballing the visible answer length. Perceived answer length can diverge from actual billed tokens.
  4. Remember there are large cross-vendor price gaps even within the “reasoning model” category. That reflects not only technical efficiency but also margin, so don’t assume “the pricier one thinks more” by default.

References

  • #ai terms
  • #reasoning model
  • #llm
  • #openai
  • #anthropic
  • #deepseek
  • #token cost