3 Million Won Monthly Token Fees: A New Reality


Tokens are changing from a unit of conversation count to a means of production for purchasing AI execution time and exploration volume.
It used to be an exaggeration, but now it's a feasible cost.

In the past, the speed of human input and verification, rather than the model itself, created an upper limit on usage.
Just a few years ago, it was hard to believe someone saying, "I spent 3 million won on AI token costs in a month." It was difficult to consistently consume that many tokens when people had to manually input questions and wait for answers.
The bottleneck at the time was not the model, but humans. The speed at which people created questions, read answers, and gave subsequent instructions set the upper limit on usage. Even with expensive models, consumption stopped when people went to sleep.
GPT-6 Astra and long-duration agents have begun to remove this upper limit. Agents read screens, call tools, identify causes of failure, and correct them. This process continues even after a person steps away. With high inference intensity and parallel agents, 3 million won per month can become a real operating cost, not just a number for show.
The core of the change isn't just that model prices have increased.
An execution structure has been created that allows work to continue, consuming tokens even while people are resting.
Related Evidence·OpenAI ChatGPT·Codex Pricing and Usage Guide·Codex 20× Limit Exhaustion User Cases
Astra-Ultra has overcome the barriers that remained up to Sol-Ultra.

Long-duration agents repeat investigation, execution, review, and correction even after a person steps away.
Sol-Ultra also used a lot of inference, but it was relatively difficult for a typical user to exhaust their subscription limit in a short time. This was because tasks would often end midway, require human intervention for the next instruction, or there were practical limits to the computational load a model could consume within a single response.
With Astra-Ultra, three things changed simultaneously. It can autonomously continue longer tasks, directly use browsers and specialized software, and repeatedly perform trial and error with high inference intensity. As long contexts and tool results continuously accumulate, the amount of information to read for a single decision increases towards the latter half of a task.
High Inference Intensity: Performs more hypotheses and reviews in a single step. — Invisible inference tokens increase.
Long-Duration Agent: Repeats execution, verification, and correction without human intervention. — Human input speed is no longer an upper limit.
Large Task Context: Files, conversations, search results, and tool outputs accumulate. — Input costs for subsequent requests can increase.
Parallel Agents: Explores multiple approaches simultaneously. — Time is reduced, but total consumption increases rapidly.
In the actual user community, there were cases where Astra exhausted the Codex 20× limit in a single day. OpenAI officials later stated that they implemented improvements to reduce subscription usage deductions by up to 3-4 times for power users' long-duration tasks. While individual experiences cannot generalize consumption for all accounts, it is clear that providers are also addressing the issue of usage in long-duration tasks.
Related Evidence·OpenAI GPT-6 Astra Announcement·Astra Long-Duration Task Usage Improvement Guide·Cases of Setting Remaining Usage as Task Budget
How true is the claim that '9 billion won' was spent on a mathematical problem?

Approximately 9 billion won is the retail equivalent calculated using public API prices, and the actual infrastructure expenditure has not been disclosed.
In September 2026, OpenAI revealed the solution to the Navier-Stokes Millennium Problem along with the scale of the effort. Approximately 10,000 agents operated simultaneously, taking about 88 hours to reach the solution. During this process, the agents exchanged 2.7 million messages and used approximately 130 billion output tokens.
Simply applying Astra's public API output price of $50 per million tokens yields $6.5 million. Assuming an exchange rate of 1,380 won per dollar, this amounts to approximately 8.97 billion won. The '9 billion won' mentioned online comes from this calculation.
130 billion output tokens ÷ 1 million × $50
= $6.5 million ≈ 8.97 billion won
However, it is inaccurate to call this the actual amount paid by OpenAI. The system that created the solution was not Astra, but a next-generation internal model that OpenAI described as "much more powerful than Astra." Astra was subsequently used for 17 hours for Lean formalization and verification. The price of the internal model and the actual infrastructure costs have not been disclosed.
Therefore, a safer statement is, "Converting the disclosed number of tokens to Astra's retail API output price amounts to approximately 9 billion won." The solution is also a result announced by OpenAI and should not be taken to mean that sufficient independent review by the mathematical community has been completed.
Related Evidence·OpenAI Navier-Stokes Solution and Project Scale·GPT-6 Astra API Pricing

Large-scale parallel exploration tests more hypotheses simultaneously but also significantly increases total token consumption.
Now, efficiency doesn't just mean 'the ability to use fewer tokens'.

In the age of agents, both token-per-unit efficiency and human-hour efficiency must be evaluated together.
Token efficiency typically referred to the ability to produce the same answer with fewer tokens. This criterion remains important because unnecessary context and repeated calls increase both cost and latency.
However, in the age of agents, other efficiencies must also be considered. Even if many tokens are used, if exploration and execution that would take a human several days can be completed in a few hours, it might be more economical for the organization as a whole. If 3 million won in token costs can replace labor worth over 10 million won or several months of exploration, it is efficient per human-hour, even if expensive per token.
Efficiency per Token: How few tokens were used to produce the same result — It might finish quickly, but the result could be useless.
Efficiency per Completed Task: How much did it cost to obtain one verified result — Failed attempts must also be included in the calculation.
Efficiency per Human-Hour: How much human intervention and waiting time were reduced — Time spent on review and recovery must also be included.
Efficiency per Exploration Scope: How many approaches were validated within a given time — Parallelization can create redundant work.
In a 'token-maxing' experiment conducted by Vals AI, some engineers reportedly used 1-2 billion tokens per day, with one engineer recording 6 billion tokens in a single day. The estimated value of tokens used over a month was approximately $1.5 million, about 10 times the employee salaries for the same period. While this is a limited experiment by a specific company and cannot be considered a general average, it demonstrates that token costs are beginning to emerge as a management item, similar to labor costs.
Related Evidence·Vals AI Token-Maxing Experiment Summary·Vals AI Interview and Usage Cases
Tokens are shifting from a unit of consumption to a means of production.

Even when using the same model, exploration capabilities vary depending on execution time, parallelism, and the budget available to tolerate failures.
Just because one can use the same model doesn't mean one can achieve the same results. One person might run a single agent for two hours, while another organization might have 10,000 agents explore different approaches for 88 hours. These two users may share the same model name, but the total amount of intelligence they can actually access differs.
This difference can be called 'token power.' Above model access, inference intensity, execution time, parallelism, and the budget to tolerate failures create another hierarchy. If the digital divide of the past was about access to the internet and software, the divide in the age of agents is about how long one can repeatedly invoke intelligence.
Accessing the same model is different from having the same exploration capabilities.
The token budget determines the number of hypotheses that can be attempted and the number of failures that can be tolerated.
However, burning through many tokens doesn't automatically guarantee good results. Redundant agents, expanded context, and incorrect completion criteria can lead to massive costs with little to show. Token power includes not only consumption capability but also the operational ability to evaluate and halt results.
Related Evidence·10,000 Agents · 130 Billion Output Tokens Case Study·Usage Guide by Model, Context, Inference, and Tools
The ability needed going forward is to design token budgets.

Powerful models should focus on problems with high-value outcomes, and costs and termination conditions must be designed before the task begins.
It's now difficult to explain productivity with just a single good prompt. Decisions must be made on which tasks to assign to Astra-Ultra and which to smaller models. It's also necessary to pre-determine when to increase parallel agents and what conditions will trigger task completion.
Enterprise AI cost management must also shift from simply counting monthly subscription fees. It's essential to track not only usage per model but also completed tasks, human-hours saved, rework rates, failed attempts, and review costs. Only then can one distinguish between teams that spent many tokens and those that delivered actual results.
The same principle applies to individuals. Routine organization and drafting are best handled by lighter models, while powerful models should be concentrated on high-value segments like complex design, research, and verification. For long-duration tasks, remaining usage or maximum cost should be set as termination conditions.
Competitiveness in the age of agents is not about using tokens without limit.
It lies in deploying sufficient inference to the most valuable problems and halting consumption that doesn't yield results.

Competitiveness is not about using tokens without limit, but about deciding where and how much intelligence to deploy.
Tokens are no longer a unit for counting conversations. They are becoming computational capital for purchasing AI labor hours and exploration volume. The change demonstrated by GPT-6 Astra is not just the emergence of smarter models, but the advent of a cost structure that transforms intelligence into actual output.
Related Evidence·OpenAI Usage Reduction Recommendations and Token/Credit Standards·GPT-6 Astra Long-Duration and Tool-Based Task Guide
Sources and References
OpenAI Announces GPT-6 Astra·OpenAI Announces Navier-Stokes Solution·GPT-6 Astra Model and API Pricing·ChatGPT Work·Codex Pricing and Usage·Astra Subscription Limit Exhaustion Cases·Astra Usage Improvement Guide·Vals AI Token-Maxing Interview Summary
Draft and Publication Notes
In this article, the approximate 9 billion won is an estimated value derived by applying Astra's public retail API price to 130 billion output tokens, and it is not OpenAI's actual expenditure. Cases of subscription limit exhaustion and token-maxing are experiences of individual users and specific organizations, and therefore should not be interpreted as the average consumption of general users. The Navier-Stokes result is a solution disclosed by OpenAI and does not imply that sufficient independent academic review has been completed.